Weaknesses
1Measured shortfalls, most urgent first
Weakness· from response times
Answers take far longer than a chat interface implies
29.3s median · 31.8s at p95
The median answer takes 29.3s and the slowest 5% take over 31.8s, against a 8s threshold past which people begin abandoning. The pipeline runs retrieval, generation and two translation hops in sequence, and nothing is streamed, so the citizen watches a blank panel for the whole duration.
What to do
Stream the response, run the two translation calls concurrently with generation where the language is already known, and cache translations of repeated questions. Measure again before adding retrieval work.
Opportunities
3Changes the data supports making
Opportunity· from ratings
Not one answer has been rated
0 of 2 answers rated
Every assistant reply carries a thumbs up/down control, but none of the 2 questions in this window produced a rating. Until some do, this dashboard can report what the pipeline believes it did and cannot report whether citizens agreed. Note also that only live web conversations record an assistant message id — imported rows can never carry a rating, by construction.
What to do
Ask for the rating rather than waiting for it: a plain "Was this helpful?" prompt under the answer reliably outperforms an unlabelled pair of icons. This is the highest-value change on this list, because it is the only one that unlocks every other feedback-based finding.
Opportunity· from language mixProvisional
The multilingual capability is barely being used
100% English · 1 language seen
100% of questions arrived in English, with no other language represented at all. Nine languages are live and billed for. Either citizens do not know they can ask in their own language, or detection is defaulting to English when it is unsure.
What to do
Put a visible language picker on the landing screen instead of relying on detection alone, and check the detector's fallback path — an unsure detection currently lands on English silently.
Opportunity· from input typesProvisional
Voice input is built but almost nobody uses it
0% of questions arrived as voice
Voice notes account for 0% of questions. The record-transcribe-review flow is shipped and working, so this is a discoverability problem rather than a capability one — and voice is the feature that matters most to the citizens least served by a text box.
What to do
Promote the microphone from an icon in the composer to an equal option on the landing screen, and confirm Speech-to-Text is not returning 403 in production — a blocked key looks identical to disinterest in this figure.
Strengths
1What is working and worth protecting
Strength· from outcomesProvisional
Most questions get a grounded answer
100% assisted
100% of questions produced a grounded answer rather than a referral to the website, at or above the 85% mark.