Insights

Public Safety

What to measure in month one of an AI non-emergency line

Colorado Springs and Dane County just published early numbers from AI agents answering their non-emergency lines. A completion rate and a call count are a start, but month one has to measure missed emergencies, handoff quality, and who fixes the failures.

September 29, 2026 · 7 min read

What to measure in month one of an AI non-emergency line

The Colorado Springs Police Department says its AI agent, S.A.R.A.H., built by Axon, completed 50% of non-emergency calls in a 54-hour test run and passed the other half to a human. Six days later, Dane County, Wisconsin reported that its assistant, AVA, had handled more than 24,000 non-emergency calls in about its first month. The county's own launch release identifies AVA as the Aurelian Virtual Assistant.

Those are honest first numbers, and both agencies deserve credit for publishing them. They are also volume numbers. Neither one tells you how many callers on the non-emergency line actually had an emergency, how long they waited for a person, or what happened to the transcript afterward. That is what month one should measure.

Anyone who has worked patrol knows people call the number they remember, not the number that matches the situation. The non-emergency line gets real emergencies. The scorecard below starts there.

Start with the missed-emergency review

The most important month-one metric is also the hardest to count: calls where the caller had an emergency and the AI did not get them to a human fast enough. You cannot read it off a dashboard. You find misses by reviewing calls.

Dane County's design shows the right instinct. Staff take turns monitoring AVA, and communicator Alison Kelly described the goal as making sure "if there is an emergency that somehow it's not catching, then we do have someone who can call someone back right away." Per the same report, AVA is trained to recognize keywords or warning signs of an emergency and forward those calls to a human dispatcher. Keyword detection catches what it was configured to catch. The review exists for everything else.

Full review may not be practical. Dane County's 24,000 calls in a month works out to roughly 800 a day. New Orleans reviewed every call its narrower 911 triage tool handled for the first three months, and that tool only engaged in a tight set of conditions. For a general non-emergency line, a workable month-one design is a fixed random sample of AI-completed calls every day, plus every call that trips a secondary signal: the same number or address calling 911 shortly after, a call for service later upgraded in priority, or a caller who hung up mid-conversation. Count the misses, write each one up, and track the rate by week.

Handoff rate is two numbers

A 50% completion rate needs a definition before it means anything. Completed could mean resolved with no human, or it could mean the caller hung up. Write the definition down in week one and hold it fixed.

Then split the handoffs by reason, because each reason tells you something different:

  1. The system detected a possible emergency.
  2. The caller asked for a person.
  3. The system could not work out where the call should go.
  4. The request was out of scope.

Dane County's public page says that when AVA cannot gather enough information to route a call, it goes to the communications center to hold for the next available communicator, behind incoming 911 calls. That is a sensible queue rule. It also means the third category should be watched closely, because those callers wait.

Handoff quality is the second number. For each handoff, check whether the summary that reached the human was accurate, whether the caller had to repeat themselves, and how many seconds passed between detection and a live voice. Measure hold time separately for emergency-detected transfers and for callers who asked for a person. The county's launch release is candid that callers requesting a live person may experience a wait while emergency calls are prioritized. Know how long that wait is.

One more reason to define your own baseline. The Colorado Springs share of non-emergency calls appears as "over half of all calls" in the news report and "nearly half of the calls" on the city's page. Either is fine for public communication. Neither is precise enough to measure a pilot against.

Test the opt-out on purpose

Both agencies say a caller can reach a person. Colorado Springs says those who want to speak to a person will continue to have that option. Dane County Deputy Director Johnny Leonard said callers can ask for a communicator and "it does pass them through."

Measure it anyway. Count the turns and seconds from "I want a person" to the transfer. Dane County's page explains that AVA first asks what the caller wants to speak to someone about, so it can route them, which is reasonable, and which is exactly the step to time. Run scripted test calls every week with different phrasing: "operator," "real person," "let me talk to somebody," a frustrated caller who says nothing useful, and the same requests in the languages your community speaks.

If your agency is in Texas, add a greeting check. Section 552.051(b) of the Texas Responsible Artificial Intelligence Governance Act requires a governmental agency that makes available an AI system intended to interact with consumers to disclose, before or at the time of the interaction, that the person is interacting with AI, in plain language. Confirm the greeting does that on every call path.

Language access is a per-language metric

A language count is a capability claim. Dane County says AVA can communicate in more than 45 languages, and Colorado Springs says the Axon system transcribes calls in dozens of languages. Month one should show what happens in each one.

Break every metric above out by the language detected: completion rate, handoff rate by reason, time to a human, and missed-emergency findings. Pull a sample of non-English calls for review by a fluent staff member or a qualified interpreter, not by the same system that handled the call. Watch for callers who switch languages mid-call, and for callers whose English is fine but whose accent the system misreads. If one language shows a much higher "could not route" rate, those callers are waiting longer, and you want to know that in week three instead of from a complaint.

Transcripts are records, and failures need an owner

Every AI-handled call produces something: audio, a transcript, a summary, a service request. The city of Colorado Springs says the Axon system automatically generates a written transcript of a call and an initial summary. Before month one ends, answer these in writing. Which retention schedule applies to each artifact? Who can read them? Does the vendor keep copies, and for how long? Can the content be used to train or tune the model? How do you answer a public records request for one call? If those records flow into CAD or your records system, map them the way I described in the CJIS AI governance gap.

Colorado Springs lists the policies that govern the system on its public page, including GO 1905, Use of Artificial Intelligence, and GO 1612, Records Security. Publishing the governing policies is a good practice to copy.

Then name who reviews failures. Leonard's read on Dane County's first month is the most useful sentence in either story: "Any of the mistakes or missed opportunities have been in the way that we've programmed the system." If most failures are configuration, the review has to include someone who can change the configuration, on the agency side and the vendor side. Keep a change log of every prompt, keyword, and routing change with a date, and rerun the scripted test calls after each one. Colorado Springs says its communications center employees will regularly review the system's performance and update its responses. Month one is when you decide what "regularly" means.

Before day 30, pull ten AI-completed calls at random and ten handoffs. Listen to all twenty with a senior call taker and write down what a person would have done differently. That list is your month-two configuration backlog.