Measure municipal AI by public outcomes
An AI pilot can generate impressive activity numbers: questions answered, documents summarized, requests classified, or minutes reportedly saved. Municipal leaders need a harder test. Did the service become more accessible, timely, accurate, equitable, and accountable for constituents?
The OECD's work on AI in public service delivery describes potential gains in efficiency, responsiveness, resource allocation, and service pathways. It also emphasizes the enablers and safeguards around the technology. A municipal scorecard should therefore evaluate outcomes, service quality, operations, and public risk together.
1. Constituent outcome
Select the result the service exists to produce. Examples include:
- a complete permit application accepted for review;
- a resident connected to the correct service on the first attempt;
- an eligible participant receiving a benefit sooner;
- a safety or maintenance issue reaching the responsible team;
- a public meeting record becoming accessible within a defined period.
Use a baseline and comparison period. Separate the AI contribution from staffing, policy, seasonal, or demand changes where possible.
2. Service quality
Measure what residents experience:
- time to first useful response;
- first-contact resolution;
- avoidable transfers and repeat contacts;
- correction and rework rates;
- accessibility and language quality;
- complaints and successful appeals;
- satisfaction, with a method that does not exclude people who use offline channels.
A faster incorrect response is not an improvement.
3. Staff operations
Track whether the system helps employees deliver the service:
- review time per completed case;
- percentage of outputs substantially corrected;
- escalation volume and reasons;
- cases abandoned because the tool was unhelpful;
- service interruptions and recovery time;
- cost per completed constituent outcome.
Reported "hours saved" should be treated carefully. Determine whether capacity was actually released and what public work used it.
4. Fairness, privacy, and safety
Record:
- materially incorrect or harmful output;
- privacy and security incidents;
- unauthorized information use;
- differences in service quality across channels or groups, where measurement is lawful and methodologically sound;
- missed escalation conditions;
- actions taken without required human authority;
- repeat failures after corrective action.
The Government of Canada's Guide on Departmental AI Responsibilities warns that poorly governed AI can amplify fairness, equity, privacy, security, legal, and reputational risks. Municipal legal duties differ, but those risk categories remain relevant to public-service management.
5. Transparency and recourse
Test whether a resident can:
- tell when they are interacting with AI;
- understand what the system is doing in plain language;
- reach a person;
- correct inaccurate information;
- contest an affected decision where applicable;
- learn which department is accountable.
Toronto's Digital Infrastructure Strategic Framework commits to ethical, accountable, transparent, and supervised digital decisions and understandable public information about City technology.
Review the scorecard as one system
Do not optimize one column in isolation. Lower handling time can coincide with more transfers. High chatbot completion can hide residents who could not reach a person. High staff adoption can reflect convenience without improving the public outcome.
Set thresholds before the pilot. Define who reviews the evidence, how often, and which results require correction, narrower use, more human review, or a pause.
Municipal AI earns expansion when evidence shows that it helps constituents and remains governable—not when it merely produces more output.