How Governments Can Govern Agentic AI

Key Takeaways

  • The Urban Institute’s playbook recommends defining each AI agent’s purpose, target population, expected outcomes, estimated costs, and intended users before deployment.

  • Internal and public-facing agents require different controls. Tools used for benefits, housing, or other resident services need more rigorous testing for accuracy, fairness, and bias.

  • The guidance identifies seven oversight roles, including an accountable owner, evaluation lead, security lead, transparency lead, responsible AI lead, contributor, and independent reviewer.

  • Agencies should maintain audit trails, collect stakeholder feedback, monitor potential harms, and continue evaluating agents after deployment.

State and local governments are beginning to test agentic AI in public health, housing, regulatory work, public assistance, and data analysis. Unlike conventional generative AI, which produces content in response to a prompt, an AI agent can plan a series of steps, use tools, make decisions, and take actions with some degree of autonomy. That expanded capability creates potential operational benefits, but it also requires agencies to reconsider how AI systems are designed, approved, and monitored.

A new Urban Institute playbook provides a framework for this. It recommends that agencies document the intended purpose of an AI agent, the population it will affect, the outcomes it should produce, and its expected cost before development begins. These details give technical teams, agency leaders, and reviewers a common basis for deciding whether the system is appropriate and for measuring its performance.

The guidance also distinguishes between internal and public-facing applications. An internal agent used to summarize operational data can often be tested on a limited scale. A resident-facing agent that supports disability claims, housing applications, or public benefits requires tighter controls because an incorrect recommendation may affect someone’s access to essential services. Agencies therefore need to define accuracy in practical terms, including whether outputs are complete, clear, consistent, and reliable when the underlying conditions change.

A system may achieve a high overall percentage of correct answers while performing poorly for a particular group. Agencies deploying public-facing agents should test whether recommendations vary across different populations. Human review alone does not resolve the problem if reviewers do not understand how the agent works or what information shaped its output.

The playbook proposes seven roles for governing these systems: an accountable owner, an evaluation lead, a security lead, a transparency lead, a responsible AI lead, a contributor, and an independent reviewer. Smaller governments may assign several responsibilities to the same person, but the underlying functions must still be covered. Agencies should know who approves the system, who tests it, who examines security risks, who documents its operation, and who can independently challenge its performance.

Early government projects show the range of possible applications. Virginia used agentic AI to review regulatory and guidance documents, reporting more than $1.4 billion in annual savings and an almost 80% reduction in permit and license processing times. Utah is evaluating an autonomous system for routine prescription renewals. Preliminary results indicated that renewal was recommended in 72% of cases, with physicians agreeing with 91% of those recommendations.

These results show why agencies are interested, but deployment cannot end with a successful pilot. Governments need audit trails, evaluation criteria, stakeholder feedback, and continuing monitoring for errors and harms. Agentic AI introduces more autonomy into public systems. Its value will therefore depend on whether agencies can establish clear responsibility for every decision the technology helps make.

QUICK HIT NEWS

  • The Administration for Children and Families is awarding $2.2 million to 14 jurisdictions to modernize the fingerprinting and background check systems used to license foster and kinship caregivers. Grants range from $54,000 to $200,000 and will fund digital live scan fingerprinting equipment, system assessments, and worker training. Recipients include Alaska, the District of Columbia, Mississippi, Nebraska, Oklahoma, Oregon, Utah, Vermont, Washington, and Wyoming, as well as several tribal governments.

  • Work for America has opened its Civic Match job platform to all job seekers interested in state and local government careers, expanding beyond its original focus on federal workers. About 14,000 people have used the platform since late 2024, and more than 500 have landed state or local roles. It now has 440 government partners in 49 states posting about 500 jobs a month. Nearly a quarter of posted jobs call for skills common in industries prone to layoffs, including data analytics, cybersecurity, and IT.

     

  • New York must replace an estimated 385,000 lead service lines by the federal deadline in 2037, but the state has no uniform policy determining who pays. Rochester plans to replace full service lines at no cost to homeowners by 2030 and has secured about $150 million from federal, state, and city sources. New York City generally treats the lines as privately owned and covers replacements only in selected areas, leaving some homeowners facing bills above $10,000. Statewide replacement could cost more than $4.8 billion.

  • Texas Tech, Old Dominion University and Tulane are moving online programs from 16-week classes to 8-week asynchronous courses designed for students balancing employment and family responsibilities. Old Dominion reported a 31.8% enrollment increase after the change, while Texas Tech’s online enrollment grew 55% over two years. The shift required universities to redesign courses, retrain faculty, shorten admissions timelines, and process financial aid every eight weeks. Old Dominion transitioned 892 courses and recorded 1,700 faculty training completions, though programs involving laboratories or fieldwork may still require longer terms.

FOR THE COMMUTE

Can a City Combat Homelessness with Improved Data (Priorities Podcast)

San Jose Mayor Matt Mahan discusses the city's new shelter system dashboard, which will track shelter use, length of stay, demographics, people housed, program exits and destinations, exits without notice, and participant deaths. Some city officials worry that the city still cannot reliably confirm whether people remain housed or sheltered 18 months after leaving its programs because post-exit information is stored in separate data systems for the city and Santa Clara County.

RESOURCES & EVENTS

From Automation to Insight (Virtual - September 30, 2026)

This webinar will examine how state and local agencies can use conversational AI to make infrastructure and operational data accessible to technical and nontechnical teams.

California Government Efficiency and Innovation Summit (Sacramento, CA - November 17, 2026)

California state and local government leaders, policymakers, technology officials, and industry representatives will meet to discuss how public agencies can improve efficiency while introducing new technology.

Report Spotlight: Michigan Public Policy Survey (University of Michigan)

A University of Michigan survey of 1,328 local governments found that 77% lack enough residents willing to run for elected office, while 75% frequently see uncontested races and struggle to fill appointed boards.

INSIGHT OF THE WEEK

Government forms lose applicants when they ask for sensitive information without saying why. At the FormFest 2026 session on designing for trust, Code for America researchers said Social Security numbers cause particular anxiety because of fraud and identity theft, and questions about language, location, income, and household size can also create hesitation. Their recommendations are to skip the number when it is not needed or make the field optional, explain why it is requested, and mask all but the last four digits. Agencies can use platform data to identify where applicants drop off, interviews to learn why, and post-submission feedback to identify wording that creates friction.