From pilot to proof: How revenue agencies can launch AI responsibly

SimpleImages via Getty Images
COMMENTARY | Leaders typically face constrained budgets, staffing gaps and complex fraud. AI can help address one of those issues, if it is tested properly and can work well under strain.
Shrinking tax revenues and persistent labor shortages are forcing states to stretch every dollar. Revenue agencies, in particular, are under pressure to close tax gaps, prevent fraud and modernize aging systems with fewer people and tighter budgets.
Artificial intelligence is emerging as part of the answer. Vermont’s Agency of Administration has identified AI tools for tax-fraud risk detection. Michigan has proposed new funding for AI-driven analytics within its Department of Treasury. Across the country, legislative activity is accelerating: states introduced more than 1,000 AI-related bills in 2025 alone, with dozens already enacted or enrolled.
States want AI’s benefits. They also want guardrails.
Against that backdrop, demand for practical AI pilots inside revenue agencies is rising quickly. But in today’s oversight environment, those pilots must be accurate, transparent and defensible from day one.
Start With the Operational Problem
In many agencies, the conversation begins with the tool. It should begin with the problem.
Revenue leaders are typically facing three realities at once: constrained budgets, persistent staffing gaps and increasingly complex fraud patterns. AI makes sense only if it directly addresses one of those pressures.
That might mean detecting duplicate refund claims, flagging reused Social Security numbers or stopping claims tied to deceased taxpayers. It might mean ensuring only eligible recipients receive refunds and that the approved amount is the exact amount issued. These are the problems revenue agencies are expected to catch before funds go out.
Clearly defining the operational objective does two things. It creates measurable outcomes for the pilot, and it gives leaders a plain-language explanation for legislators and auditors. In an environment where AI-related oversight is expanding rapidly, clarity is protection.
Build Validation Into the Pilot From the Beginning
Revenue systems are among the most sensitive environments in state government. Many were built to reliably move billions of dollars in an environment where even small errors carry public consequences.
As agencies introduce AI, they must plan for ongoing change. Fraud detection models will be updated. Tax engines will be patched. Surrounding applications will evolve. Each update can alter how data moves and how decisions are calculated. Without continuous testing, those changes can quietly affect outcomes.
In revenue operations, even a small drift can carry financial consequences. Overpayments may go unnoticed. Underpayments will not. A model may miss a new fraud pattern because it was not revalidated after an upgrade. A calculation may fall out of alignment with statute because a rule changed upstream.
Repeatable, documented testing catches these issues early. It validates fraud scenarios, confirms calculations and ensures data flows accurately from intake to payment. It also requires retesting after every significant update, not just at rollout.
Beyond being accurate, fraud systems must perform under peak demand, when most filings arrive just before the deadline. If models slow down or fail under heavy load, risk increases when transaction volume is highest.
Validation does not require ripping out existing platforms. For example, agencies using platforms such as GenTax can introduce AI-enabled capabilities around fraud detection, analytics or validation while still maintaining the reliability of the core tax administration system. That enables teams to test new use cases, measure outcomes and document performance before expanding AI more broadly.
Continuous validation turns AI from an experiment into a controlled, resilient capability.
Document Decision Logic Before You are Asked
People, including legislators, still have many questions about AI. That means scrutiny, especially when the technology influences financial outcomes.
Revenue agencies should expect to explain how their systems work: what data the model uses, how ris\k determinations are made, how false positives are tracked and where human review steps in.
During the pilot phase, agencies can strengthen their position by logging model inputs and outputs, documenting decision rules and maintaining clear audit trails of automated actions. Just as important is monitoring how AI systems behave over time and ensuring they stay within defined parameters.
Human accountability remains essential, particularly in enforcement and financial determinations. Clear escalation paths for flagged cases reinforce that AI supports decision-making rather than replacing it.
Transparency can build credibility with lawmakers and the public.
From Experimentation to Public Confidence
AI can help revenue agencies operate more effectively with limited staff and tighter budgets. It can strengthen fraud detection, improve validation and reduce manual review work. In constrained environments, those gains matter.
In revenue, the system has to do more than run. It has to stand up to questions.
In revenue, financial errors become public problems quickly. Legislators will ask how the system works. Auditors will look for documentation. Agencies will need clear answers.
That is why the pilot phase matters. It is where agencies prove not only that a system functions, but that it is tested, monitored and governed. Clear objectives, continuous validation and documented decision logic are part of operating responsibly in a high-visibility environment.
When AI is introduced deliberately — with validation and oversight built in — it becomes a tool agencies can rely on. More importantly, it becomes a system they can explain.
Allan Troup is senior director of public sector SLED at Tricentis. He has dedicated over 20 years to improving state and local government operations, and is a dedicated advocate for ensuring taxpayer dollars are used effectively. He focuses on equipping states with testing solutions that minimize organizational effort, enhance outcomes and deliver measurable savings.
NEXT STORY: New York Is struggling to track its own AI use




