AI for Tax and Revenue Agencies
A revenue agency is where algorithmic selection meets people who cannot decline the relationship, which is exactly where it has gone worst.
Tax administration is the clearest case in this sector of a service nobody opts into, which is why the design of what happens after a risk flag matters more than the accuracy of the flag itself.
AI for tax and revenue agencies covers return and document processing, taxpayer enquiry handling and guidance navigation, compliance risk flagging for human investigation, debt and payment analytics, customs classification support, and the appeal and explanation design that makes any of it defensible.
The Dutch lesson, in detail
The childcare benefits scandal is the reference case for this vertical and it is usually summarised too quickly. Roughly 26,000 families were wrongly accused of fraud between 2005 and 2019 and the third Rutte cabinet resigned over it on 15 January 2021.
- The selection was the start, not the harm. Being flagged for investigation is survivable; what followed was not.
- Consequences applied before findings. Payments stopped and repayment demands issued while the matter was still, in substance, an allegation.
- Nobody could explain a decision. Families asking why could not be told, which made challenge impossible in practice as well as in fact.
- Nationality and dual citizenship received structural attention in the administration's handling, which the data protection authority found unnecessary and improper.
- The appeal route did not function. This is the failure that turned a flawed process into a national scandal, and it is a design question rather than a policy one.
Design the appeal first, then the model
The single most useful thing a revenue agency can do before building a compliance risk system is to design what happens to a person who is flagged wrongly: how they find out, what they are told, who they can ask, how long it takes, and what does not happen to them in the meantime. If that route is weak, the model makes it worse at scale no matter how accurate it is. We treat this as the first deliverable rather than a governance annexe, because it determines what the system is allowed to trigger.
Compliance risk, designed so a flag stays a flag
A flag is a request for attention, not a finding
The only defensible architecture routes a flagged case to a human investigator who reaches a conclusion. Nothing consequential — suspension, demand, enforcement — should follow from the model output alone.
Selection bias makes validation circular
Your compliance outcome data reflects who you previously investigated. A model trained on it learns your historical selection, and confirms itself through the cases it generates. Random-sample audit programmes are the only clean source of validation data, and agencies that run them have a real advantage.
Proxies carry the discrimination risk
Nationality, ethnicity and related characteristics need not be inputs for a model to select on them. Geography, name patterns, family composition and income source all correlate, and outcome testing across groups is the only way to know.
Explainability is a legal requirement here
A taxpayer is generally entitled to know why they are being examined. A score with no articulable basis cannot support that, which is why transparent methods win in this setting even at some cost in accuracy.
The uncontroversial work, which is most of the value
| Application | Fit | Note |
|---|---|---|
| Return and document processing | Strong | Very high volume, structured, and no decision about anyone |
| Taxpayer enquiry handling | Strong | Guidance navigation is genuinely hard for people; must never state a liability |
| Guidance and legislation retrieval | Strong | For staff, with dates and legislative period attached |
| Debt and payment analytics | Good | Predicting who will pay with support beats enforcement escalation |
| Customs classification support | Good | Suggests a code for an officer; duty consequences require human confirmation |
| Demand and workload forecasting | Strong | Aggregate resourcing; no individual involved |
| Compliance risk flagging | With care | Only where the appeal route and human investigation are designed first |
| Automated assessment or suspension | We decline | This is the Dutch failure, mechanised |
Payment support prediction beats enforcement prediction
Predicting who is likely to fall behind so that an early, easy payment arrangement can be offered collects more revenue than predicting who to escalate against, and it does so with a fraction of the harm and the political exposure. It is also a far easier system to defend: the output triggers an offer a person can accept or ignore, rather than an action taken against them. Agencies that have compared the two approaches on their own data have generally been surprised by the collection difference.
How an engagement runs
The appeal route designed before the model, not alongside it.
Scope, appeal design and classification
What a wrongly flagged person experiences, and what Annex III and Article 27 require.
Data assessment
Whether compliance outcomes reflect selection, and whether random-audit data exists for validation.
Build
Document processing and enquiry handling, or risk flagging with human investigation designed in.
Trial and outcome testing
Against investigator judgement, with selection rates across groups measured and documented.
Operation
Outcome testing on a schedule, appeal volumes tracked as a quality signal, explanations audited.
What you receive
Processing volume handled, and any selection system built so a person can challenge it.
Return and document processing
High-volume extraction with arithmetic reconciliation and exception-only review.
Taxpayer enquiry handling
Guidance navigation that never states a liability or an entitlement as a decision.
Compliance risk flagging
Routed to human investigation, with nothing consequential triggered by the model alone.
Appeal and explanation design
What a wrongly flagged person is told, by whom, how fast, and what does not happen meanwhile.
Outcome testing harness
Selection rates across groups, run on a schedule and retained.
Payment support prediction
Early arrangements offered rather than enforcement escalated.
Is this the right starting point?
Worth being direct. There are situations in tax and revenue agencies where custom AI work is the wrong spend, and those are listed rather than buried.
Worth doing if
- Return and document processing volume is a resourcing constraint.
- Taxpayer enquiry volume is high and guidance is genuinely hard to navigate.
- A compliance risk model exists and nobody can explain a selection to a taxpayer.
- You run a random-sample audit programme and have never used it to validate selection.
- Debt recovery is escalation-led and early support has never been modelled.
Do something else if
- You want model output to trigger suspension, demand or enforcement without human investigation.
- The appeal route is known to be weak and there is no appetite to fix it first.
- Compliance outcome data reflects selection and no clean validation source exists or is planned.
- Outcome testing across groups is not permitted.
Frequently asked questions
Marked up with FAQPage schema so these answers can surface directly in search results and inside AI assistant responses.
Is compliance risk modelling defensible at all?
Yes, provided a flag stays a flag. The defensible architecture routes a flagged case to a human investigator who reaches a conclusion, with nothing consequential — suspension, demand, enforcement — following from the model output alone. The Dutch childcare benefits scandal, roughly 26,000 families wrongly accused with a cabinet resigning in January 2021, was not primarily a story about a bad model: consequences applied before findings, nobody could explain a decision, and the appeal route did not function. Fix those three things first and the model becomes a resourcing tool rather than a hazard.
How do we validate a selection model when our data reflects past selection?
Through random-sample audit data, which is the only clean source. Compliance outcomes record who you previously investigated and what you found, so a model trained on them learns your historical selection and then confirms itself through the cases it generates — the validation is circular. Agencies running random enquiry programmes have a genuine advantage here and frequently have not connected the two. If no such programme exists, establishing one is the prerequisite rather than an enhancement, and we would say so before building.
Our model doesn't use nationality. Is that sufficient?
No. Geography, name patterns, family composition, income source and language of correspondence all correlate with nationality and ethnicity, so a model with a clean input list can still select disproportionately. The Dutch data protection authority found structural and unnecessary attention to nationality and dual citizenship in that administration's handling, and the lesson generalises: the only way to know what your system does is to measure selection rates across groups, on a schedule, and to document the results. That requires demographic data you may not hold, which is itself a project.
What should we build if not risk models?
Return and document processing first — very high volume, structured, and touching no decision about anyone. Then taxpayer enquiry handling, because tax guidance is genuinely hard for people to navigate and the contact volume reflects that, with the strict constraint that the system explains process and never states a liability. Both deliver measurably while the harder governance questions around selection are settled properly, which is the right sequence rather than a delaying tactic.
Is there a better use of prediction than enforcement?
Predicting who is likely to fall behind, so an early and easy payment arrangement can be offered. It generally collects more than predicting who to escalate against, it causes a fraction of the harm, and it is far easier to defend because the output triggers an offer a person can accept or ignore rather than an action taken against them. Agencies that have compared the two on their own data have usually been surprised by the collection difference, and it changes the political profile of the whole programme.
Related verticals
Organisations in tax and revenue agencies usually share data, buyers or regulators with these. All fourteen are listed on the Public Sector & Education page.
AI for Federal and National Government
Case extraction, backlog triage and enquiry handling — built to the impact assessment and appeal standard from the start.
Read more →AI for State and Local Government
Service request triage, planning and licensing, and asset condition — with a clear line around social care prediction.
Read more →AI for Judiciary and Court Services
Listing, transcription and case file preparation — with judicial decision support left where it belongs.
Read more →Tell us what the problem looks like.
Thirty minutes, no charge, no deck. We will tell you whether this is an AI problem, a data problem, or a process problem — and we will say when the honest answer is to buy something rather than build it.