EU AI Act transparency duties apply now; high-risk duties from December 2027. Check your exposure
Insights About us Careers
Contact us
Public Sector & Education

AI for Tax and Revenue Agencies

A revenue agency is where algorithmic selection meets people who cannot decline the relationship, which is exactly where it has gone worst.

Tier 3
Our depth here
26,000 families
The Dutch case
Flag
Not a finding

Tax administration is the clearest case in this sector of a service nobody opts into, which is why the design of what happens after a risk flag matters more than the accuracy of the flag itself.

In one paragraph

AI for tax and revenue agencies covers return and document processing, taxpayer enquiry handling and guidance navigation, compliance risk flagging for human investigation, debt and payment analytics, customs classification support, and the appeal and explanation design that makes any of it defensible.

The Dutch lesson, in detail

The childcare benefits scandal is the reference case for this vertical and it is usually summarised too quickly. Roughly 26,000 families were wrongly accused of fraud between 2005 and 2019 and the third Rutte cabinet resigned over it on 15 January 2021.

  • The selection was the start, not the harm. Being flagged for investigation is survivable; what followed was not.
  • Consequences applied before findings. Payments stopped and repayment demands issued while the matter was still, in substance, an allegation.
  • Nobody could explain a decision. Families asking why could not be told, which made challenge impossible in practice as well as in fact.
  • Nationality and dual citizenship received structural attention in the administration's handling, which the data protection authority found unnecessary and improper.
  • The appeal route did not function. This is the failure that turned a flawed process into a national scandal, and it is a design question rather than a policy one.
Worth knowing

Design the appeal first, then the model

The single most useful thing a revenue agency can do before building a compliance risk system is to design what happens to a person who is flagged wrongly: how they find out, what they are told, who they can ask, how long it takes, and what does not happen to them in the meantime. If that route is weak, the model makes it worse at scale no matter how accurate it is. We treat this as the first deliverable rather than a governance annexe, because it determines what the system is allowed to trigger.

Compliance risk, designed so a flag stays a flag

A flag is a request for attention, not a finding

The only defensible architecture routes a flagged case to a human investigator who reaches a conclusion. Nothing consequential — suspension, demand, enforcement — should follow from the model output alone.

Selection bias makes validation circular

Your compliance outcome data reflects who you previously investigated. A model trained on it learns your historical selection, and confirms itself through the cases it generates. Random-sample audit programmes are the only clean source of validation data, and agencies that run them have a real advantage.

Proxies carry the discrimination risk

Nationality, ethnicity and related characteristics need not be inputs for a model to select on them. Geography, name patterns, family composition and income source all correlate, and outcome testing across groups is the only way to know.

A taxpayer is generally entitled to know why they are being examined. A score with no articulable basis cannot support that, which is why transparent methods win in this setting even at some cost in accuracy.

The uncontroversial work, which is most of the value

ApplicationFitNote
Return and document processingStrongVery high volume, structured, and no decision about anyone
Taxpayer enquiry handlingStrongGuidance navigation is genuinely hard for people; must never state a liability
Guidance and legislation retrievalStrongFor staff, with dates and legislative period attached
Debt and payment analyticsGoodPredicting who will pay with support beats enforcement escalation
Customs classification supportGoodSuggests a code for an officer; duty consequences require human confirmation
Demand and workload forecastingStrongAggregate resourcing; no individual involved
Compliance risk flaggingWith careOnly where the appeal route and human investigation are designed first
Automated assessment or suspensionWe declineThis is the Dutch failure, mechanised
Worth knowing

Payment support prediction beats enforcement prediction

Predicting who is likely to fall behind so that an early, easy payment arrangement can be offered collects more revenue than predicting who to escalate against, and it does so with a fraction of the harm and the political exposure. It is also a far easier system to defend: the output triggers an offer a person can accept or ignore, rather than an action taken against them. Agencies that have compared the two approaches on their own data have generally been surprised by the collection difference.

Process

How an engagement runs

The appeal route designed before the model, not alongside it.

Weeks 1 to 4

Scope, appeal design and classification

What a wrongly flagged person experiences, and what Annex III and Article 27 require.

Weeks 5 to 8

Data assessment

Whether compliance outcomes reflect selection, and whether random-audit data exists for validation.

Weeks 9 to 16

Build

Document processing and enquiry handling, or risk flagging with human investigation designed in.

Weeks 17 to 22

Trial and outcome testing

Against investigator judgement, with selection rates across groups measured and documented.

Ongoing

Operation

Outcome testing on a schedule, appeal volumes tracked as a quality signal, explanations audited.

Deliverables

What you receive

Processing volume handled, and any selection system built so a person can challenge it.

01

Return and document processing

High-volume extraction with arithmetic reconciliation and exception-only review.

02

Taxpayer enquiry handling

Guidance navigation that never states a liability or an entitlement as a decision.

03

Compliance risk flagging

Routed to human investigation, with nothing consequential triggered by the model alone.

04

Appeal and explanation design

What a wrongly flagged person is told, by whom, how fast, and what does not happen meanwhile.

05

Outcome testing harness

Selection rates across groups, run on a schedule and retained.

06

Payment support prediction

Early arrangements offered rather than enforcement escalated.

Fit check

Is this the right starting point?

Worth being direct. There are situations in tax and revenue agencies where custom AI work is the wrong spend, and those are listed rather than buried.

Worth doing if

  • Return and document processing volume is a resourcing constraint.
  • Taxpayer enquiry volume is high and guidance is genuinely hard to navigate.
  • A compliance risk model exists and nobody can explain a selection to a taxpayer.
  • You run a random-sample audit programme and have never used it to validate selection.
  • Debt recovery is escalation-led and early support has never been modelled.

Do something else if

  • You want model output to trigger suspension, demand or enforcement without human investigation.
  • The appeal route is known to be weak and there is no appetite to fix it first.
  • Compliance outcome data reflects selection and no clean validation source exists or is planned.
  • Outcome testing across groups is not permitted.
Questions

Frequently asked questions

Marked up with FAQPage schema so these answers can surface directly in search results and inside AI assistant responses.

Is compliance risk modelling defensible at all?

Yes, provided a flag stays a flag. The defensible architecture routes a flagged case to a human investigator who reaches a conclusion, with nothing consequential — suspension, demand, enforcement — following from the model output alone. The Dutch childcare benefits scandal, roughly 26,000 families wrongly accused with a cabinet resigning in January 2021, was not primarily a story about a bad model: consequences applied before findings, nobody could explain a decision, and the appeal route did not function. Fix those three things first and the model becomes a resourcing tool rather than a hazard.

How do we validate a selection model when our data reflects past selection?

Through random-sample audit data, which is the only clean source. Compliance outcomes record who you previously investigated and what you found, so a model trained on them learns your historical selection and then confirms itself through the cases it generates — the validation is circular. Agencies running random enquiry programmes have a genuine advantage here and frequently have not connected the two. If no such programme exists, establishing one is the prerequisite rather than an enhancement, and we would say so before building.

Our model doesn't use nationality. Is that sufficient?

No. Geography, name patterns, family composition, income source and language of correspondence all correlate with nationality and ethnicity, so a model with a clean input list can still select disproportionately. The Dutch data protection authority found structural and unnecessary attention to nationality and dual citizenship in that administration's handling, and the lesson generalises: the only way to know what your system does is to measure selection rates across groups, on a schedule, and to document the results. That requires demographic data you may not hold, which is itself a project.

What should we build if not risk models?

Return and document processing first — very high volume, structured, and touching no decision about anyone. Then taxpayer enquiry handling, because tax guidance is genuinely hard for people to navigate and the contact volume reflects that, with the strict constraint that the system explains process and never states a liability. Both deliver measurably while the harder governance questions around selection are settled properly, which is the right sequence rather than a delaying tactic.

Is there a better use of prediction than enforcement?

Predicting who is likely to fall behind, so an early and easy payment arrangement can be offered. It generally collects more than predicting who to escalate against, it causes a fraction of the harm, and it is far easier to defend because the output triggers an offer a person can accept or ignore rather than an action taken against them. Agencies that have compared the two on their own data have usually been surprised by the collection difference, and it changes the political profile of the whole programme.

Tell us what the problem looks like.

Thirty minutes, no charge, no deck. We will tell you whether this is an AI problem, a data problem, or a process problem — and we will say when the honest answer is to buy something rather than build it.