EU AI Act transparency duties apply now; high-risk duties from December 2027. Check your exposure
Insights About us Careers
Contact us
Public Sector & Education

AI for Federal and National Government

National agencies hold the largest document backlogs in any sector and the strictest obligations about what may be decided automatically.

Tier 1
Our depth here
3 Apr 2025
M-25-21 and M-25-22
Appeal
A minimum practice

National government runs on documents and queues, which makes the highest-value AI work here unglamorous, measurable, and entirely separate from the decisions that carry the regulatory weight.

In one paragraph

AI for federal and national government covers case file and document extraction, backlog triage and routing, citizen enquiry handling, translation and accessibility, demand and workload forecasting, and the impact assessment, oversight and appeal design required where a system informs decisions about individuals.

The obligations that now frame a federal build

OMB memoranda M-25-21 and M-25-22, both issued on 3 April 2025, replaced the previous federal AI policy. They are more permissive in tone about adoption and specific about what high-impact systems require.

  • High-impact is defined by consequence. Where the output is a principal basis for decisions or actions with a legal, material, binding or significant effect on rights or safety.
  • Pre-deployment testing is a minimum practice. Not a launch check — a documented assessment before the system touches real cases.
  • Impact assessments run before and during deployment. Which means the assessment is a live artefact, not a procurement deliverable filed once.
  • Human oversight must include intervention. An official who can see the output and change the outcome, not one who countersigns it.
  • Remedies or appeals for affected individuals. Named explicitly as a minimum practice, and the requirement that most often changes the design.
Worth knowing

Procurement is where this is actually enforced

M-25-22 governs how agencies buy AI, covering data portability, protection against vendor lock-in and performance monitoring across the contract lifecycle. In practice the procurement terms bind earlier and harder than the policy does: an agency that cannot extract its own data from a vendor, or cannot monitor whether a model's performance has degraded, has a problem no governance board can resolve after signature. Getting these terms right at procurement is cheaper than any remediation afterwards.

Backlogs, which are the sector's defining operational problem

Extraction before triage

Most national backlogs are document backlogs: applications, correspondence, evidence, forms. Turning them into structured data is the prerequisite for everything else and touches no decision about anyone.

Ordering a queue is defensible; removing someone from it is not

Prioritising cases by urgency, completeness or statutory deadline is administrative efficiency. Closing or rejecting a case automatically is a decision about a person with an entitlement attached, and the two should not be built as one feature.

Completeness checking returns immediately

A large share of case delay is applications missing information, discovered weeks after submission. Checking at intake and telling the applicant what is missing removes delay for the citizen and the caseworker simultaneously.

Measure the queue, not the model

Time to decision, cases awaiting information, and rework rate are the numbers a committee will ask about. Model accuracy is not, and a project reported in model metrics tends to lose its funding. See intelligent document processing.

The explanation problem, which precedes every AI statute

Administrative law generally requires that a decision affecting a person can be explained to them and challenged. This constraint is older than every AI framework, and it decides more federal projects than any of them.

SystemPositionNote
Document extractionClearProduces facts a caseworker verifies; no decision made
Queue prioritisationClearOrders work; every case still reaches a person
Completeness checkingClearTells the applicant what is missing; helps rather than decides
Eligibility pre-assessmentCarefulMust be presented as indicative, never as a determination
Risk flagging for investigationCarefulThe flag is not a finding; what happens next is the whole design
Automated determinationRarely defensibleRequires an explanation and appeal route most systems cannot supply
Worth knowing

If it cannot go in the letter, it cannot go in the decision

The practical test we apply is simple: can an official write the reason into a decision letter in terms the recipient could challenge? A model output that reduces to a score with no articulable basis fails that test regardless of its accuracy, and no amount of governance documentation fixes it. This is why we frequently recommend a simpler, more transparent method in government work — not because it performs better, but because it produces something that can lawfully be acted on.

Process

How an engagement runs

The decision boundary and the appeal route settled before anything is built.

Weeks 1 to 4

Scope, classification and impact assessment

Whether the system is high-impact or high-risk, what assessment is required, and where the human decision sits.

Weeks 5 to 8

Data assessment

Case file formats, decision record quality, and whether outcomes can be linked to cases reliably.

Weeks 9 to 16

Build

Extraction and triage, or enquiry handling, with explanation and appeal designed in rather than added.

Weeks 17 to 22

Trial and testing

On live caseload against caseworker judgement, with outcome testing across affected groups documented.

Ongoing

Operation

Performance monitored against the contract, assessments updated, appeal volumes tracked as a quality signal.

Deliverables

What you receive

Backlogs moving, and systems that can be explained to the person they affect.

01

Case file extraction

Structured data from applications, correspondence and evidence, verified by caseworkers.

02

Backlog triage

Queues ordered by urgency, completeness and statutory deadline, with every case still reaching a person.

03

Intake completeness checking

Missing information identified at submission rather than discovered weeks later.

04

Citizen enquiry handling

Service and process navigation, with entitlement questions routed to an official.

05

Impact assessment

Documented to the standard the applicable regime requires, kept live rather than filed.

06

Explanation and appeal design

Reasons an official can put in a letter, and a route the person can actually use.

Fit check

Is this the right starting point?

Worth being direct. There are situations in federal and national government where custom AI work is the wrong spend, and those are listed rather than buried.

Worth doing if

  • Document backlogs are the binding constraint on service delivery.
  • A large share of delay is caused by incomplete applications found late.
  • Enquiry volume consumes caseworker time that should go to decisions.
  • You need an impact assessment done properly rather than retrospectively.
  • A vendor system is in place and you cannot monitor whether it still performs.

Do something else if

  • You want automated determinations without an explanation or appeal route.
  • Decision records cannot be linked to cases and outcome analysis is impossible.
  • The procurement is structured so you cannot extract your own data later.
  • The intended system falls under a prohibited practice and there is pressure to proceed anyway.
Questions

Frequently asked questions

Marked up with FAQPage schema so these answers can surface directly in search results and inside AI assistant responses.

What counts as high-impact under the federal policy?

Where the AI output is a principal basis for decisions or actions with a legal, material, binding or significant effect on rights or safety. That framing turns on consequence rather than on technique, so a simple rules-based scoring system informing a benefits decision can be high-impact while a sophisticated model extracting text from forms is not. The minimum practices that follow are pre-deployment testing, impact assessments before and during deployment, human oversight with genuine intervention capability, remedies or appeals for affected individuals, and ceasing use of non-compliant systems.

Can we automate decisions to clear a backlog?

You can automate almost everything around the decision and should be cautious about the decision itself. Extraction, completeness checking, queue ordering, evidence assembly and drafting all move throughput substantially and touch nobody's entitlement. An automated determination requires an explanation an official can put in a decision letter and an appeal route the person can use, and most systems proposed for this cannot supply either. In our experience the surrounding automation delivers most of the backlog gain without the exposure, which is why we scope it first.

Why do you recommend simpler models for government work?

Because the output has to survive being written into a decision letter. Administrative law generally requires that a decision affecting a person can be explained and challenged, and a score with no articulable basis fails that test whatever its accuracy. A transparent method that says which factors drove the outcome, and by how much, produces something an official can lawfully act on. This is one of the few settings where we actively recommend the less accurate approach, and it changes the technical design rather than just the documentation.

How should we handle the procurement side?

Treat it as the real control point, because it binds earlier than policy does. M-25-22 covers data portability, protection against vendor lock-in and performance monitoring through the contract lifecycle, and those are the terms that determine whether you can act later. An agency unable to extract its own data from a vendor system, or unable to tell whether the model has degraded since acceptance, has a problem no governance board can fix after signature. We would review the terms before the pilot, not after.

What should we build first?

Case file extraction and intake completeness checking. Both are high volume, directly measurable in time-to-decision, and neither makes a decision about anyone — which means they deliver while the harder governance questions are being resolved properly. Completeness checking in particular has an unusual property: it improves the citizen's experience and the caseworker's throughput at the same time, because an application returned at intake is far cheaper than one abandoned halfway through assessment.

Tell us what the problem looks like.

Thirty minutes, no charge, no deck. We will tell you whether this is an AI problem, a data problem, or a process problem — and we will say when the honest answer is to buy something rather than build it.