AI for Federal and National Government
National agencies hold the largest document backlogs in any sector and the strictest obligations about what may be decided automatically.
National government runs on documents and queues, which makes the highest-value AI work here unglamorous, measurable, and entirely separate from the decisions that carry the regulatory weight.
AI for federal and national government covers case file and document extraction, backlog triage and routing, citizen enquiry handling, translation and accessibility, demand and workload forecasting, and the impact assessment, oversight and appeal design required where a system informs decisions about individuals.
The obligations that now frame a federal build
OMB memoranda M-25-21 and M-25-22, both issued on 3 April 2025, replaced the previous federal AI policy. They are more permissive in tone about adoption and specific about what high-impact systems require.
- High-impact is defined by consequence. Where the output is a principal basis for decisions or actions with a legal, material, binding or significant effect on rights or safety.
- Pre-deployment testing is a minimum practice. Not a launch check — a documented assessment before the system touches real cases.
- Impact assessments run before and during deployment. Which means the assessment is a live artefact, not a procurement deliverable filed once.
- Human oversight must include intervention. An official who can see the output and change the outcome, not one who countersigns it.
- Remedies or appeals for affected individuals. Named explicitly as a minimum practice, and the requirement that most often changes the design.
Procurement is where this is actually enforced
M-25-22 governs how agencies buy AI, covering data portability, protection against vendor lock-in and performance monitoring across the contract lifecycle. In practice the procurement terms bind earlier and harder than the policy does: an agency that cannot extract its own data from a vendor, or cannot monitor whether a model's performance has degraded, has a problem no governance board can resolve after signature. Getting these terms right at procurement is cheaper than any remediation afterwards.
Backlogs, which are the sector's defining operational problem
Extraction before triage
Most national backlogs are document backlogs: applications, correspondence, evidence, forms. Turning them into structured data is the prerequisite for everything else and touches no decision about anyone.
Ordering a queue is defensible; removing someone from it is not
Prioritising cases by urgency, completeness or statutory deadline is administrative efficiency. Closing or rejecting a case automatically is a decision about a person with an entitlement attached, and the two should not be built as one feature.
Completeness checking returns immediately
A large share of case delay is applications missing information, discovered weeks after submission. Checking at intake and telling the applicant what is missing removes delay for the citizen and the caseworker simultaneously.
Measure the queue, not the model
Time to decision, cases awaiting information, and rework rate are the numbers a committee will ask about. Model accuracy is not, and a project reported in model metrics tends to lose its funding. See intelligent document processing.
The explanation problem, which precedes every AI statute
Administrative law generally requires that a decision affecting a person can be explained to them and challenged. This constraint is older than every AI framework, and it decides more federal projects than any of them.
| System | Position | Note |
|---|---|---|
| Document extraction | Clear | Produces facts a caseworker verifies; no decision made |
| Queue prioritisation | Clear | Orders work; every case still reaches a person |
| Completeness checking | Clear | Tells the applicant what is missing; helps rather than decides |
| Eligibility pre-assessment | Careful | Must be presented as indicative, never as a determination |
| Risk flagging for investigation | Careful | The flag is not a finding; what happens next is the whole design |
| Automated determination | Rarely defensible | Requires an explanation and appeal route most systems cannot supply |
If it cannot go in the letter, it cannot go in the decision
The practical test we apply is simple: can an official write the reason into a decision letter in terms the recipient could challenge? A model output that reduces to a score with no articulable basis fails that test regardless of its accuracy, and no amount of governance documentation fixes it. This is why we frequently recommend a simpler, more transparent method in government work — not because it performs better, but because it produces something that can lawfully be acted on.
How an engagement runs
The decision boundary and the appeal route settled before anything is built.
Scope, classification and impact assessment
Whether the system is high-impact or high-risk, what assessment is required, and where the human decision sits.
Data assessment
Case file formats, decision record quality, and whether outcomes can be linked to cases reliably.
Build
Extraction and triage, or enquiry handling, with explanation and appeal designed in rather than added.
Trial and testing
On live caseload against caseworker judgement, with outcome testing across affected groups documented.
Operation
Performance monitored against the contract, assessments updated, appeal volumes tracked as a quality signal.
What you receive
Backlogs moving, and systems that can be explained to the person they affect.
Case file extraction
Structured data from applications, correspondence and evidence, verified by caseworkers.
Backlog triage
Queues ordered by urgency, completeness and statutory deadline, with every case still reaching a person.
Intake completeness checking
Missing information identified at submission rather than discovered weeks later.
Citizen enquiry handling
Service and process navigation, with entitlement questions routed to an official.
Impact assessment
Documented to the standard the applicable regime requires, kept live rather than filed.
Explanation and appeal design
Reasons an official can put in a letter, and a route the person can actually use.
Is this the right starting point?
Worth being direct. There are situations in federal and national government where custom AI work is the wrong spend, and those are listed rather than buried.
Worth doing if
- Document backlogs are the binding constraint on service delivery.
- A large share of delay is caused by incomplete applications found late.
- Enquiry volume consumes caseworker time that should go to decisions.
- You need an impact assessment done properly rather than retrospectively.
- A vendor system is in place and you cannot monitor whether it still performs.
Do something else if
- You want automated determinations without an explanation or appeal route.
- Decision records cannot be linked to cases and outcome analysis is impossible.
- The procurement is structured so you cannot extract your own data later.
- The intended system falls under a prohibited practice and there is pressure to proceed anyway.
Frequently asked questions
Marked up with FAQPage schema so these answers can surface directly in search results and inside AI assistant responses.
What counts as high-impact under the federal policy?
Where the AI output is a principal basis for decisions or actions with a legal, material, binding or significant effect on rights or safety. That framing turns on consequence rather than on technique, so a simple rules-based scoring system informing a benefits decision can be high-impact while a sophisticated model extracting text from forms is not. The minimum practices that follow are pre-deployment testing, impact assessments before and during deployment, human oversight with genuine intervention capability, remedies or appeals for affected individuals, and ceasing use of non-compliant systems.
Can we automate decisions to clear a backlog?
You can automate almost everything around the decision and should be cautious about the decision itself. Extraction, completeness checking, queue ordering, evidence assembly and drafting all move throughput substantially and touch nobody's entitlement. An automated determination requires an explanation an official can put in a decision letter and an appeal route the person can use, and most systems proposed for this cannot supply either. In our experience the surrounding automation delivers most of the backlog gain without the exposure, which is why we scope it first.
Why do you recommend simpler models for government work?
Because the output has to survive being written into a decision letter. Administrative law generally requires that a decision affecting a person can be explained and challenged, and a score with no articulable basis fails that test whatever its accuracy. A transparent method that says which factors drove the outcome, and by how much, produces something an official can lawfully act on. This is one of the few settings where we actively recommend the less accurate approach, and it changes the technical design rather than just the documentation.
How should we handle the procurement side?
Treat it as the real control point, because it binds earlier than policy does. M-25-22 covers data portability, protection against vendor lock-in and performance monitoring through the contract lifecycle, and those are the terms that determine whether you can act later. An agency unable to extract its own data from a vendor system, or unable to tell whether the model has degraded since acceptance, has a problem no governance board can fix after signature. We would review the terms before the pilot, not after.
What should we build first?
Case file extraction and intake completeness checking. Both are high volume, directly measurable in time-to-decision, and neither makes a decision about anyone — which means they deliver while the harder governance questions are being resolved properly. Completeness checking in particular has an unusual property: it improves the citizen's experience and the caseworker's throughput at the same time, because an application returned at intake is far cheaper than one abandoned halfway through assessment.
Related verticals
Organisations in federal and national government usually share data, buyers or regulators with these. All fourteen are listed on the Public Sector & Education page.
AI for State and Local Government
Service request triage, planning and licensing, and asset condition — with a clear line around social care prediction.
Read more →AI for Tax and Revenue Agencies
Return processing, enquiry handling and compliance risk — designed around the appeal route from the start.
Read more →AI for Judiciary and Court Services
Listing, transcription and case file preparation — with judicial decision support left where it belongs.
Read more →Tell us what the problem looks like.
Thirty minutes, no charge, no deck. We will tell you whether this is an AI problem, a data problem, or a process problem — and we will say when the honest answer is to buy something rather than build it.