Document Summarisation
Summaries produced for a defined reader making a defined decision, checked against the source for anything invented and, more importantly, for anything consequential left out.
Summarisation demonstrates beautifully and deploys badly, because the demo is judged on whether the summary reads well and production is judged on whether anything important was left out. Those are different questions, and only one of them is hard.
Document summarisation is the automated production of shorter text that conveys the content of a longer source, built for a specific audience and purpose, with faithfulness to the source verified and omission of material facts treated as the primary failure mode.
Omission is the failure that matters
Fabrication is the well-known risk and the easier one to test for: a claim in the summary that does not appear in the source is detectable. Omission is harder and more dangerous, because a fluent, accurate summary that silently drops the one clause your reader needed looks exactly like a good summary.
So we define, with your experts, what must never be omitted for each document type: the termination clause, the adverse finding, the deadline, the caveat on the recommendation. Those become explicit checks rather than hopes, and the summary is evaluated on coverage of them as well as on faithfulness.
Write the summary for a named reader
'Summarise this document' has no correct answer. A summary of the same contract for a lawyer, a procurement manager and an operations lead should contain different material. We specify the reader and the decision they are making, which turns an unanswerable request into a testable one.
How we build summarisation
Define the summary contract
Length, structure, required elements, prohibited content and tone, per document type and per reader. A structured summary with defined sections is easier to evaluate, easier to scan and considerably more useful than free-form prose, and it makes coverage measurable.
Handle length properly
Long documents exceed what fits usefully in one pass, and naive chunk-and-merge produces summaries that repeat themselves and lose the thread. Hierarchical summarisation with section awareness, or retrieval of the relevant sections when the summary has a specific purpose, both work; which one depends on the document and the reader.
Cite back to the source
Every statement links to the passage it came from, so a reader can verify in one click. This is what makes a summary usable for anything consequential, and it changes the summary from something to be trusted into something to be checked quickly.
Check faithfulness and coverage automatically
Claim-level verification against the source for anything unsupported, plus coverage checks for the required elements. Both run on every summary, and both are scored on a set your experts have reviewed.
Preserve the hedges
Summaries have a tendency to convert 'the evidence suggests, subject to further testing' into 'the evidence shows'. Uncertainty, conditionality and qualification are content, not padding, and we test explicitly for their removal.
Evaluate against your experts
Your specialists write reference summaries for a sample and review the system's output. Their judgement sets the bar, and their disagreements become the evaluation set that gates every subsequent change, following the practice in LLM evaluation.
Where summarisation is worth it
| Setting | Reader and decision | What must never be omitted |
|---|---|---|
| Contract review | Legal, deciding whether to sign | Liability, termination, change of control |
| Case and claim files | Handler, deciding next action | Prior decisions, deadlines, vulnerability flags |
| Research and reports | Analyst, deciding what to read fully | Method limitations, sample size, conflicts |
| Meeting and call records | Attendees, deciding follow-up | Commitments made, owners, dates |
| Clinical documentation | Clinician, deciding care | Allergies, adverse events, contraindications |
| Regulatory correspondence | Compliance, deciding response | Deadlines, required actions, penalties |
In every row the third column is the design specification. Where nobody can fill it in, the summary has no testable definition of correct and we would rather establish that before building.
How the engagement runs
The required-elements list is agreed first; it is what makes the output testable.
Reader, purpose and contract
Audience and decision defined per document type, summary structure and required elements agreed with your experts.
Reference set
Experts write reference summaries for a sample, and the evaluation criteria and coverage checks are built from them.
Summarisation pipeline
Length handling, section awareness, citation back to source, hedge preservation, scored against the reference set.
Verification and review
Automated faithfulness and coverage checks, low-confidence routing, review interface with source highlighting.
Handover
Evaluation set, pipeline and the process for adding document types.
What you receive
Summaries that are quick to verify and testable against what must not be missed.
Summary contract
Structure, length, required elements and prohibited content per document type and reader.
Summarisation pipeline
Length-aware generation with citations to source passages, deployed as code.
Faithfulness checking
Claim-level verification with unsupported statements flagged or removed.
Coverage checking
Automated verification that required elements are present, per document type.
Expert evaluation set
Reference summaries and review criteria, wired in as a release gate.
Review interface
Summary alongside source with highlighting, and correction capture.
Is this the right engagement?
Worth being direct. Document Summarisation is the wrong spend in some situations, and those are listed rather than buried.
Good fit if
- People read long documents to extract the same handful of things.
- The reader and the decision can be named per document type.
- Experts can specify what must never be omitted.
- Verification against the source is acceptable and expected.
- Document volume makes reading everything impractical.
Choose something else if
- Nobody can say what a correct summary would contain.
- The documents are short enough to read directly.
- The output would be acted on with no verification, on consequential material.
- The real need is search rather than summarisation. See semantic and vector search.
Frequently asked questions
Marked up with FAQPage schema so these answers can surface directly in search results and inside AI assistant responses.
What is the biggest risk with automated summarisation?
Omission, not fabrication. A fluent, accurate summary that silently drops the termination clause or the adverse finding looks exactly like a good summary. We agree what must never be omitted per document type and check coverage automatically, in addition to checking that nothing was invented.
How do you stop it inventing things?
Every statement cites the passage it came from and is verified against it, with unsupported claims flagged or removed before the summary is shown. Citation also makes verification fast for the reader, which matters more than the absence of errors.
Can it summarise very long documents?
Yes, using hierarchical summarisation with section awareness or by retrieving the sections relevant to the reader's purpose. Naive chunking produces repetitive summaries that lose the thread, so the approach is chosen per document type rather than applied uniformly.
Will it keep the caveats and uncertainty?
It is tested for exactly that, because summarisation systems tend to convert hedged findings into confident ones. Qualification is content rather than padding, and its removal is a defect we measure rather than a style choice.
How do you know the summaries are good enough?
Your specialists write reference summaries and review system output on a sample. Their judgement sets the bar and their disagreements become the evaluation set, which then gates every subsequent change to prompts, models or pipeline.
Often paired with this
Most clients combine two or three engagements from the Natural Language Processing pillar. These are the ones that most often run immediately before or after.
Contract and Legal Document Analysis
Clause extraction, deviation flagging and obligation tracking, cited back to the page for fast verification.
Read more →Semantic and Vector Search
Search that finds the right thing, measured on your real queries with hybrid retrieval and reranking.
Read more →Sentiment and Intent Analysis
Aspect-level sentiment and intent tuned to your domain, calibrated against your own reviewers.
Read more →Is this the right engagement?
Tell us what you are trying to build. If a different service fits better, or if you do not need us at all, we will say so.