EU AI Act high risk obligations are now enforceable. Check your exposure
Insights About us Careers
Contact us
Natural Language Processing

Document Summarisation

Summaries produced for a defined reader making a defined decision, checked against the source for anything invented and, more importantly, for anything consequential left out.

5 to 9 weeks
Typical build
Fixed fee
Commercial model
Faithfulness
Checked

Summarisation demonstrates beautifully and deploys badly, because the demo is judged on whether the summary reads well and production is judged on whether anything important was left out. Those are different questions, and only one of them is hard.

In one paragraph

Document summarisation is the automated production of shorter text that conveys the content of a longer source, built for a specific audience and purpose, with faithfulness to the source verified and omission of material facts treated as the primary failure mode.

Omission is the failure that matters

Fabrication is the well-known risk and the easier one to test for: a claim in the summary that does not appear in the source is detectable. Omission is harder and more dangerous, because a fluent, accurate summary that silently drops the one clause your reader needed looks exactly like a good summary.

So we define, with your experts, what must never be omitted for each document type: the termination clause, the adverse finding, the deadline, the caveat on the recommendation. Those become explicit checks rather than hopes, and the summary is evaluated on coverage of them as well as on faithfulness.

Worth knowing

Write the summary for a named reader

'Summarise this document' has no correct answer. A summary of the same contract for a lawyer, a procurement manager and an operations lead should contain different material. We specify the reader and the decision they are making, which turns an unanswerable request into a testable one.

How we build summarisation

Define the summary contract

Length, structure, required elements, prohibited content and tone, per document type and per reader. A structured summary with defined sections is easier to evaluate, easier to scan and considerably more useful than free-form prose, and it makes coverage measurable.

Handle length properly

Long documents exceed what fits usefully in one pass, and naive chunk-and-merge produces summaries that repeat themselves and lose the thread. Hierarchical summarisation with section awareness, or retrieval of the relevant sections when the summary has a specific purpose, both work; which one depends on the document and the reader.

Cite back to the source

Every statement links to the passage it came from, so a reader can verify in one click. This is what makes a summary usable for anything consequential, and it changes the summary from something to be trusted into something to be checked quickly.

Check faithfulness and coverage automatically

Claim-level verification against the source for anything unsupported, plus coverage checks for the required elements. Both run on every summary, and both are scored on a set your experts have reviewed.

Preserve the hedges

Summaries have a tendency to convert 'the evidence suggests, subject to further testing' into 'the evidence shows'. Uncertainty, conditionality and qualification are content, not padding, and we test explicitly for their removal.

Evaluate against your experts

Your specialists write reference summaries for a sample and review the system's output. Their judgement sets the bar, and their disagreements become the evaluation set that gates every subsequent change, following the practice in LLM evaluation.

Where summarisation is worth it

SettingReader and decisionWhat must never be omitted
Contract reviewLegal, deciding whether to signLiability, termination, change of control
Case and claim filesHandler, deciding next actionPrior decisions, deadlines, vulnerability flags
Research and reportsAnalyst, deciding what to read fullyMethod limitations, sample size, conflicts
Meeting and call recordsAttendees, deciding follow-upCommitments made, owners, dates
Clinical documentationClinician, deciding careAllergies, adverse events, contraindications
Regulatory correspondenceCompliance, deciding responseDeadlines, required actions, penalties

In every row the third column is the design specification. Where nobody can fill it in, the summary has no testable definition of correct and we would rather establish that before building.

Process

How the engagement runs

The required-elements list is agreed first; it is what makes the output testable.

Week 1

Reader, purpose and contract

Audience and decision defined per document type, summary structure and required elements agreed with your experts.

Weeks 2 to 3

Reference set

Experts write reference summaries for a sample, and the evaluation criteria and coverage checks are built from them.

Weeks 4 to 6

Summarisation pipeline

Length handling, section awareness, citation back to source, hedge preservation, scored against the reference set.

Weeks 7 to 8

Verification and review

Automated faithfulness and coverage checks, low-confidence routing, review interface with source highlighting.

Week 9

Handover

Evaluation set, pipeline and the process for adding document types.

Deliverables

What you receive

Summaries that are quick to verify and testable against what must not be missed.

01

Summary contract

Structure, length, required elements and prohibited content per document type and reader.

02

Summarisation pipeline

Length-aware generation with citations to source passages, deployed as code.

03

Faithfulness checking

Claim-level verification with unsupported statements flagged or removed.

04

Coverage checking

Automated verification that required elements are present, per document type.

05

Expert evaluation set

Reference summaries and review criteria, wired in as a release gate.

06

Review interface

Summary alongside source with highlighting, and correction capture.

Fit check

Is this the right engagement?

Worth being direct. Document Summarisation is the wrong spend in some situations, and those are listed rather than buried.

Good fit if

  • People read long documents to extract the same handful of things.
  • The reader and the decision can be named per document type.
  • Experts can specify what must never be omitted.
  • Verification against the source is acceptable and expected.
  • Document volume makes reading everything impractical.

Choose something else if

  • Nobody can say what a correct summary would contain.
  • The documents are short enough to read directly.
  • The output would be acted on with no verification, on consequential material.
  • The real need is search rather than summarisation. See semantic and vector search.
Questions

Frequently asked questions

Marked up with FAQPage schema so these answers can surface directly in search results and inside AI assistant responses.

What is the biggest risk with automated summarisation?

Omission, not fabrication. A fluent, accurate summary that silently drops the termination clause or the adverse finding looks exactly like a good summary. We agree what must never be omitted per document type and check coverage automatically, in addition to checking that nothing was invented.

How do you stop it inventing things?

Every statement cites the passage it came from and is verified against it, with unsupported claims flagged or removed before the summary is shown. Citation also makes verification fast for the reader, which matters more than the absence of errors.

Can it summarise very long documents?

Yes, using hierarchical summarisation with section awareness or by retrieving the sections relevant to the reader's purpose. Naive chunking produces repetitive summaries that lose the thread, so the approach is chosen per document type rather than applied uniformly.

Will it keep the caveats and uncertainty?

It is tested for exactly that, because summarisation systems tend to convert hedged findings into confident ones. Qualification is content rather than padding, and its removal is a defect we measure rather than a style choice.

How do you know the summaries are good enough?

Your specialists write reference summaries and review system output on a sample. Their judgement sets the bar and their disagreements become the evaluation set, which then gates every subsequent change to prompts, models or pipeline.

Is this the right engagement?

Tell us what you are trying to build. If a different service fits better, or if you do not need us at all, we will say so.