EU AI Act transparency duties apply now; high-risk duties from December 2027. Check your exposure
Insights About us Careers
Contact us
Professional Services

AI for Translation and Localisation Services

This is the professional services vertical where AI arrived first and reshaped the economics most completely, which makes it the sector others should study.

Tier 3
Our depth here
Quality estimation
The routing decision
Per-word
Already broken

Language services have already been through what other professional services are approaching: a technology that changed the unit economics faster than the pricing model adapted. The lessons are directly transferable.

In one paragraph

AI for translation and localisation covers machine translation quality estimation, effort-based routing and pricing, terminology and domain adaptation, translation memory optimisation, high-consequence content handling, and workflow automation across multilingual production.

Quality estimation, which decides everything downstream

The most valuable model in a language service business is not the translation engine — it is the one that predicts how much human effort a given machine output will need, because that decides routing, pricing and margin on every job.

  • Effort varies enormously by content type. Product descriptions and legal clauses in the same language pair need entirely different treatment, and a uniform process overpays for one and underserves the other.
  • Segment-level estimation beats document-level. Most documents contain a mix, and routing at segment level concentrates the human effort where it changes the outcome.
  • Your post-editing data is the training set. How much editors actually changed, by content type and language pair, is exactly the signal — and most providers hold it and never use it.
  • Confidence must be honest at the low end. A system that cannot say it is unsure on a critical segment is worse than one with lower average quality.
  • Effort-based pricing follows. Once effort is predictable, pricing on predicted effort rather than per word aligns the commercial model with the actual work.
Worth knowing

Per-word pricing stopped describing the work some time ago

A word requiring no change and a word requiring complete retranslation price identically under per-word rates, which means the model systematically misprices both directions. Providers that moved to effort-based pricing needed a defensible effort prediction first, which is why quality estimation is the enabling investment rather than a technical nicety. Other professional services facing the same pressure should look closely at how this vertical resolved it, because language services are three or four years ahead of the rest.

Where machine output must not be trusted

Consequence, not difficulty, sets the process

A mistranslated marketing headline is embarrassing. A mistranslated dosage instruction, safety warning, contractual obligation or legal filing causes harm. The process should be set by what happens when it is wrong, not by how hard the text looks.

Fluent errors are the dangerous ones

Machine translation errors are grammatically fluent, which means a reviewer skimming for readability will pass them. High-consequence content needs verification against the source, by someone qualified, not a fluency check.

Regulated content carries named requirements

Medical device documentation, pharmaceutical labelling, legal filings and safety instructions frequently carry qualification or certification requirements for the translator. Those apply regardless of what produced the first draft.

Back-translation is a control, not a formality

For the highest-consequence content, independent back-translation and reconciliation remains the strongest available check, and it should be reserved for where it matters rather than applied uniformly.

Terminology, memory and the assets clients pay for

AssetAI applicationNote
Terminology basesExtraction and consistency enforcementConsistency is what clients notice; enforcement is more valuable than generation
Translation memoriesCleaning, deduplication and quality scoringMost memories contain years of accumulated errors nobody has audited
Domain adaptationTuning on client corporaWhere genuine differentiation sits for a specialist provider
Style guidesAutomated conformance checkingTurns a document nobody reads into an enforced constraint
Post-editing effort dataQuality estimation trainingThe asset most providers hold and never exploit
Reviewer feedbackTargeted engine improvementCloses the loop between review and output quality
Worth knowing

Audit the translation memory before adapting on it

Translation memories accumulate over years and contain errors, superseded terminology, inconsistent register and content translated under different quality standards. Adapting a model on an unaudited memory teaches it the accumulated mistakes and then reproduces them at scale with confidence. Cleaning and scoring the memory is unglamorous work that determines whether domain adaptation improves output or entrenches old errors, and it is routinely skipped in favour of the adaptation itself.

Process

How an engagement runs

Quality estimation first, because routing and pricing both depend on it.

Weeks 1 to 2

Scope and content classification

Content types by consequence, and where uniform process is currently misapplied.

Weeks 3 to 6

Effort data assessment

Post-editing history by content type and language pair, and whether it supports estimation.

Weeks 7 to 12

Build

Segment-level quality estimation with honest low-end confidence, and routing rules by consequence.

Weeks 13 to 16

Trial

Against actual post-editing effort on live work, with high-consequence routing verified specifically.

Ongoing

Operation

Retraining as content and clients change, memories audited rather than accumulated.

Deliverables

What you receive

Effort predicted, work routed by consequence, and pricing that matches the job.

01

Quality estimation

Segment-level effort prediction with honest confidence at the low end.

02

Consequence-based routing

Process set by what happens if the translation is wrong, not by how the text looks.

03

Effort-based pricing model

Pricing on predicted effort rather than per word, with the prediction defensible to a client.

04

Terminology enforcement

Consistency checked automatically, which is what clients actually notice.

05

Translation memory audit

Errors, superseded terminology and inconsistent register identified before adaptation.

06

Domain adaptation

Tuned on cleaned client corpora, where a specialist provider's differentiation sits.

Fit check

Is this the right starting point?

Worth being direct. There are situations in translation and localisation where custom AI work is the wrong spend, and those are listed rather than buried.

Worth doing if

  • A uniform machine-then-edit process is applied to content with very different consequences.
  • Post-editing effort data exists and has never been used to predict effort.
  • Per-word pricing no longer reflects the work on a growing share of jobs.
  • Translation memories have accumulated for years without an audit.
  • Terminology consistency is a recurring client complaint.

Do something else if

  • You want machine output on regulated content without qualified verification.
  • Post-editing effort has never been recorded and cannot be reconstructed.
  • Clients will not accept any pricing basis other than per word and margin is already gone.
  • Translation memories are too degraded to clean and there is no appetite to rebuild.
Questions

Frequently asked questions

Marked up with FAQPage schema so these answers can surface directly in search results and inside AI assistant responses.

What is the most valuable model for a language service provider?

Quality estimation, not translation. Predicting how much human effort a machine output will need decides routing, pricing and margin on every job, and it is trainable on data you already hold — how much your editors actually changed, by content type and language pair. Segment-level estimation matters more than document-level, because most documents mix content that needs nothing with content that needs everything. Once effort is predictable, effort-based pricing becomes defensible to a client, which is the commercial move the per-word model has been blocking.

Where should machine translation never be used unsupervised?

Anywhere the consequence of an error is harm rather than embarrassment: medical and pharmaceutical documentation, safety instructions, legal filings, contractual obligations and regulatory submissions. The specific danger is that machine translation errors are grammatically fluent, so a reviewer skimming for readability passes them — high-consequence content needs verification against the source by someone qualified, not a fluency check. Regulated content also frequently carries translator qualification requirements that apply regardless of what produced the first draft.

Should we adapt a model on our translation memories?

After auditing them, not before. Memories accumulate over years and contain errors, superseded terminology, inconsistent register and content translated under quality standards you no longer accept. Adapting on an unaudited memory teaches the model your accumulated mistakes and reproduces them at scale, confidently. Cleaning and scoring the memory is the unglamorous step that determines whether adaptation improves your output or entrenches old errors, and it is the step most often skipped.

Our per-word rates are collapsing. What actually works?

Effort-based pricing, and it needs a defensible effort prediction underneath it. Per-word rates price a word requiring no change identically to one requiring complete retranslation, which misprices in both directions and hands the whole efficiency gain to the client. Providers that moved successfully built quality estimation first, so they could quote on predicted effort and show the client the basis. The other viable moves are specialising in high-consequence content where verification is the service, and domain adaptation that makes your output measurably better in a specific field.

What can other professional services learn from this vertical?

That the pricing model breaks before the profession decides how to respond, and that waiting is the worst option. Language services went through this three or four years ahead of law, accounting and consulting: capability improved sharply, the per-unit pricing basis stopped describing the work, and clients priced the efficiency in before providers had a new model ready. The providers who did well moved the pricing basis early, specialised where consequence made verification the product, and invested in the enabling measurement — which is the same sequence now facing everyone else in this sector.

Tell us what the problem looks like.

Thirty minutes, no charge, no deck. We will tell you whether this is an AI problem, a data problem, or a process problem — and we will say when the honest answer is to buy something rather than build it.