Transept
LocalizationMTPEPost-editingAI translation

Machine Translation Post-Editing (MTPE): What It Is, How It Works, and How to Do It Better

A practical guide to machine translation post-editing (MTPE): light vs. full post-editing, why raw MT still needs a human, how quality is measured — and why the usual batch-then-scrub model is often the wrong shape, plus the iterative, memory-fueled workflow that beats it on both quality and cost.

Vitalii Vlasiuk
Vitalii Vlasiuk19 min read
Two-panel comic: under "Traditional Machine-First MTPE" the Literess mascot wrestles a furious printer spewing pages covered in red corrections; under "Adaptive Human-First MTPE" she calmly checks off a clipboard of green-marked pages while a small deflated printer wonders why it is sad
On this page

I became an AI engineer partly by trying to translate my own novel. That is where I ran into the thing everyone in localization already knows by a different name: you run the text through a machine, then a human fixes it. Formally, it's machine translation post-editing, or MTPE, and it's the dominant way professional translation gets done in 2026.

Conceptually, it sounds smart. In the era of capable LLMs, however, the way it's usually practiced has got stale. Having a full machine pass first and then having a human scrub the whole thing afterward spends human time and compute in the wrong order.

There is a cheaper, better workflow hiding underneath it.

This guide covers both. First, the honest version of what MTPE is and how it works, because you should understand the traditional ISO way to get hired to sell your B2B SaaS. Then I will explain where the standard batch model falls down and how to do it better by putting the human in the loop earlier.

What is machine translation post-editing (MTPE)?

MTPE, also written PEMT (post-edited machine translation), is a three-step pipeline:

  1. Machine translation (MT). A machine translation engine or large language model produces a first-draft translation of the source text. This draft is called raw MT.
  2. Post-editing (PE). A human linguist reads the raw MT against the source and revises it, correcting errors, enforcing terminology, and adjusting tone until it meets the target quality bar.
  3. Quality assurance. A final check for consistency, formatting, and the errors that are easy to miss on a first pass.

As it's usually run, that's a batch handoff: the machine translates the whole document, then a human edits the whole document. Hold onto that shape; it's the part I'll argue with later.

The reason MTPE exists, rather than "just translate it" or "just use the machine," is that both extremes are wrong for most content.

Human-from-scratch translation is slow and expensive at scale. Raw MT is fast and cheap, but decent MT has existed for less than a decade, while some MT is decades old. And even with the best LLMs and workflows, the risk of accepting their output blindly is too high, while humans' ability to translate beautifully is much greater.

MTPE, hence, is the pragmatic middle: let the machine do the mechanical 70%, and spend human attention on the 30% that decides whether the translation is trustworthy.

(That, at least, is how agencies and corporations would want it to be. In real life, humans usually either engage more to provide better results (but do not get paid more) or do not engage at all, leading to suboptimal outputs.)

Regardless of that, MTPE of some sort is the default in the translation industry as of the 2020s.

Light vs. full post-editing: the two levels

TL;DR: this does not really exist in reality, but the industry demands that you be aware of this dichotomy.

"Post-editing" is not one thing. The single most important decision in an MTPE project is which level you're editing to, because it determines how much time each segment costs and what you're allowed to leave alone.

The mistake teams make is editing everything to full quality out of habit (including the low-visibility content where light editing would have been fine) and burning the time savings MTPE was supposed to deliver. The mistake is usually not the translators' alone, but also that of editors and managers who set expectations too high for light post-editing.

The opposite mistake is worse: light-edit a landing page and ship something that reads almost right and erodes trust. That is what most translations with Claude Code are: people report seeing conversion rates for their apps drop after delivering full-MT or light-MTPE localizations.

In real life, the line is blurred, and actual MTPE is somewhere in the middle. Depending on the platform or agency, the effort level is set arbitrarily or per project. There are a lot of in-house pipelines and rituals too, with QA reports and approval chains.

This blurriness of definition is what makes a lot of translators dislike working with MTPE. It does not bring the satisfaction of work done well; yet doing work well is not rewarded monetarily.

Why raw machine translation still needs a human

The obvious objection in 2026: modern LLMs are fluent. Isn't raw MT good enough now?

Fluency is exactly the trap. Older machine translation produced output that was visibly broken, so nobody trusted it unedited. Modern models produce output that reads beautifully, which gives readers and reviewers a false sense of trust.

What's worse, it may produce output that preserves meaning but has a style that reeks of AI. It makes the overall brand feel cheap, and while this particular problem is easily solvable, many fall into this exact trap.

Claude insisted I should include this:

  • Reputation damage. Engines have no linguist's grasp of cultural nuance or of what is appropriate or tactless to say in a given market. A phrasing that's neutral in the source can land as clumsy, presumptuous, or offensive in the target, or just weird; those are the "AI tells" that push people off.
  • Misinformed customers. Even the best-trained model can silently drop a clause or add a word that wasn't in the source. In a fluent, confident-sounding paragraph, an accidental omission or factual drift is very hard to spot, and it can carry real consequences in legal, medical, or financial content. Once again, this is technically solvable (Smart Proofread exists to catch exactly these omissions), but it still happens all the time.
  • A diluted brand. Raw LLM output rarely reflects a brand's specific voice and terminology. When rendered a dozen slightly different ways across a site, a product name or a signature phrase stops being recognizable, and the brand's public image gets fuzzier every time. You can solve it via styleguides and glossaries, but older MT pipelines do not make them work that well.

So the human is not optional, at least not when the text should do the work: convert customers or have legally binding power. The real question, which the standard workflow answers badly, is when the human should show up.

What actually makes MTPE hard: four real bottlenecks

In practice, post-editors are not mainly defeated by quotation marks or date formats. Those are QA rules. The hard part is evaluating deceptively fluent output, deciding how much intervention a segment needs, and doing that under a commercial model that assumes every machine draft saves the same amount of time.

Research on post-editing effort separates the job into cognitive, technical, and temporal effort. Those dimensions do not move together neatly: a sentence may need two keystrokes but several minutes of verification, while another may need a full rewrite whose solution is obvious.

The recurring bottlenecks are these.

  1. Detecting errors that sound correct but need context to fix. Modern MT can produce polished target text while omitting a clause, shifting the meaning, or choosing a plausible but wrong term. Fluency removes the warning signs, so the editor has to keep checking the source, document context, and real-world facts. That all takes time and effort. Research consistently finds that mistranslations, coherence problems, and structural errors are among the strongest drivers of post-editing effort.
  2. Deciding whether to accept, repair, or retranslate. Each segment demands triage: keep it, make a minimal correction, rewrite it, or discard the machine version and translate from scratch. Vague light-versus-full instructions make that decision worse. Editors either under-edit to satisfy “change as little as possible” or over-edit because their professional quality bar is higher than the brief.
  3. Escaping the machine’s framing. Post-editors do not approach a blank page. They read a proposed answer first, and that answer anchors vocabulary, syntax, and interpretation. Studies of professional workflows have found priming effects: errors and awkward choices in the MT can survive into the post-edited text or shape the correction itself. The better the sentence sounds, the harder it can be to notice that its framing is wrong. That is particularly common when LLMs perform the MT step.
  4. Working under false productivity assumptions. MTPE is often priced as if machine output creates a predictable discount in human effort. It does not. Difficulty changes by segment, language pair, domain, and error type, while speed-focused pricing can discourage terminology research even when clients still expect perfection. The translator absorbs the variance: easy segments justify the lower rate, but hard ones quietly erase the savings.

Formats, terminology constraints, locked segments, and locale rules still matter, but they are controls the workflow should enforce upstream, such as via glossaries, styleguides, or translation memory.

The real MTPE problem is where judgment gets spent: spotting fluent errors, choosing the right level of intervention, resisting machine anchoring, and doing all three without letting the pricing model decide the quality.

That is why I soon came to the conclusion that a human-early, interactive workflow has a more positive impact than a batch cleanup screen. But hasn't the industry arrived at that already?

How to measure MTPE quality: BLEU and beyond

MTPE has a measurement problem: quality is partly subjective. The industry's best-known automatic metric is BLEU (Bilingual Evaluation Understudy), which compares a machine translation with one or more high-quality human reference translations.

It works by breaking the machine output into short word sequences, usually one to four words long, and measuring how many also appear in the references. Repeated matches are capped so the system cannot inflate its score by repeating a correct word, and a brevity penalty lowers the result if the translation is suspiciously short.

BLEU combines those overlap measurements into a score between 0 and 1, commonly displayed as 0–100. Higher means closer wording to the references, not necessarily better meaning.

In my personal opinion, it is not really a useful metric in 2026, and here is why:

BLEU scores are only meaningfully comparable when engines are tested on the same dataset with the same scoring setup, so there is no universal threshold where “50 is good, 10 is bad.” Their practical value is in controlled before-and-after comparisons.

In one TAUS case study, training on 172,980 French–German segments from a narrow legal domain added 7.23 BLEU points, a 19% relative improvement. In a separate Russian–English aviation case, a cleaned translation memory containing one million segments raised Globalese from 23.6 to nearly 51 BLEU, a 115.5% relative gain. The scale of improvement varies sharply by domain, language pair, baseline engine, and data quality, but both cases show why relevant, carefully cleaned translation memories can outperform generic training data.

But BLEU is a similarity score, not a truth score. A fluent, confident mistranslation can score well; a perfectly good alternative phrasing that happens not to match the reference can score badly. So BLEU sets the floor (it tells you the raw MT is in decent shape), while human review still decides whether the translation is right, on-brand, and safe to ship.

Teams increasingly supplement it with LLM-as-judge quality estimation, but that alone does not make much sense.

The trouble with MTPE as it's usually practiced

Here is where I stop describing and start arguing. The thesis is:

A decision you make on page 2 — this character stays formal, redo this pun to reference a local politician, don't latinise that place name — can be fed back in to regenerate the rest of the document.

Examples and multi-shot prompting do wonders for LLMs even without advanced context engineering, even if the context is not a 100% match. Including relevant genre examples even in *another *language produced significant improvements in internal Transept benchmarks.

Run a full machine pass first, and you've spent compute credits on generating pages the human will partly discard, then spent human time bending cold output into shape.

Moreover, the model is far faster than the human. You don't need all 10,000 words drafted before the human starts.

A better MTPE workflow is to generate the first couple thousand words, let the translator make the decisions that matter, then continue with those decisions and the translation memory in context while the human is still on page five. In Transept's tests, even sparse *comments *on the untranslated source that captured the translator's overall direction measurably improved later output.

The later pages arrive already shaped by the earlier decisions, instead of arriving as a uniform slab of slop for someone to fix.

None of this means MTPE is wrong. For low-visibility bulk, a full pass followed by a light edit is the efficient choice, though not the most pleasant work.

But when the content matters, the batch model leaves quality and money on the table. And it does so because iterating feels more expensive in the short term.

It usually isn't.

A better way: put the human in early

Everything above points at the same move. Get the human's judgment into the system before the machine commits to a full draft, and let memory carry those judgments forward. Four practices do most of the work.

Play editor first, translator second. Before generating anything, read the source and leave opinionated notes on the passages that carry weight: this joke has to survive; this term is critical; here we can deviate from the source to preserve meaning. Then generate a draft grounded in those notes plus your memory.

Post-editing a draft that already knows your intentions is far less tedious than fine-tuning a cold one. I say that as someone who did both to my own book, many times.

Fuel the memory with your best work (if your MT can use it). Translation memory is underrated for AI. For a hard passage or a tricky language pair, translate a few snippets fully by hand and feed them in; the model picks up voice and register from real examples better than from any amount of instruction.

TM alone can hold the style of a long text together. Finding the right past sample (semantic plus fuzzy search) is an art of its own, but producing a good one is harder. It also needs human direction, because there are too many statistically valid ways to render a line and only you know which one is yours.

Not all the tools can use TM to generate good prose, though. In Transept, we did extensive research (written up in our essay on translation memory for AI localization) to ensure that human-approved segments, and the context and work history behind them, make LLM translation better.

Use several models in sequence, not one big one. One result has held up in both academic research and our internal research: a cascade of AI roles beats a single strong pass on quality and cost at once.

Draft with a cheap, fast model. Have a strong model read the draft and leave comments on what to improve. Have a fast model apply those comments.

That turns out to be cheaper than asking the strong model to do everything. Moreover, it's usually better. This one was a big surprise to me, as I could not believe that a combination of strong and weak models could produce output better than a strong big model could produce alone.

But it works. My hypothesis is that critiquing and generating are different jobs and activation patterns. Separating them plays to each model's strength and eases "cognitive load", helping to focus the effort.

Basically, that's the backbone of the entire multi-agent paradigm, where the model can fix errors made by itself if instantiated with different settings and tasks.

Set the plumbing up properly. Real performance comes from the API, not chat apps, whose long system prompts and product scaffolding get in the way.

LLM output is often underappreciated by translators because it's cooked wrong, just like a tuna steak with a dark soy and sesame sauce is nothing like the canned, tasteless fish you know tuna as.

Manage the input so it's clean and cache-friendly, and both cost and latency drop. Test if the model can handle extra context (TM, decisions, styleguides), then saturate it up to the point where benchmarks drop. You need to have good benchmarks, in multiple languages, to really see if you are doing the right thing to the LLM.

We at Transept spent three years researching and experimenting on our own writing to figure this out.

The tool I wished I had while translating my novel: Transept’s answer to MTPE

A disciplined translator with a technical background can recreate the workflow with model APIs, a translation memory, and patience. It takes quite a lot of research, however, to make human-in-the-loop MTPE work *beautifully, *bringing joy and good results alike.

We built Transept because assembling those pieces by hand was the part we kept resenting while translating our own long-form work.

We turned the vision above into one workspace: direct the model early, preserve the decisions, run specialized passes, and review locally.

Direct before you generate. The document editor keeps the source, translation, and surrounding context together. A translator can comment on a block, regenerate one sentence with a direction, compare alternatives, or edit by hand without disturbing the rest. Comments and review threads stay attached to the text, so “keep the metaphor” remains a decision instead of disappearing into a chat log.

Even the alternatives you reject are kept: a discarded phrasing, and the note on why you passed on it, become part of what the next pass can reuse. See translation variants and how translation memory treats past decisions as context.

Turn decisions into memory. Transept’s translation memory brings approved translations and their surrounding decision context back into later work. Glossaries pin names, product terms, and fixed renderings; styleguides carry tone, register, rhythm, and conventions. Both can be automatically generated from your previous work you already trust, or client brand materials, then reviewed before they become used by AI.

Separate the passes. Instead of asking one model to translate, judge, and polish in a single attempt, Transept can sequence those jobs. Smart Proofread rereads the translation against the source, glossary, and styleguide, then surfaces omissions, terminology drift, and register slips as reviewable fixes.

We also provide workflows like Translate, proofread & polish that bundle translation, proofreading, polishing, and QA when orchestration matters more than manual control. It works for several language versions of the same document, too!

For long documents, we also developed the way for multiple agents to work together on the same document. They share and update styleguides and glossaries, and have special gradient memory to sync their decisions.

Review locally; scale globally. Post-editing stays at the block or sentence level: compare variants, regenerate a line, accept a suggested fix, or rewrite it yourself. Literess can help run the workflow and flag drift, while batch translation applies shared context, glossaries, styleguides, and QA across many files without turning the project into a spreadsheet.

And when the built-in passes aren't the sequence you want, you can build your own workflow step by step: pick the steps, gate any of them for review, and run it with the cost shown first.

I am honestly proud of the work we've done in Transept. These features help to maximize the impact of precious human time and talent on how LLMs work.

It's still surprising to me that only a few professional translators thought all of it was possible to achieve. Technically, building an interactive MTPE workflow was doable back in 2023, and we are still early.

MTPE best practices: a checklist

Even if you are not impressed with Transept (wrongly so), here is a list of everything you need to learn about MTPE and how to make it less soul-crushing.

  • Put the human in early. Leave direction on the passages that matter before you generate, not after.
  • Fuel the memory with real examples. Hand-translate a few of your hardest snippets and let the model learn voice from them.
  • Cascade your models. Cheap draft, strong critique, fast apply. It beats one big model on quality and cost.
  • Tier your content. Full batch MTPE for low-visibility bulk; iterate where the brand is on the line.
  • Enforce terminology upstream. Do-not-translate and forbidden-target rules stop the most common and most damaging error at the source.
  • Watch locale and register drift. Regional variant, formality, and pronoun consistency are what separate native-feeling translation from assembled-by-committee.
  • QA fluent text harder, not softer. The better the raw MT reads, the easier its omissions are to miss.

MTPE is not going away, and for a great deal of content it's the right tool. But the batch-then-fix pipeline is a wrong turn in the history of translation technology.

It's conceptually easy to fix. Put the human in early, fuel the context with real examples and memory, let cheap models draft while strong ones critique, and you get better translation for less. It is exactly the outcome that would benefit both agencies and translators.

If you want to see the human-early loop for yourself, start with a document; the Free plan covers a first translation with no card.

If you'd rather ask questions first, ask Literess or find me wherever this guide reached you.

The author

Vitalii Vlasiuk
Vitalii VlasiukCo-founder

Co-founder of Transept, writing as “Mevkh.” A Language and Literature degree, then a turn into software: senior AI engineer shipping production LLM features to 50,000+ users — RAG, agentic tools, LLM-as-judge evaluation. A novelist on the slow path, with 120,000 words of satirical romance fantasy in a drawer. The friction between AI translation and his own prose is what set this whole thing in motion.