LocalizationAI translationTranslation memoryterminology

Translation and localization glossary: the terms that actually come up

Translation memory, fuzzy match, MTPE, XLIFF, TMX, leverage, transcreation. A plain-language definition of every term you will meet in a quote, a CAT tool, or a vendor conversation, and how they relate to each other.

Mariia Ivakhnenko
Mariia Ivakhnenko9 min read
On this page

Translation has a vocabulary problem. Half the words are acronyms, several of them mean nearly the same thing, and a few mean different things depending on who is talking. A quote arrives promising "80% leverage with 85% fuzzy threshold" and it is not obvious whether that is good news.

This is a working glossary: the terms you meet in a quote, a CAT tool, or a vendor conversation, defined in the order they tend to come up rather than alphabetically. Where a term has a Transept-specific meaning, that is marked as such.

What is the difference between a glossary, a term base, and translation memory?

These three get used interchangeably in conversation and mean different things in a tool.

Glossary

A list of terms with their approved translations, plus notes on when each one applies. Product names, character names, legal vocabulary, the words a client insists on. A glossary answers one question: when this term appears, what must it become?

In Transept, glossaries are pinned into the prompt for every run and checked again afterwards, so a term settled on page 1 still holds on page 500.

Term base (TB)

The CAT-tool name for a glossary, usually with more metadata per entry: part of speech, domain, a forbidden-variants list, sometimes a definition. If someone says term base and you hear glossary, you have understood them.

Translation memory (TM)

A store of segments you have already translated, source paired with target, reused when similar text comes up again. A glossary works at the level of a term; memory works at the level of a whole sentence.

Classic memory matches on text. Transept's translation memory also matches on meaning and keeps the reasoning around a decision (the alternatives that were rejected, the review comments), which is what lets it help on a sentence that is phrased differently but solves the same problem. We wrote up what that changed.

Fuzzy match

A memory hit that is close but not identical, scored as a percentage. 85% is a common threshold: above it the stored translation is offered as a starting point, below it the segment is treated as new. The number describes textual similarity, not whether the old translation was any good.

Exact match, 100% and 101%

An exact or 100% match means the source segment is identical to one already in memory. A 101% match (also called an in-context exact match, or ICE) means the segment is identical and the segments around it match too, so the context is the same and the stored translation is far more likely to still fit.

Leverage

How much of a new job memory already covers, as a percentage. Leverage drives money: agencies discount heavily-leveraged words because a 100% match costs a review rather than a translation. When a quote quotes leverage, it is telling you how much of the work it expects to skip.

What do the machine translation acronyms mean?

Machine translation (MT)

Any automatic translation, whatever the underlying technology. The umbrella term.

Neural machine translation (NMT)

Machine translation by neural networks trained specifically for translation. This is the generation of engines that includes DeepL and Google Translate. NMT is good at producing a fluent sentence and has no way to take an instruction: you cannot tell it to be more formal partway through.

LLM translation

Translation by a general-purpose large language model. The practical difference from NMT is that you can talk to it: supply a glossary, a style guide, the surrounding document, an instruction about tone. Transept runs on frontier LLMs and treats the model as a swappable backend.

MTPE (machine translation post-editing)

A human editing machine output into a finished translation. Sometimes written PEMT. It is the dominant working model in the industry, and the one LLMs changed most: older MT was obviously rough, so errors announced themselves, while LLM output can read beautifully and still have dropped a clause.

Our MTPE guide covers how to run it so that human judgment arrives before the draft rather than after it.

Light and full post-editing

Two service levels. Light post-editing aims for accurate and understandable, and tolerates flat phrasing. Full post-editing aims for publication quality, indistinguishable from a translation written from scratch. The words matter on an invoice: they are the difference between two very different amounts of human time.

How does a translation tool cut a document up?

Segment

The unit a tool splits text into, almost always a sentence. Memory is stored per segment, matches are scored per segment, and word counts are computed per segment.

Source and target

Source is the language you translate from, target the language you translate into. A source segment and its target segment together make one memory entry. "Source words" is the usual basis for pricing, including ours.

Block

Transept's unit, and deliberately larger than a segment: a paragraph, heading, list item, or table cell. A block keeps a sentence next to the sentences it belongs with, so the model can see that a pronoun three sentences down refers to something at the top. Sentence-level control still exists inside a block.

Alignment

Pairing an existing translation with its source, segment by segment, so that work done before a tool existed can become memory inside it. Alignment is also the check that runs when you bring finished translations in: if the two sides do not line up one to one, something has been added or dropped.

Which file formats get passed between tools?

XLIFF

XML Localization Interchange File Format. The standard container for moving translatable content between systems: each unit carries its source, its target, a state, and often a note for the translator. If a client sends you a bilingual file, it is probably XLIFF. We ship a free XLIFF converter for reading one in a spreadsheet.

TMX

Translation Memory eXchange. The portable format for a translation memory, so a memory built in one tool can move to another. TMX is why a memory is an asset you own rather than something a vendor holds.

TBX

TermBase eXchange. The same idea for glossaries and term bases. Less widely used than TMX, because a glossary is small enough that people email spreadsheets instead.

PO and gettext

The string format used by a great deal of open-source software, WordPress, and Python and PHP projects. A .po file pairs each source string with its translation and carries translator comments; .pot is the empty template.

SRT and VTT

Subtitle formats: timing plus text, one cue at a time. Translating them means respecting reading speed and line length as well as meaning, because a cue that is faithful but too long is unreadable on screen.

Resource files

The per-platform containers an app keeps its interface strings in: .strings on Apple platforms, .resx on .NET, .arb for Flutter, plain JSON or YAML almost everywhere else. Translating software means translating these rather than documents, which is why string-table workflows look different from document ones.

What do the quality and process terms mean?

Locale

A language plus a region plus the conventions that come with it: date order, decimal separator, currency, address shape, sometimes a different script. pt-BR and pt-PT are both Portuguese and are not interchangeable in a product. Picking a locale rather than a language is what stops a Brazilian reader getting European phrasing.

i18n and l10n

Numeronyms, counting the letters between first and last. Internationalization (i18n) is the engineering work of making a product translatable at all: no text baked into images, no sentences assembled from fragments, room in the layout for German. Localization (l10n) is adapting it for one locale. i18n happens once; l10n happens per language.

Style guide

The rules that are not terminology: formality, tone, how to address the reader, punctuation and quotation conventions, how to handle numbers and units. Style guides are what keep a translation sounding like one voice when several people and several runs contributed to it.

LQA (linguistic quality assurance)

Structured review: a reviewer scores a sample against error categories (accuracy, terminology, style, locale conventions) with a severity per error, producing a number rather than an impression. Used to accept or reject delivery, and to compare vendors.

Transcreation

Rewriting for effect rather than translating for meaning. A slogan that puns in English will not pun in Japanese, so the brief becomes "make this land" rather than "make this accurate". Priced by the hour or by the concept, never by the word.

Back-translation

Translating the target back into the source language, with a different translator who has not seen the original, to check that the meaning survived. Standard in clinical trials and regulated materials, where being wrong is expensive. It tests meaning, never style.

FAQ

Which of these terms do I actually need on day one?

Four: glossary, translation memory, segment, and locale. Those describe how the work is stored and split, and they are enough to read a quote. The acronyms are mostly vocabulary for talking to vendors and tools, and you can look them up when one appears.

Can I bring a translation memory from another tool into Transept?

You can bring the translations themselves, and they are seeded free: no model runs and no words are billed, because the work was already done. The route is the string-table import described in Bring your existing translations, or seeding an already-translated document. A raw .tmx file is not ingested directly today. Our free TMX Cleaner converts one to Excel or CSV in your browser (nothing is uploaded), which is usually the fastest way to get an old memory into a shape you can work from.

Do glossaries and translation memory conflict?

They do different jobs and are used together. A glossary fixes what a single term must be; memory recalls how a whole segment was handled last time. A tool that uses both applies the glossary to the terms and the memory to the phrasing around them, so a recalled sentence still gets your current terminology.

Is post-editing still necessary with modern models?

Yes, and the reason changed. Older machine translation failed visibly, so errors were easy to spot and fix. A modern model produces fluent text that can still drop a clause, flatten a register, or pick the wrong sense of a word. The editing is less about repairing grammar and more about checking that the meaning and the voice actually survived.

What is the difference between translation and localization?

Translation is the text. Localization is everything else that has to change for a product to work in another market: dates and currency, name and address formats, images and colour choices, legal notices, and the layout adjustments needed when the same sentence runs 30% longer. A translated product can still feel foreign; a localized one does not.

The author

Mariia Ivakhnenko

Co-founder of Transept. Three degrees in English Language and Literature — Kyiv, Ostrava, and a year in Salzburg — and a Ukrainian native who lives most of her writing life in English. Came into AI as a prompt engineer, then product and lifecycle marketing. She writes semi-fictional stories about real people, and keeps circling the question of what gets lost between languages.