Transept
Free tool

Clean and convert your TMX file

Tidy up a translation memory without a full CAT tool: drop duplicate and empty entries, strip broken formatting tags, or move the whole TMX into a spreadsheet to review. Free, no sign-up, and processed entirely in your browser.

Your files never leave your browser.

Step by step

How to clean a TMX file

Four steps, and the file never leaves your browser.

  1. Upload your TMX

    Drop your .tmx file onto the tool. It opens in your browser, and nothing is sent to a server.

  2. Review the memory

    See the language pair, how many segments it holds, how many are duplicates, and how many still carry formatting tags.

  3. Choose what to remove

    Tick the duplicates and empty entries you want gone. You can also strip formatting tags to leave plain text. Each option shows how many entries it affects.

  4. Download the clean file

    Download a cleaned TMX with your choices applied, or export the memory to Excel or CSV to edit the entries and rebuild the TMX afterwards.

What is a TMX file?

TMX (Translation Memory eXchange) is the standard format for moving a translation memory between tools. A translation memory stores every sentence you have translated together with its translation, so a CAT tool such as Trados, MemoQ, or OmegaT can reuse past work on new projects. TMX is how that memory travels from one tool to another.

Over years of use a memory collects clutter: the same segment saved twice, entries with no translation, and broken inline tags left behind by an import. That clutter inflates the file and pollutes matches when the memory is reused. This tool removes it, and if you would rather work in a spreadsheet, turns the memory into an editable grid you can translate and rebuild.

FAQ

  • It removes the entries that make a memory bigger and less reliable: exact duplicate segments, entries missing a source or a target, and broken formatting tags. The result is a smaller TMX that gives cleaner matches when you reuse it.
  • A segment is a duplicate when both its source and its target match an earlier entry exactly. The first one is kept and the later copies are removed. If you also strip formatting tags, entries that differed only by their tags collapse into one.
  • Stripping is optional and off by default. When you turn it on, inline tags such as bpt and ph are removed while the words they wrapped are kept, which is what you want for a plain-text glossary or a clean re-import. If your workflow needs those tags, leave the option unticked.
  • No. The whole tool runs in your browser, so your translation memory never leaves your device and nothing is stored on a server.
  • It reads standard TMX 1.4 exported by tools like Trados, MemoQ, OmegaT, Wordfast, and Smartcat, including multilingual memories. Bilingual memories are the common case and work best.
  • Yes. Export the memory to an Excel or CSV sheet with one row per segment: the ID, source, target, and any note. Any cleaning options you picked apply to the sheet too.
  • Export to a spreadsheet, edit the target column in Excel or any editor, then switch to Spreadsheet to TMX, add your original file and the edited sheet, and download a new TMX with your changes written back in. The original stays the template, so everything else is left untouched.

Building a memory you can trust?

Transept is a translation editor with a memory that reuses every decision you make, so your terminology and phrasing stay consistent across projects without the cleanup afterwards.