Interoperability in practiceL'interopérabilité en pratique
Standards promise interoperability interoperabilityinteropérabilité exchange formatsformats d'échange The ability of two or more systems, applications, or tools to exchange information and make use of what is exchanged.Capacité de deux systèmes, applications ou outils ou plus à échanger de l'information et à exploiter ce qui est échangé. ; the reality is more mixed. Here are the most frequent problems.
Inline tags
TMX provides standard inline tags inline tagbalise interne CAT toolsTAO A marker inside a segment standing for formatting or a code of the original document, a bold run or a footnote reference for instance, rather than for text to translate.Marqueur placé à l'intérieur d'un segment qui représente une mise en forme ou un code du document d'origine, par exemple un passage en gras ou un appel de note, et non du texte à traduire. , <bpt>, <ept>
and <ph>, precisely so that formatting can travel from one tool to another.
Those elements elementélément XMLXML The building block of an XML document: an opening tag, the content it delimits and a closing tag, or a single empty-element tag when there is no content.Unité de base d'un document XML : une balise ouvrante, le contenu qu'elle délimite et une balise fermante, ou une seule balise d'élément vide quand il n'y a pas de contenu. vary in their content, which is the
native code of the tool that wrote the file.
<bpt i="1"><strong></bpt>Another tool reads the <bpt> tag tagbalise markupbalisage A marker that delimits and identifies an element in a markup document, usually working as a pair: an opening tag and a closing tag around the content they label.Marqueur qui délimite et identifie un élément dans un document balisé, fonctionnant généralement par paire : une balise ouvrante et une balise fermante autour du contenu qu'elles désignent. without trouble, but the code inside it
means nothing to it. So it imports it as a generic marker, or as text, or it
simplifies it. Segments arrive, formatting arrives more or less, and the way
tags are displayed changes from one tool to the next.
The tools know this and ask you about it. When importing a TMX file, memoQ
offers different settings depending on where the file comes from: from memoQ,
from Trados Translator’s Workbench 2007 or earlier, from somewhere else. In the
last two cases you have to tell it to convert <ut> and custom tags into memoQ
tags, and its documentation ties that setting to the quality of the matches you
get. An import done without those settings lets the segments through but
degrades what the memory will give back to you.
Different segmentation
Tools do not cut the text into segments segmentsegment CAT toolsTAO A portion of translatable content that a CAT tool treats as a discrete unit for translation, alignment, storage and match retrieval.Portion de contenu à traduire qu'un outil de TAO traite comme une unité distincte aux fins de traduction, d'alignement, de stockage et de recherche de correspondances. in the same places. Segmentation rules are language-specific, and a good part of them is a list of the abbreviations the tool knows, so that a period after “M.”, “art.” or “etc.” is not read as the end of a sentence. Those lists differ from one tool to the next, and so do the rules around line breaks, headings and list items: the same paragraph can come out as three segments in one tool and four in another. When memories are exchanged, cuts that do not line up lose you matches: the segment you are after exists, but not in the same shape.
Metadata loss
Some information can be lost in the export-import round trip: notes specific to the tool, revision history, quality or status attributes attributeattribut XMLXML In XML, a name-value pair written in an element's start-tag or empty-element tag to provide additional information about that element, such as its language, identifier, or creation date.En XML, paire nom-valeur inscrite dans la balise ouvrante d'un élément, ou dans sa balise d'élément vide, pour fournir de l'information supplémentaire sur cet élément, comme sa langue, son identifiant ou sa date de création. . Whether it survives depends on the two tools involved and on how much each one puts in extensions of its own.
“Compatible” does not mean trouble-free
A tool can advertise support for TMX or XLIFF XML Localization Interchange File Format (XLIFF)XML Localization Interchange File Format (XLIFF) exchange formatsformats d'échange The XML-based bilingual format for passing content through a localization workflow: source and target segments travel together with workflow metadata such as translation status, rather than exchanging linguistic resources.Format bilingue fondé sur XML pour faire circuler du contenu dans un flux de localisation : segments source et cible voyagent ensemble, avec des métadonnées de flux de travail comme le statut de traduction, contrairement aux formats d'échange de ressources linguistiques. while implementing only part of the standard, adding extensions of its own, or reading it differently from another tool. The standard fixes what is written in the file, not what each program does with it.
Les normes promettent l’interopérabilité interoperabilityinteropérabilité exchange formatsformats d'échange The ability of two or more systems, applications, or tools to exchange information and make use of what is exchanged.Capacité de deux systèmes, applications ou outils ou plus à échanger de l'information et à exploiter ce qui est échangé. ; la réalité est plus nuancée. Voici les problèmes les plus fréquents.
Les balises internes
TMX prévoit des balises internes inline tagbalise interne CAT toolsTAO A marker inside a segment standing for formatting or a code of the original document, a bold run or a footnote reference for instance, rather than for text to translate.Marqueur placé à l'intérieur d'un segment qui représente une mise en forme ou un code du document d'origine, par exemple un passage en gras ou un appel de note, et non du texte à traduire. normalisées,
<bpt>, <ept> et <ph>, précisément pour que la mise en forme voyage d’un
outil à l’autre. Ces éléments elementélément XMLXML The building block of an XML document: an opening tag, the content it delimits and a closing tag, or a single empty-element tag when there is no content.Unité de base d'un document XML : une balise ouvrante, le contenu qu'elle délimite et une balise fermante, ou une seule balise d'élément vide quand il n'y a pas de contenu. varient par leur
contenu, qui est le code natif de l’outil ayant écrit le fichier.
<bpt i="1"><strong></bpt>Un autre outil lit sans peine la balise tagbalise markupbalisage A marker that delimits and identifies an element in a markup document, usually working as a pair: an opening tag and a closing tag around the content they label.Marqueur qui délimite et identifie un élément dans un document balisé, fonctionnant généralement par paire : une balise ouvrante et une balise fermante autour du contenu qu'elles désignent. <bpt>, mais le code qu’elle contient
ne veut rien dire pour lui. Il l’importe donc comme un marqueur générique, ou
comme du texte, ou il le simplifie. Les segments arrivent, la mise en forme
arrive plus ou moins, et l’affichage des balises change d’un outil à l’autre.
Les outils le savent et vous le demandent. À l’import d’un fichier TMX, memoQ
propose des réglages différents selon la provenance du fichier : venu de memoQ,
venu de Trados Translator’s Workbench 2007 ou antérieur, venu d’ailleurs. Dans
les deux derniers cas, il faut lui dire de convertir les <ut> et les balises
maison en balises memoQ, et sa documentation lie ce réglage à la qualité des
correspondances obtenues. Un import fait sans ces réglages laisse passer les
segments, mais dégrade ce que la mémoire vous rendra.
Segmentation différente
Les outils ne découpent pas le texte en segments segmentsegment CAT toolsTAO A portion of translatable content that a CAT tool treats as a discrete unit for translation, alignment, storage and match retrieval.Portion de contenu à traduire qu'un outil de TAO traite comme une unité distincte aux fins de traduction, d'alignement, de stockage et de recherche de correspondances. aux mêmes endroits. Les règles de segmentation sont propres à chaque langue, et une bonne part d’entre elles tient dans la liste des abréviations que l’outil connaît, pour qu’un point après « M. », « art. » ou « etc. » ne soit pas lu comme une fin de phrase. Ces listes ne sont pas les mêmes d’un outil à l’autre, ni les règles entourant les retours à la ligne, les titres et les listes : un même paragraphe peut donner trois segments dans un outil et quatre dans un autre. À l’échange de mémoires, des découpages qui ne coïncident pas font perdre des correspondances : le segment cherché existe, mais pas sous la même forme.
Perte de métadonnées
Certaines informations peuvent se perdre à l’aller-retour export-import : les notes propres à l’outil, l’historique de révision, les attributs attributeattribut XMLXML In XML, a name-value pair written in an element's start-tag or empty-element tag to provide additional information about that element, such as its language, identifier, or creation date.En XML, paire nom-valeur inscrite dans la balise ouvrante d'un élément, ou dans sa balise d'élément vide, pour fournir de l'information supplémentaire sur cet élément, comme sa langue, son identifiant ou sa date de création. de qualité ou de statut. Ce qui survit dépend des deux outils en présence et de ce que chacun range dans ses propres extensions.
« Compatible » ne veut pas dire sans problème
Un outil peut annoncer qu’il prend en charge TMX ou XLIFF XML Localization Interchange File Format (XLIFF)XML Localization Interchange File Format (XLIFF) exchange formatsformats d'échange The XML-based bilingual format for passing content through a localization workflow: source and target segments travel together with workflow metadata such as translation status, rather than exchanging linguistic resources.Format bilingue fondé sur XML pour faire circuler du contenu dans un flux de localisation : segments source et cible voyagent ensemble, avec des métadonnées de flux de travail comme le statut de traduction, contrairement aux formats d'échange de ressources linguistiques. tout en n’appliquant qu’une partie de la norme, en y ajoutant ses propres extensions, ou en l’interprétant autrement qu’un autre outil. La norme fixe ce qui est écrit dans le fichier, pas ce que chaque logiciel en fait.
Sources
- O'Hagan, M. (Ed.) (2020). The Routledge handbook of translation and technology. Routledge. https://doi.org/10.4324/9781315311258
- OSCAR (LISA) (2008). SRX 2.0 - OSCAR Recommendation. Localization Industry Standards Association. https://www.ttt.org/oscarStandards/srx/srx20.html
- Andrä, S. C., & Schütz, J. (2011). The semantically-enriched translation interoperability protocol. Proceedings of the Workshop on Language Resources, Technology and Services in the Sharing Paradigm (IJCNLP 2011), 25–30. https://aclanthology.org/W11-3304.pdf
- GALA (Globalization and Localization Association) (2005). TMX 1.4b Specification. https://resources.gala-global.org/tbx14b/
- memoQ (2026). Translation memory TMX import settings (dialog). memoQ 12.4 help. https://docs.memoq.com/current/en/Workspace/translation-memory-tmx-import-settings.html
- Voita, E., Sennrich, R., & Titov, I. (2019). When a good translation is wrong in context: Context-aware machine translation improves on deixis, ellipsis, and lexical cohesion. Proceedings of the 57th Annual Meeting of the ACL, 1198–1212. https://aclanthology.org/P19-1116/
- Läubli, S., Sennrich, R., & Volk, M. (2018). Has machine translation achieved human parity? A case for document-level evaluation. Proceedings of EMNLP 2018, 4791–4796. https://aclanthology.org/D18-1512/
- Karpinska, M., & Iyyer, M. (2023). Large language models effectively leverage document-level context for literary translation, but critical errors persist. Proceedings of the Eighth Conference on Machine Translation (WMT 2023). https://aclanthology.org/2023.wmt-1.41/
- Ramos, M. M., Fernandes, P., Agrawal, S., & Martins, A. F. T. (2025). Multilingual contextualization of large language models for document-level machine translation. Second Conference on Language Modeling (COLM 2025). https://openreview.net/forum?id=Ah0U1r5Ldq
- Hu, H., Vamvas, J., & Sennrich, R. (2025). Source-primed multi-turn conversation helps large language models translate documents. Findings of the Association for Computational Linguistics: EMNLP 2025, 23702–23712. https://aclanthology.org/2025.findings-emnlp.1289/
- Maruf, S., Saleh, F., & Haffari, G. (2021). A survey on document-level neural machine translation: Methods and evaluation. ACM Computing Surveys, 54(2), article 45. https://doi.org/10.1145/3441691
- Peng, Z., Bawden, R., & Yvon, F. (2025). Investigating length issues in document-level machine translation. Proceedings of Machine Translation Summit XX: Volume 1. https://aclanthology.org/2025.mtsummit-1.3/
- Iranzo-Sánchez, J., Iranzo-Sánchez, J., Giménez, A., Civera, J., & Juan, A. (2024). Segmentation-free streaming machine translation. Transactions of the Association for Computational Linguistics, 12, 1104–1121. https://aclanthology.org/2024.tacl-1.61/
- Naveen, P., & Trojovský, P. (2024). Overview and challenges of machine translation for contextually appropriate translations. iScience, 27(10), 110878. https://www.cell.com/iscience/fulltext/S2589-0042(24)02103-5
Sources
- O'Hagan, M. (Ed.) (2020). The Routledge handbook of translation and technology. Routledge. https://doi.org/10.4324/9781315311258
- OSCAR (LISA) (2008). SRX 2.0 - OSCAR Recommendation. Localization Industry Standards Association. https://www.ttt.org/oscarStandards/srx/srx20.html
- Andrä, S. C., & Schütz, J. (2011). The semantically-enriched translation interoperability protocol. Proceedings of the Workshop on Language Resources, Technology and Services in the Sharing Paradigm (IJCNLP 2011), 25–30. https://aclanthology.org/W11-3304.pdf
- GALA (Globalization and Localization Association) (2005). TMX 1.4b Specification. https://resources.gala-global.org/tbx14b/
- memoQ (2026). Translation memory TMX import settings (dialog). memoQ 12.4 help. https://docs.memoq.com/current/en/Workspace/translation-memory-tmx-import-settings.html
- Voita, E., Sennrich, R., & Titov, I. (2019). When a good translation is wrong in context: Context-aware machine translation improves on deixis, ellipsis, and lexical cohesion. Proceedings of the 57th Annual Meeting of the ACL, 1198–1212. https://aclanthology.org/P19-1116/
- Läubli, S., Sennrich, R., & Volk, M. (2018). Has machine translation achieved human parity? A case for document-level evaluation. Proceedings of EMNLP 2018, 4791–4796. https://aclanthology.org/D18-1512/
- Karpinska, M., & Iyyer, M. (2023). Large language models effectively leverage document-level context for literary translation, but critical errors persist. Proceedings of the Eighth Conference on Machine Translation (WMT 2023). https://aclanthology.org/2023.wmt-1.41/
- Ramos, M. M., Fernandes, P., Agrawal, S., & Martins, A. F. T. (2025). Multilingual contextualization of large language models for document-level machine translation. Second Conference on Language Modeling (COLM 2025). https://openreview.net/forum?id=Ah0U1r5Ldq
- Hu, H., Vamvas, J., & Sennrich, R. (2025). Source-primed multi-turn conversation helps large language models translate documents. Findings of the Association for Computational Linguistics: EMNLP 2025, 23702–23712. https://aclanthology.org/2025.findings-emnlp.1289/
- Maruf, S., Saleh, F., & Haffari, G. (2021). A survey on document-level neural machine translation: Methods and evaluation. ACM Computing Surveys, 54(2), article 45. https://doi.org/10.1145/3441691
- Peng, Z., Bawden, R., & Yvon, F. (2025). Investigating length issues in document-level machine translation. Proceedings of Machine Translation Summit XX: Volume 1. https://aclanthology.org/2025.mtsummit-1.3/
- Iranzo-Sánchez, J., Iranzo-Sánchez, J., Giménez, A., Civera, J., & Juan, A. (2024). Segmentation-free streaming machine translation. Transactions of the Association for Computational Linguistics, 12, 1104–1121. https://aclanthology.org/2024.tacl-1.61/
- Naveen, P., & Trojovský, P. (2024). Overview and challenges of machine translation for contextually appropriate translations. iScience, 27(10), 110878. https://www.cell.com/iscience/fulltext/S2589-0042(24)02103-5