Bibliography Bibliographie
Every source the site draws on, in one list: the works cited at the bottom of content pages and on glossary entries. 172 references. Toutes les sources du site en une seule liste : les ouvrages cités au bas des pages de contenu et dans les entrées du glossaire. 172 références.
- Ahmad, I. S., Anastasopoulos, A., Bojar, O., Borg, C., Carpuat, M., et al. (2024). Findings of the IWSLT 2024 evaluation campaign. Proceedings of the 21st International Conference on Spoken Language Translation (IWSLT 2024), 1–11. https://aclanthology.org/2024.iwslt-1.1/
- American National Standards Institute (1990). American National Standard for Information Systems - Dictionary for Information Systems (ANSI X3.172-1990). American National Standards Institute. https://nvlpubs.nist.gov/nistpubs/Legacy/FIPS/fipspub11-3.pdf
- American Psychological Association (2018). Metacognition. APA Dictionary of Psychology. https://dictionary.apa.org/metacognition
- Andrä, S. C., & Schütz, J. (2011). The semantically-enriched translation interoperability protocol. Proceedings of the Workshop on Language Resources, Technology and Services in the Sharing Paradigm (IJCNLP 2011), 25–30. https://aclanthology.org/W11-3304.pdf
- Basile, V., Fell, M., Fornaciari, T., Hovy, D., Paun, S., Plank, B., et al. (2021). We Need to Consider Disagreement in Evaluation. Proceedings of the 1st Workshop on Benchmarking: Past, Present and Future. https://doi.org/10.18653/v1/2021.bppf-1.3
- Bengio, Y., Courville, A., & Vincent, P. (2013). Representation Learning: A Review and New Perspectives. IEEE Transactions on Pattern Analysis and Machine Intelligence, 35(8). https://doi.org/10.1109/TPAMI.2013.50
- Benkler, Y., Faris, R., & Roberts, H. (2018). Epistemic Crisis. Network Propaganda: Manipulation, Disinformation, and Radicalization in American Politics (Oxford University Press). https://doi.org/10.1093/oso/9780190923624.003.0001
- Bolukbasi, T., Chang, K.-W., Zou, J., Saligrama, V., & Kalai, A. (2016). Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings. Advances in Neural Information Processing Systems 29 (NIPS 2016). https://proceedings.neurips.cc/paper/2016/hash/a486cd07e4ac3d270571622f4f316ec5-Abstract.html
- Bowker, L. (2002). Computer-Aided Translation Technology: A Practical Introduction. University of Ottawa Press. https://doi.org/10.1353/book6554
- Bowker, L., & Fisher, D. (2010). Computer-aided translation. Handbook of Translation Studies (John Benjamins Publishing Company). https://doi.org/10.1075/hts.1.comp2
- Bowker, L., & Pearson, J. (2002). Working with Specialized Language: A Practical Guide to Using Corpora. Routledge. https://doi.org/10.4324/9780203469255
- Brezina, V. (2018). Statistics in Corpus Linguistics: A Practical Guide. Cambridge University Press. https://doi.org/10.1017/9781316410899
- Brown, M. A., Gruen, A., Maldoff, G., Messing, S., Sanderson, Z., & Zimmer, M. (2025). Web scraping for research: Legal, ethical, institutional, and scientific considerations. Big Data & Society. https://doi.org/10.1177/20539517251381686
- Brysbaert, M., Stevens, M., Mandera, P., & Keuleers, E. (2016). How Many Words Do We Know? Practical Estimates of Vocabulary Size Dependent on Word Definition, the Degree of Language Input and the Participant's Age. Frontiers in Psychology, 7, 1116. https://doi.org/10.3389/fpsyg.2016.01116
- Canadian Intellectual Property Office (2021). Copyright infringement. Innovation, Science and Economic Development Canada. https://ised-isde.canada.ca/site/canadian-intellectual-property-office/en/copyright-infringement
- Chandra, B., Dunietz, J., Roberts, K., Lee, Y., Fontana, P., & Awad, G. (2024). Reducing Risks Posed by Synthetic Content: An Overview of Technical Approaches to Digital Content Transparency. NIST Trustworthy and Responsible AI, NIST AI 100-4, National Institute of Standards and Technology. https://doi.org/10.6028/NIST.AI.100-4
- Chiang, W.-L., Zheng, L., Sheng, Y., Angelopoulos, A. N., Li, T., Li, D., et al. (2024). Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference. Proceedings of the 41st International Conference on Machine Learning (PMLR 235). https://proceedings.mlr.press/v235/chiang24b.html
- Common Crawl (n.d.). About. Common Crawl. https://commoncrawl.org/about
- Competition Bureau Canada (2024). Artificial intelligence and competition: discussion paper - March 2024. Government of Canada / Competition Bureau Canada. https://publications.gc.ca/site/eng/9.935280/publication.html
- Council of Europe, Steering Committee for Media and Information Society (CDMSI) (2021). Content moderation: Best practices towards effective legal and procedural frameworks for self-regulatory and co-regulatory mechanisms of content moderation. Guidance note. Council of Europe. https://rm.coe.int/content-moderation-en/1680a2cc18
- DCMI (Dublin Core Metadata Initiative) (n.d.). Metadata Basics. https://www.dublincore.org/resources/metadata-basics/
- Dell'Acqua, F., McFowland, E., III, Mollick, E., Lifshitz, H., Kellogg, K. C., Rajendran, S., et al. (2026). Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality. Organization Science, 37(2). https://doi.org/10.1287/orsc.2025.21838
- Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). https://doi.org/10.18653/v1/N19-1423
- Edman, L., Schmid, H., & Fraser, A. (2024). CUTE: Measuring LLMs' Understanding of Their Tokens. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 3017–3026. https://aclanthology.org/2024.emnlp-main.177/
- Firth, J. R. (1957). A synopsis of linguistic theory 1930–1955. Studies in Linguistic Analysis (Special volume of the Philological Society), Blackwell, 1–32. https://quoteinvestigator.com/2022/09/18/word-company/
- Gabriel, I. (2020). Artificial Intelligence, Values, and Alignment. Minds and Machines. https://doi.org/10.1007/s11023-020-09539-2
- GALA (Globalization and Localization Association) (2005). TMX 1.4b Specification. https://resources.gala-global.org/tbx14b/
- Galibert, O., Camelin, N., Deléglise, P., & Rosset, S. (2016). Estimation de la qualité d'un système de reconnaissance de la parole pour une tâche de compréhension. Actes de la conférence conjointe JEP-TALN-RECITAL 2016, vol. 1, 274–282. https://aclanthology.org/2016.jeptalnrecital-jep.31/
- Geng, S., Cooper, H., Moskal, M., Jenkins, S., Berman, J., Ranchin, N., et al. (2025). JSONSchemaBench: A Rigorous Benchmark of Structured Outputs for Language Models. arXiv preprint arXiv:2501.10868. https://doi.org/10.48550/arXiv.2501.10868
- Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep Learning. MIT Press. https://www.deeplearningbook.org/
- Google (n.d.). Refine Google searches. Google Search Help. https://support.google.com/websearch/answer/2466433
- Google (n.d.). Function calling with the Gemini API. Google AI for Developers. https://ai.google.dev/gemini-api/docs/function-calling
- Government of Canada (1985). Copyright Act (R.S.C., 1985, c. C-42). Justice Laws Website, Department of Justice Canada. https://laws-lois.justice.gc.ca/eng/acts/C-42/
- Gray, M. L., & Suri, S. (2019). Ghost Work: How to Stop Silicon Valley from Building a New Global Underclass. Houghton Mifflin Harcourt. https://search.worldcat.org/title/1052904468
- Griot, M., Hemptinne, C., Vanderdonckt, J., & Yuksel, D. (2025). Large Language Models lack essential metacognition for reliable medical reasoning. Nature Communications, 16. https://doi.org/10.1038/s41467-024-55628-6
- Grobelnik, M., Perset, K., & Russell, S. (2024). What is AI? Can you make a clear distinction between AI and non-AI systems?. OECD.AI Policy Observatory. https://oecd.ai/en/wonk/definition
- Han, T., Wang, Z., Fang, C., Zhao, S., Ma, S., & Chen, Z. (2025). Token-Budget-Aware LLM Reasoning. Findings of the Association for Computational Linguistics: ACL 2025. https://doi.org/10.18653/v1/2025.findings-acl.1274
- HeyGen (2026). Privacy Policy. https://www.heygen.com/privacy
- Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., et al. (2022). An empirical analysis of compute-optimal large language model training. Advances in Neural Information Processing Systems 35 (NeurIPS 2022). https://proceedings.neurips.cc/paper_files/paper/2022/hash/c1e2faff6f588870935f114ebe04a3e5-Abstract-Conference.html
- Holtzman, A., Buys, J., Du, L., Forbes, M., & Choi, Y. (2020). The Curious Case of Neural Text Degeneration. International Conference on Learning Representations (ICLR 2020). https://openreview.net/forum?id=rygGQyrFvH
- Hu, H., Vamvas, J., & Sennrich, R. (2025). Source-primed multi-turn conversation helps large language models translate documents. Findings of the Association for Computational Linguistics: EMNLP 2025, 23702–23712. https://aclanthology.org/2025.findings-emnlp.1289/
- International Organization for Standardization (2015). Translation services — Requirements for translation services. International Organization for Standardization. https://www.iso.org/standard/59149.html
- International Organization for Standardization (n.d.). Carbon footprint: Measuring and reducing our environmental impact. ISO. https://www.iso.org/renewable-energy/carbon-footprint
- International Organization for Standardization & International Electrotechnical Commission (2022). Information technology — Artificial intelligence — Artificial intelligence concepts and terminology (ISO/IEC 22989:2022). ISO/IEC. https://www.iso.org/standard/74296.html
- International Organization for Standardization & International Electrotechnical Commission (2017). Information technology — Cloud computing — Interoperability and portability (ISO/IEC 19941:2017). ISO/IEC. https://www.iso.org/standard/66639.html
- International Organization for Standardization & International Electrotechnical Commission (2021). Information technology — Artificial intelligence (AI) — Bias in AI systems and AI aided decision making. ISO/IEC TR 24027:2021, International Organization for Standardization. https://www.iso.org/standard/77607.html
- International Telecommunication Union (2025). Measuring What Matters: How to Assess AI's Environmental Impact. ITU publication S-GEN-GDA.001-2025. https://www.itu.int/hub/publication/s-gen-gda-001-2025/
- International Telecommunication Union (2022). Environmental efficiency and impacts on United Nations Sustainable Development Goals of data centres and cloud computing. ITU-T L.Sup55 (10/2022), International Telecommunication Union. https://www.itu.int/rec/T-REC-L.Sup55
- Iranzo-Sánchez, J., Iranzo-Sánchez, J., Giménez, A., Civera, J., & Juan, A. (2024). Segmentation-free streaming machine translation. Transactions of the Association for Computational Linguistics, 12, 1104–1121. https://aclanthology.org/2024.tacl-1.61/
- ISO (2019). ISO 30042:2019. Management of terminology resources — TermBase eXchange (TBX). International Organization for Standardization. https://www.iso.org/standard/62510.html
- ISO (2024). ISO 21720:2024. XLIFF (XML Localization Interchange File Format) — Version 2.1. International Organization for Standardization. https://www.iso.org/standard/87344.html
- Jakubíček, M., Kilgarriff, A., Kovář, V., Rychlý, P., & Suchomel, V. (2013). The TenTen Corpus Family. 7th International Corpus Linguistics Conference (CL 2013), Lancaster, 125–127. https://www.sketchengine.eu/wp-content/uploads/The_TenTen_Corpus_2013.pdf
- Jensen, B. (2024). Exploring the Complex Ethical Challenges of Data Annotation. Stanford Institute for Human-Centered Artificial Intelligence. https://hai.stanford.edu/news/exploring-complex-ethical-challenges-data-annotation
- Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., et al. (2023). Survey of Hallucination in Natural Language Generation. ACM Computing Surveys, 55(12). https://doi.org/10.1145/3571730
- Joshi, P., Santy, S., Budhiraja, A., Bali, K., & Choudhury, M. (2020). The State and Fate of Linguistic Diversity and Inclusion in the NLP World. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. https://aclanthology.org/2020.acl-main.560/
- Jung, K.-H. (2023). Uncover This Tech Term: Foundation Model. Korean Journal of Radiology, 24(10). https://doi.org/10.3348/kjr.2023.0790
- Jurafsky, D., & Martin, J. H. (2026). Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition with Language Models. Online draft, 3rd edition. https://web.stanford.edu/~jurafsky/slp3/
- Kahneman, D. (2011). Thinking, Fast and Slow. Farrar, Straus and Giroux. https://us.macmillan.com/books/9780374533557/thinkingfastandslow/
- Karpathy, A. (2023). [1hr Talk] Intro to Large Language Models. YouTube (Andrej Karpathy channel). https://www.youtube.com/watch?v=zjkBMFhNj_g
- Karpinska, M., & Iyyer, M. (2023). Large language models effectively leverage document-level context for literary translation, but critical errors persist. Proceedings of the Eighth Conference on Machine Translation (WMT 2023). https://aclanthology.org/2023.wmt-1.41/
- Kilgarriff, A., & Grefenstette, G. (2003). Introduction to the Special Issue on the Web as Corpus. Computational Linguistics, 29(3). https://doi.org/10.1162/089120103322711569
- Kilgarriff, A., Baisa, V., Bušta, J., Jakubíček, M., Kovář, V., Michelfeit, J., Rychlý, P., & Suchomel, V. (2014). The Sketch Engine: ten years on. Lexicography, 1(1), 7–36. https://doi.org/10.1007/s40607-014-0009-9
- Koehn, P. (2005). Europarl: A Parallel Corpus for Statistical Machine Translation. Proceedings of Machine Translation Summit X, Phuket, 79–86. https://aclanthology.org/2005.mtsummit-papers.11/
- Koenecke, A., Nam, A., Lake, E., Nudell, J., Quartey, M., Mengesha, Z., Toups, C., Rickford, J. R., Jurafsky, D., & Goel, S. (2020). Racial disparities in automated speech recognition. Proceedings of the National Academy of Sciences, 117(14), 7684–7689. https://doi.org/10.1073/pnas.1915768117
- Kommers, C., Duede, E., Gordon, J., Holtzman, A., McNulty, T., Stewart, S., et al. (2026). Why Slop Matters. ACM AI Letters. https://doi.org/10.1145/3786777
- Koster, M., Illyes, G., Zeller, H., & Sassman, L. (2022). Robots Exclusion Protocol. RFC 9309, Internet Engineering Task Force. https://doi.org/10.17487/RFC9309
- Läubli, S., Sennrich, R., & Volk, M. (2018). Has machine translation achieved human parity? A case for document-level evaluation. Proceedings of EMNLP 2018, 4791–4796. https://aclanthology.org/D18-1512/
- LeCun, Y., Bengio, Y., & Hinton, G. (2015). Deep learning. Nature, 521(7553), 436–444. https://doi.org/10.1038/nature14539
- Lee, M. H. J., Montgomery, J. M., & Lai, C. K. (2024). Large Language Models Portray Socially Subordinate Groups as More Homogeneous, Consistent with a Bias Observed in Humans. Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency (FAccT '24). https://doi.org/10.1145/3630106.3658975
- Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., et al. (2020). BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.acl-main.703
- LTAC Global (2021). Summary of changes in ISO 30042. Introduction to TermBase eXchange (TBX). https://www.tbxinfo.net/summary-of-changes-in-iso-30042/
- Luccioni, A. S., Viguier, S., & Ligozat, A.-L. (2023). Estimating the Carbon Footprint of BLOOM, a 176B Parameter Language Model. Journal of Machine Learning Research, 24(253), 1-15. http://jmlr.org/papers/v24/23-0069.html
- Manning, C. D., Raghavan, P., & Schütze, H. (2008). Introduction to Information Retrieval. Cambridge University Press. https://nlp.stanford.edu/IR-book/information-retrieval-book.html
- Maruf, S., Saleh, F., & Haffari, G. (2021). A survey on document-level neural machine translation: Methods and evaluation. ACM Computing Surveys, 54(2), article 45. https://doi.org/10.1145/3441691
- Maynez, J., Narayan, S., Bohnet, B., & McDonald, R. (2020). On Faithfulness and Factuality in Abstractive Summarization. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.acl-main.173
- McCallister, E., Grance, T., & Scarfone, K. (2010). Guide to Protecting the Confidentiality of Personally Identifiable Information (PII). National Institute of Standards and Technology (NIST SP 800-122). https://doi.org/10.6028/NIST.SP.800-122
- McCorduck, P. (2004). Machines Who Think: A Personal Inquiry into the History and Prospects of Artificial Intelligence. 2nd ed., A K Peters. https://www.routledge.com/Machines-Who-Think-A-Personal-Inquiry-into-the-History-and-Prospects-of-Artificial-Intelligence/McCorduck/p/book/9781568812052
- McKenzie, I. R., Lyzhov, A., Pieler, M., Parrish, A., Mueller, A., Prabhu, A., et al. (2023). Inverse Scaling: When Bigger Isn't Better. Transactions on Machine Learning Research. https://openreview.net/forum?id=DwgRm72GQF
- MDN Web Docs (n.d.). Regular expressions. MDN JavaScript Guide, Mozilla. https://developer.mozilla.org/en-US/docs/Web/JavaScript/Guide/Regular_expressions
- MDN Web Docs (2025). Disjunction: |. MDN Web Docs, JavaScript reference, Regular expressions. https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Regular_expressions/Disjunction
- memoQ (2026). Resource Console: translation memories. memoQ 12.4 help. https://docs.memoq.com/current/en/Workspace/resconsole-tms.html
- memoQ (2026). Translation memory TMX import settings (dialog). memoQ 12.4 help. https://docs.memoq.com/current/en/Workspace/translation-memory-tmx-import-settings.html
- Miao, F., & Holmes, W. (2023). Guidance for generative AI in education and research. UNESCO. https://unesdoc.unesco.org/ark:/48223/pf0000386693
- Microsoft (2026). Regular expression language: quick reference. .NET documentation, Microsoft Learn. https://learn.microsoft.com/en-us/dotnet/standard/base-types/regular-expression-language-quick-reference
- Microsoft (2026). Azure OpenAI reasoning models. Microsoft Learn. https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/reasoning
- Microsoft Learn (n.d.). Localization file formats. https://learn.microsoft.com/en-us/globalization/localization/localization-file-formats
- Microsoft Learn (n.d.). Exchanging localizable resources. https://learn.microsoft.com/en-us/globalization/localization/exchanging-localizable-resources
- Microsoft Learn (n.d.). Maintain translation memories. https://learn.microsoft.com/en-us/globalization/localization/translation-memories
- Microsoft Support (n.d.). Dictate your documents in Word. https://support.microsoft.com/en-us/office/dictate-your-documents-in-word-3876e05f-3fcc-418f-b8ab-db7ce0d11d3c
- Microsoft Support (n.d.). Present with real-time, automatic captions or subtitles in PowerPoint. https://support.microsoft.com/en-us/office/present-with-real-time-automatic-captions-or-subtitles-in-powerpoint-68d20e49-aec3-456a-939d-34a79e8ddd5f
- Mikolov, T., Yih, W., & Zweig, G. (2013). Linguistic Regularities in Continuous Space Word Representations. Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. https://aclanthology.org/N13-1090/
- Mollick, E. (2023). Centaurs and Cyborgs on the Jagged Frontier. One Useful Thing. https://www.oneusefulthing.org/p/centaurs-and-cyborgs-on-the-jagged
- Moorkens, J. (2011). Translation memories guarantee consistency: Truth or fiction?. Proceedings of Translating and the Computer 33 (Aslib). https://aclanthology.org/2011.tc-1.17/
- National Institute of Standards and Technology (n.d.). Deterministic Algorithm. NIST Computer Security Resource Center Glossary. https://csrc.nist.gov/glossary/term/deterministic_algorithm
- Neuberger, C., Bartsch, A., Fröhlich, R., Hanitzsch, T., Reinemann, C., & Schindler, J. (2023). The digital transformation of knowledge order: a model for the analysis of the epistemic crisis. Annals of the International Communication Association, 47(2). https://doi.org/10.1080/23808985.2023.2169950
- NIST Big Data Public Working Group, Definitions and Taxonomies Subgroup (2019). NIST Big Data Interoperability Framework: Volume 1, Definitions (Version 3). NIST Special Publication 1500-1r2, National Institute of Standards and Technology. https://doi.org/10.6028/NIST.SP.1500-1r2
- O'Hagan, M. (Ed.) (2020). The Routledge handbook of translation and technology. Routledge. https://doi.org/10.4324/9781315311258
- O'Shaughnessy, D. (2025). Speaker diarization: A review of objectives and methods. Applied Sciences, 15(4), 2002. https://doi.org/10.3390/app15042002
- OASIS (2025). XLIFF Version 2.2. Organization for the Advancement of Structured Information Standards, Committee Specification 01. https://www.oasis-open.org/standard/xliff-v2-2-cs01/
- OECD (2019). OECD Skills Outlook 2019: Thriving in a Digital World. OECD Publishing. https://doi.org/10.1787/df80bc12-en
- OECD (2021). Methodologies to Measure Market Competition. OECD Competition Law and Policy Working Papers, OECD Publishing. https://doi.org/10.1787/29bf31c1-en
- OECD (1999). Oligopoly: Key findings, summary and notes. OECD Competition Law and Policy Working Papers, OECD Publishing. https://doi.org/10.1787/4f85ebf0-en
- Office québécois de la langue française (2003). Expression rationnelle. Vitrine linguistique, Grand dictionnaire terminologique. https://vitrinelinguistique.oqlf.gouv.qc.ca/fiche-gdt/fiche/8359703/expression-rationnelle
- Office québécois de la langue française (1999). extraction de données. Vitrine linguistique — Grand dictionnaire terminologique. https://vitrinelinguistique.oqlf.gouv.qc.ca/fiche-gdt/fiche/8873867/extraction-de-donnees
- Okapi Framework (n.d.). Open standards. https://okapiframework.org/wiki/index.php/Open_Standards
- Open Source Initiative (2025). Open Weights: not quite what you've been told. Open Source Initiative. https://opensource.org/ai/open-weights
- OpenAI (n.d.). Function calling. OpenAI Developers (API documentation). https://developers.openai.com/api/docs/guides/function-calling
- OpenAI (2023). Custom instructions for ChatGPT. OpenAI. https://openai.com/index/custom-instructions-for-chatgpt/
- OSCAR (LISA) (2008). SRX 2.0 - OSCAR Recommendation. Localization Industry Standards Association. https://www.ttt.org/oscarStandards/srx/srx20.html
- Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., et al. (2022). Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems 35 (NeurIPS 2022). https://proceedings.neurips.cc/paper_files/paper/2022/hash/b1efde53be364a73914f58805a001731-Abstract-Conference.html
- Ouyang, S., Zhang, J. M., Harman, M., & Wang, M. (2025). An Empirical Study of the Non-Determinism of ChatGPT in Code Generation. ACM Transactions on Software Engineering and Methodology, 34(2). https://doi.org/10.1145/3697010
- Park, C. R., Heo, H., Suh, C. H., & Shim, W. H. (2025). Uncover This Tech Term: Application Programming Interface for Large Language Models. Korean Journal of Radiology. https://doi.org/10.3348/kjr.2025.0360
- Peng, Z., Bawden, R., & Yvon, F. (2025). Investigating length issues in document-level machine translation. Proceedings of Machine Translation Summit XX: Volume 1. https://aclanthology.org/2025.mtsummit-1.3/
- Pennington, J., Socher, R., & Manning, C. (2014). GloVe: Global Vectors for Word Representation. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). https://doi.org/10.3115/v1/D14-1162
- Peters, M. E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., et al. (2018). Deep Contextualized Word Representations. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). https://doi.org/10.18653/v1/N18-1202
- Peterson, A. J. (2025). AI and the problem of knowledge collapse. AI & Society, 40(5), 3249-3269. https://doi.org/10.1007/s00146-024-02173-x
- Petrov, A., La Malfa, E., Torr, P., & Bibi, A. (2023). Language Model Tokenizers Introduce Unfairness Between Languages. Advances in Neural Information Processing Systems 36 (NeurIPS 2023). https://proceedings.neurips.cc/paper_files/paper/2023/hash/74bb24dca8334adce292883b4b651eda-Abstract-Conference.html
- Peykani, P., Ghanidel, S., Javadi-Sisi, I., Snasel, V., & Mirjalili, S. (2026). A Holistic Review of Agentic AI Frameworks, Applications, and Research Trajectories. Archives of Computational Methods in Engineering. https://doi.org/10.1007/s11831-026-10675-8
- Phillipson, R. (1992). Linguistic Imperialism. Oxford University Press (Oxford Applied Linguistics). https://books.google.com/books?id=4jVeGWtzQ1oC
- Polonioli, A. (2025). Moving LLM evaluation forward: lessons from human judgment research. Frontiers in Artificial Intelligence. https://doi.org/10.3389/frai.2025.1592399
- Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C., & Sutskever, I. (2023). Robust speech recognition via large-scale weak supervision. Proceedings of the 40th International Conference on Machine Learning (PMLR 202), 28492–28518. https://proceedings.mlr.press/v202/radford23a.html
- Radford, A., Narasimhan, K., Salimans, T., & Sutskever, I. (2018). Improving Language Understanding by Generative Pre-Training. OpenAI (preprint). https://cdn.openai.com/research-covers/language-unsupervised/language_understanding_paper.pdf
- Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., & Sutskever, I. (2019). Language Models are Unsupervised Multitask Learners. OpenAI. https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf
- Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., & Finn, C. (2023). Direct Preference Optimization: Your Language Model is Secretly a Reward Model. Advances in Neural Information Processing Systems 36 (NeurIPS 2023). https://proceedings.neurips.cc/paper_files/paper/2023/hash/a85b405ed65c6477a4fe8302b5e06ce7-Abstract-Conference.html
- Ramnath, K., Zhou, K., Guan, S., Mishra, S. S., Qi, X., Shen, Z., et al. (2025). A systematic survey of automatic prompt optimization techniques. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. https://doi.org/10.18653/v1/2025.emnlp-main.1681
- Ramos, M. M., Fernandes, P., Agrawal, S., & Martins, A. F. T. (2025). Multilingual contextualization of large language models for document-level machine translation. Second Conference on Language Modeling (COLM 2025). https://openreview.net/forum?id=Ah0U1r5Ldq
- Raschka, S. (2024). Machine Learning Q and AI. No Starch Press. https://nostarch.com/machine-learning-q-and-ai
- Raya, R. (2005). XML in localisation: Reuse translations with TM and TMX. Maxprograms. https://www.maxprograms.com/articles/tmx.html
- Raya, R. (2004). XML in localisation: Use XLIFF to translate documents. Maxprograms. https://www.maxprograms.com/articles/xliff.html
- Rothwell, A., Moorkens, J., Fernández-Parra, M., Drugan, J., & Austermuehl, F. (2023). Translation Tools and Technologies. Routledge. https://doi.org/10.4324/9781003160793
- Roukos, S., Graff, D., & Melamed, D. (1995). Hansard French/English (LDC95T20). Linguistic Data Consortium. https://catalog.ldc.upenn.edu/LDC95T20
- Sager, P. J., Meyer, B., Yan, P., von Wartburg-Kottler, R., Etaiwi, L., Enayati, A., et al. (2026). A Comprehensive Survey of Agents for Computer Use: Foundations, Challenges, and Future Directions. Journal of Artificial Intelligence Research. https://doi.org/10.1613/jair.1.19490
- Salvi del Pero, A., Wyckoff, P., & Vourc'h, A. (2022). Using Artificial Intelligence in the workplace: What are the main ethical risks?. OECD Social, Employment and Migration Working Papers, No. 273, OECD Publishing. https://doi.org/10.1787/840a2d9f-en
- Schaeffer, R., Miranda, B., & Koyejo, S. (2023). Are Emergent Abilities of Large Language Models a Mirage?. Advances in Neural Information Processing Systems 36 (NeurIPS 2023). https://proceedings.neurips.cc/paper_files/paper/2023/hash/adc98a266f45005c403b8311ca7e8bd7-Abstract-Conference.html
- Schick, T., Dwivedi-Yu, J., Dessi, R., Raileanu, R., Lomeli, M., Hambro, E., et al. (2023). Toolformer: Language Models Can Teach Themselves to Use Tools. Advances in Neural Information Processing Systems 36 (NeurIPS 2023). https://proceedings.neurips.cc/paper_files/paper/2023/hash/d842425e4bf79ba039352da0f658a906-Abstract-Conference.html
- Schwartz, R., Vassilev, A., Greene, K. K., Perine, L., Burt, A., & Hall, P. (2022). Towards a standard for identifying and managing bias in artificial intelligence. NIST Special Publication 1270, National Institute of Standards and Technology. https://doi.org/10.6028/NIST.SP.1270
- Sennrich, R., Haddow, B., & Birch, A. (2016). Neural Machine Translation of Rare Words with Subword Units. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). https://doi.org/10.18653/v1/P16-1162
- Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., & Gal, Y. (2024). AI models collapse when trained on recursively generated data. Nature, 631. https://doi.org/10.1038/s41586-024-07566-y
- Sketch Engine (n.d.). frTenTen: French corpus from the web. sketchengine.eu. https://www.sketchengine.eu/frtenten-french-corpus/
- Sketch Engine (2020). Gutenberg Corpora 2020. sketchengine.eu. https://www.sketchengine.eu/gutenberg-corpora-2020/
- Snell, C., Lee, J., Xu, K., & Kumar, A. (2025). Scaling LLM Test-Time Compute Optimally Can be More Effective than Scaling Parameters for Reasoning. Proceedings of the Thirteenth International Conference on Learning Representations (ICLR 2025). https://proceedings.iclr.cc/paper_files/paper/2025/hash/1b623663fd9b874366f3ce019fdfdd44-Abstract-Conference.html
- Stanford HAI (n.d.). What is Attention Mechanism?. Stanford HAI, AI definitions. https://hai.stanford.edu/ai-definitions/what-is-attention-mechanism
- Stanford HAI (n.d.). What is a Context Window?. Stanford HAI, AI definitions. https://hai.stanford.edu/ai-definitions/what-is-a-context-window
- Stanford HAI (n.d.). What is an Open-Weight Model?. Stanford HAI, AI definitions. https://hai.stanford.edu/ai-definitions/what-is-an-open-weight-model
- Suresh, H., & Guttag, J. V. (2021). A Framework for Understanding Sources of Harm throughout the Machine Learning Life Cycle. Proceedings of the 1st ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization (EAAMO '21). https://doi.org/10.1145/3465416.3483305
- Tiedemann, J. (2012). Parallel Data, Tools and Interfaces in OPUS. Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC'12), Istanbul, 2214–2218. https://aclanthology.org/L12-1246/
- Unicode Consortium (n.d.). The Unicode Standard: A Technical Introduction. https://www.unicode.org/standard/principles.html
- Unicode Consortium (n.d.). Glossary of Unicode Terms. The Unicode Consortium. https://www.unicode.org/glossary/
- United States (1976). Copyright Act of 1976, 17 U.S.C. § 107 (Limitations on exclusive rights: Fair use). United States Code. https://uscode.house.gov/view.xhtml?req=granuleid:USC-prelim-title17-section107&num=0&edition=prelim
- Vassilev, A., Oprea, A., Fordyce, A., Anderson, H., Davies, X., & Hamin, M. (2025). Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations. National Institute of Standards and Technology (NIST AI 100-2 E2025). https://doi.org/10.6028/NIST.AI.100-2e2025
- Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., et al. (2017). Attention Is All You Need. Advances in Neural Information Processing Systems 30 (NIPS 2017). https://papers.nips.cc/paper_files/paper/2017/hash/3f5ee243547dee91fbd053c1c4a845aa-Abstract.html
- Vig, J. (2019). A Multiscale Visualization of Attention in the Transformer Model. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, 37–42. https://aclanthology.org/P19-3007/
- Voita, E., Sennrich, R., & Titov, I. (2019). When a good translation is wrong in context: Context-aware machine translation improves on deixis, ellipsis, and lexical cohesion. Proceedings of the 57th Annual Meeting of the ACL, 1198–1212. https://aclanthology.org/P19-1116/
- W3C (2008). Extensible Markup Language (XML) 1.0 (Fifth Edition). W3C Recommendation. https://www.w3.org/TR/xml/
- W3C (2002). Localization-Related Formats. https://www.w3.org/2002/02/01-i18n-workshop/LocFormats
- W3C (2009). Namespaces in XML 1.0 (Third Edition). W3C Recommendation. https://www.w3.org/TR/xml-names/
- W3C (n.d.). Character encodings for beginners. W3C Internationalization. https://www.w3.org/International/questions/qa-what-is-encoding
- W3C (n.d.). Choosing & applying a character encoding. W3C Internationalization. https://www.w3.org/International/questions/qa-choosing-encodings
- Wang, L., & Sun, S. (2023). Dictating translations with automatic speech recognition: Effects on translators' performance. Frontiers in Psychology, 14, 1108898. https://doi.org/10.3389/fpsyg.2023.1108898
- Wang, Y., Wang, M., Manzoor, M. A., Liu, F., Georgiev, G. N., Das, R. J., et al. (2024). Factuality of Large Language Models: A Survey. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. https://aclanthology.org/2024.emnlp-main.1088/
- Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., et al. (2022). Emergent Abilities of Large Language Models. Transactions on Machine Learning Research. https://openreview.net/forum?id=yzkSU5zdwD
- Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., et al. (2022). Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. Advances in Neural Information Processing Systems 35 (NeurIPS 2022). https://proceedings.neurips.cc/paper_files/paper/2022/hash/9d5609613524ecf4f15af0f7b31abca4-Abstract.html
- WHATWG (2026). HTML Standard. Living Standard. https://html.spec.whatwg.org/multipage/
- Wu, C.-J., Raghavendra, R., Gupta, U., Acun, B., Ardalani, N., Maeng, K., et al. (2022). Sustainable AI: Environmental Implications, Challenges and Opportunities. Proceedings of Machine Learning and Systems 4 (MLSys 2022). https://proceedings.mlsys.org/paper_files/paper/2022/hash/462211f67c7d858f663355eff93b745e-Abstract.html
- Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., et al. (2023). ReAct: Synergizing Reasoning and Acting in Language Models. The Eleventh International Conference on Learning Representations (ICLR 2023). https://arxiv.org/abs/2210.03629
- Yuan, R., Sun, S., Li, Y., Wang, Z., Cao, Z., & Li, W. (2025). Personalized Large Language Model Assistant with Evolving Conditional Memory. Proceedings of the 31st International Conference on Computational Linguistics. https://aclanthology.org/2025.coling-main.254/
- Zekpa, N., & Peter, A. (2025-01-07). Evaluate large language models for your machine translation tasks on AWS. AWS Machine Learning Blog. https://aws.amazon.com/blogs/machine-learning/evaluate-large-language-models-for-your-machine-translation-tasks-on-aws/
- Zhao, J., Wang, T., Yatskar, M., Ordonez, V., & Chang, K.-W. (2017). Men Also Like Shopping: Reducing Gender Bias Amplification using Corpus-level Constraints. Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing. https://doi.org/10.18653/v1/D17-1323
- Zhong, W., Guo, L., Gao, Q., Ye, H., & Wang, Y. (2024). MemoryBank: Enhancing Large Language Models with Long-Term Memory. Proceedings of the AAAI Conference on Artificial Intelligence, 38(17). https://doi.org/10.1609/aaai.v38i17.29946
- Ziemski, M., Junczys-Dowmunt, M., & Pouliquen, B. (2016). The United Nations Parallel Corpus v1.0. Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16), Portorož, 3530–3534. https://aclanthology.org/L16-1561/
- Zuboff, S. (2019). The Age of Surveillance Capitalism: The Fight for a Human Future at the New Frontier of Power. PublicAffairs. https://www.hachettebookgroup.com/titles/shoshana-zuboff/the-age-of-surveillance-capitalism/9781610395694/