Saltar al contenido principal

Escribir un comentario

PREreview del Translating the Untranslatable: An Operationalizable Ontology for Untranslatability

Publicado
DOI
10.5281/zenodo.21893031
Licencia
CC BY 4.0

This review is the result of a virtual, collaborative live review discussion organized and hosted by Translate Science. The discussion was joined by three people. The authors of this review have dedicated additional asynchronous time over the course of two weeks to help compose this final report using the notes from the Live Review. We thank all participants who contributed to the discussion and made it possible for us to provide feedback on this preprint.

We had a good experience reading and discussing this paper, finding it easy to understand (even for those of us without expert knowledge of machine translation or linguistics), logically structured, and enriched with many specific and relevant examples. The abstract starts with a concise, provocative statement and clearly outlines the work.

The paper defines untranslatability as an area of research in its own right for natural language processing (NLP) with specific interest for its subfield machine translation (MT). By doing this, the authors challenge the notion of translation as a solved problem. At the same time, they turn to a problem which has often been neglected: Implicitly, in MT the task of translation is seen as something that is a priori generally feasible, that is, translatability is simply assumed and the problem of untranslatability unnoticed. To define untranslatability types and compensation strategies, the authors review a limited amount of literature from Translation and Interpreting Studies (TIS). Based on their conclusions from this literature, they construct a reusable data set of source sentences in Spanish and Japanese in which at least one key aspect is designed to be untranslatable. They also identify potentially relevant domains such as grammar, language use, formal constraints of a language or cultural specificities of a language community as well as the translation contexts ‘textbook’ and ‘movie’. To test how these factors may affect the acceptability of compensation strategies, the sentence pairs from the data set are assessed by crowd-sourced participants for the acceptability of the compensation strategies used and the result of their assessments is analysed according to the effect of the previously defined domains and translation contexts on acceptability.

The paper does not explicitly state any research questions. If the authors intended to address specific research questions, such as "Can an ontology of types of untranslatability improve machine translation?" it would be good to state those explicitly early in the paper. As reviewers, we generally approached this preprint as having goals rather than answering research questions.

List of major concerns and feedback:

Literature and theory

The authors communicated an ontology of untranslatability as economic value that would require prior cross-linguistic work. However, we do not find that the authors guided us through this prior work sufficiently to prepare us to accept the validity of their ontology. Appendix A provides examples of the authors’ untranslatable types, but does not support it with prior linguistic evidence. This lowers our confidence that the authors’ interpretation of the strategy preference is properly contextualized in the scholarship of translation. For example in section 6.3, the authors conclude that a difference in cType preference from different source languages reflects on cultural distance. However, this difference could be explained by variance between lingua-cultural domains, or between between different groups of annotators and their backgrounds, preferences and knowledge.

The rather formal view of the aspects that are assumed to be untranslatable should be complemented by the functional view which stresses the notion of the adequacy of a message in order to reach an intended purpose rather than clinging to the formal identity of messages. This view provides access to more methodological research approaches and, in practice, results in more options for contextually adequate compensations becoming available. Nord (2006) presents a brief overview especially on functionalism.

The study focuses very much on the sentence level, leaving out the possibility for co-textual strategies such as replacing a pun or a metaphor in one place with a functional equivalent in a different place in order to reach the intended effect (see, e. g., Samaniego Fernández's research on metaphor translation, Samaniego Fernández 2013; strictly speaking, co-text means the surrounding textual material whereas context refers to the greater situational and socio-historic embedding of an utterance).

Similarly, the study does not consider the large body of work on cognitive approaches to translation, some of which offer a formalisation of the conceptual content both in semantics (e. g., Czulo 2017, also for an overview) and pragmatics (Triesch-Herrmann and Czulo 2024), the former of which has been used in machine translation evaluation (Czulo et al. 2019).

The omission of functional and cognitive approaches should be acknowledged as limitations of the study.

LLMs and Machine Translation

The study uses LLMs to generate examples. This may result in effects such as "accent" (Papadimitrou et al. 2023) and "epistemological persistence" (Barkin 2025), both creating a strong bias in LLMs towards English when it comes to generating linguistic structures and cultural content. These problems should be acknowledged more explicitly and earlier in the paper rather than addressed generally as bias in the Limitations section. Further, the authors of source texts are often considered a relevant audience for translations, especially for literary texts. Using some human-generated texts (possibly from text corpora) would avoid these biases and also potentially allow for authors to play a role in evaluating translations. Also, in steps that mention use of an LLM, make it clear exactly which model was used.

We would appreciate more explanation of the limitations of the MT model the authors are seeking to improve. In section 3.3 they posit that the predicted translation of a target sentence can be improved by considering a prediction of the compensation strategy to use and the untranslatability type. In this model, both the target sentence and compensation strategy prediction are based on the translation medium/context. This raises the question: how is translation medium/context known? It could be the result of a predictive function as well. Further discussion of the conceptual MT model along with details of potential implementations and attention paid to connecting the MT model to concepts in linguistics and translation would be very helpful to understand better the importance of this work.

The study does not clarify its relation to the MQM model (https://themqm.org) which served as basis for ISO 5060 on the evaluation of translation (including machine translation). While directed at quality assessment, the study may benefit from connecting some of the supercategories and error types to untranslatability dimensions or compensation strategies.

Tables and Figures

The inclusion of many specific text examples in charts helps make the study’s message clear, interesting, and memorable. We do have several recommendations to improve the effectiveness of the charts. Some of these recommendations apply throughout the paper: use black, white, and grayscale with adequate color contrast so that figures are legible in print; if a few patterns must be used, choose patterns that are very easy to tell apart; minimize use of abbreviations; and, when abbreviations are used, spell them out on first use so that each table or figure stands alone. Also, number tables and figures in the order that they are referenced in the text and number tables and figures in appendixes separately.

We especially believe that Figure 2 could be clarified. We note that the colors are practically indistinguishable in print, and you may not need color to create a clear figure. The colored legend at the bottom seems to be redundant with the numbered steps. The graphic icons could be omitted. However, if you are going to use icons, use a computer icon rather than a human head in step 2, because this step is done by a computer. Also, step 3 would be an opportunity to signal visually that each example is reviewed by three experts. In the heading for step 3, refer to Human experts rather than just Experts to be consistent with steps 1 and 4. The current placement of the accept and reject arrows in step 3 is confusing, as the arrow does not branch from the matching word. It would also be helpful to use the standard flow chart convention of a diamond to show a decision.

We have a few additional specific recommendations related to other charts. In Figure 3, it is not clear if the subcategories are in a meaningful order. If they are in order, explain the order. If they are not in order, we recommend placing them in alphabetical order. The small patterns in Figure 4 and others like it are difficult to read in print. For Figure 4, the chart would be more intuitive if the most popular strategy (adaptation) were graphed in white and the no strategy option was graphed in black. The three adjacent patterns based on evenly spaced diagonal lines are especially hard to distinguish. If patterns must be used, choose patterns that are very distinct, especially where they are adjacent. In the caption for Figure 5, it would be good to mention the decrease in the no strategy option’s rank. In Figure 7, it would be helpful to show all untranslatability types and all compensation strategies rather than collapsing them. We found that the shadowed boxes in Figure 7 reduced readability without adding information and recommend removing them. In Tables 7 and 8, report the percentage column as a percentage rather than a proportion.

Transparency, Replicability, and Limitations

There are several ways the transparency and replicability of the study could be improved. We appreciate the intent to provide an open dataset. However, we were unable to locate it as the GitHub link in the preprint gave a 404 error. It would also be good to include more information about how unsuccessful examples were handled. Was refinement always successful? Were you able to get exactly 18,200 translations? We might have a better understanding of this if we had been able to access the dataset.

We also appreciate the inclusion of a Limitations section. We recommend placing the Limitations section before the conclusion, because conclusions should be conditional on the limitations of the study. We did not find information about whether the study had been through human subjects research review by an institutional review board (IRB) or determined to be exempt from that requirement. In the interest of transparency, we recommend referencing the initial intent to include Chinese in the body of the paper, perhaps in the limitations section, rather than only in Appendix B. The fact that your approach only worked for some languages is a relevant limitation. Including the information about the intent to include Chinese would also help clarify why some of the examples are in Chinese. Additional limitations of the evaluators should be included. We would have found it helpful to know more information about how the evaluators from Prolific were selected, assigned, and compensated. They are described as bilinugal, but are we sure they are not multilingual? To perform the evaluation, they are asked to respond as if they are monolingual English speakers, but that is not the same as being monolingual English speakers. Also, the audience for English translations includes many other people, not only monolingual English speakers.

List of minor concerns and feedback:

We noted a couple of additional minor issues and list them here for reference:

  • Rather than combining religion and philosophy in the ontology, consider making them separate classes or using an inclusive term based on what you see as their shared characteristics, such as "moral-ethical frameworks."

  • In section 6.2 some discussion about why the contextual dataset is smaller and explanation as to why the analysis can't be broken up by uType would be helpful.

  • The Spanish term sobremesa does not exist in English, but there are similar terms in other languages (for example, Dutch).

Overall Assessment

We recognize that as a conference paper it may not have been feasible to include a comprehensive literature review, although we believe that such a review would contribute meaningfully to the study. The preprint was highly readable and engaging and we believe it would be effective for use in class discussions around the concept of untranslatability as well more general discussions of translation strategies.

Concluding remarks

We thank the authors of the preprint for posting their work openly for feedback.

References

* Barkin, Gareth. 2025. Epistemological persistence in multilingual ai: The illusion of locality in large language models. International Review of Modern Sociology 51(1–2). https://www.academia.edu/download/129301512/Epistemological_Persistence_in_Multilingual_AI_The_Illusion_of_Locality_in_Large_Language_Models_PRINT.pdf.

* Czulo, Oliver. 2017. Aspects of a primacy of frame model of translation. In Silvia Hansen-Schirra, Oliver Czulo & Sascha Hofmann (eds.), Empirical modelling of translation and interpreting (Translation and Multilingual Natural Language Processing 6), 465–490. Berlin: Language Science Press. https://langsci-press.org/catalog/view/132/1053/917-1.

Czulo, Oliver, Tiago Timponi Torrent, Ely Edison da Silva Matos, Alexandre Diniz da Costa & Debanjana Kar. 2019. Designing a frame-semantic machine translation evaluation metric. In Proceedings of the human-informed translation and interpreting technology workshop (HiT-IT 2019), 28–35. https://aclanthology.org/W19-8704.pdf. Nord, Christiane. 2006. Translating for communicative purposes across culture boundaries. Journal of Translation Studies 9(1). 43–60. https://www.academia.edu/download/104128518/nord-2006translatingforcommpurposes_jots-935-eng.pdf.

* Papadimitriou, Isabel, Kezia Lopez & Dan Jurafsky. 2023. Multilingual BERT has an accent: Evaluating English influences on fluency in multilingual models. In Findings of the Association for Computational Linguistics: EACL 2023, 1194–1200. Dubrovnik, Croatia: Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.findings-eacl.89.

* Samaniego Fernández, Eva. 2013. The impact of Cognitive Linguistics on Descriptive Translation Studies: novel metaphors in English-Spanish newspaper translation as a case in point. In Ana Rojo & Iraide Ibarretxe-Antuñano (eds.), Cognitive linguistics and translation: advances in some theoretical models and applications (Applications of Cognitive Linguistics 23), 159–198. Berlin: De Gruyter Mouton.

* Triesch-Herrmann, Susanne & Oliver Czulo. 2024. A frame-based analysis of the pragmatics and semantics of “bekanntlich” in English-German translation. trans-kom. Zeitschrift für Translationswissenschaft und Fachkommunikation 17(1). 130–146. https://www.trans-kom.eu/bd17nr01/trans-kom_17_01_08_Triesch-Herrmann_Czulo_bekanntlich.20240628.pdf.

Competing interests

The authors declare that they have no competing interests.

Use of Artificial Intelligence (AI)

The authors declare that they did not use generative AI to come up with new ideas for their review.

Puedes escribir un comentario en esta PREreview de Translating the Untranslatable: An Operationalizable Ontology for Untranslatability.

Antes de comenzar

We will ask you to log in with your ORCID iD. If you don’t have an iD, you can create one.

What is an ORCID iD?

An ORCID iD is a unique identifier that distinguishes you from everyone with the same or similar name.

Comenzar ahora