Aller directement au contenu principal

Rédiger un PREreview

Os LLMs compreendem expressões idiomáticas? Evidências para uma teoria da Competência Fraseológica Artificial Simulada

Publié
Serveur de preprints
SciELO Preprints
DOI
10.1590/scielopreprints.16937

Introduction: Large Language Models (LLMs) have achieved substantial advances in Natural Language Processing, yet the processing of phraseological units, particularly idiomatic expressions (IEs), remains theoretically challenging. Their multiword, conventionalised and frequently non-compositional nature requires more than the recognition of lexical patterns, as successful interpretation may involve semantic, pragmatic, cultural and sociocognitive knowledge. This study therefore distinguishes between human phraseological competence, grounded in embodied experience, sociocultural participation and pragmatic inference, and the functional performance of artificial systems, which is primarily based on statistical-distributional regularities. Objective: The study aims to investigate whether a large language model exhibits an emergent operational form of phraseological competence when processing Brazilian Portuguese idiomatic expressions and to determine the extent to which different prompting strategies affect its performance. To this end, the paper proposes the concept of simulated artificial phraseological competence, understood as the functional capacity to recognise, interpret and reproduce phraseological units without implying the presence of sociocognitive mechanisms equivalent to those underlying human phraseological competence. Methodology: A mixed-method, quasi-controlled experiment was conducted using Claude Haiku 4.5. The experimental corpus comprised 140 Brazilian Portuguese phraseological units, equally distributed across opaque idioms, somatic expressions, culturally marked idioms and fabricated expressions designed as experimental counter-evidence. Each unit was submitted to three independent prompting conditions: zero-shot, contextualised few-shot, and Chain-of-Thought combined with role-playing. Performance was assessed across four dimensions: idiomaticity detection (D1), semantic accuracy (D2), pragmatic appropriateness (D3), and cultural sensitivity (D4). Results: The findings indicate high semantic accuracy for conventionalised expressions, with mean scores ranging from 2.56 to 3 across different subsets. Pragmatic appropriateness was more heterogeneous, whereas cultural sensitivity consistently represented the weakest dimension, particularly in the identification of culturemes, cultural symbolism and sociocultural motivations. The analysis also revealed recurrent figurative hallucination, whereby non-existent expressions were accepted and explained as genuine idiomatic units. By contrast, few-shot prompting and, particularly, Chain-of-Thought produced qualitative improvements. The CoT condition achieved a with a high degree of correct rejection rate for fabricated expressions. Conclusion: The findings support the hypothesis that the model exhibits an emergent operational form of phraseological competence, but not a sociocognitive competence equivalent to that of human speakers. The system can simulate certain outcomes traditionally associated with phraseological competence through statistical patterning, yet remains vulnerable to failures involving conventionalisation, pragmatic inference and cultural anchoring. At the same time, the results demonstrate that prompt engineering can partially compensate for these limitations, particularly when prompts incorporate examples, contextualisation, empirical validation and explicit rejection criteria.

Vous pouvez rédiger un PREreview de Os LLMs compreendem expressões idiomáticas? Evidências para uma teoria da Competência Fraseológica Artificial Simulada. Un PREreview est une évaluation d'un preprint et peut varier de quelques phrases à un rapport détaillé, semblable à un rapport d'évaluation par les pairs organisé par une revue.

Avant de commencer

Nous vous demanderons de vous connecter avec votre identifiant ORCID iD. Si vous n'en avez pas, vous pouvez en créer un.

Qu’est-ce qu’un ORCID iD ?

Un ORCID iD est un identifiant unique qui vous distingue de toute personne ayant le même nom ou nom similaire.

Commencer maintenant