Skip to content

lingua intelligenza artificiale

If you don’t speak English, it’s better not to ask the AI. Economist Report

Why large artificial intelligence models are less accurate, more expensive, and potentially risky for non-English speakers, and what is being done to bridge this digital divide. The Economist article.

 

To get the most accurate answer from a large language model (LLM), it is crucial to query it in the right language. An English-speaking user asking a world-leading model what to do about swollen legs in late pregnancy, for example, might be warned to watch out for pre-eclampsia. A future mother who speaks Swahili, however, would be more likely to be told not to worry – writes the Economist.

ACCURACY DIFFERENCES BETWEEN LANGUAGES

This illustrates a widespread problem: even when the English version of a model passes safety tests, it can still generate hallucinations and dangerous misinformation in other languages. […] In research published in October 2025, scholars found that accuracy in non-English languages was about 12-29 percentage points lower than in English, depending on the model used.

URGENCY OF THE PROBLEM IN NON-ENGLISH-SPEAKING REGIONS

The problem is becoming urgent with the acceleration of LLM usage in non-English-speaking regions. If tools intended for medical diagnosis and triage do not account for the language gap, they may not be up to the task. Two researchers working to establish the extent of this gap are Tuka Alhanai from New York University Abu Dhabi and Mohammad Ghassemi from Michigan State University. In February 2025, they released a “benchmark”: a test for LLMs’ ability to understand other languages.

IMPACT OF ENGLISH DOMINANCE IN DATA

[…] The dominance of English-language data not only affects the responses provided by LLMs but also shapes how they operate. Before processing text, models break it down into small units known as tokens. Models trained predominantly in English often fragment texts in other languages inefficiently, requiring more tokens to express the same meaning. Since developers pay for model access based on the number of tokens processed, the same command can cost up to five times more in another language than in English.

LIMITS OF MULTILINGUAL MODELS

Even explicitly multilingual models succumb to these pressures. Research from May 2025 shows that the model often answers non-English questions by first retrieving facts in English and translating the response only at the final stage. Adding such steps introduces further opportunities for error.

POSSIBLE SOLUTIONS AND FUTURE OUTLOOK

Fortunately, adding even small amounts of non-English data to training can help improve performance. Dr. Alhanai and her team found that fine-tuning a model with a small number of high-quality samples increases accuracy in that language by over five percentage points. A more intensive approach involves redesigning how models break text into tokens. For now, however, Dr. Alhanai states, “the people who would benefit most are those least able to use these tools.”

(Excerpt from the foreign press review curated by eprcomunicazione)

Back To Top