As Australian and global businesses push further into new markets, the language layer underneath that expansion is under more scrutiny than it has ever been.
Artificial intelligence has made translation faster and cheaper than at any point in the industry’s history.
But new comparative research is revealing a harder truth for businesses relying on the technology. Not all AI models translate to the same standard.
The gap becomes even wider when content moves beyond the world’s most commonly used languages.
The scale of the shift is already visible in the data. Intento’s ninth annual State of Translation Automation report, which evaluated 46 machine translation engines and large language models across 11 language pairs, found no single off-the-shelf model held the top spot across the board.
Instead, workflows that combined and cross-checked multiple models produced the best result in nine of the eleven language pairs tested, a result the report’s authors say reflects a broader move away from single-engine translation toward layered, verification-based approaches.
That pattern shows up clearly in comparative testing run by Tomedes, a professional translation company, which benchmarked leading AI models against the same source material.
On a mixed batch of technical and marketing text, one model produced the most natural, human-sounding output for major European languages, while a second model handled technical instruction manuals more precisely but struggled with marketing idiom and slang.
A third model corrected typos reliably but fabricated facts in two separate sentences when the source content involved recent events, an error type distinct from the grammar mistakes that used to define machine translation.
In a related test on current-affairs content, one model scored 92% on factual accuracy against a rival’s 85%, largely because the weaker model’s training data had not caught up with recent developments.
The benchmarking gets more pronounced once language complexity is factored in. Single AI models plateau at roughly 84 to 87% accuracy on French, German and Spanish, according to the research, largely due to formatting errors and terminology drift that persist even in advanced systems.
For a morphologically complex language like Polish, that figure drops to 76%.
When outputs from multiple models are reconciled against each other rather than taken at face value, accuracy climbs to 93 to 95% for major Western and Southern European languages, and to 88% for Polish, a gap the researchers attribute to the fact that different AI models tend to make different kinds of mistakes rather than the same ones.
Ofer Tirosh, CEO of Tomedes, says the findings point to a mistake many companies are still making by default.
“The assumption a lot of businesses carry into international expansion is that one capable AI model will perform consistently no matter which market they are entering,” Tirosh said.
“What the data actually shows is that performance is highly language-dependent. A model that looks reliable in French can lose ten points of accuracy or more the moment you move into a market with a more complex grammar structure.”
“Businesses that check AI output against more than one model, rather than trusting a single result by default, are the ones catching the errors before they reach a client or a regulator.”
The stakes attached to getting this wrong are commercial, not just linguistic.
CSA Research’s global survey of more than 8,700 consumers found 76% prefer to buy products described in their own language, and 40% say they will never buy from a website in another language at all.
For a business entering a new market, a translation that is subtly wrong, rather than obviously broken, is arguably the more expensive failure: it does not bounce a customer immediately, it slowly erodes the trust the brand needed to close the sale.
That risk is reshaping how companies budget for AI translation in the first place.
Rather than defaulting to a single subscription or AI model for every job, more businesses are matching their translation approach to the project in front of them.
The shift in how businesses are using AI to go global reflects the way companies increasingly use translation on demand.
For many businesses, translation is becoming a project-driven service rather than a constant background utility.
The direction of travel across the research is consistent: the businesses gaining ground are not the ones betting on a single AI model being good enough everywhere, but the ones treating cross-model verification as standard practice before content reaches a customer, a contract or a regulator.
As global businesses compete to enter new markets faster, the difference between checking one AI model and comparing several is becoming increasingly important.
Cross-checking translations across multiple models is emerging as a higher professional standard, particularly where accuracy matters.
Businesses that are slow to adapt could face costly mistakes when poor translations reach customers, partners or regulators.

