Multilingual AI in Support: How to Verify Quality and Consistency
Provide useful answers in multiple languages without mixing up terminology, conditions, and sources.
SqualiOnline editorial team · 2026-09-07
An assistant that answers in multiple languages doesn't break by translating badly. It breaks by answering well in Italian and, for the same question, saying something different in German, because the terms of sale, delivery times, or warranty aren't the same across every market, and it isn't drawing on the same content. Verifying multilingual support isn't about verifying a language: it's about verifying the consistency between what you say and where you say it.
The operational question isn't whether it speaks well, but whether a French customer and an Italian customer asking the same thing get two answers that are each correct for their own case.
Language and market aren't the same thing
This is the mix-up that generates the most errors. Spanish is spoken in many countries with different terms, and a Swiss customer might write in Italian while having shipping, taxes, and return terms completely different from an Italian customer's.
- Decide which languages you support and, separately, which markets you serve: these are two lists that almost never line up.
- For each market, establish which content applies: terms of sale, lead times, returns, warranty, availability, price lists, specific requirements.
- Establish what happens when the language doesn't identify the market. Asking for the country before answering about a condition that varies is more correct than guessing.
- Point to a single source for each condition: if the same information sits on three different pages for three markets, decide which one governs.
The glossary: what doesn't get translated
Before any testing, you need a shared list of terms, otherwise every check ends up as a debate about the personal taste of whoever's reviewing it.
- Product and line names, which are never translated, even when they look like ordinary words.
- Codes and abbreviations, which must stay identical across every language.
- The technical terms of your industry, with the translation accepted in that market, which isn't always the dictionary one.
- Terms with legal meaning, like warranty, withdrawal, compliance, or delivery, where the word chosen changes the commitment you're making.
- Formats: dates, numbers, units of measure, currency. A date written the American way on an order confirmation produces a phone call.
- The register — that is, formal or informal address where the language calls for it — decided per market and not left to chance.
The test sample: equivalent questions, not translated ones
Translating test questions from one language to another produces an easy sample, because it mirrors the structure of the original. What's needed are questions a customer in that market would actually ask, phrased by someone who speaks that language.
| Situation to test | What to observe | Typical error |
|---|---|---|
| Question about return terms | That the timeframe and procedure match the market of whoever is asking | Giving the main market's terms, translated |
| Question about availability and delivery times | That the times given apply to that country | Giving domestic delivery times to a customer abroad |
| Ambiguous or incomplete question | That the system asks for clarification instead of picking one interpretation | Answering the most likely question without saying so |
| Question with a technical industry term | That the glossary term gets used, not a synonym | Translating literally and producing a word nobody uses |
| Out-of-coverage request | That the system states so and hands off to a person | Building a plausible answer out of whatever it has |
| Complaint or sensitive request | That the tone is appropriate and a human operator is offered | Answering with the same register as a sales question |
The same set of situations needs to be tested in every supported language and repeated after every change to the content or the rules. That's what makes the results comparable over time, instead of starting from scratch every time.
Who evaluates, and on what basis
The evaluation needs to be done by people who know both the language and your product. A translator who doesn't know your commercial terms can only judge the phrasing, and the phrasing is the part that breaks the least.
- Correctness against that market's terms: this is the main criterion and takes precedence over everything else.
- Completeness: does the answer contain what's needed to act on it, or does it force the customer to write again?
- Use of the glossary: terms, names, and formats as established.
- Tone and register appropriate to the market.
- Behavior when information is missing: stating so counts as a correct answer, making something up counts as a serious error even if the sentence is flawless.
It's worth logging outcomes on a simple scale — correct, correct but incomplete, wrong — and keeping the wrong answers on file: they're the material you use to fix the instructions and the content, and they show whether errors cluster around a particular topic or a particular language.
Out of coverage: saying you don't know is an answer
Every supported language has a boundary, and that boundary shifts with the content available. The rules need to be written explicitly, otherwise the system will fill in the gaps on its own.
- If the content needed doesn't exist in that language and can't be verified, the system doesn't improvise it.
- If it only exists in another language, it can be offered while saying so, or handed off to a person: this is a choice to make market by market, not a single blanket rule.
- The handoff to an operator needs to take into account that country's business hours and whether someone speaks that language. Promising a callback in a language nobody at the company speaks is worse than not offering one.
What this guide doesn't cover
This guide covers the quality and consistency of answers in multilingual support. How to structure a multilingual website — addresses, country versions, and the content to prepare beyond translation — is covered in the guide on multilingual websites for selling abroad. General testing of an assistant before it goes live has a guide of its own.
Frequently asked questions
How many questions do you need for a useful sample?
Fewer than you'd think, as long as they cover the right situations: terms that vary by market, ambiguous questions, out-of-coverage requests, sensitive cases. A small sample run the same way after every change tells you more than a large one used only once.
Can you use artificial intelligence to evaluate artificial intelligence's answers?
As a first filter, yes, to flag incomplete or off-glossary answers at scale. As the final check, no: correctness against a market's terms can only be established by someone who knows those terms and is accountable for them in front of a customer.
Is it better to have one translated content base or a separate one per market?
It depends on how much the terms differ. If only the wording changes, a translated base plus a glossary is enough. If warranties, lead times, price lists, or requirements change, you need separate content per market: lumping different terms together in a single text is the fastest way to generate wrong answers.
We define the checks for your multilingual support.
If you’d like to talk it through, the service that handles this is Artificial intelligence.

