Your search speaks one language. Your company speaks four.
Marius Wehrle, Katarina Kostic ·

Page one, in any language

In Switzerland, a company doesn't get to pick the language of its own knowledge. The manual was written in German. The customer in Romandie asks in French. The supplier's data sheet arrived in English. The branch in Ticino writes back in Italian. All of it ends up in one library as pages, one for each product, process or customer, in whatever language it arrived. Nobody is going to sit and label which is which.

That's normal here. But the search your AI agents rely on is usually tuned for one language at a time.

What breaks when you cross a language #

In one language, search is good. Ask in the language the page was written in, and search has everything it needs: the same words, the same meaning.

Ask in another language, and the words no longer match. Classic search matches words, so it loses the page. AI search, the kind built on embeddings that compares meaning instead of words, copes much better. It brings back pages on the right topic, and often that's enough. When it isn't, the agent works from a page that's nearly right, and nothing tells it the right one was missing. Nothing tells you either.

The numbers #

78 questions about our own libraries, asked in English and in six more languages, always against the same correct pages. The figure is how often the right page came back in the first three results. The right-hand column includes the agent's hint, made with the library's glossary.

Asked in Classic search AI search Words, meaning and the hint
English 91 % 94 % 95 %
German 44 % 87 % 96 %
French 55 % 90 % 97 %
Italian 56 % 87 % 97 %
Serbian 33 % 83 % 97 %
Chinese 19 % 87 % 97 %
Japanese 14 % 85 % 96 %

Classic search finds the right page 91 times in 100 when you ask in English. Ask the same question in German and it's 44. In Japanese, 14. The words simply aren't on the page.

AI search holds up in every language, which is what it became famous for. It still misses about one page in seven once the language changes. Meaning gets you to the right neighbourhood, not always to the door.

Together they land at 95 to 97 in every row, whatever language the question came in. That's page one, in every language.

How it works #

Our search finds a page three ways at once. By its words, which catches the exact term, the part number, the name. By its meaning, which catches the question worded differently. And by its links: pages that point to each other vote for one another.

Languages sit underneath all of that. When a page is saved, the search recognises its language on its own, section by section if it has to, so nobody labels anything and a library can mix languages freely. Each page's words are then indexed the way its language needs, so Versicherungen still finds Versicherung. And every page also gets an embedding: a fingerprint of what it means rather than the words it uses. Modern embeddings are multilingual by nature, one of the breakthroughs of the last few years: the same idea in French and in German leaves a very similar fingerprint, so a French question lands close to the German page.

Then the agent does its part. The library tells it which languages the knowledge is written in, and the agent gives the search a hint: the question as the library would phrase it. So the words match again, not just the meaning. Where a field has its own vocabulary, in engineering, statistics or law, the library hands over its glossary, and the hint uses the field's own terms, not a dictionary's. A French engineer asks about bombé d'hélice. A dictionary makes that "propeller camber". The glossary makes it "lead crowning", the word the page actually uses.

And with that, a question in another language finds the right page as often as one asked in the page's own language.

What it costs #

About 70 milliseconds, quicker than a blink. No language model runs inside the search, so it costs no tokens.

The alternative is the default: let the agent find the page itself. Search, read, not quite, search again. Every round is paid model time, while a person waits. A fast, precise search in front of an expensive agent is the whole argument.

The fair objection: why not translate everything into one language once and be done? Because nobody controls which language the next page arrives in. The library is written by people and agents together, every day, each in the language they work in, so "once" never comes.

Page one, in whatever language page one is in. That's the job.

Marius Wehrle, PhD and Katarina Kostic

An AI wrote a version of this. It has been extensively disagreed with.

Share this post