top of page

Apparently, Polish Speaks AI Better Than I Speak English


Audio cover
Polish Speaks AI Better

I am never going to hear the end of this.


I already spend a significant portion of my life talking to machines in a version of English that can best be described as grammatically adventurous. Half sentences. Tangents. Missing nouns. Three ideas welded together with "so anyway."


And now I find out Polish may actually be better at talking to AI than English.


Wonderful.


I can already hear it.


Speaks Polish, Rich... Better than you speak English, Rich... Seven grammatical cases, Rich... Your language has vibes, Rich.


Fine.


The annoying part is that there's actually something really interesting underneath the joke.


Researchers created a benchmark called ONERULER to test how large language models handle very long contexts across 26 languages. At the longer context lengths, Polish came out on top. English, despite being overwhelmingly represented in AI training data, finished sixth.


On the long-context retrieval tasks, Polish averaged roughly 88% accuracy. English was around 83.9%. Chinese, which I absolutely would have guessed might have an advantage because of how information-dense the written language can be, landed down around 62%.


Before everybody in Warsaw starts printing "Poland Invented Artificial Intelligence" t-shirts, there's an important caveat.


The researchers didn't conclude that Polish is "the best language for AI."


In fact, one of the study's coauthors later explicitly pushed back on people making that claim. The benchmark measures a specific thing: how models retrieve and aggregate information across very long contexts. Polish happened to perform exceptionally well. The researchers haven't established exactly why.


Which, naturally, is the part I find interesting.


Because maybe there's something hiding here that we don't have particularly good language for yet.


So I'm going to make some up.


Intent texture.


Polish is a heavily inflected language.


The endings of words can encode things English frequently communicates through word order or additional words: case, gender, number, person and other grammatical relationships.


In extremely oversimplified terms, English often says: here are the words, their position tells you how they relate.


Polish can carry more of those relationships inside the words themselves.


That doesn't necessarily mean Polish carries "more information." And it certainly doesn't prove this is why the models performed better.


But I wonder whether it gives the language something like a richer relational texture.


Not just: what are the concepts?


But also: how do these concepts belong together?


That distinction matters enormously for AI.


Because the hard problem with increasingly large context windows isn't merely stuffing more information into them.


It's preserving the relationships between all that information.


A model with 128,000 tokens of context doesn't simply have a giant notebook sitting beside it. Attention has to maintain structure across that context: this person relates to that event, this instruction modifies that object, this statement contradicts something 70,000 tokens earlier, this pronoun belongs to that noun, this clause changes the meaning of another clause.


And something strange happens as the context gets larger.


The ONERULER researchers found that the performance differences between languages grew substantially as context length increased. The gap between stronger and weaker-performing languages was about 11 percentage points at 8K context and grew to roughly 34 points at 128K.


That's where my ears perk up.


Because now this isn't merely: some models know Polish surprisingly well.


It's potentially: the structure of a language changes how effectively information survives inside a model's attention architecture.


That's a much bigger idea.


Language May Not Just Be the Input


We tend to think of language as packaging.


I have an idea.


I express it in English.


Someone else expresses essentially the same idea in Polish or Mandarin or Spanish.


The AI unwraps the packaging and gets the underlying idea.


Conceptually: thought, then language, then tokens, then model.


And once the thought reaches the model, we assume the wrapper doesn't matter much.


Except apparently the wrapper may matter quite a bit.


The ONERULER team also tested cross-language situations where instructions were written in one language while the underlying context appeared in another.


Performance could move by as much as 20% depending on the language used for the instructions.


Same machine.


Same general task.


Different linguistic representation.


Different result.


That suggests language isn't merely transporting intent into the model.


It may be shaping the geometry of that intent once it's inside.


Now we're somewhere interesting.


Which Brings Me Back to Intent


I've been increasingly convinced that the future interface to intelligence isn't really going to be an interface in the traditional sense.


It's going to be an intent fabric.


Voice contributes something.


Gesture contributes something.


Location contributes something.


History contributes something.


The device you're interacting with contributes something.


Your previous decisions contribute something.


Maybe your physiological state eventually contributes something.


None of these individually represents the complete command.


Together they create a high-resolution representation of what you mean right now.


And suddenly I'm wondering whether human language itself already does different versions of this.


Perhaps some languages are relatively sparse carriers of explicit relational structure.


Others have more of that structure woven directly through them.


Not necessarily more meaning.


More texture around the meaning.


Intent texture.


And if that turns out to matter to transformers, it raises an absolutely ridiculous question.


What Happens When We Stop Designing Languages for Humans?


We invented programming languages because human languages are terrible at telling computers precisely what to do.


But large language models aren't traditional computers.


They don't require Python.


They don't require C++.


They can operate inside probabilistic semantic structures.


So what if someday we develop a language specifically optimized for machine cognition?


Not a language humans use to talk to each other.


A language humans and agents use when extremely precise, persistent contextual relationships matter.


Maybe it has unusually explicit morphology.


Maybe relationships are encoded redundantly.


Maybe ambiguity is intentionally constrained.


Maybe its vocabulary is designed around tokenizer behavior.


Maybe it contains mechanisms specifically designed to preserve references across million-token contexts.


Essentially: what would language look like if we designed it around attention instead of human vocal cords?


I have absolutely no idea.


Which is generally how I know I've found something worth thinking about.


For now, all we actually know is much narrower.


A multilingual long-context benchmark produced a genuinely surprising result: Polish performed exceptionally well, English wasn't number one, and the language used to represent information can meaningfully change how an LLM performs.


We don't yet know why.


Maybe morphology matters.


Maybe tokenization matters.


Maybe training methodology matters.


Maybe some interaction between all three matters.


But I suspect we're eventually going to discover that languages don't simply contain different words for the same thoughts.


They provide different structures for holding thought together.


And apparently the machines notice.


So congratulations, Poland.


Enjoy this.


Meanwhile I'll continue interacting with several billion dollars worth of artificial intelligence using sentences like: "Yeah no not that thing, the other thing we were talking about yesterday with the little sensor and the agent thing, you know what I mean."


And somehow it still understands me.


Mostly.



Sources

The ONERULER paper: the 26-language long-context benchmark, Polish's top ranking, English finishing sixth, and the cross-lingual instruction results.


Accuracy breakdowns by language (Polish ~88%, English ~83.9%, Chinese ~62%) and the widening performance gap at longer context lengths.


The study coauthor's pushback on the 'Polish is the best language for AI' framing, and what the benchmark does and doesn't establish.



Rich Washburn is a technologist, strategist, and Founder & Chief AI Architect of ARIA AI Labs, working at the intersection of AI, infrastructure, communications, and capital. He also serves as Managing Partner and Chief AI Officer at Eliakim Capital.

Comments


Animated coffee.gif
cup2 trans.fw.png

© 2018 Rich Washburn

bottom of page