Worked about as well as expected :D
A few months ago I went down this same rabbit hole trying to determine if using Chinese has any inherent advantages due to the way more context is encoded into Chinese characters versus what you find in English.
The total token size may not take into account any benefits gained by Chinese characters potentially encoding more context.
I have been meaning to see if any Chinese native tokenizers have been created to account for this feature of the language.
Unfortunately, many Chinese and Japanese characters can have many different meanings depending on how they are used. As a result they do not break down as well to be tokenized as a language like English.
Maybe some kind of hybrid intermediate language would be most optimal.
The only thing there that seems like it's relevant is the finding that the models they used had worse results when prompted in Chinese than when prompted in English. But ...
I am fairly sure I would do substantially worse work if I had to do it in French rather than in English, even if the work itself was all mathematics and programming and the like. Maybe that in some sense indicates that I am a stochastic parrot but it clearly doesn't indicate that I'm a stochastic parrot in some way worse than human beings are since I happen to be a human being myself.
... so what am I missing, that makes the models' worse performance in Chinese an indication that they're stochastic parrots in any interesting sense?
(I take it that "they're stochastic parrots just like we are" is not a very interesting sense. I mean, actually it would be quite interesting to understand better how much of human thinking can be reasonably described as stochastic-parroting -- fairly clearly the amount isn't zero -- but I think it would be interesting as a finding in human psychology, not as a fact about LLMs.)
Even if we didn't know french, it is straightforward to use deterministic tools (I.e. a dictionary) to parse the instruction, writing out the code in English, then refactor as many words as possible to french. The paper did not give any indication that this was happening (would be a novel result indeed) and so I would have to conclude that it is running like a stochastic parrot, more English codebase, better performance in English only.
I think this might still be in the realm of "maybe about as stochastic-parrot-y as human beings are"; I can easily imagine doing a worse job of remembering relevant stuff that happened to be in English if I had to do my work in some other language. Human memory is surprisingly context-dependent. But it would be interesting to see what happens if you ask an LLM to write code while talking to it in some language that has waaaaay less programming-related stuff on the internet. Swahili, perhaps. Is it much worse than when you prompt it in English, or a little worse, or what?
If the LLMs are very stochastic-parrot-y -- just piecing together bits of code associated with the words you wrote in English-or-Chinese-or-Swahili -- then I would expect them to get catastrophically worse when prompted in a language in which there's very little programming content on the internet. (For what it's worth, I think I also think this degree of stochastic-parrot-ness seems rather incompatible with what they are able to do. But others may disagree.) On the other hand, if they're more like humans -- somewhat better at remembering relevant things when they're in the same language as they're working in, etc., but operating at a conceptual level as well as pushing words around -- then I would expect the loss to be much more moderate.
(It seems somewhat relevant that the insides of a transformer network operate on embedding vectors rather than literal tokens; presumably those embedding vectors are much less language-specific.)
As for tokenizers - I reckon that most of them are not optimized for CJK; they're good enough, and AI performance in CJK languages, so far, haven't been the top priority matter.
Here's the token efficiency of a corpus translated into various languages and tokenized with the latest OpenAI one:
Source: "Tokenizer Fairness in 2026", a reproduction/extension of Petrov, La Malfa, Torr & Bibi, "Language Model Tokenizers Introduce Unfairness Between Languages" (NeurIPS 2023), using FLORES-200.https://github.com/partyfly/tokenizer-fairness-2026