Sanjay C./Contribution Infrastructure
Note · International Translation Day · September 2026

Being understood includes the software

The theme this year

The International Federation of Translators has chosen “Linguistic Diversity and Language Rights: The Power of Being Understood” as this year’s theme for International Translation Day, on 30 September. Most of the day’s conversation will rightly be about interpreters, terminologists and literary translators.

I want to add a smaller point from a different corner. When someone reads a menu, a setting or an error message in their own language, a translator made that possible — and then the translation had to get into the software. In open source, that second step is where the care often runs out.

What the research says

A study presented at ICSE 2026, the main software engineering research conference, analysed 9.14 billion GitHub issues, pull requests and discussions across 62,500 repositories from 2015 to 2025. It found that open source is steadily becoming more multilingual — especially in Korean, Chinese and Russian — but that non-English projects receive less visibility and participation. In the authors’ words, such projects “may struggle to attract attention even when active and well-maintained” (Bhuiyan, Bala Kumar and Staicu).

The study measured 30 natural languages, eight of them Indian — Hindi, Bengali, Tamil, Telugu, Marathi, Gujarati, Punjabi and Urdu. In its published data they barely register: by our count, together about 0.03% of the non-English messages it recorded. Hindi appears roughly 1,100 times across eleven years; Korean, about 11 million. That is part of why our own evidence starts with Hindi: how contributions in Indian languages fare upstream is still largely undocumented. The problem itself is not specific to any language — it appears wherever maintainers are asked to approve words they cannot read.

One pull request

In April I opened a pull request on Open WebUI, a widely used open-source interface for AI models, correcting ten Hindi strings. Some were plainly wrong: the “Light” theme label had been translated with a word meaning “listen”. A maintainer merged it in under four hours. There were no review comments; the only human word on the thread was “Thanks!” (PR #23745).

I am glad it merged, and users are better off. But the speed cuts both ways. No one on the project reviewed the Hindi, so the same door would have let a poor translation through just as fast. The outcome depended on the contributor happening to be careful.

Maintainers see this too. Closing a Hindi pull request on another project in June — after the files it targeted had been removed while it waited — a Kilo Code maintainer wrote that language review “does depend on contributor/maintainer bandwidth right now” (comment). That is not a criticism of anyone. Most maintainers are volunteers doing their best with what they can read. It is a gap in shared infrastructure.

What we have built so far

Code contributions get linters, automated tests and a second reviewer. Translations mostly get trust. We have been documenting that difference since earlier this year — first in Communications of the ACM, then in a DevOps.com field report that followed five localization pull requests through upstream review.

There is also a security side that translators are rarely told about. A translated string can carry an invisible character that makes displayed text differ from what is stored, a fragment of script that runs when shown on a web page, or a broken placeholder that crashes the program. We have released a small open-source checker, i18n-security-lint, that looks for these problems in translation files for any language. It is early work, and reports of what it gets wrong are the most useful thing anyone can send.

The same gap shows up in accessibility. We scored seven accessibility pull requests from well-known projects against a review rubric; none reached its top band (write-up in Bootcamp).

An invitation

If you translate, review or maintain software in any language, we would like to hear how translation review works on your project — and where it does not. Evidence from languages other than Hindi is especially welcome. The research, the tools and the open tasks are public in the oss-language-inclusion repository and on the initiative’s contribute page, and corrections are always welcome.

Correction, 26 September 2026: An earlier version of this note said Indian languages were not among the languages the ICSE 2026 study measured. They were; the paragraph above now gives their share of the study’s data. We also replaced “Nobody on the project could check the Hindi” with “No one on the project reviewed the Hindi”, which is what the record shows: the contributor checked every string, and the project merged the change with no reviews.

Disclosure: For transparency, I used Claude to help structure this piece. The final content, observations, and conclusions are my own.

Sanjay C. builds evidence and small tools for open source localization and language inclusion — including how translation contributions are reviewed upstream. He writes and researches independently from Bengaluru; more at ecogetaway.github.io.