The Press Releases Promise African-Language AI. The Data Pipeline Is the Real Product.
Tether's 800M-parameter model beat a 122B giant at African translation by deleting 96% of its training data. Google's WAXAL spent three years recording 11,000 hours of speech with African universities. The pattern is now unmistakable: in African-language AI, whoever builds the data pipeline wins. The announcements are still louder than the datasets.

On 2 September 2026, Tether AI Research released a family of open-source translation models for African languages. The smallest, TranslatePsy-AfriSLM, has 800 million parameters and runs offline on a phone. According to the accompanying paper, accepted at EMNLP 2026, it outperformed Alibaba's Qwen3.5-122B-A10B, Google's TranslateGemma-27B, and Meta's NLLB-3.3B across three translation benchmarks. A model roughly 150 times smaller won.
The headline is the size gap. The story is the method. Tether's team used a quality-estimation filtering technique that removed up to 96% of the low-quality open-source training data they started with. They won by throwing data away, carefully. That is the whole thesis of this piece, and I will keep returning to it: in African-language AI, the data pipeline is the product. The models are downstream of it.
The 800M model that beat the 122B giant
An opinion without numbers is just a mood, so here are the numbers. The researchers evaluated on three benchmarks: FLORES-200, the field's standard translation test; BOUQuET, a newer linguist-curated benchmark; and SMOL, built from professionally translated sentences. Scores use SSA-COMET, a quality metric built for Sub-Saharan African languages, where higher is better on a 0 to 1 scale.
The 800M-parameter AfriSLM scored 0.5944 on FLORES-200, 0.6223 on BOUQuET, and 0.4973 on SMOL. Qwen3.5-122B-A10B, with 122 billion total parameters, scored 0.5505, 0.5716, and 0.4574. Roughly 8 to 9 percent better, at roughly one one-hundred-fiftieth the size. The 4B version of AfriSLM extended the lead further. The weights are public, so anyone can reproduce the benchmarks, and the paper passed EMNLP peer review, which tempers the usual skepticism about vendor-reported numbers without removing it.
The model covers 19 languages: Hausa, Amharic, Yoruba, Lingala, Swahili, Igbo, Zulu, Somali, Oromo, Malagasy, Kinyarwanda, Xhosa, Afrikaans, Wolof, Luganda, Nyanja, Shona, Tswana, and Southern Sotho. Tether notes these represent roughly half of Africa's population. Read that again: 19 languages, half the population. The other half speaks the other two thousand.
Eleven thousand hours, three years, one dataset
Now the opposite direction. In February 2026, Google officially launched WAXAL, a speech dataset for African languages, after three years of development. The numbers: more than 11,000 hours of recorded speech, nearly 2 million individual recordings, covering 21 Sub-Saharan African languages including Hausa, Yoruba, Luganda, and Acholi. On top of that, more than 20 hours of high-quality studio recordings for building natural-sounding synthetic voices.
What makes WAXAL interesting is not the size. It is the ownership structure. Universities such as Makerere in Uganda and the University of Ghana led the collection. The local partners retain ownership of the datasets, released open source under licenses that allow commercial use. Within days of release, the dataset passed 4,000 downloads.
Three years. Eleven thousand hours. Twenty-one languages. That is the actual speed of building speech data properly: slow, institutional, and unglamorous. Nobody holds a press conference for the second year of recording sessions. The press release arrives at the end, and it reads as if the dataset simply appeared.
The coverage math nobody puts in the announcement
Africa has more than 2,000 languages. FLORES-200, the most-used translation benchmark in the field, covers 52 African languages. That is roughly 2.6 percent. Nearly every announcement that says "now supporting African languages" is talking about a thin slice of the top of the distribution: the same dozen or so languages with existing web text, existing benchmarks, and existing commercial incentive.
And even the benchmark data is shakier than it looks. A 2024 paper had native speakers review the FLORES evaluation sets for four African languages, Hausa, Northern Sotho, Xitsonga, and isiZulu, and found inconsistencies and inaccuracies serious enough to require corrections across the board. The authors recommend involving native speakers at every stage of future translation efforts. If the gold-standard evaluation set needed fixing by the people who actually speak the languages, imagine the state of the training data nobody checks.
Then there is the tokenization tax. A 2026 study measured how frontier tokenizers handle 20 African languages and found every one carries a cost premium over English: from 1.29x for Swahili up to 8.92x for N'Ko. In money terms: a deployment handling a million Amharic queries a month would cost about $1.35 million a year in inference, versus $183,000 for the equivalent English deployment. A 7.4x structural surcharge lands on the invoice before model quality or product-market fit enters the picture. Coverage on a benchmark does not mean affordable deployment.
The volunteers were right all along
None of this is news to the people who have been doing the work. Masakhane, the grassroots African NLP community, has spent years building machine-translation benchmarks for more than 30 African languages through participatory research: volunteers collecting and verifying data because they want their languages to survive online. Their work showed that multilingual models could be adapted to new African languages with as few as 2,000 sentences of good data. A little clean data goes a long way. Tether just proved the same point from the other side: a lot of dirty data goes nowhere.
Lelapa AI's InkubaLM made the same bet earlier: a 0.4-billion-parameter model trained from scratch on 1.9 billion tokens across five African languages (isiZulu, Yoruba, Swahili, isiXhosa, Hausa), released openly. Small, focused, built on collected data rather than scraped data.
I write this as a practitioner, not an observer. My foundation's SaloneLearn effort is collecting teacher-verified question-and-answer pairs in Krio, Mende, and Temne, written by students in Sierra Leone and verified by their teachers, to be released open and free. The stated goal is 500 pairs, and collection is underway. I can tell you from the inside what the announcements leave out: the dataset is the project. The model, when we get to it, will be the easy part. Everything slow and expensive sits upstream: finding the schools, designing the questionnaires, typing up paper forms, checking every pair with a teacher who actually speaks the language. There is no API for that.
The honest counterweight
An opinion has to survive its own objections, so here are mine.
First, Tether's numbers are vendor-reported, and 19 languages is still the head of the distribution. The next 19 will be harder, and the 19 after that harder still. Nobody has a plan for the long tail that does not involve the same slow community collection WAXAL and Masakhane do.
Second, benchmarks are translation, and translation is not the product. The gap that matters most is voice: hundreds of millions of Africans interact with technology by speaking, not typing. NKENNEAi's September launch of Swahili speech-to-text and text-to-speech is aimed at exactly this, and it is one language. The farmer who needs a voice advisory in Temne is not helped by a COMET score.
Third, the announcements keep coming faster than the deployments. The UNDP and GSMA partnership announced on 2 October promises African-led AI with African languages as a priority, and the release names no budget, no timeline, and no pilot criteria. I covered that announcement here and said the same thing: watch what gets built, not what gets announced. A promise is not a pipeline.
What would make me wrong
I want this piece to be falsifiable, because an opinion you cannot disprove is just branding. Three things would prove me wrong:
- A frontier lab ships a general model with genuinely good support for long-tail African languages, Krio, Mende, Temne, N'Ko, Acholi, with no community-collected data behind it. If scale alone solves it, the data pipeline was never the product.
- The press-release models show up in real deployments at national scale within two years: clinics, classrooms, agricultural extension, in local languages, used daily. Benchmarks becoming products would mean the pipeline I am describing is already built.
- The tokenization tax disappears through architecture alone, with no data work. If vocabularies fix themselves, one of my two scarcities was imaginary.
I do not expect any of the three. I would be glad to be wrong about all of them.
Sources
- Tether AI Research announcement, 2 September 2026
- The Profiler: Tether's open offline translation models, EMNLP 2026 acceptance
- NaijaTechGuide: benchmark table and SSA-COMET scores, arXiv 2608.18655
- TechCabal: Google's WAXAL speech dataset, February 2026
- TechCabal: NKENNEAi Swahili speech models, September 2026
- InkubaLM: A small language model for low-resource African languages, arXiv 2408.17024 and the Hugging Face model card
- Correcting FLORES Evaluation Dataset for Four African Languages, arXiv 2409.00626
- The African Language Tax, arXiv 2606.24460
- Science: "AI often mangles African languages. Local scientists and volunteers are taking it back to school"
- Participatory Research for Low-Resourced Machine Translation, arXiv 2010.02353
Related on Everyday Data Science: UNDP and GSMA Bet on African-Language AI. The Announcement Is Not the Plan. · Africa's 12,000 GPUs and the Compute Gap
If you had to build one dataset for your own language, what would it be, and who would verify it?
About the writer
Data Scientist & AI Researcher
Data scientist and AI researcher at Pace University. I coined Artificial Frictional Unemployment, and built the first machine learning model for crop yield prediction in Sierra Leone. Author of Understanding Agentic AI. I write about agentic systems and applied ML, with a bias toward what actually works, and who gets left out when it doesn't.
Found this useful? Passing it on to someone who builds is the best way to help the publication grow.
Built something worth sharing? Write it up for us →