Abstract illustration of text in many scripts flowing across Africa into a small compact model

Olaverse Releases 608-Language ID Models Built African-First — the Smallest Is 37 MB

CurratedBrief Editorial Team
13 Min Read
Disclosure: This website may contain affiliate links, which means I may earn a commission if you click on the link and make a purchase. I only recommend products or services that I will personally use and believe will add value to my readers. Your support is appreciated!
- Advertisement -

🎙️ Listen to this post: Olaverse Releases 608-Language ID Models Built African-First — the Smallest Is 37 MB

0:00 / --:--
Ready to play
Abstract illustration of text in many scripts flowing across Africa into a small compact model

Last updated: 7 October 2026. Benchmark figures below come from Olaverse Lab’s own announcement and Hugging Face model cards and have not been independently reproduced. Disclosure: Olaverse Lab and CurratedBrief share an owner.

The 60-second version

  • On 6 October, Olaverse Lab released two open-weight language identification (LID) models covering 608 languages, built “African-first”: lid-lite-608, a 37 MB fastText model, and lid-neural-608, a 140M-parameter ModernBERT classifier fine-tuned from mmBERT-small.
  • The small model is roughly 1/45th the size of GlotLID v3 (1,687 MB) and about 1/32nd the size of Meta’s fastText LID-218 (1,176 MB), and runs at around 6,800 texts per second on a single CPU thread.
  • The headline claim is accuracy on short, messy text: on Olaverse’s own held-out test, lid-lite-608 scores 0.672 on 1–5 word inputs against 0.486 for GlotLID, and averages 0.839 F1 across 20 major African languages against GlotLID’s 0.757.
  • Both models are Apache-2.0 and trained only on commercially usable data — a practical difference from Meta’s LID-218, which is licensed CC-BY-NC-4.0.
  • GlotLID still wins on FLORES+, the long, clean, professionally translated benchmark. The gains are concentrated on the short, conversational input real products actually receive.

Key numbers

lid-neural-608 lid-lite-608 GlotLID v3 Meta LID-218
Size 140M params (~560 MB) 37 MB 1,687 MB 1,176 MB
Licence Apache-2.0 Apache-2.0 Apache-2.0 CC-BY-NC-4.0
Held-out test accuracy (240k samples, 608 languages) 0.867 0.858 0.783 0.280*
Short input (1–5 words), web-weighted, traffic mode 0.906 0.829 0.808 0.764
Everyday phrases (“good morning”, “sign in”) 0.943 0.838 0.857 0.867
Tatoeba (conversational, 34k) 0.930 0.922 0.868 0.557
FLORES+ devtest (202 languages) 0.942 0.934 0.957 0.874

*Meta’s model covers only 218 languages, so it is penalised on a 608-language test. Restricted to the 208 languages it covers, the lid-lite-608 model card reports 0.822 for Meta against 0.900 for lid-lite-608 and 0.885 for GlotLID. Source for all figures: Olaverse Lab announcement and model cards.

Why language identification still isn’t solved

Language identification is the unglamorous first step of nearly every multilingual AI pipeline. Before a system can translate a message, filter a web crawl for training data, route a support ticket or pick a speech model, it has to know what language the text is in. Open tools such as GlotLID and Meta’s fastText LID already cover hundreds of languages and perform well on long, clean paragraphs, which is why the problem is often treated as finished.

Real input rarely looks like that. It is a two-word greeting, a WhatsApp message typed without tone marks, a product review with an emoji, a URL. Olaverse’s announcement argues that this is precisely where African languages fall through the cracks: Kinyarwanda, Xhosa, Wolof and Lingala are routinely confused with neighbouring languages, and Nigerian Pidgin is frequently misread as English. When that first step fails, everything downstream inherits the error — the wrong translation model is called, low-resource text is dropped from training corpora, or a user is answered in the wrong language.

- Advertisement -

What Olaverse released

The two models share the same 608 language labels (567 distinct ISO 639-3 codes across 36 scripts, labelled as language plus script, e.g. yor_Latn or srp_Cyrl) and the same two operating modes. They differ in where they run:

  • lid-lite-608 is a quantised fastText classifier using character n-grams and word bigrams. At 37 MB and ~6,800 texts per second on one CPU thread, it is designed for bulk filtering and edge or CPU-only deployment.
  • lid-neural-608 is a 140M-parameter transformer that, per its card, processes around 1,100 texts per second on an A100 GPU (roughly 40–60 per second on an Apple M4 laptop). It is the more accurate of the two on short and conversational text.

Olaverse suggests running both: the lite model as a fast first pass over everything, with only short or low-confidence inputs escalated to the neural model. According to the model card, lid-lite-608 was trained from scratch on 7.86 million samples drawn from FineWeb-2, MADLAD-400, the Aya Dataset, WURA, MasakhaNews and Tatoeba — all sources whose licences permit commercial use. Religious text, which dominates many low-resource corpora and skews vocabulary, was deliberately filtered out, along with wrong-script and mislabelled documents.

The African-language results

Across the 20 major African languages Olaverse tracks, lid-lite-608 averages 0.839 F1 against GlotLID’s 0.757. The largest gaps are in exactly the languages the announcement flags as commonly confused:

Language lid-lite-608 F1 GlotLID v3 F1
Kinyarwanda 0.758 0.371
Xhosa 0.840 0.602
Wolof 0.809 0.600
Lingala 0.814 0.642
Yoruba 0.906 0.827
Swahili 0.808 0.719

It is not a clean sweep. The model card shows GlotLID ahead on Kirundi (0.724 vs 0.687) and Zulu (0.761 vs 0.743) — both close neighbours of languages lid-lite-608 improves on, which suggests the hardest confusions within language clusters are improved rather than eliminated.

The most transferable idea: one model, two ways to read it

The more interesting design choice is not model size but a second decoding mode. A classifier trained on balanced data treats all 608 languages as equally likely. That is right when mining a crawl for low-resource text, where every short Lingala sentence matters. It is wrong for a chat box, where “Bonjour mon ami” is overwhelmingly likely to be French rather than one of the smaller languages sharing those words.

- Advertisement -

So both models ship with a coverage mode (every language equally likely, the default) and a traffic mode, which shifts each language’s score by how common it is in real web text — formally, adding 0.2 × log(real-world share ÷ training share) to each language’s log-probability. In Olaverse’s own example, coverage mode labels “Bonjour mon ami” as a minority language (dhv_Latn) while traffic mode returns French. The company reports traffic mode adds 6.6 points of accuracy on 1–5 word inputs for the lite model and 5.5 for the neural model, using the same weights plus one extra vector. The technique is essentially prior correction, and it is applicable well beyond this release: any classifier trained on deliberately rebalanced data can be recalibrated to the deployment distribution this way.

The “Good morning” bug, and why it matters beyond this model

The announcement’s most candid section describes an early build that scored well on every benchmark — then failed to recognise “Good morning” as English. Precision for English on short phrases was 0.19. The cause was in the data: web pages in small languages are littered with English boilerplate (“Read more”, “Sign in”, cookie banners), and each fragment had been labelled with the page’s language, teaching the model that short English phrases were evidence for hundreds of other languages.

The fix combined a foreign-text filter, rebalancing of high-traffic languages, about 1.6 million short Tatoeba sentences, and a 107-phrase “probe test” of everyday phrases in 11 common languages that is excluded from training and checked on every build. Olaverse’s conclusion — that aggregate benchmarks can hide a failure every user hits in their first minute — is a lesson that applies to far larger AI systems than a language identifier.

- Advertisement -

What to be cautious about

Several caveats apply before treating these numbers as settled. First, the headline held-out test was built by Olaverse, from websites not seen in training but sampled by the same pipeline that produced the training data; models trained elsewhere are naturally at a disadvantage on it. The fairer comparisons are the public benchmarks — Tatoeba, UDHR-LID and FLORES+ — where the margins over GlotLID are smaller (0.922 vs 0.868 on Tatoeba, 0.910 vs 0.908 on UDHR-LID) and where GlotLID leads on FLORES+. Second, Meta’s 0.280 held-out score largely reflects language coverage rather than quality, as the restricted 208-language comparison shows. Third, no third party has yet reproduced these results. And at the time of writing the models have just been released, with no public track record in production.

For builders and publishers

  • Check your licence before your accuracy. If you use Meta’s LID-218 in a commercial product, its CC-BY-NC-4.0 licence is a constraint; both Olaverse models and GlotLID are Apache-2.0.
  • Match the mode to the job. Use coverage mode for corpus building and low-resource data mining; use traffic mode for chat, search and request routing.
  • Test on your own short inputs. If your traffic includes greetings, UI strings or messages without diacritics, build a small probe set of those cases and score any LID model against it before switching.
  • Filter junk explicitly. Both models include a noise class (zxx_Zxxx) for numbers, URLs and code; lid-lite-608 reports 0.961 F1 on it against 0.620 for GlotLID, which matters for anyone cleaning scraped data.

Both models are installable via the olaverse Python package (pip install "olaverse[lid]" for the lite model) and also work with plain fasttext and transformers; usage code is on the model cards.

What we still don’t know

  • How the models perform on independent, third-party evaluations, particularly African-language test sets not built by Olaverse.
  • How well traffic mode’s web-derived language frequencies match specific deployments — a Nigerian fintech app and a European e-commerce site have very different language mixes.
  • How the models handle code-switched text, such as messages mixing Yoruba and English, which is common in the markets the release targets; the announcement does not report a code-switching benchmark.
  • Whether the African-language gains hold up for the hundreds of languages outside the 20 tracked in the published F1 table.

FAQ

What is language identification?

Language identification (LID) is the task of automatically detecting which language a piece of text is written in. It is typically the first step in translation, content moderation, search, speech and AI training-data pipelines.

Are the Olaverse LID models free to use commercially?

Yes. Both lid-lite-608 and lid-neural-608 are released under Apache-2.0, and Olaverse says they were trained only on data whose licences allow commercial use.

Which model should I use?

lid-lite-608 suits high-volume, CPU-only or edge use. lid-neural-608 is more accurate on short user input but needs considerably more compute. Olaverse recommends using the lite model first and escalating only short or uncertain inputs to the neural model.

Is it better than GlotLID?

On short and conversational text and on most of the African languages measured, Olaverse’s reported numbers favour its models. On long, clean text such as the FLORES+ benchmark, GlotLID scores higher. All comparisons are self-reported by Olaverse.

Sources

Please follow and like us:
Pin Share
- Advertisement -
Share This Article
Leave a Comment