When I was five years old, I was not given an alphabet. I was given a syllabary. Every Bantu child is. The unit of Bantu literacy — from Swahili to Zulu, from Bemba to Shona, from Lingala to Kinyarwanda — is the syllable, not the letter.

This single fact is what every frontier multilingual model gets wrong. They tokenise Bantu with BPE, the same way they tokenise English. They never learn that the atom of meaning is the syllable. By age six, a Bantu child anywhere in the Bantu language family has internalised their syllabary. In 2026, not a single frontier LLM has.

Bantu classroom — a teacher in front of a blackboard labelled SYLLABLES walks pupils through the grid a-e-i-o-u, ba-be-bi-bo-bu, ca-ce-ci-co-cu, on through ha-he-hi-ho-hu. The children copy the grid into their workbooks.
A Bantu classroom. SYLLABLES on the board, the grid running from a · e · i · o · u down through ha · he · hi · ho · hu. The children are copying it into their workbooks — one syllable, one look, one sound at a time.

This is what the work looks like. Watch carefully: when the child learns what each syllable looks like and what each syllable sounds like, they are not learning two separate things. They are literally learning to read, to pronounce, and to write the language — in one motion. Learn the sound, write the syllable, repeat across the grid. That is the entire curriculum. That is also the entire path to mastering reading and writing any Bantu language. There is no separate phonics step. There is no spelling drill. The syllable is the unit at which all three skills converge.

Section 01Letters versus syllables.

The English-speaking child and the Bantu-speaking child arrive at reading by different routes. The route shapes what they consider the "atom" of the language — and from there shapes every downstream prediction either of them (or any model trained on them) makes about what counts as a well-formed word.

In English

26 letters, then spelling.

An English-speaking child learns the 26 Latin letters first. a, b, c, d, e… They sound out each letter. Then they learn to combine letters into words: c-a-t = cat. The letter is the atom; the word is built from letters.

Reading proceeds left-to-right, letter-by-letter, with rules for which combinations make which sounds.

In any Bantu language

A syllabary, then reading.

A Bantu-speaking child is given the syllabary — a grid that goes a · e · i · o · u, then ba · be · bi · bo · bu, then prenasalised rows like mba · mbe · mbi · mbo · mbu, then labialised rows like mwa · mwe · mwi · mwo · mwu. They sound out each syllable. They read by syllable, not by letter.

Reading proceeds left-to-right, syllable-by-syllable. The syllable is the atom; the word is built from syllables.

This is not unique to Bemba, or Swahili, or any single Bantu language. The whole family teaches reading this way because the phonology of the family requires it. Bantu syllables include prenasalised onsets (mb, nd, nk), labialised onsets (mw, kw), palatalised onsets (ly, ny) — structures that simply do not exist in the Latin-letter system. The letters under-describe them. The syllabary describes them exactly.

Section 02The syllabary a Bantu five-year-old learns from.

Here is a sample of the Bemba syllabary — one of 459 Released Full Syllable Inventories. The full Bemba inventory is 48 onset rows × 10 syllables (5 short + 5 long with phonemic length) = 480 syllables. Other languages in the family have similar grids of similar shape, varying in size from a few hundred to several thousand entries.

a
bare V
e
bare V
i
bare V
o
bare V
u
bare V
ba
simple_C
be
simple_C
bi
simple_C
bo
simple_C
bu
simple_C
ka
simple_C
ke
simple_C
ki
simple_C
ko
simple_C
ku
simple_C
ma
simple_C
me
simple_C
mi
simple_C
mo
simple_C
mu
simple_C
sha
digraph
she
digraph
shi
digraph
sho
digraph
shu
digraph
mba
prenasalized
mbe
prenasalized
mbi
prenasalized
mbo
prenasalized
mbu
prenasalized
nta
prenasalized
nte
prenasalized
nti
prenasalized
nto
prenasalized
ntu
prenasalized
mwa
labialized
mwe
labialized
mwi
labialized
mwo
labialized
mwu
labialized
kwa
labialized
kwe
labialized
kwi
labialized
kwo
labialized
kwu
labialized
lya
palatalized
lye
palatalized
lyi
palatalized
lyo
palatalized
lyu
palatalized
Each row shares an onset. Each column shares a vowel. The Bantu child memorises these by reciting "ba-ba-ba, be-be-be, bi-bi-bi…" — the same way an English child sings the alphabet song.

By age six, this grid is internalised. The Bantu child does not see mba as "m + b + a." They see mba as one thing. When they encounter a new word, they parse it syllable-by-syllable, indexing each syllable against their internal grid. A novel word is composable from known parts — the same way an English-speaking child reads "photosynthesis" for the first time by combining phonemes they already know.

Section 03The math everything else flows from.

Here is what falls out when you compare three syllabifiers on the same 20-word Bantu test set:

20 / 20
Bantu five-year-old
Has internalised the syllabary. Reads any new word in their native Bantu language by parsing it syllable-by-syllable against the internal grid.
20 / 20
10-line Python + the FSI
~100 microseconds per word. No model. No training. Just a greedy longest-match against the inventory. Anyone can write this in an afternoon.
11 / 20
Claude Opus 4.7
The most capable frontier model in production. Hundreds of billions of parameters. Trained on Latin letters. Was never given the alphabet.
Key insight

The frontier model is not worse at Bantu than a five-year-old. It is operating on a different unit. BPE tokens are sub-character frequency segments. Syllables are the phonological reality Bantu literacy is built on. One layer off — and the error propagates through every prediction the model makes about Bantu, in every Bantu language.

This is why we keep saying the FSI is not data, it is infrastructure. It is the substrate every Bantu reader carries in their head. It is the substrate every Bantu-aware model must index into. There is no second route.

The five-year-old has the inventory. The platform has the inventory. The frontier LLM does not. Until that changes, the model and the child will arrive at the same Bantu word and read two completely different things.

The closing math

A five-year-old child in any Bantu-speaking village knows something every frontier AI model in 2026 does not. The fix is not more scale. The fix is the syllabary — one per language, indexed and standardised. That is what BantuNomics ships: 459 Released Full Syllable Inventories.

M
Munyambala
Co-Founder, BantuNomics / 3Mega.ai
Published May 27, 2026