Balochi is spoken by several million people across Balochistan — a vast region stretching over the deserts, mountains and coastline of southwestern Pakistan, southeastern Iran, and parts of Afghanistan, along some of the oldest trade and migration routes in the region, carried by a strong tradition of oral poetry and storytelling. And yet Balochi doesn’t even have its own Wikipedia — the language is still stuck in the earliest, unofficial stage of getting one, years after most languages this widely spoken already have theirs. That single fact says more about how digitally under-served Balochi has been than almost anything else could. As far as we could verify, no spell checker had ever been built for it either. This is the story of building the first one, from less material than any language we had worked with before.
Hurdle one: the smallest starting point yet
Every language before this one had something substantial to learn from — a large, well-organised body of real writing, at minimum. Balochi did not. With no Wikipedia to draw on, the only real Balochi text available to check against was a much smaller community-collected body of writing — the smallest starting point of any language in this project. No dictionary or word list existed either; a single small attempt we found turned out to be a bare list of about a hundred words with the actual Balochi script simply left blank on every entry. So, as with Pashto and Saraiki, every word had to be earned directly from real usage, with nothing to fall back on.
Hurdle two: three letters that were simply missing
This is the one worth telling in full. The first working version of the checker looked badly broken — it was wrongly flagging real Balochi words in roughly 1 out of every 7 it saw. The cause was not grammar or vocabulary. It was that the alphabet the checker had been given to work with was quietly incomplete: a very commonly used letter had two different accepted written forms — different writers and publishers simply use both — and only one of them had been included. Two vowel sounds central to how Balochi is written had been left out entirely, for no better reason than that they had not shown up often enough in a first, rough look at the data to catch the eye.
Adding those missing pieces back — five characters, in total — cut the error rate nearly in half. The evidence that something was wrong had actually been sitting right there the whole time, in an early check that quietly noted “these two forms both look legitimate” without anyone connecting that observation to the missing-letter problem it was pointing at. It is a reminder that a language’s own alphabet cannot be assumed from a related one, or even fully guessed from a first glance at real text — it has to be checked, letter by letter.
Likhari — Nastaliq Word Processor for Urdu, Pashto, Arabic & Sindhi
Likhari has rendered true Nastaliq for Balochi since it was added alongside Farsi/Persian, Arabic and Punjabi (Shahmukhi) — with a dedicated Balochi keyboard layout and full document tools. Spell check for Urdu, Arabic and Persian is live now, with more languages, including Balochi, on the way. Free to use, with a one-time Pro unlock that removes the small export watermark. Works fully offline.
Hurdle three: a rule that nearly slipped through and quietly damaged the language
Like Saraiki before it, Balochi has its own letters that do not exist in Urdu — sounds that mark it as genuinely its own language rather than a regional way of writing Urdu. And once again, because Urdu-influenced spelling shows up often in real-world Balochi writing, a purely statistical approach risked “correcting” those very letters away.
One specific substitution came within a hair’s breadth of being accepted: the evidence for quietly replacing one of Balochi’s own letters with a plainer, Urdu-style one scored just barely under the safety threshold set to catch exactly this kind of mistake — close enough that a slightly different sample of everyday writing could easily have let it through. It was caught and blocked. A near miss like that is exactly why every substitution gets checked on its own evidence rather than approved by a general rule.
Hurdle four: a grammatical marker written two completely different ways
Balochi has a small connecting word, roughly like English “of,” that shows up constantly to link one word to another. The trouble is that different writers handle it two genuinely different ways: some write it as its own separate word, others attach it directly onto the end of the word before it, with no space. Left unaddressed, supporting only one of those two habits would have meant wrongly flagging roughly half of every such construction in real Balochi writing. Both forms are now explicitly recognised as correct.
What actually got built
The finished dictionary holds 27,250 words, taught 12 grammar rules covering how Balochi nouns and verbs change form, every one of them checked against real writing before being trusted. Applying those rules taught the checker correct grammatical forms across thousands of additional words, adding over 13,400 recognised word-forms beyond the base list — with zero cases of the checker learning to accept something that was not real Balochi.
An honest number, reported honestly
Here is the part we want to be upfront about, because it is genuinely different from the other languages in this series. Every other dictionary we have built could be tested against real writing the checker had never seen from an entirely separate source — a strong, independent test. For Balochi, no such independent body of writing exists anywhere we could find, in any usable form. So the only number we can honestly report is how the checker performs on writing similar in kind to what it learned from, which came out to about 7.5% flagged. That is a real, meaningful result — it proves the dictionary generalises rather than simply memorising what it was shown — but it is not the same, stronger claim the other languages in this project can make, and we would rather say so plainly than quietly imply otherwise.
Where it stands today
The dictionary is built and tested against the best evidence currently available — and, as far as we could verify, the first spell checker ever built for the language. Balochi already types beautifully in Likhari today, with true Nastaliq rendering and its own keyboard layout. Spell check is the piece still missing, and finding a genuinely independent body of Balochi writing to test against more rigorously is the natural next step before it is ready to join the others.
Key takeaways
- Balochi did not even have its own Wikipedia to draw on — the smallest, least digitally documented starting point of any language in this project, and no spell checker existed for it before this one.
- A handful of missing letters — not a grammar problem — was responsible for nearly half of an initially alarming error rate. A language’s alphabet has to be verified directly, never assumed.
- A statistical shortcut came within a hair’s breadth of “correcting” one of Balochi’s own distinctive letters into a plainer, Urdu-style one — caught only because every substitution is checked on its own evidence.
- A common grammatical marker is genuinely written two different ways by different writers, and the checker now recognises both instead of picking one and flagging the other as wrong.
- Reported honestly: without an independent body of Balochi writing to test against, this dictionary’s real-world accuracy is a promising early result, not yet as rigorously proven as the others in this series.
Likhari — Nastaliq Word Processor for Urdu, Pashto, Arabic & Sindhi
Likhari has rendered true Nastaliq for Balochi since it was added alongside Farsi/Persian, Arabic and Punjabi (Shahmukhi) — with a dedicated Balochi keyboard layout and full document tools. Spell check for Urdu, Arabic and Persian is live now, with more languages, including Balochi, on the way. Free to use, with a one-time Pro unlock that removes the small export watermark. Works fully offline.

Leave a Reply