Step 1 of 3 · choose

Where do you want to start?

Pick the family you care about. You can add the others at any time — the letters and words you learn in one stay with you.

Hindi

Sanskrit → Hindi

Hindi descends from Sanskrit through the Prakrits, and the traffic between them never stopped: the formal vocabulary is Sanskrit largely unchanged (tatsama), and there is a second layer that drifted into its present shape (tadbhava). Sanskrit also hands you the Devanagari script and the general shape of the grammar.

20,000 sentences scored and ready · Sanskrit material is about 73% of Hindi running text once you count both layers

The map ahead →

Spanish

Latin → Spanish

Spanish is Latin after fifteen hundred years of drift. Roughly four fifths of the running text of ordinary Spanish is Latin material, some identical, some worn down (Latin causa to cosa, oculus to ojo). It is the shallowest root-to-descendant drop of the three.

20,000 sentences scored and ready · Latin material measured at 79% of running text in this corpus

The map ahead →

Japanese

Middle Chinese → Japanese

Japanese takes its writing system and most of its learned vocabulary from Chinese, in waves: Go-on from the fifth century, Kan-on from Tang Chang'an, Tō-on from the Song and Ming. The Chinese characters carry meaning straight across, and the readings are fossils of which wave a word arrived in. What does not cross is the grammar.

20,000 sentences scored and ready · Sino-Japanese material is 49% of dictionary words and 18% of running speech

The map ahead →

Latin

Old Latin → Latin

Latin is the root of this family, and this app can now feed it to you: the Vulgate, the Aeneid, the Gallic War, the Catilinarian orations. Latin is highly inflected, so a "word" here is a form rather than a dictionary entry — knowing amavit does not tick amare, and the page does not pretend otherwise. Reading the root is the cheapest way into the whole family: Spanish, Portuguese, Italian, French and Catalan all descend from it, and the vocabulary you meet here reappears in all of them.

4,042 sentences scored and ready · 4,056 sentences of real Latin: the Clementine Vulgate, Vergil, Caesar and Cicero (text from Latin Wikisource, CC BY-SA)

The map ahead →

What is not here yet

Three families have real corpora behind them today, listed above. There are a dozen more worth building (Slavic, Arabic and the Semitic family, Turkic, the Nguni and Bantu languages, Dravidian) and the honest status is that their corpora are catalogued but not measured, so the pages for them do not exist yet. Choosing a family you don't see is not possible today, and it is better to say that than to show you a dead button.