Kinyarwanda language intelligence

Kinyarwanda,
understood from the inside.

RURIMI is building a linguistic foundation for Kinyarwanda — connecting lexicon, morphology, grammar, phonology, tone and speech into one coherent language system.

Research-driven Native-speaker informed Evidence-first

Illustrative pass of one Kinyarwanda form through RURIMI's layers
umwana
Illustrative

One form, read through each layer. Where the underlying and surface shapes of a morpheme differ, a boundary process applied. A recorded snapshot of engine output — not a live query, and not a claim of coverage beyond this example.

Linguistic structure Demo data

Explore how RURIMI represents linguistic structure.

Illustrative analysis using representative Kinyarwanda examples. Each form below was checked against the current engine, and what you see is its recorded output — including where the engine returns more than one reading, or cannot yet establish a value.

The case for infrastructure

Kinyarwanda deserves infrastructure built around its language.

Kinyarwanda is structurally rich. Reliable language technology requires more than translation and surface-level tokenization — it requires understanding morphology, grammatical structure, lexical variation, phonology, tone, and the relationships between them.

Morphology

Kinyarwanda words can encode substantial grammatical information inside a single surface form. Splitting on whitespace throws most of it away.

Lexicon

A useful language system needs structured lexical knowledge — senses, stems, classes, provenance — rather than a flat word list.

Grammar

Agreement, tense, aspect, mood, argument structure and verbal extensions must be represented explicitly, not inferred from surface patterns.

Speech

Pronunciation-oriented processing requires phonological and phonetic representations — tone and vowel length among them — rather than ordinary spelling alone.

Architecture

One language. One connected linguistic pipeline.

RURIMI treats these as connected layers rather than isolated utilities. Each layer answers to the one above it, and a decision made in the lexicon is still visible by the time a form reaches speech.

Lexicon

Structured lexical knowledge: lexemes, senses, stems, noun classes, and the provenance behind each.

Morphology

Segmentation and generation over roots, extensions and inflection — in both directions, so what is analysed can be rebuilt.

Grammar

Agreement, tense–aspect–mood and the conditions under which a derivation is licensed.

Boundary phonology

What happens where morphemes meet: glide formation, palatalization, vowel coalescence, Dahl's Law.

Surface

The written form, reconciled with the underlying structure that produced it.

G2P

Phonological and phonetic representation for speech, carrying tone and vowel length where they are known — and saying so where they are not.

Evidence first

Language data should know what it knows.

RURIMI preserves evidence and uncertainty instead of turning incomplete linguistic knowledge into false certainty. Every claim the system makes carries the tier of evidence behind it — and a great deal is still open.

✓ Known

Established and recorded, with a source. The system may rely on it.

◇ Source-derived

Induced from the project's own data with its supporting counts, rather than asserted from memory.

~ Hypothesis

The engine's proposal. Recorded, testable, and deliberately not adopted as fact.

☖ Native review required

A question only a native speaker can settle. It stays open until one does.

? Unknown

Not yet elicited. Rendered as unknown — never filled in with a plausible guess.

Why this matters for tone

Standard Kinyarwanda orthography writes neither tone nor vowel length. A system that quietly guesses them produces confident, plausible, wrong pronunciation. RURIMI keeps known tone, known toneless and unknown tone as three different states, and reports the third rather than resolving it.

The workspace

A linguistic instrument, not a lookup box.

Behind the surface is a workbench for reading a form the way a linguist would: competing analyses side by side, the evidence behind each, and an explicit account of what is still unresolved.

Values shown are recorded output for this one form. The workspace is under active development alongside the linguistic engine.

For developers

Build on Kinyarwanda linguistic infrastructure.

The engine exposes its analysis over HTTP. These endpoints exist and run locally today; a hosted public API is not yet open, and nothing here is a live service.

Illustrative request shapes against a locally running engine. Response fields are real; the host is a placeholder.

Who it is for

Built for researchers, developers, and language technology teams.

RURIMI is a long-term project. It is most useful to people who need Kinyarwanda structure to be explicit and traceable.

Researchers

Explore Kinyarwanda linguistic structure and the evidence behind each analysis.

Developers

Build applications on structured Kinyarwanda language technology instead of ad-hoc tokenizers.

AI & Speech

Work with morphology, phonology, pronunciation and language-aware representations.

Institutions

Explore future language infrastructure and integration possibilities for Kinyarwanda.

Follow the development of RURIMI.

RURIMI is being developed as a long-term Kinyarwanda language technology project. Follow its progress as the linguistic foundation grows.

There is no sign-up list yet — an email reaches a person rather than a form.