Tonify
Mandarin pronunciation, graded.

Problem
Mandarin tones are the part of the language least served by reading practice. A learner can know a word’s meaning and characters and still be unintelligible, because tone carries as much information as the consonants, and self-study gives no signal about which ones are wrong.
The apps that do grade speech score it with black-box accuracy numbers, which cannot say which syllable was wrong or how. A learner is told they failed without being told what to change.
How it was solved
Montreal Forced Aligner stretches the known word’s acoustic models over the recording, so measurement happens inside exact per-syllable spans rather than guessing where boundaries fall.
Parselmouth tracks pitch inside those spans, gated to 75 to 400 Hz and scaled in semitones against a per-speaker baseline, since tones are relative: a deep voice’s high tone sits below a high voice’s low tone.
Defining features, not templates. Each tone is graded on the feature that defines it (rise, level, fall, low dip) rather than distance from an exemplar curve, because genuinely correct tones fail template conformance.
Consonant separation. A wav2vec2 CTC network scores the expected syllable against its confusable neighbor as a forced choice, in a layer structurally excluded from tone differences, so a wrong tone can never masquerade as a wrong consonant.
Results
Ran at tonedrill.com through August 2026, open to the internet: 14 learners, 254 graded attempts, 503 syllables measured across 16 words, behind 614 passing tests.
Three users praised the pitch chart and its acceptance region without being asked. The share of syllables the grader had to skip as unmeasurable fell from 38% on an early tester to 14% across the public run, after recording guidance and capture timing were reworked in response. One learner kept coming back across three separate weeks, making 52 attempts in total.