Pronunciation · 声与调
What a Chinese tone checker is actually measuring
A Chinese tone checker or speaking-app score tells you whether a recognizer matched your sound to the words it expected. That is not the same as a listener hearing the right tones.
Read September 28, 2026. No recording of our own is on this page.
A green check from a Chinese tone checker, or from the speaking step in a learning app, is evidence that software could map your sound onto the words it was expecting. It is not evidence that a person listening would hear the tones you meant. Those two jobs part company most clearly on a whole sentence, which is also where most apps give you the most practice.
People meet the gap from both sides. In January 2018 a HelloChinese user wrote that you can almost say anything at times and the app will call it perfect. On the same thread another user kept failing and wrote, “Gah I am saying it!” The score can flatter you. It can also refuse you. Neither result, by itself, tells you which.
What is the recognizer doing?
A general speech recognizer is built to recover the words you probably meant. It listens to the sound, then leans on what usually comes next in the language, and on what people often say to a phone. Hacking Chinese put the job in one sentence: speech recognition “is not meant to give you a fair assessment of your pronunciation, it’s designed to understand what you want to say.”
That is why a Chinese tone checker in the search results and a speaking exercise in an app can look like the same tool and still be doing different work. Some pages play a contour and ask you to hear a contrast. Some compare your pitch to a reference. Some, the ones inside most consumer apps, ask a recognizer whether it can reconstruct the prompt. A reconstructed prompt is not a tone grade.
Why is a whole sentence the loosest test?
The longer the prompt, the more the system can guess. Hacking Chinese, after playing student recordings into phone dictation, wrote that speech recognition “is probably too lenient for sentences.” The next sentence is the mechanism: “Provided that the sentence is fairly common, you need to make several errors at once to derail the speech recognition algorithm.”
So a pass on “请问,我可以进来吗?” can mean the recognizer landed on a polite request that people actually say. It need not mean each tone was right. The same site’s warning is the one that matters for practice: “You can not assume that your pronunciation is good just because your phone writes out the right sentence.”
Why is a single syllable the hardest?
A lone syllable gives the recognizer almost nothing to guess from. Hacking Chinese called speech recognition “next to useless for single-syllable words,” unless you only want to check something already close to a native production, and even then a dipping third tone can be split into two syllables. In the first part of that series, a native teacher’s 耳 came back as 二二 on Android and as 嗯 on iOS.
That is a false fail on a correct sound. It is also why a red mark on one character in an app is a poor reason to rerecord twenty times. You may be feeding a system that is weak on short items, not proving that your tone collapsed.
What have published measurements shown?
Two numbers get quoted as if they described the same machine. They do not.
| What was measured | What kind of tool | What the paper said | What that cannot decide |
|---|---|---|---|
| Wrong tones on shadowed words (Weenink et al., Interspeech 2007) | SpeakGoodChinese, a research system that compares your production to a synthetic pitch reference from pinyin | “less than 15% acceptance rate on incorrectly produced tones”; 6% rejection on reference readings of six words | The pass rate of any named learning app. The authors call the figures preliminary. |
| Tone diagnosis on L2 read-aloud sentences (Wu, Speech Prosody 2024) | Two unnamed automated rating APIs, one from a Chinese company and one from an international company, scored against human raters | “the overall accuracy of tonal diagnosis reached 80%, as determined through the calculation of false rejection and false acceptance rates.” Pause scores correlated 0.6 with human detection of unnatural boundaries. | Your score in HelloChinese, Duolingo, or any other unnamed app. The 0.69 to 0.81 figures in that paper belong to earlier SpeechRater studies, not to this tone test. |
A pitch-tracking research tool can refuse a wrong tone most of the time and still mis-reject a correct one. A commercial rating API can reach about 80% diagnostic accuracy on classroom read-alouds and still be the wrong instrument for a learner who wants to know whether tone 3 on one syllable was right. Mixing the two numbers into one “tone tech is 80% now” sentence is a category error.
Wu’s table of diagnostic accuracy by tone does not support a further claim that the third tone is uniquely worse. Diagnostic accuracy there sits between 73% and 85% across the tones the paper lists, and the authors write that they did not observe significant differences among tones, tone sandhi, and pauses.
What did earlier tests leave out?
Hacking Chinese measured phone dictation on a OnePlus 6 and an iPhone 6. The first article says so, and then sets learning apps aside: “There is actually a third question here about using speech recognition in various apps for learning Chinese, some of which might be purpose-built, but I’ll leave that for a possible follow-up article later.” Neither part plants an intentional wrong tone and then records what the system does. The errors in part 2 are the errors the students already made.
People have still tried a cruder version of that missing test. On 7 February 2018, Juraj Tomek wrote on Chinese Stack Exchange that HelloChinese tone recognition seemed inaccurate: “I can say an entire sentence with the same tone, which is clearly wrong, and still the app would show half of the tones as correctly pronounced.” That is one person’s report, with no recording on the page. It is the closest published attempt we have found to injecting a tone error on purpose. It is not a controlled sample.
This page is not the first to notice the problem. It is also not a substitute for a recorded, syllable-by-syllable contrast. We have not run that contrast.
Why do a developer and a published test sound as if they disagree?
They sound opposite if you treat “consistent” as “good at judging pronunciation.” They can both be true if you keep the two jobs apart.
On 26 January 2018, Chong, posting as a HelloChinese developer, wrote: “Regarding the voice recognition, short voices like a single character can sometimes be easily affected, while technically it’s quite consistent for long ones like sentences.” That is a claim about recognition stability: a sentence is easier for the system to lock onto than a single character.
Hacking Chinese found the same shape from the other side of the desk. Single syllables were the least trustworthy. Sentences were where the system became too kind to use as a pronunciation grade. Consistency at recovering a common sentence is exactly what makes the score a poor tone checker.
Chong also confirmed a different mechanism that readers often fold into “the recognizer is loose.” After a user said you can keep repeating the same mistake and eventually pass, Chong answered that this is a “secret feature”: when only one question is left in a session and you keep failing, the app lets you through, “in case a user gets irritated and throws his/her phone away.” That is a session rule. It is not a statement that the recognizer accepts a wrong tone as a right one.
So how should I use the score?
Use it as a rough signal that the system heard something like the prompt. Do not use it as a verdict on your tones, and do not use a single red mark as a verdict either.
If the sentence is ordinary and the app still cannot land on it after a clean recording, something in the sound is probably off, or the microphone is. If the sentence is ordinary and the app passes you, you have learned almost nothing about the tones. The useful next step is not another try at pleasing the same recognizer. It is a check the recognizer is bad at faking.
What can I try without treating the app as a judge?
You do not need our recording to run a contrast. You need two takes of the same prompt.
- Say the sentence as written, once, in a quiet room. Note the mark.
- Say it again with every syllable forced onto first tone. If you were practicing 你好, that second take is nī hāo, not nǐ hǎo. Note the mark.
- If both takes pass, the score is not checking tone. If only the first pass, the system is at least sensitive to a gross contour change. If both fail, you have a recording or recognition problem, not a tone lesson.
That is a household version of what Tomek reported in 2018. It still will not give you a percentage. It will tell you whether this particular exercise is a Chinese tone checker or a sentence-guessing engine. Eight lines with both traps written out are on the tone sentence bank.
A person who can hear the difference remains the standard this site is named for. The editorial policy is the rule for how a later page may claim a test. This page does not claim one.
If the number you are using to measure yourself is 2,200 class hours, that figure is an ILR score of 3, not a tone check. The matching note is how long it takes to learn Chinese.
Questions that still sit on the table
Does a Chinese tone checker check tones?
Some tools try to. A speaking-app score often does not. If the system is a general recognizer, a pass can mean it reconstructed the expected words from context. That can happen even when the tones were wrong. A research tool that compares your pitch to a reference contour is doing a different job. The two results are not interchangeable.
Why did the app mark me wrong when I think I was right?
A red mark can be a real tone problem, a recording problem, or a recognizer that is weak on a short syllable. Hacking Chinese found single-syllable dictation unreliable even with a native teacher on some third-tone items. A developer of HelloChinese has also said a single character can be easily affected. One red mark is not a diagnosis.
Can I trust HelloChinese speaking exercises?
You can trust them as practice prompts. You should not trust them as a tone verdict. Users have reported both a pass after saying almost anything and a fail when they thought they were close. The developer confirmed a pass after repeated failure on the last item in a session. That is a frustration valve, not a claim that the recognizer accepts wrong tones. This page does not give HelloChinese a pass rate.
What should I use instead of the green check?
Use a person when you can. If you cannot, compare your recording to a native recording of the same sentence, or open the tone sentence bank and run the same-tone check on one line.
Sources
- Hacking Chinese, “Using speech recognition to improve Chinese pronunciation, part 1,” https://www.hackingchinese.com/using-speech-recognition-to-improve-chinese-pronunciation-part-1/ (read 28 September 2026).
- Hacking Chinese, “How good is voice recognition for learning Chinese pronunciation?” https://www.hackingchinese.com/using-speech-recognition-to-improve-chinese-pronunciation-part-2/ (read 28 September 2026).
- Yao Wu, “The assessment of automated rating of L2 Mandarin prosody in lexical tone recognition and pauses,” Speech Prosody 2024, 250–254, https://www.isca-archive.org/speechprosody_2024/wu24_speechprosody.html (read 28 September 2026).
- David Weenink et al., “Learning tone distinctions for Mandarin Chinese,” Interspeech 2007, 2341–2344, https://www.isca-archive.org/interspeech_2007/weenink07_interspeech.html (read 28 September 2026).
- Juraj Tomek, “Chinese Pronunciation App,” Chinese Language Stack Exchange, 7 February 2018, https://chinese.stackexchange.com/questions/28745 (read 28 September 2026).
- Hayden05 and Chong, thread 49944, page 3, Chinese-Forums, 25–26 January 2018, https://www.chinese-forums.com/forums/topic/49944-hellochinese/?page=3 (read 28 September 2026).
Version 1. Recheck when the KCI 2023 HelloChinese paper is in hand, if Hacking Chinese publishes the promised app follow-up, or if this site records a tone-trap contrast. Otherwise every six months.