Is memorising 1,000 Chinese characters hard? What that number actually buys
2026-09-08 · 7 мин чтения

"Learn 1,000 characters and you can read a newspaper" is the most repeated number in Chinese learning, and it is close enough to true to be dangerous. A thousand characters does cover most of what you meet. It also leaves you unable to read.
Both things are measurable, so let's measure them.
What 1,000 characters actually covers
Take the 9,005 example sentences in our own library — one written for every word in the HSK 3.0 syllabus, across all nine levels. That's 92,428 character tokens drawn from 2,857 distinct characters. Rank the characters by frequency and add up how much of the text each block covers:
| Characters known | Share of all character tokens |
|---|---|
| 50 | 30.2% |
| 100 | 42.1% |
| 300 | 67.2% |
| 500 | 79.6% |
| 1,000 | 92.8% |
| 1,500 | 97.1% |
| 2,000 | 98.9% |
The curve is brutally steep at the start. Fifty characters — 的, 这, 了, 我, 在, 一, 他, 人, 个, 家 and forty friends — carry nearly a third of everything written. That is why the first month of Chinese feels miraculous.
Then it flattens. Going from 500 to 1,000 characters, a doubling of effort, buys you 13 percentage points. Going from 1,000 to 1,500 buys four.
Why 92.8% is not enough to read
Research on second-language reading keeps landing on the same threshold: to read comfortably without a dictionary you need to know about 98% of what is on the page. At 95% comprehension starts to break down; below that you are decoding.
92.8% is one unknown character in every fourteen. In a ten-character sentence — which is roughly the length of an HSK sentence at any level — that is an unknown character in most sentences you meet.
To reach 98% on this corpus you need 1,698 characters. For 99%, 2,067. The famous thousand gets you to the point where the text stops being opaque and starts being annoying, which is a real milestone and not the one people think they are buying.

The multiplication problem nobody mentions
Character coverage and word coverage are different numbers, and the gap between them is where the "I know the characters but not the sentence" feeling comes from.
A word is only readable if you know every character in it. Most Chinese words are two characters, so at 92.8% per character the chance of clearing a two-character word is roughly 0.928 × 0.928 ≈ 86% — and that is before you deal with words whose meaning is not the sum of their parts.
Run it against the actual syllabus rather than the arithmetic:
| Characters known | HSK words you can fully read |
|---|---|
| 100 | 677 of 11,000 (6.2%) |
| 300 | 2,400 (21.8%) |
| 500 | 3,987 (36.2%) |
| 1,000 | 6,939 (63.1%) |
A thousand characters and more than a third of the vocabulary still has a hole in it.
So is memorising 1,000 characters hard?
Hard is the wrong axis. The honest answers are:
It is very achievable. At ten new characters a day it is a hundred days. The frequency curve is on your side the whole way — every character you learn early is worth more than every character you learn late, which is the opposite of most skills.
It is not the finish line. The number that unlocks unassisted reading is closer to 1,700, and the second thousand is much slower than the first because you have already taken all the high-frequency ones.
Recognition is the easy half. Recognising 汉 in context and producing it from memory are different abilities, separated by months. Anyone quoting a character count is quoting the recognition number.
The trap in the back half
Here is the part that makes the second thousand disproportionately painful. Of the 2,857 characters used across the whole syllabus, 916 — nearly a third — appear in exactly one word.
They are not building blocks. They are one-offs: a character you learn to read a single word and then may not meet again for months. 爸 (only in 爸爸), 饺 (only in 饺子), 苹 (only in 苹果). Learning them as characters is close to wasted effort; they only ever appear inside their word, and the word is what you actually need.
That is the practical split. High-frequency characters are worth learning as characters, because they recombine endlessly. Low-frequency ones are worth learning as parts of the word they live in, and not otherwise.
Which thousand, though?
"A thousand characters" is silent on which thousand, and the choice is worth about ten percentage points of coverage.
Working through the HSK syllabus in order, here is where you stand:
| Through | Distinct characters | Coverage |
|---|---|---|
| HSK 1 | 248 | 49.2% |
| HSK 2 | 371 | 59.6% |
| HSK 3 | 655 | 76.8% |
| HSK 4 | 1,096 | 90.1% |
| HSK 5 | 1,527 | 95.8% |
So the thousand-character mark lands roughly at the end of HSK 4.
Now compare each of those against the same number of characters chosen by pure frequency instead:
| Characters | HSK order | Frequency order |
|---|---|---|
| 248 | 49.2% | 62.5% |
| 371 | 59.6% | 72.5% |
| 655 | 76.8% | 85.4% |
| 1,096 | 90.1% | 94.0% |
| 1,527 | 95.8% | 97.3% |
Early on the syllabus is measurably less efficient than raw frequency — thirteen points behind at the 248-character mark. That is not a flaw. HSK 1 is built to let you say things, so it spends characters on 你好, 谢谢, 再见 and numbers, which are socially essential and statistically unremarkable. Frequency order would have you reading sooner and speaking later.
The gap closes as you climb — 1.5 points by HSK 5 — because both routes eventually pick up the same high-frequency core. If your goal is reading specifically, frequency order pays for the first few hundred characters and stops mattering after that.
What to do instead of chasing a number
Front-load ruthlessly, and stop counting early. The first 300 characters buy 67% coverage. Nothing else in the language has that return, so nothing should compete with it for your attention.
Switch from characters to words at around 500. Past that point, coverage grows faster through words than through isolated characters, because you start meeting characters you already know in new combinations.
Measure coverage, not count. Paste something you want to read into a text analyzer and look at the fraction above your level. That number tells you whether you can read this — which is what you actually wanted to know when you asked about a thousand characters.
Check the frequency curve is still on your side. If you are learning characters that appear in one word, you have crossed into the flat part of the curve and your hours are buying much less than they did.
A thousand characters is a real and worthwhile milestone. It is roughly two thirds of the way to reading — not the arrival.
Готовы начать?
Определите свой уровень HSK 3.0 за 5 минут — бесплатно и без карты.
Пройти тест на уровень