← Todos los artículos
Characters

Is memorising 1,000 Chinese characters hard? What that number actually buys

2026-09-08 · 7 min de lectura

A coverage curve showing that the first 1,000 Chinese characters cover 92.8% of text but 98% needs 1,698

"Learn 1,000 characters and you can read a newspaper" is the most repeated number in Chinese learning, and it is close enough to true to be dangerous. A thousand characters does cover most of what you meet. It also leaves you unable to read.

Both things are measurable, so let's measure them.

What 1,000 characters actually covers

Take the 9,005 example sentences in our own library — one written for every word in the HSK 3.0 syllabus, across all nine levels. That's 92,428 character tokens drawn from 2,857 distinct characters. Rank the characters by frequency and add up how much of the text each block covers:

Characters knownShare of all character tokens
5030.2%
10042.1%
30067.2%
50079.6%
1,00092.8%
1,50097.1%
2,00098.9%

The curve is brutally steep at the start. Fifty characters — 的, 这, 了, 我, 在, 一, 他, 人, 个, 家 and forty friends — carry nearly a third of everything written. That is why the first month of Chinese feels miraculous.

Then it flattens. Going from 500 to 1,000 characters, a doubling of effort, buys you 13 percentage points. Going from 1,000 to 1,500 buys four.

Why 92.8% is not enough to read

Research on second-language reading keeps landing on the same threshold: to read comfortably without a dictionary you need to know about 98% of what is on the page. At 95% comprehension starts to break down; below that you are decoding.

92.8% is one unknown character in every fourteen. In a ten-character sentence — which is roughly the length of an HSK sentence at any level — that is an unknown character in most sentences you meet.

To reach 98% on this corpus you need 1,698 characters. For 99%, 2,067. The famous thousand gets you to the point where the text stops being opaque and starts being annoying, which is a real milestone and not the one people think they are buying.

Knowing 1,000 characters leaves 37% of HSK words still unreadable, because a word fails if any single character in it is unknown

The multiplication problem nobody mentions

Character coverage and word coverage are different numbers, and the gap between them is where the "I know the characters but not the sentence" feeling comes from.

A word is only readable if you know every character in it. Most Chinese words are two characters, so at 92.8% per character the chance of clearing a two-character word is roughly 0.928 × 0.928 ≈ 86% — and that is before you deal with words whose meaning is not the sum of their parts.

Run it against the actual syllabus rather than the arithmetic:

Characters knownHSK words you can fully read
100677 of 11,000 (6.2%)
3002,400 (21.8%)
5003,987 (36.2%)
1,0006,939 (63.1%)

A thousand characters and more than a third of the vocabulary still has a hole in it.

So is memorising 1,000 characters hard?

Hard is the wrong axis. The honest answers are:

It is very achievable. At ten new characters a day it is a hundred days. The frequency curve is on your side the whole way — every character you learn early is worth more than every character you learn late, which is the opposite of most skills.

It is not the finish line. The number that unlocks unassisted reading is closer to 1,700, and the second thousand is much slower than the first because you have already taken all the high-frequency ones.

Recognition is the easy half. Recognising 汉 in context and producing it from memory are different abilities, separated by months. Anyone quoting a character count is quoting the recognition number.

The trap in the back half

Here is the part that makes the second thousand disproportionately painful. Of the 2,857 characters used across the whole syllabus, 916 — nearly a third — appear in exactly one word.

They are not building blocks. They are one-offs: a character you learn to read a single word and then may not meet again for months. 爸 (only in 爸爸), 饺 (only in 饺子), 苹 (only in 苹果). Learning them as characters is close to wasted effort; they only ever appear inside their word, and the word is what you actually need.

That is the practical split. High-frequency characters are worth learning as characters, because they recombine endlessly. Low-frequency ones are worth learning as parts of the word they live in, and not otherwise.

Which thousand, though?

"A thousand characters" is silent on which thousand, and the choice is worth about ten percentage points of coverage.

Working through the HSK syllabus in order, here is where you stand:

ThroughDistinct charactersCoverage
HSK 124849.2%
HSK 237159.6%
HSK 365576.8%
HSK 41,09690.1%
HSK 51,52795.8%

So the thousand-character mark lands roughly at the end of HSK 4.

Now compare each of those against the same number of characters chosen by pure frequency instead:

CharactersHSK orderFrequency order
24849.2%62.5%
37159.6%72.5%
65576.8%85.4%
1,09690.1%94.0%
1,52795.8%97.3%

Early on the syllabus is measurably less efficient than raw frequency — thirteen points behind at the 248-character mark. That is not a flaw. HSK 1 is built to let you say things, so it spends characters on 你好, 谢谢, 再见 and numbers, which are socially essential and statistically unremarkable. Frequency order would have you reading sooner and speaking later.

The gap closes as you climb — 1.5 points by HSK 5 — because both routes eventually pick up the same high-frequency core. If your goal is reading specifically, frequency order pays for the first few hundred characters and stops mattering after that.

What to do instead of chasing a number

Front-load ruthlessly, and stop counting early. The first 300 characters buy 67% coverage. Nothing else in the language has that return, so nothing should compete with it for your attention.

Switch from characters to words at around 500. Past that point, coverage grows faster through words than through isolated characters, because you start meeting characters you already know in new combinations.

Measure coverage, not count. Paste something you want to read into a text analyzer and look at the fraction above your level. That number tells you whether you can read this — which is what you actually wanted to know when you asked about a thousand characters.

Check the frequency curve is still on your side. If you are learning characters that appear in one word, you have crossed into the flat part of the curve and your hours are buying much less than they did.

A thousand characters is a real and worthwhile milestone. It is roughly two thirds of the way to reading — not the arrival.

¿Listo para empezar?

Descubre tu nivel de HSK 3.0 en 5 minutos: gratis y sin tarjeta.

Hacer el test de nivel