WORDS ARE POWER

There are more than 600,000 words in the English language. An educated native speaker knows perhaps 30,000 of them. For a second language learner, vocabulary can look like an impossible mountain.

It isn't — because words are not equally useful.

Not all words are created equal

A small number of words do most of the work in English. The most frequent word, the, accounts for about 6% of everything you read or hear. The top ten account for 30%. The first 100 reach 58%. The first 1,000 reach 82%. This steep drop-off is known as Zipf's law, and it is the single most useful fact in vocabulary learning.

The consequence is dramatic. Learn the 2,809 words of the New General Service List and you understand roughly 92% of everyday English. Getting from 92% to 100% would mean learning another 597,000 words. Almost none of them would ever appear in what you read.

Why coverage matters

Research on reading gives us thresholds. At 90% coverage a learner has basic understanding, but guessing unknown words from context is hard and reading is slow and discouraging. At 95%, comprehension is good and context clues start to work. At 98%, reading becomes comfortable enough to be done for pleasure — which is where real vocabulary growth begins.

That's why our lists target the coverage they do. The core list plus one special-purpose list gets most learners to the mid-90s or higher in their chosen domain — with under 4,000 words.

How the lists are built

The lists are not all built the same way, because they do not all answer the same question. The New General Service List came out of the Cambridge English Corpus, which runs to more than two billion words — though we used 265 million of them, choosing the sub-corpora closest to the English a second language learner is actually likely to meet. A corpus that is merely enormous will hand you words nobody needs. Where no suitable corpus existed for a special-purpose list, we built one around what learners in that domain have to read and hear; the New Academic Word List drew on a collection of existing academic corpora together with material we added ourselves.

We also tried to follow Michael West, who built the original General Service List in 1953. West combined quantitative and qualitative evidence, and so did we. The corpus tells you what is frequent. It does not tell you what belongs. Experienced teachers and corpus linguists, ourselves included, went through the results by hand — adding words that learners plainly need, removing words that had no business being there.

The coverage targets were a decision, not an accident. Reaching 95% on every special-purpose list would have meant roughly 30% more words in each one, and that extra load would have bought comparatively little. So we set 90% as a minimum and in practice usually beat it — several of the special-purpose lists reach 97% to 99% when combined with the core list. The lists are revised as they get used, too: the NGSL moved from version 1.01 to 1.2 in 2023, taking in changes suggested by linguists and teachers with years of experience working with it.

Everything we make is free

All the word lists are open source under Creative Commons, including for commercial use. Most of the tools are free too. There is no paid version sitting behind them and there never has been. I have paid for the site and the tools out of my own pocket for thirteen years, because these lists are more use in the hands of learners, teachers, materials developers and researchers than in mine.

Beyond the lists, I've spent four decades in language teaching, teacher training and materials development, and co-developed ER-Central, a free extensive reading platform with more than 1,000 graded readers and a free teacher LMS, with Rob Waring.

Charles Browne
Professor of Applied Linguistics and TESOL, Meiji Gakuin University
Director, CBC Consulting