Word Frequency Calculator
Analyze word frequency distributions using type-token ratio, hapax legomena estimation, and Zipf's law frequency ranking.
About this calculator
This calculator estimates corpus-linguistics statistics from three summary counts, rather than analyzing actual text -- enter your corpus's total word count, its number of distinct word forms, and how many times one target word occurs, and it derives the rest from established approximations. Type-Token Ratio is simply Unique Words divided by Corpus Size, a standard measure of lexical diversity that falls as a corpus grows (common words like "the" repeat far more than rare ones do, so the ratio of new-word-forms to total words shrinks as you add more text). Estimated Hapax Legomena approximates how many word forms appear exactly once, using the common rule of thumb that roughly half of a corpus's unique words are hapax legomena.
Frequency Rank and Percentile estimate where your Target Word Frequency would fall in a Zipfian frequency distribution, using a harmonic-number approximation of Zipf's law rather than an actual ranked frequency list -- treat these as ballpark estimates for corpora that roughly follow Zipf's law, not exact counts from your specific text. Target Word Frequency never affects Type-Token Ratio, since that ratio depends only on Corpus Size and Unique Words. Frequency Rank is bounded to the vocabulary you entered -- it can never be finer than rank 1 (the most frequent word) or coarser than rank N (the least), and Percentile is derived from that bounded rank, so a very common word always reads near 100% and a very rare one near 0%.
Inputs
Results
Type-Token Ratio
0.1
How to Use This Calculator
- Enter Corpus Size -- the total number of words (tokens) in your text.
- Enter Unique Words -- the number of distinct word forms (types) in that same text.
- Enter Target Word Frequency -- how many times one word of interest occurs in the corpus.
- Review Type-Token Ratio and Estimated Hapax Legomena for a snapshot of the corpus's lexical diversity.
- Check Frequency Rank and Percentile to see roughly where your target word falls in the corpus's frequency distribution.
How the result changes with Corpus Size (total words)
| Corpus Size (total words) | Type-Token Ratio |
|---|---|
| 25,000 | 0.2 |
| 37,500 | 0.1333 |
| 75,000 | 0.0667 |
| 125,000 | 0.04 |
What each input means
- Corpus Size (total words)
- Total number of words (tokens) in your text corpus.
- Unique Words (types)
- Number of distinct word forms (types) in the corpus. Cannot exceed Corpus Size.
- Target Word Frequency
- Number of times your target word appears in the corpus.
How this is calculated
Worked example, using the default values
- Identify Input ParametersCorpus Size (total words) = 50000, Unique Words (types) = 5000, Target Word Frequency = 100 = 3 input(s) provided
- Calculate Type-Token RatioType-Token Ratio0.1 = 0.1
- Calculate Estimated Hapax LegomenaEstimated Hapax Legomena2500 = 2500
- Calculate Frequency RankFrequency Rank = max(155 = 55
Engine last updated . Checked against 1 independently-derived test — how we verify calculators. Built by Paul Gunder, a software engineer, not a licensed financial, medical, or legal professional.
Frequently Asked Questions
Do I need to paste in my actual text?
No -- this calculator works from three summary numbers you already know or have counted: total corpus size, number of unique word forms, and how many times your target word occurs. It estimates the rest using standard corpus-linguistics approximations rather than analyzing raw text.
Why does Type-Token Ratio get smaller for larger corpora?
As a corpus grows, common words like "the" and "and" repeat far more often than new word forms are introduced, so the ratio of unique words to total words shrinks. A short text can have a high ratio simply because it hasn't had room to repeat words yet.
How accurate is the Estimated Hapax Legomena figure?
It uses the common approximation that roughly half of a corpus's unique word forms occur exactly once -- a pattern seen often in natural-language corpora, but not a fixed law. The real hapax legomena share varies by genre, corpus size, and language, so treat this as an estimate rather than an exact count.
When should I be skeptical of the Frequency Rank and Percentile estimates?
They come from a Zipf's-law approximation of your two summary counts, not an actual ranked word list, so they hold up best for general-purpose text of the kind Zipf's law was fit to. Be more skeptical for a short corpus, or specialized text -- technical jargon, a single author's idiosyncratic vocabulary -- where a handful of words dominate more than average; treat the numbers as a rough placement rather than a precise one. Type-Token Ratio, by contrast, is calculated purely from Corpus Size and Unique Words and never moves with Target Word Frequency, so it isn't subject to this same uncertainty.
Related Calculators
The questions that sit next to this one — chosen by subject, including calculators filed under a different category.
Text Complexity Calculator
Analyze vocabulary diversity, type-token ratio, and overall text complexity using linguistic metrics like Guiraud's index.
Linguistics & TranslationPhonetic Transcription Calculator
Estimate phoneme count, transcription time, and complexity score for IPA transcription based on word characteristics.
Linguistics & TranslationReadability Score Calculator
Calculate multiple readability indices including Flesch Reading Ease, Flesch-Kincaid Grade, Gunning Fog, Coleman-Liau, and Automated Readability Index.
More in Education & Academics.