Skip to main content
Calcimator

Word Frequency Calculator

Analyze word frequency distributions using type-token ratio, hapax legomena estimation, and Zipf's law frequency ranking.

About this calculator

This calculator estimates corpus-linguistics statistics from three summary counts, rather than analyzing actual text -- enter your corpus's total word count, its number of distinct word forms, and how many times one target word occurs, and it derives the rest from established approximations. Type-Token Ratio is simply Unique Words divided by Corpus Size, a standard measure of lexical diversity that falls as a corpus grows (common words like "the" repeat far more than rare ones do, so the ratio of new-word-forms to total words shrinks as you add more text). Estimated Hapax Legomena approximates how many word forms appear exactly once, using the common rule of thumb that roughly half of a corpus's unique words are hapax legomena.

Frequency Rank and Percentile estimate where your Target Word Frequency would fall in a Zipfian frequency distribution, using a harmonic-number approximation of Zipf's law rather than an actual ranked frequency list -- treat these as ballpark estimates for corpora that roughly follow Zipf's law, not exact counts from your specific text. Target Word Frequency never affects Type-Token Ratio, since that ratio depends only on Corpus Size and Unique Words. Frequency Rank is bounded to the vocabulary you entered -- it can never be finer than rank 1 (the most frequent word) or coarser than rank N (the least), and Percentile is derived from that bounded rank, so a very common word always reads near 100% and a very rare one near 0%.

Inputs

Results

Type-Token Ratio

0.1

Estimated Hapax Legomena2,500
Frequency Rank#55
Percentile98.92%
How to Use This Calculator
  1. Enter Corpus Size -- the total number of words (tokens) in your text.
  2. Enter Unique Words -- the number of distinct word forms (types) in that same text.
  3. Enter Target Word Frequency -- how many times one word of interest occurs in the corpus.
  4. Review Type-Token Ratio and Estimated Hapax Legomena for a snapshot of the corpus's lexical diversity.
  5. Check Frequency Rank and Percentile to see roughly where your target word falls in the corpus's frequency distribution.

How the result changes with Corpus Size (total words)

Corpus Size (total words)Type-Token Ratio
25,0000.2
37,5000.1333
75,0000.0667
125,0000.04

What each input means

Corpus Size (total words)
Total number of words (tokens) in your text corpus.
Unique Words (types)
Number of distinct word forms (types) in the corpus. Cannot exceed Corpus Size.
Target Word Frequency
Number of times your target word appears in the corpus.

How this is calculated

Worked example, using the default values

  1. Identify Input Parameters
    Corpus Size (total words) = 50000, Unique Words (types) = 5000, Target Word Frequency = 100 = 3 input(s) provided
  2. Calculate Type-Token Ratio
    Type-Token Ratio
    0.1 = 0.1
  3. Calculate Estimated Hapax Legomena
    Estimated Hapax Legomena
    2500 = 2500
  4. Calculate Frequency Rank
    Frequency Rank = max(1
    55 = 55

Engine last updated . Checked against 1 independently-derived test — how we verify calculators. Built by Paul Gunder, a software engineer, not a licensed financial, medical, or legal professional.

Frequently Asked Questions

Do I need to paste in my actual text?

No -- this calculator works from three summary numbers you already know or have counted: total corpus size, number of unique word forms, and how many times your target word occurs. It estimates the rest using standard corpus-linguistics approximations rather than analyzing raw text.

Why does Type-Token Ratio get smaller for larger corpora?

As a corpus grows, common words like "the" and "and" repeat far more often than new word forms are introduced, so the ratio of unique words to total words shrinks. A short text can have a high ratio simply because it hasn't had room to repeat words yet.

How accurate is the Estimated Hapax Legomena figure?

It uses the common approximation that roughly half of a corpus's unique word forms occur exactly once -- a pattern seen often in natural-language corpora, but not a fixed law. The real hapax legomena share varies by genre, corpus size, and language, so treat this as an estimate rather than an exact count.

When should I be skeptical of the Frequency Rank and Percentile estimates?

They come from a Zipf's-law approximation of your two summary counts, not an actual ranked word list, so they hold up best for general-purpose text of the kind Zipf's law was fit to. Be more skeptical for a short corpus, or specialized text -- technical jargon, a single author's idiosyncratic vocabulary -- where a handful of words dominate more than average; treat the numbers as a rough placement rather than a precise one. Type-Token Ratio, by contrast, is calculated purely from Corpus Size and Unique Words and never moves with Target Word Frequency, so it isn't subject to this same uncertainty.

The questions that sit next to this one — chosen by subject, including calculators filed under a different category.

More in Education & Academics.