Skip to main content
Calcimator

Information Gain Calculator

Calculate information gain, gain ratio, Gini gain, and feature importance for decision trees.

About this calculator

Decision trees choose which feature to split on by measuring how much a candidate split reduces uncertainty about the target class, and this calculator works through the entropy-based version of that measurement directly. Weighted child entropy multiplies each child node's entropy by its share of the parent's samples and sums the two, representing the average remaining uncertainty after the split. Information gain is simply the parent's entropy minus that weighted child entropy — the bits of uncertainty the split actually resolved, which is exactly the criterion the ID3 decision-tree algorithm uses to pick the best feature at each node. Information gain ratio (the criterion C4.5 uses instead, to correct ID3's tendency to favor splits with many small branches) normalizes information gain by split information, which penalizes a split that divides samples very unevenly or into many tiny groups.

Feature importance and entropy reduction both express the same information gain as a fraction or percentage of the parent's original entropy, useful for comparing gains across splits with different starting entropy levels. The Gini Gain figure works differently from the rest of this calculator and deserves a direct caveat: it's computed from the same left/right split-weight inputs used for entropy weighting, not from a separate class-label distribution within each node the way real per-node Gini impurity requires — as a result it doesn't behave like the standard splitting-criterion metric used in CART decision trees, and can come out negative even for a split that achieves the maximum possible information gain. Treat Gini Gain here as illustrative at best, and rely on the entropy-based figures — information gain, gain ratio, feature importance, and entropy reduction — for an accurate read on split quality.

Inputs

bits
bits
bits

Results

Information Gain

0.38 bits

Information Gain Ratio

0.391

Feature Importance

0.38

Gini Gain-0.24
Entropy Reduction38%
How to Use This Calculator
  1. Enter the Parent Entropy (in bits) before the split.
  2. Enter the Left/Right Child Entropy and Left/Right Weight (sample proportions) for each side of the split.
  3. Review the calculated Information Gain, Information Gain Ratio, and Gini Gain from the split.
  4. Select the feature with the highest information gain for the next decision tree split.
  5. Use Gini impurity as an alternative splitting criterion for faster computation in large datasets.

How the result changes with Parent Entropy

Parent EntropyInformation GainInformation Gain RatioFeature Importance
0.5-0.12 bits-0.124-0.24
0.750.13 bits0.1340.173
1.50.88 bits0.9060.587
2.51.88 bits1.9360.752

What each input means

Parent Entropy
Entropy of parent node
Left Child Entropy
Entropy of left child node
Right Child Entropy
Entropy of right child node
Left Weight
Proportion of samples in left child
Right Weight
Proportion of samples in right child

How this is calculated

Formula

IG = H(Parent) - Σ(Weight × H(Child))

Worked example, using the default values

  1. Identify Input Parameters
    4 parameters
    Parent Entropy = 1, Left Child Entropy = 0.5, Right Child Entropy = 0.8, Left Weight = 0.6 = 5 input(s) provided
  2. Calculate Information Gain
    Information Gain
    0.38 = 0.38
  3. Calculate Information Gain Ratio
    Information Gain Ratio
    0.391 = 0.391
  4. Calculate Feature Importance
    Feature Importance
    0.38 = 0.38
  5. Calculate Gini Gain
    Gini Gain
    -0.24 = -0.24
  6. Calculate Entropy Reduction
    Entropy Reduction
    38 = 38%

Engine last updated . Checked against 2 independently-derived tests — how we verify calculators. Built by Paul Gunder, a software engineer, not a licensed financial, medical, or legal professional.

Frequently Asked Questions

Why does Information Gain Ratio sometimes differ noticeably from raw Information Gain?

Information Gain Ratio divides raw information gain by split information, a penalty term that grows larger when a split divides samples very unevenly between branches. This correction exists because raw information gain alone tends to favor splits that create many small, uneven branches — the C4.5 algorithm introduced the ratio specifically to counteract that bias.

What does Feature Importance actually measure here?

It expresses information gain as a fraction of the parent node's original entropy, which normalizes the gain so splits starting from different entropy levels can be compared on the same scale. A feature that resolves 40% of a low-entropy parent's uncertainty and one that resolves 40% of a high-entropy parent's uncertainty both show a Feature Importance of 0.4, even though the underlying entropy reduction differs in absolute terms.

Why might the Gini Gain value look unreliable or even negative?

This calculator's Gini Gain is computed from the same left/right split-weight proportions used for the entropy calculation, not from an actual class-label distribution within each child node, which is what real per-node Gini impurity requires. Because of that mismatch, Gini Gain here doesn't behave like the standard CART splitting metric and can come out negative even for a split that achieves the maximum possible information gain — treat the entropy-based figures as the reliable measures of split quality instead.

How do I use this calculator to choose between two candidate splits?

Run each candidate feature's split through the calculator using its own parent entropy, child entropies, and split weights, then compare the resulting Information Gain (or Information Gain Ratio, if branch sizes are very uneven) — the feature producing the higher value reduces more uncertainty about the target class and is the better choice for that node.

What does it mean if Information Gain comes out as a small number close to zero?

A near-zero information gain means the split barely reduced uncertainty about the target class — the two child nodes are almost as mixed as the parent was, suggesting that feature isn't very predictive at that point in the tree. A decision tree algorithm would generally prefer a different feature with a higher information gain for that split, if one is available.

The questions that sit next to this one — chosen by subject, including calculators filed under a different category.

More in Technology & Computing.