Skip to main content
Calcimator

Feature Importance Calculator

Estimate how many features to keep, overfitting risk, and expected variance retention based on dataset size, model type, and correlation threshold.

About this calculator

Feature selection has to balance two competing pressures: keeping enough features to capture the real signal in your data, and dropping enough to avoid overfitting on a dataset that isn't large enough to support them. This calculator's recommended feature count comes from a model-specific heuristic rather than one universal rule, because different model families tolerate feature counts very differently — linear and regularized models are recommended to keep roughly a 10-to-1 ratio of samples to features to stay statistically stable, tree-based ensembles use the classic square-root-or-log2-of-feature-count rule that random forests popularized, and neural networks get a more generous allowance since their regularization and architecture can absorb more features relative to data. Overfitting risk is estimated from samples-per-feature alone (your dataset size divided by total features) — a ratio under 2 is flagged as very high risk, since a model with nearly as many parameters as data points can effectively memorize noise, while a ratio above 20 is considered comfortably safe.

Estimated variance retained models how much of a dataset's total information the recommended feature subset captures, using a diminishing-returns curve where each additional feature contributes progressively less — the curve's steepness (how concentrated importance is in a few dominant features) varies by model type and by your correlation threshold, since a higher threshold implies more redundant, replaceable features exist among the total. Note that the Target Variance Explained input is not currently used to solve for a feature count — the recommendation always comes from the model-specific heuristics above regardless of what target you enter, so treat Target Variance as a benchmark to compare against the reported Est. Variance Retained rather than a constraint the calculator optimizes for.

Inputs

Results

Recommended Features to Keep

50

Est. Variance Retained

100%

Features to Remove0
Samples per Feature20
Min Samples Needed750
Overfitting Risk8%
How to Use This Calculator
  1. Enter Total Features (number of columns in your dataset).
  2. Enter Dataset Size (number of rows/samples).
  3. Select your Model Type — each type has different feature capacity.
  4. Review Recommended Features to Keep and Overfitting Risk.
  5. Check if your dataset meets the Minimum Samples Needed.
  6. Use the chart to see the keep vs. remove split.

How the result changes with Total Features

Total FeaturesRecommended Features to KeepEst. Variance Retained
2525100%
3838100%
7575100%
12510098.9%

What each input means

Total Features
Total number of features in the dataset
Dataset Size (rows)
Number of samples/rows in the dataset
Model Type
Model type affects feature selection heuristics and data requirements
Target Variance Explained (%)
Desired percentage of variance to retain after feature selection
Correlation Threshold
Features with correlation above this are considered redundant

How this is calculated

Worked example, using the default values

  1. Assess Data-to-Feature Ratio
    effectiveFeatureRatio = datasetSize ÷ totalFeatures
    1000 ÷ 50 = 20 = 20 samples per feature (Very Low overfitting risk)
  2. Calculate Recommended Features
    recommendedFeatures = min(totalFeatures, datasetSize ÷ 10)
    min(50, 1000 ÷ 10) = min(50, 100) = 50 features (Linear/Regularized)
  3. Estimate Variance Retained
    V(k) = [1 − (1 − k/n)^α] × 100
    [1 − (1 − 50/50)^1.8] × 100 = 100% variance retained with 50 of 50 features
  4. Calculate Minimum Samples Needed
    minSamples = totalFeatures × samplesPerFeature (15× for Linear/Regularized)
    50 × 15 = 750 samples needed (you have 1000 ✓)

Engine last updated . Checked against 2 independently-derived tests — how we verify calculators. Built by Paul Gunder, a software engineer, not a licensed financial, medical, or legal professional.

Frequently Asked Questions

Why does setting a different Target Variance Explained not change the Recommended Features output?

The recommended feature count is generated entirely from fixed, model-specific heuristics — sample-to-feature ratios for linear models, square-root or log2 rules for tree ensembles, a looser ratio for neural networks — rather than being solved backward from your target variance figure. Compare your Target Variance against the reported Est. Variance Retained to see whether the heuristic recommendation happens to meet your target, rather than expecting the target itself to drive the recommendation.

Why do tree-based models get a completely different feature-count rule than linear models?

Linear models have no built-in mechanism to ignore irrelevant features, so keeping too many relative to your sample size directly destabilizes the coefficient estimates, which is why the sample-to-feature ratio rule applies there. Tree ensembles instead select a random subset of features at each split, so the relevant heuristic is about how many features to consider per split for good diversity across trees, which is why it follows the square-root or log2 convention popularized by random forest research instead of a strict sample-ratio rule.

What does it mean if my dataset doesn't meet the Minimum Samples Needed figure?

It means your current dataset size falls short of what's generally recommended to reliably support the number of features you have for your chosen model type, which raises the risk of an unstable or overfit model regardless of which specific features you keep. Options include collecting more data, more aggressively reducing your feature count, or choosing a model type with lower sample requirements per feature for your data volume.

How should I interpret the Est. Variance Retained figure?

It estimates what fraction of the dataset's total useful information the recommended feature subset captures, based on a diminishing-returns model where early features contribute more and later features contribute progressively less. It's a heuristic estimate rather than a measurement of your actual data's variance structure, so use it as a planning signal for how aggressive your feature reduction is, and verify against real cross-validated model performance before finalizing a feature set.

The questions that sit next to this one — chosen by subject, including calculators filed under a different category.

More in Technology & Computing.