Skip to main content
Calcimator

Cross-Validation Calculator

Calculate k-fold cross-validation splits, train-test splits, and data utilization for machine learning.

About this calculator

Cross-validation splits a dataset so a model can be trained and evaluated without ever testing on the same rows it trained on. This calculator layers two splits together: a train-test split reserves Test Set Size percent of the full dataset for a held-out test set (never used in training or k-fold), and of what remains, Validation Set Size percent is further set aside as a validation set, leaving Final Training Size as what the model actually trains on before you even get to cross-validation.

K-Folds then determines how the calculator estimates fold-level statistics: Fold Size is simply Dataset Size divided by K-Folds, and Training Size per Fold is what's left over -- so more folds means a smaller held-out fold each round and a larger training portion per fold, which is exactly why Data Utilization (the share of the full dataset used to train in any given fold) rises as K-Folds increases, approaching but never reaching 100%. There's a real tradeoff behind that number, though this calculator only reports the sizes, not the tradeoff itself: more folds means more training data per fold (usually better model quality) but also more separate training runs and less independent held-out data per fold (a noisier per-fold estimate) -- 5 and 10 are the most common choices in practice as a balance between the two.

Inputs

Results

Fold Size

2,000

Training Size per Fold

8,000

Test Set Size

2,000

Validation Set Size800
Final Training Size7,200
Data Utilization80%
How to Use This Calculator
  1. Enter the Dataset Size (total samples) and K-Folds (typically 5 or 10) for the cross-validation split.
  2. Set the Test Set Size percentage to reserve a held-out test set never used in training.
  3. Set the Validation Set Size percentage, taken from what remains after the test split.
  4. Review Fold Size, Training Size per Fold, and Final Training Size to see how your data divides up.
  5. Check Data Utilization to see what share of the dataset trains the model in a typical fold.

How the result changes with K-Folds

K-FoldsFold SizeTraining Size per FoldTest Set Size
2.54,0006,0002,000
3.752,6677,3332,000
7.51,3338,6672,000
137699,2312,000

What each input means

Dataset Size
Total number of samples
K-Folds
Number of folds for cross-validation
Test Set Size
Percentage for test set
Validation Set Size
Percentage for validation set (from training set)

How this is calculated

Formula

Fold Size = Dataset Size / K

Worked example, using the default values

  1. Identify Input Parameters
    4 parameters
    Dataset Size = 10000, K-Folds = 5, Test Set Size = 20, Validation Set Size = 10 = 4 input(s) provided
  2. Calculate Fold Size
    Fold Size
    2000 = 2000
  3. Calculate Training Size per Fold
    Training Size per Fold
    8000 = 8000
  4. Calculate Test Set Size
    Test Set Size
    2000 = 2000
  5. Calculate Validation Set Size
    Validation Set Size
    800 = 800
  6. Calculate Final Training Size
    Final Training Size
    7200 = 7200

Engine last updated . Checked against 2 independently-derived tests — how we verify calculators. Built by Paul Gunder, a software engineer, not a licensed financial, medical, or legal professional.

Frequently Asked Questions

Why does K-Folds affect Fold Size and Training Size per Fold in opposite directions?

Fold Size is Dataset Size divided by K-Folds, so more folds means each individual fold is smaller. Training Size per Fold is what's left over after removing one fold (Dataset Size minus Fold Size), so as Fold Size shrinks with more folds, Training Size per Fold grows correspondingly closer to the full dataset -- the two outputs move in opposite directions because they're two halves of the same split.

What does Data Utilization actually measure?

Data Utilization is the percentage of your full dataset that ends up in the training set for any single fold -- Training Size per Fold divided by Dataset Size. It rises toward (but never reaches) 100% as K-Folds increases, since a larger K means each fold holds out a smaller slice for validation and trains on a correspondingly larger share of the remaining data.

Why is there both a Test Set and a Validation Set?

The Test Set is held out entirely and used only once, at the very end, to report a final unbiased performance figure -- it should never influence any training or tuning decision. The Validation Set is carved out of what remains after the test split and is used during development (for example, to tune hyperparameters or decide when to stop training), which keeps those tuning decisions from leaking information into the untouched Test Set.

Should I use more folds or fewer?

More folds (like 10) train each model on a larger share of the data per fold, which this calculator's Data Utilization figure reflects directly, but it also means more separate training runs and a smaller, noisier held-out fold each time. Fewer folds (like 3) train faster and average over larger held-out chunks but use less data per training run. 5 and 10 are the most common defaults in practice as a balance between the two, though this calculator doesn't pick one for you -- it just shows you the resulting split sizes.

The questions that sit next to this one — chosen by subject, including calculators filed under a different category.

More in Technology & Computing.