Statistical Power & Sample Size calculator

Calculate the sample size required to detect an effect or determine the statistical power of your completed study online.

What it does

Works out the sample size needed to detect an effect of a given size, at a given alpha and power — or the smallest effect a planned sample could detect.

When to use it

Before collecting data. This is a design tool.

Cautions

Alternatives

How to read the output

n per group
The result is the number needed in EACH group, so the total is twice it. Equal allocation is assumed. Check what the reported n counts before quoting it. This tab labels the unit for each family, and reading a per-group figure as a total is the error that halves a study. Unequal allocation costs power: a 2:1 design needs more people in total than a 1:1 design for the same power.
Effect size d
The difference in means divided by the pooled SD. Cohen's conventions are 0.2 small, 0.5 medium, 0.8 large. d needs a standard deviation you do not yet have. Getting it from a small pilot is unreliable — a pilot's SD is itself a noisy estimate, and underestimating it makes the study look cheaper than it is.
The effect size you assumed
Every number here follows deterministically from it. Power analysis does not discover the effect — it works out what a study would need IF the effect were the size you nominated. Choosing the effect size to make a feasible sample size come out is the commonest abuse of this tool, and it is invisible in the output. State where the number came from — a pilot, a published estimate, or the smallest difference that would change a decision — and prefer the last of those, since it is the only one that makes the study's result meaningful either way.
Significance level (α) and target power
α is the false-positive rate you accept; power is the chance of detecting the assumed effect if it is real. 0.05 and 0.80 are conventions, not laws. 80% power means a one-in-five chance of missing a real effect of exactly the size you assumed — worse for anything smaller. For a study that will not be repeated, 90% is often the more defensible choice.
The power curve
Power against sample size at your assumed effect and α. The curve is steep at small n and flattens as it approaches 1. The flattening is the practical message: past a point, extra participants buy very little power. It also shows how fragile the plan is — if the curve is still climbing steeply at your chosen n, a slightly smaller true effect costs a great deal of power.
What this is not
A-priori planning. This is a-priori power for an effect you assume. Post-hoc power — recomputed from the effect you observed — is a different thing and is not worth reporting: it is a one-to-one function of the p-value, so it always says a non-significant result was underpowered and adds nothing the p-value did not.

How this calculator is validated · Which statistical test should I use?