Skip to main content

How to Do a Sample Size Calculation: A Worked Example and What Reviewers Need to See

A sample size calculation needs four inputs: the smallest difference worth detecting, the variability, alpha and power. Here is the arithmetic for two groups, worked through, and the one paragraph that lets a reviewer check it.

Editorial processThe Manusights editorial team researches and maintains these guides using source review, field-specific analysis, and our documented editorial process.How we work

Readiness scan

Find out if this manuscript is ready to submit.

Run the Free Readiness Scan before you submit. Catch the issues editors reject on first read.

Check my manuscriptPrivate API processing. Your manuscript is not used to train models.See example reports
Working map

How to use this page well

These pages work best when they behave like tools, not essays. Use the quick structure first, then apply it to the exact journal and manuscript situation.

Question
What to do
Use this page for
Building a point-by-point response that is easy for reviewers and editors to trust.
Start with
State the reviewer concern clearly, then pair each response with the exact evidence or revision.
Common mistake
Sounding defensive or abstract instead of specific about what changed.
Best next step
Turn the response into a visible checklist or matrix before you finalize the letter.

Quick answer: For two groups, a sample size calculation takes four inputs: the smallest difference worth detecting, the standard deviation (or the control-group rate), alpha, and power. For a continuous outcome, n per group = 2 × (1.96 + 0.84)² × SD² / difference². Detecting a 5-unit difference with an SD of 10 at a two-sided alpha of 0.05 and 80% power gives 63 per group; then divide by the proportion you expect to retain, so 15% dropout means recruiting 75 per group.

Most rejections over sample size are not about the arithmetic. They come from a missing paragraph: reviewers cannot check a number when the inputs are not reported.

Check my methods section before submitting once the draft exists. The scan is free, card-free and quick, and reports the one problem in your draft it rates most serious.

Evidence basis: The formulas and the constants 1.96 and 0.842 follow Das, Mitra and Mandal, "Sample size calculation: Basic principles" (Indian Journal of Anaesthesia, 2016). The reporting items come from the CONSORT 2010 checklist and the ARRIVE guidelines 2.0. We worked the example numbers ourselves and read all three sources on September 23, 2026. These are normal-approximation formulas for two independent groups; clustered, repeated-measures, survival and non-inferiority designs need different methods, and nothing here replaces a statistician for those. The red flags below are Manusights judgment about how reviewers read these paragraphs, not a rule any journal applies mechanically.

The five steps

  1. Name one primary outcome and its type. Continuous (blood pressure, tumor volume, test score) and binary (event yes or no) outcomes use different formulas. If the paper has three co-primary outcomes, you need a calculation, or a correction, for each.
  1. Choose the smallest difference that would matter. This is a clinical or scientific judgment, not a guess at what you will find. A difference picked because it makes the numbers affordable is the first thing an experienced reviewer suspects.
  1. Estimate the variability. For a continuous outcome you need the standard deviation; for a binary one, the expected rate in the control group. Take it from a pilot study or from published work in a similar population, and cite the source.
  1. Set alpha and power. A two-sided alpha of 0.05 (z = 1.96) and 80% power (z = 0.842) are the common defaults; 90% power uses z = 1.282. One-sided tests need a stated reason.
  1. Calculate, round up, and inflate for loss. Round every fraction of a participant up, then divide by the expected retention rate.

A worked example: comparing two means

Suppose the primary outcome is a symptom score, the smallest difference that would change practice is 5 points, and a pilot study found an SD of 10.

n per group = 2 × (1.96 + 0.842)² × 10² / 5²

= 2 × 7.851 × 100 / 25

= 62.8, so 63 per group, 126 in total.

Now change one input at a time and watch what moves:

Change from the base case
Per group
Total
Base: difference 5, SD 10, alpha 0.05, 80% power
63
126
90% power instead of 80%
85
170
Difference halved to 2.5
252
504
Base, with 15% expected dropout (63 / 0.85)
75
150

Source: our arithmetic using the two-means formula in Das et al., 2016; z = 1.96 for alpha, 0.842 for 80% power, 1.282 for 90% power

The row that surprises people is the third. The difference is squared in the denominator, so halving it roughly quadruples the sample. That is why the choice in step 2 matters more than any software setting.

The Das paper works the same formula on a real design: detecting a 15 mmHg difference in mean arterial pressure with an SD of 20 gives 27.9, so 28 per group.

A worked example: comparing two proportions

For a binary outcome the formula is:

n per group = (1.96 + 0.842)² × [p1(1 − p1) + p2(1 − p2)] / (p1 − p2)²

If 40% of control patients have the event and you want to detect a fall to 25%:

n = 7.851 × (0.40 × 0.60 + 0.25 × 0.75) / 0.15²

= 7.851 × 0.4275 / 0.0225

= 149.1, so 150 per group

Software that uses exact or continuity-corrected methods will give a slightly larger number. Report which one you used.

The sentence reviewers expect in your methods

The CONSORT 2010 checklist asks randomized trials to report "how sample size was determined" (item 7a). ARRIVE 2.0 places sample size in its Essential 10 for animal studies, and its authors cite evidence that sample size justification was reported in fewer than 10% of the animal publications sampled. The gap is the paragraph, not the maths.

A paragraph a reviewer can check looks like this:

We calculated that 63 participants per group would give 80% power to detect a 5-point difference in the primary outcome (symptom score at 12 weeks), assuming an SD of 10 from our pilot study (reference), with a two-sided alpha of 0.05. Allowing for 15% loss to follow-up, we planned to enroll 75 per group.

Every number in it can be traced, and anyone can redo the arithmetic in a minute.

Red flags reviewers look for

The difference comes from nowhere. "A clinically meaningful difference" without a number, or a number with no reason, suggests the target was chosen to fit the budget.

The SD is borrowed from a different population. A pilot in healthy volunteers used to size a trial in patients will usually underestimate the variability.

The calculated number and the enrolled number disagree without comment. If you recruited 60 per group against a target of 75, say why, and report the confidence interval for the primary outcome so readers can judge what the data can exclude.

Power calculated after the fact from the observed effect. It is a restatement of the p-value and adds nothing. If there was no prospective calculation, say that plainly.

One calculation for a paper with several primary outcomes. Either justify one primary outcome or account for the multiplicity.

What Manusights checks and what it cannot

The free readiness scan gives a submission-risk signal and the single biggest blocker in the draft; if that blocker is an unjustified sample, it tells you. A $39 Full Review adds a section-by-section read with methods and reviewer-risk feedback, and the methods section is where this paragraph lives. Neither can tell you the right difference to power for, because that depends on your field and your patients. For designs beyond two independent groups, talk to a statistician before recruitment, not after the reviews come back.

If your statistics are the part you are least sure of, the guide to statistical review before journal submission covers the other checks worth running.

Readiness check

Run the scan to see how your manuscript scores on these criteria.

See score, top issues, and what to fix before you submit.

Check my manuscriptPrivate API processing. Your manuscript is not used to train models.See example reports

Think twice before you

  • power the study for the effect you hope to see rather than the smallest one that would matter.
  • add the dropout percentage on top instead of dividing by the retention rate; 63 × 1.15 is 73, not 75.
  • quote an SD without a source. If the pilot was tiny, say so and consider a more conservative value.
  • drop the calculation from the paper because it was in the protocol. Reviewers read the manuscript, not your ethics file.

Frequently asked questions

Per group, n = 2 x (z for alpha + z for power) squared x SD squared / difference squared. With a two-sided alpha of 0.05 (z = 1.96) and 80% power (z = 0.84), detecting a 5-unit difference when the SD is 10 needs 2 x 2.80 squared x 100 / 25, about 62.8, so 63 per group before any allowance for dropout.

Four: the primary outcome and its type (continuous or binary), the smallest difference that would matter, an estimate of variability or of the control-group proportion, and your alpha and power. Most studies use a two-sided alpha of 0.05 and 80% or 90% power. The variability usually comes from a pilot study or published work.

Divide the calculated number by the proportion you expect to keep. If the calculation gives 63 per group and you expect 15% loss, 63 / 0.85 = 74.1, so recruit 75 per group. Adding 15% on top (63 x 1.15 = 73) undershoots.

It roughly quadruples, because the difference is squared in the denominator. In the worked example, 5 units needs 63 per group and 2.5 units needs 252.

Enough for a reviewer to redo the arithmetic: the primary outcome, the difference to be detected and why it matters, the variability or baseline rate and its source, alpha and whether the test is one- or two-sided, power, the resulting number per group, and any inflation for dropout. The CONSORT 2010 checklist asks trials to report how sample size was determined.

Calculating power from the effect you observed adds nothing, because it is a restatement of the p-value. If there was no prospective calculation, say so plainly, and report the confidence interval so readers can see which effects the data can and cannot rule out.

References

Sources

  1. Das S, Mitra K, Mandal M. Sample size calculation: Basic principles. Indian J Anaesth. 2016;60(9):652-656
  2. Schulz KF, Altman DG, Moher D. CONSORT 2010 Statement: updated guidelines for reporting parallel group randomised trials
  3. Percie du Sert N, et al. The ARRIVE guidelines 2.0: Updated guidelines for reporting animal research. PLOS Biology. 2020

Final step

Find out if this manuscript is ready to submit.

Run the Free Readiness Scan. See score, top issues, and journal-fit signals before you submit.

Check my manuscript

Private API processing. Your manuscript is not used to train models.

See example reports

Put the guidance to work

Move from draft advice to a practical manuscript pass.

Choose the reporting framework first, then use the final checklist when the manuscript and submission files are ready to travel together.

Internal navigation

Where to go next