Skip to main content

The Statistical Mistakes That Get Papers Rejected (and How to Catch Them)

Reviewers rarely reject a paper for a wrong t-test. They reject it for pseudoreplication, uncorrected multiple comparisons, and reading a non-significant result as proof of no effect, and those three hide in papers that look statistically fine.

Author contextFounder, ManusightsView profile

Next step

Choose the next useful decision step first.

Use the guide or checklist that matches this page's intent before you ask for a manuscript-level diagnostic.

Open Journal Fit ChecklistAnthropic Privacy Partner. Your manuscript is never used to train any model.Run Free Readiness Scan

Quick answer: The statistical mistakes that get papers rejected are rarely a wrong test. They are design and interpretation errors: pseudoreplication (counting non-independent measurements as independent), uncorrected multiple comparisons (running many tests and reporting the wins), and reading a non-significant result as proof of no effect. Each can sink a careful study, and each hides in a paper whose individual tests look correct. Catch them before submission, because they are cheap to fix at your desk and expensive to defend in a review.

The reviewer is not checking your arithmetic. They are checking whether your design earns the claim.

Why the test is never the real problem

Begin with why, because the whole checklist rests on it. By the time a reviewer opens your paper, the choice between a t-test and a Mann-Whitney is settled and rarely fatal on its own. What determines the outcome is whether the analysis matches the design: whether your units are independent, whether your significance survives the number of tests you ran, and whether your interpretation goes further than the data allow.

Those are structural questions, and they are the ones a good reviewer asks first. A paper can have every p-value computed correctly and still fail all three, because the numbers are right and the reasoning behind them is not. That is why a statistical rejection so often surprises the author. The stats "worked." The design did not.

The published literature agrees on which errors recur. A widely-cited eLife guide catalogs ten common statistical mistakes reviewers watch for, and the three below are the ones we see end papers most often.

The three common patterns that sink a study

You learn this from reading a great many analyses. In our pre-submission review work with manuscripts, the statistical problems that trigger rejection almost always fall into one of three shapes, and you can test your own study against each one before anyone else does.

The first is pseudoreplication. You take fifty measurements from three animals and analyze them as fifty independent data points. What we see is that this is the most common inflator of false significance in bench and field work, because it turns a study with an n of three into one that looks like an n of fifty. The fix is to identify the true independent unit, usually the animal or the subject, and analyze at that level, often with a mixed-effects model.

The second is the uncorrected comparison battery. You run twenty tests, three come back significant, and you report those three. Without a correction for multiple comparisons, that is roughly what chance alone would hand you, and a reviewer knows the arithmetic. In practice it means you either pre-register your primary comparison, correct for the number of tests, or frame the exploratory ones honestly as hypothesis-generating.

The third is the absence-of-evidence slip. Your p-value comes back at 0.08, and you write that the groups did not differ. A non-significant result is not proof of no effect; it is a failure to detect one, which could mean the effect is absent or that you were underpowered. To claim equivalence you need an equivalence test or a stated power analysis, not a p-value that missed the threshold.

None of these three is exotic. All three routinely appear in papers with otherwise clean methods, which is exactly why they are worth a deliberate pass.

A pre-submission statistics check

Run this before you submit. It is not a substitute for a statistician, but it catches the three patterns above and most of their relatives.

Question to ask
What a bad answer looks like
What is the independent unit of replication?
The measurement, not the animal or subject
How many tests did you run in total?
Many, with no correction reported
Does any conclusion rest on a non-significant result?
"No difference" from p above 0.05, no equivalence test
Are effect sizes and confidence intervals reported?
Only p-values, no magnitude
Was the primary hypothesis set before the data?
Chosen after seeing which test worked

Manusights pre-submission statistical checklist; see the linked eLife guide for the full ten.

The check that matters most is the first one. If you cannot say cleanly what your independent unit is, stop and resolve that before anything else, because it changes every number downstream.

When it is a real problem, and when it is not

Fix it before you submit if: your sample size counts non-independent measurements; you ran many tests and reported only the significant ones; a conclusion depends on a non-significant p-value; or you report significance without effect sizes. These are the ones reviewers reject for, and they are cheap to fix now.

It is probably fine if: your analysis is at the correct unit of replication, your comparisons are corrected or pre-specified, your absence claims use equivalence testing, and you report magnitude alongside significance. A defensible design survives a skeptical statistical reviewer even when the result is modest.

The one move that never works is reaching for a fancier test to rescue a flawed design. The test is downstream. The design is the thing being judged.

Readiness check

Run the scan while the topic is in front of you.

See score, top issues, and journal-fit signals before you submit.

Get free manuscript previewAnthropic Privacy Partner. Your manuscript is never used to train any model.See example reports

What we see across manuscripts

Surviving statistical review is not about the most sophisticated model. It is about a clear unit of replication, honest comparisons, and claims that stop where the data stop. Turn to this page when your conclusions rest on the statistics and you want to catch the design problems before you submit. The statistics literature we reviewed, including the eLife guide, is cited below. If you want a read that checks the analysis against the design, our pre-submission review flags pseudoreplication and uncorrected comparisons alongside the rest, and the journal directory helps you match the finished paper to a venue.

What to do before you submit

The pattern in every statistical rejection is the same: the fix was cheap before submission and expensive after. Working the checklist below at your desk costs an afternoon. Finding the same problem in a reviewer's report costs a revision cycle, and sometimes the paper. So treat this as the last thing you do before the manuscript goes out, after the writing is done but while the analysis is still yours to change. Run the checks in order, because the first one changes every number below it.

  • Name your independent unit of replication out loud, and confirm your sample size counts that unit rather than repeated measurements taken from it.
  • Count every statistical test you ran, including the ones that never made the paper, and either correct for the total or pre-register the primary comparison.
  • Find any conclusion that rests on a non-significant result, and either soften it to "we did not detect an effect" or back it with an equivalence test.
  • Report an effect size and a confidence interval next to every p-value, so a reader sees magnitude and not significance alone.
  • Confirm your primary hypothesis was set before you saw the data, not chosen after you saw which test happened to work.
  • Re-read your strongest claim and check that it stops where the data stop, with no reach into "proves" or "establishes" that the design cannot support.
  • If the statistics carry the paper and any of these feels shaky, have a statistician or a pre-submission review look before a reviewer does.

None of these is exotic, and none needs new data. Each is a way of making the analysis say only what it can actually defend.

Frequently asked questions

Three recur far more than a wrong test choice: pseudoreplication, where non-independent measurements are counted as independent and inflate the sample size; uncorrected multiple comparisons, where running many tests without adjustment manufactures false positives; and interpreting a non-significant result as proof of no effect. Each can sink an otherwise careful study, and each hides in a paper that looks statistically fine at a glance.

Counting measurements that are not statistically independent as if they were. Taking fifty readings from three animals and analyzing them as n equals fifty, rather than n equals three, is the classic case. It inflates your effective sample size and your significance, and a careful reviewer catches it immediately by asking what the true independent unit of replication is.

If you run many tests and report the ones that reached significance, yes. Without correction, testing twenty independent hypotheses at the 0.05 level gives you roughly a two-in-three chance of at least one false positive. Reviewers know this arithmetic, so an unadjusted battery of tests reads as p-hacking even when it was not intended that way.

No, and treating it as such is one of the most common errors reviewers flag. A non-significant result means you did not detect an effect, which could be because there is none or because your study was underpowered to find it. To claim equivalence or absence, you need an equivalence test or a power analysis, not a p-value above 0.05.

For any paper where the statistics carry the conclusion, it is worth it. Most rejections on statistical grounds are for design and interpretation problems that are cheap to fix before submission and expensive to argue about after. A pre-submission check that flags pseudoreplication or an uncorrected comparison saves a review cycle.

References

Sources

  1. Makin and Orban de Xivry, Ten common statistical mistakes to watch out for (eLife, 2019): https://elifesciences.org/articles/48175
  2. Lazic, The problem of pseudoreplication in neuroscientific studies: https://bmcneurosci.biomedcentral.com/articles/10.1186/1471-2202-11-5
  3. Nature, Statistics for Biologists collection: https://www.nature.com/collections/qghhqm/

Before you upload

Choose the next useful decision step first.

Move from this article into the next decision-support step. The scan works best once the journal and submission plan are clearer.

Use the scan once the manuscript and target journal are concrete enough to evaluate.

Anthropic Privacy Partner. Your manuscript is never used to train any model.

Internal navigation

Where to go next