How to Check Your Manuscript for Plagiarism Before You Submit
The similarity score is the most misread number in publishing. It counts your quotes, your reference list, and your own prior methods as heavily as real overlap, and most authors cannot run the journal's checker before they submit.
While you wait
Waiting on a decision? Get your next move ready.
The wait is out of your hands; the next move isn't. Scan your next manuscript free, or run this paper through the scan to see what reviewers typically push back on, so the revision response is ready when the decision lands.
How to use this page well
These pages work best when they behave like tools, not essays. Use the quick structure first, then apply it to the exact journal and manuscript situation.
Question | What to do |
|---|---|
Use this page for | Getting the structure, tone, and decision logic right before you send anything out. |
Most important move | Make the reviewer-facing or editor-facing ask obvious early rather than burying it in prose. |
Common mistake | Turning a practical page into a long explanation instead of a working template or checklist. |
Next step | Use the page as a tool, then adjust it to the exact manuscript and journal situation. |
Quick answer: There is no universal similarity score that is too high. Most journals start asking questions above 15 to 20 percent, but the number alone is close to meaningless, because it counts your quotes, your reference list, and your own prior methods as heavily as real overlap. Worse, most authors cannot run the journal's checker before they submit. So the useful question is not what number is too high. It is what your number is made of.
Why the percentage is the wrong thing to look at
The similarity score is a single number standing in for a page of nuance, and that is exactly why it misleads. A 25 percent score can be completely clean. A 12 percent score can sink a paper.
Here is what actually feeds the number. Direct quotes you cited correctly. Your reference list, which matches every paper you cite. Standard methods sentences that thousands of papers share word for word. Common academic phrasing. And, buried among all of that, the thing the journal actually cares about: unattributed overlap with someone else's text, or with your own earlier papers. The checker weights all of it the same. A human has to separate the harmless matches from the real ones, and that human is usually a reviewer or an editor who has already formed an impression of your paper.
So a score is a prompt to read, not a grade. The mistake is treating it as a grade.
What the bands actually mean
Different journals set different thresholds, so treat this as a map, not a rule. The numbers below are the common reading of a Turnitin or iThenticate similarity index.
Similarity | Common reading | What to do |
|---|---|---|
0 to 14% | Generally fine | Skim the matches; usually nothing to fix |
15 to 25% | Review | Open the report, resolve any single large match |
26 to 50% | Concerning | Real work needed; find the source of the overlap |
50%+ | Integrity review likely | Do not submit until you understand every match |
Source: Turnitin similarity-report guidance and common journal editorial policy (accessed August 2026).
The band matters less than the shape. One 18 percent match to a single paper is a bigger problem than fifty matches of one sentence each that add up to 30 percent.
The three common patterns behind almost every flag
This is the part you only see from the inside, and it changes what you look for. In our pre-submission review work with manuscripts, a similarity flag almost never comes from misconduct. It comes from three quiet, honest sources, and each is a specific failure pattern worth checking your own draft against.
The first is self-overlap. You reuse your own methods section, or a paragraph of your introduction, from an earlier paper. It is your writing, so it feels safe, but the checker flags it as text recycling, and journals take that seriously. The fix is to cite your earlier work where you reuse it and to rewrite the reused passage rather than paste it.
The second is boilerplate methods. Whole fields describe a standard assay or instrument in nearly identical language, because there is only one accurate way to say it. This inflates the raw number and is mostly harmless, but a large single match to one recent paper hiding inside the boilerplate is not, so it still has to be read.
The third is the paraphrase that stayed too close. You rewrote a source in your own words, but the sentence structure and half the vocabulary survived. This is the one that actually gets papers into trouble, because it reads as an attempt to disguise borrowing rather than to explain an idea. What we see is that this pattern hides inside a perfectly ordinary-looking 12 or 15 percent score, which is why the number alone will not catch it.
All three are fixable in an afternoon if you find them before you submit. None of them are findable after, when the report is already in front of an editor.
How to actually check, and the access problem
Here is the frustrating part. The tool the journal uses, iThenticate, is sold to publishers and institutions, not to individual authors. The journal runs it at submission, which is after you have already committed. There is no free, universal way to run your manuscript against the same database the journal will use. That is the gap the whole problem lives in.
Method | What it catches | Effort | Cost |
|---|---|---|---|
Read your own draft for reuse | Self-overlap and close paraphrase you already suspect | Low | Free |
Free web plagiarism checkers | Surface matches against the open web only, not the journal databases | Low | Free, unreliable |
Paid per-check iThenticate services | Close to the journal's own check | Medium | Paid per document |
A pre-submission review that flags similarity-risk writing | The reuse patterns above, so you can fix them before you run the tool | Low | Free to paid |
The free web checkers are the trap. They match against the open internet, not the subscription literature the journal actually searches, so a clean result from a free tool tells you very little. What actually happens is that authors run a free checker, see a green result, submit, and get flagged anyway. The only tool that mirrors the journal's own check is iThenticate, and that is the one sold per document rather than given away.
A pre-submission review does not replace that iThenticate check, and it should not claim to. What our pre-submission review does is read the manuscript for the three reuse patterns above, the copied methods boilerplate, the undisclosed self-overlap, the unquoted lifted definition, so you can fix the writing that drives a high score before you pay for the tool that measures it.
Readiness check
While you wait, scan your next manuscript.
The scan takes about 1-2 minutes. Use the result to decide whether to revise before the decision comes back.
When to worry, and when it is fine
Use these two cuts to decide fast, without staring at the percentage.
Fix it before you submit if: a single match to one source is large, regardless of the total; the overlap is a paraphrase that kept the original structure; the flagged text is someone else's result or argument stated as if it were yours; or you reused your own prior text without citing it.
It is probably fine if: the matches are your correctly cited quotes and your reference list; the overlap is standard methods or instrument language shared across the field; or the total looks high but decomposes into many one-sentence matches with no single meaningful source.
The one move that is never fine is paraphrasing mechanically to push the number down. Editors and reviewers recognize evasive rewriting on sight, and it usually makes the paper read worse while leaving the real problem, the uncited idea, exactly where it was.
What we see across manuscripts
The score is a symptom, not the disease. Editors routinely screen for similarity at submission, but what they are really reading for is intent: did this author try to pass off borrowed text, or did the number just pile up from honest reuse and citations? A clean, well-cited paper with a 22 percent score clears that read. A paper with a 12 percent score built from one disguised paragraph does not. Check the shape of your matches, fix the real overlap, and cite what you borrowed, including from yourself.
Behind this page sit the current journal integrity and text-recycling policies we checked, all listed below, read against the reuse patterns that most often drive a high score before you submit.
Your similarity-check checklist
Before you submit, walk your draft through these:
- Read your methods and introduction for text reused from your own earlier papers, and cite that work wherever you kept the wording.
- Put every direct quotation in quotation marks with a citation, so the checker attributes it rather than flagging it.
- Rewrite any genuinely borrowed passage in your own framing instead of paraphrasing it word by word.
- Confirm your reference list and quotations can be excluded from the count, if the tool your journal uses allows it.
- Find any single source that supplies more than a couple of percent of your text, and reduce or properly attribute it.
- Treat a clean result from a free web checker as inconclusive, because it does not search the subscription literature the journal does.
Fix the overlap these surface, not the number. A score you understand is one you can explain to an editor in a sentence.
Frequently asked questions
There is no universal threshold, but as a rough guide most journals start asking questions above 15 to 20 percent, treat 26 to 50 percent as concerning, and open an integrity review at 50 percent or more. The number matters far less than what it is made of: a 25 percent score built from quotes and your reference list is fine, while a 12 percent score built from one uncited paraphrased paragraph is not.
Usually not directly. iThenticate is sold to journals and institutions, and the journal runs it at submission, after you have already sent the paper. Some paid per-check services exist, but there is no free, universal way for an individual author to pre-check against the same database the journal uses. That gap is exactly why the problem persists.
No. A high score most often comes from properly cited direct quotes, your reference list, standard methods language, common academic phrasing, or overlap with your own earlier papers. A low score does not prove the opposite either. The score is a starting point for a human to read the flagged matches, not a verdict.
Do not paraphrase mechanically to dodge the checker; that reads as evasive and often makes the writing worse. Instead, quote and cite direct borrowings properly, rewrite genuinely overlapping passages in your own framing, cite your own prior work where you reuse methods, and exclude the reference list and quotes from the count if the tool allows. Fix the real overlap, not the number.
Yes. Text recycling from your own published papers, especially in methods and introduction sections, is one of the most common sources of a flag and one journals take seriously. Cite your earlier work when you reuse it, and rewrite reused passages rather than pasting them.
Sources
- Turnitin, Understanding the similarity score: https://guides.turnitin.com/hc/en-us/articles/23435833938701-Understanding-the-similarity-score
- iThenticate for publishers and researchers: https://www.ithenticate.com/
- COPE guidance on text recycling and plagiarism: https://publicationethics.org/guidance/endorsed-guidance/text-recycling-guidelines-editors
Final step
Done interpreting the status? Put the wait to work.
The decision will arrive on the journal's clock. What you control is what's next: scan your next manuscript free, or run this paper through the scan so the likely reviewer pushback is mapped before the revision request lands.
Free scan, no card needed.
Anthropic Privacy Partner. Your manuscript is never used to train any model.