Reading Time: 13 minutes

Similarity screening has become a common part of academic journal workflows. Before a manuscript reaches peer reviewers, editorial teams may compare its text with published articles, websites, repositories, conference papers, theses, and other available sources.

The resulting report often includes one highly visible number: the overall similarity score. Because the number appears precise, it can seem like an objective measure of originality. A lower score may appear safe, while a higher one may appear suspicious.

This interpretation is misleading. A similarity percentage shows how much text overlaps with material found by the system under particular settings. It does not establish whether the overlap is legitimate, properly cited, academically significant, or evidence of misconduct.

For journal screening, the explanation behind the score matters more than the score itself. Editors need to know which passages match, which sources produced those matches, where the text appears, and how it functions within the manuscript.

What a Similarity Score Actually Measures

A similarity score estimates the proportion of submitted text that resembles material available to the comparison system. The system identifies matching strings or sufficiently similar sequences and links them with possible sources.

The score may include direct quotations, bibliographic entries, standard terminology, institutional names, methodological descriptions, and genuinely problematic copying.

Its value depends on several technical factors. These include the databases searched, the minimum length of a match, whether quotations are excluded, whether references are included, and how overlapping sources are counted.

The same manuscript may therefore receive different percentages under different configurations or when checked against different collections.

The score is useful as a screening signal. It is not a complete interpretation of the manuscript’s originality.

Similarity Is Not the Same as Plagiarism

Textual similarity describes a relationship between passages. Plagiarism is a judgment about unattributed or misleading use of another person’s work.

A correctly quoted sentence can produce a clear match without creating an integrity problem. A standard methodological expression may resemble hundreds of published papers because researchers describe the same established procedure.

At the same time, a manuscript can contain problematic borrowing with a relatively low overall score. A short copied conclusion, unique interpretation, or central argument may represent only a small percentage of the complete article.

Editors should therefore avoid treating similarity as a substitute for plagiarism assessment. The tool locates possible overlap. A qualified person determines what that overlap means.

What the Raw Percentage Cannot Explain

An overall score removes the details needed for editorial judgment. It does not show whether the matching text comes from one source or fifty unrelated sources.

It does not distinguish between a quoted definition and an uncited copied conclusion. It does not explain whether the overlap appears in the methods section, abstract, literature review, results, or discussion.

The percentage also cannot determine whether the author has cited the source accurately. A citation may appear near the matched passage but fail to indicate how much wording was reproduced.

Without the report’s source links and highlighted passages, the editor sees a number without its academic context.

Why Universal Similarity Thresholds Fail

Some editorial workflows use a fixed percentage to divide acceptable and unacceptable manuscripts. This simplifies triage, but it can produce inconsistent and unfair decisions.

Disciplines differ in how they use language. Experimental research may require standardized descriptions of instruments, procedures, or reporting frameworks. Legal, historical, and theoretical writing may contain longer quotations that are essential to the argument.

Article types also differ. A review article naturally contains extensive discussion of previous scholarship. A methods paper may repeat technical terminology. A short editorial can receive a high percentage from only a few matched sentences because the total text is small.

A percentage that appears acceptable in one manuscript may conceal a serious issue in another. The meaning depends on distribution, source, section, and attribution.

One Large Match and Many Small Matches Are Different

Two manuscripts can receive the same overall score while presenting very different editorial risks.

In the first manuscript, the similarity may consist of dozens of short phrases spread across references, standard descriptions, and commonly used terminology.

In the second, most of the similarity may come from several consecutive paragraphs copied from one article.

The numerical total may be identical, but the second pattern deserves much closer examination.

Report pattern Possible interpretation Editorial response
Many short matches across many sources Common language, terminology, or fragmented overlap Review representative matches and critical sections
One long match from one source Possible direct copying or insufficient quotation Examine the complete passage and attribution carefully
High overlap in the reference list Bibliographic similarity rather than reused argument Exclude references and recalculate where appropriate
Repeated methods language Standard procedure or reused description Check necessity, citation, and journal policy
Overlap in results or conclusions Possible reuse of original findings or interpretation Prioritize detailed editorial review
Low score with one distinctive copied passage Serious local issue hidden by the overall percentage Evaluate the importance of the passage, not only its length

Location Changes the Meaning of Similarity

The section containing the match often matters more than the size of the overall score.

Some overlap in a methods section may be understandable when authors use a standardized procedure. Even then, extensive reuse may require citation, quotation, or rewriting depending on the journal’s rules.

Similarity in the abstract deserves closer attention because the abstract summarizes the manuscript’s central contribution. Reused wording may indicate that the new paper is not clearly distinguished from earlier work.

Matches in the results, discussion, and conclusion are particularly sensitive. These sections should present the study’s findings, interpretation, and contribution.

A raw percentage treats every word as equal. Editorial judgment recognizes that different passages perform different scholarly functions.

The Source of the Match Matters

An explainable report identifies the likely origin of each match. This information can change the editorial interpretation completely.

A passage matching an article by another research group raises different questions from a passage matching the submitting author’s dissertation.

A match with a public reporting guideline may reflect required terminology. A match with a commercial website may suggest inappropriate reuse or may simply involve a standard product description.

Editors should examine whether the source is scholarly, public, unpublished, retracted, translated, or connected with the authors.

The source relationship also helps identify duplicate publication, manuscript recycling, or incomplete disclosure of previous dissemination.

Acceptable Similarity in Academic Manuscripts

Not all matching text requires correction. Academic writing depends on shared terminology, formal names, and references to previous work.

Properly marked quotations will normally match their original sources. Titles of publications, laws, research instruments, organizations, and standard classifications may also be identical.

Bibliographic entries frequently produce extensive technical matches. This overlap usually provides little evidence about the originality of the manuscript itself.

Some journals and disciplines permit limited reuse of methods language when the procedure cannot be described accurately in many different ways. The source may still need to be cited.

Explainability allows editors to separate these legitimate matches from passages that misrepresent authorship.

Problematic Similarity Patterns

Concern increases when a manuscript reproduces distinctive language without clear attribution. Long continuous passages, unusual phrasing, or identical analytical sequences require attention.

Close paraphrasing can also be problematic. Changing a few words while preserving the source’s sentence structure and argument may conceal dependence rather than demonstrate original writing.

Another warning pattern is repeated borrowing from several sections of one source. Each individual match may appear short, but together they may reconstruct a substantial part of the original work.

Similarity in the interpretation of findings is more significant than the repetition of generic background statements. Editors should evaluate intellectual importance, not only word count.

Explainability Supports Proportionate Decisions

Journal screening does not always produce a simple accept-or-reject outcome. Different problems require different responses.

A missing quotation mark may be corrected before review. Poor paraphrasing may require revision and clearer citation. Extensive unattributed copying may justify rejection or referral under the journal’s research integrity process.

A raw score cannot distinguish these situations. An explainable report gives editors enough detail to choose a proportionate action.

This prevents minor technical issues from receiving the same treatment as substantial misrepresentation.

Similarity Reports Should Show Context

A useful report highlights the matched passage within the manuscript and displays the corresponding source text.

Editors need to see the sentences before and after the match. Context reveals whether the author introduced a quotation, cited the source, or used the passage as part of a larger argument.

The report should also show whether several sources contain the same wording. The first listed source is not necessarily the original source.

Direct access to the source allows the editor to compare structure, meaning, and attribution rather than relying on colored highlights alone.

Source-Level Percentages Are More Useful

An overall score combines all detected matches. Source-level percentages show how much text is associated with each individual source.

This helps editors identify concentrated borrowing. A manuscript with a moderate overall score may depend heavily on one earlier publication.

Source-level information also reveals whether the report is inflated by repeated copies of the same material across repositories and websites.

Editors should still inspect the passages. A low percentage from one source can remain significant when it contains a distinctive conclusion or original analytical claim.

Settings Must Be Visible

Explainability includes transparency about how the report was produced.

The editor should know whether the system included references, direct quotations, short matches, and small sources. These settings can substantially change the final number.

If one editor excludes the bibliography while another includes it, their thresholds cannot be compared fairly.

Journals should define consistent screening settings and document any manual changes made during review.

The goal is not to create one universal configuration. It is to ensure that editorial decisions are based on a known and repeatable process.

Excluding the Bibliography

Reference lists commonly contain exact titles, author names, journal names, and publication details. These elements naturally match other documents.

Including the bibliography may increase the overall score without revealing anything important about the manuscript’s originality.

Many workflows therefore exclude references from the main assessment. The editor may still need to inspect them for citation accuracy, unusual duplication, or fabricated entries.

Bibliography exclusion should be a deliberate setting, not an attempt to force the score below a preferred threshold.

Excluding Quotations

Proper quotations can be excluded to reduce noise, but this setting also requires caution.

A manuscript may use an excessive amount of quoted material even when every passage is attributed correctly. The issue may be weak original synthesis rather than plagiarism.

Quotation detection also depends on formatting. Poorly marked quotations may remain in the score, while unusual formatting may prevent accurate identification.

Editors should use exclusion settings to improve interpretation, not to avoid reading the relevant sections.

Minimum Match Length

Short strings often produce meaningless matches. Common expressions such as “the results of this study” appear in many publications.

A minimum match length can reduce this noise and make the report easier to review.

However, setting the minimum too high may hide a series of shorter copied phrases or close paraphrases.

The appropriate setting depends on language, discipline, document length, and the journal’s screening objective.

Explainable systems should allow editors to see and adjust this parameter while preserving a record of the configuration used.

Self-Similarity Requires Context

A manuscript may match the authors’ previous publications, thesis, conference paper, preprint, protocol, or repository deposit.

This does not automatically make the reuse acceptable or unacceptable.

Authors may legitimately develop earlier work into a full article. They may also need to repeat limited methodological information. The new manuscript should clearly disclose and cite the earlier version when relevant.

Problems arise when substantial text, data, or conclusions are presented as entirely new without explanation.

An explainable report helps editors determine what was reused and whether the relationship between the documents is transparent.

Preprints and Repository Versions

Preprints create a common source of high similarity. A submitted manuscript may closely resemble a publicly available version posted by the same authors.

This similarity is expected when journal policy permits preprints. The editor should verify authorship and confirm that the manuscript is not a duplicate publication from another journal.

Repository copies, accepted manuscripts, and conference versions can create similar situations.

A raw score may classify the manuscript as highly similar. Source inspection reveals that the match comes from an openly disclosed earlier version of the same work.

Translated and Multilingual Overlap

Text reuse is not always visible through exact matching. A manuscript may translate material from another language without attribution.

Basic similarity systems may assign a low score because the wording has changed completely.

This demonstrates why a low percentage cannot prove originality. Editorial knowledge, citation review, and subject expertise remain important.

Where multilingual comparison is available, the report should explain how translated similarity was identified and how confident the system is.

A Low Score Can Create False Reassurance

Editors may spend less time reviewing a manuscript when the similarity percentage appears low. This creates false reassurance.

A short but central passage may have been copied. The author may have translated text, changed sentence structure, or borrowed ideas without retaining enough wording for direct detection.

Similarity screening is strongest at locating textual overlap. It is less capable of evaluating idea appropriation, data manipulation, fabricated citations, or misleading interpretation.

A low score should therefore mean that little matching text was found under the chosen settings. It should not be described as proof of academic integrity.

A High Score Can Create False Suspicion

High scores can also be misleading. References, quotations, templates, legal statements, and earlier versions may produce extensive overlap.

Automatically rejecting such manuscripts can unfairly penalize authors whose work is legitimate but structurally similar to existing material.

This risk may be greater for authors working with standardized reporting formats or writing in a language that offers fewer alternative technical expressions.

Explainable review protects authors by showing why the score is high and whether the matched passages are academically problematic.

Explainability Improves Fairness

Authors should not receive a rejection based only on a percentage they cannot interpret.

When a journal raises a similarity concern, it should be able to identify the relevant passages and explain which policy may have been violated.

This gives authors an opportunity to correct citation, clarify the relationship with previous work, or respond to a possible system error.

Transparent reasoning also supports consistent treatment. Editors can compare similar cases using documented criteria rather than personal impressions.

Explainability Supports Appeals

Editorial decisions can be disputed. A transparent similarity assessment creates a record that can be reviewed.

The journal can show which source was examined, which passage caused concern, and why the overlap affected the decision.

Without this detail, an appeal becomes a disagreement over one unexplained number.

Explainability protects both authors and journals because it connects the outcome with evidence and policy.

The Editor’s Role Cannot Be Replaced

Similarity software can process large numbers of documents consistently. It can locate matches that would be difficult to identify manually.

It cannot determine the complete scholarly meaning of those matches. It does not know the journal’s expectations unless those rules have been translated into a human workflow.

Editors must consider citation practice, disciplinary convention, manuscript type, intellectual significance, and author disclosure.

The tool supports editorial judgment. It should not become the unnamed decision-maker behind rejection.

Screening Before Peer Review

Early similarity screening can protect reviewer time. Obvious duplication or extensive unattributed copying can be addressed before experts spend hours evaluating the research.

However, screening should remain proportionate. A moderate or high score should trigger inspection rather than automatic rejection.

The editor may request a technical correction, ask for clarification, refer the manuscript to an integrity specialist, or continue with peer review when the matches are legitimate.

The purpose of screening is to identify risk efficiently, not to replace the complete editorial process with a threshold.

Different Sections Need Different Attention

Editors can review reports more efficiently by prioritizing sections according to scholarly importance.

The title and abstract should clearly distinguish the manuscript from previous work. Extensive similarity in these sections may indicate weak originality or undisclosed duplication.

The introduction and literature review will naturally discuss previous scholarship, but wording should be attributed and synthesized rather than copied.

The methods section may contain standardized language, especially for established protocols. The results should normally describe the current study’s observations.

The discussion and conclusion deserve close review because they contain interpretation, implications, and claims of contribution.

Explainability Helps Identify Citation Problems

A source may appear in the reference list while the manuscript still uses it improperly.

The author might reproduce several sentences without quotation marks. A citation placed at the end of a paragraph may not make the extent of copied wording clear.

Alternatively, the manuscript may paraphrase too closely while citing the source accurately.

An explainable report allows the editor to compare wording and citation placement. The problem can then be classified more accurately as missing attribution, inadequate quotation, weak paraphrasing, or acceptable use.

Do Not Confuse Similarity Screening with AI Detection

Similarity screening and AI-generated text detection address different questions.

A similarity system compares a manuscript with available sources. An AI detector attempts to estimate whether text resembles machine-generated writing patterns.

A passage can be original but AI-assisted. It can also be copied from a human-written source while receiving no AI-related signal.

Combining the outputs into one general suspicion score weakens interpretation.

Each tool should have a defined purpose, known limitations, and a separate review process.

Documenting Editorial Decisions

Editors should record why a manuscript was cleared, returned for correction, escalated, or rejected.

The note can identify the relevant source, location of the match, type of overlap, and policy applied.

Documentation improves consistency across editors and creates examples for future training.

It also helps the journal review whether its screening rules produce unnecessary delays or unequal outcomes.

A Practical Screening Workflow

Begin by confirming that the report used the journal’s approved settings. Check whether references, quotations, and small matches were treated consistently.

Review the overall distribution rather than reacting immediately to the percentage. Determine whether similarity is concentrated or fragmented.

Open the largest source matches and compare the passages in context. Check citation, quotation, and the relationship between the source and submitting authors.

Prioritize the abstract, results, discussion, and conclusion. Then examine methods overlap according to disciplinary expectations.

Classify the issue. It may be acceptable similarity, a correctable citation problem, unclear reuse of earlier work, or a serious integrity concern.

Record the reason for the decision and communicate specific concerns to the authors where appropriate.

Questions Editors Should Ask

Where is the matching text located? Is it central to the manuscript’s claimed contribution?

Does the report identify one dominant source or many minor sources? Is the source written by the same authors?

Is the wording quoted, cited, paraphrased, or presented without attribution?

Could the match result from a preprint, thesis, repository copy, reporting guideline, or standard method?

Would the editorial concern remain serious if the overall percentage were hidden?

Can the decision be explained to the author using specific passages and a clear journal policy?

Common Journal Screening Mistakes

The most common mistake is rejecting manuscripts automatically when they cross a numerical threshold.

Another is clearing every manuscript below the threshold without opening the report.

Journals may apply different settings across editors, making scores impossible to compare fairly.

Some workflows ignore the location of matches and treat references, methods, results, and conclusions as equivalent.

Another mistake is describing the score as a plagiarism percentage. This language assigns a judgment the system has not made.

Finally, editors may contact authors with only the total number and no explanation of the passages causing concern.

Building a Better Journal Policy

A journal policy should define similarity screening as a review process rather than a pass-or-fail test.

It should explain the standard settings, the sections receiving priority, the factors editors consider, and the possible outcomes.

The policy may include internal guidance for preprints, dissertations, conference papers, methods reuse, and author self-similarity.

Editors need examples showing why similar percentages can lead to different decisions.

Authors should receive enough public information to understand that manuscripts are assessed by context and attribution, not by one universal acceptable score.

Training Editorial Teams

Editors and editorial assistants need practical training in interpreting similarity reports.

Training should include legitimate overlap, direct copying, close paraphrasing, source duplication, repository versions, and standard language.

Teams can review anonymized example reports and compare their conclusions. Differences reveal where the journal needs clearer policy.

Regular calibration is important because software settings, publication practices, and available source databases change.

Evaluating the Screening Process

Journals should evaluate whether their workflow identifies meaningful problems without creating unnecessary work.

Useful indicators include the number of manuscripts escalated, the reasons for escalation, the proportion resolved through correction, and the frequency of disputed decisions.

The journal can also review whether certain article types, disciplines, languages, or author groups receive disproportionately high scores for legitimate reasons.

Assessment should focus on decision quality rather than simply reducing average similarity percentages.

Why Explainability Builds Trust

Authors are more likely to accept an editorial concern when the journal can show the relevant evidence.

Reviewers and readers also benefit from knowing that manuscripts were evaluated through a reasoned process rather than an arbitrary threshold.

Explainability makes technology accountable. It reveals what the system found, which assumptions shaped the report, and where human judgment entered the decision.

This transparency is essential when screening can delay publication, affect an author’s reputation, or trigger a research integrity review.

Conclusion

Raw similarity scores are useful for directing attention, but they are poor substitutes for editorial interpretation.

The percentage does not explain where the matching text appears, which sources produced it, whether it is cited, or how important it is to the manuscript’s contribution.

Explainable reports provide the details required for fair decisions. They show matched passages, source relationships, section location, exclusions, and configuration settings.

Editors can then distinguish acceptable quotations and standard language from close paraphrasing, undisclosed reuse, and substantial unattributed copying.

Similarity screening works best when the score is treated as a signal rather than a verdict. The quality of journal screening depends not on producing one decisive number, but on understanding what the detected overlap means within the scholarly work.