Large language models can help with drafting and language polishing, but the negative effects on academic papers are now well documented across the publication pipeline. The most important pattern is not a single spectacular failure. It is the combination of fluent prose, weak source grounding, ease of mass production, and patchy disclosure. That mix makes bad papers easier to write, harder to detect, and faster to spread. Publisher policies from Springer Nature, Elsevier, and the ICMJE now explicitly require human accountability, disclosure, and careful checking because AI output can be incorrect, incomplete, biased, plagiarized, or hallucinated.
The evidence is already strong enough to support a clear conclusion for general academic readers: when LLM use is undisclosed or lightly supervised, papers become more vulnerable to hallucinated claims, fabricated references and data-like details, shallow literature reviews, paraphrastic plagiarism, methodological shortcuts, overconfident presentation, and reproducibility failures. Those harms scale because LLM-assisted writing is no longer rare. A Nature Human Behaviour study estimated LLM-modified text in up to 22% of computer science papers, while a Science Advances analysis estimated that at least 13.5% of biomedical abstracts published in 2024 had been processed with LLMs, reaching 40% in some subcorpora.
AI detectors do not solve this. Their false positives can unfairly implicate honest writers, especially non-native English writers, while simple paraphrasing and "humanizing" tools can bypass them. Official guidance from TEQSA says an AI score alone is insufficient evidence of misconduct and notes that humanization add-ons can evade detectors, and OpenAI withdrew its own text classifier in July 2023 because of low accuracy.
Why the Risk Rose So Fast
The speed of adoption matters because academic publishing was already vulnerable to superficial reviewing, paper mills, and overloaded editors before LLMs arrived. LLMs lower the cost of producing plausible looking text, references, abstracts, and reviewer-facing cover language. They also lower the cost of iterating variants until one passes superficial checks. The problem is intensified in crowded areas, short papers, rapid review venues, and environments that reward publication volume more than verification. The Nature Human Behaviour analysis found higher estimated LLM modification in more crowded research areas and shorter papers, and the Nature survey on peer review reported that more than half of 1,600 academics had used AI for peer review, often against journal guidance.
That pressure changes incentives in two directions at once. On the author side, low-effort production becomes easier. On the reviewer side, careful checking becomes costlier relative to the flood of submissions. The result is what many researchers now call "slop": text that is grammatical and professionally formatted but weakly verified, derivative, or fabricated. The shift is visible both in publisher policy changes and in the empirical literature on scientific writing style, citation integrity, and retractions for randomly generated content. A Scientometrics study on randomly generated content retractions identified 3,540 retractions up to May 28, 2024, with annual counts rising from 3 in 2010 to 2,302 in 2023.
Also Read: [HOT] How to Cite Sources in academic work & Avoid Plagiarism?
Failure Modes Across the Paper Pipeline
The table below summarizes the common failure modes. The sections that follow give a concise explanation, dated examples, impact estimates where available, and practical mitigation steps for authors, reviewers, and publishers.
| Failure mode | What it looks like | Evidence and scale | Best mitigation |
|---|---|---|---|
| Hallucinations | Plausible but false factual claims, titles, author names, venues, dates, or identifiers | Walters and Wilder found fictitious references in 55% of ChatGPT bibliographies; only 6.4% were fully accurate. Chelli and colleagues found hallucination rates of 39.6% for GPT-3.5, 28.6% for GPT-4, and 91.4% for Bard in systematic-review search tasks. | Require source-by-source verification before submission; use DOI and database checks; desk reject unverifiable bibliographies |
| Fabricated data and references | Invented cohort sizes, results, figures, trials, or entirely fake papers presented as real | npj Digital Medicine showed ChatGPT could generate believable abstracts with completely generated data. The HKS Misinformation Review found 139 papers with undeclared or fraudulent GPT use in Google Scholar, including 19 in indexed journals. | Demand underlying data, preregistration where appropriate, and auditable provenance for citations, tables, and figures |
| Superficial literature reviews | Fast but shallow syntheses with broad coverage and weak contextual judgment | JMIR AI found GPT-4 quicker and broader than humans but worse on depth, contextual understanding, accuracy, and transparency. EMNLP 2025 reported continuing hallucinated references and variable semantic coverage and factual consistency. | Use librarians, structured search logs, PRISMA-style methods, and human synthesis rather than free-form prompting alone |
| Plagiarism and mosaic copying | Paraphrased borrowing that obscures source dependence without proper acknowledgment | Gupta and Pruthi found a considerable fraction of LLM-generated research documents were identified by experts as paraphrased or significantly borrowed from existing work, and common detectors often missed them. Tripto and colleagues show repeated paraphrasing can distort authorship signals. | Check idea provenance, not only string overlap; require citation of source frameworks; use qualitative review alongside similarity reports |
| Methodological errors | Bad search strategies, mistaken inclusion criteria, flawed meta-analysis, or fabricated scholarly scaffolding | JMIR reported GPT-4 precision of 13.4% and recall of 13.7% in a systematic-review replication task. A Nature Portfolio retraction note dated April 22, 2026 retracted a ChatGPT-in-education meta-analysis over discrepancies undermining confidence in the analysis and conclusions. | Force methods sections to name tools, versions, prompts, search strings, and validation steps; use methodological review before acceptance |
| Overconfidence and weak uncertainty | Fluent claims without calibration, caveats, or abstention when evidence is thin | TrustNLP 2024 reported that LLMs were overconfident most of the time in verbalized uncertainty tasks. A preregistered 2026 calibration study found model confidence exceeded accuracy on average. | Require explicit uncertainty statements, verification notes, and abstention rules for unsupported claims |
| Reproducibility harms | Undisclosed prompts, version drift, missing logs, and non-repeatable model behavior | Angermeir and colleagues examined 85 LLM-centric ICSE and ASE 2024 papers. Of 18 studies with artifacts and OpenAI models, only 5 were sufficiently complete and executable, and none were fully reproduced. | Archive prompts, model versions, temperatures, seeds, outputs, and evaluation scripts; disclose enough detail to replicate |
| Low-effort slop papers | Formulaic prose, shallow novelty, inflated volume, and reviewer overload | Nature Human Behaviour found widespread LLM use and faster growth in crowded fields. Nature reported hidden prompts in 18 preprints in July 2025, aimed at gaming AI-assisted peer review. | Raise submission friction for review-style papers, verify references automatically, and align incentives with quality over output count |
Hallucinations
In academic papers, a hallucination is not merely a wrong answer. It is usually a fluent, citation-shaped error that looks scholarly enough to survive cursory inspection. This matters because citations and literature summaries are the scaffolding of a paper. If the scaffolding is fake, the rest of the argument may still look coherent while being impossible to verify. Walters and Wilder tested 42 topic bibliographies generated by ChatGPT and reported that 55% of references were fabricated and 43% had substantial errors, leaving only 6.4% fully accurate. In a more task-based setting, the JMIR comparative analysis found that GPT-4 retrieved papers from the relevant human systematic reviews only 13.4% of the time and still hallucinated 28.6% of its references.
Also Read: How Gradescope Upholds Academic Integrity?
Documented examples
In June 2023, Jerome Goddard's cautionary note warned biomedical researchers that ChatGPT could provide persuasive but false biomedical content. In 2024, the JMIR paper supplied direct quantitative evidence from systematic-review tasks. In 2026, a large audit reported in eBioMedicine and summarized by Nature found rising numbers of invalid biomedical references, reaching roughly 1 in 458 papers in 2025 and 1 in 277 in the first seven months of 2026.
Scale and impact
The impact is cumulative. Hallucinated references waste reader time, contaminate downstream reviews, pollute citation graphs, and can directly damage clinical, legal, or policy writing when a fake source is relied on as authority. A 2026 preprint on large-scale evidence from biomedical publications conservatively estimated 146,932 hallucinated citations in 2025 alone.
Mitigation
Authors should verify every AI-suggested citation against Crossref, PubMed, Google Scholar, or the publisher site before submission. Reviewers should spot-check high-leverage references rather than assuming bibliographic correctness. Publishers should run automated citation verification and treat unverifiable references as a formal integrity issue, as their own policies now increasingly imply. Springer explicitly warns that AI use can create non-existent citations.
Fabrication of Data and References
LLMs do not need to fabricate an entire paper to damage one. They can fabricate the part of a paper that most readers will not re-run, such as cohort sizes, p-values, trial registration details, or a miniature literature that never existed. That is why even polished, professional text can be methodologically empty. In the npj Digital Medicine experiment published on April 26, 2023, ChatGPT produced believable medical abstracts with completely generated data. Human reviewers correctly recognized only 68% of the generated abstracts and falsely labeled 14% of real abstracts as generated.
Documented examples
The HKS Misinformation Review research note found 139 papers in Google Scholar with undeclared or fraudulent GPT use, 19 of them in indexed journals, and about 57% in policy-relevant areas such as health, environment, and computing. A separate 2025 case study on false authorship reported a fabricated AI-generated article falsely attributed to a real scholar, and suggested that at least 48 of 53 papers examined in one venue appeared AI-generated.
Scale and impact
Fabricated papers do not stay isolated once indexed. The HKS study emphasized that many questionable papers were copied across repositories, archives, academic networking sites, and search infrastructure, increasing the chance that they could be retrieved by students, journalists, policymakers, or rapid reviewers.
Mitigation
Authors should not use LLMs to invent placeholders that they plan to "fix later," because such placeholders routinely survive into final drafts. Reviewers should request raw data, protocols, and provenance when quantitative claims are central. Publishers should require auditable links between every table, figure, and numeric claim and a real dataset, registry, or source document.
Also Read: Avoiding Self-Plagiarism with Turnitin: 5 Key Tips for Students
Superficial Literature Reviews
LLMs are unusually good at producing the surface signals of synthesis: topic framing, contrastive phrases, tidy transitions, and short summaries of schools of thought. They are much weaker at the hard parts of reviewing: transparent search, selection logic, boundary setting, contextual nuance, and knowing when the literature is thin or internally contradictory. The JMIR AI comparative study concluded that GPT-4 was fast and broad but inferior to human researchers on accuracy, depth, contextual understanding, transparency, and consistency. The EMNLP 2025 evaluation paper similarly concluded that even advanced models still hallucinate references and have uneven semantic coverage and factual consistency in review composition.
Documented examples
In the 2024 JMIR AI study, GPT-4 produced an extensive list of factors in seconds, but the authors found it lacked deep contextual understanding and sometimes generated irrelevant or incorrect content. In the 2024 JMIR systematic review comparison, Bard retrieved none of the ground-truth papers and hallucinated 91.4% of references.
Scale and impact
The risk is greatest when LLMs are used as the primary literature search instrument instead of an assistant layered on top of databases and human selection. That turns what should be a documented method into an opaque conversation with a stochastic system.
Mitigation
Authors should treat LLMs as drafting tools, not literature databases. Every review should preserve search strings, databases, inclusion criteria, exclusion criteria, and screening logs. Reviewers should ask whether the cited corpus reflects a real search strategy or merely a generated synthesis. Publishers should require reporting structures modeled on PRISMA or domain-specific review standards.
Plagiarism and Mosaic Copying
Classical plagiarism detection looks for overlapping strings. LLM-assisted plagiarism increasingly works by preserving someone else's structure, method, or core claim while paraphrasing the language enough to avoid obvious text overlap. That is precisely what makes LLM-era plagiarism more "mosaic" than verbatim. The Gupta and Pruthi preprint reported that experts judged a considerable fraction of evaluated LLM-generated research documents to be paraphrased or significantly borrowed from existing work, and that common automated detectors often failed to catch deliberately plagiarized ideas. The Ship of Theseus study also showed why repeated paraphrasing destabilizes authorship attribution and weakens detector performance.
Documented examples
Gupta and Pruthi's 2025 study found Turnitin and OpenScholar detected none of the sampled deliberately plagiarized proposals in their smaller test subset, while stronger LLM-based methods only worked well in unrealistic "oracle access" conditions where the correct source paper was already supplied.
Scale and impact
This is especially damaging in methods-heavy fields where contribution is often a recombination of prior ideas. A paraphrased method section or novelty claim can be academically consequential even if no paragraph is copied verbatim.
Mitigation
Authors should cite conceptual parents, not just proximate textual sources. Reviewers should compare the claimed contribution against familiar prior work at the level of method mapping and idea lineage. Publishers should combine similarity screening with subject-expert review and explicit novelty statements grounded in cited literature.
Methodological Errors
LLMs commonly fail in places where academic method is fussy, rule-bound, and sensitive to hidden assumptions. Search criteria, coding protocols, eligibility rules, statistical aggregation, and causal wording are all vulnerable. The JMIR systematic-review comparison is a direct methodological warning because it measured retrieval performance against real systematic reviews and found very low recall and precision. A different warning came on April 22, 2026, when a Nature Portfolio journal retracted a ChatGPT-in-education meta-analysis over discrepancies in the meta-analysis that undermined the editor's confidence in the conclusions.
Documented examples
In the npj Digital Medicine study, generated abstracts often imitated the scale of real cohort sizes and journal style but used invented data. In 2023, Retraction Watch documented papers showing telltale interface phrases such as "Regenerate response," which signaled not only undeclared AI use but also a lack of method-level checking before publication.
Scale and impact
Methodological error is more serious than stylistic error because it can survive peer review and propagate into secondary syntheses, media coverage, or policy recommendations. A polished but flawed meta-analysis can be more damaging than a visibly bad paper.
Mitigation
Authors should disclose where the model entered the method pipeline, not just the prose pipeline. Reviewers should ask whether core steps were performed by humans, by software with auditable rules, or by free-form prompting. Publishers should expand methods checklists to include model version, prompt text, sampling settings, and validation procedures. The ICMJE manuscript guidance now says AI use in conducting a study should be described in the methods in sufficient detail to enable replication, including tool, version, and prompts where applicable.
Overconfidence and Lack of Uncertainty
One of the most academically dangerous features of LLM output is that it often sounds more certain than the evidence warrants. That matters in discussion sections, review articles, abstracts, and policy-facing papers, where boundary conditions and caveats are a key part of honest scholarship. The Peters and Chin-Yee study on summarization of scientific research found that LLM summaries were nearly five times more likely than human-authored summaries to contain overly broad generalizations, with some models overgeneralizing in 26% to 73% of cases. The TrustNLP 2024 paper reported poor uncertainty estimation and frequent overconfidence, and a 2026 preregistered study found that model confidence exceeded accuracy on average.
Documented examples
Walters and Wilder observed that ChatGPT often gave inaccurate answers even when asked to verify the legitimacy of its own citations, which is exactly the kind of confident self-repair that can mislead an author into false reassurance.
Scale and impact
Overconfidence does not always look like fabrication. Often it appears as the removal of qualifiers, the inflation of a tentative result into a general claim, or the omission of disciplinary disagreement. That makes it especially dangerous in general-interest reviews, editorials, and educational syntheses.
Mitigation
Authors should require models to produce confidence and evidence notes, then independently verify both. Reviewers should look for missing uncertainty language, not just factual mistakes. Publishers should encourage structured uncertainty reporting in abstracts and discussion sections, especially when AI tools were used in summarization.
Reproducibility Harms
If a paper is partly shaped by a changing commercial model, without preserved prompts, versions, temperatures, seeds, and outputs, the writing process itself becomes harder to reconstruct. That is a reproducibility problem even before any scientific result is tested. The Angermeir reproducibility study found that among 18 LLM-centric software engineering studies with artifacts and OpenAI models, only 5 were sufficiently complete and executable, none were fully reproducible, 2 were partially reproducible, and 3 did not seem reproducible. Emerging reporting guidelines such as the GAMER Statement and the CHAT-GUIDE development notice stress prompt and output documentation precisely because reproducibility otherwise collapses.
Documented examples
The ICMJE now explicitly asks authors to report AI use in sufficient detail to enable replication, including prompts where applicable. That change is itself evidence that hidden prompting and undocumented model intervention had become a serious reporting problem.
Scale and impact
Reproducibility harms extend beyond computational disciplines. If a literature review, coding frame, or qualitative memo was materially shaped by an LLM, future researchers need to know how. Otherwise they can only reproduce the published text, not the decision process that created it.
Mitigation
Authors should archive prompts, outputs, model identifiers, dates, settings, and post-edit history in supplementary materials when AI materially shaped the work. Reviewers should ask for that record when AI use is disclosed or suspected. Publishers should standardize AI disclosure fields and machine-readable metadata.
Incentives That Encourage Low-Effort Slop Papers
LLMs fit too well with existing academic incentives: fast drafting, pressure to publish, reviewer overload, and bibliometric reward systems that count outputs more readily than they assess epistemic care. This does not imply that most AI-assisted papers are poor. It does imply that the marginal cost of producing apparently competent but weakly checked manuscripts has fallen sharply. The Nature Human Behaviour study found higher estimated LLM modification in crowded research areas and shorter papers. A Nature news report in July 2025 documented 18 preprints containing hidden prompts aimed at gaming AI-assisted peer review, which is a vivid example of how incentives can push from mere convenience into strategic manipulation.
Documented examples
The hidden-prompt incident involved instructions in white text or tiny font such as requests for positive review only, exploiting automated or semi-automated review habits. Nature also reported that offending preprints would be withdrawn from servers.
Scale and impact
The risk is systemic because reviewers are also using LLMs. In the Nature survey, more than 50% of 1,600 academics reported using AI for peer review, often despite publisher guidance. When both sides automate, shallow papers and shallow reviews can lock into each other.
Mitigation
Authors need clearer norms for acceptable assistance and stronger sanctions for deceptive prompting or fabricated references. Reviewers need lower review loads and better editorial triage. Publishers and funders need to reward transparent, verifiable contributions rather than sheer output volume. This last point is partly an inference from the convergence of adoption, overload, and manipulation evidence.
Why AI Detectors Often Worsen the Problem
AI detectors are attractive because they promise a cheap answer to a messy integrity problem. In practice, they create three new problems: false positives, gaming, and chilling effects. The false-positive problem is empirically strong. The Patterns study by Liang and colleagues found that GPT detectors frequently misclassified non-native English writing as AI-generated, and one result reported that 97.8% of TOEFL essays were flagged by at least one detector. That means the population most likely to use legitimate language support can also be the population most likely to be wrongly accused.
False Positives
Even detector vendors acknowledge reliability limits. Turnitin says false positives exist and that its lower-range scores are less reliable, while its AI writing report guidance warns that results between 1% and 19% are more prone to false positives. Official advice is even more restrictive. TEQSA says the AI score alone is insufficient to allege misconduct.
Gaming
Detectors are easy to bypass. LREC-COLING 2024 showed that minor perturbations and paraphrasing can evade AI-text detectors, and a 2025 adversarial paraphrasing study reported that detection rates can be drastically reduced with slight or no degradation in text quality. TEQSA now explicitly notes that "humanisation" add-ons can bypass detectors.
Chilling Effects
Once institutions attach misconduct procedures to unreliable scores, honest writers can become more cautious about legitimate writing support, formal prose, or even certain stylistic choices. This is an inference, but it is strongly supported by the bias evidence and by official recommendations against using detector scores as standalone proof. Several universities now advise against relying on detectors for disciplinary action, including the University of Pittsburgh, which says AI detection tools are not accurate enough to prove violations.
Detector Failure Is Not New
OpenAI removed its own AI classifier on July 20, 2023 because of its low rate of accuracy. That is a particularly important signal because it came from a developer with strong incentives to make such a tool work.
The deeper problem is conceptual. Detectors ask "who wrote this text?" while research integrity often depends on a harder question: "what claims, methods, and sources can be verified?" A flawed detector can therefore punish harmless drafting assistance while missing the more serious problem of fabricated evidence or AI-aided mosaic copying.
Timeline of Warnings, Failures, and Policy Responses
The timeline shows how quickly the problem evolved from early citation warnings in 2023 to large-scale contamination, peer-review gaming, and formal reporting guidance by 2025 and 2026. The dates and incidents are drawn from primary studies, retraction notices, publisher policies, and official guidance.
2023
- April 2023 — npj Digital Medicine shows ChatGPT can generate believable abstracts with fabricated data.
- April 2023 — Patterns study shows detector bias against non-native English writers.
- July 2023 — OpenAI withdraws its AI text classifier for low accuracy.
- October 2023 — Retraction Watch documents papers showing undeclared ChatGPT traces such as "Regenerate response."
2024
- August 2024 — JMIR AI study finds human literature reviews deeper and more accurate than GPT-4 reviews.
- November 2024 — Science Advances excess-vocabulary analysis estimates at least 13.5% of 2024 biomedical abstracts processed with LLMs.
2025
- May 2025 — Royal Society Open Science study reports overgeneralization bias in science summaries.
- July 2025 — Nature reports hidden prompts in 18 preprints to game AI-assisted peer review.
- November 2025 — Nature Human Behaviour estimates LLM modification in up to 22% of computer science papers.
2026
- April 2026 — Nature Portfolio retracts ChatGPT-in-education meta-analysis over analytical discrepancies.
- May 2026 — Large-scale preprint estimates 146,932 hallucinated citations in 2025.
- July 2026 — TEQSA says AI scores alone are insufficient evidence and humanizers can bypass detectors.
One striking feature of the timeline is asymmetry. Writing tools spread first. Reliable governance, reproducibility standards, and editorial countermeasures followed later. That gap helps explain why detector-first responses have felt tempting but inadequate. The infrastructure for source verification and transparent reporting has lagged behind the infrastructure for generating plausible prose.
Mitigations for Authors, Reviewers, and Publishers
A workable response is not "ban AI everywhere," because publisher and journal policies already recognize legitimate uses such as language support and structured assistance. The workable response is narrower and more rigorous: treat AI use as a potentially material part of scholarly method, and require the same transparency, traceability, and verification that good scholarship already demands in every other part of the paper. The ICMJE, Springer Nature, and Elsevier all now place the burden of responsibility on human authors and require disclosure when AI materially contributes to submitted work.
For Authors
Use LLMs, if at all, for bounded tasks such as language smoothing, brainstorming, or formatting templates, not for unsupervised citation gathering, evidence synthesis, or quantitative claims. Keep a log of prompts, outputs, edits, and model versions. Verify every reference, every number, and every summary claim against an external source. Disclose material use in the acknowledgments or methods as required by journal policy. ICMJE preparation guidance and the GAMER Statement are good minimum baselines.
For Reviewers
Stop asking only "does this sound polished?" and ask "can I audit the evidence path?" Spot-check references, look for oversmoothed prose that removes uncertainty, ask whether the literature search is documented, and request method-level disclosure when the text looks heavily templated or unusually generic. Avoid outsourcing the review itself to chatbots, especially where journal policy forbids uploading manuscripts into external AI systems. Springer Nature explicitly asks peer reviewers not to upload manuscripts into generative AI tools.
For Publishers and Editors
Move from detector-centered governance to verification-centered governance. Add automated citation validation, structured AI disclosure metadata, model-and-prompt reporting requirements, and sanctions for fabricated references or deceptive hidden prompting. Keep AI detectors out of standalone disciplinary use. TEQSA is right that an AI score alone is insufficient evidence, and OpenAI's own withdrawal shows why.
For the Academic System as a Whole
Reduce the incentives that favor fast, generic output over slow, auditable scholarship. That means fewer rewards for sheer paper count, more support for librarians and methodologists, stronger editorial triage, and clearer distinctions between acceptable assistance and deceptive substitution. The empirical literature increasingly suggests that quality assurance has to happen before publication, not only after contamination has already entered the literature and search infrastructure.
The Bottom Line
The central lesson is simple. LLMs do not merely introduce new kinds of error. They amplify old weaknesses in scholarly publishing by making low-cost fluency abundant. Academic papers are harmed when fluency outruns verification. The safest response is not stylistic suspicion, and it is not detector panic. It is disciplined transparency, source verification, and institutional incentives that reward checked knowledge over polished text.