Ways to Spot AI Writing in Classrooms
AI Writing

Ways to Spot AI Writing in Classrooms

Shadab Sayeed
Written by Shadab Sayeed
September 18, 2026
Calculating…

Teachers should not try to “prove AI” from a detector score or a stylistic tic. The strongest approach is triangulation: notice a cluster of writing anomalies, compare the work with the student’s established writing, inspect drafts and revision evidence, verify sources, and hold a brief, neutral oral follow-up.

Independent research shows why caution matters. Hadra, Cambridge and Mesbah (2026) found Turnitin accuracy of 86% on humanities texts but only 51% on science texts; Originality.ai fell from 96% to 58%. Liang et al. (2023) found an average 61.3% false-positive rate across seven detectors on 91 TOEFL essays by non-native English writers. Krishna et al. (2023) showed that paraphrasing could drive DetectGPT accuracy from 70.3% to 4.6% at a fixed 1% false-positive rate. Detectors can therefore be useful screening signals, but not verdicts.

What the research says

The research picture is remarkably consistent: detector performance changes with genre, writer population, text length, model, and editing. In a balanced benchmark of 192 texts, Hadra, Cambridge and Mesbah (2026) reported overall accuracy of 61% for Turnitin and 69% for Originality.ai. Both struggled especially with mixed human-and-AI writing, and the authors concluded that neither tool was sufficiently reliable for high-stakes academic-integrity judgments.

Bias is particularly important for teachers. Liang et al. (2023), published in Patterns, tested seven widely used detectors on 91 TOEFL essays and 88 U.S. eighth-grade essays. The non-native English essays produced an average false-positive rate of about 61.3%; after the researchers used ChatGPT to enhance word choice, the rate fell to 11.6%. In other words, relatively predictable or constrained English can resemble the statistical patterns detectors associate with AI.

Evasion is the other major weakness. Krishna et al. (NeurIPS 2023) found that automated paraphrasing reduced DetectGPT accuracy from 70.3% to 4.6%. Weber-Wulff et al. (2023), after evaluating 12 public detectors plus Turnitin and PlagiarismCheck, concluded that the tested tools were neither accurate nor reliable and that text-obfuscation techniques worsened performance.

A real-assessment experiment by Perkins et al., published in the Journal of Academic Ethics in 2024, inserted 22 GPT-4-generated submissions into university assessment. Turnitin identified 91% of those submissions as containing at least some AI, yet marked only 54.8% of the generated content; faculty escalated 54.5% of the experimental submissions to misconduct procedures.

Also Read: What Are Common Phrases That AI Uses?

Practical signs teachers can check

Look for patterns rather than “AI words.” The following are useful reasons to investigate, not proof of misconduct:

  • A sudden, unexplained change from the student’s normal vocabulary, sentence complexity, organization, or demonstrated subject knowledge.
  • Smooth but generic prose that repeatedly restates the thesis, gives balanced-sounding claims without concrete course details, or uses examples the student cannot explain.
  • Citations that do not exist, do not support the stated claim, or contain incorrect authors, titles, page numbers, or publication details.
  • Internal mismatches: terminology changes midway through the paper, an unusually sophisticated paragraph sits beside much weaker prose, or the conclusion introduces reasoning that was never developed.

None of these signals establishes AI authorship. Punctuation such as em dashes, flawless grammar, formal transitions, or simply sounding “too polished” are especially weak evidence by themselves. The strong genre effects and false-positive findings in independent studies make style-only judgments unsafe.

Process evidence is usually more informative. Compare the submission with earlier work or a short in-class writing sample. Ask for outlines, notes, source annotations and drafts. Where privacy and school rules permit, examine document revision history. A large pasted block can justify a question, but it should remain circumstantial evidence because revision history does not directly identify who or what composed the text.

Next, use a short oral follow-up: ask the student to explain the thesis in ordinary language, describe why a source was chosen, walk through one paragraph’s reasoning, or revise a sentence while explaining the change. A 2025 study of 24 master’s-level EFL students found that teacher-led interviews encouraged students to account for their reasoning and writing decisions rather than relying on surveillance alone.

Assignment design can reduce uncertainty before it occurs. Use staged checkpoints, class-specific sources, local examples, short in-class planning, annotated bibliographies, reflection on revision choices, and an AI-use statement. Tell students in advance whether brainstorming, grammar assistance, translation, summarization, coding help, or generative drafting is permitted. Traditional plagiarism matching and AI-authorship detection should also be treated as different questions: textual similarity does not establish who produced a passage.

Also Read: How LLMs Can Degrade Academic Papers?

How detection methods compare

Name Approach Accuracy reported in research Important limitations
Turnitin AI Writing Proprietary statistical/classification system Hadra et al. (2026): 86% accuracy on humanities text versus 51% on science text. Perkins et al.: 91% of experimental submissions had some AI flagged, but only 54.8% of generated content was identified. Strong genre effects; mixed human/AI authorship is difficult; a score is not proof of misconduct.
Originality.ai Proprietary text classifier using statistical language signals Hadra et al. (2026): 96% accuracy on humanities text versus 58% on science text; 69% overall in the study’s balanced dataset. Poor performance on hybrid writing; results varied by text type and writer population; independent results may differ from vendor claims.
DetectGPT Research method based on probability curvature around language-model output Krishna et al. (2023): 70.3% detection accuracy fell to 4.6% after paraphrasing, at a constant 1% false-positive rate. Highly vulnerable to paraphrasing in the experiment; designed as a research detection method rather than a conclusive classroom authorship test.

These figures are study-specific, not universal accuracy ratings. Detectors and language models change over time, and results from one dataset should not be assumed to transfer to another subject, grade level or student population. Hadra et al. (2026) explicitly found significant effects from genre and text length.

Also Read: [DIRECT ANSWER] Is iThenticate the same as Turnitin?

False positives, ethics, and teacher actions

Student complaints on Reddit illustrate the stakes, although they are self-reported anecdotes, not verified population evidence. In r/education, around May 2024, a high-school senior said a teacher escalated a detector result involving five sentences and a reported 36% probability to the vice principal and threatened the student’s grade. The student pointed to Google Docs history as evidence of the writing process.

In r/AskTeachers, around January 2026, a student said Turnitin labeled a personal essay 70% AI; after the student removed em dashes and resubmitted, the reported score became 40%, despite the student also offering timestamped version history and an in-person explanation. The student later reported that the English department chair acknowledged the system was not functioning as intended in that case.

These stories reinforce what the controlled research already shows: punctuation or a detector percentage should prompt inquiry, not punishment. False accusations may fall particularly heavily on multilingual writers; Liang et al.’s results provide direct evidence of this risk.

UNESCO’s guidance recommends a human-centered, safe, equitable approach to generative AI in education and specifically highlights privacy and ethical validation. In practice, teachers should collect only the evidence they need, avoid uploading student work to unapproved third-party services, discuss concerns privately, and give students a meaningful opportunity to respond.

When AI use is suspected, document the specific concerns; separate possible AI use from ordinary plagiarism or citation problems; check what the assignment actually permitted; compare process evidence; and speak with the student. If corroborating evidence remains strong, follow the school’s established integrity process and appeal route. If it does not, do not convert uncertainty into guilt. The defensible classroom standard is multiple independent forms of evidence, never a detector score alone. That position is consistent with recent independent detector evaluations.

About the Author
Shadab Sayeed

Shadab Sayeed

CEO & Founder · DecEptioner
Dev Background
Writer Craft
CEO Position
View Full Profile

Shadab is the CEO of DecEptioner — a developer, programmer, and seasoned content writer all at once. His path into the online world began as a freelancer, but everything changed when a close friend received an 'F' for a paper he'd spent weeks writing by hand — his professor convinced it was AI-generated.

Refusing to accept that, Shadab investigated and found even archived Wikipedia and New York Times articles were being flagged as "AI-written" by popular detectors. That settled it. After months of building, DecEptioner launched — a tool built to defend writers who've been wrongly accused. Today he spends his days improving the platform, his nights writing for clients, still driven by that same moment.

Developer Content Writer Entrepreneur Anti-AI-Detection