Teachers vs. AI: Welcome to the Classroom Detective Era

A cartoon dog character dressed in an unsettling rainbow-wigged clown costume, grinning a little too wide, arms thrown up like it's been caught red-handed
Jun 14, 2026 10 min read AI in educationAcademic IntegrityCritical Thinking

Do AI detectors actually work? Real research on false positives, what the IB's AI policy really says, and why punishing AI use may not be the answer.

Somewhere in the last couple of years, teachers leveled up from “person who grades your essay” to “person who reads your essay like it’s a ransom note, looking for clues.” Suspiciously perfect topic sentence? Noted. A semicolon used correctly, out of nowhere, in week three of the semester? Case file opened. This is not paranoia for its own sake: it’s a reasonable response to a tool that got very good, very fast. But it does mean your English teacher now has slightly more detective energy than your average detective.

Think About It

Before we get into it: if a teacher pulled up your last essay right now and asked you to explain, out loud, why you chose one specific sentence in it, could you? Not whether it sounds like you: could you defend the choice? That gap, if there is one, is basically the whole topic of this post.

So, what are teachers doing about it?

Turns out “just ban it and hope” didn’t really work as a strategy, so most schools have moved on to something closer to actual forensics. A few of the moves showing up in classrooms right now:

  1. Baseline writing samples. An in-class, no-devices, handwritten essay early in the term, purely so the teacher has a fingerprint of how you write (typos, weird phrasing, favorite sentence structures and all) to compare later work against.
  2. Version history as an alibi. Google Docs and Word both timestamp every keystroke in the background. A document that appears fully formed in one paste, at 11:58 p.m., is a very different story from one that was built up sentence by sentence over three evenings.
  3. Live oral defense. “Walk me through this paragraph” is a surprisingly brutal test. It’s hard to explain a rhetorical choice you didn’t make.
  4. The vocabulary tell. Certain words show up in AI output so often they’ve basically become a genre: delve, tapestry, moreover, in today’s society, it is important to note. One of them isn’t proof of anything, but all of them stacked in a single paragraph you can’t explain raises an eyebrow.
  5. In-class writing days. Old-fashioned, low-tech, extremely effective: pencil, paper, no wifi, go.

None of these are foolproof, and a determined student can route around most of them. That’s kind of the point, though. The goal was never a perfect trap. It’s making the effortless version of cheating require actual effort.

Quick check

A teacher notices a student's essay uses 'delve' twice, 'moreover' three times, and has zero typos, in a class where the student's handwritten baseline sample had several. What's the fair read here?

  • It's obviously AI-written — those words are basically a signature

    Suspicious, sure. Proof, no. Plenty of humans genuinely love the word 'delve' and had a good day. A pattern is a reason to ask a follow-up question, not a verdict on its own.

  • It's worth a closer look (maybe a quick conversation about the essay) rather than an automatic accusation

    This is the move most schools are landing on. Patterns are a prompt to investigate, not a courtroom verdict: a two-minute conversation about the paragraph settles it far more fairly than a word-frequency count ever could.

  • It doesn't matter at all — word choice proves nothing so there's nothing to check

    Also too far the other way. It's not proof, but a sharp, sudden shift from someone's normal writing voice is exactly the kind of thing worth a quick, low-stakes conversation.

The awkward truth about AI detectors

Here’s the part that makes all of the above matter more, not less: the software that’s supposed to just settle this automatically is nowhere near as reliable as it sounds. A widely cited Stanford analysis ran 91 real, entirely human-written TOEFL essays (from non-native English speakers) through seven commercial AI detectors. On average, 61.3% were flagged as machine-written. Native-English essays in the same test got flagged about 5.1% of the time. Nearly all of the non-native essays got flagged by at least one detector, and about a fifth were unanimously called “AI” by all seven, despite being written entirely by humans (Liang et al., 2023).

That’s not a rounding error. That’s a tool penalizing people for the way they learned English, not for using AI at all.

Myth Busters

Myth: an AI detector gives you a clean, science-backed yes-or-no answer. Fact: it gives you a probability, generated by a system that’s frequently wrong in both directions (missing real AI text and flagging real human text), and the humans it flags incorrectly most often are exactly the students who can least afford to be wrongly accused.

A detector flagging an essay isn’t nothing, but treating its output as a verdict (rather than a “maybe go double-check this one”) is how you end up punishing your most careful writers and your least confident English speakers at the same time. Not exactly a Magic 8-Ball you’d want holding your GPA hostage.

Meanwhile, the IB is saying: relax, it’s basically just the new calculator

Zoom out, though, and one of the more measured voices in this whole conversation is coming from the International Baccalaureate. The IB’s official position is that it will not ban AI tools outright: its reasoning, more or less, is that banning useful technology has never worked as a long-term strategy. Instead, the IB frames AI the way earlier generations framed spell-checkers, translation software, and yes, calculators: something that’s about to just be part of everyday life whether anyone votes on it or not, meaning schools need to adapt their teaching and assessment around it rather than pretend it isn’t there (International Baccalaureate, statement on AI).

Which, fair. Nobody in 1985 seriously argued that calculators were going to rot humanity’s ability to add. Though, respectfully, some of us could still use the practice.

Here’s the part that doesn’t get to relax, though: the IB is explicit that AI-generated text, images, or graphs are not considered the student’s own work, and if any of it shows up, it has to be clearly identified and properly referenced, same as quoting a book or a website. The calculator comparison covers whether AI exists in the classroom. It says nothing about whether you can pass its output off as yours without saying so. Those are two completely different questions, and it’s easy to accidentally answer the first one while ignoring the second.

Pro Tip

Simple test: if you used AI for a sentence, a structure, an idea, or even just to “clean up” your phrasing, could you say out loud, specifically, what it did? If the honest answer is “I’m not totally sure anymore,” that’s usually a sign it’s time to disclose it, not hide it.

Quick check

A student uses AI to help brainstorm three possible essay topics, then writes the entire essay themselves. Under an IB-style academic integrity approach, what's the right move?

  • No need to mention it — brainstorming isn't "real" writing

    Close, but the IB's guidance leans on transparency, not on deciding for yourself, unannounced, which uses are too minor to count. A short note on how AI was used is the safer, more honest move, and it's a pretty low bar to clear.

  • Briefly disclose that AI was used to help brainstorm topics, even though the writing itself is fully the student's own

    This is the IB's approach: the emphasis is on transparency about how AI was used, not on a strict citation format for every single interaction. A one-line disclosure covers it.

  • Treat it exactly like using AI to write the whole essay, since any use at all is equally serious

    That flattens a pretty important distinction. Brainstorming with a tool and having a tool write your sentences for you are very different levels of involvement, and treating them identically doesn't match how these policies are written.

Okay, but say a detector could catch everything. Should it?

Here’s a genuinely uncomfortable follow-up question. Imagine, hypothetically, that AI detection somehow became perfect tomorrow: zero false positives, zero misses, catches every AI-assisted sentence with total accuracy. Would the right move be to just… punish every instance, every time?

Probably not, for a fairly practical reason. A system built entirely around catching and punishing teaches students to get better at not getting caught, a completely different skill from not needing to cheat in the first place. Punitive-only systems are really good at producing better evasion and really bad at producing better learners.

The more useful move, and the one this whole post has been circling, is incentivizing students to not reach for AI on things that don’t require it, not out of fear of a detector, but because doing the basic thinking themselves is the entire point of being there. And this isn’t just a vibe. Leaning on AI heavily and often has measurable costs to how people think:

  • A systematic review of AI dialogue system use in education found that over-reliance is associated with measurably weaker independent problem-solving and critical thinking, especially when students use AI for tasks they were fully capable of doing themselves (Zhai, Wibowo & Li, 2024).
  • Research on ChatGPT’s cognitive effects found associations between heavy reliance and reduced depth of engagement in learning and memory tasks: when the tool does the work, the brain does less of it (Bai, 2023).
  • A 2025 randomized controlled trial protocol studying generative AI and college students was designed specifically because of concern that AI assistance may reduce the cognitive effort students invest in analytical writing tasks, even when task performance looks fine on the surface (Chen et al., 2025).

Put those together and a pattern shows up. The risk isn’t really “AI exists.” It’s outsourcing the exact reps (finding information, weighing whether a source is credible, structuring an argument) that build the muscle in the first place. Skip the reps enough times and, same as any muscle, it doesn’t show up when you need it.

So maybe the real goal isn’t a smarter detector or a stricter policy. It’s students who don’t reach for AI to summarize a paragraph they could’ve read themselves, who still know how to sit with a confusing source and figure out, on their own, whether it holds up, instead of asking a chatbot to decide for them. That’s not a tech problem or a policy problem. That’s just… learning, the regular kind, the kind that was always going to take some effort.

If you want the toolkit for that last part (telling a credible source from one that just sounds credible), that’s credibility, validity, and reliability. And if you want to understand why skipping the thinking feels so tempting in the first place, that’s cognitive bias 101.

Frequently asked questions

Do AI detectors actually work?

Not reliably enough to be the final word. A widely cited Stanford analysis found AI detectors flagged 61.3% of genuine, human-written essays from non-native English speakers as AI-generated, compared to about 5.1% for native speakers — a huge false-positive gap, not a clean yes-or-no answer.

Does the IB (International Baccalaureate) ban AI tools?

No. The IB has said it won't ban AI, comparing it to how calculators and spell-checkers eventually just became part of everyday schoolwork. It does require that any AI-generated text, image, or graph be clearly disclosed and referenced — it's never counted as the student's own work.

Why do AI detectors flag non-native English speakers more often?

Detectors often measure how 'predictable' word choices are. Non-native English writers tend to use more restricted, predictable vocabulary for reasons that have nothing to do with AI, which research has found makes them statistically more likely to be falsely flagged.

Should schools punish every instance of AI use they catch?

Purely punitive systems tend to teach students how to avoid getting caught rather than how to think independently. Research on AI over-reliance links heavy AI use to weaker critical thinking, which suggests incentivizing students not to reach for AI on tasks they don't need it for works better long-term than punishment alone.

Ready to think like this on purpose?

Alekhni's curriculums and writing tools are built to turn this kind of thinking into a habit.

Keep reading