Why the tools built to identify AI-generated text are selling false certainty — and what a more honest approach looks like.
News

AI Detectors Are Already Broken — And It’s Going to Get Much Worse

Why the tools built to identify AI-generated text are selling false certainty — and what a more honest approach looks like.

In June 2026, writing this, I want to mark the date deliberately. Not as a rhetorical device — as a technical timestamp. Because in a few months, the question this entire industry is built around — did a human or an AI write this? — will be functionally unanswerable. Not difficult. Not probabilistic. Unanswerable. Text, images, audio, video, code, poetry, music. The boundary is dissolving faster than the tools designed to enforce it can adapt, and most of the people selling you detection software know it.

The fundamental flaw nobody advertises

AI detectors do not see AI. That's the first thing to understand, and it's the thing no vendor homepage leads with. What they measure are statistical proxies — primarily perplexity (how predictable each word is given the preceding context) and burstiness (how much that predictability varies across the document). The working assumption is that LLM output is smooth and consistent, while human writing is irregular and unpredictable.

That assumption holds — sometimes. It fails badly in others.

A non-native English speaker writing carefully, deliberately, in a register they're not fully comfortable with will produce text with low perplexity and low burstiness. Not because they used ChatGPT. Because they're being careful. A student who generates a draft with an LLM and then rewrites it sentence by sentence may produce output that scores as fully human. The detector cannot distinguish between these cases. It is measuring a proxy. The proxy is imperfect. And the tools reporting "94% AI" as if it were a forensic result are, charitably, overstating their confidence.

I've read enough flagged student essays — shared by teachers who didn't know what to do with the result — to say this plainly: a score in the 40–75% range means nothing actionable. It means the text has statistical properties that overlap with a population of AI-generated samples. So does a lot of human writing under specific conditions.

Who gets hurt by false positives

This is where the conversation usually stops being abstract. The students most likely to be falsely flagged are non-native speakers, students with formal or controlled writing styles, students who write short sentences, students who use technical vocabulary consistently. In other words: the students who are already navigating the most barriers. The detector doesn't know the difference between "this person used ChatGPT" and "this person writes careful, low-variation English because it's their third language." It returns a percentage. The teacher acts on it. The student has to defend themselves against a statistical artifact.

That's not a edge case. That's a structural problem with the entire approach.

The market is selling certainty it doesn't have

There are now dozens of AI detection tools. Turnitin, GPTZero, Compilatio, Winston AI, Copyleaks, Originality.ai — each with accuracy claims, each with testimonials, almost none with independently verified benchmarks on diverse real-world text. The accuracy numbers cited on product pages are typically measured against clean datasets: pure GPT-4 output versus pure human writing, no paraphrasing, no editing, no mixing. Those conditions don't exist in actual student submissions.

What's being sold is the appearance of a solution to a problem that doesn't have one yet. That's not cynicism — it's the honest technical position. And it matters, because institutions are building disciplinary processes around these tools.

A different approach: use AI to analyze style, not assign verdicts

Here's what actually makes sense given the current state of the technology. Instead of asking a detector "is this AI?", ask a language model to do something it's genuinely good at: identify specific stylistic signals that correlate with generated text, and flag them for human review — without issuing a verdict.

The difference is significant. A well-constructed prompt can ask a model to look for symmetric paragraph openings, uniformly formal register across the entire document, closing statements that summarize rather than conclude, absence of syntactic irregularity, absence of personal voice or concrete personal detail. These are signals. They are not proof. But they're more interpretable than a percentage score, and they give the reader — the teacher, the editor, the reviewer — something to actually examine.

A basic prompt structure for this kind of analysis looks like this: instruct the model to act as a stylistic analyst, not a detector. Ask it to identify passages where the register is unusually consistent, where sentence length variation is low, where conclusions feel summarized rather than earned. Ask it to note what's absent — idiosyncratic word choices, self-interruption, tonal shifts — rather than what's present. Ask for observations, not a score.

The output will be imperfect. It will also be more honest about its imperfection than any tool reporting "87% AI-generated."

What's actually coming — and why the date matters

The reason I marked June 2026 at the top of this article is that the trajectory is visible and it moves in one direction. Multimodal generation — text, image, audio, video produced together, coherently, at scale — is already here in early form. The watermarking initiatives from OpenAI, Google, and others are real efforts, but they depend on voluntary adoption and are trivially bypassed by anyone motivated to do so. The C2PA content provenance standard is promising at the infrastructure level, but requires camera manufacturers, platforms, and publishers to implement it end to end. That's a five-year project minimum, optimistically.

In the meantime: a poem, a song, an image, an essay, a piece of code. Within months, not years, the question of origin will be technically unresolvable for most content produced and shared online. The detection market will continue to exist — there's money in selling false certainty — but the gap between what these tools claim and what they can actually do will widen, not close.

The adaptation that matters isn't finding a better detector. It's rethinking what authorship, originality, and attribution mean in an environment where the signal is permanently mixed. That's a harder conversation. It's also the right one.

Need help with this solution or looking for custom development?

Visit my Stay In Touch page to connect and discuss your project. Discover a range of web development services and specialized WordPress solutions tailored to your needs. Let's work together to enhance your digital presence.