Do AI detectors work? OpenAI shut down its own
I took a text from 100% AI to 100% human by changing two characters. Spare a thought for the people this gets backwards.
Prove me wrong, but I think the idea that AI detectors don’t work is common sense. I may have to eat these words in the future, but I don’t think I will.
And to be clear, I’m not against Substack for the Pangram partnership. I understand the objective, and users were asking for it. The problem is that it doesn’t work the way everyone expects it to work.
AI labeling and detection is important. Think of fake news, images, videos, and impersonation. They matter so people don’t get fooled by what they see, and also for managing expectations - the gap between what a reader thinks they are getting and what is actually there. Chris Best, CEO of Substack, calls that gap Claudefishing.
So I’m not against them. What I want to emphasize is that they don’t always do their job, and worse, that sometimes they punish humans doing real work.
What Substack actually shipped
Substack partnered with Pangram, so you can now scan notes, replies, comments and posts longer than 100 words, and get an estimate of how much of the text was written by hand and how much with AI.
It’s worth reading what Chris Best wrote:
He calls the tool “not perfect”.
He notes it detects whether AI was used, not whether any human care went into the work.
He says it’s there to help people make their own judgments.
Substack also shipped a “How I make this” statement, so you can explain your process instead of letting readers guess.
And this is valid argument to start running some experiments on the platform:
But one thing we do know is that we don’t want to wait until your Substack app turns into LinkedIn before we start to learn and make progress.
How AI detectors work
They don’t read text the way you and I read it. They scan for statistical patterns and calculate a probability that a machine produced it.
Two metrics do most of the work:
Perplexity measures how predictable the text is. An AI is a prediction engine picking the most likely next word, so its output tends to be less surprising than ours.
Burstiness measures the variation in sentence length and rhythm. Human writing is bumpy, mixing short sentences with long complicated ones, while generated text tends to keep an even pace.
On top of that, some detectors look at style: average sentence length, punctuation choices, how often you reach for a complicated word. And some AI providers embed invisible watermarks in their output, which is the one case where detection is on solid ground.
The thing that always bothered me
At first I gave the doubt the benefit of the doubt. But something never sat right with me. How do you tell a text was written by AI, when it’s just text?
Ok, the model can use certain words more than we do, or build sentences in a way people typically don’t. But models change and the model that says 80% AI today can say something else tomorrow.
Detectors are trained on datasets that go stale, and as Adobe puts it: a detector trained on one model is often less effective against a newer one. Two vendors can look at your paragraph and disagree completely.
So, if two thermometers disagree by forty degrees, you don’t argue about the temperature, you stop trusting the thermometers.
What happened when OpenAI tried this
OpenAI built a classifier for exactly this, launched it in January 2023, and reported that it correctly identified 26% of AI-written text while flagging human writing as AI 9% of the time. They said it should not be used as a primary decision-making tool.
Six months later, they killed it:
As of July 20, 2023, the AI classifier is no longer available due to its low rate of accuracy.
In their guidance for educators, they answer directly whether AI detectors work:
In short, no, not in our experience. Our research into detectors didn’t show them to be reliable enough given that educators could be making judgments about students with potentially lasting consequences.
And on what went wrong:
When we at OpenAI tried to train an AI-generated content detector, we found that it labeled human-written text like Shakespeare and the Declaration of Independence as AI-generated.
They also note that students can make small edits to evade detection even when the tool gets it right.
And since people try this too: asking ChatGPT whether it wrote something doesn’t work either. OpenAI says those answers are “random and have no basis in fact”.
It fails in both directions
I asked Claude to write a text, scanned it through Pangram, and the result was 100% AI, which was correct. Then I replaced the curly apostrophes, and swapped the em dashes for commas and hyphens from my keyboard. Scanned it again: 100% human.
Nothing else changed, the words, sentences, or ideas, only the two characters.
The text was still entirely machine-written, and the tool now certified it as mine.
And it works in reverse. Sam Illingworth ran his 2010 PhD thesis through a detector and got 70% AI. As he put it, either he has a time machine, or these tools just don’t work.
Who the errors land on
Students, who can have work thrown out because a detector had a bad day. And people who don’t write English as a first language, like me. Not because we don’t know English - some write in it every day, have meetings, etc. - but because AI does help us find the better word or say the thing more clearly.
OpenAI found their own detector could disproportionately students, along with anyone whose writing is formulaic or concise.
So the tool is most likely to be wrong about the people least able to argue back, as you can’t prove you wrote something. There’s no evidence to produce. It’s your word against a percentage.
While we’re here, this is my AI statement on Substack:
The ideas and the research are mine. Some articles I write, others I narrate out loud. AI helps me edit. Either way, the substance arrives before the AI does.
English is my second language. I’m Portuguese. So I work with AI to find the word I’m looking for, or to say a sentence in a way you will understand better. It touches the words, never the argument.
What works instead
OpenAI’s own suggestion for teachers is the best I’ve read: have students share their actual conversations with the AI. You then get to see the thinking. What did they ask? Did they push back on the answer? Did they check it? Did they notice the bias in it?
It can’t be faked by swapping an em dash, and teachers will be grading judgment, which probably is what they want to measure.
The same applies outside school. If you want to know whether someone did the work, ask them about the work.
Has a detector ever got you wrong? Share about it in the comments.
Thanks for reading!
See you next week,
Nelson




