Neurocourse
How to tell if a text was written by AI: the signs, the detectors, and 24 real accusations

How to tell if a text was written by AI: the signs, the detectors, and 24 real accusations

17 min read

In short: reliably determining whether a text was written by AI is impossible today — neither by eye nor by detector. But we have something more interesting than theory. While mapping the AI course market, we pulled student reviews off Udemy and Coursera and hit a rare seam: real people publicly accusing the author of a paid course of shipping machine-made material. Not one of them ran a detector. We looked at what actually gave it away — and it turned out to be nothing that detectors measure.

Why the problem is unsolvable in the first place

A language model was trained on human text and is optimised to produce statistically "normal" text. Good machine text is, by definition, indistinguishable from average human writing. The task "tell AI from human" runs into the fact that the real boundary isn't machine versus human — it's distinctive style versus averaged style. And people write in an averaged way too, especially when tired, cautious, or working in a second language.

Add that almost nobody publishes raw model output: text gets edited, extended, blended with your own. Pure "human" and pure "AI" cases barely exist — there's a continuum. The course we dissect below is exactly that hybrid: the script looks machine-written, but a human assembled the slides, recorded the voice and sold the thing. "Did AI write this" doesn't even have a single answer here. And that's the normal case, not the exotic one.

Our data: 24 accusations of machine authorship, zero detectors

We pulled the reviews for the largest ChatGPT and generative-AI courses — our own market research, measured 17.07.2026. The by-product turned out to be worth more than the main result. Buried in the negative reviews is a recurring charge: a machine made this. And it's levelled by people who paid for the course and sat through it.

The bluntest version lands on Prompt Engineering with ChatGPT Masterclass by RPATech — 118,803 students, 57,836 reviews, rating 4.48. That is not a fringe product; it's one of the biggest courses in the niche:

"Written by AI, delivered by AI. The slides are way too crowded to be useful and it really doesn't help to have the bot read them out word-for-word"
— Bradley M., 16.04.2026, 1★

And on the same course, about as terse as a verdict gets:

"AI voice, not recommended. didn't learned anything"
— Ashvin D., 22.06.2026, 1★

Notice what isn't in those reviews. No percentage. No mention of GPTZero or ZeroGPT. No talk of perplexity, sentence rhythm, or "not just X, but Y" phrasing. Nobody ran a detector — two things were enough: the slide was overloaded, and the spoken words matched the written words exactly. Those aren't clues about style. They're clues about process: that's what text looks like when nobody rewrote it for the ear, because between generation and recording there was no human.

What actually gives a machine away: not style, but absent labour

We sorted every negative review into clusters, and a pattern fell out. The single most common complaint about video courses isn't about AI at all — it's about reading off the slide:

"why read straight from the slide? I can do that. This was not a helpful course at all"
— Janie I., 02.07.2026, 1★ (Justin Barnett's course)

"50% of this course is reading script like a robot from the slides. …they are just reading text from the slides which you can also you from any good website"
— Shashank T., 13.04.2026, 1★ (The Complete AI Guide)

"They are just reading the prompter sometimes without even knowing the point. Hating myself after purchasing it."
— Sachin S., 13.07.2026, 1★ (The Complete AI Guide)

Look at Sachin's phrasing: "sometimes without even knowing the point". That is the sharpest definition of machine text I've seen, and it has nothing to do with statistics. Machineness is what happens when no one who understood the text stood between it and the reader. A model can simulate understanding inside every individual sentence. What it can't simulate is the decision to throw half of them out — because that decision requires knowing why you're writing.

Which produces the second cluster: filler and repetition.

"I finished week 2. And all the lessons could be in 1 lesson. I jumped to week 3, 4, 5 and 6, and I see some concept reapeted. Very bad time spending for me."
— Shai Mizrachi, 22.07.2023, 1★ (Vanderbilt)

"The course is far too drawn out… padding the course with filler and making it longer than it needs to be."
— Victoria X., 13.02.2026, 2★ (The Complete AI Guide)

"The majority of the class he just rambles around unimportant subject… There's really no 'meat' in this course. A major waste of time and money!"
— Kevin K., 20.02.2026, 1★ (ChatGPT for Work)

And the third: emptiness passed off as content.

"The content of this course is incredibly simple to the point of uselessness, it is essentially stating that generative AI exists and listing a bunch of example models."
— Nicholas Munford, 1★ (Generative AI: Introduction and Applications)

"Didnt meet expectations, only theory is discussed which we already know."
— Aditya Nagavolu, 19.11.2023, 1★ (Generative AI for Everyone)

Put it together. Readers reach the verdict "machine" not from sentence rhythm but from three things: the text was never cut, the text decides nothing, the author isn't in the room. All three are properties not of a generator but of missing editing. A human on a deadline with no appetite for the job produces exactly the same artefact. Which is precisely why the accusation lands and misses at the same time.

An honest caveat against ourselves: this is not an experiment. We don't know whether a model actually wrote that course — RPATech never commented, and Bradley M.'s review remains one viewer's opinion. Our claim is narrower: we know what makes real paying humans deliver the verdict. This is data about the detector called "a person", not about the fact of generation.

What people mistake for traces of AI — and why they're wrong

Now look at the canonical list of "signs" and watch it fail to describe a machine. Here's what typically gets presented as evidence:

  • Suspicious smoothness. Every sentence roughly the same length and complexity, the rhythm never varies.
  • Everything in threes. Lists of exactly three items, symmetrical paragraphs, sections of equal size.
  • The "not X, but Y" construction. "This isn't just a tool, it's a way of thinking" — five times a page.
  • Zero specifics. No names, dates, numbers or lived examples. True of anything and therefore useful for nothing.
  • Marker vocabulary. "Landscape", "unlock the potential", "in today's digital era", "it's important to note".

Not one of these separates a machine from a human. They separate formulaic writing from living writing — and humans write to a template too: when tired, cautious, revising from a study guide, or working in a second language. This list will happily "convict" a tidy person. And conversely: a model's output under a good prompt shows none of them — ask for specifics, a real example and a stance, and the "machineness" evaporates (you can see how in the formula for a strong prompt).

Test that right here. The prompt below makes the model write the same paragraph twice — once "machine-like", once alive — and then rule on its own output. Both versions come from the machine, so any detector verdict of "AI / not AI" is wrong by construction half the time:

Write the same 120-word paragraph about why remote onboarding fails for new hires, twice.

Version A: the way a tired person writes to fill a page — no names, no numbers, no cases, every sentence roughly the same length.

Version B: with one invented but specific company, one concrete first-week failure, one number, and sentences of wildly uneven length.

Then answer in three bullets: which version a perplexity-based AI detector would flag as machine-written, why, and whether that verdict would be correct given that you wrote both.

How a detector works inside — and why it lies

An AI detector doesn't "recognise" text. It measures statistical properties — how predictable each next word is (perplexity) and how unevenly sentence complexity is distributed (burstiness). The logic: a model by construction picks the likely continuation, so its text is on average more predictable and more even. The logic is sound. The problem is what follows from it.

  • False positives on humans. Someone who writes plainly and predictably looks statistically like a model. A narrower vocabulary yields low perplexity — and a narrow vocabulary belongs to a schoolchild, a cautious lawyer, and anyone writing in a language they learned rather than absorbed.
  • False negatives. Ask the model to write with more variety, or lightly rewrite by hand, and the detector goes quiet. It catches predictability, not provenance.
  • No calibration. The percentage looks scientific but can't be checked: detectors publish no transparent method, and accuracy drops on text unlike their training data. "An 87% probability" and "our instrument emitted the number 87" are different statements — you're sold the first and given the second.
  • Short texts are guesswork. A single paragraph simply doesn't carry enough statistics.
  • A race the detector loses by definition. Every improvement to a model makes its output resemble average human writing more closely. The detector isn't catching up with a target — the target is leaving.

The conclusion is blunt but necessary: a detector's percentage cannot be used as grounds for an accusation — not at a university, not in an editorial office, not in hiring. It's a tool with an unknown error rate, unevenly distributed, that lands hardest on people already in a weaker position.

On second-language writers: what we can prove and what we can't

The claim "detectors punish people writing in a second language" sounds right, and the perplexity mechanism predicts it directly. But we have no measurement of the detectors themselves, and we're not going to invent one.

What we do have is an adjacent, solid fact. Our market research leans on a randomised controlled trial by Bälter, Kann, Mutimukwe & Malmström (Applied Linguistics Review 15(6): 2373–2396, 2024), n = 2,263 Swedish students, one online programming course delivered in two languages:

  • Dropout in Swedish: 57%. In English: 71% (φ = 0.2; p < 0.00001).
  • Mean score among those who stayed: 16.9 versus 14.3 — a negligible difference (δ = 0.085).

Read it carefully: English didn't damage comprehension — it made people leave. And, most relevant here, self-rated fluency did not protect: students who considered their English good dropped out just the same.

That says nothing about detectors directly. Here's what it does say: the cost of working in a non-native language is real, measurable, and invisible to the person paying it. A detector that penalises low perplexity is one more layer of that same tax, applied to the people already paying it. That's an argument, not a proof, and I'm not passing it off as the second.

The evidence that actually works: check the facts

The strongest thing we found is a review where someone didn't guess — they caught it:

"Stable diffusion, Dalle and Midjourney are not GAN architecture powered. They're diffusion models."
— Jose C., 12.05.2026, 1★ (Justin Barnett's course)

The error sat not only in the lecture but in the quiz — meaning the wrong answer was scored as correct. That is evidence. Not "87% probability" but a specific false claim you can put on the table, one the author can't explain because they didn't come up with it.

No detector would have found any of this — the lecture's prose is flawless. A person who knew the subject found it. The rule is simple: if a text contains facts, check the facts, not the style. A link that doesn't resolve, a confused architecture, a number with no source — those are hard evidence, and they're also the fingerprint of an AI hallucination.

The prompt for that check runs right here:

Check these four claims and mark each TRUE, FALSE, or UNVERIFIABLE-WITHOUT-SOURCES, with one sentence of reasoning each:

1. Midjourney, DALL-E and Stable Diffusion are built on GAN architecture.
2. Perplexity, in language modelling, measures how predictable a text is to a model.
3. A text scored "100% AI" by a detector proves the author used AI.
4. Watermarking can detect the output of any model.

Then say which of these four errors you would expect to find in a hastily made online course, and why those specifically.

What about watermarks?

It's technically possible to embed an imperceptible statistical mark during generation and check for it later: the model nudges its word choices by a secret rule, and the nudge can be recognised afterwards. Google DeepMind promotes such an approach under the name SynthID. The method's limits are structural, not implementation details:

  • The mark exists only where someone put it. A model without watermarking — including any model run locally — simply isn't in the game.
  • The mark doesn't survive everything. Paraphrasing, round-tripping through translation, or dense manual editing wash the statistical nudge out.
  • Nobody is obliged. There's no mandatory industry standard, and a voluntary one identifies the honest.
  • Absence of a mark means nothing. Even a perfect watermark answers "yes, this is our model" — never "no, this is a human".

It's a useful technology for platforms, not a universal detector for you.

What to do instead of using a detector

If you're a teacher, editor or manager, "did AI write this" is almost always the wrong question. The real ones are "does the author understand what they submitted" and "does the text do its job". Both are testable.

  • A conversation. Two follow-up questions about the content answer everything no detector can.
  • The process. Drafts, version history, interim submissions. You see the path, not just the result.
  • A task AI can't close alone. Local context, personal experience, data from the session, a tie to a specific case.
  • Fact-checking. See Jose C. above: one checkable claim beats ten percentage points of confidence.
  • Open rules. Say in advance what's allowed: "AI is fine for drafting and checking, not for the final conclusions." A transparent rule removes half the problem.

The obvious objection: a conversation doesn't scale, and a detector is one click. Fair objection. But look at what happens when you build the whole system around what scales:

"you send your assignments and immediatly you got your results: 100% correct. I am still speechless. I put all that effort in and have no idea whether I my answer was correct or not."
— Shinysheep, 21.09.2023, 1★ (Prompt Engineering, Vanderbilt — 698,444 enrolled)

"quizzes give unhelpful feedback for incorrect answers and just say 'watch the video again'"
— Cory Covino, 04.05.2024, 2★ (IBM)

"Videos are too short and superficial so you end up memorizing sentence by sentence to pass quizzes. Not a learning experience."
— Laurie J Phillips, 12.04.2024, 3★ (IBM)

That's the same disease as the detector, seen from the other side: automated checking that measures the wrong thing because measuring the right thing is expensive. The automatic "100% correct" and the "87% AI" are relatives. Both hand you a number instead of a judgement.

And it isn't platform laziness — it's a structural limit. To check a prompt, you need a model inside the lesson. Coursera doesn't have one: Vanderbilt's Prompt Engineering Specialization requires the student to separately buy a paid ChatGPT subscription and do the exercises elsewhere. The platform physically cannot see what came out, so it stamps an automatic pass. That's exactly why our AI course keeps the model inside the lesson — and the same reason the prompt blocks in this article have a Run button instead of asking you to copy them somewhere else.

If you're the one accused

It's a real and unpleasant situation: a detector produced a percentage, and the text is yours. What helps:

  • Version history. Google Docs, Word, any editor with autosave. It's the best evidence there is — the writing process is harder to fake than the result. Turn autosave on before you need it.
  • Drafts and notes. Keeping them is tedious; they're also what saves you.
  • Readiness to explain. If you understand your own text, a conversation closes the matter in two minutes.
  • A calm dismantling of the tool. Ask to be shown not a percentage but the specific passage that triggered it and the method behind the number. A detector gives you neither — and that absence is itself the answer.
  • Changing the question. "Ask me about the content" is the one sentence that converts the conversation from guesswork into a check. It's awkward for an accuser to refuse.

The prompt that prepares you for that conversation isn't "beat the detector" — it's "check whether you actually hold your own text":

Here is a student's paragraph:

"Urban green spaces improve mental health outcomes. Studies have shown that access to parks reduces stress and improves overall wellbeing. In today's digital era, it is important to note that city planners must holistically consider the landscape of urban development to unlock the potential of green infrastructure."

Do not judge whether AI wrote it. Instead, write five questions I could ask this author in a two-minute conversation that would reveal whether they understand what they submitted. For each question, say what a knowledgeable answer would contain and what a bluffing answer would sound like.

How to write with AI without sounding like a machine

And the most practical takeaway. If you use AI honestly — as a draft — "machineness" is removed not by trickery but by substance. Here's the quote worth remembering if you're hunting for a magic trick:

"I've studied best practices to avoid it generating content that sounds like AI, but I'm not having any success. … It is reported as 100% written by AI according to zerogpt."
— OpenAI forum, thread "Prompt to insert content without sounding like AI"

"the email writing tools just seems to strip out my personal voice making me sound like I'm writing unsolicited marketing spam."
— comment on Hacker News

Caveat: that's three sources from forums populated by advanced users, not a mass audience. The sample is small, and what comes out of it is a hypothesis, not a measurement. But the hypothesis fits everything else: the person in the first quote studied every "best practice against AI style" and failed. Because he fought the style. Style is the symptom. The disease is emptiness.

Here is a draft. Swap in your own if you have one.

"AI is transforming the way teams work. It is important to note that organisations must unlock the potential of these tools. In today's digital era, companies that fail to adapt risk falling behind their competitors in an ever-changing landscape."

Task: make it more concrete without making it longer.
Rules:
- delete every sentence that would still be true if you swapped "AI" for any other technology;
- every surviving claim needs an example or a number, or it gets deleted;
- alternate long sentences with short ones.
Finish with a separate list titled "Needs the author's experience": the places only a human can fill.

That last rule is the key. A text stops being machine-like exactly when it contains what the model doesn't have: your specific case, your number, your stance. That's why this article carries names, dates and star ratings — not for gravitas, but because they can't be generated.

The main thing

AI detectors aren't a truth sensor — they're a probabilistic guess with an unknown error rate that lands hardest on people who write plainly and in a second language. The "signs of machine style" separate formulaic writing from living writing, not a machine from a human. And the people who really do catch machine text — Bradley M. and Jose C. in our data — don't catch it by style: they catch it by text that was never cut, by a factual error, by the absence of an author in the room. So that's what's worth checking. The only reliable way to know whether a person stands behind a text is to talk to them about it.

Next on this topic: AI in school and cheating — what a teacher can actually do; AI hallucinations — where the non-existent links that serve as real evidence come from; the formula for a strong prompt — removing "machineness" at the input rather than the output; AI at work — on drafts and editing.

🧠Go deeper — in the courseNeural networks for beginners

FAQ

How accurate are AI detectors?

Accuracy is unknown and highly text-dependent. Detectors don't 'recognise' AI — they measure predictability (perplexity) and rhythm unevenness (burstiness), an indirect signal that plain-writing humans produce too. They publish no transparent method behind the percentage, and error grows on short, edited or second-language texts. You cannot verify a vendor's claimed accuracy from outside — that is the core problem.

Can I be accused of using AI when I wrote it myself?

Yes. Detectors err more often on texts by people who write plainly and predictably — including second-language writers: a narrower vocabulary yields low perplexity, which is exactly what the detector treats as evidence. The best insurance is your document's version history and drafts: the writing process is harder to fake than the result. The second is asking to be shown not a percentage but the specific passage and the method behind the number.

Can an AI detector be fooled?

Technically yes — light manual editing or asking the model to write with more variety often clears the flag. That's precisely why detectors don't work as proof: they catch stylistic predictability, not the fact of AI use. The reverse doesn't work, though: a poster on the OpenAI forum studied every 'best practice against AI style' and still got '100% written by AI according to zerogpt' — because he was fighting the style, not the emptiness.

What do people actually use to spot machine text?

We went through reviews of paid AI courses where students accuse the authors of machine authorship outright. Not one reviewer ran a detector or mentioned sentence rhythm. The wording is different: 'Written by AI, delivered by AI' (Bradley M., 16.04.2026, 1★) — about spoken words matching the slide verbatim; 'reading the prompter sometimes without even knowing the point' (Sachin S., 13.07.2026, 1★). They're spotting not a style but the absence of a human who would have cut the text and known why it exists.

What should a teacher do instead of using a detector?

Check understanding, not provenance: two questions about the content, a requirement for drafts and version history, tasks tied to the session and to personal experience. And most reliably — check the facts. Our review data has a model case: a student caught a course claiming Midjourney and Stable Diffusion run on GANs ('They're diffusion models', Jose C., 12.05.2026, 1★), with the same error baked into the quiz. The lecture's prose was flawless — a detector would have found nothing.

Do watermarks on AI text work?

Partly, and not for the job you need. Google DeepMind promotes SynthID — a statistical mark embedded at generation time. But only models that implemented it carry the mark, it doesn't survive paraphrasing or dense editing, and there's no mandatory standard. The decisive limit is structural: even a perfect watermark answers 'yes, this is our model' — never 'no, this is a human'.