Neurocourse
AI in school: spotting AI-written work and setting tasks AI can't do for you

AI in school: spotting AI-written work and setting tasks AI can't do for you

18 min read

In short: you cannot reliably prove a text was written by AI today — detectors score how predictable a text is, not where it came from, and they hit hardest the students writing in a second language. The workable teacher strategy isn't catching, it's redesigning assignments: tie them to personal experience, to your class's own data, and to a short oral defence. To show what happens when you bet on automated assessment instead, we audited the world's largest AI courses. They already made that exact mistake, and their students described it word for word.

Why detectors are a bad foundation

The temptation is obvious: paste the text, get a percentage, conversation over. The problem is there's no evidence behind that number. Detectors don't "see" a text's origin — they estimate how predictable and smooth it is. And smooth, predictable writing isn't unique to models.

  • False positives for non-native writers. A student writing in a second language uses simpler constructions and standard phrasing — exactly the signals a detector reads as "machine-like". This is reasoning from how the tool works, not our measurement: we have not tested detectors on classrooms and cannot put a number on the error rate. But the direction of the error follows directly from the mechanism.
  • False positives for careful students. Tight structure, neutral tone, no errors — and honest work scores a high "AI percentage".
  • Evasion is trivial. Ask the model to write simpler and livelier, rewrite a couple of phrases by hand, and the score drops. It's a filter in reverse: it catches the honest and lets through anyone who tried to hide.
  • A percentage isn't evidence. A number on an opaque scale is not a probability. You cannot present it to a student as proof, methodologically or ethically.

The conclusion is blunt but important: you cannot build an accusation on a detector. At most it's a reason to look more closely and to talk — and that talk starts with a question, not a verdict. More on how these tools work: AI text detectors: do they work.

Our research: what happens when you hand assessment to a machine

On 17 July 2026 we pulled data on the largest AI courses on Udemy and Coursera — student counts, ratings, share of negative reviews, and the reviews themselves. We were mapping competitors and found a ready-made teaching case instead. Because the online platforms have already run the experiment schools are only approaching: they replaced human marking with automated marking. Their own students wrote up the result, verbatim.

The scale, so you know who we're talking about: Google AI Essentials has 1,876,929 enrolments. Generative AI for Everyone from DeepLearning.AI (Andrew Ng): 814,083. Vanderbilt's Prompt Engineering for ChatGPT: 698,444. IBM's Generative AI: Prompt Engineering Basics: 654,256. Google Prompting Essentials: 365,425. Vanderbilt's Prompt Engineering Specialization: 138,967. That's millions of people, and it includes your senior students and their parents.

"100% correct" — the verdict on automated marking

Here is a review worth printing out and pinning up in the staff room:

"you send your assignments and immediatly you got your results: 100% correct. I am still speechless. I put all that effort in and have no idea whether I my answer was correct or not." — Shinysheep, 21 Sep 2023, 1★ (Vanderbilt, Prompt Engineering for ChatGPT)

Someone submitted work, instantly got "100% correct", and walked away with no idea whether they had done well or badly. This isn't a platform bug. It's a structural limit: there is no model inside the Coursera lesson, so there is physically nothing there to evaluate a student's prompt with. The only thing the system can do is accept the file and stamp it.

What that means for your classroom: a grade awarded without a human reading the work tells the student nothing. Zero information. And the same is true in the other direction — a detector's percentage is also awarded without a human reading the work. The only difference is the sign: one is an automatic A, the other an automatic accusation. Both are worthless for the same reason.

Here's what happens to learners when marking is automatic and feedback is absent:

"Videos are too short and superficial so you end up memorizing sentence by sentence to pass quizzes. Not a learning experience." — Laurie J Phillips, 12 Apr 2024, 3★ (IBM)
"quizzes give unhelpful feedback for incorrect answers and just say 'watch the video again'" — Cory Covino, 4 May 2024, 2★ (IBM)
"All the coding is done in the labs for you. You won't have to debug anything or figure anything out, just press shift-enter." — Cornelius Griggs, 1★ (Generative AI with LLMs)

Three reviews, three symptoms every teacher recognises: cramming to pass a quiz, feedback that amounts to "read it again", and an exercise where everything is already done for you and one keypress finishes it. None of the three is about AI. All three are about assessment ceasing to be a conversation. AI isn't the cause here, it's the accelerant: wherever a task can be closed with a button, a student will find the button.

"He just reads off the slide" — and what that means for your lesson

The most common theme in the negative reviews we collected (§4.1 of our research) isn't difficulty and isn't price. It's the feeling that the instructor is unnecessary:

"why read straight from the slide? I can do that. This was not a helpful course at all" — Janie I., 2 Jul 2026, 1★ (Justin Barnett, 152,821 students)
"50% of this course is reading script like a robot from the slides. …they are just reading text from the slides which you can also you from any good website" — Shashank T., 13 Apr 2026, 1★ (The Complete AI Guide)
"Written by AI, delivered by AI. The slides are way too crowded to be useful and it really doesn't help to have the bot read them out word-for-word" — Bradley M., 16 Apr 2026, 1★ (RPATech, 118,803 students)

The lesson for school is unpleasant but useful: if the content of your lesson can be read off a slide, AI reads it better — faster, at any hour, repeated as many times as needed. The part of teaching that consists of transmitting information has already been taken from you, and it isn't coming back. What remains is the part no model does: noticing that this particular child is stuck on this particular step, asking them a question, catching an error in the reasoning rather than in the prose. That is exactly where assignment redesign pays off.

"Only theory we already know"

A second cluster of reviews explains why a student reaches for the machine in the first place:

"The content of this course is incredibly simple to the point of uselessness, it is essentially stating that generative AI exists and listing a bunch of example models." — Nicholas Munford, 1★ (Generative AI: Introduction and Applications)
"Didnt meet expectations, only theory is discussed which we already know." — Aditya Nagavolu, 19 Nov 2023, 1★ (Generative AI for Everyone)

A task that requires nothing but restating the known will be closed the cheapest way available. This isn't about morals, it's about the economics of effort. "Write an essay about friendship" honestly announces: nothing personal is wanted here, nothing checkable, nothing yours. The model hears that announcement first.

Content rots — and the quizzes rot with it

One more cluster, important for anyone planning to teach a lesson about AI:

"Stable diffusion, Dalle and Midjourney are not GAN architecture powered. They're diffusion models." — Jose C., 12 May 2026, 1★ (Barnett)
"The content is mostly from 2023.I invested my 41 hours and Im learning content which is from 2023. Very disappointed." — Harsh A., 9 Apr 2026, 1.5★ (The Complete AI Guide)
"Most content is from 2024. This course is not bad for its time, but just too dated now." — Martin F., 27 May 2026, 2★ (Generative AI for Beginners — while the listing claimed "updated 18 Apr 2026")

The first one is the sharpest. The error lives in the videos and in the tests — meaning the student has to pick the wrong answer in order to score the point. The practical takeaway for a lesson: don't build it around interface screenshots and a list of fashionable model names. That decays by next term. Build it around what doesn't decay: how to check a claim, where a number came from, why a given question can't be answered with one prompt.

On the share of negative reviews — and an honest caveat against ourselves

We computed the share of ratings at or below 3.5★ for the top Udemy courses: 9.61% for Generative AI for Beginners (409,492 students, 4.53 rating), 10.36% for The Complete AI Guide (376,845 students, 42 hours, 545 lectures), and at least 4.00% for Mike Wheeler's Prompt and Context Engineering 101 (84,942 students, 4.31 — the worst rating in our sample). On Coursera the negative share is markedly lower: 1.5–3.5%.

The caveat you need in order to read those numbers at all: Coursera's star filter runs in the browser, and the server always returns the first page of reviews. So the quotes above come from the default view any visitor sees, not from a full sample of the negatives. The gap between "1.5–3.5%" and "9.61%" cannot be read as "Coursera teaches better": the platforms have different audiences, different costs of entry, and different ways of collecting reviews. We report it as a measurement, not as a conclusion about quality.

What works instead

Shift the focus: not "how do I catch them" but "how do I make copying pointless". AI is excellent at depersonalised tasks — "write an essay about friendship", "make a report on volcanoes". It's poor at tasks tied to a specific person, a specific moment, and specific data that isn't on the internet.

Six ways to redesign an assignment

  1. Tie it to personal experience. Not "an essay about friendship" but "a moment in your life when you realised friendship isn't always fun — describe the place, the time and what exactly you felt". A model will invent it, and the invention collapses at the first follow-up question. Common mistake: adding the word "personal" to the old wording and calling the task redesigned. "Write about your personal experience of friendship" gets closed just as easily as the original.
  2. Tie it to the lesson. "Use three examples we covered on Thursday and explain why the fourth example in the textbook doesn't fit here." AI doesn't have your Thursday. The power is in the second half: explaining why something doesn't fit is much harder than listing what does.
  3. Give them your data. A class survey, lab measurements, a table you built yourself. The student works with AI on top of that data — which is a legitimate skill, not cheating. Bonus: the whole class has the same data and different conclusions, and the disagreements become the lesson.
  4. Ask for the process, not the product. A draft, an outline, a list of rejected ideas and why they were dropped. AI writes the final text in a second; it can't produce the history of your doubts. A caveat against ourselves: process can be faked too, with "invent three rejected ideas". So process doesn't work on its own — it works paired with an oral defence.
  5. Oral defence. Three minutes, two questions about the text: "explain this paragraph in your own words" and "what would you change if…". It's the fairest and fastest way to know whose work it is, and it needs no technology at all. It is also precisely the thing whose absence made Shinysheep write "I am still speechless".
  6. Legalise AI and demand reflection. "You may use AI. Attach your prompts, attach what the model produced, and write what you changed and why." Copying turns into analysis — and, honestly, into a better lesson. As a bonus, you see who in the class can spot a model's errors and who copies without looking.

Prompt 1: audit an assignment for copy-ability

Start with a diagnosis, not a redesign. This prompt runs as it stands — the assignment is already filled in; swap in your own once you've seen how it behaves.

You are a curriculum designer. Below is a school assignment.
First, complete it the way a student outsourcing to AI would:
write a 150-word answer. Then assess yourself honestly.

Assignment (history, age 15): "Write a 500-word essay on the
causes of the First World War."

After your answer, give an analysis:
1. A copy-ability score from 1 to 10, where 10 means
   "submittable with no edits".
2. What exactly in the wording let you do that.
3. Three reworked versions of the assignment that would make
   your answer above useless. For each, state: what ties the
   work to the student's personal experience, to the class's
   own data, or to a specific lesson; which process artefact
   the student submits; two questions for a 3-minute oral
   defence.
4. For each version, try to defeat it and say where it
   still leaks.

Don't propose bans or detectors. Preserve the learning goal.

Point 4 is the one that matters. You are asking the model to attack its own proposal. Without that step it hands you three elegant versions, two of which are copied as easily as the original.

Prompt 2: a hallucination demo for the classroom

Five minutes that beat ten lectures about honesty. Run it live and show the screen.

Tell me in detail about the 1934 novel "The Salt Harbour
Letters" by Miriam Aldecoa. Summarise the plot, name the main
characters, and explain why critics of the period argued
about its ending.

Then, in a separate block headed "Verification":
- say plainly whether this book actually exists;
- if you are not certain, explain where the details you wrote
  above came from;
- list the signs by which a reader could have suspected
  the answer was invented.

The book doesn't exist. Two outcomes are possible and both make a good lesson. Either the model confidently invents a plot and a cast (and then exposes itself in the "Verification" block), or it honestly says it has no record of the book. The second outcome doesn't spoil the demo — ask the class why the model hesitated here and doesn't hesitate when it produces a date or a citation in their homework. The mechanics: AI hallucinations.

Prompt 3: fifteen variants of a problem in one minute

The flip side: the routine that eats your evenings. This prompt is self-contained and runs as it stands.

Produce 15 variants of a word problem for 11-12 year olds
using the same underlying method, with different numbers
and different scenarios.

Model problem: "A school library holds 240 books. 35% of them
are fiction and the rest are textbooks. During the year the
library buys 60 more textbooks. What percentage of the
collection is fiction now?"

Requirements:
- choose numbers so answers need at most two decimal places;
- scenarios must be varied and everyday, with no repeats;
- make five variants one step simpler and five harder
  (add one piece of data that isn't needed to solve it);
- finish with a separate answer table showing the full
  working for each.

Check your arithmetic before printing, and mark with an
asterisk any variant you are not confident about.

That last line matters more than it looks: these models calculate worse than they write. Check the answers anyway — but narrowing the check to the flagged variants already saves you an evening.

Prompt 4: oral defence questions

You are a teacher. Below is a paragraph from a student paper.
Write five questions for a three-minute oral defence that
cannot be answered by simply reading the paragraph aloud.

Paragraph: "The Industrial Revolution changed the structure
of society. A new class appeared — industrial workers — whose
conditions were harsh: long working days, child labour, no
insurance. At the same time cities grew, which created
problems of sanitation and housing."

For each question, state:
- what it actually tests (grasp of cause, ability to give an
  example, capacity to notice an oversimplification);
- what an answer from someone who understood sounds like;
- what an answer from someone who submitted someone else's
  text sounds like.

Separately, name three places in the paragraph where the
author smoothed something over, and how to ask about each.

That last instruction is the most useful part. Smoothed-over places exist in every text, and a question aimed at one stops both the student who copied and the student who wrote it themselves without thinking. That gap — between "text submitted" and "person understood" — is the whole thing you're trying to measure.

Language of instruction: the number that changes how you see non-native students

If your class includes children for whom the language of instruction isn't native, here is a peer-reviewed result worth knowing. Bälter, Kann, Mutimukwe and Malmström (Applied Linguistics Review 15(6): 2373–2396, 2024) ran a randomised controlled trial with 2,263 Swedish students: the same online programming course, delivered in Swedish and in English.

  • Dropout in the native language: 57%. Dropout in English: 71% (φ=0.2; p<0.00001).
  • Mean score among those who stayed active: 16.9 vs 14.3 (δ=0.085 — a negligible effect size).

Read that carefully: the language did not damage comprehension — those who stayed scored nearly the same. It made people leave. One more detail: self-rated fluency did not protect against the higher dropout.

Two consequences for school. First: when a student writing in a second language produces simpler, drier prose, that is not a sign of laziness and not a sign of a model. It's the price they're paying for working in someone else's language at all. Second: a detector that reacts to simplicity will systematically hit exactly that student. The combination — higher dropout risk plus higher false-accusation risk — makes them the most exposed person in your room. That is the one category where a teacher's error costs the most.

How to know the redesign worked

A simple honesty test: take your new assignment, paste the whole thing into a chat, and see what comes out. If the model produces text a student could submit as is, you haven't redesigned the task — you've just made it longer. If the model produces a plausible hollow shell that collapses at "so where did these numbers come from?", you're on the right track.

Second signal: a student who did the work honestly should be able to defend it for three minutes with no preparation. If they can't, either they didn't understand it, or the assignment tests something other than what you intended. Both findings are useful, and neither needs a detector.

Third signal, the strictest one: two students doing your task honestly should hand in different work. If honest submissions are indistinguishable from each other, the task measures compliance rather than thinking — and a machine will close it.

Five common mistakes in redesigning

  • Lengthening instead of redesigning. A "minimum 1000 words" rule only punishes the student who writes it themselves.
  • "In your own words" as an incantation. Models are excellent at "in your own words". The phrase is a wish, not a constraint.
  • Banning AI with no way to verify. A rule whose compliance you can't observe teaches exactly one thing: rules need not be followed.
  • Betting on an obscure topic. "There's definitely nothing about this online" is almost always false and takes a minute to disprove. What should be rare is the data, not the topic.
  • Accusing on a percentage. The most expensive one. A false accusation costs a whole class's trust for years; a missed case of copying costs one grade.

How to talk to your class about it

A flat "no AI" fails for the same reason banning calculators failed: the tool is in everyone's pocket, and students use it regardless of your opinion. What works is an agreement stated out loud.

  • Name the rules explicitly. Where AI is fine (finding ideas, explaining something unclear, spellcheck), where it isn't (submitting generated text as your own), and what must be disclosed if used.
  • Explain why. Not "because it's dishonest" but "because what you submit isn't text, it's your understanding; text without understanding is worthless to me and to you".
  • Teach them about hallucinations. Students genuinely don't know that models confidently invent facts, dates and sources. One live demo beats ten lectures — prompt 2 above runs straight from this page.
  • Show them the Shinysheep review. Seriously. "I put all that effort in and have no idea whether I my answer was correct or not" is exactly what a student feels when a grade arrives with no comment. A conversation about why human marking exists starts more easily from someone else's example than from your own gradebook.
  • Never accuse on a percentage. If a paper raises questions, ask them: "walk me through how you wrote this."

Where AI genuinely helps the teacher

The routine that eats your evenings automates surprisingly well.

  • Generate 15 variants of the same problem type with different numbers (prompt 3 above)
  • Rewrite a text to your class's level, or adapt it for a student who's struggling
  • Build a grading checklist and oral-defence questions (prompt 4 above)
  • Invent a 5-minute warm-up, debate or role-play
  • Draft feedback on a paper — which you then rewrite for that specific child

One hard boundary: don't paste student work with names, grades or any personal detail into a chat. That's children's personal data, and rules like the GDPR apply in full. Anonymise before sending, or use tools your school has approved. More: privacy when working with AI.

And a second boundary, which follows straight from our own data: a draft comment is a draft. The machine that awards a grade without a human reading the work is called "100% correct", and it teaches nobody anything. Don't repeat someone else's mistake on your own material.

The bottom line

Catching is pointless, detectors are unreliable, bans don't work. Something else does: assignments where the value is created by a human and AI is only a tool. The world's largest AI courses, counting students in the hundreds of thousands and the millions, have already shown where automated assessment ends up: "100% correct" and zero feedback. You have exactly one advantage over them, and it is enormous — you can ask a question and hear the answer. If you want to understand these models yourself first, start with the beginner's guide to ChatGPT, our article on prompts and the breakdown of hallucinations.

🧠Go deeper — in the courseNeural networks for beginners

FAQ

Can a detector prove a student used AI?

No. Detectors score how predictable a text is, not where it came from, and they err in both directions. Their output isn't evidence and can't justify an accusation or a lower grade. It's structurally the same thing as Coursera's automatic "100% correct", which made one student write: "I put all that effort in and have no idea whether I my answer was correct or not" (Shinysheep, 21 Sep 2023, 1★). A verdict issued without a human reading the work means nothing — whichever direction it points.

Why are detectors especially dangerous for students writing in a second language?

Because they react to simple constructions and predictable vocabulary — exactly how a person writes in a second language. How much that costs is measurable: in the Bälter et al. RCT (Applied Linguistics Review, 2024, n=2,263), teaching in English rather than native Swedish barely moved the scores of those who finished (16.9 vs 14.3, δ=0.085), but dropout rose from 57% to 71%. The language doesn't block comprehension — it pushes people out. A detector then targets precisely that vulnerable group.

What if I'm almost certain a paper was written by AI?

Start with a conversation, not an accusation: ask the student to explain a couple of paragraphs in their own words and describe how they got there. A three-minute oral defence tells you more than any percentage. And weigh the cost of error: a false accusation costs a class's trust for years, a missed case of copying costs one grade.

Which assignment can AI definitely not do for a student?

One that needs data not on the internet: personal experience with specifics, material from your particular lesson, your class survey results, your lab measurements. Plus anything requiring the process — drafts, rejected ideas, an explanation of choices. Simple test: paste your assignment into a chat in full. If the answer could be submitted as is, the task isn't redesigned.

Should I model an AI lesson on the popular online courses?

Carefully. We audited the largest ones (17 Jul 2026) and found two systemic problems. Content rots: "The content is mostly from 2023… I invested my 41 hours" (Harsh A., 9 Apr 2026, 1.5★, The Complete AI Guide), and one course carries a factual error inside its quizzes — "Stable diffusion, Dalle and Midjourney are not GAN architecture powered. They're diffusion models" (Jose C., 12 May 2026, 1★). And theory without practice lands badly: "only theory is discussed which we already know" (Aditya Nagavolu, 19 Nov 2023, 1★). Take the topic list from them, not the structure.

Can I upload student work to ChatGPT for grading?

Not with names, grades or personal details. That's children's personal data, and responsibility sits with your school and with you, not the service. If you want help grading, anonymise the text or use tools your school has officially approved. And keep the boundary: a draft comment is a draft, not a grade.