
How to research and fact-check with AI: deep research mode
In short: deep research mode is when AI doesn't answer instantly but runs dozens of searches on its own, opens pages and assembles a sourced report in 5–30 minutes. It saves hours — but it doesn't remove verification: the model can still misread a source or cite a page that doesn't exist. One rule: open the links yourself. What follows isn't theory. It's a walk through our own market study — what we dug up, what we re-checked, and what we honestly failed to get.
How it differs from ordinary chat with search
Ordinary chat with web search fires one or two queries and answers in seconds. Research mode works differently: it splits your question into sub-questions, searches each, reads what it finds, notices contradictions, searches for the gaps — and only then writes a structured report with footnotes. It's a "search → read → refine" loop, spun many times without you.
Hence the difference in feel: a normal answer is the opinion of a well-read companion; a deep research report is a draft analyst's memo. With a bibliography you can check. Or fail to check — and step on a rake.
Here's the thing worth grasping up front. The mode changes the economics of research, not its nature. Visiting forty pages used to cost you a day. Now it costs ten minutes of waiting. But the decision about which forty pages were worth visiting, and what follows from them, is still yours. The tool made collection cheap and did precisely nothing to verification. Verification stayed exactly where it was — which is why the rest of this article is mostly about verification.
Where you'll find it
- ChatGPT — Deep Research mode, typically on paid plans, with a monthly run limit.
- Gemini — Deep Research, built into the interface, shows its research plan before starting.
- Claude — Research mode with web search and multi-source work.
- Perplexity — search-first by design, with a dedicated research mode.
Names and limits change almost monthly, so memorise the principle rather than the brand: if a tool shows a source list and spends minutes rather than seconds, that's the mode. How the models differ overall is covered in ChatGPT vs Claude vs Gemini.
How to brief it so the report is useful
The classic mistake is throwing in a one-line topic ("the EV market in Europe") and expecting magic. The mode will run and hand you fluff. Constrain the task on four axes: the question, the boundaries (period, region, source types), the format and what to do with uncertainty.
Rather than describe that abstractly, watch it happen. Hit Run below and the model turns a vague topic into a real brief in front of you. Swap in your own topic afterwards.
You are an editor of research briefs. Here is a request exactly as a client sent it: "I need an overview of the EV market in Europe." That is a bad brief. Rewrite it into a working one, along four axes: 1. THE QUESTION. Break the topic into 4-6 concrete questions that can be answered with facts rather than reasoning. Bad question: "what are the market's prospects". Good question: "how has the EV share of new registrations moved by country over three years". 2. BOUNDARIES. Propose a period, a region, and a priority order for source types. Say explicitly which sources count as primary for this topic and which are retellings. 3. FORMAT. Describe the report structure and the citation rule: what must sit next to every number. 4. UNCERTAINTY. Write the section the researcher is obliged to fill in: where sources disagree, what could not be found, which claims are weakly supported. Finish with 5 "traps of this topic" — the typical numbers and claims that travel from article to article in this field without a primary source. Explain why each deserves re-checking.
That last axis — uncertainty — is the most valuable one and almost nobody asks for it. It turns the report from "confident text" into a map of what's known, showing where the ground is solid and where it's swamp. For request design fundamentals see what a prompt is, and worked examples in our prompt library.
The main trap: sources that don't exist
Language models invent sources. Not occasionally, and not only "bad models" — it's systemic: the model generates a plausible continuation of text, and a plausible URL with a plausible paper title is just as easy to produce as a real one. The mechanics are unpacked in AI hallucinations: why models make things up.
Research mode lowers the risk because it genuinely visits pages. It does not remove it. Three ways fabrication leaks into a finished report:
- A dead link. The source either moved or never existed — in the report both look equally respectable.
- The link works, the claim isn't in it. The most common and most treacherous case: the page opens, looks relevant, but the number from the report simply isn't there — the model filled in the emphasis.
- A chain of retellings. The report cites a blog, the blog cites a news piece, the news cites "a study" nobody has seen. Formally there's a footnote. Factually there's no primary source.
The uncomfortable but honest conclusion: a footnote isn't verification, it's a promise of verification. The verifying is on you.
Six cases from our own research
What follows are not textbook examples. We ran a study of the AI-course market ourselves (data pulled 17 July 2026): we scraped prices and course metrics out of an internal Udemy API, gathered the enrolment statistics on Coursera, and read what unhappy students actually wrote. It's a convenient body of material for showing every rake at once, because we stepped on all of them.
Case 1. The primary source versus "everybody knows"
Common knowledge says Udemy courses list at $200 and are permanently on sale at 90% off. Ask any model — it will repeat this confidently, because thousands of articles say so. We went into the raw API response (the price_text endpoint) and saw this:
You are an analyst who reads raw API responses and does not take
marketing at its word. Here is a fragment of the API response
for a course page (we pulled it on 17 July 2026):
{
"price": {"amount": 24.99, "currency": "EUR"},
"list_price": {"amount": 24.99, "currency": "EUR"},
"saving_price": {"amount": 0.0},
"has_discount_saving": false,
"discount_percent": 0
}
Your tasks:
1. State what this fragment proves and what it does not.
Keep the two strictly apart.
2. The common claim is: "this platform is permanently on sale;
a 200-euro course goes for 20." Does the fragment refute that
claim? Fully, or only partly?
3. Give three alternative explanations under which both the
fragment and the common claim could still be compatible.
4. Name three further fields or requests you would demand to
settle the question for good. For each, say which specific
doubt it removes.
5. Write the conclusion the way it could actually be printed:
one sentence, no overstatement, with an explicit caveat about
the limits of the data.
What we found: there is no discount. saving_price: 0.0, has_discount_saving: false, discount_percent: 0. The course costs €24.99, and €24.99 is also its full price. Of the twelve top courses we captured, eleven are priced at €19.99; one (The Complete AI Guide) at €24.99. No "was €199.99" anywhere.
That's the whole lesson about primary sources in one paragraph. A mass belief existed, had been published thousands of times, and sounded entirely plausible. No amount of deep research would have overturned it: the model would have found those same thousands of articles and dutifully summarised them. The only thing that refutes it is a machine-readable server response that somebody had to go and fetch. Deep research is superb at finding the consensus and nearly incapable of finding what the consensus missed.
So before you launch a run, ask yourself an unpleasant question: is my answer likely to be written down somewhere, or does it live in a database, a filing, an API, a price list, a court record? If it's the latter, the report will hand you a beautifully-sourced summary of what everyone believes, and you will have learned nothing.
Case 2. The number we couldn't get
Our own study contains a line I like better than any of the numbers in it: the exact price of ChatGPT Plus — unconfirmed, every OpenAI domain returned 403. We know that price roughly, the way everyone does. But the primary source refused to give it to us, so in our internal file it's marked as second-hand and banned from use as a fact.
The Udemy Personal Plan price carries the same mark: €20.00/month, €10.00/month on promo — taken off the landing page, because udemy.com returned 403 to our crawler. The number is in the report, but its provenance stands right next to it.
This is what deep research reports do worst. The model hates coming back empty-handed. If the primary source is unavailable, it will quietly substitute a figure from a retelling and say nothing about it. The distinction between "we verified this" and "we found a mention of this" vanishes without trace inside its prose. Demand a "what I could not find" section explicitly — it will never appear on its own.
And notice the shape of the failure. It isn't a hallucination. The number is probably right. What's missing is the label. An unlabelled correct number and an unlabelled wrong number look identical on the page, which is exactly why the label matters more than the number.
Case 3. An authoritative source with a broken method
While preparing to localise, we ran into the EF English Proficiency Index — the world's most-cited ranking of English proficiency. It puts Romania 11th in the world. Any AI report about languages will drag it in first: it's everywhere, it looks like statistics, media cite it.
We threw it out entirely. Here's why. Its sample is self-selected — EF says so itself in its methodology: the test is taken by people who decided to check their own English on a language school's website. The average test-taker is 26. Now compare sources with probability samples: Eurobarometer 540 (n = 26,523) and Eurostat AES 2022 both say Romania is the worst in the EU — 51.2% of the population speaks no foreign language at all, and 25% speak English. Eleventh in the world and worst in the EU is not a spread of estimates. Those are incompatible statements, and one of them was manufactured by a broken method.
Hence a rule worth more than any checklist: look at how a number was produced, not at who published it. Deep research ranks sources by authority and prominence — and a broken method is exactly what makes a source prominent, because it produces clean, quotable numbers. A careful survey reports ranges and confidence intervals; a self-selected quiz reports a league table. Guess which one wins the citation race.
Case 4. A study that doesn't say what it appears to say
The one peer-reviewed experiment we found on our topic: Bälter, Kann, Mutimukwe & Malmström, Applied Linguistics Review 15(6): 2373–2396, 2024. A randomised controlled trial, n = 2,263 Swedish students, an online programming course, randomly assigned to the Swedish or the English version.
The result: dropout of 57% in Swedish versus 71% in English (φ = 0.2; p < 0.00001). Mean score among those who stayed: 16.9 versus 14.3, effect size δ = 0.085 — negligible.
Now look at what happened there. On scores, there's effectively no difference. The tempting write-up is "language doesn't affect outcomes" — the number exists, the significance doesn't. But on dropout the gap is enormous. English didn't damage comprehension; it made people leave. Those are different claims, and you only see the second one if you look at what the study measured and on whom. Self-rated fluency, incidentally, offered no protection against dropping out.
Deep research reports break down exactly here, and reliably. They summarise the conclusion from the abstract rather than the construction of the study. An abstract is itself a retelling — just one written by the authors. If a claim matters to you, the two things worth reading are the method section and the table. Everything else in a paper is prose.
Case 5. A number that looks like an answer and isn't
Sizing the market by language, we hit Udemy's course counters: so many courses in Spanish, zero in Swedish. A lovely number, a ready-made slide. Except it measures supply, not demand. Zero Swedish courses is either an empty niche you should sprint into, or proof that there's no market there at all. One number, two opposite conclusions, and no data to choose between them. So that's what we wrote down: ambiguous.
A useful habit: before you put a number in a conclusion, ask which opposite decision that same number would also support. If it supports both, it isn't an argument — it's decoration. Deep research produces this kind of decoration in bulk, because a number with a source attached passes every superficial test there is.
Case 6. Effect size versus headline size
The localisation industry promises "+40–70%" growth. We went looking for independent confirmation and found exactly two measurements. eBay, on rolling out machine translation: +10.9% exports (Brynjolfsson, Hui & Liu, Management Science 65(12), 2019) — and the effect declined as item price rose. A Wikimedia A/B test: +3.64%. The "+40–70%" promise has nothing independent behind it.
The gap between 3.64% and 70% isn't a disagreement about a number. It's the difference between a measurement and an advertisement. And note the detail about the effect shrinking as price rises: it exists only in the paper and disappears from every retelling, because it ruins the headline. Caveats die first — they're the first thing amputated in a chain of retellings, long before the number itself gets distorted.
Verify a report in 5 minutes
- List the claims something depends on. Usually 3–7 per report, not forty. The rest is context, where an error costs nothing.
- Open the link for each one. Does it load? Then Ctrl+F for the number or keyword. Not there? The claim is unsupported. Full stop.
- Walk back to the primary source. If the footnote leads to news about a report, find the report. Numbers mutate on the way and caveats fall off.
- Check the date. A "fresh" report is easily assembled from five-year-old articles if you didn't bound the period.
- Check the method, not the brand. How was the number produced? Who was asked, and how were they chosen? Our EF case is exactly this.
- Ask for the opposite. Run a second pass: "find arguments against conclusion X and data that contradicts it." If nothing turns up at all, the search was probably shallow.
The first five minutes pay for themselves by stopping you quoting a fabrication. The sixth step pays for itself by stopping you quoting a truth that means nothing.
Here's a ready-made auditor for steps 1–3. The prompt is self-contained — the report is baked into it, so run it as is.
You are a fact-checker. You have been handed a fragment of an AI report. Your job is not to check the facts (you have no internet) — it is to build a verification plan and rank the claims by risk. REPORT FRAGMENT: "The online AI-course market is booming. According to industry estimates the market will reach $47bn by 2030 (source: a marketing agency blog, 2025). Courses sell at an average discount of 87% — platforms are practically giving them away. The largest course has attracted over 400,000 students. Experts agree that video remains the preferred format for most learners. Research has shown that localising content raises conversion by 40-70%." Do four things: 1. DECOMPOSE. Write out each claim on its own line. Label each one: measurable fact / forecast / estimate / restated opinion / marketing cliche. 2. RISK. Give each claim a risk of being fabricated or distorted: high / medium / low. Justify in one phrase. Pay particular attention to numbers with no date and no method, and to the constructions "experts agree" and "research has shown". 3. PLAN. For the three riskiest claims, write down which specific primary source would settle the question, and the single request you would make to it. 4. REWRITE. Rewrite the fragment so that only what can be stood behind remains, and everything else is explicitly flagged as unverified. Invent nothing: where there is no data, say so.
When your sources are people
Half of our study isn't numbers, it's reviews. We collected 26 quotes from unhappy students verbatim, with authors and dates (plus two from forums, which we count separately, because a forum post is not a course review). On video courses, for example:
"why read straight from the slide? I can do that. This was not a helpful course at all" — Janie I., 02.07.2026, 1★
"50% of this course is reading script like a robot from the slides. …they are just reading text from the slides which you can also you from any good website" — Shashank T., 13.04.2026, 1★
What matters is what we did next. We did not write "students are massively dissatisfied". We counted the share of negative reviews (≤3.5★): Generative AI for Beginners — 9.61%; The Complete AI Guide — 10.36%. So roughly one student in ten. The quotes are vivid, but they illustrate a minority, and that has to stand right next to them in the text.
The second honest caveat: on Coursera, filtering reviews by star rating happens in the browser — the server HTML always returns the first page. Which means our quotes from there come from the default page any visitor sees, not from a systematic sample. That weakens the claim, so we wrote it down.
The rule: a quote is evidence that an opinion exists, not evidence of how common it is. A deep research report stitched together from reviews always reads like a verdict — because shouting is louder than silence, and people write negative reviews far more readily than neutral ones. Demand the denominator. If the report gives you five furious quotes and no percentage, it has told you nothing about the world; it has told you that angry text is easy to find.
Telling a good report from a pretty one
Deep research reports almost always look convincing: headings, bullets, footnotes, an even analytical tone. Beauty of form says nothing about quality of substance.
- How many distinct sources back the key claims. If the whole report rests on two pages, it isn't research — it's a retelling of two pages.
- Whether source types vary. A report built only from blogs and media is weak. One with primary material (raw API responses, statistics, documentation, filings, papers) is strong.
- Whether dates are stated. An undated claim in a fast-moving topic is close to useless.
- Whether the method is stated. Who was asked, how many, how they were selected. Without that, "a survey showed" is a turn of phrase, not evidence.
- Whether uncertainty is admitted. A good report says somewhere "the data diverges" or "no confirmation found". A report where everything is unambiguous is suspicious: reality doesn't look like that.
- Whether facts are separated from conclusions. "Sales grew X%" and "the market is heading for a boom" are different classes of statement.
Stress-test your own conclusion with this. It runs without internet — it attacks the logic of the text rather than the facts.
You are a devil's advocate. Your job is to try to destroy the conclusion below. Do not be polite and do not seek balance. CONCLUSION: "Video loses to text for teaching digital skills to adults, because video cannot be skimmed at the reader's own pace, whereas text comes with a transcript by definition." THE EVIDENCE THE CONCLUSION RESTS ON (real reviews of top-selling video courses, collected 17 July 2026): - "It is too fast for a beginner... until you try to see and understand it the next screen pops up" — Amit A., 29.04.2026, 2 stars, on The Complete AI Guide - "I am very frustrated that I can not easily access a transcript to reference later. The instructors are talking so quickly" — Vince D., 16.03.2026, 1 star, course not recorded - share of reviews at 3.5 stars or below for The Complete AI Guide: 10.36% Do five things: 1. Name the weakest point in this argument. 2. Explain what these reviews prove and what they do not. Separately: what does the 10.36% figure do to the strength of the conclusion — support it or undermine it? 3. Build the strongest counter-argument: under what conditions does video beat text, and for whom specifically? 4. Invent the data that would refute the conclusion outright. Describe the study that would produce it. 5. Rewrite the conclusion so that it is honest about the evidence available. It will probably become narrower and more boring. That is fine.
On the chain of retellings specifically
The quietest failure of the lot. The link is live, the page opens, the number is on it — and there is no primary source anywhere, because everyone is citing everyone else in a circle. Deep research is helpless here: formally, every footnote is valid. Nothing in the report's structure can tell you that the whole edifice rests on a press release that cited nothing.
Here's that same "+40–70%" laid out step by step. The AI report claims "localisation raises conversion by 40–70%" and links to a translation platform's blog. The blog, 2025: "studies show an uplift of 40–70%" — linking to an industry association's press release. The press release, 2024: "in a survey of market participants, companies report growth of up to 70%" — with no link at all.
Watch where it broke. Between the press release and the blog, "companies report growth of up to 70%" became "studies show growth of 40–70%". Self-reporting by interested parties turned into "studies", "up to 70%" turned into a range of "40–70%", and the lower bound came from nowhere. By the third retelling it reads like a measurement. And not one link in that chain is formally false — which is exactly why machine-checking footnotes doesn't save you here.
Meanwhile the real measurement sat right there and never made it into the chain: eBay's +10.9%, with the effect declining as price rose. The primary source lost to the retelling because it offered a worse number.
When research mode is the wrong tool
- You need one fact. A rate, a date, a definition — plain search is faster.
- The topic is under a day old. Indexing and access to fresh pages are uneven; read breaking news with your own eyes.
- The data is behind a wall. Paid databases, internal documents, paywalled papers — the model can't get in and, worse, will build an overview from abstract retellings. Our 403s from Udemy and OpenAI are exactly this class of problem.
- You need what the consensus missed, not the consensus. The Udemy discount story marks the boundary of the genre: what isn't written down anywhere, the mode cannot find, by definition.
- The stakes are medical or legal. A report is fine as a terrain map and a source list, not as the basis for a decision.
How to fit it into your workflow
The working pattern: deep research maps the field (who says what, and where) → you spend 5 minutes verifying the key claims → you ask the model to rewrite the report keeping only what's confirmed and to list separately what was dropped. The output is a short text you can stand behind line by line.
That is precisely how the file behind this article is built. It has a section called "what we don't have": the ChatGPT Plus price (403), the Coursera review filter (client-side), Reddit (unavailable to our crawler), and a note that our three sources on one particular question are advanced users, not the beginners we care about — a small sample, and we said so. That section is unpleasant, because it publicly lists our holes. It is also the only reason the rest of the numbers deserve any trust at all: with an author who admits three gaps, the fourth is easier to find than with an author for whom everything neatly added up.
Ask your report for the same thing. Not "are you sure?" — the model will say yes. Ask: which claims here would you drop if you had to bet money on each one? The answers are usually specific, and usually the same three claims you were about to build a decision on.
Keep basic hygiene in mind too: don't feed sensitive or other people's data into research mode — privacy when working with AI applies exactly as it does in ordinary chat. If you're still getting comfortable with chat itself, start with the beginner's guide to ChatGPT. And if your report went wrong in its reasoning rather than its facts, that's a story about how models compound errors across steps.
FAQ
What is deep research mode in plain words?
It's a mode where the AI runs many searches on its own, opens pages and assembles a sourced report. It spends minutes instead of seconds and returns a structured text with references rather than a short answer from memory. What changes is the economics of collecting information, not its reliability: the verifying is still yours.
Can I trust the links in the report?
Not on faith. Models can invent plausible links and paper titles, and even more often attribute to a real source something it doesn't say. Open the link and find the claim on the page with a text search. If it's not there, treat the claim as unsupported. Check separately whether the footnote leads to a retelling of a retelling: a live link and a primary source are different things.
Doesn't research mode solve hallucinations?
It reduces them but doesn't solve them, and it adds a failure of its own: it is superb at finding the consensus and nearly incapable of finding what the consensus missed. Our example: the popular belief in permanent Udemy discounts is refuted only by their raw API response (saving_price: 0.0, discount_percent: 0), not by anything written on the web — in the texts, the discount 'exists'.
How do I tell a source is bad when it looks authoritative?
Look at the method, not the brand. The EF English Proficiency Index is the most-cited English proficiency ranking and puts Romania 11th in the world. But its sample is self-selected (the test is taken by visitors to a language school's site, average age 26), while Eurobarometer 540 (n = 26,523) and Eurostat AES 2022 show Romania is the worst in the EU: 51.2% speak no foreign language at all. Those are incompatible statements, and one was produced by a broken method.
The report cites a peer-reviewed paper — is that enough?
No — what matters is what the paper actually measured. From our study: the Bälter et al. RCT (Applied Linguistics Review 15(6): 2373–2396, 2024, n = 2,263) found a negligible gap in scores (16.9 vs 14.3, δ = 0.085) but an enormous gap in dropout — 57% vs 71% (φ = 0.2; p < 0.00001). A summary written off the abstract easily turns that into 'language has no effect'. It has one; just not on the thing being looked at.
How do I make the report admit what it didn't find?
Demand the section explicitly in the brief: where sources disagree, what could not be found, what is weakly supported. It won't appear by itself — the model hates coming back empty-handed and will silently substitute a figure from a retelling. In our own file such spots are marked in place: the ChatGPT Plus price is unconfirmed (OpenAI returned 403), the Udemy Personal price was taken off the landing page (udemy.com returned 403). The number is there — and so is its provenance.
Can I use such a report for study or work?
As a draft and a source map — yes, that's its strength. As finished text under your name — no: you'll be accountable for every number. Many schools and companies also have disclosure rules for AI use; check them in advance.