Why AI Invents Sources — and What Changes When It Works From Your Uploads

The essay is due tomorrow and the bibliography looks excellent. Twelve sources, properly formatted, plausible journals, plausible years, author names you half-recognize. You did what a lot of students do: asked for sources on your topic, formatted them, and moved on.

Two of them do not exist. Not "hard to find" — they were never published. The journal is real, the authors are real, the volume number is real, and the article is not. Your grader will find this in about ninety seconds, because checking a suspicious reference is the first thing an experienced marker does.

This is the single most damaging error in AI-assisted academic work, and it's damaging precisely because it doesn't look like an error. Here's the mechanism, and what actually changes it.

1. Why this happens, and why it's not a bug in one product

A language model generates text by producing what plausibly comes next. A citation is text. "Smith, J. (2019). Journal of Educational Psychology, 111(4), 623–641" is a highly patterned string, and producing a well-formed one requires no knowledge that the article exists — only knowledge of what citations look like.

This is why the failure is so convincing. A fabricated reference is structurally perfect. Real journal, plausible volume, plausible page range, an author who genuinely publishes in that area. Everything about it is right except the fact of it.

Three consequences worth internalizing:

It affects all of them. ChatGPT, Gemini and Claude are all capable of this, because it's a property of generating text rather than a flaw in one company's product. Asking a different assistant for a second opinion doesn't help — you get another plausible string.

Asking "are you sure?" doesn't work. The model will often apologize and produce a different plausible citation. You now have two references and no more information than before.

It gets worse the more obscure the request. Ask for well-known foundational papers in a large field and you'll mostly get real ones. Ask for "three recent studies on the effect of X on Y in Z population" — a narrow, specific, recent slice — and the probability that the exact paper you described exists drops sharply, while the probability of getting a confident answer does not.

2. The one habit that prevents all of it

Every reference in your paper should be one you have personally opened.

Not "found in a search result." Opened, and confirmed it says what you're claiming it says. That single rule eliminates the entire category, and it costs about ninety seconds per source.

The verification sequence, in order of speed:

  1. Search the exact title in Google Scholar or your library. A real paper appears immediately. A fabricated one returns nothing, or returns something with a similar title by different authors — which is its own trap.
  2. Check the DOI resolves. Paste it into doi.org. Fabricated DOIs either don't resolve or resolve to something unrelated.
  3. Open it and find your claim. People stop after confirming existence and skip this one. A real paper cited for something it doesn't say is still a citation error, and it's more common than outright fabrication.
  4. Check the year and the venue. Real paper, wrong year or wrong journal is the most common half-error, and it looks careless in exactly the way markers notice.

Step 3 deserves emphasis. Once students learn about fabricated citations they often over-focus on existence and under-focus on accuracy. "This source is real" and "this source supports my claim" are different checks.

3. Two ways of working, compared

The deeper fix is structural: change where the material comes from, so there's nothing to fabricate.

Asking a general assistant for sources Working from documents you uploaded
Where references come from Generated as plausible text Files you already possess
Can a source not exist? Yes — this is the core risk Not for uploaded material; the file is on your machine
Can a claim still be misattributed? Yes Yes — you still verify what it says
Good for Finding directions to search in Working with sources you've already gathered
Bad for Producing a bibliography Discovering literature you don't have yet

The row that matters is the second. When a summary is built from a PDF sitting in your own folder, the question "does this source exist?" is answered by the folder. What remains is the ordinary scholarly obligation of checking that the source says what you think — which you'd owe anyway.

This is the property behind working from uploads. AI Notes under AI Study builds notes from a PDF you provide, so what you're reading traces back to a file on your own machine rather than to a model's general recollection of documents like it.

Cheatsheet goes one step further in a way that's directly useful here: its source-reference setting can show the page or section each point came from, or include short quotes from your material. That turns a study artifact into something you can audit — you can see which page a claim came from and go check it, which is the discipline this whole piece is arguing for. Neither removes your responsibility to read the source. Both remove the failure mode where the source was never real.

Note what the table also says: general assistants are genuinely useful for the discovery half. "What are the main debates in this area, and which researchers are associated with each?" is a reasonable question, and the answers give you search terms. The rule is that what comes out of that conversation is a lead, not a citation. You go find the actual paper, open it, and cite what you read.

4. A workflow that keeps both

  1. Use a general assistant to orient. What are the debates, who's associated with which position, what terminology should I be searching. Fast, and low-risk because you're not citing any of it.
  2. Go find the real papers in your library database or Google Scholar, using those terms. This is the step that converts leads into sources.
  3. Upload what you actually obtained and work from there — notes per paper, and a compressed sheet across the set once you have several. Now everything downstream traces to files you hold.
  4. Verify each citation once, at the point it enters your document. Title, DOI, year, and the specific claim. Ninety seconds, done once, never revisited.
  5. Keep the PDFs. If a marker ever queries a reference, "here is the paper" ends the conversation immediately.

The order is the whole design: generate leads, obtain real documents, work from the documents, verify at the boundary.

5. Five habits that keep your bibliography clean

<br>

Frequently Asked Questions

Q1: Does ChatGPT make up references?

It can, and so can Gemini and Claude, because producing a well-formed citation doesn't require the article to exist — a citation is a highly patterned string, and generating plausible text is what these systems do. The output is convincing precisely because it's structurally perfect: real journal, plausible volume, plausible authors. Treat any reference you didn't personally open as unverified, regardless of which tool produced it or how confident it sounded.

Q2: How do I check whether a citation is real?

Search the exact title in Google Scholar or your library catalogue — a real paper appears immediately. Then check that the DOI resolves at doi.org. Then open the paper and confirm it actually says what you're citing it for, which is the step people skip once they've confirmed existence. Real source cited for a claim it doesn't make is still an error, and it's more common than outright fabrication.

Q3: What happens if a fabricated citation ends up in my submitted work?

It depends on your institution and on whether it reads as carelessness or as deception, and both outcomes are bad. At minimum it undermines the credibility of everything else in the paper; at many institutions it can trigger an academic integrity process, since a fabricated source is a false claim about your research. Markers find these quickly, checking a suspicious reference takes about ninety seconds, so the habit of opening every source you cite is the cheapest insurance available.

Q4: Is it safer to work from PDFs I've already downloaded?

For the specific problem of non-existent sources, yes, structurally: if the summary is built from a file on your machine, the source demonstrably exists. What that doesn't remove is the need to check that the source says what you claim, and the ordinary obligation to read enough of it to represent it fairly. Use general assistants to find search directions, obtain the real papers yourself, then work from the documents you actually hold.

Continue Your Learning with Sovi.AI

Sovi.AI is your free AI study buddy for step-by-step explanations, document-based learning, and AP exam prep. Put what you just read into practice with the tools below:

  • Ask Sovi — upload a photo to open the Ask Sovi chat and get a clear, step-by-step AI homework explanation across math, science, and writing.
  • AI Study — upload your draft or source PDFs to outline arguments, generate cheatsheets, and revise faster.
  • AP Test Prep — drill timed AP questions with full mock exams and unit-level practice across every AP subject.
  • Practice Resources — browse expert-verified study guides across Math, Biology, Chemistry, History, and more.

Looking for more guides like this one? Visit the Sovi.AI Blog for writing tips, grammar walkthroughs, and study strategies.