How to Check Your Students' Citations for AI Hallucinations
TL;DR: A plagiarism checker tells you if text was copied. It says nothing about whether the works cited actually exist. Students using ChatGPT or Claude to draft a literature review can hand in a bibliography full of perfectly formatted, entirely fake sources — and a similarity score alone won't catch it. This is a step-by-step walkthrough of bulk-checking a stack of student papers for hallucinated and mismatched citations in one pass.
Why this is a different problem than plagiarism
A similarity checker compares a student's text against a corpus of existing writing. It's built to answer "did this come from somewhere else?" It was never built to answer a much narrower, newer question: "does the source cited on page 4 actually exist, and does it say what the student claims?"
That second question matters more than it used to, because LLMs don't fail at citations the way a student copying from Wikipedia fails. An LLM asked to draft a reference list will produce entries that are stylistically perfect — correct APA or MLA formatting, plausible author names, a journal that sounds real — while citing a paper that was never published, or attaching the wrong year and publisher to a real one. None of that trips a plagiarism check, because none of it is copied from anywhere. It's fabricated from the model's sense of what a citation should look like.
What to actually check for
There are three distinct failure modes worth looking for, and they call for different responses:
- Hallucinated: the citation doesn't exist in any indexed academic database. No DOI, no matching title, nothing. This is the clearest sign a source was never read — possibly never checked by the student at all.
- Mismatch: something real is cited, but a detail is wrong — the year, the author list, the journal, the DOI. Sometimes this is sloppy manual transcription. Sometimes it's an LLM blending two real papers into one citation.
- Verified: the source exists and the metadata lines up. No further action needed on that entry.
The point isn't to accuse every student with a hallucinated citation of academic dishonesty — sometimes it's a rough draft, a citation manager error, or a source they meant to double-check and forgot to. But you can't have that conversation at all if you don't know which citations to look at first.
Step-by-step: checking a stack of papers at once
- Collect the papers as PDFs. AccuraCite's Bulk Verification plan accepts up to 30 PDFs in a single scan — a full section's worth of essays or a batch of thesis chapters in one job, instead of one paper at a time.
- Drop them into a bulk scan. From your workspace, open the Bulk Verification tab and upload the set. AccuraCite's parser extracts each paper's reference list automatically, regardless of which citation style the student used.
- Let it cross-check against real sources. Every extracted citation is queried against seven independent academic databases at once — OpenAlex, Crossref, Semantic Scholar, PubMed, DBLP, arXiv, and CORE — and a consensus vote decides the result, so a single noisy or wrong database entry doesn't produce a false "verified." Read more on the How It Works page.
- Read the per-file breakdown, not just a score. Each file in the run shows a verified/total ratio and expands into the individual citations, so you can go straight to the paper with 6 hallucinated sources out of 12 instead of skimming every bibliography by hand.
- Follow up on what's flagged, not what's clean. A hallucinated or mismatched citation is a prompt to ask the student where the source actually came from — not an automatic verdict. Treat it the way you'd treat a plagiarism-checker flag: a lead, not a conclusion.
What this doesn't replace
This isn't a substitute for a plagiarism/similarity checker, and it isn't a grading tool. It answers one specific question — "is this reference list real?" — that similarity checking was never designed to answer. Most instructors get the most value running both: a similarity check for copied text, and a citation check for fabricated sources, since an LLM-assisted paper can pass the first and fail the second cleanly.
It's also not a guarantee that a "verified" citation was actually read or used correctly — only that the source exists and the metadata matches. A student can cite a real paper about something the paper doesn't actually say. AccuraCite confirms the reference is real; it doesn't (yet) confirm the argument built on top of it is honest.
FAQ
Do I need every student's paper to be a clean PDF? Yes — the bulk scanner parses PDFs directly, extracting the reference list regardless of citation style (APA, MLA, Chicago, IEEE, and others are auto-detected).
What if a citation is real but formatted incorrectly? That's flagged as a mismatch, not a hallucination — the tool distinguishes "this doesn't exist" from "this exists but something's off," so you're not treating a typo the same as a fabricated source.
Can I check papers one at a time instead of in bulk? Yes — the Pro plan supports single-PDF scans (PDF parsing isn't included on the free tier, which works from pasted citation text instead). Bulk scanning (up to 30 PDFs per job) is built for exactly this use case — checking a whole class set or stack of submissions at once — and is available on the Bulk Verification plan or the 7-Day Bulk Pass for a shorter grading window.
Does this tell me if a student used AI to write the paper? No — it tells you whether the citations in the paper are real. That's a narrower, more concrete signal than AI-writing detection, and it doesn't rely on guessing at writing style, which is part of why it's harder to fool: a citation either exists in the indexed literature or it doesn't.