September 8, 2026

We Benchmarked AccuraCite's Citation Finder Against Raw AI — Here's What We Found

By Thu Tran

TL;DR: On claims about research an AI couldn't have memorized, AccuraCite's Find Citations feature returned a real, on-topic source 56.2% of the time per citation — nearly double a naive title match (24.3%) and well ahead of asking an AI directly with no search (32.1%), which invented a citation that doesn't exist more than half the time.

Two ways to read a number like "56.2%" below: per citation looks at every individual source returned and asks how many were good (56.2% here means 18 good sources out of 32 returned). Per claim asks a simpler question — did the search return at least one good source for this claim, yes or no. We report both, since they answer different questions: per-citation tells you how much to trust an individual result; per-claim tells you how often the search comes back empty.

How we built the test

Verifying a citation and finding one are different problems. Verification checks a citation you already have. Finding starts from a claim with no citation at all and has to search for something that actually supports it — a harder, more open-ended task.

We took 55 real, published papers with real abstracts and wrote one factual claim from each abstract — the same kind of sentence you'd type into AccuraCite's Find Citations box. 40 claims each need one citation; 15 are compound claims bundling several distinct facts into one sentence, needing multiple citations at once. The claims span four fields:

Test claims by field Chemistry 7 Psychology 14 Medicine 16 Biology 18

Every paper was published between April and September 2026 — no more than five months old — with fewer than 10 citations to its name. That's deliberate: a paper old and famous enough to be well-known might already be sitting in an AI model's training data, so asking that model to name a source for it isn't testing search at all, just its memory. Testing on brand-new, barely-cited work removes that shortcut — see "Why we didn't test on well-known papers" below for what happens when you skip this step.

For every citation any method returned, we checked two separate things: is it real (matched against actual databases, not invented), and does it actually support the claim. "Support the claim" is a strict, either/or bar, not a sliding scale — a paper is sorted into exactly one of three buckets: it directly backs the specific thing the claim says (SUPPORTS), it's on the same general topic but doesn't actually address that specific point (TANGENTIAL), or it's unrelated (UNRELATED). Only SUPPORTS counts as "good" in the numbers below. A paper that's merely in the right subject area — related, adjacent, reminiscent — does not count, even though it might look like a reasonable partial match at a glance. We check both real and SUPPORTS separately because a made-up citation is built to look like a real one — plausible title, plausible author, plausible journal — so a check that only asks "does this sound relevant?" will call some fake citations relevant too.

Three methods, tested the same way

We ran the same 55 claims through three methods:

  1. AccuraCite's real pipeline — search three academic databases (Semantic Scholar, Crossref, OpenAlex), then have an AI read each candidate paper's full abstract and keep only the ones that actually support the claim.
  2. Title-match only — the same database search returns a handful of candidate papers, each with its own real title. Skip the abstract check and just keep whichever candidate's title reads closest to the claim's own wording. This isolates what checking the abstract is actually worth, on top of the search itself.
  3. Ask an AI directly, no search — paste the claim into an AI chatbot and ask it to name a source, with no database lookup at all. We check two things about its answer: does the citation it names actually exist (checked against real databases), and does it genuinely support the claim.

The results

Per-citation real and relevant rate by method Title-match only 24.3% Ask an AI directly 32.1% AccuraCite 56.2%

Per citation — of every individual source a method returned, how good was it:

Method Fabricated (doesn't exist) Real & at least on-topic Real & relevant (precise match)
Title-match only (no abstract check) 0%* 66.2% (49/74) 24.3% (18/74)
Ask an AI directly (no search) 53.1% (43/81) 46.9% (38/81) 32.1% (26/81)
AccuraCite (search + abstract check) 0%* 100% (32/32) 56.2% (18/32)

*AccuraCite and title-match search real databases and never invent a candidate, so 0% here is true by construction — there's no step where a citation could be made up, not something separately checked the same way the 53.1% figure was (that one comes from an actual verify_citation() lookup against real databases for every citation the AI proposed).

Every single citation AccuraCite ever returned was real and at least on the right topic — never fabricated, never flatly unrelated. Checking the abstract is doing real work beyond that baseline too: AccuraCite's 56.2% precise-match rate is well ahead of both alternatives.

Per claim — of all 55 test claims, how often each method actually answered, and how good that answer was:

Method Attempted an answer Hit rate when it attempted Claims with at least one good citation
Title-match only (no abstract check) 92.7% (51/55) 31.4% (16/51) 29.1% (16/55)
Ask an AI directly (no search) 92.7% (51/55) 43.1% (22/51) 40.0% (22/55)
AccuraCite (search + abstract check) 40.0% (22/55) 63.6% (14/22) 25.5% (14/55)

This is the one worth being upfront about rather than glossing over: asking an AI directly actually names at least one good citation for more claims than AccuraCite does (40.0% vs. 25.5%) — but that comparison isn't as fair as it looks. AccuraCite only attempts an answer for 40.0% of claims, declining the rest rather than guess; the other two almost never hold back, attempting 92.7% of the time. A method that always answers will rack up more "at least one good citation" claims purely by taking more swings. Conditioned on actually attempting, AccuraCite's hit rate (63.6%) beats both alternatives — and 43 of the AI's 81 proposed citations, 53.1%, don't exist at all, matched against nothing in any real database. AccuraCite would rather return nothing than guess, which is exactly why the per-citation table above — how trustworthy an individual answer is, not just whether something came back — is the one we lead with.

The full test set and every method's raw output is public: github.com/AccuraCite/finding-benchmark. Run compute_metrics.py yourself and check every number above directly.

What this looks like on one real claim

Numbers like these can feel abstract, so here's one actual claim from the test set and what each method returned for it:

Claim: "Integrating a neuro-symbolic, multi-agent artificial intelligence platform with an oncology-specific knowledge graph can reduce the median per-patient screening time for clinical trial eligibility from 120 minutes to approximately 30 minutes."

This is one case, not the average — see the aggregate numbers above for the full picture — but it's a clean illustration of the two separate failure modes the per-citation table quantifies: a title match can lock onto a real paper that's still the wrong one, and an AI asked directly will name something that reads as real and relevant when, in this test, 53.1% of the time it wasn't real at all.

Why we didn't test on well-known papers

Our first attempt at this benchmark picked papers a normal topic search would surface — no restriction on how well-known they were. Search engines rank by relevance and popularity, so that surfaces famous, heavily-cited papers first. On that version of the test, asking an AI directly scored almost as well as AccuraCite's actual search:

Method (on well-known papers) Per-citation, real & relevant Claims with a good citation
Ask an AI directly (no search) 74% (25/34) 68% (15/22)
AccuraCite (search + abstract check) 61% (20/33) 77% (17/22)

That's not because asking an AI is actually a good strategy. 35% of its answers (12 of 34) were an exact, word-for-word match to the very paper each test claim had been written from — it wasn't finding anything, it was recalling something it had already read during training. Any well-known paper is likely somewhere in an AI model's training data, which makes "just ask it" look deceptively strong on exactly the claims where it doesn't need to search at all. That's not what happens when someone asks about their own specific claim about recent or lesser-known work — which is most real citation needs. So we rebuilt the test around recent, barely-cited papers instead, which is the version reported above.

What this benchmark doesn't cover

  • Small sample. 55 claims is enough to see a clear, consistent direction, not enough to treat the exact percentages as precise. We'd want several hundred for that.
  • This is the hardest case, not the typical one. Restricting the test to recent, barely-cited papers was necessary to get an honest read on the AI baseline, but it means 56.2%/25.5% is a lower bound, not an average. On well-established literature — most real citations — the same pipeline scored noticeably higher (see the table above). We haven't yet built one benchmark that mixes realistic proportions of both, and didn't want to publish a single blended number until we have.
  • Very new papers aren't always indexed yet. We checked a sample of the claims where AccuraCite found nothing: most of the time, the paper genuinely isn't in any of the three databases we search yet, not because our search missed something that was there. That's a real, current limit on how fast academic databases catch up with brand-new research, separate from how trustworthy AccuraCite's results are once it does find something.

Try it yourself

Find sources for your own claim free and see what comes back.

FAQ

Why does "real AND relevant" matter more than just "relevant"? A citation that looks relevant but doesn't exist is worse than no citation at all — it's the kind of mistake that's invisible until someone checks it by hand. A relevance check working only from a title and author list can't always tell a fabricated citation from a real one; both can read as on-topic. We only count something as good when it passes both checks.

Why test on recent, barely-cited papers instead of typical claims? Because that's the only way to get an honest read on what asking an AI directly is actually worth. Well-known papers are already memorized by most AI models, which makes a no-search answer look artificially strong on exactly the claims it doesn't need to search for. Testing on obscure, recent work removes that shortcut.

Does this mean AccuraCite only finds a citation a quarter of the time? On this specific, deliberately hard test — yes, 25.5% of claims got at least one citation we could confirm was both real and genuinely on-topic. That's the floor, not the typical case: on well-established literature, the same pipeline scored 77% in a separate test (see the table above).

Topics AI citation finder accuracy citation finding benchmark does AI citation search work find citations for a claim accuracy AI hallucinated citations benchmark