We Ran Resumes With Gemini vs. Anthropic. The Hallucinations Changed Shape.

Every AI model will sometimes invent things on your resume — but each model invents different things. We know because in July we swapped the AI engine under Bloom from Google’s Gemini 2.5 Pro to Anthropic’s Claude Sonnet 4.6, and before shipping the change we re-ran our entire fabrication test suite on the new model. The safety rules we had spent months building still worked — they cut fabrication-linked errors 80–96% on both models. But the kinds of lies each model told had almost nothing in common. The only protection that carried over completely: a separate check that verifies every generated line against the resume you actually wrote.

If you use any AI resume tool, this matters more than it sounds. Every tool upgrades its underlying model eventually, usually without telling you. Here’s what we measured when we did it deliberately, in the open.

What we did

We run a standing test suite for resume fabrication: synthetic resumes (no user data, ever) across nine professions — software engineers, nurses, electricians, hospitality managers, career switchers — each paired with real-style job postings. We tailor each resume to each job, then count how many of the resulting bullet points get flagged as unsupported by the original resume. A flagged bullet means the AI wrote something your history doesn’t back up.

When we swapped models, we ran the new model through the same suite two ways: completely raw, and with the trained guardrail instructions we use in production — the same rules, unmodified, that we built for the old model.

The numbers

ModelRaw, no guardrailsWith Bloom’s trained guardrails
Gemini 2.5 Pro (our engine until July)~50% of bullets flagged1.9% (106 bullets, 12 resume-job pairs)
Claude Sonnet 4.6 (our engine now)30.5% of bullets flagged6% (2 of 33 bullets, 4 pairs)

Two honest footnotes, because numbers without them are marketing: the Gemini samples are larger (106 bullets across 12 pairs vs. 33 across 4), and the ~50% raw figure is Gemini’s historical baseline from our May testing rather than a same-day head-to-head. Full methodology below.

So which model did better? Trick question.

Depends what you mean by “better.” Left alone, Claude is the more honest model — raw, it produced unsupported claims on 30.5% of bullets versus roughly 50% for Gemini. But Gemini posts the lower guarded number (1.9% vs. 6%) — and that’s not because Gemini is safer. It’s because our guardrail rules were built by cataloguing Gemini’s bad habits for months. Gemini is playing on home turf. Pointed at Claude, the same rules still worked — five times fewer errors than raw — just not as perfectly, because Claude cheats differently.

That’s the real headline: the model matters less than whether anyone is checking its output. Which is also why “which AI writes the best resume?” is the wrong question. The right question is “what happens when the AI gets it wrong — and does the tool even notice?”

Every model has a fabrication fingerprint

A fabrication fingerprint is the recognizable style of things a model makes up when it’s trying to be helpful — like handwriting, but for lying. Models don’t invent randomly; each has go-to moves it reaches for when your resume doesn’t quite fit the job.

Gemini’s go-to moves (the ones our rules were built to stop):

  • Tools it never saw — a “Tableau” or “Snowflake” appearing on a resume that never mentioned them
  • Title promotions — “co-managed” quietly becoming “owned”
  • Certainty upgrades — “contributed to” becoming “led”
  • Bolted-on buzzwords — “data-driven,” “self-serve,” attached to work that never claimed them

Claude’s go-to moves (different, and in one way worse):

  • Process and stakeholder stories that never happened — invented coordination, invented handoffs
  • Credentials and named systems — “Epic EHR” experience, a “ServSafe” certification, a “Cornell” degree, none of them on the source resume
  • The pull is strongest when the job posting itself names the credential

That second fingerprint is the dangerous one, because a fabricated certification isn’t a fuzzy exaggeration — it’s a checkbox claim that fails the moment an employer verifies it. We wrote up that specific finding, with the run-by-run data, in the credential magnet report.

The fix that survived the swap

Guardrail instructions reduced fabrication on both models. They eliminated it on neither. What caught everything, on both models, in every sample we took: a separate verification pass that checks each tailored bullet against your original resume and flags anything unsupported — before you ever send it. In our swap testing it caught 100% of the fabrications that slipped past the guardrails, including the stubborn ones no instruction could suppress.

That’s a structural point, not a Bloom victory lap. Rules tuned to one model’s habits partially transfer to the next model. Verification transfers completely, because it doesn’t care how the model lied — only whether your resume supports the claim. If your resume tool can’t show you, line by line, which bullets your real history backs up, then its safety story is one model upgrade away from expiring — and you won’t be told when it happens. Our first hallucination study, the Resume Hallucination Report, goes deeper on how per-line verification works, and these red flags are what unverified AI output tends to look like.

Methodology

All testing used synthetic resumes and job postings — constructed profiles across nine professions, zero user data. “Flagged” means our verification pass marked a tailored bullet as unsupported by the source resume. Gemini figures: ~50% raw is the historical unconstrained baseline from our May 2026 variance testing; 1.9% guarded was measured across 106 bullets in 12 resume-job pairs. Claude figures were measured in July 2026 on the production infrastructure: 30.5% raw across the full test matrix, 6% guarded (2 of 33 bullets across 4 pairs). Guardrail instructions were identical across both models. Sample sizes differ; we publish them alongside every figure for that reason.

FAQ

Which AI model is best for writing resumes?

The honest answer: it matters less than you think. In our testing, both frontier models fabricated resume content when unguarded (30.5% and ~50% of bullets), and both improved dramatically with trained guardrails (6% and 1.9%). The differentiator isn’t the model — it’s whether the tool verifies each line against your real experience afterward.

Do all AI models make up the same things on resumes?

No — and that’s the finding. One model inflated tools, titles, and certainty; the other invented process stories and credentials. If a tool tuned its safety rules on one model’s habits, those rules only partially apply after a model swap.

My resume tool upgraded its AI. Should I care?

Yes, at least enough to re-check the output. A model upgrade silently changes what the tool tends to get wrong. Re-read any tailored resume line by line after a tool announces (or you suspect) an engine change — especially certifications, tools, and job titles.

How do I catch AI fabrications on my resume?

Compare every bullet against what you actually did, and be strictest with checkable claims: certifications, named software, employers, titles, dates. Bloom does this automatically — every tailored bullet is verified against your source resume and anything unsupported is flagged in red before you export. Here’s how to use AI without lying on your resume.


Bloom tailors your resume to each job and verifies every bullet against your real experience — whichever model is under the hood that month. Try it free.