Frontier AI Models Disagree on Most Fact-Checks

  • AI
  • May 28, 2026
  • 0 Comments

Ask five of the world’s most capable AI models whether a claim is true, and you will get different answers more often than not. A research paper published this month put numbers on the problem: on 67% of 1,000 real-world fact-checking claims, a panel of five frontier large language models failed to agree — at least one model dissented from the majority verdict, or no majority formed at all.

The claims were not synthetic benchmarks with published answer keys. They were recent claims that real users submitted for verification to a fact-checking platform, the messy material of everyday misinformation: a rumor about a politician, a viral health claim, a disputed statistic. Each model was given the same claim and asked to pick a verdict from a four-bucket rubric: true, mostly true, misleading, or false. Exactly one bucket can be correct per claim, so every disagreement means at least one model was label-inconsistent.

The five models — GPT-5.4, Claude Opus 4.7, Gemini 3 Pro, Gemini 3 Pro with search, and Sonar Pro — are the strongest systems their makers produce. They scored comparably on public benchmarks, which is what makes the disagreement striking: models that look interchangeable on tests diverge sharply on the messier questions of fact.

The paper, by researchers at Lenz Research, found that 34% of claims produced a substantive split — a gap of two or more verdict buckets between the two most distant models. On those claims, one model said true while another said misleading, not a calibration quibble but a fundamental difference in reading the evidence. A statistical measure of inter-rater agreement, Krippendorff’s alpha, came in at 0.639 on an ordinal scale — structured agreement, researchers said, but far short of what would justify treating the panel as a single reliable judge.

Worse for users, confidence is no guide. The models rated 76% of their answers 9 or 10 on a 10-point confidence scale, yet they agreed with each other on confidence far less than on the verdicts themselves, with an agreement score of 0.44. A model that is wrong can be just as certain as one that is right — and all five models were highly confident almost everywhere.

The pattern held even when models were given retrieval tools. Gemini 3 Pro with search disagreed with its parametric counterpart on a large share of claims — the retrieval-augmented variant did not automatically converge with the base model, and in some pairs agreement was as low as 53%. The finding undercuts the comforting assumption that giving a model access to the web fixes its factual judgment.

The models also differ in temperament. Gemini 3 Pro concentrated its verdicts at the poles, calling more than half its claims plainly true, while Claude Opus 4.7 spread its answers across the middle buckets. Sonar Pro leaned toward mostly true and misleading. Two models could be systematically wrong in opposite directions — one too credulous, one too skeptical — and a naive reader of either would absorb its bias.

For the AI industry, the findings complicate the claim that models can police facts at scale. Automated fact-checking systems increasingly lean on LLMs to judge claims, and platform moderation pipelines route disputed content through model verdicts. If the models themselves cannot agree, the output of those pipelines inherits the disagreement. For enterprises, the stakes are practical: companies deploying AI assistants to answer customer questions, summarize contracts, or vet vendor claims cannot know which model’s answer is authoritative.

The paper’s authors recommend ensembles — a majority vote across diverse models eliminates many cases of dissent — but note that even after voting, 45% of claims still had at least two models disagreeing. Human oversight remains necessary, they conclude, which is another way of saying the automation has not arrived. The findings also feed the growing scrutiny of AI reliability from regulators and enterprise buyers: a chatbot that confidently asserts a false fact in a high-stakes setting is a liability, and the EU’s AI Act now requires providers to document model limitations.

None of this means the models are useless at fact-checking. Majority verdicts in the study aligned with human judgments in most cases, and disagreement concentrates in the middle of the rubric — claims that resist clean adjudication. The definitive poles of true and false drew near-unanimous verdicts about half the time, versus roughly one in ten for intermediate verdicts.

The practical lesson for anyone building on these models is blunt: pick one, test it against the others, and do not assume the consensus answer is right — assume it is a majority. The paper’s headline number — two-thirds of claims produce disagreement — is the industry’s clearest evidence yet that factual consistency, the property that would let a model serve as a neutral arbiter of truth, is still the hardest thing for AI to deliver.

Related Posts

  • September 6, 2026
  • 10 views
Anthropic Moves Its IPO Filing to Late September

The bankers and lawyers running Anthropic’s initial public offering had told investors to expect the company’s registration documents as soon as this week. The calendar has moved. Anthropic now plans…

  • September 6, 2026
  • 11 views
OpenAI Quietly Revises GPT-6 Astra Scores After Launch

When OpenAI released GPT-6 Astra on Sept. 3, the launch post carried the usual furniture of a modern model debut: coding results, speed comparisons and a figure for how often…