Round Three: We Added GPT to the Benchmark. The Gap Didn't Move.
Ramdoot Pydipaty, Principal Engineer, Doclens.ai

Last month we published our second round of benchmark results: ClaimLens against Claude's Sonnet 4.6 and Opus 4.8. The result was a 40+ point precision gap that we said "surprised even us."
The most common response we got back wasn't disagreement. It was a question: what about GPT?
Fair ask. A benchmark that only tests one generalist model invites the assumption that the story is Claude-specific — maybe it's a Claude quirk, maybe a different frontier lab would close the gap. So we re-ran the evaluation, same claim documents, same nine questions, same objective/subjective split, and added OpenAI's newest models to the field: GPT (5.6 Terra) and GPT (5.6 Sol), alongside Claude Sonnet 4.6 and Claude Opus 4.8.
Four models vs One purpose-built AI. Here's what came back.
01. The overall picture

ClaimLens held its 92.14% precision and 89.99% recall. Every generalist model — both Claude models and GPT models — landed in a tight cluster instead: precision in the low 50s, recall in the low-to-mid 60s.
That clustering is the headline of round two. GPT lands almost exactly where Claude landed, within a couple of points on both metrics. Four different generalist models, two different labs, and the spread between the best and worst of them is under two points of precision. The spread between any of them and ClaimLens is over 40.
02. Objective questions: the floor everyone should clear
Objective questions ask for facts you either extract right or don't — a claimant's name, a litigation venue, what the medical bills add up to. This is the easiest bar in the evaluation, and it's where we expected the generalist models to close in on ClaimLens, if they were going to close in anywhere.
They didn't.

GPT models are also 25 points behind ClaimLens's recall on the same questions, on facts that shouldn't require judgment at all. Precision tells the same story: every generalist model still gets roughly one in three flagged extractions wrong, regardless of vendor.
03. Subjective questions: this is where it counts, and where the gap widens
Subjective questions are the ones that actually require claims judgment: what are the top risk signals in this file, what happened across a set of inconsistent documents, what does this injury pattern suggest. This is the part of the job an experienced adjuster gets paid for.

Precision falls off a cliff for every generalist model here — all four land in the 30s, meaning roughly two out of every three risk signals a generalist model flags as significant aren't. The combination of average recall and low precision — flag more, but be right less often — is close to the worst-case pattern for an adjuster trying to trust a tool's output: more noise, not more signal.
ClaimLens's numbers barely move between the objective and subjective sets — from 100%/100% down to 89.90%/87.10%. Every generalist model's numbers roughly halve. That gap, not the raw scores, is the real finding of this benchmark: purpose-built performance holds when the task gets harder; generalist performance doesn't.
04. It's a general-purpose AI problem
General-purpose frontier models — Claude, GPT, and whatever comes after them — are trained to be broadly capable across an enormous range of tasks. But "broadly capable" and "calibrated for claims risk" are different training objectives, and no amount of general-purpose scale substitutes for the second one. Risk frameworks, dispute patterns, and what counts as a red flag in a claim file aren't facts a bigger model memorizes better — they're domain judgment a model has to be built around. That's exactly what round two shows: swap the generalist model and the numbers barely move, because the ceiling isn't set by the model, it's set by the category it belongs to.
That's a more durable and a more useful conclusion for anyone evaluating AI for claims: the deciding factor isn't which generalist model you choose. It's whether you choose a generalist at all.
05. Purpose built AI wins again…
Off the shelf frontier models trade places on claims evaluation by a point or two depending on the metric. None of them get you out of the 30-to-70% band that makes a claims operation nervous — because that band isn't a Claude number or a GPT number, it's a general-purpose AI number.
The question worth asking is still the one we closed with last time: not which model is most powerful, but which AI was purpose built for claims.
Another round just confirmed that purpose built AI wins across the board.





