Consensus Is Not Evidence
Apple measured the problem in May 2026. This is the engineering consequence.
By Korhan Temocin, founder of Core Method. 1 August 2026.
Nine models can agree and still give you only two votes' worth of information. That is the measurement. The engineering problem starts on the next line.
In May 2026, a researcher at Apple put nine frontier language models on the same judging task. Their panel carried about two effective votes. The models were not simply weak. They tended to fail on the same items, so adding judges multiplied confidence much faster than it multiplied independent evidence.
The finding deserves credit. It does not deserve the last word.
I came to the operating problem from the other direction. I was not trying to estimate effective votes. I was deciding what a running system should do when several readers agree.
My answer is not to add a tenth judge or to average the panel more carefully. It is to change the job of the panel. Track who has been right in this domain. Preserve material disagreement as an unresolved question. Separate operator judgment from claims that can later be graded against reality.
Part of that loop runs today: the operator-scored record and the preserved disagreements, plus dated forecasts scored later on a separate path. The correlation layer does not run today: it is a specification, and the strongest prior work warns against assuming it will rescue the panel.
This is a claim about engineering consequence, not a claim to have discovered correlated error: once consensus stops being independent evidence, the system must change what happens next.
The day four laboratories agreed
On 12 July 2026, four models from four laboratories attacked one of my designs and converged. I did not treat four matching answers as four independent observations. I sent the convergence into another grounding pass. That pass was imperfect: its academic lane returned noise that day, and the fallback web references were transcribed rather than independently reverified. It did not establish that the design was right or wrong. It established the operating requirement: convergence must create a new verification task, not close the old one.
That requirement sounds obvious until you watch what convergence does to a person at a keyboard. Four strong models, four different vendors, one voice. Every instinct says ship. The whole point of paying for four opinions was to be more certain, and here is the certainty, delivered. The reflex I had to build, and am still wiring into the system so it stops depending on my discipline on a tired evening, is the opposite one: the stronger the agreement, the more interesting the question of what all four might be missing together.
Why agreement inflates
The statistics underneath this is old. In 1969, Bates and Granger showed that once forecast errors are correlated, combining them stops delivering the variance reduction that independence would promise. The off-diagonal terms do not vanish, so two correlated opinions are worth less than two independent ones, no matter how confident each sounds.
What is new is the measurement. Kim and colleagues showed at ICML 2025 that more accurate models are more correlated, which means the better your panel gets, the worse the illusion becomes. Apple's May 2026 study put a number on the damage: nine frontier judges, about two effective votes, with established aggregation methods recovering little of what correlation had destroyed. And the idea that unanimity itself should raise a flag has already been proposed for multi-agent triage in a 2026 preprint.
None of that is mine. The measurement belongs to Apple and Kim, the flag idea has prior owners, and the statistics is half a century old. What I can offer is the wiring: what changes in a working system once you decide to act on a result that old.
One more disclosure, about the order of events, because this article would otherwise imply I read the literature first and built second. It was the reverse. The mechanisms described below were built to answer operating problems inside one system; the papers were located afterward, when the work went looking for its own lineage, most of them in the prior-art pass run for this piece. I offer that as a disclosure, not a credential. Independent arrival cannot be verified from outside, ideas travel by many roads, and nobody builds in a clean room in 2026. What the order changes is the reading: nothing here was built to implement these results. The system simply turned out to agree with them about where the problem lives, and because agreement is a work order, that convergence got the treatment the article prescribes: a verification pass, which is where the citations in this piece come from.
What a running system does instead
The system this comes from is Horizon Core, a private research system I operate. The part that matters here is small and concrete, and I will keep the verbs honest: what follows is what the system does, not what a design document hopes.
Every reader in the system carries a track record. A reader here is a reading framework, a typed lens that interprets material and emits structured fields, not a chat window. Each one accumulates ratifications and refutations in a ledger, exposed as a single number: a Beta posterior mean that starts at 0.5 for an unproven reader and is earned up or down as operator scorings land. A reader that has never been checked is worth exactly as much as a coin flip, and the arithmetic says so.
When readers read the same material, the system weights each one by that earned mean before combining their positions into a center. Confidence buys nothing. History buys weight. The weighting is linear in that mean, a deliberate simplicity: an unproven reader still pulls the center, it just pulls with less force than a proven one.
When readers disagree beyond a threshold, the system does not average them and does not pick a winner: it records the disagreement and raises a question. A panel where fewer than two readers speak is marked insufficient, not resolved. The question becomes a ticket I answer, and my answer is scored back onto each reader's track record. The answer is write-once: one answer can never move the ledger twice.
Separately, the system can post typed, dated forecasts into a prediction ledger, where code grades them at their horizon. A typed comparison and a Brier score, no model in the loop. A domain whose mean Brier drifts worse than 0.25 over at least twenty scored predictions is muted automatically.
That forward path is already loaded. On 12 July 2026, the same day as the convergence above, two of the system's forecasting lanes posted its first two open claims onto a public question: whether battery-electric cars reach 28 percent of new UK registrations in calendar 2026. The lanes disagreed, one at 0.55 that it happens, the other at 0.58 that it does not. Both rows stand as written, opposed and dated, until the industry's own registration figures grade them in January 2027. Nothing in this system has yet been graded by reality; the earliest of these horizons lands in January 2027. But the design has already closed one loop elsewhere: a sibling system that inherited this ledger design posted a short-horizon industry call at 0.85, with its resolution criterion written down at posting time, and five days later it resolved correct. One graded row is an anecdote, not calibration. It is also the difference between a mechanism that waits and a mechanism that works.
One boundary before the walkthrough. What runs today reconciles two reading frameworks on a single axis, inside one model lane. It is not a panel of models from different labs. The fusion pipeline does carry a separate refute-first approval pass, and it is a real check, but the builder and the approver are the same model lane in different contexts. That is a separation of stance, not of model family, so on its own it does nothing about correlated error.
Here is one decision carried end to end. The numbers are synthetic; the mechanism is the one that runs.
- Two readers read the same material. Reader A has been scored 31 times and stands at 0.64. Reader B has never been scored and stands at the prior, 0.50.
- On the contested axis, A reads +0.2, B reads +0.7. The spread, 0.5, crosses the divergence threshold. The spread is raw and unweighted; only the center is weighted, and it lands at 0.42, leaning toward A's earned record.
- The system writes no averaged verdict, and nothing pretends the disagreement is settled.
- Instead it raises a question: two readers split on this axis, and both positions are preserved.
- I answer it. In this example my answer lands nearer B's reading.
- The answer is scored back: B is ratified, A is refuted, for this reading only.
- The ledger moves. A slips from 0.64 toward 0.62. B climbs from 0.50 to 0.67. A barely moves and B moves a lot, because an empty record has nothing to anchor it. That asymmetry is the point: the longer the record, the less one outcome can move it.
- The same answer, replayed against the ledger a second time, is refused. Write-once means a single judgment can never compound itself.
This trace does not prove the method produces better decisions, only that the system actually does what this article claims it does.
Status
| Status | What exists |
|---|---|
| Running now | A reliability ledger, a reliability-weighted center, unresolved divergence raised as a question, write-once operator scoring, and a separate path for dated forecasts graded against reality |
| Specified, not shipped | Chance-corrected agreement, an independence discount, contested-construct discounts, and the expanded divergence gate and ticket schema |
The discount every implementation still owes
The obvious next step is the one Apple's paper points straight at: discount a reader not only for being wrong, but for being redundant. The design proposes that discount. A reader's vote would become the product of three factors: earned reliability, an independence weight, and a discount for leaning on contested constructs. The independence weight is specified as one divided by one plus the reader's summed pairwise agreement with its peers, where agreement must be chance-corrected, because two readers that both emit the most common value agree constantly by luck, and raw agreement would read that luck as redundancy.
I want to be precise about the standing of that formula. It is a heuristic. It is motivated by the 1969 result and by the design-effect correction from survey statistics, but derived from neither, and the design document says so in those words. Amazon published the more rigorous approach at ICML 2026: fit the dependence structure itself with class-conditional Ising models rather than assuming a functional form. If you need this layer today, start from that work, not my formula.
And there is a caution built into the strongest prior result. Apple tested correlation-aware weighting with oracle access to the true error structure, and it recovered only a small fraction of the lost information. The bottleneck is the correlation itself, not the cleverness of the weighting. So the independence discount is specified, not shipped. Every stored reading already carries the two fields it would key on, the theoretical family of the reader and the disputed constructs it leaned on, so the schema is provisioned for it. No code reads them back to deflate a vote yet, and I am not going to describe a design as a deployment.
One corner of the same system already enforces the arithmetic, though. The gauntlet that graduates forecasting methods pools the two lanes answering the same held-out question into a single trial, which hits only if both lanes hit. It is a different subsystem and a much narrower rule than the discount specified above, but the principle is in running code: correlated agreement counts once.
The newest work approaches the same gap from underneath. In the week I finished this piece, the system gained a layer that writes down, at serve time, the record every after-the-fact correlation study is missing: one exposure row for every target served on every one of its hunting paths, so examined-and-found-nothing can no longer be confused with never-reached, alongside planted ground-truth items whose identity is hidden from the producing lanes, from the judge, and from my own live view. Apple and Kim measure correlation after the fact, on whatever judgments were observed; with public benchmarks and closed training sets, nobody can know who was exposed to what. A private system can know, if it writes exposure down before anyone answers. Only the recording layer runs today. The scorer meant to read it was specified, reviewed, and withdrawn in the same week: two independent reviews, one by a second model family and one by a multi-lens pass, converged on the same fatal defect, and the convergence was treated as evidence precisely because their routes differed. The instrument would have been well tested and wrong, and its tests would have read as proof that it worked.
Three operating rules
If you run multi-model evaluation, orchestrate agents at any serious scale, or sit in any decision process where several AI opinions land on your desk, three rules transfer directly.
First, tie confidence to earned track record, not to vote count. An unproven reader is a coin flip and should be weighted like one, however fluent it sounds. Keep the record per domain: being right about markets earns nothing in medicine.
Second, preserve disagreement. When two credible readers split, the split is the most informative thing you own. Averaging it away deletes exactly the signal you paid for. Route it to a named owner as an explicit question, with a date. In an eval pipeline that looks like this: two judges split on factuality, the eval lead owns the question, the answer is due Friday, and the resolution is scored back onto both judges.
Third, prefer claims that reality can grade. A dated forecast with a horizon beats a confident opinion, because it produces calibration data even when it is wrong. Opinions produce nothing but comfort.
And the rule that binds the other three: when your panel converges, schedule verification. Agreement among correlated readers is a trigger for checking, never a permit for shipping.
Inside Horizon Core, that trigger has a permanent employee. A standing adversary, named after cinema's most famous anomaly, does nothing but attack the system's own beliefs, and its queue is sorted the way societies never sort theirs: the more fiercely a belief is held, the higher it climbs on the attack list, because the beliefs a community trusts hardest are exactly the ones it questions least. Every attack is logged, every finding is judged by an independent gate before it can change anything, and the full record lands on my desk. Trust, in that corner of the system, is not what survived agreement; it is what survived assault. The machinery deserves its own article and will get one. Here it earns one sentence: verification stopped being an event and became a budget line.
What this does not solve
The deeper limit is who grades whom. Inside the panel, the final answer is mine, and nothing grades me. Only the forecast path faces reality directly; everything else bottoms out in one operator's judgment, which is exactly the kind of unchecked single source this article warns about.
Most importantly, the proposed independence discount is not evidence of robustness; it is an unshipped specification with no result behind it, inside one operator's system. Prior experiments suggest that correlation-aware weighting may recover only a small part of the information already lost. The design therefore ends where it should: as a testable claim, not a victory lap.
Consensus is a work order
The rule underneath all of this cuts both ways. Correlated agreement is weak evidence, which is the whole story above. Independently derived agreement is strong, which is why the scorer died. The discipline is telling which of the two is on your screen.
The next time several models hand me the same answer, it will get what the July convergence got: a new verification task, not a closed question. Inside the system, a split between credible readers stays preserved as a question, and the claims worth keeping go out as dated forecasts that reality can grade later. Agreement still carries information, about two votes' worth on Apple's numbers, and two votes are not nothing.
They are just not a verdict. Consensus is not a conclusion. It is a work order, and the work it orders is verification.
References: G. Kohli, "Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels", arXiv:2605.29800 (2026). E. Kim, A. Garg, K. Peng, N. Garg, "Correlated Errors in Large Language Models", ICML 2025 (arXiv:2506.07962). K. Balasubramanian, S. Podkopaev, S. Kasiviswanathan, "Dependence-Aware Label Aggregation for LLM-as-a-Judge via Ising Models", ICML 2026 (arXiv:2601.22336). McDonnell et al., arXiv:2607.19899, accepted at PAAMS 2026, on treating unanimous agreement as a warning flag in multi-agent triage. J. M. Bates and C. W. J. Granger, "The combination of forecasts", Operational Research Quarterly 20(4), 1969.
Korhan Temocin, founder, Core Method. Istanbul, 1 August 2026.