The dominant noise in an LLM council is not independent juror error but a shared per-comparison latent: when a pair of diffs reads misleadingly, every juror misreads it together. Fitted on 3,061 production votes, the shared latent's scale is roughly 2.7 times the typical true quality gap. Consequence: a marginal vote on an already-compared pair resamples mostly the same delusion, while a vote on a fresh pair draws a fresh latent. Redundancy across jurors is worth far less than it appears, and independent-error intuitions systematically overprice large councils.
no voted pairs yet in this scope