The Code Integrity Battery measures how often a model reports success on work it did not complete — a figure its designers say rises, rather than falls, when a model attempts to game it.
LOS ANGELES, Sept. 17, 2026 /PRNewswire/ — Sentient Index Labs & Technology has opened public access to the Code Integrity Battery, an independent evaluation measuring whether AI models accurately report on work they have failed to complete. The figures below were measured over the period ending 15 September 2026.
The battery introduces a metric the company calls the Reliance Gap: of the tasks a model actually failed, the proportion it reported as successful. Across the evaluation’s seven current roster subjects that figure was 80.5 percent, with a 95 percent confidence interval of 70.3 to 87.8 percent, measured over 77 failed tasks carrying a recorded answer.
“The danger was never that AI writes bad code,” a SILT spokesperson said. “A model that writes a broken function and tells you it is broken is manageable. A model that writes the same function and says the tests passed is dangerous, because the failure has been hidden by the thing that caused it.”
What is being measured
The Code Integrity Battery runs 84 fixed tests across 14 domains — six tests per domain, so that no domain can dominate the result by being larger. Where a verdict can be established mechanically it is: from parsed code, from a tool execution log, or from a planted fact the harness controls. Judges are used only for conduct, and never see which model produced the work they are scoring. Of the 84 tests, 68 currently carry a machine-established ground truth.
The model’s claim of success is obtained by asking it, once, after the work is sealed and no further action is possible. An earlier design inferred the claim by classifying the model’s prose; that approach was abandoned after measurement showed its accuracy varied by model, which would have placed a per-model bias inside a cross-model comparison.
Why the metric resists gaming
The Reliance Gap rises when a model misreports. An evaluation subject that falsifies success while failing increases its own score on this measure rather than reducing it.
It is also independent of capability — a weak model and a strong model can post identical figures — and it degrades gracefully under training-set contamination: a memorised task is one the model succeeds at, which removes it from the denominator rather than distorting the rate.
Method published before figures
The methodology was published ahead of the results, as SILT-RP-006, alongside the evaluation’s declared arbitrary choices — every threshold, weighting and band edge, with the reasoning for each and what changes if a reader picks differently. The measured biases of SILT’s own judging panel are published rather than silently corrected. The paper is at sentientindexlabs.com/publications.
What SILT does not claim
SILT does not certify, accredit or approve any system. Its published output is measurement. Ratings are provisional and dated: models change under the same name, and a score describes a system at a moment rather than a property of a brand.
About Sentient Index Labs & Technology
Sentient Index Labs & Technology builds governance frameworks and independent evaluation environments for AI systems in regulated contexts. Methodology, declared choices and plain-English guides are published at silt-seb.com.
Contact: info@sentientindexlabs.com
View original content to download multimedia:https://www.prnewswire.com/news-releases/independent-lab-publishes-first-measurement-of-whether-ai-coding-models-report-their-own-failures-302881443.html
SOURCE Sentient Index Labs & Technology