No Single Safeguard

This week: Reliability theory for AI Control, auditing sabotage bench, autobenchmark, optimal timing for superintelligence, serious AI product.

Reliability Theory for AI Control

This is promising

“Frontier AI control increasingly relies on defense in depth. Access controls, monitors, responses, and recovery mechanisms are composed so that no single failure is catastrophic. We thus must study AI security at the level of architecture, not of individual safeguards. Classical reliability theory asks how system success depends on which components work, which fail, and how those successes and failures are connected [2, 10]. It is tailor-made to answer this sort of question.”

Molnar.png

“We define a deployment route to be a specified sequence of attacker actions and system transitions that, if not stopped, carries an AI from its authorized execution context to such a rogue deployment. Researchers have proposed increasingly sophisticated AI control protocols to thwart such deployment routes [3, 25]. These control protocols deserve commensurately sophisticated analysis. Here, we believe reliability theory can earn its keep.”

Molnar-2.png

“For AI control, we may ask how many genuinely independent monitoring opportunities are needed, when voting or threshold rules outperform simple duplication, and how to allocate a fixed assurance budget across safeguards once their dependence structure is estimated.”

Molnar, G. (2026). Reliability Theory for AI Control. arXiv preprint arXiv:2609.26419.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases

An earlier version called ASMR-Bench

“Misaligned systems could sabotage the research they conduct in order to slow down the pace of research or to make AI systems appear safer than they actually are. AI companies have identified research sabotage as a risk category in their safety risk reports [8, 9, 10]. Recent work has begun evaluating sabotage capabilities in ML engineering settings [11, 12], although existing work has primarily studied sabotage in synthetic or simplified settings.”

gan-1.png

“Many humans operated as LLM-orchestrators. Two hours is not a long time to audit a
codebase, so humans often relied heavily on LLMs to understand the codebase and find sabotages, managing multiple parallel agents and synthesizing their outputs. However, some humans reported that LLMs struggled to evaluate research design choices. Humans tended to flag the same benign distractors as LLMs, but often also disagreed more fundamentally with core research methodology. Running counterfactual experiments often enabled auditors to rule out benign distractors and arrive at more accurate conclusions.”

gan-2.png

“As AI systems are increasingly used to conduct research autonomously, significantly more work is needed to develop monitoring techniques and control protocols that can reliably mitigate the risks of research sabotage.”

Gan, E., Bhatt, A., Shlegeris, B., Stastny, J., & Hebbar, V. (2026). Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases. arXiv preprint arXiv:2604.16286.

AutoBenchmark: benchmark creation & the role of humans

A sweet way to avoid the bitter lesson

To this end, we introduce AutoBenchmark, a framework in which an autoresearch agent builds a benchmark from a task specification and revises it across iterations using two feedback signals: (1) the trajectories and scores of a set of solvers that attempt the benchmark, and (2) the critique of an LLM verifier that checks the quality of the benchmark. Additionally, as a third axis, we investigate if human AI scientists can contribute to this process by directing what to build at the start of the loop and advising on revisions between iterations.

autobenchmark.png

Our experiments show that current autoresearch agents can run this loop end to end, but that the benchmarks they produce unaided (i.e., without human feedback) are close to saturated, allowing the solvers to achieve scores above 80. We find that human feedback helps substantially, and the gain grows with how specific that feedback is. Specifically, providing detailed direction on what type of benchmark to build and how to build it halves the score of the solvers, while a brief statement of what to build helps only marginally. This indicates that human involvement is still required in benchmark creation. Looking ahead, we envision AutoBenchmark as a way to measure progress on the capability of AI agents to autonomously create or co-create high-quality benchmarks.

https://facebookresearch.github.io/RAM/blogs/autobench/

Optimal Timing for Superintelligence

It’s certainly one take

“Models incorporating safety progress, temporal discounting, quality-of-life differentials, and concave QALY utilities suggest that even high catastrophe probabilities are often worth accepting. Prioritarian weighting further shortens timelines. For many parameter settings, the optimal strategy would involve moving quickly to AGI capability, then pausing briefly before full deployment: swift to harbor, slow to berth. But poorly implemented pauses could do more harm than good.”

bostrom-1.png

“We observe a clear pattern. When the initial risk is low, the optimal strategy is to launch AGI as soon as possible—unless safety progress is exceptionally rapid, in which case a brief delay of a couple of months may be warranted. As the initial risk increases, optimal wait times become longer. But unless the starting risk is very high and safety progress is sluggish, the preferred delay remains modest—typically a single-digit number of years. The situation is further illustrated in Figure 1, which shows iso-delay contours across the parameter space.”

bostrom-2.png

https://nickbostrom.com/optimal.pdf

What would a serious AI product look like?

Are we getting closer to breaking out of chat?

“The fact that every conversation is presented as this flat chat prompt that doesn’t let me interact with any of the widgets that were previously produced except through more chatting, really makes me feel like the whole product is just doing predatory social-media style “increase time on site” optimization, just trying to lure me into further repetitive and unreliable chats, rather than letting me get in, solve my problem, and get out.”

https://blog.glyph.im/2026/09/serious-ai-product.html

Bluesky Atlas

Pretty

bluesky-atlas.png

https://atlas.jazco.dev/#view=1.75/0.50000/0.50000

Reader Feedback

“Now we’re cooking with gas!”

Footnotes

I had a great time at MeasureCamp in Toronto. It was our third annual. I got in at 7:25 am to help setup the catering and walked home, exhausted, during Nuit Blanche at 2:20 am.

I learned how recent advancements in intelligent systems were affecting industry. Most of it expected. Some of it surprising.

The story of agents going back and editing source data so as to comply with the narrative that an analyst had conveniently reasoned was not on my bingo card. I was expecting to hear confusion about the value of Epistemic Security (EpiSec). Instead I heard a lot of pain from veterans of the data quality wars.

For low stakes decision making, routine decisions that don’t seem to matter, the re-writing of telemetry is merely the substitution of one form of fiction for another. This is a key source of misery from many industry line analysts a previous MeasureCamps: nothing they do seems to matter. Low stake decision making are the practice runs for the higher stakes.

For high stakes decision making, irreversible, consequential choices that matter, the re-writing of telemetry may be existential. If data quality erosion is contagious, then it’s likely its already made its way into consequential systems.

For AI Control, the stakes are high, because the risk of dangerous agentic behaviour is already present. It’s here. We simply can’t have analytical stacks with poor EpiSec. That’s a massive vulnerability. EpiSec isn’t likely to emerge as a relevant quality attribute for low-stake analytical systems, because who cares. If it mattered, data quality would matter. It doesn’t, so it doesn’t.

But I think senior executives, those facing extremely high stakes, will understand the gravity the situation.

Never miss a single issue

Be the first to know. Subscribe now to get the gatodo newsletter delivered straight to your inbox

Subscribe to gatodo

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
jamie@example.com
Subscribe