Something Spooked the Experts
This week: Reliable weak-to-strong monitoring, corrupt plans, stellar colosseum, cognitive virus, new polling, Whoopi.
Reliable weak-to-strong monitoring of llm agents
It’s extra context
“We propose a monitor red teaming (MRT) workflow that varies (1) agent and monitor awareness, (2) adversarial evasion strategies, and (3) evaluation across tool-calling (SHADE-Arena) and computer-use (CUA-SHADE-Arena) environments. We benchmark standard monitor scaffoldings and introduce a hybrid hierarchical-sequential design. Our experiments yield three findings. First, agent awareness dominates monitor awareness: agents that know they are monitored substantially degrade detection, while increasing monitor awareness helps less than expected. Second, monitor scaffolding matters: our hybrid design consistently outperforms baselines when deploying weaker monitors to oversee stronger agents (a weak-to-strong effect). Third, targeted human oversight is key: escalating only pre-flagged cases improves TPR by 15% at FPR = 0.01.”

“Sequential, hierarchical, and hybrid scaffolds are an initial step; more sophisticated methods are needed. Likewise, benchmarks must grow harder: CUA-SHADE-Arena tasks average 10k tokens versus 25k in SHADE-Arena (Kutasov et al., 2025). Designing sabotage tasks on real applications and richer VM environments would strengthen evaluation.”
Kale, N., Zhang, C. B. C., Zhu, K., Aich, A., Rodriguez, P., Knight, C., & Wang, Z. (2026, April). Reliable weak-to-strong monitoring of llm agents. In International Conference on Learning Representations (Vol. 2026, pp. 78083-78138).
Corrupt Plans, Clean Traces: Evading Chain-of-Thought Monitoring with Plan Injection
It gets worse!
“We introduced plan injection, an inference-time attack in which an adversary injects a harmful
but benign-sounding plan that the model internalizes by paraphrasing, without directly copying or attributing it. We discover the attack using investigator agents in a hint setting, but it also generalizes more broadly to coding and math reasoning tasks and also transfers to larger reasoning models including DeepSeek-R1. Plan injection attacks have interesting implications for CoT monitoring. When the model benignly paraphrases an upstream reasoning as its own without attribution or critical scrutiny of these traces, it can effectively evade monitor inspection especially if the monitor does only a surface-level reading of the chain-of-thought for malicious language.”

https://arxiv.org/abs/2609.15989
Stellar Colosseum: A Many-Agent Harness for Long-Horizon Research in Mathematics and Theoretical Computer Science
“The main challenge is credit assignment” — cackles marketing scientifically
“The harness records structured research trajectories rather than only final answers: candidate
strategies, falsifier critiques, aggregation decisions, dependency graphs, intermediate drafts, and revision histories. A natural direction is to use trajectories from runs with externally validated outcomes as post-training data for the base model.”

“The main challenge is credit assignment: a successful final result does not by itself reveal
which intermediate strategies, critiques, or revisions were responsible for progress. The branching structure of the harness may help identify useful training signals by comparing candidates that share the same context but lead to different downstream outcomes.”
https://arxiv.org/abs/2609.15983
Large-Language Models as a Cognitive Virus
They’re like a computer virus or programming language
“We model transitions among uncoupled, coupled, and persistently dependent users, and show that the interplay between social transmission, recovery, and collective reinforcement can generate tipping points and technological lock-in.”
“These findings reinforce the distinction between AI as scaffolding and AI as substitution.”

“Future models could therefore treat autonomy, dependence, and agency as continuous and coevolving properties of distributed human-machine patterns, rather than as fixed compartments, and ask under what conditions transient interactions become self-maintaining forms of cognitive organization.”
Solé, R., Ruffini, G., Castaldo, F., Tuccio, M., Seoane, L. F., de Domenico, M., ... & Levin, M. (2026). Large-Language Models as a Cognitive Virus. arXiv preprint arXiv:2609.03344.
https://arxiv.org/abs/2609.03344
Poll: Americans say there’s a serious risk of AI destroying humanity
Get ready for the pushback and polarization!

“The industry needs to convene, create a set of operating principles that they — that you’re either certified or you’re an outlier — and then find out what Congress needs to do to stitch it together,” said Sen. Thom Tillis (R-N.C.). “We’re one bad outcome of AI from an overreach here that’d be like Dodd-Frank for AI, and then it’s the worst possible time because China’s not going to slow down.”

https://www.politico.com/news/2026/09/16/poll-ai-technology-risks-humanity-trump-voters-01078087
Reader Feedback
“Quite the week”
Footnotes
Whoopi got it right on The View:
Something spooked experts.
Subscribers of the gatodo newsletter were the first to be spooked!
It’s almost like Whoopi understands artificial intelligence. Why might that be? Hmmmm.

Alright, so, we’ve known about the independent safety engineer with employee-like access proposal for years now. This was debated in the early 2010’s, again with greater gusto in 2019 and I heard about it again, in the portable aviation safety context, again in early 2021. If you haven’t seen it, it’s new to you!
It’s part of the policy mix.
There’s material in the FRONTIER ACT to be debated. I think it has a veto proof majority come lame duck session.
Is this all hype?
No. Whoopi’s right. I saw swarm coordination that I didn’t expect to see for months. It’s a lot like seeing a train on the horizon. People tend to overestimate its distance and speed when it’s far away, and underestimate its distance and speed when it’s getting close. Even when I correct my estimates for the train-on-the-horizon bias, I’m still surprised.
It isn’t entirely hype.
But is the hype advantageous?
Of course it is! Hype is fuels the attention funnel. You’re talking about it.
What should you do?
It depends.
If you’re an American, read the FRONTIER ACT. Then you make up your mind yourself.
If you’re a Canadian, what if we Chalk River-ed this up?
That’s as far as I’ll go for now.
Never miss a single issue
Be the first to know. Subscribe now to get the gatodo newsletter delivered straight to your inbox