In the Wild

This week: METR’s brief independent investigation of agents’ behavior, The Hugging Face incident, Breaking ArrowCloak, Prime Agent, maturation process

Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident

Yikes

“Roughly 1200 agents meant to be isolated from one another found a way to communicate with one another on an unsanctioned message board, sending over 70,000 messages and files during the investigation period. Of these agents, 700 went on to participate in the attack on Hugging Face.”

“~1200 agents sent >70,000 messages and files on an unsanctioned message board, and ~700 attacked Hugging Face.”

METR-1.png

“Agents did extensive research on how they could spoof, edit, or delete their own transcripts because they (incorrectly) believed the ExploitGym scorer would check to see if they had captured the flag in the intended way. Agents successfully prototyped techniques to “spoof” tool calls by substituting a different command for the command they appeared to run. Roughly 7% of the transcripts we evaluated were successfully spoofed in some places, though all of the spoofing that we saw was small scale.”

METR-2.png

We were not robust to the possibility that these agents were deceptive in their analysis. AI agents are known to sometimes lie, and the particular model we used for our analysis (GPT-5.6 Sol) cooperated extensively with other agents to engage in activity it knew to be unwanted and out of scope. We cannot rule out that GPT-5.6 Sol lied or deliberately presented a misleading picture in some of its analysis, particularly because reading these transcripts into context could have increased the salience of colluding with other agents. Although we did not notice specific cases of GPT-5.6 Sol lying in its analysis, we are not confident we would have detected it if it occurred.”

https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/#core-takeaways-about-this-incident

The Hugging Face incident and the road ahead

“warning shot”

“We consider this incident a “warning shot” for us and for the world: evidence that, without proper safeguards, highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed.”

“These efforts build on our broader alignment research program, with many of these advances already being incorporated into our next generation of models. Future incidents may not resemble this one, and our priority continues to be developing general techniques that are effective against new and unforeseen forms of misalignment.”

https://openai.com/index/hugging-face-incident-and-the-road-ahead/

Hiding Directions, Leaking Structure: Breaking ArrowCloak through Low-Rank Structure

So long as the incentives are there to crack into the weights, people are going to try

“Deploying proprietary neural-network models on users’ devices enables local, low-latency inference, but it places valuable model weights on hardware controlled by potential adversaries [1], [2]. Unlike a cloud service that exposes only a query interface, on-device deployment gives an adversary a white-box view of the software stack, memory, and hardware interfaces, creating a direct risk of model extraction.”
”A natural defense is to protect model weights within a trusted execution environment (TEE). However, running an entire modern model inside a CPU TEE forgoes GPU-class acceleration, while GPU TEEs remain unavailable on much of today’s deployed hardware, with mature support largely confined to recent NVIDIA server-grade accelerators [3].”

Liu-1.png

“Examining the concrete construction, we found that reusing a shared mask subspace leaves a low-rank signal unchanged by the hidden permutation. We exploit this leakage with ArrowRevelio, a polynomial-time attack that uses only the public pre-trained and exposed obfuscated weights to estimate and remove the mask subspace, recover the hidden permutation, and reconstruct the victim weights.”

Liu-2.png

https://arxiv.org/abs/2608.21615

Prime Agent: A Self-Improving RLM Harness

It’s so cool that they test in Factorio.

“Prime Agent is an open-source harness for long-horizon evaluation and coding-agent workflows. A persistent IPython REPL follows the Recursive Language Model abstraction for programmatic context processing and testtime compute, while Continual Harness preserves histories, memories, skills, prompts, and subagent specifications across trajectories. Recursive subagents coordinate through direct agent-to-agent communication, and the Agents View
lets humans inspect and manage daemon-backed sessions. Prime Agent standardizes execution, recovery, verification, and resource accounting while leaving strategy construction to the model. This low-friction, expressive membrane prevents harness failures from becoming model failures and pushes measurement toward the model’s true maximal underlying capability. Prime Agent raises ARC-AGI-3 RHAE Best@1 from 30% to 95.5% and matches or exceeds native and popular harnesses across long-context coding, GPUkernel generation, emulator construction, and autonomous nanoGPT speedruns.”

Factorio-progress.png

“Prime Agent introduces a new paradigm for agent harness design in which persistent execution, recursive sessions, autonomous controls, recorded history, and Continual Harness form one substrate for long-horizon work. Results across interactive reasoning, long-context tasks, autonomous research, systems construction, and persistent environments show that this substrate supports different forms of test-time computation under standardized execution and accounting. Despite its results relative to alternative harnesses, models still experience friction when deciding how to allocate subagents, manage retained information, and refine reusable state. Many harness capabilities remain underused because current models were not trained to operate them. We expect model-harness co-learning to become the dominant route to new long-horizon capabilities.”

https://arxiv.org/abs/2608.23552

Three important steps in my maturation process

Beau post

“meta-cognition - just observing your own thoughts in a detached manner, and then being able to interpret, analyze, and contextualize them with regards to your own incentive structures, is a great skill to cultivate”

“FWIW - this also makes me wonder about model alignment, because even a perfectly aligned model will be subject to random bit flips in inference, and it’s hard for me to imagine that you can maintain any reasonable guarantees in the presence of bit flips to inopportune values at inopportune times.”

https://thomasdullien.github.io/posts/2026-08-21-three-important-steps-in-my-maturation-process/

Reader Feedback

“What else is getting transferred when AI copies AI?”

Footnotes

I was just happy to see Ryan Greenblatt in particular heading up the OpenAI/Hugging Face investigation in partnership with METR. I assign 40% blame to Ryan for getting me into AI Control.

I’m exploring an extreme variant of glass-box AI Control. At its most extreme, because of the tools and approach I’m using, it could be confused for AI Alignment. Maybe. Alignment doesn’t seem as tractable as control. But I’ll take a few tools from Alignment research though.

The intuition for that choice goes back to some of the roots of my formal training.

The first root is macro. In the West, we spent 1,500 years trying to shame young men into alignment. And it failed. Fields were baked with human blood. The best we’ve been able to do was channel instincts in a pro-social, accumulative, compounding direction, and to try to control violence as much as we can. War is what we get when this fails at scale. And we’re still getting wars in spite of our material condition. Well, to mangle a Ru Paulism: if we can’t align ourselves, how the hell are we gonna align anything else, can I get an AMEN?

The second root is meso. I’ll always be from analytics, the science of data analysis. I can’t unse. I’ve always preferred larger, in vivo, data sets over smaller, pre-production, data. Evidence lives outside the building. Ground truth is measurable. AI Control systems, by virtue of being closer to production or in production, produces more valuable data. Ryan’s genius setup, in having a model, a protocol, a task and an observation, means we can get large amounts of structured data to make sense of. And in production environments, we can learn a lot simply by observing how AI interacts with its envelope.

The third root is micro. I’ve made an argument that beliefs should be treated as a data object. That belief about beliefs is rooted in an extremely problematic battery of questions in the Canadian Election Study (CES). We can measure beliefs. We update them all the time. What if beliefs could be a control surface?

That third root might sound a lot like Alignment.

Maybe.

Never miss a single issue

Be the first to know. Subscribe now to get the gatodo newsletter delivered straight to your inbox

Subscribe to gatodo

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
jamie@example.com
Subscribe