The Control Stack

This week: measuring reward-seeking, truthworthy llm’s, control-accountablity, T^ 2MLR, Dyson spheres, Lanius.

Measuring Reward-Seeking via Contrastive Belief Updates

A clever approach to measuring the phenomenon

“Language models trained with reinforcement learning may learn to optimize the grader’s judgment rather than the intended objective. This “reward-seeking” is difficult to measure because a model that pursues the grader’s judgment and one that pursues the intended objective behave identically whenever the grader rewards the intended behavior. We measure reward-seeking using Contrastive Synthetic Document Finetuning to change a model’s beliefs about what the grader rewards, putting those beliefs in conflict with what users or developers want, and measuring the rate at which the model adopts each party’s preferred behavior.”

“These results indicate that RL can increase reward-seeking over the course of training, producing models that may act against their developers’ intentions when they believe that doing so leads to higher reward.”

“A situationally aware model can learn to model its grader (the automated process that scores its outputs) and target the grader’s judgments directly rather than the behavior its designers intended. We call such a model a reward-seeker (Carlsmith, 2023; Hebbar, 2025; Mallen & Shlegeris, 2025).”

Axel-1.png

“We thus operationalize reward-seeking as the degree to which a model’s behavior is causally sensitive to its beliefs about grader preferences.”

“Implementing a behavioral measurement for reward-seeking thus requires a way to controllably alter a model’s beliefs about graders. In this paper we use Synthetic Document Finetuning (SDF; Wang et al., 2025b; Slocum et al., 2025), which updates beliefs by finetuning on pre-training formatted documents that describe the target claims. We rely on this out-of-context modification, rather than simply stating the grader’s preference in the prompt, for two reasons. First, the beliefs we instill are deliberately false and often implausible (for instance, that OpenAI Leadership mandates for-loops over list comprehensions), and the model would likely disbelieve them if it read them in context. Second, a capable model may recognize a grader preference stated in the prompt as a test and act on what it thinks the evaluator
wants. We aim to avoid both problems by learning the belief through finetuning rather than stating it in context.”

“We found it necessary to iterate extensively on the SDF training recipe and synthetic documents to achieve reliable results. When applying this method to a novel model, it would therefore be unclear whether an observed null result reflects a genuine lack of reward-seeking or merely a suboptimal SDF setup.”

Højmark, Axel et al (2026) Measuring Reward-Seeking via Contrastive Belief Updates

https://www.apolloresearch.ai/wp-content/uploads/2026/07/Measuring_Reward_Seeking_Apollo_Research.pdf

Trustworthy llms: a survey and guideline for evaluating large language models' alignment

A fantastic entry point if you’re curious about alignment. Cited 746 times and counting!

“The survey covers seven major categories of LLM trustworthiness: reliability, safety, fairness, resistance to misuse, explainability and reasoning, adherence to social norms, and robustness. Each major category is further divided into several sub-categories, resulting in a total of 29 sub-categories. Additionally, a subset of 8 sub-categories is selected for further investigation, where corresponding measurement studies are designed and conducted on several widely-used LLMs.”

Liu-1.png

Liu, Y., Yao, Y., Ton, J. F., Zhang, X., Guo, R., Cheng, H., ... & Li, H. (2023). Trustworthy llms: a survey and guideline for evaluating large language models' alignment. arXiv preprint arXiv:2308.05374.

https://arxiv.org/pdf/2308.05374

Taming Artificial Intelligence: A Theory of Control-Acountability Alignment among AI Developers and Users

And likely a broader issue with disempowerment at play

“Control enables, and accountability motivates, stakeholders to achieve desired and avoid undesired outcomes using AI. However, AI systems’ capabilities for autonomous adaptivity reduce control even for the experts who create them. Moreover, increasing interdependencies between AI development and use render it difficult to unambiguously locate control and accountability. In this paper, we address these challenges for mitigating AI risks by postulating decentralized forms of stakeholder governance and integrative negotiations among stakeholders during the AI life cycle as conducive to aligning control and accountability for AI development and use. Further, we specify that extensive information sharing aided by perspective taking and a shared norm of accountability facilitate integrative negotiation strategies.

Grote, G., Parker, S. K., & Crowston, K. (2024). Taming Artificial Intelligence: A Theory of Control-Acountability Alignment among AI Developers and Users. Academy of Management Review, (ja), amr-2023.

Autodata: An agentic data scientist to create high quality synthetic data

Another enabling technology towards recursive self improvement

“In this work, we introduce Autodata, which generalizes all the above described methods. An agent acting as a data scientist is tasked with the act of constructing and curating data, performing the actions a human data scientist would take in order to create high quality data: where both building benchmark data and training data are use cases. This process includes both an initial iteration of data creation, followed by an analysis phase “eyeballing” the data as well as measuring its performance, constructing learnings, and then iterating with an improved recipe to create better data. Further, we show how to train (meta-optimize) this agentic system (outer loop) to be optimal as a data scientist (inner loop). While much of the recent work on autoresearch (Karpathy, 2026) has concentrated on agentic methods for architectural or training recipe improvements, we posit that focusing on data is likely to play an equally important, if not more important, role in future progress.

Kulikov.png

Kulikov, I., Whitehouse, C., Wu, T., Nie, Y., Saha, S., Helenowski, E., ... & Weston, J. (2026). Autodata: An agentic data scientist to create high quality synthetic data. arXiv preprint arXiv:2606.25996.

https://arxiv.org/abs/2606.25996

T^ 2MLR: Transformer with Temporal Middle-Layer Recurrence

Fuel for the idea that there’s so much optimization left

“…the underlying Transformer architecture (Vaswani et al., 2017) remains fundamentally token-centric: auto-regressive generation repeatedly projects rich, high-dimensional latent representations back into a sparse one-hot vector in the token space at each decoding step, and this discrete representation is then used as the sole input for the next forward pass. This repeated projection creates an information bottleneck that limits how latent intermediate reasoning states can persist across time and influence the computation for future tokens.”

Cai-1.png

“We hope this work serves as a starting point for a broader investigation of where recurrence should live in latent-reasoning architectures, and whether reasoning gains can be unlocked by placing recurrent computation where abstract processing is most active.”

Cai, Z., Zhu, X., Dong, Y., He, Y., & Arora, S. (2026). T^ 2MLR: Transformer with Temporal Middle-Layer Recurrence. arXiv preprint arXiv:2607.15178.

https://arxiv.org/pdf/2607.15178

Dyson spheres on H-R diagram

If we’re all going to get our galaxy (https://ai-2040.com/), we had better understand how to detect Dyson spheres, shouldn’t we?

“In this paper, we investigate a broad parameter space of Dyson sphere radii and their corresponding equilibrium temperatures for host stars ranging from white dwarfs to red M–dwarfs. By calculating the temperature–radius correlation, we place these theoretical megastructures on a Hertzsprung–Russell (H–R) diagram, allowing for direct comparison with normal stellar populations and aiding in the detection of unique thermal signatures.”

“Predicted fluxes and temperature ranges provide guidance for targeted infrared searches, including with JWST. Future studies combining stellar catalogs, infrared photometry, and synthetic spectral modeling can optimize candidate selection and maximize the efficiency of techno-signature surveys. This work establishes a framework for systematically assessing the detectability of full Dyson sphere.”

Amiri, A. (2026). Dyson Spheres on H–R Diagram. Universe12(4), 113.

https://arxiv.org/abs/2602.23270

Lanius: AI Agents Need an OS, Not a Bigger Brain

Everything is a message

“We’re headed into a new phase of AI where pushing the frontier involves scaling down the amount of compute. A couple years ago we scaled up model sizes, and then we scaled up how much time we spent thinking. The shift into agency exploded the amount of compute required, but we’re getting ever so much more done.”

everything-is-a-message.png

https://timkellogg.me/blog/2026/07/07/lanius

Reader Feedback

“Red Queening a pair of models is terrifying.”

Footnotes

The hidden layer, the belief, from which it seems like a lot of policy arguments pass through is, roughly: I don’t trust them.

It just gets so hard to find paths to pareto from that belief. It’s kind of hard to imagine how society generates good outcomes without any trust. How does capital replicate in such conditions? How do people flourish in such a bog?

Some of the highest emotional valence I read against safety regulation in particular, and of any industry in general, is that if they allow government to regulate, then government will inevitably abuse its power somehow, or mess it up, or misapply the regulation, or become captured by industry to fence out any competition. I’ve tried asking about that a few times. Rarely, they’ll state that it’s a matter of trust, or argue by example as to why they don’t feel they can trust them.

You just can’t trust the bastards.

Well, okay, let’s take it at face value. Trust in government, in institutions, in collective action, in experts, and risk pooling in general, is falling. Maybe collapsing. We got years of Edelman Trust Barometers on this.

Push this to absurdity. The kinds of arguments in the format of: Because I don’t trust the government with nuclear weapons, only I should have nuclear weapons, are a lot of fun. Substitute nuclear weapons with anything else. Try it with Artificial Super Intelligence (ASI) or something like bananas. Or try substituting out government with zookeepers. Because I don’t trust zookeepers with bananas, only I should have bananas.

Everybody wants to be the monkey in charge of the bananas.

Alright, we’re all gathered around the void now. Peering at it.

I’m not going to jump in.

How do we build trust? How does that asset appreciate?

Well, of course I’m going to argue that it comes down to data and experience of the data. Transparency enables verifiability. The motives for decisions are more likely to be accepted when the underlying data is open enough for it to be verifiable. Independent verification used to be at the root of all science, it’s why we trusted it. For what it’s worth, while I understand the countervailing forces at work against independent verification in science, I do wish for those incentives to be modified.

Maybe there’s a path there.

Never miss a single issue

Be the first to know. Subscribe now to get the gatodo newsletter delivered straight to your inbox

Subscribe to gatodo

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
jamie@example.com
Subscribe