The World Inside

This week: Mental world modeling, execution trajectories, benchmark saturation, social environment design, omnicidal futures, pandemic risk in shared socioeconomic pathways

Mental World Modeling

Consider this…from the monitor’s perspective

“We formulate Mental World Modeling (MWM), a generic theoretical framework that makes mental variables core components of a world model rather than posthoc rationales: MWM maintains a coupled physical-mental world state, renders a target-specific partial observation, and simulates how candidate actions jointly update both components. We instantiate the framework in MENTIS, a training-free and fully inspectable baseline that decomposes the process into state parsing, target-observation generation, action decomposition, coupled physical and mental transition, and branch-level value evaluation.”

Fei-1.png

“Mental World Modeling is a target-centric, action-conditioned world-modeling framework in which an external simulator (i) maintains a joint physical-mental world state, (ii) renders a first-person partial observation for a specified target agent, and (iii) simulates the next joint state conditioned on the target’s action. The state space factors as S = S phy × Sment, where S phy stores entities, relations, and environmental conditions, and S ment stores latent mental-social variables such as beliefs, attention, goals, intentions, emotions, preferences, norms, role relations, and atmosphere.”

“The minimal object of simulation for social decision-making is not a physical trajectory alone and not a theory-of-mind answer alone. It is an action-conditioned transition over a coupled state: physical variables constrain what can happen, while mental variables determine what the same happening means to the agents involved.”

Fei, H., & Zhao, Y. (2026). Mental World Modeling. arXiv preprint arXiv:2607.27201.

https://arxiv.org/pdf/2607.27201

Agent Skills Matter: Inferring Proprietary Skills from Execution Trajectories

Skill Leakage

Geng-1.png

“In such black-box deployments, skills are formalized as core intellectual property (IP) with strictly proprietary attributes. Providers retain these artifacts in controlled backends while exposing only the resulting agent capabilities through service interfaces, thereby preserving both their technical privacy and commercial viability.”

“Consequently, backend retention does not conceal a skill’s behavioral effects. An
adversary may exploit these traces to infer and independently deploy a functional approximation,
enabling unauthorized replication or redistribution and undermining providers’ intellectual property and competitive advantage. Yet this trajectory based threat remains underexplored. As illustrated in Figure 1, we define Skill Leakage as the unauthorized inference of proprietary skill content from execution trajectories observed through black-box interactions with the target agent.”

“(1) We formulate Skill Leakage as a trajectory based threat whereby adversaries infer proprietary skills from unlabeled trajectories elicited by benign diagnostic queries, and provide empirical evidence that distinct functional profiles and trajectory-level skill signatures make agent trajectories a behavioral side channel for skill inference.
(2) We propose SigLeak, a two-stage framework that constructs benign diagnostic probes and infers reusable skill instructions through contrastive comparison of paired skill-enabled and skill disabled trajectories, without access to the original skill or correctness annotations.”

Geng, J., He, R., Fei, Z., Yi, B., Wang, R., Liu, Z., ... & Zeng, Q. (2026). Agent Skills Matter: Inferring Proprietary Skills from Execution Trajectories. arXiv preprint arXiv:2607.25560.

https://arxiv.org/pdf/2607.25560

When ai benchmarks plateau: A systematic study of benchmark saturation

Undifferentiated

“A benchmark is saturated if the evaluated models can not be reliably distinguished by their performance scores and any further improvements are not statistically distinguishable under the evaluation protocol. Formally, saturation is characterized by: (1) statistically alike performance among different top performing models (2) top performing models are approaching the benchmark’s empirically inferred ceiling.”

“Saturation is widespread. Of the 60 benchmarks analyzed, 29 exhibit high or very high saturation (Sindex ≥ 0.7), out of which 14 fall into the very high category (Sindex ≥ 0.9). These benchmarks show strong compression among top-performing models, indicating limited discriminative power at the frontier. Across benchmarks, larger test sets are associated with lower saturation indices. Benchmarks with more test items show less score compression among top models, consistent with lower evaluation uncertainty and higher resolution. This relationship persists in joint regression (Section 4.2), suggesting that measurement scale impacts discriminative power.”

Akhtar-1.png

Akhtar, M., Reuel, A., Soni, P., Ahuja, S., Ammanamanchi, P. S., Rawal, R., ... & Solaiman, I. (2026). When ai benchmarks plateau: A systematic study of benchmark saturation. arXiv preprint arXiv:2602.16763.

https://arxiv.org/abs/2602.16763

Social environment design

What happens when the players are a mix of synthetic and organic?

“This paper proposes a new research agenda … by introducing Social Environment Design, a general framework for the use of AI in automated policy-making that connects with the Reinforcement Learning, EconCS, and Computational Social Choice communities.”

Zhang-1.png

“In our framework, illustrated in Figure 1, we suggest addressing the concern of a misaligned
policy-maker with “Voting on Values (Hanson, 2013),” coupled with a Principal policy-maker (or environment designer) who seeks to achieve suggested policy goals. We capture the complexity of a general economic environment whilst maintaining computational tractability by modeling the economy as a Partially Observable Markov Game (POMG), which maintains a fixed observation space for each agent. Finally, we structure our framework as repeatedly finding Stackelberg Equilbria, enabling theoretical understanding by allowing reduction to simpler subproblems.”

“Computational social choice is an interdisciplinary field combining computer science and social choice theory, focusing on the application of computational techniques to social choice mechanisms (such as voting rules or fair allocation procedures) and the theoretical analysis of these mechanisms with computational tools (Brandt et al., 2016). A fundamental component of the field is the study of manipulative behavior in elections and other collective decision making processes, as well as the design of systems resistant to manipulation (Elkind et al., 2010; Procaccia, 2010). This area of study will likely inform the development of the Voting Mechanism. Additionally, computational social choice attempts to optimize the fair distribution of resources, often involving complex allocation problems (Thomson, 2016; Procaccia, 2016).”

Zhang, E., Zhao, S., Wang, T., Hossain, S., Gasztowtt, H., Zheng, S., ... & Chen, Y. (2024). Social environment design. arXiv preprint arXiv:2402.14090.

https://arxiv.org/pdf/2402.14090

A Taxonomy of Omnicidal Futures Involving Artificial Intelligence

But how?

“Hundreds of leading scientists, developers, and other public figures have warned that human extinction is a potential result of artificial intelligence development. Even so, many ask in response “But how?””.

“Our taxonomy is simple. For any hypothetical AI-driven omnicide, one can ask if the omnicide was somehow undergirded by an intent to kill. If no, it is a case of unintentional omnicide (Section 1). If yes (Section 1), one can subdivide into four cases (2a-1d) based on the seat of that intention: human states, human institutions, human individuals, or AI itself.”

“These five categories are exhaustive — any omnicidal event will fall into at least one of them. They can also be made mutually exclusive by categorizing a scenario into the earliest category that admits it as a case.”

critch-1.png

Critch, A., & Tsimerman, J. (2025). A Taxonomy of Omnicidal Futures Involving Artificial Intelligence. arXiv preprint arXiv:2507.09369.

https://arxiv.org/abs/2507.09369

Pandemic risk in the Shared Socioeconomic Pathways

We’re connected

“For over a decade, the Shared Socioeconomic Pathways (SSPs) have served as the principal framework for quantitative modeling of the socioeconomic dimensions of global environmental change. The SSP scenarios describe many of the ecological and social processes thought to shape pandemic risk, including the emergence of novel pathogens (accelerated by processes such as deforestation, livestock intensification, and land-use change) and their subsequent spread (mediated by factors such as inequality, human mobility, and health system capacity).”

Lavelle-1.png

“These findings also reinforce a growing consensus in disease ecology: escalating pandemic risk should not be considered separately from climate change and biodiversity loss, but instead, as another manifestation of the same underlying polycrisis of social-ecological systems.”

Lavelle, T., Sanchez, C., Andrijevic, M., Becker, D. J., Gibb, R., Gonsalves, G. S., ... & Carlson, C. J. (2026). Pandemic risk in the Shared Socioeconomic Pathways. medRxiv, 2026-07.

https://www.medrxiv.org/content/10.64898/2026.07.29.26359250v1

Reader Feedback

“LLM’s are kind of like us right? Multiple personas all sharing the same space.”

Footnotes

I’m giving a talk at Trajectory Labs this Thursday, August 13. It’s titled “A Rough Introduction to AI Control” and you can register here:

https://luma.com/trajec-r4gd?tk=oHRV4w

The fun, tricky, bit is that it’s a mix of people. The audience is a mix of members from TL, many of them experts in AI Safety, folks from tech, journalism, entrepreneurs, academics and the curious. So I’ve broken down the talk into three learning objectives.

By the end of this talk, you’ll be able to:

Explain what AI Control is;

Explain what kind of risk AI Control can mitigate;

Explain core concepts of AI Control.

And that ought to be able to take care of both the front and the back of the room. Experts won’t get too bored with the basic material, and there’ll be a nice ramp for the public towards the end.

I’ll see some of you there.

Never miss a single issue

Be the first to know. Subscribe now to get the gatodo newsletter delivered straight to your inbox

Subscribe to gatodo

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
jamie@example.com
Subscribe