Hidden Curriculum This week: Mechanism design for alignment and control, you are what you read, cheating and whistleblowing in autonomous research swarms, confidence shortcut, course design in the age of AI.
The Lathe This week: If God Arrived As a Chatbot, Late-Stage Pressure States in Long-Horizon Tool-Use Agents, AI Control safety case, Questionnaire, The Coming Storm
In the Wild This week: METR’s brief independent investigation of agents’ behavior, The Hugging Face incident, Breaking ArrowCloak, Prime Agent, maturation process
Boundaries This week: open-weight genome language model safeguards, watermark localization, AI guardrail survival, stealing reasoning traces, A US strategy to prevent the creation of mirror life
Large Worlds This Week: Genome language models, weak verifiers, agentic inequality, cognitive commons, cooperation in large worlds, cooperative AI, genesys open models
The World Inside This week: Mental world modeling, execution trajectories, benchmark saturation, social environment design, omnicidal futures, pandemic risk in shared socioeconomic pathways
Masks and Monitors This week: Model belief, role-typed credit assignment, personascope, distillation and personas, monitoring the monitor
The Control Stack This week: measuring reward-seeking, truthworthy llm’s, control-accountablity, T^ 2MLR, Dyson spheres, Lanius.
The Intelligence Begins Watching Itself This week: SAD, distributed attacks in persistent-state AI control, metacognitive reasoning, CALIBER, the red queen godel machine, subjective self-experience, the economics of recursive self improvement
Widening the Bottleneck This week: Haiku to Opus in just 10 bits, QKV variants, next-latent prediction transformers, ai moral status, motivated reasoning, emotion concepts, AI designed radio chips, plan A