Large Worlds
This Week: Genome language models, weak verifiers, agentic inequality, cognitive commons, cooperation in large worlds, cooperative AI, genesys open models
Generative design of bacteriophages with genome language models
Wild! Wait. No. The opposite. Synthetic!

“In this work, we leveraged genome language models to achieve the first generative design of complete bacteriophage genomes. We established a computational framework for specifying our design goals, including the development of a new gene annotation method and diverse scoring metrics, allowing us to controllably design toward a target genomic architecture and host tropism. In particular, our design template was based on ΦX174, a tractable, safe, and historically significant model genome (Barrell et al., 1976; Goulian et al., 1967; Kirchberger & Ochman, 2023; Sanger et al., 1977; Smith et al., 2003). We systematically evaluated thousands of computationally generated sequences and experimentally tested nearly 300 designs, resulting in 16 viable phages containing substantial evolutionary diversity and enabling a phage cocktail that rapidly overcame bacterial resistance. Multiple generated phages exhibited increased fitness or faster lytic dynamics relative to ΦX174, demonstrating the ability of generative models to efficiently evolve high-fitness genomes.”


King, S. H., Driscoll, C. L., Li, D. B., Guo, D., Merchant, A. T., Brixi, G., ... & Hie, B. L. (2026). Generative design of bacteriophages with genome language models. Science, 393(6811), eaec2657.
https://www.biorxiv.org/content/10.1101/2025.09.12.675911v1.full.pdf
Shrinking the generation-verification gap with weak verifiers
Fantastic efficiency
“A core challenge in deploying language models (LMs) is verification: determining the quality or correctness of a model’s response. This problem arises across various components of the LM pipeline, including dataset curation, model alignment, and inference-time decision-making. Verification relies on verifiers—functions that score responses. When combined with repeated sampling—generating multiple candidate responses from a LM—a perfect verifier can be used to select a correct candidate response, significantly enhancing model capability on tasks such as math, code, and reasoning [75, 6, 58].”

“We ask: to what extent can we leverage weak verifiers to improve accuracy in the repeated sampling regime?”
“First, we establish that weighted aggregation of weak verifiers substantially outperforms both individual verifiers and majority voting across reasoning and mathematics tasks, with weighted aggregation exceeding majority voting by an average of 12.3% across all tasks explored (Table 1; Figure 2). Second, we developed a principled approach for unsupervised estimation of verifier accuracies using weak supervision, enabling effective ensemble weighting without fine-tuning on costly ground-truth annotations. This allows us to close the generation-verification gap by 12.8% for 8B models and 16.0% for 70B models (Table 3). By leveraging Weaver with 70B models, we marginally outperform frontier closed-source models such as OpenAI’s o3-mini (87.7% vs. 86.7%) on average across the tasks explored (Table 1). Third, we improve the accuracy-compute trade-off by distilling Weaver into lightweight 400M-parameter cross encoders. These distilled models retain 98.2% of Weaver’s performance while reducing inference compute by 99.97% (Section 6). This enables highthroughput, cost-efficient verification without sacrificing accuracy, and demonstrates that scalable verification can be achieved without repeatedly querying large models.”
Saad-Falcon, J., Buchanan, E. K., Chen, M. F., Huang, T. H., McLaughlin, B., Bhathal, T., ... & Ré, C. (2025). Shrinking the generation-verification gap with weak verifiers. arXiv preprint arXiv:2506.18203.
https://arxiv.org/abs/2506.18203
Agentic inequality
I’d prefer to do more than hope that they don’t pull the ladder up behind themselves…
“Autonomous AI agents capable of complex planning and action mark a shift beyond today’s generative tools. As these systems enter political and economic life, who can access them, how capable they are, and how many can be deployed will shape distributions of power and opportunity. We define this emerging challenge as “agentic inequality”: disparities in power, opportunity, and outcomes arising from unequal access to, and capabilities of, AI agents.”
“Agentic inequality has three dimensions: availability, quality, and quantity. Each captures a distinct and analytically separable source of differential advantage; together, they determine the extent to which an actor can derive value from agentic AI.”
“A second programme is normative. Once empirical patterns are clearer, the next task is determining which disparities in agent availability, quality, and quantity are socially or ethically unacceptable. Key questions here include: Which participatory methods, such as citizen assemblies, are best suited to eliciting public preferences about unacceptable capability gaps? How can these preferences be translated into concrete design principles or regulatory guardrails governing disparities in agent quality?”
Sharp, M., Bilgin, O., Gabriel, I., & Hammond, L. (2025). Agentic inequality. arXiv preprint arXiv:2510.16853.
https://arxiv.org/pdf/2510.16853
The Tragedy of the Cognitive Commons: How AI Could Disrupt the Regeneration of Professional Expertise
Potentially huge implications for professional cooperation
“This conceptual paper introduces the Cognitive Commons framework, integrating commons theory, HRD scholarship, and distributed cognition to explain how rational AI adoption decisions can deplete the shared expertise pool professions require for renewal. The framework distinguishes Internalized Mastery (deep domain knowledge from sustained practice) from Distributed Mastery (orchestrating human-AI systems), and develops the Validation Tether: effective AI oversight depends on the expertise AI adoption may undermine.”
“Organizations do need workers who can collaborate effectively with AI systems, and reskilling initiatives deserve continued investment. Alone, however, they are insufficient responses to the deeper structural problem AI adoption can create. Wiles et al.’s (2024) experimental evidence is instructive: when AI assistance was provided during a skill-building task, participants performed significantly better during the access period, but this performance advantage did not transfer to subsequent unassisted performance.”
“This requires HRD to expand its mandate beyond organizational boundaries to
encompass profession-level expertise ecosystems. The field must develop governance mechanisms enabling collective action despite free-rider incentives, prepare practitioners to contend with genuine tensions rather than assume false harmony, and advocate for developmental pathways that maintain both Distributed Mastery and Internalized Mastery.”
Lovett, N. (2026). The Tragedy of the Cognitive Commons: How AI Could Disrupt the Regeneration of Professional Expertise. Human Resource Development Review, 15344843261470602.
https://arxiv.org/abs/2607.29380
Evidential Cooperation in Large Worlds
What would the 19th copy of you do?
“Evidential Cooperation in Large worlds (ECL) refers to the idea that, since the universe may be very large, it likely contains many agents very similar to oneself, and that by taking actions that are good for these agents, we learn that they take actions that are good for us. The idea of ECL was introduced in Oesterheld, 2017 as MSR (Multiverse-wide superationality).”

Cooper, Emery (2024) Evidential Cooperation in Large Worlds. URL:
https://www.lesswrong.com/w/evidential-cooperation-in-large-worlds
https://longtermrisk.org/media/Multiverse-wide-Cooperation-via-Correlated-Decision-Making.pdf
Cooperative AI
This exists! Neat!
“Supporting research that will improve the cooperative intelligence of advanced AI for the benefit of all”
https://www.cooperativeai.com/
Genesis Open Models
New institution building? Pretty cool that the DOE is trying
“The U.S. Department of Energy (DOE), in collaboration with industry partners, has announced a new class of open-weight foundation models designed specifically to accelerate scientific discovery as part of DOE’s broader Genesis Mission initiative. These Genesis models aim to provide researchers, national laboratories, industry partners, and the broader open science community with powerful, transparent, and extensible AI models that can be adapted to a wide range of scientific domains and use cases. The first model in this class will be Genesis-Science-1, developed in partnership with Arcee.
By releasing models with open-weights, DOE seeks to galvanize the scientific and AI communities around shared infrastructure for science: enabling new workflows in materials discovery, energy systems, earth systems modeling, fusion, biology, high-energy physics, and beyond. This effort is grounded in the principles of open science, reproducibility, and responsible AI, with the goal of lowering barriers to advanced AI capabilities for the public good.
As part of this initiative, we are collecting information from interested participants who would like to contribute to and benefit from this new ecosystem:
- Open weight models: Organizations to provide open-weight models, both as base models to support downstream fine-tuning and for immediate deployment. Models should have transparent provenance.
- Pretraining contributions: Organizations and researchers with access to high-quality, domain-specific scientific data who are interested in contributing to future pretraining cycles (e.g., curated datasets, benchmarks, or specialized corpora).
- Fine-tuning efforts: Teams interested in shaping downstream, domain-adapted versions of Genesis-Science-1 (and future models) for particular scientific fields, applications, or mission needs (e.g., lab assistants, simulation surrogates, scientific copilots).”
https://genesisopenmodels.anl.gov/
Reader Feedback
“Skill leakage also happens when employees leave the company. Or at least it used to?”
Footnotes
I gave a talk last Thursday at Trajectory Labs. You can catch it here.
It went well.
What’s next are a few more grants applications to push out the door: because I’d prefer to do the next wave of research out in public. Even with a system, these cost time.
There are a few pretty severe information asymmetries in this field. Proposals and proponents have different quality attributes. Different funders differentiate their allocation process against different quality attributes. And there are all the typical issues with discernment and goal drift. It isn’t easy for anyone.
Wish me luck!
Never miss a single issue
Be the first to know. Subscribe now to get the gatodo newsletter delivered straight to your inbox