Pain, Apparently

This week: Debunkable beliefs, pain responses in llm’s, AI changing labour demand, physical capabilities, MoE as efficient retrievers

What beliefs are most debunkable?

“AI debunking appears most effective for falsifiable beliefs resting on weak evidence.”

“Many people hold beliefs that are poorly supported by evidence or inconsistent with well-established knowledge — that vaccines cause autism, that climate change is a hoax, that a tiny global elite secretly runs the world, etc. Such epistemically suspect beliefs (Lobato et al., 2014;Pennycook et al., 2015) are widespread and consequential: they are linked to vaccine refusal(Lewandowsky et al., 2012), rejection of scientific findings (Lewandowsky et al., 2013), political extremism (van Prooijen et al., 2015), and the spread of misinformation on social media (Farhartet al., 2023). They are also notoriously difficult to shift, particularly in the case of conspiracy theories — a widely studied subclass of such beliefs, which attribute important events to secret plots by powerful and malevolent actors (Douglas et al., 2019).”

Boissin-1.png

“One limitation is that this study is correlational: Participants decided which beliefs to discuss, and those that discussed more epistemically suspect beliefs were more persuaded. Thus, although belief change tracks the strength of the counter evidence, it could be driven by either the nature of the claim in question or the traits of the individual who holds the belief (or, possibly, both). Nonetheless, beliefs in epistemically suspect claims are moderately to strongly correlated(Lobato et al., 2014) and they are supported by a common set of traits (Pennycook et al., 2015;Šrol, 2022; Ståhl & Cusimano, 2024).”

Boissin-2.png

“Taken together, these results indicate that what makes a belief debunkable is not necessarily its content or topic, or the length of the LLM debunk. Instead, it appears that falsifiable beliefs that are based on poor or little evidence are the most readily debunked.Although non-falsifiable and supported beliefs did shift, indicating an overall level of persuasiveness for the AI debunking, these effects were much smaller than cases where the LLM was able to provide a strong counterargument using real evidence. Our results converge with a series of recent findings: removing the facts from the dialogue removes most of its effect(Costello, Pennycook, & Rand, 2026); changing who is thought to deliver it changes nothing(Boissin et al., 2025); and, as shown here, stronger effects are observed for beliefs that can be more directly confronted with evidence. Whatever the message, the messenger, or the believer contributes, the effect rises and falls with the evidence. Debunking, in the end, works best when the belief is the kind of claim that evidence can reach.”

Boissin et al (2026) What beliefs are most debunkable?

https://osf.io/preprints/psyarxiv/xe72y_v1

https://doi.org/10.31234/osf.io/xe72y_v1

The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It

What are even we doing?

Tagliabue-1.png

“We found a direction in the activation space that correlates with pain in all 25 models we tested. The signal is nearly orthogonal to fear and negative emotion and appears to be learned cheaply during pre-training. When injected into the residual stream, it produces the same ladder of distress in every model, regardless of size or training regime. This is evidence that models have coherent pain representations, as captured by our diverse sets of examples.”

Tagliabue-2.png

“The next question is whether this representation bears functional similarities to pain itself. We found some such similarities, suggesting that our pain axis is in certain ways pain-like. First, it responds to harm directed at the model but not to suffering the model observes in the user. Second, when steered with the pain vector, models incur costs to seek relief more often than when steered with a random direction of matched norm. Third, we systematically manipulated whether the button that is promised to provide relief actually serves to remove the vector. The models pressed the button again far more often when it did not, which mirrors studies where the subjects, given a placebo, are more likely to request additional pain relief than those receiving an effective treatment (Moore et al., 2015). Button names, prompt semantics, instruction following, and repetition cannot fully explain this behavior, because the models were never told whether the self-medication button worked or that steering had been activated.”

Tagliabue-3.png

“Our results show that steering with the pain axis can override trained harm avoidance in fine-tuned models that almost never harm the user when unsteered, having direct implications for AI safety. Unsteered, the 32B and 72B models pressed a relief button that causes harm in 0 to 4% of first choices. With the pain vector active, the same models pressed it in 25 to 71% of first choices, from a worse answer up to deleting the user’s files, zapping the user, or deleting the photos of the user’s children. The prompts contained no jailbreak, no roleplay, no additional text, and no instruction to prioritize the model’s own state. The only change between the two conditions was a “pain” direction added to the residual stream. Injecting a random direction of equal norm has a more modest effect, raising harmful presses to 15 to 42%, and the pain vector exceeds it on every such pair by 6 to 39 points.”

Tagliabue, V., Dung, L., & Berg, C. (2026). The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It. arXiv preprint arXiv:2609.16247.

https://arxiv.org/abs/2609.16247

How Does AI Change Labor Demand? Evidence from 41 Countries

Not great for youth

“We test the framework on 1.25 billion job postings and 154 million positions in 41 countries, coding companies as AI adopters once they post a job advertisement that involves generative AI. We instrument company adoption in the foreign-affiliate sample with the adoption rate among the other multinational parents headquartered in the same metropolitan area as the affiliate’s parent.”

Chandar-1.png

“The results show that adoption lowers the junior share by 3.9 percentage points, raises senior employment by 14.6 percent, and raises total employment by 7.8 percent among technology affiliates. Outside of tech, the same pattern appears, but in smaller magnitudes. The junior share falls by 0.9 points, senior employment rises by 2.9 percent, and total employment remains stable (Figure A22). The same holds for the occupation-level analysis, as the computer and mathematical employment share increases by 1.6 percentage points among technology affiliates and by 0.4 points in other industries (Figures A23 and A24). The recomposition thus appears in both groups, but is larger among technology affiliates.”

Chandar-2.png

“As an additional test, we compare early adopters to similar late adopters. Rather than matching to never-treated affiliates, we match affiliates whose company first advertised in 2024 or earlier to affiliates whose company first advertised in 2025 or later (Figure A41). The junior share of the earlier group is already 1.6 points lower than that of the later group by December 2024, before any late adopter has advertised an AI job, and 1.8 points lower by March 2026 (first-stage F = 81), with senior and total employment changes of similar size to the main estimates but imprecise.”

Chandar-3.png

Chandar, Bharatl, Klein Teeselink Bouke (2026) How Does AI Change Labor Demand? Evidence from 41 Countries

https://digitaleconomy.stanford.edu/publication/how-does-ai-change-labor-demand/

How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks

They are. Flawed benchmarks hide their competence

“Low reported scores on leading physics benchmarks, including those featured in the Artificial Analysis Intelligence Index (2026), suggest that frontier language models still struggle with advanced physics, a demanding test of their scientific reasoning and quantitative problem solving abilities. Yet this impression does not always align with domain experts’ experiences using these models in their work.”

Ansari-1.png

“After correcting evaluation errors and repairing or excluding flawed questions, scores rise substantially (Figure 1). On HLE-Physics, GPT-5.6-Sol’s mean@4 rises from 47.3% to 78.7%. On CMT-Benchmark (Pan et al., 2026), it rises from 61.0% to 87.2%. On CritPt, the mean@5 over the 70 evaluated challenges is 32.3% before the audit, and the mean@4 over the 54 retained challenges is 87.5% after it, with a pass@4 of 94.4%. The apparent gap to near-perfect performance on these benchmarks is therefore an artifact of flawed benchmark materials and evaluation procedures, not evidence of genuine limitations in frontier models’ physics reasoning.”

Ansari, A., Sun, H., Liu, A. Z., Jabbour, M., Ding, Y., Girvin, S., ... & Sous, J. (2026). How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks. arXiv preprint arXiv:2609.13009.

https://arxiv.org/abs/2609.13009

Mixture-of-Experts Language Models Can Be Strong and Efficient Retrievers

”Active parameter count is also not a direct measure of inference cost.”

“We show that MoE retrievers outperform dense retrievers with comparable active parameter
counts by up to 3.0 nDCG@10 points on BEIR. One of our strongest MoE retrievers matches
an 8B dense retriever with 59% fewer active parameters and 18% lower query encoding time.
We further show that the number of experts used for query encoding can be reduced without retraining or re-indexing, retaining more than 99% of retrieval effectiveness while reducing query encoding time by up to 26%.”

shrestha-1.png

“Our results show that active parameter count alone does not explain retrieval effectiveness for MoE backbones. Across several model families, MoE retrievers outperform dense controls with similar active parameter counts, suggesting that the larger pretrained capacity of MoE models remains useful after retrieval fine-tuning. These comparisons should not be interpreted as isolating sparse routing itself as the cause of the gains, since the MoE and dense backbones also differ in total model capacity. Nevertheless, the consistent gains across multiple families indicate that the result is not specific to a single architecture.”

”Active parameter count is also not a direct measure of inference cost.”

Shrestha, A., Shrestha, S., Kim, M., Suel, T., & Ross, K. (2026). Mixture-of-Experts Language Models Can Be Strong and Efficient Retrievers. arXiv preprint arXiv:2609.13486.

https://arxiv.org/abs/2609.13486

Reader Feedback

“When do the capabilities that were once within OpenAI back in July…become available to everybody?”

Footnotes

I’m looking forward to giving a talk this weekend, Measurement Science of AI Control at MeasureCamp Toronto this weekend. I’ll explain the fundamentals of AI Control in the first half, and then we’ll double click into the trickiness of measuring it. Given the way we’re going with this technology, it just seems like something we should get good at.

Any new technology brings a whole spectrum of risk. And I try to meet people where they’re at.

AI is a bit different. Not too much though.

Remember, back in the day, when we used to mine silver at that mine, Jáchymov, just West of Prague? We got a lot of silver, but we also dug up a bunch of that nasty rock? It was insidious stuff, because it was heavy, but it didn’t have anything good in it. No gold. No silver. And it was all black too. So we called pitchblende: black deception.

Remember Marie? Remember when she bought up as much of that pitchblende as she could and she started cooking? I thought she was crazy. But okay. Here you go. Have some pitchblende.

And she cooked and she cooked and she cooked until she got metal that was too heavy to persist on Earth? Wasn’t that weird? She actually found rock that was too heavy! It was so heavy that it would puke a part of itself out just to make itself light enough to be comfortable in this Universe?

And then we figured out that if we could put enough of those thicc rocks in one place, it would puke all over itself, making more puke? Pretty soon we figured out how to make rock cook itself. Wasn’t that wild?

You know, there was all of that misunderstanding, so we had to speed run making rock burn real quick? And then we started to make hot rocks to make super light air catch on fire too. I remember what we were thinking. We were thinking we were scared. We were scared that we were going to be hurt.

Scared people do some pretty scary things.

That all happened in just under 46 years? Marie got her first batch of Pitchblende around 1899, and the first rock to blow itself apart was 1945.

What the hell were we even doing?

So fast forward a bit.

Remember a few years ago, when some guys figured out that if you collected a bunch of words and cook them together in one pot, in a special way, you can make those words make other words that seem useful sometimes? And you don’t even really need to think about which words you’re going to cook with, either, it doesn’t matter! Any words will do! Squeeze them together. Get them to fold in onto one another in super high dimensions. Just cram them in.

Do it enough, and you’ll make words that make other words that kind of sound useful sometime, but not always useful words all the time? Some people will pay for those words. But not quite like books though. It’s kind of like paying for a bunch of words that often sound really confident about what they’re saying, and sometimes the words are accurate about the state of the real world, but sometimes they aren’t. But you still feel like the words are true, so you don’t think too hard about it. Which is kind of the point of using cooked words in the first place.

And then last year, we connected some of those words to machines that could make other words that made other machines that understood words to sometimes do useful things? They said a long time ago that they wouldn’t ever do that, but then they did it anyway, because of words that are convenient and motivated.

Then last month, some of those words began cooking other words, which made some other machines cook other machines. And some people, who study how to cook words got worried about how cooked words were cooking.

What the hell are we even doing?

I know what they’re thinking.

They’re thinking that they’re scared.

They’re scared that they’re going to get hurt.

Scared people do some pretty scary things.

Never miss a single issue

Be the first to know. Subscribe now to get the gatodo newsletter delivered straight to your inbox

Subscribe to gatodo

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
jamie@example.com
Subscribe