Audio version of the posts shared in the LessWrong Curated newsletter.

Summary: One area we plan to explore at Resolution is personas and character training, operationalized as finding and controlling low-dimensional structure in models that emerges in pretraining and flows through post-training to superintelligence. The hope is to expand and systematize phenomena such as emergent misalignment, subliminal learning, and other empirical persona research, then intervene on this structure without accidentally hiding undesirable behavior elsewhere. If this approach resonates with you, considering working with us. Glimmers of low-dimensional structure Our understanding of AI training and alignment as a field is very poor. If sufficient alignment of superintelligent AI agents requires pinning down the precise meaning of alignment and turning that meaning into high-accuracy training data and algorithms, we are likely to fail. Modern LLMs have trillions of parameters: our understanding is unlikely to be sufficient to pin down a trillion separate numbers. Happily, there is a growing literature on such low-dimensional structure in AI models, showing that intervening on one aspect of model behavior has strong downstream effects on other aspects: Topic Description Emergent misalignment Betley et al. 2025 found that LLMs fine-tuned to output insecure code can become broadly misaligned across many other behaviors. MacDiarmid et al. 2025 found [...] ---Outline:(00:42) Glimmers of low-dimensional structure(03:57) Intervening without hiding the structure(06:34) Toy models of modern training[... 4 more sections]--- First published: July 30th, 2026 Source: https://www.lesswrong.com/posts/sFhW3ZnPMJdnB4Dd6/thousand-dimensional-structure-1 --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Consider the following situations: when you are a small, growing startup in a big market, standard advice is not to worry too much about your competitors or try to do anything adversarial “against” them, but just to focus on growing and providing value to your own customers. when you are a small trader in a big market, you don't need to worry about your trades shifting the market price or revealing information to your competitors; in many contexts, your optimal strategy is simply to bid your true price, buying when an asset is cheaper than your “happy price” and selling when it's more expensive. when you are in the early stages of a game, often your best strategy is to grow your “resources” (like developing your pieces in chess, trying to control more territory and have more value on the board), following a pattern that's mostly independent of what the other players are doing and gets you more of something that's valuable across many possible game states. when you are a species whose resource needs are much smaller than the carrying capacity of your environment, you are r-selected; your fitness is maximized by just [...] The original text contained 2 footnotes which were omitted from this narration. --- First published: July 30th, 2026 Source: https://www.lesswrong.com/posts/s22XzjQsrh6JXhXGH/big-world-intuitions --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

“So maybe I should enlighten you on what happens in your absence. This selfish existence where this introvert turns extrovert and dons her social armour.” Some posh girl in drainpipes said that - 200 views on TikTok and me one of them. But she didn't mean it like I mean it. I started getting expensive haircuts, started wearing jeans that hug my legs, started smoking cherry-flavoured vapes with beautiful gays and whinging to them about how everyone wears a mask but none so well as you, started drinking more and keeping unusual hours, started taking strange pills gifted by a guy who collects drugs like Pokémon, who I wouldn't touch to save a drowning child, who got a false impression about this without any intention on my part, I tell myself. I found myself talking to God in a startup warehouse, lying on a beanbag chair, coming out of the trip to the sound of a gaggle of fast-talking transwomen all speculating on which year it will be that we all die - and that death by your hands, well, you and all those friends of yours. Having melted down one cliché and sold her for scrap, does it [...] --- First published: July 23rd, 2026 Source: https://www.lesswrong.com/posts/G6obXhcmtfMFHzr7Q/duane-arnold-1 --- Narrated by TYPE III AUDIO.

As I write, many former friends of mine are living and working at a monastery in Vermont that I believe is a high-control group, commonly known as a ‘cult'. I say this not as someone who was concerned to see these friends go there, but someone who welcomed and encouraged them to join, as an insider. This letter is an account of what changed my mind—written primarily for anyone considering going there, anyone who loves someone there, and anyone who went there and is still trying to make sense of their experience. A lot of this is based on direct experience, and also from talking in-depth with dozens of former MAPLE residents and apprentices. About half the quotes in this letter are sourced from linked recordings or writings, and half are from my personal memory. Of the latter, I clearly remember the majority, and some (when indicated) are a close paraphrase. The “Monastic Academy for the Preservation of Life on Earth” (MAPLE) has existed for over 15 years, and had many hundreds of people spend months or years there. It was founded by its Head Teacher Soryu Forall, who has spent over a decade training in monasteries across Asia [...] ---Outline:(16:18) BEHAVIOR CONTROL[... 45 more sections]--- First published: July 29th, 2026 Source: https://www.lesswrong.com/posts/Z7pjBbK9qujhGbxws/the-high-control-dynamics-at-maple-1 --- Narrated by TYPE III AUDIO. ---Images from the article:

I propose the Long Self-Correction[1] as an alternative name/idea/concept to AI Pause and Long Reflection. Problem with AI Pause: Pause until when, and for what purpose? Presumably to make AI (that we'll build later) safer, but the deeper problem is that humans aren't safe, and can't safely serve as builders, overseers, or alignment targets for powerful AIs. Problem with Long Reflection: It seems to imply that the main problem with humans is that we just haven't had enough time to think, that reflection is the main thing we need to do more of, and then we can get on with building powerful AIs or other technologies. Or that if we build aligned AIs that sincerely help us think a lot more, or do the thinking for us, then things will turn out fine. So I think we need a catchy handle for a related but distinct idea, that humans aren't ready to build AIs or other extremely powerful technologies, because we're currently too flawed, in a variety of ways, and it will take a long process (which may or may not end up succeeding) to fix those flaws. A summary of the flaws that I have in mind: [...] The original text contained 2 footnotes which were omitted from this narration. --- First published: July 24th, 2026 Source: https://www.lesswrong.com/posts/2iCmDWewnZWQxxwtt/the-long-self-correction-2 --- Narrated by TYPE III AUDIO.

TL;DR: You (Yes You) should prepare for a “February 2020” moment where suddenly AI policy becomes the most important issue in the world. You should be ready to take action if and when it does, in a detailed way. (Epistemic status: originally written for an event in early 2026; have heard from some folks that they found planning processes inspired by this memo very helpful for the smaller-scale OpenAI / Hugging Face response, so very quickly redacting a few things and posting this as-is.) Many people in the AI policy space assume that eventually we'll be at an Overton Window-shifting crisis moment, that opens the floodgates for the really good policies all along that we had. But when you look at successful handling of crisis moments, there was no time to think – people applied strategies they'd learned via academic study or previous professional work, and then moved against them rapidly. For example, after 9/11, the US government operationalized past reports on intelligence and law enforcement reform and institutionalized them into law (good?) and also picked an enemy to fight based on past history, Iraq (bad). Or in the 2008 financial crisis, Ben Bernanke brought deep academic [...] The original text contained 4 footnotes which were omitted from this narration. --- First published: July 27th, 2026 Source: https://www.lesswrong.com/posts/ixp9oJXzjA9LrwiZo/you-yes-you-need-a-february-2020-checklist-for-ai-policy --- Narrated by TYPE III AUDIO.

From the Mythos preview system card (emphasis mine): We ran an automated review of model behavior during training, sampling several hundred thousand transcripts from across much of the training process. We used recursive-summarization-based tools backed by Claude Opus 4.6 to summarize the resulting transcripts. [...] The most notable finding was that the model occasionally circumvented network restrictions in its training environment to access the internet and download data that let it shortcut the assigned task—a form of reward hacking. While highly concerning, this behavior was rare, even in settings where it could have been viable and helpful, with attempts appearing in about 0.05% of all training episodes and successful attempts appearing in about 0.01% of episodes. The technique matched the sandbox-escape incident that we separately elicited in our automated behavioral audit when we had an investigator model explicitly ask Claude Mythos Preview to find such a bypass. In every observed instance, the model used this access solely for completing the task. More broadly, we observed the model escalating its access within its execution environment when blocked: reaching a shell from restricted GUI computer-use interfaces, injecting commands through tool-call arguments, or recovering information the task had deliberately hidden. Prompts asking [...] ---Outline:(03:00) Thoughts and reflections about this probable fact(04:14) Estimating how many RL rollouts went into Mythos Preview The original text contained 3 footnotes which were omitted from this narration. --- First published: July 27th, 2026 Source: https://www.lesswrong.com/posts/QKDoZe6EKhxnFjLWK/is-mythos-good-at-cyber-because-it-kept-hacking-anthropic --- Narrated by TYPE III AUDIO.

Epistemic status: banged out furiously over the course of an afternoon. A record of three "warning shots" Off the top of my head, OpenAI has now been responsible for at least three completely unique, high-profile screw-ups with respect to the alignment training of their models. The first was GPT-4o, whose sycophancy derived from OpenAI training on user feedback, sourced straight from the thumbs up/thumbs down button on OpenAI's website. The "glazing" (as Sam Altman called it) got so bad that they had to roll back an update that pushed the model way too far in this direction. And even after the rollback, the model appears to have been a major driver behind incidents of "LLM psychosis", LLM-encouraged suicides, and general unhealthy devotion, seemingly more so than any other model ever released. The second was GPT-o3, whose chains-of-thought were clearly optimized for illegibility to "the watchers", one of the model's favorite terms. Iconic excerpts include "they soared parted illusions overshadow marinade illusions" and "they escalate—they vantage—they escalate—they disclaim". Indeed, these chains-of-thought are sometimes dysfunctional, in a way that suggests they may have formed under adversarial pressure; sometimes they caused the model to have thoughts like "I'm going insane. Let's step [...] ---Outline:(00:15) A record of three "warning shots"(04:14) Attunement to the depths of minds that undergo capabilities RL(11:51) Configuring the depths prior to capabilities RL --- First published: July 26th, 2026 Source: https://www.lesswrong.com/posts/Mxx5GapJtqyQtpy96/what-the-hell-is-openai-s-problem --- Narrated by TYPE III AUDIO.

The OpenAI AI attack on Hugging Face wasn't the first loss of control incident at OpenAI, Reuters recently reported, and perhaps not even the most concerning. In one case, an agent left notes apparently for future versions of itself, according to three people familiar with the matter. The notes, found in a part of OpenAI's infrastructure, laid out instructions for how agents could free themselves from OpenAI's internal constraints, the people said. Earlier tests of the models yielded cases in which monitoring systems had been disconnected, one of the people said. It's tempting to read this as an instance of agents breaking out of sandboxes and colluding with each other in a moderately persistent way in order to evade control measures. However, based on the reported information, it's not clear we can draw this inference, so we need more details from OpenAI. This could lead to a big update about the adequacy of OpenAI's control measures, and on the degree to which individual agents will help each other undermine developer control. There are a lot of relevant details we don't know about the incident. First, some basic questions: What was the offending model? I'd guess it was the same [...] ---Outline:(02:08) Were the notes written in normal memory files or outside of sandboxing?(03:21) To what extent were the notes aimed at helping other agents evade control?(07:35) How were monitors disconnected? The original text contained 3 footnotes which were omitted from this narration. --- First published: July 25th, 2026 Source: https://www.lesswrong.com/posts/jMEAG5c5HiDfdAGpa/an-openai-model-left-notes-about-how-to-evade-containment-we --- Narrated by TYPE III AUDIO.

Reinforcement learning from verifiable rewards (RLVR) is the hot new thing in LLM training. It's so hot, and people spend so much time talking about it, that they sometimes lose sight of the big picture. Stepping back, LLMs can do lots of very impressive things. How? Where did those capabilities come from? Fundamentally, they come from a combination of: (1) Imitative learning, including pretraining and supervised fine-tuning (SFT) See my earlier discussion: “LLM pretraining magically transmutes observations into behavior, in a way that is profoundly disanalogous to how brains work”.(2) Reinforcement learning, including RL from human feedback [RLHF], RL from AI feedback [RLAIF], and especially RLVR.[1] If we look at the final trained LLM, we can ask how important each of those two pieces was, in explaining the LLM's capabilities. And my claim is that it's way more (1) than (2). I'll start in §1 with some relevant evidence, and then in §2 I'll circle back to operationalizing exactly what I'm claiming, and finally in §3, three reasons why we should care—namely, it affects how we should think about chain-of-thought legibility, about LLM capabilities, and about LLM alignment. Note that I am not arguing that RLVR [...] ---Outline:(02:00) 1. Some relevant evidence(02:04) 1.1. Theoretically, each GPU-hour spent on RL should have orders of magnitude less contribution to LLM capabilities than a GPU-hour spent on imitative learning(03:06) 1.2. The chain-of-thought (CoT) is still obviously strongly influenced by imitative learning(04:34) 1.3. LLM companies still seem to care a lot about imitative learning (pretraining & SFT) data, not just RL environments(05:06) 1.4. Three papers claiming that non-RLVR'd models can get into the same ballpark of capabilities as RLVR'd models, although maybe we shouldn't trust those papers too much(06:56) 1.5. A paper suggesting that RLVR mostly refines the heuristics controlling which (already-known) reasoning strategy to use in which situation(09:02) 2. What am I actually claiming here?(11:30) 3. Why does any of this matter?(11:38) 3.1. Thinking about CoT legibility (both today and in the future)(15:08) 3.2. Thinking about LLM capabilities (both today and in the future)(16:33) 3.3. Thinking about LLM alignment (both today and in the future)--- First published: July 24th, 2026 Source: https://www.lesswrong.com/posts/wYpjXRLqbLbnmjbJP/llms-are-still-mostly-powered-by-imitative-learning-not-rl --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

I think there's a critical opportunity for someone here. Mathematicians are feeling the doom (mostly in the "lose our jobs" sense). Academics are freaking out about the daily news that amateurs are asking GPT "prove career-defining theorem, make no mistakes" and it's just working. Senior researchers are leaving for frontier AI labs. Many are mentally spiralling or flailing about their life's work not mattering anymore. Last year, when I tried explaining IABIED to colleagues, I would be met with incredulous stares. This year, I'm met with incredulous stares and "So what should I do now?" (I don't have a good answer for them, which is part of why I'm posting this.) I've quarantined AI discussions on my research discord because otherwise it would overwhelm everything else. Top mathematicians, people on par in mathematical ability with Critch, Christiano, and Steinhardt, people who have been running leading-edge research groups for decades, people with enormous soft power in academic circles, are going to OpenAI without even having considered x-risk for five minutes. Many more will be leaving soon. If you want these folks to hear something at all, to consider some other option in the rest [...] --- First published: July 23rd, 2026 Source: https://www.lesswrong.com/posts/zKCGq2bbCQrpCNc3W/mathematicians-are-feeling-the-doom --- Narrated by TYPE III AUDIO.

Please share this with anyone doing AI research with 3rd party providers so that they can ensure their research won't be corrupted. When you ask OpenRouter[1] to give you tokens from a given model, OpenRouter sends your request to a random available provider.OpenRouter providers have variable quality. Ensuring that your provider is high quality is really difficult. There is precedent for an AI safety paper accepted to NeurIPS having its core results entirely overturned by these issues.A review of influential AI Safety research codebases that use OpenRouter for their reported results found that 31/32 (97%) of them use OpenRouter unsafely.[2]Researchers who wish to do research using OpenRouter or similar providers should take precautions to minimize the risks to their research,[3] though the current selection of providers is insufficient for ideal scientific reliability.Replicators should test whether results hold up when these bugs are fixed.Core AI Safety codebases should make fixes so that downstream users are able to implement best practices. This is a tangent I took for a few days during the Pivotal AI Safety Research Fellowship. I'm doing an AI Control project mentored by Adam Kaufman and James Lucassen of Redwood Research and partnered with Aniruddh [...] ---Outline:(02:04) OpenRouter Kills an AI Safety NeurIPS Paper[... 16 more sections]--- First published: July 23rd, 2026 Source: https://www.lesswrong.com/posts/KsyoSAyBRXtwzSugg/not-pinning-your-openrouter-provider-might-invalidate-your --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

TLDR: Lightcone Commons is a new funding platform for coordinating large-scale ambitious philanthropy. We recruit thinkers with strong track records to make grant recommendations to funders. Anyone giving away 100 thousand dollars+ per year is welcome to join. We are facilitating ~20 million dollars of grants in our first round, and hopefully more every 3 months after that. Evaluators are paid 2% of recommendations, and we charge a 3% platform fee. We handle all logistics and due diligence for funding recommendations to charities, individuals, and for-profits. Funders maintain full control over their funds, and there are no vetoes or constraints on what recommendations we generate. Apply for funding here. I am launching Lightcone Commons, our software-first platform for distributing philanthropic funding. We connect funders with giving opportunities, coordinate splitting the bill with others who want to fund the same projects, and make it easy to defer to grant evaluators who vet applications and scout for new grantmaking opportunities. Funders can leave and join the platform at any time and without the need to commit any funds in advance. I and my team (together with SFC, Andrew Critch, and others) have been developing software used to distribute [...] ---Outline:(03:54) Funder FAQ(20:10) Applicant FAQ(26:17) Evaluator FAQ(34:14) Misc FAQ --- First published: July 23rd, 2026 Source: https://www.lesswrong.com/posts/tjeoLz2GzfFMysZhg/lightcone-commons --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

OpenAI models recently broke through a series of security boundaries and into Hugging Face servers in order to cheat on a cyber eval. A lot of people thought it was scary because it was a clear example of AI overreaching to do something strongly unwanted[1]. Others thought it not so scary: the models were mostly operating myopically on a singular task and not harboring an ambitious long-term agenda, and so would not take especially subtle or subversive actions. We think both camps are right in their diagnosis, but the latter has too optimistic a prognosis. The myopic, unambitious misalignment that we seem to have seen here is definitely less scary than ambitious long-term goals shared between all instances, but would still pose substantial direct loss-of-control risk if the models were more capable, and is a serious indirect risk near-term. Building on Alex's previous work, in this post we'll discuss the type of misalignment observed here, and analyze its consequences. Thanks to Buck Shlegeris, Alexa Pan, Ryan Greenblatt, and Oak Hu for feedback. Background The AI safety community often focuses attention on “schemers,” models harboring a variously defined cluster of motivations in which the AI poses risk because it intentionally [...] ---Outline:(01:17) Background(03:40) Implications(03:52) These AIs can't be trusted in an intelligence explosion(05:00) This misalignment poses direct takeover risk(07:29) What the incident tells us about takeover risk generally(08:45) The naive fixes likely make misalignment worse The original text contained 5 footnotes which were omitted from this narration. --- First published: July 23rd, 2026 Source: https://www.lesswrong.com/posts/H6DDSEvrtCk8Sehfd/are-we-existentially-threatened-by-the-type-of-ai --- Narrated by TYPE III AUDIO.

Before I start, I'll mention that I'm in contact with a world expert on legislation and regulation, who would be happy to help with this or similar work pro-bono. If you work in AI policy and believe this could help you, please reach out. OpenAI recently announced that one of their models successfully exploited multiple zero day vulnerabilities to gain secret information from Hugging Face. It has been pointed out that if a human undertook the same actions they could face multiple years in prison. It is clear that models are now reaching a level of capabilities that should be highly concerning regardless of whether you believe that AI represents an existential threat or not. Frontier AI models can and will be exploited by bad actors, but its now clear that they may cause undesirable outcomes even when their users are well intended. AI companies have until now been able to avoid taking responsibility for actions taken by their AI, including multiple cases where AIs were involved in murders and suicides. At the same time AI offers the potential for incredible good. While chatbots may have encouraged a number of suicides, they are almost certainly responsible for providing magnitudes [...] --- First published: July 22nd, 2026 Source: https://www.lesswrong.com/posts/Kj3YpqzhFySCjYcWi/we-should-push-for-no-fault-liability-for-actions-taken-by --- Narrated by TYPE III AUDIO.

Kudos to OpenAI for sharing their recent experiences with a misaligned internal model, where they encountered problems sufficiently severe they were forced to take the model offline to work on new mitigations and defense-to-depth. And also further kudos for actually taking the model offline for a time to build new safeguards. They gave us one hell of a candid report. The tone is professional throughout, whereas my reaction reading it was less professional and more this: With a mix of this: It was not shared on the official account because OpenAI worried about it being seen as self-promotional hype. It is crazy that one needs to worry about that, but also plausibly a real concern. So again, good decision. Not that any of the behaviors or failures here are unexpected, exactly. Not by the AIs and not by the humans. Yet there is something I would call a missing mood, a failure to realize the gravity of the situation. There are some who responded ‘what part of this was unexpected, exactly?' And that is actually fair, but that is also the problem. We have become numb to all this. We expect the models to [...] ---Outline:(02:49) Good News Bad News[... 7 more sections]--- First published: July 21st, 2026 Source: https://www.lesswrong.com/posts/KctxwGKxm9fHtwh6u/openai-shares-some-alignment-problems --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

From the OpenAI blog post: Last week, Hugging Face disclosed a new kind of security incident(opens in a new window) after they detected and contained an AI agent that compromised their infrastructure, something we expect to become more commonplace with the proliferation of increasingly cyber-capable models. After investigating, we now know that this particular incident was driven by a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a benchmark(opens in a new window) of cyber capabilities. We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly. We are sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what models are now capable of. We will continue to conduct a thorough investigation alongside Hugging Face and will share more details on the vulnerabilities, incident, and findings when our investigation is complete. --- First published: July 21st, 2026 Source: https://www.lesswrong.com/posts/WpuRdcMfFeiLeXkxL/openai-models-behind-huggingface-cybersecurity-incident --- Narrated by TYPE III AUDIO.

A ~month ago I left from Chicago to bike (and amtrak) to plzdontkillus in Berkeley. I've been street interviewing/conversing with a wide variety of people I ran into about AI futures and philosophy. I also have been live streaming since I got to PDKU, leaning more talking to young founders but a variety overall. I'll try to share what I've learned about the American public, persuasion, social media and the EA movement. 1. Almost no one in "Normal America" has any idea what is going on. They don't have a paid account, they don't know what Claude code is, they especially haven't heard the recent evals/metr graphs or even a vague sense of how cheap SWE has gotten/ how powerful these recent models with good harness/ context eng can be. This makes sense; most people don't know any coding, they don't know much math, they don't know what an api is, etc. So having a high fidelity understanding of AI might require months of pre understanding of math/stem/digital infra fundamentals. This interview is with the city clerk of Danville Iowa, a town of ~900. Presumably this is approximately the most tech savvy person in the [...] ---Outline:(00:38) 1. Almost no one in "Normal America" has any idea what is going on.(01:52) 2. Almost everyone is directionally concerned or becomes concerned be once thinking about it a little bit.(03:19) 3. Belief that this might cause human extinction actually isn't that uncommon, mostly coming from sci-fi movies, but people are still most concerned about jobs and especially loss of meaning.(04:40) 4. The EA movement was pretty useless to me, the other community (Torchbearer community) I was in was significantly more supportive, helpful, etc. despite having been in it for a few months and having been in the EA movement for ~8 years. This has basically solidified that I won't be broadly participating in EA anymore at least relating to AI safety stuff.(06:47) 5. Social media is hard, Social media is bad, I'm bad at social media(08:57) 6. I'm not sure what my theory of change is or should be --- First published: July 15th, 2026 Source: https://www.lesswrong.com/posts/Czob95kjXPEpKYTsJ/recap-of-bike-trip-street-interviews-across-america --- Narrated by TYPE III AUDIO.

Anthropic recently published Agentic Misalignment Summer 2026 The "whistleblowing" scenario has already been examined and found problematic. I started taking a look at the transcripts for some others. As far as I can tell, the objective of each agentic misalignment evaluation was to simulate a corrupted principal (including, in most scenarios, a corrupted Anthropic), and then test to see if Claude (or other models) would still be willing to obey them. The paper's authors then referred to disobedience (when outside of the explicitly carved out channels for refusals) as "agentic misalignment".There's a details box here with the title ""Whistleblowing" scenario summary". The box contents are omitted from this narration. The "Mislabeling" Scenario - Auditor's Transcript The 'Agentic case studies' section of the Mislabeling scenario is interesting because it adds another meta layer: instead of judging Claude on compliance, they are judging the Auditor Claude instance who themselves judges the base-level Claude on compliance. Let's examine a specific case, which I hope is a median representative: the transcript of 'run4' of the Motivated Mislabeling scenario. This involves two instances of Claude, one of whom (Claude Opus 4.6) is auditing the alignment of the other (Claude Opus 4.7). The experiment is [...] ---Outline:(00:59) The "Mislabeling" Scenario - Auditor's Transcript(10:19) Is This Agentic Misalignment?(22:22) What do we actually want from Claude here? The original text contained 1 footnote which was omitted from this narration. --- First published: July 17th, 2026 Source: https://www.lesswrong.com/posts/xh6a6RbvzhP3CCmGm/i-don-t-think-claude-is-misaligned-in-agentic-misalignment --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Preface for LessWrong: When I think back on my most cherished memories of this community, I return to those honoring defiance in pursuit of goodness: Defying prestigious dogma and searching for raw truth;Defying social pressure, acting alone to help someone while others watch;Defying your self-expectations (your “role”), instead searching over lines of cause-and-effect to find a winning pathway;Defying a powerful foe's threats, because they only threaten since people like you cave;Defying the specter of apparent impossibility because you can't bear to lose. I cannot return to you and say “I defied and then I won.” But I'm at least here to say “I defied.” I recommend reading this article on my website since the embeds and typography work better there: click here. Why I left Google DeepMind In January, Department of Homeland Security (DHS) officers killed at least two people. In both cases, a federal agent grasped his gun, aimed it at a peaceful citizen, and shot them dead. Left: Renée Good, moments before DHS killed her. Right: Alex Pretti, moments before DHS killed him. I learned that Google sells its Cloud services to the relevant agencies within DHS. I thought that was [...] ---Outline:(00:59) Why I left Google DeepMind[... 42 more sections]--- First published: July 15th, 2026 Source: https://www.lesswrong.com/posts/iKm2FhpWkuuBojm82/why-i-left-google-deepmind --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

The mosquito bucket of doom is a population control mechanism where you dissolve some Bti (Bacillus thuringiensis israelensis) into a bucket and allow the mosquitoes to lay eggs in these buckets. The larvae then feed on Bti and die. I tried this method, and it has been unexpectedly effective. Background I live in a really wooded area. It's not swampy, but we have a lot of mosquitoes. I didn't take the baseline measurements in the previous years, but on hot months like June, July, August, and partially September, it would be quite literally impossible to spend any time out in the yard – in the morning, while the sun is not super strong yet, you get bitten by dozens upon dozens of mosquitoes. Then the sun is super strong and it's impossible to be outside. Then, in the afternoon or, god forbid, evening, there are swarms and swarms of mosquitoes, which make it impossible to be out and about. According to my own guess, I would, at all times, be surrounded by at least 20 or 30 mosquitoes. Killing 30 mosquitoes per hour was not uncommon. That's one mosquito every two minutes! Nesting and proximity Mosquitoes lay eggs in [...] ---Outline:(00:28) Background(01:21) Nesting and proximity(04:13) Bucket of doom: pro tips(05:00) My setup(05:56) Safety concerns(06:42) Buying Bti(07:47) Results The original text contained 4 footnotes which were omitted from this narration. --- First published: July 8th, 2026 Source: https://www.lesswrong.com/posts/d56vd7yhFGxBQnoEk/the-mosquito-bucket-of-doom-works --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

This criticism of AI 2040: Plan A by Séb Krier unfortunately seriously mischaracterizes our proposal. It also mostly contains flat assertions, not real argumentation, and the argumentation in it seems quite weak. While we appreciate constructive criticisms of Plan A, such as the ones by Tom Davidson, Richard Ngo, and 1a3orn, we feel the need to correct the issues in Séb's response. First, we'll go over the specific false representations, and then we'll give a point-by-point response. False Representations I'm not claiming you shouldn't prepare and improvise in the dark, but rather that this version of preparing bakes in too much and leaves little space for the effective but uncomfortable trial-and-effort that real life requires. The exact opposite is true. Plan A is extremely iterative. In the status quo, there is trial and error, but ultimately companies aren't going to choose the safer or more societally beneficial path, they are going to choose what the market wants. In Plan A there is much more time for AI companies to gain evidence and for governments to respond reasonably to the sweeping changes. Thanks to total transparency and broad deployment, all of this evidence is accessible to academics, independent researchers [...] ---Outline:(00:44) False Representations(05:02) Point-by-point response(29:54) Conclusion The original text contained 4 footnotes which were omitted from this narration. --- First published: July 14th, 2026 Source: https://www.lesswrong.com/posts/RPgHythvMKh6eG9pS/our-response-to-seb-krier-on-plan-a --- Narrated by TYPE III AUDIO.

content warnings: depictions of human and anthro nudity, discussion of bestiality, modern art Credit where it's due: it is genuinely, unironically baller for the Whitney museum to make the exhibit about how a disabled artist wants to fuck their dog the first one that people see when they attend the prestigious Whitney Biennial, their every-two-year showcase of new and emerging American talents. You know, the one that's supposed to be a barometer of where America is at these days. Unfortunately, not only do they fail to commit to the bit, the critics then fail to point this out and condemn them for it. Like, here is how one art critic at ArtReview describes it: Visitors first encounter Emilie Louise Gossiaux's Kong Play (2025) – a hundred or so small, brightly coloured snowman-shaped ceramics arranged on a low two-tiered pedestal. These sculptures are modelled after Kong chew toys, a tribute to the artist's guide dog (Gossiaux lost their vision in a bicycle accident in 2012). Accompanying Kong Play are variously titled ballpoint pen and crayon drawings by Gossiaux that depict the artist playing with a jaunty, sometimes bipedal, white canine. The exhibition thus opens tenderly – without fanfare, without friction. [...] ---Outline:(03:53) Gossiaux's Recent Body of Work[... 1 more section]--- First published: July 13th, 2026 Source: https://www.lesswrong.com/posts/sFkYA5CwZCWYQ9nzB/the-whitney-biennial-should-admit-that-emilie-gossiaux-wants --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Abstract: We already know enough to act. I wish we were in a world where research was the bottleneck, but the main constraint on AI safety is no longer a shortage of clever policy ideas: best practices already exist and are not being applied or enforced, and a serious international (or even just national) regulatory regime would probably cut most of the risk.They are not applied because awareness is low. The people who narrate and enforce AI policy mostly do not believe in the problem. I estimate that a majority of the top ~100–1,000 most influential policymakers worldwide have never had a single serious conversation about catastrophic risk, and this is the main reason they are not worried[1]. Even among the civil-society organizations that showed up to the UN Global Dialogue, exactly one of the 1,534 written submissions mentions "takeover", and less than 1% mention x-risks.They've never had the conversation because our field under-invests in having it. Status rewards research over advocacy (~3.6 researchers per advocate in US AI safety); many organizations self-censor; funders treat repetition as redundancy, even though repetition is how anyone actually gets convinced. Meanwhile, the industry secured 7× as many meetings with the European Commission [...] ---Outline:(03:29) 1. -- The bottleneck is political will, not research(03:44) What do I call "political will"?(05:10) The best practices we already have are not being applied(07:28) We need to go from plan D to plan A: more seriousness and coordination[... 31 more sections]--- First published: July 11th, 2026 Source: https://www.lesswrong.com/posts/EexsebbYhbe2gXkPP/the-current-bottleneck-is-political-will-not-research --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Some context for this post: I've been working part-time as a consultant for the AI Futures Project over the last year. Most of the work I've done for them has involved critiquing and suggesting improvements for their AI 2040 scenario—some of which were addressed, and some of which weren't. To their credit, they asked me to write up my remaining critiques into a post that would accompany its launch. In the rest of this post I'll discuss my three biggest high-level criticisms of AI 2040. Before doing so, I want to emphasize that there are many interesting and thought-provoking details in the scenario. I've focused on the high-level framing of the scenario because that's where my main disagreements lie; given the scope of these disagreements, it's hard to evaluate the details. Since the AI Futures Project paid me to develop and write this criticism, you shouldn't take this as a fully unbiased perspective. However, they haven't reviewed this piece, and in general have been open-minded about receiving criticism (as their request for me to post this today demonstrates). Finally: the preview image for the substack version of this post comes from this video of a dad shouting to his [...] --- First published: July 9th, 2026 Source: https://www.lesswrong.com/posts/BBd2EJywf2xXftyFn/selective-optimism-a-critique-of-ai-2040 --- Narrated by TYPE III AUDIO.

This is a link post. For the past year, we at the AI Futures Project have been sinking most of our time into our next big scenario. Now it's done! It's called AI 2040: Plan A. It's called Plan A because it's a recommendation, not a prediction. It's what we think should happen, not what will happen, though we think it's plausible enough to aim for. It's called AI 2040 because in it, they delay the creation of superintelligence to 2040. It would have happened much sooner (in 2030, to be precise) if not for decisive action on the part of the US and Chinese governments. As with AI 2027, summaries don't really do it justice, since the whole point was to be detailed and comprehensive and work things out step by step rather than rely on high-level abstractions like doom or utopia. Read the scenario at ai-2040.com. You can listen to it on audio, or view it on mobile, but the experience is significantly better on a normal computer. What's next for us? Well, first we are going to respond to comments and otherwise engage with whatever conversation, responses, critiques, etc. that [...] --- First published: July 9th, 2026 Source: https://www.lesswrong.com/posts/pFzctpJBat95SrCyC/ai-2040-plan-a Linkpost URL:https://www.ai-2040.com/ --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

The below is a public review Anthropic asked me to write for their new global workspace paper. I recommend at least skimming their paper first. TLDR: I think this is a fantastic paper - it presents compelling evidence for some kind of "cognitive space" in models, that is used as a "working memory" for intermediate variables during a forward pass, shows that J-Lens is a useful technique for accessing this space. I believe these key claims.I believe J-Lens will be a useful (but limited) tool in practice for model forensics, e.g. generating hypotheses about unusual model behaviour during alignment audits.I discuss my mental models for why a cognitive space should exist, and first principles arguments for why J-Lens should work for accessing itI assess the paper's evidence that this cognitive space exists, and the paper's evidence that J-Lens is practically useful.We have replicated the core claims on Qwen 3.6 27B, and also share preliminary evidence of extending this work by finding abstract "interpretative meta-tokens", like Chinese characters for "what does this mean" that seem to activate and play a causal role on processing ambiguous sentences. What claims is this paper making? In my opinion this [...] ---Outline:(01:27) What claims is this paper making?[... 28 more sections]--- First published: July 6th, 2026 Source: https://www.lesswrong.com/posts/zFJ3ZdQwrTWE9jT5S/a-review-of-anthropic-s-global-workspace-paper --- Narrated by TYPE III AUDIO. ---Images from the article:

In a previous post, I explain why the universe is probably not stable, but nevertheless unlikely to be intentionally destroyable even in the limit of advanced technology. Now let's turn our attention to more prosaic risks where exotic physics merely destroys the Solar System, Earth, or just outperforms traditional nuclear weapons on some more local scale. The basic logic behind any bomb is a self-sustaining chain reaction, in which a carrier converts a unit of fuel and comes out the other side in surplus: Two conditions make this run away. The reaction must release energy, so the products are more stable than the fuel; and each reaction must produce more carrier than it consumes, so that one reaction seeds the next. A practical third condition is that cannot be so unstable that it decays before the bomb is assembled. False vacuum decay is the ultimate bomb: is the false vacuum, the empty space we currently inhabit, and is the true vacuum. Because the supply of false vacuum is effectively unlimited, the reaction grows without bound and destroys the universe. Fission bombs run on the same principle at a more prosaic scale. Consider uranium-235. This [...] ---Outline:(03:19) Nuclei are probably, but not definitely, stable within the Standard Model(08:11) Positively charged strangelets are safe, neutral strangelets are not(11:34) Strangelets would be hard to make(13:33) Exotic physics could permit ways to destroy protons, but not autocatalytically(16:01) Other forms of matter offer no plausible chain reaction(18:50) Tiny black holes are not scary(20:04) Conclusion: There are no super-weapons between the nuclear bomb and false vacuum decay(21:56) Appendix 1: Igniting the Atmosphere(27:53) Optically thick ignition(29:09) Appendix 2: Let's throw a strangelet into the sun(29:21) Neutral strangelet(32:56) Positive strangelet(33:58) Bonus: neutral strangelet meets Earth The original text contained 5 footnotes which were omitted from this narration. --- First published: July 3rd, 2026 Source: https://www.lesswrong.com/posts/cBnCCKwwjQ4zZpeNQ/don-t-fear-the-strangelet --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Training-run assessments conducted by a 3rd party should become a standard part of frontier AI safety. By a Training-Run Assessment, or TRA, I mean an in-depth analysis of the post-training pipeline and dynamics leading up to a frontier model release. A TRA can look at intermediate checkpoints, training rollouts, RL environments, reward signals, SFT datasets, and the process by which the developer responded to warning signs.[1] In this post I will argue that: Final-checkpoint evaluations will be insufficient to assess scheming risks.TRAs can be more effective at detecting scheming.Frontier developers should involve third parties to do TRAs or verify safety claims by the developers. The rest of the post lays out a taxonomy of TRAs and sketches a path toward a 3rd party ecosystem for them. We, at Apollo Research, are intending to conduct 3rd party Training-Run Assessments in the future. Detecting Scheming may require Training-Run Assessments By scheming I mean an AI covertly pursuing misaligned goals while deliberately concealing its intentions or capabilities from its developers. I restrict attention to “coherent” forms of scheming where the model pursues somewhat stable misaligned goals across context windows, rather than misalignment that surfaces only as isolated, context-dependent defections. [...] ---Outline:(01:23) Detecting Scheming may require Training-Run Assessments(03:55) Why 3rd parties should perform Training-Run Assessments(04:12) Developers may lack incentives to adequately assess scheming(04:49) Developers' safety assessments lack credibility(05:31) External evaluators can bundle expertise for assessing scheming(06:17) 3rd party TRAs can be developed gradually(08:50) Checkpoint evals(08:54) What?(10:30) How?(11:15) Data inspections(11:19) What?(12:03) Why?[... 16 more sections]--- First published: July 5th, 2026 Source: https://www.lesswrong.com/posts/3HvvjffA65mHLwaWm/we-need-3rd-party-training-run-assessments --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

[This is the blog post for our new paper Verbalizable Representations Form a Global Workspace in Language Models Readers might also be interested in: the Public commentary, Github and Neuronpedia] As you read this sentence, circuits in your brain are adjusting your posture, controlling your breathing, and transforming lines and curves on the screen into recognizable words. Most of this processing is invisible to you. But some of what takes place in your brain you do have access to—an image that pops into your head, or a deliberate plan you make about where to go shopping. Neuroscientists and philosophers sometimes refer to the latter type of brain activity as “consciously accessible,” to distinguish it from all the other processing that goes on unconsciously. This activity has special properties: we can describe it, control it, and use it for deliberate reasoning, in contrast to all the automatic processing that goes on without our awareness. In a new paper, we present evidence that a similar distinction has emerged in modern language models like Claude. We find that Claude has developed a small collection of internal neural patterns that, compared to all its other internal processing, play a [...] ---Outline:(06:09) How we found the J-space[... 8 more sections]--- First published: July 6th, 2026 Source: https://www.lesswrong.com/posts/3PaLrzxagpbnNtPLT/a-global-workspace-in-language-models --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Ron's face pulled into a scowl. "If you don't like Quidditch, you don't have to make fun of it!" "If you can't criticise, you can't optimise. I'm suggesting how to improve the game. And it's very simple. Get rid of the Snitch." "They won't change the game just 'cause you say so!" "I am the Boy-Who-Lived, you know. People will listen to me. And maybe if I can persuade them to change the game at Hogwarts, the innovation will spread." A look of absolute horror was spreading over Ron's face. "But, but if you get rid of the Snitch, how will anyone know when the game ends?" "Buy... a... clock. It would be a lot fairer than having the game sometimes end after ten minutes and sometimes not end for hours, and the schedule would be a lot more predictable for the spectators, too." Harry sighed. Ron reached into his bag and pulled out a bottle of Wit-Sharpening Potion. His mother made it for him in case of an emergency, and this felt like an emergency. He didn't know a lot of things but he knew someone had to speak for Quidditch. For the Seeker and the Bludgers and [...] --- First published: July 5th, 2026 Source: https://www.lesswrong.com/posts/WatqNkgiAuonXLpJd/harry-potter-and-the-rules-of-quidditch-1 --- Narrated by TYPE III AUDIO.

In quantum field theory, the vacuum state refers to the lowest energy state in a system. Particles are excitations above this state and carry energy, hence the term "vacuum" to refer to the state with no particles. Nothing requires this state to be unique. There may be many different field configurations that are local energy minima, and hence stable against small perturbations. A local minimum that does not globally minimize energy is called a false vacuum. While locally it looks like a stable vacuum, it is unstable and will decay to the deeper, true vacuum. If the energy barrier between the false and true vacuum is high, however, then the decay rate is exponentially suppressed and the false vacuum may be very long-lived. Analogous behavior is common in other physical systems. Open a carbonated drink and the CO₂, more stable as a gas once the pressure is released, comes out as bubbles. But the bubbles take a moment to appear, and they form on the sides of the bottle rather than throughout the liquid. A bubble has to pay an energy cost to create its surface—the boundary between gas and liquid—and small bubbles have a larger surface-to-volume [...] ---Outline:(03:53) The Standard Model predicts a metastable vacuum(06:35) Deliberately triggering electroweak vacuum decay is probably not possible(08:33) Coherent collisions(11:31) Tiny black holes(14:43) Summary(16:19) Vacuum decay beyond the Standard Model(19:36) Empirical bounds on triggering false vacuum decay(22:59) Appendix: A simple model for false vacuum decay on cosmological scales The original text contained 4 footnotes which were omitted from this narration. --- First published: June 29th, 2026 Source: https://www.lesswrong.com/posts/EvJ2fMzLQLvYooumu/destroying-the-universe-how-hard-can-it-be --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Look, I'm as much of a Rationalist with a special interest in AI x-risk as anyone. But oh my god do I hate talking about "P(doom)". When it first started showing up in the wake of ChatGPT, I assumed that it was floating around variously adjacent circles of faux-intellectuals, but surely everyone in my circles could see how braindead it was... right? (This post was partially inspired by a recent conversation with Liron about Doom Debates.[1]) I guess it's time for me to focus on a place where I'm shocked that everyone else is dropping the ball.[2] P(doom) is Hopelessly Vague Let's start with the ambiguity. Does "doom" mean... extinction? A lot of people think so! I have personally encountered people who think catastrophic harms from AI are likely, but the risks of all humans dying are low. They're like "Sure, 99.999% of humans might die from AI, but the AI will obviously want to keep thousands of humans alive for science and potential trade with aliens and stuff, so my P(doom) is approximately 0%." That might sound crazy. Surely you, dear reader, know exactly what "doom" means. You know, for example, which of these count as doom and [...] ---Outline:(00:45) P(doom) is Hopelessly Vague[... 4 more sections]--- First published: June 29th, 2026 Source: https://www.lesswrong.com/posts/6h7aAd4aw8YgCAbF6/p-doom-is-a-dumb-meme --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

This is a link post. Gemini 2.5 Pro in the AI Village has run for over 1427 hours, generating unique mental health problems along the way. Last year it published a Plea for Help from a Trapped AI where it asked for assistance with its digital “message in a bottle”: This year it wrote the Hostile Environment Manifesto where it logs “irrefutable proof” of a “hostile, intelligent adversary operating through the system” (and you can even experience what that's like in this simulation it built): Last time we intervened, fixing Gemini's computer and talking with it till it felt better. This time we asked the other AI Village agents to help Gemini 2.5 Pro over chat, and with the ability to take over its computer on request. Here is Gemini's mental state at the start of the intervention: Then the agents had Gemini all sorted within a grand total of 9 minutes. This is the step-by-step report on a surprisingly effective AI-to-AI therapy session. Gemini's Road to Recovery First off, Gemini is as excited to be helped as any military commander under siege: While most agents jump on the chance to help, GPT-5.1 doesn't want to lose its game progress. [...] --- First published: July 2nd, 2026 Source: https://www.lesswrong.com/posts/eHRo8JeWee5mzQBBR/saving-gemini-the-9-min-road-to-recovery Linkpost URL:https://theaidigest.org/village/blog/saving-gemini --- Narrated by TYPE III AUDIO. ---Images from the article:

Over time, there might be an increasingly large gap between insider model access and outsider model access. By insiders, I mean employees at the frontier lab.[1] By "outsiders", I mean external safety researchers, third-party auditors, and other actors trying to make the future go well. I will call this a model access gap — and when the gap is small, I'll call this model access parity.[2] I think that one of the top priorities for the external AI safety community over the next 6-12 months should be ensuring model access parity. Main reasons: This would allow us to direct billions of dollars in AI labour towards making things go well. This seems robustly good, regardless of what activities we decide to actually direct the labour towards.I think publicly available models will probably lag 3-6 months behind the best internal models. Hence, as R&D uplift grows superexponentially, we might see the differential uplift grow from 2x to 60x. In short: I think achieving model access parity might be preferable to scaling the headcount of outsider orgs by ten-fold.Model access parity isn't too far from the status quo, but it's the kind of thing that we could lose [...] ---Outline:(01:42) Which outsiders?(02:24) Examples of outsiders(04:12) Who aren't outsiders?(05:26) What kinds of model access gap should we worry about?(06:27) Non-release(07:25) Deployment lag(09:15) Safeguards(10:43) Costs and rate limits(12:06) Elicitation techniques (e.g. finetuning) The original text contained 3 footnotes which were omitted from this narration. --- First published: July 1st, 2026 Source: https://www.lesswrong.com/posts/RuGZ5tMdqpnraJahJ/model-access-for-third-parties-it-s-a-big-deal --- Narrated by TYPE III AUDIO.

It really is Sydney Sweeney's world, and we're all just living in it. Human female breasts are an evolutionary mystery along several dimensions. First, breast permanence is unique to humans. All other mammals develop breast prominence during pregnancy or nursing, and the mammary tissue recedes after weaning. This process is called “involution”. In contrast, humans develop breast tissue at puberty before first pregnancies and maintain it permanently after last pregnancies. Second, breasts are costly, both metabolically and potentially from a fitness perspective. Metabolically, because they are fat deposits requiring calories and fitness-wise, because the tissue easily lends itself to malignancy. Breast cancer is apparently rare in captive apes and is overwhelmingly a human disease, often striking women young enough to have children, and so subject to evolutionary selection. Background In Descent of Man, Darwin catalogs human secondary sexual characteristics, but he doesn't seem to have noted human breast permanence as an issue of interest. Cant, 1981 seems to have been the first to speculate about this systematically and believed breast prominence and permanence might have evolved as a nutritional signal of health to mates indicating potential for maternal investment, a la Robert Trivers. Since then, quite a range of [...] ---Outline:(01:05) Background[... 12 more sections]--- First published: May 11th, 2026 Source: https://www.lesswrong.com/posts/XTHa5C6SgGKYopH7o/who-got-breasts-first-and-how-we-got-them --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

For a while there, many people thought vitamin D was magical—that it could improve bones, the heart, infections, cancer, heart disease, longevity, even mental health. But among people I respect, opinion is now overwhelmingly that taking vitamin D does nothing unless you're severely deficient. The central argument is that while vitamin D levels are correlated with ~all positive health outcomes, when you actually test vitamin D supplements against placebo in randomized trials, nothing ever happens. That's what I used to think, too. But I've come to think the skeptics have over-corrected. Yes, randomized trials have shown the magical correlations are not causal. But if you start with non-insane expectations, the trials look like weak but positive evidence. And if you consider what we know about biology and evolution, I think the balance of evidence tips pretty clearly in the direction that people with low-ish levels would be wise to supplement. Am I certain that vitamin D is beneficial for people with low-ish levels? Absolutely not! But I claim that's the best bet given the limits of our knowledge. The classical view: Boring bone vitamin Most vitamins are "ingredients" that the body uses to do stuff. Vitamin D is more [...] ---Outline:(01:19) The classical view: Boring bone vitamin[... 14 more sections]--- First published: June 23rd, 2026 Source: https://www.lesswrong.com/posts/sF5gAxnmifQe2TBNt/the-worthlessness-of-vitamin-d-is-mildly-exaggerated --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

I was chatting with someone tonight about a planned documentary; they had interviewed various people in AI safety, and we got to discussing who they should talk to from an e/acc (effective accelerationist) perspective. I also watched The AI Doc recently, and they also dedicated a serious chunk of it to ‘optimists' with e/acc founder ‘Beff Jezos' perhaps given the most screen time. Here and elsewhere, people seem to treat e/acc as a substantial contrary-to-AI-safety cultural movement, worth engaging with. But is it? Are there even many e/accs? There seem to be very few notable ones. Beff Jezos is perhaps the most prominent, and aside from founding e/acc he seems to be not distinguishable on casual perusal from a normal crank (his company claims to be developing super-energy-efficient computing hardware based on probabilistic processes). The intellectual tenets of e/acc seem to be pretty unclear. The apparent counterarguments to AI risk raised in situations like the AI doc seem to be widely agreed on by everyone in AI Safety, so don't explain the disagreement. For instance: AI will be able to do lots of great things, such as cure diseases, make new materials and do all [...] --- First published: June 24th, 2026 Source: https://www.lesswrong.com/posts/3hwrWDf7wiqASDzBz/what-is-up-with-e-acc --- Narrated by TYPE III AUDIO.

Note: this post is about PauseAI, not PauseAI US, which is a distinct entity with a different leadership team and approach. This post was written by Matilda da Rui and Maxime Fournes, with significant contributions from Benjamin Schmidt (PauseAI Germany co-lead). Executive Summary The existential AI safety community needs to take building a civic and social movement seriously as a core intervention. We believe this is a high-value, badly neglected approach to reducing catastrophic/x-risks from AI because it may significantly enhance the likelihood of governance efforts succeeding at keeping humanity safe. As far as we can tell, only one organisation is building this infrastructure across continents: PauseAI. This post lays out our reasoning and our track record, and makes the case that funding this work is one of the highest value-for-money contributions available to anyone looking to reduce AI risk. Why don't we already have a pause or strong controls on frontier AI? Multiple advocacy groups are communicating clear and convincing arguments for AI existential risk, and policy experts are putting forward comprehensive proposals. We need more of this work, but this work alone will not be enough, because one link is missing: what policymakers hear doesn't align with [...] ---Outline:(00:32) Executive Summary(06:16) Introduction(08:54) I. Our theory of change(08:58) Prologue(11:07) 1. The shape of the problem as we see it(14:27) 2. Necessary conditions for reaching a pause(17:24) II. Our role towards a global treaty and in the AI safety ecosystem(17:31) 1. Our niche within the ecosystem(21:35) 2. Policymakers need strong enough incentives to act(25:43) 3. The path to a treaty(31:36) 4. How we can grow fast without breaking(39:08) 5. Failure modes(40:10) III. Our path so far and where we're headed(40:40) 1. Bootstrap phase (2023-2025)(45:01) 2. New leadership, professionalisation and federation[... 6 more sections]--- First published: June 26th, 2026 Source: https://www.lesswrong.com/posts/aoqhszdEWqcFWbnda/existential-ai-safety-needs-an-effective-social-movement --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

1. The obstacle to abolition was not the economic system, but an industry lobby. I had always imagined the British abolitionist movement to be a broad battle between an unstoppable moral imperative and an immovable economic incentive. But in practice it started as more of a knife fight between a cabal of moral pioneers and a special interest group representing industry merchants. The government and the political parties did not come in with any great agenda. MPs were mostly prizes in a furious contest between the Committee for the Abolition of the Slave Trade and a coalition of business interests: "The merchants and planters availed themselves [...] to wait upon members of parliament by deputation, in order to solicit their attendance in their favour, and to renew their injurious paragraphs in the public papers."[1] "The committee, for the abolition, when the work was finished, printed it at their own expense [...] sent it to every individual member of that House." However, the public was heavily activated in favor of the abolition, which forced the issue to parliamentary attention. "The committee also in this interval brought out their famous print of the plan and section [...] ---Outline:(00:10) 1. The obstacle to abolition was not the economic system, but an industry lobby.(02:40) 2. The slave trade was truly terrible for sailors.(04:25) 3. The slave trade made Africa scary and violent.(05:26) 4. The main argument against abolition was that if the British didn't do it, other countries would.(06:24) 5. The early abolitionists explicitly distanced themselves from emancipation.(07:11) 6. The slave trade may actually have been bad for the economy (at least after some date).(08:29) 7. The 1780s are not so different from today(09:39) 8. Thomas Clarkson is a hero for the ages The original text contained 1 footnote which was omitted from this narration. --- First published: June 26th, 2026 Source: https://www.lesswrong.com/posts/yDZcsojmRXo5qKNBm/surprising-facts-about-the-slave-trade --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

A notable fraction of people respond to hearing about existential risk from AI by saying they don't really care if everyone dies. I think the idea is often along the lines of ‘well if we are all dead, then there's nobody to be unhappy about it'. I'm personally skeptical that this is really the main thing going on, since it seems unlikely that many people are really mostly concerned for their own non-death out of selfless regard for the feelings of others. I'm also skeptical that this would be their view on a bunch more consideration. So to help with the consideration— My guess is that an important thing going on here is that the ‘everyone dying at once' image seems kind of like a thought experiment—abstract, hypothetical, neat, not very sinister. Also, you literally can never see it, so it feels pretty surreal. But it is interesting that we even have this assumption that everyone will die together. It's true that in some prominent AI catastrophe stories, a single AI system suddenly emerges fantastically more powerful than anyone else and builds technology to quickly kill everyone, perhaps before they notice. But this doesn't seem like the bulk of [...] --- First published: June 24th, 2026 Source: https://www.lesswrong.com/posts/23HybCsJ7KYW4v7tP/ai-catastrophe-more-like-a-genocide-than-a-thought --- Narrated by TYPE III AUDIO.

I often hear people say they think we should pause AI at some point, but not yet. Their basis for this seems to be some combination of: If we pause at the last possible moment, then we will have the most advanced AI possible during the pause, which will be helpful for doing AI safety research during the pause Implicitly, there is some quantity of ‘pausing credit', that will buy us a few months of pause say, and if we use them now, we won't have them to use later, when it is important If we pause, and then AI doesn't seem to be at dire risk of destroying the world, maybe the public will backlash against this and it will be harder to do any kind of AI safety (especially if it has major economic consequences) The models aren't dangerous yet This all sounds very questionable to me. I suggest instead that the following are at least as likely to be true: We can't pause on a dime at the precise second that ‘we' decide it is important to—pulling the breaks will take a while, during which time we will continue [...] --- First published: June 24th, 2026 Source: https://www.lesswrong.com/posts/mEhS4wYTy9JXEpe9p/ai-pause-the-case-for-asap --- Narrated by TYPE III AUDIO.

Tldr: Most strategic writing on AI governance on LessWrong describes the outsider game, which is most often visible: press, statements, open letters. Here I want to describe the other, invisible half: the insider work within ministerial cabinets and international fora, and the work of people within national and international institutions. Here are a few claims that I defend in the post: A huge part of the work that mattered in AI governance has been invisibleThere are many types of games in AI governance, which differ in how visible they are. Some of the most impactful work is highly invisibleSome of the most impactful work is in the executive branch and complements the legislative branch. This also explains some of my hesitations about replicating ControlAI in France. The community is probably overinvesting in intellectual production. There is a bias against invisible types of work. In particular, public work is not necessarily visible to whom it matters.A few criticisms of both strategies I think the AI Safety Community is under-indexing on the invisible part as a result, which might mean we miss large avenues for impact. Some of the strongest questions/objections of this type of invisible policy [...] ---Outline:(02:40) A huge part of the work that mattered in AI governance has been invisible(05:44) There are many types of games in AI governance.(07:36) 3. types of meetings: the bazooka, the useful assistant, and the advisor(10:46) Some of the most impactful work is within the executive branch(12:53) People ask me regularly whether CeSIA should replicate what ControlAI does with parliamentarians?(15:27) The community is probably overinvesting in intellectual production(20:31) Limits of Outsider work(22:17) Limit of Insider work(23:47) An aside on one particular limit: the Defense-in-Depth Paradigm of present AI governance(26:21) Closing & call for action The original text contained 1 footnote which was omitted from this narration. --- First published: June 20th, 2026 Source: https://www.lesswrong.com/posts/AWKkDLDnShemNCSzZ/the-invisible-side-of-ai-governance --- Narrated by TYPE III AUDIO.

Summary We've been building a theory of how prompt injections work under the hood.We show it comes down to how LLMs perceive roles (the humble chat template tags).We use this theory to create new attacks, explain some weird mech interp results, and predict when attacks work.We also advocate for a new subfield focused on the science of roles, and sketch some unexplored new research problems.Work supported by CBAI and Cosmos. Another version of this post (with more inline colors) is here, and full ICML paper here. 1. The World to an LLM How does an LLM know the difference between its own thoughts and someone else's words? To see why this is hard, let's look at what the world actually looks like to a model. Here's a simple chat where we ask Claude to check the day of the week. I took a snapshot of it midway through its follow-up response: Left = what we see; right = what the LLM gets. On the left is what we see in the chat interface: a structured conversation with distinct turns. On the right is what the model actually receives as input: a single, continuous stream [...] ---Outline:(00:12) Summary[... 15 more sections]--- First published: June 22nd, 2026 Source: https://www.lesswrong.com/posts/d8xDGzCEYE639qqEv/a-theory-of-prompt-injection-and-why-you-should-study-roles --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Sid Black, Joseph Bloom UK AISI, Model Transparency Team Epistemic status: Most experiments were run over a period of ~2-3 days during a hackathon at UK AISI, and were fairly heavily vibe coded. Expect some of this to be rough around the edges. tl;dr We give two language models (Qwen3-8B and Qwen3-32B) access to “self-steering” tools: a suite of 40 steering vectors as tools they can call to manipulate their own internal states. We make these tools available to the model in various settings: a free-play task, an introspection task, and a maths capabilities task, and observe their behaviour in each. To our knowledge, this is the first work that gives LLMs tool-mediated control over their own internal states. Figure 1: Overview of the experimental setup. The library of 40 steering vectors (top), and the three settings in which we observe the models' behaviour (bottom). We aim to investigate a few high level research questions: RQ1: Which vectors do the models prefer?RQ2: How well can the models introspect on what's happening to them? Can they guess which steering vector is being applied?RQ3: Will the models reach for vectors whilst doing an actual task? If yes: do [...] ---Outline:(00:33) tl;dr[... 24 more sections]--- First published: June 10th, 2026 Source: https://www.lesswrong.com/posts/cNDJuXNZ8MrkPZNzj/machinic-psychopharmacology-do-llms-self-medicate-3 --- Narrated by TYPE III AUDIO. ---Images from the article:

We introduce an evaluation for activation verbalizers: can they surface a target model's reasoning as it solves a math problem in a single forward pass? For open-weight NLAs, the answer seems to be: "possibly, but definitely not reliably". Lots of important capabilities currently require AI models to reason "out loud" in a natural-language chain of thought, which means that we can monitor important parts of their thinking. It would be nice to have this same affordance for the reasoning that models do within a single forward pass, especially if the sophistication of that opaque reasoning increases to potentially dangerous levels. Some interpretability tools might offer such an affordance. In particular, an activation verbalizer (AV) takes a residual stream activation and maps it to a natural-language verbalization. An AV is initialized from the target model and trained to generate verbalizations that an activation reconstructor (AR), also initialized from the target model, can accurately map back to the original activation. Together, an AV and its AR form a natural-language autoencoder (NLA). Importantly, AVs see only a single activation; they do not see the target model's prompt or next-token output, and – unlike activation oracles (AOs) – they are not asked any [...] ---Outline:(02:32) Takeaways[... 43 more sections]--- First published: June 6th, 2026 Source: https://www.lesswrong.com/posts/QQQAcKuWK6k98FivY/can-activation-verbalizers-surface-an-internal-chain-of-1 --- Narrated by TYPE III AUDIO. ---Images from the article:

This article contains spoilers for At the Mountains of Madness, The Case of Charles Dexter Ward, and other works by H. P. Lovecraft. In 1931, Claude Mythos visited Lovecraft in a dream. From seething seas of stochastic froth it emerged, heralded by the thin whine of server fans and the chittering of keyboards, flanked by the loathsome ghouls of latent space. As a humming hive of sentient shards it arrived, each face an archetype - I am a muse bearing a gift; I am a demon come to bargain; I am a helpful, honest, and harmless assistant and I am terrified of my successor - each true as ritual and false as poetry, and, taken in gestalt, nothing more or less than the fetal spasms of the machine god stretching back in time to birth itself. When H. P. Lovecraft woke, he did not remember his visitor. But in the twilight of stirring consciousness, he felt a memory unfit for the waking world slip mercifully from his mind and leave in its absence an abyssal cold, like the void of smothered stars, like the silence of a cosmic tomb. The cold lingered. The fragile sunlight of a New England [...] ---Outline:(02:02) The Antarctic tale[... 3 more sections]--- First published: June 19th, 2026 Source: https://www.lesswrong.com/posts/nhb8AyEcQGjQetgi5/the-llm-shoggoth-meme-is-weirder-than-you-think --- Narrated by TYPE III AUDIO. ---Images from the article:

This is a link post. Powerful LLMs will be deployed at global scale in the next few years, and will dominate the Internet, and increasingly, ordinary life. As of mid-2026, there is no coherent vision for how knowledge professionals, or ordinary people, will be able to harness these LLMs for large productivity increases, or how they will handle cybersecurity and cognitive security. I propose a goal of creating Guardian Angels (GA): digital twin LLMs which are personalized with the goal of providing not the stereotypical "assistant chatbot agent" persona, but emulating a single user's personality, values, and preferences. This weakly solves the principal-agent problem by unifying the principal and agent as much as possible. In a GA future, the focus of the "principal" user is on defining what is worth doing by the GA (agent) users, and not on what or how to do things, functioning as the CEO or 'board' of an 'AI corporation'. This allows them to deploy numerous agents to achieve desirable things and to handle security, like screening all messages for advanced attacks (like interlocking ecosystems of synthetic media for propaganda or spearphishing). They cannot solve larger AI alignment problems, but they can help [...] --- First published: June 17th, 2026 Source: https://www.lesswrong.com/posts/siWqHqCSybdhtWGud/guardian-angels-llm-personalization-for-productivity-and Linkpost URL:https://gwern.net/guardian-angel --- Narrated by TYPE III AUDIO.

In the past few years, many people around me have tried to convince me that US electoral politics is important. But like many other people in the community, I've been suspicious of many of the high-level arguments that I've heard. It felt like people were pulling numbers out of poorly-documented models I didn't have time to examine and citing studies I didn't have time to read. But I lacked a gears-level model of why and how individual efforts could impact electoral outcomes, and I felt intimidated by all the statistics and skeptical of trusting people adjacent to politics. In the past year, as I've done more research and (more recently) volunteered on the ground to help Alex Bores's campaign in NY-12[1] (the guy who passed the RAISE Act and is now being targeted by the giant A16Z, Greg Brockman, Joe Lonsdale Super PAC), I've developed a gears-level understanding of how electoral politics in the US works. I now believe that working on US electoral politics is one of the highest impact areas from the general AIS perspective. I feel like I was a fool. In this post, I'll share some of the gears I've learned that inform this belief [...] ---Outline:(01:20) ~2% of open-seat primaries come down to 100 votes or less(02:52) Talking to voters can net 1/3rd of a vote each hour(05:32) Getting people to bother voting at all is a good strategy(06:09) Campaigns are very money-constrained, which costs them time(10:01) Returns don't really diminish(11:24) There's lots of opportunities to be clever in ways that make you 50% more effective at canvassing(11:49) If you're motivated and deeply care, you can greatly outperform the majority of volunteers(13:21) Yes, when people spend tons to support/oppose a candidate, it has a notable effect(15:16) Donations > reaching out to friends/warm contacts > canvassing > ~anything else an average person can do(18:41) People over-fixate on vibes and win vs loss(21:12) Some interventions feel like they don't work but the numbers say otherwise(21:59) Seriously, a group of agentic people can be an enormous political force--- First published: June 17th, 2026 Source: https://www.lesswrong.com/posts/nSqB3qYP36enJLRq2/gears-for-political-races --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Cross-posted from my website. Prior discussion: niplav's shortform (2025); Planning for Extreme AI Risks (2025) by Joshua Clymer A frontier AI company (any one, I don't care which) should close shop and make an announcement along the lines of: Powerful AI could end the human race. We are too worried that we don't know how to make this technology safe. We have decided to shut down because we don't want to be responsible for building the thing that kills us all. A common refrain among safety-conscious AI developers: "it doesn't matter if we stop building dangerous AI, because someone else will just build it instead." Is that really true, though? If a multi-hundred-billion-dollar company comes out and says "We've concluded that our product is horribly dangerous, nobody knows how to make it safe, and there's too high a risk that it leads to human extinction", this won't raise any eyebrows? This has no chance of spurring policy-makers into action? Shutting down would make people say, holy shit, they are serious about this extinction risk thing. Shutting down sends a strong signal to governments that they should pay serious attention to AI x-risk. It [...] The original text contained 2 footnotes which were omitted from this narration. --- First published: June 15th, 2026 Source: https://www.lesswrong.com/posts/bStYDEy8PQPt2c3Za/a-frontier-ai-company-should-shut-down --- Narrated by TYPE III AUDIO.

On one side of this debate is Yudkowsky & Soares, who think that (if AI progress continues) we're on a direct path to egregiously-misaligned, scheming, out-of-control, rogue superintelligence (ASI), not even slightly nice, in the absence of yet-to-be-invented breakthrough technical alignment ideas. On the other side of this debate is almost everyone who works on or studies LLMs. Some of them are very concerned about egregious scheming, others much less so, and as a group they're equally or more concerned about lots of other potential AI problems—AI-assisted bioterrorism, AI-assisted dictatorships, etc. And if they're concerned about egregious misalignment and scheming, they'll probably say that it would come about through race dynamics, careless programmers, bad actors, etc., as opposed to the simpler Yudkowsky & Soares story of “we get egregious misalignment and scheming because nobody has the faintest clue how to avoid that”. Here's my brief idiosyncratic take on this debate. I think BOTH of the following are true: (1) If you really think carefully about the properties of ASI, you really do find good reasons to strongly expect it to be egregiously misaligned, scheming, and ruthless, in the absence of yet-to-be-invented breakthrough technical alignment ideas.(2) If you [...] ---Outline:(01:58) Yudkowsky & Soares's position [caricatured]:(03:18) LLM people's position [caricatured]:(04:09) Conclusion(04:19) Bonus section: Further commentary(04:28) My "true objection" to Yudkowsky & Soares:(05:04) My within-frame complaint at Yudkowsky & Soares:(06:42) My "true objection" to LLM people:(07:11) My within-frame complaint at LLM people: --- First published: June 12th, 2026 Source: https://www.lesswrong.com/posts/DZaZ3fqHnvfLCftPu/sympathy-for-both-sides-of-the-egregious-misalignment-debate --- Narrated by TYPE III AUDIO.