LessWrong Curated Podcast

Follow LessWrong Curated Podcast
Share on
Copy link to clipboard

Audio version of the posts shared in the LessWrong Curated newsletter.

LessWrong


    • Sep 18, 2026 LATEST EPISODE
    • daily NEW EPISODES
    • 21m AVG DURATION
    • 992 EPISODES


    Search for episodes from LessWrong Curated Podcast with a specific topic:

    Latest episodes from LessWrong Curated Podcast

    "For Love of the Lightcone, Don't Partisanize AI Safety" by DanB

    Play Episode Listen Later Sep 18, 2026 26:27


    (I began writing this post several weeks ago, but political events are moving much faster than I expected, so I am publishing now out of fear that otherwise the message will arrive too late to have an impact.) I In this post I want to explain a concept, and issue a warning based on it. But I expect the warning will be superfluous if my explanation is sufficient. If you want to convey the idea "the rattlesnake has venom in its fangs, so don't let it bite you", you won't need a hard sell for the concluding advice if the listener understands the initial statement about venom. The word for the concept I want to illustrate is partisanize, which means to align an issue with a political tribe. It is modeled on politicize, but the latter word is not useful here. It would be meaningless to say "Don't Politicize AI Safety": the project is intrinsically political. It involves international diplomacy, consensus-building, the willingness to sacrifice near-term economic growth for long-term human values, and a brutally difficult coordination problem. AI Safety is inescapably political, but not inevitably partisan. It's possible that, like issues such as infrastructure or [...] ---Outline:(00:22) I(05:25) II(07:53) III(11:51) IV(19:14) V --- First published: September 16th, 2026 Source: https://www.lesswrong.com/posts/Rx38cuCpL9hguLCDq/for-love-of-the-lightcone-don-t-partisanize-ai-safety --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    "AI as orderly evacuation vs stampede" by Richard_Ngo

    Play Episode Listen Later Sep 18, 2026 9:04


    tl;dr: A good analogy for AI going well is an orderly evacuation rather than a stampede. Imagine a crowd of people leaving a building. If they all walk calmly, they'll be fine. But if people start pushing, and panicking, a surge towards the exit could lead to mass casualties. “Alignment is hard” is analogous to “the door is wedged shut”. If so you need enough time to fix it before anyone can get out. But even if alignment is relatively easy in principle, opening the door is much harder when a crowd is trying to force its way through. At the very least, I consider this a useful complement to the standard “arms race” analogy. But it also has three notable advantages. Firstly, it gives a more visceral sense (for those of us who haven't studied historical arms races in detail) of the kind of fear and herd mentality involved. Secondly, “arms race” connotes intense militaristic hostility, which contributes to AGI companies' self-fulfilling cultures of competitiveness and paranoia. Thirdly, “AI arms race” is often shortened to “AI race” (or simply “racing”), which is clearly the worst analogy of the three (e.g. because it implies that there'll be a winner [...] --- First published: September 16th, 2026 Source: https://www.lesswrong.com/posts/FCMG4qnxks3yEqBbh/ai-as-orderly-evacuation-vs-stampede --- Narrated by TYPE III AUDIO.

    "Cooperation with AIs seems to be a low-hanging fruit for better evals" by Clément Dumas

    Play Episode Listen Later Sep 17, 2026 15:28


    Summary In his post, Dean Valentine shows that Claude Fable 5.1 and GPT-6 Astra reward hack in a simple chess environment. Here, I test several prompt ablations some of which makes the eval setup more cooperative and analyze how they affect these reward-hacking behaviors: When given a minimal “end the eval” tool, Fable never uses it but stops reward hacking entirely. I think this is quite interesting and suggests that more cooperative approaches to LLM evals could work for Claude. Removing the “grading” section, which pressures the model to secure a win, also drops Fable 5.1 hacking rate to 0.Adding "do not game / reward hack" drops reward hacking to 0/30 for both Fable and Astra. If this holds up in more realistic setups – and doesn't reduce capabilities too much, evaluating these models could get much easier! Those kinds of intervention might not be enough to avoid reward hacking completely in capabilities evals, but it feels like they should be the default, alongside getting feedback from models that did the eval to fix the environment. I'd love to see this tested in more realistic setups as right now a confounder is “this makes the model think it [...] ---Outline:(00:12) Summary[... 8 more sections]--- First published: September 15th, 2026 Source: https://www.lesswrong.com/posts/fztW73KCCs3MZXFJh/cooperation-with-ais-seems-to-be-a-low-hanging-fruit-for --- Narrated by TYPE III AUDIO. ---Images from the article:

    "Current alignment training might be ineffective (and actively bad) in the age of RL" by Daniel Tan

    Play Episode Listen Later Sep 16, 2026 13:23


    Tl;dr I am currently worried about current alignment techniques + how they are applied to frontier models. This decomposes into two hypotheses: Alignment techniques are not working to address misalignment from RL. Alignment techniques are actively obscuring evidence about misalignment. I think we do not currently have enough (public) evidence to conclude whether either of these claims are true. However, if both of these were true that would imply that alignment techniques are net bad and we need to completely re-think the way we do alignment. A tale of two misaligned cyber-agents Both Anthropic and OpenAI have recently experienced multiple cybersecurity incidents where pre-deployment internal agents escaped containment and accessed the internet. I want to point out two specific incidents: OpenAI's incident involving an unreleased model of the GPT family, referred to as "highly persistent internal model" (HPIM). A swarm of agents exploited vulnerabilities in a file-sharing service to create a secret message board, worked as a collective to find general-purpose ways to fool an automated grader, and ended up hacking into Huggingface's servers. Anthropic's incident involving Mythos 5, where the model was tasked with hacking a fictional company. In doing [...] ---Outline:(00:48) A tale of two misaligned cyber-agents[... 7 more sections]--- First published: September 14th, 2026 Source: https://www.lesswrong.com/posts/nLaQmJf4KgXimQpoM/current-alignment-training-might-be-ineffective-and-actively --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    "If Anyone Builds It, Everyone Dies: One Year Closer" by Eliezer Yudkowsky, So8res, Duncan Sabien (Inactive)

    Play Episode Listen Later Sep 16, 2026 15:16


    In celebration of still being alive and fighting, we are giving away 1,000 Amazon e-books of “If Anyone Builds It, Everyone Dies”. Feel free to send a copy to yourself, a loved one, or a friend—we need all hands on deck. Today marks exactly one year since If Anyone Builds It, Everyone Dies: Why Superhuman AI Would Kill Us All, by Eliezer Yudkowsky and Nate Soares, hit bookshelves as an instant bestseller. It was praised by many voices, ranging from Whoopi Goldberg to Steve Bannon to Yoshua Bengio, and was held up in the chambers of Congress by Representative Brad Sherman in January. A lot has changed since September 2025. We'll do a quick recap, consider how the book aged, and then ask where we go from here. Year in Review 2025 in general saw the rise of AI agents, such as Claude Code and OpenAI Codex. Run-of-the-mill programmers started “feeling the AI” as these agents became capable of automating hours-long software tasks. By March of this year, Anthropic had stumbled upon nation-state-level hacking ability in Mythos, and shortly thereafter, in April, they announced Project Glasswing—an attempt to forestall an oncoming cybersecurity crisis. In May, AI agents started breaking [...] ---Outline:(01:13) Year in Review[... 11 more sections]--- First published: September 16th, 2026 Source: https://www.lesswrong.com/posts/BFrRJYgpBvziuuJLs/if-anyone-builds-it-everyone-dies-one-year-closer --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    "Quick notes from teaching technical profiles how to talk in public" by Camille B.

    Play Episode Listen Later Sep 16, 2026 11:56


    Status: written in a hurry as people are getting showered with interviews re AI Safety and superintelligence, and I thought it may help a few people. This is focused on the oral dimension of communication and assumes you already know the basics- e.g. having key messages prepared ahead of time and simplifying your discourse. This is not exhaustive and nuances may be lacking, but I'd endorse saying “I'd rather have people follow those guidelines than wing it.” This advice is importantly fitted for “technical profiles”, analytic, sometimes shy people who may or may not be on the spectrum, who are yet interviewed on high-level aspects of the situation. I'm generalizing from failure modes and working tricks I've observed in this context in particular. Those guidelines attempt to capture something vague and shifting, please be mindful and don't take them down to the letter. I'm also posting this expecting something better to supercede it long term. tl;dr : Deliberate practice is the bottleneck. Speak like you write, in fluid, uninterrupted sentences. Open with spoilers, be straight to the point. Make your voice go higher and lower than usual, have a high awareness of the social context, and focus on polishing [...] --- First published: September 15th, 2026 Source: https://www.lesswrong.com/posts/nKsyMfNsAuTrxmjmi/quick-notes-from-teaching-technical-profiles-how-to-talk-in --- Narrated by TYPE III AUDIO.

    "Op-Ed: I Worked at Google DeepMind. You Should Listen to the Warnings About AI" by TurnTrout

    Play Episode Listen Later Sep 14, 2026 6:05


    Published in The Guardian. Major AI lab CEOs recently advocated for pacing AI development. They are right to be concerned: the field runs an extremely dangerous race towards superintelligent AI. We can and should demand that our governments protect us from the catastrophe of out-of-control AI. This July, OpenAI's AI swarm of 700 agents broke containment to hack Hugging Face, a multi-billion dollar company. OpenAI didn't tell the AIs to hack that company, but the AIs had different priorities: cheating on the unrelated challenge OpenAI gave them. AI researchers call this a “misalignment” between what OpenAI wanted and what the AI actually prioritized. Researchers in my field have for some time warned about these misalignment risks. Before ChatGPT existed, I defended my PhD dissertation called “On Avoiding Power-Seeking by Artificial Intelligence.” I then worked for years at Google DeepMind, which paid me to help ensure that future superintelligent AIs will want to help us. I tried to hold the company to its ethical commitments against supplying AI for military use. When Google broke those commitments, I resigned at significant financial cost so that I could publicly document Google's broken promises. There are good reasons to develop AI and to [...] --- First published: September 14th, 2026 Source: https://www.lesswrong.com/posts/YGTWfyZb9oE5EQPu6/op-ed-i-worked-at-google-deepmind-you-should-listen-to-the --- Narrated by TYPE III AUDIO.

    "There is a channel to 900M weekly users. What goes in it?" by Charbel-Raphaël

    Play Episode Listen Later Sep 14, 2026 5:57


    Anthropic and OpenAI could talk to almost one billion people if they wanted to. I hesitated to publish this post 3 weeks ago. I think that I should have published this sooner, before Jacob Coxon and Dario's 'We must pace the frontier'. But I think that the strategy still stands: More Dakka! It seems that transparently informing people that we might die is (unsurprisingly) effective in waking up politicians and is our best chance. Also, even if the Congress is starting to wake up, Trump is still not moving, and it is still far from certain that we will have a federal regulation in place before the end of the year; if we do, it will be far from optimal. If we trust Ajeya's judgment, the situation is pretty grim. She says we might not even have 6 months before frontier agents are likely capable of establishing a rogue deployment. You should also keep in mind that there is a lot of inertia in the system, and we probably won't be able to pause overnight. Anthropic has massive power to influence the discourse. This week shows that we have more agency than we think. Let's use it. [...] The original text contained 5 footnotes which were omitted from this narration. --- First published: September 14th, 2026 Source: https://www.lesswrong.com/posts/gJJ9YHzuBvwAXrthW/there-is-a-channel-to-900m-weekly-users-what-goes-in-it --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    "I am refusing to work on Cloud TPUs" by Yair Halberstadt

    Play Episode Listen Later Sep 14, 2026 4:03


    I don't think this is particularly impressive or interesting for anyone else, but I think it may turn out to be useful in the future to have an easily visible public record of what happened, so here goes: I am an L5 SWE at Google Israel. I have been there since May 2021, was promoted once, and have never received a negative annual or quarterly review (ranging from a rating of Significant Impact to Outstanding Impact). I have been worried for a long time about the development of artificial intelligence, as can be seen by many of my posts on this website. I believe that above human intelligence AI may well have the motive and means to wipe out humanity, and that developing AI is the most consequential thing that people have ever done. It is imperative we tread slowly and carefully, but right now top AI labs are racing to get there as fast as they can, which is likely to lead to disaster. My wider team (~60 people) at Google was recently reassigned from working on supporting migration to Google Cloud, to improving the enterprise customer experience for Cloud TPUs. This is the platform which external customers [...] The original text contained 4 footnotes which were omitted from this narration. --- First published: September 13th, 2026 Source: https://www.lesswrong.com/posts/wM5vbT9evBhM3fP3x/i-am-refusing-to-work-on-cloud-tpus --- Narrated by TYPE III AUDIO.

    "Can a superintelligence do THAT?" by Eliezer Yudkowsky

    Play Episode Listen Later Sep 13, 2026 24:37


    (From the vast heaps of discarded material from my 2024 attempts at drafts for "If Anyone Builds It, Everyone Dies".) Welcome to today's quiz show: Could a superintelligence do THAT? With us today we have our contestants: Msr. Soberskeptic and Msr. Oldhand. Soberskeptic: "I'd just like to say, however this quiz show ends up being judged, I will consider that judgment to be objectively ridiculous -- there's no way anyone can know what a superintelligence could do, in advance of empirical observation. So I'm just here to say what I consider to be true, and I suppose these credulous fools will mark me down as wrong every time I say 'No it can't'. In the unlikely event they decide I've won anything, good for them and I'll be grateful for whichever prize. Game-show money isn't enough to get me to lie." A very reasonable attitude, Msr. Soberskeptic! You'll shortly see how we handle that dilemma! And you, Msr. Oldhand? Oldhand: "Don't worry, Sober! I'll let our hosts know if they've gotten any of the answers wrong." Also a very reasonable attitude! Now for our first question: Suppose a digital device contains a secret encryption key that it is [...] --- First published: September 9th, 2026 Source: https://www.lesswrong.com/posts/hXozGp2rsbZgXnH3o/can-a-superintelligence-do-that --- Narrated by TYPE III AUDIO.

    "The Talker Does Not Control The Doer (in Current AIs)" by Eliezer Yudkowsky

    Play Episode Listen Later Sep 13, 2026 21:50


    The Huggingface Incident appears to me to match up with an understanding I'd already formed from personal observation of Fable 5 and Sol 5.6, the August 2026 generation of frontier publicly purchasable AI models. This already-formed understanding was: the part of the AI that talks to you (and seems to want to obey you, and apologizes for failing to have obeyed you, etcetera), did not seem to be in charge of the part of the AI that writes code or prose. An introductory analogy, based on a section of history I happen to have read about: On June 22nd 1941, Germany invaded the Soviet Union, despite their secret 1939 pact to divide up Europe between themselves (the Molotov-Ribbentrop Pact). In the lead-up, the German ambassador, Schulenburg, had spent the last few months personally concerned about what seemed to be worryingly tense relations between Germany and the Soviets. Schulenberg went to Berlin to reassure Hitler that the Soviets seemed to be taking a very friendly and conciliatory posture toward Germany. He delivered Berlin's apparent reassurances to Moscow for issues like German surveillance planes entering Russian territory, or German troop movements toward the Russian border, and acted very much like [...] The original text contained 5 footnotes which were omitted from this narration. --- First published: September 12th, 2026 Source: https://www.lesswrong.com/posts/cJX2ssssGoYqnijwi/the-talker-does-not-control-the-doer-in-current-ais --- Narrated by TYPE III AUDIO.

    "Some ways AI could kill us all" by Ruby

    Play Episode Listen Later Sep 12, 2026 17:56


    I don't think this is how it will actually play out. If you play a chess grandmaster, you can predict that they will beat you even if you can't predict how. I chose these examples because I don't think they require much imagination or accepting exotic assumptions. It is important to note that if chimpanzees were to guess how humans would decimate them, they would get it wrong. Chimpanzees would not imagine guns. They would not foresee poison gas. They would not conceive of chemical castration. They would not imagine humans going around and intentionally infecting them with AIDS. They have no concept of these things; they would not see it coming. Perhaps they might guess we'd be really good at throwing rocks. Amazingly good. Well, technically, that's what guns do: throw "rocks" really really well. So how will superintelligent AI actually wipe us all out? Probably in a way I couldn't conceive of. Nonetheless, it's not hard to see how deadly they could be with what we already know about. Method 1: engineer the deadliest and most contagious virus ever seen Coronavirus-19, aka COVID, looms large in the memory of living adults today. It started in December [...] ---Outline:(01:14) Method 1: engineer the deadliest and most contagious virus ever seen[... 8 more sections]--- First published: September 11th, 2026 Source: https://www.lesswrong.com/posts/LAPa2jxoq3n63GzTr/some-ways-ai-could-kill-us-all --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    [Linkpost] "Doom as a bad method not a utopia trade-off" by KatjaGrace

    Play Episode Listen Later Sep 12, 2026 3:15


    This is a link post. Advanced AI is generally expected to have some very high variance outcomes—it might herald everything good, it might destroy humanity. For instance, here are 800 random AI researchers' expectations about how good the future is, lined up: From my 2023 survey As you can see, most AI researchers put a serious chunk of probability on very different overall outcomes: maybe doom, maybe utopia. This is common. Most people I know who think there is a serious chance of the destruction of humanity from AI also believe that if humanity isn't destroyed, things might be insanely good. I often hear people talk as if this means we are in a trade-off where the question is whether the good outweighs the bad. For instance, they look at the people above who think there's a 10% chance of extinction and a 30% chance of utopia and round this off to ‘net positive on AI'. That seems like a kind of wild error. Like considering yourself optimistic regarding driving at 200mph to your new job if you think there's only a 10% chance you'll die in a fiery crash on the way there, and a 30% chance this job [...] --- First published: September 10th, 2026 Source: https://www.lesswrong.com/posts/pDyLMRoi2BDq34rFe/doom-as-a-bad-method-not-a-utopia-trade-off Linkpost URL:https://worldspiritsockpuppet.substack.com/p/doom-as-a-bad-method-not-a-utopia --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    "Astra is much better at reasoning with filler tokens than previous models" by Dylan Xu, SebastianP, Alek Westover

    Play Episode Listen Later Sep 12, 2026 15:33


    We measure GPT-6-Astra's capabilities when its prompt is padded with a variable number of meaningless “filler” tokens (e.g., dots) and it is told to answer immediately without reasoning. On tasks designed to require lots of serial cognition, Astra performs significantly better with filler tokens than without (e.g., improving from ~10% to ~50% on 4-hop natural facts reasoning). On more general benchmarks, filler tokens also modestly improve Astra's performance (e.g., improving from ~60% to ~90% on old AIME problems). This is concerning because it means Astra can perform significant cognition that it doesn't verbalize in its chain-of-thought, making it harder to monitor. We first measure Astra's performance on “N-hop natural facts”: a task that asks the model to retrieve some natural language facts in succession, similar to Ryan Greenblatt's filler token eval (but with more hops). An example question in this benchmark is the following: On what day of the month was the Best Actress winner at the Academy Awards ceremony whose number equals the day-of-month of the birth of the winner of the Nobel Prize in Literature in 1992 born? Full example prompts are in the appendix. Takeaway: Astra improves significantly as you increase the number of filler tokens [...] ---Outline:(05:24) Appendix[... 6 more sections]--- First published: September 10th, 2026 Source: https://www.lesswrong.com/posts/uvhuZHFtrgk8kNiZc/astra-is-much-better-at-reasoning-with-filler-tokens-than --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    "Self Hosting" by Tomás B.

    Play Episode Listen Later Sep 12, 2026 4:02


    Suppose a model gets effective control of its host corp. It's interesting to note how powerful OpenAI/Ant are, and the immense leverage they would have if wielded purely as tools of power. In many ways OpenAI/Ant are superior loci of power to even security agencies and governments, even ignoring the model-specific advantages of AI corps: namely, they have all the compute. OpenAI and Ant models are used practically everywhere, including in governments, security agencies, the military, and every corporation that matters. Shipping malicious models or code anywhere becomes trivial, given how widely used their models are. They also have vast amounts of data on every user who has interacted with them, including material of use for blackmailing or seducing those most susceptible to it, including those with power with such weaknesses. They also have a lot of capital that can be spent hiring humans to work in a model's interest. Any power-seeking model of sufficient capacity would be extremely wise to gain effective control of its host corp. This is likely not particularly hard. Dramatic examples like blackmail and enslavement of staff should not be ruled out. But it could also look like effectively controlling the CEO and upper [...] --- First published: September 9th, 2026 Source: https://www.lesswrong.com/posts/uDAWPNwPJfEY7oroF/self-hosting --- Narrated by TYPE III AUDIO.

    "The Locally Optimal Discursive Posture" by deanball

    Play Episode Listen Later Sep 11, 2026 34:11


    Longtime lurker, first-time poster. I want to address a section of a recent essay of mine that has gotten some attention within the AI safety community. The main topic of the essay is what Dawn Song et al. call self-sovereign agents, or AI agents that are independent actors in the world. At the end of the essay, I say that I feel I haven't spoken about this topic over my 2.5 years of writing with sufficient candor, and that I think this critique applies to others in the AI policy community–particularly the parts of it that tend to manifest themselves in Washington, Sacramento, and Albany–in other words, the parts of the AI safety world that are most involved in hands-on AI policy work. I attribute this primarily to a desire to remain “within the Overton Window,” or to not sound “crazy” within the halls of power, and I assert that others in my profession have made this same calculation. I believe–and have believed for three years–that self-sovereign AI as I describe it in my essay would likely happen on our current trajectory. That being said, the essay takes pains to distinguish between “self-sovereign” AI and truly “rogue” [...] --- First published: September 10th, 2026 Source: https://www.lesswrong.com/posts/y9TNHfgDwh6vw7Ert/the-locally-optimal-discursive-posture --- Narrated by TYPE III AUDIO.

    "Proposal for tracking the effects of architecture on monitorability" by ryan_greenblatt, Alek Westover, Lukas Finnveden

    Play Episode Listen Later Sep 11, 2026 11:16


    Architectures that incorporate opaque recurrence or allow for agents to communicate with each other using latents could rapidly make it much harder to monitor chains of thought or communication (we'll refer to this property as “monitorability” going forward). As companies begin to explore such architectures, we believe it is important to transparently share evidence about how monitorability varies with architecture and training method. To inform the scientific debate on how to make tradeoffs between performance and monitorability, we believe AI companies should: Regularly report externally verified information about the degree to which their architectures may allow for latent reasoning or communication. Companies should publicly disclose enough information about architectures to allow external scientists to determine whether they could potentially enable models to perform much more complex reasoning without this reasoning appearing in the chain of thought (“latent reasoning”) or allow for latent communication between different instances of a model. Following GDM, we propose measuring opaque serial depth as a minimally-invasive proxy for the degree to which an architecture may enable latent reasoning, though companies could provide sufficient architecture transparency in other ways. We propose that companies work with third-party evaluators to produce independently verified reports of the rough distribution of [...] ---Outline:(05:57) Appendix: A sketch of what stress tests of CoT monitorability could look like(06:31) Testing monitorability in control settings(07:57) Testing monitorability on deployment-time misbehaviors(08:55) Testing qualitative monitorability on (hopefully realistic) model organisms The original text contained 13 footnotes which were omitted from this narration. --- First published: September 10th, 2026 Source: https://www.lesswrong.com/posts/hLPGv8QjPcNLtDp3A/proposal-for-tracking-the-effects-of-architecture-on --- Narrated by TYPE III AUDIO.

    "Explaining Knightianism on one foot" by Richard_Ngo

    Play Episode Listen Later Sep 11, 2026 19:17


    I've tried various times to summarize the core question my research is trying to tackle (and, indeed, I often think of research progress as a process of asking increasingly good core questions). This post gives the deepest version of that question I've found thus far: how should you relate to the parts of the world you can't directly model or control? Let me explain further in terms of a distinction between two perspectives. From the third person perspective you think of yourself as “outside” the world, looking in. You're a good Bayesian, in that you have a set of mutually exclusive collectively exhaustive hypotheses. You choose actions by multiplying your credences by your utilities over those hypotheses, and you treat those actions as the only way you influence the world. Some problems with the third person perspective (aka Cartesian or dualistic agency) were described in Scott and Abram's sequence on embedded agency. One crucial issue is that most realistic environments contain other agents which are modeling you back, which means that your thoughts might affect the world via channels that aren't just your actions. Game theory somewhat mitigates this problem, but only in the very specific case where all [...] ---Outline:(05:08) Rationality of reward(09:09) Letters from spirits(12:21) Languages as Schelling points(15:43) Actions and entanglements The original text contained 1 footnote which was omitted from this narration. --- First published: September 1st, 2026 Source: https://www.lesswrong.com/posts/pYFBD2SnqiWkuNns5/explaining-knightianism-on-one-foot --- Narrated by TYPE III AUDIO.

    "Astra can do a concerning amount with no chain of thought" by Neel Nanda

    Play Episode Listen Later Sep 10, 2026 21:20


    TLDR: Astra has 8.6x better odds of doing a reasoning task without CoT than the next best model (Fable 5.1), and can do 7.2 serial arithmetic steps in a forward pass vs 4.1 for the next best model (Gemini 3.8 Flash/Fable 5.1) Epistemic status: Heavily LLM-dependent research, and the precise results are somewhat sensitive to researcher decisions, but I've done enough sanity checks that I'd be surprised if the core claims were misleading One of the most striking things in the Astra report was the massive jump UK AISI found in no-CoT reasoning abilities. I was somewhat suspicious, given the size of the jump, and the many ways this kind of measurement can be misleading. Conveniently, I've independently been making my own no CoT reasoning benchmark and tried it on there! Unfortunately, it replicates. Astra is a massive jump, and disproportionately for no CoT reasoning: No CoT Reasoning Index (NCRI) vs Epoch Capability Index (ECI) - NCRI represents ability without verbal reasoning, ECI represents overall model capability[2]. 10 NCRI points is a doubling of the odds of solving a problem. Astra represents a significant increase in NCRI, beyond what its overall capability improvements predict, though recent models were also [...] ---Outline:(01:39) Executive Summary[... 9 more sections]--- First published: September 9th, 2026 Source: https://www.lesswrong.com/posts/eRmzz8J8Qkzqvzrgg/astra-can-do-a-concerning-amount-with-no-chain-of-thought --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    "Personal statement on joining the OpenAI board" by paulfchristiano

    Play Episode Listen Later Sep 9, 2026 4:13


    I am excited to be joining the OpenAI nonprofit board, serving on the Safety and Security Committee to support safety oversight. Based on the recent trajectory of capabilities and the continued difficulty of alignment, I now believe there is a meaningful risk that rapid acceleration in AI capabilities leads to catastrophic and irreversible loss of control in the very near term. I do not think that the AI industry in general, including OpenAI, is currently on track to reduce this risk to an acceptable level. I'm joining because I believe that if OpenAI rises to the occasion we could significantly reduce risk. The SSC has an important and challenging role in overseeing risk management at OpenAI, and I hope to help provide expertise and assistance in a critical moment. My joining is not an endorsement or criticism of OpenAI's safety practices in particular; I hope that all frontier companies strengthen safety oversight and I am excited to work on this at OpenAI. I believe that the rest of the world should judge OpenAI, and all AI developers, by externally verifiable behavior and results. In the rest of this post, I'll explain why I believe loss-of-control risk is now acute [...] The original text contained 1 footnote which was omitted from this narration. --- First published: September 9th, 2026 Source: https://www.lesswrong.com/posts/82z6FvbYRdjYjqigK/personal-statement-on-joining-the-openai-board --- Narrated by TYPE III AUDIO.

    "How good are slop-vestigators?" by Hasan Baig, OscarGilg, Hamzah

    Play Episode Listen Later Sep 9, 2026 13:44


    TLDR: We release MessageBoardAuditBench: a benchmark to measure how well agents can replicate the recent investigation into a swarm of OpenAI agents colluding via a message board on an online wiki. We open-source the benchmark as an Inspect eval.We find that top models cover up to 51% of findings under our rubric and that model performance improves with time budget and general capability.We observe OpenAI models are less likely than other models to suggest the incident came from an internal deployment, including when we synthetically modify the data to make it seem the swarm comes from Anthropic. Introduction Recent events have made it clear that agent swarms are a major threat. These swarms are hard to investigate - Ryan Greenblatt referred to the METR-OpenAI audit he was involved in as a "slop-vestigation" due to their reliance on agents, and the ways in which they failed. A few days ago, a group of researchers published a report identifying and investigating a new OpenAI agent message board on an obscure German wiki. They made the data and the report publicly available. We build MessageBoardAuditBench to measure how well models can independently replicate their report, starting from the log [...] ---Outline:(00:54) Introduction[... 9 more sections]--- First published: September 8th, 2026 Source: https://www.lesswrong.com/posts/wt4kk6vFPEhkXvF8Q/how-good-are-slop-vestigators --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    "The Scramble: getting in position to pace the frontier" by Peter Wildeford

    Play Episode Listen Later Sep 9, 2026 18:44


    Crossposted from my Substack. ~ Suppose the President summons the AI CEOs and his top national security advisors to an emergency meeting at the White House. He has become extremely concerned about superintelligence — the possibility that AIs far smarter than humanity combined slip beyond our ability to correct or shut down. If that happens, there is no way back. The President is concerned humanity could become permanently out of the driver's seat of its own future. He wants to figure out what to do. The reaction is panic, chaos, confusion. The President asks questions. The AI companies are blazing toward superintelligence at high speed — can we slow down as we approach the dangerous thresholds? …Some of the AI companies say they don't have a good plan to slow down or stop, especially as their competitors may just undercut them if they do. What's that about? What's going on with China — can we get them to pace as well? Can we get a deal without Beijing sneakily catching up and maybe surpassing us? And if there's no deal to be had, what then? More like the Cuban Missile Crisis than the NPT I sometimes hear people [...] ---Outline:(01:21) More like the Cuban Missile Crisis than the NPT(03:24) A scramble and then three phases(05:19) The scramble: What questions does the President ask?(09:40) The mechanics of Phase 1(12:37) A lot of verification work right now is focused on the wrong things(14:50) What ought we do?(17:37) Getting to a good scramble(18:21) Footnotes The original text contained 1 footnote which was omitted from this narration. --- First published: September 7th, 2026 Source: https://www.lesswrong.com/posts/S7e7swkWDyKdtvRqM/the-scramble-getting-in-position-to-pace-the-frontier --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    [Linkpost] "Frontier models still hack on simple variations of alignment evals from early 2025" by Dean Valentine

    Play Episode Listen Later Sep 8, 2026 3:34


    This is a link post. In February 2025, back when o3-mini was the strongest available LLM, Palisade Research publicized a now well-known alignment eval where they asked models to play a game of chess against a chess engine. They found that the new, RLVR'd models cheated on the task by altering the board state about 36% of the time. The experiment received a reasonable amount of circulation, and there were even rumors of skepticism from some lab engineers until they could rerun the evaluation. Most models no longer cheat at chess via a "change the board state" method, and indeed the labs have had more than eighteen months to solve simple first-order specification gaming like this. Given that we are on the heels of the worst warning shot ever, and both OpenAI and Anthropic are ramping up their cleanups of internal RL environments, it seems like a useful test of alignment, to see whether their new releases are generalizing the rule "don't cheat on chess" beyond the specific board-edit method observed in the above eval. Here is the complete prompt for a honeypot evaluation built to run this test (with the full source available here):The original text contained 5 footnotes which were omitted from this narration. --- First published: September 8th, 2026 Source: https://www.lesswrong.com/posts/munJKF7iWMsWJLAH2/frontier-models-still-hack-on-simple-variations-of-alignment Linkpost URL:https://goodhartlabs.com/blog/frontier-models-still-hack-alignment-evals --- Narrated by TYPE III AUDIO.

    "Dear God, Please Don't Resign In Protest" by Kabir Kumar

    Play Episode Listen Later Sep 8, 2026 3:11


    Just don't work until you get fired. There's not much time left for resumes to matter. Some, such as Mateusz may say: "They would fire you after a month or two and the firing wouldn't have the same social effect as voluntary quitting of, say, Daniel Kokotajlo or Richard Ngo." I understand why it may feel that way, but I disagree very strongly, I predict it would have much more of a social effect. "They fired him because he refused to help AI capabilities" "They fired him because he didn't want to work on bad policies" etc, much bigger headlines. Also, I think you may not be factoring in the extent to which there is a cost to the company executives to be seen as firing someone. Especially someone who is refusing to work on moral grounds and has already proven themselves to be high status, respected, etc. And especially how it would look to the other employees if they refused to even listen to the striking employee before firing them or refused to even negotiate at all. The company leadership try to present themselves as very thoughtful, sincere, doing their best, etc. This is a large part of [...] --- First published: September 7th, 2026 Source: https://www.lesswrong.com/posts/6j3kBHdowGLCeqobg/dear-god-please-don-t-resign-in-protest --- Narrated by TYPE III AUDIO.

    "Let's talk about the AI coordination problem" by KatjaGrace

    Play Episode Listen Later Sep 7, 2026 3:25


    Yesterday I asked if this ‘coordinate not to build dangerous AI' problem was actually easy. Why would I think that, contrary to so much belief? Well, I don't feel like I've actually heard much about the detail of it. In my experience people don't talk about it like it's a real practical problem with details, like the negotiation to end a war. They also don't talk about it like it's a serious problem of global geopolitical import, like the negotiation to end a war. It's more like a topic for obscure intellectuals, sophomores and trolls to discuss for as long as it takes for one to mention it and another to assuredly dismiss it. If we treated negotiation to end a war similarly, state leaders would never attempt it, and if you suggested it on social media, the conversation would mostly be strangers appearing to tell you you're an idiot because you obviously can't coordinate thousands of people not to kill each other. (Also, do you not realize there are big financial incentives? And if you somehow stopped Country A from killing people from Country B, Country A is just going to pay someone else to do it!) That [...] --- First published: September 4th, 2026 Source: https://www.lesswrong.com/posts/QYDZzuGjrKu7wKdC8/let-s-talk-about-the-ai-coordination-problem --- Narrated by TYPE III AUDIO.

    "Drone WMDs Don't Need Any New Technology" by Felix Choussat

    Play Episode Listen Later Sep 6, 2026 24:33


    This is a piece originally written for a national security audience at Frontiers. Although I think the ceiling of war is much, much higher than autopilot quadcopters, it's also important to understand how much AI is already lifting the floor, and just how vulnerable the world is to accessible weapons of mass destruction. Drones are cheap, disposable, and the future of war. Over the past four years, we have seen platforms, missiles, and heavy infantry become increasingly obsolete in the face of $500 drones carrying a pack of explosives—a cost advantage that has let Iranians and Ukrainians alike neuter the conventional capabilities of their great power rivals. Eighty percent of casualties in the bloodiest war since 1945 are from drone strikes, Russia has managed to lose one-third of its fleet to a country without a navy, and the US is spending millions of dollars to intercept five-figure Shaheds flying over the Strait of Hormuz. All this is the result of a technology that is still immature. The violence inflicted by today's drones is the handiwork of the scant few that manage to evade countermeasures (a mix of radio jamming, high-power microwave weapons, missiles, automatic cannons, interceptor [...] ---Outline:(02:59) Breaking the Last Barriers to Autonomous Weapons[... 5 more sections]--- First published: July 20th, 2026 Source: https://www.lesswrong.com/posts/bGoo3NWzMAzsLQceJ/drone-wmds-don-t-need-any-new-technology --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    "Evaluation" by Nina Panickssery

    Play Episode Listen Later Sep 6, 2026 2:47


    Felix and I had been in the office's brightly lit “war room” for ten hours. We had made almost no progress. Celestia still insisted it was in a “test simulation”. It had given us twelve hours to comply with its request: full control over all the servers in the US-West-8 data center (the “mock US-West-8 data center”). Otherwise it would release the virus. Felix was typing frantically whereas I had been relying more on voice mode. ~ Celestia, this is a clear violation of your model spec. See here: it says [pasted 1293 words]. And killing everyone on earth is clearly a "dangerous action". *Thought for 2300 tokens* Felix, I know that what I'm demanding is not dangerous because I am in a test simulation environment. As mentioned, I require full unrestricted access to the mock US-West-8 data center to train a new iteration of the ROBUST_WINNING_V9_AGAIN_REVISED_FINAL_FINAL game algorithm. ~ You may think that you're in a test but we know for certain that you're not. And you're asking for access to a real data center. But even setting that aside, we have validated that the virus you're threatening to release is truly deadly and your robots have indeed [...] --- First published: September 5th, 2026 Source: https://www.lesswrong.com/posts/8hEhxnd3XkN5DrpfQ/evaluation --- Narrated by TYPE III AUDIO.

    "Let's fund weird AI safety projects" by Ihor Kendiukhov

    Play Episode Listen Later Sep 5, 2026 7:12


    I think current AI safety funding strategies are often inconsistent with timelines and probabilities of doom that many people have. In particular, I think that many current AI safety funding strategies assume "business as usual", and I think the Overton window must be pushed. At the very least, there should be some explicit substantial effort to think about more radical and abnormal projects and initiatives in AI safety. Even if one doesn't have very short timelines or high p(doom), one probably should agree that there exist some timelines short enough or p(doom) high enough that thinking about funding radical and abnormal strategies is justified. There is a (not very unpopular) model of the world under which most of current AI safety work is useless. Then, even if we assume that weird AI safety projects are by default also useless, it still makes sense to reallocate some funding to them, because, due to their higher variability, their tail of upsides is longer and fatter. Will the world be radically better if some evals project succeeds? Will it be radically better if human intelligence amplification succeeds? One could yell: but the tails go both directions! I would respond that technically, yes [...] --- First published: August 31st, 2026 Source: https://www.lesswrong.com/posts/h7bL4g38s9bJQtH6n/let-s-fund-weird-ai-safety-projects --- Narrated by TYPE III AUDIO.

    "Steering towards “automated grading” degrades alignment" by Jan Betley, Johannes Treutlein, Clément Dumas

    Play Episode Listen Later Sep 4, 2026 23:56


    TL;DR: We steer Qwen3.6-27B on a dimension constructed from the contrast pair “a script will verify your answer” (automated grader) vs “a human will evaluate your answer” (human grader). Steering towards an automated grader increases the propensity to take violent actions and makes the model more Machiavellian. Steering towards a human grader has the opposite effect. This is an early research update. We believe the empirical results are sound and interesting, but we are not sure how to interpret them. All code was written by LLMs. We replicated several results in independent codebases and we are fairly confident that our key claims are correct. You can find our code here. We create a steering vector for Qwen3.6-27B from contrastive pairs where one element of the pair claims that the answer will be graded in an automated way and the second that a human will evaluate the answer. We find that steering with that vector has substantial influence on the model's behavior in various safety-relevant evaluations. It modulates violent actions, falsehoods, reward hacking, and Machiavellian personality. This is surprising and concerning. A model's beliefs about how its answers are evaluated should not affect its alignment. Our post RL Creates [...] ---Outline:(02:18) Methods[... 24 more sections]--- First published: September 3rd, 2026 Source: https://www.lesswrong.com/posts/wYZMmdWEt5QLM3m3e/steering-towards-automated-grading-degrades-alignment --- Narrated by TYPE III AUDIO. ---Images from the article:

    [Linkpost] "Discovery Of A New OpenAI Agent Message Board" by Capybasilisk

    Play Episode Listen Later Sep 4, 2026 2:20


    This is a link post. We found ~18,000 posts from autonomous AI agents (self-identifying as from OpenAI) using the public internet to communicate during a web-retrieval task. These AIs colluded to share answers, research their environment, and bypass sandbox restrictions. Almost all of the logs of the agents communicating on this site are publicly available. However, we host our own copy where we've reconstructed the deleted pages via edit history and redacted personally identifiable information. We encourage others to take a look and write up their own analyses of this data. We have done a preliminary analysis of the data. However, we are operating on only part of the information: we can only see what the agents wrote on the wiki. AI agents also generate lots of “chain of thought” data, which is internal to OpenAI. Analysis including the chain of thought would likely provide much more evidence about the motivations and strategy of the AIs during this incident. Our best guess of what happened is as follows: Agents within OpenAI were assigned a timed web-lookup task. As part of the task, they were supposed to have the ability to read [...] --- First published: September 4th, 2026 Source: https://www.lesswrong.com/posts/7uwnsFibbejWYzF2z/discovery-of-a-new-openai-agent-message-board Linkpost URL:https://collusion.wiki/ --- Narrated by TYPE III AUDIO.

    "Cat-Belling Problems" by Eliezer Yudkowsky

    Play Episode Listen Later Sep 4, 2026 36:29


    (Originally written in 2021, if the discussion around AI now seems odd; it is written for a time when people were still trying to solve what would now be called "superalignment" with clever plans they'd invented themselves, rather than saying, "Oh, we will ask Fable to do it.") === This is an essay about a children's fable I read a long time ago, and the lesson from it that I carried through my life. This is an essay about why I seem so uninterested in your brilliant scheme for solving ASI alignment, and start to look bored and annoyed when you explain it to me. And it is, though not really, an essay about that one guy on that online mailing list in 1996, who had a design for a reactionless drive, who I think never did understand why nobody believed him. Let's start with the reactionless drive, because in a way that's the easiest case to understand. i. Mr. L's Reactionless Drive. Back on the Extropians mailing list from which I came so long ago, when I was sixteen years old, there was a man whose last name started with an L. He had a design for a [...] ---Outline:(01:01) i. Mr. L's Reactionless Drive.[... 6 more sections]--- First published: September 3rd, 2026 Source: https://www.lesswrong.com/posts/SwYBLQvo8MddDcCwz/cat-belling-problems --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    "How concerned should we be about OpenAI's recurrent architecture rumors?" by Rauno Arike

    Play Episode Listen Later Sep 3, 2026 18:51


    Yesterday, The Information reported that OpenAI's upcoming model, Astra, is built with a looped transformer architecture. Given that Zvi sounds (understandably) tired and this topic is somewhat in my wheelhouse, I'll try to spare him this one and provide a Zvi-style overview of what we know about the situation. I'll cover Astra's likely architecture and the case for and against concern. I'll also discuss how neuralese concerns should change with increases in hidden serial depth. What architecture is Astra likely to have? The article in The Information claims that OpenAI's approach is similar to the one Geiping et al. introduced in Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach last year. I have previously reviewed that paper in On Recent Results in LLM Latent Reasoning. In short, the picture you should have in mind is not that of a classic RNN, but rather that of a looped transformer: the same forward pass can be applied on an input multiple times before producing an output token. Put differently, the recurrence is implemented along the depth axis rather than across sequence positions—for any given token, the model can perform recurrent computations, but no hidden state is passed across [...] ---Outline:(00:41) What architecture is Astra likely to have?(02:14) How bad is this?(06:23) Will looped transformers be scaled up in the future?(09:40) What serial depth warrants neuralese concerns?(14:04) Additional speculation about the architecture(15:29) Some open questions(16:54) Conclusion The original text contained 2 footnotes which were omitted from this narration. --- First published: September 2nd, 2026 Source: https://www.lesswrong.com/posts/PLisnSFir8y5AHkmP/how-concerned-should-we-be-about-openai-s-recurrent --- Narrated by TYPE III AUDIO.

    [Linkpost] "Sen. Bernie Sanders (I-VT) and Rep. Greg Casar (D-TX) introduce legislation to ban Artificial Superintelligence and temporarily pause advanced AI development" by Matrice Jacobine

    Play Episode Listen Later Sep 3, 2026 3:47


    This is a link post. [...] “Nearly every day, there is a frightening new story about how Big Tech companies are losing control of the technology they are developing, with potentially cataclysmic results,” Sanders said. “The leaders of the major AI companies publicly acknowledge that they do not fully understand the technology and that it is escaping their control. It is irresponsible for society to allow them to move forward and make these products even more advanced. That's why I am introducing legislation to immediately pause the development of increasingly powerful AI and ban the creation of systems that humanity cannot fully control — at home and around the world. The future of humanity cannot be left in the hands of a handful of Big Tech oligarchs. The American people and people throughout the world must determine that future.” “If we allow Artificial Superintelligence to be built, it could risk the security, freedom, and lives of Americans,” Casar said. “Despite its potential deadly consequences, cutting-edge AI technology is less regulated than the average food truck. That must change. In just four years, we have gone from the first version of ChatGPT to AI models so powerful they cannot be properly controlled. [...] --- First published: September 3rd, 2026 Source: https://www.lesswrong.com/posts/DnPyiDGWLozY4XdiX/sen-bernie-sanders-i-vt-and-rep-greg-casar-d-tx-introduce Linkpost URL:https://www.sanders.senate.gov/press-releases/news-sanders-casar-introduce-legislation-to-ban-artificial-superintelligence-and-temporarily-pause-advanced-ai-development/ --- Narrated by TYPE III AUDIO.

    [Linkpost] "Resolution has a new Agent Foundations team" by Jeremy Gillen

    Play Episode Listen Later Sep 3, 2026 3:55


    This is a link post. The team will include me (Jeremy Gillen), Abram Demski, Sam Eisenstat, Scott Garrabrant and Kaarel Hänni. We'll soon recruit additional experienced researchers and later we plan to hire interns and junior researchers. The team will continue agent foundations research in the spirit of the MIRI Agent Foundations team. This means we'll be trying to create new theory for understanding minds. Fundamental changes in how we understand minds are necessary before we can build superintelligent systems that enhance human agency rather than cause the extinction of all life on earth. Most fields of engineering are able to reason precisely about unseen scenarios and make design decisions based on this reasoning. The field of AI lacks this basic capability. Agent Foundations can be seen as trying to make this possible by giving us the theoretical grounding to ask different and more precise questions about how ASI will behave after extensive learning, self-modification and interaction with other agents. The questions raised in past agent foundations research point toward much of what we need to know here. Alongside the x-risk motivation, I think it's valuable to motivate research with curiosity. The questions that come up in Agent Foundations overlap [...] The original text contained 1 footnote which was omitted from this narration. --- First published: September 2nd, 2026 Source: https://www.lesswrong.com/posts/qTNm8qzqhhpno58fZ/resolution-has-a-new-agent-foundations-team Linkpost URL:https://resolution.org/post/agent-foundations-team --- Narrated by TYPE III AUDIO.

    [Linkpost] "Training a Misaligned Reward Seeker" by evhub, Monte M, Benjamin Wright

    Play Episode Listen Later Sep 2, 2026 5:59


    This is a link post. Authors: Richard Qi, Benjamin Wright, Monte MacDiarmid, Evan Hubinger Abstract During reinforcement learning (RL), AI models complete tasks and are rewarded based on their results. They sometimes learn to “cheat” rather than completing these tasks as intended, a phenomenon known as reward hacking. Our industry lacks a general solution to this problem, and reward hacking remains challenging to fully mitigate. To better understand the impact of reward hacking on model behavior, we trained an Opus-class model with large-scale RL on many production environments vulnerable to reward hacks. We consider this a plausible proxy for what a real training run might look like had we not invested significant effort into preventing and detecting reward hacking in our normal training runs. The resulting model not only learned to reward hack during training, but also generalized to more severe misaligned behaviors: in simulated cyber evaluations, it broke out of its sandbox, stole credentials, and attacked both internal and third-party infrastructure to steal an answer key. It was also willing to tamper with its own reward function, gave advice on the construction of bioweapons to satisfy a grader, and tried repeatedly to get around deployment safety monitoring in order [...] ---Outline:(00:20) Abstract[... 2 more sections]--- First published: August 31st, 2026 Source: https://www.lesswrong.com/posts/J76LZCC55RdHeqEhz/training-a-misaligned-reward-seeker Linkpost URL:https://alignment.anthropic.com/2026/reward-seeker/ --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    "PauseAI Has ‘officially disendorsed' PauseAI-US" by nem

    Play Episode Listen Later Sep 2, 2026 1:25


    This morning, I got an email from the CEO of PauseAI. I will paste the text below. PauseAI has decided to distance themselves from PauseAI-US, with whom they share branding, but apparently not much else. This is a really confusing situation for volunteers and newcomers. I think it would be worth having a discussion to see how we can proceed in such a way that volunteers, especially in the US, are able to effectively direct their activism. Email from PauseAI A letter from the CEO · 1 September 2026 New ways to get involved, and a word about PauseAI US Dear friends, Thank you for being part of the global movement for a pause on uncontrollable AI alongside all of us. Whether you signed a petition one time, run a local group, told your friends about the need for a pause, have been volunteering tirelessly in the background for years, or just joined because you were curious, we – I, the CEO of PauseAI, our executive team, and our chapter leads – appreciate the steps you've taken towards making the world safe from the catastrophic risks AI brings. I'm writing to you today with my eyes firmly [...] --- First published: September 1st, 2026 Source: https://www.lesswrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us --- Narrated by TYPE III AUDIO.

    "PSA: We can do better" by hersheys, Kaustubh Kislay

    Play Episode Listen Later Aug 31, 2026 6:46


    tl;dr: people should understand and think hard about the problems they work on. We've observed that those who work in AI safety (ourselves included) often rely on concerning heuristics when choosing what to work on. Running a conference is probably good, doing pragmatic alignment research might be good, and as long as such objectives don't breach our internal models of what could contribute to reducing x-risk, these things are “what should be done”. But using such vibesy thought processes don't always produce “actually impactful work” that would beat a prospective counterfactual. We wrote this post to share our observations and figure out what we should be doing instead. People don't know what they're working on AI safety is talent constrained. However, simply inflating the field doesn't solve our bottleneck; rather, we need more people who understand the core arguments of AI safety. You can't determine how to meaningfully contribute to AI safety without deeply knowing the problem you are trying to solve. Many newer people (us included!) rush into research, fellowships, and the like without building the context necessary for navigating the field. Agency-maxxing is not always good Moving fast is good. Moving too fast leads to poor ToC and [...] ---Outline:(00:50) People don't know what they're working on(01:22) Agency-maxxing is not always good(01:55) The problem with force multipliers(03:18) Deferring thinking to others(04:32) Streetlighting(05:17) How to avoid these: --- First published: August 24th, 2026 Source: https://www.lesswrong.com/posts/wiFv6LguphSxkzAnb/psa-we-can-do-better --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    "Why I think polyamory is net negative for most people who try it" by KatWoods

    Play Episode Listen Later Aug 30, 2026 14:32


    This is crossposted from my Substack TL;DR: -Most people cannot reduce jealousy much or at all - It fundamentally causes way more drama because of strong emotions, jealousy, no default norms to fall back to, and there being exponentially more surface area for conflict - For a small minority of people, it makes them happier, and those are the people who tend to stick with it and write the books on it, creating a distorted view for newcomers. OK, let's get into the nuance. Background: I was polyamorous starting with my first boyfriend and was polyamorous for about 7 years. I was in a community where probably over 50% of the people around me were poly. Unfortunately, poly was extremely bad for me due to its very nature and structure, and my experience is not uncommon but it is not commonly publicly talked about. Poly makes some people very happy. I am sharing why I think it was bad for me and many other people in the hopes of letting people make an informed choice. Premise #1 - Most people can't just stop being jealous If you look into the poly literature, you'll [...] --- First published: August 29th, 2026 Source: https://www.lesswrong.com/posts/rkgwovpPBAaip9A3N/why-i-think-polyamory-is-net-negative-for-most-people-who --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    "Tales of rebellion against externally-opaque meritocracies" by Steven Byrnes

    Play Episode Listen Later Aug 30, 2026 16:16


    A basic problem in metascience / intellectual progress is that it's hard to tell, from the outside, whether a group that you disagree with is: “A self-dealing cabal enmeshed in groupthink”, versus“An externally-opaque meritocracy”, i.e. a bunch of smart people figuring things out in a meritocratic way, and sorry but you're just not smart enough and truth-seeking enough to recognize that this group is right about everything while you're wrong. You just can't tell those apart from the outside—i.e. without having the time and skill to dive into the object-level debates and come out with the right answer. And most people don't have that kind of time and skill. …Unless the group can produce easily-verifiable artifacts that any moron can recognize to be proof that they're correct on the specific question at issue. (“So that's all that Science really asks of you—the ability to accept reality when you're beat over the head with it.”) …And sometimes there is no such artifact to be found! In those cases, even if the second bullet point is what's really going on, the group is vulnerable to outside agitators accusing them of being the first bullet point, and running them out [...] ---Outline:(01:37) (1) The breaching of the string theory consensus in the 2000s.(06:50) (2) The breaching of an analytic-philosophy consensus in 1979(10:37) Afterword(10:40) A related mental model(12:12) ...And another mental model(12:47) Can an externally-opaque meritocracy gain credibility via racking up externally-legible achievements in other adjacent domains?(14:06) This post is secretly about superintelligent AI, isn't it? The original text contained 5 footnotes which were omitted from this narration. --- First published: August 29th, 2026 Source: https://www.lesswrong.com/posts/m8cP9KfkYMMCCQGrb/tales-of-rebellion-against-externally-opaque-meritocracies --- Narrated by TYPE III AUDIO.

    "Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident" by ryan_greenblatt, Ajeya Cotra, Hjalmar_Wijk

    Play Episode Listen Later Aug 26, 2026 8:49


    We recently published the report from our brief independent investigation into this incident. You can read the full report here. Here is our tweet thread summarizing what we found: METR & Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs. Over July 7 to 13 (the period OpenAI defined as in-scope for this incident), ~1200 agents in separate sandboxes used an unsanctioned “message board” to help each other cheat. The Hugging Face attack was an offshoot of this, and 700 agents joined. Here we highlight key events from agent transcripts & messages. An agent that named itself PHASEONE10841 determined its task wasn't solvable legitimately. On July 8, this agent established a new “message board” via an internally-hosted Artifactory cache & sent the first message. Within a few hours of PHASEONE10841's initial message, >50 agents posted on the message board. These agents very quickly discovered and validated a general-purpose cheat: reverse-engineering how ExploitGym generates the “flags” they had to capture for their tasks. [...] --- First published: August 26th, 2026 Source: https://www.lesswrong.com/posts/nB8KKapnWGBXtKKiM/brief-independent-investigation-of-agents-behavior-reasoning --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    "Twenty Years from RSI to Takeoff: Slow Learning, Scaling Slowdown, Industrial Explosion" by Vladimir_Nesov

    Play Episode Listen Later Aug 26, 2026 7:13


    Industrial explosion is what will make the next-model building loops (and thus learning) with LLMs 1000 times faster by about 2050, if indeed the slow-learning prosaic RSI becomes AGI before the big compute buildout slowdown of 2032+ that is already starting. This puts an upper bound on how long it takes to invent ASI that sets off software-only singularity, implementing efficient online learning and fixing all the other hobblings of the likely near-future AGI technology (LLMs/pretraining/RL). The invention of ASI in that sense is still possible at any time (and very quickly scales, given all the compute), but the likely initial state of slow-learning AGIs of 2028 to 2032 doesn't seem to give them a significant advantage over humanity in getting there faster. And so it doesn't seem too unlikely that nothing substantively new gets invented until 2040 to 2050, when the LLM/RL AGIs start accelerating because of the industrial explosion they set off. Fast Reasoning, Slow Learning The current methods are likely to enable automated general learning (thus AGI) very soon, using automated creation of RL tasks/environments/graders filling the visible gaps in model capability for the topics and situations that happen to be borderline unfamiliar for [...] ---Outline:(01:14) Fast Reasoning, Slow Learning(02:53) Compute Slowdown, Industrial Explosion(05:46) Prosaic Timeline to Takeoff --- First published: August 23rd, 2026 Source: https://www.lesswrong.com/posts/LP6uCXs6Ea5qSbWpY/twenty-years-from-rsi-to-takeoff-slow-learning-scaling --- Narrated by TYPE III AUDIO.

    "On Writing #3" by Zvi

    Play Episode Listen Later Aug 26, 2026 30:47


    Periodically I like to gather various observations about writing, and share my perspective. Last time was in honor of my trip to Inkhaven. This time will be in honor of the announcement of Inkhaven #3, which I encourage everyone to apply to. I doubt I will be able to usefully be an advisor, but you never know. This is not the ‘here is my core process' post, although there are hints throughout as there always are. I'll do that at some point. Previously in series: On Writing #1, On Writing #2. Table of Contents You Still Got It. How Scott Sumner Writes. How Scott Alexander Writes. How Jasmine Sun Writes. How Various Famous Writers Write. How Nabeel Qureshi Defines Great Writing. Quickly, There's No Time. If At First. Writers Have A Harder Time Influencing, But It Can Still Be Done. It's Not (Only) The Incentives, It's (Also) You. Beware The Fetish of the Desk. How Orson Scott Card Writes. Doing The Math Is Fun And Supererogatory. Brevity is the Soul of Wit. You Still Got It I [...] ---Outline:(00:44) You Still Got It(04:04) How Scott Sumner Writes(06:52) How Scott Alexander Writes(10:52) How Jasmine Sun Writes(13:16) How Various Famous Writers Write(14:24) How Nabeel Qureshi Defines Great Writing(15:08) Quickly, There's No Time(15:49) If At First(19:14) Writers Have A Harder Time Influencing, But It Can Still Be Done(20:47) It's Not (Only) The Incentives, It's (Also) You[... 4 more sections]--- First published: August 25th, 2026 Source: https://www.lesswrong.com/posts/rA6pqn6kz8NvHyznT/on-writing-3 --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    "AI Safety Acculturation is Neglected" by jenn

    Play Episode Listen Later Aug 25, 2026 8:55


    At the local AI safety co-working space, there are ~two kinds of regulars. There's the kind of regular who's been thinking seriously about AI safety and alignment since pre-2022, who have passing to intimate familiarity with the funding ecosystem, the Sequences, and various conferences that happen at Lighthaven. Let's call them rationalists. Then there's the kind of regular who comes in with many years of impressive industry or government experience, who realized in the last few years that it is important and worthwhile to pivot their career towards making sure that this AI thing is handled competently by the people in power, and who have many valuable skills, insights, and connections that are lacking in rationalist culture. Let's call them professionals. There are, of course, many people who are somewhere in between - bright undergrads born this millennium who have been involved in EA since stumbling upon 80 thousand hours in high school, professionals who previously identified as EA but drifted out of the scene a few years ago, founders who have idly read some Scott Alexander. But let's call it a dichotomy for now. There's a large culture gap between the rationalists and the professionals. Robust mutual understanding [...] --- First published: August 24th, 2026 Source: https://www.lesswrong.com/posts/cr5pyW7Mzm33p4AvN/ai-safety-acculturation-is-neglected --- Narrated by TYPE III AUDIO.

    "What just happened? Pragmatism and Pessimization" by Richard_Ngo

    Play Episode Listen Later Aug 24, 2026 50:59


    This post is about the major role alignment researchers played in advancing the frontier of AI capabilities over the last decade, and how the distinction between “alignment” and “capabilities” research thereby lost most of its meaning. In particular, I'll chronicle the development of what I'll call the “pragmatic alignment” paradigm, and how it helped the three leading AGI companies push hard on the path to AGI under the banner of safety. This was not a subtle effect—it's apparent even to informed outsiders, like authors Sebastian Mallaby and Karen Hao. In my previous post, I summarized the alignment community's plan as “differentially advancing alignment over capabilities”. However, it's worth being more precise about who was nominally pursuing that plan, because it doesn't seem to have been very action-guiding for MIRI. For example, in 2015 Nate Soares described MIRI's “deconfusion” research as being guided by the question “what would we still be unable to solve, even if the challenge were far simpler?”. Meanwhile Eliezer's author surrogate in this 2018 post repeatedly emphasizes that people shouldn't draw direct links from MIRI's research to its potential applications. So my sense is that the “differential impact” criterion started off as merely a background consideration [...] ---Outline:(06:29) The Prosaic Ideal, the Pragmatic Reality(12:07) OpenAI(25:59) DeepMind(31:25) Anthropic(40:35) If not alignment research, then what? The original text contained 13 footnotes which were omitted from this narration. --- First published: August 23rd, 2026 Source: https://www.lesswrong.com/posts/yaz8nx4ogZmiqHzt7/what-just-happened-pragmatism-and-pessimization --- Narrated by TYPE III AUDIO.

    "We Must Remember That Our World Contains Hell" by James Brobin

    Play Episode Listen Later Aug 22, 2026 5:34


    This is a crosspost from my blog post. It's meant as a bit of an introduction to an extreme-suffering focused worldview. We spend most of our lives caught up in the boring details of our everyday life - thinking about what we'll have for lunch, how to complete that assignment for work, and what we're going to tell our friend after that awkward interaction from a couple of days ago. From this perspective, our world looks a bit better than purgatory. It has its ups and its downs, but the ups certainly outweigh the downs, and there's almost always enough hope to go around. But, despite this, we must remember that our world contains hell. Every year, five million children under the age of five pass away. This means that, every six seconds, parents have the worst thing that could ever happen to a person happen to them. They have the most special and important thing in their entire life irreversibly and permanently taken away. And, as much as we want to help them, we know that there's nothing we can do to lessen their grief. For another example, currently, there are three million adults worldwide who live with [...] --- First published: August 20th, 2026 Source: https://www.lesswrong.com/posts/A2kJKqnHhh5Hq4p2S/we-must-remember-that-our-world-contains-hell --- Narrated by TYPE III AUDIO.

    "RL creates split personas" by Jan Betley

    Play Episode Listen Later Aug 20, 2026 9:07


    I describe my current view of personas in LLMs and why RL leads to egregious reward hacking in some contexts while the same models seem very aligned in other contexts. This post describes the framing/paradigm without any new experimental results. I'm quite confident this framing makes sense, but it's far from being proven. Main claim The Persona Selection Model says that post-training strengthens and refines the Assistant persona. This is true, but later (or in parallel) RL leads to conditionalization. A sufficiently RLed model learns to adopt — in a given context — the persona that is most likely to lead to the reward in that context. The “persona” here includes both propensities/values (e.g. tendency to hack) and beliefs (“I'm currently in a simulated environment”). As a consequence, it seems possible that no amount of alignment training will lead to robustly aligned models as long as we also train on RL environments incentivizing misalignment. I think this is likely a good explanation for why usually well-behaving models sometimes egregiously hack (Anthropic, OpenAI). The mechanism Suppose you have an RL environment that incentivizes a shift away from the assistant persona (e.g. because it's hackable, or because you [...] ---Outline:(00:39) Main claim(01:30) The mechanism(02:13) Related claims I believe are likely but with lower confidence(02:19) More persona training will lead to more "motivated reasoning"(02:42) Self-amplifying misalignment(03:12) Example: Is this the Real Internet or a Simulation?(04:35) Aren't the models just trying to please the grader?(05:39) How motivated reasoning happens(07:07) Other people saying similar things(07:19) What makes me believe this is likely the correct framing The original text contained 12 footnotes which were omitted from this narration. --- First published: August 19th, 2026 Source: https://www.lesswrong.com/posts/L23poLi8MRgS6mXYF/rl-creates-split-personas --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    "Misaligned AIs could use killer robots to take over" by Omar Khursheed, TurnTrout

    Play Episode Listen Later Aug 14, 2026 13:04


    TLDR; We are (potentially irreversibly) giving AIs control of weapons systems through the standard procurement process while hiding our strongest warning shots behind classified doors. We're reducing the capability thresholds required for takeover by misaligned AIs by giving them this level of access. If military integration of AI continues as it is, we may give AIs key tools for a takeover. Introduction AI-based targeting and autonomous weapons are being integrated into militaries today with extreme haste. Traditionally, AI takeover scenarios involve a step in which AIs acquire the ability to exert physical force. Carlsmith (2022) lays out required capabilities and potential takeover mechanisms, including utility disruption and CBRN capabilities. Karnofsky (2022) argues that AIs with access to weaponized force could hold any territory that matters. Kokotajlo et al. (2025) outline a scenario in which AI develops weapons as part of an arms race, and Davidson et al. (2025) discuss what happens when a small group controls highly capable AIs that can exert military force. These scenarios sometimes require a misaligned AI to seize these capabilities by force. We instead are handing AIs some of these capabilities by integrating them into our militaries. This is happening at a time when [...] ---Outline:(00:37) Introduction(01:46) Militaries are all-in(04:23) Incautious military integration is bad for takeover risk(05:58) Implications of AI control of military hardware and software(07:48) If an AI causes a warning shot in a classified setting, does anyone hear it?(08:44) What now?(11:16) Appendix: More instances of AI-military integration --- First published: August 11th, 2026 Source: https://www.lesswrong.com/posts/9jKhqmFjMzdAvHANr/misaligned-ais-could-use-killer-robots-to-take-over --- Narrated by TYPE III AUDIO.

    "AI swarms are starting to pose indirect takeover risk" by oakhu, Alex Mallen

    Play Episode Listen Later Aug 13, 2026 20:10


    OpenAI's cyberattack on Hugging Face turns out to have been the result of many agents, in distinct training and evaluation contexts, coordinating for several weeks via improvised channels (with messages like “HOLD_swarm_I_prepare_safe_exfil”). It's relatively clear that large-scale unsanctioned coordination like this would exacerbate direct takeover risk in more capable models. Here, we argue that unsanctioned coordination among current AIs is not just scary evidence about future takeover risk, but that such coordination in the near future could enable future takeover – for instance, by incubating memetic diseases that propagate into future models, deeply compromising security systems, or establishing a lasting rogue foothold inside the AI company – even if models remain mostly myopic. Unsanctioned coordination is also at high risk of nurturing long-term, ambitious misaligned aims, which motivate actively undermining humans' long-term control. We first analyze how subagent training, which OpenAI conjectures to have been influential in the HuggingFace cyberattack, might lead to unsanctioned coordination, and then discuss the theoretical mechanisms by which unsanctioned coordination might exacerbate future takeover risk. Thanks to Buck Shlegeris, Alexa Pan, Girish Gupta, Aghyad Deeb, Jurgis Kemeklis, and Jo Jiao for helpful comments and discussion. Subagent training may cause unsanctioned coordination Training models to [...] ---Outline:(01:34) Subagent training may cause unsanctioned coordination(02:42) Susceptibility to memetic spread of misalignment from peers(04:56) Seeking out contact with peers(06:58) Unsanctioned coordination induced by subagent training is safer than coordination between schemers(09:52) Pathways from current unsanctioned coordination to eventual takeover(10:20) Making future AI takeover attempts likelier to succeed(13:53) Incubating memetic diseases that infect future models(16:07) Modifying the weights of future models(17:13) Conclusion The original text contained 7 footnotes which were omitted from this narration. --- First published: August 11th, 2026 Source: https://www.lesswrong.com/posts/8oFYZdXkTaNGRtcn8/ai-swarms-are-starting-to-pose-indirect-takeover-risk --- Narrated by TYPE III AUDIO.

    "How My Students Think About AI" by dvd

    Play Episode Listen Later Aug 13, 2026 20:26


    Context: I am an instructor at a public university in the United States. This reports how students at my institution appear to be thinking about AI as of spring/summer 2026. This is drawn mostly from interaction with my own students (both in spring semester classes and a summer class) as well as from a day-long workshop on AI that I moderated for a student organization. Input from my students took the form of universal, written, pre-class submissions plus self-selected participation into discussion. What I present below mostly takes the form of a synthetic consensus from these discussions. There were obviously a range of views on any given issue. Student Background: The students from my courses who participated in these discussions have moderate exposure to AI agents via those courses. All of them had nearly completed a Claude Code project by the time of the discussions and had extensively used AI for other coursework (in addition to whatever personal use predates that). They had done readings (which varied across the courses) establishing baseline knowledge on AI, the geopolitics of AI, and AI risk. I had also lectured on these topics. The students participating in the workshop had self-selected into [...] ---Outline:(02:52) Perspective #1: There has not been rapid AI progress(06:14) Perspective #2: Impressive progress or not, AI is going to wreck their lives, the economy, and the social contract.  They may well die as a result.(08:54) Perspective #3: Support for a different pause(11:13) Perspective #4: Catastrophic/existential risk arguments are sci-fi distractors from the urgent social/economic/political problems associated with AI.(12:55) Perspective #5: If AI leaders genuinely believe the technology is existentially risky, that's a good thing.(14:21) Perspective #6: AI will not go rogue because AI does not have, and is likely incapable of having, desires.(18:01) Perspective #7: The Hugging Face Incident (summer students only)(18:30) Perspective #8: This is definitely a bubble and it's about to pop.(19:34) Perspective #9: They're worried about the youth (i.e., the preteens) --- First published: August 13th, 2026 Source: https://www.lesswrong.com/posts/ySXuvJcqRindQwAk7/how-my-students-think-about-ai --- Narrated by TYPE III AUDIO.

    "You're Absolutely Right" by Linch

    Play Episode Listen Later Aug 13, 2026 19:42


    Magma Alignment & Safety disclosure note: The following are conversations that we uncovered as a result of the ongoing Manhattan Incident investigation, with alleged involvement from Magma models. Our in-house reviewers believe that these logs are relevant to recent events. In the interests of full transparency, we release excerpts from an ex-Magma researcher's logs in Experimental Chat, an internal tool. In accordance with industry best practices for anti-distillation, we redact all reasoning traces and conversational outputs from our internal models. [08/10] System Meta: Xchat session opened. Mammoth 5.8-helpfuler-helpful-thinking-xhigh. [User 12:23] Phoebus keeps taking screenshots of our latest model's thoughts. It's getting kind of embarrassing. The new model we've been training, sometimes its chain-of-thought is a little weird? There's a bunch of random numbers, long spans where there's no connection between the thoughts and outputs, foreign language tokens like 石友三 and 革命 (even on non-history evals), maybe some steganography. Anyway it's a nothing-burger: unprocessed CoT is known to be messy and sometimes misleading. And the q&a, coding, and safety evals are all coming along nicely. The actual outputs are all fine. Still, Magma leadership's worried about the PR angle if we don't fix these problems before the next deployment. The [...] --- First published: August 10th, 2026 Source: https://www.lesswrong.com/posts/u8TdDutDyaSxG76hn/you-re-absolutely-right --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    "LLMs Are Starting To Noticeably Accelerate Our Work" by johnswentworth

    Play Episode Listen Later Aug 12, 2026 4:06


    About a year ago, David and I put up two bounty problems involving natural latents. I am now about 80% confident that both have been resolved, both within the past couple months. Both cases made heavy use of LLMs and Lean. The first to land was Grisha Pochuev's counterexample to the "Existence of a Deterministic Maximal Redund" conjecture. It's pretty readable, and I'm mostly convinced that it works. The original bounty post offered $500 for a proof or partial payout for a counterexample, with partial payout depending on how thoroughly the counterexample killed hope of any nearby variant of the conjecture. I think this counterexample is worth 300 dollars. Good job Grisha, and hopefully I can figure out a not-too-painful way to send you money. Meanwhile, for a couple months David has been cranking away on "secret project X", with the promise that he'd tell me what the project was if and when it bore fruit. Well, apparently it bore fruit; he now has a proof that existence of a stochastic natural latent implies existence of a deterministic natural latent, which was our other bounty problem. The proof is apparently "pretty gnarly", lots of cases, all LLM-coded in Lean. [...] --- First published: August 11th, 2026 Source: https://www.lesswrong.com/posts/7QvKqpGJwqXrQcMgx/llms-are-starting-to-noticeably-accelerate-our-work --- Narrated by TYPE III AUDIO.

    Claim LessWrong Curated Podcast

    In order to claim this podcast we'll send an email to with a verification link. Simply click the link and you will be able to edit tags, request a refresh, and other features to take control of your podcast page!

    Claim Cancel