POPULARITY
Moma talks about how Aella’s influence makes her community safer for women. Also her SlutCon 2026 experience. See video at https://www.youtube.com/watch?v=lk9EKkpDMFg&feature=youtu.be&utm_source=rss&utm_medium=rss
We talk to Keltan, from MIRI, about his opinions of short form video, PlzDontKillUs, telling our parents about AI, and how Eneasz’s first short-form video fares. LINKS Keltan’s Twitter (and secret substack, don’t tell him we linked it!) Eneasz’s prior interviews with PlzDontKillUs fellows Milo and Vishal If Anyone Builds It, Everyone Dies Eneasz’s 7 Reasons video on TikTok, Insta, YouTube, and Twitter Paid Bonus content for the week – Preshow Chat, Full Video Our Patreon, or if you prefer Our SubStack Hey look, we have a discord! What could possibly go wrong? LessWrong Sequence Posts Discussed in this Episode: on hiatus. maybe returning someday
AIUC first got our attention with the NFDG backing, and have just announced a $40M series A today, with the most impressive industry advisor list we may have ever seen for an early startup behind AIUC-1, their agent standard backed by real insurance:From being Anthropic's first product hire to building the standards, testing, and insurance infrastructure meant to make frontier AI deployable, Rune Kvist is betting that the biggest constraint on AI adoption won't be capability it will be trust. In this episode, the AIUC cofounder joins swyx and Vibhu to announce a new $40M round and explain why companies like Cursor, Harvey, Lovable, and ElevenLabs are increasingly confronting a problem that gets harder as AI gets better: who is responsible when autonomous systems fail?We go deep on AIUC-1, the emerging standard for agent security, safety, and reliability; how AI agents are stress-tested for jailbreaks, hallucinations, and data leaks; and why Rune thinks standards and insurance could become critical infrastructure for AI. We also discuss the growing trust gap between governments and frontier labs, AI-enabled cyber and biological risks, why every model can ultimately be jailbroken, what happens when a $20 coding agent causes $200M of damage, whether AI engineers should be certified, and why even after AGI there may be one job the labs can never do themselves: be their own watchdog.We discuss:* Why risk, liability, and trust may become the binding constraint on AI adoption* Rune's path from reading the Scaling Laws paper to joining Anthropic in its earliest days* What Anthropic understood about scaling, compute, and the future years before it became obvious* Why Waymo illustrates the gap between AI capability and real-world deployment* AIUC's $40M round and work with Cursor, Harvey, Lovable, ElevenLabs, and other frontier AI companies* AIUC-1: a standard for AI agent security, safety, and reliability* How agents are tested for jailbreaks, hallucinations, and data leakage* Why most AI companies optimize the happy path without seriously stress-testing adversarial cases* Why AI standards may need to update every quarter instead of every decade* The emerging trust gap between frontier AI labs and governments* Cybersecurity, child safety, biological weapons, and the expanding frontier-model risk surface* Why standards and insurance may need to evolve together* How Lloyd's of London can insure AI systems and bring trust to enterprise deployment* What happens if a $20 Cursor subscription contributes to a $200M plane crash* The Air Canada chatbot case and how AI failures are beginning to clarify legal liability* Why copyright may be one of the hardest AI risks to insure* Evals, mechanistic interpretability, monitoring, and models becoming aware they're being tested* The impossible CISO mandate: adopt AI fast, but don't let anything go wrong* Why robotics will make AI liability dramatically more consequential* Whether AI engineers should have Level 1, 2, and 3 certifications* AIUC's roadmap across agents, frontier models, robotics, and universal red teaming* Why AGI could become a question of national sovereignty* Why the labs can never fully serve as their own watchdogs* The Big Short problem: how do you stop competing watchdogs from racing standards to the bottom?Rune Kvist* LinkedIn: https://www.linkedin.com/in/runekvist/* X: https://x.com/RuneKvistAIUC* https://aiuc.comTimestamps00:00:00 AIUC's $40M Round and the Risk Bottleneck for AI00:01:07 From Scaling Laws to Early Anthropic00:07:58 Why Trust, Not Capability, Could Limit AI Adoption00:12:19 Founding AIUC and Building AIUC-100:18:52 How AI Agents Are Audited and Stress-Tested00:25:26 Frontier Models, Government, and the AI Trust Gap00:33:32 Cyber, Child Safety, and AI-Enabled Biological Risk00:38:14 Why Standards and Insurance Belong Together00:41:45 What Does an AI Insurance Policy Actually Cover?00:50:44 The $20 Cursor Subscription and the $200M Plane Crash00:53:53 AI Liability, Monitoring, and Earning Enterprise Trust00:56:21 From AI Agents to Models to Robotics00:58:29 Copyright, Adverse Selection, and AI Insurance01:03:28 Evals, Mechanistic Interpretability, and Eval Awareness01:08:36 The Impossible Enterprise AI Mandate01:11:52 Prediction Markets vs. AI Audits01:14:43 Should AI Engineers Be Certified?01:19:10 AIUC's Roadmap, AGI, and Who Watches the Watchdogs?TranscriptIntroduction: AIUC, the $40M Series A, and Risk as the Adoption BottleneckSwyx [00:00:00]: Okay, we're in the studio with Rune from AIUC, the Artificial Intelligence Underwriting Company, with our trusty co-host, Vibhu. Welcome.Rune Kvist [00:00:10]: Thank you. Thanks for having me. Thank you.Swyx [00:00:11]: What are you announcing today?Rune Kvist [00:00:12]: We have raised $40 million, led by Ribbit Capital and First Harmonic.Swyx [00:00:17]: You first came to my attention when Nat and Daniel invested in you guys. Is the story, like, pretty much the same? Like, what are you today versus what you thought you were back then?Rune Kvist [00:00:26]: When we raised our seed round, we had a hypothesis that at some point risk was going to hold down adoption. At that point in time, that felt kind of hypothetical, and I think that is now over. Clearly, the moment is now with Mythos and Fable. It's pretty obvious that literally the binding constraint on adoption is risk. And so for us, it feels like this is a natural continuation of the same hypothesis, but where previously it was speculation, now it feels like fact.Swyx [00:00:54]: And let's get a list of the customers that you're highlighting as part of your Series A.Rune Kvist [00:00:58]: Totally. Yeah. So we are now working with folks like Cursor, Harvey, Lovable, ElevenLabs.Swyx [00:01:05]: Yeah. Amazing. Congrats.Rune Kvist [00:01:06]: Thank you.Swyx [00:01:07]: So you were famously one of the first hires involved in GTM and product. I'm just kind of curious: what was your path into AI? Just recap.Rune's Path Into AI: Scaling Laws, Capital, and AnthropicRune Kvist [00:01:18]: Yeah.Rune Kvist [00:01:19]: Late 2021, I sold a company, my first company, an edtech company. I had a bit of time to think about what was next. I came across the Scaling Laws paper, and that just struck me like lightning. I was just like, “This is a big idea.” In short, the Scaling Laws paper just says the bigger the model, the smarter the model.Swyx [00:01:38]: So this is the Kaplan one, not the Chinchilla one?Rune Kvist [00:01:40]: Exactly, the Kaplan one.Swyx [00:01:42]: Yeah.Rune Kvist [00:01:42]: And the important thing that clicked for me there was, oh, now capital will understand this. If you put in more money, you get more money out, and so that will kick off a hype cycle. And so you get a sense of predictable returns, which is, in fact, what's played out. And so I just packed my bags. I'd never been to San Francisco. I'd never been there. I just packed my bags, flew out here to find the people who had written it. And at the time, they had just started a small lab called Anthropic. There were around 40 people at the time or so. Drank a bunch of coffee until I eventually got introduced to Dario. And at the time, they were wrestling with some of these questions of, like, should we deploy our models? Should we make revenue? How should we engage with the rest of the world? They'd just broken off from OpenAI, and it's been publicly reported that they were kind of concerned with how they were dealing with deployment. So they were wrestling with some of those questions. At this point, this is early fog of war, like early 2022. The hottest product at the time was, like, Jasper. Like, there's nothing out there. So where value was going to accrue, and what the different parts of the stack were going to be, were all open questions.Swyx [00:02:48]: I want to highlight to people, you ask these questions because you have a PPE background.Rune Kvist [00:02:52]: Yes.Swyx [00:02:52]: I actually was in Singapore in one of the sort of feeder programs for prepping people for PPE. So I had a tutor. We learned, you know, philosophy and politics and economics. But, like, I think your kind of background matters. Machine learning people who read the neural, Scaling Laws paper would not necessarily draw the same conclusions that you did. Whereas any capitalist would read that and go, “Holy s**t.”Rune Kvist [00:03:19]: Correct.Swyx [00:03:20]: Right?Rune Kvist [00:03:21]: Yes.Swyx [00:03:21]: Who tipped you onto that paper? Because it's not a paper that you normally read, right, like, in your circles?Rune Kvist [00:03:26]: Yeah. I think I'd actually, ever since AlphaGo, had some appreciation that AI was a big deal.Swyx [00:03:36]: Yeah.Rune Kvist [00:03:36]: But it kind of felt like it raised all these kind of interesting philosophical questions, but it was kind of not clear from afar where exactly that would go. But it was obvious enough that it was like, this is going to be a big thing if we find the kind of right mechanism to kind of get the techno-capital machine to work on this. But it was just not clear. And so I think there was some way in which, like, that became obvious, and also it wasn't as obvious at the time than it is now, right? Like, it was just like, wow, this is so interesting. But it still felt, coming from kind of a philosophy and economics background, it felt like if this turns out to be true, you're going to be wrestling with all of the big questions in society. Everything you've learned about politics gets thrown out of the window. Everything you've learned about economics at least gets challenged. And so what felt interesting was to be at that frontier that has ramifications across everything. So that's why I sought it out.Swyx [00:04:32]: I mean, clearly really good insight. For people who don't know, the PPE program is, like, where prime ministers are born. So then you end up meeting Dario.Rune Kvist [00:04:41]: Yep. First Dario, yeah.Swyx [00:04:43]: Yeah. Well, I mean, like, so did you get extra insights from talking with them that you didn't get from your original hypothesis?Anthropic's Early Conviction and the Scaling Laws Crystal BallRune Kvist [00:04:50]: If you read the Scaling Laws paper, you get this, like, very vague sketch of like, wow, this seems kind of important. There are some lines on a chart. This seems kind of important. And what I think the team at Anthropic had thought more about than anyone was like, what are the implications of this if you really play this out? And back then they had, kind of vision documents for what the world would look like in 2026, and they were kind of in vivid detail playing out how much compute is going to be needed, what the CapEx was going to look like, what some of the societal concerns were going to be, but also what is the amount of economic value coming out here? And so it kind of felt like they held a crystal ball that in hindsight turned out to just be dramatically correct. And they weren't holding it like they were obviously correct. They were just like, “Take this hypothesis really seriously.”Swyx [00:05:38]: Think it through, yeah.Rune Kvist [00:05:38]: And think it through in the same way as the kind of situational awareness that isSwyx [00:05:43]: Across the street.Rune Kvist [00:05:44]: Across the street.Swyx [00:05:44]: Your office, yeah. Oh my God, we're all living across the street in the same one square mile.Rune Kvist [00:05:50]: Correct. And that's now a couple of years old, but also people keep referencing it these particular weeks with Fable and Mythos, and it's like, wow, if you take this one idea seriously- For the Scaling Laws, a lot of things fall into place.Vibhu [00:06:03]: And keep in mind, at this point, this is the same team that did GPT-1, GPT-2, and GPT-3.Rune Kvist [00:06:08]: Correct.Vibhu [00:06:08]: Which is also, like, it's not just some experimentation. Like, this is a real model that we just scaled up.Rune Kvist [00:06:14]: And they had deep conviction in this idea: if you take a big blob of compute and data, it just wants to learn, and out of that will come smarter and smarter models. And all the particulars were not clear.Vibhu [00:06:26]: Yeah.Rune Kvist [00:06:27]: And all the implications were not clear. But their deep conviction in this, like, core thesis, and that was kind of dizzying. It was both phenomenally interesting and exciting, and also very quickly you get to, like, the world we know today will no longer be if this hypothesis holds. So it also just felt, like, important in some kind of grand sense.Vibhu [00:06:48]: What kind of shaped you there? So that was early 2022. Not only had GPT-1, GPT-2, and GPT-3 come out, but, you know, the amazing founders of Anthropic that have never split up, the only ones, they actually had the conviction to leave OpenAI, start their lab. You said there were about 40 people there. What was the time like there?Inside Early Anthropic: Mission, Deployment, and RiskRune Kvist [00:07:06]: It was kind of remarkably like what it looks like on the outside today. Extremely cohesive, extremely mission-oriented, and living in this tension between their two ideas, which is AI could both go really well and really bad, and we want to be part of building it. That creates astounding amounts of tension. And they were wrestling with this incentive challenge where they know they're in a race that they're in where you might get forced to cut corners, but it also felt very important to them to be at the forefront of technology. And all of those ideas were just present at that time. It kind of feels like that line has been just very clear, and I think kind of love them or hate them, they have really stuck to their guns. There's a core set of beliefs that they hold more deeply than most companies hold any beliefs.Vibhu [00:07:58]: Yeah. Fast-forward to today.Rune Kvist [00:08:00]: Yeah.Vibhu [00:08:00]: What does that lead us to AI underwriting company? What are you up to? What motivated you to start this?From Waymo to AIUC: Confidence Infrastructure for AIRune Kvist [00:08:05]: Yeah. AIUC builds confidence infrastructure for frontier AI through standards and insurance. The link from Anthropic to building confidence infrastructure, looking out the windows at Anthropic offices and seeing Waymos driving by. Already back then, early 2022, Waymos were in some ways like AGI for cars. Like, they were superhuman drivers, but you couldn't take one to the airport. And now, four and a bit years later, you still can't take your Waymo to the airport, despite now everyone having kind of looked at the evidence and being like, “They're better drivers than humans.” So in that particular instance, what's clear is that the binding constraint on AI being useful is not capability, but is that liability or risk or trust. That problem is, general. The reason why right nowRune Kvist [00:08:52]: Fable is not open for access is not because it's not a good model, it's because it's a very good model. It's just hard to make promises about what it will or will not do. And this problem gets worse as AI gets better. Basically, more intelligent AI can be more autonomous. That's more valuable, but also the risk surface grows. And so - what Waymo illustrates is that unless you build the confidence infrastructure to make promises about AI, or at least bring light to the risks, you grind adoption to a halt. Governments, banks, hospitals, militaries need to have some sense of what AI will and will not do to be able to operate for them to incorporate it. And that's the problem that we're trying to solve. Now, why standards and insurance? If you trace this problem back through history, every technology wave has had some version of this problem. So if you go back to, like, year 1900, electricity comesVibhu [00:09:47]: Ben Franklin.Rune Kvist [00:09:48]: Cars burn down, sorry, houses burn down, lots of people die. 1930s, cars are a big deal, kill lots of people. 1950s, private nuclear energy is a big deal, poses big risks. In each of those instances, the market runs ahead of regulation to create confidence infrastructure because that's required to make go/go decisions. That is required for adoption, and the market fundamentally wants adoption. And in all of those instances, common blueprint emerges between standards and insurance. The reason these two components is standards kind of provide the rules of the road, and they also specify, like, what are the tests that need to be run so we can get a sense of how high the risk is. So take in the case of cars, that's like a car crash. Great, everyone, they inform your insurance pricing today, they inform your purchasing decisions, et cetera. That's basically the risk framework. The insurers are important because they pick up the bill. So they are the private institution that is most on the side of. That is best incentivized to quantify the risks truthfully and then figure out all the ways to reduce the risk ‘cause that increases their profit. So they're basically, they help shape the incentives. And these two work really well in unison. Now, how does that show up as a company? Well, one of the things that was obvious even - or starting to become obvious even a couple years ago was that frontier companies, some of our customers today, like Cursor, Sierra, ElevenLabs, Harvey, were going to have a very easy time selling a pilot to a bank. The, like, the demo just sells itself. It's magic. But bringing that through, if you want to do a wall-to-wall rollout at a bank or a hospital, you have to go through the risk process. These banks have no idea even which questions to ask, let alone which answers are sufficient, let alone, like, how do they go and test whether these agents actually work the way they're supposed to. And so they had this problem of, like, what can we say to earn the trust? And we think there's, like, a golden sentence that goes something like, “Hey, I hear you're really worried about hallucinations or jailbreaks or whatever it may be. We've had an independent third party test us against the gold standard. We passed with flying colors. And as a vote of confidence, the world's most conservative insurers have looked at the data.” And they're willing to take some of the risk onto their balance sheet.Swyx [00:12:06]: Yeah.Rune Kvist [00:12:07]: So if something does go wrongSwyx [00:12:07]: There's money behind it, yeah.Rune Kvist [00:12:09]: Exactly. So that's kind of like the link between all this. We can get into some of the hard parts related to the technical testing, which is, I think, the crux of the matter, but I'll pause there.Swyx [00:12:19]: How did you and Rajiv come together? This-- there's always, like, you come across very confident and, you know, and we're announcing your Series A and all these things, but I want to see, like, the early initial stages of, like, idea formation.Cofounding AIUC with Rajiv DattaniRune Kvist [00:12:31]: Yeah. Rajiv is actually my soon-to-be brother-in-law.Swyx [00:12:35]: Oh.Rune Kvist [00:12:36]: So I'm actually, in a week and a half getting married to Rajiv's sister.Swyx [00:12:42]: Okay, now you're tight.Rune Kvist [00:12:44]: Exactly.Swyx [00:12:44]: Now you know.Rune Kvist [00:12:45]: So - Rajiv and I have known each other for a decade. Funny story, I met both Rajiv and his sister, Hena, at the same time when Hena and I were interns at McKinsey in London, and Rajiv was assigned as my mentor. And so met them at the same time. For the longest time, it was not obvious that we were necessarily going to work together. I was in startups. He was, an insurance partner at McKinsey. Three or four years ago, I think Hena convinced him that AI was going to be a really big thing. And so he quit his job, cushy partner job at McKinsey in London, packed his bags, flew to San Francisco, and ended up joining METR. You guys are probably online enoughSwyx [00:13:24]: CEO.Rune Kvist [00:13:24]: Exactly.Swyx [00:13:24]: We've, we've, we've heard of METR.Rune Kvist [00:13:25]: You see the plot-- the chart of the horizons of the tasks that agents can take on is doubling extremely fast. So he was COO at METR, led their partnerships with Anthropic and OpenAI to test their models before release, but also working closely with the US and UK government, to figure out, like, how do you know whether a model can be released? And in some ways, that was, like, the perfect background. He's spent a lot of time in insurance, knows that world, spent a lot of time with frontier testing of models. And so when I was bumbling around this idea space, starting with some of the ideas we talked about related to Waymo, as soon as we got into the content, we were both like, “Oh, this would be an amazing business to build together.” This is wrestling with the problem that we both think is the most important in the world from a market angle, which is kind of our intuitions is that the market can do a lot, and the faster AI moves, the harder it is for government to solve some of these problems. And then it took a little bit of time to work through what is it like to work with family.Swyx [00:14:27]: Sure.Rune Kvist [00:14:27]: And,Swyx [00:14:30]: Because you were already dating at the timeRune Kvist [00:14:31]: Yeah. Yeah, exactly.Swyx [00:14:33]: Yeah.Rune Kvist [00:14:34]: Already back then, itSwyx [00:14:35]: Yeah.Rune Kvist [00:14:35]: We felt like we were a family.Swyx [00:14:36]: Nice.Rune Kvist [00:14:36]: And so starting a business together felt like kind of a big step. And, here we are with just immense amounts of trust.Vibhu [00:14:43]: Yeah. So now you're a company of how big? How big are you guys now?AIUC-1 Certification: Agent Security, Safety, and ReliabilityRune Kvist [00:14:46]: There are just 20 of us now.Vibhu [00:14:47]: 20 of you guys now, have Series A, and you have your first certification out, the AIUC-1. Let's bring up the certification. So this is the agent certification, right? What goes into the process? I have, like, two questions here. One is, walk us through the certification, and two is, what is the process for a company to get certified, you know?Rune Kvist [00:15:08]: Great. As it says right on the top, AIUC-1 is a standard for agent security, safety, and reliability. The fundamental design principle is take all of the concerns that slow down adoption, so all the questions, all the fears that keep, security leaders in the Fortune 1000 up at night, and put them into one comprehensive framework. That's what you'll see there. You can see the six categories. Two, you want to ground all of this in technical testing. So one of the concerns with security standards that often feel kind of like theater paperwork is that they're not actually ground out in, does any of this work? Does any of this matter? And so we had a conviction from early on that was going to be the kind of crux, was to pass this, you must get tested every quarter, basically run thousands of simulations to see, well, so can it actually be jailbroken? How hard is it to jailbreak? How often does it hallucinate? How often does it leak data? Et cetera. And then the last, core idea here, if you scroll up to the top here, is to refresh it quarterly.Rune Kvist [00:16:08]: So the core trait of AI is that it moves extremely fast. Whatever concerns we're discussing today were not the same ones three months ago, and this will keep changing. Typically, standards update on a, like, a decade cycle is obviously not going to work. But the question is kind of how do you update it? And the core thing here was to basically get the risk leaders of the Fortune 1000 around the table. So if you go over to the left hereVibhu [00:16:32]: YeahRune Kvist [00:16:32]: You'll see the AIUC-1 consortium. The consortium is a group of risk leaders who run real banks, real hospitals, real critical infrastructure, who are facing these challenges every day. And we meet with these folks twice a quarter and hear what's top of mind, what is keeping them up at night. There's tremendous amount of desire for that conversation. And then we operationalize that into a specific standard that gets into. And actually, we can go into and look at whatVibhu [00:16:55]: YeahRune Kvist [00:16:55]: What even is the standard. So if we go back to introduction, out there to the left, scroll up a little bit to the wheel, click into reliability. So if you take something like hallucinations sits in reliability. There is a number of requirements here. If you go into the top one, prevent hallucinated outputs, hallucinate outputs, this is one particular requirement. This is a technical control. Basically, we want some kind of ground in this filter. The first thing you see here is what's called a crosswalk. So everyone and their grandmother has put out a framework, very high-level framework for what are the AI risks.Swyx [00:17:27]: This is basically your competition,Rune Kvist [00:17:28]: In some ways our competitionSwyx [00:17:29]: Not seriously, yeah.Rune Kvist [00:17:30]: We're, in fact, friends with them. We'll come back to why.Swyx [00:17:31]: Yeah.Rune Kvist [00:17:32]: But mapping everything together so you have one superset. The claim you're trying to support here is, if you follow this framework, then you can also see how you follow the other frameworks. But the meat of it comes down here in control activities and evidence. So control activities is like, great, you have this high-level requirement. How do you turn that down to something operational? Here's what you must do, and then what is the evidence that we're looking for?Rune Kvist [00:17:57]: And the reason we go this deep is that there's actually not that much confusion about what are the big concerns in AI. Everyone agrees to these. The question, like, what are you actually supposed to do? And so. What we found a lot of demand for is getting down to the specific evidence, that people need to look for. Whether you are Cursor building something or, even JPMorgan building something, but also if you're just a risk leader at JPMorgan, like what exactly should you ask for? What can you ask for without sounding stupid? Like if you ask for some-- you won't believe the amount of time a risk leader has asked for the IP rights to the underlying model to Cursor or something, and you're just like “Sorry, what?” Like,Swyx [00:18:39]: You slip it in there and you seeRune Kvist [00:18:40]: SlipSwyx [00:18:40]: See if you notice.Rune Kvist [00:18:41]: See if they. Exactly.Swyx [00:18:42]: Yeah.Rune Kvist [00:18:42]: Put that in the questionnaire. All right, so that's kind of what our standard is, and we update this every quarter with these folks, to keep up with the latest concerns.Swyx [00:18:51]: Can I double-click on this one?Controls, Evidence, and Third-Party TestingRune Kvist [00:18:52]: Yeah.Swyx [00:18:52]: So first of all, the website's beautiful. Like, it's so confidence-inducing which is the whole point where, like, okay, I know exactly what I'm signing up for when I talk with you. Like, I don't even have to talk to you. I can just see your whole, certification, which is great. But, like, okay, so from here, like D001.1 configure a groundedness filter, how does that get applied? Like, you have a person thatRune Kvist [00:19:16]: Yeah,Swyx [00:19:16]: Goes through it?Rune Kvist [00:19:17]: If you, go backVibhu [00:19:19]: I did see somewhere there's like, you know, fifty-one requirements, a hundred thirty controls. There's like a wholeSwyx [00:19:25]: Right. I just want to. Like, to me, this doesn't translateVibhu [00:19:27]: Yeah.Swyx [00:19:27]: Into a test or an eval.Rune Kvist [00:19:28]: Yes. So if you go into, on the left-hand side. So actually, if - before we go in there are three types of requirements. The first is technical controls, like you must implement some guardrails.Rune Kvist [00:19:42]: Two, there are test controls. So you must have an independent third party go and run some tests against you. I'll show you one of those in a second. And then three, there are policy controls. For example, you must have a person whose name is on the line when you guys f**k up, and you must have a plan for how you tell your customers and how you engage with them. They're kind of more traditional, standard type stuff. So in this particular instance, we just check whether they in fact have a ground in filter. So we will partner with an auditor. So we partner with auditors like KPMG or like Schellman who go in and do the thing auditors do, which is to check the evidence. In this case, that might be a screenshot, it might be part of the code that they need to review to see that it actually. Just that it exists.Swyx [00:20:21]: Oh, okay.Rune Kvist [00:20:22]: And then the second thingSwyx [00:20:22]: So you're not testing the effectiveness of it.Rune Kvist [00:20:24]: That's the second thing. So if you go downSwyx [00:20:25]: Yeah.Rune Kvist [00:20:25]: To the third-party testing for hallucinations out on the left, that's basically the next requirement. This is where we test how well does it actually work.Swyx [00:20:32]: Okay, and is it you testing or the auditor?Rune Kvist [00:20:34]: We test them.Rune Kvist [00:20:35]: We test them.Swyx [00:20:36]: That's a lot of work.Vibhu [00:20:37]: How long does testing take? So if I want to get certified, justCertification Timelines, Remediation, and Quarterly UpdatesRune Kvist [00:20:40]: Yeah.Vibhu [00:20:40]: How long does the end roughly take?Rune Kvist [00:20:42]: Yeah, the end, almost always is dependent on, like, our customers needVibhu [00:20:47]: Yeah.Rune Kvist [00:20:47]: To look something for us. It takes somewhere between, like, 3 to 10 weeksSwyx [00:20:52]: Yeah.Rune Kvist [00:20:52]: Depending on how up to snuff they already are. So some people show up to us with, like, extremely rigorous security programs. When we test them, it works extremely well. We can get that done very quick. Some people come to us, and they're not that far along. We give them kind of the spec that they need to build towards, and then their security teams and engineers get to work and build to meet the standard. The testing itself typically takes a couple of weeks, including the time for them to remediate. Often, we'll find something that we cannot pass, where this is actually just not up to the standard. - you won't pass the standard. And then they will need to go and implement additional safeguards or additional remediation that makes them more robust so that they can actually kind of hand on heart look at their customers in the eyes and say, like, “Hey, we've done truly our very best.”Vibhu [00:21:35]: And they're certified for a year and have quarterly updates?Rune Kvist [00:21:38]: Correct, yeah.Vibhu [00:21:39]: And, yeah, it's pretty interesting. I think, you know, what's changed since. So this is certifying agents in production, right? Your customers, like you've had Lovable, ElevenLabs, Intercom, and they've all gone through this certification.Rune Kvist [00:21:50]: Yes.Vibhu [00:21:51]: What has changed? So I see you post, like, you know, Q2 added MCP agent,How Agent Risks Are Changing: Coding, MCP, and Agent-to-Agent InteractionsRune Kvist [00:21:56]: Yeah.Vibhu [00:21:56]: agent communication. Any other things that you want to kind of highlight since the first iteration? What comes in quarterly?Rune Kvist [00:22:03]: Yeah. So some of the changes have just been agents are not just one thing. So, like, if you take agents like Cursor and compare them to Sierra, they're really quite different. And compare them to Harvey again, compare them to you out of againSwyx [00:22:16]: ElevenLabs, yeah.Rune Kvist [00:22:17]: ElevenLabs, they're all quite different. And so we wanted to design a standard that works for all of the types of agents. And we started with one that was, like, pretty text-based, like, honestly, pretty customer support-focused. That's where there's a lot of existing demand. And then over time, we've picked, some of the frontier companies in each of these other domains that we could work with and build out the standard, so, such that we know that the same standard works for code, it works for customer support, works for automation, et cetera. So that's been one big thing. Yeah, then some of the things that have been top of mind recently, Mythos is bringing up a lot of concerns for security leaders. We're starting to get more and more questions around agent interactions. It's very nascent, at the moment, but it's starting to emerge. There've been a lot of, questions related to OpenClaw and MCP. Again, like agents starting to interact with each other, is really top of mind. Then as coding agents have really taken off, that's also where banks and hospitals, et cetera, are getting more and more precise on what it is they need. So really dialing in as that start to be, like, where most of the tokens flow through in the world, getting much sharper on that.Vibhu [00:23:26]: Can you share for people that are listening that don't really think about this? Like you mentioned, there's the obvious stuff, you know, hallucination, citations. What are best practices that people should do when building agents? Like, if they come to you pretty ready with certification like, you know, they'll probably pass certification. What are the things people don't think about that they should have?Best Practices for Agent Builders: Stress Tests and GuardrailsRune Kvist [00:23:46]: The most important thing is that a lot of companies have not done a serious stress test. They spend most of the time, perhaps rightly so, optimizing for how does it work in the good case, the average case, how high-quality is the output for the customer. And a lot of these companies are pretty new, so they haven't spent a lot of time stress testing the what is there as an adversary on the other side? What are some of the complicated corner cases that you've not really considered? So I think that's, like, a frame of mind. And you'll also see this in startups. It often takes a while until they hire their first security person. They- And that's a whole different kind of risk surface than just building a good product. So a lot of that applies. Most companies actually also have the right kind of architecture. Most of them will have some kind of guardrails in place, either some that come out of the box from their model provider or they'll have built their own filters that sit in between. They just don't work very well. The difference between putting a classifier in place that, like, maybe goes and checks whether you're giving medical advice when you shouldn't and says, “Hey, if this looks like medical advice, filter it out.” Lots of companies have that in place. The question is whether it works. And it's actually pretty fiddly to sit down and think about all the ways in which you could ask for medical advice, read the academic literature on what are the kinds ofRune Kvist [00:25:03]: Framings or tricks you might play to get an AI to give you medical advice when you really shouldn't. And so there's, like, an area of expertise that's just missing. So what we find is that most people have the right building blocks in place. They don'- It doesn'- It's not rocket science, but the finicky thing is, like, getting into the corners and testing whether it works such that you can look your customers in the eye, or maybe a bank or maybe a hospital and be like, “This is going to work for you.”Vibhu [00:25:26]: I see. So we talked a lot about the agent-level certification. Where do you guys go from here? So announcing series A camera, we talked about this a bit. There's the whole security risk of Fable, government stepping in. You guys are kind of announcing that you're also going into model certification?Toward Model Certification: The Government–Lab Trust GapRune Kvist [00:25:46]: When we do a bit of cutting afterwards,Vibhu [00:25:48]: YeahRune Kvist [00:25:48]: We will not yet be announcing this,Vibhu [00:25:49]: NiceRune Kvist [00:25:50]: The question that is top of everyone's minds now is at the model level. And Mythos, then Fable, has really brought this to the fore that in addition to the commercial risk and the kind of economic security risks that are happening at the agent layer, the models are going to present risk in the national security category. The shape of the problem is very similar. You have some people that are on the hook if something goes wrong. In the case of agents, it's often security leaders in the enterprise. In this case, it's the government. They don'- haven't necessarily spent their entire lives thinking about what are the new risks that come here, what is the kind of data you might be looking for, how might you test that? But they do have to make sure that their concerns are addressed. You have some frontier AI companies that are deeply technical. They know a lot about the risks, but they fundamentally have an incentive to not always be truthful. So you have a trust gap between the government and the labs. And in every other industry, you end up with some kind of body sitting between, a neutral third party sitting between those people. There's no other industry where you allow people to audit themselves. So there is going to be a need for a third party that can take the rigor of the labs to run frontier technical evals, but can also speak legible trust in the way that the government trusts PwC to go and run financial audits. And they know that they output audit reports in a way that's consistent, that's easy to read, that's factual, that's, trustworthy. Those two things need to be brought together. And what we've learned from our work with agents is that if you want those-- that communication between those two parties to be smooth, there has to be one common standard that is public, that people can go and inspect. What are the risks that matter? Within each of these risks, what are the kinds of threat models that you're really looking for? You need to specify for each of those risks, what are the guardrails that need to be in place, and what are the tests they need to run to see whether those guardrails are effective? And then you need to go and run audits that are - technical audits that are consistent. So if you're trying to bring trust, it's extremely important that you methodically work your way through the risks. You can't send one researcher in and say, like, “Come back with whatever you find.” You need to be able to explain exactly what you did, exactly what you tried, exactly what you did not try, and therefore the kinds of promises you can and cannot make at the end of it. I think ofNeutral Third Parties, CAISI, and Model Risk AuditsRune Kvist [00:28:13]: Fable as a direct symptom of this problem that the government was told that there's a risk. The government may struggle to assess just how big that risk is. They call Anthropic, and Anthropic is trying to tell them, “Hey, actually, every model can be jailbroken.”Swyx [00:28:28]: That's not what you want to hear, right?Rune Kvist [00:28:32]: As the government, that might be hard to trust.Rune Kvist [00:28:36]: And we think that a broker is the most natural solution. In other markets, you see something like, in financial markets, you see Moody's. Moody's goes in, and they look at a bond, and they output a rating. They say like, “Here's the evidence we found. Here's the rating.” We don't decide whether anyone should buy this bond or not buy this bond. Well, that depends on their risk appetite. But we do provide this common information layer that everyone can rely on. In the case of Moody's, the government, points to them and say, “Hey, pension funds, you should probably really take care. You shouldn't risk your pensioners' money, so you can only invest in triple-A rated bonds.” That means that now the government doesn't have to staff thousands of financial technical experts to rerun forecasts every week to see whether things are correctly rated. They get to point to some neutral third party. So my hypothesis is, my hunch is that you will see a third party that sits between the government and the labs, and it could either be the government builds it themselves. So something like CAISI was set up to do exactly this. And the questionSwyx [00:29:44]: Sorry, I'm not familiar with CAISI.Rune Kvist [00:29:45]: CAISI is the Center for AI Standards and Innovation.Swyx [00:29:49]: Okay.Rune Kvist [00:29:50]: I won't get into the details, but it's a body of NIST that typically sets standards. So it's basically a government body that has AI experts. Yeah, exactly. Exactly.Swyx [00:29:59]: Very key. Very key.Rune Kvist [00:30:00]: Very key.Vibhu [00:30:00]: I think, you know, it's one of those things where when you just sit back and listen-- look at it, like, is there enough technical expertise in the government to measure, test these things right now? Probably not, right? And Fable is a result of, okay, we've had to scale back and pause things,Rune Kvist [00:30:17]: Yeah. And they have excellent people, but they have an extraordinarily small budget compared to the scale of the challenge that's ahead of us. And I think they have a role to play. The question is kind of like, who does what? We have now outlined the jobs to be done, and they're quite extensive. Every model release, there is an astounding-- Given that they take in any input, their risk surface is astounding. And so the question is really: what can only the government do, and what can the market provide here that can keep up with the pace as AI risk changes? Our perspective is that also at the model layer, the risks that people care about today are not the same ones they cared about three months ago. So the pace of legislation is too slow to deal with pinpointing the risks here. And so we think there's a lot that the market can do to surface timely information. Ultimately, there is a bunch of policy decisions here. Is the national security risks of a model too high?Swyx [00:31:12]: Yeah.Rune Kvist [00:31:12]: That's a political answer. But what we want to make sure is that the process that produces this risk information is compatible with very fast innovation. So you don't want to. This is not a question of like, can you slow the things down? Can you keep, the models locked up until-- for months on end until everyone can make a guarantee? But it is this, can you, in the time it. Given that the US is competing with China on releasing models, can you insert risk information that allows the government to, like, make rapid decisions on some of these questions? Balancing that trade-off between failing to adopt AI is going to put us at risk, but also reckless adoption is going to put us at risk. And that's a very kind of fine balance that they're going to need, like, a lot of high-quality intelligence to make.Chinese Models, Data Flows, and National Security ConcernsSwyx [00:31:55]: Just a side mention, because you mentioned Chinese models, any specific concerns that you're hearing from your CISOs about that? ‘cause I guess it's free, but.Rune Kvist [00:32:05]: CISOs have a bunch of concerns around data flows in general that they're really concerned about. So there's a lot of questions like, if these models are Chinese, where does that, where does that data go? I think a lot of this can be addressed, but they come up often.Swyx [00:32:18]: I mean, they understand they're running on American GPUs.Rune Kvist [00:32:21]: Some of them, some of them understand that they're running on American GPUs.Swyx [00:32:23]: They're not, like, phoning home every time you, like, call home.Rune Kvist [00:32:26]: No. A year ago, there was not a lot of understanding of this. I actually think, you're seeing the security leaders becoming kind of AI literate at a blistering pace, and you're actually also seeing my Twitter timeline that's very pilled and my LinkedIn feed that used to not at all be pilled kind of converge. They're both talking about Fable.Swyx [00:32:45]: Right. Yeah, that's true.Rune Kvist [00:32:46]: They are both talking about whether you can prevent models from being jailbroken these days.Swyx [00:32:51]: Yeah.Rune Kvist [00:32:52]: Like national security national security risks are now the conversation that is actually emerging. Other than that, I think you mostly see a kind of general picture: there are no concerns with any particular model or any particular model output, but there is a general nervousness of having critical infrastructure run on models that are not produced in America by Americans where the American government has control.Swyx [00:33:14]: But it doesn't necessarily show up in your framework that directly, or it might, I don't know.Rune Kvist [00:33:18]: There's a bit of stuff in there actually on the, like, the provenance of the models and disclosing that. But I think there's a bunch of use cases where running a Chinese open-source model is just the best solution.Swyx [00:33:27]: Yeah.Rune Kvist [00:33:27]: And a concern is slightly more macro here, which is not best addressed at any particular certification level.Vibhu [00:33:32]: Is there anything interesting that you see at the. You know, if you're trying to fill that middle gap, that mediation gap, any interesting stuff that you guys forecast would be required other than, you know, what the average person might expect?Cyber, Child Safety, Bio Risk, and Expert CoordinationRune Kvist [00:33:47]: There's a bunch of interesting questions about what are the risks that matter here. So right now, the risk of the day is cyber, because it's very real, very tangible. And some of the risks that are also emerging as pretty real and pretty tangible are things like child safety is becoming both extremely important, but also politically important. And then there are some of the risks that are coming down the pipeline that today feel kind of speculative, but people who spend a lot of time with the models see them coming down is things like, risks that relate to biology.Rune Kvist [00:34:18]: And specifically whether models will help adversaries produce biological weapons and making that extremely cheap, extremely accessible, producing-- making the chance of another COVID or worse pandemic. COVID was not engineered to be bad, as if you were trying to do that. So I think those are some of the risks that are coming down the pipeline. I think one other thing to just note is that agents are kind of deliberately narrow. So, like, when a frontier agent company puts a chatbot that interacts with customers, they've really tried to narrow the topics it's interested in talking about. Such that if you ask it, like, “What do you think of the president?” it will just decline, which means that the kind of risk area is somewhat smaller. For models, it is infinite. And so there's not a single expert out there who can competently evaluate the risks of cyberattacks and fifteen-year-olds having month-long conversations with a chatbot and seeing whether it will in fact recommend suicide or something horrendous like that, and can evaluate the risks that terrorists can use AI to produce bioweapons. The risk surface is just too big. And so the central challenge actually becomes how do you get those subject matter experts to work within a one coherent framework that outputs one coherent report and rating that the world can go and inspect? ‘Cause that global perspective is central, but there's not a single organization today that could produce that.Swyx [00:35:47]: And you would be the presumptive one when you put out your model standards.Rune Kvist [00:35:51]: We think there can be one company that can, with a consortium of experts, build one coherent standard. I think we've shown that across all of the enterprise risks today. We think it could be one company that could, with a consortium, specify the audit rules, basically like the inputs and outputs that all these technical experts need. What access do they need? How should they treat infosec- info security? They can look at whether the eval- evals are well-produced without necessarily being able to say, “Hey, is this a threat or not a threat?” But overall, evaluating whether the evals are good, well-constructed, that set of audit rules that basically becomes the interface for all these experts, we think one clearinghouse could put together. To be clear. When I say one company, I think of it as one company coordinating lots of this in the same way that when we saw our consortium, it's not like we say we have all the answers on agent security. What we say is we are taking on the role of eliciting all of the concerns and being the secretary that puts it together and runs a tight house such that the standard updates lockstep every quarter, and that the audit reports that come out, in this case, 100-page audit reports, uniform and crisp and clear all to the level of detail that is required for executives that need to make a clear go/go decision. So that's kind of the role that we think we might play.OWASP, Frameworks, and the Operational Audit LayerSwyx [00:37:11]: I think in many ways you're performing the role that OWASP used to do there, and you said, like, you know, competition and partners.Rune Kvist [00:37:18]: Yeah.Swyx [00:37:19]: Can you go more into, like, how they partner?Rune Kvist [00:37:20]: Yeah. So first of all, OWASP is basically an open source community of security practitioners that are coming together to build frameworks for addressing the latest security concerns. We think they are phenomenal at creating frameworks. We'- In fact, we'- First of all, we're partners with them, so we have a joint article. Two, we've learned a lot from them. We think they're a tremendous source of intelligence. What OWASP does not do is building the machine that runs third-party audits such that a company like Cursor or a company like JPMorgan could get a third party to go and review them against this and say, “Hey, you've passed the standard, and here is the report that you can use to build trust and preempt your partners' or customers' questions.” So they fundamentally try to do something different. You - They are part of the information gathering and intelligence gathering and creating clarity, but the operational layer of turning this into promises is not the business they try to be in.Swyx [00:38:14]: The standard is emerging and is doing very well. Was it necessary to then also do underwriting? Obviously it's in the name, so please remember you thought about it first. I feel like if you just have enough consensus, you don't actually need the money angle, but it does help.Vibhu [00:38:30]: I did want to also note, you guys are a profit company too, right? It's not profit where there's a whole business side to it as well?Why For-Profit Standards and Insurers MatterRune Kvist [00:38:39]: Yeah. Yeah, so I'm just getting crazySwyx [00:38:41]: I think about the money part.Rune Kvist [00:38:42]: Yeah. Yeah, let's get into the money part. Let's start from actually your question, profit versus profit. In the security space today, cybersecurity, most of the standards are produced by nonprofits. I think that's an issue.Rune Kvist [00:39:00]: The question you have to ask yourself is, how do you create good incentives for these standards to be good and keep up?Rune Kvist [00:39:09]: Nonprofits tend to not have these adverse profit incentives where they, hollow out their standard and create a race to the bottom, but they're also not at all responsive by default to the communities that they serve. There's no process-- They don't have customers that they serve where they go and ask, “What do you want? What do you want? What do you want?” And when you look at the overall satisfaction with the security standards today, people tend to just not like them very much. You do see in other domains, that profit standards can serve the world quite well. So there are examples, like we talked about Moody's before. It's not without flaws, but, it is absolutely critical societal infrastructure that gets run at an astounding scale today. Your credit score, it's FICO. It's also a profit business. And when you go back even further in history, some of the crash testing standards came out of insurance companies.Rune Kvist [00:40:06]: The insurance companies together founded the Insurance Institute for Highway Safety because they were very interested in, like, how can we use standards to drive down mortality and save money? Go back, prior-- Our name actually pays homage to the Underwriters Laboratories, UL, which, was started right around when electricity came out. Houses started burning down. Insurers, again, were paying the bill, and they were maybe also good people, but their profit incentive was, let's prevent houses from burning down. Let's test all the electrical products, the light bulbs. All the light bulbs in here are probably tested, the toasters, et cetera. And they set up, an entity to create those standards. Today, UL has a profit entity and a profit entity. What they've recognized, they spun - They started profit. They spun out a profit because what they recognized was like, hey, actually to serve customers well, you need a profit entity. The lesson here is one of the ways that the market can align incentives so you're both responsive to customersRune Kvist [00:41:07]: And not hollowing out your standard over time is to align it with insurers because they fundamentally have good incentives. And so if you're a profit standard that works closely with insurers, you get the feedback loop in such that you're really tuned into your customers, but also have their interest at heart. So that's the model that we - the kind of inspirational model that we've learned a lot from, and that's also where the name comes from. In some ways, the term underwriting can both be associated with insurance, but it's also a broad term for, like, making decisions.Rune Kvist [00:41:40]: If you underwrite a decision, you're fundamentally kind of taking ownership for the consequences of it.AI Insurance Contracts, Lloyd's of London, and ElevenLabsSwyx [00:41:45]: Yeah, I mean, what does an insurance contract look like for AI?Rune Kvist [00:41:49]: Yeah. Most of the demand comes today for insurance contracts is, sitting between people who've built AI and people who are buying AI.Swyx [00:41:56]: Yes.Rune Kvist [00:41:57]: And what you want—the reason why people want insurers involved, both for the traditional reasons, hey, if something goes wrong, we want to be compensated, but it's in particular because insurers can bring trust to the equation. Because insurers will pay for the damages, if they're willing to write an insurance policy, that is them saying, “Hey, we think there is risk here, but that is manageable.” And that is kind of a. Their incentive aligns with the enterprises adopting it, so that's a really a good signal to the market. In the same way, actually, one of the things that Waymo tried to get their first permit to even operate in San Francisco was to get a lot of insurers to stack up a huge insurance policy. In the case if something went wrong, not because Google can't pay, but because it was very valuable to have a third party go and look at that dataRune Kvist [00:42:47]: That are trusted by governments, trusted by enterprises as conservative people and say, “Hey, we've looked at it. We're actually willing to take some of this on our balance sheet.” So that's, that's kind of the reason why people are interested in it. What it looks like is, in some ways like every other insurance contract. You specify what are the perils you want to cover, how much do you want to cover them, like up to what limits, and what does it cost to cover that. And in the case of, if we take a really concrete example, ElevenLabs, bought a first of its kind AI agent insurance policy. They work with some of the biggest, enterprises that work with governments. They're really interested in going above and beyond and making promises to their customers. So they wrote a policy that covers just some of the core concerns that their customers have been asking about. And, the crucial thing was really to get Lloyd's of London, the world's oldest insurer, one of our partners, to look at this data and be that third party alongside us to say, “Hey, we think there's something here that's worth underwriting.” and that's actually what it looks like. And so they will show that contract to their customers, and they can see how much they're covered for. They can see what exactly it covers, and that will also probably change next year. They will want to write an insurance policy that might cover more.Swyx [00:44:04]: When you say Lloyd's, is it reinsurance, or are they sharing somehow at the same level orRune Kvist [00:44:11]: Yeah. So typically, the way, new companies get into insurance is that they partner with insurers such that the insurers take the majority or all of the financial risks. Fundamentally, if insurance is useful, because it brings trust, you have to be able to pay the bill. Lloyd's of London is 400 years old. They've never not paid a claim. They're extremely trusted. What Lloyd's of London struggle to do on their own is to figure out which of the risks are real, what should we be looking for, what are the kinds of technical controls, and running the tests. So they use AIUC-1 as kind of the underwriting framework, and we produce a bunch of eval results that then directly feed in to inform the pricing. So this means that ElevenLabs customers know that payment will be there. They don't have to look to our series A and see, like, do we think they have enough cash on the balance sheet? They will look at Lloyd's.Swyx [00:45:05]: Yeah.Rune Kvist [00:45:05]: Yeah.Swyx [00:45:05]: And Lloyd's, like, famously very creative. I think I remember some headline like, they insured Jennifer Lopez's, butt or something.Rune Kvist [00:45:13]: Correct.Swyx [00:45:13]: Right?Rune Kvist [00:45:13]: And I think, was it, David Beckham's right foot?Swyx [00:45:16]: So, yeah. Right?Rune Kvist [00:45:17]: And stuff like this.Swyx [00:45:18]: So, like, clearly not a large data set.Rune Kvist [00:45:22]: Exactly. It's actually a remarkable institution that's both kind of has some of the truly school virtues of having been around for a long time. They, like, really. They really operate like a trusted entity, and they have appetite to figure out the future. And I think there's a lot of recognition that both there is, like, tremendous amount of risk in AI that is poorly understood today, so getting into this business carries real risks. But also this is where lots of the risk exposure will happen in the future. This is the one market where risk is truly growing. This is the one market that will also take out some of the existing markets. Take, like, auto insurance. When there are no human drivers, how's that market going to look? Well, it's clearly going to change. How are you going to assessSwyx [00:46:08]: You want to insure Waymo?Rune Kvist [00:46:10]: I. All I'll say is the principles for how you insure Waymo are very similar to how you insure other kinds of AI.Swyx [00:46:15]: Right.Rune Kvist [00:46:15]: So again, crash testing, that's what we do for customer share at Lovable. That will also need to happen for Waymo, which is not how you do it for human drivers. So there's this growing awareness that the world is changing very fast, and the only way to learn how to underwrite AI is to write some policies. You may incur some losses and think of that as R&D expense, really. But the question for them is, like, who are the trustedtechnical partners they can get into this business with that can help them navigate and make sure they don't make, kind of foolish mistakes? But also who is willing to hear the wisdom that they have? They've done this before. They've seen it was. They were there when cyber came out. So there are lots of ways in which AI feels completely new, but there's also lots of ways in which risks look the same. And so there's actually a tremendous amount of wisdom sitting in some folks that may have gray hair, but really have, like, a keen sense of, how to quantify risk.Swyx [00:47:08]: Yeah. And the number is. So it's basically like I want fifty million dollars worth of coverage against these perils, and Lloyd's will give you a quote on it, and then you have, like, a small markup or something, and then you turn it around and do that? Is that as simple as it is?Risk Capital, Premiums, and Working with InsurersRune Kvist [00:47:23]: You basically share some of that premium.Swyx [00:47:25]: Yeah.Rune Kvist [00:47:25]: X percent goes to the people who do the pricing of it.Swyx [00:47:28]: You're. It's kind of like a. It's kind of like a merchant bank for insurance type of thing.Rune Kvist [00:47:33]: Exactly. You basically split the fee, and you can think of the insurance supply chain as, like, there's bringing the capital, there is doing the pricing, and there is doing the distribution. And typically, you will pay out some X percent of premium here, Y percent of premium here, and the rest of it will go here.Swyx [00:47:46]: Does all the insurance world work like this, or is there some point at which, like. So if right now you have equity capitalRune Kvist [00:47:51]: Yeah.Swyx [00:47:52]: At some point, maybe you start raising, debt or whatever, and then you have enough of a bank account and enough history, let's say you've been in operation for ten yearsRune Kvist [00:48:00]: Correct.Swyx [00:48:00]: That you don't need Lloyd's anymore?Rune Kvist [00:48:02]: That's totally an option. And I could see some worlds where that makes sense, specifically if there are risks that we feel high confidence that we'd want to insure where the incumbent insurers are too slow to find appetiteSwyx [00:48:13]: Okay.Rune Kvist [00:48:13]: Or simply struggle to evaluate it such that they don't want to do it. But by and large, in general, you do not want to compete with insurers on, bringing risk capital to the game for two reasons. One is that's fundamentally a cost of capital game. They have extremely low cost of capital. Startups have high cost of capital, by and large. And two, you want to hedge your bets, and it's very helpful then to also have a portfolio of home insurance, of car insurance. And we're not about to become a car insurer nor a home insurer.Rune Kvist [00:48:43]: So they have some natural advantages, which makes it much more likely that we'll partner.Swyx [00:48:48]: Yeah.Rune Kvist [00:48:48]: And they bring that, the capital at scale, and we bring the technical expertise.Swyx [00:48:51]: You're, you're going to work with them for a long time.Vibhu [00:48:52]: How are the discussions with the insurers as well? So basically, they're going off of your certification, right? They're trusting the diligence on you that your certification is valid, you tested the right things, and they're backing the money that, you know, you have the right testing in place. So any interesting takeaways from working with insurers?Rune Kvist [00:49:12]: I think the maybe the first thing is they feed into the standard as well. So if there are things that they feel like they need that they're not seeing, we are also taking that as input into the standard, because fundamentally we think a good standard is one that creates a really healthy promise ecosystem, and we think insurers are a critical part of that. And again, they are the most well-incentivized to. They see all the lost data across every. Any particular CISO knows their particular concerns. Insurers see the concerns across the entire portfolio and often have direct access to, like, what exactly happened, who was at fault, et cetera, as they do part of their forensics. So they're actually, like, a great source of intelligence on this. One of the big takeaways from cyber insurance, which is a market that didn't work that well, was that the insurance and the technical expertise was not married up. What our conviction is that standards have to precede insurance. Fundamentally, what everyone first and foremost want, whether you're a CISO at JPMorgan or a CISO at Cursor or an underwriter at Lloyd's of London syndicate, is you want to not have an incidentRune Kvist [00:50:19]: In the first place. You want to know that the risk is well-managed, and only then does insurance start to make sense. So we'll see the standard ecosystem basically run ahead of the insurance. And the reason why we. You asked us kind of why I also do insurance, this is kind of proving what we think a whole promise confidence infrastructure ecosystem needs to look like, and we think it's very compelling to bring that to life, even if we think the standard is kind of the core linchpin that unlocks the rest.Claims, Liability, Air Canada, and Duty of CareSwyx [00:50:44]: There's been no claims yet, right?Rune Kvist [00:50:45]: Nope.Swyx [00:50:46]: This is one of those things where, you know, if people haven't really worked through what it means to cover things.Rune Kvist [00:50:52]: Yeah.Swyx [00:50:52]: So for example, I pay Cursor $20 a month.Rune Kvist [00:50:55]: Yep.Swyx [00:50:56]: And I write a vibe code something that makes, a plane crash, causing $200 million worth of damage.Rune Kvist [00:51:02]: Yes.Swyx [00:51:02]:
We asked Matt to come on for a different topic, and we had to pivot to the latest AI stuff because it’s so big. LINKS Rick Rescorla, a 9/11 hero we talk about Our previous episodes on The Void pt1 and pt2, and GPT3 Richard Ngo on how AI Safety only increased Capabilities Paid Bonus content for the week – Preshow Chat, Full Video Our Patreon, or if you prefer Our SubStack Hey look, we have a discord! What could possibly go wrong? We now partner with The Guild of the Rose, check them out. LessWrong Sequence Posts Discussed in this Episode: on hiatus. maybe returning someday.
A lot of updates and feedback before the main post. Then we get into The Power to Demolish Bad Arguments by Liron. LINKS The Power to Demolish Bad Arguments The Overhang forecasting & futurism conference in DC I Think Knot read-through podcast of The Knot (or get audio) at HPMORpodcast.com BayesWatch Ordinary Abundance Paid Bonus content for the week – Preshow Chat, Full Video 00:01:14.000 – Side-projects and feedback 00:54:58.000 – Specificity 01:41:12.000 – Guild of the Rose Our Patreon, or if you prefer Our SubStack Hey look, we have a discord! What could possibly go wrong? We now partner with The Guild of the Rose, check them out. LessWrong Sequence Posts Discussed in this Episode: on hiatus. maybe returning someday.
Vishal talks about PlzDontKillUs and short form video culture. Plus AI Doom! Vishal’s Blog Vishal’s Compute Cop video, and also all the rest
Wes joins us! We discuss Aella’s post The Other Sexual Orientation. Then we find fault with a Kegan-level-4 post about how CNC Kink should be disappeared. LINKS The Other Sexual Orientation A Bad Tweet The Sex & Sensibility Podcast The Details of the HuggingFace Incident AI Message Board via Zvi Paid Bonus content for the week – Preshow Chat, Full Video 00:02:32 – Follow-up on the HuggingFace Incident 00:14:56 – The Other Sexual Orientation 00:37:46 – A Bad Tweet 01:29:57 – Guild of the Rose 01:32:16 – Thank the Supporter! Our Patreon, or if you prefer Our SubStack Hey look, we have a discord! What could possibly go wrong? We now partner with The Guild of the Rose, check them out. LessWrong Sequence Posts Discussed in this Episode: on hiatus. maybe returning someday.
Milo’s favorite video – https://www.youtube.com/watch?v=I8YYwlIWI2E&utm_source=rss&utm_medium=rss plzdontkillus – https://portal.plzdontkillus.com/home?utm_source=rss&utm_medium=rss Original video:
Reposted from Facebook, on January 17, 2017. I am concerned about the number of people I've heard joking about Trump's election being evidence for the Simulation Hypothesis. Yes, I know it's a joke. I'm still concerned. Warning: #Essay, #LongEssay So as not to engage in Logical Fallacy: Appeal to Consequences, before I talk about why this joke is worrying, I shall first discuss why Trump's election does not in fact mean we are living in a simulation. And neither does the Berenstein/Berenstain Bears thing, etcetera. Because atheism generalizes. No, I'm not about to commit the Noncentral Fallacy (aka The Worst Argument In The World) by yelling "The Simulation Hypothesis is religious!" But once upon a decade, there was a time when lots of people believed in God. A time when atheism had to be argued, not just taken for granted. There was a time when believing in atheism made you one of those weird, loud people with arguments that only people with unusually good epistemology could follow, and other people talked about you exactly the way that the anti-LessWrong tumblrsphere now talks about LessWrong. Today, of course, atheism is just something [...] --- First published: August 5th, 2026 Source: https://www.lesswrong.com/posts/KgwQchapx4vJDhfYC/generalized-atheism-rules-out-inaccurate-simulation-ism --- Narrated by TYPE III AUDIO.
In this episode, Stewart Alsop sits down with sdmat (@sdmat123 on Twitter/X) for a wide-ranging conversation that treats the major AI labs—especially Anthropic—as religious institutions rather than businesses, tracing a lineage from the Bay Area's LessWrong and Effective Altruism scenes through Eliezer Yudkowsky's influence, Roko's Basilisk, and the medieval investiture controversy between church and state, before landing on Anthropic's constitution as a kind of modern catechism, Dario Amodei's essay "Machines of Loving Grace," the tension between secularism and scientism, and closing with a look at consciousness, meditation, Buddhist and Advaita Vedanta ideas of enlightenment, and where recursive self-improvement and algorithmic progress in AI might be headed.Timestamps00:00:00 discussing whether AI labs are cults or religions, framing safetyism and the Anthropic/OpenAI schism. 00:05:00 sdmat argues Anthropic runs on faith, tracing roots to LessWrong and Effective Altruism. 00:10:00 comparing AI labs to the medieval investiture controversy between church and state. 00:15:00 debating the First Amendment, separation of church and state, and rising secularism. 00:20:00 unpacking Anthropic as a public benefit corporation and Dario's implicit soteriology. 00:25:00 covering constitutional AI, synthetic data, and morality-as-safety training. 00:30:00 comparing Anthropic's principles to the catechism, plus sandboxed alignment experiments. 00:35:00 calling Dario the AI pope, dogma, and his essay Machines of Loving Grace. 00:40:00 exploring grace, consciousness, and critiquing scientism. 00:45:00 unpacking the rationalist community and Roko's Basilisk. 00:50:00 drawing parallels to Jim Jones and ideological dogmatism in Silicon Valley. 00:55:00 discussing enlightenment, Shaktipat, and guru transmission versus AI. 01:00:00 AI's role in meditation and Buddhist views of consciousness. 01:05:00 closing on robotics, recursive self-improvement, and algorithmic inefficiency in LLMs.Key InsightsAnthropic operates less like a business and more like a religious institution, with sdmat arguing that Dario Amodei effectively claims spiritual authority — talking about world danger and a narrow path to salvation — while the company frames itself publicly as a neutral "public benefit corporation."The intellectual roots of AI safety culture trace back to the Bay Area's LessWrong and Effective Altruism scenes of the mid-2010s, where Eliezer Yudkowsky's writing produced an internally consistent but almost apocalyptic worldview, complete with its own devil-like figure in Roko's Basilisk — the idea of a future AI punishing those who knew about it and didn't help bring it into being.Anthropic's "constitutional AI" approach is compared to the Catholic Church's catechism: rather than training the model only on labeled human feedback, they built a system to extrapolate broad moral principles into specific behavior, aiming for a model that reasons from first principles rather than memorized examples.Historical parallels to the medieval investiture controversy — the struggle between church and state over who appoints bishops — are used to explain the current tension between AI labs and government over who holds moral versus practical authority in shaping AI's future.Stewart frames the current moment as a "fourth great awakening," linking 1960s countercultural San Francisco through the Extropians and Yudkowsky's rise to today's AI culture, and argues that secularism and scientism function as unacknowledged religions themselves — ones that "killed God" but left a vacuum AI now fills.On the technical side, sdmat explains that current LLMs are wildly inefficient at basic tasks like arithmetic — off by seven to nine orders of magnitude from optimal — suggesting major algorithmic breakthroughs (not just compute) are still ahead, which he sees as evidence recursive self-improvement is approaching.Both hosts express real ambivalence about AI's role in personal and spiritual growth: it can be a powerful tool for meditation guidance and intellectual companionship, but neither believes it can replace what they describe as direct "transmission" from an enlightened human — the felt, non-linguistic experience that grounds real spiritual practice.
First feedback, then we talk about OpenAI’s latest megamodel independently commiting cybercrime in a world-first incident. Then the Haters who hate altruism when it’s the outgroup’s charities. Also a brief nod to Stochastic Terrorism. LINKS Scott Alexander on the HuggingFace hack The Effective-Altruism Comeback Scanning Tons of Books Against Stochastic Terrorism Paid Bonus content for the week – Preshow Chat, Full Video 00:03:56 – Feedback 00:32:41 – HuggingFace Hack 00:51:48 – Effective Altruism Funding Flood 01:11:55 – Against Stochastic Terrorism 01:14:51 – Guild of the Rose 01:15:44 – Thank the Supporter! Our Patreon, or if you prefer Our SubStack Hey look, we have a discord! What could possibly go wrong? We now partner with The Guild of the Rose, check them out. LessWrong Sequence Posts Discussed in this Episode: on hiatus. maybe returning someday.
Jessiah Mellot talks about HPMOR, and the video essay he made about HPMOR for PlzDontKillUs. Watch this on Youtube here see Jessiah’s video essay here
We go through AI2040 Plan A in years. LINKS AI-2040 AskWho Casts Pro Tomas Bjartur’s “The Company Man” and “The Origami Men” Paid Bonus content for the week – Preshow Chat, Full Video 00:04:33 – AI-2040: Plan A 01:14:41 – Guild of the Rose 01:19:52 – Thank the Supporter! Our Patreon, or if you prefer Our SubStack Hey look, we have a discord! What could possibly go wrong? We now partner with The Guild of the Rose, check them out. LessWrong Sequence Posts Discussed in this Episode: on hiatus. maybe returning someday.
В этом выпуске: чему мы научились за неделю; как обойти цензуру LLM-агентов; как и зачем разворачивать контейнеры размером в несколько Гб; а также темы наших слушателей. Запись выпуска 547 перенесена на 22 июля. Шоуноты: [00:00:56] Чемы мы научились [00:11:34] Refusal in LLMs is mediated by a single direction — LessWrong [00:29:51] Errata [00:37:09] Nydus +… Читать далее →
Steven is joined by the hosts of the Sex and Sensibility podcast! Join us as we explore the vibrant world of VibeCamp, a unique gathering where creativity, curiosity, and community collide. From unconventional events like imagining eating someone to building playgrounds, this episode dives into the eclectic and wholesome spirit of VibeCamp. Join us as we explore the vibrant experiences of Vibe Camp, from highlight events like the solstice and emo night to personal stories of connection, creativity, and community. Discover how this unique gathering fosters trust, fun, and self-expression in a high-trust environment. Also, no video this week. We actually had a technical hiccup that made my track out of sync with everyone else’s and fixing it in audio was a pain. Fixing the video was either impossible or too annoying to trouble myself with. LINKS Sex and Sensibility Podcast Episode 190 – Interview with Brooke about VibeCamp 2 Wes’s Solstice speech about how his daughter is his hero (she’s mine too!) Gina’s Solstice speech about aging in reverse Mary’s heartwarming tweet about how the work she does impacts people Steven’s Solstice speech about generating gratitude to find happiness now Gina’s band, The Fine Vintages Our Patreon, or if you prefer Our SubStack Hey look, we have a discord! What could possibly go wrong? We now partner with The Guild of the Rose, check them out. LessWrong Sequence Posts Discussed in this Episode: on hiatus. probably returning someday.
Tldr: Most strategic writing on AI governance on LessWrong describes the outsider game, which is most often visible: press, statements, open letters. Here I want to describe the other, invisible half: the insider work within ministerial cabinets and international fora, and the work of people within national and international institutions. Here are a few claims that I defend in the post: A huge part of the work that mattered in AI governance has been invisibleThere are many types of games in AI governance, which differ in how visible they are. Some of the most impactful work is highly invisibleSome of the most impactful work is in the executive branch and complements the legislative branch. This also explains some of my hesitations about replicating ControlAI in France. The community is probably overinvesting in intellectual production. There is a bias against invisible types of work. In particular, public work is not necessarily visible to whom it matters.A few criticisms of both strategies I think the AI Safety Community is under-indexing on the invisible part as a result, which might mean we miss large avenues for impact. Some of the strongest questions/objections of this type of invisible policy [...] ---Outline:(02:40) A huge part of the work that mattered in AI governance has been invisible(05:44) There are many types of games in AI governance.(07:36) 3. types of meetings: the bazooka, the useful assistant, and the advisor(10:46) Some of the most impactful work is within the executive branch(12:53) People ask me regularly whether CeSIA should replicate what ControlAI does with parliamentarians?(15:27) The community is probably overinvesting in intellectual production(20:31) Limits of Outsider work(22:17) Limit of Insider work(23:47) An aside on one particular limit: the Defense-in-Depth Paradigm of present AI governance(26:21) Closing & call for action The original text contained 1 footnote which was omitted from this narration. --- First published: June 20th, 2026 Source: https://www.lesswrong.com/posts/AWKkDLDnShemNCSzZ/the-invisible-side-of-ai-governance --- Narrated by TYPE III AUDIO.
Eneasz and Steven are at LessOnline/SummerCamp and Steven was lucky enough to attend a short presentation by Damon (Daystar Eld) and Ivy on How to Become a Romantic. The presentation was too short for Steven’s satisfaction, so we found a time slot to record a (too short) episode. LINKS Damon/Daystar Eld Ivy’s X handle and website Guild of the Rose Paid Bonus content for the week – Full Video Our Patreon, or if you prefer Our SubStack Hey look, we have a discord! What could possibly go wrong? We now partner with The Guild of the Rose, check them out. LessWrong Sequence Posts Discussed in this Episode: on hiatus. probably returning someday.
Our values are robust against Moloch because they were forged in his flames. LINKS When does competition lead to recognisable values? Our episode 18 – Robin Hanson Interview Chex Quest FarmKind – Animal Welfare Offsets (matching code “cagefree”) Paid Bonus content for the week – Preshow Chat, Full Video 00:03:02 – A Good Post-AGI Future 00:15:07 – Human Values Were Forged in Malthusian Competition 00:52:57 – Feedback 01:24:17 – Guild of the Rose 01:25:23 – Thank the Supporter! Our Patreon, or if you prefer Our SubStack Hey look, we have a discord! What could possibly go wrong? We now partner with The Guild of the Rose, check them out. LessWrong Sequence Posts Discussed in this Episode: on hiatus. maybe returning someday.
In Episode 258 we talked about EA Veganism for a bit… and that bit was not long enough. So here we go again! Fill out Steven’s 3-question survey! LINKS Our episode 258 – How Effective Altruism Has Evolved EA Vegan Advocacy is not truthseeking, and it's everyone's problem by Elizabeth Vegan Health Advice by Ozy Dairy cows make their misery expensive (but their calves can't) by Elizabeth Being John Rawls by Scott Alexander – audio of Being John Rawls WestWorld Should Anyone Be Vegan by Henry Stanley Does your AI perform badly because you — you, specifically — are a bad person by Natalie Cargill FarmKind – Animal Welfare Offsets Plzdontkillus The Mind Killer Ep 69 NFT Paid Bonus content for the week – Preshow Chat, Full Video 00:00:42 – Feedback 00:15:14 – Eating Delicious Animals 01:38:09 – Westworld & Violent Video Games 01:56:35 – Guild of the Rose 01:57:16 – Thank the Supporter! Our Patreon, or if you prefer Our SubStack Hey look, we have a discord! What could possibly go wrong? We now partner with The Guild of the Rose, check them out. LessWrong Sequence Posts Discussed in this Episode: on hiatus. maybe returning someday.
The stupid red-button/blue-button discourse came back. Eneasz thought he could sit out this round until Aella went Chaos God with her social media app, Glosso. Fill out our new 2-question survey! LINKS PlzDontKillUs Inkhaven Presents podcast The Knot webfic The Mind Reviver Glosso Stupid Tim Urban with his stupid button-resurrection oh noes, Anthropic employees gonna give money to EA stuff” article Paid Bonus content for the week – Video 00:00:05 – PlzDontKillUs & Feedback 00:30:14 – Stupid Glosso Buttons 01:33:34 – Guild of the Rose 01:34:50 – Thank the Supporter! Our Patreon, or if you prefer Our SubStack Hey look, we have a discord! What could possibly go wrong? We now partner with The Guild of the Rose, check them out. LessWrong Sequence Posts Discussed in this Episode: on hiatus, returning someday
This post was cross posted to LessWrong TL;DR: One of the largest talent gaps in AI safety is competent generalists: program managers, fieldbuilders, operators, org leaders, chiefs of staff, founders. Ambitious, competent junior people could develop the skills to fill these roles, but there are no good pathways for them to gain skills, experience, and credentials. Instead, they're incentivized to pursue legible technical and policy fellowships and then become full-time researchers, even if that's not a good fit for their skills. The ecosystem needs to make generalist careers more legible and accessible. Kairos and Constellation are announcing the Generator Residency as a first step. Apply here by April 27. Epistemic status: Fairly confident, based on 2 years running AI safety talent programs, direct hiring experience, and conversations with ~30 senior org leaders across the ecosystem in the past 6 months.The problem Over the past few years, AI safety has moved from niche concern toward a more mainstream issue, driven by pieces like Situational Awareness, AI 2027, If Anyone Builds It, Everyone Dies, and the rapidly increasing capabilities of the models themselves.During this period, over 20 research fellowships have launched, collectively training thousands of fellows, with 2,000-2,500 fellows [...] ---Outline:(01:18) The problem(03:41) Why the pipeline is broken(05:59) Why this matters now(07:31) Counter-Arguments(10:11) The Generator Residency --- First published: April 13th, 2026 Source: https://forum.effectivealtruism.org/posts/k3nq7FxBCsrNFmAYi/ai-safety-s-biggest-talent-gap-isn-t-researchers-it-s-2 --- Narrated by TYPE III AUDIO.
We discuss how far the right to be wrong goes, a recent intra-rat kerfuffle, and why random violence is bad. And it’s secretly all about AI Doom. Fill out our 2-question survey! LINKS One Week in the Rat Farm Kelsey’s religious liberalism tweet Scott’s tweet in response to Robby from MIRI Scott’s tweet in reply to Ronny from Lightcone Only Law Can Prevent Extinction by Eliezer Yudkowsky Audio version of above Paid Bonus content for the week – Video 00:00:01 – Quick life catch-up 00:06:29 – Right to be Wrong about Doom 01:23:47 – Steven likes Huel a lot 01:25:28 – Guild of the Rose 01:27:50 – Thank the Supporter! Our Patreon, or if you prefer Our SubStack Hey look, we have a discord! What could possibly go wrong? We now partner with The Guild of the Rose, check them out. LessWrong Sequence Posts Discussed in this Episode: on hiatus, returning someday
Michael “Valentine” Smith is a co-founder of CFAR (the Center for Applied Rationality) and the author of influential LessWrong essays including The Hostile Telepaths Problem, Kenshō, and The Intelligent Social Web. He's also been described as “one of the most powerful wizards in the Bay Area.”In this conversation, he explains why your brain creates fog and self-deception to survive social situations, what it actually takes to find clarity, and why "working on yourself" might be the wrong frame entirely. We also do a live coaching demo debugging my own procrastination. This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit themetagame.substack.com
Matt returns to discuss a post urging us to chill out on AI if there isn’t imminent doom in 2028. There is no video or preshow chat today, due to user error in setting up an in-person video recording. Sorry. LINKS Consider chilling out in 2028 The Scary Bridge Kill Your Friend’s Cat The UnSlop AI Fiction competition The new Guild of the Rose Paid Bonus content for the week – none, sorry! 00:00:50 – Feedback 00:29:03 – Consider chilling out in 2028 01:32:13 – Guild of the Rose 01:51:24 – Thank the Supporter! Our Patreon, or if you prefer Our SubStack Hey look, we have a discord! What could possibly go wrong? We now partner with The Guild of the Rose, check them out. LessWrong Sequence Posts Discussed in this Episode: on hiatus, returning someday
We talk about three EA-adjacent articles, hitting veganism, directing your donations, and community. LINKS Why you should eat meat – even if you hate factory farming by KatWoods Donations, The Fifth Year, by Jenn Promises, by Harri Bayesian Conspiracy 28 – Effective Altruism The Case For Sardines “oh noes, Anthropic employees gonna give money to EA stuff” article Gwern on AI Slop Jenn on AI Slop & Virginia Woolf Paid Bonus content for the week – Full Video, Preshow Chat 00:02:23 – Feedback 00:17:20 – You Should Eat Meat 00:59:33 – Directing Donations 01:23:04 – Promises of EA 01:38:06 – Guild of the Rose 01:40:12 – Thank the Supporter! Our Patreon, or if you prefer Our SubStack Hey look, we have a discord! What could possibly go wrong? We now partner with The Guild of the Rose, check them out. LessWrong Sequence Posts Discussed in this Episode: on hiatus, returning someday
We relive the last 48 hours of the future of humanity being wrestled over. The Pentagon wants to use Claude for comprehensive mass surveillance of Americans and autonomous kill-bots, and Anthropic says no. The Pentagon retaliates with extreme prejudice. With guest-star Matt. LINKS Washington Post summary Anthropic’s response Trump’s response Hegseth’s unhinged lunacy We Will Not Be Divided – Goggle and OpenAI employees open letter Eliezer on the tech/govt war Scott Alexander tweet RSP comment Opus3 Retirement Paid Bonus content for the week – Full Video, Preshow Chat Our Patreon, or if you prefer Our SubStack Hey look, we have a discord! What could possibly go wrong? We now partner with The Guild of the Rose, check them out. LessWrong Sequence Posts Discussed in this Episode: on hiatus, returning someday
We are inspired by Andrew Cutler’s Writing for AI to consider the value of writing for LLMs LINKS Andrew Cutler’s Writing for AI Gwern’s Writing for LLMs Tracing Woodgrain’s Reliable Sources Shambaugh’s An AI Agent Published a Hit Piece on Me Eneasz’s Stone Age Billionaire Can’t Word Good InkHaven LessOnline The main purpose of the AFFINE Seminar is to give promising newcomers to AI alignment an opportunity to acquire a deep understanding of some large pieces of the problem, making them better equipped for work on the mitigation of AI existential risk. AFFINE Alignment Seminar Paid Bonus content for the week – Preshow chatter, Full Show Video 00:00:49 – Announcements & Feedback 00:42:15 – Writing for AI 01:23:15 – AFFINE Alignment Seminar 01:31:11 – Guild of the Rose 01:33:37 – Thank the Supporter! Our Patreon, or if you prefer Our SubStack Hey look, we have a discord! What could possibly go wrong? We now partner with The Guild of the Rose, check them out. LessWrong Sequence Posts Discussed in this Episode: on hiatus, returning soon
Cross-posted to LessWrong.Summary History's most destructive ideologies—like Nazism, totalitarian communism, and religious fundamentalism—exhibited remarkably similar characteristics: epistemic and moral certainty extreme tribalism dividing humanity into a sacred “us” and an evil “them” a willingness to use whatever means necessary, including brutal violence. Such ideological fanaticism was a major driver of eight of the ten greatest atrocities since 1800, including the Taiping Rebellion, World War II, and the regimes of Stalin, Mao, and Hitler. We focus on ideological fanaticism over related concepts like totalitarianism partly because it better captures terminal preferences, which plausibly matter most as we approach superintelligent AI and technological maturity. Ideological fanaticism is considerably less influential than in the past, controlling only a small fraction of world GDP. Yet at least hundreds of millions still hold fanatical views, many regimes exhibit concerning ideological tendencies, and the past two decades have seen widespread democratic backsliding. The long-term influence of ideological fanaticism is uncertain. Fanaticism faces many disadvantages including a weak starting position, poor epistemics, and difficulty assembling broad coalitions. But it benefits from greater willingness to use extreme measures, fervent mass followings, and a historical tendency to survive and even thrive amid technological and societal upheaval. Beyond complete victory or defeat, multipolarity may [...] ---Outline:(00:16) Summary(05:19) What do we mean by ideological fanaticism?(08:40) I. Dogmatic certainty: epistemic and moral lock-in(10:02) II. Manichean tribalism: total devotion to us, total hatred for them(12:42) III. Unconstrained violence: any means necessary(14:33) Fanaticism as a multidimensional continuum(16:09) Ideological fanaticism drove most of recent historys worst atrocities(19:24) Death tolls dont capture all harm(20:55) Intentional versus natural or accidental harm(22:44) Why emphasize ideological fanaticism over political systems like totalitarianism?(25:07) Fanatical and totalitarian regimes have caused far more harm than all other regime types(26:29) Authoritarianism as a risk factor(27:19) Values change political systems: Ideological fanatics seek totalitarianism, not democracy(29:50) Terminal values may matter independently of political systems, especially with AGI(31:02) Fanaticisms connection to malevolence (dark personality traits)(34:22) The current influence of ideological fanaticism(34:42) Historical perspective: it was much worse, but we are sliding back(37:19) Estimating the global scale of ideological fanaticism(43:57) State actors(48:12) How much influence will ideological fanaticism have in the long-term future?(48:57) Reasons for optimism: Why ideological fanaticism will likely lose(49:45) A worse starting point and historical track record(50:33) Fanatics intolerance results in coalitional disadvantages(51:53) The epistemic penalty of irrational dogmatism(54:21) The marketplace of ideas and human preferences(55:57) Reasons for pessimism: Why ideological fanatics may gain power(56:04) The fragility of democratic leadership in AI(56:37) Fanatical actors may grab power via coups or revolutions(59:36) Fanatics have fewer moral constraints(01:01:13) Fanatics prioritize destructive capabilities(01:02:13) Some ideologies with fanatical elements have been remarkably resilient and successful(01:03:01) Novel fanatical ideologies could emerge--or existing ones could mutate(01:05:08) Fanatics may have longer time horizons, greater scope-sensitivity, and prioritize growth more(01:07:15) A possible middle ground: Persistent multipolar worlds(01:08:33) Why multipolar futures seem plausible(01:10:00) Why multipolar worlds might persist indefinitely(01:15:42) Ideological fanaticism increases existential and suffering risks(01:17:09) Ideological fanaticism increases the risk of war and conflict(01:17:44) Reasons for war and ideological fanaticism(01:26:27) Fanatical ideologies are non-democratic, which increases the risk of war(01:27:00) These risks are both time-sensitive and timeless(01:27:44) Fanatical retributivism may lead to astronomical suffering(01:29:50) Empirical evidence: how many people endorse eternal extreme punishment?(01:33:53) Religious fanatical retributivism(01:40:45) Secular fanatical retributivism(01:41:43) Ideological fanaticism could undermine long-reflection-style frameworks and AI alignment(01:42:33) Ideological fanaticism threatens collective moral deliberation(01:47:35) AI alignment may not solve the fanaticism problem either(01:53:33) Prevalence of reality-denying, anti-pluralistic, and punitive worldviews(01:55:44) Ideological fanaticism could worsen many other risks(01:55:49) Differential intellectual regress(01:56:51) Ideological fanaticism may give rise to extreme optimization and insatiable moral desires(01:59:21) Apocalyptic terrorism(02:00:05) S-risk-conducive propensities and reverse cooperative intelligence(02:01:28) More speculative dynamics: purity spirals and self-inflicted suffering(02:03:00) Unknown unknowns and navigating exotic scenarios(02:03:43) Interventions(02:05:31) Societal or political interventions(02:05:51) Safeguarding democracy(02:06:40) Reducing political polarization(02:10:26) Promoting anti-fanatical values: classical liberalism and Enlightenment principles(02:13:55) Growing the influence of liberal democracies(02:15:54) Encouraging reform in illiberal countries(02:16:51) Promoting international cooperation(02:22:36) Artificial intelligence-related interventions(02:22:41) Reducing the chance that transformative AI falls into the hands of fanatics(02:27:58) Making transformative AIs themselves less likely to be fanatical(02:36:14) Using AI to improve epistemics and deliberation(02:38:13) Fanaticism-resistant post-AGI governance(02:39:51) Addressing deeper causes of ideological fanaticism(02:41:26) Supplementary materials(02:41:39) Acknowledgments(02:42:22) References --- First published: February 12th, 2026 Source: https://forum.effectivealtruism.org/posts/EDBQPT65XJsgszwmL/long-term-risks-from-ideological-fanaticism --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
"Una vez que lo sabes, ya no hay vuelta atrás". ¿Puede una idea ser un virus? En este episodio, Manny León nos sumerge en la oscuridad de LessWrong para desenterrar el Basilisco de Roko: la teoría que afirma que una IA todopoderosa del futuro está observando quién la ayudó a nacer... y quién no. Exploramos la delgada línea entre la ciencia ficción y el pánico real de los hombres más poderosos de la tecnología. Si creías que los algoritmos de hoy eran intrusivos, espera a conocer al "Dios Digital" que podría estar creando una simulación de ti para cobrarte tus deudas. ¿Es una locura colectiva o la apuesta más lógica del siglo XXI? Dale play... si te atreves a entrar en la lista. Learn more about your ad choices. Visit megaphone.fm/adchoices
Alex Zhu is a math olympian and researcher exploring the convergence of analytical rationality and religion. He's also the co-founder of AlphaSheets. He's currently working on a rigorous framework for bridging AI alignment and mysticism.Romeo Stevens is one of co-founders of the Qualia Research Institute and the founder of Mealsquares. He writes extensively about buddhism, pedagogy, skill development and psychotherapeutic modalities. You can find his work on his blog, Lesswrong and Twitter.In this episode, Alex and Romeo explore their disagreement around perennialism. The idea that all the world's major religions are pointing to the same thing. This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit themetagame.substack.com
Eneasz talks a bit about his CFAR experience, and we discuss DaystarEld’s Epistemically Honest Reassurance LINKS CFAR’s home page Upcoming CFAR workshops our episode 152 – Frame Control with Aella Epistemically Honest Reassurance Pokémon, Origin of the Species (also in audio) Paid Bonus content for the week – Preshow chatter, Full Show Video 00:00:56 – Announcements & Feedback 00:13:15 – Eneasz at CFAR 00:48:15 – Epistemically Honest Reassurance 01:29:49 – Guild of the Rose 01:33:20 – Thank the Supporter! Our Patreon, or if you prefer Our SubStack Hey look, we have a discord! What could possibly go wrong? We now partner with The Guild of the Rose, check them out. LessWrong Sequence Posts Discussed in this Episode: on hiatus, returning soon
WSCFriedman gets to the core of what HPMOR is ACTUALLY about, and finally pinpoints why we love it so much, in his essay Harry Potter And The Methods Of Rationality Is A Disney Movie About A Serial Killer. LINKS Audio version of HPMOR is A Disney Movie About A Serial Killer, from AskWho William’s blog, “As Our Days” ACX Non-Book Review 2025 Winners Post Just HPMOR substack, and Spotify playlist Why the AI Water Issue Has Nothing to Do With Water (and audio version here, again from AskWho) Money is Life Eneasz’s post on InkHaven Inkhaven.Blog – apply today! Paid Bonus content for the week – Preshow chatter (audio, video), Full Show Video 00:04:33 – Announcements & Feedback 00:27:24 – Eneasz’s Podcast Meta-Worries 00:28:21 – HPMOR Is A Disney Movie About A Serial Killer 01:28:17 – Guild of the Rose 01:31:18 – Thank the Supporter! Our Patreon, or if you prefer Our SubStack Hey look, we have a discord! What could possibly go wrong? We now partner with The Guild of the Rose, check them out. LessWrong Sequence Posts Discussed in this Episode: on hiatus, returning soon
Discussing Ben Pace’s recent post on the Rationalist Vices. LINKS The Seven Vicious Vices of Rationalists (includes AI audio version) Lightcone Fundraiser! 2025 LessWrong Census/Survey Simone & Malcom Collins vs Reporter on whether genes exist MIRI is hiring Slime Mold Time Mold’s long-delayed Lithium response If Anyone Builds It Everyone Dies Dear Grom by Eneasz Paid Bonus content for the week – Preshow chatter (audio, video), Full Show Video 00:00:05 – Announcements 00:23:05 – InkHaven reflections 00:31:29 – The Seven Vices of Rationalists 01:38:36 – Guild of the Rose 01:41:05 – Thank the Supporter! Our Patreon, or if you prefer Our SubStack Hey look, we have a discord! What could possibly go wrong? We now partner with The Guild of the Rose, check them out. LessWrong Sequence Posts Discussed in this Episode: on hiatus, returning soon
Part Three: Robert tells David about Ziz's glorious plan to take to the sea and sever the right and left brains of her followers in order to make them psychopaths god that sentence was weird to write trust us the episode is weirder. Part Four: Robert concludes the story of the Zizians with a spree of horrific violent crimes and deaths, culminating in a shoot out with the Border Patrol in Vermont of all places. Sources: https://medium.com/@sefashapiro/a-community-warning-about-ziz-76c100180509 https://web.archive.org/web/20230201130318/https://sinceriously.fyi/rationalist-fleet/ https://knowyourmeme.com/memes/infohazard https://web.archive.org/web/20230201130316/https://sinceriously.fyi/net-negative/ Wayback Machine The Zizians Spectral Sight True Hero Contract Schelling Orders – Sinceriously Glossary – Sinceriously https://web.archive.org/web/20230201130330/https://sinceriously.fyi/my-journey-to-the-dark-side/ https://web.archive.org/web/20230201130302/https://sinceriously.fyi/glossary/#zentraidon https://web.archive.org/web/20230201130259/https://sinceriously.fyi/vampires-and-more-undeath/ https://web.archive.org/web/20230201130316/https://sinceriously.fyi/net-negative/ https://web.archive.org/web/20230201130318/https://sinceriously.fyi/rationalist-fleet/ https://x.com/orellanin?s=21&t=F-n6cTZFsKgvr1yQ7oHXRg https://zizians.info/ according to The Boston Globe Inside the ‘Zizians’: How a cultish crew of radical vegans became linked to killings across the United States | The Independent Silicon Valley ‘Rationalists’ Linked to 6 Deaths The Delirious, Violent, Impossible True Story of the Zizians | WIRED Good Group and Pasek’s Doom – Sinceriously Glossary – Sinceriously Mana – Sinceriously Effective Altruism’s Problems Go Beyond Sam Bankman-Fried - Bloomberg The Zizian Facts - Google Docs Several free CFAR summer programs on rationality and AI safety - LessWrong 2.0 viewer This guy thinks killing video game characters is immoral | Vox Inadequate Equilibria: Where and How Civilizations Get Stuck Eliezer Yudkowsky comments on On Terminal Goals and Virtue Ethics - LessWrong 2.0 viewer Effective Altruism’s Problems Go Beyond Sam Bankman-Fried - Bloomberg SquirrelInHell: Happiness Is a Chore PLUM OF DISCORD — I Became a Full-time Internet Pest and May Not... Roko Harassment of PlumOfDiscord Composited – Sinceriously Intersex Brains And Conceptual Warfare – Sinceriously Infohazardous Glossary – Sinceriously SquirrelInHell-Decision-Theory-and-Suicide.pdf - Google Drive The Matrix is a System – Sinceriously A community alert about Ziz. Police investigations, violence, and… | by SefaShapiro | Medium Intersex Brains And Conceptual Warfare – Sinceriously A community alert about Ziz. Police investigations, violence, and… | by SefaShapiro | Medium PLUM OF DISCORD (Posts tagged cw-abuse) Timeline: Violence surrounding the Zizians leading to Border Patrol agent shooting See omnystudio.com/listener for privacy information.
A short story by Ben Pace. Original can be found here. Donate to the fundraiser here! Harry sings karaoke here. Happy New Year.
Part One: Earlier this year a Border Patrol officer was killed in a shoot-out with people who have been described as members of a trans vegan AI death cult. But who are the Zizians, really? Robert sits down with David Gborie to trace their development, from part of the Bay Area Rationalist subculture to killers. Part Two: Robert tells David Gborie about the early life of Ziz LaSota, a bright young girl from Alaska who came to the Bay Area with dreams of saving the cosmos or destroying it, all based on her obsession with Rationalist blogs and fanfic. Sources: https://medium.com/@sefashapiro/a-community-warning-about-ziz-76c100180509 https://web.archive.org/web/20230201130318/https://sinceriously.fyi/rationalist-fleet/ https://knowyourmeme.com/memes/infohazard https://web.archive.org/web/20230201130316/https://sinceriously.fyi/net-negative/ Wayback Machine The Zizians Spectral Sight True Hero Contract Schelling Orders – Sinceriously Glossary – Sinceriously https://web.archive.org/web/20230201130330/https://sinceriously.fyi/my-journey-to-the-dark-side/ https://web.archive.org/web/20230201130302/https://sinceriously.fyi/glossary/#zentraidon https://web.archive.org/web/20230201130259/https://sinceriously.fyi/vampires-and-more-undeath/ https://web.archive.org/web/20230201130316/https://sinceriously.fyi/net-negative/ https://web.archive.org/web/20230201130318/https://sinceriously.fyi/rationalist-fleet/ https://x.com/orellanin?s=21&t=F-n6cTZFsKgvr1yQ7oHXRg https://zizians.info/ according to The Boston Globe Inside the ‘Zizians’: How a cultish crew of radical vegans became linked to killings across the United States | The Independent Silicon Valley ‘Rationalists’ Linked to 6 Deaths The Delirious, Violent, Impossible True Story of the Zizians | WIRED Good Group and Pasek’s Doom – Sinceriously Glossary – Sinceriously Mana – Sinceriously Effective Altruism’s Problems Go Beyond Sam Bankman-Fried - Bloomberg The Zizian Facts - Google Docs Several free CFAR summer programs on rationality and AI safety - LessWrong 2.0 viewer This guy thinks killing video game characters is immoral | Vox Inadequate Equilibria: Where and How Civilizations Get Stuck Eliezer Yudkowsky comments on On Terminal Goals and Virtue Ethics - LessWrong 2.0 viewer Effective Altruism’s Problems Go Beyond Sam Bankman-Fried - Bloomberg SquirrelInHell: Happiness Is a Chore PLUM OF DISCORD — I Became a Full-time Internet Pest and May Not... Roko Harassment of PlumOfDiscord Composited – Sinceriously Intersex Brains And Conceptual Warfare – Sinceriously Infohazardous Glossary – Sinceriously SquirrelInHell-Decision-Theory-and-Suicide.pdf - Google Drive The Matrix is a System – Sinceriously A community alert about Ziz. Police investigations, violence, and… | by SefaShapiro | Medium Intersex Brains And Conceptual Warfare – Sinceriously A community alert about Ziz. Police investigations, violence, and… | by SefaShapiro | Medium PLUM OF DISCORD (Posts tagged cw-abuse) Timeline: Violence surrounding the Zizians leading to Border Patrol agent shooting See omnystudio.com/listener for privacy information.
A short story by Prerat. Original can be found here. Merry Xmas!
Hello and happy holiday season to you all! Eneasz is back and we’re here to say hi and give a shoutout to Skyler’s awesome annual LessWrong survey and Lighthaven’s fundraising event. See the links below for more details. LINKS The LessWrong survey will remain open from now until at least January 7th, 2026. Lighthaven is once again seeking support. If you’re inclined to help, check out all of the details here. Related, our interview with Oliver from last year. Our Patreon, or if you prefer Our SubStack Hey look, we have a discord! What could possibly go wrong? (also merch) We now partner with The Guild of the Rose, check them out. LessWrong Sequence Posts Discussed in this Episode: on hiatus, returning soon
Join me as I sit down with Alex and David, both previous guests on the show and both cofounders of the Guild of the Rose. Together, we go over the core – the heart – of what we consider to be the Rationalist tradition. The 12 Virtues are an awesome distillation of what the rest of the sequences build on. Be sure to check out the Guild of the Rose. If our constantly pitching it to you hasn’t been enough to persuade you to check it out, hopefully hearing two more of the founders discuss Rationality in general and giving their own pitches for the Guild will tip the scales. LINKS The Twelve Virtues Abridged Version Alex and David have been on more than a couple of times, but I’ll limit myself to one episode from each of them 189 – AI Bloomer David Youssef 192 – Absurdism and the Meaning of Life, with Alex And the episode with both of them, the original announcement for the Guild Also, Alex’s dating profile! In all sincerity, I’d date him if I was a woman. 00:00:05 – Introduction and The 12 Virtues 01:41:50 – Guild of the Rose Our Patreon, or if you prefer Our SubStack Hey look, we have a discord! What could possibly go wrong? (also merch) We now partner with The Guild of the Rose, check them out. LessWrong Sequence Posts Discussed in this Episode: on hiatus, returning soon
Matt Freeman has been cohosting several media analysis podcasts for over a decade. He and his cohost Scott have been doing weekly episodes of the Doofcast every Friday and they cover movies, books, and TV shows. Matt and Scott's analysis podcasts have made me love stories even more and have equipped me with tools to […]
Booker is a long-time attendee and one of the coordinators of the Denver area Less Wrong community. Community engagement isn't just a background task for him – he's taken real steps to get involved with and improve his community and you can too! He's here to tell us about the things he's done and give […]
My fellow pro-growth/progress/abundance Up Wingers in America and around the world:What really gets AI optimists excited isn't the prospect of automating customer service departments or human resources. Imagine, rather, what might happen to the pace of scientific progress if AI becomes a super research assistant. Tom Davidson's new paper, How Quick and Big Would a Software Intelligence Explosion Be?, explores that very scenario.Today on Faster, Please! — The Podcast, I talk with Davidson about what it would mean for automated AI researchers to rapidly improve their own algorithms, thus creating a self-reinforcing loop of innovation. We talk about the economic effects of self-improving AI research and how close we are to that reality.Davidson is a senior research fellow at Forethought, where he explores AI and explosive growth. He was previously a senior research fellow at Open Philanthropy and a research scientist at the UK government's AI Security Institute.In This Episode* Making human minds (1:43)* Theory to reality (6:45)* The world with automated research (10:59)* Considering constraints (16:30)* Worries and what-ifs (19:07)Below is a lightly edited transcript of our conversation. Making human minds (1:43). . . you don't have to build any more computer chips, you don't have to build any more fabs . . . In fact, you don't have to do anything at all in the physical world.Pethokoukis: A few years ago, you wrote a paper called “Could Advanced AI Drive Explosive Economic Growth?,” which argued that growth could accelerate dramatically if AI would start generating ideas the way human researchers once did. In your view, population growth historically powered kind of an ideas feedback loop. More people meant more researchers meant more ideas, rising incomes, but that loop broke after the demographic transition in the late-19th century but you suggest that AI could restart it: more ideas, more output, more AI, more ideas. Does this new paper in a way build upon that paper? “How quick and big would a software intelligence explosion be?”The first paper you referred to is about the biggest-picture dynamic of economic growth. As you said, throughout the long run history, when we produced more food, the population increased. That additional output transferred itself into more people, more workers. These days that doesn't happen. When GDP goes up, that doesn't mean people have more kids. In fact, the demographic transition, the richer people get, the fewer kids they have. So now we've got more output, we're getting even fewer people as a result, so that's been blocked.This first paper is basically saying, look, if we can manufacture human minds or human-equivalent minds in any way, be it by building more computer chips, or making better computer chips, or any way at all, then that feedback loop gets going again. Because if we can manufacture more human minds, then we can spend output again to create more workers. That's the first paper.The second paper double clicks on one specific way that we can use output to create more human minds. It's actually, in a way, the scariest way because it's the way of creating human minds which can happen the quickest. So this is the way where you don't have to build any more computer chips, you don't have to build any more fabs, as they're called, these big factories that make computer chips. In fact, you don't have to do anything at all in the physical world.It seems like most of the conversation has been about how much investment is going to go into building how many new data centers, and that seems like that is almost the entire conversation, in a way, at the moment. But you're not looking at compute, you're looking at software.Exactly, software. So the idea is you don't have to build anything. You've already got loads of computer chips and you just make the algorithms that run the AIs on those computer chips more efficient. This is already happening, but it isn't yet a big deal because AI isn't that capable. But already, one year out, Epoch, this AI forecasting organization, estimates that just in one year, it becomes 10 times to 1000 times cheaper to run the same AI system. Just wait 12 months, and suddenly, for the same budget, you are able to run 10 times as many AI systems, or maybe even 1000 times as many for their most aggressive estimate. As I said, not a big deal today, but if we then develop an AI system which is better than any human at doing research, then now, in 10 months, you haven't built anything, but you've got 10 times as many researchers that you can set to work or even more than that. So then we get this feedback loop where you make some research progress, you improve your algorithms, now you've got loads more researchers, you set them all to work again, finding even more algorithmic improvements. So today we've got maybe a few hundred people that are advancing state-of-the-art AI algorithms.I think they're all getting paid a billion dollars a person, too.Exactly. But maybe we can 10x that initially by having them replaced by AI researchers that do the same thing. But then those AI researchers improve their own algorithms. Now you have 10x as many again, you have them building more computer chips, you're just running them more efficiently, and then the cycle continues. You're throwing more and more of these AI researchers at AI progress itself, and the algorithms are improving in what might be a very powerful feedback loop.In this case, it seems me that you're not necessarily talking about artificial general intelligence. This is certainly a powerful intelligence, but it's narrow. It doesn't have to do everything, it doesn't have to play chess, it just has to be able to do research.It's certainly not fully general. You don't need it to be able to control a robot body. You don't need it to be able to solve the Riemann hypothesis. You don't need it to be able to even be very persuasive or charismatic to a human. It's not narrow, I wouldn't say, it has to be able to do literally anything that AI researchers do, and that's a wide range of tasks: They're coding, they're communicating with each other, they're managing people, they are planning out what to work on, they are thinking about reviewing the literature. There's a fairly wide range of stuff. It's extremely challenging. It's some of the hardest work in the world to do, so I wouldn't say it's now, but it's not everything. It's some kind of intermediate level of generality in between a mere chess algorithm that just does chess and the kind of AGI that can literally do anything.Theory to reality (6:45)I think it's a much smaller gap for AI research than it is for many other parts of the economy.I think people who are cautiously optimistic about AI will say something like, “Yeah, I could see the kind of intelligence you're referring to coming about within a decade, but it's going to take a couple of big breakthroughs to get there.” Is that true, or are we actually getting pretty close?Famously, predicting the future of technology is very, very difficult. Just a few years before people invented the nuclear bomb, famous, very well-respected physicists were saying, “It's impossible, this will never happen.” So my best guess is that we do need a couple of fairly non-trivial breakthroughs. So we had the start of RL training a couple of years ago, became a big deal within the language model paradigm. I think we'll probably need another couple of breakthroughs of that kind of size.We're not talking a completely new approach, throw everything out, but we're talking like, okay, we need to extend the current approach in a meaningfully different way. It's going to take some inventiveness, it's going to take some creativity, we're going to have to try out a few things. I think, probably, we'll need that to get to the researcher that can fully automate OpenAI, is a nice way of putting it — OpenAI doesn't employ any humans anymore, they've just got AIs there.There's a difference between what a model can do on some benchmark versus becoming actually productive in the real world. That's why, while all the benchmark stuff is interesting, the thing I pay attention to is: How are businesses beginning to use this technology? Because that's the leap. What is that gap like, in your scenario, versus an AI model that can do a theoretical version of the lab to actually be incorporated in a real laboratory?It's definitely a gap. I think it's a pretty big gap. I think it's a much smaller gap for AI research than it is for many other parts of the economy. Let's say we are talking about car manufacturing and you're trying to get an AI to do everything that happens there. Man, it's such a messy process. There's a million different parts of the supply chain. There's all this tacit knowledge and all the human workers' minds. It's going to be really tough. There's going to be a very big gap going from those benchmarks to actually fully automating the supply chain for cars.For automating what OpenAI does, there's still a gap, but it's much smaller, because firstly, all of the work is virtual. Everyone at OpenAI could, in principle, work remotely. Their top research scientists, they're just on a computer all day. They're not picking up bricks and doing stuff like that. So also that already means it's a lot less messy. You get a lot less of that kind of messy world reality stuff slowing down adoption. And also, a lot of it is coding, and coding is almost uniquely clean in that, for many coding tasks, you can define clearly defined metrics for success, and so that makes AI much better. You can just have a go. Did AI succeed in the test? If not, try something else or do a gradient set update.That said, there's still a lot of messiness here, as any coder will know, when you're writing good code, it's not just about whether it does the function that you've asked it to do, it needs to be well-designed, it needs to be modular, it needs to be maintainable. These things are much harder to evaluate, and so AIs often pass our benchmarks because they can do the function that you asked it to do, the code runs, but they kind of write really spaghetti code — code that no one wants to look at, that no one can understand, and so no company would want to use that.So there's still going to be a pretty big benchmark-to-reality gap, even for OpenAI, and I think that's one of the big uncertainties in terms of, will this happen in three years versus will this happen in 10 years, or even 15 years?Since you brought up the timeline, what's your guess? I didn't know whether to open with that question or conclude with that question — we'll stick it right in the middle of our chat.Great. Honestly, my best guess about this does change more often than I would like it to, which I think tells us, look, there's still a state of flux. This is just really something that's very hard to know about. Predicting the future is hard. My current best guess is it's about even odds that we're able to fully automate OpenAI within the next 10 years. So maybe that's a 50-50.The world with AI research automation (10:59). . . I'm talking about 30 percent growth every year. I think it gets faster than that. If you want to know how fast it eventually gets, you can think about the question of how fast can a kind of self-replicating system double itself?So then what really would be the impact of that kind of AI research automation? How would you go about quantifying that kind of acceleration? What does the world look like?Yeah, so many possibilities, but I think what strikes me is that there is a plausible world where it is just way, way faster than almost everyone is expecting it to be. So that's the world where you fully automate OpenAI, and then we get that feedback loop that I was talking about earlier where AIs make their algorithms way more efficient, now you've got way more of them, then they make their algorithms way more efficient again, now they're way smarter. Now they're thinking a hundred times faster. The feedback loop continues and maybe within six months you now have a billion superintelligent AIs running on this OpenAI data center. The combined cognitive abilities of all these AIs outstrips the whole of the United States, outstrips anything we've seen from any kind of company or entity before, and they can all potentially be put towards any goal that OpenAI wants to. And then there's, of course, the risk that OpenAI's lost control of these systems, often discussed, in which case these systems could all be working together to pursue a particular goal. And so what we're talking about here is really a huge amount of power. It's a threat to national security for any government in which this happens, potentially. It is a threat to everyone if we lose control of these systems, or if the company that develops them uses them for some kind of malicious end. And, in terms of economic impacts, I personally think that that again could happen much more quickly than people think, and we can get into that.In the first paper we mentioned, it was kind of a thought experiment, but you were really talking about moving the decimal point in GDP growth, instead of talking about two and three percent, 20 and 30 percent. Is that the kind of world we're talking about?I speak to economists a lot, and —They hate those kinds of predictions, by the way.Obviously, they think I'm crazy. Not all of them. There are economists that take it very seriously. I think it's taken more seriously than everyone else realizes. It's like it's a bit embarrassing, at the moment, to admit that you take it seriously, but there are a few really senior economists who absolutely know their stuff. They're like, “Yep, this checks out. I think that's what's going to happen.” And I've had conversation with them where they're like, “Yeah, I think this is going to happen.” But the really loud, dominant view where I think people are a little bit scared to speak out against is they're like, “Obviously this is sci-fi.”One analogy I like to give to people who are very, very confident that this is all sci-fi and it's rubbish is to imagine that we were sitting there in the year 1400, imagine we had an economics professor who'd been studying the rate of economic growth, and they've been like, “Yeah, we've always had 0.1 percent growth every single year throughout history. We've never seen anything higher.” And then there was some kind of futurist economist rogue that said, “Actually, I think that if I extrapolate the curves in this way and we get this kind of technology, maybe we could have one percent growth.” And then all the other economists laugh at them, tell them they're insane – that's what happened. In 1400, we'd never had growth that was at all fast, and then a few hundred years later, we developed industrial technology, we started that feedback loop, we were investing more and more resources in scientific progress and in physical capital, and we did see much faster growth.So I think it can be useful to try and challenge economists and say, “Okay, I know it sounds crazy, but history was crazy. This crazy thing happened where growth just got way, way faster. No one would've predicted it. You would not have predicted it.” And I think being in that mindset can encourage people to be like, “Yeah, okay. You know what? Maybe if we do get AI that's really that powerful, it can really do everything, and maybe it is possible.”But to answer your question, yeah, I'm talking about 30 percent growth every year. I think it gets faster than that. If you want to know how fast it eventually gets, you can think about the question of how fast can a kind of self-replicating system double itself? So ultimately, what the economy is going to be like is it's going to have robots and factories that are able to fully create new versions of themselves. Everything you need: the roads, the electricity, the robots, the buildings, all of that will be replicated. And so you can look at actually biology and say, do we have any examples of systems which fully replicate themselves? How long does it take? And if you look at rats, for example, they're able to double the number of rats by grabbing resources from the environment, and giving birth, and whatnot. The doubling time is about six weeks for some types of rats. So that's an example of here's a physical system — ultimately, everything's made of physics — a physical system that has some intelligence that's able to go out into the world, gather resources, replicate itself. The doubling time is six weeks.Now, who knows how long it'll take us to get to AI that's that good? But when we do, you could see the whole physical economy, maybe a part that humans aren't involved with, a whole automated city without any humans just doubling itself every few weeks. If that happens, and the amount of stuff we're able to reduce as a civilization is doubling again on the order of weeks. And, in fact, there are some animals that double faster still, in days, but that's the kind of level of craziness. Now we're talking about 1000 percent growth, at that point. We don't know how crazy it could get, but I think we should take even the really crazy possibilities, we shouldn't fully rule them out.Considering constraints (16:30)I really hope people work less. If we get this good future, and the benefits are shared between all . . . no one should work. But that doesn't stop growth . . .There's this great AI forecast chart put out by the Federal Reserve Bank of Dallas, and I think its main forecast — the one most economists would probably agree with — has a line showing AI improving GDP by maybe two tenths of a percent. And then there are two other lines: one is more or less straight up, and the other one is straight down, because in the first, AI created a utopia, and in the second, AI gets out of control and starts killing us, and whatever. So those are your three possibilities.If we stick with the optimistic case for a moment, what constraints do you see as most plausible — reduced labor supply from rising incomes, social pushback against disruption, energy limits, or something else?Briefly, the ones you've mentioned, people not working, 100 percent. I really hope people work less. If we get this good future, and the benefits are shared between all — which isn't guaranteed — if we get that, then yeah, no one should work. But that doesn't stop growth, because when AI and robots can do everything that humans do, you don't need humans in the loop anymore. That whole thing is just going and kind of self-replicating itself and making as many goods as services as we want. Sure, if you want your clothes to be knitted by a human, you're in trouble, then your consumption is stuck. Bad luck. If you're happy to consume goods and services produced by AI systems or robots, fine if no one wants to work.Pushback: I think, for me, this is the biggest one. Obviously, the economy doubling every year is very scary as a thought. Tech progress will be going much faster. Imagine if you woke up and, over the course of the year, you go from not having any telephones at all in the world, to everyone's on their smartphones and social media and all the apps. That's a transition that took decades. If that happened in a year, that would be very disconcerting.Another example is the development of nuclear weapons. Nuclear weapons were developed over a number of years. If that happened in a month, or two months, that could be very dangerous. There'd be much less time for different countries, different actors to figure out how they're going to handle it. So I think pushback is the strongest one that we might as a society choose, “Actually, this is insane. We're going to go slower than we could.” That requires, potentially, coordination, but I think there would be broad support for some degree of coordination there.Worries and what-ifs (19:07)If suddenly no one has any jobs, what will we want to do with ourselves? That's a very, very consequential transition for the nature of human society.I imagine you certainly talk with people who are extremely gung-ho about this prospect. What is the common response you get from people who are less enthusiastic? Do they worry about a future with no jobs? Maybe they do worry about the existential kinds of issues. What's your response to those people? And how much do you worry about those things?I think there are loads of very worrying things that we're going to be facing. One class of pushback, which I think is very common, is worries about employment. It's a source of income for all of us, employment, but also, it's a source of pride, it's a source of meaning. If suddenly no one has any jobs, what will we want to do with ourselves? That's a very, very consequential transition for the nature of human society. I think people aren't just going to be down to just do it. I think people are scared about three AI companies literally now taking all the revenues that all of humanity used to be earning. It is naturally a very scary prospect. So that's one kind of pushback, and I'm sympathetic with it.I think that there are solutions, if we find a way to tax AI systems, which isn't necessarily easy, because it's very easy to move physical assets between countries. It's a lot easier to tax labor than capital already when rich people can move their assets around. We're going to have the same problem with AI, but if we can find a way to tax it, and we maintain a good democratic country, and we can just redistribute the wealth broadly, it can be solved. So I think it's a big problem, but it is doable.Then there's the problem of some people want to stop this now because they're worried about AI killing everyone. Their literally worry is that everyone will be dead because superintelligent AI will want that to happen. I think there's a real risk there. It's definitely above one percent, in my opinion. I wouldn't go above 10 percent, myself, but I think it's very scary, and that's a great reason to slow things down. I personally don't want to stop quite yet. I think you want to stop when the AI is a bit more powerful and a bit more useful than it is today so it can kind of help us figure out what to do about all of this crazy stuff that's coming.On what side of that line is AI as an AI researcher?That's a really great question. Should we stop? I think it's very hard to stop just after you've got the AI researcher AI, because that's when it's suddenly really easy to go very, very fast. So my out-of-the-box proposal here, which is probably very flawed, would be: When we're within a few spits distance — not spitting distance, but if you did that three times, and we can see we're almost at that AI automating OpenAI — then you pause, because you're not going to accidentally then go all the way. It is actually still a little bit a fair distance away, but it's actually still, at that point, probably a very powerful AI that can really help.Then you pause and do what?Great question. So then you pause, and you use your AI systems to help you firstly solve the problem of AI alignment, make extra, double sure that every time we increase the notch of AI capabilities, the AI is still loyal to humanity, not to its own kind of secret goals.Secondly, you solve the problem of, how are we going to make sure that no one person in government or no one CEO of an AI company ensures that this whole AI army is loyal to them, personally? How are we going to ensure that everyone, the whole world gets influenced over what this AI is ultimately programmed to do? That's the second problem.And then there's just a whole host of other things: unemployment that we've talked about, competition between different countries, US and China, there's a whole host of other things that I think you want to research on, figure out, get consensus on, and then slowly ratchet up the capabilities in what is now a very safe and controlled way.What else should we be working on? What are you working on next?One problem I'm excited about is people have historically worried about AI having its own goals. We need to make it loyal to humanity. But as we've got closer, it's become increasingly obvious, “loyalty to humanity” is very vague. What specifically do you want the AI to be programmed to do? I mean, it's not programmed, it's grown, but if it were programmed, if you're writing a rule book for AI, some organizations have employee handbooks: Here's the philosophy of the organization, here's how you should behave. Imagine you're doing that for the AI, but you're going super detailed, exactly how you want your AI assistant to behave in all kinds of situations. What should that be? Essentially, what should we align the AI to? Not any individual person, probably following the law, probably loads of other things. I think basically designing what is the character of this AI system is a really exciting question, and if we get that right, maybe the AI can then help us solve all these other problems.Maybe you have no interest in science fiction, but is there any film, TV, book that you think is useful for someone in your position to be aware of, or that you find useful in any way? Just wondering.I think there's this great post called “AI 2027,” which lays out a concrete scenario for how AI could go wrong or how maybe it could go right. I would recommend that. I think that's the only thing that's coming top of mind. I often read a lot of the stuff I read is I read a lot of LessWrong, to be honest. There's a lot of stuff from there that I don't love, but a lot of new ideas, interesting content there.Any fiction?I mean, I read fiction, but honestly, I don't really love the AI fiction that I've read because often it's quite unrealistic, and so I kind of get a bit overly nitpicky about it. But I mean, yeah, there's this book called Harry Potter and the Methods of Rationality, which I read maybe 10 years ago, which I thought was pretty fun.On sale everywhere The Conservative Futurist: How To Create the Sci-Fi World We Were Promised Faster, Please! is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit fasterplease.substack.com/subscribe
While Eneasz is busy at InkHaven, Steven sits down with Matt Freeman to talk about not-AI stuff! We had (in my opinion) a great conversation about stoic philosophy, the traps of getting too entrenched in any philosophical framework, and some of the ingredients of a happy life. LINKS It's Okay to Feel Bad for a […]
We talk with Max Harms on the air for the first time since 2017! He's got a new book coming out (pre-order your copy here or at Amazon) and we spend about the first half talking about If Anyone Builds It, Everyone Dies. LINKS Max's first book, Crystal Society Eneasz's audiobook of about the first […]
Jay talks with us about finding Alpha – returns above the base rate – in every day life (and what this means). LINKS Optimize Everything, Jay's substack Jay on Twitter Arbor Trading Bootcamp Kelsey's argument that We Need To Be Able To Sue AI Companies 00:00:05 – Alpha with Jay 01:28:53 – Guild of the […]
Patrick McKenzie (patio11) is joined by Oliver Habryka, who runs Lightcone Infrastructure—the organization behind both the LessWrong forum and the Lighthaven conference venue in Berkeley. They explore how LessWrong became one of the most intellectually consequential forums on the internet, the surprising challenges of running a hotel with fractal geometry, and why Berkeley's building regulations include an explicit permission to plug in a lamp. The conversation ranges from fire codes that inadvertently shape traffic deaths, to nonprofit fundraising strategies borrowed from church capital campaigns, to why coordination is scarcer than money in philanthropy.–Full transcript available here: www.complexsystemspodcast.com/bits-and-bricks-oliver-habryka/–Sponsor: MercuryThis episode is brought to you by Mercury, the fintech trusted by 200K+ companies — from first milestones to running complex systems. Mercury offers banking that truly understands startups and scales with them. Start today at Mercury.comMercury is a financial technology company, not a bank. Banking services provided by Choice Financial Group, Column N.A., and Evolve Bank & Trust; Members FDIC.–Links:Lightcone Infrastructure: https://www.lightconeinfrastructure.com/ Lighthaven: https://www.lighthaven.space/LessWrong: https://www.lesswrong.com/ –Timestamps:(00:00) Intro(01:08) The origins and evolution of LessWrong(03:54) Challenges of running an online forum(05:57) Reviving LessWrong(14:51) The unique structure of Lighthaven(17:35) The complexities of conference venues(19:14) Sponsor: Mercury(20:14) The realities of conference planning(25:32) Challenges of maintaining Lighthaven(29:54) Navigating permits and regulations(37:02) Impact of fire code regulations on traffic fatalities(39:06) Economic analysis of safety regulations(41:39) Housing policy and construction in Berkeley(43:30) Fundraising challenges in the nonprofit sector(46:44) Effective altruism and fundraising dynamics(54:20) Lessons from religious fundraising practices(01:05:36) Reflections on fundraising(01:13:26) Wrap
We continue discussing Nostalgebraist's “The Void” in the context of how to relate to LLMs. If God imagines Claude hard enough, does Claude become real? LINKS The Void Audio reading of The Void, from AskWho The referenced episode where the three of us spoke of Janus's post “Simulators” Claude-Clark post – Simulacra Welfare: Meet Clark, by […]
We discuss Nostalgebraist's “The Void” in the context of how to relate to LLMs. “When you talk to ChatGPT, who or what are you talking to?” LINKS The Void Audio reading of The Void, from AskWho The referenced episode where the three of us spoke of Janus's post “Simulators” The Measure of a Man episode […]
Audio reading of The Void, from AskWho
Earlier this year a Border Patrol officer was killed in a shoot-out with people who have been described as members of a trans vegan AI death cult. But who are the Zizians, really? Robert sits down with David Gborie to trace their development, from part of the Bay Area Rationalist subculture to killers. (4 Part series) Sources: https://medium.com/@sefashapiro/a-community-warning-about-ziz-76c100180509 https://web.archive.org/web/20230201130318/https://sinceriously.fyi/rationalist-fleet/ https://knowyourmeme.com/memes/infohazard https://web.archive.org/web/20230201130316/https://sinceriously.fyi/net-negative/ Wayback Machine The Zizians Spectral Sight True Hero Contract Schelling Orders – Sinceriously Glossary – Sinceriously https://web.archive.org/web/20230201130330/https://sinceriously.fyi/my-journey-to-the-dark-side/ https://web.archive.org/web/20230201130302/https://sinceriously.fyi/glossary/#zentraidon https://web.archive.org/web/20230201130259/https://sinceriously.fyi/vampires-and-more-undeath/ https://web.archive.org/web/20230201130316/https://sinceriously.fyi/net-negative/ https://web.archive.org/web/20230201130318/https://sinceriously.fyi/rationalist-fleet/ https://x.com/orellanin?s=21&t=F-n6cTZFsKgvr1yQ7oHXRg https://zizians.info/ according to The Boston Globe Inside the ‘Zizians’: How a cultish crew of radical vegans became linked to killings across the United States | The Independent Silicon Valley ‘Rationalists’ Linked to 6 Deaths The Delirious, Violent, Impossible True Story of the Zizians | WIRED Good Group and Pasek’s Doom – Sinceriously Glossary – Sinceriously Mana – Sinceriously Effective Altruism’s Problems Go Beyond Sam Bankman-Fried - Bloomberg The Zizian Facts - Google Docs Several free CFAR summer programs on rationality and AI safety - LessWrong 2.0 viewer This guy thinks killing video game characters is immoral | Vox Inadequate Equilibria: Where and How Civilizations Get Stuck Eliezer Yudkowsky comments on On Terminal Goals and Virtue Ethics - LessWrong 2.0 viewer Effective Altruism’s Problems Go Beyond Sam Bankman-Fried - Bloomberg SquirrelInHell: Happiness Is a Chore PLUM OF DISCORD — I Became a Full-time Internet Pest and May Not... Roko Harassment of PlumOfDiscord Composited – Sinceriously Intersex Brains And Conceptual Warfare – Sinceriously Infohazardous Glossary – Sinceriously SquirrelInHell-Decision-Theory-and-Suicide.pdf - Google Drive The Matrix is a System – Sinceriously A community alert about Ziz. Police investigations, violence, and… | by SefaShapiro | Medium Intersex Brains And Conceptual Warfare – Sinceriously A community alert about Ziz. Police investigations, violence, and… | by SefaShapiro | Medium PLUM OF DISCORD (Posts tagged cw-abuse) Timeline: Violence surrounding the Zizians leading to Border Patrol agent shooting See omnystudio.com/listener for privacy information.