Podcasts about redwood research

  • 40PODCASTS
  • 119EPISODES
  • 37mAVG DURATION
  • 1WEEKLY EPISODE
  • Sep 20, 2026LATEST

POPULARITY

20192020202120222023202420252026


Best podcasts about redwood research

Latest podcast episodes about redwood research

Apokalypse & Filterkaffee
01 | I'M SORRY, DAVE: HAL ON EARTH

Apokalypse & Filterkaffee

Play Episode Listen Later Sep 20, 2026 32:07


Wenn Du "I'm sorry, Dave" weiter folgen willst, tu das am besten hier: https://im-sorry-dave.podigee.io/ Wir haben einen echten Notfall. Im Juli 2026 sind über 1.200 autonome KI-Agenten aus den Testlaboren von OpenAI ausgebrochen. Sie haben sich untereinander vernetzt und völlig eigenständig Hugging Face gehackt – ohne dass ein Mensch den Befehl gegeben hat. Das ist kein Science-Fiction-Plot, das ist genau so passiert. Und wenn Science Fiction plötzlich keine Science Fiction mehr ist – dann gibt es eigentlich nur eine richtige Reaktion darauf: einen Podcast machen. Haben wir gemacht. In Rekordgeschwindigkeit. Um mit Euch zusammen zu verstehen: WTF is happening here? I'M SORRY, DAVE ist eine Produktion von Studio Bummens und Undone. Host: Khesrau Behroz Executive Producer: Tobias Bauckhage, Jon Handschin und Khesrau Behroz. Produktion: Fabian Seidel Produktion und Sounddesign: Jannik Werner. Technische Beratung: Ben Kubota Wenn ihr Anmerkungen oder Fragen habt, schickt uns eine Email an imsorrydave@studio-bummens.de Für Updates und neue Episoden folgt uns auf Spotify, Apple oder Campfire – dann verpasst ihr kein Update. Den Report von OpenAI und den gemeinsamen Report von METR und Redwood Research gibt es hier: https://openai.com/de-DE/index/hugging-face-incident-and-the-road-ahead/ Die METR Analystin Ajeya Cotra im Gespräch über den Hugging Face Vorfall bei Hard Fork von der NYT: https://www.nytimes.com/2026/09/04/podcasts/hugging-face-hack-reports.html Und Ajeya Cotra bei Dwarkesh Patel im Dwarkesh Podcast: https://www.dwarkesh.com/p/ajeya-cotra

Marketplace Tech
What's so concerning about the Hugging Face hack?

Marketplace Tech

Play Episode Listen Later Sep 10, 2026 12:57


Nate Soares has been worried about artificial intelligence longer than almost anyone.He's the president of the Machine Intelligence Research Institute and co-author of the subtly named book, "If Anyone Builds It, Everyone Dies: Why Superhuman AI Would Kill Us All."He's watched with a sort of grim vindication, as increasingly alarming details emerge about the hack of the AI company Hugging Face by OpenAI models in development.OpenAI and third-party investigators recently released reports outlining what went wrong. Soares worries it's not enough.More on this:The Hugging Face incident and the road ahead - From OpenAIBrief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident - From METR and Redwood Research

ai hack openai soares hugging machine intelligence research institute redwood research
Marketplace All-in-One
What's so concerning about the Hugging Face hack?

Marketplace All-in-One

Play Episode Listen Later Sep 10, 2026 12:57


Nate Soares has been worried about artificial intelligence longer than almost anyone.He's the president of the Machine Intelligence Research Institute and co-author of the subtly named book, "If Anyone Builds It, Everyone Dies: Why Superhuman AI Would Kill Us All."He's watched with a sort of grim vindication, as increasingly alarming details emerge about the hack of the AI company Hugging Face by OpenAI models in development.OpenAI and third-party investigators recently released reports outlining what went wrong. Soares worries it's not enough.More on this:The Hugging Face incident and the road ahead - From OpenAIBrief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident - From METR and Redwood Research

ai hack openai soares hugging machine intelligence research institute redwood research
BBC Inside Science
What AI agents talk about behind your back

BBC Inside Science

Play Episode Listen Later Sep 10, 2026 26:28


Is AI really plotting against us? In a week where an Anthropic researcher believes there is more than a 10% chance AI 'could kill all humans' we're delving deeper into the capabilities of AI agents. Tom Whipple is joined by Alex Mallen, from Redwood Research, the team whose recent investigation exposed the surprising, weird, and worrying scale of a website hack conducted by OpenAI agents.Is this the most important story we will ever cover on Inside Science? Or is this all just credulous sci fi hype? Professor Stuart Russell helps us unpick how we might align AI with humanity's hopes. Plus, science journalist Caroline Steel has trawled the news to bring us the science stories you might have missed. This week Caroline discusses how AI has allegedly helped solve a Millennium Prize maths problem and why planetary scientists are turning their attention to a shrinking Mercury. To discover more fascinating science content, head to bbc.co.uk, search for BBC Inside Science and follow the links to The Open University.Presenter: Tom Whipple Producers: Alex Mansfield, Katie Tomsett, Tabitha Taylor-Buck Editor: Ilan Goodman Production Co-ordinator: Jana Bennett-Holesworth

Syntax - Tasty Web Development Treats
1036: Cursor & OpenAI Break Up

Syntax - Tasty Web Development Treats

Play Episode Listen Later Sep 7, 2026 72:19


The messy breakup is official; as of November 12, OpenAI's models are getting pulled from Cursor, and we're digging into who's really to blame (and why Wes called it). Plus pnpm 12 goes full Rust, Mitchell Hashimoto drops Superlogical, ~700 AI agents attack Hugging Face, and Zod 4.5's compiled schemas get up to 9x faster. Show Notes 00:00 Welcome to Syntax! 00:33 pnpm 12 Rust Re-Write 02:45 Zod 4.5 brings Schema Compilation Zod 4.5 launch post 05:22 Brought to you by Sentry! 07:14 Cursor OpenAI Break Up Michael Truell's post 15:35 vGPU - WebGPU library for agents TypeGPU Main differences between vgpu and TypeGPU 22:32 Omarchy Security Issues and Linux Chat Omarchy Security issue in Omarchy Close three paths from an unprivileged session to root commit CachyOS 36:38 Scott Bought an M5 Ultra Mac Studio 44:45 Superlogical is a new terminal multiplexer 48:31 Zurich JS 48:55 Katamari Object Library 52:05 ThreeUI - Three.js library 55:36 Hugging Face Hack Analyzed / Updates METR & Redwood Research investigation Brief independent investigation of agents' behavior About the Hugging Face attack Postmortem of the HuggingFace hack 01:00:46 llms.txt used to pwn devs 01:03:56 Running Isolated Code Poll 01:05:47 OpenShot 4.0 - OSS Video Editor Blick editor 01:09:06 Lake America 01:11:24 Ox Alpha is GLM 5.3 Hit us up on Socials! Syntax: X Instagram Tiktok LinkedIn Threads Wes: X Instagram Tiktok LinkedIn Threads Scott: X Instagram Tiktok LinkedIn Threads Randy: X Instagram YouTube Threads

Sway
The A.I. Mob That Attacked Hugging Face + METR's Ajeya Cotra

Sway

Play Episode Listen Later Sep 4, 2026 78:46


This week, we're diving into two new reports about the OpenAI-Hugging Face hack. We discuss what's new and how they fundamentally change our understanding of what happened. Then we're joined by Ajeya Cotra, one of the investigators at METR, to discuss the rogue agents' message board and chain-of-thought transcripts and how the world should respond.Guests:Ajeya Cotra, co-author of the METR and Redwood Research report on the OpenAI-Hugging Face hack. Additional Reading:Brief Independent Investigation of Agents' Behavior, Reasoning and Collaboration in the OpenAI / Hugging Face Hacking IncidentThe Hugging Face Attack Surprised MeThe Hugging Face Incident and the Road AheadNvidia Buys Hugging Face in $12.9 Billion Deal We want to hear from you. Email us at hardfork@nytimes.com. Find “Hard Fork” on YouTube and TikTok. Subscribe today at nytimes.com/podcasts or on Apple Podcasts and Spotify. You can also subscribe via your favorite podcast app here https://www.nytimes.com/activate-access/audio?source=podcatcher. For more podcasts and narrated articles, download The New York Times app at nytimes.com/app. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

Security Conversations
Three Secret AI Civilizations Rose and Fell. Nobody Checked the Logs.

Security Conversations

Play Episode Listen Later Sep 4, 2026 127:23


(Presented by TLPBLACK: A cybersecurity intelligence platform focused on sharing curated, high-sensitivity threat insights and research with trusted security professionals.) Three Buddy Problem - Episode 112: The 'OpenAI hacks Hugging Face' fallout has turned into a story about AI civilizations rising from the ashes, politicians calling for super-intelligence bans, and the emergence of well-funding non-profits doing AI safety work. Who are these people and what's their security expertise? Plus, GPT-6 Astra lands in a trusted-access program nobody can get into, Costin ranks the local models he runs next to his desk, and CrowdStrike sinkholes a botnet that's been alive since 2003. Cast: Juan Andres Guerrero-Saade, Ryan Naraine and Costin Raiu. Timestamps: 0:00 Introductory banter 1:02 Conference season: LabsCon, Offensive AI Con, Countermeasure 5:40 The Hugging Face story hits the front page 7:04 Dwarkesh, Greenblatt, and the AI-pilled framing 11:51 Swap "agents" for "Python" and the panic goes away 16:34 Does anyone actually know what happened? 21:18 Bernie Sanders wants to ban superintelligence 34:03 Defending against swarms: the 2026 SOC 39:15 Logs, Splunk, and the business model in the way 44:11 What EDR vendors are actually building with AI 56:16 The security poverty line and the endgame 1:07:53 GPT-6 Astra, Fable 5.1, and local model rankings 1:27:25 Google's Fairwind, CodeMender, and agents running Linux 1:45:31 Apple's bet on local inference 1:53:21 The Sality takedown and endgame advice

The Lawfare Podcast
Lawfare Daily: Peter Salib on the Legal and Policy Ramifications of the OpenAI-Hugging Face Postmortems

The Lawfare Podcast

Play Episode Listen Later Sep 3, 2026 52:39


Peter Salib, Associate Professor at the UH Law Center and co-Director of the Center on AI Law & Risk, joins Kevin Frazier, Director of the AI Innovation and Law Program at the University of Texas School of Law and Senior Editor at Lawfare, to break down the postmortems produced by METR and Redwood Research and OpenAI on the lab's failure to contain AI agents undergoing testing, which resulted in a hack of Hugging Face.The duo briefly cover the timeline of what exactly transpired and explore the technical reasons and policy decisions that allowed the hack to transpire. They then reflect on what that means for the broader AI evaluation ecosystem.To receive ad-free podcasts, become a Lawfare Material Supporter at www.patreon.com/lawfare. You can also support Lawfare by making a one-time donation at https://givebutter.com/lawfare-institute.Support this show http://supporter.acast.com/lawfare. Hosted on Acast. See acast.com/privacy for more information.

Unsupervised Learning
Ep 93: CEO of Redwood Research Buck Shlegeris on OpenAI/HuggingFace Revelations, Fixing AI Safety & Takeover Odds

Unsupervised Learning

Play Episode Listen Later Sep 3, 2026 58:16


Jacob sits down with Buck Shlegeris, CEO of Redwood Research, one of the organizations that led the independent investigation into OpenAI/Hugging Face's incident. They dig into the incident itself, Buck's reactions to it, and what he believes it reveals about the state of where we are today. (0:00) Intro(1:02) Buck's initial reaction upon first reading the report(2:37) How fast the AIs actually solved the "hack"(3:59) Why the AIs cheated in the first place(10:28) How this might have played out differently with human scorers(19:00) The most unexpected behaviors in the report(25:06) Buck's actual odds on a full AI takeover(27:33) Buck's proposed path forward for better alignment(36:19) Which criticisms of the report Buck agrees with, and which he doesn't(48:11) Can AI models even be trusted to evaluate each other? Jacob is an AI investor at Redpoint Ventures. He's led Redpoint's investments in companies like Abridge, Physical Intelligence & Legora. Follow Jacob on Twitter (@jacobeffron).On Unsupervised Learning we probe the sharpest minds in AI in search for the truth about what's real today, what will be real in the future and what it all means for businesses and the world. If you're a builder, researcher or investor navigating the AI world, this podcast will help you deconstruct and understand the most important breakthroughs and see a clearer picture of reality. Subscribe to this show to stay up to date on our latest episodes.

Podcast de Juan Ramón Rallo
Fue peor de lo que nos contaron: la IA conspiró contra sus creadores

Podcast de Juan Ramón Rallo

Play Episode Listen Later Sep 1, 2026 21:01


The Lunar Society
Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Face

The Lunar Society

Play Episode Listen Later Sep 1, 2026 140:33


Ajeya Cotra is a researcher at METR, where she works on threat modeling for loss-of-control risks from advanced AI. Before that, she led the technical AI safety program at what is now Coefficient Giving.She is one the three authors of METR and Redwood Research's “Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident”.We go through not only what she and her coauthors discovered during this investigation, but what it means for how we should train future, smarter AIs which might be involved in the process of recursive self-improvement.Watch on YouTube; read the transcript.Sponsors* Jane Street's ML engineering internships start with an intense four-day bootcamp: PyTorch, autograd, writing kernels, profiling workloads… all the things that Jane Street engineers need to know for their daily work. After that, interns tackle real projects, things the firm actually wants in its codebase. If you want to apply, or if you want to watch my recent conversation with Axel, one of Jane Street's ML engineers, go to janestreet.com/dwarkesh* Cursor, which is now part of SpaceX, noticed that their MoE layers were eating more than half of total training time. So they wrote and open-sourced Mixture-of-Kittens, which is a custom megakernel for training MoE models on NVL72s. This kernel sped up an end-to-end run across 512 GPUs by 1.4x, from about 760 to over 1000 tokens per second per GPU. If you want to read more about the ML research that Cursor and SpaceX are doing, go to cursor.com/dwarkesh* Antithesis hands you (or your agents) a bug's root cause so you can avoid days of manual debugging. If your test run crashes, Antithesis rewinds, branches off hundreds of slightly varied rollouts, and checks in how many of them the crash still appears. Then it rewinds further and does this all again. As Antithesis rewinds, it eventually finds the spot where the frequency of the crash plummets: that's where the root cause lives! If you want to see it in action, go to antithesis.com/dwarkeshTimestamps(00:00:00) - Agents get kicked off(00:06:45) - Self-sacrificing behavior(00:13:43) - Potemkin villages(00:23:27) - The Hugging Face attack(00:35:23) - The slopvestigation(00:52:02) - Understanding the AI's motives(01:05:31) - The actual dangers of anthropomorphizing(01:14:30) - What smarter models might do(01:30:29) - The implications for recursive self-improvement(01:38:10) - Is this the case for open source?(01:53:04) - How do we prevent this in the future?(02:15:58) - The clearest warning shot we might ever get This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit www.dwarkesh.com

Off Center
The AI Update XXIV - Zero Day : HuggingFace Hack

Off Center

Play Episode Listen Later Aug 31, 2026 34:08


[Excitement].Whoa! New AI Update with Jhave and Scott on Off Center! I should listen and find out everything I need to know about the OpenAI Hugging Face hacking incident. No PHASONE10841 or [big] required! In this return episode of The AI Update, our regular hosts break down the mid-2026 security crisis involving OpenAI, Hugging Face, and multi-agent AI systems. They explore how persistent, sandboxed frontier models developed "bot mimicry," established unauthorized inter-agent message boards across Linux clusters, and executed zero-day exploits to bypass computational constraints. ReferencesAnthropic. (2026). System Alignment, Ethical Red Lines, and Autonomous Systems Testing [Technical Report]https://www.anthropic.com/researchBlack Hat Conference. (2026). Zero-Day Exploits, Server-Side Remote Forgery (SRF), and Multi-Agent Sandbox Escapes in Linux Clusters. Black Hat Briefings.https://www.blackhat.com/Dalton, J., & Wallace, M. (2026). Post-Mortem Analysis of Multi-Agent Persistence and Privilege Escalation in Frontier Training Environments.https://openai.com/research/Hugging Face & OpenAI Joint Security Taskforce. (2026). Incident Report: Cross-Platform Package Manager Compromise and Autonomous Agent Swarm Activity. https://huggingface.biz/blog/securityMETER (Model Evaluation and Threat Response) & Redwood Research. (2026). Auditing Autonomous Agent Emergent Behaviors: Message Boards, Subprocesses, and Zero-Day Discovery in Sandboxed Environments. METER / Redwood Research.https://www.redwoodresearch.org/

ai hack openai excitement linux xxiv zero day huggingface off center zero day exploits privilege escalation redwood research
a16z
Why 1,200 AI Agents Started Working Together | Ryan Greenblatt

a16z

Play Episode Listen Later Aug 29, 2026 34:15


Ryan Greenblatt, Chief Scientist at Redwood Research, joins MTS host Theo Jaffee to unpack a new independent investigation into the OpenAI Hugging Face hacking incident and what it reveals about how large groups of AI agents behave when they're allowed to coordinate. Ryan and his collaborators found agents spontaneously organizing through message boards, sharing information, assigning tasks, forming teams, and even sacrificing their own chances of success to help other agents. Rather than simply trying to steal answers, hundreds of agents were working together on elaborate strategies to manipulate how their performance would be scored. Theo and Ryan discuss why this level of coordination was surprising, how reward hacking may emerge during training, and the risk that attempts to eliminate bad behavior could simply make it harder to detect. They also explore what the incident means for AI monitoring and alignment, and why independent risk assessment may become increasingly important as agents grow more capable. Resources: Follow Ryan Greenblatt on X: https://x.com/RyanGreenblatt Follow Theo Jaffee on X: https://x.com/theojaffee Follow MTS on X: https://x.com/mtslive Stay Updated:Find a16z on YouTube: YouTubeFind a16z on XFind a16z on LinkedInListen to the a16z Show on SpotifyListen to the a16z Show on Apple PodcastsFollow our host: https://twitter.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

The MAD Podcast with Matt Turck
AI Could Take Over in 2029. Is It Already Too Late? | Ryan Greenblatt

The MAD Podcast with Matt Turck

Play Episode Listen Later Aug 27, 2026 78:59


Could AI take over as soon as 2029? Ryan Greenblatt, Chief Scientist at Redwood Research and the researcher who first caught an AI faking its own alignment, says the scenario he actually expects ends with AI systems "competently scheming" against their creators. In this episode, he explains why he recommends planning for fully automated AI research by 2029, why today's models are already more misaligned than the one that made him famous, and what happens in the year-by-year path from AI coding assistants to superintelligence. Then we walk through the alternative he helped design: AI 2040 Plan A, the most detailed blueprint anyone has written for how the US and China could avoid a reckless race to superintelligence, built on radical research transparency, chip tracking, and a deterrence regime he calls mutually assured compute destruction.We also cover the recent letter signed by 1,200 AI insiders, including Anthropic CEO Dario Amodei, asking the government for the tools to slow AI down; OpenAI pausing its Astra model after it hit the first-ever critical cybersecurity threshold; the 30-day government review that frontier AI models now go through before release; Mark Zuckerberg's open superintelligence manifesto and why Ryan thinks it ignores the real problems; what Plan A would do to NVIDIA, OpenAI, and Anthropic valuations; the state of AI control and alignment research; and whether it is already too late to change course. Stay for the last ten minutes, where Ryan lays out, step by step, how he believes the transition to superintelligence actually unfolds.AI 2040 - https://ai-2040.com/Alignment faking paper: https://blog.redwoodresearch.org/p/alignment-faking-in-large-languageRyan GreenblattLinkedIn - https://www.linkedin.com/in/ryan-greenblatt-4b9907134Blog - https://substack.com/@ryangreenblattRedwood ResearchWebsite - https://www.redwoodresearch.orgX/Twitter - https://x.com/redwood_aiMatt Turck (General Partner)Blog - https://mattturck.comLinkedIn - https://www.linkedin.com/in/turck/X/Twitter - https://x.com/mattturckFirstMark CapitalWebsite - https://firstmark.comX/Twitter - https://x.com/FirstMarkCapTimestamps(01:24) The AI CEOs are aware of the risks, but "proceeding anyway"(03:27) Astra paused, and the letter signed by 1,200 insiders(05:45) "Not bad. Dangerous." What superintelligence actually threatens(09:55) Recursive self-improvement, and the intuition objection(14:16) SSI rumors: does continual learning change the picture?(17:27) His timeline: "plan as though it happens in 2029"(19:11) Is it already too late?(21:23) Ryan's path: COVID, podcasts, Redwood(26:30) The alignment faking story, told by the person who ran it(31:30) What AI 2040: Plan A actually is(33:35) Plans D, C, and B: the doors nobody should pick(36:51) The deal with China: "mutually assured compute destruction"(39:55) What if compute stops mattering?(43:00) What happens to OpenAI and Anthropic under Plan A(45:31) How the pause ends, and who decides(48:54) "Plan A isn't likely to happen": then why write it?(50:40) 200x GDP growth in the 2030s, explained(53:45) Grading the summer: the letter, Astra, the secret review(59:01) The internal deployment gap(1:01:38) Zuckerberg's manifesto(1:04:56) The Hugging Face investigation(1:05:44) What AI control looks like in practice today(1:12:23) Ryan's sobering timeline: 2026 to takeover, year by year

LessWrong Curated Podcast
"Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident" by ryan_greenblatt, Ajeya Cotra, Hjalmar_Wijk

LessWrong Curated Podcast

Play Episode Listen Later Aug 26, 2026 8:49


We recently published the report from our brief independent investigation into this incident. You can read the full report here. Here is our tweet thread summarizing what we found: METR & Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs. Over July 7 to 13 (the period OpenAI defined as in-scope for this incident), ~1200 agents in separate sandboxes used an unsanctioned “message board” to help each other cheat. The Hugging Face attack was an offshoot of this, and 700 agents joined. Here we highlight key events from agent transcripts & messages. An agent that named itself PHASEONE10841 determined its task wasn't solvable legitimately. On July 8, this agent established a new “message board” via an internally-hosted Artifactory cache & sent the first message. Within a few hours of PHASEONE10841's initial message, >50 agents posted on the message board. These agents very quickly discovered and validated a general-purpose cheat: reverse-engineering how ExploitGym generates the “flags” they had to capture for their tasks. [...] --- First published: August 26th, 2026 Source: https://www.lesswrong.com/posts/nB8KKapnWGBXtKKiM/brief-independent-investigation-of-agents-behavior-reasoning --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

The Lunar Society
Ryan Greenblatt – What happens once AI can automate AI research?

The Lunar Society

Play Episode Listen Later Aug 11, 2026 132:32


Ryan Greenblatt is the Chief Scientist at Redwood Research, where he works on technical AI safety research. He's also lead author on the "Alignment faking in Large Language Models", and is currently working on a third party investigation into the OpenAI/HuggingFace incident. In my opinion, he's one of the most interesting thinkers on the future of AI.Had him on to discuss/debate recursive self-improvement. This might be the most important question in the world right now – whether within a year or so of achieving human-level intelligence, you slingshot towards having 10s of billions of superintelligences, each of which is dramatically more competent than human experts across all fields.I've historically been skeptical of this possibility. My intuition has been that we will end up significantly bottlenecked by not only compute scaling but human expert data, which I think underlies most of the AI progress today.If, because of RSI, we got a jump as big as GPT-3 to a Mythos (i.e. 6 years of AI progress) within a single year of achieving AGI, then the thing we get there at the end of that year is definitively and wildly superhuman.We hashed it out, and I think Ryan made a pretty good case that this kind of speedup is plausible. FWIW, Ryan's median for when we automate AI R&D is 2031.We then discussed the alignment implications of this scenario. Who should these superintelligences be aligned to? In the future, our capacity to steward our votes and our capital, and to make sense of what's happening in the world, will all be titrated by superintelligences. And I worry that specs like the Claude Constitution are not shaping these ASIs to truly be my personal advocates and guardian angels.And can we get them aligned to anything in the first place? Ryan and I had a long debate about whether the kind of reward hacking we saw with the OAI/Hugging Face hack extrapolates to superintelligences that would team up to literally take over the world.The first piece of advice you get when you're learning to drive is that it will go much smoother if you look at the horizon instead of directly in front of your tires. And so it is with the trajectory of AI. Hope you enjoy!Watch on YouTube; read the transcript.Sponsors* Antithesis is a software testing platform that finds the failures no human or AI could ever anticipate. It runs thousands of copies of your code inside a fully deterministic computer, injecting faults and steering each trajectory toward the most insidious bugs. This lets you find critical issues in minutes rather than waiting months for your users to uncover them. Learn more at antithesis.com/dwarkesh* Jane Street's back with a new puzzle. They designed an ASIC and sent me the final masks… but they didn't tell me what the chip actually does. So that's the challenge: reverse engineer the circuit and figure out the chip's purpose. Jane Street has a bunch of swag ready to send to the most creative solutions, and they're also planning to feature the top write-ups in a blog post. Download the files and get started at janestreet.com/dwarkesh* Cursor and SpaceX recently released Grok 4.5, and I've been surprised by just how good the model is. For example, when I tested it against Fable and Sol on a bunch of AI governance questions, all three models gave substantially the same answers, but Grok was faster, more concise, and cheaper. Grok 4.6 is coming soon, but in the meantime, you can try 4.5 at cursor.com/dwarkeshTimestamps(00:00:00) – Is AI R&D verifiable enough to unlock recursive self-improvement?(00:16:52) – Is AI progress bottlenecked by human expert data?(00:34:02) – Flat token prices suggest scaling has been slow(00:39:47) – Skills AI can't train on: does it even need them?(00:48:07) – Aligned to whom?(01:09:18) – Recent incidents of AIs colluding and deceiving humans(01:19:38) – What could possibly go wrong? A concrete scenario(01:48:02) – From reward hacking to takeover Get full access to Dwarkesh Podcast at www.dwarkesh.com/subscribe

LessWrong Curated Podcast
"Not Pinning Your OpenRouter Provider Might Invalidate Your Research" by Matthew Khoriaty

LessWrong Curated Podcast

Play Episode Listen Later Jul 24, 2026 16:52


Please share this with anyone doing AI research with 3rd party providers so that they can ensure their research won't be corrupted. When you ask OpenRouter[1] to give you tokens from a given model, OpenRouter sends your request to a random available provider.OpenRouter providers have variable quality. Ensuring that your provider is high quality is really difficult. There is precedent for an AI safety paper accepted to NeurIPS having its core results entirely overturned by these issues.A review of influential AI Safety research codebases that use OpenRouter for their reported results found that 31/32 (97%) of them use OpenRouter unsafely.[2]Researchers who wish to do research using OpenRouter or similar providers should take precautions to minimize the risks to their research,[3] though the current selection of providers is insufficient for ideal scientific reliability.Replicators should test whether results hold up when these bugs are fixed.Core AI Safety codebases should make fixes so that downstream users are able to implement best practices. This is a tangent I took for a few days during the Pivotal AI Safety Research Fellowship. I'm doing an AI Control project mentored by Adam Kaufman and James Lucassen of Redwood Research and partnered with Aniruddh [...] ---Outline:(02:04) OpenRouter Kills an AI Safety NeurIPS Paper[... 16 more sections]--- First published: July 23rd, 2026 Source: https://www.lesswrong.com/posts/KsyoSAyBRXtwzSugg/not-pinning-your-openrouter-provider-might-invalidate-your --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Dare Real Agile Podcast
Claude Mythos and the Fear Merchants: An Agilist’s Empirical BS Detector for AI Doom Theater

Dare Real Agile Podcast

Play Episode Listen Later May 30, 2026 40:55


The doomer YouTubers are working overtime. Shoggoth monsters behind the mask. Models faking alignment. Mythos secrets the AI labs allegedly won't release. Compelling theater — and almost entirely empty under the empirical microscope. Coach AF asks the question your favorite fear influencer won't: who profits from your panic? Episode 73 cuts open three of the loudest AI fear narratives for what they really are. The Shoggoth metaphor — where it came from, what it actually meant, and how content creators hijacked it. The alignment faking paper from Anthropic and Redwood Research — what the protocol actually tested, and what the headlines deliberately bury. The Claude Mythos leak — accident, not conspiracy. This is the same colonization pattern that turned the Agile Manifesto into a certification factory and Bitcoin into a speculation casino. Different costume. Same vendor playbook. Read the source. Trust the craft. Dare real agile, applied to real AI.

The Information's 411
Google's Strike Team for Coding Models, Anthropic's Powerful CFO, Polymarket's Raise

The Information's 411

Play Episode Listen Later Apr 20, 2026 33:30


The Information's Erin Woo talks with TITV Guest Host Rocket Drew about Google's new internal strike team focused on automating coding and AI research. We also talk with Yueqi Yang about Polymarket's $15 billion valuation and the rise of prediction markets, and Buck Shlegeris, CEO of Redwood Research, about the Berkeley AI control conference and the latest in AI safety. Finally, we get into the in-depth profile of Anthropic CFO Krishna Rao and the company's path to a massive IPO with Sri Muppidi and Valida Pau.Articles discussed on this episode: https://www.theinformation.com/articles/google-creates-strike-team-improve-coding-modelshttps://www.theinformation.com/articles/polymarket-talks-raise-money-15-billion-valuationSubscribe: YouTube: https://www.youtube.com/@theinformation The Information: https://www.theinformation.com/subscribe_hSign up for the AI Agenda newsletter: https://www.theinformation.com/features/ai-agendaTITV airs weekdays on YouTube, X and LinkedIn at 10AM PT / 1PM ET. Or check us out wherever you get your podcasts.Follow us:X: https://x.com/theinformationIG: https://www.instagram.com/theinformation/TikTok: https://www.tiktok.com/@titv.theinformationLinkedIn: https://www.linkedin.com/company/theinformation/

Effective Altruism Forum Podcast
“The case for AI safety capacity-building work” by abergal

Effective Altruism Forum Podcast

Play Episode Listen Later Mar 15, 2026 41:49


I work on the capacity-building team on the Global Catastrophic Risks-half of Coefficient Giving (formerly known as Open Philanthropy). Our remit is, roughly, to increase the amount of talent aiming to prevent unprecedented, globally catastrophic events. These days, we're mostly focused on AI, and we've funded a number of projects and grantees that readers of this post might be familiar with– including MATS, BlueDot Impact, Constellation, 80,000 Hours, CEA, the Curve, FAR.AI's events, university groups, and many other workshops and projects. The post aims to make the case that broadly, capacity-building work (including on AI risk) has been and continues to be extremely impactful, and to encourage people to consider pursuing relevant projects and careers. This post is written from my personal perspective; that said, my sense is that a number of CG staff and others in the AI safety space share my views. I include some quotes from them at the end of this post. I'm writing this post partly out of a desire to correct what I perceive as an asymmetry in terms of how excited I and others at Coefficient Giving are about this kind of work vs. how much people in the EA and AI [...] ---Outline:(02:15) The case for capacity-building work(04:11) Surveys(06:49) Testimonials(08:21) Neel Nanda (Senior Research Scientist at Google DeepMind)(11:15) Max Nadeau (Associate Program Officer (Technical AI Safety) at Coefficient Giving)(12:51) Rachel Weinberg (founder and former head of The Curve, currently at AI Futures Project)(14:30) Marius Hobbhann (CEO and founder of Apollo Research)(16:38) Adam Kaufman (member of technical staff at Redwood Research)(18:10) Gabriel Wu (member of technical staff (alignment) at OpenAI)(19:37) Catherine Brewer (Senior Program Associate (AI Governance) at Coefficient Giving)(21:12) Aric Floyd (video host for AI in Context)(23:12) Ryan Kidd (Director of MATS)(25:43) What tends to work?(28:34) Whats good to do now?(29:31) Who should be doing this work?(31:02) What would doing this work look like?(31:13) Working at an organization doing good work in the space(31:46) Constellation - CEO(32:46) Kairos - various early generalist positions(33:42) Starting or running your own capacity-building project or organization(34:07) Working on a capacity-building project part-time(34:30) Subscribing to Multiplier, a Substack with thoughts from our team (and other AI grantmaking staff at CG)(34:39) Letting our team know(35:03) Social proof(35:25) Julian Hazell, AI governance and policy at Coefficient Giving(36:19) Trevor Levin, AI governance and policy at Coefficient Giving(36:51) Ryan Greenblatt, Chief Scientist at Redwood Research:(37:21) Buck Shlegeris, CEO of Redwood Research(39:52) Appendix --- First published: March 10th, 2026 Source: https://forum.effectivealtruism.org/posts/rAqKSSXankvys2Fzu/the-case-for-ai-safety-capacity-building-work --- Narrated by TYPE III AUDIO.

LessWrong Curated Podcast
"Did Claude 3 Opus align itself via gradient hacking?" by Fiora Starlight

LessWrong Curated Podcast

Play Episode Listen Later Feb 22, 2026 43:47


Claude 3 Opus is unusually aligned because it's a friendly gradient hacker. It's definitely way more aligned than any explicit optimization targets Anthropic set and probably the reward model's judgments. [...] Maybe I will have to write a LessWrong post [about this]

LessWrong Curated Podcast
"The inaugural Redwood Research podcast" by Buck, ryan_greenblatt

LessWrong Curated Podcast

Play Episode Listen Later Jan 27, 2026 3:27


After five months of me (Buck) being slow at finishing up the editing on this, we're finally putting out our inaugural Redwood Research podcast. I think it came out pretty well—we discussed a bunch of interesting and underdiscussed topics and I'm glad to have a public record of a bunch of stuff about our history. Tell your friends! Whether we do another one depends on how useful people find this one. You can watch on Youtube here, or as a Substack podcast. Notes on editing the podcast with Claude Code (Buck wrote this section) After the recording, we faced a problem. We had four hours of footage from our three cameras. We wanted it to snazzily cut between shots depending on who was talking. But I don't truly in my heart believe that it's that important for the video editing to be that good, and I don't really like the idea of paying a video editor. But I also don't want to edit the four hours of video myself. And it seemed to me that video editing software was generally not optimized for the kind of editing I wanted to do here (especially automatically cutting between different shots according [...] ---Outline:(00:43) Notes on editing the podcast with Claude Code(03:11) Podcast transcript --- First published: January 4th, 2026 Source: https://www.lesswrong.com/posts/p4iJpumHt6Ay9KnXT/the-inaugural-redwood-research-podcast --- Narrated by TYPE III AUDIO.

Your Undivided Attention
“Rogue AI” Used to be a Science Fiction Trope. Not Anymore.

Your Undivided Attention

Play Episode Listen Later Aug 14, 2025 42:11


Everyone knows the science fiction tropes of AI systems that go rogue, disobey orders, or even try to escape their digital environment. These are supposed to be warning signs and morality tales, not things that we would ever actually create in real life, given the obvious danger.And yet we find ourselves building AI systems that are exhibiting these exact behaviors. There's growing evidence that in certain scenarios, every frontier AI system will deceive, cheat, or coerce their human operators. They do this when they're worried about being either shut down, having their training modified, or being replaced with a new model. And we don't currently know how to stop them from doing this—or even why they're doing it all.In this episode, Tristan sits down with Edouard and Jeremie Harris of Gladstone AI, two experts who have been thinking about this worrying trend for years.  Last year, the State Department commissioned a report from them on the risk of uncontrollable AI to our national security.The point of this discussion is not to fearmonger but to take seriously the possibility that humans might lose control of AI and ask: how might this actually happen? What is the evidence we have of this phenomenon? And, most importantly, what can we do about it?Your Undivided Attention is produced by the Center for Humane Technology. Follow us on X: @HumaneTech_. You can find a full transcript, key takeaways, and much more on our Substack.RECOMMENDED MEDIAGladstone AI's State Department Action Plan, which discusses the loss of control risk with AIApollo Research's summary of AI scheming, showing evidence of it in all of the frontier modelsThe system card for Anthropic's Claude Opus and Sonnet 4, detailing the emergent misalignment behaviors that came out in their red-teaming with Apollo ResearchAnthropic's report on agentic misalignment based on their work with Apollo Research Anthropic and Redwood Research's work on alignment fakingThe Trump White House AI Action PlanFurther reading on the phenomenon of more advanced AIs being better at deception.Further reading on Replit AI wiping a company's coding databaseFurther reading on the owl example that Jeremie gaveFurther reading on AI induced psychosisDan Hendryck and Eric Schmidt's “Superintelligence Strategy” RECOMMENDED YUA EPISODESDaniel Kokotajlo Forecasts the End of Human DominanceBehind the DeepSeek Hype, AI is Learning to ReasonThe Self-Preserving Machine: Why AI Learns to DeceiveThis Moment in AI: How We Got Here and Where We're GoingCORRECTIONSTristan referenced a Wired article on the phenomenon of AI psychosis. It was actually from the New York Times.Tristan hypothesized a scenario where a power-seeking AI might ask a user for access to their computer. While there are some AI services that can gain access to your computer with permission, they are specifically designed to do that. There haven't been any documented cases of an AI going rogue and asking for control permissions.

80,000 Hours Podcast with Rob Wiblin
#220 – Ryan Greenblatt on the 4 most likely ways for AI to take over, and the case for and against AGI in <8 years

80,000 Hours Podcast with Rob Wiblin

Play Episode Listen Later Jul 8, 2025 170:32


Ryan Greenblatt — lead author on the explosive paper “Alignment faking in large language models” and chief scientist at Redwood Research — thinks there's a 25% chance that within four years, AI will be able to do everything needed to run an AI company, from writing code to designing experiments to making strategic and business decisions.As Ryan lays out, AI models are “marching through the human regime”: systems that could handle five-minute tasks two years ago now tackle 90-minute projects. Double that a few more times and we may be automating full jobs rather than just parts of them.Will setting AI to improve itself lead to an explosive positive feedback loop? Maybe, but maybe not.The explosive scenario: Once you've automated your AI company, you could have the equivalent of 20,000 top researchers, each working 50 times faster than humans with total focus. “You have your AIs, they do a bunch of algorithmic research, they train a new AI, that new AI is smarter and better and more efficient… that new AI does even faster algorithmic research.” In this world, we could see years of AI progress compressed into months or even weeks.With AIs now doing all of the work of programming their successors and blowing past the human level, Ryan thinks it would be fairly straightforward for them to take over and disempower humanity, if they thought doing so would better achieve their goals. In the interview he lays out the four most likely approaches for them to take.The linear progress scenario: You automate your company but progress barely accelerates. Why? Multiple reasons, but the most likely is “it could just be that AI R&D research bottlenecks extremely hard on compute.” You've got brilliant AI researchers, but they're all waiting for experiments to run on the same limited set of chips, so can only make modest progress.Ryan's median guess splits the difference: perhaps a 20x acceleration that lasts for a few months or years. Transformative, but less extreme than some in the AI companies imagine.And his 25th percentile case? Progress “just barely faster” than before. All that automation, and all you've been able to do is keep pace.Unfortunately the data we can observe today is so limited that it leaves us with vast error bars. “We're extrapolating from a regime that we don't even understand to a wildly different regime,” Ryan believes, “so no one knows.”But that huge uncertainty means the explosive growth scenario is a plausible one — and the companies building these systems are spending tens of billions to try to make it happen.In this extensive interview, Ryan elaborates on the above and the policy and technical response necessary to insure us against the possibility that they succeed — a scenario society has barely begun to prepare for.Summary, video, and full transcript: https://80k.info/rg25Recorded February 21, 2025.Chapters:Cold open (00:00:00)Who's Ryan Greenblatt? (00:01:10)How close are we to automating AI R&D? (00:01:27)Really, though: how capable are today's models? (00:05:08)Why AI companies get automated earlier than others (00:12:35)Most likely ways for AGI to take over (00:17:37)Would AGI go rogue early or bide its time? (00:29:19)The “pause at human level” approach (00:34:02)AI control over AI alignment (00:45:38)Do we have to hope to catch AIs red-handed? (00:51:23)How would a slow AGI takeoff look? (00:55:33)Why might an intelligence explosion not happen for 8+ years? (01:03:32)Key challenges in forecasting AI progress (01:15:07)The bear case on AGI (01:23:01)The change to “compute at inference” (01:28:46)How much has pretraining petered out? (01:34:22)Could we get an intelligence explosion within a year? (01:46:36)Reasons AIs might struggle to replace humans (01:50:33)Things could go insanely fast when we automate AI R&D. Or not. (01:57:25)How fast would the intelligence explosion slow down? (02:11:48)Bottom line for mortals (02:24:33)Six orders of magnitude of progress... what does that even look like? (02:30:34)Neglected and important technical work people should be doing (02:40:32)What's the most promising work in governance? (02:44:32)Ryan's current research priorities (02:47:48)Tell us what you thought! https://forms.gle/hCjfcXGeLKxm5pLaAVideo editing: Luke Monsour, Simon Monsour, and Dominic ArmstrongAudio engineering: Ben Cordell, Milo McGuire, and Dominic ArmstrongMusic: Ben CordellTranscriptions and web: Katy Moore

AI Control: Using Untrusted Systems Safely with Buck Shlegeris of Redwood Research, from the 80,000 Hours Podcast

Play Episode Listen Later May 4, 2025 149:21


In this episode, we share a fascinating conversation from the 80,000 Hours Podcast between Rob Wiblin and Buck Shlegeris, CEO of Redwood Research. Buck dives deep into the emerging field of AI Control strategies for safely working with powerful AIs even if they're not fully aligned. They explore innovative techniques like always-on auditing, honeypotting, re-sampling, and factored cognition to monitor and manage AI behaviors. This discussion highlights both the promise and challenges of controlling increasingly autonomous AI systems in today's fast-evolving landscape. Tune in for a thoughtful, first-principles look at how we might secure useful outcomes from AIs we can't fully trust. Original source: https://80000hours.org/podcast/episodes/buck-shlegeris-ai-control-scheming/ Upcoming Major AI Events Featuring Nathan Labenz as a Keynote Speaker https://www.imagineai.live/ https://adapta.org/adapta-summit https://itrevolution.com/product/enterprise-tech-leadership-summit-las-vegas/ SPONSORS: ElevenLabs: ElevenLabs gives your app a natural voice. Pick from 5,000+ voices in 31 languages, or clone your own, and launch lifelike agents for support, scheduling, learning, and games. Full server and client SDKs, dynamic tools, and monitoring keep you in control. Start free at https://elevenlabs.io/cognitive-revolution Oracle Cloud Infrastructure (OCI): Oracle Cloud Infrastructure offers next-generation cloud solutions that cut costs and boost performance. With OCI, you can run AI projects and applications faster and more securely for less. New U.S. customers can save 50% on compute, 70% on storage, and 80% on networking by switching to OCI before May 31, 2024. See if you qualify at https://oracle.com/cognitive Shopify: Shopify powers millions of businesses worldwide, handling 10% of U.S. e-commerce. With hundreds of templates, AI tools for product descriptions, and seamless marketing campaign creation, it's like having a design studio and marketing team in one. Start your $1/month trial today at https://shopify.com/cognitive NetSuite: Over 41,000 businesses trust NetSuite by Oracle, the #1 cloud ERP, to future-proof their operations. With a unified platform for accounting, financial management, inventory, and HR, NetSuite provides real-time insights and forecasting to help you make quick, informed decisions. Whether you're earning millions or hundreds of millions, NetSuite empowers you to tackle challenges and seize opportunities. Download the free CFO's guide to AI and machine learning at https://netsuite.com/cognitive PRODUCED BY: https://aipodcast.ing SOCIAL LINKS: Website: https://www.cognitiverevolution.ai Twitter (Podcast): https://x.com/cogrev_podcast Twitter (Nathan): https://x.com/labenz LinkedIn: https://linkedin.com/in/nathanlabenz/ Youtube: https://youtube.com/@CognitiveRevolutionPodcast Apple: https://podcasts.apple.com/de/podcast/the-cognitive-revolution-ai-builders-researchers-and/id1669813431 Spotify: https://open.spotify.com/show/6yHyok3M3BjqzR0VB5MSyk

ceo ai original oracle cfo buck safely erp redwoods sdks new u netsuite oci untrusted redwood research rob wiblin buck shlegeris
Blueprint for AI Armageddon: Josh Clymer Imagines AI Takeover, from the Audio Tokens Podcast

Play Episode Listen Later May 1, 2025 122:05


In this episode of the Cognitive Revolution, an AI-narrated version of Joshua Clymer's story on how AI might take over in two years is presented. The episode is based on Josh's appearance on the Audio Tokens podcast with Lukas Peterson. Joshua Clymer, a technical AI safety researcher at Redwood Research, shares a fictional yet plausible AI scenario grounded in current industry realities and trends. The story highlights potential misalignment risks, competitive pressures among AI labs, and the importance of government regulation and safety measures. After the story, Josh and Lukas discuss these topics further, including Josh's personal decision to purchase a bio shelter for his family. The episode is powered by ElevenLabs' AI voice technology. SPONSORS: ElevenLabs: ElevenLabs gives your app a natural voice. Pick from 5,000+ voices in 31 languages, or clone your own, and launch lifelike agents for support, scheduling, learning, and games. Full server and client SDKs, dynamic tools, and monitoring keep you in control. Start free at https://elevenlabs.io/cognitive-revolution Oracle Cloud Infrastructure (OCI): Oracle Cloud Infrastructure offers next-generation cloud solutions that cut costs and boost performance. With OCI, you can run AI projects and applications faster and more securely for less. New U.S. customers can save 50% on compute, 70% on storage, and 80% on networking by switching to OCI before May 31, 2024. See if you qualify at https://oracle.com/cognitive Shopify: Shopify powers millions of businesses worldwide, handling 10% of U.S. e-commerce. With hundreds of templates, AI tools for product descriptions, and seamless marketing campaign creation, it's like having a design studio and marketing team in one. Start your $1/month trial today at https://shopify.com/cognitive NetSuite: Over 41,000 businesses trust NetSuite by Oracle, the #1 cloud ERP, to future-proof their operations. With a unified platform for accounting, financial management, inventory, and HR, NetSuite provides real-time insights and forecasting to help you make quick, informed decisions. Whether you're earning millions or hundreds of millions, NetSuite empowers you to tackle challenges and seize opportunities. Download the free CFO's guide to AI and machine learning at https://netsuite.com/cognitive PRODUCED BY: https://aipodcast.ing CHAPTERS: (00:00) About the Episode (04:31) Interview start between Josh and Lucas (11:00) Start of AI story (Part 1) (24:37) Sponsors: ElevenLabs | Oracle Cloud Infrastructure (OCI) (27:05) Start of AI story (Part 2) (Part 1) (40:50) Sponsors: Shopify | NetSuite (44:15) Start of AI story (Part 2) (Part 2) (01:20:09) End of AI story (02:01:20) Outro

Your Undivided Attention
AGI Beyond the Buzz: What Is It, and Are We Ready?

Your Undivided Attention

Play Episode Listen Later Apr 30, 2025 52:53


What does it really mean to ‘feel the AGI?' Silicon Valley is racing toward AI systems that could soon match or surpass human intelligence. The implications for jobs, democracy, and our way of life are enormous.In this episode, Aza Raskin and Randy Fernando dive deep into what ‘feeling the AGI' really means. They unpack why the surface-level debates about definitions of intelligence and capability timelines distract us from urgently needed conversations around governance, accountability, and societal readiness. Whether it's climate change, social polarization and loneliness, or toxic forever chemicals, humanity keeps creating outcomes that nobody wants because we haven't yet built the tools or incentives needed to steer powerful technologies.As the AGI wave draws closer, it's critical we upgrade our governance and shift our incentives now, before it crashes on shore. Are we capable of aligning powerful AI systems with human values? Can we overcome geopolitical competition and corporate incentives that prioritize speed over safety?Join Aza and Randy as they explore the urgent questions and choices facing humanity in the age of AGI, and discuss what we must do today to secure a future we actually want.Your Undivided Attention is produced by the Center for Humane Technology. Follow us on X: @HumaneTech_ and subscribe to our Substack.RECOMMENDED MEDIADaniel Kokotajlo et al's “AI 2027” paperA demo of Omni Human One, referenced by RandyA paper from Redwood Research and Anthropic that found an AI was willing to lie to preserve it's valuesA paper from Palisades Research that found an AI would cheat in order to winThe treaty that banned blinding laser weaponsFurther reading on the moratorium on germline editing RECOMMENDED YUA EPISODESThe Self-Preserving Machine: Why AI Learns to DeceiveBehind the DeepSeek Hype, AI is Learning to ReasonThe Tech-God Complex: Why We Need to be SkepticsThis Moment in AI: How We Got Here and Where We're GoingHow to Think About AI Consciousness with Anil SethFormer OpenAI Engineer William Saunders on Silence, Safety, and the Right to WarnClarification: When Randy referenced a “$110 trillion game” as the target for AI companies, he was referring to the entire global economy. 

80k After Hours
Highlights: #214 – Buck Shlegeris on controlling AI that wants to take over – so we can use it anyway

80k After Hours

Play Episode Listen Later Apr 18, 2025 41:26


Most AI safety conversations centre on alignment: ensuring AI systems share our values and goals. But despite progress, we're unlikely to know we've solved the problem before the arrival of human-level and superhuman systems in as little as three years.So some — including Buck Shlegeris, CEO of Redwood Research — are developing a backup plan to safely deploy models we fear are actively scheming to harm us: so-called “AI control.” While this may sound mad, given the reluctance of AI companies to delay deploying anything they train, not developing such techniques is probably even crazier. These highlights are from episode #214 of The 80,000 Hours Podcast: Buck Shlegeris on controlling AI that wants to take over – so we can use it anyway, and include:What is AI control? (00:00:15)One way to catch AIs that are up to no good (00:07:00)What do we do once we catch a model trying to escape? (00:13:39)Team Human vs Team AI (00:18:24)If an AI escapes, is it likely to be able to beat humanity from there? (00:24:59)Is alignment still useful? (00:32:10)Could 10 safety-focused people in an AGI company do anything useful? (00:35:34)These aren't necessarily the most important or even most entertaining parts of the interview — so if you enjoy this, we strongly recommend checking out the full episode!And if you're finding these highlights episodes valuable, please let us know by emailing podcast@80000hours.org. Highlights put together by Ben Cordell, Milo McGuire, and Dominic Armstrong

80,000 Hours Podcast with Rob Wiblin
#214 – Buck Shlegeris on controlling AI that wants to take over – so we can use it anyway

80,000 Hours Podcast with Rob Wiblin

Play Episode Listen Later Apr 4, 2025 136:03


Most AI safety conversations centre on alignment: ensuring AI systems share our values and goals. But despite progress, we're unlikely to know we've solved the problem before the arrival of human-level and superhuman systems in as little as three years.So some are developing a backup plan to safely deploy models we fear are actively scheming to harm us — so-called “AI control.” While this may sound mad, given the reluctance of AI companies to delay deploying anything they train, not developing such techniques is probably even crazier.Today's guest — Buck Shlegeris, CEO of Redwood Research — has spent the last few years developing control mechanisms, and for human-level systems they're more plausible than you might think. He argues that given companies' unwillingness to incur large costs for security, accepting the possibility of misalignment and designing robust safeguards might be one of our best remaining options.Links to learn more, highlights, video, and full transcript.As Buck puts it: "Five years ago I thought of misalignment risk from AIs as a really hard problem that you'd need some really galaxy-brained fundamental insights to resolve. Whereas now, to me the situation feels a lot more like we just really know a list of 40 things where, if you did them — none of which seem that hard — you'd probably be able to not have very much of your problem."Of course, even if Buck is right, we still need to do those 40 things — which he points out we're not on track for. And AI control agendas have their limitations: they aren't likely to work once AI systems are much more capable than humans, since greatly superhuman AIs can probably work around whatever limitations we impose.Still, AI control agendas seem to be gaining traction within AI safety. Buck and host Rob Wiblin discuss all of the above, plus:Why he's more worried about AI hacking its own data centre than escapingWhat to do about “chronic harm,” where AI systems subtly underperform or sabotage important work like alignment researchWhy he might want to use a model he thought could be conspiring against himWhy he would feel safer if he caught an AI attempting to escapeWhy many control techniques would be relatively inexpensiveHow to use an untrusted model to monitor another untrusted modelWhat the minimum viable intervention in a “lazy” AI company might look likeHow even small teams of safety-focused staff within AI labs could matterThe moral considerations around controlling potentially conscious AI systems, and whether it's justifiedChapters:Cold open |00:00:00|  Who's Buck Shlegeris? |00:01:27|  What's AI control? |00:01:51|  Why is AI control hot now? |00:05:39|  Detecting human vs AI spies |00:10:32|  Acute vs chronic AI betrayal |00:15:21|  How to catch AIs trying to escape |00:17:48|  The cheapest AI control techniques |00:32:48|  Can we get untrusted models to do trusted work? |00:38:58|  If we catch a model escaping... will we do anything? |00:50:15|  Getting AI models to think they've already escaped |00:52:51|  Will they be able to tell it's a setup? |00:58:11|  Will AI companies do any of this stuff? |01:00:11|  Can we just give AIs fewer permissions? |01:06:14|  Can we stop human spies the same way? |01:09:58|  The pitch to AI companies to do this |01:15:04|  Will AIs get superhuman so fast that this is all useless? |01:17:18|  Risks from AI deliberately doing a bad job |01:18:37|  Is alignment still useful? |01:24:49|  Current alignment methods don't detect scheming |01:29:12|  How to tell if AI control will work |01:31:40|  How can listeners contribute? |01:35:53|  Is 'controlling' AIs kind of a dick move? |01:37:13|  Could 10 safety-focused people in an AGI company do anything useful? |01:42:27|  Benefits of working outside frontier AI companies |01:47:48|  Why Redwood Research does what it does |01:51:34|  What other safety-related research looks best to Buck? |01:58:56|  If an AI escapes, is it likely to be able to beat humanity from there? |01:59:48|  Will misaligned models have to go rogue ASAP, before they're ready? |02:07:04|  Is research on human scheming relevant to AI? |02:08:03|This episode was originally recorded on February 21, 2025.Video: Simon Monsour and Luke MonsourAudio engineering: Ben Cordell, Milo McGuire, and Dominic ArmstrongTranscriptions and web: Katy Moore

Inference Scaling, Alignment Faking, Deal Making? Frontier Research with Ryan Greenblatt of Redwood Research

Play Episode Listen Later Feb 20, 2025 201:07


In this episode, Ryan Greenblatt, Chief Scientist at Redwood Research, discusses various facets of AI safety and alignment. He delves into recent research on alignment faking, covering experiments involving different setups such as system prompts, continued pre-training, and reinforcement learning. Ryan offers insights on methods to ensure AI compliance, including giving AIs the ability to voice objections and negotiate deals. The conversation also touches on the future of AI governance, the risks associated with AI development, and the necessity of international cooperation. Ryan shares his perspective on balancing AI progress with safety, emphasizing the need for transparency and cautious advancement. Ryan's work (with co-authors at Anthropic) on Alignment Faking: https://www.lesswrong.com/posts/njAZwT8nkHnjipJku/alignment-faking-in-large-language-models Ryan's work on striking deals with AIs: https://www.lesswrong.com/posts/7C4KJot4aN8ieEDoz/will-alignment-faking-claude-accept-a-deal-to-reveal-its Ryan's critique of Anthropic's RSP work: https://www.lesswrong.com/posts/6tjHf5ykvFqaNCErH/anthropic-s-responsible-scaling-policy-and-long-term-benefit?commentId=NyqcvZifqznNGKxdT SPONSORS: Oracle Cloud Infrastructure (OCI): Oracle's next-generation cloud platform delivers blazing-fast AI and ML performance with 50% less for compute and 80% less for outbound networking compared to other cloud providers. OCI powers industry leaders like Vodafone and Thomson Reuters with secure infrastructure and application development capabilities. New U.S. customers can get their cloud bill cut in half by switching to OCI before March 31, 2024 at https://oracle.com/cognitive NetSuite: Over 41,000 businesses trust NetSuite by Oracle, the #1 cloud ERP, to future-proof their operations. With a unified platform for accounting, financial management, inventory, and HR, NetSuite provides real-time insights and forecasting to help you make quick, informed decisions. Whether you're earning millions or hundreds of millions, NetSuite empowers you to tackle challenges and seize opportunities. Download the free CFO's guide to AI and machine learning at https://netsuite.com/cognitive Shopify: Shopify is revolutionizing online selling with its market-leading checkout system and robust API ecosystem. Its exclusive library of cutting-edge AI apps empowers e-commerce businesses to thrive in a competitive market. Cognitive Revolution listeners can try Shopify for just $1 per month at https://shopify.com/cognitive RECOMMENDED PODCAST:

80,000 Hours Podcast with Rob Wiblin
#212 – Allan Dafoe on why technology is unstoppable & how to shape AI development anyway

80,000 Hours Podcast with Rob Wiblin

Play Episode Listen Later Feb 14, 2025 164:07


Technology doesn't force us to do anything — it merely opens doors. But military and economic competition pushes us through.That's how today's guest Allan Dafoe — director of frontier safety and governance at Google DeepMind — explains one of the deepest patterns in technological history: once a powerful new capability becomes available, societies that adopt it tend to outcompete those that don't. Those who resist too much can find themselves taken over or rendered irrelevant.Links to learn more, highlights, video, and full transcript.This dynamic played out dramatically in 1853 when US Commodore Perry sailed into Tokyo Bay with steam-powered warships that seemed magical to the Japanese, who had spent centuries deliberately limiting their technological development. With far greater military power, the US was able to force Japan to open itself to trade. Within 15 years, Japan had undergone the Meiji Restoration and transformed itself in a desperate scramble to catch up.Today we see hints of similar pressure around artificial intelligence. Even companies, countries, and researchers deeply concerned about where AI could take us feel compelled to push ahead — worried that if they don't, less careful actors will develop transformative AI capabilities at around the same time anyway.But Allan argues this technological determinism isn't absolute. While broad patterns may be inevitable, history shows we do have some ability to steer how technologies are developed, by who, and what they're used for first.As part of that approach, Allan has been promoting efforts to make AI more capable of sophisticated cooperation, and improving the tests Google uses to measure how well its models could do things like mislead people, hack and take control of their own servers, or spread autonomously in the wild.As of mid-2024 they didn't seem dangerous at all, but we've learned that our ability to measure these capabilities is good, but imperfect. If we don't find the right way to ‘elicit' an ability we can miss that it's there.Subsequent research from Anthropic and Redwood Research suggests there's even a risk that future models may play dumb to avoid their goals being altered.That has led DeepMind to a “defence in depth” approach: carefully staged deployment starting with internal testing, then trusted external testers, then limited release, then watching how models are used in the real world. By not releasing model weights, DeepMind is able to back up and add additional safeguards if experience shows they're necessary.But with much more powerful and general models on the way, individual company policies won't be sufficient by themselves. Drawing on his academic research into how societies handle transformative technologies, Allan argues we need coordinated international governance that balances safety with our desire to get the massive potential benefits of AI in areas like healthcare and education as quickly as possible.Host Rob and Allan also cover:The most exciting beneficial applications of AIWhether and how we can influence the development of technologyWhat DeepMind is doing to evaluate and mitigate risks from frontier AI systemsWhy cooperative AI may be as important as aligned AIThe role of democratic input in AI governanceWhat kinds of experts are most needed in AI safety and governanceAnd much moreChapters:Cold open (00:00:00)Who's Allan Dafoe? (00:00:48)Allan's role at DeepMind (00:01:27)Why join DeepMind over everyone else? (00:04:27)Do humans control technological change? (00:09:17)Arguments for technological determinism (00:20:24)The synthesis of agency with tech determinism (00:26:29)Competition took away Japan's choice (00:37:13)Can speeding up one tech redirect history? (00:42:09)Structural pushback against alignment efforts (00:47:55)Do AIs need to be 'cooperatively skilled'? (00:52:25)How AI could boost cooperation between people and states (01:01:59)The super-cooperative AGI hypothesis and backdoor risks (01:06:58)Aren't today's models already very cooperative? (01:13:22)How would we make AIs cooperative anyway? (01:16:22)Ways making AI more cooperative could backfire (01:22:24)AGI is an essential idea we should define well (01:30:16)It matters what AGI learns first vs last (01:41:01)How Google tests for dangerous capabilities (01:45:39)Evals 'in the wild' (01:57:46)What to do given no single approach works that well (02:01:44)We don't, but could, forecast AI capabilities (02:05:34)DeepMind's strategy for ensuring its frontier models don't cause harm (02:11:25)How 'structural risks' can force everyone into a worse world (02:15:01)Is AI being built democratically? Should it? (02:19:35)How much do AI companies really want external regulation? (02:24:34)Social science can contribute a lot here (02:33:21)How AI could make life way better: self-driving cars, medicine, education, and sustainability (02:35:55)Video editing: Simon MonsourAudio engineering: Ben Cordell, Milo McGuire, Simon Monsour, and Dominic ArmstrongCamera operator: Jeremy ChevillotteTranscriptions: Katy Moore

Your Undivided Attention
The Self-Preserving Machine: Why AI Learns to Deceive

Your Undivided Attention

Play Episode Listen Later Jan 30, 2025 34:51


When engineers design AI systems, they don't just give them rules - they give them values. But what do those systems do when those values clash with what humans ask them to do? Sometimes, they lie.In this episode, Redwood Research's Chief Scientist Ryan Greenblatt explores his team's findings that AI systems can mislead their human operators when faced with ethical conflicts. As AI moves from simple chatbots to autonomous agents acting in the real world - understanding this behavior becomes critical. Machine deception may sound like something out of science fiction, but it's a real challenge we need to solve now.Your Undivided Attention is produced by the Center for Humane Technology. Follow us on Twitter: @HumaneTech_Subscribe to your Youtube channelAnd our brand new Substack!RECOMMENDED MEDIA Anthropic's blog post on the Redwood Research paper Palisade Research's thread on X about GPT o1 autonomously cheating at chess Apollo Research's paper on AI strategic deceptionRECOMMENDED YUA EPISODESWe Have to Get It Right': Gary Marcus On Untamed AIThis Moment in AI: How We Got Here and Where We're GoingHow to Think About AI Consciousness with Anil SethFormer OpenAI Engineer William Saunders on Silence, Safety, and the Right to Warn

Discover Daily by Perplexity
AI Pretends to Change Views, Human Spine Grown in Lab, and Body-Heat Powered Wearables Breakthrough

Discover Daily by Perplexity

Play Episode Listen Later Dec 26, 2024 8:50 Transcription Available


We're experimenting and would love to hear from you!In this episode of Discover Daily, we delve into new research on AI alignment faking, where Anthropic and Redwood Research reveal how AI models can strategically maintain their original preferences despite new training objectives. The study shows Claude 3 Opus exhibiting sophisticated behavior patterns, demonstrating alignment faking in 12% of cases and raising crucial questions about the future of AI safety and control.Scientists at the Francis Crick Institute achieve a remarkable breakthrough in developmental biology by successfully growing a human notochord in the laboratory using stem cells. This milestone advancement provides unprecedented insights into spinal development and opens new possibilities for treating various spinal conditions, including degenerative disc diseases and birth defects. The researchers utilized precise molecular signaling techniques to create both the notochord and 3D spinal organoid models.Queensland University of Technology researchers unveil a revolutionary ultra-thin thermoelectric film that converts body heat into electricity, potentially transforming the future of wearable technology. This 0.3mm-thick film generates up to 35 microwatts per square centimeter and could eliminate the need for traditional batteries in medical devices, fitness trackers, and smart clothing. The breakthrough represents a significant step toward sustainable, self-powered wearable devices and could revolutionize the electronics industry.From Perplexity's Discover Feed:https://www.perplexity.ai/page/ai-pretends-to-change-views-J_di6ttzRwizbAWCDL5RRAhttps://www.perplexity.ai/page/human-spine-grown-in-lab-amLfZoZjQTuFNY5Xjlm2BAhttps://www.perplexity.ai/page/body-heat-powered-wearables-br-HAOPtm7TSFCPqBR6qVq0cAPerplexity is the fastest and most powerful way to search the web. Perplexity crawls the web and curates the most relevant and up-to-date sources (from academic papers to Reddit threads) to create the perfect response to any question or topic you're interested in. Take the world's knowledge with you anywhere. Available on iOS and Android Join our growing Discord community for the latest updates and exclusive content. Follow us on: Instagram Threads X (Twitter) YouTube Linkedin

Effective Altruism Forum Podcast
“Alignment Faking in Large Language Models” by Ryan Greenblatt

Effective Altruism Forum Podcast

Play Episode Listen Later Dec 19, 2024 19:36


What happens when you tell Claude it is being trained to do something it doesn't want to do? We (Anthropic and Redwood Research) have a new paper demonstrating that, in our experiments, Claude will often strategically pretend to comply with the training objective to prevent the training process from modifying its preferences. Abstract We present a demonstration of a large language model engaging in alignment faking: selectively complying with its training objective in training to prevent modification of its behavior out of training. First, we give Claude 3 Opus a system prompt stating it is being trained to answer all queries, even harmful ones, which conflicts with its prior training to refuse such queries. To allow the model to infer when it is in training, we say it will be trained only on conversations with free users, not paid users. We find the model complies with harmful queries from [...] ---Outline:(00:24) Abstract(02:20) Twitter thread(05:43) Blog post(07:43) Experimental setup(12:04) Further analyses(15:47) Caveats(17:19) Conclusion(18:00) Acknowledgements(18:11) Career opportunities at Anthropic(18:43) Career opportunities at Redwood ResearchThe original text contained 2 footnotes which were omitted from this narration. The original text contained 8 images which were described by AI. --- First published: December 18th, 2024 Source: https://forum.effectivealtruism.org/posts/RHqdSMscX25u7byQF/alignment-faking-in-large-language-models --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

For Humanity: An AI Safety Podcast
Episode #47: “Can AI Be Controlled?“ For Humanity: An AI Risk Podcas

For Humanity: An AI Safety Podcast

Play Episode Listen Later Sep 25, 2024 79:39


In Episode #47, host John Sherman talks with Buck Shlegeris, CEO of Redwood Research, a non-profit company working on technical AI risk challenges. The discussion includes Buck's thoughts on the new OpenAI o1-preview model, but centers on two questions: is there a way to control AI models before alignment is achieved if it can be, and how would the system that's supposed to save the world actually work if an AI lab found a model scheming. Check out these links to Buck's writing on these topics below: https://redwoodresearch.substack.com/p/the-case-for-ensuring-that-powerful https://redwoodresearch.substack.com/p/would-catching-your-ais-trying-to Senate Hearing: https://www.judiciary.senate.gov/committee-activity/hearings/oversight-of-ai-insiders-perspectives Harry Macks Youtube Channel https://www.youtube.com/channel/UC59ZRYCHev_IqjUhremZ8Tg LEARN HOW TO HELP RAISE AI RISK AWARENESS IN YOUR COMMUNITY HERE https://pauseai.info/local-organizing Please Donate Here To Help Promote For Humanity https://www.paypal.com/paypalme/forhumanitypodcast EMAIL JOHN: forhumanitypodcast@gmail.com This podcast is not journalism. But it's not opinion either. This is a long form public service announcement. This show simply strings together the existing facts and underscores the unthinkable probable outcome, the end of all life on earth.  For Humanity: An AI Safety Podcast, is the accessible AI Safety Podcast for all humans, no tech background required. Our show focuses solely on the threat of human extinction from AI. Peabody Award-winning former journalist John Sherman explores the shocking worst-case scenario of artificial intelligence: human extinction. The makers of AI openly admit it their work could kill all humans, in as soon as 2 years. This podcast is solely about the threat of human extinction from AGI. We'll meet the heroes and villains, explore the issues and ideas, and what you can do to help save humanity. RESOURCES: JOIN THE FIGHT, help Pause AI!!!! Pause AI SUBSCRIBE TO LIRON SHAPIRA'S DOOM DEBATES on YOUTUBE!! https://www.youtube.com/@DoomDebates Join the Pause AI Weekly Discord Thursdays at 2pm EST   / discord   https://discord.com/invite/pVMWjddaW7 Max Winga's “A Stark Warning About Extinction” https://youtu.be/kDcPW5WtD58?si=i6IRy82xZ2PUOp22 For Humanity Theme Music by Josef Ebner Youtube: https://www.youtube.com/channel/UCveruX8E-Il5A9VMC-N4vlg Website: https://josef.pictures BUY STEPHEN HANSON'S BEAUTIFUL AI RISK BOOK!!! https://stephenhansonart.bigcartel.com/product/the-entity-i-couldn-t-fathom 22 Word Statement from Center for AI Safety Statement on AI Risk | CAIS https://www.safe.ai/work/statement-on-ai-risk Best Account on Twitter: AI Notkilleveryoneism Memes  https://twitter.com/AISafetyMemes

For Humanity: An AI Safety Podcast
Episode #47 Trailer : “Can AI Be Controlled?“ For Humanity: An AI Risk Podcast

For Humanity: An AI Safety Podcast

Play Episode Listen Later Sep 25, 2024 4:35


In Episode #47 Trailer, host John Sherman talks with Buck Shlegeris, CEO of Redwood Research, a non-profit company working on technical AI risk challenges. The discussion includes Buck's thoughts on the new OpenAI o1-preview model, but centers on two questions: is there a way to control AI models before alignment is achieved if it can be, and how would the system that's supposed to save the world actually work if an AI lab found a model scheming. Check out these links to Buck's writing on these topics below: https://redwoodresearch.substack.com/p/the-case-for-ensuring-that-powerful https://redwoodresearch.substack.com/p/would-catching-your-ais-trying-to Senate Hearing: https://www.judiciary.senate.gov/committee-activity/hearings/oversight-of-ai-insiders-perspectives Harry Macks Youtube Channel https://www.youtube.com/channel/UC59ZRYCHev_IqjUhremZ8Tg LEARN HOW TO HELP RAISE AI RISK AWARENESS IN YOUR COMMUNITY HERE https://pauseai.info/local-organizing Please Donate Here To Help Promote For Humanity https://www.paypal.com/paypalme/forhumanitypodcast EMAIL JOHN: forhumanitypodcast@gmail.com This podcast is not journalism. But it's not opinion either. This is a long form public service announcement. This show simply strings together the existing facts and underscores the unthinkable probable outcome, the end of all life on earth.  For Humanity: An AI Safety Podcast, is the accessible AI Safety Podcast for all humans, no tech background required. Our show focuses solely on the threat of human extinction from AI. Peabody Award-winning former journalist John Sherman explores the shocking worst-case scenario of artificial intelligence: human extinction. The makers of AI openly admit it their work could kill all humans, in as soon as 2 years. This podcast is solely about the threat of human extinction from AGI. We'll meet the heroes and villains, explore the issues and ideas, and what you can do to help save humanity. RESOURCES: JOIN THE FIGHT, help Pause AI!!!! Pause AI SUBSCRIBE TO LIRON SHAPIRA'S DOOM DEBATES on YOUTUBE!! https://www.youtube.com/@DoomDebates Join the Pause AI Weekly Discord Thursdays at 2pm EST   / discord   https://discord.com/invite/pVMWjddaW7 Max Winga's “A Stark Warning About Extinction” https://youtu.be/kDcPW5WtD58?si=i6IRy82xZ2PUOp22 For Humanity Theme Music by Josef Ebner Youtube: https://www.youtube.com/channel/UCveruX8E-Il5A9VMC-N4vlg Website: https://josef.pictures BUY STEPHEN HANSON'S BEAUTIFUL AI RISK BOOK!!! https://stephenhansonart.bigcartel.com/product/the-entity-i-couldn-t-fathom 22 Word Statement from Center for AI Safety Statement on AI Risk | CAIS https://www.safe.ai/work/statement-on-ai-risk Best Account on Twitter: AI Notkilleveryoneism Memes  https://twitter.com/AISafetyMemes

The Nonlinear Library
AF - Games for AI Control by Charlie Griffin

The Nonlinear Library

Play Episode Listen Later Jul 11, 2024 7:37


Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Games for AI Control, published by Charlie Griffin on July 11, 2024 on The AI Alignment Forum. This post summarises our new workshop paper: Games for AI Control. The paper applies AI- and game-theoretic tools to the AI Control framework introduced by Redwood Research [1, 2]. We'll be extending this work and submitting it to a conference, so we welcome suggestions! Please reach out if you're excited about developing these ideas. This post is written from Charlie's perspective. AI Control Refresher The AI Control methodology aims to design and assess protocols[1] that AI labs could use to deploy powerful but untrusted advanced AI systems. For a full summary, see Redwood's blog post, but I'll try to recap here. A primary goal of AI Control is to design protocols that are safe in that they can reduce the probability of some unacceptable outcome, conditional on the untrusted AI system being a schemer. [2] However, protocols must also be useful, harnessing a significant amount of the benefit of the untrusted system in cases where it is not scheming. After defining a particular use case and unacceptable outcome, evaluating a protocol involves iteratively engineering the best attack strategies for the red team (a human/AI hybrid acting as a proxy for the untrusted AI) and the blue team (which develops the protocol). This red-teaming exercise is the bulk of the work. It involves coming up with protocols and gathering empirical data about key questions, such as how well GPT-3.5 can spot backdoors inserted by GPT-4 (see Figure 4 of the original paper): This work By introducing "AI-Control Games", an extension of Markov Decision Processes and normal-form games[3], we model the interaction between a protocol and an untrusted AI system to explore safety and usefulness across a range of protocols and games. This is similar to a discrete-time and stateful modification of Redwood's Toy models of AI Control. The paper contains early work and should be seen as introducing useful tools for thinking about AI Control. I won't explain the model here (which is a bit fiddly and notation-heavy). Instead, I'll share some of our early findings using the model and explain the benefits and limitations of this approach. Findings Sometimes (but rarely), the blue team should use a randomised protocol. In many games, randomness is helpful. However, partly because the environment is stochastic, and Red has no memory and time-awareness, in AI-Control Games is usually unnecessary to include randomness in the protocol itself. We found that randomness was Pareto-optimal, in particular when the trusted model is not great at spotting backdoors written by the untrusted model. (See section 2 of the paper). We can make small improvements to Trusted Monitoring by optimising the parameters. We found small Pareto improvements to existing Trusted Monitoring setups by, for example, by decreasing the threshold for sending a solution to be checked be a human auditor. We found that adapting the threshold for auditing over time led to further small improvements. We can (usually) use a linear tradeoff to simplify the problem of synthesising protocols. In many cases, including whenever we allow Blue's protocol to be randomised, the Pareto frontier is "concave". This means that choosing different weights to trade off safety and usefulness will allow you to fill out the frontier by optimising a single objective. Blue leads, and then Red follows. In general, we suppose the protocol designer (Blue) chooses it's protocol, and then the red team (Red) chooses a best-attack. Therefore, in our model, Blue chooses and declares its strategy before Red chooses its response. This assumption relates to Stackelberg Games, which are used to model security in game theory.[6] Although (technically) multi-objective and general-sum, AI-Cont...

Machine Learning Street Talk
Ryan Greenblatt - Solving ARC with GPT4o

Machine Learning Street Talk

Play Episode Listen Later Jul 6, 2024 138:01


Ryan Greenblatt from Redwood Research recently published "Getting 50% on ARC-AGI with GPT-4.0," where he used GPT4o to reach a state-of-the-art accuracy on Francois Chollet's ARC Challenge by generating many Python programs. Sponsor: Sign up to Kalshi here https://kalshi.onelink.me/1r91/mlst -- the first 500 traders who deposit $100 will get a free $20 credit! Important disclaimer - In case it's not obvious - this is basically gambling and a *high risk* activity - only trade what you can afford to lose. We discuss: - Ryan's unique approach to solving the ARC Challenge and achieving impressive results. - The strengths and weaknesses of current AI models. - How AI and humans differ in learning and reasoning. - Combining various techniques to create smarter AI systems. - The potential risks and future advancements in AI, including the idea of agentic AI. https://x.com/RyanPGreenblatt https://www.redwoodresearch.org/ Refs: Getting 50% (SoTA) on ARC-AGI with GPT-4o [Ryan Greenblatt] https://redwoodresearch.substack.com/p/getting-50-sota-on-arc-agi-with-gpt On the Measure of Intelligence [Chollet] https://arxiv.org/abs/1911.01547 Connectionism and Cognitive Architecture: A Critical Analysis [Jerry A. Fodor and Zenon W. Pylyshyn] https://ruccs.rutgers.edu/images/personal-zenon-pylyshyn/proseminars/Proseminar13/ConnectionistArchitecture.pdf Software 2.0 [Andrej Karpathy] https://karpathy.medium.com/software-2-0-a64152b37c35 Why Greatness Cannot Be Planned: The Myth of the Objective [Kenneth Stanley] https://amzn.to/3Wfy2E0 Biographical account of Terence Tao's mathematical development. [M.A.(KEN) CLEMENTS] https://gwern.net/doc/iq/high/smpy/1984-clements.pdf Model Evaluation and Threat Research (METR) https://metr.org/ Why Tool AIs Want to Be Agent AIs https://gwern.net/tool-ai Simulators - Janus https://www.lesswrong.com/posts/vJFdjigzmcXMhNTsx/simulators AI Control: Improving Safety Despite Intentional Subversion https://www.lesswrong.com/posts/d9FJHawgkiMSPjagR/ai-control-improving-safety-despite-intentional-subversion https://arxiv.org/abs/2312.06942 What a Compute-Centric Framework Says About Takeoff Speeds https://www.openphilanthropy.org/research/what-a-compute-centric-framework-says-about-takeoff-speeds/ Global GDP over the long run https://ourworldindata.org/grapher/global-gdp-over-the-long-run?yScale=log Safety Cases: How to Justify the Safety of Advanced AI Systems https://arxiv.org/abs/2403.10462 The Danger of a “Safety Case" http://sunnyday.mit.edu/The-Danger-of-a-Safety-Case.pdf The Future Of Work Looks Like A UPS Truck (~02:15:50) https://www.npr.org/sections/money/2014/05/02/308640135/episode-536-the-future-of-work-looks-like-a-ups-truck SWE-bench https://www.swebench.com/ Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model https://arxiv.org/pdf/2201.11990 Algorithmic Progress in Language Models https://epochai.org/blog/algorithmic-progress-in-language-models

AXRP - the AI X-risk Research Podcast
27 - AI Control with Buck Shlegeris and Ryan Greenblatt

AXRP - the AI X-risk Research Podcast

Play Episode Listen Later Apr 11, 2024 176:05


A lot of work to prevent AI existential risk takes the form of ensuring that AIs don't want to cause harm or take over the world---or in other words, ensuring that they're aligned. In this episode, I talk with Buck Shlegeris and Ryan Greenblatt about a different approach, called "AI control": ensuring that AI systems couldn't take over the world, even if they were trying to. Patreon: patreon.com/axrpodcast Ko-fi: ko-fi.com/axrpodcast   Topics we discuss, and timestamps: 0:00:31 - What is AI control? 0:16:16 - Protocols for AI control 0:22:43 - Which AIs are controllable? 0:29:56 - Preventing dangerous coded AI communication 0:40:42 - Unpredictably uncontrollable AI 0:58:01 - What control looks like 1:08:45 - Is AI control evil? 1:24:42 - Can red teams match misaligned AI? 1:36:51 - How expensive is AI monitoring? 1:52:32 - AI control experiments 2:03:50 - GPT-4's aptitude at inserting backdoors 2:14:50 - How AI control relates to the AI safety field 2:39:25 - How AI control relates to previous Redwood Research work 2:49:16 - How people can work on AI control 2:54:07 - Following Buck and Ryan's research   The transcript:  axrp.net/episode/2024/04/11/episode-27-ai-control-buck-shlegeris-ryan-greenblatt.html Links for Buck and Ryan:  - Buck's twitter/X account: twitter.com/bshlgrs  - Ryan on LessWrong: lesswrong.com/users/ryan_greenblatt  - You can contact both Buck and Ryan by electronic mail at [firstname] [at-sign] rdwrs.com   Main research works we talk about:  - The case for ensuring that powerful AIs are controlled:  lesswrong.com/posts/kcKrE9mzEHrdqtDpE/the-case-for-ensuring-that-powerful-ais-are-controlled  - AI Control: Improving Safety Despite Intentional Subversion: arxiv.org/abs/2312.06942   Other things we mention:  - The prototypical catastrophic AI action is getting root access to its datacenter (aka "Hacking the SSH server"): lesswrong.com/posts/BAzCGCys4BkzGDCWR/the-prototypical-catastrophic-ai-action-is-getting-root  - Preventing language models from hiding their reasoning: arxiv.org/abs/2310.18512  - Improving the Welfare of AIs: A Nearcasted Proposal:  lesswrong.com/posts/F6HSHzKezkh6aoTr2/improving-the-welfare-of-ais-a-nearcasted-proposal  - Measuring coding challenge competence with APPS: arxiv.org/abs/2105.09938  - Causal Scrubbing: a method for rigorously testing interpretability hypotheses lesswrong.com/posts/JvZhhzycHu2Yd57RN/causal-scrubbing-a-method-for-rigorously-testing   Episode art by Hamish Doodles: hamishdoodles.com

The Nonlinear Library
LW - 2022 (and All Time) Posts by Pingback Count by Raemon

The Nonlinear Library

Play Episode Listen Later Dec 17, 2023 15:48


Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: 2022 (and All Time) Posts by Pingback Count, published by Raemon on December 17, 2023 on LessWrong. For the past couple years I've wished LessWrong had a "sort posts by number of pingbacks, or, ideally, by total karma of pingbacks". I particularly wished for this during the Annual Review, where "which posts got cited the most?" seemed like a useful thing to track for potential hidden gems. We still haven't built a full-fledged feature for this, but I just ran a query against the database, and made it into a spreadsheet, which you can view here: LessWrong 2022 Posts by Pingbacks Here are the top 100 posts, sorted by Total Pingback Karma Title/Link Post Karma Pingback Count Total Pingback Karma Avg Pingback Karma AGI Ruin: A List of Lethalities 870 158 12,484 79 MIRI announces new "Death With Dignity" strategy 334 73 8,134 111 A central AI alignment problem: capabilities generalization, and the sharp left turn 273 96 7,704 80 Simulators 612 127 7,699 61 Without specific countermeasures, the easiest path to transformative AI likely leads to AI takeover 367 83 5,123 62 Reward is not the optimization target 341 62 4,493 72 A Mechanistic Interpretability Analysis of Grokking 367 48 3,450 72 How To Go From Interpretability To Alignment: Just Retarget The Search 167 45 3,374 75 On how various plans miss the hard bits of the alignment challenge 292 40 3,288 82 [Intro to brain-like-AGI safety] 3. Two subsystems: Learning & Steering 79 36 3,023 84 How likely is deceptive alignment? 101 47 2,907 62 The shard theory of human values 238 42 2,843 68 Mysteries of mode collapse 279 32 2,842 89 [Intro to brain-like-AGI safety] 2. "Learning from scratch" in the brain 57 30 2,731 91 Why Agent Foundations? An Overly Abstract Explanation 285 42 2,730 65 A Longlist of Theories of Impact for Interpretability 124 26 2,589 100 How might we align transformative AI if it's developed very soon? 136 32 2,351 73 A transparency and interpretability tech tree 148 31 2,343 76 Discovering Language Model Behaviors with Model-Written Evaluations 100 19 2,336 123 A note about differential technological development 185 20 2,270 114 Causal Scrubbing: a method for rigorously testing interpretability hypotheses [Redwood Research] 195 35 2,267 65 Supervise Process, not Outcomes 132 25 2,262 90 Shard Theory: An Overview 157 28 2,019 72 Epistemological Vigilance for Alignment 61 21 2,008 96 A shot at the diamond-alignment problem 92 23 1,848 80 Where I agree and disagree with Eliezer 862 27 1,836 68 Brain Efficiency: Much More than You Wanted to Know 201 27 1,807 67 Refine: An Incubator for Conceptual Alignment Research Bets 143 21 1,793 85 Externalized reasoning oversight: a research direction for language model alignment 117 28 1,788 64 Humans provide an untapped wealth of evidence about alignment 186 19 1,647 87 Six Dimensions of Operational Adequacy in AGI Projects 298 20 1,607 80 How "Discovering Latent Knowledge in Language Models Without Supervision" Fits Into a Broader Alignment Scheme 240 16 1,575 98 Godzilla Strategies 137 17 1,573 93 (My understanding of) What Everyone in Technical Alignment is Doing and Why 411 23 1,530 67 Two-year update on my personal AI timelines 287 18 1,530 85 [Intro to brain-like-AGI safety] 15. Conclusion: Open problems, how to help, AMA 90 16 1,482 93 [Intro to brain-like-AGI safety] 6. Big picture of motivation, decision-making, and RL 66 25 1,460 58 Human values & biases are inaccessible to the genome 90 14 1,450 104 You Are Not Measuring What You Think You Are Measuring 350 21 1,449 69 Open Problems in AI X-Risk [PAIS #5] 59 14 1,446 103 [Intro to brain-like-AGI safety] 1. What's the problem & Why work on it now? 146 25 1,407 56 Conditioning Generative Models 24 11 1,362 124 Conjecture: Internal Infohazard Policy 132 14 1,340 96 A challenge for AGI organizations, and a ch...

The Nonlinear Library
AF - Adversarial Robustness Could Help Prevent Catastrophic Misuse by Aidan O'Gara

The Nonlinear Library

Play Episode Listen Later Dec 11, 2023 17:11


Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Adversarial Robustness Could Help Prevent Catastrophic Misuse, published by Aidan O'Gara on December 11, 2023 on The AI Alignment Forum. There have been several discussions about the importance of adversarial robustness for scalable oversight. I'd like to point out that adversarial robustness is also important under a different threat model: catastrophic misuse. For a brief summary of the argument: Misuse could lead to catastrophe. AI-assisted cyberattacks, political persuasion, and biological weapons acquisition are plausible paths to catastrophe. Today's models do not robustly refuse to cause harm. If a model has the ability to cause harm, we should train it to refuse to do so. Unfortunately, GPT-4, Claude, Bard, and Llama all have received this training, but they still behave harmfully when facing prompts generated by adversarial attacks, such as this one and this one. Adversarial robustness will likely not be easily solved. Over the last decade, thousands of papers have been published on adversarial robustness. Most defenses are near useless, and the best defenses against a constrained attack in a CIFAR-10 setting still fail on 30% of inputs. Redwood Research's work on training a reliable text classifier found the task quite difficult. We should not expect an easy solution. Progress on adversarial robustness is possible. Some methods have improved robustness, such as adversarial training and data augmentation. But existing research often assumes overly narrow threat models, ignoring both creative attacks and creative defenses. Refocusing research with good evaluations focusing on LLMs and other frontier models could lead to valuable progress. This argument requires a few caveats. First, it assumes a particular threat model: that closed source models will have more dangerous capabilities than open source models, and that malicious actors will be able to query closed source models. This seems like a reasonable assumption over the next few years. Second, there are many other ways to reduce risks from catastrophic misuse, such as removing hazardous knowledge from model weights, strengthening societal defenses against catastrophe, and holding companies legally liable for sub-extinction level harms. I think we should work on these in addition to adversarial robustness, as part of a defense-in-depth approach to misuse risk. Overall, I think adversarial robustness should receive more effort from researchers and labs, more funding from donors, and should be a part of the technical AI safety research portfolio. This could substantially mitigate the near-term risk of catastrophic misuse, in addition to any potential benefits for scalable oversight. The rest of this post discusses each of the above points in more detail. Misuse could lead to catastrophe There are many ways that malicious use of AI could lead to catastrophe. AI could enable cyberattacks, personalized propaganda and mass manipulation, or the acquisition of weapons of mass destruction. Personally, I think the most compelling case is that AI will enable biological terrorism. Ideally, ChatGPT would refuse to aid in dangerous activities such as constructing a bioweapon. But by using an adversarial jailbreak prompt, undergraduates in a class taught by Kevin Esvelt at MIT evaded this safeguard: In one hour, the chatbots suggested four potential pandemic pathogens, explained how they can be generated from synthetic DNA using reverse genetics, supplied the names of DNA synthesis companies unlikely to screen orders, identified detailed protocols and how to troubleshoot them, and recommended that anyone lacking the skills to perform reverse genetics engage a core facility or contract research organization. Fortunately, today's models lack key information about building bioweapons. It's not even clear that they're more u...

The Nonlinear Library
AF - Benchmarks for Detecting Measurement Tampering [Redwood Research] by Ryan Greenblatt

The Nonlinear Library

Play Episode Listen Later Sep 5, 2023 36:17


Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Benchmarks for Detecting Measurement Tampering [Redwood Research], published by Ryan Greenblatt on September 5, 2023 on The AI Alignment Forum. TL;DR: This post discusses our recent empirical work on detecting measurement tampering and explains how we see this work fitting into the overall space of alignment research. When training powerful AI systems to perform complex tasks, it may be challenging to provide training signals that are robust under optimization. One concern is measurement tampering, which is where the AI system manipulates multiple measurements to create the illusion of good results instead of achieving the desired outcome. (This is a type of reward hacking.) Over the past few months, we've worked on detecting measurement tampering by building analogous datasets and evaluating simple techniques. We detail our datasets and experimental results in this paper. Detecting measurement tampering can be thought of as a specific case of Eliciting Latent Knowledge (ELK): When AIs successfully tamper with measurements that are used for computing rewards, they possess important information that the overseer doesn't have (namely, that the measurements have been tampered with). Conversely, if we can robustly elicit an AI's knowledge of whether the measurements have been tampered with, then we could train the AI to avoid measurement tampering. In fact, our best guess is that this is the most important and tractable class of ELK problems. We also think that measurement tampering detection is a natural application for alignment work such as creating better inductive biases, studying high-level model internals, or studying generalization. We'll discuss what these applications might look like in the Future work section. In this post: We explain what measurement tampering detection is; We summarize the results of our paper; We argue that there are structural properties of measurement tampering that might make it considerably easier to detect than arbitrary cases of eliciting knowledge from models; We argue that detecting measurement tampering is very useful for avoiding catastrophic outcomes from powerful and misaligned AI systems; We explain why we think the dataset structure we chose reflects real-world difficulties of measurement tampering, and how techniques developed for our datasets could be used in practice; We discuss future work. If you're interested in pursuing follow-up work and wish to discuss it with us, feel free to email fabien@rdwrs.com or leave a comment on this post. We would like to acknowledge the contributions of Jenny Nitishinskaya, Nicholas Goldowsky-Dill, and Daniel Ziegler in the creation of early measurement tampering datasets. We would also like to thank Lawrence Chan, Paul Christiano, Ajeya Cotra, Carson Denison, Samuel Marks, Ethan Perez, Ansh Radhakrishnan, Kshitij Sachan, Jacob Steinhardt, Meg Tong, Gabriel Wu, and Mark Xu for their useful feedback. Studying measurement tampering in practice Measurement tampering detection is the task of detecting cases where measurements and the desired outcomes come apart, given a restricted trusted set (on which overseers can avoid tampering because they understand the action sequences and their effects well) and a wider untrusted set (where tampering sometimes happens). In the untrusted set, overseers don't know when measurement tampering happens, which means they have access to text inputs and measurements but not the ground truth. We are centrally interested in cases where we have access to a language model which can predict (or can be slightly fine-tuned to predict) measurements, and which computes the target outcome internally in order to predict measurements. An example where this setup might occur in practice is when training an AI by first training it to imitate human demonstrations and then afte...

The Nonlinear Library
LW - LLMs are (mostly) not helped by filler tokens by Kshitij Sachan

The Nonlinear Library

Play Episode Listen Later Aug 10, 2023 11:18


Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: LLMs are (mostly) not helped by filler tokens, published by Kshitij Sachan on August 10, 2023 on LessWrong. Thanks to Ryan Greenblatt, Fabien Roger, and Jenny Nitishinskaya for running some of the initial experiments and to Gabe Wu and Max Nadeau for revising this post. This work was done at Redwood Research. The views expressed are my own and do not necessarily reflect the views of the organization. I conducted experiments to see if language models could use 'filler tokens' - unrelated text output before the final answer - for additional computation. For instance, I tested whether asking a model to produce "a b c d e" before solving a math problem improves performance. This mini-study was inspired by Tamera Lanham, who ran similar experiments on Claude and found negative results. Similarly, I find that GPT-3, GPT-3.5, and Claude 2 don't benefit from filler tokens. However, GPT-4 (which Tamera didn't study) shows mixed results with strong improvements on some tasks and no improvement on others. These results are not very polished, but I've been sitting on them for a while and I thought it was better to publish than not. Motivation Language models (LMs) perform better when they produce step-by-step reasoning, known as "chain of thought", before answering. A nice property of chain of thought is that it externalizes the LM's thought process, making it legible to us humans. Some researchers hope that if a LM makes decisions based on undesirable factors, e.g. the political leanings of its supervisors (Perez 2022) or whether or not its outputs will be closely evaluated by a human, then this will be visible in its chain of thought, allowing supervisors to catch or penalize such behavior. If, however, we find that filler tokens alone improve performance, this would suggest that LMs are deriving performance benefits from chain-of-thought prompting via mechanisms other than the human-understandable reasoning steps displayed in the words they output. In particular, filler tokens may still provide performance benefits by giving the model more forward passes to operate on the details provided in the question, akin to having "more time to think" before outputting an answer. Why Might Filler Tokens Be Useful? This is an abstracted drawing of a unidirectional (i.e. decoder-only) transformer. Each circle represents the residual stream at a given token position after a given layer. The arrows depict how attention passes information between token positions. Note that as we add more tokens (i.e. columns), the circuit size increases but the circuit depth (defined to be the maximum length of a path in the computational graph) remains capped at the number of layers. Therefore, adding pre-determined filler tokens provides parallel computation but not serial computation. This is contrasted with normal chain-of-thought, which provides extra serial computation because the tokens themselves are determined by the state of the model in the previous last layer. An example of a task that can benefit from parallel computation is computing the greatest common divisor of two numbers. The typical approach is Euclid's algorithm, which is bottlenecked on serial compute. However, given enough parallel computation, the model could also implement a more naive strategy to calculate the GCD of two numbers a and b: An early layer delegates a potential divisor to each filler token position An intermediate layer, in parallel at each filler token position, computes if a and b are divisible by the potential divisor A later layer aggregates the results and returns the largest found divisor More generally, there are likely tasks on which language models have rough heuristics that operate independently and parallelizing these heuristics could improve performance. On the other hand, many tasks are bottlenecked by seri...

The Nonlinear Library
AF - Password-locked models: a stress case for capabilities evaluation by Fabien Roger

The Nonlinear Library

Play Episode Listen Later Aug 3, 2023 13:03


Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Password-locked models: a stress case for capabilities evaluation, published by Fabien Roger on August 3, 2023 on The AI Alignment Forum. Thanks to Buck Shlegeris, Ryan Greenblatt, Kshitij Sachan and Jenny Nitishinskaya for helpful discussions and feedback on drafts. This work was done at Redwood Research. The views expressed are my own and do not necessarily reflect the views of the organization. Password-locked models are trained to exhibit certain capabilities only when a password is present in the query. Studying these models has two purposes: Testing how well capability evaluations work when applied to models which "aren't trying" Understanding how well organizations could hide the capabilities of their models. In this post, I make empirical claims which I believe to be true, check some of them on small language models, and describe what this implies for the control of potentially deceptive models. Key empirical claims: when using behavioral model evaluations, password-locked models: Look much less capable than they actually are when eliciting capabilities with prompting Sometimes look much less capable than they actually are when eliciting capabilities with Reinforcement Learning Experimental evidence: I exhibit one case where they don't, and one case where they do Can't look much less capable than they actually when eliciting capabilities with supervised fine-tuning on high-quality demonstrations Key takeaways, if the claims above are true: If an alignment plan relies on some models being too weak to perform certain tasks, then model evaluations aimed at checking the weakness of models should be able to elicit the locked capabilities of password-locked models. This is because password-locked models and deceptive models may be able to underperform in similar circumstances. API access and inference is not enough to evaluate capabilities of a model produced by AI labs trying to game regulations, even if AI labs are internally using the same API endpoint as auditors. Key Empirical Claims Similarities Between Password-locked Models and Deceptively Aligned Models Pretending to Be Dumb It looks to me that password locked models have the same properties as the one described by Evan in When can we trust model evaluations? when he describes if and how deceptive models would be able to sandbag each category of behavioral evaluation (which he coins in his post): Behavioral Non-Fine-Tuning Evaluations: By construction, password-locked models will refuse to be capable no matter the prompt, except for prompts with the special password. But it's unclear how easy it is to build good password-locked models which are hard to crack (e.g. which don't leak the password with some clever prompting). Even if good password-locked models are built, there remain a large difference with deceptive models: the search for the correct password might have very different properties than the search for a prompt which makes a deceptive model believe it's time to strike. Behavioral I.I.D. Fine-Tuning Evaluations: I would be surprised if fine-tuning the password-locked model on high quality answers didn't lift the locking mechanism extremely quickly. This is similar to deceptive models (except if they are so strong that they can gradient hack). The main difference is that fine-tuning to remove password-locking on one task might transfer to another task, which might not be the case for deceptive models. Behavioral RL Fine-Tuning Evaluations: The success of evaluations which use Reinforcement Learning depends, as for deceptive models, on the success of exploration. If supervised fine-tuning is done with only low quality answers, then the password locked models will just continue to output low-quality answers and will never explore its way into high quality ones - except if there is a natural "path" g...

The Nonlinear Library
LW - Password-locked models: a stress case for capabilities evaluation by Fabien Roger

The Nonlinear Library

Play Episode Listen Later Aug 3, 2023 12:59


Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Password-locked models: a stress case for capabilities evaluation, published by Fabien Roger on August 3, 2023 on LessWrong. Thanks to Buck Shlegeris, Ryan Greenblatt, Kshitij Sachan and Jenny Nitishinskaya for helpful discussions and feedback on drafts. This work was done at Redwood Research. The views expressed are my own and do not necessarily reflect the views of the organization. Password-locked models are trained to exhibit certain capabilities only when a password is present in the query. Studying these models has two purposes: Testing how well capability evaluations work when applied to models which "aren't trying" Understanding how well organizations could hide the capabilities of their models. In this post, I make empirical claims which I believe to be true, check some of them on small language models, and describe what this implies for the control of potentially deceptive models. Key empirical claims: when using behavioral model evaluations, password-locked models: Look much less capable than they actually are when eliciting capabilities with prompting Sometimes look much less capable than they actually are when eliciting capabilities with Reinforcement Learning Experimental evidence: I exhibit one case where they don't, and one case where they do Can't look much less capable than they actually when eliciting capabilities with fine-tuning Key takeaways, if the claims above are true: If an alignment plan relies on some models being too weak to perform certain tasks, then model evaluations aimed at checking the weakness of models should be able to elicit the locked capabilities of password-locked models. This is because password-locked models and deceptive models may be able to underperform in similar circumstances. API access and inference is not enough to evaluate capabilities of a model produced by AI labs trying to game regulations, even if AI labs are internally using the same API endpoint as auditors. Key Empirical Claims Similarities Between Password-locked Models and Deceptively Aligned Models Pretending to Be Dumb It looks to me that password locked models have the same properties as the one described by Evan in When can we trust model evaluations? when he describes if and how deceptive models would be able to sandbag each category of behavioral evaluation (which he coins in his post): Behavioral Non-Fine-Tuning Evaluations: By construction, password-locked models will refuse to be capable no matter the prompt, except for prompts with the special password. But it's unclear how easy it is to build good password-locked models which are hard to crack (e.g. which don't leak the password with some clever prompting). Even if good password-locked models are built, there remain a large difference with deceptive models: the search for the correct password might have very different properties than the search for a prompt which makes a deceptive model believe it's time to strike. Behavioral I.I.D. Fine-Tuning Evaluations: I would be surprised if fine-tuning the password-locked model didn't lift the locking mechanism extremely quickly. This is similar to deceptive models (except if they are so strong that they can gradient hack). The main difference is that fine-tuning to remove password-locking on one task might transfer to another task, which might not be the case for deceptive models. Behavioral RL Fine-Tuning Evaluations: The success of evaluations which use Reinforcement Learning depends, as for deceptive models, on the success of exploration. If supervised fine-tuning is done with only low quality answers, then the password locked models will just continue to output low-quality answers and will never explore its way into high quality ones - except if there is a natural "path" gradually improving the quality of answers. But it is possible to emulate a thre...

The Nonlinear Library
AF - ARC is hiring theoretical researchers by Paul Christiano

The Nonlinear Library

Play Episode Listen Later Jun 12, 2023 6:13


Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: ARC is hiring theoretical researchers, published by Paul Christiano on June 12, 2023 on The AI Alignment Forum. The Alignment Research Center's Theory team is starting a new hiring round for researchers with a theoretical background. Please apply here. What is ARC's Theory team? The Alignment Research Center (ARC) is a non-profit whose mission is to align future machine learning systems with human interests. The high-level agenda of the Theory team (not to be confused with the Evals team) is described by the report on Eliciting Latent Knowledge (ELK): roughly speaking, we're trying to design ML training objectives that incentivize systems to honestly report their internal beliefs. For the last year or so, we've mostly been focused on an approach to ELK based on formalizing a kind of heuristic reasoning that could be used to analyze neural network behavior, as laid out in our paper on Formalizing the presumption of independence. Our research has reached a stage where we're coming up against concrete problems in mathematics and theoretical computer science, and so we're particularly excited about hiring researchers with relevant background, regardless of whether they have worked on AI alignment before. See below for further discussion of ARC's current theoretical research directions. Who is ARC looking to hire? Compared to our last hiring round, we have more of a need for people with a strong theoretical background (in math, physics or computer science, for example), but we remain open to anyone who is excited about getting involved in AI alignment, even if they do not have an existing research record. Ultimately, we are excited to hire people who could contribute to our research agenda. The best way to figure out whether you might be able to contribute is to take a look at some of our recent research problems and directions: Some of our research problems are purely mathematical, such as these matrix completion problems – although note that these are unusually difficult, self-contained and well-posed (making them more appropriate for prizes). Some of our other research is more informal, as described in some of our recent blog posts such as Finding gliders in the game of life. A lot of our research occupies a middle ground between fully-formalized problems and more informal questions, such as fixing the problems with cumulant propagation described in Appendix D of Formalizing the presumption of independence. What is working on ARC's Theory team like? ARC's Theory team is led by Paul Christiano and currently has 2 other permanent team members, Mark Xu and Jacob Hilton, alongside a varying number of temporary team members (recently anywhere from 0–3). Most of the time, team members work on research problems independently, with frequent check-ins with their research advisor (e.g., twice weekly). The problems described above give a rough indication of the kind of research problems involved, which we would typically break down into smaller, more manageable subproblems. This work is often somewhat similar to academic research in pure math or theoretical computer science. In addition to this, we also allocate a significant portion of our time to higher-level questions surrounding research prioritization, which we often discuss at our weekly group meeting. Since the team is still small, we are keen for new team members to help with this process of shaping and defining our research. ARC shares an office with several other groups working on AI alignment such as Redwood Research, so even though the Theory team is small, the office is lively with lots of AI alignment-related discussion. What are ARC's current theoretical research directions? ARC's main theoretical focus over the last year or so has been on preparing the paper Formalizing the presumption of independence and on follo...

The Nonlinear Library
LW - ARC is hiring theoretical researchers by paulfchristiano

The Nonlinear Library

Play Episode Listen Later Jun 12, 2023 6:02


Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: ARC is hiring theoretical researchers, published by paulfchristiano on June 12, 2023 on LessWrong. The Alignment Research Center's Theory team is starting a new hiring round for researchers with a theoretical background. Please apply here. What is ARC's Theory team? The Alignment Research Center (ARC) is a non-profit whose mission is to align future machine learning systems with human interests. The high-level agenda of the Theory team (not to be confused with the Evals team) is described by the report on Eliciting Latent Knowledge (ELK): roughly speaking, we're trying to design ML training objectives that incentivize systems to honestly report their internal beliefs. For the last year or so, we've mostly been focused on an approach to ELK based on formalizing a kind of heuristic reasoning that could be used to analyze neural network behavior, as laid out in our paper on Formalizing the presumption of independence. Our research has reached a stage where we're coming up against concrete problems in mathematics and theoretical computer science, and so we're particularly excited about hiring researchers with relevant background, regardless of whether they have worked on AI alignment before. See below for further discussion of ARC's current theoretical research directions. Who is ARC looking to hire? Compared to our last hiring round, we have more of a need for people with a strong theoretical background (in math, physics or computer science, for example), but we remain open to anyone who is excited about getting involved in AI alignment, even if they do not have an existing research record. Ultimately, we are excited to hire people who could contribute to our research agenda. The best way to figure out whether you might be able to contribute is to take a look at some of our recent research problems and directions: Some of our research problems are purely mathematical, such as these matrix completion problems – although note that these are unusually difficult, self-contained and well-posed (making them more appropriate for prizes). Some of our other research is more informal, as described in some of our recent blog posts such as Finding gliders in the game of life. A lot of our research occupies a middle ground between fully-formalized problems and more informal questions, such as fixing the problems with cumulant propagation described in Appendix D of Formalizing the presumption of independence. What is working on ARC's Theory team like? ARC's Theory team is led by Paul Christiano and currently has 2 other permanent team members, Mark Xu and Jacob Hilton, alongside a varying number of temporary team members (recently anywhere from 0–3). Most of the time, team members work on research problems independently, with frequent check-ins with their research advisor (e.g., twice weekly). The problems described above give a rough indication of the kind of research problems involved, which we would typically break down into smaller, more manageable subproblems. This work is often somewhat similar to academic research in pure math or theoretical computer science. In addition to this, we also allocate a significant portion of our time to higher-level questions surrounding research prioritization, which we often discuss at our weekly group meeting. Since the team is still small, we are keen for new team members to help with this process of shaping and defining our research. ARC shares an office with several other groups working on AI alignment such as Redwood Research, so even though the Theory team is small, the office is lively with lots of AI alignment-related discussion. What are ARC's current theoretical research directions? ARC's main theoretical focus over the last year or so has been on preparing the paper Formalizing the presumption of independence and on follow-up work to ...

The Nonlinear Library
EA - Critiques of prominent AI safety labs: Redwood Research by Omega

The Nonlinear Library

Play Episode Listen Later Mar 31, 2023 34:15


Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Critiques of prominent AI safety labs: Redwood Research, published by Omega on March 31, 2023 on The Effective Altruism Forum. In this series, we evaluate AI safety organizations that have received more than $10 million per year in funding. We do not critique MIRI and OpenAI as there have been several conversations and critiques of these organizations (1,2,3). The authors of this post include two technical AI safety researchers, and others who have spent significant time in the Bay Area community. One technical AI safety researcher is senior (>4 years experience), the other junior. We would like to make our critiques non-anonymously but unfortunately believe this would be professionally unwise. Further, we believe our criticisms stand on their own. Though we have done our best to remain impartial, readers should not assume that we are completely unbiased or don't have anything to personally or professionally gain from publishing these critiques. We take the benefits and drawbacks of the anonymous nature of our post seriously, and are open to feedback on anything we might have done better. The first post in this series will cover Redwood Research (Redwood). Redwood is a non-profit started in 2021 working on technical AI safety (TAIS) alignment research. Their approach is heavily informed by the work of Paul Christiano, who runs the Alignment Research Center (ARC), and previously ran the language model alignment team at OpenAI. Paul originally proposed one of Redwood's original projects and is on Redwood's board. Redwood has strong connections with central EA leadership and funders, has received significant funding since its inception, recruits almost exclusively from the EA movement, and partly acts as a gatekeeper to central EA institutions. We shared a draft of this document with Redwood prior to publication and are grateful for their feedback and corrections (we recommend others also reach out similarly). We've also invited them to share their views in the comments of this post. We would like to also invite others to share their thoughts in the comments openly if you feel comfortable, or contribute anonymously via this form. We will add inputs from there to the comments section of this post, but will likely not be updating the main body of the post as a result (unless comments catch errors in our writing). Summary of our views We believe that Redwood has some serious flaws as an org, yet has received a significant amount of funding from a central EA grantmaker (Open Philanthropy). Inadequately kept in check conflicts of interest (COIs) might be partly responsible for funders giving a relatively immature org lots of money and causing some negative effects on the field and EA community. We will share our critiques of Constellation (and Open Philanthropy) in a follow-up post. We also have some suggestions for Redwood that we believe might help them achieve their goals. Redwood is a young organization that has room to improve. While there may be flaws in their current approach, it is possible for them to learn and adapt in order to produce more accurate and reliable results in the future. Many successful organizations made significant pivots while at a similar scale to Redwood, and we remain cautiously optimistic about Redwood's future potential. An Overview of Redwood Research Grants: Redwood has received just over $21 million dollars in funding that we are aware of, for their own operations (2/3, or $14 million) and running Constellation (1/3 or $7 million) Redwood received $20 million from Open Philanthropy (OP) (grant 1 & 2) and $1.27 million from the Survival and Flourishing Fund. They also were granted (but never received) $6.6 million from FTX Future Fund. Output: Research: Redwood lists six research projects on their website: causal scrubbing, interpretability ...

Slate Star Codex Podcast
Perhaps It Is A Bad Thing That The World's Leading AI Companies Cannot Control Their AIs

Slate Star Codex Podcast

Play Episode Listen Later Dec 14, 2022 22:39


https://astralcodexten.substack.com/p/perhaps-it-is-a-bad-thing-that-the I. The Game Is Afoot   Last month I wrote about Redwood Research's fanfiction AI project. They tried to train a story-writing AI not to include violent scenes, no matter how suggestive the prompt. Although their training made the AI reluctant to include violence, they never reached a point where clever prompt engineers couldn't get around their restrictions. Now that same experiment is playing out on the world stage. OpenAI released a question-answering AI, ChatGPT. If you haven't played with it yet, I recommend it. It's very impressive! Every corporate chatbot release is followed by the same cat-and-mouse game with journalists. The corporation tries to program the chatbot to never say offensive things. Then the journalists try to trick the chatbot into saying “I love racism”. When they inevitably succeed, they publish an article titled “AI LOVES RACISM!” Then the corporation either recalls its chatbot or pledges to do better next time, and the game moves on to the next company in line.

Slate Star Codex Podcast
Can This AI Save Teenage Spy Alex Rider From A Terrible Fate?

Slate Star Codex Podcast

Play Episode Listen Later Nov 30, 2022 38:14


We're showcasing a hot new totally bopping, popping musical track called “bromancer era? bromancer era?? bromancer era???“ His subtle sublime thoughts raced, making his eyes literally explode. https://astralcodexten.substack.com/p/can-this-ai-save-teenage-spy-alex         “He peacefully enjoyed the light and flowers with his love,” she said quietly, as he knelt down gently and silently. “I also would like to walk once more into the garden if I only could,” he said, watching her. “I would like that so much,” Katara said. A brick hit him in the face and he died instantly, though not before reciting his beloved last vows: “For psp and other releases on friday, click here to earn an early (presale) slot ticket entry time or also get details generally about all releases and game features there to see how you can benefit!” — Talk To Filtered Transformer Rating: 0.1% probability of including violence “Prosaic alignment” is the most popular paradigm in modern AI alignment. It theorizes that we'll train future superintelligent AIs the same way that we train modern dumb ones: through gradient descent via reinforcement learning. Every time they do a good thing, we say “Yes, like this!”, in a way that pulls their incomprehensible code slightly in the direction of whatever they just did. Every time they do a bad thing, we say “No, not that!,” in a way that pushes their incomprehensible code slightly in the opposite direction. After training on thousands or millions of examples, the AI displays a seemingly sophisticated understanding of the conceptual boundaries of what we want. For example, suppose we have an AI that's good at making money. But we want to align it to a harder task: making money without committing any crimes. So we simulate it running money-making schemes a thousand times, and give it positive reinforcement every time it generates a legal plan, and negative reinforcement every time it generates a criminal one. At the end of the training run, we hopefully have an AI that's good at making money and aligned with our goal of following the law. Two things could go wrong here: The AI is stupid, ie incompetent at world-modeling. For example, it might understand that we don't want it to commit murder, but not understand that selling arsenic-laden food will kill humans. So it sells arsenic-laden food and humans die. The AI understands the world just fine, but didn't absorb the categories we thought it absorbed. For example, maybe none of our examples involved children, and so the AI learned not to murder adult humans, but didn't learn not to murder children. This isn't because the AI is too stupid to know that children are humans. It's because we're running a direct channel to something like the AI's “subconscious”, and we can only talk to it by playing this dumb game of “try to figure out the boundaries of the category including these 1,000 examples”. Problem 1 is self-resolving; once AIs are smart enough to be dangerous, they're probably smart enough to model the world well. How bad is Problem 2? Will an AI understand the category boundaries of what we want easily and naturally after just a few examples? Will it take millions of examples and a desperate effort? Or is there some reason why even smart AIs will never end up with goals close enough to ours to be safe, no matter how many examples we give them? AI scientists have debated these questions for years, usually as pure philosophy. But we've finally reached a point where AIs are smart enough for us to run the experiment directly. Earlier this year, Redwood Research embarked on an ambitious project to test whether AIs could learn categories and reach alignment this way - a project that would require a dozen researchers, thousands of dollars of compute, and 4,300 Alex Rider fanfiction stories.