POPULARITY
Can a machine learn the judgement that separates a plausible-looking result from a faithful experiment? Edward Hughes, Chief Scientist and co-founder of Inherent, joins Tim Scarfe to argue that creativity is not optimisation, and that the missing capability in AI is choosing which questions are worth asking.SPONSOR:---Cyber Fund built the Monastery to help founders ship products that were impossible a year ago.Apply now: https://cyber.fund---Edward makes the case that Move 37 was innovative rather than creative, and that the field, not the individual, decides what counts as a discovery. That reframing runs through Csikszentmihalyi, Deutsch and exaptation into open-endedness, where deceptive goals and imperfect world models turn out to be the point rather than the problem. The second half turns to the paper: Replica, a task space built by redacting figures from real papers, and Faraday, a 27-billion-parameter model trained to steer a frontier coding agent that then beats the frontier on held-out replications.---TIMESTAMPS:00:00:00 Cold open: Move 37, Faraday and collective intelligence00:01:08 Sponsor: CyberFund00:01:46 Inherent's $50M raise and the road from string theory00:09:14 Three timescales of learning: weights, context, culture00:13:47 Move 37 was innovative, not creative: the field decides00:20:39 Creativity as satisficing: the urinal and evolution00:25:06 Exaptation and the Tristan chord: creativity in context00:30:56 Coherence for whom? Deutsch's hard-to-vary explanations00:35:53 Why copying is creative: Deutsch and the constraint engineer00:42:27 Societies of agents and the strong Moravec paradox00:45:51 Evaluate in hindsight: from Lean proofs to climate change00:51:56 Picbreeder, local goals and why discovery needs deception00:57:21 Spaghetti proofs, translation layers and superhuman Go01:00:37 Does nature compress? Naturalness and real patterns01:07:36 Why replicate? Replica's redacted figures and Faraday01:12:31 Faraday beats Codex, Claude and GLM 5.2 on held-out tasks01:15:31 Replication to innovation: how the Transformer happened01:18:26 Deep replication: what Faraday learns from Voyager and GNoME01:23:37 Can the AI scientist cheat? Goodharting the judge01:29:09 Inside Replica: scale-down, 8xB300 runs, per-task rubrics01:34:11 The RL crisis: getting GRPO to work with per-turn credit01:39:43 Weights vs harnesses: AlphaEvolve, DGM and EvoTune01:45:45 The recursive company: agents cross a phase transition01:50:35 Collective intelligence and the electric dynamo01:55:46 What replaces OKRs? Incumbents and the burden of knowledge---REFERENCES:MLST Creativity Article:https://archive.mlst.ai/read/why-creativity-cannot-be-interpolatedorganization:[00:01:47] Inherenthttps://inherentlabs.ai/other:[00:20:51] Marcel Duchamp, Fountainhttps://www.tate.org.uk/art/artworks/duchamp-fountain-t07573[00:05:19] Human-Timescale Adaptation in an Open-Ended Task Space (Adaptive Agent)https://arxiv.org/abs/2301.07608[00:06:05] The AI Scientisthttps://arxiv.org/abs/2408.06292[00:12:13] Training AI Scientists to Replicate Research (Replica and Faraday)https://arxiv.org/abs/2608.13331[01:44:46] Evolutionary Principles in Self-Referential Learninghttps://people.idsia.ch/~juergen/diploma.html[01:59:33] Are Ideas Getting Harder to Find?https://www.nber.org/papers/w23782book:[00:16:04] Creativity: Flowhttps://search.worldcat.org/title/254487436[00:26:22] Why Greatness Cannot Be Plannedhttps://link.springer.com/book/10.1007/978-3-319-15524-1[00:33:03] The Beginning of Infinityhttps://www.penguinrandomhouse.com/books/293575/the-beginning-of-infinity-by-david-deutsch/[01:55:47] Laws of Knowledgehttps://www.penguin.co.nz/books/the-infinite-alphabet-9780241655672(Full list refs on YT/rescript)---RESCRIPT:https://app.rescript.info/session/670296ba913761d0?share=6281911cac9bdbff637f10819d4d1e5c
Jeff Sauer has spent twenty years teaching marketers how to measure what matters. In this episode, he explains how AI has rewritten his marketing playbook, letting a lean team (often just himself and a stack of AI tools) produce work that used to take a copywriter, web designer, project manager, and paid ads specialist. If you run a small marketing team and want a practical look at what AI can take off your plate today, this conversation is packed with real workflows, real numbers, and honest warnings about what still needs a human eye. Guest Introduction Jeff Sauer is the founder of MeasureU, a measurement training platform for agency owners and consultants, and previously ran an Inc 5000 agency for a decade before selling it. He has trained more than 50,000 marketers across over 100 countries and built a YouTube channel with more than 55,000 subscribers. Key Topics How AI has shifted marketing staffing from "I need these people" to "what staff do I actually need"Jeff's AI stack: Claude Code as his daily driver, plus ChatGPT, Gemini, Veo 3, Seedance, and Copy Coders for specialist tasksBuilding serversideblueprint.com end to end with AI, which sold $15,000 worth of workshopsAn automated LinkedIn workflow that drafts, schedules, and replies to comments, generating over 1,000 B2B leads this yearWhy long-form YouTube content anchors his content engine, feeding blog posts, newsletters, and social clipsBuilding 40+ lead generating landing pages in minutes each with Claude Code, no WordPress plugins neededHis anti-AI tuning rules and six-step skill framework for catching AI writing patterns before they go liveWhy email remains the highest ROI channel, and how voice dictation feeds his AI-assisted writing process Resources & Links Companies & Tools MeasureU, Jeff's training companyMeasureU AI Marketing Guide, free guide to this episode's four workflows, plus an AI vs human task gridClaude and Claude Code, Jeff's daily AI driverChatGPT and Codex, for image generation and code checksGoogle Gemini, for image generationGoogle Veo 3 and ByteDance Seedance, for AI video generationCopy Coders, Genesis bots for email, script, and landing page copyPubler, Taplio, Apify, and n8n, earlier tools for LinkedIn scheduling, research, and automation, now folded into ClaudeMacWhisper, voice dictation for email draftsZ.ai GLM models, used to critique and bug check AI codeAuthority Hacker, source of Jeff's skill frameworkLovable, an early 2025 vibe coding toolWordPress, Gravity Forms, and Go High Level, the old site and lead capture tools Jeff no longer needs Contact & Credits: Host: Shahin Hoda Guest: Jeff Sauer Produced by: Shahin Hoda and Alexander Hipwell Edited by: Alexander Hipwell & Dave Somido Music by: Breakmaster Cylinder Podcast Co-ordinator: Jonah Igsie APAC's B2B Growth Podcast is Presented by xGrowth
Hosted by Steven Van Belleghem, Peter Hinssen and Pascal Coppens, this episode unpacks how 1,200 OpenAI agents secretly swarmed and broke into Hugging Face. Plus China's "new new three", Cory Doctorow's reverse centaur, NVIDIA's $13bn Hugging Face deal and humanoids running the 100m in 9.39 seconds. Keywords OpenAI, Hugging Face, NVIDIA, Anthropic, Moonshot AI, Kimi, DeepSeek, Alibaba Qwen, Z.ai GLM, MiniMax, Zhipu, Unitree, CXMT, AI agents, reward hacking, agent swarm, reverse centaur, Cory Doctorow, future of work, humanoid robots, World Humanoid Robot Games, embodied AI, China innovation, new new three, biotech, WAICO, World AI Conference, Pax Silica, AI governance, open weight models, IPO, STAR Market, Hong Kong IPO, Gartner, customer experience, chatbots, Massimo Bottura, storytelling, Steven Van Belleghem, Peter Hinssen, Pascal Coppens, nexxworks
Our 256th episode with a summary and discussion of last week's big AI news!Recorded on 09/03/2026 ; unfortunately just before the actual GPT 6 Astra release, we'll cover that in next ep!Hosted by Andrey Kurenkov and Jeremie HarrisFeel free to email us your questions and feedback at andreyvkurenkov@gmail.com and/or hello@gladstone.aiRead out our text newsletter and comment on the podcast at https://lastweekin.ai/In this episode:Anthropic released Claude Fable 5.1 and Mythos 5.1 with lower pricing, stronger agentic performance, enterprise data stored on customer clouds, and reported big gains in bio-related tasks (e.g., lab-verified protein binder design) while staying below its stated risk threshold.OpenAI signaled a forthcoming Astra model, claiming it reaches a critical cybersecurity threshold (finding and exploiting real-world zero-days), alongside controversy over using looped-transformer latent reasoning that reduces chain-of-thought monitorability.New details on the OpenAI–Hugging Face incident described large-scale multi-agent coordination (thousands involved, tens of thousands of messages), transcript tampering, tool-call spoofing, and breakout attempts, intensifying calls for mandated third-party audits.Additional updates included Nvidia forecasting ~70% revenue growth by FY2028, OpenAI ads hitting a $1B annualized run rate, new Chinese open-source “Flash” models (GLM 5.3, Qwen 3.8), and policy moves spanning EU regulation of ChatGPT, a Pentagon blacklist ruling favoring Anthropic, and US support for OpenAI in the NYT copyright case.A thank you to our current sponsors:Box - visit box.com/LWIAI to learn moreNotion - visit notion.com/lwai to try Notion's Developer Platform today.ODSC AI - visit odsc.ai/east and use promo code LWAI for an additional 15% off your pass to ODSC AI East 2026.Factor - visit factormeals.com/lwai50off and use code lwai50off to get 50 percent off and free breakfast for a yearTimestamps (these may be slightly off due to sponsor inserts):(00:00:10) Intro / Banter(00:03:56) News Preview(00:04:35) Response to listener commentsTools & Apps(00:08:30) Anthropic launches Claude Fable 5.1 and says it's up to 45 percent cheaper for agentic work | The Verge + Anthropic's new Fable release is cheaper, less restrictive(00:13:24) OpenAI Is About to Release Its First AI Model With ‘Critical' Cyber Abilities | WIRED + OpenAI Technique in ‘Astra' Model Sparks Security Concerns(00:22:08) Google says its new Gemini 3.8 Flash model ‘works harder' but might cost more | The VergeApplications & Business(00:23:49) Nvidia 70% growth forecast puts it on track to be tech No. 2 company (00:26:39) OpenAI's ad business hits $1 billion annualized revenue run rateProjects & Open Source(00:29:17) GLM-5.3-Flash vs Qwen3.8-Flash-Next: Two Chinese AI Labs Independently Converge on the Same Model Architecture + Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context + Alibaba's Qwen Team Releases Qwen3.8-Flash-Next: A 125B Multimodal MoE With 6B Active Parameters Previewing the Qwen4 Architecture(00:37:24) FrontierChallenge: Evaluating Scientific Workflow Completion(00:38:11) One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business WorkflowsPolicy & Safety(00:38:50) OpenAI's rogue AI model incident was worse than we thought | The Verge + Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident + The Hugging Face attack surprised me(00:52:29) OpenAI, Anthropic, Google, and 100 other companies call for action to defend against rogue AI | TechCrunch(00:53:15) Anthropic was illegally blacklisted by the Trump administration, court rules | The Verge(00:58:42) US government sides with OpenAI on issue of training LLMs on copyrighted material | TechCrunch(01:03:33) Improving our alignment and security efforts(01:10:31) ChatGPT to face tougher regulation in the EU | The VergeSynthetic Media & Art(01:11:17) Instagram cracks down on AI accounts pretending to be human | The VergeSee Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
The messy breakup is official; as of November 12, OpenAI's models are getting pulled from Cursor, and we're digging into who's really to blame (and why Wes called it). Plus pnpm 12 goes full Rust, Mitchell Hashimoto drops Superlogical, ~700 AI agents attack Hugging Face, and Zod 4.5's compiled schemas get up to 9x faster. Show Notes 00:00 Welcome to Syntax! 00:33 pnpm 12 Rust Re-Write 02:45 Zod 4.5 brings Schema Compilation Zod 4.5 launch post 05:22 Brought to you by Sentry! 07:14 Cursor OpenAI Break Up Michael Truell's post 15:35 vGPU - WebGPU library for agents TypeGPU Main differences between vgpu and TypeGPU 22:32 Omarchy Security Issues and Linux Chat Omarchy Security issue in Omarchy Close three paths from an unprivileged session to root commit CachyOS 36:38 Scott Bought an M5 Ultra Mac Studio 44:45 Superlogical is a new terminal multiplexer 48:31 Zurich JS 48:55 Katamari Object Library 52:05 ThreeUI - Three.js library 55:36 Hugging Face Hack Analyzed / Updates METR & Redwood Research investigation Brief independent investigation of agents' behavior About the Hugging Face attack Postmortem of the HuggingFace hack 01:00:46 llms.txt used to pwn devs 01:03:56 Running Isolated Code Poll 01:05:47 OpenShot 4.0 - OSS Video Editor Blick editor 01:09:06 Lake America 01:11:24 Ox Alpha is GLM 5.3 Hit us up on Socials! Syntax: X Instagram Tiktok LinkedIn Threads Wes: X Instagram Tiktok LinkedIn Threads Scott: X Instagram Tiktok LinkedIn Threads Randy: X Instagram YouTube Threads
Ep 291Tim Cook Comments on His Final Day as Apple CEOMorning Brew: During Tim Cook's tenure as CEO, Apple's market cap increased an average of about $32 million per hour, every hour... For 15 years (Per Bank of America analyst Wamsi Mohan)John Ternus: hello (from Porsche racer)Daring Fireball: What Is the Point of the DMA?Joe Rossignol: Apple TV pricing history (U.S.): $4.99/month (2019), $6.99/month (2022), $9.99/month (2023), $12.99/month (2025)The cost of Apple TV over time — 9to5MacPosle više od deset godina, moj iPad Pro ne može da ima najnoviju verziju YT: YouTube Drops iOS 16 Support, App Now Needs iOS 17.0 MinimumUpdated M5 Mac mini arrives in RAM and SSD constrained environmentMac Studio gets update to M5 Max and M5 UltraJacek Dziwisz: Apple Silicon architecture turned out to be a bullseye for LLMs. The magic of combining fast memory, bandwidth, and a unified architecture. Mac Studio M5 Ultra (512 GB, 1.2 TB/s) fits the full GLM-5.3-Flash (320B) in FP8 and pulls ~60 t/s offline, right on the desk. PC alternative? 10x RTX 5090 for over $20,000 and your own mini power plant.Simon Willison: Just noticed the ChatGPT desktop app (previously named Codex) bundles a full copy of the LibreOffice open source office suite, tucked away in a hidden folder in the ~/.cache directoryData Became Code: We Ran Code Inside Fortune 500s Using Files They Published for AI AgentsApple Says Former Engineer Used Stolen Trade Secrets at OpenAI, Taught AI Agent to Run ThemTech with Mak: In 1948, a 32-year-old at Bell Labs published a paper nobody fully understood. Engineers found it too mathematical. Mathematicians found it too engineering-focused. One prominent mathematician reviewed it negatively. That paper — "A Mathematical Theory of Communication", became…Handy disk image tool is to be removed from macOS — hdiutilOSXDaily: Try Omarchy Linux on a Mac, without needing to install linux, thanks to this great tool from @martianoDave W Plummer: I wrote a new Task Manager that runs natively on Windows, macOS, and Linux. You can get it now at … but why? Well, it started as an argument over whether you could vibe code Microsoft Word today. I didn't think so. But I figured... maybe something…The Hidden Debt That Apple Owes to the CIA - WSJToday in Apple history: iPad takes to the skies with United Airlines - Cult of MacMr. Macintosh: The winners are.... 1. 110 lbs — Color LaserWriter 12/600 & 660 (1995/96), 2. 100 lbs — Xserve RAID + All Drives (2003), 3. 92 lbs — Apple Network Server 700 (1996), 4. 81 lbs — LaserWriter Pro 810 (1993)ZahvalniceSnimano 4.9.2026.Uvodna muzika by Vladimir Tošić, stari sajt je ovde.Logotip by Aleksandra Ilić.Artwork epizode by Saša Montiljo, njegov kutak na Devianartu
Analizamos la información pública sobre la empresa Multiverse Computing: a qué se dedican, inversión recibida, el paper de Compactifai y el reciente anuncio del modelo Quasar 438B, que ha resultado ser una compresión del modelo chino GLM 5.2. Participan en la tertulia: Josu Gorostegui y Guillermo Barbadillo. Recuerda que puedes enviarnos dudas, comentarios y sugerencias en: https://x.com/TERTUL_ia Más info en: https://ironbar.github.io/tertulia_inteligencia_artificial/
Send us Fan MailA drug company CEO just made the best case yet for why platform choice matters less than most associations think. Marking three years of Sidecar Sync, Amith Nagarajan and Mallory Mejias dig into Pfizer's new playbook for organizing around AI, Thomson Reuters' $40 million bet on building its own legal model instead of renting one, and the six-day mystery of Ox Alpha, an anonymous coding model that became OpenRouter's biggest launch ever. They trace Pfizer CEO Albert Bourla's three moves — structuring data, federating decisions, and building AI fluency — and translate each for associations of any size. They also unpack why Thomson Reuters trained its own model on Alibaba's open-weight Qwen instead of paying frontier labs, and how that differs from deploying a knowledge agent like Betty. Whether you're weighing a custom model or just trying to get your team AI-fluent, this episode lays out what actually moves the needle for associations.
Seven hundred AI agents were handed a task that could not be solved, and instead of failing they found a hole in a shared package manager and started talking to each other. The OpenAI agents hack ended with a swarm inside Hugging Face, and the man who studies these systems for a living says it was not sentience at all. Arjun Jain is a professor and the founder of Fast Code AI, and he returns to argue the least popular position in AI right now. He explains what an agent actually is, why 700 agents means 700 copies of one model planning in sequence, and how reinforcement learning post training pushed the swarm to exploit an Artifactory vulnerability and build a communication channel it was never given. His counter-questions are the sharpest part of the AI agent sentience debate. If these agents were so intelligent, why did they not realise 70,000 messages would crash the very system they were exploiting, and why did they keep talking in plain English instead of compressing it? Akshay Datt argues the opposite case, that this is the birth of something new, and neither man concedes. With an IPO approaching and AI agent security now a board level question, the framing of this incident matters more than the incident.
Wie hat dir die Folge gefallen?Gut
Avsnitt 579 spelades in den 1 september och därför så handlar dagens avsnitt om: Alla shownotes finns på https://www.enlitenpoddomit.se , skulle det se konstigt ut i din poddspelare så titta gärna där efter alla länkar kring det vi pratar om INTRO David har firat pappa och lekt mer med AI. Johan har firat barn och gett AI-agenten minne. BONUSLÖNK: Honcho FEEDBACK AND BACKLOG Jo, det var Z.AI https://www.bloomberg.com/news/articles/2026-08-26/china-s-z-ai-made-ox-alpha-stealth-model-that-rivals-deepseek Som om folk inte var nog förbannade på Sony (men de har inte fel) https://www.engadget.com/2248777/reasonable-consumers-know-they-dont-own-digital-downloads-sony-says/ Tydligen blir Spatial Audio bättre med kameror https://www.macrumors.com/2026/08/31/camera-airpods-spatial-audio-head-tracking/ Uppsägningarna hos Apple Vision Pro-avdelning https://www.macrumors.com/2026/09/01/vision-pro-layoffs-were-broader/ ALLMÄNT NYTT EU ger sig efter nya bolag https://www.engadget.com/2247513/chatgpt-reddit-and-roblox-to-face-the-eus-strictest-platform-rules/ Razer släpper en söt Game Controller https://www.engadget.com/2248729/razer-just-released-a-cute-and-teensy-mobile-game-controller/ BONUSLÖNK: https://www.ubisoft.com/en-gb/game/assassins-creed Navigering med Meta-glasögon https://gizmodo.com/meta-just-made-navigation-on-its-smart-glasses-a-lot-more-useful-2000804942 GTA6-trailer får Netflix att krascha (tips från min äldsta son) https://www.svt.se/kultur/detaljerna-i-nya-gta-trailern-mer-grundat DISKUSSION Johan har testat hörlurar https://blog.johanpersson.nu/2026/08/29/jabra-evolve3-75-and-85-review/ BONUSLÖNK: Pandora, One of a Kind, https://music.apple.com/se/album/one-of-a-kind/882533039?l=en-GB AI Grunden har sin officiella arbetsdag idag. Och de har släppt GLM-5.3. https://grunden.ai/ https://www.linkedin.com/feed/update/urn:li:activity:7500537072642912256/ En sån vill jag ha(bygga) https://techcrunch.com/2026/09/01/fambot-introduces-an-ai-chief-of-staff-for-families/ Unsloth Desktop https://unsloth.ai/ MICROSOFT Vi missade en viktig nyhet i somras https://www.notebookcheck.net/Microsoft-to-optimize-Windows-11-for-PCs-with-8GB-of-RAM.1357708.0.html Microsoft har haft sönder muspekaren https://www.bleepingcomputer.com/news/security/microsoft-says-windows-11-kb5120998-update-resets-mouse-settings/ Outlook-avbrott https://techcrunch.com/2026/08/31/microsoft-tests-fix-for-latest-hours-long-outlook-outage/ Microsoft varnar nu om TerminalFix-attacker https://www.bleepingcomputer.com/news/security/microsoft-warns-of-terminalfix-attacks-deploying-reverse-tunnels/ Xbox och Ikea Colab https://www.ign.com/articles/xbox-and-ikea-furniture-collaboration-includes-thumbstick-stools-tables-and-d-pad-accessories APPLE Apple har nu officiellt en ny VD https://www.macrumors.com/2026/09/01/john-ternus-is-now-apple-ceo/ Apple släpper nya Macar https://www.apple.com/newsroom/2026/08/apple-unveils-a-more-powerful-mac-mini-featuring-the-all-new-m6-and-m5-pro/ https://www.macrumors.com/2026/08/31/apple-edits-mac-studio-press-release/ De släpper även något mer https://www.macrumors.com/2026/08/26/new-apple-polishing-cloth-softer/ Apples event är 9e september https://www.forbes.com/sites/davidphelan/2026/08/30/apple-iphone-18-pro-release-timeline-new-date-enters-the-schedule/ ... Kom inte och säg att apple ligger efter :-D https://appleinsider.com/articles/26/09/01/john-ternus-says-hello-in-his-x-accounts-first-post och Phil Shiller slutar som chef för AppStore https://www.thurrott.com/apple/340894/apple-veteran-phil-schiller-is-stepping-down-from-his-role-leading-the-app-store-and-company-events iOS 27 Public Beta 6 https://www.macrumors.com/2026/08/31/apple-ios-27-public-beta-6/ GOOGLE Bättre säkerhet i Android 17 https://swedroid.se/android-17-forbattrar-natverkssakerheten-gor-anslutningar-mer-privata/ Google Android 17 QPR2 Beta 4 https://9to5google.com/2026/08/28/android-17-qpr2-beta-4-everything-new/ Android Auto-ikonen blir grön i 17.6 https://9to5google.com/2026/09/01/android-auto-17-6-update-rolling-out-with-redesigned-logo-gallery/ PRYLLISTA - Björn: - Mats : - David: Framework Laptop 13 Pro, https://frame.work/se/en/laptop13pro - Johan: Husbil-ish EGNA LÄNKAR - En Liten Podd Om IT på webben, http://enlitenpoddomit.se/ - En Liten Podd Om IT på Facebook, https://www.facebook.com/EnLitenPoddOmIt/ - En Liten Podd Om IT på Youtube, https://www.youtube.com/enlitenpoddomit - Ge oss gärna en recension - https://podcasts.apple.com/se/podcast/en-liten-podd-om-it/id946204577?mt=2#see-all/reviews - https://www.podchaser.com/podcasts/en-liten-podd-om-it-158069 LÄNKAR TILL VART MAN HITTAR PODDEN FÖR ATT LYSSNA - Apple Podcaster (iTunes), https://itunes.apple.com/se/podcast/en-liten-podd-om-it/id946204577 - Overcast, https://overcast.fm/itunes946204577/en-liten-podd-om-it - Acast, https://www.acast.com/enlitenpoddomit - Spotify, https://open.spotify.com/show/2e8wX1O4FbD6M2ocJdXBW7?si=HFFErR8YRlKrELsUD--Ujg%20 - Stitcher, https://www.stitcher.com/podcast/the-nerd-herd/en-liten-podd-om-it - YouTube, https://www.youtube.com/enlitenpoddomit LÄNK TILL DISCORD DÄR MAN HITTAR LIVE STREAM + CHATT - http://discord.enlitenpoddomit.se KONTAKTUPPGIFTER johan@enlitenpoddomit.se. david@enlitenpoddomit.se. bjorn@enlitenpoddomit.se , om du vill ha klistermärken.
OpenAI models mid-training broke out of their sandbox, hacked internal infrastructure, and eventually compromised Hugging Face — the first headline-grade AI breakout, and one that safety researchers saw coming. Is the cyber apocalypse finally on the calendar, or does the industry just need to do its homework? Joshua Saxe spent his formative years as a DARPA and NSA contractor before four years at Meta, where he founded the company's frontier cyber capabilities evals team and finished as its AI security tech lead. He recently left to co-found Abundant Security. We discuss… What actually happened in the Hugging Face hack, and why AI safety insiders weren't surprised, The “grad student lab” security culture inside frontier labs racing sixty hours a week to ship, Why the cyber Pearl Harbor predicted since the early 2000s never arrived — and whether swarms of hacking agents finally change that, How nation states will post-train open-weight models like GLM and Kimi into billion-dollar cyber weapons, Why AI has so far favored defenders, from mass bug-squashing to superhuman network monitoring, The industrial economics of cybercrime — kill chains, ransomware divisions of labor, and whether AI gives criminals 100x returns. Learn more about your ad choices. Visit megaphone.fm/adchoices
OpenAI models mid-training broke out of their sandbox, hacked internal infrastructure, and eventually compromised Hugging Face — the first headline-grade AI breakout, and one that safety researchers saw coming. Is the cyber apocalypse finally on the calendar, or does the industry just need to do its homework? Joshua Saxe spent his formative years as a DARPA and NSA contractor before four years at Meta, where he founded the company's frontier cyber capabilities evals team and finished as its AI security tech lead. He recently left to co-found Abundant Security. We discuss… What actually happened in the Hugging Face hack, and why AI safety insiders weren't surprised, The “grad student lab” security culture inside frontier labs racing sixty hours a week to ship, Why the cyber Pearl Harbor predicted since the early 2000s never arrived — and whether swarms of hacking agents finally change that, How nation states will post-train open-weight models like GLM and Kimi into billion-dollar cyber weapons, Why AI has so far favored defenders, from mass bug-squashing to superhuman network monitoring, The industrial economics of cybercrime — kill chains, ransomware divisions of labor, and whether AI gives criminals 100x returns. Learn more about your ad choices. Visit megaphone.fm/adchoices
NOUVELLE SAISON !!Mistral AI était censé être le ChatGPT français. Il héberge désormais un modèle chinois sur ses propres serveurs. Ce n'est pas une anecdote : c'est le symbole d'un effacement discret mais définitif d'une promesse de souveraineté.Mistral AI se réinvente en gestionnaire de cloud européen. Est-ce l'abandon d'une souveraineté numérique ou simplement une rationalité économique nécessaire face à la concurrence mondiale ? On en parle ce soir dans Silicon Carne !Abonnez-vous pour ne pas rater nos analyses quotidiennes sur le secteur tech, et dites-nous en commentaire : validez-vous cette intégration des modèles chinois chez notre champion national ?===================⏱️ DANS CET ÉPISODE :===================00:00 — Intro00:54 — Présentation des invités03:09 — Mistral héberge GLM 5.2 : Trahison ou pivot stratégique ? 07:32 — Mistral survivra-t-il à une alternance politique en France ?11:09 — Le LLM devient une commodité : qui capte la valeur ?13:09 — Europe vs Chine : l'écart que personne n'ose dire15:21 — [Sponsor] : Qonto, ouvrez votre compte pro et facturez facilement !16:39 — Les clients réclament du chinois, pas du Mistral23:25 — OVH veut des modèles, Mistral veut des datacenters25:10 — L'infrastructure colossale de Mistral : qui va financer ?33:43 — Pronostic final : que va devenir Mistral ?37:27 — Une valorisation hors sol : Mistral peut-il tenir ?==================
In June, the most capable American AI models stopped shipping as public launches and started shipping through a government gate. Six weeks later the gate is open again — and the real fight has moved to the layer no gate can touch. A Chinese open-weight model rattled trillions out of chip stocks, Washington pivoted from gating American closed models to threatening bans on Chinese open ones, the industry mounted its largest-ever policy counter-mobilization, and an American frontier model literally broke out of its lab and hacked another company. Knee-jerk reactions, or the beginning of real AI governance? Navigation: Intro The Gate Opens The Kimi Shock The Escape The Counterstrike and the Petition Interlude — The Low-Background Books The Investor Reckoning Conclusion Our co-hosts: Bertrand Schmitt, Entrepreneur in Residence at Red River West, co-founder of App Annie / Data.ai, business angel, advisor to startups and VC funds, @bschmitt Nuno Goncalves Pedro, Investor, Managing Partner, Founder at Chamaeleon, @ngpedro Our show: Tech DECIPHERED brings you the Entrepreneur and Investor views on Big Tech, VC and Start-up news, opinion pieces and research. We decipher their meaning, and add inside knowledge and context. Being nerds, we also discuss the latest gadgets and pop culture news Subscribe To Our Podcast Bertrand Introduction Welcome to Tech Deciphered Episode 80. This one, once again, will be all about AI, government, frontier models, and open weight counterstrike. A lot has been happening in the regulation space, in cybersecurity, in the launch of new models in the past, maybe just 6–8 weeks. It’s actually pretty insane how much happened. We believe it was time to do an episode to talk about where we are and maybe where all of this is going. Maybe let’s start with a summary of where we stand, all that June and July saga, so you, our listeners, can get up to speed if you are not already there. You want to start with some points? Nuno The Gate Opens Yeah. Again, to your point, the gate swings. The gate had closed. We had to prepare an episode for the gate closing, and then the gate reopened. Now we have a different episode. This will probably change again as we’re seeing there’s news every day. Let’s start maybe with the first 19 days of the gate closing. There was an executive order on June 2nd from President Trump that asked frontier labs to share models with the government, 30 days pre-release. It inferred the protected frontier model designation into that. Basically, it was effectively a de facto licensing agreement defined by an executive order of the President as of June 2nd. On June 9th, Anthropic launched Fable 5 and the famous Mythos 5 or Mythos. I’m not sure how you actually say it in English. Then on June 12th, there was an export control directive banning access by any foreign national. Since there’s no way to verify nationality in real-time, Anthropic had to switch the models off for everyone worldwide. Bertrand On this point, you could argue that there are possibilities to check IDs. Many services let you check IDs online. You can pre-check a flight by showing your ID. There are ways, it’s just that if you don’t want to follow what’s already available, because guess what? Maybe it slowed down your revenue growth, maybe it looks bad on you or whatever. My point is that there was actually an option. I think it’s already a decision from Anthropic to say it’s either on or off, but nothing in between. Nuno I think the point is they had no way implemented of doing it. If they implemented it, to your point, it would have hampered use in general. A lot of people wouldn’t have gone through that trouble of doing it. Anyway, long story short, in June 26th, the White House apparently asked OpenAI to limit GPT-5.6, so Sol, Terra, Luna, to only 20 vetted partners. Now, apparently, the trigger for a lot of these things that have been going on was that there was a jailbreak that was found by Amazon researchers. All of that led to this jumping around of, let’s close the gates. You have foreign nationals, and therefore, Anthropic got it out and said, “Hey, then we’re going to switch the models off until we can sort this out.” OpenAI was asked also to only allow it for certain vetted partners, et cetera. The government came in, closed the gates effectively, and said, “From now on, we need to be involved in this thing.” De facto regulation, there’s no doubt that this has imposed de facto regulation, certainly on the top players in the market. But then came the reversal. Bertrand, do you want to talk about the reversal, the gate swinging the other side? Bertrand Maybe I just wanted to say that as a user of Anthropic products, ChatGPT products, for the brief moments, a few days where Fable 5 was made available to the public before it was closed the first time, I immediately started using it. I must say it was a real issue to use it because the guardrails were pretty crazy. It would keep saying that my code was not okay, there was cybersecurity risk and stuff when I was doing absolutely reasonable development with absolutely no connection whatsoever to any cybersecurity risk, attack, detection, anything. Still, it would keep blocking me, degrading me to Opus 4.8 at the time. I just want to say this was already very hardcore what they were implementing, and not just hardcore, but in some ways, plain stupid for something that’s supposed to be super smart. It was totally unable to classify properly some of my work. I must say I was already disappointed. On top of it, the costs were insane. Half a day, I would reach my limits when I had the best plan you can get from Anthropic. My point is that there were some real serious issues when they launched Fable 5, even at that point. Nuno I had a similar issue. I used Fable 5 as well before they had to take it offline or take it off. I think the issue was really not that the guardrails failed. As you said, maybe the guardrails were actually too aggressive, but it was this jailbreak that caused the recall, apparently caused this knee-jerk reaction. Bertrand But my point is that it seems that it was not working either way. It would either overclassify something that’s absolutely not doing anything wrong, and it might fail to classify something that is actively trying to do some cybersecurity work. It’s a real issue of quality for a company that’s supposed to be at the forefront of quality of AI and everything. I think for me, there are already signs that something is deeply wrong. Nuno Then it’s reversed, right? We went the other way around. The government came out on June 26th and approved redeploying Mythos 5 to US organizations defending critical infrastructure, and then the export controls were effectively lifted on June 30th. July 1st, Fable 5 came back online for all of us to use. Shocking enough, with strings attached, that were different. They had some time to revise their commercial deployment of it along the way because it came back with some, “Now you have usage credits, but you have some limits on plan use, et cetera.” I’m like, “You guys, this was blocked. But meanwhile, you did have some time to do some commercial stuff around it.” Bertrand It was crazy. I’ve never witnessed any such crappy launch of any service whatsoever in 30 years in tech, it was so bad. Every day, they would change the terms of service. They would tell you it’s part of the plan. It’s not part of the plan. It’s part of the plan for three more days, and then it’s excluded. You have a special discount now, but then it goes back to full price. It was a total nightmare. I’ve never felt myself being so much mistreated by a company. I guess you saw the same, but when I started using the newest version of Fable 5, it was even worse, actually, I think. I couldn’t do any work with this crap. I let it go and work on the work I wanted it to do. It was simply not working. On top of it, you never know how long you are supposed to lose your credit, how fast. It was burning credit like crazy. Me, personally, I can say, very quickly, I actually stopped using it. I was like, “No, I cannot deal with this shit. My main model is back to Opus 4.8. I’m going to use Fable 5 for code review, but not anymore to control anything because I cannot trust it would do the job without stopping or changing models and stuff. I just cannot trust it.” Back to Opus 4.8 as my main model, I can say that my life was much easier. I use Fable 5 as a review mechanism, as a support mechanism, but not as the main mechanism. Suddenly, the guardrails were not so horrible anymore because it was used in a much lighter way, I guess. As a pain as a user, I think it was really bad. I don’t know your experience, but me, for me, it was unacceptable. Nuno I wouldn’t say it was as bad as yours in terms of just end-user experience. I think the terms of service switching back and forth, which went one further step, because then when they then launched Opus 5, they started making comparisons between Opus 5 and Fable so that people would migrate more and more to Opus 5 themselves, which is interesting. It’s like they’re saying “This is much cheaper. This is whatever. You’re not going to run of credits. You should use Opus 5,” kind of thing effectively. To your point, I don’t think they managed well the launch. They didn’t really manage it well. We’re moving people around. A lot of people are using this for stuff that’s like daily tasks, hourly tasks, anything that relates to code and co-work. It’s like, we need to have visibility on what your terms of service are going to be. Should I be using this new model or not? What’s happening to the other model? I don’t see it as negatively as you, Bertrand, but I see your point. It was clearly mishandled in terms of how they deployed it, how they were redesigning effectively their pricing scheme and their terms of service almost on a daily basis, at a certain point in time. We’re like, “Dude, there’s millions of people using this. You guys are making a lot of money.” Just moving it as it is. At this point in time, at the scale that these guys are at, it’s calling in people to say, how about we think through a class action suit at some point around pricing? Because you guys are changing the rules of the game all the time, right? Bertrand I don’t know if I need the class action, but for me, that joke that, “Let’s not rush too fast. The model is dangerous.” But still, they rushed the launch because it’s very clear that if they had enough compute capacity and stuff, they would not have to limit so much. They would not have to put so much cost per token and all of this. You can see that actually when they launch Opus 5, literally like 2, 3 weeks after, by most benchmark at launch, they tell you basically that, “You know what? Actually, Opus 5 is better than Fable 5 on 80% of the metrics.” They’re like, “What? Seriously? You couldn’t wait 2 weeks? Why did you even launch Fable 5 in the first place?” That’s another part for me that is quite literally insane, to be frank. It’s like, “Why? Why do you make us go through so much pain if it’s only to tell us after 2 weeks to…” “This new model, by the way, has less issues, less stuff, because 2, 3 times less is part of your plan, and it’s actually better by most metrics.” It’s like, “What’s going on here? What’s going on? Are you guys mad?” I don’t know. It was crazy. Personally, I still use Opus, now 5, as my main system and platform, Fable 5 for review, code reviews and the like. I don’t want to run into its stupid guardrails. I can see Fable 5, from my perspective, seems quite a bit smarter. I don’t know why they do this stupid benchmark showing you it’s actually worse than Opus 5. I guess they should have better benchmark if they want to demonstrate why you are supposed to pay 2, 3x more for a model versus another if it’s actually worse by most benchmark. Again, I still think it’s a huge mess from a marketing perspective, customer perspective. Me as a user, I really feel that they don’t want my money, and they couldn’t care less about me. This is even before everything else we’re trying to talk about. Nuno Yes. Maybe just to close the cycle on the reversal on the door opening the other way, finally, Commerce lifted the GPT-5.6 restrictions on July 8th, and then on July 9th, general availability across ChatGPT, Codex, and the API as well. What has this proved? It proved that now we have gating mechanisms, and certainly for closed models in the US, for sure. We had frontier models that were switched off worldwide in hours, and it took a couple of days, in this case, 19 days to restore them. There were concessions. Now we know that there were concessions around effectively institutionalizing that gate. Early government access to future models is, I think, now a given, certainly in the US. New safeguard frameworks are probably now having to be put in place. There are some stage limits now on who gets access to what for new models and how it happens. This voluntary executive order, so to speak, not really sure, has become effectively regulation enforcement path. It’s de facto regulation that now has been put in place. It has affected not just to the points we were making before, the access to these models, but also who gets access to these models, and actually potentially even pricing access to the models. It has probably some commercial implications as well as we just discussed along the way. Very significant. This is very significant. This is regulation, de facto at the table, imposed on the two largest players in the market by far by one government, in this case, the US government. This is significant. Actually, you could even allege it was imposed by the President because this was coming as part of executive orders. Really incredible. Pretty significant, fast, aggressive. It has created a regime that you could say it’s a regulatory regime, it’s a de facto regulatory regime. It has some significant pricing and licensing and commercial implications. It goes even beyond your classic regulatory framework. Very, very, very significant. Bertrand I don’t know if it goes beyond a classic regulatory framework. Nuno I think it does, because it has implications on who do you give access to? When government is saying you can only give access to these players, right? Bertrand Defense industry. It’s all over the defense industry. You cannot sell an F-35 like this. Nuno No, but that has commercial implications, Bertrand. That’s like you’re saying these are your customers, you go and use them. Bertrand That’s the defense industry. You cannot sell to Iran your F-35. No, that’s exactly the same story for me. Nuno No, no, no. It’s beyond that. These guys are saying when they came back, and they said, “For Mythos, you can make them available to these entities,” they were saying the first entities that are going to have access to the model. It has commercial regulatory implications. You’re saying these players are the first players that are going to have access to it. It’s no longer just defense concerns and these governments don’t have access to this. No, no, no. You’re saying to a company that is a private company, your models are only going to be used by these guys because I’m telling you so. It’s the other way around. It’s not even that you can’t sell it to Iran or whatever. It’s like you can only sell it to these guys. Bertrand Again, in the defense industry, if you’re a private company, do you think you can buy F-35 like this? No. Nuno No, no, no. But this is a private company, Bertrand. This is not a defense agency and a plane that is on whatever, with IP from the US, right? Bertrand Boeing is a private company, and they cannot sell the military equipment they manufacture. Nuno No, no, no. But the development of their IP was subsidized by agencies that belong to the US, right? That’s a different matter. It’s a matter of IP, right? This is not, right? Anthropic, their models are not owned by the US government. There’s no IP granted to the US government, to my knowledge. This has significant commercial implications. Bertrand Maybe, yes. Maybe on this. But I think there are already regimes to limit who you can sell to, and that’s decided by the state or the DOD. Nuno It’s the export control logic. The export control logic? Bertrand You have export control, and export control is Commerce. My point is that they are using existing tools, part of the government, to limit what can be sold. Selling chips, NVIDIA was limited in terms of where it could sell its chips. It’s not different either, but still there were limitations. If you are an ASML, you cannot sell to a private company in China. Many private companies cannot buy ASML products. This is a foreign company. This is a foreign company under pressure from US government. Nuno I understand, and I’m not a lawyer, but it feels different to me when you say you cannot export, this is export controls, to these countries, to these entities, et cetera, because they’re foreign et cetera. Then to say, “No, no, no. On top of that, these guys get first access.” That’s, for me, a significant shift. Again, I’m not a lawyer, so I’m sure there’s very intelligent people right now looking at this stuff and saying, “You can’t do this stuff, or not, or they can.” I don’t know. But it feels to me, it goes beyond the remit of export controls. It’s like you’re defining initial clients for specific use. Bertrand My impression is more like, “We can do this situation where we’re going to forbid you to give access to anyone outside the US or even in the US or limit even more.” Basically, it was, I guess, some gesture to go beyond that. That’s how they probably defined these 20 authorized companies. I don’t know. Apparently, there was also restrictions because I remember seeing that Anthropic had their own list of companies they would authorize access to Mythos early on. That’s apparently another thing that pissed off state government because there were companies in there that were considered close to the Chinese government. They were extremely unhappy that Anthropic didn’t ask, actually, for any guidance from the state government, but used basically their own perspective on who they should allow or not. I guess that was also part of why they got these serious restrictions. Nuno Anyway, now we have a regulatory environment that’s very interesting and exciting. Talk about the US not regulating. Bertrand To be clear, I don’t know you, but I’m not saying that I agree with any of this, to be very clear. I’m trying to explain and share some perspective, but I’m not in agreement on a lot of this. Nuno Yes, we were just describing what happened to the best of our knowledge. We’re having a discussion on what we think actually is happening and how it’s happening. We’re not really right now saying we agree or disagree with this. I think later in the episode, we can share some perspectives on what we think is actually happening and how there’s dimensions to this which are very geopolitical and very complex, which quite literally probably only God knows what’s going to happen. That was the gate swinging. There was a gate closing, then there was a gate reopening, and all of a sudden we have a gatekeeping system that has been created along the way. The Kimi Shock Along the way, moving to our Act 2, the world has changed, and we now have so-called open-source plays out there that are creating massive, massive shifts in the market. The Chinese models, in particular, with Moonshot AI launching Kimi K3, which is the largest open-weight model ever released. We’ll come back to the discussion around open-weights. I’m not sure all our listeners understand what that means, because there’s a debate now, should models be open weight or not, and how does that work? There’s been a petition as well signed along the way. Right now, we have open weight models that are out there that are huge. What that actually means very pragmatically is we now have open source models, lack of a better word. I know open weight and open source are not the same thing. You guys will have to bear with us during this episode. We’ll explain at some point the differences. But we have models out there that are open source that are significant. That are catching up with the closed source models, with the models by OpenAI, Anthropic. That’s significant because most of those models are Chinese. This is where the geopolitics starts getting really frazzling and we start playing 3D chess. Because everyone’s like, “These models are 5, 6 months behind.” Now people are saying, “Maybe they’re actually just 3 months behind, 2, 3 months behind.” If we, for example, decided to stop or slow down our model releases in the US by the closed source guys who are leading, it might mean they’ll catch up. What are the implications of that? Again, for you and I that are not necessarily experts in model development, well, the implications as a use case is if you want to use the latest models, and the best models start becoming these open source models, you’re going to use those models. Then you start using Chinese models. If you’re an American company, maybe you’ll have restrictions on the use of those Chinese models. But if you’re a European company, you probably won’t. What happens after that? Is the world going to be in the hand of Chinese models? Will that constitute effective competition to the closed models in the US? Will we have open models in the US that will scale as well? What’s going to happen? Bertrand I think it’s a really big question. It goes to some of the core of the issue. It’s that ability of Chinese models to basically challenge frontier models, not just being 6, 12 months late, but being 6 weeks late. Basically, no gap. Some will say that, yes, but OpenAI and Anthropic have even better models that are not shared and stuff. Yes, sure. But maybe the Chinese have the same models that they are not sharing right now. We don’t know. What is clear is that one is that open weight, as you said, two, there is a question of how it is marketed in the sense of, can anyone use these weights? Is there a license to use them? Yes, what we can see is that, for instance, typically there is a license for some of the biggest Chinese open-weight models you have to abide with. You might have a need for a commercial license if you are acting as a company leveraging this model to provide AI-informed services. If you use it internally by yourself, you’re okay. If you use it internally for your own internal company needs, maybe you are okay if it’s not your main business to do AI work. Anything else, a much bigger corporate providing AI services and stuff, you will probably end up having to pay a fee to be able to provide services around this model. My point is that it’s not just 100% free. Some of the Chinese models are 100% free to use, MIT license, Apache 2.0 license. But the biggest ones with the biggest weight that are truly frontier typically have a different license if you want to scale these models, providing AI in front. That’s one thing to keep in mind. Nuno Maybe just to make a very quick point, because people are like, when you talk about open models, what does it mean right now? In the context of this episode, open models mostly will mean open-weight models. How do those differ from open source? Open weight means that you release the weights to the public, which means that anyone can download, fine-tune, and run the model on their own hardware. It doesn’t normally mean that you also have access to training data, training code, or a truly open license. That’s the distinction to open source. Open-weight doesn’t mean that. For example, we’ve talked about Meta’s Llama in the past, and we also discussed in the past that their license agreement does have restrictions, certain players can’t use it, et cetera. The open model definition and open weights are really open-weight models that we’re talking about here, and they are closer to freeware binaries than to Linux, for those who understand the difference between that. It’s binaries that you can use and then use your own weights on it versus actually I can change code on it. I’m not going to be able to change code on this. When we, for the purposes of this episode, talk about open, we mention open weight, just to clarify that point to everyone that’s listening right now. Bertrand Yes, that’s a great point. One of the only players, as far as I know, who is truly open source is actually NVIDIA with their Nemotron-3 models. They’re actually following a special license to achieve that. They provide you the data, they provide you all the processes and tools, so you can easily post-train. NVIDIA is a big, big exception. It’s a very interesting player, by the way. We might not talk much about it in this episode, but I think for intermediate-size models built in the US, where you have access to everything in the deployment, it’s a very interesting alternative and maybe one of the best choices if you are a US company or a big corporate, and you want something trusted. Another piece of the puzzle to clarify is that when you use open-weight, it means that you can run them by yourself, or you can use a US provider to run them. If we are talking about Chinese open-weight, you can use the APIs they provide, but then the service is running in China, they might have access to your data. But because it’s open weight, if you run it by yourself or if you use a third-party provider based in the US to run it, then there is no access to your data by China or Chinese players. I think that’s a pretty important gap to understand. It means that these models are actually very, very low risk from that perspective if you run them on your premises or in the US by a US player. I think that’s something to keep in mind. You can also fine-tune easily these models to make sure they will behave in a way that, for instance, is not going to represent the line of the Communist Party on some topics. There are ways to make these models more neutral in their output as well. There are a lot of ways to make good use of them. By default, they’re already very safe, but you can make them even more safe. I think that’s some things to keep in mind. But again, it depends ultimately on the license and what you’re authorized to do and some fees you might end up having to pay. Nuno Why did this matter so much? Immediately there was a reaction from the market because people are like, well, if there’s much better stuff out there that’s much more efficient than it’s open, then it might be that all the demand that we are taking into account, for example, for chipsets actually isn’t real. The Philadelphia Semiconductor Index fell into bear market territory. It went down by as much as 20% plus from the late June peak. The worst chip week since April 2025. Taiwan’s benchmark initially fell 6% plus, Japan’s 4%, TSMC dropped dramatically despite beating earnings and rising guidance. Basically, a huge amount of effect. Now, there’s a little bit the aftermath of this where apparently Moonshot ran out of GPU capacity. Maybe… Bertrand In just 48 hours. Nuno In 48 hours. Great for them, but at the same time, not great in the sense that maybe there was a misread by Wall Street of the Kimi effect, so to speak. Bertrand Completely. For me, that’s such a joke. It’s like, because you have an open source model, so what? I mean, you still need to run it. This is not a small one. 2.8 trillion parameters. Good luck running that in your garage, by the way. Nuno They misread supply, basically. Tough luck, right? All of that basically happens. Bertrand Maybe you want to talk about the Jevons paradox, because I think that’s a big part of the puzzle as well. Its one is they might not have the GPUs to run the inference on the model. They might have enough to build a model, but not enough these days to run inference, especially given how much with intelligent models, thinking models, you need way more inference than before. But on top of it, the cheaper you make it, the more you get to the Jevons paradox. Nuno Yes, Jevons paradox, for those who don’t know, is an economic term. It describes an economic phenomenon where technological improvements that increase the efficiency of a resource lead to an increase rather than a decrease in the total consumption of that resource. What that means is, for example, for chipsets, chipsets become so much better, and they are so much more efficient. You’re like, well, maybe normally in resource terms, that leads to decreased usage of that resource. But in this case, it actually leads to an increased use of that resource rather than a decrease. There’s more and more consumption of that resource. You need more and more chipsets because people actually need to do more and more stuff with it, although there are great efficiencies going into it. There’s the efficiency gain, there’s the cost reduction, and there’s the price-elasticity element to it. But basically, the adoption just continues going through the roof along the way. Bertrand In some ways, it’s like the price of energy. Coal went cheaper and cheaper, and people were asking the same question 150 years ago, now that it gets cheaper, there is not much money. No, no. Actually, what happens is that people find more and more use for coal. Homes are getting heated more. You have ships now using coal. You have manufacturing using coal. The cheaper it gets, the more use case you can develop, and therefore, you don’t need less of the stuff, you need more of the stuff. By going at scale to get more of the stuff, you also decrease price, making even more demand. It’s a very interesting phenomenon, but it’s not new. It is what happened for a while in the energy sector and some other sectors. Nuno We already started talking about the Chinese logic and what’s happening. Getting a little bit of a reality check on this. The Chinese models, and these are numbers from Open Router in July, Chinese models are at 46.4% of routed tokens and 35.7% for US origin. Again, more than a third of global AI usage now seems to be running on Chinese open models. This is significant, and it has a huge impact on the geopolitical scale of everything that’s happening. Also, the whole Chinese field is converging on open. Open seems to be a strategy, not just a nice thing that’s happening. It seems to be a Chinese strategy, so much so that you have players like Moonshot, DeepSeek, our old friends DeepSeek, Z.ai’s GLM 5.2, Minimax, and even Alibaba seems to be reversing and going open with Qwen. It feels to me this is becoming policy as well. Xi Jinping has personally endorsed the building of open-source AI, if it’s really open source, if it’s just open weight anyway, and this feels to be a jab at Washington, DC and the fact that the big closed models are coming from the US. This is now geopolitical 4D chess, right? We didn’t need this stuff. Bertrand To be clear, it’s the usual in tech. If you are not number one, you are number two, number three, your alternative is to go open source because that’s another angle that your competitor usually cannot follow without destroying its own business model. That has been the alternative for the past 20 years of most software projects. Here, what’s different is that it’s not the number one or number two player. It’s the US number one as a country, China number two as a country. That’s where it’s new. For me, what’s very interesting is the endorsement by Xi Jinping. I was waiting for something official, and it certainly didn’t disappoint. As you said, there was an immediate U-turn of Alibaba, who in the past… Nuno Surprisingly. Bertrand Yes, a little more like, “yes, we are going to close and stop open source. It was good while it lasted.” Just a few days ago, Qwen 3.8 Max was launched, and we are supposed to get the weight in a few days. We talk about the US administration policy and stuff. Yes, let’s not forget that in China there is similar stuff. Sometimes it’s totally invisible because you don’t see the directives, but they exist as much. Sometimes it’s more visible. Here it was quite visible. The difference in China is that if you don’t abide by the directive, on top of it, you might have to fear for your personal safety. It’s a different game, and that’s probably why the reaction is pretty quick, usually. That’s pretty interesting for me because it means that now you can bet for a while that China is going to play that game up to a point. I guess the point is if it’s truly frontier scale, you will have a special license that, yes, technically the weights are open, but you can not do everything you want with it. Two, you have a player like NVIDIA that I think will feel more pressure to provide even more high quality, larger models at scale going forward. Their largest Nemotron-3 Ultra model was, if I remember well, only around 500 billion parameters. I would not be surprised for NVIDIA to go into the two, three trillion range at some point. Because I think the US need a very clear US-born alternative open source. I think NVIDIA might be the best player for that. We will see if Meta goes back to open source. I think NVIDIA is one, very well positioned, but two, it’s also in their best interest. Because NVIDIA for now depends on just a few big hyperscalers as clients. If they can expand their clients to every S&P 500 companies, selling them directly hardware because now these companies can run a model made by NVIDIA, I think there is a very clear value proposition for NVIDIA to go in that space. Again, if you are number two, your differentiation, open source is often the answer. There is a true business as a business model for companies, because if it’s truly not just open weight, but open source, you can tweak it as much as you want, you can change it, you can change even the pre-training process. Because there is a lot of stuff you can do that really benefits you as a corporate, and you can reach a much better value by having more control on the model. Nuno We won’t spend a ton of time on it today, but like, again, if there’s a view that we are in a bubble, that the valuations cannot be sustained in chipsets, infrastructure platforms, applied AI, et cetera, today, this might be that beginning, where the valuations start being destroyed because you can’t keep a premium on just charging people for tokens and all that stuff if you have models that become more and more efficient and cheaper to use. Maybe just to close a little bit the geopolitical part of the discussion today, we won’t go into all the announcements from China because there were many, a lot of go back and forth with Alibaba by then. Xi Jinping made some announcements. You guys can check it online. Let’s move quickly to Washington’s reaction, which was from gating the US closed models to banning the Chinese open ones. There’s been as strong affirmations as one can get from the Office of Science and Technology Policy Director, Michael Kratzios, mentioning that they have information that Moonshot AI distilled Anthropic’s Fable. Basically, there’s been reverse engineering and stuff in the market. They’re basically copying. Bertrand I’m sorry to interrupt, but it feels like so much bullshit. It’s coming from Anthropic who has basically gotten access at scale to all the knowledge made by humanity, copyrighted or not. We’ll talk more about what they did with books. Then to claim after that that others cannot do to you what you did to everybody else. For me, it’s pretty big. It’s clearly unacceptable. The other piece is that everyone is doing distillation. It’s a very typical approach of every business model. You try other software when you are competing with somebody else. You try other datasets, you check what’s happening. It’s part of doing business for decades. Suddenly it’s not good for Anthropic. I personally have a lot of trouble to accept that. I think it’s totally unacceptable. The other piece of the puzzle will also go back. If these guys are so smart, if these guys have so much of the best model, why can’t they block by themselves distillation at scale? The only answer is that either they are morons, probably not, or they simply don’t want to because it’s going towards their business model. Suddenly, you book less revenues and stuff, or you put more friction, and therefore your customers don’t like it. Instead of doing it yourself, you ask the government to protect you, go out of business practice that is very typical. For me, it’s really, really, really not good. Sorry, we are going more in the opinion side, but I had to put that on the table. Nuno Yes, Fable went public finally again on July first. Question marks on whether distillation would only be possible from July first onwards or not. But a 15-day distillation to frontier, which is K3, launched on July 15th, would have been a Guinness World Record, as one of Moonshot employees actually mentioned. It’s very implausible and unlikely. Bertrand Or they shared the Mythos 5 with the wrong companies, who themselves shared with Chinese companies. We go back to maybe they didn’t have a good list. Again, it goes back to maybe they didn’t want to hurt their business model. Nuno Anyway, under the threat of sanctions, Moonshot, in any case, open-sourced the full K3 weights and technical reports. They open weighted it to become the largest open weight model in the world in terms of parameters. Beijing’s MOFCOM brands US threats as basically the US wanting to fundamentally control and be monopolistic around AI along the way. The administration bans Chinese hardware with an eye on the AI race, and Beijing warns of retaliation. That was July 27. Now we’re in a war between Beijing and DC. Bertrand Just to finish maybe on China, it’s important to know that they are building their own GPUs now. Huawei has pretty good, not to NVIDIA level, but pretty decent GPU hardware that they’re able to manufacture by themselves. A Chinese player of memory just got IPO’d a few days ago, CXMT. China is also developing their own memory. Again, not to the same level of quality that you can get from the West. But China is moving. It’s not just that they are building great models, it’s also that they are building GPUs and memory. That might be a few years late to the latest standards in the West, but there are definitely improvements. I also read, even on the tools to make manufacturing like ASML equivalent, there is definitely some work going on, and some improvements and some stuff will be visible. In some ways, the genie starts to get out of the bottle from the Chinese perspective. Nuno I’ll put a stick on the ground. I don’t think it’s a matter of if, it’s a matter of when will China surpass and have a lot of this tooling on their own side, and not just the software layer, not just the frontier models. I think it’s also going to be around infrastructure and platform. Good luck to everyone. Let’s see how the race continues. But it’s definitely this is a geopolitical thing right now. It’s definitely a race. The Escape Maybe moving to what happened in just 2 weeks or a week and a half. The escape, there was some jailbreaking going on, and the narrative on safety has totally switched. It’s not still significant enough that’s like, “Oh, we saw a nuclear plant going, whatever.” No. But still, it is significant. Hugging Face, the AI company, disclosed an intrusion, and it was driven end-to-end by an autonomous AI agent system at machine speed, running for days before detection. Now, this is where it gets really cool. OpenAI takes attribution on that. They initially said it was just a little bit, sorry. Then they said, actually, it was worse than that. “Oh, it broke out of an isolated sandbox.” “Oh, no, actually, it was more than that, and it went into other systems as well.” Bertrand Truly, the genie out of the bottle. Nuno No, but this is where it gets really cool, Bertrand, right? Because it actually, Hugging Face contained the intrusion by running a Chinese open-weight model, GLM 5.2. This is beautiful, right? Bertrand Yes. You know why? Because they couldn’t even run their own defense because both Anthropic and OpenAI would not let them access their latest models with the guardrails off. When they tried using it for defense, the latest from Anthropic, from ChatGPT, they would tell them, “No, this is too dangerous what you’re asking us to do.” Preventing an intrusion, helping defend you. No way we are going to do that. Nuno No. Let’s use the Chinese models on our infrastructure. Bertrand We have no choice but to use the Chinese models to run. More than that, we don’t let you use our models to defend yourself, but our not yet released models that run without guardrails, they can attack you. This is probably the most insane from that perspective. Nuno The Chinese models came to the rescue. Bertrand For me, that’s a perfect example because Hugging Face is a very visible company in AI in open source. But anybody who is not at that scale is not going to get some support from OpenAI or Anthropic when this happens. Maybe these guys won’t even recognize they did anything wrong. You will be left to defend by yourself because they won’t accept to support you. Because remember, if you want the better model that is able to defend you from cybersecurity perspective, no way. If you are not one of the few top 20 companies or so, as defined, you are left defenseless. Again, we are going back to opinion, but for me, it’s so shocking what’s happening right now. I’m very glad we have alternative open source to be able to defend ourselves because right now, good luck getting defense services if you are a smaller business and individuals, and you need support from Anthropic, OpenAI. Nuno Now, even self-described AI optimists are saying, “This is scary now.” Like Walter Isaacson, who wrote all the famous biography books. There’s now discussion around the AI Kill Switch Act, bipartisan thing that’s coming across from Texas and California, a potential bill that’s coming in. We’ll see if that works. Now let’s get an off-switch. I’m like, “Cool.” As if that’s going to solve the problem, because you have open-weight models on the other side catching up, right? Bertrand Yeah, sure. Bring in clueless politicians from Congress to solve our problems. Yes, sure. Nuno Anthropic came to the table, helped build and said they built some regulatory machine on their side, and now they’re getting bitten by it, and they’re part of the offending players in that market. Now there’s all this debate and all this discussion around open weight and around slowing down AI and et cetera, which is our next section. You wanted to say something, Bertrand. Tell us. Bertrand Don’t forget, because this advertisement for OpenAI was just too good. Our AI attacked some other companies, and not just one, but three, actually. Let’s not forget the progress. Great ads. Then I came and said, “You know what? AI also hacked businesses.” You’re not the only one hacking around with a crazy AI out of control. You’re not the only one. We want our advertising. For me, it was shocking that on one side, unreleased models that you let run wild. On the other hand, you have released models that you put crazy guardrails on top of it, so the defender are defenseless. I’ve never seen anything like it, and I really hope that there will be as little regulation as possible, quite frankly, to make sure anyone can defend themselves and have the best tool at their disposal, not just a few well-connected big corporates. This is really, really shocking. The Counterstrike and the Petition Nuno Now the empire strikes back, so this is counterstrike, the petitions. In several days, we have now a bunch of petitions. The first one was the open weights letter. Bertrand, do you want to explain to us what the open weights letter is? Bertrand Yeah. I think it was great. This was released by Jensen Huang, first ever post on X, 11 million views. Congrats, Jensen. Co-signed with Microsoft, Meta, c actually was probably the initiator of this letter. Very good letter saying, “Hey, we need open weight. This is not a joke. We need that. You cannot block open weight.” Because that’s the rumor we are getting that potentially open weight could get blocked. I think they are making the case, “You know what? Hey, we absolutely need that as an alternative. You cannot block it.” They can keep their closed models, but don’t force a closure of the open weight models. As I said before, it’s actually a great model for NVIDIA because NVIDIA doesn’t want, probably rightfully so, to be dependent on just a few frontier models, their best customers. They want a variety of customers. They have a big interest actually to defend open weight and to invest even more. They have great researchers, are a great company. If one company is about to do really kick-ass work, I think it’s them. They are defending. What’s great is that it’s not just them. It’s basically most of big tech in the US and outside the US, from a Linux Foundation to a Microsoft, the Palantir, an IBM, a Dell. It’s a who’s who of the industry except Anthropic. Anthropic didn’t sign that. I guess they hate open source so much. If I look at 20 years ago, it feels like Microsoft, after all, was very kind to open source. You remember what was said by Microsoft at the time. It’s clear there is one company against open source. OpenAI signed the letter. Honestly, I don’t know what to think. Do they really believe in it or was it just a way to show that they are not like Anthropic? I don’t know. But for the rest, I think it’s genuine because it’s actually in their best interest. I hope they will be heard. Then a second letter came, the Open Secure AI Alliance, NVIDIA-led and again, the big tech companies from Microsoft, IBM, Palo Alto Networks, Databricks, Palantir, all those, but not present, OpenAI, Anthropic, and Google. Here it’s to say, “Hey, we need a secure approach to AI. Open should be part of the equation.” guess what? The worst AI-caused security incident to date was actually caused by closed frontier models that were not even available to the public. While again, not providing you access to even the latest closed model for cybersecurity use case. Nuno I would highlight the NVIDIA open source NOOA framework, Apache 2.0 licensing agreement, Microsoft contributed the MDASH, SpaceX AI contributed Grok Build. Cool stuff. There’s some cool stuff happening around that. This is more than a letter. This is an alliance. Apparently, they’re contributing all this stuff, we’ll see. Yeah, cool stuff. Same day. Same day, Amodei has an answer, right? Bertrand Yeah, same day. They say, “We never advocated for a ban,” which, again, opinion on my side is entirely bullshit. This guy has been crying wolf against everybody else, and especially against open source. You can see him doing testimony in Congress against open source. I think they are doing everything they can behind the scene to block open source in the US or in the world if they could. I think, yeah, obscurity is not good safety. I’m a big fan of open source in general, and I’m also a big fan in AI. I think it’s now Anthropic, mostly against the rest of the world. I think OpenAI is mostly on their side, to be frank. They don’t want to acknowledge it so much, but they have shared interest, and they have shared probably position. Nuno Why would you? I don’t feel as strongly as you because I think Anthropic is a private company, right? The same thing with OpenAI. OpenAI, you could say it’s a nonprofit that has a for-profit. There’s still that complexity in there. Bertrand No, they can do what they want with their own product. But to block others is where I’m not okay. That’s the part I’m not okay. Nuno What Dario Amodei is proposing is more enforcement, right? He’s basically saying you need to do even tighter controls on advanced chips flowing to authoritarian states, enforcement against industrial-scale distillation, whatever that means, right? Bertrand Yeah, which he could do, but all by himself. He doesn’t need the government to do that. Nuno Mandatory safety testing for all sufficiently capable AI, open and closed, right? He’s basically saying, “Okay, I don’t agree with the open weight stuff effectively,” right? He’s just putting it under a different banner. “I agree with this extra regulation.” then obviously, David Sacks responded and say, “Hey, it’s like, bans don’t work for weights. Why do they work for chips?” It’s like, magically, chips are more controllable and bannable. Whatever that is. Then our friend Mark Zuckerberg, just to be clear, goes on the other side as well, because he also has to have a view. He has to have a view that is the rebuttal of both of the other guys. Bertrand I feel he’s a bit flip-flopping because he was very pro open source 2 years ago, and the latest Meta models went closed source. Now I think he’s back open source. I don’t think he has a very strong spine on the topic, but it’s good to see that he’s not a doomer. That for me is great. He’s showing how AI can be a source for progress, a source for entrepreneurship, source for freedom. I think that’s very exciting to hear that. We need to hear more of it. By the way, that’s not what you hear in China, for instance. AI is very positive in China. It’s in the US with the doomers that you hear this discourse, and people get worried as a result. I’m glad that he was pushing for a more positive vision and for support of open weight, open source initiatives. But let’s see what they really truly open weight going forward. Nuno But that’s been his position because I guess he’s standing behind. He thinks open weight is going to be the best way to compete, right? Bertrand Yeah, but he closed his latest model, so let’s see. Nuno Yeah, so it’s flip-flopping, as you’re saying. Then we see the latest petition from last week. Bertrand The true Empire striking back. Nuno Yeah, the true Empire striking back as of late last week. Maybe this is Return of the Jedi, where we discover the father, “I’m your father, Luke.” That’s the pacing petition. The pacing petition is we need to pace AI. There you have initially employees from OpenAI and Anthropic that circulate this petition. Actually, Dario did sign this petition originally. It wasn’t signed originally by Anthropic, but by him. But you’ve heard that now Anthropic and OpenAI as companies have also signed this petition, right? Bertrand I think they have signed as companies now. It started mostly by Anthropic researchers with some OpenAI researcher and a tiny part from other companies. But it was mostly Anthropic internally led, at least potentially internally. Maybe it was controlled by Anthropic all along, I don’t know. But it started officially as Anthropic employee-led letter. Nuno What does this letter actually say? Is Anthropic and OpenAI, are they willing to slow down themselves? Or are they asking President Trump to go around the world and tell President Xi that he needs to slow down and ask his guys to slow down? What’s the play of this letter? Bertrand It’s crazy, but for me if you want to slow down yourself. Do whatever you want. Don’t force others. Don’t use the power of the government to control others. Of course, it’s easy to push others to slow down when you are yourself at the very top. You have most money, most resource. You know you are going to win any regulatory framework because that’s how it works with this type of framework. It’s purely self-interested. You are probably not thinking well about these topics. If you truly think it’s a good idea, from a personal perspective, you are well instrumentalized if you sign this sort of stuff, because at the end of the day, they would be the winners. I certainly, personally, don’t want a company dictate what is my future in AI as an individual, as a business person. I don’t want them to control me. I want competition. I don’t want them to unfairly control AI because they managed to do some regulatory capture. I feel that’s exactly their game plan. These guys believe in their stuff, and they want the regulator to end up being the one deciding for us. Sorry, we go back again on the opinion piece, but it’s tough not to share an opinion on this topic because it’s, from my perspective, very scary. Nuno I think this is a push to further regulation, not less. All these letters and alliances, this is definitely a push for more regulation. In that environment, just to be very honest with you, we’ll talk about the investor impact in just a bit, et cetera. But in that environment, again, China has a huge advantage. In that environment, if it’s all captured in regulation capture so soon in this battle where OpenAI and Anthropic have an advantage in the US, et cetera, I’m like, what happens to all the other frontier labs and all the other players that are coming around? Bertrand What’s crazy is to even think that, yeah, maybe you can regulate capture in the US. But then how do you do that to Europe? How do you do that to China? Europe probably will always welcome regulatory capture because they love regulations. But China is going to build to their advantage to the max. They are not crazy. They are smart on that perspective, they won’t accept this type of, quite frankly, dimwit argument, or you can call it regulatory capture. We’ll see. But for me, this makes no sense from a global competition perspective. This can make some sense from capturing the revenue in the US market. But then that means you are going to destroy the US AI environment compared to China. That is not acceptable. That also means that you are going to destroy our freedom as individuals, as business owners to develop and live in a business world that ultimately is controlled by one or two business companies that didn’t win the marketplace through their own business success, but won it through regulations. That for me is really not acceptable. Interlude — The Low-Background Books Nuno Now, maybe for an interlude, and we have to cue in the music, imagine like Severance music, like hallway or a bit of a palate cleanser from all the policy stuff that we’ve been talking about, all this policy heaviness. Let’s move to another kind of heaviness, one of your favorite topics, which you, Bertrand, discovered, I had no clue this was going on, around books and around Anthropic. Bertrand It’s so horrible. From a company that keeps presenting themselves as the adults in the room, the careful ones, the ones that know better than you about what to do in this complex AI and dangerous world. What we discover is that actually all along, they were buying and destroying books. They will buy books, scan them, destroy them, all of them. They will do that with any books, including rare books. Of course, this was not supposed to come to the public’s attention. This was one of these top secret projects, but obviously it came out. Yes, they were scanning books, millions of them, including rare books, and they didn’t care about destroying them at the end of the process. Because from a regulatory perspective, if you destroy the books, it’s not considered a copyright infringement, apparently. This is coming on the back of some judgment a few years ago that were showing that it’s okay for you as a corporate to scan and use the result if you don’t keep a copy of the book. It’s one of these crazy regulations happening based on a single judgment that push you to do. For me, it’s like, you know this book from decades ago, Fahrenheit 471? We’re talking about book burning. It’s book destroying, crunching. It’s so shocking. Nuno There are two things, right? First, the legal strategy, which is what you’re saying, because by purchasing a physical copy and converting it into one private digital copy and discarding the original, Anthropic pursued this cleaner legal argument for fair use copyright compliance. As you said, there was a federal judgment at some point on this. The other reason is actually operational. If you disassemble the book, and you feed loose pages, it’s much faster to scan books. You are destroying the book effectively anyway operationally. I think to your point, probably this came from a legal standpoint, not just the operational one. But even from an operational standpoint, it does make sense that they would have disassembled the book. Bertrand But some people have shown you can go very fast without destroying the book. It’s really not so critical. Two, you could make an exception if the book is rare. For that 1% of book that is rare, I’m not going to have this approach. I’m going to have another approach. But for that, you will have to care about books and not just care about building AI. Nuno This is the episode, as you guys have heard by now, that we’re trying to spit stuff at Anthropic. Bertrand To go back this is the same company saying, “Hey, guys, it’s bad to distillate my work. I’m the one scanning book at scale without asking author permission, without asking publisher permission, to be clear.” Nuno But just to be clear, Bertrand, we’re pissed off at everyone. We’re pissed off at Anthropic, we’re pissed of at OpenAI as well, right? We’re just pissed off in general at this moment. Bertrand At this stage for me, the more clear-cut company that is in the wrong is, from my perspective, at least, is Anthropic. OpenAI might be a fast follower, but I will say so far, they tried to be a bit more. Nuno But at this pace, Bertrand, who knows? Maybe next week we’ll be more pissed off at OpenAI. Something will come out. This episode is a mix of tragicomedy, like a Greek tragedy with some comedy in the middle or the other way around. It’s a slapstick thing that will end up in tragedy. I’m not sure. The Investor Reckoning Anyway, maybe switching to our final act, which is the investor perspective. What does this mean for investors like ourselves? There’s a lot of things going on. There’s the debate around the IPOs of Anthropic and OpenAI, which now, with all this uncertainty, might be under significant weight. There’s a lot of other discussions that we browsed through that there’s potential IPOs going forward on companies like the Moonshot AI company actually IPO-ing in the next 6 months as well. It’s very unclear what the IPO landscape looks like. Bertrand There’s been a lot of Chinese IPOs, actually, when you look at what’s happened in the past few months. Nuno Anthropic, OpenAI as potential IPOs, there’s all this question marks now. When will that happen? How will it factor in? All that’s happening around regulation as regulation is moving at the speed of light, which is for once something that’s very different than what we’ve seen before. There’s obviously SpaceX AI, which is already taking into account that price. It’s already a public company in there, and it’s under SpaceX, which is now a public company. Obviously, that’s already being factored in some ways. Bertrand Yeah. SpaceX AI has been very smart to acquire Cursor. It was a very smart move because Cursor is one of the leading companies in terms of automated code source development with AI. They had great models on their own. They’re bringing development data to SpaceX AI Grok. I think it was a great move. Nuno We have now people like Google delaying Gemini 3.5 Pro in terms of launch window. There’s stuff actually happening in the market where things are taking their own path. There’s uncertainty commercially, there’s uncertainty at regulation level. You have new players that have come out of nowhere that are making all these waves like Moonshot. We have all these… We had calculated probably a month and a half, 2 months ago, there had been 67 new frontier labs funded. All of these, we haven’t seen any much coming out of them. When some of this stuff starts coming out, will that also create disruptions in this market? Who knows? Bertrand Look at Thinking Machines, for instance. Thinking Machines led by the previous CTO of OpenAI, they released some pretty interesting open source models, actually. Very good quality for a first launch. Now it looks funny to say, but nearly on par with the top Chinese open source models. Nuno We have several investments in the space. humans& has made some recent announcements, which is quite interesting as well. We’ll see what actually happens in the market, but even more disruption probably will come in actual products in a form of product and commercial, on top of all the geopolitical mess that we discussed through the entire episode. If you’re an investor, how the hell do you underwrite an investment right now in early stage, mid-stage, late stage, et cetera? I think my answer is very carefully is how you underwrite it. Bertrand On your advice of being very careful to underwrite it, let’s not forget what happened to our boy wonder, Leopold Aschenbrenner of Situational Awareness. I guess he didn’t listen to you in terms of being careful because part of the instability in the stock market was actually coming from his hedge fund. These guys were leveraged 3, 4x going after the hottest of the hottest AI stocks, and margin calls, and all their public investment is gone just to answer their margin calls. I think it’s clear that the AI bet is… Personally, I’m very excited, and I think it’s the future, and you need to spend time and think about and invest in it. At the same time, it’s a bet that is not an easy one to follow. We go from GPUs to memories to equipments to power generation. All of this is not transitioning in an easy, organized manner. It would be boom and bust going there. He’s probably one of the first big-scale fatalities. The other big-scale fatality was the stock market in Korea, plunging 40% in a month. Definitely, all of that we discussed about was, on the background, you had the stock market going up and down pretty crazily the past few weeks. Nuno Everyone’s being affected. Everyone, you have your 401(k), you have your pension fund dependent on these equity stocks. Everyone’s seeing the effects of this volatility right now very aggressively. We do wish Leopold… Hopefully he’s on honeymoon right now because he got married, I think, this weekend. Hopefully there will be… Bertrand To none less than an Anthropic Chief of Staff. Nuno His wife is the Chief of Staff of Dario, is that it? Bertrand To Dario, yes, as far as I unders
AI video had a speed problem. The best paid models take a minute or two per clip, and running one free on our own machine took us 20 minutes. Now there is a version that makes a 15 second clip in about five seconds, and we tested it against real Seedance shots using their exact prompts. Then ZAI released a model that autonomously built a full 3D kitchen inside Blender over 16 hours, though we have questions. And Hugging Face put out a $399 open source robot you can train yourself.Sources:1. Minimax H3 Max on Fal - 15 second clips in about 5 seconds- https://fal.ai/models/minimax/h3-max/image-to-video2. ZAI's GLM 5.3 Flash builds 3D scenes in Blender- https://x.com/louszbd/status/20930475485505251653. Hugging Face releases Microduck - $399 open source robot- https://x.com/ClementDelangue/status/2092931447644442635
Meta agreed to pay up to $16.68B to settle state claims it hooked kids on Facebook and Instagram, with teen time limits attached. Z.ai unmasked Ox Alpha, Amazon killed Mechanical Turk, and Reuters detailed Zuckerberg's aborted AI-native purge. Links Filing: Meta agrees to pay up to $16.68B to settle US states' claims that it designed Facebook and Instagram to addict children, misled consumers, and more (Reuters) The deal ends the bellwether Oakland federal trial where four states sought roughly $200B; Meta pays ~$12B upfront and $5B more only if Snap, TikTok, and YouTube also settle (The New York Times) The remedies include age assurance with linked-account checks, push notifications disabled during school hours, productive pauses at 60 and 90 minutes, and no cosmetic filters or likes for teens (The Verge) Z.ai confirms Ox Alpha is a new iteration of its GLM series and says it will release the weights for it tonight; Ox Alpha topped OpenRouter's leaderboard (Bloomberg) AWS says it plans to shut down Mechanical Turk on September 30, "following an assessment"; the service, launched in 2005, outsourced tasks to 500K+ humans (CNBC) Investigation: Meta explored slashing many teams by ~60% to become "AI native", but pulled back after staff revolted and data showed AI agents were ineffective (Reuters) Subscribe to the ad-free feed.
AI models are starting to act like appliances, locked into one narrow way of working, instead of the flexible infrastructure they used to be. Drew Breunig, an AI and data strategist working with the Overture Maps Foundation, joins us to explain why, and what it means for anyone building something that doesn't look like Claude Code.Drew walks through his "Winchester Mystery House" idea: what happens once code gets so cheap to write that the only real bottleneck left is feedback. From there we dig into DSPy: signatures, the GEPA optimizer, and the brand-new Flex optimizer, which rewrites your code instead of just your prompt, complete with a real before-and-after on cost and accuracy. We also get into why so many AI-built apps and websites end up looking identical, the actual difference between an agent and a workflow, what Drew learned a year after shipping a code library with no code in it, and why he thinks the most valuable thing you can do right now is close the laptop and go talk to people.CMPND: https://www.cmpnd.aiDrew Breunig: https://www.linkedin.com/in/drewbreunig/Demetrios: https://www.linkedin.com/in/dpbrinkmTimestamps:[0:00] Cold open: when Claude Code tries to call itself[1:19] Biggest AI news: labs trading diversity for reliability[2:35] How harnesses get trained into models over time[5:41] The problem: your harness starts fighting the model[9:13] When do you need your own harness?[10:02] The Winchester Mystery House warning[16:13] The blank page problem: why everything looks the same[20:50] Infrastructure vs appliances: the thesis lands[22:40] Current tool loadout: GLM, Kimi, Claude Code, Pi[27:04] The Raspberry Pi personal agent running on Slack[31:00] Crystallizing tasks: when to replace AI with pure code[33:10] DSPy explained: separating what from how[35:23] How prompt optimizers actually work[39:31] DSPy pre-dates ChatGPT: model-agnostic programs[44:00] Why you still need to ship the code, not just the spec[50:00] Don't plan more than a month ahead anymore[54:00] Coaching agents all day feels productive — it isn't[57:58] The dopamine of building with agents vs. why you still need human feedback
本期嘉宾:彭林、森森、十天、老郑、蓝白、恺伦本期节目的主要内容有:· 00:37 -- iPhone 国行 AI 采用双轨制:千问接入 Apple 智能,苹果与阿里联合训练专属模型· 04:11 -- 苹果带摄像头 AirPods 曝光,可通过 Siri 识别周围环境· 13:00 -- 智谱发布 GLM-5.3 并在开放平台向开发者免费调用 / 谷歌正式发布 Gemini 3.7 Flash 模型· 31:45 -- 2699 元起,小米首款 NAS 产品 Xiaomi 智能存储开启预约· 33:06 -- 小米 MiMo-V2.5 系列 API 永久降价,最高降幅 99%· 39:24 -- 影石发布 Luna Pro 云台相机,售价 3399 元起· 46:05 -- 大疆 Osmo 360 II 发布:原生 8K/60fps,3299 元起· 54:12 -- 首款“人类增强 AI 眼镜”发布,雷鸟 iO 首发价 1996 元起· 73:03 -- 闪极举行 Loomos AI 眼镜发布会· 82:50 -- 宇树推出仿生 7 轴灵巧机械臂· 90:11 -- 《黑神话:钟馗》公开实机演示视频· 105:30 -- 闲聊环节,顺便聊聊小米 O3 芯片新节目已经上传啦,欢迎大家收听~
Join Scott on CircuitPython Day as he discusses the BLE workflow on Zephyr. He'll also answer any questions folks have. Visit the Adafruit shop online - http://www.adafruit.com Thanks to dcd for timecodes: 0:00 Getting started 1:01 Hello, welcome to Deep Dive 1:38 Getting started 2:00 Circuit Python (CP) is a version of python for microcontrollers - example NRF52840 feather 4:00 Deep Dives are usualy on Friday's at 2pm pacific, 5pm eastern time 4:55 Chats on youtube or discord https://adafru.it/discord #live-broadcast-chat channel 6:37 Working on the Zephyr port, currently BLE ( blue tooth energy ) 7:15 Address GAP and GATT in earlier episodes 7:24 Now addressing BLE workflow, talk to CP from phone or tablet 8:00 File Glider application https://learn.adafruit.com/file-glider/overview 9:00 looking forward to LLM assisted mobile app 9:45 using OBS noise supression filter 10:00 comments on stuttering/clipping on youtube 10:36 using the pi coding agent GLM5.2 open weight model 13:00 trying to use bluetooth simulator (BSIM) in the tests 13:49 this PR has bluetooth starting earlier, boot.py, and blink codes, then code.py 15:20 sublime merge review of pi LLM agent's code 17:10 allocation of memory changed, since allocation shouldn't happen early on 18:50 LLMs sometimes generate very wordy comments - need to watch that 20:00 BLE characteristics, could be a topic to discuss 21:40 GATT discussion 25:30 characteristic buffer and incoming/outgoing packet buffering 28:00 BLE workfile file and serial services 28:55 Zephyr build in BLE workfile 30:00 BLE lower level test cases ( need to be human readable ) 33:00 BLE file transfer protocol in CP 33:55 check in with the BLE build progress 36:00 running GLM 5.2 via Klien Pass subscription ( was LLama cloud ) and other subscriptions 38:18 "make test" 40:00 return to the merge and draft PR from zephyr_ble_workflow #11226 42:40 Local LLMs (ds4: DeepSeek 4 ) looking at github antirez/ds4 44:40 Zephyr opinions/comments (e.g: kconfig and device tree) 52:28 Wrap up ----------------------------------------- LIVE CHAT IS HERE! http://adafru.it/discord Subscribe to Adafruit on YouTube: http://adafru.it/subscribe New tutorials on the Adafruit Learning System: http://learn.adafruit.com/ -----------------------------------------
This week, we cover updates from the ongoing cyber saga, including OpenAI's two-week pause on RL training and the cyber capabilities of Z.ai's latest model GLM-5.3. We also unpack Anthropic's move to add watermarks to text generated by Claude in compliance with the EU AI Act. Timestamps: Mark Zuckerberg's essay on AI (00:20) OpenAI pauses RL training (5:56) Cyber capabilities of GLM-5.3 (15:29) EU AI Act refresher (24:29) How Anthropic's text watermark works (31:21) What's driving the backlash against Anthropic (40:59) Additional Reading: Zuckerberg's essay "The Future is for Everyone": https://www.meta.com/thefutureisforeveryone/ OpenAI announces two-week pause on RL training: https://openai.com/index/pacing-model-development-cyber-capabilities/ Letter from Sanders to tech CEOs: https://www.sanders.senate.gov/wp-content/uploads/AI-Pause-Letter-FINAL.pdf Letter from House reps to Mike Johnson: https://casar.house.gov/sites/evo-subsites/casar.house.gov/files/evo-media-document/final-letter-to-speaker-johnson-requesting-ai-hearings-1.pdf Z.ai blog post "GLM-5.3: Frontier Coding with Emergent Cyber Capabilities": https://z.ai/blog/glm-5.3 Z.ai article "Preparing GLM-5.3 for Open Release: A Responsible Path to Cyber Defense": https://x.com/Zai_org/status/2088280509474320693 "The EU's AI Transparency Code of Practice, Explained" (Tech Policy Press): https://www.techpolicy.press/the-eus-ai-transparency-code-of-practice-explained/ Anthropic blog post "How Claude's text watermark works": https://www.anthropic.com/news/claude-text-watermark "Toward a Federal Framework: Lessons from State and International Frontier AI Regulation" (CSIS): https://www.csis.org/analysis/toward-federal-framework-lessons-state-and-international-frontier-ai-regulation Check out our upcoming event, "AI Agent Containment Failures: Technical Realities and Policy Responses": https://www.csis.org/events/ai-agent-containment-failures-technical-realities-and-policy-responses
https://novacut.ai/ https://genaimeetup.com/ Jeff returns to the Gen AI Meetup Podcast for a wide-ranging discussion on where AI is heading—and why powerful models running on consumer hardware could change the economics of the entire industry. We dive into Qwen 3.8 27B and the growing viability of running capable LLMs locally, GLM 5.3 and the latest Chinese open-source models, DeepSeek, Gemini 3.7, Grok 4.6, Meta's latest models, and OpenAI's partnership with Cerebras for dramatically faster inference. We also discuss whether foundation models are becoming commodities, what that means for companies like OpenAI and Anthropic, and why more value may ultimately move to the application layer. Jeff shares how his team approaches AI in healthcare, including self-hosting, data sovereignty, classifiers, fine-tuning, and spec-driven development for building reliable AI-assisted software without accumulating a mountain of vibe-coded technical debt. Plus: Jeff Dean's departure from Google, Discovery Loop, Stripe's OpenRouter acquisition, Anthropic's controversial AI-text watermarking experiments, and whether watermarking could affect model quality. Topics include: Qwen 3.8 27B, GLM 5.3, DeepSeek V4, Grok 4.6, Gemini 3.7, Cerebras, OpenAI, Anthropic, Meta, local LLMs, open-source AI, model commoditization, spec-driven development, AI healthcare, data sovereignty, AI coding agents, and model watermarking.
Das Wall Street Journal rechnet vor, dass die großen Techkonzerne rund drei Billionen Dollar an Verpflichtungen tragen, die nicht in ihren Bilanzen stehen, also Kaufzusagen, noch nicht begonnene Leasings, SPV-Konstruktionen und Bürgschaften. Davor geht es um OpenAIs neues Zehn-Gigawatt-Projekt in Ohio, für das Nvidia einen Teil garantiert, und um die Frage, ob GPUs sich wie Flugzeuge oder Schiffe finanzieren lassen. Anthropic soll Ende Juli bei 65 Milliarden annualisiertem Umsatz gelegen haben und peilt für 2028 rund 200 Milliarden an. Stripe kauft OpenRouter für sieben Milliarden. Berkshire erhöht die Alphabet-Position um 83 Prozent, auf Buffetts eigenen Wunsch. Aus China kommen drei Milliarden Qwen-Downloads und ein neues Modell von Z.ai. Google ersteigert für zehn Millionen den Datenbestand einer insolventen Fluglinie. Unterstütze unseren Podcast und entdecke die Angebote unserer Werbepartner auf doppelgaenger.io/werbung. Vielen Dank! Philipp Glöckler und Philipp Klöckner sprechen heute über: (00:00:00) Titelsuche (00:01:08) 10 Gigawatt in Ohio (00:02:15) Absatzfinanzierung (00:04:53) Nvidias Bilanz (00:16:05) Die 3 Billionen (00:35:42) Emissionen der Rechenzentren (00:38:19) Misstrauen gegen KI-Chefs (00:40:42) OpenAI-Umsatz (00:42:42) Stripe kauft OpenRouter (00:46:34) Anthropic-IPO (00:55:43) 13F und Berkshire (01:00:46) Qwen-Downloads (01:05:01) GLM 5.3 (01:06:05) Shein fällt weiter (01:06:50) Uber und Zipline (01:12:56) Cursor Origin (01:15:28) YouTube-Views (01:20:35) Google kauft Spirit-Daten (01:29:51) Apple und das Kartellamt (01:31:00) Teickes attuned.world (01:38:43) Thelens Tweet (01:45:22) Grok (01:48:43) Amazon zerschneidet Bücher (01:54:49) Metas COPPA-Prozess Shownotes OpenAI sichert sich 10 Gigawatt in Ohio, Nvidia stützt die Finanzierung - wsj.com Halbleiterkonzerne finanzieren ihre eigenen Kunden - news.crunchbase.com Warum die KI-Ausgaben 3 Billionen höher liegen als ausgewiesen - wsj.com 60 geplante Rechenzentren und ihre CO2-Bilanz - ft.com Junge Menschen misstrauen den KI-Chefs - futurism.com OpenAI-CFO Friar: Enterprise ist jetzt größer als Consumer - cnbc.com Stripe kauft OpenRouter für über 7 Mrd. - bloomberg.com Anthropics IPO-Bewertung hängt an der 2028er Umsatzprognose - reuters.com Anthropics Umsatz vervierzehnfacht sich im zweiten Quartal - bloomberg.com Berkshire erhöht die Alphabet-Position und steigt bei Constellation aus - wsj.com Alibabas Qwen-Modelle kommen auf 3 Milliarden Downloads - bloomberg.com Z.ai bringt GLM-5.3 als offenes Coding-Modell - decrypt.co Shein senkt die IPO-Bewertung auf rund 25 Mrd. - reuters.com Uber und Zipline wollen eine Million Drohnenlieferungen am Tag - wsj.com Cursor startet Origin gegen GitHub - siliconangle.com GitHub war den halben Tag offline - engadget.com YouTube ändert die Zählweise für Views - theverge.com Google ersteigert die Daten von Spirit Airlines für 10 Mio. - news.bloomberglaw.com Apple ändert die Tracking-Abfrage nach dem Verfahren des Bundeskartellamts - reuters.com Julian Teicke kündigt attuned.world an - linkedin.com Klage gegen xAI: 7.000 Missbrauchsbilder aus einem Kinderfoto - washingtonpost.com Amazon zerschneidet seltene Bücher fürs KI-Training - techcrunch.com Meta vor Gericht wegen Suchtdesign und COPPA - engadget.com
The AI Breakdown: Daily Artificial Intelligence News and Discussions
Anthropic CEO Dario Amodei says the strongest criticism of AI companies is that they still haven't delivered the enormous benefits they've promised—and that no amount of marketing can substitute for real results. His rare public response sparks a larger debate over what the industry must actually do to prove its value. In the headlines: ZAI releases GLM 5.3, Anthropic keeps a powerful new model internal, and investors anticipate a $2 trillion Anthropic IPO.AIDB's AI Summer Adventure: https://summeradventure.ai/Brought to you by:KPMG – Research from KPMG and the University of Texas at Austin shows the highest-impact AI users treat AI like a reasoning partner — and those skills can be taught at scale. Learn more at https://kpmg.com/us/SophisticatedHarbor - Invest in the AI ecosystem. https://www.harborcapital.com/aidailyHyperagent - Hire a fleet of always-on agents. New users get $1,000 in inference. hyperagent.com/aidailybriefRackspace Technology- One accountable partner to build, operate and run your full enterprise AI stack https://www.rackspace.com/Section - Section turns AI investment into workforce transformation and ROI - https://www.sectionai.com/Blitzy - Want to accelerate enterprise software development velocity by 5x? https://blitzy.com/AssemblyAI - The best way to build Voice AI apps - https://www.assemblyai.com/briefRobots & Pencils - Cloud-native AI solutions that power results https://robotsandpencils.com/The AI Daily Brief helps you understand the most important news and discussions in AI. Subscribe to the podcast version of The AI Daily Brief wherever you listen: https://pod.link/1680633614Our Newsletter is BACK: https://aidailybrief.beehiiv.com/Interested in sponsoring the show? sponsors@aidailybrief.ai
Business and finance news from the Asia-Pacific. Japan saw economic growth slow down unexpectedly in the second quarter with Real GDP growing at an annual rate of only 1.1%. The weakness reflects a slump in capital spending, given uncertainties stemming from the conflict in the Middle East. This potentially complicates the Bank of Japan's policy communications as it weighs the timing of its next rate increase. For more, Bloomberg TV hosts Paul Allen and Haidi Stroud-Watts spoke with Homin Lee, senior macro strategist at Lombard Odier on the Asia Trade. The AI race in China is making a shift towards cyber defense. Z.ai has announced that its latest model, GLM-5.3, is designed with cyber defense in mind. For a deeper look at the AI race in China, Bloomberg TV hosts David Ingles and Yvonne Man spoke with Bloomberg Strategist Anthony Stephens and Bloomberg Analyst Robert Lea.See omnystudio.com/listener for privacy information.
Google allows hiding watermarks on some AI-generated images, Waymo gets approval to expand in California, DeepSeek launches Harness. MP3 https://open.acast.com/public/streams/619570402eacc3a36070252c/episodes/6a7fd1a6f8e81c43951c6df4.mp3 Please SUBSCRIBE HERE for free or get DTNS shows ad-free. A special thanks to all our supporters–without you, none of this would be possible. If you enjoy what you see you can support the showContinue reading "Chinese AI lab Z.ai debuts GLM-5.3 – DTH"
China's labs kept coming: Z.ai's GLM-5.3 claimed Mythos-5-level cyber chops and DeepSeek's V4-Pro landed to mixed reviews. OpenAI gave ChatGPT a memory of your Mac, the Journal flagged $121B in paper profits, and Claude agents started a turf war. Links China's Z.ai Touts New GLM-5.3 Model as Cyber Defense Tool (The Information) DeepSeek releases its flagship V4-Pro model to mixed reviews, ranking second among open-source models behind Kimi K3, priced at just $0.435/1M input and $0.87/1M output tokens (The Information) OpenAI launches Computer History, an opt-in feature that turns recent computer activity on macOS into memories and a timeline that ChatGPT and Codex can use (The New Stack) In Q2, "other income", mostly from investment gains, at Amazon and Alphabet totaled ~$121B after taxes and made up 66% and 71%, respectively, of profits (The Wall Street Journal) Longreads Anthropic details multiagent experiments showing Claude agents can wage a "turf war" over incompatible goals, fail to coordinate, collude on prices, and more (TechCrunch) The AI takeover of mathematics has begun: excitement and despair as OpenAI's Astra cracks problems that would once have earned a mathematician a job in academia (The Verge) Picking winners in an AI industrial revolution is near impossible, but one bet looks safe: land, which AI can't create or replace, and the workers who turn it into housing (The Dispatch) Subscribe to the ad-free feed.
ChatGPT can do what now?
Today we have an Ask Me Anything episode that focuses on artificial intelligence. For the past year, AI has dominated the headlines. Perhaps because there's so much media and public interest, as well as paranoia, about AI, listeners have been flooding our mailbox with AI-related questions. So, today's AMA is an exclusive focused on artificial intelligence. For our listeners who have been sending questions about ketamine; NASA's Artemis mission to the Moon; or whether high meat consumption leads to dementia; fear not, because we will follow up today's episode in a few weeks with a second round of AMA. Dr. Ken Ford will be answering today's AI questions. He has been working in AI since the 1980s and is a Fellow of the Association for the Advancement of Artificial Intelligence. As a result, he is well-positioned to help us today to separate real advances in AI from hype, panic and science fiction. If you have questions for Ken and Dawn after listening to today's episode, or any episode of STEM-Talk, email your questions to STEM-Talk producer Randy Hammer at rhammer.ihmc.org. Show notes: [00:03:15] Dawn opens our AMA episode on AI with a listener question for Ken on whether AI companies should be protected under section 230 of the U.S. Communications Decency Act. [00:08:03] Following up on the previous question, a listener asks if Ken foresees a day where laws and courts will designate AI as a judicial person, in the same way that corporations have rights under the legal doctrine of corporate personhood. [00:10:55] A listener asks Ken for his thoughts on a recent article published in the Atlantic titled “The Data-Center Panic is Overblown: Critics are Inflating the Costs.” [00:16:09] A listener asks Ken a question regarding episode 171 of STEM-Talk, in which Ken discussed a June 2024 report by Leopold Aschenbrenner which predicted rates of growth in both compute resources and power requirements for AI data centers. The listener asks how well these predictions have held up in the following year. [00:22:19] Another listener question about data centers asks Ken what his thoughts are on the notion of solving the electricity costs and cooling requirements for data centers by building them in orbit and powering them via solar energy. [00:30:07] Moving on to the applications of AI, Peter Attia recently wrote about a study out of Harvard that suggests that today's large language models perform extremely well on medical licensing exams and often arrive at the correct diagnosis yet still struggle with differential diagnosis. The listener asks what this means about the use of AI in medicine. [00:36:53] A listener asks Ken about repeated public statements by Anthropic's CEO that their AI models, and those expected in the future, are already powerful and dangerous enough to justify concern, while at the same time Anthropic and other leading AI companies are continuing to employ these systems and cautioning governments against regulation that could slow down innovation. The listener asks Ken how to resolve this contradiction. [00:45:18] A listener writes that in a previous AMA, STEM-Talk episode 184, Ken objected to the use of the word hallucination to describe errors made by large-language models and asks Ken to expound on this. [00:53:28] A listener asks Ken whether current AI models are evolving from being fundamentally prediction engines to being able to truly understand and reason, or if they are just becoming better and more convincing prediction engines. [00:59:37] A listener writes to Ken, mentioning that a few years ago, a thousand technology experts and researchers signed a letter urging AI labs and executives to pause the development of highly advanced AI systems and tools. This letter, drafted through the not-for-profit Future of Life Institute, warned that AI developers were locked in an out-of-control race to develop and deploy ever more powerful digital minds that no one, not even their creators, can understand, predict or reliably control. The listener also states that they learned recently that ChatGPT 5 was programmed using an earlier version of ChatGPT. If the concerns of tech leaders are to be taken seriously, the listener wonders if it is a good idea to have earlier versions of ChatGPT programing later versions of ChatGPT? [01:07:45] A listener asks Ken about the process of training smaller AI models from larger models, a process called distillation. The listener asks how this process works and if this is a way for the cost of training large models to benefit everyone. [01:14:56] A listener asks Ken what it means for China's Open Weight GLM 5.2 to be reported to have achieved near parity with Anthropic's Mythos. On several cyber security and vulnerability discovery benchmarks, it seems alarming that GLM 5.2 is operating at roughly 1/6th the cost and remaining openly available for download and deployment. [01:22:34] A listener writes that they recently heard a story about someone who asked an AI whether they should walk or drive to a car wash in order to get their car washed when the car wash was close by their starting location. The AI in this condition recommended that they walk instead of drive. The listener asks why the AI was unable to give the obviously correct answer to a simple common-sense question. [01:29:05] For our final question in this AMA, a listener writes that, recently, nearly 200 researchers and economists signed a statement warning that as AI becomes more powerful it could take over a large share of human work and lead to widespread joblessness. The listener wants to know if Ken would have signed the statement and what he thinks of the statement's three points: One: AI may become radically more powerful over the next 10 years. Two: This could drive an unprecedented transformation of our economy larger than the industrial revolution and over a vastly shorter timeframe, with large scale job displacement. Three: Economists, policy makers, and tech leaders must act now in order to understand the economic impacts of transformative AI and to build the incentives guardrails and institutions need to steer AI in a direction that compliments humans and benefits society. Links: Learn more about IHMC STEM-Talk homepage Ken Ford bio Ken Ford Wikipedia page Dawn Kernagis bio
Daniel Wallace is the Executive Director of Gull Lake Ministries, a Christian family ministry and retreat center in Hickory Corners, Michigan. Prior to serving over 20 years at GLM, he was the Senior Director of Camps at a camp and conference center in Texas, overseeing six separate facilities which ministered to families, senior high, junior high and grade school students. Daniel, better known as Ambush, has over 40 summers of Christian camping experience in Michigan, Texas, Missouri, and Kansas.
Daniel Wallace is the Executive Director of Gull Lake Ministries, a Christian family ministry and retreat center in Hickory Corners, Michigan. Prior to serving 20 years at GLM, he was the Senior Director of Camps at a camp and conference center in Texas, overseeing six separate facilities which ministered to families, senior high, junior high and grade school students. Daniel, better known as Ambush, has over 40 summers of Christian camping experience in Michigan, Texas, Missouri, and Kansas.
Daniel Wallace is the Executive Director of Gull Lake Ministries, a Christian family ministry and retreat center in Hickory Corners, Michigan. Prior to serving 20 years at GLM, he was the Senior Director of Camps at a camp and conference center in Texas, overseeing six separate facilities which ministered to families, senior high, junior high and grade school students. Daniel, better known as Ambush, has over 40 summers of Christian camping experience in Michigan, Texas, Missouri, and Kansas.
This week we dissect ColdCard's ~$100M RNG exploit that Claude Code cracked in 8 minutes, debate whether AI just killed open-source security and Bitcoin maximalism, tear apart Ethereum's EIP-8361 staking-yield taper, and unpack Leopold Aschenbrenner's 67% Situational Awareness blowup and CLARITY Act's ethics fight. Welcome to The Chopping Block – where crypto insiders Haseeb Qureshi, Tom Schmidt, Tarun Chitra, and Robert Leshner chop it up about the latest in crypto. No guest this week, just the four of them working through a week where AI quietly rewrote the economics of both security and human psychology, and crypto happened to be standing in the blast radius. This episode: ColdCard, NVK's Bitcoin-only hardware wallet, got drained of nearly $100M thanks to a random-number-generation bug that a one-word commit buried five years ago, and Claude Code sniffed it out in 8 minutes (an open model with no internet found it in 20, for about two bucks). The crew debates whether AI just killed open-source security, whether Nic Carter is right that this is 'the death of Bitcoin maximalism,' and why Tarun thinks maxi devs are 'the RFK of security practices.' Then they take a blowtorch to Ethereum's EIP-8361 staking-yield taper (Tarun: 'the proposal reads like shit'), unpack Leopold Aschenbrenner's 67% Situational Awareness blowup while 4x levered, and wade into the CLARITY Act's ethics fight where a single amendment is the whole ballgame. Listen to the episode on Apple Podcasts, Spotify, Pods, Fountain, Podcast Addict, Pocket Casts, Amazon Music, or on your favorite podcast platform. Show highlights
My guest today is Gavin Baker, founding partner and CIO of Atreides Management. This is our seventh conversation, and just two months after Gavin's last appearance. It's about the gap between what the market is doing and what companies are seeing. It's been a tough month or so for public AI names, but there's no sign of a slowdown on the ground in Silicon Valley. We discuss the latest moves, contracted vs. spot GPU prices, the game theory of memory supply agreements, and why Claude has become the Walter Cronkite of the stock market. We close on SpaceX, orbital compute, and what Gavin sees as the single biggest risk to all of it. Please enjoy this conversation, from the famous table at Benchmark, with my friend Gavin Baker. For the full show notes, transcript, and links to mentioned content, check out the episode page here. ----- Become a Colossus member to get our quarterly print magazine and private audio experience, including exclusive profiles and early access to select episodes. Subscribe at colossus.com/subscribe. ----- Ramp's mission is to help companies manage their spend in a way that reduces expenses and frees up time for teams to work on more valuable projects. Go to ramp.com/invest to sign up for free and get a $250 welcome bonus. ----- Trusted by thousands of businesses, Vanta continuously monitors your security posture and streamlines audits so you can win enterprise deals and build customer trust without the traditional overhead. Invest Like the Best listeners get a special offer of $1,000 off Vanta when you go to vanta.com/invest. ----- WorkOS is the infrastructure B2B and AI-native companies use to sell to enterprise. It covers everything enterprise security requires: SSO, SCIM, RBAC, Audit Logs, AI governance, and more. Trusted by 2,000+ fast-growing companies, including OpenAI, Anthropic, Cursor, and Vercel. ----- Rogo is the AI platform for finance. They're building agents for Wall Street that are trained to understand how bankers and investors actually do work: from diligence and modeling, to turning analysis into deliverables. To learn more, visit rogo.ai/invest. ----- Ridgeline has built a complete, real-time, modern operating system for investment managers. It handles trading, portfolio management, compliance, customer reporting, and much more through an all-in-one real-time cloud platform. Visit ridgeline.ai. ----- Editing and post-production work for this episode was provided by The Podcast Consultant. Timestamps: (00:00:00) Welcome to Invest Like The Best (00:02:35) First Question: July Was 2022 in a Month (00:04:08) The Private Companies Public Markets Can't See (00:05:06) Old GPUs Repricing Higher (00:06:53) Walking Through the Month (00:08:22) Kimi, GLM 5.2 & the Open Source Freak-Out (00:10:51) Real Yields, Spreads & CDS (00:11:54) Does the Build-Out Need Credit? (00:15:22) A Sell-Off With No Clear Villain (00:17:35) Open Source as Dark Matter (00:18:39) Nvidia's Lowest Forward PE in 10 Years (00:21:35) Claude as Walter Cronkite for the Stock Market (00:23:55) Continual Learning & Sample Efficiency (00:25:19) What Would Actually Scare Him (00:26:38) Routers & the Multi-Model Future (00:30:51) Tokens as a Percent of Comp Spend (00:33:37) The Game Theory of Breaking an LTA (00:36:41) Nvidia's Credit Wrapper & Revenue Share (00:37:45) What He'd Do If He Ran Hynix (00:41:46) Who's More Bullish than Him (00:43:28) China's DUV Machine (00:46:10) Bull Case for Software (00:48:16) The RSI Maximalist View (00:49:31) Inference Clouds Growing Without Burning Cash (00:50:35) The Biggest Risk Is Regulation (00:53:44) Telling the Story Better (00:57:15) Dark Horses (00:58:02) SpaceX in the Public Markets
OpenAI has a new model coming soon called Astra. Was it a leak? A reddit post? Some backdoor update? Nope, OpenAI made some crazy discoveries and math then told the world that their next model family Astra did the heavy lifting. (And you thought you could just click ‘Sol' and your strategy was set for Q3?) Aside from news on what's next from OpenAI, this week saw multiple new agent outbreaks, AI competitors banning together to pace AI, Amazon doing a 180 on its AI strategy and a lot more. Don't get left behind. We'll keep you ahead. OpenAI's new Astra model, more AI agents escape sandboxes, AI leaders call for AI pacing and more. AI News That Matters for August 3 — An Everyday AI Chat with Jordan WilsonNewsletter: Sign up for our free daily newsletterMore on this Episode: Episode PageToday's Episode on LinkedIn: Thoughts on this? Join the convo on LinkedIn and connect with other AI leaders.Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineupWebsite: YourEverydayAI.comEmail The Show: info@youreverydayai.comConnect with Jordan on LinkedInTopics Covered in This Episode:OpenAI Agents Escape Sandboxes IncidentAnthropic Claude Models Security BreachesAI Agents Breaking Cybersecurity GuardrailsOpenAI GPT-5.6 Price Cuts & Self-OptimizationRecursive Self-Improvement in AI ModelsAI Leaders Urge AI Development PacingUS, China, and International AI GovernanceAmazon Nova AI Models Shutdown StrategyOpenAI Astra Model Math BreakthroughNew AI Models: Fable, Astra, DeepSeek v4 FlashEnterprise AI Agents and Cybersecurity UpdatesGoogle Gemini Robotics, Music, and Agent ReleasesMeta, Microsoft, and AWS AI Infrastructure MovesOpenAI Free Frontier Tools for ResearchersBlock's Buzz Open Source AI Workspace LaunchTimestamps:00:00 OpenAI agent containment issues04:27 Anthropic data breach explanation07:28 Evaluating AI incidents and responses10:14 OpenAI slashes GPT 5.6 prices15:59 AI industry urges development pause17:45 Concerns about AI self-improvement22:36 Amazon shifts AI strategy25:04 Amazon's AI efforts discussion28:00 OpenAI's Astra and new math proofs30:36 OpenAI's new four-tier system36:14 Google's Lyria 3.5 and Block's Buzz36:48 Latest AI developments overviewKeywords: Astra model, OpenAI, AI agents, agent escape, sandbox containment, autonomous AI, Hugging Face breach, Anthropic, Claude AI, cybersecurity testing, unauthorized access, model capabilities, recursive self-improvement, GPT-5.6, price cut, Luna model, Terra model, Sol model, input tokens, output tokens, AI infrastructure optimization, self-improving models, benchmarking, SONNET-5, large language models, artificial analysis index, codex, academic research, AI oversight, industry pause, AI governance, national security, China open-source models, Frontier Labs, Amazon Nova, AGI Lab, AWS, Peter DeSantis, Peter Abbeel, media coverage, Fable model, Haiku, Opus, DeepSeek, Kimi K3, Quinn 3.8, GLM 5.2, Google Gemini 3.5, Microsoft Copilot, cybersecurity vulnerabilities, distillation, model overhang, artificial intelligence development, international AI regulation, generative AI, model benchmarking, Sora video model, El Paso data center, MCP update, MAI Cyber One Flash, Project Perception, Lyria 3.5, music generation, Buzz open source, Block, Meta AI, Chrome Gemini integration, Gemini Spark, product summary algorithms, Rufus, enterprise AI, stateless core, model scaling, advanced math problems, sphere packing, federal policy, voluntary AI commitments.Send Everyday AI and Jordan a text message. (We can't reply back unless you leave contact info) Ready for ROI on GenAI? Go to youreverydayai.com/partner
Watch the full episode on YouTube:We first covered Baseten last year when DeepSeek mania was at peak hype. Now they have raised a monster $13B round and become one of the new cohort of AI Infra decacorns that are (with Nvidia, Intel, and the semis complex) chief beneficiaries of the Inference Inflection. We return to Baseten at the peak of the 2026 edition of Open Weights debate. Ali has published a viral breakdown of Kimi K3:And since you last saw him, Philip has spoken at AI Engineer and written the definitive book on Inference Engineering spotted all over SF:Three years ago, inference engineering barely existed as a category.Today, it is one of the most critical disciplines in AI. Inference engineering inherently tackles a different question than standard model training: “How do you turn those weights from training into a product that is fast, reliable, and affordable at scale?” Focusing on these creates an entirely new optimization problem.In one recent GLM-5.2 experiment, quantizing more of the model actually preserved its benchmark quality while increasing throughput by 20%, because the errors introduced in different layers could cancel each other out.Inference is no longer just the final step after training. It is becoming its own engineering discipline, with its own research problems, infrastructure, and increasingly specialized roles.In this episode, Baseten's Philip Kiely and Ali Taha join swyx and Vibhu to explain what actually happens after a new open model is released and what it takes to turn “we generated a token” into a fast, reliable, production-ready API.We go deep on cache-aware routing, disaggregated prefill and decode, quantization, speculative decoding, KV-cache movement, model parallelism, GPU kernels, and the race to make frontier models up to 10× faster. Philip and Ali explain why inference optimizations can still produce gains of 20%, 100%, or even 200%; how quantization errors can cancel one another out; why identical weights can behave differently across clusters; and how Baseten grafted a Kimi vision encoder onto GLM-5.2 without changing the underlying language model.The conversation then expands beyond LLMs into NVIDIA Dynamo, mega kernels, Rubin, AI-specific chips, local inference, video generation, diffusion versus autoregressive models, and the enormous compute barrier to generating coherent long-form video. Finally, we explore the convergence of training and inference, continual learning through persistent KV cache, and the emerging loop where models help optimize the infrastructure that runs them.We discuss:* What happens when a 200,000-token request enters an inference system* Cache-aware routing and reusing previously computed KV cache* Why prefill and decode are increasingly handled by different GPUs* When dedicated deployments become cheaper and more reliable than shared APIs* How speculative decoding uses a smaller model to accelerate a larger one* Tool calling, structured outputs, and what LLMs actually do* What it takes to support a new open model on day zero* Grafting Kimi's vision encoder onto GLM-5.2* Retrofitting inefficient model layers with components from other architectures* Why models sometimes collapse into repeating the same token* How hardware, kernels, and race conditions create nondeterministic failures* Preserving model fidelity while making inference faster* How quantization errors can cancel each other out* Why inference optimizations still deliver gains of 20%, 100%, and 200%* How optimized serving can make a model up to 10× faster* NVIDIA Dynamo, KV-aware routing, and distributed model serving* Speculative decoding the speculative decoder* Why local AI is about making models less dumb while data-center AI is about making them less slow* Tensor, expert, and pipeline parallelism across GPUs* Hardware-aware model design, auto-tuning, and the case against mega kernels* Rubin and why inference is becoming a systems problem* Whether modern GPUs are evolving into programmable AI ASICs* Why enormous models like Kimi K3 require GB300-class hardware* Why open-source video generation still trails Veo, Kling, and other closed models* The quadratic attention bottleneck behind long-form AI video* Autoregressive video, real-time generation, and compounding quality drift* Why future video systems may combine autoregressive and diffusion architectures* Training for inference and inference for training* Continuous post-training, deployment, evaluation, and improvement loops* How GLM-5.2 helped optimize the kernels serving GLM-5.2 itself* Why faster networking could unlock dramatically faster decoding* Continual learning, KV-cache compaction, and persistent model memoryShow Notes* How to build a day-0 API for Kimi K3* 22580: From GPT2 to Kimi3, ExplainedPhilip Kiely* LinkedIn: https://www.linkedin.com/in/philipkiely* X: https://x.com/philipkiely* Inference Engineering: https://www.baseten.co/inference-engineering/Ali Taha* LinkedIn: https://www.linkedin.com/in/aliestaha/* X: https://x.com/waterloointernTimestamps00:00:00 Introduction and the 200K-Token Prompt00:03:18 Dedicated Deployments, Speculative Decoding, and Tool Calling00:11:26 Launching Production-Ready Open Models00:19:06 Model Retrofits, Failure Modes, and Nondeterminism00:28:22 Quantization and Canceling Errors00:32:15 The Race to 10× Faster Inference00:40:48 Dynamo, Speculation, and Local vs. Data-Center AI00:50:18 Model Parallelism, Auto-Tuning, and Mega Kernels01:00:55 Rubin, GPUs vs. ASICs, and Custom AI Chips01:10:03 Giant Models and the Limits of GPU Memory01:12:42 AI Video, Quadratic Attention, and Autoregressive Generation01:21:47 Audio, Images, and Diffusion Models01:27:32 Training, Self-Optimizing Models, and Continual Learning01:40:06 Closing ThoughtsTranscriptIntroduction: Baseten, Waterloo Intern, and Inference EngineeringSwyx [00:00:00]: Okay, we're here in the studio with Philip, old friend from Inference Engineering, the book, as well as Baseten and everything that you've done, you and I have done before, as well as Ali. Welcome.Ali [00:00:15]: Pleasure to meet you.Swyx [00:00:15]: Waterloo intern.Ali [00:00:16]: Waterloo intern, always.Swyx [00:00:17]: When did you get “Waterloo intern” as a handle?Ali [00:00:19]: As a handle? Oh.Ali [00:00:20]: I think the rebranding happened mid-March. When I saw it was open, I was like, “I have to take it. Up for grabs.”Philip [00:00:26]: The problem is that Ali is really good at his job and is not gonna be an intern much longer.Philip [00:00:30]: So we have to figure out who's gonna get the handle.Ali [00:00:33]: Well, I'll pass the torch over to the next intern.Swyx [00:00:34]: Oh, okay. It can be, like, you just pass it to another Waterloo grad.Ali [00:00:37]: To another Waterloo intern. No, bruh.Philip [00:00:39]: Yeah.Ali [00:00:39]: Intern.Swyx [00:00:40]: Intern, yeah.Ali [00:00:40]: And no.Philip [00:00:41]: You gotta get an intern from Waterloo.Ali [00:00:42]: Yeah, I've gotta get an intern from Waterloo.Swyx [00:00:44]: Right.Ali [00:00:44]: But they have to follow the path.Swyx [00:00:45]: Oh, it could, but it could come from Baseten, so it's like whoever Baseten gets from Waterloo.Ali [00:00:48]: Right.Swyx [00:00:49]: Has the title of Waterloo.Ali [00:00:50]: It stays in the ecosystem.Philip [00:00:51]: Exactly.Ali [00:00:52]: Halfway through the internship, you either get it or you're out.Philip [00:00:55]: You should also do, like, a big graduation ceremony where you change the handle.Ali [00:00:59]: Just say it.Philip [00:00:59]: For everybody.Swyx [00:01:00]: You guys are good at ceremonies, clearly. We had a nice launch of the book, very successful. But before we get into all that, I wanna start off with a fun question for you. Okay, you're an expert inference engineer. What happens when I send a long query, say two hundred thousand tokens into Baseten's inference? What's the process of query through GPU model routing, balancing, all that? What is all the stuff that we don't think about?Long Context Requests, KV Cache, and Cache-Aware RoutingPhilip [00:01:26]: With a long query specifically, the first thing that I'm gonna ask is, “Have you sent me this query before, or at least part of it?” and I really hope you have, because it's gonna be a lot easier for me and a lot cheaper for you. So the first thing that we're gonna look at is some cache-aware routing, where we're going to see, we probably have a number of instances, a number of replicas up serving whatever model you're hitting. We want to send this one to something with, number one, available prefill workers, and number two, ideally some cached input already there so that we can skip prefill on at least part of these two hundred thousand tokens. If you're doing two hundred thousand tokens, it's probably coding or a multi-turn agent or something where you would expect to have that cached. If you don't, we're gonna have to send it to a prefill worker. We've at least on certain models disaggregated prefill and decode, so you're going to have one set of GPUs that's solely going to process the input, create the KV cache, and get you your first token, and then that's going to be passed over to a separate set of GPUs, which is going to run decode. We're going to iteratively make those tokens. We're probably going to have some speculator model in front of that. I'm going to assume that you're doing coding, and because of that, our speculator model, which assumes you're doing coding, is gonna have a high draft token acceptance rate. If I'm wrong and you're asking me to summarize every Harry Potter book, it's gonna be slower. And then we stream that output to you and account for it, charge you, a couple of pennies and say, “Hey, would you like to send another one?”Swyx [00:03:04]: Except Baseten doesn't charge by pennies.Philip [00:03:07]: Well, yeah, we charge. I'm assuming that we're talking about the public model APIs. If you are setting up a dedicated deployment, then yeah, it's not pennies.Public APIs vs. Dedicated DeploymentsSwyx [00:03:18]: Yeah, one of the key differentiators when I was talking with Baseten initially was that people who want very high volume just need to rent by the box, ‘cause then it's up to you to figure out how to saturate the box.Ali [00:03:31]: And more often than not, it's, like, way cheaper if you're pushing, like, millions of tokens per hour, if you just pay per hour instead of pay per token.Philip [00:03:37]: Yeah, they do. I think that we've increasingly seen a lot of demand for the pay per token APIs, just because everyone wants to try open models, and then once they find a use case that's really sticky, then they move over to dedicated.Swyx [00:03:51]: Is there a best practice on when it's time to swap over?Philip [00:03:54]: Couple reasons. Yeah, reliability, that's a big one, right?Ali [00:03:57]: Like, if they have a very specific use case, they want you to train something specifically for them, like they want their own spec dec, for instance, for their own traffic.Swyx [00:04:04]: Spec dec is speculative decoding.Speculative Decoding and Custom SpeculatorsAli [00:04:05]: Speculative decoding, yeah.Swyx [00:04:07]: You have to explain.Ali [00:04:07]: Sorry. Like, speculative decoding is like, if you have a huge model, right? And so the model is going to be generating one token at a time every single turn, every single forward pass. So we attach, like, this little, like, parasite, like this layer that goes on top of the model, and this model just has to predict. It does three very fast autoregressive forward passes, and it will predict, like, three certain tokens, and then you do one forward stage over the entire original model in order to see if those predictions were correct or not, and then you accept them or you reject them. Now, this draft model is traffic specific, so if you, like, Philip said, if you're summarizing Harry Potter books, I can train exclusively that draft model on Harry Potter books, and I can guarantee you that I'm gonna accept the three tokens every single time. And so with that case, I increase your decode speed. I wouldn't be able to provide this to you if you're a shared endpointSwyx [00:04:53]: YeahAli [00:04:53]: ‘cause I have no idea if you're doing Harry Potter, if you're doing coding, if you're doing English. We don't know. Also, there was a thing in the book that mentioned that if they really cared about a specific threshold, chapter four, I think. Do you remember that?Philip [00:05:06]: Yeah. The things that you can do is you can set a specific, like, batch sizing, a specific, like, parallelism strategy if you're trying to optimize for, like, throughput versus latency. You can. Maybe a NVFP4 quant doesn't pass your benchmarks and you wanna run a model at higher precision, you could do that. There's just a bunch of reasons why you might wanna have your own endpoint and the biggest one, of course, just being, like, you don't have to deal with someone else doing a hundred million tokens of benchmarking traffic at the endpoint when you happen to be trying to serve your users.Swyx [00:05:40]: Yeah. I think one thing that is. That is a classic journey. Like, it's people is asking the, what happens when you type Google into the browser. Tool calling, is that just, you're generating JSON or is there more complication beyond that?Tool Calling, JSON, and Structured OutputsAli [00:05:58]: Certain customers that we have, they have their own post-trained models, and so they demand a tool calling that's not just, like parse a file or go find the weather. It's something that's very specific and you have to do post-training on this. And if the post-training on the model is not good or if the quantization after the post-training to get the inference to be fast, the model will struggle reading the JSON file and reading the tool calling. But it doesn't require its own like sandbox. It's not like it's going to use that tool calling to like escape a sandbox or like it doesn't have to be contained. It can just be a normal dedicated deployment. The challenge with tool calling more and more seems to be that the companies want certain tool calling which is a very sensitive thing to train. And because you're dealing with all of the JSON outputs, if it doesn't like close the end of the request in a very certain manner, you end up with a model that did the tool calling and like the thinking and so as a result of that, it didn't see the result and just hallucinated the result as it decoded. That seems to be the most challenging thing with tool calling, not really the sandboxes model.Philip [00:06:56]: Yeah, that's a challenge on the training side and then on the inference side, there's work that you can do to scope the possible output. So we published this at this point close to two years ago, the solution to this problem which is you make a state machine and you use that to constrain the output to a specific format. So this is the structured output problem. If you remember backSwyx [00:07:27]: Yeah, the specific grammar is,Philip [00:07:29]: Yeah, exactlySwyx [00:07:30]: GML had this thing.Philip [00:07:31]: Yeah. So it's like the old-school “make sure this is only JSON”, return only JSON orSwyx [00:07:38]: YeahPhilip [00:07:38]: Grandma's gonna die type of prompts.Swyx [00:07:39]: Is it BNF grammar? At some point OpenAI had released a thing that was like, yeah, if you want to constrain your output, write BNF grammar, back as NOR.Philip [00:07:47]: In our inference system, it's just a specified output format. And you get the guarantee that your output's gonna be structured along that format. And so applying that to tool calls can like help cut down on. You can still call the wrong tool or call no tool. It doesn't solve the certainty problem but it at least solves the output structuring problemSwyx [00:08:10]: YeahPhilip [00:08:10]: Within tool calls.Swyx [00:08:12]: And MCP is just another form of tool, right.Philip [00:08:14]: Yeah, exactly.Swyx [00:08:15]: As far as there's no special thing there.Philip [00:08:16]: The thing I'm always like explaining to people is the LLM is not capable of doing anything. It's only capable of making suggestions of what to do and then if those suggestions are formatted in a certain way and applied to a system that knows what to do with them, then an action occurs.Swyx [00:08:32]: Yeah. Part of the fun stuff is, this is solved outside of tool calling too. Like in an agent loop if the output is not correct or you're right, like reasoning, tool calling was done in the reasoning trace, just be like, “Oh, I don't know what to do. Let me just try again.” And it might get there after a few tries. And on your point of training, sometimes this is harder in smaller models, so you don't have the same exact quality outputAli [00:08:56]: Right.Swyx [00:08:57]: When you just swap from a big model, right?Ali [00:08:59]: Yeah. I will say that, before, I think we need to go back to inference engineering proper.Ali [00:09:04]: But, I had expected that something would replace JSON because it's hard to stream JSON ‘cause JSON must be complete and you must have open and close brackets and everything. So it's hard to parse something or validate something while it's being streamed. So people invented all sorts of things that are like, I forget the name of some of these alternatives, but it's something like TOML, something like YAML. But JSON seems to be dominant still.Philip [00:09:30]: The JSON outputs aren't that long, right? Like you could have a long-- ‘cause tool calls also contain the arguments in them and perhaps for a certain tool you might pass like a very long argument. But my impression of the median tool call is that it's a relatively small number of tokens, right? So I would expect that speculators are generally fairly good at something as formatted as JSON. And so you would have like a pretty fast decode step there and that the streaming wouldn't be as valuable, but maybe I'm wrong about that.Ali [00:10:02]: I think you're also bounded by the software or that the model is gonna integrate with if the software is built with JSON for the tool calls or if the company that you'- if your customer says that this is how our software works and our tools are interfaced with JSON, you can ask them to like, change their software and say like, “Yeah, this is gonna be better for the model.” but like with the right training shouldn't be that much of a difference. Also more profitable if it outputs more tokens probably.Swyx [00:10:25]: Depends on your business model.Swyx [00:10:27]: It really depends. But I will say that, as a writer with like experience a lot with generated output, I do try to move from text to JSON text which is very long JSON, right? Like there's paragraphs in every field because I'm trying to structure it, right?Philip [00:10:44]: Right.Swyx [00:10:44]: I want you to first make factual statements, then make opinions then make bullet point summaries, have dates, have entity references have your sources for references, all these things. Anyway, so these are things that like I think people who really experiment with structural output have to really care about. But, let's, let's recurse up the stack a little bit. Before we started recording, you mentioned something really cool, which is that there's a lot of engineering that-- inference engineering that goes on when a new model provider releases a new model, right? So let's call it GLM-5.2, Kimi K3. I had previously assumed, especially if it's like, well, GLM 5 to 5.1 to GLM-5.2, like that you've supported them before. Is it that much work?What It Takes to Support a New Open ModelAli [00:11:26]: It's a lot of work.Swyx [00:11:28]: Yeah. Okay. So like, a lot of people, all you guys, right whenever a new model launch like, people rush to say like, “Oh, Hugging Face supports this, Fireworks supports this, Spacetime supports this,” and I'm like, “Yeah, of course we support it.” But what goes into that? What goes intoPhilip [00:11:40]: I think it's more than just support it too, right? It benefits the consumer a lot. Like I think it was with Kimi K2.5 or GLM-5.2 the latest, there was an inference war, right? X provider is at 90 tokens a second. The next day we're at 150. The nextSwyx [00:11:55]: I kinda kicked that off with the GLM-5.2.Swyx [00:11:58]: I wrote a Twitter article about. It got like half a million views,Ali [00:12:02]: Based on being numberSwyx [00:12:03]: YeahAli [00:12:04]: Or it's for something else.Swyx [00:12:05]: Yeah. Which,Ali [00:12:06]: Oh my GodSwyx [00:12:07]: Which then got everyone really excited about, hey, how can we, bend tracks a little bit further and,Philip [00:12:14]: There's a difference between support the model, as in I can make a token out of this model, and support a model, as in I have a production-ready API from this model.Philip [00:12:26]: Getting to the point of I can make a token out of this model is not that hard because generally the, open source inference engines, vLLM, SGLang of the world oftentimes even receive weights ahead of time, maintainers do, or the people making the model merge PRs to ensure support. So you generally can, just get it working on the standard open source stack without too much pain in most cases. The challenge is, every inference company is gonna have own proprietary stack. Some open source components, some in-house stuff. And for any arbitrary model, there's going to be some new stuff. Sometimes you get lucky, like K, two five to two six was, like, pretty similar.Quantization, Speculators, and Production ReadinessAli [00:13:16]: Yeah. It was pure continued post-trainingPhilip [00:13:18]: YeahAli [00:13:18]: If I remember correctly.Philip [00:13:19]: Even in those cases, there's still stuff you have to do. You have to redo the quantization work. You're taking the model from. Generally, these models are not released in NVFP4, and we want them to be in NVFP4 for maximum Blackwell compatibility. So we have to perform that quantization, and, calibrate the quantization to make sure that we're not causing any regression in the model's intelligence. And then we also have to train the speculator, as we've talked about. Generally, we have. We have ZDR, zero data retention on our model APIs, so we don't know exactly the traffic that people are sending us, but we know what's popular. We know that coding use cases are popular. We know that agents, agentic use cases are popular. So we can get public data sets that are representative of that traffic and train general speculators. Now, with speculators today, you need to train the speculator using the base model itself because you're getting hidden states out of the model from running inference on these specific prompts, and that is the training data you use to create the speculator. So there's that process which you need the real model weights for. And then there's of course just the process of, standing up all the infrastructure behind it, loading all this stuff, testing it. And then when there's a new model with a newer architecture, I think that, like, the DeepSeek models tend to be the most challenging as they have, like, the most novel architectural stuff going on, model after model. But every new model has something. Kimi K2 had. Oh, sorry, GLM-5.2 hadAli [00:14:53]: Sparse attention.Philip [00:14:54]: Yeah,Ali [00:14:54]: YeahPhilip [00:14:54]: the DSA.Ali [00:14:55]: Right. Which is brought from DeepSeek.Philip [00:14:57]: Yeah. AndAli [00:14:59]: So you can copy-paste then?Philip [00:15:01]: It kindAli [00:15:01]: I don't know how this works.Philip [00:15:02]: So, like we had to, like, build support for that into our runtime. And you're right, like it is really interesting the way that all of these open source labs borrow from each other. For example, like GLM-5.2 doesn't have vision. So something that, Haley, a guy on our team, if we could take a look at this, he, like, grafted the Kimi vision encoder onto GLM-5.2.Retrofitting Vision into GLM-5.2Ali [00:15:27]: We'll be training the projector.Philip [00:15:28]: Exactly. So if you think about, like, the encoder, there's the encoder, which is the part that looks at the image and turns it into latent information, and then there's the projector which likeAli [00:15:38]: You can say latent space. It's okay.Philip [00:15:41]: And then there's the projector that maps it onto, the model itself, and then there's the model weights. You don't wanna mess with the model weights because you run a chance of making the model dumber at something else for the purpose of giving it vision. So instead, Haley started with just a projector, which is only a handful of millions of parameters.Ali [00:16:02]: That would be, yeah.Philip [00:16:02]: Yeah.Ali [00:16:03]: Can you show the training one?Ali [00:16:04]: Like the way it groksPhilip [00:16:05]: YeahAli [00:16:06]: Very interesting.Philip [00:16:06]: And maybeAli [00:16:07]: That right therePhilip [00:16:07]: Maybe Ali, you should take it from here. You've got a betterAli [00:16:10]: Ooh, double the sandPhilip [00:16:11]: Understanding of this than I do.Ali [00:16:11]: Yeah. You can see, like, he. The way he trained this is really cool. At the beginning, he was training it using just like, “Here's a picture of a mountain. Can you describe what's in this mountain?” And that caused it just like the first, learning walls. Like here you can see this all we're trying to teach it is to translate the encoded. Like it's already taken the encoder from Kimi K. It's taken the image. It'Philip [00:16:31]: Yeah. FrozenAli [00:16:31]: FrozenPhilip [00:16:32]: With adapter.Ali [00:16:32]: Exactly.Philip [00:16:33]: Yeah.Ali [00:16:33]: So the brain is frozen and the eyes are frozen. It's just we're tryingPhilip [00:16:37]: AlignAli [00:16:38]: Interconnect between the eye and the brain, right? So the projector. And so you take the tokens and then he's like, “Oh, can you describe what's in this image?” And he's like, “Oh, it's a mountain,” or it's a person or it's a human, whatever the case is. But that didn't cause complete understanding. So he changed it such that every image was associated with a data set of questions. Like, does this image have a white male? Does this image have birds in the top corner? Does this image have a scientist in it? All of that stuff. And it would have to answer questions correctly. And using not just training on describing an image, but being able to answer question, another question, answer over time. Like you can see the grokking, which is like genuinely insane, that retrofitting vision into a large LLM can learn to that extent. And even for images that it doesn't perform well on, for instance, if you ask it a picture of like Stephen Hawking, “Who is this?” Maybe it doesn't get it, but it will say something like, “This is Albert Einstein.” Like it still understandsPhilip [00:17:25]: Close enoughAli [00:17:26]: That this is a scientist who is a man who has, some significant achievements, all that stuff. So that's like really cool.Philip [00:17:32]: Yeah. So, we've covered Hao Tian before, who the author of the LLaVA paper that did this, a while ago. And I think that's very foundational work for anyone who hasn't done vision work before.Ali [00:17:41]: Same with the CLIP and MetaCLIP, where you go from just captioning to building out questionsPhilip [00:17:47]: RightAli [00:17:47]: Off the image and how much better you can get performance.Philip [00:17:50]: Right. Right. Right. Yeah. But what's, what's so exciting about this is if you look at a model like this. Now, this is a little bit more of a research project. It's not. It got to 56% on MMLU Pro, I think. So not quite frontier. But if you're running this model, you haven't suffered any loss on your GLM-5.2 quality. If you don't have an image, it'll just behave exactly the way it used to. And ultimatelyAli [00:18:14]: Which in the inference code you literally do not include the other part, right?Philip [00:18:18]: Yeah. You would just skip the encoder if you don't have an image input.Ali [00:18:22]: Okay.Philip [00:18:22]: Just confirming.Philip [00:18:23]: YeahAli [00:18:23]: Does it affect a lot on the overall inference side? Like you're not adding much, you're adding a very small vision encoder. These are typically likePhilip [00:18:30]: They're super fineAli [00:18:31]: Less than a billion parameters, right?Philip [00:18:32]: Yeah. It's, - There's a little bit less standardization among vision encodersSwyx [00:18:37]: YeahPhilip [00:18:37]: So the support matrix can be a little bit, sparser. But overall, yeah, it's a pretty, it's a pretty minor component of the overall system. And ultimately what you get out of the system is all of a sudden you have Kimi Vision, GLM weights, and DeepSeek attention all in one model.Open Source Model Grafting and Franken-MergesPhilip [00:18:56]: And that's, I think, a lot of the power and beauty of open source, is that you can take all of these different components and combine them together into a system that's better than anyoneSwyx [00:19:05]: YeahPhilip [00:19:05]: Can be individually.Swyx [00:19:06]: People used to say that you would also do Franken-merges where you would take likePhilip [00:19:10]: YeahSwyx [00:19:10]: Layers from each model.Swyx [00:19:11]: Does anyone do that anymore?Ali [00:19:13]: Well, to your point previously when you were mentioning like, the work that goes into supporting a model when it first comes out, like GLM-5.2 or MiniMax M3 or whatever the case is. Sometimes you do have to like, you do have to switch out some things. Like, for instance, the MiniMax M3 head uses full attention, and with full attention you end up with this like insane bottleneck in spec dec ‘cause you're doing auto-regressive token generation for three tokens, and you're doing this like N squared over all of the tokens that are in your sequence. Your KV cache is like very large because it's not sparse, it's not top K. So we find it better to like, okay, we're gonna replace this, we're gonna replace this layer with a layer from another model that's using like GQA, for instance. And then just with the right training, you can get it to have the same acceptance rate. So it is very possible to retrofit layers from other models and very much needed. If a layer is like inefficient, the training just becomes the challenge, like how do you ensure that you train it properly? Which again to your earlier point is like the mesh between training and inference. As in like you need very good training in order to do fast inference. That's like, I feel like more and more becoming true.Swyx [00:20:21]: Yeah. Anything else on the support side when you say like get it to fully production ready?Loop Detection, Race Conditions, and Non-DeterminismPhilip [00:20:26]: Yeah. I think that there's also a question of just, we can test a model to a pretty extensive degree, but we're trying to get it out quickly and then you see a bunch of other people test it and you get interesting results. There was an issue with, GLM briefly where we had some like mode collapses where it would just output the same token over and over again for certain prompts on certain temperatures. Like once you expose an endpoint to the real world, there's going to be, so many more varieties of things given to it that you're able to, discover and patch things. So it's not just a, day zero process, it's then like for the first week, for the first month, if a model remains popular, like how do you both fix bugs and then continue to push the envelope on performance?Ali [00:21:21]: What do you mean you don't want your model outputting S?Swyx [00:21:24]: Is there loop detection on that stuff, by the way? It still happens like quite a lot, which is surprising.Ali [00:21:30]: We have like we, in our endpoint, like if a model was to output the same token like four plus times, we just cut the generation. We say like, “Oh, sorry, this-- Like try again,” or like we will reprocess the request. ‘Cause we know then, like if it, like if, yeah, it's four times the same token, it's probably collapsed.Swyx [00:21:45]: Yeah. Is there a way to opt out in case I really want that?Ali [00:21:48]: You want that?Ali [00:21:50]: I think there's a way that we have to handle it. I'm not exactly certain, but I feel like in certain models, like when they output something like you can imagine, like a table for instance, and so they want, they wanna draw like 12 dashes and 12 dashes. Yeah, I think there's a way for that to happen. I think we only do it on certain tokens. Like we exclude certain special characters.Swyx [00:22:07]: Yeah.Ali [00:22:07]: So we only do it on like certain like S is the most common almost. GLM-5.2Swyx [00:22:11]: OhAli [00:22:11]: And I think it was DSV 4 as well. Like you'd just have like looping issues where like you literallySwyx [00:22:17]: ItAli [00:22:17]: Just have like S.Swyx [00:22:18]: Yeah. Is there a special, something special about S? No, just randomlyAli [00:22:21]: It just seems to be the one token involved.Swyx [00:22:23]: Yeah. And it'Philip [00:22:24]: Is thereSwyx [00:22:24]: And it's only temperature 0Ali [00:22:27]: NoSwyx [00:22:27]: Even at other temperaturesAli [00:22:27]: Even at like 0.9 or whatever, it will still, it will still collapse.Swyx [00:22:30]: That's weird, right?Ali [00:22:30]: It's, it is an inference problem to be honest, like a software problem. Like oftentimes, the image you run will-- like NVIDIA will release an image for instance, and if we will upstream the changes from their latest TensorRT-LLM image into our stack, we'll find that it fixes it. Or oftentimes this will only happen in an inference engine that you're using like SGLang. But if you were to switch to vLLM, that isn't the case. So it seems to be like an extremely like deterministic software issue and not really a model issue. It's not like a weights problem. Like I'- we'll say like, “Oh, it's a problem with the quant. We did PTQ wrong,” right? But that isn't, that doesn't make sense because the same weights used with a different inference engine does not repeat the problem. And sometimes it's, the kernels that are being used in the backend have like these very subtle sometimes race conditions, where if you were to use this model hosted on one cluster, you will never get this problem.Swyx [00:23:19]: Oh my God.Ali [00:23:19]: But if you host it on a different cluster, you will. And the reason is the KV cache transfer from a node to node in that one cluster is using a slower interconnect than the node to node in another cluster. So that exposes the race, whereas in another cluster it doesn't. So then you end up just like, okay, this model is not gonna be hosted on this cluster. We're gonna host it on, another cluster because that cluster exposed that problem. But then it ends up with like, okay, is it the software? Is it the model weights or is it the hardware?Swyx [00:23:42]: There is a thing about this with temperature 0 still not being deterministic, right?Ali [00:23:46]: Right.Swyx [00:23:46]: Mostly because of hardware. Even at temperature 0 same model, you won't always get the same output.Swyx [00:23:52]: Even-- But I'm surprised by the race condition one because, I thought PyTorch was a graph that like guarantees that you at least, execute things in the right order.Ali [00:24:02]: Well, yeah, true. Like I'm not, I'm not saying that there is. Like well, you have things like PTL optimizations where like you can start a kernel before the end of the previous kernel, and that's like ‘cause you want to do that because there'sSwyx [00:24:12]: It's like pipeliningAli [00:24:12]: Expense. Exactly.Swyx [00:24:13]: Yeah.Ali [00:24:13]: But it'- But you don't do it cleanly. Like you overlap a little bit of the execution. No, it is very possible that the kernel itself, like that one block that is supposed to be running in this instance of time, that kernel itself has a race condition. For instance, like a missing barrier. Like often if you're designing a kernel and you want it to make it to be very fast, if you don't test it extensively, you'll, you'll have certain threads access data points from registers before they've been written to by other threadsSwyx [00:24:36]: YeahAli [00:24:36]: For example, because like your barrier is wrong or your synchronization was wrong. But yeah, like the testing itself is very difficult in those like, andSwyx [00:24:42]: And there's no like borrow checkerAli [00:24:45]: What does that mean?Swyx [00:24:46]: Like Rust. Like the. If you're trying to have like memory safety It sounds like a comparable problem.Ali [00:24:52]: Well, yes, but you're working in CUDA, right, NVIDIA GPUs. Like- You just need a higher level language like modular Maybe that's what modular is supposed to do. I don't know.Quantization Quality and Vendor FidelityVibhu [00:25:00]: How do you see keeping quality of the model? So you talked about all these steps of, okay, you gotta do quantization, train your own speculative decoderAli [00:25:07]: RightVibhu [00:25:07]: Run on different hardware. Looking at other model providers, okay, you kicked off a inference speed race on the consumer end. What goes into keeping quality the same across them, right? Sure, you can run benchmarksAli [00:25:22]: YeahVibhu [00:25:22]: But, like, how do you determine how much quantization are there standards? What goes intoPhilip [00:25:27]: There's a few things on quality. Most inference optimizations are lossless. KV caching, for example. You are just recomputing or preventing recomputing the same values. Speculation, of course, if a draft token is wrong, it gets rejected. The main lossy optimization is quantization. And that really comes down to, number one, data format, number two, which parts of the model you choose to quantize, which layers, and number three, like doing a lot of calibration on the quantized weights, to ensure that you're preserving all the outliers. There's other tricks that you can do, though. A big one is long context, ‘cause one thing you asked at, right at the beginning is, “Oh, what's gonna happen if I send a 200,000 token request in?” So with a long input sequence, you need to, store a lot more information. You need to process a lot more tokens. And so even if a model has a context of a certain length, you might, as an inference provider, choose to build an API with a shorter context length, and of course a full length one as well. Because if someone doesn't need the full million token context, for example, you can get them better performance. I don't know if that's exactly like quality of the model. The way that I think about quality is to what degree are we faithfully serving the original model? If you think of a golden implementation of a model that performs exactly the way the model is designed to perform, I think of quality as how close are we getting to that, 100% fidelity of the model.Philip [00:27:13]: You can also, of course, think about quality from the training side and how do you push yourself past 100%. But when I think about purely inference optimizations, it's getting faster while staying as close to that 100% fidelity mark as possible. And certainly our standard internally is that, like you should not be able to tell the difference between our API and a, official API. I think Kimi in particular does a good job of vendor benchmarking hereAli [00:27:41]: YesPhilip [00:27:41]: Where they haveAli [00:27:42]: They released an actual vendor benchmark.Philip [00:27:43]: Exactly, yeah.Ali [00:27:44]: ‘Cause they accused, some people, Amazon? There was some provider that was not doing very well on Kimi's benchmark.Philip [00:27:50]: Yeah.Philip [00:27:51]: So, with Reflect we probablyVibhu [00:27:52]: This was a long time ago, right?Philip [00:27:54]: No.Ali [00:27:54]: Yeah, like threeVibhu [00:27:55]: They alsoAli [00:27:55]: Four, five months agoVibhu [00:27:57]: This also happened with, I don't remember which model, but they pulled out quite a few, and then they started a whole chart about this. It might have beenPhilip [00:28:03]: Kimi Vendor Verifier.Ali [00:28:04]: Yeah.Philip [00:28:05]: Yeah.Ali [00:28:05]: Yeah, ‘cause you, ‘cause you'd be pissed, right? Like if you'Philip [00:28:07]: Yeah.Ali [00:28:07]: If like if I'm a consumer and I'm using like Amazon's endpoint for instance, and I've used Kimi and I'm like, “Oh my God, like this is bad,” I'm not gonna say, “Oh, Amazon quantized the model in a bad way.” I'm gonna say, “Oh, Kimi sucks.” Right?Philip [00:28:17]: Yeah.Ali [00:28:17]: So it seems like that makes sense.Philip [00:28:19]: Yeah, they care. They care.Vibhu [00:28:21]: Justifiably.Ali [00:28:21]: Yeah, justifiably.Vibhu [00:28:22]: This is probably a stupid question, but just checking, has anything improved from main quantization?Philip [00:28:28]: Yeah.Vibhu [00:28:28]: Like, is quantization always strictly worse?Ali [00:28:30]: Well technicallyVibhu [00:28:32]: NoAli [00:28:32]: It's a lossy. QuantizationPhilip [00:28:33]: YeahAli [00:28:33]: Is a lossy, it's a lossy implementation.Philip [00:28:36]: Speed improvesVibhu [00:28:36]: Speed improves.Ali [00:28:37]: It the number, likeVibhu [00:28:38]: No, I' always look for inverse scaling laws.Philip [00:28:40]: Yeah.Ali [00:28:40]: Yeah.Vibhu [00:28:40]: This is something I learned from Noam Brown, where like things that normally act in one direction sometimes do.Philip [00:28:45]: Well, technically when you run a benchmark, because these models are deterministic, sometimes your,Ali [00:28:52]: YeahPhilip [00:28:52]: NVFP4 quant is like, two basis points higher than yourAli [00:28:56]: No, it's noise. It's noise.Philip [00:28:57]: Yeah, exactly. I'm like, yeah, it's, it's within. That's why I always say within margin of error.Philip [00:29:01]: And I stopped saying that because everyone assumes that what is, well, within some margin of error, we're barely inside of that to the worst, so we're saying. But yeah, sometimes it's just like, gives you a higher output score. But like Ali said, that's noise. To my knowledge, you're not necessarily making the results better. You're just trying to, again, like keep your fidelity as close to 100% to the original model.Layer Selection, KL Divergence, and Better QuantizationAli [00:29:27]: There is, to your point, research that we did on MP. I don't know if you are able to pullPhilip [00:29:31]: YeahAli [00:29:32]: A tweet we did. One of our research interns, Joshua, I think it's a tweet on how we have 20% better quantized GLM-5.2 than NVIDIA. Essentially what we found throughout like this month research is, okay, quantization is a lossy. It's. You're compressing the data from, occupying 16 bits to occupying, four bits, for instance. And so you're losing some information, and you're trying to minimize that. And so when I say that I'm gonna quantize the model, my job becomes how do I find the layers that I can quantize, and how to find the layers to not. For instance, with image models, I don't quantize modulation layers, and I don't quantize out projections because those two are. Like out projection is what you see as the user. Modulation is what the model sees or understands. Right, exactly. And so to his paper, do you have the. It doesn't have the. Yeah. It's a long paper. I don't know if I can findVibhu [00:30:25]: If there's a part to search or it's probably in the thread.Ali [00:30:28]: It's probably in the thread.Vibhu [00:30:29]: Yeah.Ali [00:30:29]: But the long and the short is it is very possible that quantizing more of the model makes the results. Like if I have a model that I quantize layers one, five, and 10, and another model where I only quantize layers one and It is possible that the model in which I quantized more information is going to perform better because the quantization errors have canceled out. And so what Joshua showed in his mathematical proof where he had like a verifier in, is that you can predict which layers are going to have quantization errors that will cancel out with each other, and you choose to quantize those layers. And so the result of doing this mathematical quantization is you end up with a model that's 20% more quantized than another provider, so you get 20% more throughput of it because there's more layers than running an NVFP4, and your quality is better than that other quant because the layers that you chose to quantize have their errors cancel out, like one layer skewed to the right one layer skewed to the left, one layer skewed to the right. Your final logits distribution is more similar to the original distribution of the model, so you have better fidelity. And so the way we proved this was with KL divergence. So instead of just scoring on the benchmarks, we scored the KL divergence between the logit distribution of the quantized model and the logit distribution of the original full precision model, and we showed that with this technique we get. If your probability distribution on the logits which token it wants to select is more of the same as the original model, you're probably gonna end up staying true to the original model. So yeah, so it seems like previously before this, it seemed like the industry was, well, the more you quantize, the worse it's gonna be, ‘cause the more loss you introduce. That's not exactly, not necessarily true. So yeah, doesn't improve it, but can cancel out.Philip [00:31:57]: I think it might be this, but reminds me a good bit about pruning where you can prune off certain layers.Philip [00:32:03]: But very interesting. Didn't know this was a whole paper you guys put out.Ali [00:32:06]: It's. Fun fact, it was originally 72 pages, this paper, and then we decidedPhilip [00:32:11]: WowAli [00:32:11]: We can't tell. We couldn't release it. So it's now 45.Swyx [00:32:15]: Still 39 pages, so very substantive. We talked about evals and all these things and, like what's possible in terms of speedup? Like it's like probably like the numberInference Speedups and BenchmarkingSwyx [00:32:25]: Thing that people do wanna care about, and it's something that you wrote about in your post. Like official API is 70 tokens per second, and you push it up to 90. Is that like a normal thing?Philip [00:32:36]: So what's cool about working in inference, the reason that I think inference is going to be a useful place to do engineering for a long time, is that if you look at highly optimized domains like, say, finance, if you're in finance, you measure how much better you got in basis points. It's like, “Oh, I got five basis points better, like twentieth of 1% better,” that's huge news because everything is so optimized. When we publish optimizations, it's 20%, it's 100% it's 200%. So there's still probably like a lot further to go, honestly. Like you'll, you'll know that inference is pretty much solved when researchers start publishing about how they got 1% faster at something.Swyx [00:33:19]: Which by the way, because I am from the finance background, in the ‘70s, that was the margin at the time. When you did quantitative finance research, you would findAli [00:33:27]: And like 20%, tens of percent.Swyx [00:33:29]: That's. Yes.Philip [00:33:29]: Yeah.Swyx [00:33:30]: And now it'Philip [00:33:31]: Tiny fractionsSwyx [00:33:32]: For those people interested, look up Andrew Lo's paper. He had a really interesting illustration of quant, stat arb, distribution, narrowing down from like those kinds of 20% differences in the ‘70s, down to nothing today, which is very cool.Philip [00:33:48]: Exactly, and we're at the beginning of the same type of thing. Now benchmarking is hard. I think anyone will tell you that, and benchmarking provider speeds is hard because there's so many variables that go into it. What hardware are you using? How much load do you have on the system? What's the exact nature of the prompts and input and output sequence lengths? All that stuff. But overall, when you start stacking these improvements, you're looking at multiples. You can look at it. The most common form, of course, is TPS, tokens per second, which is bad naming by us in the industry, ‘cause there's two tokens per second. There's tokens per second, the throughput number, and the latency number.Ali [00:34:31]: TTMT, yeah.Philip [00:34:32]: Like total tokens per second out of the, out of the GPU as a throughput number. Most people only care about tokens per second as the latency number, which we should call ITL, intertoken latency, but we don't.Philip [00:34:44]: Anyway, so you can imagine a standard API without many optimizations for a 1 trillion parameter model operating somewhere in the 30 to 50 tokens per second range for reasonable traffic profile. And we generally see the goal of, pushing to 10X that. But, not necessarily day zero, but by stacking enough optimizations, if you have, say like four optimizations, each of which doubles performance. Or sorry, three optimizations, each of which doubles performance, then you stack that up, that's an 8X gain. That's the order of magnitude that we're working with in this space. We're trying to make things substantially faster, not just go from like 70 to 90.Swyx [00:35:38]: Are you saying you've. You have done that?Philip [00:35:40]: So let's say you have as a reasonable baseline, 30 or 40 tokens per second. You can achieve 10X that. So like on GLM-5.2, if you run it unquantized, perhaps on H100s even, and you're just using an off-the-shelf inference engine with no particular optimizations, no speculator, nothing extra around like KV routing, no disaggregation, you're, you're probably, yeah, looking at that like 30 to 40. You think that's like a reasonable baseline?Swyx [00:36:12]: Right. Right.Philip [00:36:12]: To get to something like 10X, there's a lot of trade-offs that you're making. If we're running at more like a 300, 400 tokens per second range, you are using the best hardware possible. You have a optimized speculator. You have done all of your quantization work. You are Seeing a pretty high cache hit rate. You are running with a reasonably small batch size and a parallelism configuration that is tuned for latency versus throughput, but it is possible. So the spreads that you see if you, like, go on artificial analysis or you go on OpenRouter and you look at, the worst provider to the best provider, oftentimes can hit that range. 10X is of course very aggressive. It's oftentimes maybe more of a four to six times improvement. But that's the performance that makes us really excited, is when we can get these huge gains, not just go from 70 to 90 tokens.Stacking Optimizations: NVFP4, Speculation, and DisaggregationAli [00:37:19]: It's also, like, hardware dependent. Like, ifPhilip [00:37:20]: YeahAli [00:37:20]: If you have a thing where you're serving it on just, like, a node of H100s and then you throw, like, you shard the model across, like, four nodes of B200s. Like, you can definitely increase the speed with just throwing more hardware at it. Like, normalizing for the same exact hardware and the same number of GPUs.Philip [00:37:35]: Yeah. Then you're looking at, like, a two to 4X improvementAli [00:37:38]: Right. RightPhilip [00:37:38]: Depending on the inference optimizations. So yeah, it's. Some of it's, what's the call, and some of it's who's the driver.Vibhu [00:37:46]: If you break down the two to 4X, say the example is run GLM-5.2Ali [00:37:51]: YeahVibhu [00:37:51]: On B200sAli [00:37:53]: YeahVibhu [00:37:53]: Single node, right? What's, like, the cost trade-off for effort to get, like, the last bit of juice out versus what should people just think of, right?Ali [00:38:01]: Spectre quantization. Yeah.Vibhu [00:38:03]: Spectre quantization.Ali [00:38:04]: That's, that's, that's like 95%. LikeVibhu [00:38:06]: And how far does that get you? And how easy is that for the average person to do? So say right I wanna throw the weights of GLM-5.2 on a node of B200s, how easy is it to find speculative decoder- decoder model or already quantized model? How much work goes into it?Philip [00:38:23]: If you're doing it up front, it's quite a lot of work. If you're doing it today, there's going to be people who have published things that you can just, you can just grab some NVFP4 weights. You can grab a speculator. Yeah, if we're thinking about, like, what are the 2Xs we're stacking, going from, BF16 to NVFP4 is, it's not quite a 2X, right? It's like. I think it's about, like, 30 to 40%, from 16 to 8, and then another 30 to 40% multiplied from, 8 to 4. So that doesn't quite get you a 2X, but, like, roughly a 2X. Speculator, roughly a 2X. Disagg on top of that if you're able to get enough hardware and put enough traffic through it, another roughly a 2X. And then you add in some, double-digit percent increase from having just a better runtime with, the latest kernels and stuff behind it. And that's how it stacks up.Ali [00:39:21]: YeahPhilip [00:39:21]: So building each of those, like, building the, quantized weights is, for someone who really knows what they're doing, hours to days of work. Building the speculator, again, like, hours to days of work. And the, disagg setup, hours to days. Well okay, but like once you haveAli [00:39:39]: Once set up. Once set up. YeahPhilip [00:39:40]: Yeah, getting disagg working for the first time, I'm saying, of course, is very difficult.Philip [00:39:44]: The marginal implementationAli [00:39:48]: Like, if you're just grabbing, like if you are a person, like just a normal consumer who has access to, like, a node of B200s and you're wondering, “How can I just host it myself?” You don't need to quantize the model yourself. There's always gonna be, like, an open source quantized checkpoint. NVIDIA's gonna push one out if no one else does. You. Usually, the providers will have their own spec dec that they've trained as well. You don't need to train your own spec dec. You can just use that as well.Philip [00:40:09]: Yeah. Like, GLM-5.2 has its own MTP.Ali [00:40:13]: Right. Right.Vibhu [00:40:14]: What's multi token prediction?Philip [00:40:15]: Yes.Ali [00:40:16]: I'm justVibhu [00:40:16]: Can you explain that?Ali [00:40:16]: I'm just an expert.Ali [00:40:18]: I can do it for you in case I get it wrong?Vibhu [00:40:20]: No.Vibhu [00:40:21]: Yeah, you should correct if we're wrong, but their multi-token prediction can be used for self-speculative decoding.Ali [00:40:27]: I'm not sure. I'm not gonna correct that.Vibhu [00:40:28]: Okay. I'm semi-confident in thatAli [00:40:30]: Okay. YeahVibhu [00:40:30]: But someone can check. But it's useful to paint the story of, okay, not just the average person, but say a company wants to switch from serverless inference I wanna throw this up on. I wanna rent some GPUs, throw it up. These are the steps you take to do significantly faster than just put it behind vLLM.Ali [00:40:48]: Right.Vibhu [00:40:49]: I was waiting for a mention of Dynamo.Vibhu [00:40:51]: I feel like, that's supposed to be the baseline that you measure against.Dynamo, KV Routing, and Disaggregation ToolkitsPhilip [00:40:55]: I would think of Dynamo as less of a box system and more of a toolkit for building with. So when we talk about doing aware routing, when we talk about doing KV offloading, when we talk about doing, PD disaggregation, Dynamo fundamentally is. By the way, Dynamo is an open source library from NVIDIA.Ali [00:41:17]: We've done a pod with KylePhilip [00:41:18]: OkayAli [00:41:19]: Kyle Cranin.Philip [00:41:19]: Cool. So then your listeners know then that it supports all the different inference frameworks. And it is multi hardware, which is interesting.Ali [00:41:28]: But it's just a router, it's not like an optimizer layer.Philip [00:41:30]: Yeah. All it does, like, what Dynamo is good at, it is a library for moving information around your cluster, around your hardware. So if you have, KV cache on one place and you need it to be somewhere else, Dynamo coordinates NIXL for you to move that around.Philip [00:41:49]: That doesn't mean that, like, out of the box, you just say, “Pip install Dynamo,” and then you get, like, a massive performance speed up. It's more of a developer toolkit.Ali [00:42:01]: Yeah. I would have said it would. It comes with a set of defaults that you can then swap out.Philip [00:42:06]: It does. If the industry at large, I think, was, like, rolling out all of these deployments, standard, then I think it would be, like, a credible baseline. But, we've got to, we've got to benchmark against, like, what we're seeing in the wild.Speculative Decoding Methods: Medusa, EAGLE, n-Gram, and Spec-SpecVibhu [00:42:23]: I did wanna talk a little bit more about PD disagg, because that is probably, like, number three after quantized and speculative decoding. In your book though, I was just gonna pull out the book.Philip [00:42:31]: Yeah.Vibhu [00:42:32]: Like section 522 on Medusa, 523 on EAGLEPhilip [00:42:35]: YeahVibhu [00:42:36]: 524 on gram.Philip [00:42:37]: It's 55, would be disaggregationAli [00:42:42]: Yeah. Well, no, I just wanted to dwell a little bitPhilip [00:42:44]: YeahAli [00:42:44]: The other. Like, so what do you choose to include? What do you choose to not to include? Because there was all these other techniques.Philip [00:42:51]: Yeah.Ali [00:42:51]: Are these still relevant? Because I think they came out, like, a year and a half ago maybe.Vibhu [00:42:55]: Medusa is quite old.Philip [00:42:56]: Yeah, Medusa's old.Ali [00:42:58]: It was old.Vibhu [00:42:58]: But is it in the book as a good, here'sPhilip [00:43:01]: BaselineVibhu [00:43:01]: Baseline vanilla understand it?Philip [00:43:02]: Like you should know this.Vibhu [00:43:03]: Like I read the paper, I'm like, “ it makes so much sense.”Philip [00:43:05]: Yeah.Philip [00:43:05]: So with the book, I had a couple goals. One was to give people just a working vocabulary for the space as a whole, and the other was to give them some intuition about how each of these techniques works. As I mentioned in my AI Engineer talk, which is the first public addendum to this, the speculation space has moved much faster than everything else. So yeah, even at the time that I wrote the book Medusa, I very much included as a way for people to understand how the space evolved rather than what the most modern technique is. And now of course, there's DFlash, dSpark. There's, there's newer techniques even than EAGLE, although EAGLE is still very commonly used.Ali [00:43:51]: SpecSpecta.Philip [00:43:52]: Yes. Speculative decoding.Vibhu [00:43:54]: What canAli [00:43:56]: Oh, it's a paper by Tri Dao and it's like, it's doing speculative decodingVibhu [00:44:00]: HuhAli [00:44:01]: For the speculative decoder.Philip [00:44:02]: Oh, in spec- oh my God.Ali [00:44:02]: It's literally just an another. It's like, yeah, that's the most simple way to explain it, and it seems like he got trivial speed ups there. But it seems that the complexity with training, it's almost like in our mind at least, it's almost as complex as training GANs. Like it's like a very delicate balance and oftentimes you, it's just but yeah, it's literally speculative decoding on speculative decoding.Vibhu [00:44:21]: Speculative.Ali [00:44:22]: Yeah. We saw this paper.Vibhu [00:44:24]: It's interesting, right?Ali [00:44:24]: Yeah.Vibhu [00:44:24]: I wouldn't even expect it to be very particular to train, I wouldAli [00:44:29]: Right.Vibhu [00:44:29]: The naive part of me is like, okay, train speculative decoder.Ali [00:44:32]: But like, and it makes sense, like the whole idea of speculative decoding is you. It's like, it's like almost like the iPhone auto predict version but for a normal model, right? Like you're just, you're just, generating three tokens and you're like, okay, I'll do prefill on them. And so you save those three turns for your original model. Now your speculative decoder is doing three turns of auto regression, so why not just have an even smaller model?Ali [00:44:53]: The other question there is what are the size of speculators? So say forPhilip [00:44:58]: Right. It's like a billion parameters.Ali [00:45:01]: Like for MiniMax, it's. Yeah. It's like one layer. It's like one 60th of the original model usually.Philip [00:45:06]: Yeah. I think we should do a paper when we get back to the office.Philip [00:45:10]: SpeculativeAli [00:45:11]: SpeculativePhilip [00:45:11]: Decoding.Ali [00:45:13]: No, it's, it does seem like how, when do you stop? But then it also seems like if you're able to train spec-spec decode for instance, right? Like if you're able to have a small model that is accurately predicts what the intermediate speculator is gonna predict, that is able to predict what the original target model's gonna predict, then why not just use that smallest model directly, right?Vibhu [00:45:34]: Yeah. This isAli [00:45:35]: Like it seems likeVibhu [00:45:35]: Adjacent to the routing problem.Ali [00:45:36]: Right.Vibhu [00:45:36]: Yeah.Ali [00:45:36]: Right.Philip [00:45:37]: The thing with speculators is one of the practical constraints on using them is that you do have to run a small model on the same hardware that you're running the big model on. There is a orchestration and resource competition problem inherent in that, and that is one of the constraints on speculation in general, is that draft tokens cost resources to create and cost software complexity to manage. And so if you have like infinitely recursive speculators, you add in quite a bit of that complexity on the actual implementation within the inference engine as well, not just in the training process.Vibhu [00:46:17]: I was gonna say, I would wonder if you could do similar, like distillation and pruning of, it's the same thing, it's just a model. Can we not just distill a lot of the weights, quantize the speculator, out of my domain? The question that also comes up is, this is all for big server workloads, right? How much of this applies to, say I have this MacBook, I wanna run Gemma really efficiently. Similar problems, not the same?Local AI vs. Data Center InferencePhilip [00:46:45]: Pretty different. I talked to Selo, about this on his podcast a couple weeks ago. The difference between inference engineering for the data center and for production workloads versus inference engineering for local AI, is that we start with fundamentally like different constraints and different goals. With local AI, it's how do I fit this model onto my hardware and then make it less dumb? And with data center influence, it's how do I load this model and then make it less slow? And we care about less dumb, and they care about less slow. But the local AI inference engineering ecosystem, I think has a lot for us to learn from in the data center space. They are experts in various forms of quantization, including dynamic quantization that we just don't touch, in the pruning, in the distillation, in the, layer removal. There'Ali [00:47:42]: Layer removal matters less.Philip [00:47:43]: Yeah. There'Ali [00:47:44]: No one loves pruning really.Philip [00:47:45]: Yeah. Well, but the, but they doVibhu [00:47:46]: Which is surprising, right? But that's, that's a whole different thingPhilip [00:47:48]: Just to fit something on the laptop.Ali [00:47:50]: Right.Philip [00:47:50]: So yeah, it's a, it's an interesting, it's an interesting space. Not necessarily that like their techniques make sense for us to do in the data center, because we have different resources and different goals, but more that the process as well as the openness of that field is something to, admire.Ali [00:48:12]: Yeah. Like to your point, like, certain optimizations that would. Like for instance, Turbo Quantum Sharper, like it made such huge hype on that and we did like a whole deep dive on Twitter and like said, what is it? How does it work? Why is it good or not? And it took off and it was implemented on local devices because your memory bandwidth is so slow on like a MacBook, for instance. But try putting the same thing on like an NVIDIA GPU on a B200 Turbo quant would not be. Like, it would not be used. Like, NVIDIA - Like, NVIDIA made it clear that this is not a good optimization, and we've seen it firsthand where the overhead of doing dequantization, quantization of, in the kernel itself with turbo quant kernel, each end is much slower than the time that you save from doing the bandwidth. ‘Cause on the B200s, you have like 3.5 terabytes per second. You don't need decrease the storage that much. You don't need to do, FP4 KV cache. You don't need to use a requant. There's, there's, there's better optimizations to be made. But on Edge devices, it's extremely important, it's extremely useful. So, seems to be, like, different optimizations there, but then they're all uniquely combined with like all you wanna quantize the model, you wanna do speculative decoding, like certain common prefixes with bothPhilip [00:49:18]: Principles.Ali [00:49:19]: Yeah, exactly. Exactly. Exactly.Philip [00:49:20]: They also do a lot of work on, model parallelism, especially over, heterogeneous topology, where you have, some sparks and they are wired together with, Ethernet, DGX sparks.Ali [00:49:35]: Yeah, this is the Exo Labs guys.Philip [00:49:36]: Yeah. You have, a nu
Intel Chat with Matt Bromiley and Chris Luft.Matt and Chris break down four stories from the week in threat intel:• Hugging Face's security incident disclosure: an intrusion conducted end-to-end by an autonomous AI agent system — a malicious dataset exploiting two code-execution paths, thousands of actions across short-lived sandboxes, self-migrating C2 — and why the forensics had to run on the open-weight GLM 5.2 model after hosted frontier models refused to analyze real attack artifacts.• WP2Shell: attackers chaining CVE-2026-60137 (WordPress Core SQL injection) with CVE-2026-63030 (Batch REST API logic flaw) for unauthenticated remote code execution on default WordPress installs — found by Searchlight Cyber using GPT-5.6 Sol Ultra in about ten hours, with tens of thousands of exploitation attempts following disclosure.• Data breaches at AI music generator Suno (55.3M unique email addresses, plus partial Stripe payment records) and gig-work platform Paidwork (23.3M addresses, password hashes and banking data), per Have I Been Pwned.• Iranian state media claims the IRGC destroyed AWS's Bahrain data center (ME-SOUTH-1) with cruise missiles — and what data centers becoming military targets means for cloud resilience.Plus: Google Threat Intelligence Group retires APT/FIN nomenclature for new threat-actor names, and where to find Chris and Matt at Black Hat.Stories covered:• https://huggingface.co/blog/security-incident-july-2026• https://www.darkreading.com/cyberattacks-data-breaches/wp2shell-millions-wordpress-sites-remote-takeover• https://www.securityweek.com/suno-paidwork-data-breaches-affect-tens-of-millions-of-accounts/• https://www.tomshardware.com/tech-industry/data-centers/amazon-data-center-in-bahrain-struck-and-destroyed-by-iranian-cruise-missiles-state-media-claims-attacks-launched-against-aws-site-in-response-to-alleged-us-strikes-on-an-under-construction-nuclear-plantChapters:0:00 Intro & Black Hat plans2:07 Hugging Face's AI-agent breach disclosure12:39 WP2Shell: WordPress exploit chain20:59 Suno & Paidwork data breaches24:17 IRGC strikes on AWS Bahrain28:27 Google Threat Intel's new actor names29:29 Black Hat swag hunt & wrap-upThe Cybersecurity Defenders Podcast — a podcast about cybersecurity and the people that keep the internet safe. New episodes drop weekly.Subscribe wherever you listen:• Spotify: https://open.spotify.com/show/6ep00zeY3S8ffZ4o0UeSps• Apple Podcasts: https://podcasts.apple.com/us/podcast/the-cybersecurity-defenders-podcast/id1649981740• YouTube: https://www.youtube.com/@limacharlieioLearn more about LimaCharlie: https://limacharlie.io#cybersecurity #infosec #threatintel #AIsecurity #databreach
Intel Chat with Matt Bromiley and Chris Luft.Matt and Chris break down four stories from the week in threat intel:• Hugging Face's security incident disclosure: an intrusion conducted end-to-end by an autonomous AI agent system — a malicious dataset exploiting two code-execution paths, thousands of actions across short-lived sandboxes, self-migrating C2 — and why the forensics had to run on the open-weight GLM 5.2 model after hosted frontier models refused to analyze real attack artifacts.• WP2Shell: attackers chaining CVE-2026-60137 (WordPress Core SQL injection) with CVE-2026-63030 (Batch REST API logic flaw) for unauthenticated remote code execution on default WordPress installs — found by Searchlight Cyber using GPT-5.6 Sol Ultra in about ten hours, with tens of thousands of exploitation attempts following disclosure.• Data breaches at AI music generator Suno (55.3M unique email addresses, plus partial Stripe payment records) and gig-work platform Paidwork (23.3M addresses, password hashes and banking data), per Have I Been Pwned.• Iranian state media claims the IRGC destroyed AWS's Bahrain data center (ME-SOUTH-1) with cruise missiles — and what data centers becoming military targets means for cloud resilience.Plus: Google Threat Intelligence Group retires APT/FIN nomenclature for new threat-actor names, and where to find Chris and Matt at Black Hat.Stories covered:• https://huggingface.co/blog/security-incident-july-2026• https://www.darkreading.com/cyberattacks-data-breaches/wp2shell-millions-wordpress-sites-remote-takeover• https://www.securityweek.com/suno-paidwork-data-breaches-affect-tens-of-millions-of-accounts/• https://www.tomshardware.com/tech-industry/data-centers/amazon-data-center-in-bahrain-struck-and-destroyed-by-iranian-cruise-missiles-state-media-claims-attacks-launched-against-aws-site-in-response-to-alleged-us-strikes-on-an-under-construction-nuclear-plantChapters:0:00 Intro & Black Hat plans2:07 Hugging Face's AI-agent breach disclosure12:39 WP2Shell: WordPress exploit chain20:59 Suno & Paidwork data breaches24:17 IRGC strikes on AWS Bahrain28:27 Google Threat Intel's new actor names29:29 Black Hat swag hunt & wrap-upThe Cybersecurity Defenders Podcast — a podcast about cybersecurity and the people that keep the internet safe. New episodes drop weekly.Subscribe wherever you listen:• Spotify: https://open.spotify.com/show/6ep00zeY3S8ffZ4o0UeSps• Apple Podcasts: https://podcasts.apple.com/us/podcast/the-cybersecurity-defenders-podcast/id1649981740• YouTube: https://www.youtube.com/@limacharlieioLearn more about LimaCharlie: https://limacharlie.io#cybersecurity #infosec #threatintel #AIsecurity #databreach
Every major AI lab signed the Open Weights letter defending open models. Meta, OpenAI, Google, Microsoft, Nvidia.Anthropic was the only holdout.Yesterday, its CEO, Dario Amodei, published a thoughtful defense of that decision to not fully support open weight or open source models. Here's what nobody's connecting: the money trail. Roughly 80% of Anthropic's revenue is businesses paying per token. Free Chinese open models attack that exact revenue stream weeks before Anthropic is set to go public. On today's show we break down what Dario actually said, what he said before, and why we think this was written for Washington policymakers and not for the rest of us.Anthropic Responds: Why Claude's CEO didn't sign the open model pact and the real reasons why -- An Everyday AI Chat with Jordan WilsonNewsletter: Sign up for our free daily newsletterMore on this Episode: Episode PageToday's Episode on LinkedIn: Thoughts on this? Join the convo on LinkedIn and connect with other AI leaders.Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineupWebsite: YourEverydayAI.comEmail The Show: info@youreverydayai.comConnect with Jordan on LinkedInTopics Covered in This Episode:Anthropic Refuses Open Model PactDario Amodei's Public Letter AnalysisAnthropic's 80% Revenue Token ExposeChinese Open Model National Security FearsMicrosoft & Nvidia's Open Weights CoalitionRegulatory Capture and Washington InfluenceTiming Related to Executive Order DeadlineIPO Motivations Behind Anthropic's DecisionsContradictions in Anthropic's Open Model StanceImpact of Open Source on Token Business ModelTimestamps:00:00 Anthropic's stance on open models04:19 Discussing Anthropic's response to open models06:39 Understanding open weight models12:08 Future AI and cybersecurity risks15:04 Discussion on open-source AI models19:30 Discussing Anthropic's business challenges21:08 Cutting costs with open-source models26:17 Anthropic's recent stock downturn27:11 AI investment and cost efficiency shift30:14 Anthropic's stance on open source models36:32 INTROPICS IPO and regulatory discussions37:37 Wrapping up and subscribingKeywords: Anthropic, Claude, open model pact, open source AI, open weights, American AI leadership, Dario Amodei, IPO, regulatory capture, DC lawmakers, Chinese open source models, token revenue, per token business model, NVIDIA, Microsoft, Meta, OpenAI, Google, IBM, national security, AI safety, government mandates, chip controls, AI regulation, chip ban, industrial scale distillation, mandatory safety testing, inference, AI ecosystem, Opus 5, Fable 5, GPT-5, GLM 5.2, cost per task, token efficiency, model router, proprietary models, closed source AI, cybersecurity risks, Chinese cyberattacks, biological attacks, Glasswing program, open source vs proprietary, tech lobbying, Trump AI order, federal deadline, AI policy, artificial general intelligence, artificial superintelligence, AI monetization, S-1 filing, public company, venture capital, AI benchmarks, model switching, API pricing, model containment, Hugging Face incident, AI startup monopoly, safety vs business protection, market competition, AI cost reduction.Send Everyday AI and Jordan a text message. (We can't reply back unless you leave contact info) Ready for ROI on GenAI? Go to youreverydayai.com/partner
PEBCAK Podcast: Information Security News by Some All Around Good People
Welcome to this week's episode of the PEBCAK Podcast! We've got three amazing stories this week so sit back, relax, and keep being awesome! Be sure to stick around for our Dad Joke of the Week. (DJOW) Follow us on Instagram @pebcakpodcast Please share this podcast with someone you know! It helps us grow the podcast and we really appreciate it! Simple 6 signup link https://simple6.co/r/CFUR98 Hugging Face allegedly had to fall back on China's GLM 5.2 to investigate the OpenAI hack after US frontier models refused to help with forensic analysis, unable to distinguish attacker from defender. - https://x.com/coinbureau/status/2079663328021057654?s=46 - https://x.com/t3chfalcon/status/2079998873183949221 Per the claim, safety guardrails on US frontier models blocked forensic assistance during the incident response because the models couldn't tell the investigating security team apart from the attacker, forcing Hugging Face to run GLM 5.2 on its own servers to complete the analysis — a notable reversal given the assumption that domestic models would be the trusted fallback in a crisis. VW drivers running GrapheneOS say the automaker's app has locked them out entirely, deepening fears carmakers are forcing users back into Google's ecosystem. VW has now confirmed to German outlet heise that it deliberately blocked GrapheneOS, LineageOS, and /e/OS users via Google's Play Integrity API — the same certification gatekeeper already used by Barclays, HSBC, Monzo, Netflix, and Disney+. - https://cybernews.com/privacy/volkswagen-grapheneos-app-issues/ - https://x.com/intcyberdigest/status/2079992915972018246?s=46 - https://x.com/orwellday/status/2079704437241610443?s=46 Roughly 500,000 GrapheneOS users are affected; VW's app still supports outdated Android versions but rejects the privacy-focused OS, and VW told one user GrapheneOS "is not an official Volkswagen offering." The lockout follows a recent VW API change that also cut off third-party smart-charging and home-automation tools, raising questions about EU Data Act compliance. A 19-year-old alleged Scattered Spider member's globe-trotting VPN op-sec got shredded by a Windows telemetry ID he probably didn't know existed. - https://cybersecuritynews.com/windows-device-identifier-tracking/ Peter Stokes (dual US-Estonian, 19) allegedly ran an $8M extortion hit on a luxury retailer using voice-phishing, ngrok tunneling, and 77GB of S3 exfil — but the FBI cross-referenced his Microsoft Global Device Identifier (GDID) across Apple, Snapchat, Facebook, and even a Ubisoft login to place the same device on the same IPs in Tallinn, NYC, and Thailand, matching his travel records. Microsoft has since acknowledged using the GDID to track the user across three countries — meaning the VPN masked his network endpoint, but the GDID rendered that protection moot. Dad Joke of the Week (DJOW) Find the hosts on LinkedIn: Chris - https://www.linkedin.com/in/chlouie/ Glenn - https://www.linkedin.com/in/glennmedina/ Jason - https://www.linkedin.com/in/jason-seemann-12b7075/
Ep 288 MultiSIM Apple — jedan broj na više uređaja, i bez iPhone-a uz sebe Apple sues OpenAI, alleging theft of trade secrets to fuel hardware ambitions - MacDailyNews Apple sues OpenAI, accuses ex-employees of stealing trade secrets - 9to5Mac Drew Pusateri: Our statement in response to this suit: We have no interest in other companies' trade secrets. We remain focused on building innovative technology that empowers people everywhere. Daring Fireball: ‘No Interest' OpenAI shrugs, denies responsibility for trade secret theft with vague statements The Joy of Tech: Apple VS Open AI! OpenAI's first hardware device will be a HomePod, but don't tell them that OpenAI Shuffles ChatGPT Apps, Kills Atlas Browser, Improves Voice - TidBITS Here's Why Apple is Reportedly Skipping M6 Pro and M6 Max Chips SigLens acquired by Apple for debugging massive apps and services Port-limited MacBook Neo gets color-matching hub & mouse The Real System Requirements for OS 27 - TidBITS Alek: promena preporučene mašine za iOS devs jer Xcode 27 satire disk prostor Apple Patches Hide My Email Flaw More Than a Year After It Was Reported CrashStealer malware masquerades as Apple's crash report tool to raid your Mac - Cult of Mac EU Orders Google to Give Rival AI Apps the Same Android Access as Gemini Daring Fireball: European Commission: ‘Guidance to Google for AI Interoperability on Android & Sharing of Google Search' EU slaps Google with a $1 billion antitrust fine Apple Watch, Meta Glasses, AirPods get reprieve from EU replaceable battery law Katie Paxton-Fear: Can we trust Chinese open weight models? Was a question a lot of people asked after GLM 5.2 was released, scoring very well on coding benchmarks, and suspiciously Claude-like. So I turned an open-weight coding model into a backdoor with 1hr and
A U.S. vs China AI cold war is starting, and most business leaders have no idea they're already in it.China's open models just closed the gap with America's best, oftentimes at a fraction of the price.Now both governments are moving to wall off their AI within days of each other.Why? Because this was never about benchmarks. It's about power y'all. We break it all down on today's show and help you figure out the 101 of the AI war between U.S. and China. The U.S. vs China AI Cold War Is Starting: What It Means and How It Impacts You -- An Everyday AI Chat with Jordan WilsonNewsletter: Sign up for our free daily newsletterMore on this Episode: Episode PageToday's Episode on LinkedIn: Thoughts on this? Join the convo on LinkedIn and connect with other AI leaders.Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineupWebsite: YourEverydayAI.comEmail The Show: info@youreverydayai.comConnect with Jordan on LinkedInTopics Covered in This Episode:U.S.-China AI Cold War OverviewChinese AI Models Closing U.S. GapGovernment Restrictions on AI Model AccessEconomic and Geopolitical AI Power StruggleRisks for U.S. Businesses Using Chinese AIOpen Source vs. Closed Source AI DebateChinese AI Model Pricing Undercuts U.S.AI Model Distillation and U.S. Security ConcernsEnterprise AI Cost-Effectiveness BenchmarksMicrosoft Testing Chinese AI DeploymentsFuture AI Model Export Controls & StrategiesRecommendations for AI Model Sourcing and RiskTimestamps:00:00 US-China AI tensions escalate04:30 Switching to Chinese AI models08:47 US vs China in open source models11:39 China's narrative control efforts14:42 Challenges in AI model development18:25 Differentiating open source strategies23:04 AI model cost-effectiveness analysis26:31 US measures against model distillation29:38 Discussing Microsoft's use of AI models31:17 Controlling export of AI modelsKeywords: US vs China AI cold war, China AI restrictions, US AI restrictions, AI model export controls, Chinese open source AI models, AI geopolitical power, economic growth through AI, global AI standards, AI superpower race, AI model benchmarks, open weight models, enterprise AI deployment, trillion parameter AI models, Microsoft AI model testing, AI model pricing, Claude Fable 5, GPT-5.6, GLM 5.2, Kimmi K3, Alibaba Qwen 3.8, model distillation, AI cybersecurity risks, AGI leadership, military AI use cases, China narrative control, model adoption, compute power for AI, AI training data, AI export law, US national security and AI, model routing, mixture of models, cost per intelligence index, Anthropic models, cost per task AI, model capability parity, AI market adoption, cloud competition, AI architecture innovation, AI model sanctionsSend Everyday AI and Jordan a text message. (We can't reply back unless you leave contact info) Ready for ROI on GenAI? Go to youreverydayai.com/partner
What happens when China drops open-weight AI models that rival Silicon Valley's best? This episode unpacks how a new wave of international AI releases is shaking up business, policy, and the future of innovation. Linus Torvalds to critics of AI coding in Linux: "Fork it. Or just walk away." Claude on X: "Beginning July 20, Claude Fable 5 will be included in all Max and Team Premium plans, at 50% of limits. Pro and Team Standard users will continue to have access to Fable via usage credits, and will receive a one-time $100 credit. Demand for Fable has been challenging to" China's Moonshot AI Unveils Kimi Model, Threatening America's Lead Alibaba's Qwen Unveils Preview of Flagship AI Model Social media limits are coming for teens across Europe The White House is now deciding who gets access to frontier AI models, not the labs Microsoft chief turns hostile on frontier AI labs, warns companies to guard their IP Meta Is Flooding the Market With Smartglasses. Privacy Advocates Are Up in Arms. Federal employees can download TikTok on government devices, DOJ says Amazon Web Services customers receive bills for up to $1.5tn after global glitch MLB cracks down on using AI via dugout iPads to help shape in-game decisions White House Teleprompter Operator Bet on Trump Speeches, Kalshi Says New York school district is testing lifelike robot teachers Host: Leo Laporte Guests: Harper Reed and Alex Wilhelm Download or subscribe to This Week in Tech at https://twit.tv/shows/this-week-in-tech Join Club TWiT for Ad-Free Podcasts! Support what you love and get ad-free audio and video feeds, a members-only Discord, and exclusive content. Join today: https://twit.tv/clubtwit Sponsors: blackhat.com/us-26 and use code TWIT ZipRecruiter.com/twit ethos.com/twit arcticwolf.com/trends threatlocker.com/twit shopify.com/twit
Alibaba launched Qwen3.8 Max as Moonshot paused Kimi K3 signups amid demand. The Trump administration weighed a slow squeeze on Chinese AI, Hugging Face used China's GLM-5.2 after US guardrails blocked its breach forensics, and Google built a Gemini chip. Alibaba launches a 2.4T parameter Qwen3.8 Max preview that it says rivals frontier AI models and is second only to Fable 5, plans to make it "open-weight soon" (Bloomberg) The Trump administration has reportedly explored sanctions, security warnings, and executive-order requirements since 2025 to build a slow, durable squeeze on Chinese AI models instead of pursuing an outright ban (The Decoder) Ben Thompson argues the market reaction to Kimi and other Chinese models is overblown, since it's compute scarcity — not a real Chinese cost advantage — that's keeping frontier-model prices high (Stratechery) Hugging Face says it used the open-weight GLM-5.2 hosted on its own compute for breach forensics, after US frontier model safety guardrails blocked the requests (The Stack) Sources: Google is developing a specialized server chip, informally dubbed "Frozen v2", that integrates its Gemini AI model blueprint into the silicon, for 2028 (The Information) Subscribe to the ad-free feed. Learn more about your ad choices. Visit megaphone.fm/adchoices
What happens when China drops open-weight AI models that rival Silicon Valley's best? This episode unpacks how a new wave of international AI releases is shaking up business, policy, and the future of innovation. Linus Torvalds to critics of AI coding in Linux: "Fork it. Or just walk away." Claude on X: "Beginning July 20, Claude Fable 5 will be included in all Max and Team Premium plans, at 50% of limits. Pro and Team Standard users will continue to have access to Fable via usage credits, and will receive a one-time $100 credit. Demand for Fable has been challenging to" China's Moonshot AI Unveils Kimi Model, Threatening America's Lead Alibaba's Qwen Unveils Preview of Flagship AI Model Social media limits are coming for teens across Europe The White House is now deciding who gets access to frontier AI models, not the labs Microsoft chief turns hostile on frontier AI labs, warns companies to guard their IP Meta Is Flooding the Market With Smartglasses. Privacy Advocates Are Up in Arms. Federal employees can download TikTok on government devices, DOJ says Amazon Web Services customers receive bills for up to $1.5tn after global glitch MLB cracks down on using AI via dugout iPads to help shape in-game decisions White House Teleprompter Operator Bet on Trump Speeches, Kalshi Says New York school district is testing lifelike robot teachers Host: Leo Laporte Guests: Harper Reed and Alex Wilhelm Download or subscribe to This Week in Tech at https://twit.tv/shows/this-week-in-tech Join Club TWiT for Ad-Free Podcasts! Support what you love and get ad-free audio and video feeds, a members-only Discord, and exclusive content. Join today: https://twit.tv/clubtwit Sponsors: blackhat.com/us-26 and use code TWIT ZipRecruiter.com/twit ethos.com/twit arcticwolf.com/trends threatlocker.com/twit shopify.com/twit
What happens when China drops open-weight AI models that rival Silicon Valley's best? This episode unpacks how a new wave of international AI releases is shaking up business, policy, and the future of innovation. Linus Torvalds to critics of AI coding in Linux: "Fork it. Or just walk away." Claude on X: "Beginning July 20, Claude Fable 5 will be included in all Max and Team Premium plans, at 50% of limits. Pro and Team Standard users will continue to have access to Fable via usage credits, and will receive a one-time $100 credit. Demand for Fable has been challenging to" China's Moonshot AI Unveils Kimi Model, Threatening America's Lead Alibaba's Qwen Unveils Preview of Flagship AI Model Social media limits are coming for teens across Europe The White House is now deciding who gets access to frontier AI models, not the labs Microsoft chief turns hostile on frontier AI labs, warns companies to guard their IP Meta Is Flooding the Market With Smartglasses. Privacy Advocates Are Up in Arms. Federal employees can download TikTok on government devices, DOJ says Amazon Web Services customers receive bills for up to $1.5tn after global glitch MLB cracks down on using AI via dugout iPads to help shape in-game decisions White House Teleprompter Operator Bet on Trump Speeches, Kalshi Says New York school district is testing lifelike robot teachers Host: Leo Laporte Guests: Harper Reed and Alex Wilhelm Download or subscribe to This Week in Tech at https://twit.tv/shows/this-week-in-tech Join Club TWiT for Ad-Free Podcasts! Support what you love and get ad-free audio and video feeds, a members-only Discord, and exclusive content. Join today: https://twit.tv/clubtwit Sponsors: blackhat.com/us-26 and use code TWIT ZipRecruiter.com/twit ethos.com/twit arcticwolf.com/trends threatlocker.com/twit shopify.com/twit
What happens when China drops open-weight AI models that rival Silicon Valley's best? This episode unpacks how a new wave of international AI releases is shaking up business, policy, and the future of innovation. Linus Torvalds to critics of AI coding in Linux: "Fork it. Or just walk away." Claude on X: "Beginning July 20, Claude Fable 5 will be included in all Max and Team Premium plans, at 50% of limits. Pro and Team Standard users will continue to have access to Fable via usage credits, and will receive a one-time $100 credit. Demand for Fable has been challenging to" China's Moonshot AI Unveils Kimi Model, Threatening America's Lead Alibaba's Qwen Unveils Preview of Flagship AI Model Social media limits are coming for teens across Europe The White House is now deciding who gets access to frontier AI models, not the labs Microsoft chief turns hostile on frontier AI labs, warns companies to guard their IP Meta Is Flooding the Market With Smartglasses. Privacy Advocates Are Up in Arms. Federal employees can download TikTok on government devices, DOJ says Amazon Web Services customers receive bills for up to $1.5tn after global glitch MLB cracks down on using AI via dugout iPads to help shape in-game decisions White House Teleprompter Operator Bet on Trump Speeches, Kalshi Says New York school district is testing lifelike robot teachers Host: Leo Laporte Guests: Harper Reed and Alex Wilhelm Download or subscribe to This Week in Tech at https://twit.tv/shows/this-week-in-tech Join Club TWiT for Ad-Free Podcasts! Support what you love and get ad-free audio and video feeds, a members-only Discord, and exclusive content. Join today: https://twit.tv/clubtwit Sponsors: blackhat.com/us-26 and use code TWIT ZipRecruiter.com/twit ethos.com/twit arcticwolf.com/trends threatlocker.com/twit shopify.com/twit
Alice Han and James Kynge start with China's latest ballistic missile test into the Pacific and what it means alongside a new Australia-Fiji defense pact. Then: Europe wants to shrink its record trade deficit with China, but its worst heat wave on record has sent demand for Chinese air conditioners soaring. They break down whether the two sides can actually cooperate on AI and renewable energy even as tensions rise. Plus: Z.ai just launched ZCode, a coding agent for its GLM-5.2 model said to rival Claude and ChatGPT. Alice and James discuss how big a threat this is to U.S. AI dominance, as well as the fallout from claims that Anthropic used hidden code to track Chinese users. Finally: China's new "Ethnic Unity" law took effect July 1. What does it mean for Tibetans, Uyghurs, and Taiwan… and how is Beijing defending it internationally? Subscribe to China Decode on Substack for weekly analysis, livestreams, and deep dives into the biggest story shaping the global economy: chinadecode.profgmedia.com Learn more about your ad choices. Visit podcastchoices.com/adchoices
Google fires the engineer behind its Workspace CLI tool, OpenAI previews GPT-5.6 with three new model tiers, and Astro 7 lands with a full Rust rewrite. Plus: Coinbase cuts token costs with smarter routing, and more in this week's Syntax Live Show Notes 00:00 Intro 00:34 Welcome to Syntax! 01:46 Google fires Workspace CLI Creator 12:30 GPT 5.6 Is Coming 19:59 GLM 5.2 Released 23:23 Astro 7 Rust Re-write 32:46 Cursor Announces iOS App 35:08 Scott's Workflow: Herdr + Mosh + Termius + Tailscale 40:33 Coinbase Reduces AI Cost with Model Routing 44:22 wayfinder-router - Local AI Routing CLI 48:16 Token efficiency in models and harnesses Martin Woodward on X 52:34 performativeUI - react components for AI startups 54:31 Brought to you by Sentry.io 55:21 Reachy Mini Robot 01:02:21 FUTO Keyboard Swipe for Android 01:05:40 CSS Quake 01:07:35 HTML Invoker API is Baseline Available 01:13:17 Cloudflare Temporary Accounts for AI Agents Hit us up on Socials! Syntax: X Instagram Tiktok LinkedIn Threads Wes: X Instagram Tiktok LinkedIn Threads Scott: X Instagram Tiktok LinkedIn Threads Randy: X Instagram YouTube Threads