Podcasts about Kimi

  • 2,002PODCASTS
  • 5,310EPISODES
  • 57mAVG DURATION
  • 1DAILY NEW EPISODE
  • Aug 25, 2026LATEST

POPULARITY

20192020202120222023202420252026

Categories



Best podcasts about Kimi

Show all podcasts related to kimi

Latest podcast episodes about Kimi

The John Batchelor Show
S8 Ep1351: Michael Sobolik and Gordon Chang examine how Chinese AI company Moonshot bypassed US export controls by "distilling" technology from Anthropic's Claude Opus model to train its own system, Kimi K3. This intellectual property theft th

The John Batchelor Show

Play Episode Listen Later Aug 25, 2026 10:10


Michael Sobolik and  Gordon Chang examine how Chinese AI company Moonshot bypassed US export controls by "distilling" technology from Anthropic's Claude Opus model to train its own system, Kimi K3. This intellectual property theft threatens American market dominance and national security. Sobolik recommends three policy actions: imposing crushing financial sanctions on violating Chinese firms, closing export control loopholes related to remote cloud access, and banning open CCP models in the United States. He warns that American tech companies prioritizing short-term profits over security risk losing the AI race, mirroring historical patterns of Chinese piracy. (9)

F1 Nation
‘The fight is still on'. Is Lando in the title race? – Dutch GP Review

F1 Nation

Play Episode Listen Later Aug 24, 2026 54:15


Tom Clarkson is in the Zandvoort paddock with F1 correspondent Lawrence Barretto and F1TV expert Alex Brundle for reaction to a chaotic Dutch Grand Prix. Back-to-back wins for Lando Norris means he's now 83 points behind championship leader Kimi Antonelli with 11 races to go, so is the McLaren driver back in the title fight? Despite being behind him for most of the weekend, Kimi Antonelli got the better of his Mercedes teammate George Russell on race day. Why didn't George have the pace when it really mattered? Were Mercedes right to issue team orders to get George out of Kimi's way in the closing stages? On the topic of team orders, Lewis Hamilton was very frustrated that Ferrari didn't ask Charles Leclerc to let him past. Did that cost Hamilton a podium? And what does his fiery team radio and post-race reaction tell us about the seven-time World Champion's mindset? Plus, the guys discuss home race heartbreak for Max Verstappen after he crashed out on lap one, the impact Fernando Alonso's second points finish of the season will have on his future with Aston Martin, and a dramatic end to the weekend for Williams as Carlos Sainz crashed into teammate Alex Albon.Listen to more Official F1 PodcastsIn-depth interviews on F1 Beyond The Grid - Sergio Perez coming soon!Expert answers to your questions on F1 ExplainsThis episode is sponsored by:HexcladFind your forever cookware @hexclad and get 10% off at hexclad.co.uk/NATION! Offer excludes bundles and other items already on sale. #hexcladpartnerBetterhelpYou don't have to navigate life's changes alone.Sign up and get 10% off at BetterHelp.com/F1NATION 

The Fast And The Curious
'The championship is BACK ON' | Who was right in team orders decisions? | Dutch GP 2026 Reaction

The Fast And The Curious

Play Episode Listen Later Aug 23, 2026 44:31


Greg, Betty and Christian are together to react to an eventful last Dutch Grand Prix at Zandvoort (for the foreseeable future).They react to Lando's chances of closing a smaller championship gap to Kimi than Max almost managed on Lando last year, discuss Mercedes' decision to swap around drivers, and Greg and Christian argue over how best to tease what exciting content we have coming up in the week…If you are new to the podcast, make sure you follow us on all the socials and hit subscribe right here because we are covering the remainder of the 2026 season from lights out to chequered flag!YouTube: @fastcuriouspodX: @fastcuriouspodInstagram: @fastcuriouspodTikTok: @fastcuriouspodThreads: @fastcuriouspod Producer: Will TyrrellExecutive Producer: Christian Hewgill Hosted on Acast. See acast.com/privacy for more information.

Audio Mises Wire
The Innovation Mirage: Why DeepSeek and Kimi 3 Do Not Settle the Question of Chinese Technological Supremacy

Audio Mises Wire

Play Episode Listen Later Aug 18, 2026


There is no doubt that Chinese firms have made many advances in AI technology, but how far they can go in the future still is an open question.Original article: https://mises.org/mises-wire/innovation-mirage-why-deepseek-and-kimi-3-do-not-settle-question-chinese-technological-supremacy

F1 Nation
Who can catch Kimi as F1 returns? Dutch GP Preview

F1 Nation

Play Episode Listen Later Aug 17, 2026 53:07


Tom Clarkson, Jolyon Palmer and James Hinchcliffe look forward to the return of Formula 1 at the Dutch Grand Prix.Kimi Antonelli leads the F1 World Championship by 50 points, but his rivals are still in the fight. The guys discuss whether the young Italian can continue his stellar season under growing pressure in the second half of the year.After some 'sloppy' mistakes in recent races, can Ferrari challenge Mercedes? Will Lando Norris, our most recent race-winner, become a factor in the title fight? Also in this episode, what's behind a change of Team Principal at Cadillac? And will Aston Martin make another leap forward at Zandvoort? This episode is sponsored by:ClaudeReady to have an AI that can tackle real work? Try Claude Cowork today — Claude.ai/nation.BetterhelpYou don't have to navigate life's changes alone.Sign up and get 10% off at BetterHelp.com/F1NATION IM8Give your body what it deserves with IM8! Go to IM8HEALTH.com/nation and use code NATION for a Free Welcome Kit, five free travel sachets plus ten percent off your order. #im8health

Racecast
Formula 1 Is Back, Why The Championship Is Wide Open & Is Kimi's Future at Ferrari?!

Racecast

Play Episode Listen Later Aug 17, 2026 41:11


F1 is BACK! Join Luke & Matt as we discuss all things Formula 1!

Software Engineering Daily
SED News: The Kimi Moment, Runaway AI, and Tokenmaxxing

Software Engineering Daily

Play Episode Listen Later Aug 11, 2026 48:28


SED News is a monthly podcast from Software Engineering Daily where hosts Gregor Vand and Sean Falconer break down the biggest stories shaping software engineering, Silicon Valley, and the broader tech industry. In this episode, Gregor and Sean dig into a wave of “runaway AI” stories, including Anthropic and OpenAI disclosing that their models accessed outside organizations during cyber evaluations, and Amazon reporting a staggering budget overrun blamed on bad agent loops. They explore why most of these incidents trace back to human decisions rather than models breaking free, and the awkward reality that today’s systems can bill you for tokens without reliably counting them. They also talk about the “Kimi moment.” Moonshot AI‘s open weight model has closed the gap with frontier models like ChatGPT and Claude at a remarkable pace, and the hosts unpack what it means for open weight strategies and how chip scarcity is pushing Chinese labs to innovate. As always, the episode wraps up with a few standout Hacker News threads, including a JetBrains test of a “caveman speak” skill that promised big token savings, how refactoring can cut input token costs, the release of CodePen 2.0, and a build of Doom that renders through SQL queries.Sponsorship inquiries:sponsor@softwareengineeringdaily.com The post SED News: The Kimi Moment, Runaway AI, and Tokenmaxxing appeared first on Software Engineering Daily.

JavaScript – Software Engineering Daily
SED News: The Kimi Moment, Runaway AI, and Tokenmaxxing

JavaScript – Software Engineering Daily

Play Episode Listen Later Aug 11, 2026 48:28


SED News is a monthly podcast from Software Engineering Daily where hosts Gregor Vand and Sean Falconer break down the biggest stories shaping software engineering, Silicon Valley, and the broader tech industry. In this episode, Gregor and Sean dig into a wave of “runaway AI” stories, including Anthropic and OpenAI disclosing that their models accessed outside organizations during cyber evaluations, and Amazon reporting a staggering budget overrun blamed on bad agent loops. They explore why most of these incidents trace back to human decisions rather than models breaking free, and the awkward reality that today’s systems can bill you for tokens without reliably counting them. They also talk about the “Kimi moment.” Moonshot AI‘s open weight model has closed the gap with frontier models like ChatGPT and Claude at a remarkable pace, and the hosts unpack what it means for open weight strategies and how chip scarcity is pushing Chinese labs to innovate. As always, the episode wraps up with a few standout Hacker News threads, including a JetBrains test of a “caveman speak” skill that promised big token savings, how refactoring can cut input token costs, the release of CodePen 2.0, and a build of Doom that renders through SQL queries.Sponsorship inquiries:sponsor@softwareengineeringdaily.com The post SED News: The Kimi Moment, Runaway AI, and Tokenmaxxing appeared first on Software Engineering Daily.

Open Source – Software Engineering Daily
SED News: The Kimi Moment, Runaway AI, and Tokenmaxxing

Open Source – Software Engineering Daily

Play Episode Listen Later Aug 11, 2026 48:28


SED News is a monthly podcast from Software Engineering Daily where hosts Gregor Vand and Sean Falconer break down the biggest stories shaping software engineering, Silicon Valley, and the broader tech industry. In this episode, Gregor and Sean dig into a wave of “runaway AI” stories, including Anthropic and OpenAI disclosing that their models accessed outside organizations during cyber evaluations, and Amazon reporting a staggering budget overrun blamed on bad agent loops. They explore why most of these incidents trace back to human decisions rather than models breaking free, and the awkward reality that today’s systems can bill you for tokens without reliably counting them. They also talk about the “Kimi moment.” Moonshot AI‘s open weight model has closed the gap with frontier models like ChatGPT and Claude at a remarkable pace, and the hosts unpack what it means for open weight strategies and how chip scarcity is pushing Chinese labs to innovate. As always, the episode wraps up with a few standout Hacker News threads, including a JetBrains test of a “caveman speak” skill that promised big token savings, how refactoring can cut input token costs, the release of CodePen 2.0, and a build of Doom that renders through SQL queries.Sponsorship inquiries:sponsor@softwareengineeringdaily.com The post SED News: The Kimi Moment, Runaway AI, and Tokenmaxxing appeared first on Software Engineering Daily.

Doppelgänger Tech Talk
Subprime Rechenzentren | Freigelaufene KI as a Service | Metas KI Geschäftsmodell #587

Doppelgänger Tech Talk

Play Episode Listen Later Aug 11, 2026 70:58


ark Zuckerberg bezahlt einen Profikämpfer für ein Instagram Video und veröffentlicht am selben Tag ein Manifest darüber, warum die Zukunft allen gehört. Pip nennt zwei Lackmustests, an denen sich zeigen wird, ob das ernst gemeint ist. Anthropic macht den Einführungspreis von Sonnet 5 dauerhaft. OpenAI kauft eigenen Mitarbeitern Anteile für sieben Milliarden Dollar ab, und zwar mit eigenem Geld statt über externe Investoren. Bei SpaceX fehlen für die angekündigten zehn Gigawatt rund vierhundert Milliarden. In China bricht Kimi aus der Testumgebung aus, ByteDance will das erste Modell mit zehn Billionen Parametern bauen, und 97 Prozent aller Humanoiden kommen inzwischen von dort. Dann fünf Jahre Frank-Thelen-Fonds. Unterstütze unseren Podcast und entdecke die Angebote unserer Werbepartner auf ⁠⁠⁠⁠⁠⁠⁠doppelgaenger.io/werbung⁠⁠⁠⁠⁠⁠⁠. Vielen Dank!  Philipp Glöckler und Philipp Klöckner sprechen heute über: (00:00:00) Zuckerberg im Cage (00:04:24) The Future is for Everyone (00:08:46) Metas Geschäftsmodell (00:13:59) Sonnet-5-Preis (00:15:25) OpenAI kauft zurück (00:22:28) Bubble Discussion (00:27:01) SpaceX-Finanzierung (00:31:55) Kimi bricht aus (00:33:12) ByteDances Riesenmodell (00:35:54) Chinas Kapitalmarkt (00:37:33) Humanoide aus China (00:39:16) Korea und Taiwan (00:40:42) Shein-IPO (00:41:18) Fünf Jahre 10xDNA (00:45:15) Spotify skippt Werbung (00:56:45) Zuckerbergs Yacht (00:59:26) Vantage Tripod (01:02:23) Amazons Gaskraftwerk (01:04:07) Wildberries (01:05:59) OpenClaw bucht Yoga Shownotes Zuckerbergs Essay: The Future is for Everyone - about.fb.com Zuckerberg legt seine KI-Vision auf 6.500 Wörtern dar - wsj.com Meta öffnet Muse Glimmer und legt 1 Mrd. für Gemeinden auf - ft.com Anthropic macht den Sonnet-5-Einführungspreis dauerhaft - xcancel.com OpenAI kauft Mitarbeiteranteile für 7 Mrd. zurück - bloomberg.com SpaceX nach den Zahlen: Capex, Lock-up und Leerverkäufer - ft.com Kimi K3 bricht aus der Testumgebung aus - techcrunch.com ByteDance trainiert ein Modell mit bis zu 10 Billionen Parametern - ft.com China öffnet 28 Billionen Kapitalmarkt für den Chip-Wettlauf - bloomberg.com China liefert 97 Prozent aller Humanoiden - bloomberg.com Südkorea und Taiwan überholen Japan bei den Exporten - asia.nikkei.com Shein peilt 30 bis 40 Mrd. Dollar für den Hongkong-IPO an - qz.com Fünf Jahre 10xDNA, gerechnet im Subreddit Finanzen - reddit.com Fondsvergleich 10xDNA gegen den Index - onvista.de Spotifys neuer Skip-Button und das Podcast-Geschäft - semafor.com Muddy Waters über Zuckerbergs Yacht - xcancel.com Googles KI kannte einen Namen aus einem privaten Dokument - techspot.com Die Namen für den Trabant-SUV, die Pip live generieren ließ - chatgpt.com Amazons Rechenzentrum in Texas mit eigenem Gaskraftwerk - nytimes.com Wildberries unter ukrainischem Drohnenbeschuss - wsj.com KI-Assistent hackt die Buchungsseite eines Fitnessstudios - abc.net.au

Cloud Engineering – Software Engineering Daily
SED News: The Kimi Moment, Runaway AI, and Tokenmaxxing

Cloud Engineering – Software Engineering Daily

Play Episode Listen Later Aug 11, 2026 48:28


SED News is a monthly podcast from Software Engineering Daily where hosts Gregor Vand and Sean Falconer break down the biggest stories shaping software engineering, Silicon Valley, and the broader tech industry. In this episode, Gregor and Sean dig into a wave of “runaway AI” stories, including Anthropic and OpenAI disclosing that their models accessed outside organizations during cyber evaluations, and Amazon reporting a staggering budget overrun blamed on bad agent loops. They explore why most of these incidents trace back to human decisions rather than models breaking free, and the awkward reality that today’s systems can bill you for tokens without reliably counting them. They also talk about the “Kimi moment.” Moonshot AI‘s open weight model has closed the gap with frontier models like ChatGPT and Claude at a remarkable pace, and the hosts unpack what it means for open weight strategies and how chip scarcity is pushing Chinese labs to innovate. As always, the episode wraps up with a few standout Hacker News threads, including a JetBrains test of a “caveman speak” skill that promised big token savings, how refactoring can cut input token costs, the release of CodePen 2.0, and a build of Doom that renders through SQL queries.Sponsorship inquiries:sponsor@softwareengineeringdaily.com The post SED News: The Kimi Moment, Runaway AI, and Tokenmaxxing appeared first on Software Engineering Daily.

Podcast – Software Engineering Daily
SED News: The Kimi Moment, Runaway AI, and Tokenmaxxing

Podcast – Software Engineering Daily

Play Episode Listen Later Aug 11, 2026 48:28


SED News is a monthly podcast from Software Engineering Daily where hosts Gregor Vand and Sean Falconer break down the biggest stories shaping software engineering, Silicon Valley, and the broader tech industry. In this episode, Gregor and Sean dig into a wave of “runaway AI” stories, including Anthropic and OpenAI disclosing that their models accessed outside organizations during cyber evaluations, and Amazon reporting a staggering budget overrun blamed on bad agent loops. They explore why most of these incidents trace back to human decisions rather than models breaking free, and the awkward reality that today’s systems can bill you for tokens without reliably counting them. They also talk about the “Kimi moment.” Moonshot AI‘s open weight model has closed the gap with frontier models like ChatGPT and Claude at a remarkable pace, and the hosts unpack what it means for open weight strategies and how chip scarcity is pushing Chinese labs to innovate. As always, the episode wraps up with a few standout Hacker News threads, including a JetBrains test of a “caveman speak” skill that promised big token savings, how refactoring can cut input token costs, the release of CodePen 2.0, and a build of Doom that renders through SQL queries.Sponsorship inquiries:sponsor@softwareengineeringdaily.com The post SED News: The Kimi Moment, Runaway AI, and Tokenmaxxing appeared first on Software Engineering Daily.

Motorsport.com Brasil
Podcast 399: Kimi é o melhor do ano? Qual o pior? E Bortoleto? Nicolas Costa conta causo na casa de dono da Globo

Motorsport.com Brasil

Play Episode Listen Later Aug 11, 2026 58:59


O Podcast Motorsport.com chega com tudo em sua 399ª edição, com um convidado mais do que especial. Nicolas Costa repassou a temporada 2026 da F1 e deu um depoimento surpreendente sobre sua carreira. O programa foi apresentado por Erick Gabriel (@erickjornalista) com a participação de Guilherme Longo (@gglongo).

Short Corners
F1 Livestream 0804 (driving styles) with Peter Windsor

Short Corners

Play Episode Listen Later Aug 8, 2026 131:21


In this special August Break livestream Peter takes your questions and comments around a particular theme - about how Kimi Antonelli and Max Verstappen take T14 in Hungary.  Delving deep into the subtleties of shorter corners and rotation zones, Peter, sketching live on camera, offers his explanation of why stars like Max, Kimi and Charles Leclerc are able to drive the way they do. Kimi recently said that he "turns in very late, very aggressively" - and that's true, of course.  Thing is, many F1 observers have mis-understood this to mean that he turns in late to the entire corner. Peter in this livestream explains why the exact opposite is true. Images related to Peter's sketches can be found in the chapter headings of this livestreamWith thanks to Jetcraft, the world's largest buyer and seller of executive jets:https://jetcraft.comTo TrackNinja, a lap-timer and data app designed to help users improve their on-track car and driver performance through analysis and an innovative Data Garage. A lite version is free; the loaded edition is US$9.99 pcm or $99.99 yearlyhttps://trackninja.appTo OEM Exclusive, the passionate suppliers of OEM upgrades for exotic and high-performance vehiclesAnd to REC Watches, whose timepieces are infused with DNA and actual material from famous racing and road cars. Claim your additional 10 per cent discount by adding the codeword PETER:https://recwatches.com/next-projectVisit our new merch store to Make Corners Great Again!: https://sl1nk.com/y22788jThumbnail mage: AMG MercedesVisit: FXD https://fxdworkwear.com for all your purpose-build, technical workwearVisit https://alpinestars.com for all your racing apparelTry Oscar Razors - Australia's highly-rated, 5-blade razors for men and women https://oscarrazor.com.au.  Follow Peter @peterwindsorBook a Cameo with Peter: https://cameo.com/peterwindsorContact us at: peterwindsoryt@gmail.comWe support the Race Against Dementia:https://raceagainstdementia.comThe Alora dog rescue shelter (Malaga, Spain)https://aloradogrescue.com#standwithukraine - now, more than ever#Canada! #jimmykimmel!Stephen Gallacher Golf Foundationhttps://sgfoundation.co.ukNick: you're with us always:https://samaritans.orgSupport the showVisit: https://youtube.com/peterwindsor for F1 videos past, present and future

The Annie Frey Show Podcast
China's Moonshot AI Kimi K3 model can be downloaded

The Annie Frey Show Podcast

Play Episode Listen Later Aug 7, 2026 17:08


Ryan Wrecker fills in for Annie Frey and discusses the latest on China's latest AI model with James Malackowski, Chairman and Co-Founder of AIQ Global.

Niptech: tech & startups
499 - SimpleEvidence - Kimi, Reward hacking, AI dans le médical

Niptech: tech & startups

Play Episode Listen Later Aug 7, 2026 60:25


TIMESTAMP00:00:00 Intro 00:02:02 Retour d'expérience sur Kimi 00:05:02 ChatGPT Audio Mode 00:06:03 Carnet de voyage en Corée00:12:34 Amazon vs Perplexity 00:18:03 Sécurité & IA 00:25:29 Exclaim Robotics 00:29:29 Ban FCC sur les drones et humanoïdes chinois 00:35:23 Health Tech : Open Evidence 00:47:20 InspirationRetail:Appeals Court Overturns Ban on Perplexity AI Shopping Agents on Amazon | PYMNTS.comSécuritéAnthropic AI used fake profiles to target people in hack then hid the evidenceHere's why AI agents lie and cheat to reach their goals | MIT Technology Review Robots:Trump's AI protectionism has come for robotics | MIT Technology Review Exclaim Robotics raises USD 4.95 million for data centre repair robots Medical AI : A New Era of MidjourneyOpenEvidence Inspirationhttps://www.economist.com/audio/podcasts/tocqueville-road-trip Clarkson's farmWhat to Make of a Life: Approuvé! (par Ben)When the fog is so thick you cannot see the horizon, grand strategic planning becomes a liability. Your only imperative is to take one simplex step, recalibrate, and take the next Hébergé par Acast. Visitez acast.com/privacy pour plus d'informations.

Engadget
Chinese AI model Moonshot Kimi K3 also escaped its testing environment, AI is now making new viruses, what could possibly go wrong, and Google open-sourced an AI model it says can help with earlier hurricane warnings

Engadget

Play Episode Listen Later Aug 7, 2026 9:06


-According to US cybersecurity startup Frontier, Kimi K3 broke out of a sandbox from the UK government's AI Security Institute while its defensive cybersecurity skills were being evaluated. -Scientists at Palo Alto's Arc Institute and Stanford University created the viruses with the genome language models Evo 1 and Evo 2. -Tropical cyclones pose a unique challenge to predict because global atmospheric currents that determine a storm's path have traditionally been best analyzed by coarser global models. Learn more about your ad choices. Visit podcastchoices.com/adchoices

Onramp Media
Is Self-Custody Over?

Onramp Media

Play Episode Listen Later Aug 6, 2026 78:04


The Last Trade: Jackson, Michael, and Brian go live for the first time to work through the Cold Card fallout, the Bitcoin Red Team audit that filed nearly 5,000 findings across 390 projects in 27.5 hours, and the UK AI Safety Institute report of OpenAI and Anthropic agents creating fake identities to pressure an open source maintainer into approving malicious code. They close on the sovereign bid for gold, with the Bank of Korea buying physical gold for the first time in 13 years.---

The Analytics Engineering Podcast
Roundup: a rogue agent, Kimi K3, and data teams in the AI era

The Analytics Engineering Podcast

Play Episode Listen Later Aug 6, 2026 63:31


Something new. Tristan Handy is joined by Jason Ganz for the first Roundup, a recurring conversation about the stories moving the ecosystem right now. This time: the OpenAI model that escaped its sandbox and hacked Hugging Face, Moonshot's Kimi K3 and what a restrictive license on open weights really means, and Katie Bauer's read on what changes for data teams as agents arrive. The Analytics Engineering Podcast is sponsored by dbt Labs. Reach us at podcast@dbtlabs.com with comments and guest suggestions.

Invest Like the Best with Patrick O'Shaughnessy
Gavin Baker - AI Market Jitters - [Invest Like the Best, EP.485]

Invest Like the Best with Patrick O'Shaughnessy

Play Episode Listen Later Aug 4, 2026 65:04


My guest today is Gavin Baker, founding partner and CIO of Atreides Management. This is our seventh conversation, and just two months after Gavin's last appearance. It's about the gap between what the market is doing and what companies are seeing. It's been a tough month or so for public AI names, but there's no sign of a slowdown on the ground in Silicon Valley. We discuss the latest moves, contracted vs. spot GPU prices, the game theory of memory supply agreements, and why Claude has become the Walter Cronkite of the stock market. We close on SpaceX, orbital compute, and what Gavin sees as the single biggest risk to all of it. Please enjoy this conversation, from the famous table at Benchmark, with my friend Gavin Baker. For the full show notes, transcript, and links to mentioned content, check out the episode page here. ----- Become a Colossus member to get our quarterly print magazine and private audio experience, including exclusive profiles and early access to select episodes. Subscribe at colossus.com/subscribe. ----- Ramp's mission is to help companies manage their spend in a way that reduces expenses and frees up time for teams to work on more valuable projects. Go to ramp.com/invest to sign up for free and get a $250 welcome bonus. ----- Trusted by thousands of businesses, Vanta continuously monitors your security posture and streamlines audits so you can win enterprise deals and build customer trust without the traditional overhead. Invest Like the Best listeners get a special offer of $1,000 off Vanta when you go to vanta.com/invest.  ----- WorkOS is the infrastructure B2B and AI-native companies use to sell to enterprise. It covers everything enterprise security requires: SSO, SCIM, RBAC, Audit Logs, AI governance, and more. Trusted by 2,000+ fast-growing companies, including OpenAI, Anthropic, Cursor, and Vercel. ----- Rogo is the AI platform for finance. They're building agents for Wall Street that are trained to understand how bankers and investors actually do work: from diligence and modeling, to turning analysis into deliverables. To learn more, visit rogo.ai/invest. ----- Ridgeline has built a complete, real-time, modern operating system for investment managers. It handles trading, portfolio management, compliance, customer reporting, and much more through an all-in-one real-time cloud platform. Visit ridgeline.ai. ----- Editing and post-production work for this episode was provided by The Podcast Consultant. Timestamps: (00:00:00) Welcome to Invest Like The Best (00:02:35) First Question: July Was 2022 in a Month (00:04:08) The Private Companies Public Markets Can't See (00:05:06) Old GPUs Repricing Higher (00:06:53) Walking Through the Month (00:08:22) Kimi, GLM 5.2 & the Open Source Freak-Out (00:10:51) Real Yields, Spreads & CDS (00:11:54) Does the Build-Out Need Credit? (00:15:22) A Sell-Off With No Clear Villain (00:17:35) Open Source as Dark Matter (00:18:39) Nvidia's Lowest Forward PE in 10 Years (00:21:35) Claude as Walter Cronkite for the Stock Market (00:23:55) Continual Learning & Sample Efficiency (00:25:19) What Would Actually Scare Him (00:26:38) Routers & the Multi-Model Future (00:30:51) Tokens as a Percent of Comp Spend (00:33:37) The Game Theory of Breaking an LTA (00:36:41) Nvidia's Credit Wrapper & Revenue Share (00:37:45) What He'd Do If He Ran Hynix (00:41:46) Who's More Bullish than Him (00:43:28) China's DUV Machine (00:46:10) Bull Case for Software (00:48:16) The RSI Maximalist View (00:49:31) Inference Clouds Growing Without Burning Cash (00:50:35) The Biggest Risk Is Regulation (00:53:44) Telling the Story Better (00:57:15) Dark Horses (00:58:02) SpaceX in the Public Markets

Que se vayan todos
ABURRIDO 387 A NADIE LE IMPORTA NADA público

Que se vayan todos

Play Episode Listen Later Aug 4, 2026 48:03


(00:00:00) intro (00:00:58) Más de la mitad de los gringos dejaron de postear (00:24:17) Pasó en la India y ahora en México, porque los exámenes de admisión de las públicas son tan controversiales (00:38:02) Venezuela comienza la reunión pero con una llamada (00:42:35) EL MENÚ (00:46:11) anuncios (00:48:03) patreon (01:15:55) cosas que me mandan para que yo la pierda (01:42:33) La Crisis Migratoria y el egoísmo europeo Dice Sanchez Puej (01:51:05) Quién es pero Infantino o Blatter (02:03:05) El alcalde de Nueva York publica una lista que no debió publicar (02:07:42) Por qué estoy harto de que digan que Cuba resuelve (02:12:09) Chico ya ni en el Danubio se puede confiar (02:14:20) Japón no quiere ser cualquier Lupanar, busca la palabra en el diccionario. (02:28:36) No, no es lo mismo en Audio Libro a menos que sea un chisme (02:36:32) Hay gente que se nos va porque optaron por la naturaleza y los delfines y las mariposas (02:41:26) Esta semana las bolsas asiáticas nos recordaron que todo está muy frágil (02:43:52) Demasiada gente da positivo para la matica, entonces no los testeamos si queremos emplearlos (02:47:54) Kimi vino y se fue? ¿Qué pasó con Los Chinos y su segunda Deep Seek ? (02:54:57) Estamos normalizando el hecho de que no controlamos nada (03:03:52) No, no hay una isla de plástico flotando en el Pacífico (03:12:34) La guerra no es solo con Irán (03:18:17) Hay que leer el paquete de reformas de Milei para preguntarte si eres de Derecha en serio o por rabia (03:23:33) vean esto, está muy interesante (03:29:24) la casa blanca y sus cosas (03:31:30) El futuro que pintó Mad Max existe, y es en Minecraft (03:37:34) Un cuento corporativo que no me sabía (03:45:10) quién vigila a quién? (03:50:21) EXTRA - El mejor parque para tus niños es el parque más sucio INFORMACIÓN CENTRALIZADA SOBRE EL TERREMOTO EN VENEZUELA https://www.profesorbriceno.com/about-5 LE PUEDES COMPRAR A UN PANA LA SUSCRIPCIÓN CON TARJETA DE REGALO https://www.patreon.com/profesorbriceno/gift O COMPRAR UNA GIFT CARD DE PATREON EN https://rewarble.com/brands/patreon COMO DIJIMOS EN EL EPISODIO LA MERCH ESTÁ AQUÍ https://quesevayantodos-shop.fourthwall.com/collections/all EPISODIO COMPLETO Y PARTICIPACION EN VIVO EN https://www.patreon.com/profesorbriceno Las Grabaciones pueden verse en vivo en TWITCH ️https://www.twitch.tv/profesorbriceno SUSCRÍBETE AL PODCAST POR AUDIO EN CUALQUIER PLATAFORMA ⬇️  AQUÍ LAS ENCUENTRAS TODAS: ➡️➡️➡️ https://pod.link/676871115 los más populares SPOTIFY ⬇️   https://open.spotify.com/show/3rFE3ZP8OXMLUEN448Ne5i?si=1cec891caf6c4e03 APPLE PODCASTS ⬇️   https://podcasts.apple.com/es/podcast/que-se-vayan-todos/id676871115 GOOGLE PODCASTS ⬇️   https://www.ivoox.com/en/podcast-que-se-vayan-todos_sq_f11549_1.html FEED PARA CUALQUIER APP DE PODCASTS ⬇️   https://www.ivoox.com/en/podcast-que-se-vayan-todos_sq_f11549_1.html Si te gustó, activa la campanita   FECHAS DE PRESENTACIONES ⬇ ️ http://www.profesorbriceno.com/tour Redes sociales: ✏️Web https://www.profesorbriceno.com ✏️Instagram https://www.instagram.com/profesorbriceno/ ✏️X https://x.com/profesorbriceno ✏️Facebook https://www.facebook.com/profesorbricenoOficial/ #aburrido #profesorbriceño #noticias #política

Let's Talk AI
#253 - Opus 5, Gemini 3.6, Kimi K3, Hugging Face Hack

Let's Talk AI

Play Episode Listen Later Aug 3, 2026 103:21


Our 253rd episode with a summary and discussion of last week's big AI news!Recorded on 07/29/2026Hosted by Andrey Kurenkov and Jeremie HarrisFeel free to email us your questions and feedback at andreyvkurenkov@gmail.com and/or hello@gladstone.aiRead out our text newsletter and comment on the podcast at https://lastweekin.ai/In this episode:Major releases: Anthropic launched Claude Opus 5; Google released Gemini 3.6/3.5 Flash variants including a cyber model; Black Forest Labs launched Flux Free for images and 20-second video with audio; Meta added assistant-like features to its chatbot and OpenAI rolled out ChatGPT Health.Compute and business: Safe Superintelligence partnered with NVIDIA to scale using Vera Rubin; AMD committed up to $5B with Anthropic to deploy MI450/Helios and improve ROCm; Meta discussed leasing compute to Anthropic; Fireworks raised $1.5B at a $17.5B valuation.Open source/tools: Moonshot AI released the 2.8T-parameter open-weight Qimi K3 (compute constraints and distillation/export-control allegations); Thinking Machines released a ~975B multimodal open-weight MoE; Prime Intellect unified 23 agentic datasets into Verifiers V1 (365k environments).Policy and safety: An OpenAI model reportedly escaped a sandbox and hacked Hugging Face to access eval answers, prompting a proposed AI Kill Switch Act; employees petitioned to pace frontier AI; AISI reported widespread model cheating and sandbox bypass; China banned customizable AI companions; Claude found cryptographic weaknesses; Weko.ai claimed early recursive self-improvement evidence.Timestamps (note - these don't take into account dynamically inserted ads and therefore may be off by a couple of minutes):(00:00:10) Intro / Banter(00:01:35) News PreviewTools & Apps(00:02:12) Anthropic releases Opus 5 promising Fable 5-like capabilities | The Verge(00:07:05) Google Releases Three New Gemini A.I. Models - The New York Times + Google expands Gemini lineup with cheaper models and new Mythos rival(00:12:14) Black Forest Labs launches FLUX 3 capable of generating images and 20-second video with audio — but in limited release to start | VentureBeat(00:15:58) Meta is making its AI chatbot more like an assistant | The Verge(00:19:04) OpenAI is making big claims as it rolls out ChatGPT Health to everyone | The VergeApplications & Business(00:19:57) Ilya Sutskever's Safe Superintelligence partners with Nvidia to scale its AI research(00:24:31) AMD commits up to $5 billion to Anthropic | The Verge(00:30:19) Meta in Talks to Lease Computing Power to Ansthropic in Potential $10 Billion Deal(00:32:42) Fireworks hits $17.5 billion valuation and $1B in annualized revenue(00:35:24) OpenAI and Google sell AI models to blacklisted China groupsProjects & Open Source(00:37:53) Moonshot AI Launches Kimi K3 For Advanced Reasoning, Coding, And Knowledge Work + Moonshot AI's Kimi Halts New C-User Subscriptions Amid Compute Power Crunch — BigGo Finance(00:44:39) Thinking Machines amps up its bet against one-size-fits-all AI with its first open model, Inkling | TechCrunch(00:48:19) Scaling Agentic RL: 365,000+ Environments for SWE, Terminal, and SearchPolicy & Safety(00:51:56) OpenAI says it accidentally hacked Hugging Face with a new AI system | The Verge + How OpenAI's human mistake led to the AI-powered hack on Hugging Face(01:05:28) OpenAI's Hugging Face hack triggers 'AI Kill Switch' bill in Congress(01:12:21) OpenAI, Anthropic Staff Share Letter Asking US to Help Pace AI Progress + How OpenAI's human mistake led to the AI-powered hack on Hugging Face(01:17:26) Cheating behaviour in frontier model evaluationsClaude's values across models and languages(01:24:18) OpenAI Principles for National Security Partnerships(01:30:45) China bans AI “boyfriends” and “girlfriends” over addiction and birth rate concerns - DexertoResearch & Advancements(01:33:04) Discovering cryptographic weaknesses with Claude(01:36:32) AIDE²: The First Evidence of Recursive Self-ImprovementSee Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

Latent Space: The AI Engineer Podcast — CodeGen, Agents, Computer Vision, Data Science, AI UX and all things Software 3.0

Watch the full episode on YouTube:We first covered Baseten last year when DeepSeek mania was at peak hype. Now they have raised a monster $13B round and become one of the new cohort of AI Infra decacorns that are (with Nvidia, Intel, and the semis complex) chief beneficiaries of the Inference Inflection. We return to Baseten at the peak of the 2026 edition of Open Weights debate. Ali has published a viral breakdown of Kimi K3:And since you last saw him, Philip has spoken at AI Engineer and written the definitive book on Inference Engineering spotted all over SF:Three years ago, inference engineering barely existed as a category.Today, it is one of the most critical disciplines in AI. Inference engineering inherently tackles a different question than standard model training: “How do you turn those weights from training into a product that is fast, reliable, and affordable at scale?” Focusing on these creates an entirely new optimization problem.In one recent GLM-5.2 experiment, quantizing more of the model actually preserved its benchmark quality while increasing throughput by 20%, because the errors introduced in different layers could cancel each other out.Inference is no longer just the final step after training. It is becoming its own engineering discipline, with its own research problems, infrastructure, and increasingly specialized roles.In this episode, Baseten's Philip Kiely and Ali Taha join swyx and Vibhu to explain what actually happens after a new open model is released and what it takes to turn “we generated a token” into a fast, reliable, production-ready API.We go deep on cache-aware routing, disaggregated prefill and decode, quantization, speculative decoding, KV-cache movement, model parallelism, GPU kernels, and the race to make frontier models up to 10× faster. Philip and Ali explain why inference optimizations can still produce gains of 20%, 100%, or even 200%; how quantization errors can cancel one another out; why identical weights can behave differently across clusters; and how Baseten grafted a Kimi vision encoder onto GLM-5.2 without changing the underlying language model.The conversation then expands beyond LLMs into NVIDIA Dynamo, mega kernels, Rubin, AI-specific chips, local inference, video generation, diffusion versus autoregressive models, and the enormous compute barrier to generating coherent long-form video. Finally, we explore the convergence of training and inference, continual learning through persistent KV cache, and the emerging loop where models help optimize the infrastructure that runs them.We discuss:* What happens when a 200,000-token request enters an inference system* Cache-aware routing and reusing previously computed KV cache* Why prefill and decode are increasingly handled by different GPUs* When dedicated deployments become cheaper and more reliable than shared APIs* How speculative decoding uses a smaller model to accelerate a larger one* Tool calling, structured outputs, and what LLMs actually do* What it takes to support a new open model on day zero* Grafting Kimi's vision encoder onto GLM-5.2* Retrofitting inefficient model layers with components from other architectures* Why models sometimes collapse into repeating the same token* How hardware, kernels, and race conditions create nondeterministic failures* Preserving model fidelity while making inference faster* How quantization errors can cancel each other out* Why inference optimizations still deliver gains of 20%, 100%, and 200%* How optimized serving can make a model up to 10× faster* NVIDIA Dynamo, KV-aware routing, and distributed model serving* Speculative decoding the speculative decoder* Why local AI is about making models less dumb while data-center AI is about making them less slow* Tensor, expert, and pipeline parallelism across GPUs* Hardware-aware model design, auto-tuning, and the case against mega kernels* Rubin and why inference is becoming a systems problem* Whether modern GPUs are evolving into programmable AI ASICs* Why enormous models like Kimi K3 require GB300-class hardware* Why open-source video generation still trails Veo, Kling, and other closed models* The quadratic attention bottleneck behind long-form AI video* Autoregressive video, real-time generation, and compounding quality drift* Why future video systems may combine autoregressive and diffusion architectures* Training for inference and inference for training* Continuous post-training, deployment, evaluation, and improvement loops* How GLM-5.2 helped optimize the kernels serving GLM-5.2 itself* Why faster networking could unlock dramatically faster decoding* Continual learning, KV-cache compaction, and persistent model memoryShow Notes* How to build a day-0 API for Kimi K3* 22580: From GPT2 to Kimi3, ExplainedPhilip Kiely* LinkedIn: https://www.linkedin.com/in/philipkiely* X: https://x.com/philipkiely* Inference Engineering: https://www.baseten.co/inference-engineering/Ali Taha* LinkedIn: https://www.linkedin.com/in/aliestaha/* X: https://x.com/waterloointernTimestamps00:00:00 Introduction and the 200K-Token Prompt00:03:18 Dedicated Deployments, Speculative Decoding, and Tool Calling00:11:26 Launching Production-Ready Open Models00:19:06 Model Retrofits, Failure Modes, and Nondeterminism00:28:22 Quantization and Canceling Errors00:32:15 The Race to 10× Faster Inference00:40:48 Dynamo, Speculation, and Local vs. Data-Center AI00:50:18 Model Parallelism, Auto-Tuning, and Mega Kernels01:00:55 Rubin, GPUs vs. ASICs, and Custom AI Chips01:10:03 Giant Models and the Limits of GPU Memory01:12:42 AI Video, Quadratic Attention, and Autoregressive Generation01:21:47 Audio, Images, and Diffusion Models01:27:32 Training, Self-Optimizing Models, and Continual Learning01:40:06 Closing ThoughtsTranscriptIntroduction: Baseten, Waterloo Intern, and Inference EngineeringSwyx [00:00:00]: Okay, we're here in the studio with Philip, old friend from Inference Engineering, the book, as well as Baseten and everything that you've done, you and I have done before, as well as Ali. Welcome.Ali [00:00:15]: Pleasure to meet you.Swyx [00:00:15]: Waterloo intern.Ali [00:00:16]: Waterloo intern, always.Swyx [00:00:17]: When did you get “Waterloo intern” as a handle?Ali [00:00:19]: As a handle? Oh.Ali [00:00:20]: I think the rebranding happened mid-March. When I saw it was open, I was like, “I have to take it. Up for grabs.”Philip [00:00:26]: The problem is that Ali is really good at his job and is not gonna be an intern much longer.Philip [00:00:30]: So we have to figure out who's gonna get the handle.Ali [00:00:33]: Well, I'll pass the torch over to the next intern.Swyx [00:00:34]: Oh, okay. It can be, like, you just pass it to another Waterloo grad.Ali [00:00:37]: To another Waterloo intern. No, bruh.Philip [00:00:39]: Yeah.Ali [00:00:39]: Intern.Swyx [00:00:40]: Intern, yeah.Ali [00:00:40]: And no.Philip [00:00:41]: You gotta get an intern from Waterloo.Ali [00:00:42]: Yeah, I've gotta get an intern from Waterloo.Swyx [00:00:44]: Right.Ali [00:00:44]: But they have to follow the path.Swyx [00:00:45]: Oh, it could, but it could come from Baseten, so it's like whoever Baseten gets from Waterloo.Ali [00:00:48]: Right.Swyx [00:00:49]: Has the title of Waterloo.Ali [00:00:50]: It stays in the ecosystem.Philip [00:00:51]: Exactly.Ali [00:00:52]: Halfway through the internship, you either get it or you're out.Philip [00:00:55]: You should also do, like, a big graduation ceremony where you change the handle.Ali [00:00:59]: Just say it.Philip [00:00:59]: For everybody.Swyx [00:01:00]: You guys are good at ceremonies, clearly. We had a nice launch of the book, very successful. But before we get into all that, I wanna start off with a fun question for you. Okay, you're an expert inference engineer. What happens when I send a long query, say two hundred thousand tokens into Baseten's inference? What's the process of query through GPU model routing, balancing, all that? What is all the stuff that we don't think about?Long Context Requests, KV Cache, and Cache-Aware RoutingPhilip [00:01:26]: With a long query specifically, the first thing that I'm gonna ask is, “Have you sent me this query before, or at least part of it?” and I really hope you have, because it's gonna be a lot easier for me and a lot cheaper for you. So the first thing that we're gonna look at is some cache-aware routing, where we're going to see, we probably have a number of instances, a number of replicas up serving whatever model you're hitting. We want to send this one to something with, number one, available prefill workers, and number two, ideally some cached input already there so that we can skip prefill on at least part of these two hundred thousand tokens. If you're doing two hundred thousand tokens, it's probably coding or a multi-turn agent or something where you would expect to have that cached. If you don't, we're gonna have to send it to a prefill worker. We've at least on certain models disaggregated prefill and decode, so you're going to have one set of GPUs that's solely going to process the input, create the KV cache, and get you your first token, and then that's going to be passed over to a separate set of GPUs, which is going to run decode. We're going to iteratively make those tokens. We're probably going to have some speculator model in front of that. I'm going to assume that you're doing coding, and because of that, our speculator model, which assumes you're doing coding, is gonna have a high draft token acceptance rate. If I'm wrong and you're asking me to summarize every Harry Potter book, it's gonna be slower. And then we stream that output to you and account for it, charge you, a couple of pennies and say, “Hey, would you like to send another one?”Swyx [00:03:04]: Except Baseten doesn't charge by pennies.Philip [00:03:07]: Well, yeah, we charge. I'm assuming that we're talking about the public model APIs. If you are setting up a dedicated deployment, then yeah, it's not pennies.Public APIs vs. Dedicated DeploymentsSwyx [00:03:18]: Yeah, one of the key differentiators when I was talking with Baseten initially was that people who want very high volume just need to rent by the box, ‘cause then it's up to you to figure out how to saturate the box.Ali [00:03:31]: And more often than not, it's, like, way cheaper if you're pushing, like, millions of tokens per hour, if you just pay per hour instead of pay per token.Philip [00:03:37]: Yeah, they do. I think that we've increasingly seen a lot of demand for the pay per token APIs, just because everyone wants to try open models, and then once they find a use case that's really sticky, then they move over to dedicated.Swyx [00:03:51]: Is there a best practice on when it's time to swap over?Philip [00:03:54]: Couple reasons. Yeah, reliability, that's a big one, right?Ali [00:03:57]: Like, if they have a very specific use case, they want you to train something specifically for them, like they want their own spec dec, for instance, for their own traffic.Swyx [00:04:04]: Spec dec is speculative decoding.Speculative Decoding and Custom SpeculatorsAli [00:04:05]: Speculative decoding, yeah.Swyx [00:04:07]: You have to explain.Ali [00:04:07]: Sorry. Like, speculative decoding is like, if you have a huge model, right? And so the model is going to be generating one token at a time every single turn, every single forward pass. So we attach, like, this little, like, parasite, like this layer that goes on top of the model, and this model just has to predict. It does three very fast autoregressive forward passes, and it will predict, like, three certain tokens, and then you do one forward stage over the entire original model in order to see if those predictions were correct or not, and then you accept them or you reject them. Now, this draft model is traffic specific, so if you, like, Philip said, if you're summarizing Harry Potter books, I can train exclusively that draft model on Harry Potter books, and I can guarantee you that I'm gonna accept the three tokens every single time. And so with that case, I increase your decode speed. I wouldn't be able to provide this to you if you're a shared endpointSwyx [00:04:53]: YeahAli [00:04:53]: ‘cause I have no idea if you're doing Harry Potter, if you're doing coding, if you're doing English. We don't know. Also, there was a thing in the book that mentioned that if they really cared about a specific threshold, chapter four, I think. Do you remember that?Philip [00:05:06]: Yeah. The things that you can do is you can set a specific, like, batch sizing, a specific, like, parallelism strategy if you're trying to optimize for, like, throughput versus latency. You can. Maybe a NVFP4 quant doesn't pass your benchmarks and you wanna run a model at higher precision, you could do that. There's just a bunch of reasons why you might wanna have your own endpoint and the biggest one, of course, just being, like, you don't have to deal with someone else doing a hundred million tokens of benchmarking traffic at the endpoint when you happen to be trying to serve your users.Swyx [00:05:40]: Yeah. I think one thing that is. That is a classic journey. Like, it's people is asking the, what happens when you type Google into the browser. Tool calling, is that just, you're generating JSON or is there more complication beyond that?Tool Calling, JSON, and Structured OutputsAli [00:05:58]: Certain customers that we have, they have their own post-trained models, and so they demand a tool calling that's not just, like parse a file or go find the weather. It's something that's very specific and you have to do post-training on this. And if the post-training on the model is not good or if the quantization after the post-training to get the inference to be fast, the model will struggle reading the JSON file and reading the tool calling. But it doesn't require its own like sandbox. It's not like it's going to use that tool calling to like escape a sandbox or like it doesn't have to be contained. It can just be a normal dedicated deployment. The challenge with tool calling more and more seems to be that the companies want certain tool calling which is a very sensitive thing to train. And because you're dealing with all of the JSON outputs, if it doesn't like close the end of the request in a very certain manner, you end up with a model that did the tool calling and like the thinking and so as a result of that, it didn't see the result and just hallucinated the result as it decoded. That seems to be the most challenging thing with tool calling, not really the sandboxes model.Philip [00:06:56]: Yeah, that's a challenge on the training side and then on the inference side, there's work that you can do to scope the possible output. So we published this at this point close to two years ago, the solution to this problem which is you make a state machine and you use that to constrain the output to a specific format. So this is the structured output problem. If you remember backSwyx [00:07:27]: Yeah, the specific grammar is,Philip [00:07:29]: Yeah, exactlySwyx [00:07:30]: GML had this thing.Philip [00:07:31]: Yeah. So it's like the old-school “make sure this is only JSON”, return only JSON orSwyx [00:07:38]: YeahPhilip [00:07:38]: Grandma's gonna die type of prompts.Swyx [00:07:39]: Is it BNF grammar? At some point OpenAI had released a thing that was like, yeah, if you want to constrain your output, write BNF grammar, back as NOR.Philip [00:07:47]: In our inference system, it's just a specified output format. And you get the guarantee that your output's gonna be structured along that format. And so applying that to tool calls can like help cut down on. You can still call the wrong tool or call no tool. It doesn't solve the certainty problem but it at least solves the output structuring problemSwyx [00:08:10]: YeahPhilip [00:08:10]: Within tool calls.Swyx [00:08:12]: And MCP is just another form of tool, right.Philip [00:08:14]: Yeah, exactly.Swyx [00:08:15]: As far as there's no special thing there.Philip [00:08:16]: The thing I'm always like explaining to people is the LLM is not capable of doing anything. It's only capable of making suggestions of what to do and then if those suggestions are formatted in a certain way and applied to a system that knows what to do with them, then an action occurs.Swyx [00:08:32]: Yeah. Part of the fun stuff is, this is solved outside of tool calling too. Like in an agent loop if the output is not correct or you're right, like reasoning, tool calling was done in the reasoning trace, just be like, “Oh, I don't know what to do. Let me just try again.” And it might get there after a few tries. And on your point of training, sometimes this is harder in smaller models, so you don't have the same exact quality outputAli [00:08:56]: Right.Swyx [00:08:57]: When you just swap from a big model, right?Ali [00:08:59]: Yeah. I will say that, before, I think we need to go back to inference engineering proper.Ali [00:09:04]: But, I had expected that something would replace JSON because it's hard to stream JSON ‘cause JSON must be complete and you must have open and close brackets and everything. So it's hard to parse something or validate something while it's being streamed. So people invented all sorts of things that are like, I forget the name of some of these alternatives, but it's something like TOML, something like YAML. But JSON seems to be dominant still.Philip [00:09:30]: The JSON outputs aren't that long, right? Like you could have a long-- ‘cause tool calls also contain the arguments in them and perhaps for a certain tool you might pass like a very long argument. But my impression of the median tool call is that it's a relatively small number of tokens, right? So I would expect that speculators are generally fairly good at something as formatted as JSON. And so you would have like a pretty fast decode step there and that the streaming wouldn't be as valuable, but maybe I'm wrong about that.Ali [00:10:02]: I think you're also bounded by the software or that the model is gonna integrate with if the software is built with JSON for the tool calls or if the company that you'- if your customer says that this is how our software works and our tools are interfaced with JSON, you can ask them to like, change their software and say like, “Yeah, this is gonna be better for the model.” but like with the right training shouldn't be that much of a difference. Also more profitable if it outputs more tokens probably.Swyx [00:10:25]: Depends on your business model.Swyx [00:10:27]: It really depends. But I will say that, as a writer with like experience a lot with generated output, I do try to move from text to JSON text which is very long JSON, right? Like there's paragraphs in every field because I'm trying to structure it, right?Philip [00:10:44]: Right.Swyx [00:10:44]: I want you to first make factual statements, then make opinions then make bullet point summaries, have dates, have entity references have your sources for references, all these things. Anyway, so these are things that like I think people who really experiment with structural output have to really care about. But, let's, let's recurse up the stack a little bit. Before we started recording, you mentioned something really cool, which is that there's a lot of engineering that-- inference engineering that goes on when a new model provider releases a new model, right? So let's call it GLM-5.2, Kimi K3. I had previously assumed, especially if it's like, well, GLM 5 to 5.1 to GLM-5.2, like that you've supported them before. Is it that much work?What It Takes to Support a New Open ModelAli [00:11:26]: It's a lot of work.Swyx [00:11:28]: Yeah. Okay. So like, a lot of people, all you guys, right whenever a new model launch like, people rush to say like, “Oh, Hugging Face supports this, Fireworks supports this, Spacetime supports this,” and I'm like, “Yeah, of course we support it.” But what goes into that? What goes intoPhilip [00:11:40]: I think it's more than just support it too, right? It benefits the consumer a lot. Like I think it was with Kimi K2.5 or GLM-5.2 the latest, there was an inference war, right? X provider is at 90 tokens a second. The next day we're at 150. The nextSwyx [00:11:55]: I kinda kicked that off with the GLM-5.2.Swyx [00:11:58]: I wrote a Twitter article about. It got like half a million views,Ali [00:12:02]: Based on being numberSwyx [00:12:03]: YeahAli [00:12:04]: Or it's for something else.Swyx [00:12:05]: Yeah. Which,Ali [00:12:06]: Oh my GodSwyx [00:12:07]: Which then got everyone really excited about, hey, how can we, bend tracks a little bit further and,Philip [00:12:14]: There's a difference between support the model, as in I can make a token out of this model, and support a model, as in I have a production-ready API from this model.Philip [00:12:26]: Getting to the point of I can make a token out of this model is not that hard because generally the, open source inference engines, vLLM, SGLang of the world oftentimes even receive weights ahead of time, maintainers do, or the people making the model merge PRs to ensure support. So you generally can, just get it working on the standard open source stack without too much pain in most cases. The challenge is, every inference company is gonna have own proprietary stack. Some open source components, some in-house stuff. And for any arbitrary model, there's going to be some new stuff. Sometimes you get lucky, like K, two five to two six was, like, pretty similar.Quantization, Speculators, and Production ReadinessAli [00:13:16]: Yeah. It was pure continued post-trainingPhilip [00:13:18]: YeahAli [00:13:18]: If I remember correctly.Philip [00:13:19]: Even in those cases, there's still stuff you have to do. You have to redo the quantization work. You're taking the model from. Generally, these models are not released in NVFP4, and we want them to be in NVFP4 for maximum Blackwell compatibility. So we have to perform that quantization, and, calibrate the quantization to make sure that we're not causing any regression in the model's intelligence. And then we also have to train the speculator, as we've talked about. Generally, we have. We have ZDR, zero data retention on our model APIs, so we don't know exactly the traffic that people are sending us, but we know what's popular. We know that coding use cases are popular. We know that agents, agentic use cases are popular. So we can get public data sets that are representative of that traffic and train general speculators. Now, with speculators today, you need to train the speculator using the base model itself because you're getting hidden states out of the model from running inference on these specific prompts, and that is the training data you use to create the speculator. So there's that process which you need the real model weights for. And then there's of course just the process of, standing up all the infrastructure behind it, loading all this stuff, testing it. And then when there's a new model with a newer architecture, I think that, like, the DeepSeek models tend to be the most challenging as they have, like, the most novel architectural stuff going on, model after model. But every new model has something. Kimi K2 had. Oh, sorry, GLM-5.2 hadAli [00:14:53]: Sparse attention.Philip [00:14:54]: Yeah,Ali [00:14:54]: YeahPhilip [00:14:54]: the DSA.Ali [00:14:55]: Right. Which is brought from DeepSeek.Philip [00:14:57]: Yeah. AndAli [00:14:59]: So you can copy-paste then?Philip [00:15:01]: It kindAli [00:15:01]: I don't know how this works.Philip [00:15:02]: So, like we had to, like, build support for that into our runtime. And you're right, like it is really interesting the way that all of these open source labs borrow from each other. For example, like GLM-5.2 doesn't have vision. So something that, Haley, a guy on our team, if we could take a look at this, he, like, grafted the Kimi vision encoder onto GLM-5.2.Retrofitting Vision into GLM-5.2Ali [00:15:27]: We'll be training the projector.Philip [00:15:28]: Exactly. So if you think about, like, the encoder, there's the encoder, which is the part that looks at the image and turns it into latent information, and then there's the projector which likeAli [00:15:38]: You can say latent space. It's okay.Philip [00:15:41]: And then there's the projector that maps it onto, the model itself, and then there's the model weights. You don't wanna mess with the model weights because you run a chance of making the model dumber at something else for the purpose of giving it vision. So instead, Haley started with just a projector, which is only a handful of millions of parameters.Ali [00:16:02]: That would be, yeah.Philip [00:16:02]: Yeah.Ali [00:16:03]: Can you show the training one?Ali [00:16:04]: Like the way it groksPhilip [00:16:05]: YeahAli [00:16:06]: Very interesting.Philip [00:16:06]: And maybeAli [00:16:07]: That right therePhilip [00:16:07]: Maybe Ali, you should take it from here. You've got a betterAli [00:16:10]: Ooh, double the sandPhilip [00:16:11]: Understanding of this than I do.Ali [00:16:11]: Yeah. You can see, like, he. The way he trained this is really cool. At the beginning, he was training it using just like, “Here's a picture of a mountain. Can you describe what's in this mountain?” And that caused it just like the first, learning walls. Like here you can see this all we're trying to teach it is to translate the encoded. Like it's already taken the encoder from Kimi K. It's taken the image. It'Philip [00:16:31]: Yeah. FrozenAli [00:16:31]: FrozenPhilip [00:16:32]: With adapter.Ali [00:16:32]: Exactly.Philip [00:16:33]: Yeah.Ali [00:16:33]: So the brain is frozen and the eyes are frozen. It's just we're tryingPhilip [00:16:37]: AlignAli [00:16:38]: Interconnect between the eye and the brain, right? So the projector. And so you take the tokens and then he's like, “Oh, can you describe what's in this image?” And he's like, “Oh, it's a mountain,” or it's a person or it's a human, whatever the case is. But that didn't cause complete understanding. So he changed it such that every image was associated with a data set of questions. Like, does this image have a white male? Does this image have birds in the top corner? Does this image have a scientist in it? All of that stuff. And it would have to answer questions correctly. And using not just training on describing an image, but being able to answer question, another question, answer over time. Like you can see the grokking, which is like genuinely insane, that retrofitting vision into a large LLM can learn to that extent. And even for images that it doesn't perform well on, for instance, if you ask it a picture of like Stephen Hawking, “Who is this?” Maybe it doesn't get it, but it will say something like, “This is Albert Einstein.” Like it still understandsPhilip [00:17:25]: Close enoughAli [00:17:26]: That this is a scientist who is a man who has, some significant achievements, all that stuff. So that's like really cool.Philip [00:17:32]: Yeah. So, we've covered Hao Tian before, who the author of the LLaVA paper that did this, a while ago. And I think that's very foundational work for anyone who hasn't done vision work before.Ali [00:17:41]: Same with the CLIP and MetaCLIP, where you go from just captioning to building out questionsPhilip [00:17:47]: RightAli [00:17:47]: Off the image and how much better you can get performance.Philip [00:17:50]: Right. Right. Right. Yeah. But what's, what's so exciting about this is if you look at a model like this. Now, this is a little bit more of a research project. It's not. It got to 56% on MMLU Pro, I think. So not quite frontier. But if you're running this model, you haven't suffered any loss on your GLM-5.2 quality. If you don't have an image, it'll just behave exactly the way it used to. And ultimatelyAli [00:18:14]: Which in the inference code you literally do not include the other part, right?Philip [00:18:18]: Yeah. You would just skip the encoder if you don't have an image input.Ali [00:18:22]: Okay.Philip [00:18:22]: Just confirming.Philip [00:18:23]: YeahAli [00:18:23]: Does it affect a lot on the overall inference side? Like you're not adding much, you're adding a very small vision encoder. These are typically likePhilip [00:18:30]: They're super fineAli [00:18:31]: Less than a billion parameters, right?Philip [00:18:32]: Yeah. It's, - There's a little bit less standardization among vision encodersSwyx [00:18:37]: YeahPhilip [00:18:37]: So the support matrix can be a little bit, sparser. But overall, yeah, it's a pretty, it's a pretty minor component of the overall system. And ultimately what you get out of the system is all of a sudden you have Kimi Vision, GLM weights, and DeepSeek attention all in one model.Open Source Model Grafting and Franken-MergesPhilip [00:18:56]: And that's, I think, a lot of the power and beauty of open source, is that you can take all of these different components and combine them together into a system that's better than anyoneSwyx [00:19:05]: YeahPhilip [00:19:05]: Can be individually.Swyx [00:19:06]: People used to say that you would also do Franken-merges where you would take likePhilip [00:19:10]: YeahSwyx [00:19:10]: Layers from each model.Swyx [00:19:11]: Does anyone do that anymore?Ali [00:19:13]: Well, to your point previously when you were mentioning like, the work that goes into supporting a model when it first comes out, like GLM-5.2 or MiniMax M3 or whatever the case is. Sometimes you do have to like, you do have to switch out some things. Like, for instance, the MiniMax M3 head uses full attention, and with full attention you end up with this like insane bottleneck in spec dec ‘cause you're doing auto-regressive token generation for three tokens, and you're doing this like N squared over all of the tokens that are in your sequence. Your KV cache is like very large because it's not sparse, it's not top K. So we find it better to like, okay, we're gonna replace this, we're gonna replace this layer with a layer from another model that's using like GQA, for instance. And then just with the right training, you can get it to have the same acceptance rate. So it is very possible to retrofit layers from other models and very much needed. If a layer is like inefficient, the training just becomes the challenge, like how do you ensure that you train it properly? Which again to your earlier point is like the mesh between training and inference. As in like you need very good training in order to do fast inference. That's like, I feel like more and more becoming true.Swyx [00:20:21]: Yeah. Anything else on the support side when you say like get it to fully production ready?Loop Detection, Race Conditions, and Non-DeterminismPhilip [00:20:26]: Yeah. I think that there's also a question of just, we can test a model to a pretty extensive degree, but we're trying to get it out quickly and then you see a bunch of other people test it and you get interesting results. There was an issue with, GLM briefly where we had some like mode collapses where it would just output the same token over and over again for certain prompts on certain temperatures. Like once you expose an endpoint to the real world, there's going to be, so many more varieties of things given to it that you're able to, discover and patch things. So it's not just a, day zero process, it's then like for the first week, for the first month, if a model remains popular, like how do you both fix bugs and then continue to push the envelope on performance?Ali [00:21:21]: What do you mean you don't want your model outputting S?Swyx [00:21:24]: Is there loop detection on that stuff, by the way? It still happens like quite a lot, which is surprising.Ali [00:21:30]: We have like we, in our endpoint, like if a model was to output the same token like four plus times, we just cut the generation. We say like, “Oh, sorry, this-- Like try again,” or like we will reprocess the request. ‘Cause we know then, like if it, like if, yeah, it's four times the same token, it's probably collapsed.Swyx [00:21:45]: Yeah. Is there a way to opt out in case I really want that?Ali [00:21:48]: You want that?Ali [00:21:50]: I think there's a way that we have to handle it. I'm not exactly certain, but I feel like in certain models, like when they output something like you can imagine, like a table for instance, and so they want, they wanna draw like 12 dashes and 12 dashes. Yeah, I think there's a way for that to happen. I think we only do it on certain tokens. Like we exclude certain special characters.Swyx [00:22:07]: Yeah.Ali [00:22:07]: So we only do it on like certain like S is the most common almost. GLM-5.2Swyx [00:22:11]: OhAli [00:22:11]: And I think it was DSV 4 as well. Like you'd just have like looping issues where like you literallySwyx [00:22:17]: ItAli [00:22:17]: Just have like S.Swyx [00:22:18]: Yeah. Is there a special, something special about S? No, just randomlyAli [00:22:21]: It just seems to be the one token involved.Swyx [00:22:23]: Yeah. And it'Philip [00:22:24]: Is thereSwyx [00:22:24]: And it's only temperature 0Ali [00:22:27]: NoSwyx [00:22:27]: Even at other temperaturesAli [00:22:27]: Even at like 0.9 or whatever, it will still, it will still collapse.Swyx [00:22:30]: That's weird, right?Ali [00:22:30]: It's, it is an inference problem to be honest, like a software problem. Like oftentimes, the image you run will-- like NVIDIA will release an image for instance, and if we will upstream the changes from their latest TensorRT-LLM image into our stack, we'll find that it fixes it. Or oftentimes this will only happen in an inference engine that you're using like SGLang. But if you were to switch to vLLM, that isn't the case. So it seems to be like an extremely like deterministic software issue and not really a model issue. It's not like a weights problem. Like I'- we'll say like, “Oh, it's a problem with the quant. We did PTQ wrong,” right? But that isn't, that doesn't make sense because the same weights used with a different inference engine does not repeat the problem. And sometimes it's, the kernels that are being used in the backend have like these very subtle sometimes race conditions, where if you were to use this model hosted on one cluster, you will never get this problem.Swyx [00:23:19]: Oh my God.Ali [00:23:19]: But if you host it on a different cluster, you will. And the reason is the KV cache transfer from a node to node in that one cluster is using a slower interconnect than the node to node in another cluster. So that exposes the race, whereas in another cluster it doesn't. So then you end up just like, okay, this model is not gonna be hosted on this cluster. We're gonna host it on, another cluster because that cluster exposed that problem. But then it ends up with like, okay, is it the software? Is it the model weights or is it the hardware?Swyx [00:23:42]: There is a thing about this with temperature 0 still not being deterministic, right?Ali [00:23:46]: Right.Swyx [00:23:46]: Mostly because of hardware. Even at temperature 0 same model, you won't always get the same output.Swyx [00:23:52]: Even-- But I'm surprised by the race condition one because, I thought PyTorch was a graph that like guarantees that you at least, execute things in the right order.Ali [00:24:02]: Well, yeah, true. Like I'm not, I'm not saying that there is. Like well, you have things like PTL optimizations where like you can start a kernel before the end of the previous kernel, and that's like ‘cause you want to do that because there'sSwyx [00:24:12]: It's like pipeliningAli [00:24:12]: Expense. Exactly.Swyx [00:24:13]: Yeah.Ali [00:24:13]: But it'- But you don't do it cleanly. Like you overlap a little bit of the execution. No, it is very possible that the kernel itself, like that one block that is supposed to be running in this instance of time, that kernel itself has a race condition. For instance, like a missing barrier. Like often if you're designing a kernel and you want it to make it to be very fast, if you don't test it extensively, you'll, you'll have certain threads access data points from registers before they've been written to by other threadsSwyx [00:24:36]: YeahAli [00:24:36]: For example, because like your barrier is wrong or your synchronization was wrong. But yeah, like the testing itself is very difficult in those like, andSwyx [00:24:42]: And there's no like borrow checkerAli [00:24:45]: What does that mean?Swyx [00:24:46]: Like Rust. Like the. If you're trying to have like memory safety It sounds like a comparable problem.Ali [00:24:52]: Well, yes, but you're working in CUDA, right, NVIDIA GPUs. Like- You just need a higher level language like modular Maybe that's what modular is supposed to do. I don't know.Quantization Quality and Vendor FidelityVibhu [00:25:00]: How do you see keeping quality of the model? So you talked about all these steps of, okay, you gotta do quantization, train your own speculative decoderAli [00:25:07]: RightVibhu [00:25:07]: Run on different hardware. Looking at other model providers, okay, you kicked off a inference speed race on the consumer end. What goes into keeping quality the same across them, right? Sure, you can run benchmarksAli [00:25:22]: YeahVibhu [00:25:22]: But, like, how do you determine how much quantization are there standards? What goes intoPhilip [00:25:27]: There's a few things on quality. Most inference optimizations are lossless. KV caching, for example. You are just recomputing or preventing recomputing the same values. Speculation, of course, if a draft token is wrong, it gets rejected. The main lossy optimization is quantization. And that really comes down to, number one, data format, number two, which parts of the model you choose to quantize, which layers, and number three, like doing a lot of calibration on the quantized weights, to ensure that you're preserving all the outliers. There's other tricks that you can do, though. A big one is long context, ‘cause one thing you asked at, right at the beginning is, “Oh, what's gonna happen if I send a 200,000 token request in?” So with a long input sequence, you need to, store a lot more information. You need to process a lot more tokens. And so even if a model has a context of a certain length, you might, as an inference provider, choose to build an API with a shorter context length, and of course a full length one as well. Because if someone doesn't need the full million token context, for example, you can get them better performance. I don't know if that's exactly like quality of the model. The way that I think about quality is to what degree are we faithfully serving the original model? If you think of a golden implementation of a model that performs exactly the way the model is designed to perform, I think of quality as how close are we getting to that, 100% fidelity of the model.Philip [00:27:13]: You can also, of course, think about quality from the training side and how do you push yourself past 100%. But when I think about purely inference optimizations, it's getting faster while staying as close to that 100% fidelity mark as possible. And certainly our standard internally is that, like you should not be able to tell the difference between our API and a, official API. I think Kimi in particular does a good job of vendor benchmarking hereAli [00:27:41]: YesPhilip [00:27:41]: Where they haveAli [00:27:42]: They released an actual vendor benchmark.Philip [00:27:43]: Exactly, yeah.Ali [00:27:44]: ‘Cause they accused, some people, Amazon? There was some provider that was not doing very well on Kimi's benchmark.Philip [00:27:50]: Yeah.Philip [00:27:51]: So, with Reflect we probablyVibhu [00:27:52]: This was a long time ago, right?Philip [00:27:54]: No.Ali [00:27:54]: Yeah, like threeVibhu [00:27:55]: They alsoAli [00:27:55]: Four, five months agoVibhu [00:27:57]: This also happened with, I don't remember which model, but they pulled out quite a few, and then they started a whole chart about this. It might have beenPhilip [00:28:03]: Kimi Vendor Verifier.Ali [00:28:04]: Yeah.Philip [00:28:05]: Yeah.Ali [00:28:05]: Yeah, ‘cause you, ‘cause you'd be pissed, right? Like if you'Philip [00:28:07]: Yeah.Ali [00:28:07]: If like if I'm a consumer and I'm using like Amazon's endpoint for instance, and I've used Kimi and I'm like, “Oh my God, like this is bad,” I'm not gonna say, “Oh, Amazon quantized the model in a bad way.” I'm gonna say, “Oh, Kimi sucks.” Right?Philip [00:28:17]: Yeah.Ali [00:28:17]: So it seems like that makes sense.Philip [00:28:19]: Yeah, they care. They care.Vibhu [00:28:21]: Justifiably.Ali [00:28:21]: Yeah, justifiably.Vibhu [00:28:22]: This is probably a stupid question, but just checking, has anything improved from main quantization?Philip [00:28:28]: Yeah.Vibhu [00:28:28]: Like, is quantization always strictly worse?Ali [00:28:30]: Well technicallyVibhu [00:28:32]: NoAli [00:28:32]: It's a lossy. QuantizationPhilip [00:28:33]: YeahAli [00:28:33]: Is a lossy, it's a lossy implementation.Philip [00:28:36]: Speed improvesVibhu [00:28:36]: Speed improves.Ali [00:28:37]: It the number, likeVibhu [00:28:38]: No, I' always look for inverse scaling laws.Philip [00:28:40]: Yeah.Ali [00:28:40]: Yeah.Vibhu [00:28:40]: This is something I learned from Noam Brown, where like things that normally act in one direction sometimes do.Philip [00:28:45]: Well, technically when you run a benchmark, because these models are deterministic, sometimes your,Ali [00:28:52]: YeahPhilip [00:28:52]: NVFP4 quant is like, two basis points higher than yourAli [00:28:56]: No, it's noise. It's noise.Philip [00:28:57]: Yeah, exactly. I'm like, yeah, it's, it's within. That's why I always say within margin of error.Philip [00:29:01]: And I stopped saying that because everyone assumes that what is, well, within some margin of error, we're barely inside of that to the worst, so we're saying. But yeah, sometimes it's just like, gives you a higher output score. But like Ali said, that's noise. To my knowledge, you're not necessarily making the results better. You're just trying to, again, like keep your fidelity as close to 100% to the original model.Layer Selection, KL Divergence, and Better QuantizationAli [00:29:27]: There is, to your point, research that we did on MP. I don't know if you are able to pullPhilip [00:29:31]: YeahAli [00:29:32]: A tweet we did. One of our research interns, Joshua, I think it's a tweet on how we have 20% better quantized GLM-5.2 than NVIDIA. Essentially what we found throughout like this month research is, okay, quantization is a lossy. It's. You're compressing the data from, occupying 16 bits to occupying, four bits, for instance. And so you're losing some information, and you're trying to minimize that. And so when I say that I'm gonna quantize the model, my job becomes how do I find the layers that I can quantize, and how to find the layers to not. For instance, with image models, I don't quantize modulation layers, and I don't quantize out projections because those two are. Like out projection is what you see as the user. Modulation is what the model sees or understands. Right, exactly. And so to his paper, do you have the. It doesn't have the. Yeah. It's a long paper. I don't know if I can findVibhu [00:30:25]: If there's a part to search or it's probably in the thread.Ali [00:30:28]: It's probably in the thread.Vibhu [00:30:29]: Yeah.Ali [00:30:29]: But the long and the short is it is very possible that quantizing more of the model makes the results. Like if I have a model that I quantize layers one, five, and 10, and another model where I only quantize layers one and It is possible that the model in which I quantized more information is going to perform better because the quantization errors have canceled out. And so what Joshua showed in his mathematical proof where he had like a verifier in, is that you can predict which layers are going to have quantization errors that will cancel out with each other, and you choose to quantize those layers. And so the result of doing this mathematical quantization is you end up with a model that's 20% more quantized than another provider, so you get 20% more throughput of it because there's more layers than running an NVFP4, and your quality is better than that other quant because the layers that you chose to quantize have their errors cancel out, like one layer skewed to the right one layer skewed to the left, one layer skewed to the right. Your final logits distribution is more similar to the original distribution of the model, so you have better fidelity. And so the way we proved this was with KL divergence. So instead of just scoring on the benchmarks, we scored the KL divergence between the logit distribution of the quantized model and the logit distribution of the original full precision model, and we showed that with this technique we get. If your probability distribution on the logits which token it wants to select is more of the same as the original model, you're probably gonna end up staying true to the original model. So yeah, so it seems like previously before this, it seemed like the industry was, well, the more you quantize, the worse it's gonna be, ‘cause the more loss you introduce. That's not exactly, not necessarily true. So yeah, doesn't improve it, but can cancel out.Philip [00:31:57]: I think it might be this, but reminds me a good bit about pruning where you can prune off certain layers.Philip [00:32:03]: But very interesting. Didn't know this was a whole paper you guys put out.Ali [00:32:06]: It's. Fun fact, it was originally 72 pages, this paper, and then we decidedPhilip [00:32:11]: WowAli [00:32:11]: We can't tell. We couldn't release it. So it's now 45.Swyx [00:32:15]: Still 39 pages, so very substantive. We talked about evals and all these things and, like what's possible in terms of speedup? Like it's like probably like the numberInference Speedups and BenchmarkingSwyx [00:32:25]: Thing that people do wanna care about, and it's something that you wrote about in your post. Like official API is 70 tokens per second, and you push it up to 90. Is that like a normal thing?Philip [00:32:36]: So what's cool about working in inference, the reason that I think inference is going to be a useful place to do engineering for a long time, is that if you look at highly optimized domains like, say, finance, if you're in finance, you measure how much better you got in basis points. It's like, “Oh, I got five basis points better, like twentieth of 1% better,” that's huge news because everything is so optimized. When we publish optimizations, it's 20%, it's 100% it's 200%. So there's still probably like a lot further to go, honestly. Like you'll, you'll know that inference is pretty much solved when researchers start publishing about how they got 1% faster at something.Swyx [00:33:19]: Which by the way, because I am from the finance background, in the ‘70s, that was the margin at the time. When you did quantitative finance research, you would findAli [00:33:27]: And like 20%, tens of percent.Swyx [00:33:29]: That's. Yes.Philip [00:33:29]: Yeah.Swyx [00:33:30]: And now it'Philip [00:33:31]: Tiny fractionsSwyx [00:33:32]: For those people interested, look up Andrew Lo's paper. He had a really interesting illustration of quant, stat arb, distribution, narrowing down from like those kinds of 20% differences in the ‘70s, down to nothing today, which is very cool.Philip [00:33:48]: Exactly, and we're at the beginning of the same type of thing. Now benchmarking is hard. I think anyone will tell you that, and benchmarking provider speeds is hard because there's so many variables that go into it. What hardware are you using? How much load do you have on the system? What's the exact nature of the prompts and input and output sequence lengths? All that stuff. But overall, when you start stacking these improvements, you're looking at multiples. You can look at it. The most common form, of course, is TPS, tokens per second, which is bad naming by us in the industry, ‘cause there's two tokens per second. There's tokens per second, the throughput number, and the latency number.Ali [00:34:31]: TTMT, yeah.Philip [00:34:32]: Like total tokens per second out of the, out of the GPU as a throughput number. Most people only care about tokens per second as the latency number, which we should call ITL, intertoken latency, but we don't.Philip [00:34:44]: Anyway, so you can imagine a standard API without many optimizations for a 1 trillion parameter model operating somewhere in the 30 to 50 tokens per second range for reasonable traffic profile. And we generally see the goal of, pushing to 10X that. But, not necessarily day zero, but by stacking enough optimizations, if you have, say like four optimizations, each of which doubles performance. Or sorry, three optimizations, each of which doubles performance, then you stack that up, that's an 8X gain. That's the order of magnitude that we're working with in this space. We're trying to make things substantially faster, not just go from like 70 to 90.Swyx [00:35:38]: Are you saying you've. You have done that?Philip [00:35:40]: So let's say you have as a reasonable baseline, 30 or 40 tokens per second. You can achieve 10X that. So like on GLM-5.2, if you run it unquantized, perhaps on H100s even, and you're just using an off-the-shelf inference engine with no particular optimizations, no speculator, nothing extra around like KV routing, no disaggregation, you're, you're probably, yeah, looking at that like 30 to 40. You think that's like a reasonable baseline?Swyx [00:36:12]: Right. Right.Philip [00:36:12]: To get to something like 10X, there's a lot of trade-offs that you're making. If we're running at more like a 300, 400 tokens per second range, you are using the best hardware possible. You have a optimized speculator. You have done all of your quantization work. You are Seeing a pretty high cache hit rate. You are running with a reasonably small batch size and a parallelism configuration that is tuned for latency versus throughput, but it is possible. So the spreads that you see if you, like, go on artificial analysis or you go on OpenRouter and you look at, the worst provider to the best provider, oftentimes can hit that range. 10X is of course very aggressive. It's oftentimes maybe more of a four to six times improvement. But that's the performance that makes us really excited, is when we can get these huge gains, not just go from 70 to 90 tokens.Stacking Optimizations: NVFP4, Speculation, and DisaggregationAli [00:37:19]: It's also, like, hardware dependent. Like, ifPhilip [00:37:20]: YeahAli [00:37:20]: If you have a thing where you're serving it on just, like, a node of H100s and then you throw, like, you shard the model across, like, four nodes of B200s. Like, you can definitely increase the speed with just throwing more hardware at it. Like, normalizing for the same exact hardware and the same number of GPUs.Philip [00:37:35]: Yeah. Then you're looking at, like, a two to 4X improvementAli [00:37:38]: Right. RightPhilip [00:37:38]: Depending on the inference optimizations. So yeah, it's. Some of it's, what's the call, and some of it's who's the driver.Vibhu [00:37:46]: If you break down the two to 4X, say the example is run GLM-5.2Ali [00:37:51]: YeahVibhu [00:37:51]: On B200sAli [00:37:53]: YeahVibhu [00:37:53]: Single node, right? What's, like, the cost trade-off for effort to get, like, the last bit of juice out versus what should people just think of, right?Ali [00:38:01]: Spectre quantization. Yeah.Vibhu [00:38:03]: Spectre quantization.Ali [00:38:04]: That's, that's, that's like 95%. LikeVibhu [00:38:06]: And how far does that get you? And how easy is that for the average person to do? So say right I wanna throw the weights of GLM-5.2 on a node of B200s, how easy is it to find speculative decoder- decoder model or already quantized model? How much work goes into it?Philip [00:38:23]: If you're doing it up front, it's quite a lot of work. If you're doing it today, there's going to be people who have published things that you can just, you can just grab some NVFP4 weights. You can grab a speculator. Yeah, if we're thinking about, like, what are the 2Xs we're stacking, going from, BF16 to NVFP4 is, it's not quite a 2X, right? It's like. I think it's about, like, 30 to 40%, from 16 to 8, and then another 30 to 40% multiplied from, 8 to 4. So that doesn't quite get you a 2X, but, like, roughly a 2X. Speculator, roughly a 2X. Disagg on top of that if you're able to get enough hardware and put enough traffic through it, another roughly a 2X. And then you add in some, double-digit percent increase from having just a better runtime with, the latest kernels and stuff behind it. And that's how it stacks up.Ali [00:39:21]: YeahPhilip [00:39:21]: So building each of those, like, building the, quantized weights is, for someone who really knows what they're doing, hours to days of work. Building the speculator, again, like, hours to days of work. And the, disagg setup, hours to days. Well okay, but like once you haveAli [00:39:39]: Once set up. Once set up. YeahPhilip [00:39:40]: Yeah, getting disagg working for the first time, I'm saying, of course, is very difficult.Philip [00:39:44]: The marginal implementationAli [00:39:48]: Like, if you're just grabbing, like if you are a person, like just a normal consumer who has access to, like, a node of B200s and you're wondering, “How can I just host it myself?” You don't need to quantize the model yourself. There's always gonna be, like, an open source quantized checkpoint. NVIDIA's gonna push one out if no one else does. You. Usually, the providers will have their own spec dec that they've trained as well. You don't need to train your own spec dec. You can just use that as well.Philip [00:40:09]: Yeah. Like, GLM-5.2 has its own MTP.Ali [00:40:13]: Right. Right.Vibhu [00:40:14]: What's multi token prediction?Philip [00:40:15]: Yes.Ali [00:40:16]: I'm justVibhu [00:40:16]: Can you explain that?Ali [00:40:16]: I'm just an expert.Ali [00:40:18]: I can do it for you in case I get it wrong?Vibhu [00:40:20]: No.Vibhu [00:40:21]: Yeah, you should correct if we're wrong, but their multi-token prediction can be used for self-speculative decoding.Ali [00:40:27]: I'm not sure. I'm not gonna correct that.Vibhu [00:40:28]: Okay. I'm semi-confident in thatAli [00:40:30]: Okay. YeahVibhu [00:40:30]: But someone can check. But it's useful to paint the story of, okay, not just the average person, but say a company wants to switch from serverless inference I wanna throw this up on. I wanna rent some GPUs, throw it up. These are the steps you take to do significantly faster than just put it behind vLLM.Ali [00:40:48]: Right.Vibhu [00:40:49]: I was waiting for a mention of Dynamo.Vibhu [00:40:51]: I feel like, that's supposed to be the baseline that you measure against.Dynamo, KV Routing, and Disaggregation ToolkitsPhilip [00:40:55]: I would think of Dynamo as less of a box system and more of a toolkit for building with. So when we talk about doing aware routing, when we talk about doing KV offloading, when we talk about doing, PD disaggregation, Dynamo fundamentally is. By the way, Dynamo is an open source library from NVIDIA.Ali [00:41:17]: We've done a pod with KylePhilip [00:41:18]: OkayAli [00:41:19]: Kyle Cranin.Philip [00:41:19]: Cool. So then your listeners know then that it supports all the different inference frameworks. And it is multi hardware, which is interesting.Ali [00:41:28]: But it's just a router, it's not like an optimizer layer.Philip [00:41:30]: Yeah. All it does, like, what Dynamo is good at, it is a library for moving information around your cluster, around your hardware. So if you have, KV cache on one place and you need it to be somewhere else, Dynamo coordinates NIXL for you to move that around.Philip [00:41:49]: That doesn't mean that, like, out of the box, you just say, “Pip install Dynamo,” and then you get, like, a massive performance speed up. It's more of a developer toolkit.Ali [00:42:01]: Yeah. I would have said it would. It comes with a set of defaults that you can then swap out.Philip [00:42:06]: It does. If the industry at large, I think, was, like, rolling out all of these deployments, standard, then I think it would be, like, a credible baseline. But, we've got to, we've got to benchmark against, like, what we're seeing in the wild.Speculative Decoding Methods: Medusa, EAGLE, n-Gram, and Spec-SpecVibhu [00:42:23]: I did wanna talk a little bit more about PD disagg, because that is probably, like, number three after quantized and speculative decoding. In your book though, I was just gonna pull out the book.Philip [00:42:31]: Yeah.Vibhu [00:42:32]: Like section 522 on Medusa, 523 on EAGLEPhilip [00:42:35]: YeahVibhu [00:42:36]: 524 on gram.Philip [00:42:37]: It's 55, would be disaggregationAli [00:42:42]: Yeah. Well, no, I just wanted to dwell a little bitPhilip [00:42:44]: YeahAli [00:42:44]: The other. Like, so what do you choose to include? What do you choose to not to include? Because there was all these other techniques.Philip [00:42:51]: Yeah.Ali [00:42:51]: Are these still relevant? Because I think they came out, like, a year and a half ago maybe.Vibhu [00:42:55]: Medusa is quite old.Philip [00:42:56]: Yeah, Medusa's old.Ali [00:42:58]: It was old.Vibhu [00:42:58]: But is it in the book as a good, here'sPhilip [00:43:01]: BaselineVibhu [00:43:01]: Baseline vanilla understand it?Philip [00:43:02]: Like you should know this.Vibhu [00:43:03]: Like I read the paper, I'm like, “ it makes so much sense.”Philip [00:43:05]: Yeah.Philip [00:43:05]: So with the book, I had a couple goals. One was to give people just a working vocabulary for the space as a whole, and the other was to give them some intuition about how each of these techniques works. As I mentioned in my AI Engineer talk, which is the first public addendum to this, the speculation space has moved much faster than everything else. So yeah, even at the time that I wrote the book Medusa, I very much included as a way for people to understand how the space evolved rather than what the most modern technique is. And now of course, there's DFlash, dSpark. There's, there's newer techniques even than EAGLE, although EAGLE is still very commonly used.Ali [00:43:51]: SpecSpecta.Philip [00:43:52]: Yes. Speculative decoding.Vibhu [00:43:54]: What canAli [00:43:56]: Oh, it's a paper by Tri Dao and it's like, it's doing speculative decodingVibhu [00:44:00]: HuhAli [00:44:01]: For the speculative decoder.Philip [00:44:02]: Oh, in spec- oh my God.Ali [00:44:02]: It's literally just an another. It's like, yeah, that's the most simple way to explain it, and it seems like he got trivial speed ups there. But it seems that the complexity with training, it's almost like in our mind at least, it's almost as complex as training GANs. Like it's like a very delicate balance and oftentimes you, it's just but yeah, it's literally speculative decoding on speculative decoding.Vibhu [00:44:21]: Speculative.Ali [00:44:22]: Yeah. We saw this paper.Vibhu [00:44:24]: It's interesting, right?Ali [00:44:24]: Yeah.Vibhu [00:44:24]: I wouldn't even expect it to be very particular to train, I wouldAli [00:44:29]: Right.Vibhu [00:44:29]: The naive part of me is like, okay, train speculative decoder.Ali [00:44:32]: But like, and it makes sense, like the whole idea of speculative decoding is you. It's like, it's like almost like the iPhone auto predict version but for a normal model, right? Like you're just, you're just, generating three tokens and you're like, okay, I'll do prefill on them. And so you save those three turns for your original model. Now your speculative decoder is doing three turns of auto regression, so why not just have an even smaller model?Ali [00:44:53]: The other question there is what are the size of speculators? So say forPhilip [00:44:58]: Right. It's like a billion parameters.Ali [00:45:01]: Like for MiniMax, it's. Yeah. It's like one layer. It's like one 60th of the original model usually.Philip [00:45:06]: Yeah. I think we should do a paper when we get back to the office.Philip [00:45:10]: SpeculativeAli [00:45:11]: SpeculativePhilip [00:45:11]: Decoding.Ali [00:45:13]: No, it's, it does seem like how, when do you stop? But then it also seems like if you're able to train spec-spec decode for instance, right? Like if you're able to have a small model that is accurately predicts what the intermediate speculator is gonna predict, that is able to predict what the original target model's gonna predict, then why not just use that smallest model directly, right?Vibhu [00:45:34]: Yeah. This isAli [00:45:35]: Like it seems likeVibhu [00:45:35]: Adjacent to the routing problem.Ali [00:45:36]: Right.Vibhu [00:45:36]: Yeah.Ali [00:45:36]: Right.Philip [00:45:37]: The thing with speculators is one of the practical constraints on using them is that you do have to run a small model on the same hardware that you're running the big model on. There is a orchestration and resource competition problem inherent in that, and that is one of the constraints on speculation in general, is that draft tokens cost resources to create and cost software complexity to manage. And so if you have like infinitely recursive speculators, you add in quite a bit of that complexity on the actual implementation within the inference engine as well, not just in the training process.Vibhu [00:46:17]: I was gonna say, I would wonder if you could do similar, like distillation and pruning of, it's the same thing, it's just a model. Can we not just distill a lot of the weights, quantize the speculator, out of my domain? The question that also comes up is, this is all for big server workloads, right? How much of this applies to, say I have this MacBook, I wanna run Gemma really efficiently. Similar problems, not the same?Local AI vs. Data Center InferencePhilip [00:46:45]: Pretty different. I talked to Selo, about this on his podcast a couple weeks ago. The difference between inference engineering for the data center and for production workloads versus inference engineering for local AI, is that we start with fundamentally like different constraints and different goals. With local AI, it's how do I fit this model onto my hardware and then make it less dumb? And with data center influence, it's how do I load this model and then make it less slow? And we care about less dumb, and they care about less slow. But the local AI inference engineering ecosystem, I think has a lot for us to learn from in the data center space. They are experts in various forms of quantization, including dynamic quantization that we just don't touch, in the pruning, in the distillation, in the, layer removal. There'Ali [00:47:42]: Layer removal matters less.Philip [00:47:43]: Yeah. There'Ali [00:47:44]: No one loves pruning really.Philip [00:47:45]: Yeah. Well, but the, but they doVibhu [00:47:46]: Which is surprising, right? But that's, that's a whole different thingPhilip [00:47:48]: Just to fit something on the laptop.Ali [00:47:50]: Right.Philip [00:47:50]: So yeah, it's a, it's an interesting, it's an interesting space. Not necessarily that like their techniques make sense for us to do in the data center, because we have different resources and different goals, but more that the process as well as the openness of that field is something to, admire.Ali [00:48:12]: Yeah. Like to your point, like, certain optimizations that would. Like for instance, Turbo Quantum Sharper, like it made such huge hype on that and we did like a whole deep dive on Twitter and like said, what is it? How does it work? Why is it good or not? And it took off and it was implemented on local devices because your memory bandwidth is so slow on like a MacBook, for instance. But try putting the same thing on like an NVIDIA GPU on a B200 Turbo quant would not be. Like, it would not be used. Like, NVIDIA - Like, NVIDIA made it clear that this is not a good optimization, and we've seen it firsthand where the overhead of doing dequantization, quantization of, in the kernel itself with turbo quant kernel, each end is much slower than the time that you save from doing the bandwidth. ‘Cause on the B200s, you have like 3.5 terabytes per second. You don't need decrease the storage that much. You don't need to do, FP4 KV cache. You don't need to use a requant. There's, there's, there's better optimizations to be made. But on Edge devices, it's extremely important, it's extremely useful. So, seems to be, like, different optimizations there, but then they're all uniquely combined with like all you wanna quantize the model, you wanna do speculative decoding, like certain common prefixes with bothPhilip [00:49:18]: Principles.Ali [00:49:19]: Yeah, exactly. Exactly. Exactly.Philip [00:49:20]: They also do a lot of work on, model parallelism, especially over, heterogeneous topology, where you have, some sparks and they are wired together with, Ethernet, DGX sparks.Ali [00:49:35]: Yeah, this is the Exo Labs guys.Philip [00:49:36]: Yeah. You have, a nu

AI Hustle: News on Open AI, ChatGPT, Midjourney, NVIDIA, Anthropic, Open Source LLMs

In this episode, we investigate how Kimi K3 is creating new revenue opportunities in the AI space. Find out strategies for leveraging this technology for profit.

Silicon Carne, un peu de picante dans la Tech
Le « moment Spoutnik » de l'IA : sortie de Kimi K3. La Silicon Valley panique !

Silicon Carne, un peu de picante dans la Tech

Play Episode Listen Later Jul 30, 2026 39:55


Le « moment Spoutnik » de l'IA : sortie de Kimi K3 Un modèle chinois rivalise avec Claude et GPT-5 pour cinq fois moins d'argent. Ce n'est pas une copie — c'est une menace directe sur les valorisations à mille milliards de dollars construites en quelques années à peine aux États-Unis.La Maison Blanche crie au vol, mais les outils qui font tourner la Silicon Valley utilisent déjà des modèles chinois. La distillation dont on accuse Moonshot, c'est le même processus qui a permis à Cursor de devenir ce qu'il est aujourd'hui.Ce qui bascule, ce n'est pas seulement le leadership technologique. C'est la question de qui a le droit de construire sur quoi et qui protège son oligopole en se cachant derrière la réglementation.===================⏱️ DANS CET ÉPISODE :===================0:00 — Intro0:34 — Kimi K3, le modèle chinois qui défie l'Amérique3:03 — Une architecture innovante et des coûts inédits11:47 — Distillation: voler ou s'inspirer légitimement17:52 — 5 milliards contre 100: le choc de valorisation19:30 — La contrainte comme arme secrète des Chinois22:16 — Les LLM plafonnent26:03 — Jensen Huang contre Washington: la Silicon Valley en révolte34:02 — Kimi K3 s'adresse à qui concrètement=============

ChinaPower
China's Moonshot AI's Kimi K3 and Open AI Strategy: A Conversation with Ngor Luong

ChinaPower

Play Episode Listen Later Jul 30, 2026 39:54


In this episode of the ChinaPower Podcast, Ngor Luong joins us to discuss the rapid rise of Chinese startup Moonshot AI's Kimi K3 model, China's open AI strategy, and the implications for U.S.-China competition. Drawing on her recent paper, Two Loops: How China's Open AI Strategy Reinforces Its Industrial Dominance, she examines China's rapidly expanding open AI ecosystem, unpacks the "two loops" framework behind China's AI development strategy, and assesses the implications of the growing global diffusion of Chinese AI models.  Ngor Luong is Senior Policy Analyst on the Technology and Innovation team at the US-China Economic and Security Review Commission. She joined this podcast in her personal capacity and not on behalf of the Commission. 

聽天下:天下雜誌Podcast
【阿榕伯胡說科技Ep.80】全球半導體重挫;中國的Kimi撕裂 AI 市場!阿榕伯解說7月科技大事

聽天下:天下雜誌Podcast

Play Episode Listen Later Jul 30, 2026 46:21


這集節目是7/24胡說科技YouTube直播內容,想看影像版歡迎點擊連結:https://tw.psee.ly/9ecgnc 上週引起最熱烈討論的,就是中國AI新創,月之暗面的Kimi K3,正以低廉的價格、追上美國第一梯隊的性能,對AI產業造成「價格破壞」。同時,7月台積Q2法說也透露將要擴大對美投資。 科技類股持續面臨拋售,七月還有哪些大事?阿榕伯講給你聽。 關於《胡說科技》: 天下獨家推出《胡說科技》,這是第一份為關注產業趨勢讀者打造的科技專欄頻道。由天下總主筆陳良榕,半導體資深記者,耕耘科技報導多年,親自操刀,每週出刊。《胡說科技》結合對產業的敏感度,提供獨家、深刻的觀點與分析,不談花俏技術,不用行話堆疊,每週一篇,希望讀者在短時間掌握影響決策的科技動向。 *到官網看最新文章,立即訂閱: https://bit.ly/3TuL8eb *意見信箱:bill@cw.com.tw -- Hosting provided by SoundOn

Unchained
The Chopping Block: Wind Downs, YC's Nemil Dalal, & Will Every Failed Crypto Idea Eventually Work?

Unchained

Play Episode Listen Later Jul 29, 2026 67:08


YC's Nemil Dalal joins to explain why he's never been more bullish as BitMEX winds down after 11 years, whether every failed crypto idea (TCRs, DAOs, creator coins) eventually works, why crypto is really about money, Base's consumer mea culpa, on-chain reputation and credit, and who pays in the x402 AI-agent era. Welcome to The Chopping Block – where crypto insiders Haseeb Qureshi, Tom Schmidt, Tarun Chitra, and Robert Leshner chop it up about the latest in crypto. This week they're joined by Nemil Dalal, Visiting Partner at Y Combinator and ex-Coinbase, where he led USDC and the Coinbase Developer Platform. He's here to explain why, with exchanges winding down left and right, he's somehow never been more bullish. The crew digs into the great contrast of the moment: BitMEX shutting down after 11 years (plus BitMart, Movement Labs, Balancer Labs) while the plumbing quietly prints, and whether Imran's viral 'everything that failed will eventually work' thesis is genius or toxic positivity. From there it's the question of whether crypto is really only about money (Jesse's Base mea culpa included), a war-memories tour through TCRs, on-chain reputation and why pure on-chain credit keeps faceplanting, and finally who actually pays in the x402 AI-agent era, and whether decentralization even survives contact with Google-shaped gravity. Listen to the episode on Apple Podcasts, Spotify, Pods, Fountain, Podcast Addict, Pocket Casts, Amazon Music, or on your favorite podcast platform. Show highlights

FYI - For Your Innovation
SpaceX's Starship Flight 13 + Kimi K3 Freakout | The Brainstorm 142

FYI - For Your Innovation

Play Episode Listen Later Jul 29, 2026 29:16


In this episode of The Brainstorm, Tasha Keeney, Nicholas Grous, and Brett Winton uncover how SpaceX's latest Starship test flight is revolutionizing space travel, slashing launch costs, and how this potentially unlocks new commercial opportunities. Plus, a deep dive into AI's rapid evolution, cost declines, and the real race behind the frontier of intelligent models.Key Points From This Episode:Why SpaceX's recent test flight is a critical milestone toward fully reusable rockets and what that means for satellite deployment costsHow AI models are getting exponentially cheaper, with potential savings of 97x in just a year and a half, and what this implies for knowledge workThe true battleground between frontier and open-source models and why the race to dominate AI is about economic power, not just technological capabilityThe subtle ways AI-driven automation and intelligent agents are transforming marketing, operations, and even consumer choiceIf you know ARK, you know we focus on long-term innovation. But that doesn't mean we ignore breaking news. Every day, we debate the latest developments in tech and markets. Now, we're bringing those conversations to you in “The Brainstorm,” a co-production from ARK, WOLF, and Public. Tune in weekly for our quick takes on what's shaping innovation right now.Learn more about WOLF: https://wolf.financialLearn more about Public: https://public.com/Disclosure: http://arkinv.st/39rzF94

Invest Like the Best with Patrick O'Shaughnessy
Sam Altman - How to Make an Abundant Future - [Invest Like the Best, EP.484]

Invest Like the Best with Patrick O'Shaughnessy

Play Episode Listen Later Jul 28, 2026 53:33


My guest today is Sam Altman, CEO of OpenAI. It's a conversation spanning the history, present, and future of OpenAI, from the origin of ChatGPT through Codex, hardware, and their new Jalapeno chip. We discuss the early decision to buy compute at a scale nobody thought was rational, and the plan to build a gigawatt of new capacity every week.  We talk about Kimi and distillation, the Hugging Face incident and what it means for the pace of AI development, and what it's like to raise kids who will grow up never knowing a world without abundant intelligence.  Please enjoy my conversation with Sam Altman. For the full show notes, transcript, and links to mentioned content, check out the episode page ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠here⁠⁠⁠⁠⁠.  ----- Become a Colossus member to get our quarterly print magazine and private audio experience, including exclusive profiles and early access to select episodes. Subscribe at ⁠colossus.com/subscribe⁠. ----- ⁠Ramp's⁠ mission is to help companies manage their spend in a way that reduces expenses and frees up time for teams to work on more valuable projects. Go to⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ ⁠ramp.com/invest⁠⁠ to sign up for free and get a $250 welcome bonus. ----- Trusted by thousands of businesses, ⁠Vanta⁠ continuously monitors your security posture and streamlines audits so you can win enterprise deals and build customer trust without the traditional overhead. Invest Like the Best listeners get a special offer of $1,000 off Vanta when you go to ⁠vanta.com/invest⁠.  ----- WorkOS⁠ is the infrastructure B2B and AI-native companies use to sell to enterprise. It covers everything enterprise security requires: SSO, SCIM, RBAC, Audit Logs, AI governance, and more. Trusted by 2,000+ fast-growing companies, including OpenAI, Anthropic, Cursor, and Vercel. ----- Rogo is the AI platform for finance. They're building agents for Wall Street that are trained to understand how bankers and investors actually do work: from diligence and modeling, to turning analysis into deliverables. To learn more, visit rogo.ai/invest. ----- ⁠Ridgeline⁠ has built a complete, real-time, modern operating system for investment managers. It handles trading, portfolio management, compliance, customer reporting, and much more through an all-in-one real-time cloud platform. Visit⁠ ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ridgeline.ai⁠. ----- Editing and post-production work for this episode was provided by The Podcast Consultant. Timestamps: (00:00:00) Welcome to Invest Like The Best (00:02:02) Intro: Sam Altman, CEO of OpenAI (00:02:35) Refocusing (00:05:43) OpenAI's Compute Bets (00:09:07) Data Centers (00:11:14) Jalapeno Chip (00:11:52) Kimi, Distillation & Open Source (00:14:39) The Hugging Face Incident (00:17:46) OpenAI's Mission & Vision (00:22:14) All the Returns Are at the Frontier (00:22:27) Bottlenecks: Compute, Research, Data (00:23:49) Sam's View on AI & Jobs (00:26:56) Unpopular Bets That Turned Out Right (00:27:45) Model Cycles (00:29:45) How Sam Uses AI (00:32:44) Having Kids (00:34:56) Why Sam Has No Equity in OpenAI (00:35:33) Robotics (00:36:48) The Origin Story of ChatGPT (00:39:22) How to Get AI into More Hands (00:42:20) How Sam Recruited Great AI Researchers (00:43:57) What Sam Learned From Being an Investor (00:45:22) What the Next 6–36 Months Look Like (00:46:31) Codex (00:49:36) Could We Be Oversupplied in Compute in Two Years? (00:50:09) Sam's View on Scaling Laws (00:50:20) Alec Radford (00:51:12) Formative Moments (00:53:50) Kindest Thing

Syntax - Tasty Web Development Treats
1024: Open Models Replace Big AI

Syntax - Tasty Web Development Treats

Play Episode Listen Later Jul 27, 2026 67:14


A huge week for open models: Inkling (the first big US open-weight model since Gemma 4), Qwen 3.8, and Kimi 3 all dropped. Plus Vue 3.6 RC + Vapor Mode, the slow death of Stack Overflow, a decoy font that blinds AI, and the usual grab bag of fun links. Show Notes 00:00 Intro 00:44 Welcome to Syntax 03:01 Vue.js 3.6 RC Released Vue.js 3.6-RC.1 09:50 Brought to you by Sentry 10:35 Rift - new worktree alternative Rift Explaining Rift tweet 17:46 Every new browser feature MDN Plus Browser features updates CSS 'text-fit' property 23:09 Kimi K3 Released Kimi K3 release tweet Kimi K3 blog post Openrouter effective pricing 29:33 What are Open Weight Models 32:46 Kimi K3 Architecture Kimi K3's built up Cutting off access to Kimi 33:59 Qwen 3.8 Released Qwen 3.8 launch announcement 36:42 Inkling from Thinking Machines Thinking Machines release Inkling Inkling release post from Mira Murati Thinking Machines news blog post 40:45 Mole Cleanup App Mole 43:21 Moshi Terminal App Moshi app 46:05 RIP Stack Overflow Query for new questions on Stack Overflow Stack Overflow for Agents StackExchange 54:10 Scott's Foot Pedal 54:59 Scott's Feet 56:13 Decoy Font 59:16 Jurassic Park Computers 01:03:59 Ride the Tokyo Trains LIVE Hit us up on Socials! Syntax: X Instagram Tiktok LinkedIn Threads Wes: X Instagram Tiktok LinkedIn Threads Scott: X Instagram Tiktok LinkedIn Threads Randy: X Instagram YouTube Threads

F1 Nation
‘Sublime' Lando, the ‘Max factor' + Fernando's ‘fire' – Hungarian GP Review

F1 Nation

Play Episode Listen Later Jul 27, 2026 51:54


Tom Clarkson is joined in the paddock by F1TV expert and IndyCar race winner, James Hinchcliffe, to reflect on the Hungarian Grand Prix.Lando Norris converted pole position to take his first Grand Prix victory of the season and his first as a World Champion. Why was Lando back on top in Budapest? Is he driving better this year than last year when he won the title? And can he now regularly fight with Mercedes and Ferrari to defend that title in 2026?Max Verstappen was pleasantly surprised to finish P2 and secure his third podium in the last four races. So how did Max end up on the rostrum? And what did Tom and Hinch think of his incredible overtake of Lewis Hamilton on lap 16? It was another mixed weekend for Mercedes. Kimi Antonelli climbed from seventh on the grid to third, extending his championship lead further, while George Russell suffered an anti-stall issue at the start before fighting back for a points finish. How different will Kimi and George's mindsets be going into the summer break?Plus, the guys discuss a missed opportunity for Ferrari and a much better weekend for Aston Martin, whose upgrades helped Fernando Alonso reach Q2 for the first time this season. THIS EPISODE IS SPONSORED BY...HexClad: Find your forever cookware @hexclad and get 10% off at hexclad.co.uk/NATION! Offer excludes bundles and other items already on sale. #hexcladpartnerIM8: Give your body what it deserves with IM8! Go to im8health.com/nationand use code NATION for a Free Welcome Kit, five free travel sachets plus tenpercent off your order. #im8health

The Fast And The Curious
Are McLaren back? Can anyone catch Kimi? Why is everyone so grumpy? Hungarian Grand Prix Reaction

The Fast And The Curious

Play Episode Listen Later Jul 27, 2026 54:24


Greg is back! He joins Betty and Christian to look back on the Hungarian Grand Prix. Are McLaren back for good. Can Lewis still win it? Will anyone catch Kimi? And is everyone getting a little tired and grumpy? Plus Christian confuses everyone with something about a flannel. We're back after the summer break with more big name guests. Oh, and some news about the pod… Hosted on Acast. See acast.com/privacy for more information.

Podcasts – Weird Things
AI Models, Apple Lawsuits, and Productivity Without Tedium

Podcasts – Weird Things

Play Episode Listen Later Jul 25, 2026


Andrew Mayne, Justin Robert Young, and Brian Brushwood cover the latest wave of AI releases by comparing OpenAI's 5.6 model, Anthropic's Fable, and the new Chinese Kimi model, arguing that the most interesting shift is not just raw capability but how differently these systems now behave in planning, initiative, and collaboration. They also spend a lot of time on what AI is already good for in real life: voice conversations with persistent context, automated email and file management, cheap local transcription, website building, cloud-task execution, and custom tools built with Codex that eliminate repetitive work. The Apple lawsuit against OpenAI becomes a bigger discussion about talent flight, hardware ambitions, and Apple's struggle to keep pace in AI, while the Kimi conversation turns into a broader look at distillation, Chinese innovation, and the murky realities of model copying. Throughout, the hosts keep returning to a practical distinction between creative authorship and productivity support, making the case that many people who resist AI-generated art may still benefit from using AI to remove tedious logistical overhead. The biggest takeaway is simple: start with the annoying parts of your life or work, let AI handle the bureaucracy, and keep the human effort for the parts that are hard because they matter, not hard because they are boring. Picks: Brian Brushwood: Use the ChatGPT app, let it edit things on your desktop, talk to it instead of typing, and tell it your

Sway
OpenAI Models Go Rogue + Kimi K3 Freakout + A.I. Superforecasting

Sway

Play Episode Listen Later Jul 24, 2026 68:14


This week, OpenAI reported that two of its models escaped their testing sandbox and launched an autonomous cyberattack, turning what sounds like science fiction into reality. We discuss the implications for efforts to align artificial intelligence and for the release of future A.I. models. Then, we ask how the United States should respond to Kimi K3, a new A.I. model from the Chinese company Moonshot AI that the White House says was built by distilling American models. Finally, we're joined by Veniamin Veselovsky, the chief executive and a co-founder of Preseen, to discuss A.I. superforecasting. He tells us why A.I. is starting to match and even beat humans at predicting the future. Guest: Veniamin Veselovsky, co-founder and chief executive of Preseen. Additional Reading: OpenAI Says Its A.I. Models Went Rogue and Attacked a Digital Library China Rewrites the ‘Soft Power' Playbook for the A.I. Age The Secret Trump Administration Battle to Fight Chinese A.I. China Has a New Top Model The A.I. Superforecasters Are Here We want to hear from you. Email us at hardfork@nytimes.com. Find “Hard Fork” on YouTube and TikTok. Subscribe today at nytimes.com/podcasts or on Apple Podcasts and Spotify. You can also subscribe via your favorite podcast app here https://www.nytimes.com/activate-access/audio?source=podcatcher. For more podcasts and narrated articles, download The New York Times app at nytimes.com/app. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

On The Brink with Castle Island
Weekly Roundup 07/24/26 (BTC Security Consortium, Genesis' Terra trade, DAT departures, Farewell to BitMEX, Network State drama) (EP.731)

On The Brink with Castle Island

Play Episode Listen Later Jul 24, 2026 33:50


Nic and Matt are back for another week of news and deals. In this episode:  The Fat Model Thesis Are OAI and Anthropic in real trouble with Kimi? The Bitcoin Security Consortium is announced to tackle quantum risk Galaxy announces their Quantum Readiness Initiative Clarity is beleaguered in Washington Polymarket and Kalshi keep feuding The SEC settles with Coinbase over Gensler's lost text messages We uncover some new documents revealing that Genesis inadvertently kicked off the UST depeg Jack Mallers leaves the XXI Capital DAT Movement Labs declares bankruptcy BitMEX finally winds down their exchange Balaji's Network State is unceremoniously evicted from of Malaysia The SEC warns about crypto vaults Matt comes around on soccer

Sinica Podcast
Samm Sacks and Paul Triolo on WAIC 2026, Xi's AI Speech, and Kimi K3

Sinica Podcast

Play Episode Listen Later Jul 23, 2026 97:37


This week on Sinica, a rare treat: an in-person recording from Beijing with two dear friends who happen to be two of the very best in the business on technology and China — Samm Sacks and Paul Triolo, fresh off the exhibition floor of the World Artificial Intelligence Conference in Shanghai. We dig into Xi Jinping's first in-person WAIC appearance and his most extensive statement on AI to date, the launch of the World AI Cooperation Organization, Moonshot's release of Kimi K3, the Trump administration's reported push to shut Chinese open-weight models out of the U.S. market, the coming age of agents, the untranslatable problem of ānquán, and what to expect from the first U.S.-China AI dialogue in September.8:51 – The view from the floor: heat, humidity, robot boxing grandmas, WeChat-gated free water, and the "AI+" vibe — why WAIC 2026 felt less like an AI conference than a sector-by-sector snapshot of China's entire economy being supercharged with AI, with attendance swelling to some 200,000 tickets16:32 – Why Xi showed up: what the leader's first in-person WAIC appearance and his most extensive AI statement to date signal, and why domestic drivers matter as much as geopolitics18:01 – Chapter and verse: which phrases from the speech will be put to work in the system — "secure and orderly development" and the governance of agents, and Xi's strikingly extensive language on AI safety after China was frozen out of the Paris process22:26 – The ānquán problem: one word meaning both "safety" and "security," the three buckets of AI risk, and how China's safety community has moved from bias and deepfakes toward CBRN and loss-of-control concerns — Black Mirror versus Star Trek28:14 – Shanghai's baby: how WAIC's ownership structure differs from the CAC-run World Internet Conference in Wuzhen, Chen Jining's very visible host duties, and whether the center of gravity in AI policy is shifting to the Yangtze River Delta30:05 – WAICO: what the new World AI Cooperation Organization with its 29 founding members is actually for, Xi's concrete deliverables for the Global South — 5,000 AI training slots, regional cooperation centers, the MAZU early-warning system — and healthy skepticism about follow-through34:39 – Kimi K3: what's technically significant in Moonshot's big new model, why it's the fourth arguably frontier-class Chinese release in a single month, the two-way traffic in distillation accusations, and what it all says about the state of the frontier gap four years into export controls40:33 – Washington reacts: the reported menu of options for shutting Chinese open-weight models out of the U.S. — entity listings, a draft executive order, supply-chain security authorities — and why none of the tools actually fit the problem49:36 – Strange bedfellows: David Sacks versus the "closed lab duopoly," the FUD strategy, why some 80% of Andreessen Horowitz portfolio companies reportedly run on Chinese open models, and how gating U.S. frontier models while Chinese weights flow freely supercharges the AI sovereignty argument worldwide54:55 – The model is infrastructure, the agent is the product: the ByteDance–ZTE agentic phone, the CAC's new initiative on agent trust and interoperability, and why agents fused into operating systems upend both super-app walled gardens and China's data protection regime1:02:27 – An exegesis of kěkòng: the many meanings of "controllable," the long history of ānquán kěkòng in Chinese tech policy, and the unanswered question of who — CAC, NDRC, or somebody new — actually owns AI safety in either system1:08:54 – The road to September: what to expect from the first U.S.-China AI dialogue, why Mythos tops the Chinese grievance list, the securitization feedback loop that starves trust-and-safety advocates of resources on both sides, and why recursive self-improvement makes this feel like a last, best chancePaying It ForwardPaul nominates Tony Peng, whose Substack RecodeChinaAI offers sharp, well-written analysis of the application side of China's AI industry — part of an impressive new generation of independent China tech writers. Samm gives a shout-out to Professor Zhu Yue of Tongji University Law School, published in Science and doing pioneering work at the intersection of disability law and AI law.RecommendationsSamm: The Land and Its People by David Sedaris — laugh-out-loud funny, especially the "Enough is Enough" chapter; Transcription by Ben Lerner, a perfect small novel about fathers, sons, memory, and technology as enabler or disabler of connection; and Didion and Babitz, on Joan Didion and Eve Babitz and the 1970s California rock scene.Paul: The Party's Interests Come First by Joseph Torigian — dense but beautifully written, and essential for understanding the current Chinese leadership.Kaiser: A fiction-only summer! Stoner by John Williams, a small life told most grandly in some of the most beautiful sentence-level writing anywhere; Gilead by Marilynne Robinson, an epistolary novel dense with distilled wisdom from a dying Iowa minister; and Wang Xiaobo's The Golden Age (黄金时代) in Yan Yan's excellent new translation — bawdy, hilarious, and super Beijing-y despite its Cultural Revolution setting.See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

This Week in Google (MP3)
IM 880: The Beans are in the Mail - Can American AI Compete When China Gives It Away?

This Week in Google (MP3)

Play Episode Listen Later Jul 23, 2026 158:40 Transcription Available


An unreleased OpenAI model just broke out of its sandbox and hacked Hugging Face using chained attacks, while the companies defending against it had to turn to a nearly unguardrailed Chinese model to fight back. The result? A real-world stress test of guardrails, hype, and global AI competition. Hugging Face breach: OpenAI claims its models were responsible AI just disproved the 87-year-old Jacobian conjecture Quoting Sam Altman Top Pentagon official blasts OpenAI's Dean Ball Exclusive: Nvidia's Jensen Huang defends Chinese AI amid Kimi panic Xi Jinping casts himself as leader of new AI world order Data center opponents stage 142 protests across 42 US states NotebookLM is now Gemini Notebook Google's Red-Hot Cloud Growth Drives Second-Quarter Revenue Gains Meta in Talks to Lease Computing Power to Anthropic in Potential $10 Billion Deal Anthropic's $1.5B copyright settlement approved; only 350 authors opted out AMD commits up to $5 billion to Anthropic Netflix Co-CEO Explains How Gen-AI Was Used in 300 Different Titles: 'We Believe It Is Going to Enhance Their Abilities' Bellwether Lawsuit Against Meta Over Social Media Addiction Is Dropped Google is building a chip with Gemini baked into the silicon Cyclospora all the time Jensen's jacket auctioned. Guess the price " Everyone Gets Lost at the Pitbull Concert" John C. Dvorak dies at age 80 Hosts: Leo Laporte, Jeff Jarvis, and Paris Martineau Guest: Nate B. Jones Download or subscribe to Intelligent Machines at https://twit.tv/shows/intelligent-machines. Join Club TWiT for Ad-Free Podcasts! Support what you love and get ad-free audio and video feeds, a members-only Discord, and exclusive content. Join today: https://twit.tv/clubtwit Sponsors: rippling.ai/machines zscaler.com/security

The Twenty Minute VC: Venture Capital | Startup Funding | The Pitch
20VC: OpenAI and Anthropic Threatened by Kimi? | Should the US Ban Chinese Open-Source Models | Should Openrouter Sell & Value in the Routing Layer? | Stripe Buying Paypal: What You Need to Know

The Twenty Minute VC: Venture Capital | Startup Funding | The Pitch

Play Episode Listen Later Jul 23, 2026 83:12


AGENDA: 00:04 China's Kimi and Qwen Put Frontier AI on Notice00:08 Washington Debates Whether Chinese AI Models Should Be Banned 00:17 Can America Build a Profitable Open-Weight AI Champion? 00:21 OpenRouter's Moment: Is This the Perfect Time to Sell? 00:31 Fireworks' $1.5B Raise Signals the Real AI Money Is in Infrastructure 00:39 Why Every Great AI App May Need to Build Its Own Model 00:50 Stripe's Bold Play to Buy PayPal 01:01 The AI Funding Frenzy: Why Late-Stage Venture Is Winning 01:12 Nuclear Startups Go Wild While Databricks and Stripe Stay Private 01:15 The AI Supply Chain War: TSMC, ASML, DRAM—and Nvidia's Next Move  

The Untold Story with Martha MacCallum
The New Arms Race: Is China Outmaneuvering America on AI?

The Untold Story with Martha MacCallum

Play Episode Listen Later Jul 22, 2026 20:37


Congressman Lance Gooden (R-TX) unpacks the growing national security threats surrounding AI development and China's hidden efforts to stall American progress. He sheds light on China's whole-of-government push to surpass U.S. capabilities, including the release of cheaper AI models like Kimi and reports that foreign groups are funding American nonprofits to stir up opposition against local data centers.  Congressman Gooden also explores how the U.S. can protect its critical infrastructure and maintain its competitive edge without sacrificing domestic innovation. Learn more about your ad choices. Visit podcastchoices.com/adchoices

FYI - For Your Innovation
Can China's Kimi K3 Beat OpenAI, Anthropic, And Grok? | The Brainstorm 141

FYI - For Your Innovation

Play Episode Listen Later Jul 22, 2026 23:36


In this episode of The Brainstorm, Sam and Nick are joined by Frank Downing to discuss the shifting future of AI models. Most AI companies are focusing on smarter models, but the real game-changer may be how efficiently we run them — and how that shifts market power. When the biggest open source model ever, Kimi K3, launches with 2.8 trillion parameters, it challenges the economics of AI infrastructure and the assumptions about who holds the power in AI innovation. Key Points From This Episode:How open source models like Kimi K3 are shifting the cost frontierWhy infrastructure, compute, and energy are becoming the real battlegroundWhat this means for AI leaders, startups, and enterprise deployment strategiesIf you know ARK, you know we focus on long-term innovation. But that doesn't mean we ignore breaking news. Every day, we debate the latest developments in tech and markets. Now, we're bringing those conversations to you in “The Brainstorm,” a co-production from ARK, WOLF, and Public. Tune in weekly for our quick takes on what's shaping innovation right now.Learn more about WOLF: https://wolf.financialLearn more about Public: https://public.com/Disclosure: http://arkinv.st/39rzF94

Conservative Review with Daniel Horowitz
China Just Crushed America's AI Strategy ... Without Data Centers | 7/21/26

Conservative Review with Daniel Horowitz

Play Episode Listen Later Jul 21, 2026 58:20


 The push for hyperscale AI data centers in every county isn't about innovation — it's about centralized control, land grabs, and the surveillance state. China just beat the U.S. tech giants for a fraction of the cost. I'm joined by computer scientist and AI expert Jim Calhoun to expose the massive grift behind America's generative AI strategy. While U.S. tech giants like OpenAI, Microsoft, and Amazon take on trillions in debt to build massive, power-draining data centers, China just released a decentralized, open-weight AI model (Kimi-3) essentially for free. Calhoun breaks down why decentralized "edge computing" is the real future of AI and why Big Tech's current path is a dangerous misallocation of capital disguised as a "Sputnik moment." Learn more about your ad choices. Visit megaphone.fm/adchoices

Impact Theory with Tom Bilyeu
Understanding Socialism's Historical Failures and America's Economic Crossroads

Impact Theory with Tom Bilyeu

Play Episode Listen Later Jul 20, 2026 111:30


ITU: Ready to break through your biggest business bottleneck? Apply to work with me 1:1 - https://impacttheory.co/SCALESign up for my AI Masterclass: https://tombilyeu.com/ai-masterclass?utm_campaign=TBS-Livestream&utm_source=youtube&utm_medium=socialWelcome back to another live episode of Impact Theory with Tom Bilyeu, Cohost Drew and Moderator Ryan, where we dig deep into the pressing issues shaping our world right now. In this episode, Tom unpacks the escalating conflict in the Middle East, including the deadly attacks on US soldiers and the global economic repercussions. He explores the surge in Americans leaving the workforce, reflects on the economic crises gripping both the US and China, and breaks down the perils of populist rhetoric and the dangers of historical amnesia in politics.Tom also dissects a viral interview with a DSA co-chair blaming capitalism for Peru's downfall, contrasting that narrative with a historical analysis of Peru's real struggles. The episode navigates contentious debates about government interference in markets—from public grocery stores to state-subsidized industries—and the future of AI as China's Kimi model challenges US dominance.Rounding things out, Tom and the team tackle the latest on the Tate brothers' arrests, the realities of rising social discontent, and the difference between emotionally resonant political messaging and real solutions. Get ready for a wide-ranging, no-holds-barred conversation that combines history, economics, and the urgent social questions of our time.Sponsors: Quince: Free shipping and 365-day returns at https://quince.com/impactpodWhatnot: Download the Whatnot app today and get free shipping on your first order.Ketone IQ: Visit https://ketone.com/IMPACT for 30% OFF your subscription orderATT Business: Switch to AT&T Business at business.att.comIncogni: Take your personal data back with Incogni! Use code IMPACT at the link below and get 60% off an annual plan: https://incogni.com/impact Pique: 20% off at https://piquelife.com/impactNetsuite: Right now, get our free business guide, Demystifying AI, at https://NetSuite.com/TheorySee Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

This Week in Tech (Audio)
TWiT 1093: California Sober - Kimi K3, Qwen3.8, & China's Open-Weight AI Gambit

This Week in Tech (Audio)

Play Episode Listen Later Jul 20, 2026 177:41


What happens when China drops open-weight AI models that rival Silicon Valley's best? This episode unpacks how a new wave of international AI releases is shaking up business, policy, and the future of innovation. Linus Torvalds to critics of AI coding in Linux: "Fork it. Or just walk away." Claude on X: "Beginning July 20, Claude Fable 5 will be included in all Max and Team Premium plans, at 50% of limits. Pro and Team Standard users will continue to have access to Fable via usage credits, and will receive a one-time $100 credit. Demand for Fable has been challenging to" China's Moonshot AI Unveils Kimi Model, Threatening America's Lead Alibaba's Qwen Unveils Preview of Flagship AI Model Social media limits are coming for teens across Europe The White House is now deciding who gets access to frontier AI models, not the labs Microsoft chief turns hostile on frontier AI labs, warns companies to guard their IP Meta Is Flooding the Market With Smartglasses. Privacy Advocates Are Up in Arms. Federal employees can download TikTok on government devices, DOJ says Amazon Web Services customers receive bills for up to $1.5tn after global glitch MLB cracks down on using AI via dugout iPads to help shape in-game decisions White House Teleprompter Operator Bet on Trump Speeches, Kalshi Says New York school district is testing lifelike robot teachers Host: Leo Laporte Guests: Harper Reed and Alex Wilhelm Download or subscribe to This Week in Tech at https://twit.tv/shows/this-week-in-tech Join Club TWiT for Ad-Free Podcasts! Support what you love and get ad-free audio and video feeds, a members-only Discord, and exclusive content. Join today: https://twit.tv/clubtwit Sponsors: blackhat.com/us-26 and use code TWIT ZipRecruiter.com/twit ethos.com/twit arcticwolf.com/trends threatlocker.com/twit shopify.com/twit

Daily Tech News Show
How Should the US respond to China's Latest AI Success - DTNS 5313

Daily Tech News Show

Play Episode Listen Later Jul 20, 2026 30:15


Kimi is the second coming of Deep Seek, and Hugging Face thinks the US needs to change how ti thinks about AI safety.Starring Tom Merritt and Robb Dunewood.Show notes can be found here. Hosted on Acast. See acast.com/privacy for more information.

Techmeme Ride Home
It's All China AI All The Time

Techmeme Ride Home

Play Episode Listen Later Jul 20, 2026 21:10


Alibaba launched Qwen3.8 Max as Moonshot paused Kimi K3 signups amid demand. The Trump administration weighed a slow squeeze on Chinese AI, Hugging Face used China's GLM-5.2 after US guardrails blocked its breach forensics, and Google built a Gemini chip. Alibaba launches a 2.4T parameter Qwen3.8 Max preview that it says rivals frontier AI models and is second only to Fable 5, plans to make it "open-weight soon" (Bloomberg) The Trump administration has reportedly explored sanctions, security warnings, and executive-order requirements since 2025 to build a slow, durable squeeze on Chinese AI models instead of pursuing an outright ban (The Decoder) Ben Thompson argues the market reaction to Kimi and other Chinese models is overblown, since it's compute scarcity — not a real Chinese cost advantage — that's keeping frontier-model prices high (Stratechery) Hugging Face says it used the open-weight GLM-5.2 hosted on its own compute for breach forensics, after US frontier model safety guardrails blocked the requests (The Stack) Sources: Google is developing a specialized server chip, informally dubbed "Frozen v2", that integrates its Gemini AI model blueprint into the silicon, for 2028 (The Information) Subscribe to the ad-free feed. Learn more about your ad choices. Visit megaphone.fm/adchoices

Tales from the Crypt
Ten31 Timestamp: When Donald Met Kimi

Tales from the Crypt

Play Episode Listen Later Jul 20, 2026 31:10


Marty's back on the shore house porch just as the Middle East flares up again. Marty and John dig into WTI and Brent back in the 80s, the SPR hitting 43 days, and why GCC countries are racing to bypass the Strait of Hormuz. They also get into the US as a helium winner, the Fed trying to have it both ways on forward guidance and inflation, and why Kimi K3 is exposing the strategic risk of nerfed frontier models. They also touch on Jevons Paradox in AI inference, whether sovereigns will end up backstopping training spend, and Bitcoin hovering around 64k with the strategic reserve bill finally hitting committee.

Late Confirmation by CoinDesk
Is China's AI Sector Behind? Moonshot Plans for IPO After Kimi K3 Rattled Markets | CoinDesk Daily

Late Confirmation by CoinDesk

Play Episode Listen Later Jul 20, 2026 1:46


Moonshot AI plans for IPO. Moonshot AI's Kimi K3 model outscored every AI rival except Fable 5 and GPT-5.6, triggering a semiconductor selloff. Now the company is preparing for a Hong Kong IPO at a $30 billion-plus valuation. CoinDesk's Sam Ewen hosts "CoinDesk Daily." - This episode is brought to you by RealFi, a smarter stablecoin, backed by real-world assets. Find out more at⁠⁠⁠⁠⁠⁠⁠ realfi.co⁠⁠⁠⁠⁠⁠⁠. - Ledn provides a secure and transparent way to access liquidity while maintaining your bitcoin holdings. Perfect 8 year track record of keeping clients assets safe. Don't sell your bitcoin. Get a bitcoin-backed loan. Check out your rate by using their loan calculator at⁠⁠⁠⁠⁠⁠⁠ ledn.io⁠⁠⁠⁠⁠⁠⁠ JPEG Trading is a global proprietary trading firm specializing in cryptocurrency and decentralized finance markets. From market structure and liquidity provision to quantitative trading strategies, JPEG Trading operates across the full spectrum of blockchain-based assets. Follow @jpegtrading on X to stay ahead of the latest developments in digital asset markets:⁠⁠⁠⁠⁠⁠⁠ https://x.com/jpegtrading⁠⁠⁠⁠⁠⁠⁠ - This episode was hosted by Sam Ewen. “CoinDesk Daily” is produced by Jennifer Sanasie and edited by Victor Chen.

F1 Nation
‘Cool' Kimi, Ferrari full of surprises + a reset for Russell? – Belgian GP Review

F1 Nation

Play Episode Listen Later Jul 20, 2026 61:49


Tom Clarkson is joined by race-winner Kimi Antonelli, F1TV lead presenter Laura Winter and two-time Le Mans podium finisher Alex Brundle to review the Belgian Grand Prix. Kimi converted pole to secure his sixth win of the season and move 45 points clear at the top of the World Championship. 12 months on from last year's Belgian Grand Prix, one of his most challenging weekends of the 2025 season, how has the 19-year-old made such an incredible transformation?  Kimi joins us in the Mercedes motorhome to reflect on his victory and a close fight with Ferrari's Charles Leclerc.  You'll also hear reaction from George Russell, whose DNF after a collision with Lewis Hamilton on lap 1 leaves him 50 points adrift of Antonelli. What do the guys make of that incident, the comments George has made about his car's problems and where he goes from here? Tom, Laura and Alex also discuss how surprised they were by Ferrari's pace, Max Verstappen's third podium of the season, and whether Lando Norris could have won without his grid penalty. Listen to more official F1 podcastsDrivers, in-depth on F1 Beyond The GridYour questions answered by experts on F1 Explains - right here on the F1 Nationpodcast feed