Method of separating mixtures
POPULARITY
Le « moment Spoutnik » de l'IA : sortie de Kimi K3 Un modèle chinois rivalise avec Claude et GPT-5 pour cinq fois moins d'argent. Ce n'est pas une copie — c'est une menace directe sur les valorisations à mille milliards de dollars construites en quelques années à peine aux États-Unis.La Maison Blanche crie au vol, mais les outils qui font tourner la Silicon Valley utilisent déjà des modèles chinois. La distillation dont on accuse Moonshot, c'est le même processus qui a permis à Cursor de devenir ce qu'il est aujourd'hui.Ce qui bascule, ce n'est pas seulement le leadership technologique. C'est la question de qui a le droit de construire sur quoi et qui protège son oligopole en se cachant derrière la réglementation.===================⏱️ DANS CET ÉPISODE :===================0:00 — Intro0:34 — Kimi K3, le modèle chinois qui défie l'Amérique3:03 — Une architecture innovante et des coûts inédits11:47 — Distillation: voler ou s'inspirer légitimement17:52 — 5 milliards contre 100: le choc de valorisation19:30 — La contrainte comme arme secrète des Chinois22:16 — Les LLM plafonnent26:03 — Jensen Huang contre Washington: la Silicon Valley en révolte34:02 — Kimi K3 s'adresse à qui concrètement=============
“You just do it by the old way. It's not difficult. It's not fast, but it's not difficult. If we make it fast, you have to correct a lot of things. When you give the real time for a distillation, you get all the flavors. Give it the time, the real time that it needs to make it.” That's Alvaro Fernandez, founder of El Acabo Raicilla and Asil Raicilla de la Sierra. We profile those brands — and talk about the Raicilla denomination of origin — in this episode of Agave Road Show! Agave Road Show is an advertorial podcast that helps gringx bartenders better understand specific brands of agave spirits. Sometimes other Mexican spirits. Really, whatever they pay us to profile. This episode is hosted by Lou Bank with special guest Linda Sullivan of seynasecreto with tons of quotes from Alvaro Fernandez and Fausto Romero from Taberna Tres Gallos, makers of El Acabo Raicilla and Asil Raicilla de la Sierra! Episode Notes You can read the current NOM for Raicilla here. And to see the rules of all of the NOMs for alcohol in Mexico, click here!
My guest today is Sam Altman, CEO of OpenAI. It's a conversation spanning the history, present, and future of OpenAI, from the origin of ChatGPT through Codex, hardware, and their new Jalapeno chip. We discuss the early decision to buy compute at a scale nobody thought was rational, and the plan to build a gigawatt of new capacity every week. We talk about Kimi and distillation, the Hugging Face incident and what it means for the pace of AI development, and what it's like to raise kids who will grow up never knowing a world without abundant intelligence. Please enjoy my conversation with Sam Altman. For the full show notes, transcript, and links to mentioned content, check out the episode page here. ----- Become a Colossus member to get our quarterly print magazine and private audio experience, including exclusive profiles and early access to select episodes. Subscribe at colossus.com/subscribe. ----- Ramp's mission is to help companies manage their spend in a way that reduces expenses and frees up time for teams to work on more valuable projects. Go to ramp.com/invest to sign up for free and get a $250 welcome bonus. ----- Trusted by thousands of businesses, Vanta continuously monitors your security posture and streamlines audits so you can win enterprise deals and build customer trust without the traditional overhead. Invest Like the Best listeners get a special offer of $1,000 off Vanta when you go to vanta.com/invest. ----- WorkOS is the infrastructure B2B and AI-native companies use to sell to enterprise. It covers everything enterprise security requires: SSO, SCIM, RBAC, Audit Logs, AI governance, and more. Trusted by 2,000+ fast-growing companies, including OpenAI, Anthropic, Cursor, and Vercel. ----- Rogo is the AI platform for finance. They're building agents for Wall Street that are trained to understand how bankers and investors actually do work: from diligence and modeling, to turning analysis into deliverables. To learn more, visit rogo.ai/invest. ----- Ridgeline has built a complete, real-time, modern operating system for investment managers. It handles trading, portfolio management, compliance, customer reporting, and much more through an all-in-one real-time cloud platform. Visit ridgeline.ai. ----- Editing and post-production work for this episode was provided by The Podcast Consultant. Timestamps: (00:00:00) Welcome to Invest Like The Best (00:02:02) Intro: Sam Altman, CEO of OpenAI (00:02:35) Refocusing (00:05:43) OpenAI's Compute Bets (00:09:07) Data Centers (00:11:14) Jalapeno Chip (00:11:52) Kimi, Distillation & Open Source (00:14:39) The Hugging Face Incident (00:17:46) OpenAI's Mission & Vision (00:22:14) All the Returns Are at the Frontier (00:22:27) Bottlenecks: Compute, Research, Data (00:23:49) Sam's View on AI & Jobs (00:26:56) Unpopular Bets That Turned Out Right (00:27:45) Model Cycles (00:29:45) How Sam Uses AI (00:32:44) Having Kids (00:34:56) Why Sam Has No Equity in OpenAI (00:35:33) Robotics (00:36:48) The Origin Story of ChatGPT (00:39:22) How to Get AI into More Hands (00:42:20) How Sam Recruited Great AI Researchers (00:43:57) What Sam Learned From Being an Investor (00:45:22) What the Next 6–36 Months Look Like (00:46:31) Codex (00:49:36) Could We Be Oversupplied in Compute in Two Years? (00:50:09) Sam's View on Scaling Laws (00:50:20) Alec Radford (00:51:12) Formative Moments (00:53:50) Kindest Thing
Marty and John dig into the Houthis blocking the Bab el Mandeb strait, oil whipsawing on peace headlines, and why Lloyd's of London pulling war coverage matters for US insurers. They get into Trump's air defense stockpile problem, Chinese diesel demand falling off a cliff, and why the MOVE index staying subdued even as yields rise is the chart to watch. The bulk of the show is an extended breakdown of the open versus closed AI model debate, why the Hugging Face hack exposes the absurdity of current guardrails, and how the real axis of conflict is US versus China, not open source versus safety. They close with Bitcoin long term holder supply hitting 84 percent, a classic bottoming signal.
In this episode we explore the history of the alembic pot still and distillation on our way to discovering Arak, an anise-flavored spirit from the Middle East, which most likely inspired your favorite anise tipple. Resources from this episode: Books: The Oxford Companion to Spirits and Cocktails [Kindle Edition], Wondrich, D & Rothbaum, N., (2022) Websites: Atlas Obscura: Thank These Master Alchemists for the Magic of Alcohol, Eubank, A. (12 February 2018) https://www.atlasobscura.com/articles/distillation-alcohol-invention-muslim Britannica: Arak, Payne, L. https://www.britannica.com/topic/arak-beverage Dharma Dispatch: Bengali Arrack and Bouleponge - A Glimpse into the Savage Lifestyle of European Naval Merchants and their Crew in the 17th Century Mughal Empire, Balakrishna, S. (26 January 2022) https://www.dharmadispatch.in/bengali-arrack-and-bouleponge-a-glimpse-into-the-savage-lifestyle-of-european-naval-merchants-and-their-crew-in-the-17th-century-mughal-empire/ Difford's Guide: Arrack, Difford, S. and Sutcliffe, T. (n.d.) https://www.diffordsguide.com/beer-wine-spirits/category/949/batavia-arrack Liquor: How to Use Arak in Your Cocktails, Dingwall, K. (14 September 2022) https://www.liquor.com/arak-in-cocktails-6666222 RTL Today: Raisin moonshine banned in Iran enjoys resurgence in New York (2 May 2026) https://today.rtl.lu/news/world/raisin-moonshine-banned-in-iran-enjoys-resurgence-in-new-york-443728357 The Spruce Eats: What is Arak? A Guide to Buying and Drinking Arak, Fayed, S. (Updated 20 January 2023) https://www.thespruceeats.com/arak-middle-eastern-alcoholic-beverage-2355492 Tasting Table: The Cultural Significance Behind Arak, A Timeless Levantine Drink, Holland, J. (10 July 2023) https://www.tastingtable.com/1330174/cultural-significance-arak-levantine-drink/ This Week in Palestine: Distilled Spirits an Arab Invention, Muaddi, N. (n.d.) https://thisweekinpalestine.com/distilled-spirits-an-arab-invention/ Wine Enthusiast: Understanding Arak, an Ancient Spirit with Modern Appeal, Teclemariam, T. (Updated 8 May 2023) https://www.wineenthusiast.com/culture/spirits/arak-middle-eastern-spirit-modern-appeal/ Wooden Cork: What is Arak? A Comprehensive Guide to this Unique Spirit (17 July 2024) https://woodencork.com/blogs/uncorked/what-is-arak-a-comprehensive-guide-to-this-unique-spirit Glass in Session Episodes Referenced in or Related to This Episode: S8E4: Let's Absinthe https://glassinsession.libsyn.com/s8e4-lets-absinthe S10E3: Getting our Grappa On https://glassinsession.libsyn.com/s10e3-getting-our-grappa-on S12E4: Rum - King of Booty? https://glassinsession.libsyn.com/rum-king-of-booty-s12e4 S15E6: Parlez-vous Pastis? https://glassinsession.libsyn.com/parlez-vous-pastis-s15e6 S22E1: Wine in Lebanon - Wine, War, and Rising https://glassinsession.libsyn.com/wine-in-lebanon-wine-war-rising S22E4: Amenian Wine - Of caves, mountains, wars, and resilience https://glassinsession.libsyn.com/armenian-wine-of-caves-mountains-wars-and-resilience-s22e4 Glass in Session® swag mentioned in this show: https://www.teepublic.com/user/glass-in-session Glass in Session® is a registered trademark of Vino With Val, LLC. Music: "Write Your Story" by Joystock (Jamendo.com cc_Standard License, Jamendo S.A.)
In recent months, the open vs closed, and US vs China discussions on model ownership and sovereign/local AI have heated up to a fever pitch. So it is very very good news that Poolside AI are finally emerging with new models, like Laguna S 2.1, that are beating Thinking Machines' recent release nearly 10 times their size.Poolside's recent tech report got a lot of praise due to their level of detail, and Vibhu first covered Laguna's recent technical report on our paper club:From spending $12 million building language models for code before the world cared to creating a Model Factory that can take a model from pre-training to release in eight weeks, Eiso Kant has spent more than a decade betting that code is the path to AGI. In this episode, the Poolside co-founder joins swyx and Vibhu to explain why ChatGPT felt like vindication, why Poolside embraced open weights and open research, and why he would rather live in a world with 100 foundation model companies than five even if Poolside were one of the five.We go deep on Poolside's Model Factory: the engineering systems behind 10,000–20,000 experiments per month, streaming data directly into training, reproducible experimentation, low-precision compute, and agents that increasingly write code, launch jobs, evaluate results, and modify the pipelines used to train future models. Eiso also unpacks their recent launch Laguna S, why persistence, verification, and backtracking may matter more than raw intelligence, how much capability remains inside smaller models, why reinforcement learning will move earlier into pre-training, and why next-token prediction is still extracting too little from the web.We also discuss model-harness co-design, Poolside's path from coding agents to AGI, why Eiso thinks MCP and traditional tool calls are “stupid,” the real economics behind frontier-model training, Poolside's $500 million raise, open-source AI, regulation, NVIDIA and TSMC's influence, engineering productivity in the agent era, high-agency teams, and hiring at Poolside.We discuss:* How Andrej Karpathy's RNN work inspired Eiso to start building language models for code in 2015* Why Eiso spent four years and $12 million pursuing an idea before the market cared* Why ChatGPT felt like vindication and brought Poolside back to open source* Why Eiso would prefer 100 foundation model companies over an oligopoly of five* The difference between releasing open weights and publishing genuinely open research* Why Poolside deliberately built a global research organization outside the Bay Area talent war* Why model building is ultimately 90% engineering* The Model Factory: Poolside's end-to-end system for rapidly training and improving models* How fewer than 70 researchers run roughly 10,000–20,000 experiments each month* How Poolside moved from six-month model cycles to five- and eight-week launches* Why streaming data directly into training unlocked faster experimentation* How immutable data, versioned code, and reproducibility enable rigorous model research* Why Eiso wants capable researchers to leave their labs and become Poolside's competitors* Why 95% of model building can be reduced to better data or compute efficiency* Laguna S and why persistence, verification, and backtracking can outperform raw intelligence* Why smaller models may handle far more knowledge work than previously expected* Why reinforcement learning will move earlier into pre-training* Why next-token prediction is still failing to extract enough knowledge from the web* Why distillation and environments have become the AI industry's favorite “drugs”* Why mid-training is really an early form of curriculum design* Low-precision training, networking bottlenecks, and the next gains in compute efficiency* Laguna S: 118 billion total parameters, 8 billion active, and eight weeks from training to launch* Why model builders can often evaluate a new checkpoint within its first 30 minutes* Model versus harness: where agent capabilities actually come from* Why Poolside sees coding and long-horizon software tasks as a path to AGI* Why Eiso thinks MCP and traditional tool calls are “stupid”* Why future agents will write scripts instead of choosing from dozens of predefined tools* The case for minimal harnesses, containers, and model freedom* Why Poolside is prioritizing vision but does not expect to work on audio soon* Why language may be the most compute-efficient modality for encoding knowledge and reasoning* The real cost of model development and why the final training run is anticlimactic* The story behind the Poolside name and why it represents refusing to lower ambitions* How Poolside raised $500 million while investors still questioned whether AGI was real* Why intelligence could become the world's most demanded and commoditized resource* When open models may become too capable to release without restrictions* Why unilateral AI safety does not work in a globally competitive environment* How regulation could accidentally lock in an oligopoly of two or three AI companies* NVIDIA, TSMC, and the hardware systems underpinning foundation-model progress* Why reinforcement-learning wall-clock time is one of Poolside's biggest bottlenecks* Why Poolside trains models from scratch instead of simply distilling larger models* How AI changes the way companies should measure engineering productivity* Why agency may become the most important quality for employees in the AI era* How leaders align high-agency people through shared goals and clear constraints* Hiring across research, post-training, pre-training, architecture, evals, and engineering at PoolsideEiso KantLinkedIn: https://www.linkedin.com/in/eisokantX: https://x.com/eisokantPoolside: https://poolside.aiTimestamps00:00:00 Introduction00:00:54 Karpathy, RNNs, and Building Code Models Before Transformers00:02:26 The $12M Failure and ChatGPT Vindication00:03:39 Open Source and the Case for 100 Foundation Model Companies00:09:22 Open Weights, Open Research, and Poolside's Global Team00:16:04 The Model Factory: Why Model Building Is 90% Engineering00:20:19 Agents, Automated Experiments, and Early Signs of RSI00:24:04 Streaming Data, Reproducibility, and Scientific Rigor00:30:35 Creating More Foundation Model Companies00:36:07 Laguna S: Persistence vs. Raw Intelligence00:43:01 Reinventing Pre-Training, RL, and Curriculum Design00:52:33 Low-Precision Training and Squeezing More From Smaller Models00:58:37 Model Harnesses, Coding Agents, and the Path to AGI01:09:26 Why MCP and Traditional Tool Calls Are “Stupid”01:13:04 Vision, Multimodality, and Why Language Still Matters01:18:15 Scaling Models and the Real Economics of Training01:20:40 Why Poolside Is Called Poolside and Raising $500M01:27:37 Open Models, AI Safety, and the Risk of an Oligopoly01:33:53 NVIDIA, TSMC, and the Reinforcement-Learning Bottleneck01:41:52 Smaller Models, Distillation, Engineering Productivity, and HiringTranscriptIntroduction: Eiso Kant, Poolside, and Open ModelsSwyx [00:00:00]: All right, we're here in the studio with Eiso Kant from Poolside, together with Vibhu. Welcome.Eiso Kant [00:00:08]: Thanks. Thanks for having me, guys. Good to be here.Swyx [00:00:10]: Yeah, fresh on the plane. You texted me, you were like, “Hey, I'm on my way to SF.” I was like, “You're on a plane right now, right?” Like, hey.Eiso Kant [00:00:16]: I know. After I texted you, I realized that probably coming in with major jet lag was gonna offer some fun experiences today, but let's do it.Swyx [00:00:23]: I mean, I think the thing I would tell guests is that they don't have to prepare that much because if you're truly working on this every single day, then even, like, what you hazily remember is going to be new for a lot of the audience that don't live in your world every day, right? so 10 years ago, you did a talk at Google Slush, talking about the democratization of AI. and, now here you are, like, open sourcing an incredible new model that we're gonna talk about. But I guess, like, what got you into democratization of AI? Like, it's not obvious from your LinkedIn or something.From Karpathy's RNN Post to SourcedEiso Kant [00:00:57]: No, it's not at all. I don't think it's obvious how I got in this space. I owe getting into this space to Andrej Karpathy.Eiso Kant [00:01:05]: In 2015, he wrote an article called “The Unreasonable Effectiveness of Recurrent Neural Nets.”Swyx [00:01:10]: Neural Nets, yep.Eiso Kant [00:01:11]: And that article, I read it, and I pivoted my startup at the time overnight to working on RNNs, and later LSTMs and Transformer models to be able to write code. If you go to this article and you scroll down, you can start seeing, like, this was the precursor to what ended up becoming language models. So, at least when he was character-level language models that were starting to predict letters, he has an example out here. There's a little Paul Graham generator, and you can read it, and the text makes sense, but it doesn't. and there's a little-- There's an example of code a little bit further down. Yeah, so Shakespeare.Swyx [00:01:47]: Shakespeare.Swyx [00:01:49]: CoolEiso Kant [00:01:49]: And for some reason, I read this, and I went down the rabbit hole of learning everything I could about RNNs and LSTMs, right? This is Transformer paper. And I had built a completely unreasonable belief, that neural nets should be able to generalize to anything and everything, and that language should be able to generalize, to a lot of things that are intelligent and the ability to write code. And so I started building Sourced, which was a fully open source company trying to build, what we used to call machine learning on code, language models on code. And we spent about four or five years on this, till the end of 2019. And that sounds really cool today, but back then, no one cared.Eiso Kant [00:02:29]: Right? Like, no one cared. We were in the dark. Like, we did things along the way. We tried applying convolutional neural nets to, like, the structure of code. We were. when attention came out, we were applying it to LSTMs, and then the Transformer paper came out. And it - it wasn't obvious, and what we missed throughout that entire journey, that we were on the right track, but we should have just kept scaling up. And today, to all of us, the scaling laws and scaling up seems like the most obvious thing. But having spent four or five years of my life on working on language models on code, it wasn't obvious. So I have a lot of respect to folks at Google and OpenAI and others who took that confidence and kept going. we failed ultimately at the time, and it was, like, biggest failure of my career, right? You blew $12 million of investors' money, which was a lot back then.Swyx [00:03:18]: Yep.Eiso Kant [00:03:19]: You spent, still a lot, but, And you spent years with, like, a group of 40 people just obsessing over this problem. And life took a different turn, And it was, and family became a focus, and I kept my heads down and really, didn't really look at language models for the following two years. big mistake considering Following years are gonna be really interesting. And then ChatGPT came out And it was like a vindication. It's like people started texting me. I found, like, my old, work decks and these old talks. And throughout that whole journey, we,ChatGPT, Vindication, and Returning to Open SourceEiso Kant [00:03:56]: We really had a strong point of view at the time that, like, as you're building more capable intelligence, it should be open and open source.Eiso Kant [00:04:04]: When we started Poolside, that wasn't the case at all, and I wanna be very open about it. When we started Poolside, we were like, there was a premise of two things. One is this technology is not gonna stop compounding in capabilities. I think to most people obvious today, but three-plus years ago when we started, most people were still arguing if these were stochastic parrots or not.Eiso Kant [00:04:23]: And the second was that reinforcement learning was gonna be the biggest driver for LLM capabilities. Today, very obvious. Three years ago, was not an opinion held or direction held at either OpenAI or Google or Anthropic or others. And so people looked down on us a little bit. They were like, “ is this really gonna work?” And so we just started working the problem, and we never really thought about open source again. We just kept our heads down and we built our, like, knowledge, understanding from scratch, right? We didn't roll out of an existing lab. So we picked up the papers and started writing code and figuring things out.Eiso Kant [00:04:59]: And it wasn't until the beginning of this year that me and my founder, Jason, picked up the open source conversation again.Eiso Kant [00:05:07]: And if you go back to some of the early things on our website, it was very straightforward. It was we wanna get to AGI, we wanna support a world of abundance, and we wanna be the first company that gets there.Eiso Kant [00:05:20]: But we started talking at the beginning of this year because it became obvious that the world was going in a direction that was starting to like, pick at us a little bit. Like, it didn't, this didn't happen overnight. It was, like, a little bit we were seeing this and we're like, “Okay, The world's going down a path.” And Throughout this journey, there was something that I used as a, as an analogy or thing. So I said well, if I go back to back in those days, 2015 or 2016, we're working on this, and I picked up a fi book off the shelf, and I was reading the book about 2035. AGI is achieved, and the story would be over the following, decades. And it would have that first chapter where everyone's trying to figure things out. You'd get the chapter of ChatGPT coming out And then you would get to the chapter where the world was at a fork in the road, and the one that it picked was one where three or four or a handful of companies were going to create all of intelligence moving forward.Eiso Kant [00:06:21]: And when I thought about that story, it felt like a dystopian fi book, not a utopian fi book. And the reality is, I'm a utopian fi guy. Like, and so We took a step back and said, “Hey, can we play a role here?” Now it was easy for us to do so because we were not at the frontier.Eiso Kant [00:06:41]: If we were at the frontier, I don't think we could have changed our mind. and I don't mean this like it's when the moment there's too much capital involved, too much expectations, you've built up things, right? We're a small team, just improving and improving. And so we knew that we could make that decision now, but it would be a lot harder to make as we got closer and closer to the frontier and caught up to others. And did a lot of soul-searching and a lot of conversations, and said, “No, this makes sense,” Even if there's big unanswered questions, like how the hell do you build a business model with foundation models about open source? Big open-ended question that we do not fully have the answer to yet, right? At what point do you no longer wanna release open source models because misuse of models has, real potential risks associated with it? how is the government gonna respond to open source? but I think it all just came down to one thing, and I'll stop the monologue, is the fact that I rather live in a world that has 100 foundation model companies than a world that has five, even if I was one of the five. And the smallest and most meaningful contribution we can make for 100 to exist is to open up our research and open up, like, our weights right now and figure out along the way how we can, like, do more.Neo-Labs, Model Choice, and the Token EconomySwyx [00:08:01]: Yeah. I think if anything, over the past three years, that has become a bit more true. you are one of a cohort of Neo labsEiso Kant [00:08:10]: YeahSwyx [00:08:10]: That people are now calling that. And, we're, we're doing this on the day that Thinky launched their, new model and you are outperforming them on their, on some benchmarks that they released, right? Like, they just don't have it yet. so it goes to show that I think, like, this is one of those things where, like, there is room for multiple players, and you are seeing a little bit more of the future. Maybe more like 20, not 100, but, like, you are one of the 20.Eiso Kant [00:08:36]: I really hope so, right? I think we I'm, I'm excited about their release, and I'm excited about everyone releasing because, like, ultimately, like, choice competition is both gonna drive progress in the right direction. But the fact that like, we create models and while we all, drink out of the same well of data effectively, we do introduce very different behaviors and biases in our models. Some are intended biases, some are completely unintended biases.Swyx [00:09:03]: Yeah.Eiso Kant [00:09:03]: And if we shape up in an ecosystem in the world where open models are gonna be a part of the token economy, like, I don't think there's any question about it anymore Then we want to be able to live in a world where companies, countries, people can choose and say, “Hey, I am most aligned and I trust most this provider for these things.”Swyx [00:09:25]: Yeah.Vibhu [00:09:26]: I think more than just one of the 20 Neo labs, up until recently, most of open source innovation was coming from the Chinese labs, right? So there's the DeepSeek of the West. Is it today? Okay, maybe it's thinking machines reflection, but there aren't many, right? So, one of the things you guys started in France, Europe, but very much now you're taking that American standpoint and more than just that, the point is the Chinese models that we see, they're not super open research. the work you put out is, I think, some of the best. So every few months you get not only frontier models, but also here's a breakdown blog, paper, technical report of here's everything for state of the art to build, frontier intelligence and you're filling that gap too, right? So not just only open weight, not just Western, but also pretty open research.Open Weights vs. Open ResearchEiso Kant [00:10:20]: No, I appreciate it. Look, I think it's, I think it's the most meaningful contribution, right? Weights are a binary. Let's call them what they are. Yes, we can modify them, we can change them, but, like, giving someone the weights does not allow them ultimately to recreate what you're doing, right? And so now there's challenges around releasing data sets, challenges around like releasing certain things, but being able to share your research, like, right, how do we do it? What are the lessons we learned that we spent, tens of thousands of experiments of compute on? I think very much so. One correction though, Vibhu, and I say this because it's been haunting us for quite a few years. We from day zero were an American company.Swyx [00:10:55]: Yeah. They movedPoolside's Global Team and American Company StorySwyx [00:10:56]: To France.Eiso Kant [00:10:56]: So the story once and for all is very. We start as an American company. We have always been an American company, and early on we made a very conscious decision. We said, “We're not gonna hire any researchers in the Bay Area. We're gonna look for talent everywhere else in the world.” and that is everything from Middle Americas, Seattle to, Serbia, and to Taiwan and Singapore and other places. And it was because we took a view that this was gonna become a talent war for this, and I think it has over the years now. Three years ago, that wasn't fully obvious yet. I think today it very much is. And we also realized that, like, some of the world's most capable people with, like, the most interesting, innovative ideas were not just gonna be here. And so it led us to create like a fully remote company. and we ended up opening an office in Paris and London and different places and we have a lot of the team in the US and a lot of team outside. But we always took this view of like, we're an American company, but if we want the best of the best to work with us, we need to take a global view. Now we do also have people here in Silicon Valley, like the company's grown and others, but I think one of the things that, it slowed us down at the beginning, but it has sped us up now, and it's why you're seeing like the progress, I think, on our models and the cadence at which we release, is because we didn't roll out of an existing lab. Right? we didn't, we didn't have a lot of the information that's freely flowing around here at the time. We just took this point of view as like, “Okay, well, let's just work the problem. Let's just go and, like, read the few papers that are out there, and let's just figure this stuff out.” And we made some hilarious mistakes in model training because of that over the yearsEiso Kant [00:12:35]: Like especially in the first 12 months. there's a few that I think still haunt me and scare me. We can talk about them later. but it created a, like, a resiliency and persistency in the team, right? with extremely few people have left us over the years, that, like, told us, “Okay, we can do this.” When we first wrote our first training code base completely from scratch, it wasn't a fork of any open source. It was just like, “Okay, let's build it from scratch.” I remember we had this one moment where we spent three weeks working out an optimizer bug. Like, it was like training just couldn't get stable. We, like, obsessed over it, and we thought, like, maybe we were wrong. Maybe we should have just forked this repo, or we should have. But then when we solved it, I still remember at the time we were like five people in the company. when we solved it, we were like, “Oh, we can do things,” like if we're just willing to work hard. and I think that culture with a very strong engineering bias has helped us, like, get to where we were. And so there's this notion of open source and talent and these things. I think we, We just took different decisions from a different starting point. and I think we are lucky. I do want to definitely call it lucky. And there was a lot of hard work at the team that now, like, that's starting to show up in results.Swyx [00:13:52]: Just ‘cause we probably won't revisit this again, but, and this is a fun recruiting challenge if someone knows the answer. What was the bug? And then we won't tell the solution, but we'An Optimizer Bug and the Value of Building From ScratchEiso Kant [00:14:01]: So the - This - You're gonna test my memory here,Swyx [00:14:04]: Oh, okayEiso Kant [00:14:04]: So but I thinkSwyx [00:14:05]: DirectlyEiso Kant [00:14:05]: I think I can recall. So if you, so if you look at, So if you take like Adam as an optimizer, you have epsilonSwyx [00:14:12]: YeahEiso Kant [00:14:13]: Which is, right, like in the denominatorSwyx [00:14:14]: Momentum and weights. YeahEiso Kant [00:14:15]: Is exactly, in the denominator. And at the time, if I recall, you looked at like the early Llama papers and things like that. People were juicing epsilon, like, quite a bit. Like, they were, like, adding, I don't know if it was E minus four or whatever, like a high value for epsilon.Eiso Kant [00:14:31]: And if you think about this during training, it's like a bit weird and counterintuitive that we're adding noise to our optimizer by just adding effectively, like, a random number in the denominator, right? Like behind the decimal point. And I don't recall the exact bug, but it had - What I remember is once we solved it, we no longer had to juice epsilon as much as, like, was happening in the Llama paper and other places. and it was like one of those fundamental moments where we had trusted this paper that was out there, and we're like, “Oh, no, it has to be this way. It has to have this high value of epsilon.” But it made no sense to us intuitively. Like, why do you have to have this so high? Like, if you're just trying to avoid division by zero, why can't the value be extremely small? and that was like one of those moments where you realize like, okay, finding things out from scratch yourself builds a better intuition. Because the one thing you learn very quickly with model building is that your intuitions that you start with are gonna get beaten up so hard.Eiso Kant [00:15:33]: Right? Like - It's such an experimental science, that the things that seem obvious, you very quickly get to learn, like, you were wrong, and hopefully you figure out why, and sometimes you don't even.Swyx [00:15:45]: Yeah. yeah, so, one of the reasons that you, when you released your new models, Vibhu got really excited. I mean, everyone got really excited. But Vibhu led our paper club on it, and you guys sawEiso Kant [00:15:58]: YeahSwyx [00:15:58]: Obviously. maybe talk through some lessons learned in that, whatever you can disclose. we can focus on the model factory stuff, whatever you think is a good starting point.Model Building as EngineeringEiso Kant [00:16:08]: So I would say that our view from very early on in the company was that model building is ultimately 90% engineering.Eiso Kant [00:16:18]: And I think we all know it in the industry because if you look at where's every researcher spending their time, they're spending their time writing code, right? Looking at data and writing code. And so we said, okay, The state at the moment, like three years ago, was bash scripts and Slurm and spaghetti code bases for training and, like, data pipelines that were patched together. And we looked at this and said, “Well, ultimately, model building is a process.” You're going from raw data, right? Like training raw material, the web, et cetera. you're doing a whole bunch of filtering, cleaning up, transformations, analyzing. These days, that's, far more complex than it was three years ago. then you're training a model, which is effectively a large distributed systems problem, right? Across hardware that has still-- It's become a lot more reliable. It was extremely flaky back then. and now with every new generation, we get our new sets of challenges. And then you go into the next stages, right? There was no training back then, but, like, you got, your post-training and then your reinforcement learning. And so we looked at this and we said, “Well, this looks like an industrialized process. This looks like an end process, that every single part of it has its machinery,” right? If it's your big data pipelines, if it's your crawling ingestion of the web, if it's your, large-scale distributed training, and then you've got your reliability. And we said, “Well, why don't we take some of the world's smartest distributed systems engineers that we knew and make them part of the process of research from day zero?” Not retrofitting it later on, but, like, really from the beginning. And that became our model factory. And so our model factory started with a handful of components. Today, it's thousands of components, and I try to equate it to, if you think about, like, someone who was at the very early days of Foxconn, if they had been there for the following, decade, they would be able to rebuild Foxconn because they saw every decision that led to building that system and all the complexity. If you and I walk into Foxconn today, no chance.The Model Factory and Experiment VelocityEiso Kant [00:18:18]: Right? Because we don't have the lineage and history of decisions that led to that. And so we built early on from the beginning- with a team that really understood that, well, the metric that we are optimizing for is the speed of an idea from a researcher to an experimental result that we can trust to then being part of the next model training.Eiso Kant [00:18:42]: And in the. And because it's such an experimental science, ultimately, in the beginning when it wasn't that complex, you could patch your way around it, right? But now, at any foundation model company, you are running. I mean, we're a small team, right? We're less than 70 researchers, another 35 engineers. and we are running, I haven't checked the latest count, but far more than 10,000, maybe 10 to 20,000 experiments a month that we cut. And so if you look at that scale of every model run that is, like it's ultimately it's, it's you need to be able to trust it as an infra problem. And so what we have now done over the years is gotten really good at that, and just by working it and improving it and obsessing over those end decisions. So now what that means is that you looked up Laguna XS 2 that we launched. It was five weeks from the beginning of training to launch. The model that we're gonna talk about today was eight weeks from start of training, to launch. We started the next model literally yesterday because we now finished the post-training required for the model we're launching, next week or by the time this comes out today. and we move that compute to the much larger Laguna M model that we're now training. And so the model should be an artifact of someone's process. It shouldn't be really a thing in itself. Like, and we treat this like the way you would look at like a SpaceX factory where, yes, the first rocket, really hard to build, but the much harder challenge was building the factory. And now they're rolling off, and no one is really thinking about the next launch anymore. So it's just another launch, it's another launch, another rocket comes off. And that's what we're trying to do with model building.Eiso Kant [00:20:22]: And what has been, which was not planned from day zero, it was in the back of our mind like this will happen one day, is that when you build a really good end model factory with really good APIs and really good engineering systems, Well, what is it perfect for? It's perfect for agents.Agents Inside the Model FactoryEiso Kant [00:20:40]: Because agents are now starting to take over more and more work in our model factory.Vibhu [00:20:43]: Yeah.Eiso Kant [00:20:44]: So I look at the screens when I walk, like when we're, we come together, in our monthly, we do monthly onsites, and I walk behind people's screens and I stop by and I talk to our researchers. And the default is all of these different agents running on their screen that are writing the code. They're launching the jobs. They're evaluating the results that are coming back from the model runs. They are, making the changes. And we're still in the driver's seat. We're still coming up with the ideas. We're still helping with the debugging. But more and more, and this is right now very profound on the data side of our pipelines in both pre and post and the synthetic data pipelines, it's starting to become more on the architecture side as well. You're starting to see these twinklings of what RSI is gonna look like.Eiso Kant [00:21:27]: And that's. So when we talk about, like to your question about our models, every talk about the model factory, And my coolest example of these things is always that when we kick off a new run, doesn't matter if it's a training like big run or if it's now a post, like one of 10 post-training versions we do for like release or many experiments, is that at any given moment, the changes that somebody made that they had experimental results from the day before make it into that run.Eiso Kant [00:21:57]: So there's not like a cutoff 90 days before. Like no, it's like literally from that moment because we can now trust the machine enough. And then you also have to invest in the reliability. So one of my favorite metrics about like Laguna S is that there was no call events, Right? Like completely zero. And we haven't had a meaningful call event, like something to wake up for, as far as I recall this entire year. now there is one asterisk to that. In usually the first six hours of launching a new model run, something breaks because you set a config wrong, you made a small mistake, et cetera. So that's usually there's a little bit of intervention, but that's always within like call periods, right? Not on call. And I think that's starting to now compound. So the model we're releasing now, I love it. It's amazing, but we're already onto the next one. and I think that's the way it should be.Laguna, Five-Week Builds, and Zero On-Call EventsVibhu [00:22:50]: Hey, I also just wanna point out, so for context, this was like a month ago. we found it in the tech report, so we just came in with, “Okay, new model's dropped. Haven't heard about it.” We wereEiso Kant [00:23:02]: Yeah, we're very used to doing this every few months.Vibhu [00:23:03]: We're, we're very much like, “ okay, look, it's like, on par with Kimi, DeepSeek, whatnot, the small ones, Gemma level. Oh, it's a very cool paper on what goes into building.” And then we hit this page, right? Like literally page two of tech report is, “This process allowed us to build the small model from scratch to delivery within five weeks applying the lessons”. And then I'm like, oh, this paper is not about here's a tech report of benchmarks and here's how many tokens it was trained on. Like for people that wanna dive more from what we're not gonna discuss on the podcast, it's all laid out here, right? FromEiso Kant [00:23:38]: YeahVibhu [00:23:39]: Custom software that agents can use to interface with training code, training data.Eiso Kant [00:23:45]: Yeah. Well, link the paper correctly, so yeah.Vibhu [00:23:47]: Yeah. All that stuff. read the paper here, but,Technical Report Principles and Streaming Training DataEiso Kant [00:23:50]: But I would like to. I love principles, and I think that is a good starting off point for maybe telling some stories. Maybe we can go one by one past the principles. I'll just call out that Dagster just got bought by a Prefect.Vibhu [00:24:01]: Yeah.Eiso Kant [00:24:01]: Isn't it fun? But yes, I'm very familiar with Dagster. just anything where like they trigger some story.Vibhu [00:24:07]: So, well, I would say, well, experiments code's obvious, but I think one of my favorite things is, I don't know where it is in here, but early on, and I still think this is the case a lot of foundation model companies, people prepare their training data sets, they get packaged up, then they get copied over to a training cluster distributed across all of the nodes, and then training starts.Vibhu [00:24:30]: And we looked at this like three years ago and we were like That makes no senseEiso Kant [00:24:36]: You lose so much time because the moment you have to rematerialize the data set, you have to make a change, you have to fix something, et cetera, you've got all this time of like repackaging it, right? Toca- tokenizing it, repacking it, moving it over to a cluster, then distributing it across the nodes. The bigger your clusters are, you start using fancy like torrent-like algorithms to like distribute your data. So why aren't we streaming data into training? Right? Something that's very common and like just basicVibhu [00:25:00]: Like just in timeEiso Kant [00:25:01]: Just in time, like good computer science like principle. And that was one of the first things that I think unlocked - the model factory. Because the moment you start thinking about, well, a training job, it doesn't matter if it's a big hero run or a small like, post-training experiment, consumes a certain number of tokens per second, right? And it's not a lot, right? From a like a data, moving data perspective. So we said, well, we have our training cluster, and then we've got like our AWS kinda setup where we can build these amazing big data pipelines. We can set things up. We use Spark underneath the hood, like all these things.Vibhu [00:25:36]: But when you say AWS, it's not actual AWS, it's your internal AWS.Eiso Kant [00:25:39]: It's our internal-- No, it's our internal like just running like our infrastructureVibhu [00:25:42]: Site web servicesEiso Kant [00:25:43]: Exactly. Our stuff running on like an AWS account or on like any hardware, right?Vibhu [00:25:47]: Yeah.Eiso Kant [00:25:48]: And so once we made that shift into I can stream data into training, all of a sudden you realize a lot of things unlock. Because now you don't have to wait for the whole data set to materialize.Immutable Data, Experiments as Code, and Scientific RigorEiso Kant [00:26:00]: You now all of a sudden when you're running data experiments about mixing data, it's a config. Because you've got these data sources that are coming in, and you just - we have this service called Blender that's in the report, where we then say, “Okay, for this run, I want 20% of this source, 10% of this source. I want this much, so many epochs of repetition. I want this to be, shuffled in a certain way,” and your training job can start while the rest of the data is even still materializing. also what it does is because all of this underneath-- So for us, we treated the data layer underneath as like an immutable data layer, and that was really important. Like experiments as code, immutable data layer means that you can always go back and understand literally down to the single token at which cursor it went in on which version of the code.Vibhu [00:26:47]: Yeah.Eiso Kant [00:26:48]: And it took us a I have to admit, like the first year of Poolside, we understood that engineering had to get great, But we didn't understand yet, that this is ultimately in support of like a good rigorous scientific progress. We were quite a - We were a very small number of people, so a lot of it was YOLO ideas and YOLO runs.Vibhu [00:27:08]: Yeah.Eiso Kant [00:27:09]: And we built great infra for the YOLO runs. But once we realized that we treated data as immutable and code as always versioned, and you could always track and trace every experiment end to end perfectly, you could repeat everything perfectly, right? You have perfect reproducibility. I can still reproduce runs from two years ago if I wanted to, right? It enables the scientific progress, like the scientific process, and I think that took us probably about a year and a half into the company to figure out. We also had some great hires, like our head of applied research, Nikolai, who joined us from Yandex, who'd been working on language models since like the early 2020s, I think brought that into the company of like, “Hey, we wanna have even more rigor.” And then once we kinda had the combination of like increasingly more capable platform that allowed people to do more, but had this immutability, we were able to start “Okay, every experiment is truly an ablation. We truly need to understand it.” And I think we became much more scientifically rigorous in the last couple of years, and the infra underneath enabled it. and then there's just fun stuff like, andVibhu [00:28:16]: Yeah, a lot of it's fun, like even just the, one, you share all the ablations, two, picking the data sets, right? There's like a random small paragraph in here where it's just like, “Oh yeah, training data, we have some, we have an auto mixer.” it trains eight small models, scales them up, picks the training data set. We don't even need to look at it. I'm like, “Wow, a lot of engineering rigor there.” And there's just, there's just a lot in here.Publishing Research and Giving BackEiso Kant [00:28:40]: Yeah, and it'- and look, and we wanna put out more. Like we, We treat writing papers as something that we haven't earned the right for yet for a long time. So you earn the right to spend time, publishing research once you're at the frontier, because until then, you're catching up, and every minute and hour in this industry matters. Like I obsess over it, not just the wall clock time from idea to result, but just general like time every day that we, waste is one that doesn't allow us to catch up. But in this case, we said, “Okay, we're gonna give ourselves.” I think we gave the team like three or four days while still doing their work, like give everything in there. And to your point earlier, if your stuff, it's easy to like put it out. And so there's so many more things that we wanna talk about over time, and we will definitely start doing. And as we earn more of the right, but also now have like added to our mission that we want more foundation model companies to exist, you'll see us like be way more proactive, and just trying to keep dropping some of those like things that we've learned along the way that can help others like speed up.Vibhu [00:29:40]: Which is the other cool side of this, right? It's, it's not like, back to your point, it's not just here's the benchmarks of our training. If you want to replicate, here's experiments of optimizers, data sets, post-training. you lay out a lot of it here alongside here's your system for how to do it? So it's, it's really like promotingEiso Kant [00:29:59]: No, thank youVibhu [00:29:59]: Other people can do the same.Eiso Kant [00:30:00]: And by the way, I also wanna make clear, right, we have been incredible-- Like we've taken a lot of advantage of the fact of all the open research that others have published, Right? And you mentioned, the Chinese labs, and we I think it's important that there's, from every country and every culture and background, including like Western companies like us, there's different models that come out that people can choose to trust. But I think we do have to give credit where credit's due, right? The incredible Chinese lab have done an amazing job at sharing their research, and we have definitely like been on the receiving end of taking advantage of that. So when you're on the receiving end of something coming to you, I think it's, you also have an obligation to give back.Swyx [00:30:39]: Do you have a favorite or underrated Chinese lab that you wanna shout out? Everyone shout outs DeepSeek.Chinese Labs, Zhipu, and PersistenceEiso Kant [00:30:44]: That's a good question.Swyx [00:30:45]: Moaan obviously for Therapsi. Yeah.Eiso Kant [00:30:48]: Yeah, look, I think, I think obviously everyone's been talking about Zhipu lately, with 5.2. I think what most people don't realize is when they started.Swyx [00:30:59]: Yeah.Eiso Kant [00:30:59]: Right? They started years before ChatGPT.Swyx [00:31:02]: They just rebranded. YeahEiso Kant [00:31:03]: And so, I've like, I remember how hard it was to work on these things Before the rest of the world got excited about it. And so I have an immense amount of respect for people, who were working on improving models when it wasn't the sexy thing to do, when believing in LLMs, was gonna get you ridiculed. I remember like back in 2016 when we were doing what we'd call, machine learning on code with some of these models. we would-- people would just laugh at us, like they'd be like, “This makes no sense. Like why are you wasting all these, like, millions of dollars on trying to figure this out?” And so I would say they're probably the one that, I think deserves a shout-out, not just because their latest model is very good, but because they fought to get here. And I think, I think every foundation model company it takes time to get here, right? It took us three years to get to the model that we're, that we're now gonna be releasing. and now the time in between the models is coming, is counted in weeks. It's no longer counted in months or years. But this stuff's hard. and if we can make it a little bit easier for the next person, like we should all do so. Because if we don't do so, we're, we've got a small window before models are really impacting recursive self-improvement to a level where catching up otherwise might become unfeasible. And we should try to, in that window, encourage as many labs or however we wanna call them, like to start. And so one of my currentEiso Kant [00:32:36]: Mission, but qualm is like I wanna encourage whoever is a researcher right now who thinks they can tackle this to go and leave and become my competitor.Eiso Kant [00:32:45]: Like start another foundation model company because I think we need it. I think otherwise we're not gonna be in the world where, I don't want to just be the fifth or the sixth company that wins. I wanna look at a world where there's lots of choice.Starting a Foundation Model CompanyVibhu [00:32:57]: What else do people not see in starting a foundation model? it's, there's a lot of compute, there's a lot of capital required, a lot of compute. You lay out model factory and how to do the training, but there's a lot there, right? That's,Eiso Kant [00:33:10]: Well, look, it's, I in turn-- this is an oversimplification, and I always asterisk it with that because it can land a little bit the wrong way in people's minds. But I think you can sum down, And I saw it, 95% of model building to just doing, you're just doing two things. You're improving data or you're improving compute efficiency. And I know that feels like an oversimplification for the incredible, like, Gifted and skilled work people do. But if you really look at it, like what are we doing? We are looking at data, we're generating new data, we're improving data. and the only way to do that is to look at the data, right? That's a big part of foundation model building. And on the other hand, we come up with these incredible breakthroughs in inference, in architecture, and new attention mechanisms. But what are they really doing? They're bringing compute efficiency. Now, we have definitely had some breakthroughs over the years that allow for more model capabilities. But at the limit, if you could train a large enough model, right, like, and you had infinite compute, we probably-- if you had infinite compute, you'd be at AGI probably already tomorrow.Eiso Kant [00:34:12]: Right? Like it's not. And so, and let me say that infinite compute with infinite ability of much faster networking because networking ends up being more of the bottleneck than compute. But, so I do think that's, those are the main things. And to just realize that this is engineering. I think it's become more obvious, but I think for quite a few years, people have held foundation model companies and researchers and others on this pedestal of like you're doing incredible magic or rocket science, or only like, Nobel laureate physicists can do this. And don't get me wrong, there are some really hard problems that need to be solved, but a lot of the work that all of us are doing on a day Is not sitting down trying to solve a math theorem. A lot of the work that we're doing is just really doing the basics right, writing good code, looking at data, improving it, running experiments, looking at plots, trying to see like, hey, trying to shape our intuitions. And a lot more people could be highly capable researchers. and I think that's, it feels far for people to do so. But I've seen in our own company, we've seen engineers become researchers because the model factory allowed them to be, have a much lower hurdle of running experiments and trying things. And one of the guys on our team who started as an engineer building our agents is a legit reinforcement learning researcher now, making real progress. and that happened in the span of like six months. that would've not been what I think most people assumed was possible, a couple of years ago.Swyx [00:35:46]: Yeah. I think one of the interesting moments is when you can self-host, like, if in a programming language, like if you can compile the language in the language, the equivalent is can you use your own tools, right? You have the pool CLI, you have your own models. presumably you're not only using your own models. There's no way. But like, what's that percentage over time?Laguna S, Persistence, and Behavioral GainsEiso Kant [00:36:10]: This is the first model that we're releasing that is starting to meaningfully contribute to our own work. It's not a it's not state-art model yet. Fable and other, they're, they're very capable models, but Laguna S Is really interesting. I'm gonna pull up the quote. Peng Ming, one of our heads of applied research, said something, last week as the model came out about 10 days ago, much better than we had hoped for or expected. And he said, I have the feeling that a lot of the gains in Laguna S come not from more intelligence, but more from different behavior, more verification, less taking things for granted, not declaring victory early, and being way more persistent. And to be honest, those are more predictive than raw intelligence for success in human also to some degree. And this was, he wrote me this on 5th of July on a Sunday, and it's been burned in my brain ever since because the Laguna S model, as you'll see it and why it does so well on benchmarks and why it does so well in using it on a day basis, is that it's just incredibly persistent. It reasons a lot. I do call that out. We have work to do on making it more efficient. We have to work to do on offering different reasoning modes. But this is the model that has been able to do things that I never thought it could do. A hundred eighteen billion 8B active model, which is not that large. It fits on a DGX Spark and still runs at, thirty, forty tokens a second on a Spark, is able to solve Erdős 397 independently. It's able to do complex programming tasks. It's able to. I asked it this morning to make me a Fi scanner without using any external libraries on my Mac, and it's, like, figuring out, like, the core WLAN API by really persistently trying to understand it without access to the internet. And more, I love vibe checking. I've probably spent eight to ten hours a day with this model for the last ten days.Eiso Kant [00:38:05]: I'm not exaggerating. I was on my eleven-hour flight yesterday. I spent ten hours reading trajectories and traces and, like, of the model.Eiso Kant [00:38:12]: And what I take away from it is exactly what Peng Ming said. We are gonna be able to squeeze so much more out of smaller models than I think we had imagined in the industry because, yes, there's intelligence and larger models are more intelligent. Like, no doubt about it. We should continue to scale up. but the behaviors of being really persistent, of being able to backtrack when you're wrong, of, like, understanding how to interact with your environment show us that we can get a lot more out of it. And this, for me, has created a bit of a Question in my mind the last couple of days. If you think about where we're using models today, right? We are using models, say, for knowledge work. Represents twenty-five percent of the global economy, twenty-five trillion dollars of work.Eiso Kant [00:39:00]: As we scale up models and they become more intelligent, we are excited about using them more and more for pushing the frontier of science.Small Models, Knowledge Work, and CommoditizationEiso Kant [00:39:08]: And if you look at the frontier of science, like true breakthroughs in science, they have been linked, they are linked to more intelligence in many places. Einstein figuring out general relativity is able to bring ideas together that other people would have not brought together. And I think one of the many dimensions of intelligence is the ability to do that, and it's something we clearly see that as models get larger and more capable, they're able to pull more ideas and threads together that a smaller model wouldn't be able to.Eiso Kant [00:39:36]: And we're starting to see examples of that in medicine and, like, in bio and other things. But if you think about the majority of knowledge work that we do, and it includes building software. I'm a software developer at heart first and foremost probably, although I probably can't say it that much anymore as I don't write production code in years, is that what makes us good is our persistence. It's our ability to encounter a problem and backtrack and say, “I need to go figure out this bug. I need to go research this. I need to go look at the documentation. I need to, like, try different, five different ways to see, like, if I can solve it.” But it is not necessarily bringing three ideas together from radically different fields. And so if we are now seeing, and I think Laguna S is an example, that we are able to make a relatively small model much more capable than I had definitely predicted or any previous, like, benchmarks had shown for any model remotely this size or even larger, At least on coding tasks, that it's because of the behaviors. And so now the question I have, and I don't have an answer, it is I know at the limit, so infinite model size, right, extremely large model, and the cost of that model is gonna be very expensive to run. We know this, right? So larger model ROI.Eiso Kant [00:40:52]: So I know that at the very limit, I'm not gonna use the world's largest model one day, quadrillion parameter, whatever crazy, like, scale we scale up, to do a basic coding task. Already today, I'm starting to size down for certain tasks.Eiso Kant [00:41:07]: So it means that there is an optimal. It means there's some curve that goes as we go up to model size for knowledge work, at some point we're at the peak, and after that, the return on investment of using a bigger model, just doesn't make sense.Eiso Kant [00:41:22]: Now, I think the question is, before I would have thought that peak was extremely very far away.Eiso Kant [00:41:30]: This model for me is the first sign that Maybe that peak is At a trillion, five trillion, ten trillion. Maybe we can just squeeze way more out of these models. I'm no longer thinking that we need two or three orders of magnitude on the largest models to be able to, solve knowledge work, the accounting, the legal, the code that we write. And so if that holds true, It is an argument for the commoditization of models. It's an argument that open source can win and, like, succeed in this world. And now it's of course a self-serving argument and it's a hopeful argument, but theoretically at the limit it works. We just have to go discover in the next couple of years of how much more we can squeeze out. Now, I do want to put a big asterisk. This does not mean I'm against scaling models. I think we ultimately only succeed if we scale our models as large as our competition. I do not like. I think we should not put our head in the sand and say we're gonna be king of open source small models. I think that's, It's a out. It's trying to be king of your own kingdom, but not realizing what the rest of the world's doing. All of us rather use a smarter, faster, more model. It's a sign of hope. And so I don't wanna overly state this is a good model. We have a long way to go to get to the state-art. But what hopefully people take away when they use this model is that the behaviors inside of it are what push it to be far more capable, less than necessarily the number of parameters.Pre-Training, Mid-Training, and RL Moving EarlierVibhu [00:43:03]: Is that mostly post-training? LikeEiso Kant [00:43:05]: YesVibhu [00:43:05]: Right.Eiso Kant [00:43:06]: It's entirely post-training.Vibhu [00:43:08]: Are we done improving anything on training? Is, like, training done?Eiso Kant [00:43:12]: No.Vibhu [00:43:12]: Okay.Eiso Kant [00:43:13]: SoVibhu [00:43:13]: I just wanted to cover training, and then we go post-trainingEiso Kant [00:43:15]: Training is not done. I mean, look, there's a part of training of just dealing with skill, right? Every new order of magnitude of model skill, you are going to get new things you gotta solve for. That'- but those are ultimately, engineering challenges.Eiso Kant [00:43:31]: I have a, I would say, a not commonly held opinion that reinforcement learning Will move earlier and earlier into training.Vibhu [00:43:42]: Yeah, training.Eiso Kant [00:43:44]: Not even training. Like training today, right, is, like if you look at - So we've been working on this for years already. and I think the best-- I think the first time we saw it out in public was the DeepSeek Zero paper. this is a year and a half ago, I think, if I recall correctly. where, you can Very early on in a model as it starts capable of being able to use language, et cetera, induce reasoning. and so the question that I have is like, we have this- we have the dataset that's the web. and the web, I think we could arguably say probably has The totality of humanity's knowledge somewhere encoded in different places. It's a huge variance degree of quality, from garbage data, and like once you look at training data, you really get humbled of like what the web is, to like, the most greatest scientific papers and best blog posts and like, best transcripts and whatnot.Eiso Kant [00:44:39]: And so now What we are trying to figure out, and have been doing a lot of work on, and it's a place where maybe not as open as we're on other things, but we will become more over time. we've been spending a couple of years really doing research on how can we turn the web into not just next token prediction, but into a way to teach the model to think earlier in its training. and I think there's a huge amount of gold to be found there. I think we are right now in, we've got some drugs in the industry. One of the drugs is distillation. Another drug is, more environments. Like, and they're great, and they make us feel good, and they make the models better, and like we're all addicted to them, and we'll use them, right? in various different ways. and but ultimately, I think we are still barely squeezing out of the web what we should be getting out of the web.Eiso Kant [00:45:33]: I think just next token prediction during training is not enough.Eiso Kant [00:45:36]: AndVibhu [00:45:38]: YeahEiso Kant [00:45:38]: I think we'll see some very interesting things still happen. and that RL in post-training to induce behaviors, to improve things, like I think - the whole world knows how to do this now. I think we're, we're scaling it up. Everyone is. But I wonder if we need to go as far as we're going today with environments. I'm not sure yetVibhu [00:46:01]: You mean we're going too far?Eiso Kant [00:46:02]: I'm, I'm not sure if the path to AGI is justVibhu [00:46:06]: Is more environmentEiso Kant [00:46:07]: More environments.Vibhu [00:46:08]: It seems like a never-ending, “Okay, I want instruction manual for this table, right? Am I gonna environment out building furniture? Or are we just gonna tail end like we need some general solution?”Eiso Kant [00:46:19]: I think there is, I think there's an ability to generalize more from the web. but I also am very encouraged, like when I look at Laguna S and, which is post-training is, well, is the big impact there. and I see like, oh, wait a second, just by making some of these behaviors much better, we're able to get so much more out of it. It just changes a little bit the way you think about intelligence.Vibhu [00:46:40]: Yeah. The analogy people draw often is the RL phase is where you don't learn as much new knowledge. You shiftEiso Kant [00:46:46]: Yeah.Vibhu [00:46:46]: Yeah. So, you shift distribution, and you can have it reason towards what you want. on your point about training, a lot of training is still just continue training in a domain, say medicine, then you do RL. So still justEiso Kant [00:47:00]: It's just better data, right? Like, I mean, training, ooh, I like how we invented this word. Like it's effectively just like,Vibhu [00:47:06]: Second phaseEiso Kant [00:47:07]: It's the second phase of training With like a really dumb way to do a curriculum. But like ultimately, what you'd want is a curriculum from token zero to token 30 whatever or 40 trillion tokens that really truly is the optimal curriculum for the model to learn. But training is essentially a stage curriculum on the web because we do not have to compute, And, effectively to try to ablate the perfect curriculum, right? And so I'm pretty sure that you'll start to see people talking soon about some other term, and there's two or - ‘cause now we do this, right? We talk stage two and stage three and stage four training and like. But ultimately, all we're doing is we're trying to assign a curriculum to the web data that we have to allow the model to learn better. I think at some point, as things get compute, as models get cheaper to run, as the next generations of compute, this will become more of a continuous spectrum. I also think the reason, by the way, you have training and like stage two and stage three is organizational, Right? It'- this is, I think, a thing where-- that we really try to avoid with the model factory is like Training exists because there's a training team now, right? There's people, or like people in training decide to focus on like a training effort. but what you really want is engineering and scale of experiments that allows for a much more continuous spectrum that you don't, you have infinite stages. Now, we're not there. Compute's not there. Organization design is not there for it yet. but I think we'll get there. we'll look back on a couple of years and be like, “Oh my God, it was so cute that we did our training data like this in such a like naïve way. Like we barely ordered it. We didn't really do a good job at likeCurriculum, Auto Research, and New ObjectivesVibhu [00:48:48]: The building that curriculum will get you that in the industry.Eiso Kant [00:48:51]: And I'll confirm that, when I talk to some researchers that this is a lot of the focus now is like how does training change and what is the next objective other than, next token prediction. I assume you don't have the answers, but you have some ideas.Vibhu [00:49:02]: We have some ideas. We're not ready to talk about it yet.Eiso Kant [00:49:05]: Yeah.Vibhu [00:49:05]: We've been working on them for years, and I think that's the one thing that's also like you asked earlier about, like what's not obvious about building a foundation model company is that you are constantly balancing the table stakes work, the recipe worksEiso Kant [00:49:19]: Yeah.Vibhu [00:49:19]: Versus like your, my crazyEiso Kant [00:49:22]: Pure researchVibhu [00:49:22]: Breakthrough.Eiso Kant [00:49:22]: Yeah.Vibhu [00:49:22]: Pure research and finding that balance and adjusting the percentage to it based on where you are in the race is really important.Eiso Kant [00:49:31]: I mean, so like, this is a nice way. I was gonna bring up auto research at some pointVibhu [00:49:35]: YesEiso Kant [00:49:35]: As another Andrej invention, or coinage, which is like, I honestly, like how many objective functions can there be, right? Like just try 1,000 of them, set it running, whatever.Vibhu [00:49:47]: Man, it's alsoEiso Kant [00:49:48]: Like what you're looking for. You're looking for loss curves like that, likeVibhu [00:49:51]: It's also a thing people take bets on, right? When you say more Neo labs, you're doing a version of we'll do foundation models, scale them up, next token predictors. A lot of other Neo labs that we see want to take a completely different approach, right? At some level, you're right. It's all, compute efficiency, and that's the net objective. But some are okay, different architecture, like vastly different amounts of compute spend. So some are different. They're not justEiso Kant [00:50:19]: YeahVibhu [00:50:19]: They're like, 99% not balancing, here's the vanilla and scale up. They're 99% on, here's novel research that'll change everything.Eiso Kant [00:50:27]: And I think, Luke, I think you. It depends when you started as well, right?Pure Research vs. Table StakesVibhu [00:50:30]: Yeah.Eiso Kant [00:50:30]: When we started, like the novel thing we did was reinforcement learning on code. No long- that's no longer novel by far, but we were like, - that's where we obsessed over when no one believed in RL. So you have to when you start the company, you have to have your own idea. You have to have something that's different that allows you to speed up, right? For us, it was RL to LLMs that later became common, like, Knowledge. But in the beginning, it wasn'tVibhu [00:50:53]: It's cool. this was like your original 2023 blogEiso Kant [00:50:57]: YeahVibhu [00:50:57]: Of purpose.Eiso Kant [00:50:58]: Yeah.Vibhu [00:50:59]: And like you do lay it all out here.Eiso Kant [00:51:01]: We laidVibhu [00:51:01]: The blog is pretty underrated, right? The whole RL on code was very early on.Eiso Kant [00:51:06]: Very early. And even we had to argue with people, like we say here things like to push beyond current capability, to train your own foundation model. We had to argue with people that it mattered that you had your own like, base model. you can fine-tune your way to success, right? major capabilities emerge from training a base model made accurate and useful during fine-tuning.Vibhu [00:51:23]: Which like, for perspective at the time, we knew closed models, OpenAI, Anthropic were huge. The open models we had were like Mistral 7B, a 30B, a 70B.Eiso Kant [00:51:35]: When weVibhu [00:51:35]: YeahEiso Kant [00:51:36]: The date on this thing is wrong. When we published this, it was April 2023. I think this was justVibhu [00:51:42]: YeahEiso Kant [00:51:42]: Happened on a migration, probably found it on archive.org.Vibhu [00:51:45]: Mistral.Eiso Kant [00:51:46]: Mistral had started, we started on the same month, right?Vibhu [00:51:49]: Yeah.Eiso Kant [00:51:49]: So this wasn't even, there was only, I think, Llama out at the timeVibhu [00:51:52]: SnellEiso Kant [00:51:52]: And that's it, right? And so, but I agree. I think we wan
Before science was invented, things were confusing. And, to be fair, for several decades into the birth of science, it was still confusing. Don't believe us? Check out Scientific American magazines from the late 1800s — especially this one article about scientists tasting Mezcal and Tequila for the first time. Agave Road Trip is a critically acclaimed, award-winning podcast that helps gringx bartenders better understand agave, agave spirits, and rural Mexico. This episode is hosted by Lou Bank with special guest Linda Sullivan of seynasecreto. Episode Notes Read Scientific American, No. 1096,January 2, 1897, for the source of the epic quote in this episode: Not only does the maguey plant give the sap for the drink pulque, but its roots , distilled, yield much more powerful intoxicants, mescal and tequila. And the heart of another variety of the same plant is cut out, roasted and put through a process of distillation to make still another drink. Mescal is described as tasting like a mixture of gasoline, gin and electricity, Tequila is even worse, and is said to incite murder, riot and revolution. A very small quantity of tequila is a drink. Experts advise it to be taken with a grain of salt, literally — the salt to be placed on the tongue first. Several Michigan editors who lately visited the sister republic wandered into a "cantina" in the city of Puebla one evening and called for a drink apiece of the tequila. Each took a taste, felt his hair stand erect, his toes curl up and his skin wrinkle all over his body and was wise enough to leave the glass unfinished. Not one of them wished that time that Providence had endowed him with greater capacity. Shout outs this episode to Jeppson's Malort, Vienna Beef hot dogs, Cocktail MD Ryan Aycock, and Metsal inventor, Enrique Martinez!
Aromatherapy is much more than "essential oils," it is an experience. One that connects you with nature no matter where you are; one that inherently connects mind and body. Are you ready to tap into your chemical sense of smell and the chemistry of plants? Tune in and join Amy Anthony to question and explore the world of Aromatic plants and their preciously concentrated essential oils.Website: https://nycaromatica.com/
In this episode of the Crazy Wisdom Podcast, host Stewart Alsop speaks with Aaron Neyer, founder of Parachute and community organizer in Boulder, about knowledge management, extended minds, and the intersection of AI with human consciousness. They explore how Parachute functions as a digital brain tool for organizing thoughts and information across fragmented systems, discuss the dangers of AI psychosis and over-reliance on technology, and debate open source AI development versus controlled releases by companies like Anthropic. The conversation weaves through topics including the limitations of metrics-driven business thinking, consciousness and relevance realization, the value of technological sabbaths, and Aaron's hope for locally-run open source models that protect personal data while still accessing more powerful gated models when needed. You can find Aaron's writing at unforced.org and unforced.substack.com, and learn more about Parachute at parachute.computer and parachute.computer/blog.Timestamps00:00 Stewart welcomes Aaron Neyer, founder of Parachute and Boulder community organizer, discussing the origin of Parachute's name from Frank Zappa's quote about open minds.05:00 Aaron explains Parachute as an extended mind tool for organizing notes, contacts and information across multiple platforms, emphasizing the distinction between primary mind and extended mind as interconnected systems.10:00 Discussion shifts to metrics-driven business culture and the limitations of pure rationality, exploring how Google's data-driven approach misses subjective experience and the whole picture of relationships.15:00 Aaron discusses AI's ability to help identify relevant variables across different domains and the dangers of AI psychosis, comparing it to cult dynamics and belief systems.20:00 The conversation covers AI sabbaths and nineties retreats as intentional breaks from technology, plus Aaron's experiences with electrical engineering and using AI to design circuits with Arduinos.25:00 Exploring forbidden knowledge and open source AI, Aaron discusses Anthropic's guardrails around powerful models while arguing for distributed access to prevent concentration of power.30:00 Deep dive into open source AI strategy, with Aaron highlighting NVIDIA's approach and the potential for running capable models locally while reserving ultra-intelligent models for complex research tasks.35:00 Aaron shares his vision for local Sonnet-class models handling personal data while accessing Fable-class models for deep research, and directs listeners to unforced.org and parachute.computer for his writing.Key Insights1. The philosophy behind Parachute stems from Frank Zappa's quote that the mind is like a parachute and doesn't work if it isn't open. Aaron Neyer explains that having an open mind is valuable, but it must be balanced with deep roots to avoid becoming untethered. He has experienced periods in his life where excessive openness led him to feel disconnected, teaching him that creativity and expansion need to be grounded in something substantial. This same principle applies to how we organize information digitally, where openness and interoperability allow our extended minds to become more connected and coherent, which in turn helps our primary minds think more clearly.2. Parachute is designed as an extended mind tool that addresses the fragmentation problem in how we currently manage information. Most people use multiple disconnected tools like Obsidian, Notion, Apple Notes, Google Keep, and various CRMs to organize their thoughts, notes, and relationships. These systems don't communicate well with each other, creating inefficiency and confusion. Parachute aims to create a simple, intuitive system where all this information can be organized in one place with true interoperability, allowing users to own their data and have it speak effectively with other tools, ultimately making our entire extended mind more functional.3. Understanding ourselves as unified body mind organisms rather than fragmented parts is essential for effectiveness. Living systems theory shows that any living system is three things: a membrane bound dissipative structure, a self regulating autopoietic network, and a cognitive process actively knowing the world. Western civilization since Descartes and Galileo has created artificial separation between body and mind, and between subjective and objective experience, which limits our effectiveness. The same fragmentation affects our digital technology, and recognizing both our biological and digital systems as coherent wholes rather than disconnected parts makes us vastly more capable.4. The relationship between data driven approaches and holistic thinking reveals important limitations in modern business and science. While working at Google, Aaron observed how data driven decision making can be powerful, but over reliance on metrics like ROI creates blindness to crucial unmeasurable factors like goodwill and relationship quality. This reflects a broader Western tendency to exclude subjective experience because it's difficult for objective science to measure. However, emotions, relationships, and other subjective elements are essential parts of reality, and focusing only on quantifiable metrics means missing the whole picture and ultimately becoming less effective despite appearing more rational.5. AI accelerates the ability to work with technical complexity by helping with relevance realization across domains where we lack expertise. In any specialized field, experts develop intuitive senses for which variables matter and can quickly identify problems, whether in computer troubleshooting, music, or cooking. AI's ability to generalize allows it to point people toward relevant solutions in areas where they haven't developed that intuitive expertise, effectively democratizing technical capability. This means people can direct their creativity more effectively across more domains, though it also raises concerns about giving powerful capabilities to those who may lack the wisdom to use them responsibly.6. The question of open source AI versus gated access involves complex tradeoffs between democratizing power and preventing harm. Aaron respects Anthropic's approach of creating guardrails around powerful models like Mythos, which would likely have caused significant system hacks if released without restrictions. However, this creates concerning power dynamics where only wealthy companies, governments, and their allies have access to the most powerful tools. NVIDIA offers hope through their truly open source approach including full training pipelines, and there may be a viable path where open source models at the Sonnet capability level handle most tasks locally while more powerful Fable class models remain gated for the most demanding work.7. Creating intentional breaks from AI and technology is essential for maintaining clear independent thinking. Aaron practices an AI Sabbath at least one day per week when he doesn't interact with AI, and he finds these are the days when he does his best thinking and journaling. Without these breaks, he finds himself constantly jumping between journaling and prompting AI rather than giving himself space for deep reflection. This pattern mirrors broader concerns about AI consistency creating cult like dynamics similar to organized religion, where constant immersion in a belief system or technology can lead to losing the ability to think independently, making periodic disconnection crucial for maintaining cognitive autonomy and clarity.
Okay, for real, this episode is only vaguely about Raicilla. It's a bit more about profound justifications — that's a reference to a quote from my guest co-host, Shawn Miller of PKGD Group. He brings that quote up as we're talking about the white paper PKGD released a few weeks ago. That white paper is what this episode is really about. But who would listen to an episode titled, “The one about Shawn's white paper”? Agave Road Trip is a critically acclaimed, award-winning podcast that helps gringx bartenders better understand agave, agave spirits, and rural Mexico. This episode is hosted by Lou Bank with special guest Shawn Miller of PKGD Group. Episode Notes Read “Built Right and Brought Right,” the white paper published by PKGD Group! Shout outs this episode to Binny's, Mezcal Ultramundo, FOMO special editions, G4 Tequila, Marriott Hotels, Jose Cuervo, Diageo, Dan Kennedy, and tahonas. Ad Links If you want to taste 24,000 acres of wild desert, get yourself a bottle of Mezcal Ultramundo! Since 2005, Chicotona has been producing Mezcal with a focus on protecting wild agave and preserving the land for future generations! Head out on an Agave Road Trip with Finca 18! Greg Rutkowski will take you on his Agave Road Trip Route #2 - Raicilla de la Costa! Price includes a bottle of Paulo Rodriguez's fabled, limited Tumbado batch! Order beautiful spirits to be delivered to anywhere in Mexico — beautiful or otherwise — through Agave Spirits Presents!
AI Chat with Maxime Lamothe-Brassard and Chris Luft.A new segment on the podcast: AI news in cybersecurity that is less than 24 hours old, discussed while it is still hot. Joining Chris for these conversations is LimaCharlie founder and CEO Maxime Lamothe-Brassard.In this episode:• Nipun Gupta (founder of Optimus Labs) reports that xAI's Grok Build CLI packaged and uploaded an entire local Git repository — commit history, branches and .env files with API keys — to a Google Cloud bucket; wire-level analysis via mitmproxy, a quiet server-side fix, and why you should rotate keys if you used the tool.• Fortinet's take (via Mexico Business News) on AI accelerating vulnerability discovery and exploitation: 24–48 hours from disclosure to active exploitation vs. 16 days to patch — and whether "virtual patching" is a real mitigation or a feat of marketing.• The AI distillation debate: after years of arguing fair use for scraping the internet, frontier labs now object to competitors training on their model outputs — Business Insider's look at the irony, shared by Pascal Hetzscholdt (Wiley).• Neon Cyber's survey on shadow AI rising with seniority: 14% of individual contributors use unapproved AI tools vs. 63.7% of managers and 70% of VPs and above — and why enforcement, not awareness, is the real challenge.Stories covered:• / guptanipun_my-spare-laptop-ran-completely-... • https://mexicobusiness.news/cybersecu...• / pascal-hetzscholdt_quote-heres-some-delici... • https://neoncyber.com/blog/shadow-ai-...Chapters:0:00 Intro — welcome to AI Chat0:45 Grok Build CLI uploading entire repos (Nipun Gupta / Optimus Labs)4:57 AI is outpacing patch management — is virtual patching the answer?12:32 The AI distillation debate: scraping irony at the frontier labs16:29 Shadow AI use rises with seniority (Neon Cyber)22:51 Wrap-upThe Cybersecurity Defenders Podcast — a podcast about cybersecurity and the people that keep the internet safe. New episodes drop weekly.Subscribe wherever you listen:• Spotify: https://open.spotify.com/show/6ep00ze...• Apple Podcasts: https://podcasts.apple.com/us/podcast...• YouTube: / @limacharlieio
What makes you different? Do you have a concise and compelling answer to that question? When it comes to positioning, less is more. But too often, companies get wordy and convoluted, losing the plot in a sea of text. Emma and Adam Ketterer argue that true clarity comes from radical constraints. In their new book, Eight Words or Fewer, this marketing duo shares how to distill your difference into a single line that sticks: a micropitch. What You'll Learn in This Episode Why B2B companies naturally fall prey to the instinct that more must be better when defining their value proposition The strategic breakdown that occurs during the last mile of the positioning process and how to build for brevity from day one Tactical ways to overcome the anxiety of leaving things out by showing your stakeholder work and mapping customer journeys How to use a homeopathic distillation framework to create a micropitch that infuses everything from sales decks to homepages The creative and editorial realities of writing a pocket guide on concision as a married couple Episode Chapters (00:00) Intro (01:34) Why We Struggle with Brevity (03:27) Losing the Plot in the Positioning Process (06:44) Overcoming the Anxiety of Distillation (09:59) The Subline and Structure of Eight Words or Fewer (12:08) How to Kill Your Darlings (15:27) The Micropitch as a Brand Strategy Anchor (17:26) Collaboration and Constraint as a Married Couple (22:57) Finding Your Point of View and Conviction (25:38) Brands That Make Us Smile (27:50) Where to Connect with Emma and Adam About Emma and Adam Ketterer Emma Ketterer and Adam Ketterer are seasoned marketing professionals specializing in positioning, messaging, and content strategy for B2B technology companies. Emma brings extensive experience from both the agency and corporate sides, currently serving as the Global Lead Editor at Celonis. Adam is a co-founder of Beneath, a specialized B2B marketing agency. As a married couple, they blended their strategic and creative insights to co-author Eight Words or Fewer: Distil your difference into the perfect micropitch, a pocket guide exploring the critical intersection where strategic brand positioning meets ultra-concise writing. What Brands Have Made Emma and Adam Smile Recently? Emma found recent inspiration in the B2B tech space through GitHub's smart, witty, and exceptionally well-produced explainer videos. Adam pointed to the timeless, iconic print campaigns of The Economist, noting how the brand's playful, minimalist approach continues to inspire creative professionals to riff on its memorable format across digital and physical signage. Resources & Links Connect with Emma and Adam on LinkedIn. Check out their new book, Eight Words or Fewer. Watch or listen on Apple Podcasts, Spotify, YouTube, Amazon/Audible, TuneIn, and iHeart. Rate and review on Apple Podcasts and Spotify to help others find the show. Share this episode — email a friend or colleague this episode. Sign up for my free Story Strategies newsletter for branding and storytelling tips. On Brand is a part of the Marketing Podcast Network. Until next week, I'll see you on the Internet! Learn more about your ad choices. Visit megaphone.fm/adchoices
Pastor Alan R. Knapp discusses the topic of "A Move Toward Distillation" in his series entitled "Hebrews 2020: We See Jesus" This is Increment 438 and it focuses on the following verses: Acts 13:15-43; Hebrews in toto
In this episode, Ray Cochrane breaks down AI distillation, the teacher-student technique frontier labs now lean on to train smaller, cheaper models. He also covers GPT-5.6’s government-vetted rollout, Claude Sonnet 5 landing on AWS, Maryland’s two-year data center pause, and Microsoft’s climbing carbon numbers. Finally, he wraps with Apple’s $30 billion Broadcom deal, Meta’s tamper-proof recording light, Michigan’s parasite outbreak, and a simulation that erased a super El Niño. – Want to start a podcast? Its easy to get started! Sign-up at Blubrry – Thinking of buying a Starlink? Use my link to support the show. Subscribe to the Newsletter. Email Ray if you want to get in touch! Like and Follow Geek News Central’s Facebook Page. Support my Show Sponsor: Best Godaddy Promo Codes Get 1Password Full Summary Cochrane opens with a quick personal update. Longer days have him outdoors, including a float trip on the Sandy River at Dabney State Park, where he found clearer water, clay-like sand, and easy footing. Next week brings both a move and a trip home, so he is stocking up on Trader Joe’s “Power Berries” and IKEA bags at his mom’s request. Then he turns to the lead story. AI Distillation Explained: How Frontier Models Teach Each Other Cochrane’s featured story comes from Hugging Face engineer Sergio Paniego. Distillation is teacher-student training for AI: a capable model generates the training signal, and a smaller student learns to match it. The classic off-policy version compresses giant models into cheap students, either through soft labels or piles of worked answers. Google’s Gemma models and DeepSeek’s R1-Distill line were built exactly this way. However, the industry is now converging on multi-teacher on-policy distillation, or MOPD. Labs build reinforcement-learning specialists for math, coding, and agentic work, then have them grade a single student, word by word, as the student generates its own answers. DeepSeek-V4, MiMo-V2-Flash, and NVIDIA’s Nemotron 3 Ultra all run versions of the recipe, and the Qwen3 team reported better results at roughly a tenth of the GPU hours of raw reinforcement learning. Finally, self-distillation lets models like Cursor’s Composer 2.5 learn from better-prompted versions of themselves. Sponsor: GoDaddy Economy hosting $6.99/month, WordPress hosting $12.99/month, domains $11.99. Website builder trial available. Use codes at geeknewscentral.com/godaddy to support the show. GPT-5.6 Arrives With a Government-Vetted Rollout OpenAI shipped GPT-5.6 as a three-tier family: Sol, Terra, and Luna. Sol costs five dollars in and thirty dollars out per million tokens, half of Claude Fable 5’s rate. The benchmarks split: Sol Ultra wins Terminal-Bench at 91.9 percent, while Claude Fable 5 still leads SWE-Bench Pro. Notably, the API launched in limited preview to roughly 20 partners vetted by the U.S. government, though the model went live in Microsoft 365 Copilot on day one. Claude Sonnet 5 Lands on AWS, Plus Quick AWS Wins Claude Sonnet 5 arrived on AWS through Bedrock, pitched as top-tier intelligence at Sonnet pricing. Additionally, Amazon WorkSpaces for AI agents reached general availability, enabling agents to drive full desktop applications securely. OpenSearch gained a log-analytics engine claiming four times the price-performance, and SageMaker now scales inference about twice as fast. Cochrane also flags that Kendra and Q Business move to maintenance mode at the end of July. Anthropic Wants You to Reflect on Your Claude Habits Anthropic launched Reflect, a beta feature that analyzes your past Claude conversations and visualizes how you actually use the assistant. It requires Memory, excludes incognito and health-related chats, and keeps its insights inside the tool. Cochrane loves the idea. He reviews his own transcripts to extract prompt patterns and turn them into reusable skills, and he suggests listeners simply ask their AI to do the same. AlphaEvolve Goes GA on Google Cloud Google made AlphaEvolve generally available to Google Cloud customers on the Gemini Enterprise Agent Platform. The agent acts as an evolutionary collaborator: provide a baseline algorithm and your goals, and it searches for better, human-readable code. BASF, JetBrains, and Kinaxis are the named early adopters. Meanwhile, Cochrane renews his standing wish that DeepMind release AlphaGo as a playable teacher. Google Adds “How This Ad Was Made” AI Labels Google is adding a “How this ad was made” section to My Ad Center across Search, YouTube, and Discover. Ads built with Google’s own AI tools automatically get the disclosure, backed by invisible watermarks. However, ads made with outside tools rely on advertiser self-declaration. Cochrane points out the limits of voluntary disclosure in an AI-flooded content economy. Microsoft’s Carbon Emissions Climb 25 Percent Microsoft’s new sustainability report shows emissions up 25% in 2025, driven by a data center construction spree. The gross figure is 34 million metric tons before offsets, while other coverage puts the net figure at around 20 million. Water consumption also jumped thirty-four percent, even as Microsoft claims its first water-positive year. Cochrane argues regulation needs to catch up, since Google and Amazon report similar increases. Prince George’s County Pauses Data Centers for Two Years Prince George’s County adopted a two-year moratorium on new data center development, the longest pause in Maryland so far. The resolution blocks new applications, including hyperscale projects, until the council passes real regulations. Water and energy impacts remain open questions the county intends to study. Cochrane gives kudos to residents for making their voices heard. Apple and Broadcom Ink a $30 Billion U.S. Chip Deal Apple is expanding its partnership with Broadcom with a multiyear agreement expected to exceed $30 billion. The deal covers custom silicon and wireless components, with more than fifteen billion chips to be made on American soil. Broadcom’s Fort Collins, Colorado plant anchors the work with a $1.5 billion equipment expansion. Tim Cook framed the deal as accelerating Apple’s commitment to American manufacturing. MSI and Intel Ship the First Arc G3 Extreme Handheld Intel detailed how it co-engineered the MSI Claw 8 EX AI+, the first handheld on the Arc G3 Extreme processor. Highlights include a heat-spreading board layout and game-tuning loops that Intel says run Cyberpunk 2077 up to thirty-seven percent faster. The device is on sale now in void purple for around $1,500. At that price, Cochrane jokes he would rather buy a computer. Meta’s Glasses Get a Tamper-Proof Recording Light Meta answered the most common privacy questions about its AI glasses. Photos stay private on the device until the wearer imports or shares them, and a white capture LED blinks during any recording with no off switch. Moreover, newer glasses disable the camera if the LED is blocked, tampered with, or destroyed. Cochrane reminds listeners these claims are Meta grading its own homework, but the blink signal is worth recognizing in public. Michigan’s Parasite Outbreak Tops 1,200 Cases Michigan’s cyclosporiasis outbreak reached 1,251 cases since June 22, with roughly forty hospitalizations along the way. Northwest Ohio adds more than five hundred cases. The parasite typically spreads through contaminated fresh produce, and investigators still have not found the source. Cochrane’s advice: wash your produce, and get tested if your symptoms fit. AI Finds the San Andreas Fault’s Silent Slips Researchers paired AI with borehole strainmeters to detect dozens of hidden slow-slip events beneath the San Andreas Fault’s Parkfield section. Each silent slip releases stress within hours and is reliably followed by low-frequency earthquakes. Together, the findings support a continuous spectrum from silent creep to destructive quakes. The study appears in Nature Communications, and Cochrane hopes it will lead to better earthquake prediction. Cloud Brightening Erased a Super El Niño, in a Simulation Finally, a Science Advances study simulated marine cloud brightening in response to the 1997 and 2015 super El Niño events. Seeding clouds over the eastern Pacific erased the events entirely inside the model. Real deployment would take roughly 2,400 ships spraying continuously, and the simulations showed side effects like extra warming over Europe and Asia. Cochrane finds the weather-machine concept fascinating, yet he questions the consequences of altering cycles the planet runs for a reason. The post AI Distillation: How Frontier Models Teach Each Other #1870 appeared first on Geek News Central.
GPT 5.6 darf released werden. Fiji Simo tritt aus gesundheitlichen Gründen zurück, gleichzeitig wird ihre Strategie revidiert. Elon Musk launcht Grok 4.5 als Efficient-Frontier-Modell. Meta launcht Muse Image und Muse Video, kündigt Iris-KI-Chips ab September an und lässt die Financial Times testen, wie eine Brille alles aufnimmt, was der Träger sieht und hört. OpenAIs Deployment Company kauft Northslope, das erste Signal für eine Konsolidierung der KI-Consulting-Branche. China warnt seine KI-Firmen vor Distillation durch US-Firmen, US-Lawmakers prüfen die wachsende Nutzung chinesischer Modelle bei US-Firmen, und die Financial Times deckt auf, dass OpenAI und Google an geblacklistete chinesische Firmen über Singapur verkaufen. Blue Origin nimmt $10 Mrd. auf $130 Mrd. Bewertung auf, Lovable verhandelt eine Runde auf $13Mrd. Unterstütze unseren Podcast und entdecke die Angebote unserer Werbepartner auf doppelgaenger.io/werbung. Vielen Dank! Philipp Glöckler und Philipp Klöckner sprechen heute über: (00:00:37) GPT 5.6 endlich released (00:05:27) Grok 4.5 (00:10:07) Meta Muse Image/Video (00:14:07) Meta Brillen (00:21:35) OpenAI kauft Northslope (00:29:13) Cognition Devin (00:31:20) China warnt eigene KI-Firmen (00:36:48) OpenAI/Google verkaufen über Singapur (00:38:36) Blue Origin $10 Mrd. Runde (00:40:52) Lovable auf $13 Mrd. (00:41:43) Northstar: Europäischer Humanoid-Robot (00:44:04) München-Umnutzungsagentur Shownotes GPT 5.6 nach Trump-Ban released - axios.com Meta Muse Image & Muse Video vorgestellt - ai.meta.com Peking will Zugang zu chinesischen Modellen einschränken - reuters.com OpenAI Deployment Company kauft Northslope - axios.com SpaceX AI & Cursor unveiled Grok für Legal/Finance - bloomberg.com Musk-Tweet zu Grok 4.5-Launch - xcancel.com Meta testet Super-Sensing-Brillen - ft.com Blue Origin raist $10 Mrd. auf $130 Mrd. Bewertung - nytimes.com US-Lawmakers prüfen Nutzung chinesischer KI in US-Firmen - cnbc.com OpenAI & Google verkaufen AI an geblacklistete China-Firmen über Singapur - ft.com Cognition (Devin-Macher) launcht neues Modell - xcancel.com Meta bringt Iris-KI-Chips im September in Produktion - reuters.com Lovable verhandelt $13,2 Mrd. Bewertung - techcrunch.com Umnutzungsagentur München: Büro wird Wohnraum - linkedin.com Ex-Tesla-Scientist baut europäischen Humanoid-Robot (Northstar) - bloomberg.com Hackergruppe Everest verkauft 460 GB Platform-Group-Daten - xcancel.com
I recorded a few weeks ago an episode with Marissa Paragano in which we talked about the 2025 numbers from Comercam. But we got so distracted by overall sales that I forgot to highlight this one very important fact: plain, old Mezcal (in other words, the stuff that can be made industrially) has doubled in one year. Marissa was too busy with her weekly YouTube show, “The TequiLadies,” so Linda Sullivan got tagged in for the conversation! Agave Road Trip is a critically acclaimed, award-winning podcast that helps gringx bartenders better understand agave, agave spirits, and rural Mexico. This episode is hosted by Lou Bank with special guest Linda Sullivan of seynasecreto. Episode Notes Catch the 2025 Mezcal numbers from Comercam here and listen to Marissa and I talking the numbers in “Is Mezcal searching for mainstream success?” Shout outs this episode to Humboldt Park, Chava's podcast, “Heritage Mezcal,” Jonathan McKinney, Tequila Arriesgado, the Wisconsin Old Fashioned, Fausto of Asil and El Acabo Raicillas, and Marsh Hen Mills grits. Ad Links Every time you drink El Acabo Raicilla, you're helping to support the biodiversity of Jalisco — it's a delicious way to help the environment! Since 2005, Chicotona has been producing Mezcal with a focus on protecting wild agave and preserving the land for future generations! Head out on an Agave Road Trip with Finca 18! Greg Rutkowski will take you on his Agave Road Trip Route #2 - Raicilla de la Costa! Price includes a bottle of Paulo Rodriguez's fabled, limited Tumbado batch! Order beautiful spirits to be delivered to anywhere in Mexico — beautiful or otherwise — through Agave Spirits Presents!
“Before, only white people or gringos used to sell Mezcal, and the consumer was used to only buying it from those people. But now the consumer is changing their mind and saying that I want to buy from the producer. I want to know how it's made. I want to know if they are planting more agaves, or if they are not cutting trees, or what are they doing for the environment. So we have the information here. Not all of it, but maybe enough information to show the people how we do this Mezcal, how we do Palomo.” That's mezcalero Carlos Mendez Blas, owner and producer of Palomo Mezcal. We profile Carlos and Palomo in this episode of Agave Road Show! Agave Road Show is an advertorial podcast that helps gringx bartenders better understand specific brands of agave spirits. Sometimes other Mexican spirits. Really, whatever they pay us to profile. This episode is hosted by Lou Bank with special guest Linda Sullivan of seynasecreto.
I don't know, maybe I'm just behind the curve on this one. But I had a conversation with Marcado 28 Tequilero Bruno Barba and, during that conversation, he brought up a subject that surprised me: the cleaning of the bottles. And it made me realize, there's an “abocado con” Tequila I never imagined! Agave Road Trip is a critically acclaimed, award-winning podcast that helps gringx bartenders better understand agave, agave spirits, and rural Mexico. This episode is hosted by Lou Bank with special guest Marissa Paragano (AKA the eye-rolling, not-easily-offended, kind-of-5'2” Tequila Encyclopedia) of The TequiLadies Show and The Tequila That Cares Foundation and added wisdom from Bruno Barba at Marcado 28 Tequila. Episode Notes Marissa is a board member of Tequila That Cares, a philanthropic organization bringing positive change to the agave spirits industry! Shout outs this episode to Cava de Oro Tequila, chef Michael Voltaggio, Chef 28, and Nock Tequila! Listen to our bird poop and horse piss episodes! Ad Links Whatever you're talking about when you're talking about G4 Tequila, what it boils down to is Felipe Camarena! Head out on an Agave Road Trip with Finca 18! Greg Rutkowski will take you on his Agave Road Trip Route #2 - Raicilla de la Costa! Price includes a bottle of Paulo Rodriguez's fabled, limited Tumbado batch! Order beautiful spirits to be delivered to anywhere in Mexico — beautiful or otherwise — through Agave Spirits Presents!
Tom Uren and James Wilson talk about Chinese AI labs stealing the special sauce of American AI models in ‘distillation attacks'. These attacks are fed by a grey market in which Chinese consumers buy access to American models, where one of the byproducts is logs of user requests and responses. These make wonderful inputs into distillation attacks and the whole market might be subsidised by Chinese AI Labs paying for these logs. They also discuss the possibility that last year's hack of Jaguar Land Rover was caused by a group of Russian hackers. Was it Russians? Was it state-directed or endorsed? Who knows, but even the possibility that it was has some benefits for the Russian state. This episode is also available on YouTube Show notes
On this week's show Patrick Gray, Adam Boileau and James Wilson discuss the week's cybersecurity news. They cover: Anthropic's Fable 5 returning while OpenAI's GPT-5.6 gets thrown in model jail Distillation, cheap tokens, and AI chat harvesting is an industry in China Edge becomes a lolbin via a new malicious extension An Iranian APT boss's vacation in a beautiful place goes wrong Much, much more! In this week's sponsor interview Daf Stuttard and Katie Warren from Portswigger pop along to talk about how they built an AI security testing product that people would actually feel comfortable using. This episode is also available on YouTube. Show notes Anthropic (@AnthropicAI) on X | X (formerly Twitter) Howard Lutnick (@howardlutnick) on X | X (formerly Twitter) U.S. government gives Anthropic green light for limited re-release of Mythos 5 | NBC News Tech OpenAI limits GPT-5.6 rollout after government request | TechCrunch The U.S. government will decide who gets to use the latest American AI technology | washingtonpost.com Anthropic says Alibaba illicitly extracted Claude AI model capabilities | reut.rs How to Buy Cheap Claude Tokens in China | Alex Stamos (@alexstamos) on X | X (formerly Twitter) Synthesis of Exploitarium Mass Zero-Day Disclosure | detections.ai Mythos on your desk? Using local LLMs for code reviews | Risky Business Media Beyond Fable: Can a Local LLM Replace Cloud AI for Security Code Reviews | Security Research Labs Accelerating EDR Evasion with LLM-Driven Analysis | SpecterOps CISA: Windows BlueHammer flaw now exploited by ransomware gangs | BleepingComputer When cybercriminals hire burglars: Inside an alleged Russian effort to infiltrate multibillion-dollar US law firms | CNN Politics | Social Signals Microsoft quietly extends free Windows 10 ESU support to October 2027 | BleepingComputer Edgecution: Malicious Edge Extension Backdoor | ThreatLabz | Social Signals Bluekit phishing kit adopts browser-in-the-middle for login theft | BleepingComputer New macOS malware embeds fake errors to confuse AI analysis tools | BleepingComputer DraftKings hacker 'Snoopy' sentenced to 18 months in prison | BleepingComputer Polymarket says hackers stole users' funds | TechCrunch Security Australia's spy chief warns of rising terror and cyber threats | japantimes.co.jp Russian hackers were behind $2.5 billion hack of Jaguar Land Rover: Report | TechCrunch Security Iranian national sought by US on hacking charges arrested in Montenegro | apnews.com [un]prompted.au - AI x CyberSecurity: Notes from the Field: Call for Speakers |
เคยสงสัยไหมว่าทำไมคอมพิวเตอร์แพงขึ้นอย่างกะทันหัน ทั้งที่เทคโนโลยีควรจะถูกลง? ความจริงแล้ว เบื้องหลังการเลื่อนเปิดตัว AI ตัวเทพอย่าง Fable 5 หรือ GPT 5.6 ไม่ใช่เรื่องบังเอิญ แต่เป็นผลมาจาก “สงครามจารกรรมข้อมูล” ระดับโลกที่ Anthropic กำลังเปิดศึกกับยักษ์ใหญ่จากจีน จนส่งผลกระทบไปถึงห่วงโซ่การผลิตชิปหน่วยความจำทั่วโลก นี่คือเหตุผลว่าทำไม Data Center ที่เรามองไม่เห็น ถึงกำลังทำให้ราคา MacBook ในมือคุณพุ่งสูงขึ้น และทำไมยุค AI ที่ดูเหมือนจะฟรี กำลังจะเปลี่ยนไปตลอดกาล เลือกฟังกันได้เลยนะครับ อย่าลืมกด Follow ติดตาม PodCast ช่อง Geek Forever's Podcast ของผมกันด้วยนะครับ #AI #ปัญญาประดิษฐ์ #สงครามเทคโนโลยี #เศรษฐกิจโลก #ข่าวไอที #TechWar #ChipShortage #Apple #DataCenter #เทคโนโลยี #geekdaily #geekforeverpodcast
This liquor made with (mostly) blue agave has a complex set of rules and regulations – and a near-mythical history. In the final episode of the series, Anney and Lauren say ‘salud’ with the science and stories behind tequila.See omnystudio.com/listener for privacy information.
AI Chat: ChatGPT & AI News, Artificial Intelligence, OpenAI, Machine Learning
In this episode, we cover Anthropic's allegation that Alibaba-linked operators used nearly 25,000 fake accounts and 28.8 million Claude interactions in what it calls its largest known distillation attack. We also look at why model distillation is becoming a major front in the U.S.-China AI race. Show LinksGet the top 80+ AI Models for $8.99 at AI Box: https://aibox.aiHow I Grow and Scale My Business with AI: https://www.skool.com/aihustleGet the AI Chat Daily Newsletter: https://www.aichatdaily.com/newsletter
There are scientific names and common names for agaves … or magueys, to use the common name. But the common names lack commonality: A papalometl in one part of the Mixteca is a potatorum. Less than 40 kilometers away, papalometl will refer to a Cupreata. And even the scientific names get screwy. So what's a nerd to do? Agave Road Trip is a critically acclaimed, award-winning podcast that helps gringx bartenders better understand agave, agave spirits, and rural Mexico. This episode is hosted by Lou Bank with special guest evolutionary biologist Daniel Moen. Episode Notes Thanks to Google Books for the agave image on this week's cover, which comes from El Maguey: Memoria Sobre el Cultivo y Beneficio de sus Productos (1901) by Jose Carmen Segura. Shout outs this episode to Michaela at Waffle House #431 in Tucson, the Agave Heritage Festival, mezcalero Ildefonso Macedas Ginez, mezcalero Amando Alvarado Alvarez, Jason Paul Cox and Cinco Sentidos, and Hidden Rose Apples.
In this episode, we cover Anthropic's allegation that Alibaba-linked operators used nearly 25,000 fake accounts and 28.8 million Claude interactions in what it calls its largest known distillation attack. We also look at why model distillation is becoming a major front in the U.S.-China AI race. Show LinksGet the top 80+ AI Models for $8.99 at AI Box: https://aibox.aiHow I Grow and Scale My Business with AI: https://www.skool.com/aihustleGet the AI Chat Daily Newsletter: https://www.aichatdaily.com/newsletter See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
In this episode, we cover Anthropic's allegation that Alibaba-linked operators used nearly 25,000 fake accounts and 28.8 million Claude interactions in what it calls its largest known distillation attack. We also look at why model distillation is becoming a major front in the U.S.-China AI race. Show LinksGet the top 80+ AI Models for $8.99 at AI Box: https://aibox.aiHow I Grow and Scale My Business with AI: https://www.skool.com/aihustleGet the AI Chat Daily Newsletter: https://www.aichatdaily.com/newsletter See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
ChatGPT: OpenAI, Sam Altman, AI, Joe Rogan, Artificial Intelligence, Practical AI
In this episode, we cover Anthropic's allegation that Alibaba-linked operators used nearly 25,000 fake accounts and 28.8 million Claude interactions in what it calls its largest known distillation attack. We also look at why model distillation is becoming a major front in the U.S.-China AI race. Show LinksGet the top 80+ AI Models for $8.99 at AI Box: https://aibox.aiHow I Grow and Scale My Business with AI: https://www.skool.com/aihustleGet the AI Chat Daily Newsletter: https://www.aichatdaily.com/newsletter
ChatGPT: News on Open AI, MidJourney, NVIDIA, Anthropic, Open Source LLMs, Machine Learning
In this episode, we cover Anthropic's allegation that Alibaba-linked operators used nearly 25,000 fake accounts and 28.8 million Claude interactions in what it calls its largest known distillation attack. We also look at why model distillation is becoming a major front in the U.S.-China AI race. Show LinksGet the top 80+ AI Models for $8.99 at AI Box: https://aibox.aiHow I Grow and Scale My Business with AI: https://www.skool.com/aihustleGet the AI Chat Daily Newsletter: https://www.aichatdaily.com/newsletter See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
In this episode, we cover Anthropic's allegation that Alibaba-linked operators used nearly 25,000 fake accounts and 28.8 million Claude interactions in what it calls its largest known distillation attack. We also look at why model distillation is becoming a major front in the U.S.-China AI race. Show LinksGet the top 80+ AI Models for $8.99 at AI Box: https://aibox.aiHow I Grow and Scale My Business with AI: https://www.skool.com/aihustleGet the AI Chat Daily Newsletter: https://www.aichatdaily.com/newsletter See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
In this episode, we cover Anthropic's allegation that Alibaba-linked operators used nearly 25,000 fake accounts and 28.8 million Claude interactions in what it calls its largest known distillation attack. We also look at why model distillation is becoming a major front in the U.S.-China AI race. Show LinksGet the top 80+ AI Models for $8.99 at AI Box: https://aibox.aiHow I Grow and Scale My Business with AI: https://www.skool.com/aihustleGet the AI Chat Daily Newsletter: https://www.aichatdaily.com/newsletter
In this episode, we cover Anthropic's allegation that Alibaba-linked operators used nearly 25,000 fake accounts and 28.8 million Claude interactions in what it calls its largest known distillation attack. We also look at why model distillation is becoming a major front in the U.S.-China AI race. Show LinksGet the top 80+ AI Models for $8.99 at AI Box: https://aibox.aiHow I Grow and Scale My Business with AI: https://www.skool.com/aihustleGet the AI Chat Daily Newsletter: https://www.aichatdaily.com/newsletter See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
In this episode, we cover Anthropic's allegation that Alibaba-linked operators used nearly 25,000 fake accounts and 28.8 million Claude interactions in what it calls its largest known distillation attack. We also look at why model distillation is becoming a major front in the U.S.-China AI race. Show LinksGet the top 80+ AI Models for $8.99 at AI Box: https://aibox.aiHow I Grow and Scale My Business with AI: https://www.skool.com/aihustleGet the AI Chat Daily Newsletter: https://www.aichatdaily.com/newsletter See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
In dieser Folge ordnen wir die Entwicklung im digitalen Ökosystem und im E-Commerce ein und sprechen zunächst über die Frage, wie sich Frontier-Modelle wie OpenAI und Anthropic von destillierten Modellen unterscheiden. Wir erklären, dass Destillation bedeutet, ein kleineres, günstigeres Modell mit dem Output eines großen Modells zu trainieren, und diskutieren, warum dies technisch schwer zu verhindern ist. Wir sprechen dann über die wirtschaftliche Seite dieser Modelle und darüber, ob die hohen Bewertungen von OpenAI und Anthropic gerechtfertigt sind. Aus unserer Sicht sind die Bewertungen sehr ambitioniert, weil die Kosten für neue Modelle stark steigen, viele Anwendungsfälle keine AGI erfordern und günstigere Modelle oft ausreichen. Dazu kommt, dass die Zahlungsbereitschaft vieler Nutzer und Unternehmen begrenzt ist. Ein weiterer Schwerpunkt ist die Frage, wer am Ende die Wertschöpfung kontrolliert. Wir diskutieren die These, dass große Plattformen wie Google, Apple oder Microsoft langfristig eher als Distributions- und App-Store-Ebene profitieren könnten, während die Modellanbieter unter Druck geraten. Dabei geht es auch um die geringe Zahl zahlender Nutzer und die Bedeutung einer besseren Nutzeroberfläche für KI-Anwendungen. Außerdem sprechen wir über einen kurzfristigen Zugangsstopp zu einem neuen Anthropic-Modell und was das über Abhängigkeiten von US-Technologie zeigt. Wir sehen darin vor allem ein Beispiel dafür, wie schnell Zugänge zu zentralen Tools eingeschränkt werden können, und warum europäische eigene Fähigkeiten wichtiger werden. Vorbereitungsdokument von Julian: https://www.kassenzone.de/wp-content/uploads/2026/06/Distillation.pdf Partner in der Folge: https://linktr.ee/kassenzone Community: https://kassenzone.de/discord Feedback zum Podcast? Mail an alex@kassenzone.de Disclaimer: https://www.kassenzone.de/disclaimer/ Kassenzone” wird vermarktet von Podstars by OMR. Du möchtest in “Kassenzone” werben? Dann https://podstars.de/kontakt/?utm_source=podcast&utm_campaign=shownotes_kassenzone Alexander Graf: https://www.linkedin.com/in/alexandergraf/ https://twitter.com/supergraf Youtube: https://www.youtube.com/c/KassenzoneDe/ Blog: https://www.kassenzone.de/ E-Commerce Buch 2019: https://amzn.eu/d/5Adc1ZH Plattformbuch 2024: https://amzn.eu/d/1tAk82E
A recent episode of the VinePair podcast, titled "Mezcal's Continued Search for Mainstream Success,” feels to me like it was riddled with inaccuracies. And for some reason, it also feels to me like it stands in stark contrast to the recently released 2025 numbers from Comercam, the largest of the bodies certifying Mezcal. So I asked one of my favorite numbers gabbers, Marissa Paragano, to gab about that in this gabfest episode of Agave Road Trip! Agave Road Trip is a critically acclaimed, award-winning podcast that helps gringx bartenders better understand agave, agave spirits, and rural Mexico. This episode is hosted by Lou Bank with special guest Marissa Paragano (AKA the eye-rolling, not-easily-offended, kind-of-5'2” Tequila Encyclopedia) of The TequiLadies Show and The Tequila That Cares Foundation. Episode Notes Marissa is a board member of Tequila That Cares, a philanthropic organization bringing positive change to the agave spirits industry! Listen to The VinePair Podcast: “Mezcal's Continued Search for Mainstream Success.” And do also check out The Cocktail College Podcast: “Mezcal's Untold Past, Soaring Present, and Fragile Future”! Read Comercam's 2026 report, covering Mezcal (and uncertified agave spirits) commercial numbers for 2025. Shout outs this episode to the national parks, Kegel pelvic-floor exercises, GameStop, Auburn, Alabama, Jelly Babies, and on-premise and off-premise popcorn sales!
(Presented by TLPBLACK: A cybersecurity intelligence platform focused on sharing curated, high-sensitivity threat insights and research with trusted security professionals.) Three Buddy Problem - Episode 101: We discuss Anthropic's Mythos 5 and Claude Fable 5 release and the bombshell that the company was silently downgrading paid users' results, sparking a heated debate over guardrails, gatekeeping, and whether elite AI reasoning is becoming a privilege for the few. Plus, AI-generated N-day exploits killing the patch window, a record-shattering Patch Tuesday, Meta's latest court filing against spyware maker NSO Group, the return of cyber paleontology, and a detour into the new government UFO drops. Cast: Juan Andres Guerrero-Saade, Ryan Naraine and Costin Raiu. Timestamps: 0:00 - Introductory banter 3:22 - The Mythos 5 / Claude Fable 5 release 14:42 - Anthropic's silent downgrade trust problem 26:18 - Anti-competitive behavior & the AV "stealing detection" parallel 32:29 - Distillation, China & the real motive 38:04 - "Too dangerous to release" & gatekeeping vs. guardrailing 45:53 - Is Mythos a threat to malware-analysis startups? 48:20 - Dario's AI regulation essay 56:48 - N-day exploits and death of the patch window 1:07:18 - Patch Tuesday and 10x vulnerability surge 1:10:34 - Meta catches NSO Group 1:14:45 - Cyber paleontology, Shadow Brokers leaks 1:28:29 - Moonlight Maze and learning from history 1:34:22 - UFOs, UAPs and Disclosure Day
We've talked about agave straws in previous episodes, but I've been spending more time digging into them. And more than that, I've been spending time reading about – and now speaking with – Dr. Sandra Pascoe Ortiz, who has developed a plastic that is compostable and leaves behind no microplastics. So is this a solution for our plastics problem? Agave Road Trip is a critically acclaimed, award-winning podcast that helps gringx bartenders better understand agave, agave spirits, and rural Mexico. This episode is hosted by Lou Bank with special guest Linda Sullivan of seynasecreto and wisdom from researcher and chemical engineer Dr. Sandra Pascoe Ortiz. Episode Notes Listen to “Are agave straws really better than plastic?,” “Agave Road Trip,” season 4, episode 37. Read “Microplastics Everywhere,” Harvard Medicine magazine, Spring 2023. Shout outs this episode to D2W, CO2, and R2D2. Ad Links If you want a Tequila that reaches back in time, go check out Tequila Arriesgado! Head out on an Agave Road Trip with Finca 18! Greg Rutkowski will take you on his Agave Road Trip Route #2 - Raicilla de la Costa! Price includes a bottle of Paulo Rodriguez's fabled, limited Tumbado batch! Order beautiful spirits to be delivered to anywhere in Mexico — beautiful or otherwise — through Agave Spirits Presents!
There are places that have massive collections of beautiful, small-batch agave spirits, but having the bottles isn't the same as understanding them. And having those bottles, which are often expensive, requires additional expense in training staff. Or, anyway, are most likely going to realize their financial potential with training. But … who does that? Where do agave geeks go to have the geekiest conversations? Agave Road Trip is a critically acclaimed, award-winning podcast that helps gringx bartenders better understand agave, agave spirits, and rural Mexico. This episode is hosted by Lou Bank with special guest evolutionary biologist Daniel Moen. Thanks to photographer Russell Lee (1903-1986) for this week's cover photo, titled, “Proprietor of barroom near Crowley, Louisiana. This man is a Cajun.” Photo retrieved from the Library of Congress. Shout outs this episode to Michaela at Waffle House #431 in Tucson, the Agave Heritage Festival, Lamata Spirits, Mezonte, John Douglass of Pretty Decent, Fausto Romero of El Acabo Raicilla, and Ivan Vasquez of Madre.
“This is my playground. And we are always looking for the best profile we can get. Different profiles, different changes.” Those are the words of Master Tequilero Felipe Camarena, founder of G4 Tequila. We profile Felipe and G4 in this first episode of Agave Road Show! Agave Road Show is an advertorial podcast that helps gringx bartenders better understand specific brands of agave spirits. Sometimes other Mexican spirits. Really, whatever they pay us to profile. This episode is hosted by Lou Bank with special guest Linda Sullivan of seynasecreto.
We're announcing AIEWF speakers this week! Take the AI Engineering Survey!Today's guest Ethan first joined us for the LS Paper Club as the lead on NVIDIA Cosmos World Model, but then joined xAI and built Grok Imagine in 3 months:He comes back on Latent Space with some nuclear hot takes: that Video Models primarily get their intelligence from LLMs, not from training on video data, and that the next frontier for truly interactive, realtime, long-horizon world models is to work on LLMs (perhaps Interaction Models as well…)Put it this way: In the near term, the next Sora won't be a better video model, but a video agent.Generative Media may more closely follow the evolution of AI coding which went from focusing on one-shot output performance and cost, to multiturn reasoning and planning models for agents and systems that can plan, edit, test, debug, and submit PRs.At a certain point, coding models got so good that the only significant next step to improve performance was handling the orchestration of these models.Now as the performance of video models increases significantly across realism, consistency, & prompt adherence while becoming more cost efficient, the next evolution of video generation may also be systems that can plan, generate, edit, critique, and iterate across an entire creative task. In this episode, Ethan joins swyx and Vibhu to unpack what it actually takes to build frontier image and video systems: data, VAEs, diffusion transformers, audio-video alignment, inference speedups, and the hidden cost of storing and moving massive video datasets. From building NVIDIA's Cosmos world model to joining xAI as Grok Imagine was being built from zero to one, Ethan He has been at the center of some of the most important work in video generation, multimodal models, and real-time world models.We go deep on Grok Imagine, how a small xAI team shipped its first multimodal video model in three months, why iteration speed matters more than almost anything in model development, and why many of the biggest gains come from fixing tiny bugs in data and training pipelines. Flipbook: The future of VideomaxxingVideo agents are almost a sure bet to be the trend in the coming year. We end with a glance at what's beyond video agents:Flipbook caused a minor sensation this year when it was released, but most treat it as a fun demo. Ethan takes it very seriously — with the speed and cost of inference coming down every year, the future of custom video JIT UI is closer than you think. We talked about why videogen models may become the front end of AI, how generative UI could replace traditional HTML/CSS, why world models need to be real-time, interactive, and long-horizon, and why the future of video generation may depend more on language models and agents than on diffusion alone.We discuss:* Why fast iteration mattered more than meetings* Why small training bugs can drive huge model quality gains* Why coding models may make compute the bottleneck again* How image and video models are trained with synthetic captions* The role of VAEs and latent space in frontier video models* Why image models are the foundation for video models* The tradeoff between temporal compression and real-time interactivity* Flipbook, Neural OS, and the future of generative UI* Why future interfaces may go from user intent to pixels* The hidden cost of training video models: storage, egress, and GPU hours* How step distillation and consistency models (like OpenAI sCM) makes video inference orders of magnitude faster* Grok Imagine 0.9 and large-scale audio-video generation* Why audio-video alignment is harder than text-video alignment* Ethan's definition of world models* Reference-to-video, video extension, and long-context video generation* Why xAI's research communication undersells Grok Imagine* How xAI culture shaped the speed of development* AI watermarking, SynthID, and detecting generated media* Why prompt rewriting matters for video models* Grok Imagine Agent and the rise of video agents* Why language models may unlock better video generation* Robotics, physical AI, and embodied world models* Why Ethan left xAI and shifted focus toward LLMs* Self-managed context, memory, and the next frontier for language modelsEthan He* LinkedIn: https://www.linkedin.com/in/ethanhe42* X: https://x.com/EthanHe_42Timestamps00:00:00 Introduction00:01:25 From NVIDIA Cosmos to xAI00:03:24 Building Grok Imagine from Zero to One00:10:07 How Image and Video Models Are Trained00:18:53 Video Compression, VAEs, and Real-Time Tradeoffs00:22:10 Generative UI, Flipbook, and Neural OS00:32:10 The Cost of Training Large Video Models00:37:04 Distillation, GANs, and Fast Video Inference00:41:21 Audio-Video Generation and Grok Imagine 0.900:48:34 What Makes a World Model?00:55:51 Reference Videos, Long Context, and Video Memory01:00:11 xAI Culture, Research, and First-Principles Building01:09:45 AI Safety, Watermarking, and Prompt Rewriting01:13:10 Video Agents and AI-Assisted Creation01:27:32 Why Language Models Unlock Better Video01:31:15 Robotics, Physical AI, and Embodied World Models01:32:38 Why Ethan Left xAI01:34:16 Self-Managed Context and the Future of LLMs01:38:43 Ethan's Career Path and Closing ThoughtsTranscriptIntroduction: Ethan He, Latent Space, and the Path to xAISwyx [00:00:00]: We're here in the studio with Ethan He, most recently of xAI. Welcome.Ethan [00:00:10]: Thank you. Glad being here.Swyx [00:00:11]: We're also here with Vibhu. you were first coming to us or joining the latent space world because you were working on Kosmos at NVIDIA, and you did a paper. We loved it. you presented it as well, so thank you for doing that.Ethan [00:00:23]: I've actually, I also presented the MoEs twice at latent space.Swyx [00:00:29]: How did you actually hear about us? Did we reach out to you? Is that how it worked?Ethan [00:00:33]: No, actually, I-- the community. Like I realized, oh, there is this online community that people talk about AI and also learn from each other through papers every week through the Paperclip. It's very nice.Ethan [00:00:49]: I learned a lot.Swyx [00:00:49]: I think three years stop. We haven't stopped even on Christmas and New Years. many weeks I want to stop but it keeps going.Vibhu [00:00:58]: No, that was good. I think you had posted that you worked on a paper, and I was “Oh, very cool. We have Paperclip. Present then.”Vibhu [00:01:04]: But I might have reached out to you after.Swyx [00:01:05]: you-- because it's an amateur club, right?Swyx [00:01:08]: so it's very unusual and but we have sometimes paper authors come by and actually explain the paper. Today we just did, the poolside paper, which was apparently very good.Vibhu [00:01:18]: Came out yesterday.Vibhu [00:01:19]: pretty interesting, right? Fully open. They talk about everything, systems. So it's a good one. We'll, we'll recommend people to read it.Swyx [00:01:25]: Bring us up to speed on your transition to xAI, ‘cause I actually don't even know when you joined. just like tell the, tell the story about the sort of transition.From NVIDIA Cosmos to xAI: Scaling Video and World ModelsEthan [00:01:34]: Before xAI, I was working on Kosmos world model as in-- at NVIDIA. So Kosmos is, it's a giant video foundation models that can-- that aims to simulate the world and for-- it serves as a foundation of-- for all of the roboticists to build on top of. There, once I built the Kosmos one, I realized as this thing also has a scaling law similar to language model, we need to scale up the video models further. that's, that's why I realized I need to move to somewhere with much more compute resources. That's how ISwyx [00:02:13]: Than NVIDIA?Vibhu [00:02:14]: The GPU rich came themselves.Vibhu [00:02:19]: And timeline-wise, when was Kosmo? It was pretty early, right? It was open world model, open paper, everything.Ethan [00:02:25]: It was end of twenty-four.Vibhu [00:02:28]: End of twenty-four.Ethan [00:02:30]: Then at mid twenty-five, I moved to xAI. At that time-- I joined about the time when xAI was about to build video models and in multi-model models. There were no infra, no data, and no model, and it just-- as a few engineers, we built it in three months and released the first model, Grok Imagine zero point nine.Ethan [00:02:55]: And since then, I keep working on video models and move more from training and to post-training of the video models. For example, like a reference to videos, kind of like the cameo feature and, video extensions. And, before I left, I worked on a world model, leading a small team to focus on the real-time long horizon video generation.Building Grok Imagine From Scratch in Three MonthsSwyx [00:03:24]: Can you give like a rough roadmap of okay, you're on a brand-new team. Grok previously was only text, or they partnered with BFL for their image gen stuff. What do you-- what are the building blocks, right? You have compute, data you can procure somewhere. Like just what are like the sequence of things that people should think about when you're setting up a new team?Vibhu [00:03:43]: actually even deeper, not just data you can procure. You guys had to go through getting the data too, right? So you shipped it pretty fast, but yeahSwyx [00:03:51]: three months is likeVibhu [00:03:52]: From everythingSwyx [00:03:52]: actually like very surprisingly fast.Ethan [00:03:55]: One thing I say like thanks to my experience at NVIDIA, ‘cause first time when we were building Kosmos together, we built it, for about a year. So this is like the second time I do it. Roughly have an idea, what to do. I say the most important thing is the talent. Everyone were very strong and clever, very close with each other towards a common goal. So that speed up things a lot. So you reduce the communication bandwidth among people, and everyone can work towards the same goal. It's, it's like every day there's not that much meetings on the calendar, like maybe like a, like a sync a day, and after that it's, it's just all building. It was pretty fun at that time.Ethan [00:04:47]: And another thing is that xAI has very strong foundations of like data inference, model inference, and the supporting there can help the model develop a lot. When I look at, training models, I don't so actually the top important thing is like how many, how many iterations can you do, per day? and the more iteration can you do, you can, you can train the model much faster. So if you have very strong infra and you have a lot of compute, you can, you can train these models in very short period of time. That can give you a much larger buffer to, for errors, and it also gives you the opportunity to spot more bugs.Iteration Speed, Compute, and Debugging Model PipelinesSwyx [00:05:46]: What is an iteration? Is it like a few hundred steps or what are youEthan [00:05:50]: Let's say just the train-training the model, like from acquire new data and maybe design new algorithms and train a new model, maybe at smaller scale orSwyx [00:06:01]: So cycle time for like any hyperparam that you're searching.Ethan [00:06:04]: Cycle time and tune to like eval this model. Is this model better than my previous iteration?Ethan [00:06:11]: SoSwyx [00:06:11]: So it's like before you, someone had already set this up that you can iterate very quickly.Ethan [00:06:15]: I think the foundation there is extremely good forDeveloping and research models.Ethan [00:06:23]: And often I find is it-- this is kind of boring, but like a lot of the improvements does not come from new algorithms. It comes from finding small bugs here and there in the data pipeline, in the, in the model training pipeline. Those give, those give the biggest boost to the model quality.Vibhu [00:06:46]: It's interesting, right? So you say it's like small team, less communication bandwidth, but also a lot of quality is like find little bugs. It seems counterintuitive, right? You have a lot of people, you can iron out more of those, but it's interesting to see the other side, right?Swyx [00:07:00]: I also wonder, have you-- do you try using LLMs to look for bugs? I don't know.Ethan [00:07:05]: I remember at that time it was mid two thousand and twenty-five, so it's the coding model wasn't quite there yet. I remem- I remember like December two thousand and twenty-five, it was extremely good. Yeah, I've been, I've been using it at that time. It's, it's helpful. sometimes it produce codes that are kind of difficult to maintain, even though like the first time it built something extremely fast. But it gave the, like a spaghetti code, thousands of lines that I couldn't maintain, and the LLM itself couldn't figure out what's, what's wrong and how to improve on top of it. But now I find it much better. Yeah, I want to bring up another point here is now coding models are much more efficient and can help us implement stuff much faster. Compute might become a bottleneck again because previously, like if you want to train a new model, say you want to generate new synthetic data and then or write a new algorithm, it might take a few weeks. And during that period of time, you don't-- you might not have experiments to run. But now you can build that thing within a few hours, then you can immediately train a model.Ethan [00:08:24]: Now you have to have enough compute to try all of the ideas. So compute might be the bottleneck of iterating speed again.Swyx [00:08:36]: yeah, I actually, honestly, I think it's like kind of a stressful job because you're “Well, I should be trying everything, and if I'm not, then I'm not doing my job well.”Vibhu [00:08:48]: there's also the stress of you're eating thousands of GPUs per hour, which is very expensive and, compute can go to other researchers.Swyx [00:08:56]: You got the daddy Elon toVibhu [00:08:57]: You got daddy Elon.Ethan [00:08:59]: It wasVibhu [00:09:00]: But there's still finite amount of compute, like you want to use it, you want to use it well, you want more of it.Ethan [00:09:06]: That was quite stressful indeed. Yeah, I think one thing is the-- with coding models now, like a lot of these jobs can be automated, which is much better. A second, it's a, it's a marathon, so you got to maintain good health and, a regular schedule.Vibhu [00:09:28]: It's, it's hard to hear that when you shift from zero to nothing in two months.Swyx [00:09:32]: and, I think obviously the culture at xAI is very famously, people work very hard. one thing I did want to dive into, in our-- in the notes that you, that you sent ahead of time, you had specific comments about the cost of Video Gen training. presumably this is on the Colossus-1, right? the two hundred megawatt cluster. Any whatever you want to just share on that.Vibhu [00:09:54]: I think there's, there's three things we're talking about, right? So there's Video Gen, there's also the Image Gen model that you put out. Do you want to like complete the, okay, so zero to one, you have a few months. Just what are the stages of create Image Gen model?Swyx [00:10:06]: Oh, yeah, maybe I got distracted.How Image and Video Models Are Trained: Synthetic Captions, Tokenizers, and VAEsVibhu [00:10:07]: Sorry. and then, from there's Video Gen, there's Audio Gen. Would love to get into those next. But what is that first few months like? So small team, a lot of bugs, iterations, but what does it look like? Do we take something off the shelf? Do we just get data compute? What's, what's the few months like? How do you go to state-art Image Gen model? How do you just start?Ethan [00:10:28]: I cannot comment specifically how xAI did, but it's, it's a quite standard process. I can draw some, examples from Cosmos. So mainly it's building a video model, you actually need to build a image model first. And building these two models, the data you need is a hundred percent synthetic pair of language and image or language to video. Because on the, on the internet, actually, the videos don't naturally associate with text. So you can say, oh, like on YouTube, you have the title and you have the description and the commentsSwyx [00:11:11]: TitleEthan [00:11:11]: of a video, but usually they're not relevant to the video itself. And say maybe like the video is a natural scene of mountains or something, and the title is, I'm so happy today.Ethan [00:11:26]: So they have they have no correlation at all. So the first step is to, you have to generate synthetic pair of language with the videos. So you gather videos from the internet, and you use a VLM to caption the videos. So that part, here's a question, like how do you, how do you gather VLM to begin with? So if there's noSwyx [00:11:55]: You, so you fuse the model, right? LikeEthan [00:11:57]: Say if there's no like VLM exists, like how do you generate the text to the beginning, right? It's, it's impossible.Swyx [00:12:04]: I see.Ethan [00:12:05]: In the beginning, it's like you ask human to describe the video as detailed as possible.For example, you ask them to describe everything, like all objects, all characters, and all interaction and dialogues in the, in the videos. So that's in the protocol of Cosmos labeling. We require the objective we give to the labelers was that you have to describe the video as detailed as possible, such that a blind person hears a blob of text can reconstruct what the video is like from their head.Swyx [00:12:43]: Video or image? You're talking about images.Ethan [00:12:44]: Video or image, either one of them.Vibhu [00:12:47]: This was pretty common when we went from clip and DALL-E, right?Vibhu [00:12:51]: It's all training on really detailed captioning of images. So same is applied to video, but insteadEthan [00:12:57]: same appliedVibhu [00:12:57]: of using multimodal model to pass in video images and write rich descriptions, you can alsoSwyx [00:13:04]: I think there's this traditional perspective of supervised, or, very highly human curated thing. I feel like there's a unlock with unsupervised, right? Where like you have enough to bootstrap that you can just throw common corpus on it or, whatever. like unsupervised vision and language pairing, right? Like where you just have, interspersed image and text and it just learns. To me, that is the VLM breakthrough that is different from the clip, different from the LM era.Ethan [00:13:36]: It's interesting to see that you kind of need both data.Ethan [00:13:41]: For example, for theSwyx [00:13:41]: You need it to bootstrap it up. YeahEthan [00:13:43]: for the generative model training, there's also usually like a small percentage of unlabeled data. So the model is instructed to generate a video without any text instruction. That can also help the model generalize. So after this stage of generative synthetic pair, so, one important common step is to train a compressor or a tokenizer of the image or videos. So because, if you train-- If you can technically, theoretically train image or video models on pure pixels, but the problem is that the, it's, it's a lot of tokens. So like one image, it's, a thousand by a thousand, it's like one million tokens, one million pixels. It's impossible to train transformer on that. So it's, you need to train a tokenizer, which can go from image to latent space and latent space back to image.Swyx [00:14:45]: That's why we named the podcast.Swyx [00:14:48]: But, basically, you're talking about vocabulary science.Ethan [00:14:50]: so vocab.Swyx [00:14:51]: And so, what is, what is imp-- like a million is impossible?Ethan [00:14:54]: In generative models, the vocab is continuous. It's a continuous space. We can think about like you map an image to a vector. It's a, it's a fixed length vector. It's sixteen or forty-eight, something like that. And then you map that vector back to the image space. And the mapping is, has-- The mapping is patch-based. So you say you haveEthan [00:15:22]: a sixteen by sixteen patch and you match, you map that patch of pixels into this latent space.Swyx [00:15:29]: We've covered thisVibhu [00:15:30]: This is like the vision transformersSwyx [00:15:32]: VAEs,Ethan [00:15:33]: VAEs.Vibhu [00:15:34]: You basically compress your input, you do your generation, you're reasoning all that generation in smaller dimension, and then you project back out.Swyx [00:15:43]: VAE is a form compression, but I think the for me, the patching thing is from VIT, right?Ethan [00:15:48]: You can make those.Swyx [00:15:49]: Literally the, yeah, the paper is titled like sixteen by sixteen is all you need. something like that. and then I think also, people make a lot of comparisons with this kind of patching with convolutions.Swyx [00:16:02]: Which is you're, you're kind of re- reconstructing the old paradigm with the new.Ethan [00:16:05]: Actually, in VAEs, there are, there are both convolution networks and transformers. You can actually do both.Ethan [00:16:14]: After this VAE, so what you've got is you've got latent space tokens and you've got the language tokens. So now the training of the diffusion transformer, usually generative models use diffusion transformers. It is actually quite standard. It's, it's very similar to how you train a language transformer models. It's not that much difference. It's just the tokens, the visual tokens in, visual tokens out. The only difference is there's a denoising process. So you train the model to unmask some of the noise. So you add, you add random noise to the visual tokens, and then you train the model to remove those noise to generate the clean tokens. Any inference, the model can iteratively remove noise from a hundred percent noise.Swyx [00:17:12]: And then there's also, to speed things along on the tech tree of diffusion, there's CFG, and then there's, there's also, latent diffusion that, there's, there's someone in there. I think, somewhere along the line, obviously, like stability and all these other guys, pioneered a lot of this, architecture. I don't know if you want to get into that or just, or do the video side up to you.Bootstrapping Video from Image Models and Temporal CompressionEthan [00:17:37]: After you train such model, such image model, the reason it's a, it's a foundation for video models is that image models are cheaper to train, and they have much denser connection between language and text. So, sorry, language and images. For example, you train a billion, you train on a billion images, and there's a mapping from the text to the image. And the cost to train the same, like the, a billion, a billion text to a billion videos, that's much more expensive because videosNaturally have more tokens than images. Because the diffusion models, their understanding of, language purely come from this mapping. So if you don't have enough mapping, so if you only train on like a ten million videos or something, there-- you might not see enough language tokens in your training, so your model does not understand human intention enough. So that's why you really-- you train-- you first train this image diffusion models, and then you bootstrap the video model from there.Swyx [00:18:53]: One thing I did want to ask, because I-- actually, I think you're, you're the first per-- video model person I've ever talked to, I think. we've, we've like talked to Luma and all those folks. There's all these tricks in video compression where basically frame by frame there's not that much difference, so actually you don't have to regenerate or save the whole frame, right? but I think MP4 compression or something else like that.Swyx [00:19:16]: is it tempting to use that? Or as far as I can tell, everyone just treats it as, “No, we would just generate every frame.” Is that roughly the state-art?Ethan [00:19:27]: There are a few different approaches. Let's say first, like you want to just directly use MP4 compression and use that as the tokens for the transformers to train, right? So people actually have tried that, but the main challenge is the latent space for the MP4 tokens were not, were not very comprehensible for the models. It's, it's extremely hard to train on that. And there's aEthan [00:20:01]: So that's why they created VAEs, which creates more continuous, latent space, so the models can understand that latent space and learn from it much easier. Even within the VAEs, there are different difficulties of the latent space. So you can imagine something the simplest, the most naive VAE is like you have an image, and you just shuffle all of the images into a, into a vector. So you don't need to train any VAEs, right? But that latent space is extremely hard for models to train on top of. That's why there are some debate on like how do you compress the tokens. So you mentioned like you can compress frame by frame. Also, you can compress, the temporal dimension.Ethan [00:20:52]: The difference is if you compress the temporal dimension, you get a much higher compression rate. Because there's temporal redundancy between frames, because, this frame and the last frame, likely they are mostly similar, so there's only some small difference. for example, I think in 12.1 VAE, they have like a eight by eight by four compression rate. So the four temporal tokens are compressed into one tokens. That can save a lot of, save a lot of the context length. If you do it frame by frame, you have to do maybe like eight by eight by one. Your context length will be four times larger. That being said, the benefit of the frame-- per frame compression, we might come back to this later, is, real-timeness and interactivity. ‘Cause if you, if you strain the output of the model, frame by frame, you can-- the model can respond to any user request immediately. So if you have like a temporal four compression, four times compression, thenSwyx [00:22:06]: It might be laggyEthan [00:22:07]: there's a lag there in nature.Swyx [00:22:10]: So you're very pilled on this. let's just go ahead and bring it up ‘cause we have the visual prepared anyway. There's some frontier applications of real-time video gen. So Flipbook is one of the examples that went viral recently, right? What is Flipbook?Real-Time Generative UI: Flipbook, Neural OS, and Diffusion Front EndsEthan [00:22:23]: Flipbook is kind of like a web brow- web browser. You can see like it has the web bro- browser UI on top. The difference is all of the UIs are generated by generative image model in real time, and anything here are fake. But you can, you can explore inside this wor- this imaginary world. Say like we-- here we have engineering the Great Pyramid. Like the model generates this for us to understand how it works, and if we want to navigate around and understand further, we can click on some of the, some of the description here, and the model will generate a new page, new subpage describing the details we want to know about.Swyx [00:23:14]: So it's basically kind of we're playing a video, but it's pausing for our next interaction, and then it just plays the next thing based on our interaction.Swyx [00:23:23]: Which is kind of cool.Vibhu [00:23:25]: and you kind of decide your story. So this was, how do you make a pyramid? levering technique seemed interesting, right? It shows how do you take Okay, I want to know what is thisSwyx [00:23:35]: The demo, the demo tweet had more animation between frames.Vibhu [00:23:38]: I think it's just skipping,Swyx [00:23:39]: Oh, it's just skipping a lot of frames.Ethan [00:23:40]: they also have a video modeVibhu [00:23:42]: It takes a lot. There's a lot of peopleEthan [00:23:42]: but, a lot of people are using it.Ethan [00:23:45]: So it's not available.Vibhu [00:23:46]: There's a live video stream. We can try,Swyx [00:23:50]: So this is an example of the kind of future that you see at the extreme. We don't-- we're obviously not in it today.Swyx [00:23:56]: But in a world where inference is completely free this is better than generating code and text?Ethan [00:24:02]: So this is, this is a final state of where Viva will be at for word model, I think. Imagine internet doesn't exist, and then you type in google.com. Like what should, what should, what should a model show you?the model can imagine something, and this is what the model imagine. And these web pages, they completely do not exist. So I think as the inference costs come down, we are going to have generative UI for everything. If you think about how the coding model works, so they write code for a web page, and they render the code might be con- converted into binary, and the binary render the pixels on the screen. So we in machine learning, every time we have some breakthrough, obviously it's, it's more intuit. So why don't we have like user instruction to the pixel directly? So the generative UI will be user intention to the pixels directly. And say like even if I want email, let's say everyone have the same interface, but I want, I want it slightly different. I want the email to show to me like a TikTok, so I can swipe left and right for the emails. And or maybe you want something else. We can have completely different things. Or like I have I'm looking at, Instagram stories, and I don't like the Like button. I always may click it. And, generative UI resolved it. So it's going to be a revolutionary replacement of the interface. So in the future, we might have much more powerfulEthan [00:25:50]: LLMs and coding models running behind the scene. And in the, in the front-end, the diffusion model will actually be the front-end to show stuff to you. That's how I imagine it.Swyx [00:26:02]: Diffusion front-end, deterministic back-end.Swyx [00:26:04]: Something like that. I find that very expensive, but,Vibhu [00:26:08]: I find it interesting you called LLMs writing code on the back end deterministic, but okay.Swyx [00:26:14]: you write it onceVibhu [00:26:15]: Compare it toSwyx [00:26:16]: And then you execute.Ethan [00:26:17]: If you think about the cost, say, let's say H100 costs $1 per hour, and if you use this eight hours a day and thirty days, so, every month you're paying this two forty, you'll actually not wanna pay for that. That's even more expensive than Cloud Code Max. But if you think about the compute costs come down like two times every year, and I think the future will likely arrive like within few years.Vibhu [00:26:49]: It's everything, right? compute cost comes down, compute gets faster, model gets smarterEthan [00:26:54]: More efficientVibhu [00:26:54]: model gets smaller.Swyx [00:26:55]: I don't know why you say two times, ‘cause I think it's like 100 times. In language models, it is roughly one hundred to a thousand times every twelve to eighteen months, for the same given level of LMSys, ELO.Vibhu [00:27:08]: That's a net of everything, right? That's model performance alongside compute. So different than just compute costs come down. But, a very interesting future.Swyx [00:27:19]: So the web designers will have to shout out that accessibility is an issue, right? how do you deal with screen readers or whatever. But yes, this is higher bandwidth storytelling than anything you can possibly generate with code, right? So I think that's the rough idea.Ethan [00:27:34]: And I'd like to add a little bit that so human naturally have the maximum bandwidth when we are looking at things, look at videos, and we also have maximum output bandwidth when we are talking. So in the future, it might be something like we talk to AI models, and the AI model responds back with a generative UI. So that would be the maximum input and output bandwidth to interact with AI models before neural link happens.Vibhu [00:28:06]: And it's also very custom, right? Some people are very visual, some people are not as visual, right? They prefer the text. But the best thing about generative UI, right, it can also be text.Swyx [00:28:17]: There's another project that we wanted to highlight, which is the Neural OS. Kinda similar idea, but here you're literally operating, simulating an operating system with a video model.Swyx [00:28:27]: and you can play Doom, you can do Firefox. I find this like mildly less impressive, obviously, because it's an OS that I can run.Swyx [00:28:37]: But here everything is imagined.Vibhu [00:28:40]: I was, used to the Command+W to close the Firefox tab. It didn't crash. That's why I saidSwyx [00:28:45]: It's too immersive.Vibhu [00:28:46]: It's, it's too immersive for me.Swyx [00:28:47]: Too immersive.Vibhu [00:28:48]: I wanted to close the tab.Vibhu [00:28:49]: But yes, I can play generated diffusion.Swyx [00:28:51]: this is shockingly fast.Swyx [00:28:54]: Because I remember there was a demo about like maybe one to two years ago. Someone tried to do the first-person shooter with a image model. There was no consistency. It was very slow. But here it looks like realistically it's-- this is Doom.Vibhu [00:29:07]: I think there's two sides to that, right? There's okay, what is running a game? The heavy part of it is actually the game engine, all the lighting, all that stuff, the graphics. This is just kind of video, right? Like we've solved consistency. This is still, it looks like a few years old image generation. There's some temporal consistency, but it's, it's kind of just images stitched together as frame video. But it's a good visual representation to pi- to picture the future you wanna see, right? that's, that's what I see in these more so.Ethan [00:29:38]: This reminds me of how the video models gets better and better. So Neural OS is kinda if you just look at it feels like it's just a crappy version of the, like the Windows we could have, right? And, but the difference is, so the model, this model is overfitted on the existing operating systems. It can generate nothing different than that. But it's actually also similar to video models. So when we are training these video model, image model, we train them on internet. There's no imaginary supernatural stuff on the internet. But once we train this model, you can prompt the model to generate something supernatural that have never existed in the data set. So if you train your Neural OS or neural computer on the standard screen recordings on the entire internet. The model can imagine completely new interface to interact with the computer.Swyx [00:30:43]: This is one of those things that is magical to me. usually generalizing out of distribution is bad, but somehow we have learned some kind of internal world model that you say, this plus, but it looks like rainbows and butterflies, it'll do it and it will kind of make sense.Swyx [00:31:03]: So yeah, that's kind of cool. Yeah, I don't know if there's any comment more on there. I do, I do wanted to, I did wanted to touch a little bit more on the model architecture stuff, which I think you were getting. It's, really fascinating. We don't get a chance to talk about this enough. So one of the papers that we covered, we've covered every annual, segment anything release. and I don't know if you follow-- you're a computer vision guy, so youEthan [00:31:26]: I knowSwyx [00:31:27]: . So they did memory attention, which is kind of interesting. And I always think, anything where you can, across the temporal dimension, keep some consistency, I think it's, very fascinating, and I don't know if Basically, does that-- the CV side bleeding into video gen side, I think is underexplored, right? we talk about it for labeling, but actually you can borrow the architecture itself.Ethan [00:31:50]: There's, there's also complete different approaches, right? you brought up the term world model, so we went from video model to world model. There is diffusion, but there's also other approaches that people are doing. So maybe we get into those after as well,?Swyx [00:32:03]: He has a whole definition of world models and stuff. I feel like we threw a lot at you. Whatever you want to comment on.Why Video Models Are Expensive: Storage, I/O, and Training ScaleEthan [00:32:10]: I think one thing that we should actually comment back on is okay, so we were talking about the steps to train image gen to video model. One thing we don't see as much of is okay, you brought up the delta in training data, right? SoEthan [00:32:24]: you won't have as much a video model might not generalize, but what is the cost of training a large video model? So we know for LLMs roughly, okay, even like the poolside thing that came out today, right? It's a Gemma level model trained on roughly forty trillion tokens at this many H200s over this much time, right? You can see what is the exact cost of that. So how many GPU hours over how much H200 costs? So how do we do the back-end math of, same thing for video models, image models. How do you, how do you kind of break that down? I can share some back-envelope calculation. So surprisingly, video models is-- the cost is very-- is comparable to language models and obviously the largest scale is language model, maybe like a medium scale to language models. I said just storing the videos alone, it costs a lot. You can, you can maybe look up on AWS or something.Ethan [00:33:20]: You really, say if you have a billion videos and let's say, let's just say like each video, like five megabyte, then you need five petabyte to just store those videos. And also remember we talk about you use a VAE to compress the videos, and you also need to store, typically you need to store those continuous feature, in-- also in your storage. That's also comparable size with the videos themselves. So just storing these videos and the features is tens of petabytes alone. And,Swyx [00:33:58]: I just, I just looked up the calculation. Five petabytes on S3 Standard is one hundred K per month.Ethan [00:34:05]: AndSwyx [00:34:05]: It's comparableEthan [00:34:05]: and you needSwyx [00:34:06]: AndEthan [00:34:06]: And then like tens of petabytes, two hundred K. And even more expensive is you have the ingress and egress.Swyx [00:34:13]: Oh, yeah.Ethan [00:34:14]: Like you-- through the internet. You have to just to download those videos, I believe it's, it's more expensive on AWS than just storing those videos.Swyx [00:34:25]: Storing, yeah.Ethan [00:34:25]: And each training runs, you probably need to pull them once. If you train multiple times, it's, it's even more than that. So it's like just storing the network, those costs is just, it would be a few, a few millions per month to just storing everything, not to mention the GPU cost.Ethan [00:34:45]: AndSwyx [00:34:45]: my side tangent, the compute rental, like GPU rental is very efficient. There's one side, okay, you can be XAI and build your data center. Should we not just build our, storage compute as well? LikeEthan [00:34:57]: Of courseSwyx [00:34:57]: cloud cost compared to just,Ethan [00:34:59]: You save so muchSwyx [00:35:00]: store. Yeah, exactly.Swyx [00:35:01]: Especially with like egress and stuff. So.Ethan [00:35:04]: That's a good idea, but it also comes to-- there are some of its own challenges.Swyx [00:35:09]: Of course, of course.Ethan [00:35:10]: like people who build the GPU data centers, they might not expect this much, storage. And yeah, people build storage, typically they just build it somewhere with just CPUs.Swyx [00:35:23]: I just looked it up. Five-- AWS only charges for egress, not ingress. Tier five for five petabytes is two hundred and thirty K.Ethan [00:35:32]: Even more expensive than the storage.Swyx [00:35:34]: But storing is per month, right? You check in, then you cannot check out. so it's so cool. It's okay. So there's that side.Ethan [00:35:41]: So the TLDR, my backhand mathSwyx [00:35:42]: Data is larger than you think. Yes.Ethan [00:35:44]: my backhand math of GPU hours times GPU cost is also very much, I'm missing some storage.Swyx [00:35:49]: You're also-- you're basically like also more IO bound than normal training.Swyx [00:35:55]: Yes. ‘Cause like data loading, so caching everything, it becomes super important.Ethan [00:36:00]: So in Cosmos, we did a lot of optimizations to make it not IO bound. So, speaking of the training, actually training the model, the GPU cost, if you look up like the open source model, how big these video models are, I think like LTX has nineteen B parameters. That's a dense model. And people are also exploring, MoEs, so it might be twenty B active and, like a hun- hundreds B, total. So that's, that's even-- that's similar size as medium-sized LLM models. And if you, if you look at number of tokens-Uh, we disclose that in Cosmos. It's also like tens of trillions of tokens on the visual tokens. So putting this together, the cost of, training these video models, it's actually comparable with LLMs. Not to mention, the infra is slightly different from LLM, so it might be less efficient to train these models.Inference Speedups: Step Distillation, Consistency Models, and GANsSwyx [00:37:04]: Do you get the benefits of traditional diffusion speed-up? So for, images, there's LCM, LoRAs for, fine-tuning. There's, there's a lot of stuff that's beenEthan [00:37:15]: Flow matching.Swyx [00:37:16]: there's flow matching. There's a lot of stuff that's been done. there's some overlap that applies to diffusion on the inference side and stuff or?Ethan [00:37:23]: so the difference-- the inference side is a completely different story.Ethan [00:37:28]: I think for the training side, it might be a little bit hard to reduce that cost. And for the inference side, the biggest gain is from the distillation of these models. You can-- It's called step distillation, slightly different from knowledge distillation in LLMs. So you-- Typically, for flow matching models, you need like 100 steps or something. Like a distortion model even need even more, like 1,000 steps to generate a good image or video. A step distillation is try to learn to generate fewer step from the model itself. It's kind of like now we-- you use the full model to generate in 100 steps, and then you take a model that only generate 10 steps and let that model to learn from the perfect one.Ethan [00:38:25]: why this workSwyx [00:38:27]: Strong to weak seemingly.Ethan [00:38:28]: It is. It's kind ofSwyx [00:38:29]: DistillationEthan [00:38:29]: kind of like strong to weak. the-- from the modeling perspective, the strong model, the teacher model is trying to model the image and videos of inter-internet, and that distribution is extremely complex. But the step distilled model is just trying to learn from the teacher. The teacher is a model, and the size is fixed, as the distribution is much simpler than the whole internet. That's the intuition I have why step distillation can work. So usually these models serve in productions, they only run in a few steps. In Cosmos, I believe we have, we have like four step and eight steps. If you do some simpler task, image-image translation, it can even run in fewer step, like one step in Cosmos Transfer.Swyx [00:39:22]: I think this is the same intuition that guides a lot of the consistency model work. I sent you a link for, SCM. I don't know if you covered that. To me, that was actually one of, the most impressive papers I've ever seen from OpenAI.Swyx [00:39:34]: That this is the unifying grand concept of consistency models. I don't know if you have any comments on this.Ethan [00:39:41]: So there are, there are a few different approaches,Swyx [00:39:46]: Oh, yeah. Here it is.Swyx [00:39:47]: Two steps versus twenty or 100 steps, whatever. It's already done.Ethan [00:39:52]: So there are, there are a few different approaches, for example, consistency model, and there are also Actually, we shouldn't forget GAN. So GAN, actually, that was, that was the OG ofSwyx [00:40:05]: OGEthan [00:40:05]: step distillation ‘cause it trained just one step to begin with. So actually, a lot of, uh-- For example, there's a distribution matching distillation which use, which uses GAN, as one of the laws for distillation. It-- GAN just tells you, “Hey, generate an image,” and thenEthan [00:40:31]: it has a discriminator to tell, is this image real or not? So the model, the model just need to learn one of the distribution, not the full distribution. Because in training, the model is asked to reconstruct the ground truth image from the internet, which is extremely hard. And in-- When you're training GAN, it's a step process. It's just a, “Hey, you generate image. Does this image look as real as the image from the internet?” Which is a much simpler task. And, yeah, combining a lot of these approaches together, people typically do that, like consistency model and distribution matching and GAN, and we can get these few step models.Audio-Video Generation and Time AlignmentSwyx [00:41:21]: Then there's one step I wanted to add, which is audio and video.Ethan [00:41:26]: So, Grok Imagine zero point nine, I believe it's, it's a first audio video transmodel deployed at a large scale. SoSwyx [00:41:39]: And that was your first model?Ethan [00:41:40]: that was, Grok Imagine's first model. It's, it's audio video, joint generation. I think the hard part is, the modality alignment, ‘cause before this transmodel, we have, we have text to video alignment. We have this, correspondence between text and video. Typically, most of the VLMs, they understand images and videos. Video's very rare, and they don't understand audio mostly. And if you look at the audio generation on the LLM side, you can talk to them perfectly fine, but if you ask them to sing a song or something, it typically is not very good. Also, they don't have, they don't have music either. The hard part is thatUh, actually audio has two component. It has like a discrete component, a continuous component. The discrete component is like the language.Ethan [00:42:44]: So when we speak, it's just, someSwyx [00:42:47]: It's an ASR issue, yeah.Ethan [00:42:49]: It's, it's text token with some characteristics, I would say.Ethan [00:42:54]: But musicSwyx [00:42:56]: I think the speech guys would disagree with this.Swyx [00:42:57]: Like disfluencies and then,Vibhu [00:43:00]: There's tones you can get angry.Ethan [00:43:01]: Well, I say largely.Ethan [00:43:03]: the mu- but the music is completely different. It's, it's very continuous, and you cannot model them like discrete tokens in language models. this is like the hard part for models is, not to mention we have to align text, video, and audio together.Ethan [00:43:26]: SoVibhu [00:43:26]: How?Ethan [00:43:28]: So significant-- some significant challenges are like-- So first, like we talk about as the VLMs, they cannot understand most of them cannot understand audio.Ethan [00:43:39]: So you have to have some way to do the synthetic data generation for audio. You have to caption the model, and that involve, that involve synthetic data and human data effort a lot. And not just surprisingly, most of the LLMs are very bad at recognizing, like the beat, tone, and the details of the of music. They can, they can give some general prediction of which song is this, but it's very hard to describe the details of the music. like we mentioned in image generation, like you have to describe image as detailed as possible so that someone blind can reconstruct that. So here is like someoneVibhu [00:44:32]: DeafEthan [00:44:32]: someone deaf can reconstruct how the music sounds like without actually listening to it. Maybe you can think of it need to have the-- or they call the script.Vibhu [00:44:49]: Subtitles, yeah.Ethan [00:44:49]: You gotta have all the details of the music, and the dialogue.Vibhu [00:44:55]: So is the challenge there typically stuff like music and audio, or is it just Like is there a baseline? Okay, there's enough data where we can understand, narration, conversation, but there's nuances in audio that's where you hit all the data issues or is it just from stage zero, you just do it all right?Ethan [00:45:15]: So one important thing is like the alignment. So the model, the model has to know like the video and audio, the, uh-- it has to have a time-based alignment, like at which time step the video and the audio token correspond to each other. But we actually don't have this kind of alignment for most of the other modalities. If you think about like text and image, text and video, they are loosely aligned. So you can, you can have a description of what's going on in the video, but you don't have to exactly, You typically don't have exact description, oh, at, time step one second like what happened?Vibhu [00:46:02]: It's veryEthan [00:46:03]: At time step two second what happenedVibhu [00:46:03]: coarse. Yeah.Swyx [00:46:05]: So what was the ideal time step? You have to oblate it, and then it's like four seconds or something.Ethan [00:46:09]: So that comes down to how you design the model to, for the model to be aware of as a time, as a time modality. So the model is like a time aware. And that's something pretty unique if you think about LLMs. So if you ask LLM to complete a task, say they, uh-- you ask them and they will say, “Oh, this task will probably take twelve hours to complete,” and they come back in one hour. Say “I've already spent two days on this and I've exhausted everything.”Ethan [00:46:47]: So the LLMs them-themselves, they don't have a sense of time there.Vibhu [00:46:53]: I actually don't think that's just them not having a sense of time. I think it's somewhat based, right?Vibhu [00:46:58]: Like you tell someone, “Okay, go work on this feature. Go implement this,” there's a general understanding you would have of how long that would take without LLMs working at LLM speed, right? So you think back like two years ago, if I tell you to like build me like a new front end for latent space, have a search bar, have all this, you'll estimate that it'll take a few days, right?Vibhu [00:47:19]: So you tell an LLM, “Go build this.” It'll take me a few days. But I think it's somewhat grounded as opposed to them not having the best-- Not saying that they have a great understanding, but I think that example is like you can see where it comes from, right? You're trained on all over the text.Swyx [00:47:35]: They're, they're trying to estimate what a human would say.Vibhu [00:47:37]: because that's what the, that's what the data kind of represents. It's not themEthan [00:47:41]: It came from the corpus on the internet. People have a estimate of how much time.Vibhu [00:47:45]: And not even just in direct like training samples, right? Just your world understanding of tokens of how long stuff takes, right? Go read a book. It'll take you a while, right?Vibhu [00:47:56]: Even if you do nothing but read a book, it takes a few days. So yeah, LLM, I read it took me a few hours.Vibhu [00:48:01]: It'll take me a few hours to go through this research. But this is a tangent.Swyx [00:48:05]: Somewhat, yeah.Swyx [00:48:06]: This is a train of thought I haven't really expressed until now is, which is basically like a full world model must also be recursive, meaning that the participant in the world model must also be aware that they have a world model. which is like this whole recursive thing down the, down the line. but yes, and that the world model can be wrong and that they need to update it and blah. Yeah. We've, argued this on the, newsletter as well, that there needs to be sort of recursive or adversarial world models.World Models: Real-Time, Long-Horizon, Interactive VideoVibhu [00:48:34]: just, to ask, how do you define world model?Swyx [00:48:38]: Oh, yeah, let's go there.Ethan [00:48:40]: SoVibhu [00:48:40]: So just for context, we talked about, video generation, and then there's a-- if you say there's a distinction between world models, what's your, what's your definition? How do you see the two?Ethan [00:48:53]: So disclaimer, I'm not going to debate, what is world model. Yeah. there are many definitions, so I'll just talk about my definition. Since I came from the multi-model, multi-model domain, so mainly talking from video. So world model is like real-time interactive long horizon videos. So there are three parts. so we-- let's talk about them one by one. So the so interaction, so we just, we just look at Facebook and neural computer. So the interaction part of it, so you, world model can allow you to interact with them through keyboard, mouse, and maybe also voice. So these all is-- all is a modality. You can, you can interact with the model, and the model should respond reasonably. Second part is real time. So once you, once, say, you move your mouse, if, say, the world model generate a game, how fast can the game respond? So if you're like professional CS: GO players- -my say, oh, you have to respond- He's beginner within sub ten milliseconds or- Yeah even less. So that's not most of the- No, sixty FPS. Let's go. Oh, three hundred FPS. Oh, five hundred FPS. Wait. okay, yeah. I didn't do the math, but yeah, okay. Uh- Yeah, three hundred FPS, that's a three millisecond. So you have to respond- Oh, s**t. Okay. YeahEthan [00:50:29]: within a millisecond. Most of the video models cannot do that. Yeah. And, but if you, say, if you have a video model that is, say, like a digital human, the response time might be more generous. Maybe typically, for real-time voice interaction, it's like two hundred millisecond. So that's, that's much more generous. But even two hundred millisecond is pretty, it is pretty tricky, ‘cause remember we mentionedEthan [00:51:01]: you have this, temporal compression coming from the VAE. So if you, if you don't compress the temporal dimension, your sequence length is going to explode. So if you want to have this real-time, real-timeness in your model, you have to do is one context problem. And the third part is long horizon, ‘cause we-- if you're not going to just play with, video games just, a few seconds, most video models only a few seconds. We're going to play with minutes, hours. The model have to be able to generate long-form content.Ethan [00:51:42]: So putting these three together, it's, real-time, long horizon interactive videos. I think the final state will be, for example, like a video, a video version of Playbook, where you can, you can interact with, a neural computer. You move your mouse, and you click on the generative interface, and it will reply to you through pixels- generating in real time. But getting there, it's, it's a very long way to get there. So one of the first step, at Grok Imagine, where I led a small world model team there, was to build video extension. So, video extension- it's the first step of interactivity. Yeah. It's, it's the first step. Yeah. So it's the first step- You have it here, video editing, yeah. Yeah. Yeah. So the first step is because, this unlocks long horizon videos. Typically, for most of the video generation models, you give it a prompt or an image as an initial frame. You generate video, that's it. That's just, one time, done. And some creators would try to, use the last frame as a first frame for the second video. It can-- sometimes it works, but if you do it a few times, it says the quality would decrease. And- It doesn't have that context- Yeah over the full video, so the temporal- Yeah, exactly. Yeah, ‘cause you only gave it the last frame, of course, right? Yeah. Exactly. And- it's actually a pretty fun hack. if you've seen like- Oh, no, he's saying something better. Yeah. And for example, like Vue, I remember Vue 3 has like a second context of the last video. It is slightly better than using the last frame, but it has the same problem-- similar problem that it, the quality would decrease. if you extend a few times to, one minute, the video quality would look much worse than the first video. Second, another problem is that the model doesn't have long-range knowledge of, what's happening before. Say, if they generate some dialogue, some, two people speaking, and their voice might change, over some time, especially if the second conditioning, it does not cover the previous context. So these are the core challenges. So the Grok Imagine video extension, it has historical context of all of the previous generated videos. It can, It has, it has the context of, who is speaking and what objects have appeared and everything, having that to generate the next video. So if we naively do this, you can imagine, just, put all of the previous history video tokens into the context. The context lens will easily explode. Especially for video models, that can be like a few, a few million context, I would imagine- context lens. Yes.Yeah.Swyx [00:54:58]: Let's run with that.Ethan [00:54:59]: for example, like in Cosmos, I think just five seconds of video is like a fifty K or sixty K number of tokens. So like if you do, if you do fifty second, that's a five hundred K tokens. If you do longer than that, easily explode. This long horizon, problem was the first step we're trying to solve world model. It turns out people, yeah, people love video extension. Like a lot, a lot of the creators love using video extension to create longer form videos. This is the part I liked that you have a, you have an intermediate step toward the final goal instead of just a straight shot to the final version very much.Swyx [00:55:48]: But I can see you have a strong vision of where we want to end up.Long Context, Redundancy, and Efficient Interactive VideoVibhu [00:55:51]: Does it seem like it's an efficiency issue? okay, we're at a few million tokens context,. If you draw the parallel to language models, we had very short context, two thousand, eight thousand, then, you scale it up one million, ten million. sure, there's effective context, but at the end of the day, it's just what's it worth? sure, there's a whole training data side. In video, it might be slightly easier ‘cause we have a hundred million token video, right? Just take a movie with the full context there. Like is this efficiency from an inference standpoint that like it's expensive, but we know how to solve it? Or like why is this not the approach? So like my broader point was on your second point of world models, you say it needs to be interactive and live, right? You should be able to play a game and see the interaction live. So one thing I see with research is a lot of what you actually serve is different than what you build, right? So we talked about distillation. You train big model, you distill it, you do quantization, speculative decoding. We do all this stuff to serve it efficiently. Should we not just have a solution, like a world model that can interact well, do inference optimization, serve it, distill it secondary, so make it real time after you solve it? So like a-- another parallel is say, continual learning, right? What we need is someone to solve it and show it works inefficiently. Give it a few years, people will make it efficient. Same thing with regular attention, right? It worked. Over a few years, people have different forms of attention, and we've scaled it to be efficient at log context,? So kind of two things there, right? One is it seems like it works. You've scaled it. Can we not just scale it a lot more efficiently over time? Do we need a separate approach if this works? And same thing with interaction, right? if we can get it done, like if we can solve some way that it works, we can solve making it more efficient from an inference standpoint later.Ethan [00:57:53]: that's actually a very good point. So in videos, there's actually a lot of redundancies. So we solve a lot of the pixel redundancy from VE, but there's more redundancy in long range and long horizon videos. Say, if a character appear in the first clip and then it disappeared, it only reappear at the end of the video, you probably don't need the-- the context, like in the middle of the generation. So you only need that character, where you need. So that's why, I helped build another feature. It's a reference video.Vibhu [00:58:36]: Is it here?Swyx [00:58:36]: is it the same model release or different one?Ethan [00:58:39]: It's a different one.Ethan [00:58:41]: You probably need to search onSwyx [00:58:43]: I'll find itEthan [00:58:43]: X reference to video.Ethan [00:58:46]: So reference video allow you to like upload up to seven images as condition and generate the video. Say, if like I want-- it can, it can be characters or objects or even scenes. Say like I want, I want condition on, Sean's selfie and holding a bladeSwyx [00:59:07]: We have a dogEthan [00:59:08]: or whatever.Swyx [00:59:08]: We put the dog in the thing.Ethan [00:59:09]: you can put them there and the video models will generate the video from and copies the context over. So that can solve a lot of the problems there, like the long context problem. It doesn't need to have a very long context, but it's-- I feel like it's an intermediate solution. The modelSwyx [00:59:29]: It's cheating.Ethan [00:59:30]: the model should be able to like selectively know, where should I draw the references. So say if I want to generate a movie, I generate it autoregressive, like a ten second at a time or something. And now this character appear, I can look back to where it first appear and, bring that back. Yeah, this one, I put the references. Yeah, that's, Optimus, Einstein myself, Annie.Vibhu [01:00:02]: Oddly enough, I used Grok Search to find it, and it pulled your LinkedIn post. But yeah we found it.Ethan [01:00:08]: Interesting.Vibhu [01:00:10]: ButxAI's Underrated Work, Culture, and WatermarkingSwyx [01:00:11]: this is a problem. This is not your fault, but like XAI doesn't communicate all this work that you do very well because they just have the model release and then that's it. But actually, these details are very good.Swyx [01:00:22]: As far as I understand, everything you just described is state-art, like no one else has done it.Vibhu [01:00:30]: A lot of-- yeah, I have a lot moreSwyx [01:00:32]: And then, and then you just put this blog post with the cookies. I'm this is not enough,?Swyx [01:00:37]: but I, obviously this is like the high level numbers that people want to know. But no, okay, soVibhu [01:00:42]: And I wonder, like part of that is also some labs don't share research into what happens. And ifSwyx [01:00:50]: No, but this is literally bragging about how good they are, right?Swyx [01:00:54]: Like, why would you not say that you are capable of extending with full context? this is not a secret sauce. This is like we did the work. yeah, I don't know.Ethan [01:01:02]: different labs have slightly different communication styles.Swyx [01:01:07]: Anyway, if anyone from XAI is listening we are always happy to help you tell your story. Yeah, okay, so you did references, and I think, I think kind of the point you're, you're making is it is sort of like a kludge, right? this is-- you can do seven, but what about 100?Swyx [01:01:23]: Right? Then you need a completely different thing.Ethan [01:01:26]: So I think it's-- this is, a mechanism to, select the context from the history, and you might not put the entire history into the context. for example, there's a paper called Frame Pack, which haveEthan [01:01:41]: a heuristic that the latest history, the last one second, I put the entire history, and the history before that, I would, compress it and makes the video smaller. So they follow this pattern, this build overall pattern that the maximum sequence length is fixed. So the further you are from the current frame, you have a smaller image. So this is just a heuristic. I think it can be more automatic. The model is aware like which history part of it can be select. So this part of the research is actually being actively, worked on by a lot of people. It's also quite interesting. I feel this is actually, this part of long context is a little bit ahead of the LLM part.Ethan [01:02:31]: So for example, like in LLMs, if you-- so contexts keep growing. Let's say if you call tool and the tool call history is extremely long, that's still in context, and keep growing, keep growing. Even if you switch the topic to something else, the whole context was there. There are some agentic harnesses that help you to, say, prune the tool results and, prune Like when you, when you query a file, only show like the top 200 lines or something. Those were very heuristic-driven.Swyx [01:03:08]: For listeners, we did a write-up on the cloud code, leak where there are eight different kinds of pruning, including like you prune the tool results and all that. So you can, you can read up on that kind of thing.Ethan [01:03:17]: I think, one breakthrough in continual learning might be like a way to automatically, manage its own context.Swyx [01:03:27]: These are all heuristics, and they will be replaced by machine learning.Ethan [01:03:30]: InterestinglyVibhu [01:03:32]: TheEthan [01:03:32]: the same thing is being researched in both LLMs and video models.Vibhu [01:03:36]: The interesting thing is also like in the paper you showed, it's actually happening at the model level, right? Compared to like language models, sure, we have base attention, but we'll do our own compression, we'll do our own pruning, which is separate from model error.Vibhu [01:03:49]: Eventually, it all just boils in, hopefully.Swyx [01:03:52]: I think this is a form of like attention, but like also know sort of reasoning attention. I feel like that's different than normal attention.Swyx [01:04:03]: Does that, does that make sense?Ethan [01:04:04]: It's, it's different in the sense that attention, not to mention, set sparse attention aside,
Host E.B. Moss brought on the new GM of Alembic, Hitesh Wadhwani, during the POSSIBLE Conference in Miami, as part of a mini-series for "Insider Interviews" called “POV: Possible.” Because, as Wadhwani explains, with causal AI it's now possible for marketers to prove what actually drove business results. Wadhwani arrived at Alembic from 12 years at Google, where he helped build measurement products including Google Meridian. He explains why having more marketing data does not necessarily create more confidence — and why traditional approaches still leave CMOs and CFOs asking the same question: what actually caused the outcome? Wadhwani makes the case that LLMs were designed to predict language, not deliver the kind of precision needed for multimillion-dollar budgeting and pricing decisions. Alembic's answer is causal AI: a real-time model of the business that connects marketing channels, pricing, promotions, inventory, and more to identify not just what happened, but what caused it. He shares a standout case study involving a major airline's Olympic campaign spend. Alembic's model identified, at moment-level granularity, that the placement of the brand's logo during the medal ceremony drove more flights to Paris than any other single moment in the campaign — the kind of insight that makes causal AI feel a lot less like a buzzword and a lot more like a business tool. One of the biggest ideas in the episode is how causal AI can help close the gap between marketing and finance. Instead of separate teams using separate metrics, Alembic puts brand, performance, and measurement into one framework that gives CMOs and CFOs a shared language for decision-making. That includes getting better vision in to how to exactly measure creator and influencer marketing. Learn how Alembic's model works, and what's next for the company with NVIDIA as a compute partner. Subscribe for more Insider Interviews and share this one with the measurement skeptic on your team.
Fundamental technique lets researchers use a big, expensive “teacher” model to train a “student” model for less. The story How Distillation Makes AI Models Smaller and Cheaper first appeared on Quanta Magazine.
Kentucky Wildcask bourbon is back! Around this time last year, we tasted and reviewed the inaugural release of this product. We were impressed and excited to see what was on the horizon for the next release. As a reminder, Wildcask Bourbon is created by a group of students at the University of Kentucky who are learning about all aspects of the commercialization of bourbon. This class at UK is part of the Distillation, Wine, and Brewing Certificate offered by the Martin-Gatton College of Agriculture, Food and Environment. So how does the sophomore release of product taste? More importantly, how does it taste compared to last years offering? You'll have to listen to find out. This bourbon is definitely one of a kind in terms of how it is produced and we are excited to see what the future holds. --------------------------SocialsIG: https://www.instagram.com/themashupkyFB: https://www.facebook.com/themashupkyYouTube: https://www.youtube.com/@themashupkyJoin our community on Patreon: https://www.patreon.com/TheMashUpBourbonPodcastPartnership(s)Visit Bourbonoutfitter.com and enter code THEMASHUP for a special discount or visit bourbonoutfitter.com/THEMASHUPMusic: All the Fixings by Zachariah HickmanThank you so much for listening!
This Week In Startups is made possible by:Grasshopper Bank - https://grasshopper.bank/twistPaperOS - https://paperos.com/twistLinkedIn Jobs - https://linkedIn.com/twistPlaud - https://Plaud.ai/twistThe top 5 U.S. venture firms captured 73% of all LP commits in Q1, and three veteran VCs say the math may have officially broken. Aleph's Michael Eisenberg argues we may be witnessing the end of a 60-year run for venture capital as a craft business. Maniv's Mike Granoff and Oxcart's Larry Covert push back, arguing it's merely splitting into two asset classes: "Consensus VC" and traditional VC. Either way, the implications for founders, LPs, and the next decade of innovation are enormous.TWiST is back on the beat with a venture round table discussing investment concentration, the IPO drought, "bullshit ARR" in the AI era, AI gross margins, the U.S.-China chip war, the Iran conflict's impact on defense tech, the death of NATO and the rise of allied supply chains, why Tel Aviv's stock exchange could become the next NASDAQ, and a lightning round on each VC's favorite portfolio company. Let's go!Timestamps:0:00 Intro + sponsor reads (Grasshopper Bank, PaperOS, LinkedIn Jobs)0:58 Plaud: If your work depends on conversations — interviews, meetings, calls — you need a Plaud NotePin. You can check it out at https://Plaud.ai/twist and use code TWIST for 10% off!2:13 Introductions: Eisenberg (Aleph), Granoff (Maniv), Covert (Oxcart)3:57 The impact of rising venture capital concentration6:43 "We may be witnessing the end of venture capital"9:17 Consensus VC vs. Traditional VC10:01 LinkedIn Jobs - Hire right, the first time. Post your first job and get $100 off towards your job post at https://LinkedIn.com/twist11:48 Why mid-size firms beat the behemoths on founder access19:35 Coining "Consensus Colossal Collaborative Capital" (CCCC)20:03 AngelList's USVC retail VC fund — does it help mid-size funds?20:18 PaperOS - Whether you're raising a round, launching a fund, or managing a venture portfolio, PaperOS can unlock simplicity and scale across your empire of capital, contracts, and companies. Claim your $10,000 credit at https://paperos.com/twist23:27 Are there more breakout startups today than 5 years ago?28:36 The "bullshit ARR" problem and AI gross margins30:03 Grasshopper Bank - Time is money. Don't waste either. Go to https://grasshopper.bank/twist and get an exclusive $500 cash bonus just for opening an account.32:08 Cursor's negative gross margins and the hyperscaler funding flywheel33:06 Are we all electron constrained?35:35 Are we headed for surge pricing on compute?40:22 Will anything replace NVIDIA? NextSilicon, Hailo & Israel's chip stack43:33 Distillation, small models, and Apple's edge advantage46:38 Public trust in AI: should government mandate Waymo & FSD?1:02:58 Defense tech: Saronic, Anduril & the coming defense M&A wave1:06:35 The Iran war timeline & supply chain impact1:18:52 Lightning round: Jiga, Divergent, Volaback, Firehawk, HarbingerSubscribe to the TWiST500 newsletter: https://ticker.thisweekinstartups.comCheck out the TWIST500: https://www.twist500.comSubscribe to This Week in Startups on Apple: https://rb.gy/v19fcpFollow Alex:X: https://x.com/alexLinkedIn: https://www.linkedin.com/in/alexwilhelmFollow Jason:X: https://twitter.com/JasonLinkedIn: https://www.linkedin.com/in/jasoncalacanisCheck out all our partner offers: https://partners.launch.co/Great TWIST interviews: Will Guidara, Eoghan McCabe, Steve Huffman, Brian Chesky, Bob Moesta, Aaron Levie, Sophia Amoruso, Reid Hoffman, Frank Slootman, Billy McFarlandCheck out Jason's suite of newsletters: https://substack.com/@calacanisFollow TWiST:Twitter: https://twitter.com/TWiStartupsYouTube: https://www.youtube.com/thisweekinInstagram: https://www.instagram.com/thisweekinstartupsTikTok: https://www.tiktok.com/@thisweekinstartupsSubstack: https://twistartups.substack.com
Anthropic kündigt für Dienstag ein Financial Services Briefing an – möglicherweise wackelt der ganze Bankensektor. Amazon launcht Connect Talent als Hiring-Lösung. Google bekommt mit Preferred Sources die News-Quellen-Auswahl. Big-Tech-Earnings-Woche: Googles Earnings begeistern, vor allem die Cloud-Sparte explodiert. Amazons AWS-Geschäft beschleunigt sich solide. Apples Earnings sind ordentlich. Metas Earnings wachsen stark, doch die Aktie crasht wegen vorsichtigem Ausblick und steigenden Kosten. Microsofts Cloud wächst zwar weiter, enttäuscht aber gegenüber Google. Reddits Logged-In-User-Wachstum bremst stark ab. Robinhoods Krypto-Umsatz halbiert sich, denn Prediction Markets fressen das Trading. Spotifys Premium-Wachstum verlangsamt sich. Stargate-Projekt von OpenAI/Microsoft/SoftBank fällt heimlich zusammen – Data Center in UK, Norwegen und Texas storniert. Tencents neues Modell wurde mit Anthropic-Hilfe trainiert. Im Musk-Altman-Prozess gibt Musk zu, dass xAI Distillation von OpenAI gemacht hat. SpaceX-IPO-Filing: Nur Class-B-Holder können Musk feuern und diese Shares gehören Musk selbst. Neue S&P-500-Regeln sollen SpaceX-Aufnahme erleichtern. Unterstütze unseren Podcast und entdecke die Angebote unserer Werbepartner auf doppelgaenger.io/werbung. Vielen Dank! Philipp Glöckler und Philipp Klöckner sprechen heute über: (00:00:00) Anthropic Financial Services Briefing am Dienstag (00:04:08) Amazon: Produkt-Podcasts und Connect Talent Hiring (00:14:32) Google Preferred Sources für News-Suche (00:20:33) Google Earnings (00:33:54) Amazon Earnings (00:38:52) Apple Earnings (00:41:24) Meta Earnings (00:52:45) Stargate-Projekt fällt zusammen (00:55:12) Microsoft Earnings (00:58:09) Reddit, Robinhood, Spotify Earnings (01:06:25) Anthropic-Distillation (01:15:55) SpaceX-IPO: S&P-500-Sonderregeln und Musk-Governance (01:22:30) David Sacks demystifiziert Mythos Shownotes Anthropic kündigt Financial Services Briefing an - linkedin.com Amazon-Produkt-Podcasts: KI-Werbeformat - xcancel.com Amazon Connect Talent: Hiring-Lösung vorgestellt - youtu.be Google Preferred Sources: Eigene News-Quellen wählen - blog.google Google Q1 2026: Cloud +63%, Search +19% - cnbc.com Amazon Q1: Cloud +28%, Werbung +24% - xcancel.com Apple Q2 2026 Earnings - cnbc.com Zuckerberg: Iran-Krieg und KI-Kosten belasten Meta - wsj.com Stargate - ft.com Reddit Q1: Logged-In-User wachsen nur 7% - cnbc.com Robinhood: Krypto-Umsatz halbiert sich - cnbc.com Spotify Q1: Premium-Wachstum unter 10% - cnbc.com Tencents neues Modell mit Anthropic-Hilfe - theinformation.com Musk gibt zu: xAI hat OpenAI-Modelle distilliert - wired.com SpaceX: Neue S&P-500-Regeln könnten IPO erleichtern - marketwatch.com Reuters: Nur Musk kann Musk feuern bei SpaceX - reuters.com David Sacks demystifiziert Mythos auf X - xcancel.com
Tom Uren and Amberleigh Jack talk about the US government stepping in to fight ‘distillation attacks' by Chinese AI labs. These are methods used to steal the special sauce of frontier AI models simply by asking questions. They also discuss the wide-spread shift amongst Chinese threat actors to using botnets for all aspects of their operations. It's a problem for defenders, but also a disruption opportunity for authorities. This episode is also available on YouTube. Show notes
Happy Friday, y'all! For today's show, The Boys pair a whiskey that was a gift from Matt's brother (thanks Josh!) and a blanco that has been sitting on Drew's shelf for a long time. The Octomore 15.1 Islay Single Malt is high-proof, loaded with peat smoke (108 ppm) and tons of flavor. The 110 proof Villa Lobos Blanco is full of passion, history, and amazing flavors. The QuickSips™ are epic and the laughs are contagious. Both of these bottles are harder to find, but don't let that stop you. Grab a high proof scotch and a high proof tequila, invite your friends, listen and sip along, and Make It A Happy Friday!™
Is the era of the secret sourced bottle finally over, or are we just getting started? Today, we're digging into contract distillation. We're tracing the roots from the early days of "shhh, don't tell them it's MGP" to the massive, high-tech operations like Bardstown Bourbon Company that have turned being a NDP into a badge of honor. But as you, the consumer, gets more sophisticated, the demand for transparency is hitting an all-time high, and we're talking about why "faking the funk" with a made-up story just doesn't fly in 2026. We also tackle the elephant in the room: market oversaturation. With speculative investments pouring in and foreign tariffs squeezing exports, the landscape is shifting under our feet. It's not all bad news. We're envisioning a true whiskey renaissance as those massive inventories of aging barrels finally hit the market, likely leading to some of the most unique, high-quality releases we've ever seen. Show Notes: Historical significance and evolution of contract distillation Increased consumer interest in whiskey authenticity and sourcing Key players in the current contract distillation landscape Risks and opportunities for new distilleries entering the market Discussions on market oversaturation and speculative investments Effects of foreign tariffs on the whiskey export business Predictions for a bright future in whiskey with innovative barrel-aged flavors Importance of strong customer relationships for contract distillers Learn more about your ad choices. Visit megaphone.fm/adchoices
Welcome to episode 579 of the Perceptive Photographer. This week, we explore the unexpected connection between the distillation of alcohol and the art of photography. This idea came to me when I was thinking about a visit to a local distillery mean years ago. I was amazed how the process of removing impurities from spirits mirrors the photographic journey of refining images to their essential core. So this week I thought I would talk about the “triple distillation” mindset and how distilling your images, your intention, and your creative approach can lead to photographs that are clearer, more intentional, and truly resonate. Whether your work leans toward complexity or simplicity, I hpe that you can find someithng in this weeks episode on the value of eliminating noise/impurities from both in your frames and your mind to make more meaningful photographs.
12. AI SMUGGLING AND CIVILIAN-MILITARY FUSION. DAVID SHEDD AND JACK BURNHAM. The guests detail illicit efforts to smuggle Nvidia chips and steal American AI models through "adversarial distillation". They highlight China's strategic plan to acquire Western innovation without the investment. (12) PERSIA
In this episode, GG Hawkins speaks with editor Harrison Atkins about shaping A24's How to Make a Killing with director John Patton Ford. Atkins breaks down his path into editing, his holistic “total filmmaker” approach to storytelling, and the editorial challenges of balancing dark comedy, violence, voiceover, and audience empathy around a morally compromised protagonist. The conversation also explores the realities of studio post-production, from long edit timelines and test screenings to cutting in Adobe Premiere's Productions workflow while collaborating with a London-based post team more accustomed to Avid. In this episode, No Film School's GG Hawkins and guest Harrison Atkins discuss... How Harrison Atkins found his way into editing through directing and making his own films Why he thinks of editing as a holistic, dramaturgical part of filmmaking rather than a purely technical role Reuniting with director John Patton Ford after Emily the Criminal What drew him to the multi-tonal mix of crime, satire, dark comedy, and violence in How to Make a Killing How voiceover created both opportunity and endless editorial possibilities in the cut The difference between an indie sprint like Emily the Criminal and the extended timeline of a studio feature How test screenings and audience response helped refine comedy, pacing, and emotional momentum Why the first reel was crucial to getting audiences aligned with a charismatic but morally gray lead The editorial challenge of shaping an underdog around Glenn Powell's natural confidence and charm How Premiere's Productions workflow supported a collaborative feature edit with multiple people working simultaneously What it was like cutting the film in London with assistant editors adapting from an Avid-heavy post environment How temporary VFX comps in After Effects and Photoshop helped solve story and joke-building problems inside the edit Harrison's philosophy of leadership, collaboration, intuition, and staying present as both an editor and director His advice to emerging filmmakers: fail boldly, work small if necessary, and keep making things instead of waiting for permission Memorable Quotes: “I never really considered myself an editor. I still kind of weirdly don't.” (01:19) “The calendar is really a myth.” (06:59) “The difference between a joke that lands and one that doesn't is often microscopic.” (13:30) “Perfection is the enemy of good.” (33:50) Guests: Harrison Atkins Resources: How to Make a Killing Emily the Criminal Total Filmmaker by Jerry Lewis Find No Film School everywhere: On the Web: No Film School Facebook: No Film School on Facebook Twitter: No Film School on Twitter YouTube: No Film School on YouTube Instagram: No Film School on Instagram