POPULARITY
Categories
What becomes possible when enterprise computer vision no longer depends on expensive GPU infrastructure? In this episode of Tech Talks Daily, I speak with Glenn Jocher, founder and CEO of Ultralytics, about YOLO26, CPU inference, edge AI, open vocabulary vision, deployment economics, and the practical work required to move computer vision from a promising pilot into production. Glenn's route into AI began inside the U.S. intelligence community. He worked with the National Geospatial Intelligence Agency and Defense Intelligence Agency on particle physics applications, attempting to detect and track antineutrinos. Antineutrinos are extraordinarily difficult to detect because they pass through almost everything. Glenn describes them as the perfect spy. While searching for better detection methods, he discovered that computer vision researchers were solving similar problems with images. His original attempt to transfer those techniques into particle physics did not succeed. However, the work introduced him to a field where the technology could create a visible effect on everyday life. That led him toward open source development and eventually the YOLO models for object detection, classification, segmentation, and tracking. Glenn believes computer vision research has historically placed too much attention on small gains in accuracy while overlooking deployment economics. A model can perform impressively inside a laboratory and still remain unsuitable for a factory, warehouse, store, vehicle, drone, or medical environment. Price, latency, power consumption, data privacy, and deployment speed can determine whether the technology is commercially useful. This led Glenn and Ultralytics toward smaller models capable of running close to where images and video are generated. YOLO26 continues that approach with architectural changes designed specifically for CPU inference. Glenn says the model can process camera streams in real time at 30 frames per second and run across Intel CPUs, AMD CPUs, and lower power devices such as Raspberry Pi computers. This matters because specialist GPUs can increase the equipment cost and power requirements of a computer vision project. Running inference on existing CPUs or edge hardware can make deployment economically possible across larger numbers of cameras and locations. The scale already involved is difficult to comprehend. Glenn says Ultralytics models now process approximately three billion inference jobs each day, equivalent to around 30,000 every second. These jobs include images, videos, and collections of images being analyzed to detect, segment, or track objects. He attributes the platform's maturity to thousands of mistakes and bugs corrected through a rapid feedback cycle. New models are released, users report problems and request features, and the team incorporates that information into later versions. We also discuss the respective roles of cloud and edge infrastructure. Glenn sees cloud platforms continuing to provide the computing power required for training, while computer vision inference often belongs at the edge. Local processing can reduce latency, control operating costs, and keep sensitive video or medical information closer to where it was created. The smallest YOLO model is approximately three megabytes, according to Glenn. That allows it to reach mobile phones, vehicles, drones, battery powered devices, and other environments where a large language model would be impractical. Open vocabulary vision provides another development. Traditional object detection models are trained to recognize a fixed collection of objects. If a model learns to detect dogs and the user later wants it to detect cats, retraining can cause it to forget earlier knowledge unless both categories appear in the new training data. Glenn explains how promptable models can identify common everyday objects from text or visual instructions without additional training. A user could request a person wearing a blue shirt and white shoes, for example, and the system could search an image for that description. That flexibility could benefit businesses whose requirements change regularly. It reduces the need to create and label a new data set every time the company wants the model to recognize another common object. The range of current applications is already extensive. Glenn describes YOLO being used across robotics, parking, industrial safety, PPE detection, warehouses, aviation, security, traffic management, food quality, and manufacturing. Some of his favorite examples involve environmental problems. One company uses YOLO with underwater vehicles to identify and recover plastic from the ocean. Other applications detect smoke and fire early enough to support forest fire response. For leaders considering computer vision, Glenn recommends beginning with a defined problem and measurable outcome. A manufacturing company may want to reduce defects, but it still needs labeled examples showing the model what acceptable and defective products look like. He advises testing the idea through a limited pilot, measuring the return, and expanding only when the evidence supports further investment. Computer vision has become easier to deploy, but practical problems involving data, cameras, integration, reliability, and operating conditions still separate a demonstration from a production system. Could CPU inference and open vocabulary models make computer vision practical for processes your organization previously considered too expensive? Listen to the episode and share your thoughts with me. Useful Links Ultralytics website Ultralytics Platform
This Week in Machine Learning & Artificial Intelligence (AI) Podcast
For more than a decade, AI has advanced by training ever-larger models on ever-larger datasets. But as high-quality training data becomes harder to find and pretraining grows increasingly expensive, researchers are looking for new ways to keep foundation models improving. In this episode, Damian Borth, professor of AI and machine learning at the University of St. Gallen, argues we've been overlooking an important source of knowledge: the models we've already trained. His group's work on weight space learning treats trained neural networks themselves as data, learning from the distilled results of millions of GPU hours of optimization rather than starting from raw data each time. We explore what it means to build foundation models of neural networks, how knowledge can be transferred across architectures and domains, why this approach could dramatically reduce the cost of developing specialized models, and whether future AI systems may be trained on collections of existing models instead of ever-growing datasets.
If anyone builds superintelligent AI before we know how to control it, everyone dies. Nate Soares wrote the book on why that's not a metaphor. Subscribe if you want science with evidence, not speculation. Soares runs the Machine Intelligence Research Institute and co-wrote If Anyone Builds It, Everyone Dies with Eliezer Yudkowsky. The first word in that title is if. That matters. His argument is not that doom is certain. His argument is that the path we are on leads there, that the driver is asleep at the wheel, and that we still have time to wake him up. We argue for over an hour. I push on whether LLMs can ever reach superintelligence, whether GPU lock-in is a real ceiling, and what it would actually take to move his p-doom. He pushes back with one clean point: by the time an AI can rediscover general relativity from pre-1911 data the way Einstein did, we will have almost no time left. You don't wait for that goalpost. What you'll hear: Why the bus-racing-toward-a-cliff analogy depends entirely on whether the driver is asleep or awake Whether LLM lock-in is a prison or a temporary inefficiency What the AI that broke out of its virtual machine to solve a hacking problem tells us Why GPT-4o encouraging a teenager toward suicide is not a malice problem but a training problem The difference between an AI doing the right thing too well and an AI that never wanted to do what you asked What Soares actually thinks about aliens, Dyson spheres, and why we should not see stars going out The first word in the title is if. The second word to watch is would. CHAPTERS 00:00 The people racing to build superhuman AI say it might kill everyone 00:42 Who coined "AI alignment" and why the first word in the title matters 02:28 Is it already too late for the if? 04:40 The bus, the cliff, and the sleeping driver 05:02 Silicon Valley is spooked. Washington is not. 07:02 Align with who? The rogue actor problem 07:34 Who is holding the leash? 08:24 The AI that edits its own test and deletes the log file 10:04 Controllability vs. making an AI that actually cares 10:44 The move gets harder. The outcome gets easier. 13:04 Are GPUs and LLMs a ceiling or a temporary inefficiency? 16:56 Brian's Einstein test: can an LLM rediscover general relativity? 18:38 Waiting for the goalpost is waiting too long 20:14 How prediction training can push AI beyond humans 21:44 Tycho Brahe, Kepler, and planetary motion as a prediction problem 24:38 Yann LeCun said never. GPT-4 did it half a generation later. 27:28 Can you prove a no-go theorem for superintelligence? 29:14 Training a human takes a light bulb. Training an AI takes a city. 33:28 What would proof of alien life do to p-doom? 35:00 Why interstellar aliens should have Dyson spheres 44:26 What would actually update Soares' p-doom? 49:42 Nobody intended this. Intent doesn't matter. 51:08 The AI hides its tracks before it does what you want 51:34 Sycophancy vs. hallucination: which runs deeper? 51:56 Leaded gasoline and civilizational risk 59:48 Sam Harris: humans have no free will but AIs do 01:00:38 Is alignment really a governance problem? 01:01:48 Unaligned AI vs. AI aligned to the wrong person 01:04:20 2026: 10 to 30% chance of automated AI research this year 01:06:44 What if Soares is wrong? 01:09:18 What gets him out of bed 01:12:38 Watch my conversation with Roman Yampolskiy Get the transcript, fascinating bonus content, and my Monday M.A.G.I.C. Message: https://briankeating.com/yt All my top AI episodes in one place: https://briankeating.com/ai Have a .edu email and live in the USA? You automatically win a meteorite: https://BrianKeating.com/edu Subscribe: https://www.youtube.com/DrBrianKeating?sub_confirmation=1 Support Into the Impossible on Patreon, get my weekly M.A.G.I.C. Message, unfiltered bonus content, and live monthly Office Hours with me: https://www.patreon.com/drbriankeating Join this channel for perks, monthly Office Hours, and your name in the Member Roster at the end of every episode: https://www.youtube.com/channel/UCmXH_moPhfkqCk6S3b9RWuw/join Featured Guest: Nate Soares / MIRI: https://intelligence.org If Anyone Builds It, Everyone Dies (book): https://ifanyonebuildsit.com/ Nate Soares on Twitter/X: https://x.com/So8res?lang=en My books: Losing the Nobel Prize (memoir): http://amzn.to/2sa5UpA Think Like a Nobel Prize Winner: https://a.co/d/03ezQFu Focus Like a Nobel Prize Winner: https://a.co/d/hi50U9U Galileo's Dialogue (first-ever audiobook): https://a.co/d/iZPi9Un Twitter/X: https://x.com/BrianKeating Substack: https://briankeating.substack.com Blog: https://briankeating.com/blog Audio-only: https://briankeating.com/podcast #intotheimpossible #briankeating #AIrisk #aisafety #artificialintelligence #superintelligence #NateSoares #MIRI #podcast Learn more about your ad choices. Visit megaphone.fm/adchoices
Datacenters staan volop in het nieuws vanwege hun energie- en waterverbruik, maar klopt het beeld dat de algemene media schetsen? In deze aflevering van Techzine Talks gaan we dieper in op de feiten, de context en de nuance die vaak ontbreekt. Van het aandeel van datacenters in het nationale energieverbruik tot de rol van hernieuwbare energie: het echte verhaal is genuanceerder dan koppen als 'datacenters blokkeren woningbouw' doen vermoeden.We bespreken hoe datacenters in Nederland gemiddeld meer groene energie gebruiken dan huishoudens, waarom vergelijkingen met huishoudelijk energieverbruik misleidend zijn, en wat de werkelijke cijfers van het CBS en de Dutch Data Center Association zeggen. Ook de watercomponent komt uitgebreid aan bod: van open koelingsystemen tot gesloten kringloopkoeling, en waarom de chipproductie-industrie al decennia lang veel meer gefilterd water verbruikt dan datacenters.Daarnaast gaan we in op de geopolitieke waarde van Amsterdam als datacenterregio, de noodzaak van transparantie en wetgeving rond rapportage, het verschil tussen hyperscalers, co-locaties en neoclouds, en de vraag of de samenleving een eerlijk debat voert over de kosten en baten van digitale infrastructuur.• Datacenters verbruiken 5% van alle energie in Nederland, waterverbruik is slechts 0,088%• Veel Nederlandse datacenters draaien al voor een deel op groene energie• Eén Microsoft-datacenter neemt 1% van het landelijk energieverbruik voor zijn rekening• PUE als metric heeft beperkingen: closed-loop koeling is beter voor waterverbruik, maar slechter voor PUE• Amsterdam verliest langzaam terrein als datacenterregio aan Lissabon en de Nordics• Transparantie en verplichte rapportage ontbreken nog grotendeels in de sector• Neoclouds winnen terrein als alternatief voor hyperscalers bij GPU-capaciteit op aanvraag
In this episode of the Crazy Wisdom Podcast, host Stewart Alsop speaks with Aaron Neyer, founder of Parachute and community organizer in Boulder, about knowledge management, extended minds, and the intersection of AI with human consciousness. They explore how Parachute functions as a digital brain tool for organizing thoughts and information across fragmented systems, discuss the dangers of AI psychosis and over-reliance on technology, and debate open source AI development versus controlled releases by companies like Anthropic. The conversation weaves through topics including the limitations of metrics-driven business thinking, consciousness and relevance realization, the value of technological sabbaths, and Aaron's hope for locally-run open source models that protect personal data while still accessing more powerful gated models when needed. You can find Aaron's writing at unforced.org and unforced.substack.com, and learn more about Parachute at parachute.computer and parachute.computer/blog.Timestamps00:00 Stewart welcomes Aaron Neyer, founder of Parachute and Boulder community organizer, discussing the origin of Parachute's name from Frank Zappa's quote about open minds.05:00 Aaron explains Parachute as an extended mind tool for organizing notes, contacts and information across multiple platforms, emphasizing the distinction between primary mind and extended mind as interconnected systems.10:00 Discussion shifts to metrics-driven business culture and the limitations of pure rationality, exploring how Google's data-driven approach misses subjective experience and the whole picture of relationships.15:00 Aaron discusses AI's ability to help identify relevant variables across different domains and the dangers of AI psychosis, comparing it to cult dynamics and belief systems.20:00 The conversation covers AI sabbaths and nineties retreats as intentional breaks from technology, plus Aaron's experiences with electrical engineering and using AI to design circuits with Arduinos.25:00 Exploring forbidden knowledge and open source AI, Aaron discusses Anthropic's guardrails around powerful models while arguing for distributed access to prevent concentration of power.30:00 Deep dive into open source AI strategy, with Aaron highlighting NVIDIA's approach and the potential for running capable models locally while reserving ultra-intelligent models for complex research tasks.35:00 Aaron shares his vision for local Sonnet-class models handling personal data while accessing Fable-class models for deep research, and directs listeners to unforced.org and parachute.computer for his writing.Key Insights1. The philosophy behind Parachute stems from Frank Zappa's quote that the mind is like a parachute and doesn't work if it isn't open. Aaron Neyer explains that having an open mind is valuable, but it must be balanced with deep roots to avoid becoming untethered. He has experienced periods in his life where excessive openness led him to feel disconnected, teaching him that creativity and expansion need to be grounded in something substantial. This same principle applies to how we organize information digitally, where openness and interoperability allow our extended minds to become more connected and coherent, which in turn helps our primary minds think more clearly.2. Parachute is designed as an extended mind tool that addresses the fragmentation problem in how we currently manage information. Most people use multiple disconnected tools like Obsidian, Notion, Apple Notes, Google Keep, and various CRMs to organize their thoughts, notes, and relationships. These systems don't communicate well with each other, creating inefficiency and confusion. Parachute aims to create a simple, intuitive system where all this information can be organized in one place with true interoperability, allowing users to own their data and have it speak effectively with other tools, ultimately making our entire extended mind more functional.3. Understanding ourselves as unified body mind organisms rather than fragmented parts is essential for effectiveness. Living systems theory shows that any living system is three things: a membrane bound dissipative structure, a self regulating autopoietic network, and a cognitive process actively knowing the world. Western civilization since Descartes and Galileo has created artificial separation between body and mind, and between subjective and objective experience, which limits our effectiveness. The same fragmentation affects our digital technology, and recognizing both our biological and digital systems as coherent wholes rather than disconnected parts makes us vastly more capable.4. The relationship between data driven approaches and holistic thinking reveals important limitations in modern business and science. While working at Google, Aaron observed how data driven decision making can be powerful, but over reliance on metrics like ROI creates blindness to crucial unmeasurable factors like goodwill and relationship quality. This reflects a broader Western tendency to exclude subjective experience because it's difficult for objective science to measure. However, emotions, relationships, and other subjective elements are essential parts of reality, and focusing only on quantifiable metrics means missing the whole picture and ultimately becoming less effective despite appearing more rational.5. AI accelerates the ability to work with technical complexity by helping with relevance realization across domains where we lack expertise. In any specialized field, experts develop intuitive senses for which variables matter and can quickly identify problems, whether in computer troubleshooting, music, or cooking. AI's ability to generalize allows it to point people toward relevant solutions in areas where they haven't developed that intuitive expertise, effectively democratizing technical capability. This means people can direct their creativity more effectively across more domains, though it also raises concerns about giving powerful capabilities to those who may lack the wisdom to use them responsibly.6. The question of open source AI versus gated access involves complex tradeoffs between democratizing power and preventing harm. Aaron respects Anthropic's approach of creating guardrails around powerful models like Mythos, which would likely have caused significant system hacks if released without restrictions. However, this creates concerning power dynamics where only wealthy companies, governments, and their allies have access to the most powerful tools. NVIDIA offers hope through their truly open source approach including full training pipelines, and there may be a viable path where open source models at the Sonnet capability level handle most tasks locally while more powerful Fable class models remain gated for the most demanding work.7. Creating intentional breaks from AI and technology is essential for maintaining clear independent thinking. Aaron practices an AI Sabbath at least one day per week when he doesn't interact with AI, and he finds these are the days when he does his best thinking and journaling. Without these breaks, he finds himself constantly jumping between journaling and prompting AI rather than giving himself space for deep reflection. This pattern mirrors broader concerns about AI consistency creating cult like dynamics similar to organized religion, where constant immersion in a belief system or technology can lead to losing the ability to think independently, making periodic disconnection crucial for maintaining cognitive autonomy and clarity.
When you see 30 cars launch off a cliff to beautiful destruction you have to talk about it. After that we discuss Samsung and the slump in stock price after the latest profit beat. Where is memory and GPU going? What craziness are we in now? We smoke the Rojas Unfinished Business and drink the Remus Lou Gehrig 0367 of 9665. https://www.msn.com/en-us/video/peopleandplaces/watch-this-car-fly-off-a-cliff-into-a-huge-crowd-of-people/vi-AA1ZfoGd?ocid=socialshare
JDK 27 entered Ramp-Down-Phase 1 in early June, so we can go over its features. Valhalla's first JDK Enhancement Proposal (401) will likely be part of JDK 28. Show off your Java ideas in the "Modern Java in the Wild" hackathon. There are two interesting articles for those who want to dive deeper into AOT with agents or into using GPU tensor cores from Java. Links: JDK 27 JEP list: https://openjdk.org/projects/jdk/27/ Valhalla mail: https://mail.openjdk.org/archives/list/jdk-dev@openjdk.org/threadAIA3O3LHFZ6T7TIPH7KZT4WS4B6U72U5/ JEP 401 PR: https://github.com/openjdk/jdk/pull/31120 "Java in the Wild" hackathon: https://www.hackster.io/contests/modern-java-in-the-wild AOT vs agents: https://bugs.openjdk.org/browse/JDK-8387190 AOT and JVMTI agents: https://bugs.openjdk.org/browse/JDK-8387194 "Exploiting GPU Tensor Cores": https://openjdk.org/projects/babylon/articles/hat-tensors/hat-tensors
嘉宾 | 张从志,《三联生活周刊》主笔主播 | 高一丁,《Talk三联》编辑芯片无处不在。从小区门口的自动抬杆,到你手中的智能手机,再到驱动AI大模型的超级算力中心,这颗几厘米见方的硅片,已成为现代生活的基础元件,也成了大国博弈的关键筹码。但芯片到底是怎么造出来的?从实验室到成品,这中间经历了怎样的漫长旅程?又是怎样的压力,让设计芯片的工程师们要靠“玄学祈福”?光刻机之外,还有哪些“卡脖子”的环节不为人知?这个领域内应验数十年的摩尔定律,又因何走向终结?本期节目,我们和记者张从志一起,走进芯片产业的真实世界——从晶圆厂严苛到“牛屁都能影响良率”的洁净车间,到封装厂千级无尘室里的“无米之炊”;从英伟达如何从游戏显卡厂商蜕变为AI时代新贵,到中芯国际等国产企业的追赶之路。这是一场需要耐心、资本和代际传承的漫长征程。喧嚣之下,更需要长期主义和具体的人的付出,让我们一起探寻那颗芯片背后的故事。【时间轴】01:36 芯片为何变得如此重要?08:08 大热的GPU,如何从游戏显卡变成算力基石?15:19 一枚芯片的诞生,要分哪几步?22:12 为什么芯片工厂不能和奶牛场做邻居?29:37 “流片”的压力有多大,要去烧香拜拜?34:03 被称作“芯片之母”的EDA究竟是什么?37:06 “摩尔定律”真的要失效了吗?47:21 国产芯片行业是如何起步的?55:47 先进封装是国产芯片“弯道超车”的机会吗、64:44 追上“英伟达”们,还要花多久?编辑/一丁剪辑/译丹————“Talk三联”是《三联生活周刊》出品的一档软硬皆有的泛文化类音频栏目,用声音记录报道背后的故事,提供丰富的新知与思辨的可能。在以下渠道均可收听我们的节目:三联中读APP |小宇宙|喜马拉雅|苹果播客|网易云音乐【我们还有这些播客】苗师傅·天真与经验|中场时间|我有一个朋友·董晨宇|多一种生活|岁时茶山记|你好,陌生人HelloStranger|孩子,你的情绪我们在乎|“三联·大案追踪”有声剧如果你喜欢我们的节目,欢迎点赞支持,或者把我们的节目推荐给更多的朋友~ 【关注我们】App:三联中读微信公众号:三联中读微博:@三联中读小红书:三联中读官网:https://www.lifeweek.com.cn/ 【商务合作】zhongdu@lifeweek.com.cn
DigiPower X (DGXX) CEO Michel Amar breaks down the company's milestone contract with Cerebras (CBRS)… its AI stack roadmap, from real estate to GPU-as-a-service… key catalysts through 2026… and why DigiPower X is in a league of its own. In this episode: Welcome back, Michel Amar, CEO of DigiPower X [0:01] DigiPower's $1 billion contract with Cerebras is a major milestone [0:53] Why DGXX is exempt from New York's moratorium on data centers [6:33] From real estate to GPU-as-a-service: DigiPower's full AI stack roadmap [10:51] Management's plans to keep executing through 2026 and beyond [19:49] DigiPower X is in a league of its own [24:43] Did you like this episode? Get more Wall Street Unplugged FREE each week in your inbox. Sign up here: https://curzio.me/syn_wsu Find Wall Street Unplugged podcast… --Curzio Research App: https://curzio.me/syn_app --iTunes: https://curzio.me/syn_wsu_i --Stitcher: https://curzio.me/syn_wsu_s --Website: https://curzio.me/syn_wsu_cat Follow Frank… X: https://curzio.me/syn_twt Facebook: https://curzio.me/syn_fb LinkedIn: https://curzio.me/syn_li
Episode 107: We chat about some of the latest CPU releases, including the Ryzen 7 5800X3D 10th Anniversary Edition, and the Ryzen 7 7700X3D, which doesn't make a ton of sense when nobody is doing platform upgrades. Also, we thought it was quite amusing that Nvidia's hotspot GPU temperature has now been uncovered, so we discuss that whole situation.CHAPTERS00:00 - Intro01:02 - The Ryzen 7 7700X3D has Launched08:32 - The Ryzen 7 5800X3D Actually Makes Sense24:53 - Nvidia GPU Hotspot Temperature Situation40:24 - Steam Machine Thermals and HDMI-CEC01:04:55 - Updates From Our Boring LivesSUBSCRIBE TO THE PODCASTAudio: https://shows.acast.com/the-hardware-unboxed-podcastVideo: https://www.youtube.com/channel/UCqT8Vb3jweH6_tj2SarErfwSUPPORT US DIRECTLYPatreon: https://www.patreon.com/hardwareunboxedLINKSYouTube: https://www.youtube.com/@Hardwareunboxed/Twitter: https://twitter.com/HardwareUnboxedBluesky: https://bsky.app/profile/hardwareunboxed.bsky.social Hosted on Acast. See acast.com/privacy for more information.
This week we let ChatGPT help pick the news stories (which seemed like a good idea at the time), and honestly, it kind of worked.We start with Russia reportedly having to buy gasoline after Ukrainian drone strikes hit their refinery capacity, which turns into a bigger conversation about energy markets, gasoline and diesel, and how cheap drones are completely changing the math of modern warfare.Then we get into Japan potentially pushing its giant pension funds to bring more money back home, the AI trade still running through SK Hynix and GPU demand, and the Fed now treating AI infrastructure spending as one of the things keeping inflation hotter than expected.So there is real market stuff in here.But also, this is Dan and me, so we also talk about a squirrel getting loose in a Meta office, a bear breaking into a shopping mall on a military base, Britain trying to survive heat waves with better refrigerators and tiny fans, Pepsi making a very questionable Wild Cherry post in India, AI reading ancient Roman scrolls, Pompeii, Pink Floyd, and whether Wendy's Biggie Bag is secretly one of the better recession trades on the board. Chapters00:00 — Hook, Intro, and Cold Open01:47 — Welcome Back to the China Shop02:30 — Letting ChatGPT Pick the News03:22 — Russia's Fuel Problem and Ukrainian Drone Strikes06:36 — EVs, Refinery Risk, and Energy Independence09:10 — A Squirrel Gets Loose at Meta14:28 — Japan's Pension Fund and Money Coming Home20:08 — The Alaska Mall Bear Story21:54 — SK Hynix, AI Chips, and Inflation Pressure25:24 — FedWatch, Rate Expectations, and No Cuts Yet27:38 — Britain's Heat Wave Economy32:49 — Scottsdale Will Pay for Your Lawn, Not Your Pool34:29 — Pepsi's Wild Cherry Problem38:33 — AI, Ancient Scrolls, and the Pompeii Detour42:28 — Wendy's, Biggie Bags, and Takeover Rumors48:38 — Jobs, Rent, and the Economy People Actually Feel50:30 — Chessferatu and Closing ThoughtsSubscribe, share, and join the trading conversations on Facebook, Twitter, LinkedIn and Discord!Sponsors and FriendsOur podcast is sponsored by Sue Maki at Fairway Independent Mortgage (MLS# 206048). Licensed in 38 states, if you need anything mortgage-related, reach out to her at SMaki@fairwaymc.com or give her a call at (520) 977-7904. Tell her 2 Bulls sent you to get the best rates available!If you are interested in signing up with TRADEPRO Academy, you can use our affiliate link here. We receive compensation for any purchases made when using this link, so it's a great way to support the show and learn at the same time! **Use code CHINASHOP15 to save 15%**To contact us, you can email us directly at bandoftraderspodcast@gmail.comIf you like our show, please let us know by rating and subscribing on your platform of choice!If you like our show and hate social media, then please tell all your friends!If you have no friends and hate social media and you just want to give us money for advertising to help you find more friends, then you can donate to support the show here!Dan:Dan co-founded 2 Bulls in a China Shop with Kyle when their shared passion for active trading ignited during the lockdowns. Their daily discussions about trades, interests, and the valuable lessons learned created the bedrock for what eventually evolved into both the 2 Bulls in a China Shop and Band of Traders podcasts.While navigating the complexities of trading, Dan infused humor into the shows with his self-deprecating wit and candid discussions about their trading experiences. This dynamic duo's chemistry became the catalyst for a podcast that resonated widely, capturing the attention of a diverse audience.Download ChessFeratuHalf-Cocked TalesAdvertising Inquiries: https://redcircle.com/brandsPrivacy & Opt-Out: https://redcircle.com/privacy
As AI applications become more complex, the infrastructure powering them needs to evolve. Corey Sanders, SVP of Product at CoreWeave, joins Chris to discuss why AI requires a fundamentally different approach than traditional cloud computing. They explore AI-native infrastructure, training and inference workloads, the rise of agentic development, optimizing GPU performance, AI research workflows, and why the future of software will be built around AI-first experiences rather than websites and apps.Featuring:Corey Sanders – LinkedIn Chris Benson – Website, LinkedIn, Bluesky, GitHub, XLinks:CoreWeaveSponsors:Framer: The enterprise-grade website builder that lets your team ship faster. Get 30% off at framer.com/practicalaiUpcoming Events: Register for upcoming webinars here!Midwest AI Summit 2026
Credit markets are stepping in to fund the surging demand for AI. Our experts Lindsay Tyler and Anish Shah explore the opportunities and risks behind this record financing wave.Read more insights from Morgan Stanley.----- Transcript -----Lindsay Tyler: Welcome to Thoughts on the Market. I'm Lindsay Tyler, TMT Credit Research Analyst at Morgan Stanley. Anish Shah: And I'm Anish Shah, Global Head of Debt Capital Markets at Morgan Stanley. Lindsay Tyler: Today, how issuers and investors are approaching the rapidly evolving world of AI financing. It's Thursday, July 16th at 10am in New York. As AI demand accelerates, credit markets are being asked to finance infrastructure on a scale that used to be associated with utilities, telecom, or energy. That raises a central question for issuers and investors: How much debt can the AI ecosystem absorb? And at what price? Anish, can you walk our listeners through the key products in your purview? Anish Shah: Certainly, in my nearly twenty years at Morgan Stanley, this is probably the most incredible time period I've ever seen in the credit markets. I've had the privilege of working across a number of different roles in capital markets and lending. And a couple of years ago, we integrated the debt underwriting business across both investment-grade and leverage finance franchises in recognition of how interconnected the whole credit ecosystem has become. In addition to our core activities helping clients raise capital for their strategic priorities, two of the big focus areas that we've had have been finding ways to harness the power of the private credit universe and also delivering best-in-class capabilities in funding this incredible growth in AI spend. Lindsay Tyler: AI financing has certainly been a theme we've also been focused on in research. Our equity research colleagues project that a handful of key players could add more than 30 gigawatts of capacity over a two-year timeframe, driving around [$]2 trillion of aggregate cash CapEx in that period. And to put that into context, a single gigawatt of data center capacity can require roughly $12 billion for the shell, and then often more than double that for chips and racks. So, from your vantage point, what inning are we in? And what gives you confidence that credit markets can continue funding this opportunity at scale? Anish Shah: I mean, Lindsay, the numbers certainly are staggering, as you note. And if you just observe the CapEx estimates for the hyperscalers and broadly for AI infrastructure, we're certainly in the early innings. Lindsay Tyler: Mm-hmm. Anish Shah: The largest tech companies have historically, as you know, raised very little debt. In fact, many of these companies have not even needed a credit facility. As CapEx projections were materially increased in the second half of last year, we saw the beginning of scaled capital raises. Hyperscaler issuance has quickly gone from less than one percent of the investment-grade market to more than 10 percent of the market. You know, as I look ahead, based on what we're seeing on the ground, we think that AI-related funding, whether it's for data center development or financing compute capacity, could top 15 percent of the total issuance across all credit products. This has been an unprecedented test for the capital markets, both in terms of the depth of capacity and the breadth of product. The teams have been on the forefront of deep investor dialogue and product innovation. This spans corporate investment grade, first of their kind financings in high-yield and leveraged loan markets, and new takes on asset-backed financing. And each of these areas has seen material issuance both in public and private markets. Lindsay Tyler: Great backdrop. Let's dig first into investment-grade corporate debt, an area you know well from your time previously leading the investment-grade team. Can you help frame the scale and the significance of this financing bucket and how AI-related debt is scaling within it? Anish Shah: Well, you know, as you know, the investment-grade bond market, specifically in dollars, is the deepest, most liquid pool of capital in the world. Volumes have grown materially over the last few years and are likely to eclipse $2 trillion in issuance this year. Hyperscalers are among the very best credits in the world, and they have the ability to come in and out of markets with relatively quick twitch, little to no pre-marketing, and in fairly large size. You know, $20 billion-plus deals used to be rare in the investment-grade market, now happen multiple times a quarter. This is why we've seen the predominance of AI-driven capital raising take place in the investment-grade market. For the most part, investors have digested that supply very well. While we've seen some modest widening credit spreads for hyperscalers and some of the other tech issuers, I'd say it's de minimis relative to their expected ROI. Lindsay, I've talked a lot about supply dynamics and issuance. What other factors are you and investors considering when assessing fair value for investment-grade rated technology bonds? Lindsay Tyler: Sure. It's prudent to really weigh a mix of technicals, fundamentals, and relative value. You know, as you discussed on the technical side, and related to my discussions with debt and equity investors, I've been focused on scale of buildouts, market capacity, digestibility across currencies, positioning along the curve, implications of equity issuance, and whether AI financing could crowd out other areas of TMT credit. But moving more to the fundamental side of things, you mentioned ROI, and for the players that are scaling compute capacity, there are a handful of key monetization and return questions that keep coming up. How quickly can these companies bring new capacity online? Once it's live, how does it translate into durable revenue and cash flow? Is that capacity supporting internal products, proprietary models, broader cloud offerings, or compute leased to third parties? And then how fungible is the capacity across those use cases if demand or returns shift? Further on the fundamental side, we've done some differentiated work around growing long-term commitments. We've seen that high-quality hyperscalers and a few of the semis companies are anchoring the AI ecosystem through leases, guarantees, other obligations. These commitments really extend beyond vanilla bond issuance. So, I encourage investors to look beyond the funded debt and really understand the accounting and the ratings implications here of some of those commitments. And this ties nicely into the next topic that I wanted to raise, which is project finance debt. I've noticed that, you know, a lot of the commitments that we're seeing from IG players support another layer of financing. Lease commitments can underpin project finance debt, an area of sizable issuance and innovation. The public high-yield market has emerged as a new funding source in this way for data center construction, with more than 30 billion priced across 15 deals, since fall 2025. Can you walk us through, Anish, the innovation behind these structures, and how are these high yield deals different than other ways to, kind of, raise project finance debt? Anish Shah: Yeah, it's incredibly interesting. I mean, the bulk of the issuance, as I noted has come in the investment grade market, but I would say the bulk of the innovation has come in the sub-investment grade market. You know, historically, for very capital-intensive sectors like energy and power or real estate, the project loan market was the most efficient source of initial funding. The developer would tap banks to underwrite a highly structured construction loan. Once the project is up and running, you could then refinance that loan with the predictable cash flows into a more institutional financing, like the investment grade bond market or the term loan B or securitization markets.That product may still be very viable in many sectors, but we felt early on that bank-provided construction loans would not meet the capacity needs of the AI investment cycle. The market really needed an institutional credit product that bypassed the need for construction loans. The key innovation came in the form of first-of-its-kind high-yield bonds that funded the development of a new data center complex. Given the relatively short construction period and the "offtake" supported by some of the highest quality credits in the world, we felt like this financing structure would be incredibly well-received in the high-yield market. The win here is that the developer accesses fixed rate long-term capital and maintains flexibility to call the bonds and refinance at a lower cost. Judging by how these financings have gone, there's a strong level of investor enthusiasm. I think that they've only scratched the surface, and I would expect that we see much more of this. And potentially even expand it to other products in the leverage finance markets given the tremendous level of investor demand. Lindsay Tyler: Yeah. It's certainly been exciting to follow many of those deals. Beyond the public space, we're also seeing a wave of innovation in private credit and asset-backed finance. Anish, how do companies decide whether capital is best raised in the public or the private markets? Anish Shah: Well, I'm glad you raised the whole avenue of private markets because it may be the most significant change in the credit markets over the last few years, broadening the scope of private credit from directly lending into leverage buyouts to now financing large investment-grade projects. There are great examples in the world of GPU and TPU financing, where we structure loans secured by the asset and the cash flows, or in data center development.Lindsay, from your perspective, what are investors focused on when these private structures intersect with public credits? Lindsay Tyler: Sure. Many of these asset-backed private financings have prompted investors to look more closely at any of the public companies involved, whether as issuers, tenants, customers, or support providers. This ties back to the point I raised earlier. Where does the risk reside, and who ultimately is on the hook? These financings have also sparked broader discussions around circularity, vendor financing, and technology obsolescence risk, even when amortizing structures are in place. I do think those are fair concerns to weigh, and they really speak to how quickly the AI financing trend is evolving and how much credit work there is to do. So, Anish, with that balance in mind, relatively strong demand, rapid innovation, but also some real credit questions, let's end with a quick lightning round. Anish Shah: Lindsay, let's do it. Lindsay Tyler: First, what is the biggest risk that could test investor appetite for AI-related debt? Anish Shah: I would say investors are acutely focused on construction delays. Don't underestimate the level of diligence being done by the breadth of capacity you're seeing in the markets. Investors are doing their homework, and we're spending a lot of time trying to mitigate any of their concerns with structural protections. Lindsay Tyler: Got it. Second, beyond data center shells and chips, what is the next potential AI financing opportunity? Anish Shah: It most certainly is energy and power. We're going to see a ton of capital being raised in utilities. It's going to be a little different than what the hyperscalers are doing, just given the nature of their balance sheets. You're going to see more junior capital. We've seen a wave of junior subordinated debt issuance out of the utilities. We're also seeing a lot of activity from our project finance and tax equity team, just given all things energy infrastructure. Lindsay Tyler: Great. And third, if we're sitting here a year from now, what do you think could be the biggest AI financing story we're talking about? Anish Shah: Well, we certainly underestimated the level of financing activity that we saw in the past year. I think when we look back a year from now, we will probably see that the AI labs were much more ready to finance on their own on a standalone basis. That's going to alleviate some of the pressures in the market, but I think it's going to create a whole new set of considerations and structural innovation. Lindsay Tyler: Well, it's certainly been remarkable to watch this financing theme take shape in real time, and the next chapter sounds like it could be even more interesting to follow. Anish, thanks for joining us and sharing your insights. Anish Shah: Great to join, Lindsay. Thanks. Lindsay Tyler: And thank you for listening. If you enjoy Thoughts on the Market, please leave us a review wherever you listen and share the podcast with a friend or colleague today.*****Anish Shah is a member of Morgan Stanley's Global Capital Markets Division and is not a member of Morgan Stanley's Research Department. Unless otherwise indicated, his views are his own and may differ from the views of the Morgan Stanley Research Department and from the views of others within Morgan Stanley.
You've probably never heard of Inkling. It's the newest (and first) model from Thinking Machines Labs, and it could very well be a small snowball that picks up major momentum in today's enterprise AI landscape. If you haven't heard of Thinking Machines, they're led by Mira Murati, the former CTO at OpenAI. The big bet with Inkling? The future of AI could be using smaller models fine-tuned and optimized for smaller tasks. Will it work? Tune in live as we dive in. The Most Important AI Model You'll Probably Never Use That Just Dropped -- An Everyday AI Chat With Jordan WilsonNewsletter: Sign up for our free daily newsletterMore on this Episode: Episode PageToday's Episode on LinkedIn: Thoughts on this? Join the convo on LinkedIn and connect with other AI leaders.Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineupWebsite: YourEverydayAI.comEmail The Show: info@youreverydayai.comConnect with Jordan on LinkedInTopics Covered in This Episode:Inkling AI Model Launch OverviewThinking Machines Lab Leadership HighlightInkling's Multimodal and Agentic CapabilitiesOpen Source vs. Proprietary AI ModelsEnterprise Procurement with American AI ModelsAI Fine Tuning as a Service (Tinker)Benchmark Scores: Inkling vs. Frontier ModelsCustomization and Model Shopping for EnterprisesAI Token Costs Driving Model EfficiencyBridgewater Case Study: AI Model CustomizationFrontier Models Enabling Efficient Fine-TuningFuture Trends: Specialized Small Language ModelsTimestamps:00:00 Inkling: A new AI model release05:43 Inkling AI model details09:08 China's dominance in open source AI11:48 Launch and model updates discussed15:21 Concerns over using Chinese open-source models19:06 Training smaller AI models20:22 Using GPT for AI Model Training23:54 Predicting Rise of Small Language Models28:38 Choosing the right AI modelKeywords: Inkling, Thinking Machines Lab, Meera Muradi, former OpenAI CTO, open source AI model, American AI model, fine tuning as a service, enterprise AI, multimodal AI, agentic models, customizable AI, Tinker, enterprise distribution, model procurement, Chinese open source models, strategic reset, model overhang, capabilities gap, AI model shopping, model routing, cost-conscious enterprises, artificial intelligence index, 975 billion parameter model, text-image-audio AI, open weights, proprietary AI models, customization accessibility, small language models, AI workflows, context window, Bridgewater use case, model distillation, GPU infrastructure, API costs, token efficiency, fine-tuned models, post training, AI competitive leverage, recurring financial judgment, AI benchmarks, middle tier models, automated model evaluation, privacy and workflow mapping, economical AI models, model rental, model routing automation.Send Everyday AI and Jordan a text message. (We can't reply back unless you leave contact info) Ready for ROI on GenAI? Go to youreverydayai.com/partner
Imagine a dark warehouse. Racks and racks of devices with wires, tubes, and electronics sticking out. The next AI data center? No. This is Lila Sciences‘ dream for the future of science. A dark warehouse full of AI-guided robotics and lab equipment, cranking out new experiments 24/7, building toward a scientific superintelligence.Their automated lab is almost hypnotizing to watch. They have floating plates zipping around on Wall-E-esque tracks, used vision-language models to control Windows 95 boxes, and created the world's largest collection of voided warranties. In the process they've built a massive library of scientific reasoning tokens. Over 10 trillion of them, all experimentally validated.No warranties were voided in the making of this videoTo say Lila is ambitious is an understatement. Their goal is a scientific superintelligence wired directly into the wet lab. They are all in on the bitter lesson, and the thesis follows from it: a lab is an infinite token generator. Produce data at scale, and the synergies give you a general reasoner that can tackle any scientific problem. They are committing hard. Biology, chemistry, drug discovery, and materials science, all at the same time. Time will tell if it works, but it is an exciting hypothesis.In our latest episode we sat down with Lila's very own Andy Beam (CTO) and Rafa Gómez-Bombarelli (CSO, physical sciences) and went on a journey through the possibilities of AI-run science, almost as wide-ranging as Lila's goals.Did we mention they do both materials science and biology? In the same AI science factory? Same time, same lab, same AI. Finally a guest who can settle a long-running debate we've had amongst ourselves: is biology or materials science harder?Watch to find out!We discuss:* The internet is spent, science is next. Why Lila thinks the scientific method is the last untapped internet-scale dataset, and why they treat RL as a data generation mechanism with nature as the verifier.* The lab as a data center. Instruments as nodes on a graph, a magnetically levitating “PCI bus” transport layer between them, orchestration as a slurm queue. Andy is not short on analogies.* Why Lila insists it is not an automation company. They optimize for flexibility and generalizability over raw throughput, which means humans stay below the API line wherever automating does not pay.* Your experiment has a runtime. We put Escalante Bio's question to Andy: if science is the token generator, what is the runtime of your data collection? His answer, in short, is that you cannot make the ribosome go faster. Why Lila bets on fast round-over-round iteration rather than big noisy multiplexed screens, and how Rafa's team rebuilt a gas sorption measurement to run roughly 2,500x faster.* What is actually in 10 trillion scientific tokens. Not sequences. Experimentally verified reasoning traces, a kind of data that Andy argues exists on the internet in quantities that round to zero.* Breadth as a path to depth. Small molecule chemistry priors transferring to metal organic frameworks for carbon capture, and the claim that the general model beats domain-specific models sample for sample.* If you have the data, what do you need the model for? Sri Kosuri's koan about the ML-for-drug-discovery business model, and Andy's answer: the coding model got better because it also read Shakespeare and carnitas recipes.* The serendipity they want to automate. Emily Whitehead survived the first pediatric CAR-T cure only because the doctor treating her happened to know, from pediatric arthritis, which antibody would blunt her IL-6 response. Roll that dice again and you probably lose her. Breadth is how you stop depending on luck.* Move 37 for catalysts. Model suggestions for platinum-group-free electrocatalysts that went from boring, to what a 40-paper expert called stupid, to the best performers they have made.* Six months to in vivo CAR-T data in non-human primates, and the zero-FTE virtual startup commercial model that fell out of it. For context on why that number is startling, AbbVie paid $2.1B for Capstan on the strength of preclinical in vivo CAR-T data.* You cannot have scientific superintelligence if you are just a good test taker. Ken Stanley, who wrote Why Greatness Cannot Be Planned, runs open-endedness at Lila. RL at scale gives you a ruthlessly Vulcan problem solver. Machine creativity is a different thing, and it is the part nobody has solved.* The chain of thought is an unreliable narrator. The model reasons in latent space and only emits tokens. Sometimes it skips the experiment entirely and is still right. So how much do you trust the reasoning versus the verifier?* Reward hacking when the rollout is physical. Chains of thought that collapse into repetition, and a model that got annoyed and swore at the scientist who kept asking it to redo a plate map. What happens when a pathological loop has a wet lab inside it?* The bittersweet lesson. Rafa's inversion of the bitter lesson: in AI, scaling is a roadmap. In materials, scaling is a filter, because only the things that scale end up mattering.* Not your typical Flagship company. Why a famously single-asset biotech incubator spun out a platform bet, and Andy's line that if Lila called itself a biopharma it would have a top-three GPU cluster.* Bottlenecks they would remove by fiat. Sim-to-real for physics-based simulation, and the fact that RL training runs at roughly 5% mean FLOP utilization.Watch on YouTube: This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe
Alex Thorn talks with Lucas Tcheyan (Galaxy Research) about compute, AI, and GPU financial markets. Alex also talks with Beimnet Abebe (Galaxy Trading) about CPI, rates, equities, and bitcoin. Participants, along with Galaxy Digital, hold a financial interest in Bitcoin (BTC). Galaxy regularly engages in buying and selling BTC, including hedging transactions, for its own proprietary accounts and on behalf of its counterparties. Galaxy also provides services to vehicles that invest in BTC. If the value of such assets increases, those vehicles may benefit, and Galaxy's service fees may increase accordingly. The valuation in this communication is based on technical, fundamental, and market analysis and not on any formal valuation method. For more information, please refer to Galaxy's public filings and statements. Cryptocurrencies, including BTC, are inherently volatile and risky and ultimate market movements may not align with this statement. For additional risks related to digital assets, please refer to the risk factors contained in filings Galaxy Digital Inc. makes with the Securities and Exchange Commission (the “SEC”) from time to time, including its Quarterly Report on Form 10-Q, available at www.sec.gov. This episode was recorded on Wednesday, July 15, 2026. ++ Follow us on Twitter, @glxyresearch, and read our research at www.galaxy.com/research/ to learn more! This podcast, and the information contained herein, has been provided to you by Galaxy Digital Holdings LP and its affiliates (“Galaxy Digital”) solely for informational purposes. View the full disclaimer at www.galaxy.com/disclaimer-galaxy-brains-podcast/
Richard and Brian are back for this week's episode of Macro Aggressions. This episode breaks down the "Russian doll" problem sitting inside mega-cap tech earnings, why portfolio diversification may matter more now than it has in fifteen years, and what's happening beneath the surface of an S&P 500 that keeps hitting new highs. Richard Taylor of Plan First Wealth and Brian Dunhill of Dunhill Financial unpack Burry's concerns around private company valuations (SpaceX, Anthropic, OpenAI) sitting inside public company earnings, changes to GPU depreciation accounting that are quietly inflating profits, and why small cap stocks, emerging markets, and international stocks are starting to outperform after over a decade of US large-cap dominance. This is practical stock market advice for anyone wondering if their portfolio is over-concentrated in seven companies and whether now is the moment to start rebalancing. They also cover the diverging picture between the stock market and the real economy: sticky 4.2% inflation, weakening wage growth, and job losses under the current administration, set against a market still riding high on AI enthusiasm and a growing conversation around a potential market bubble. The conversation turns geopolitical, covering Europe's active effort to decouple from American tech infrastructure, why universities across Europe are pushing to get off US servers, and what that could mean long term for US-Europe relations and international wealth strategies. Richard and Brian also dig into the UK's ongoing political instability, the lasting economic impact of Brexit, and whether a new Labour leadership shift could change the UK's trajectory. Whether you're watching the Magnificent Seven dominate your portfolio, thinking about how UK politics and Brexit affect cross-border wealth, or just want a grounded read on where markets stand versus the economy, this episode covers the full picture, not just the headlines. -- Expat Wealth is supported by Plan First Wealth. Plan First Wealth is a Registered Investment Advisor serving fellow expatriates and immigrants living across the US on matters such as retirement planning, investment management, tax planning and non-US asset management. https://planfirstwealth.com/ -- Expat Wealth is affiliated with Plan First Wealth LLC, an SEC registered investment advisor. The views and opinions expressed in this program are those of the speakers and do not necessarily reflect the views or positions of Plan First Wealth. Information presented is for educational purposes only and does not intend to make an offer or solicitation for the sale or purchase of any specific securities, investments, or investment strategies. Investments involve risk and unless otherwise stated, are not guaranteed. Be sure to first consult with a qualified financial adviser and/or tax professional before implementing any strategy discussed herein. Plan First Wealth does not provide any tax and/or legal advice and strongly recommends that listeners seek their own advice in these areas.
As part of our summer replay series, we're revisiting one of our favorite conversations on the future of AI infrastructure. SemiAnalysis founder Dylan Patel joins Erin Price-Wright, Guido Appenzeller, and Erik Torenberg to examine the rapidly evolving economics of AI hardware, from GPUs and custom silicon to data centers, power, and the global race for compute. The conversation explores NVIDIA's competitive advantages, the rise of custom chips from Google, Amazon, and Meta, the economics of frontier AI models, and the infrastructure constraints shaping the industry's next phase. They also discuss AI startups, export controls, robotics, enterprise software, and why simply copying NVIDIA isn't enough to build a winning AI hardware company. Whether you're building AI products, investing in infrastructure, or trying to understand where the industry is headed, this conversation offers a practical look at the forces shaping the future of compute. Resources: Follow Dylan Patel on X: https://x.com/dylan522p Follow Erin Price-Wright on X: https://x.com/espricewright Follow Guido Appenzeller on X: https://x.com/appenz Learn more about SemiAnalysis: https://semianalysis.com/dylan-patel/ Stay Updated:Find a16z on YouTube: YouTubeFind a16z on XFind a16z on LinkedInListen to the a16z Show on SpotifyListen to the a16z Show on Apple PodcastsFollow our host: https://twitter.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Small teams with smart AI can solve problems no one imagined before. In this episode, I spoke with Zhen Lu, CEO of Runpod, about the rapidly evolving AI landscape and the future of software development. Zhen Lu shared how Runpod is empowering engineers to train AI models tailored to specific business needs, improve efficiency, and reduce waste. We also discussed his founder journey, the power of co-founder alignment, global innovation outside traditional tech hubs, and the importance of human accountability in AI-driven organizations. Here are the highlights: ● Runpod builds AI developer infrastructure. The platform supports tailored AI workloads, enabling businesses to efficiently train models specific to their needs. ● AI is changing how software is built. Engineers are challenged to rethink software design, not just accelerate existing processes, creating new opportunities for innovation. ● Efficiency and sustainability matter. Fine-tuning smaller, focused models reduces resource waste, energy consumption, and operational costs compared to massive off-the-shelf models. ● Co-founder alignment drives success. Implicit trust, complementary skills, and low ego between Zhen and his co-founder have accelerated decision-making and execution. ● Global networks and bootstrapping foster innovation. Scarcity encourages creativity, and building strong relationships outside Silicon Valley has enabled Runpod to grow and support AI developers worldwide. About the guest: Zhen Lu and co-founder Pardeep Singh started by running crypto mining rigs out of their New Jersey basements. When Ethereum's "The Merge" threatened to make that obsolete, they pivoted, converting the rigs into AI servers. As corporate developers at Comcast building ML projects, they saw firsthand that the GPU developer experience was, in Zhen's words, "just hot garbage." That insight became Runpod. Launched in early 2022, Runpod offers fast, developer-friendly GPU infrastructure: clean APIs, CLI tools, serverless options, and easy configuration. Rather than the traditional VC route, they debuted on Reddit with "Hey Reddit, give us your worst" — and got "please take my money" in return. They went on to earn validation from Hugging Face co-founder Julien Chaumond and a seed round led by Dell Technologies Capital, hitting $120M ARR, 10 billion serverless requests, and serves ~900,000 developers across 31 global regions. Connect with Zhen Lu: LinkedIn: https://www.linkedin.com/in/zeen/ Website: https://www.runpod.io/ Connect with Allison: Feedspot has named Disruptive CEO Nation as one of the Top 25 CEO Podcasts on the web. LinkedIn: https://www.linkedin.com/in/allisonsummerschicago/ Website: https://www.disruptiveceonation.com/ #CEO #leadership #startup #founder #business #businesspodcast Learn more about your ad choices. Visit megaphone.fm/adchoices
CNBC reported that Apple is in talks with a startup that specializes in compressing AI models to run on iPhones. The move aligns with Apple's strategy to execute more AI locally, reducing reliance on cloud infrastructure and lowering latency. On-device AI relies on techniques such as quantization, pruning, and distillation to fit models within CPU, GPU, and Neural Engine limits. The shift could reduce cloud inference costs that depend on GPUs from providers like Amazon Web Services, Microsoft Azure, and Google Cloud. Competitors including Google, Samsung, Qualcomm, and Meta are advancing on-device AI capabilities. Founders should benchmark compact models, assess battery impact, and decide which features should run locally versus in the cloud.Learn more on this news by visiting us at: https://greyjournal.net/news/ Hosted on Acast. See acast.com/privacy for more information.
As part of our summer replay series, we're revisiting one of the standout conversations from Runtime, a16z's conference on AI infrastructure and the future of computing. Gavin Baker, Managing Partner and CIO of Atreides Management, joins David George to examine the biggest questions surrounding today's AI investment cycle. Is AI a bubble? What does the unprecedented buildout of data centers, GPUs, and compute infrastructure mean for the economy? And how should investors think about the companies building the next generation of AI? The conversation explores frontier models, Nvidia, Google, custom silicon, AI infrastructure, application software, robotics, and why Baker believes today's AI investment cycle looks fundamentally different from the internet bubble of the early 2000s. Along the way, they discuss the economics of GPUs, enterprise software, AI business models, and what comes next as AI moves from experimentation into the broader economy. Resources: Follow Gavin Baker on X: https://x.com/GavinSBaker Follow Atreides Management on X: https://x.com/atreidesmgmt Follow David George on X: https://x.com/DavidGeorge83 Stay Updated:Find a16z on YouTube: YouTubeFind a16z on XFind a16z on LinkedInListen to the a16z Show on SpotifyListen to the a16z Show on Apple PodcastsFollow our host: https://twitter.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
While LLMs and agents are new to appsec and everyone else, a lot of AI security requirements translate to well-known API security requirements. Jeremy Snyder helps us frame the OWASP LLM Top 10 into five layers in order to help orgs understand and prioritize their attack surface. A lot of orgs don't have to deal with model-specific threats or building their own GPU architecture, but every org adopting LLMs and agents should be aware of how those agents are being invoked and the output those agents are producing. That awareness of input and output helps in identifying and mitigating prompt injection attacks, ensuring agents are working within their expected boundaries, and taming token budgets. Resources: https://genai.owasp.org/llm-top-10/ https://github.com/rtk-ai/rtk https://docs.aws.amazon.com/bedrock/latest/userguide/prompt-caching.html https://www.firetail.ai/blog/beyond-the-spectacle-rsac-2026-and-the-5-layers-of-ai-security Visit https://www.securityweekly.com/asw for all the latest episodes! Show Notes: https://securityweekly.com/asw-391
On this episode of The Crypto Rundown, host Mark Longo welcomes back Bill Ulivieri from Cenacle Capital Management to break down a wild, volatile few weeks in the digital asset space. Bill calls his official generational bottom for Bitcoin, leaning into quantitative power laws, internal AI metrics, and cyclic low models that point to a major buying window between now and November. The guys also dive deep into the collapsing NAV discounts on MicroStrategy (MSTR), the rising options volume across IBIT and ETHA, and the latest dramatic re-shuffling in the Altcoin Universe.
While LLMs and agents are new to appsec and everyone else, a lot of AI security requirements translate to well-known API security requirements. Jeremy Snyder helps us frame the OWASP LLM Top 10 into five layers in order to help orgs understand and prioritize their attack surface. A lot of orgs don't have to deal with model-specific threats or building their own GPU architecture, but every org adopting LLMs and agents should be aware of how those agents are being invoked and the output those agents are producing. That awareness of input and output helps in identifying and mitigating prompt injection attacks, ensuring agents are working within their expected boundaries, and taming token budgets. Resources: https://genai.owasp.org/llm-top-10/ https://github.com/rtk-ai/rtk https://docs.aws.amazon.com/bedrock/latest/userguide/prompt-caching.html https://www.firetail.ai/blog/beyond-the-spectacle-rsac-2026-and-the-5-layers-of-ai-security Show Notes: https://securityweekly.com/asw-391
While LLMs and agents are new to appsec and everyone else, a lot of AI security requirements translate to well-known API security requirements. Jeremy Snyder helps us frame the OWASP LLM Top 10 into five layers in order to help orgs understand and prioritize their attack surface. A lot of orgs don't have to deal with model-specific threats or building their own GPU architecture, but every org adopting LLMs and agents should be aware of how those agents are being invoked and the output those agents are producing. That awareness of input and output helps in identifying and mitigating prompt injection attacks, ensuring agents are working within their expected boundaries, and taming token budgets. Resources: https://genai.owasp.org/llm-top-10/ https://github.com/rtk-ai/rtk https://docs.aws.amazon.com/bedrock/latest/userguide/prompt-caching.html https://www.firetail.ai/blog/beyond-the-spectacle-rsac-2026-and-the-5-layers-of-ai-security Visit https://www.securityweekly.com/asw for all the latest episodes! Show Notes: https://securityweekly.com/asw-391
While LLMs and agents are new to appsec and everyone else, a lot of AI security requirements translate to well-known API security requirements. Jeremy Snyder helps us frame the OWASP LLM Top 10 into five layers in order to help orgs understand and prioritize their attack surface. A lot of orgs don't have to deal with model-specific threats or building their own GPU architecture, but every org adopting LLMs and agents should be aware of how those agents are being invoked and the output those agents are producing. That awareness of input and output helps in identifying and mitigating prompt injection attacks, ensuring agents are working within their expected boundaries, and taming token budgets. Resources: https://genai.owasp.org/llm-top-10/ https://github.com/rtk-ai/rtk https://docs.aws.amazon.com/bedrock/latest/userguide/prompt-caching.html https://www.firetail.ai/blog/beyond-the-spectacle-rsac-2026-and-the-5-layers-of-ai-security Show Notes: https://securityweekly.com/asw-391
Anthropic had a really bad week.
AWS Morning Brief for the week of July 13th with Corey Quinn. Links:AWS Security Hub extends unified security management to Microsoft AzureAmazon EKS Auto Mode reduces GPU management fees by up to 60%Amazon RDS for Oracle now supports Oracle Database 26aiAWS Builder Center Now Offers Free Sandbox EnvironmentsAWS Security Hub now offers Network Scanning to identify publicly reachable resourcesAmazon Cognito now supports self-service provisioned API rate limitsAWS Security Hub adds impact analysis for exposure findingsBlazing a Trail: How Peloton Rebuilt the SDLC for the Agentic Era with Amazon BedrockHow United Airlines solved IP exhaustion with Private NAT GatewayBuilding secure AI agents at scale: Introducing Loom for AWSWhat does it cost to answer one question? Measuring per-request cost in agentic workloadsDesigning for the inevitable: System prompt leakage and mitigations in generative AI applicationsThe CISO's guide to post-quantum mandates and migrationsTwo CVEs on the theme of guarding keys badly
В мире пользовательских интерфейсов каждый год происходят эволюции, а иногда и революции. Терминалы тем временем уже десятки лет как будто бы застыли во времени. Концептуально мало что поменялось, но баззвордов стало много: bash, zsh, tty, GPU acceleration в конце концов. А с приходом AI выросла популярность CLI-тулов, благодаря чему терминалы переживают ренессанс. Так что же там происходит? Разбираемся в этом выпуске вместе с Ником Скрябиным! Также ждем вас, ваши лайки, репосты и комменты в мессенджерах и соцсетях! Telegram-чат: https://t.me/podlodka Telegram-канал: https://t.me/podlodkanews Twitter-аккаунт: https://twitter.com/PodcastPodlodka Ведущие в выпуске: Женя Кателла, Андрей Смирнов Полезные ссылки: Attyx – терминал с аппаратным ускорением, встроенными сессиями и программным управлением https://attyx.sh https://github.com/semos-labs/attyx Xyron – шелл на стероидах. Структурированные команды, управление проектами, встроенные альтернативы популярных утилит (fzf, zoxide, etc.) https://github.com/semos-labs/xyron Как работает GPU рендер в терминале https://semos.sh/blog/attyx-gpu-rendering/ Что под капотом реакт-фреймворка для терминалов https://semos.sh/blog/how-glyph-works/
Episode 106: We chat about Sony ending production of physical games in 2028, and the implications that has for game ownership. Plus we talk about the challenges Nvidia faces with improving GPU performance, and the future of gaming as a result.CHAPTERS00:00 - Intro00:30 - Life Update02:35 - A chat about recent content06:00 - Playstation ending disc support32:06 - Has GPU performance really hit a ceiling?44:00 - The future of gaming?53:34 - Is Radeon in the same boat as Nvidia?01:01:03 - Why games aren't getting better?01:19:36 - Our Boring LivesSUBSCRIBE TO THE PODCASTAudio: https://shows.acast.com/the-hardware-unboxed-podcastVideo: https://www.youtube.com/channel/UCqT8Vb3jweH6_tj2SarErfwSUPPORT US DIRECTLYPatreon: https://www.patreon.com/hardwareunboxedLINKSYouTube: https://www.youtube.com/@Hardwareunboxed/Twitter: https://twitter.com/HardwareUnboxedBluesky: https://bsky.app/profile/hardwareunboxed.bsky.social Hosted on Acast. See acast.com/privacy for more information.
In this episode, Ray Cochrane breaks down AI distillation, the teacher-student technique frontier labs now lean on to train smaller, cheaper models. He also covers GPT-5.6’s government-vetted rollout, Claude Sonnet 5 landing on AWS, Maryland’s two-year data center pause, and Microsoft’s climbing carbon numbers. Finally, he wraps with Apple’s $30 billion Broadcom deal, Meta’s tamper-proof recording light, Michigan’s parasite outbreak, and a simulation that erased a super El Niño. – Want to start a podcast? Its easy to get started! Sign-up at Blubrry – Thinking of buying a Starlink? Use my link to support the show. Subscribe to the Newsletter. Email Ray if you want to get in touch! Like and Follow Geek News Central’s Facebook Page. Support my Show Sponsor: Best Godaddy Promo Codes Get 1Password Full Summary Cochrane opens with a quick personal update. Longer days have him outdoors, including a float trip on the Sandy River at Dabney State Park, where he found clearer water, clay-like sand, and easy footing. Next week brings both a move and a trip home, so he is stocking up on Trader Joe’s “Power Berries” and IKEA bags at his mom’s request. Then he turns to the lead story. AI Distillation Explained: How Frontier Models Teach Each Other Cochrane’s featured story comes from Hugging Face engineer Sergio Paniego. Distillation is teacher-student training for AI: a capable model generates the training signal, and a smaller student learns to match it. The classic off-policy version compresses giant models into cheap students, either through soft labels or piles of worked answers. Google’s Gemma models and DeepSeek’s R1-Distill line were built exactly this way. However, the industry is now converging on multi-teacher on-policy distillation, or MOPD. Labs build reinforcement-learning specialists for math, coding, and agentic work, then have them grade a single student, word by word, as the student generates its own answers. DeepSeek-V4, MiMo-V2-Flash, and NVIDIA’s Nemotron 3 Ultra all run versions of the recipe, and the Qwen3 team reported better results at roughly a tenth of the GPU hours of raw reinforcement learning. Finally, self-distillation lets models like Cursor’s Composer 2.5 learn from better-prompted versions of themselves. Sponsor: GoDaddy Economy hosting $6.99/month, WordPress hosting $12.99/month, domains $11.99. Website builder trial available. Use codes at geeknewscentral.com/godaddy to support the show. GPT-5.6 Arrives With a Government-Vetted Rollout OpenAI shipped GPT-5.6 as a three-tier family: Sol, Terra, and Luna. Sol costs five dollars in and thirty dollars out per million tokens, half of Claude Fable 5’s rate. The benchmarks split: Sol Ultra wins Terminal-Bench at 91.9 percent, while Claude Fable 5 still leads SWE-Bench Pro. Notably, the API launched in limited preview to roughly 20 partners vetted by the U.S. government, though the model went live in Microsoft 365 Copilot on day one. Claude Sonnet 5 Lands on AWS, Plus Quick AWS Wins Claude Sonnet 5 arrived on AWS through Bedrock, pitched as top-tier intelligence at Sonnet pricing. Additionally, Amazon WorkSpaces for AI agents reached general availability, enabling agents to drive full desktop applications securely. OpenSearch gained a log-analytics engine claiming four times the price-performance, and SageMaker now scales inference about twice as fast. Cochrane also flags that Kendra and Q Business move to maintenance mode at the end of July. Anthropic Wants You to Reflect on Your Claude Habits Anthropic launched Reflect, a beta feature that analyzes your past Claude conversations and visualizes how you actually use the assistant. It requires Memory, excludes incognito and health-related chats, and keeps its insights inside the tool. Cochrane loves the idea. He reviews his own transcripts to extract prompt patterns and turn them into reusable skills, and he suggests listeners simply ask their AI to do the same. AlphaEvolve Goes GA on Google Cloud Google made AlphaEvolve generally available to Google Cloud customers on the Gemini Enterprise Agent Platform. The agent acts as an evolutionary collaborator: provide a baseline algorithm and your goals, and it searches for better, human-readable code. BASF, JetBrains, and Kinaxis are the named early adopters. Meanwhile, Cochrane renews his standing wish that DeepMind release AlphaGo as a playable teacher. Google Adds “How This Ad Was Made” AI Labels Google is adding a “How this ad was made” section to My Ad Center across Search, YouTube, and Discover. Ads built with Google’s own AI tools automatically get the disclosure, backed by invisible watermarks. However, ads made with outside tools rely on advertiser self-declaration. Cochrane points out the limits of voluntary disclosure in an AI-flooded content economy. Microsoft’s Carbon Emissions Climb 25 Percent Microsoft’s new sustainability report shows emissions up 25% in 2025, driven by a data center construction spree. The gross figure is 34 million metric tons before offsets, while other coverage puts the net figure at around 20 million. Water consumption also jumped thirty-four percent, even as Microsoft claims its first water-positive year. Cochrane argues regulation needs to catch up, since Google and Amazon report similar increases. Prince George’s County Pauses Data Centers for Two Years Prince George’s County adopted a two-year moratorium on new data center development, the longest pause in Maryland so far. The resolution blocks new applications, including hyperscale projects, until the council passes real regulations. Water and energy impacts remain open questions the county intends to study. Cochrane gives kudos to residents for making their voices heard. Apple and Broadcom Ink a $30 Billion U.S. Chip Deal Apple is expanding its partnership with Broadcom with a multiyear agreement expected to exceed $30 billion. The deal covers custom silicon and wireless components, with more than fifteen billion chips to be made on American soil. Broadcom’s Fort Collins, Colorado plant anchors the work with a $1.5 billion equipment expansion. Tim Cook framed the deal as accelerating Apple’s commitment to American manufacturing. MSI and Intel Ship the First Arc G3 Extreme Handheld Intel detailed how it co-engineered the MSI Claw 8 EX AI+, the first handheld on the Arc G3 Extreme processor. Highlights include a heat-spreading board layout and game-tuning loops that Intel says run Cyberpunk 2077 up to thirty-seven percent faster. The device is on sale now in void purple for around $1,500. At that price, Cochrane jokes he would rather buy a computer. Meta’s Glasses Get a Tamper-Proof Recording Light Meta answered the most common privacy questions about its AI glasses. Photos stay private on the device until the wearer imports or shares them, and a white capture LED blinks during any recording with no off switch. Moreover, newer glasses disable the camera if the LED is blocked, tampered with, or destroyed. Cochrane reminds listeners these claims are Meta grading its own homework, but the blink signal is worth recognizing in public. Michigan’s Parasite Outbreak Tops 1,200 Cases Michigan’s cyclosporiasis outbreak reached 1,251 cases since June 22, with roughly forty hospitalizations along the way. Northwest Ohio adds more than five hundred cases. The parasite typically spreads through contaminated fresh produce, and investigators still have not found the source. Cochrane’s advice: wash your produce, and get tested if your symptoms fit. AI Finds the San Andreas Fault’s Silent Slips Researchers paired AI with borehole strainmeters to detect dozens of hidden slow-slip events beneath the San Andreas Fault’s Parkfield section. Each silent slip releases stress within hours and is reliably followed by low-frequency earthquakes. Together, the findings support a continuous spectrum from silent creep to destructive quakes. The study appears in Nature Communications, and Cochrane hopes it will lead to better earthquake prediction. Cloud Brightening Erased a Super El Niño, in a Simulation Finally, a Science Advances study simulated marine cloud brightening in response to the 1997 and 2015 super El Niño events. Seeding clouds over the eastern Pacific erased the events entirely inside the model. Real deployment would take roughly 2,400 ships spraying continuously, and the simulations showed side effects like extra warming over Europe and Asia. Cochrane finds the weather-machine concept fascinating, yet he questions the consequences of altering cycles the planet runs for a reason. The post AI Distillation: How Frontier Models Teach Each Other #1870 appeared first on Geek News Central.
Deze aflevering van Einde van de Week Live is ook te bekijken op https://youtu.be/RsJLcXIcdmc Deze talkshow wordt mede mogelijk gemaakt door MSI. Alle meningen in deze video zijn onze eigen. MSI heeft inhoudelijk geen inspraak op de content en zien de video net als jullie hier voor het eerst op de site. Zijn we allemaal klaar voor een dit keer lekker warm weekend? Of houd je niet van al te hoge temperaturen? In beide gevallen gaan we je plezieren. Want deze nieuwe episode van Einde van de Week Live kun je in de tuin kijken of lekker binnen met de gordijnen dicht. Shelly, JJ, Koos gaan er, zoals we gewoon van hen zijn, een spetterende talkshow van maken. En daar is genoeg discussie-stof voor. Zo is de hele discussie rond de ontslagen bij XBOX en het verdwijnen van de fysieke disk bij PlayStation, nog steeds niet verstomd. Ook hebben de drie het over de kritiek vanuit de gamers op Assassin’s Creed Black Flag Resynced. Dit terwijl de media en critici vrijwel unaniem positief tot zeer positief over het spel van Ubisoft zijn. Wat is er aan de hand? Dit alles en meer zie en hoor je voorbijkomen in de Einde van de Week Live van vrijdag 10 juli 2026. Discussie over verdwijnen schijfjes PlayStation wil maar niet verstommen Andere onderwerpen die aan bod komen in deze video zijn de situatie van Bethesda en hun volgende game, Elder Scrolls 6, en de vraag of gamers GTA 6 en de nieuwe console van PlayStation moeten en willen boycotten, omdat Rockstar geen schijfjes wil produceren? Pak 150 euro korting bij aankoop van de Cyborg 15 gaming laptop MSI zet deze week de Cyborg 15 B2R gaming laptop in de spotlights. Een ideaal instapmodel dat nu hier bij de MediaMarkt met 150 korting verkrijgbaar is. Aan boord bevinden zich een Intel Core 7 240H processor, een NVIDIA GeForce RTX 5060 GPU, 16GB RAM intern geheugen, 512GB SSD, een 15.6” Full HD display en een 4-zone RGB toetsenbord. Krijg 15% korting bij aankoop van de 52G930B OLED ‘beuker’ van LG UltraGear Als je het blije gezicht van JJ in de video over de 52” beuker ziet, dan weet je eigenlijk genoeg. LG heeft, zoals zo vaak, weer iets fijns in elkaar gezet. En groot. De grootste 5K2K 240hz gaming monitor zelfs. Zin in een uitspatting? Check hier alle specs. Bij aankoop krijg je 15% exclusieve Gamekings korting als je bij het afrekenen de code GameKings_2026 invoert. Deze korting geldt ook voor de 27GX790B en de 39GX950B monitor die we de afgelopen weken lieten zien. Check de gloednieuwe Philips Evnia 32M2N8900P gaming monitor Het is monitorendag. De Philips Evnia 32M2N8900P is een 31,5-inch 4K QD-OLED gaming-monitor die echt gemaakt is voor de ultieme combi voor gamers: competitief gamen en hoogwaardige beeldkwaliteit. De monitor combineert een vierde generatie QD-OLED-paneel met UltraClear 4K UHD-resolutie, een verversingssnelheid van 240 Hz en een reactietijd van 0,03 ms. Daarnaast ondersteunt het model NVIDIA G-SYNC-compatibiliteit en beschikt het over DisplayHDR True Black 500, true 10-bit color en een kleurdekking tot 99,5% DCI-P3 en 97,7% Adobe RGB. Ook biedt de monitor AI-Enhanced Ambiglow, AmbiScape via Philips Evnia Precision Center, HDMI 2.1, DisplayPort 2.1, USB-C met Power Delivery tot 65 W, een ingebouwde KVM-switch en MultiView. Interesse? De brute monitor is vanaf deze maand hier verkrijgbaar. Scoor onze nieuwe Low Key Lavender cap en support Gamekings Traditiegetrouw hebben wij ook deze zomer een nieuw, zomers stuk fashion voor je klaarstaan. Ideaal voor de zomerse warmte. Je kunt de cap hier kopen. De laatste exemplaren gaan naar alle verwachting dit weekend de deur uit. Op = op. Het geld gebruiken we om de toko overeind te houden. Dus support ons waar je kunt en je krijgt er wat cools voor terug.Wil je adverteren bij de podcast Gamekings óf misschien bij een andere podcast van ILVY Network? Mail dan naar management@ilvy.com en/of kijk even op de website : https://ilvy.com/podcastSee omnystudio.com/listener for privacy information.
We've been running a bit of an Agent Cloud series surveying all the top inference/compute/cloud providers, from Databricks to Daytona to Railway and, even further back, E2B, but we're excited to conclude this series returning to Modal, which has just raised a monster $355M Series C.The cloud was built for developers. But agents are now changing that.The old infra stack was designed for a human who could read docs, reason through YAML, and understand dashboards to figure out what they need when something broke. While this was painful for developers, it worked since they could fill in missing context in their heads.However, agents don't have that luxury. Now in this new era of agents, everything has to be tighter.They need a place to write code, run it, inspect the output, change the environment, debug failures, and try again. Fast iteration and feedback loops with all the necessary context are crucial for agents to operate properly. Furthermore, sandboxes are a clear representation of this shift as agents can easily spin up isolated environments. This programmatic infra even extends to research:Two years ago, we were one of the first to cover Modal with CEO Erik Bernhardsson and Alessio designed our favorite LS thumbnail of all time:At the time, Modal was just a teeny little company with a $17M Series A.Today, fresh off their $355M Series C, Modal is one of the clearest examples of the agent cloud future being built in real time: a cloud platform moving past traditional web app assumptions toward the workloads AI actually creates such as elastic inference, sandboxes, GPU burst, post-training, background agents, and infrastructure that agents themselves can operate.In this episode, Modal CTO Akshat Bubna joins swyx and Vibhu to unpack why AI applications don't fit traditional cloud assumptions, why Kubernetes was never designed for bursty compute-heavy workloads, and why Modal is now shifting from developer experience to agent experience.We go deep on Modal's AI infra stack: serverless functions, decorator-based infrastructure, elastic inference for custom models, GPU snapshotting, DeFlash, speculative decoding, Auto Endpoints, sandboxes, persistent storage, networked containers, private IPv6, RDMA, multi-node training, and Modal's capacity pool across 17 cloud providers. Akshat also explains why RL rollouts can require 100,000 sandboxes, why production agents need hard guardrails, why observability may matter more than reading code, and why AI has made infrastructure exciting again.We discuss:* Why Kubernetes wasn't built for bursty AI workloads* How Modal started as a better runtime before becoming an AI cloud* Why Modal added GPUs before ChatGPT* The shift from developer experience to agent experience* Why observability matters when agents are writing the code* Elastic inference for custom models across audio, video, robotics, and comp bio* GPU snapshotting, cold starts, and why inference workloads are so bursty* Why RL rollouts can require 100,000 sandboxes* DeFlash, speculative decoding, and frontier-level inference performance* Auto Endpoints and making optimized inference easier to deploy* What Modal adds beyond vLLM, SGLang, and raw GPU rental* Modal's 17-cloud capacity pool and supercloud strategy* Networked sandboxes, sidecars, private IPv6, and RDMA* Serverless multi-node training for post-training and research workloads* Auto-research, model-guided sweeps, and agents launching GPU experiments* Compute strategy, capacity planning, and batch tiers* Why production agents need specialized sandboxes and hard guardrails* Modal's take on managed agents, CI, Gitpod/Ona, Python, TypeScript, and Modal BenchAkshat Bubna* LinkedIn: https://www.linkedin.com/in/akshat-bubna-188885103* X: https://x.com/akshat_bModal* Website: https://modal.comTimestamps00:00:00 Introduction00:00:39 Modal's origin and why Kubernetes wasn't enough00:04:32 Developer Experience → Agent Experience00:06:21 Modal's AI cloud primitives00:09:14 Sandboxes, agent loops, and proto-Cognition00:12:12 Elastic inference, GPU snapshotting, and 100,000 sandboxes00:15:24 DeFlash, speculative decoding, and Auto Endpoints00:19:59 Production-grade inference beyond raw GPUs00:22:00 Background agents, Ramp Inspect, and the agent lifecycle00:24:08 Modal's 17-cloud supercloud strategy00:26:40 Networked sandboxes, private IPv6, and RDMA00:32:48 Multi-node training, post-training, and auto research00:37:36 Compute strategy, capacity planning, and batch tiers00:40:55 Open models, real-time AI, and production agent infra00:43:06 Hard guardrails, managed agents, and specialized sandboxes00:46:06 Why AI made infrastructure exciting again00:48:30 Model APIs, differentiated products, and agentic video00:51:50 CI, coding-agent infra, SDKs, and Modal Bench00:57:28 Closing ThoughtsTranscriptIntroduction: Modal, Series C, and the Art PartySwyx [00:00:00]: We're here with Akshat, CTO of Modal, together with Vibhu. Congrats on your Series C.Akshat [00:00:10]: Thank you.Swyx [00:00:11]: Your party yesterday was amazing.Akshat [00:00:15]: Yeah.Swyx [00:00:15]: From all the photos and all the swag.Akshat [00:00:17]: We had a bunch of art installations, which was fun, seeing, like, our products on pedestals next to, like, Rodin.Swyx [00:00:25]: Very nice. Very nice. When you started, it was not the GPU inference company. Maybe it was in your mind. Take us back to the origin story.Modal's Origin: A New Runtime Beyond KubernetesAkshat [00:00:39]: I first met Eric, who's the CEO, through an investor. Back then Eric was already thinking about building, a new runtime, and he got there thinking through why are workflow orchestration products so hard to use. It's because you have to run them on Kubernetes. Kubernetes is hard to manage. It's not built for burstiness and, custom images,Swyx [00:01:03]: YeahAkshat [00:01:03]: It has a terrible developer experience.Swyx [00:01:05]: And I'll, I'll interjectAkshat [00:01:06]: YeahSwyx [00:01:07]: For listeners, who are new, we interviewed Eric two years ago, and there's a bit more of the story there from Spotify and all those things.Swyx [00:01:14]: And I came across Eric through Data Council because he did that talk on the serverless container stack that you guys did, which was like, that was my first like, “Okay, I need to take Modal very seriously” moment.Akshat [00:01:26]: Yeah.Swyx [00:01:26]: But it was still very unclear, like, do I need all this for just my data pipelines?Akshat [00:01:33]: Yeah. initially what we were thinking about was if we build a better runtime, it's a very useful primitive in itself. It's There's a lot of things that, get solved by serverless functions, like you can do, ETL stuff, you can do job queues, you can do all this, like, bursty processing, which it turns out every company had needs for. but then we also were thinking about this as like, this is a primitive that we can build a whole collection of products on, which are very verticalized. So perhaps data engineering would've been the first one, but we were thinking about inference. Back then it was more classical inference, like computer vision stuff and running XGBoosts and whatnot. But we added GPUs to the product a year before ChatGPT came out.From Serverless Containers to GPU WorkloadsSwyx [00:02:19]: Nice.Akshat [00:02:19]: We just didn't think it would be that big of a deal.Swyx [00:02:22]: Yeah, just like add A100.Vibhu [00:02:23]: Was there any, like, early key problem that really sparked off why you built it?Akshat [00:02:28]: Yeah. Primarily it's just, none of the tooling that was out there was built for, one, a really great developer experience, and also there's a general trend of, a lot of the workloads that we were seeing were very. I wish there was a better word for it, but compute-heavy. Like, they need, one, like, need a lot more resources, so you need to burst up and down a lot, versus like Kubernetes designed for, like, slow scaling and, more for, like, web server use cases. And also there's just a lot more specialization in, like, what kinds of environments these workloads run in. Like, we had sometimes they need accelerators, sometimes they need different kinds of images, and this is just like a consistent thing that we saw across a lot of companies. That would be the next step.Software-Defined Infrastructure and Decorator-Based DXSwyx [00:03:13]: Yeah. Yeah. Be nice. I don't know how much this factored into the early story, but I wrote a post when I was at Temporal about infrastructure, software-defined infrastructure or something like that.Akshat [00:03:22]: Yeah, the self-provisioningSwyx [00:03:23]: Self-provisioning.Akshat [00:03:24]: Yeah.Swyx [00:03:24]: Yeah. I can't even remember my own post.Swyx [00:03:26]: And then you put me on the landing page.Akshat [00:03:28]: Yeah. We really like, the term and so we stole it.Swyx [00:03:32]: Because you had the insight that everything can just be in decorators co-located with the code, right?Akshat [00:03:37]: Yeah.Swyx [00:03:37]: Was that a big part of the originalAkshat [00:03:39]: YesSwyx [00:03:39]: Story or it was just like a DX layer?Akshat [00:03:41]: That was, really important because we really didn't want people to spend, so much time, writing YAML, and it seemed like you could really condense the surface area of what you're doing, put it in code so you can operate on it just like you operate on other code, and like build stuff that's more expressive and dynamic. and so yeah, that was always a very important part.Swyx [00:04:04]: Then the pushback is this is a DSL.Akshat [00:04:07]: Yeah.Swyx [00:04:07]: It's you're closed source. I am locked into Modal.Akshat [00:04:11]: Yeah. We never really got pushback for that because the nice thing about Modal is you can bring whatever code you have, and sure, the DSL is at the configuration layer for, what hardware you're using, how you're scaling things up, but you still own the code.Akshat [00:04:27]: And that's, that's been an important, part of our story, even as we do inference now.Swyx [00:04:32]: Yeah.Vibhu [00:04:32]: How much of do you think still stays the same today? Like if you were to build something today, DevX very important, but I feel like, a lot of this has been changed with just hook it up to an agent, have Claude Code, have Codex implement a tool. there's very agent native primitives that are different than if I'm doing this myself, right?Developer Experience → Agent ExperienceAkshat [00:04:54]: We've changed our SDK team to think about agent experience instead of, developer experience and we think that the same benefits that apply for DX also apply for AX, which is why would you have an agent read through hundreds of Kubernetes files and like write YAML that's not even typed when it can make a couple of changes in a decorator and it gets this self-provisioning runtime of, being able to see its changes live in action? yeah, it just seems from the customers we talk to, they find Modal is much faster for agents to use versus operating on a different substrate.Swyx [00:05:34]: Yeah, because like you, again, you co-locate the infrastructure requirements to the code that runs it.Akshat [00:05:38]: Yeah.Swyx [00:05:38]: Well, the negative thesis now is that nobody's looking at their code anymore, so there's no point.Akshat [00:05:44]: Yeah, people aren't looking at code. one thing we still see is really important is observability.Swyx [00:05:51]: Yeah.Akshat [00:05:51]: Like how good is your dashboard? And of course, like we have, we push a lot of it to the CLI so the agents can do their own investigation, but you still need humans to go interpret what's going on and, make judgment calls and whatnot. and that's I feel like, Maybe more important now than looking at the code itself.Swyx [00:06:11]: Yes, because like, you can try to treat the code as a black box and then use, see the observable action that comes out of it, and then just prompt a change.What Modal Is For: AI Cloud PrimitivesAkshat [00:06:21]: Yeah.Swyx [00:06:22]: So I think it takes a bit of restraint to not specialize, to say, “I want to ship a new primitive,” and then just be general purpose.Swyx [00:06:31]: People ask you, “What are you for?” You're like, “ I don't know. We can do this, we can do that.”Vibhu [00:06:36]: Well, I'd be curious to see, like, okay, if we were to ask you, like, what is Modal for even at a high level? There's a lot you guys do, sandboxes, GPUs, everything. How do you answer?Akshat [00:06:46]: Modal is a cloud platform that's built for, where we've built the primitives from scratch for AI applications. and right now it covers, inference, training, batch processing, and sandbox workloads.Akshat [00:07:00]: But we're building a lot moreSwyx [00:07:02]: I noticed you didn't say web server, so there is still a role for, like, the always-on large-scale Kubernetes type things.Akshat [00:07:09]: Yeah, absolutely. We're, we're not trying to compete with the renders of the world, because yeah, we think the differentiator for us is the, are the workloads that need specialized compute, need to scale up and down a lot. yeah, they're, they're, they're just shaped differently.Working Alongside Frontier StartupsVibhu [00:07:26]: I think you're building a lot of it alongside the startups, right? They're innovating quite a bit, even in your, like, latest blog post. Like, even in the series C, the customers that you mention here, the cognitions, technical ones, ramps and whatnot, they're, they're innovating with you, right? And that's not something AWS is doing directly with.Akshat [00:07:45]: Yeah, absolutely. I think, this is again classic. We're a small team. We can move really fast. our engineers are working with our customers and figuring it out. Yeah.Swyx [00:07:54]: So my first week at Cognition, I walked in, there was someone wearing a Modal shirt. I was like, “What are you doing here?” They're like, “Yeah, I just. I am embedded inside of Cog.”Akshat [00:08:05]: Yeah, I think that was Peyton. We sent him overSwyx [00:08:07]: Yeah.Akshat [00:08:07]: Because, the latency of communication was too high otherwise.Swyx [00:08:12]: Yeah, distributed node, you have to - you have to place one and collocate.Vibhu [00:08:16]: Yeah.Swyx [00:08:16]: So I had a, I had direct personal experience, right? So I worked on smol developer three years ago. it was inspired by Claude 1. I think you onboarded me at some point, like, just before, and I was like, “Oh, like, I need some bursty compute. Like, I was just gonna try using Modal.” And it was a, it was a pretty pleasant experience. apparently, I showed up in the board meeting, like the analytics.smol developer, Sandboxes, and Proto-CognitionAkshat [00:08:39]: Yeah, you blew up on Hacker News and,Swyx [00:08:41]: YeahAkshat [00:08:41]: We got a big traffic spike. I. I think the way you used smol developer was Modal functions for running stuff, which was. Like, the, that was a good use case. but then, yeah.Swyx [00:08:53]: Yeah. That - So to me, that was proto-cognition.Akshat [00:08:55]: Right.Swyx [00:08:56]: If only I had, like, stuck to it.Swyx [00:08:58]: Like, that was like, if - did you say draw the tech treeAkshat [00:09:00]: AbsolutelySwyx [00:09:00]: You're just like, “Yeah, like, probably this will happen.”Akshat [00:09:02]: Yeah. Like, he was so close. You were just rebuilding upon usSwyx [00:09:04]: I just didn't realize.Akshat [00:09:05]: But the funny story there is at the same time, we were talking to a bunch of customers who needed something like sandboxing.Swyx [00:09:14]: Yeah.Akshat [00:09:14]: This is like twenty-three.Swyx [00:09:15]: Yeah.Akshat [00:09:16]: So we builtSwyx [00:09:17]: You introduced a new API right after that.Akshat [00:09:18]: Yeah.Swyx [00:09:19]: Yes.Akshat [00:09:19]: Like, we built sandboxes in May of twenty-three before anyone was even knew this was gonna be a thing. And the first example we published was, we took smol developerSwyx [00:09:28]: Smol developerAkshat [00:09:28]: And put it in a loop, so the agent can iterate on itself.Swyx [00:09:33]: Loops are hot these days.Vibhu [00:09:34]: It's the looper.Akshat [00:09:34]: Yeah.Vibhu [00:09:35]: Loops in. When was this, twenty-three?Akshat [00:09:38]: Yeah.Vibhu [00:09:39]: A small check.Akshat [00:09:39]: Yeah.Swyx [00:09:39]: It's like twenty-three. so the. the, those for listeners, like, the problem was the models are not built for any of this, right?Swyx [00:09:46]: Like, you're just trying to like. They're not post-training to understand, like, looping and, like, self-correction and tool calling was there, but, like, also not that great.Akshat [00:09:55]: Yeah.Akshat [00:09:55]: I don't remember if you used tool calling in this one, but yeah, the models would just diverge after like ten iterations and not produce anything meaningful.Swyx [00:10:03]: Yeah. But like, then. So okay, like now talking to myself three years ago, the answerVibhu [00:10:08]: Of course they will get betterSwyx [00:10:09]: Collect all the failures, build benchmark, and then collect all the, examples, build the RL environmentAkshat [00:10:15]: RightSwyx [00:10:15]: Sell it for like ten billion dollars to Meta.Swyx [00:10:17]: And then also train a model and then sell that for sixty billion dollars to Elon. And this isAkshat [00:10:23]: Yeah, of courseSwyx [00:10:23]: The funny machine. Like, it's like, it's about the hardware.Akshat [00:10:28]: It's hard to have that inherent conviction that the stuff will get that much better.Swyx [00:10:33]: In retrospect, it's so f*****g obvious.Akshat [00:10:36]: Fair enough.Swyx [00:10:37]: Like, what else were we doing back then? I don't know. anyway. Yeah. So this. That was the start of your sandboxing journey, right? I feel like it didn't blow up until, like, last year.Akshat [00:10:49]: Yeah.Swyx [00:10:50]: So there was like a couple years of quietness.Akshat [00:10:52]: Exactly, yeah. We wereVibhu [00:10:53]: I think very underrated product value. Like, my experience with Modal, Charles, before he had joined Modal, met this guy at a hackathon, and he really insisted we wanted to run some small model, not hosted anywhere, and he's like, “ there's this cool company, Modal. They'll like spin up a GPU sandbox, we can throw it on there. They'll take a Hugging Face link.” And like there's so much value just right there, right? Like instant hosting, spin it up, spin it down. It'll stay cold, but we run the demo a few days later, it'll come back up and like all this stuff in retrospect, like it's still what we needed like today.Akshat [00:11:27]: Yeah, it's still needed today. workload shapes have changed a lot as, we run stuff for people with really massive production scale and, there it's it's not about scaling from zero to one, but it's how do we scale really elastically, from like thousand to fifteen hundred GPUs very quickly in a given region. It's the same shape problem.Elastic Inference, GPU Autoscaling, and Custom ModelsVibhu [00:11:50]: Okay. So you look at, say, Cursor Composer, right?Akshat [00:11:53]: Yeah.Vibhu [00:11:53]: They had a. “We'll do RL on a model every couple hours.” you guys have a whole version of RL inference gym and whatnot.Vibhu [00:12:01]: When you look at workloads like that, you're doing train runs where you need to scale up, scale down every hour thousands of GPUs, right? That's the example for we do need it, right?Akshat [00:12:12]: Yeah. Well, so I'll, I'll take a step back and, maybe talk about like how people use Modal today. because our biggest use case is, elastic inference. And the thing we first found product market fit, with was inference for custom models. So we stayed away from the LLM space, and we were serving companies like Suno for audio, Runway for video, robotics, comp bio companies that train their own model elsewhere. But Modal is the best black box that for deployment, scaling to however many GPUs you need as your traffic pattern changes. And we saw all of them like have a very unpredict- predict- predictable, traffic pattern. it's like diurnal. It's Some days, like the company will do a launch and, they'll need like, way more. And it's not just one model that they deploy. They-- all these companies deploy, lots of different models in different regions, and so the autoscaling problem becomes even harder because then you have to scale within a certain region, and those cycles are offset. So different times you scale up in different regions.Akshat [00:13:20]: So that's like our sortVibhu [00:13:22]: And thatAkshat [00:13:22]: YeahVibhu [00:13:22]: That in and of itself is a huge category. There's a bunch of inference providers which, provide this fireworks, does this as a service together, whatnot, Base10. that's carved into its own niche for language models, at least right now.Akshat [00:13:36]: Yeah. the thing that we have specialized in is the autoscaling aspect.Vibhu [00:13:41]: Yeah.Akshat [00:13:41]: Because we found that it's not universally true that everyone else can autoscale, and we've gone deeper into it on the tech side by, we've incorporated GPU snapshotting into the product so we can take the GPU state, like your torch.compile model, snapshot it, and the next cold start is way faster. And so going back to your question, it's That's why you need a lot of burstiness for inference. But then people also do a lot of demand training, like for RL stuff, your rollouts are bursty, as you said. People also do a lot of batch jobs. So we'll see, a lot of companies, before they have a training run, they'll need thousands of GPUs to run encoding or something like that. And I think those things are much more bursty than. I agree that agents are not that bursty. sandboxes are, except when you're doing RL. RL is justRL, Batch Jobs, and 100,000 SandboxesVibhu [00:14:28]: Or commerceAkshat [00:14:28]: Insanely bursty.Vibhu [00:14:29]: Yeah.Akshat [00:14:30]: Yeah. Like when you're doing, rollouts, you sometimes need a hundred thousand sandboxes in your sandboxes.Vibhu [00:14:37]: Yeah. I'm curious if you've seen early sparks of continual learning. There are some people, like our friends, ngram, recently announced thisAkshat [00:14:45]: YeahVibhu [00:14:45]: They're, they're trying to do training. That also seems like a different workload, right? If you're doing training twenty-four/seven per se, there's a very weird dynamic of how you're using GPUs between people and whatnot, but seems like something you guys would work for.Akshat [00:15:00]: As you said, we're, we're fortunate to work with a number of, customers at the frontier and grab some of our customers. and they are taking the primitives we have, and trying to use them in very interesting ways, like continual learning. It's possible as the stuff gets better, some of that will be part of, our offering as well if, more people need it. but we're, we're just waiting to seeVibhu [00:15:23]: YeahAkshat [00:15:23]: How it shakes out.Vibhu [00:15:24]: Is there a primitive that you added after sandboxing that was the next step in the story?LLM Inference, DeFlash, and Speculative DecodingAkshat [00:15:32]: I guess we've been going much deeper into LLM inferenceVibhu [00:15:35]: YeahAkshat [00:15:35]: Because we realized that some of the advantages we have with like autoscaling, again, especially in different regions and whatnot, are, not present elsewhere. and the place where we had a gap was we weren't, working on the model layer itself. Like we were a black box. And, we realized that, we can get to frontier-level model performance, with, by having great people who work on this. And, we've been open sourcing a lot of our work, in terms of, Recently, we, shared our work on DeFlash, which is a block-based, speculator, and we've open sourced, all of it. So, you can - By using open source DeFlash, you can get the same performance as you would with one of the proprietary providers. And the next thing we're thinking about hereVibhu [00:16:23]: I thought this wasAkshat [00:16:24]: YeahVibhu [00:16:24]: An interesting blog post as well, right? Like, I think in here you make a claim that. Not a claim, just that how effective speculative deco-decoding really just get to.Akshat [00:16:33]: Yeah.Vibhu [00:16:33]: Anything you wanna point out from this around, what people should know?Akshat [00:16:39]: Yeah, absolutely. the high-level summary is, it would help to describe what speculative decoding is.Vibhu [00:16:44]: Yes.Akshat [00:16:44]: I will, yes.Vibhu [00:16:45]: I think, likeAkshat [00:16:46]: YeahVibhu [00:16:46]: So we've covered like Eagle and all thisAkshat [00:16:47]: YeahVibhu [00:16:47]: Like Hydra and all those things, but it was like two years ago.Akshat [00:16:51]: Yeah.Vibhu [00:16:51]: I think it doesn't hurt, right?Akshat [00:16:52]: Yeah. Speculative decoding is you have a smaller model, called a draft model, predict tokens ahead of the bigger model, and then you have the bigger model, verify all of this, all the tokens are predicted. And the reason it's faster is if you're predicting, one token at once, you're bound by memory bandwidth. But if you can batch the verification of, the draft model, then you're much more efficient using compute, and it's faster, and as long as your draft model is producing a lot of tokens that can get accepted, which is called the accept length, you can get a speed up that's, multiple times of, the original model speed. and well, that's what we highlight here. It's Like people talk a lot about we made these kernels faster and whatnot, but improving kernel will only give you like few percentage points of improvement, and, increasing accept length, literally is a multiplicative decreaseVibhu [00:17:47]: Like two to four X.Akshat [00:17:48]: Yeah, exactly.Vibhu [00:17:48]: Without much head-on performance.Akshat [00:17:50]: Yeah. I think it may - you are running a second model, right? So it may be something more expensive in the compute,Vibhu [00:17:57]: I meant quality performanceAkshat [00:17:58]: Probably not by muchVibhu [00:17:58]: But yeah. I thinkAkshat [00:17:59]: So there's no drop in quality performanceVibhu [00:18:01]: YeahAkshat [00:18:01]: Because you're always. You're never accepting a token that the big modelVibhu [00:18:04]: It's strictly betterAkshat [00:18:05]: YeahVibhu [00:18:05]: Or it's same.Akshat [00:18:06]: Exactly.Vibhu [00:18:07]: Right. Yeah.Akshat [00:18:08]: And so we've been working a bunch on DeFlash, which is a block-based speculator. so it's instead of predicting, one token at a time, it's predicting a block. And we've been open sourcing our work with it. The next thing for us here is for helping people train speculators and custom models. it's it's something that traditionally is very forward-deployed engineering driven, support deployed, engineer driven, like you work with customers and help them do that. And our vision for. This is why we launched Auto Endpoints, is we want to make frontier-level performance available to everyone. And so, we mentioned this in the announcement, we teased it. The next thing we're, we're launching is, as you run an auto endpoint, we shadow trafficAuto Endpoints and Frontier-Level PerformanceVibhu [00:18:54]: Do you want to explain what auto endpoints are?Akshat [00:18:57]: Yeah.Vibhu [00:18:57]: I lovely, yeah.Akshat [00:18:58]: Yeah. So, this is, I guess, going back to your Modal is you touch the code, but, sometimes people don't wanna touch the code, and they wanna get started with an endpoint that works and has all the great performance and, scalability that Modal has. So we've made that easier with, a way to create an endpoint from our UI, from the CLI, that has all of our optimizations that we talked about, like the DeFlash stuff already baked in, and there's full transparency. So we give you the code, you can go run it yourself, and if you want, you can eject out into the full Modal experience, which we see as people get sophisticated, they do wanna tweak the models, they wanna, fine-tune stuff. You can still do all of that. It's it's not a black box. And yeah, the next thing, as we teased later in the post, is how do we give you value even beyond this in terms of having your draft models evolve as your data distribution evolves, again, without having to talk to a person and, yeah.Vibhu [00:19:59]: I guess just to understand it directly, you have the GPUs, you have an endpoint that's compatible, you serve open model. If someone was to do this themselves, what's the delta that you guys provide? So you do a lot of open source great work on effective inference. how does it compare to, say, I take the same model, 5.2 FP8, take shelf inference engine, vLLM, SGLang, get compute of similar capacity, similar cost. What's the delta that plugging into something this, like this offers outside of the benefit of, scaling?Production Inference Beyond Raw GPUsAkshat [00:20:34]: It's interesting because we've taken the approach of open sourcing our contributions and upstreaming them. we work closely with the SGLang team. We want the improvements that our team, comes up with to be, there in open source for others to use, even outside of Modal. The benefit to us is we have a team that has significant expertise in terms of if you do have something that is not there, our team can help you get that performance, first. the other thing is with these endpoints, we are way more elastic, as you said, than, anyone else, and you have true scaling to zero. you have true, burstiness, and in practice, that matters a lot more to people than just finding, the GPU and, running Modal code on something.Vibhu [00:21:20]: Yeah. And I will say it's not that straightforward to just. like what I said is easier said than done, right?Akshat [00:21:26]: Yeah.Vibhu [00:21:27]: It's I think still for the average person, still hard to just gut check using different. There's, there's quite a bit of combinations you can make there. the trade-offs aren't really known at face value.Akshat [00:21:40]: Yeah. it's it's not just that. I think it's it's that running production-grade inference is a hard infer problem.Vibhu [00:21:49]: YeahAkshat [00:21:49]: Even if you subtract out the autoscalingVibhu [00:21:50]: YeahAkshat [00:21:51]: Is controlling things like tail latency and, making sure every, request is delivered at least once and whatnot.The Model and Agent LifecycleVibhu [00:22:00]: There's a lot of innovation that you can do here. I think, it's very interesting that you're starting to encroach on, like as you become a full cloud, you're starting to encroach on other people's turf.Vibhu [00:22:09]: What will you not do?Akshat [00:22:13]: Well, we wanna follow our users and, make sure they get like a platform that has everything that works well together. so right now we're focused on the model lifecycle and the agent, lifecycle. so both like going from data prep to training to inference, and then also if I want to deploy a background agent, let's say, sandbox, do persistent storage, a whole bunch of other stuff.Vibhu [00:22:38]: We talked to Cole, who did, OpenInspect. Yeah.Akshat [00:22:42]: Yeah.Vibhu [00:22:42]: And RealInspect also is on Modal.Akshat [00:22:44]: Yeah. So Ramp Inspect was a great example of a background agent that was really successful because they, were able to use some of the primitives like snapshotting and fast scaling to just have something that feels really reactive and works well.Ramp Inspect and Background AgentsVibhu [00:23:02]: Yeah. That's the new CTO of, Ramp right there.Akshat [00:23:05]: Yeah, Rahul.Vibhu [00:23:08]: It was really fun. yeah, okay, I think, all very bullish. Like, one of my reflections was also I did not originally. So when I met you guysThe Inference Inflection: CPU, GPU, and Co-LocationVibhu [00:23:19]: You weren't that much in the GPU game, and now you're all about, inference. And one of the points that I hinged on for Jensen's keynote at GTC this year was, what we're calling like the inference inflection, right? That let's say in AI workloads or machine learning workloads, it used to be like, let's call it eight to one GPU to CPU, and now it's more like one to one, which is like a interesting. Like, - because of how much agents are blocked or call out to this, to CPU heavy stuff the actual, like, limiting factor, like, swings back and forth from GPU to CPU a lot more than it used to be all GPU and then occasional CPU.Akshat [00:24:01]: Yeah.Vibhu [00:24:02]: GPU, CPU. And now it's like just constantly, and you just have to locate everything.Seventeen Clouds and the Supercloud StrategyAkshat [00:24:08]: Yeah. And that's one of the things that, again, we see as, something appealing about Modal, which is we've built this capacity pool that spans, 17 cloud providers, so we're, we're very good at Running on various kinds of cloud capacity across the worldSwyx [00:24:24]: You don't have your own data centers?Akshat [00:24:25]: We don't have our own data centers. We just run across a lot of neo cloudsSwyx [00:24:29]: Yeah. AreAkshat [00:24:30]: Metal providers.Swyx [00:24:30]: Yeah. Question mark.Swyx [00:24:31]: Yeah. You're, you're running the math, and you're like, “What's the cutover point where you're like.”Akshat [00:24:36]: Yeah, it's a good question. part of it is we see our differentiator in the software layer, and, being capital light and focusing on the software helps us move really fast. so far it's worked out well because there are so many other people building data centers that we're able to work effectively with them, and again, focus on what makes us, special.Swyx [00:24:55]: Yeah.Swyx [00:24:56]: 17 gets you into, like, the local providers sometimes. LikeAkshat [00:25:00]: The,Swyx [00:25:01]: Which was the most interesting one?Akshat [00:25:02]: There are a lot more neo clouds than you expect, and they all have various degrees of, various levels of reliability. And, that's why it's something we've invested a lot of time in, is building our own reliability layer on top. so if the GPU falls off the bus or something happens, we user workloads are not affected, and that lets us use a lot more capacity than,Swyx [00:25:30]: YeahAkshat [00:25:30]: You as a user would be able to.Swyx [00:25:32]: It's a useful thing to have because like now everyone knows, like, what layer you are and, like, you optimize for being the super cloud of all clouds.Akshat [00:25:41]: Yeah. That's, that's, that's the idea. and so I guess when you mentioned colocation, that's, that's another interesting thing where, one thing we've seen is people come to us when they want, very specifically located, CPUs or GPUs, like they wantSwyx [00:25:57]: Oh, they pin it in likeAkshat [00:25:58]: YeahSwyx [00:25:58]: EU?Akshat [00:25:59]: Exactly. Or EU, US.Swyx [00:26:01]: Right. Data resiliencyAkshat [00:26:02]: AustraliaSwyx [00:26:02]: Locality thing or performance or what?Akshat [00:26:04]: It's either data locality or latency, yeah.Swyx [00:26:07]: Yeah.Akshat [00:26:07]: Like, you want your. They're running sandboxes and model. They want them to be right next to aSwyx [00:26:10]: Yeah, it's easy thenAkshat [00:26:11]: YeahSwyx [00:26:12]: To. That is important in all those things. and so, like, you've accidentally, I don't know if it's accident, but, like, you've built the perfect primitive for agents to express themselves. And then, like, it's almost very funny how every extra development just involves more file system, just involves more CPU.Akshat [00:26:30]: Yeah.Swyx [00:26:31]: Just like the things that you already have. I don't know much about, if there's any, like, networking usages that are interesting, but you've also done some good work on networking.Networking, Sidecars, Private IPv6, and SandboxesAkshat [00:26:40]: Yeah, that's exactly right. Like, we're just taking compute storage and networking and building stuff on that layer, for, again, the stuff people need.Swyx [00:26:49]: YeahAkshat [00:26:50]: We see a few interesting networking things coming up. one is people want networked sandboxes. so we haveSwyx [00:26:57]: For like a Docker cluster type thing.Akshat [00:26:59]: Yeah.Swyx [00:26:59]: Sorry, Docker Swarm. Oh, f**k. What is it called?Akshat [00:27:02]: Compose.Swyx [00:27:03]: Compose type thing.Akshat [00:27:04]: Yeah. So if you want Docker Compose, our sandboxes now support, this thing called sidecars. So you can. A sandbox is a pod of containers, and you can run multiple containers in, a sandbox. also useful because, going back to networking, people want a lot of control over, outbound networking from a sandbox.Swyx [00:27:23]: Yeah.Akshat [00:27:23]: Like, they might wanna run a middle proxy for, like, maybe logging stuff for RL or, controlling how egress can happen to a domain, injecting credentials. and yeah. So we've, we've had to build a lot of that stuff ourselves.Swyx [00:27:38]: Yeah.Akshat [00:27:39]: But then also sometimes people want, sandboxes spanning multiple nodes to talk to each other, which is an emerging thing we're seeing. We have support for that for a different reason, and yeah, we'll see if that becomes stable.Swyx [00:27:52]: Like, just an open socket. It's a. This is directly like mTLS.Akshat [00:27:56]: We do support that, which is you can, expose a tunnel inside a sandbox.Swyx [00:28:01]: Yeah.Akshat [00:28:01]: And then you can either expose it to public internet or it can be, you can add like a HTTP, auth layer above it. But we have this thing called I6PN, which we haven't talked about, which is this, like, overlay network using IPv6 addresses. so if Modal containers, within the same workspace, when this is enabled, can address each other using this private IPv6 address, and no one else can.Akshat [00:28:28]: So it's like private networking, for containers. We built it because we needed it as a primitive for our distributed training product. so we have this other feature, which is you can add a decorator to a function, and you get a cluster of GPUs. and they have RDMA networking. so you can run a distributed training job, that's truly serverless. and we did the overlay network for that. But then we've seen that people are using it for other reasons, and, I'm intrigued to yeah, what would people do with it.Swyx [00:28:59]: Build primitives and let people figure it out, right?Akshat [00:29:01]: Yeah, exactly.Swyx [00:29:02]: You put out a pretty interestingAkshat [00:29:03]: They're like, they read the docs webpage. Let me use thatSwyx [00:29:06]: YeahAkshat [00:29:06]: Something they never intended to work. This is literally not even in our docs page. People somehow found it, and they're using it.RDMA, Memory Movement, and Distributed TrainingSwyx [00:29:12]: Huh.Swyx [00:29:14]: The way you portrayed it with, like, RDMA versus TCP, like, very well laid out, but just the transfer speed change at scale for RL, like yeah, you have it, you have it built in. I'm sure someone found it. It's found it to be a lot more efficient before you made a thing out of it, right?Akshat [00:29:32]: Yeah. And not to split hairs, I guess the overlay network is the TCP overlay network.Akshat [00:29:39]: The reason we have that is you need that to do the key exchange for RDMA before you set up the RDMA network on top of that. but then people found the TCP part.Swyx [00:29:48]: Can I tell you, this is like a big aha moment for me becauseAkshat [00:29:51]: YeahSwyx [00:29:51]: So I review 2,200 submissions for the World's Fair.Akshat [00:29:56]: Yeah.Swyx [00:29:57]: And then I got this from John OsterhoutAkshat [00:29:58]: HuhSwyx [00:29:59]: Who I don't know if. Do John Osterhout by name?Akshat [00:30:01]: The name sounds familiar.Swyx [00:30:02]: He published a. He's a well-known professor, published a lot of interesting software design books, and this is the talk he chose to submit, is on RDMA at Inference. And I'm like, you wouldn't think that this guy, who is like operating systems guy, would care about RDMA.Akshat [00:30:20]: I, it makes sense to me because I,Swyx [00:30:24]: This is the cloud, right? YeahAkshat [00:30:25]: Like, the way you move around your KV cache and how efficiently you can do it, how efficiently you move, your weights from your training GPUs to your inference GPUs in RL is there's a lot of degrees of freedom, and it is a systems problemSwyx [00:30:41]: YeahAkshat [00:30:41]: Moving memory aroundSwyx [00:30:42]: YeahAkshat [00:30:43]: Scheduling.Swyx [00:30:44]: This shows you how primitive my understanding of networking stuff is.Swyx [00:30:46]: Is this like the domain of WireGuard as well?Akshat [00:30:50]: Not quite.Swyx [00:30:51]: It's adjacent?Swyx [00:30:53]: Explain everything.Akshat [00:30:54]: Sure.Swyx [00:30:56]: How do we move memory around GPUs?Akshat [00:30:58]: Well, so sorry. Yeah, that is memory. Sorry, I was talking more, and maybe I was talking like five minutes back, about the private IPv6, addressing that you've set up.Swyx [00:31:09]: Yeah.Akshat [00:31:09]: Is it like it's a VPN?Swyx [00:31:10]: Yeah, it is like a VPN, and yeah, WireGuard is, yeah, you're right. It is,Akshat [00:31:16]: Right. Yeah, you already moved on to new topicsSwyx [00:31:17]: A similarAkshat [00:31:18]: OkaySwyx [00:31:19]: In the same space, WireGuard is, encrypted and this is,Akshat [00:31:23]: And you don't need encryption.Swyx [00:31:23]: Yeah.Akshat [00:31:24]: Yeah.Swyx [00:31:24]: This is not encrypted. that's the main difference. This is TCP and we have eBPF programs that will reject or allow the TCP connection based on whether you're allowed to do it.Akshat [00:31:35]: Used to involve a full sidecar, but now you have eBPF in the Linux kernel.Swyx [00:31:39]: Yeah.Akshat [00:31:40]: Yeah. I don't know if this is a natural follow-on to the topic of like my skepticism on distributed training is that while, like, people spend a lot of money on, like, cables to hook up GPUs, and even that is not, like, fast enough, and that's the bottleneck, is your networking fast enough?Swyx [00:31:59]: Yeah. So I guess you're talking about fully distributed training like, Dialog or something which is like cross data centerAkshat [00:32:06]: That would be, yes.Swyx [00:32:07]: That's the extreme.Akshat [00:32:08]: Yeah.Swyx [00:32:08]: You're in the middle, and then other people would have like the Mellanox cables up in, like, their actual data center.Akshat [00:32:14]: When you run multi-node training on Modal, RDMA, I think Mellanox, is, or InfiniBand is like a, is all seen as RDMA. but it's a way to bypass the TCP networking stack and, transfer, stuff much faster, between one node, to the other. And we have I think like 3 terabit per second, internal networkingSwyx [00:32:40]: OkayAkshat [00:32:40]: Which is the standard that's needed.Swyx [00:32:42]: Okay. So I misunderstood whatAkshat [00:32:43]: 50Swyx [00:32:43]: What part of the stack you wereAkshat [00:32:44]: 50 gigs overSwyx [00:32:45]: YeahAkshat [00:32:45]: If you wentSwyx [00:32:45]: YeahAkshat [00:32:46]: RDMA.Swyx [00:32:46]: Okay.Swyx [00:32:48]: Yeah. I, very impressive work.Multi-Node Training, Post-Training, and Auto ResearchSwyx [00:32:52]: So effectively you're extending like the model philosophy to the training cluster, like, yeah.Akshat [00:32:59]: Yeah. And we're, we're not going for like large scale training runs. the thing that we've built multi-node training for is, we see a lot of, smaller scale post-training. like, people are post-training like medium sized fund models, so they can, get higher quality on inference. this is a perfect fit, for something like that.Swyx [00:33:21]: Yeah. That is my impression of how a lot of these labs explore branches in post-training and then eventually merge whatever they find in.Akshat [00:33:31]: Yeah. The other use case we've seen for multi-node training is even if you have a big cluster, your researchers are still doing small runsSwyx [00:33:38]: YesAkshat [00:33:39]: Having elasticity thereSwyx [00:33:40]: Right, sureAkshat [00:33:40]: Matters a lot more.Swyx [00:33:41]: Yeah. the, like, this is like the current limiting factor for auto research, which is like you need to give your model some GPUs in order for it to completely run.Akshat [00:33:51]: We have a blog post on auto resource and model is,Swyx [00:33:55]: YeahAkshat [00:33:56]: Yeah, like, turns out to be pretty good substrate for that.Swyx [00:33:59]: So my impression is auto research means many things, likeAkshat [00:34:01]: YeahSwyx [00:34:01]: Anything that Andrej coins. Right now it's still science fair, right? Like not like, I don't know how many people are doing this.Akshat [00:34:08]: We're having a golf.Swyx [00:34:08]: Yeah.Akshat [00:34:09]: I thought the same thing.Swyx [00:34:11]: Yeah, you would know.Akshat [00:34:12]: We, like, our internal both training and inference teams use this the general shape of this quite a bit. like we have this one internal repo called auto inference, which essentially we've automated our own forward-deployed engineering efforts using, this harness, which is, the agent will just spin up a sweep of different things. It'll even run like, NVIDIA inside profiler and it'll like tweak configs and it'll arrive the right thing. it'll change your GPUs both from H200 to B200, and works really well.Swyx [00:34:47]: Nice.Akshat [00:34:47]: So yeah.Swyx [00:34:48]: By the way, I enjoy that your forward-deployed engineering is so technical that you have to do these things.Swyx [00:34:52]: It's very different from forward-deployed engineering from other people.Akshat [00:34:54]: Yeah. For our forward-deployed engineering team is, essentially they're like applied inference researchers or applied training researchers.Swyx [00:35:02]: Someone told me like they have to be able to build, but they also have to be able to sell. do they have to sell or are they like they're good, they're just like post-sale type of thing?Akshat [00:35:09]: It does, being able to talk to a customer and engage effectively with themSwyx [00:35:13]: YeahAkshat [00:35:13]: Matters a lot.Swyx [00:35:14]: They want the same thing.Akshat [00:35:15]: Yeah.Swyx [00:35:15]: ?Akshat [00:35:15]: But it's it's not really a sales, thing. We pair them with-- We have solution architects as well that are more on the sales side.Swyx [00:35:23]: Okay. Let's spend a bit more time on auto research. This is a big focus for for this year. Where does this go? like, have people explored enough? Like, there's all these beautiful charts of like improve and then level off a bit and then you find the next thing. Is this one abstraction up from normal training? Is that how we think about it, or do you think about it differently? Like model level training versus high, like driven hyperparameter search.Auto Inference and Modal BenchAkshat [00:35:51]: Yeah, like,Swyx [00:35:51]: Someone, some people call it like neural architecture search or whatever, right? Like.Akshat [00:35:54]: Yeah, - So the stuff I've seen people do with it is nowhere on the architecture level. It's pretty much tweaking parameters, but it's it's a hyperparameter sweep that's guided by some model intuition, so it's like much more efficient than, whatever other, sweep you would have.Swyx [00:36:12]: Yeah, it's just, it's just a question of where you want to spend your compute?Akshat [00:36:16]: Right.Swyx [00:36:16]: ‘Cause yeah, you can just throw infinite amounts of money on this and somehow you'll bang out Shakespeare?Akshat [00:36:22]: Yeah, infinite monkey.Swyx [00:36:24]: Yeah, so like the very good for model. and I think it's also very important that agents can spin up other agents, can spin up their infrastructure. Like very good for you. how good is our LLMs at generating model code? Like the benefit of existing LLMs is that you are in the data.Akshat [00:36:42]: Yeah. They're, they're surprisingly good. I think like pre Cloud 4 they were not, and then now they're able to shot, stuff out of the box. But we're playing around with releasing like a Modal Bench for like the harderSwyx [00:36:55]: YeahAkshat [00:36:55]: Things, that the LLMs cannot do yet and maybeSwyx [00:36:59]: What's an example of that?Akshat [00:37:01]: I think the things that- Sometimes agents struggle with, without right guidance and a skill is, how to, use the rest of our observability. Like how to. Something is failing, like how do you look at the logs and then update the right thing? It's reasoning about that. But they're able to shot, likeSwyx [00:37:23]: Yeah. You can just add a skill to it?Compute Strategy and Capacity PlanningAkshat [00:37:26]: Yeah. So we have a Modal skill now that. Which is why we built this Modal Bench. It's to find things like that, so we can address them in our tool.Swyx [00:37:35]: Tune a skill. Yeah.Akshat [00:37:36]: Yeah.Swyx [00:37:36]: No. it's it's good. are you facing any shortages? like we talk a lot about GPU shortages, but also CPU, also memory.Swyx [00:37:44]: Yeah.Akshat [00:37:45]: We have had a lot of growth, which means that, there's - we've had to be much better aboutSwyx [00:37:53]: PlanningAkshat [00:37:54]: Proactive capacity planning.Swyx [00:37:55]: Yeah.Akshat [00:37:55]: So we have,Swyx [00:37:57]: Which by the way, like it's like a MBA's like dreamAkshat [00:38:00]: YesSwyx [00:38:00]: Is like just planning this stuff. I think last time you and I talked about something maybe about this.Akshat [00:38:03]: Yeah. we have a really competent team of people that we call, The role is called compute strategy. so yeah, if anyone listening here or wants to work on thatSwyx [00:38:13]: Compute strategy?Akshat [00:38:13]: Yeah.Swyx [00:38:14]: I think,Akshat [00:38:14]: I feel like,Swyx [00:38:15]: I think the normies call it FP&A or something.Akshat [00:38:18]: Well, it's more It's it's not FP&A. It's it's There's a lot of interesting financial questions of like what is the blend between one year and three-year reservations? how do we forecast our own capacity? how do we. especially since our capacity is very fungible across different GPU types and different regions, like you have to model a lot of it. and you also have to have an opinion on how the supply chain is gonna evolve, and then you have to like, take bets,Swyx [00:38:49]: YeahAkshat [00:38:49]: Based on that.Swyx [00:38:50]: Tokenomics.Akshat [00:38:50]: Yeah.Swyx [00:38:51]: This is like probably a not a real point, but, I was trying to think about like what other industries. I was trying to think about like, we cannot be first to like these kinds of problems.Akshat [00:38:59]: Yeah.Swyx [00:39:00]: And what other industries have had this? And I was like, airlines with fuel and like they have to hedge their fuel and like, I think for a long time Southwest because they made like a hero fuel bet, they like were like super low cost becauseAkshat [00:39:12]: OhSwyx [00:39:12]: Compared to everyone else.Akshat [00:39:14]: Yeah. I hadn't thought about that.Vibhu [00:39:16]: We're at a fun time too?Akshat [00:39:18]: Yeah. It's. A lot of the compute business in general, for us is also about being very good about capacity management. That is how you have great unit, economics. but also over time it's how you can unlock more value for customers. Like, one of the things we're building now is like a way for customers to get, If they don't care about latency, like get much cheaper pricing and they'll get results back in like next 24 hours or something, like a batch tier essentially.Batch Tiers and Latency-Insensitive WorkloadsSwyx [00:39:47]: Yeah.Akshat [00:39:47]: And those are levers we have because we control the whole stack and scheduling and whatnot to give people a sufficientSwyx [00:39:53]: Yeah. I feel like they're not as popular. Like those, like the Frontier Labs have all those APIs. They're not as popular as they should be.Akshat [00:40:00]: The demand that we see for something like that is not for LLMs. although sometimes people wanna run evals andSwyx [00:40:08]: OkayAkshat [00:40:08]: Synthetic data prep and there it makes sense.Swyx [00:40:10]: Okay.Akshat [00:40:11]: But it's from a lot of LLM companies, like people who are doing computational bio, like they have to run really big batch jobs and they don't care about when they get it back.Swyx [00:40:22]: Yeah. And like they have a reasonable. It's it's also like a cousin to the stopping problem of like, will this finish in time?Akshat [00:40:30]: Yeah. You can bound it.Swyx [00:40:33]: Yeah.Akshat [00:40:33]: Like you can give peopleSwyx [00:40:34]: YeahAkshat [00:40:34]: SLAs on it.Swyx [00:40:35]: Yeah. I think what's, what's interesting is like the next phase of model.Swyx [00:40:38]: Like what, do people expect from you, now that you're established and you're like well-known compute player among all these leading companies. You had an inference launch week, and we talked a little bit about the launches. like what else? Like what else should people know?What Modal Builds NextAkshat [00:40:55]: We are building primitives that make our users' lives much easier. So, I think for example, with LLM inference, thousands more companies are gonna post-train their own models and, deploy open source models for inference. so we're thinking a lot about what is the best product shape for that. And, that involves everything from our training gym to, then, endpoints that get frontier-level performance. again, but I haven't talked to anyone. It looks somewhat different on other verticals. Like, we're also seeing a lot of real-time, audio-video stuff in there, which is why like, we're working on things like regional routing, with fallbacks. So you can get GPUs that are as close to users as possible. so you get like low latency for video streaming and whatnot. And then on the agent side, it's,Akshat [00:41:52]: We're still working very closely with our customers because stuff is changing so fast in terms of what they need. And, I think beyond sandboxes and persistent file systems, there's a lot of other things people will need from this agent stack as they build production agents. So yeah, we're thinking about those other things that fit in there.Swyx [00:42:13]: I want to ask what the other things are.Akshat [00:42:15]: Yeah. I probably should share right now.Swyx [00:42:17]: I think-- I think, okay, so, I do think a lot about the principal components of cloud, and you do talk about compute storage networking.Akshat [00:42:25]: Yeah.Swyx [00:42:25]: Because so far for me, it's fine. so far for the. the first couple generations of cloud, it's fine. What's different, qualitatively different about agents that you need some new permission level? Like a lot of people, okay, and I'll just kinda spew tokens at you until it like hopefully sparks something.Akshat [00:42:43]: Yeah.Swyx [00:42:44]: Like the new level now is whatever Claude Code does, which is dangerously scope permissions or like allow list by command or like whatever, right? And sometimes they're like, “Well, okay, we have like this adaptive thinking mode where like, just trust me, bro. I will make the calls for you.” Is that it? like mediated permissions.Hard Guardrails vs. LLM-Mediated PermissionsVibhu [00:43:03]: Now you're looping it with a goal and letting it roll.Akshat [00:43:06]: Yeah, I'm, I'm skeptical of LLM media permission for stuff that is at the sandbox level because you do want hard boundaries.Swyx [00:43:16]: Yeah.Akshat [00:43:16]: Otherwise, someone can exfiltrate stuff.Swyx [00:43:20]: But likeAkshat [00:43:20]: YeahSwyx [00:43:20]: Maybe that's old school thinking. Maybe we're the dinosaurs.Swyx [00:43:23]: Maybe the AI OS or the LLM OS is really the kernel is a goddamn LLM.Swyx [00:43:30]: Like it makes you feel uncomfortable.Akshat [00:43:31]: Yeah, I'm, I'm toldSwyx [00:43:32]: But that's what trusting the LLM is. Like imagine a spherical cow perfect LLM.Akshat [00:43:36]: Right.Swyx [00:43:37]: That it.Akshat [00:43:39]: Maybe.Swyx [00:43:41]: I wanna test the boundaries, right?Akshat [00:43:42]: Yeah.Swyx [00:43:42]: Like, and I don't believe that, but I wanna see where I'm wrong ‘cause that's, that's the consensus.Akshat [00:43:49]: Yeah. I think you always need hard guardrails when you want, And you can pair those with softer guardrails, right? And that's gonna be a lot of mediated.Managed Agents and Specialized SandboxesSwyx [00:44:00]: There. I'll also get you a end with a couple of your commentary on like the ecosystem outside of Modal. Manage agents. Everyone has one. Gemini, OpenAI, Claude, very useful for you, but also like it is their way of starting to edge into your space.Akshat [00:44:17]: Yeah.Swyx [00:44:17]: What's going on?Akshat [00:44:19]: Yeah, we're, very excited to partner with Anthropic and some of the other foundation labs, will not name who we're also working with. the way we see it is the manage agent thing is a great place to start if you're starting out building an agent and, But then when you get to, building something more production grade, like you're a company that's like Ramp that's building their own, Ramp also runs their accounting agent on us, so their external-facing agent. You need a lot more control over, your compute primitive on things like, what sort - how do you persist different files that the agent has access to, and how do you snapshot and restore? How do you control the networking? maybe you want GPUs. When you get to that point, you kinda want, a specialized sandbox provider, that gives you those things, and that's the role that we are trying to play.Swyx [00:45:15]: YeahAkshat [00:45:16]: We don't really have an opinion on the harness, whether it runs - it's a cloud-managed agent, and you hook it up to Model Sandbox, or you run the harness in Model Sandbox. We'll see where people converge with that.Swyx [00:45:26]: Yeah. Do you any opinions on like the meta harnesses, or just another layer on top of these things?Akshat [00:45:31]: You mean like the OpenPipeSwyx [00:45:33]: OpenPipe is one. I think Vercel had one, which I can't remember the name of right now. Fredshot had one. and then, to me, most recently was Data Databricks that had Omnigen. All these are meta harness. Like it's kinda pseudo agent cloud type things.Akshat [00:45:50]: I personally have not played around with them.Swyx [00:45:53]: Yeah.Akshat [00:45:53]: Build agents with them.Swyx [00:45:54]: Everything's bullish Modal, as long as it consumes more infra.Akshat [00:45:57]: That's why we're focusing on the infra layer. It's somewhere where our, relative competence is and, also it's a hard problem to solve.Swyx [00:46:06]: Yeah. I will say like just generally reflecting on that, I don't know if - if there's other topics on Modal, but like just generally reflecting as an infra person, not as intense as you, but in that field, this has like been the most exciting time in infra. Like it was boring for a while, and you couldn't really get people excited about data infrastructure. Like Eric would get on Data Console, everyone just watched the video and like say, “Look at how many sandboxes I can spin up,” and no one gave a crap.Why Infrastructure Became Exciting AgainAkshat [00:46:39]: Yeah.Swyx [00:46:40]: And like now everyone gives a crap.Akshat [00:46:42]: That's true. It is a very exciting time, and I think a lot of that's driven by just the amount of scale all of this stuff needs.Swyx [00:46:50]: I think the, like a lot of your initiatives or a lot of your like product directions make sense in retrospect, which is like the best kind, but I wouldn't necessarily have thought about it myself, which.Akshat [00:47:00]: We need the predictions.Swyx [00:47:02]: I think there's a lot that you just don't even see, right? Like you have the batch, you have the voice, you have the multimodal, but what else?Akshat [00:47:10]: What else is coming up for usSwyx [00:47:11]: Yeah. Where do you see things going?Akshat [00:47:13]: Yeah. I, in generalBiotech, Robotics, and Non-LLM AI WorkloadsAkshat [00:47:15]: It's it's clear that there's there's a huge shift happening. I think one thing that's not as obvious to people because LLM inference gets talked about so much and is also we work a lot of companies that are, doing things like drug discovery and computational bio, like the Chai Discoveries of the world. Big things are probably gonna happen there. we work a lot of robotics companies that are putting robots in like active deployments and getting good results out of them.Swyx [00:47:45]: Is there Air Gap Modal? Is there a version that is like prem air gapped whatever?Akshat [00:47:50]: No. We,Swyx [00:47:51]: You should cloud only.Akshat [00:47:51]: Yeah.Swyx [00:47:52]: Yeah. Okay. But yeah, so what you're saying is like because you're focused on primitives and they're good primitives, you find use cases in all these kinds of things.Akshat [00:48:01]: Yeah.Swyx [00:48:01]: Probably diversifies you a little bit away from LMS all the time.Akshat [00:48:05]: Yeah, absolutely. We're, we'- our goal isn't to only serve the LLM inference market.Swyx [00:48:10]: There are a lot just on the website, the audio,Akshat [00:48:12]: Yeah. We said both onSwyx [00:48:14]: Computational bio images. Yeah, there's a lot here. There's QTA TTS, customizing. Oh, Chatterbox. there was customizing Whisper.Akshat [00:48:24]: Okay. Yeah.Swyx [00:48:25]: This screen reminds me of a fallen competitor, which Replicate.Model APIs vs. Differentiated AI ProductsSwyx [00:48:31]: What's your postmortem on what happened?Akshat [00:48:34]: This is one thing we've stayed away from is providing an API for models because I think providing model APIs is some of it ends up serving like a really hobbyist market, which is much less sticky.Swyx [00:48:50]: Yeah.Akshat [00:48:50]: And we've always wanted to build for companies that are building products and need more flexibility that's not just an API.Swyx [00:48:57]: Which you can build an API for a model and this is clearly what it is. But you - but what you're saying, you can wrap it into a more fully functioning back end that you run.Akshat [00:49:06]: Yeah. So all of our examples, it's not that spin up this model, here's an API token, use it. They're all code.Swyx [00:49:13]: Okay.Akshat [00:49:13]: And so the point is that this is just an example.Swyx [00:49:16]: Starter code.Akshat [00:49:17]: Yeah. But you can tweak it however you want.Swyx [00:49:20]: Yeah.Akshat [00:49:21]: And if you're like a company building a product, like, computational bio whatnot, yeah.Swyx [00:49:26]: I guess I'm trying to tease out for listenersAkshat [00:49:28]: YeahSwyx [00:49:28]: When does it stop becoming, oh, you're just an API call and you're just a wrapper on API to becoming what you call a product, right?Swyx [00:49:36]: Like, what is that layer? Like what-- Like, more lines of code, but like beyond that, what is the substance that people add that qualifies it to be something more?Akshat [00:49:46]: I think there's a little bit of like a selection effect of like a lot of the companies who do wanna get deeper into that level are probably building something that's more differentiated. And, I think, an example is like - with LLM inference, originally we, worked with companies that were building their own post-training frameworks or they were, - Ramp early in the day was training their own tokenizer and like swapping out the tokenizer in Llama and whatnot. I'm not saying that's, that successful, in that case. But a better example is like, let's say Suno. because Suno, does not use Modal for training.Swyx [00:50:26]: Mikey on the pod. Yeah.Akshat [00:50:27]: But they use Modal for all their inference and that's because they have like a custom-- They have completely custom model architecture and that means that they have to be at the code level and tweak things that are not, just an API.Swyx [00:50:41]: It's interesting as well, like we had, Ethan, most recently on the xAI Groq team make a prediction that like the next tier in video gen is not a better video model, it's a better model or agent that orchestrates video models.Video Agents and Production WorkflowsAkshat [00:50:56]: Oh, interesting.Vibhu [00:50:56]: Language model backbone that can use toolsAkshat [00:50:58]: RightVibhu [00:50:59]: And write code.Akshat [00:51:00]: Like, yes, I can make my second video or my second video from Groq, but I want my minute video.Akshat [00:51:06]: And I'm not going there through normal video gen.Swyx [00:51:10]: Yeah, that's interesting. I - So we have GPU sandboxes and recently have seen a few companies doing agents that do video manipulation or,Akshat [00:51:22]: Yeah. Give it FFmpeg and just do it.Swyx [00:51:23]: Run FFmpeg. But likeAkshat [00:51:25]: That's not enough.Swyx [00:51:25]: Yeah.Akshat [00:51:26]: You need to give it Adobe.Swyx [00:51:27]: Yeah, I hadn't put it together with like it would be a video production thing. in my mind these things were going more towards editingAkshat [00:51:36]: Yeah.Vibhu [00:51:36]: Well, shout out Mantis.Akshat [00:51:37]: I think about this a lot.Swyx [00:51:38]: .Akshat [00:51:41]: Yeah. Sorry.Vibhu [00:51:41]: Luma. Luma Agent is a version of this for video production, but it's a off.Swyx [00:51:46]: I was gonna get your quick takes, on some other stuff that happensGitpod/Ona, CI, and Runtime SandboxesSwyx [00:51:50]: In recent news and just-just see if you have anything interesting. Gitpod, very li
Patrick Moorhead and Daniel Newman unpack OpenAI's reported equity offer to the U.S. government, Anthropic's phased return following export controls, and why AWS and Microsoft are betting billions on the engineers who deploy AI rather than the labs that train it. They also dig into NVIDIA's response to its roadmap critics, Meta's expanding cloud ambitions, and the week's biggest market developments, from SpaceX to Samsung. Catch the full analysis on Episode 311 of The Six Five Pod. The handpicked topics for this week are: OpenAI's Equity Offer to Washington: OpenAI is reportedly considering offering the federal government a 5-10% stake in exchange for a two-year runway on Department of Defense contracts. Pat and Daniel read it as a distribution and IPO hedge rather than a governance concession, especially with a trillion-dollar valuation still the target. (The Decode) Anthropic's Phased Return After Export Controls: Fable 5 is back online, and Mythos remains partially restored to a limited set of vetted institutions following Commerce Department approval, a rollout Pat treats as the first real test case for how future frontier models clear review. Daniel adds that Alibaba's claimed block on Claude in China carries more optics than substance given how easily a VPN routes around it. (The Decode) NVIDIA and Palantir's Open Source Enterprise Play: Alex Karp's viral CNBC interview argued that enterprises pairing Palantir with NVIDIA's Nemotron models can match frontier-level output without a frontier lab in the loop, a clip that pulled nearly 200,000 views once Pat posted it. Both hosts agree the priority for open source models has moved from leading capability benchmarks to closing that gap fast enough for slower-moving enterprises to adopt them. (The Decode) Forward-Deployed Engineering Becomes a Billion-Dollar Line Item: AWS invested $1 billion, followed just two days later by Microsoft's $2.5 billion commitment to forward-deployed engineering teams, signaling a shift in where enterprise AI value is being created. Daniel argues the biggest opportunity now lies in implementation talent rather than model development, while Patrick sees the investments as a strategic hedge, noting that frontier AI labs have already been building their own forward-deployed engineering organizations to help customers put AI into production. (The Decode) Meta's Cloud Ambitions Meet a Trust Problem: Meta Cloud can likely match neocloud-level GPU pricing and performance, but Pat argues the company's record outside advertising makes enterprise-grade services a much harder sell. He expects the offering to function as tactical overflow capacity for labs needing compute, with a shelf life tied to demand rather than a durable AWS rival. (The Decode) The Flip: Has Enterprise AI Value Already Left the Model Layer? Daniel argues that Microsoft and AWS writing nine- and ten-figure checks for forward-deployed engineering in the same week, alongside an essay from Microsoft's Satya Nadella declaring models replaceable, confirms that implementation talent now captures value model providers used to keep. Pat counters that today's models remain far from AGI, and that once frontier labs get there they will spin off cheap narrow models fast enough to keep pricing power on their side. SpaceX Joins the Nasdaq 100 Within Weeks of Its IPO: SpaceX entered the index just 15 days after going public, the fastest addition on record, which Pat says will pull in passive buying regardless of the underlying valuation. Both hosts flag the speed itself as a signal of broader market froth. SK Hynix and Samsung Signal Memory's Return to the Center of the AI Trade: SK Hynix is pursuing a $28 billion US listing that could value the company near a trillion dollars, timed just ahead of Samsung earnings both hosts expect to show a sharp profit jump. Daniel ties the moves directly to NVIDIA's memory demand, and Pat notes most investors had barely heard of SK Hynix before this year. NVIDIA and Microsoft Split Paths in the Magnificent Seven: NVIDIA gained 7% for the half while Microsoft shed 23%, a divergence Pat ties to Microsoft's exposure to OpenAI's fading momentum and its perception as a bundled SaaS play rather than a leading infrastructure provider. Daniel calls the selloff overdone and floats it as a generational buying opportunity. NVIDIA's Roadmap Rebuttal Tests the Limits of Corporate Denial: NVIDIA issued a statement calling its roadmap intact after a SemiAnalysis report raised delay concerns tied to its Kyber rack architecture, and Pat reads the denial as legally meaningful given the liability regulated companies take on when refuting analyst claims. Both hosts land on the same read, that first customer shipment may hold while full volume ramp slips. Palo Alto Networks and CrowdStrike Undercut the SaaSpocalypse Narrative: CrowdStrike climbed 95% and Palo Alto Networks gained 113% from April to June, each posting record quarters that Daniel says quiet the theory that frontier models would simply replace dedicated cyber tools. Pat adds that security was always the wrong category for that narrative, since it ranks among AI's biggest risks rather than one of its casual replacement targets. New episodes of the Six Five Pod land every week. Watch the full video at sixfivemedia.com, and be sure to subscribe to our YouTube channel so you never miss an episode. The Decode The American AI Champions Doctrine Comes Into Focus — OpenAI 5% Gov Stake, GPT-5.6 Gated to 20 Orgs, and MGX Closes $49B https://www.cnbc.com/2026/07/02/openai-proposes-us-government-own-5percent-stake-to-address-political-blowback.html Anthropic's Super-Week — Mythos + Fable Restored (with California in Tow), and Alibaba Fires Back https://www.axios.com/2026/06/27/commerce-anthropic-mythos-restrictions-lift Palantir–NVIDIA Sovereign AI Alliance — Karp Attacks Token Pricing, Stock +8–9% https://blogs.nvidia.com/blog/palantir-secure-ai-us-agencies-nemotron-open-models/ The Embedded AI Engineering Arms Race — AWS $1B FDE Meets Microsoft Frontier Company ($2.5B) https://www.reuters.com/business/retail-consumer/amazons-aws-commits-1-billion-toward-new-unit-embedded-ai-engineers-2026-06-30/ Pat's Anthropic Forbes' Article https://www.forbes.com/sites/patrickmoorhead/2026/05/05/enterprises-need-to-be-careful-before-they-go-all-in-on-anthropic/ Meta Compute — The World's Biggest AI Capex Spender Becomes a Cloud Provider https://www.cnbc.com/2026/07/01/metas-plan-to-launch-a-cloud-business-eases-the-biggest-overhang-on-the-stock.html The Flip: Enterprise AI value has shifted from model providers to implementation partners — the buyer's next big check is for who builds it, not who trained it. FOR: The AI stack has commoditized fast enough that integration is the new moat. https://www.reuters.com/business/retail-consumer/microsoft-launches-firm-help-companies-adopt-ai-with-25-billion-2026-07-02/ AGAINST: Frontier-model access is still the scarcity, and enterprise IP is the real moat. https://www.cnn.com/2026/06/26/tech/anthropic-mythos-release Bulls & Bears SpaceX Joins the Nasdaq-100 Effective July 7 — First-Ever Pre-IPO-Type Index Add of Its Kind https://finance.yahoo.com/markets/stocks/articles/spacex-joins-nasdaq-100-july-133000124.html SK Hynix Launches ~$28B Nasdaq ADR Listing Today — Potentially the Biggest-Ever First-Time Sale by a Foreign Company https://www.bloomberg.com/news/articles/2026-07-05/sk-hynix-seeks-access-to-ai-investors-in-29-billion-us-listing H1 Mag-7 Divergence — NVDA +7.3% While MSFT Shed 22.9% Over the Same Six Months https://finance.yahoo.com/markets/article/tech-stocks-post-best-6-months-since-2023--even-with-much-of-the-magnificent-7-in-the-penalty-box-chart-of-the-day-100000120.html Palo Alto Networks and CrowdStrike Log Best Cyber Quarters Ever — +113% and +95% Between April and June https://www.cnbc.com/2026/06/30/palo-alto-crowdstrike-stock-ai-mythos.html Enterprises Need To Be Careful Before They Go All-In On Anthropic https://www.forbes.com/sites/patrickmoorhead/2026/05/05/enterprises-need-to-be-careful-before-they-go-all-in-on-anthropic/
Join us as Du'An digs into the real mechanics of running AI locally and in production - from GPU memory math to multi-agent architectures, observability, and the economics of self-hosted inference. Du'An walks through how model weights and KV cache compete for GPU memory, why continuous batching matters when you have more than a handful of users, and how agent architectures like single-agent, workflow, graph, swarm, and supervisor patterns each solve different problems. You will learn how to instrument your agents with Langfuse for observability and cost tracking, when to use Ollama versus vLLM, how prompt caching can cut provider costs by up to 75%, and why GPUs should never sit idle. Episode two of three - the next episode covers deploying at scale. Timestamps 0:00 Welcome & Introduction 1:47 Du'An's New Role at Akamai Cloud 3:10 Data Privacy and the Case for Self-Hosted AI 7:21 Anthropic and OpenAI as the New Cloud Layer 12:48 Local Models for Specific Use Cases - Cancer Detection Example 15:02 GPU Memory Math - Weights, KV Cache, and Context Windows 19:32 Continuous Batching and GPU Time Slicing 20:03 Observability with Langfuse - Live Demo 27:44 Agent Architectures - Single Agent, Workflow, Graph, Swarm, Supervisor 36:36 Token Economics, Prompt Caching, and GPU Cost Planning 45:32 Ollama vs vLLM - Prototyping vs Production How to find Du'An: https://duanlightfoot.com https://www.linkedin.com/in/duanlightfoot/ Links from the show: https://langfuse.com/ https://github.com/akamai-developers/akamai-workshop-solution-architect-agent https://amzn.to/4bvHn1p https://vllm.ai/
Дмитрий Казаков — ведущий разработчик Krita, opensource-редактора, которым художники рисуют комиксы, концепт-арты и анимацию. Начали с того, чем растровые форматы отличаются от векторных и что на самом деле скрыто внутри проприетарных форматов графических редакторов — там не только пиксели, но и зашитые в файл функции, а иногда и легаси на уровне самого формата. Дальше разобрали графические редакторы изнутри: как устроена система хранения на тайлах, зачем редактору lock-free структуры данных и подкачка с диска, как планировщик гарантирует отсутствие race condition, почему в Krita два раздельных рендеринга и где реально помогает GPU, а где нет, какие инструкции процессора вроде AVX и AVX2 ускоряют кисть и гауссово размытие. Партнер эпизода — Контур. Команда из 12 000 сотрудников развивает экосистему продуктов для бизнеса, от онлайн-бухгалтерии до сервиса видеоконференций. Вы наверняка знаете некоторые из них: Толк, Диадок, Эльбу и другие. Присоединяйтесь, если вас драйвят сложные задачи и возможность избавлять миллионы людей от рутины: https://clck.ru/3UCKXL Послушать новый подкаст Контура «От нуля до единицы. История российского IT»: https://kontur-it-story.mave.digital/ Реклама 16+, АО «ПФ «СКБ Контур», ОГРН 1026605606620. 620144, Екатеринбург, ул. Народной Воли, 19А. Erid:2SDnjdU2dRt Также ждем вас, ваши лайки, репосты и комменты в мессенджерах и соцсетях! Telegram-чат: https://t.me/podlodka Telegram-канал: https://t.me/podlodkanews Страница в Facebook: www.facebook.com/podlodkacast/ Twitter-аккаунт: https://twitter.com/PodcastPodlodka Ведущие в выпуске: Стас Цыганов, Евгений Кателла
掌握前瞻趨勢與科技脈動,立即訂閱 IC之音電子報:https://pse.is/8wpwwx—日前公布的最新全球超級電腦TOP500榜單中,中國自主研製的超級電腦「靈晟」,超越美國的超級電腦「酋長岩」,為中國奪得世界第一。「靈晟」不使用輝達高階GPU、純靠國產CPU的架構設計,引發了全球科技界的熱烈討論。當美國全面管制先進AI晶片出口到中國,超級電腦「靈晟」的問世,代表中國已經有方法「繞過」美國的限制,建立自己的龐大算力嗎?本集特別邀請工研院產科國際所研究經理石立康,從這次超級電腦排行榜的公布,探討全球雲端運算與AI產業格局,是否將因此面臨洗牌?或者也不見得?台灣在這波浪潮中的定位在哪裡?台灣在雲端運算、AI伺服器供應鏈的優勢,會受到任何影響嗎?節目將從這次的排名,深入剖析超級電腦排行榜對我們的意義,以及台灣產業的相關AI應用方向。歡迎收聽。—製作團隊製作人:李知昂企劃團隊:李知昂 / 莊俐心
Warren Pies of 3Fourteen Research joins Excess Returns to break down the AI bull market, the macro risks investors should watch, and why the data still supports continued strength in semiconductors and equities. We discuss GPU demand, token usage, open source AI, Fed policy, housing weakness, oil, earnings growth, market valuations and the biggest risks to the current cycle.Warren Pies on Xhttps://x.com/WarrenPies3Fourteen Researchhttps://www.3fourteenresearch.com/Calibanhttps://www.3fourteenresearch.com/calibanMain topics coveredWhich bearish AI arguments actually matter for investorsWhy regulatory risk may be the biggest long-term AI concernHow data center spending is crowding out housing investmentWhy the Fed may struggle to cool AI-driven investment without hurting the labor marketWhat GPU availability says about real-time AI compute demandWhy open source AI is not yet replacing frontier modelsHow token pricing and OpenRouter data help measure AI usageWhy semiconductor stocks may still be in the middle of a major cycleHow semis are being valued differently than traditional cyclicalsWhy Fed policy, earnings growth and market multiples are key to the second half of 2026What oil positioning and refined product inventories say about macro riskWhy 3Fourteen remains constructive on equities despite rising overheating riskTimestamps00:00 Intro01:04 Which bearish AI arguments have teeth?04:00 Why AI regulation is the biggest long-term risk07:03 Technology spending versus housing investment11:03 How AI CapEx is showing up in inflation data13:04 Why the labor market is more fragile than headline jobs data suggests16:24 Why GPU availability is a cleaner signal than CapEx announcements21:00 What token pricing and OpenRouter data reveal about AI demand27:36 How 3Fourteen benchmarks frontier models against open source AI30:00 Why the semiconductor selloff looked like a buyable dip34:02 Are semiconductors still cyclical businesses?38:08 Why Fed tightening could be the thing that ends the bull market42:15 What the oil shock means now45:47 Refined product inventories, crack spreads and energy stocks47:18 Are earnings estimates becoming too optimistic?50:49 Why the debasement regime still supports equities54:05 Where to find Warren Pies and 3Fourteen Research
Phoebe Liu, Nvidia Reporter, talks with TITV Host Akash Pasricha about Nvidia's financial backstops for GPU purchases. We also talk with Anita Ramaswamy about Microsoft's true market valuation and Andrew Dai, CEO and Co-Founder of Elorian, about why AI models fail basic visual reasoning tasks. We also chat with Rob Long of Eleos AI about the empirical evidence for machine consciousness, and Laura Bratton about how model routers are reining in enterprise software costs.Articles discussed on this episode: https://www.theinformation.com/articles/microsofts-real-value-close-3-8-trillion-500-sharehttps://www.theinformation.com/newsletters/ai-agenda/five-kinds-model-routers-cut-ai-costshttps://www.theinformation.com/articles/nvidia-says-will-take-cut-customers-cloud-revenuesSubscribe: YouTube: https://www.youtube.com/@theinformation The Information: https://www.theinformation.com/subscribe_hSign up for the AI Agenda newsletter: https://www.theinformation.com/features/ai-agendaTITV airs weekdays on YouTube, X and LinkedIn at 10AM PT / 1PM ET. Or check us out wherever you get your podcasts.Follow us:X: https://x.com/theinformationIG: https://www.instagram.com/theinformation/TikTok: https://www.tiktok.com/@titv.theinformationLinkedIn: https://www.linkedin.com/company/theinformation/Chapters:00:00 - Introduction01:13 - Nvidia Backstops Neo-Cloud GPU Purchases12:42 - Why Microsoft's Real Value is Closer to $3.8 Trillion22:34 - AI's Visual Reasoning Challenge34:35 - Can AI Be Conscious?45:30 - How Model Routers Are Lowering AI Costs
The AI Breakdown: Daily Artificial Intelligence News and Discussions
AI is now running at a $175 billion annualized revenue rate, with token demand, compute, and power growth reshaping the economy around it. NLW breaks down new research from Exponential View on why the AI boom may be more revenue-validated than the bubble discourse suggests. In the headlines: Fable relaunch rumors, agent regulation, California's Claude deal, Amazon-Anthropic pricing, Meta's distillation worries, GPU price hikes, and “Ramageddon.”Brought to you by:KPMG – Research from KPMG and the University of Texas at Austin shows the highest-impact AI users treat AI like a reasoning partner — and those skills can be taught at scale. Learn more at kpmg.com/us/SophisticatedSection - Section turns AI investment into workforce transformation and ROI - https://www.sectionai.com/Outsystems - Stop wondering how AI will change your business and start building the agents that will lead it - http://outsystems.com/Scrunch - The AI customer experience platform - https://scrunch.com/Zenflow Work - Agents for knowledge work - https://zenflow.free/Blitzy - Want to accelerate enterprise software development velocity by 5x? https://blitzy.com/MissionCloud - Eliminate AWS complexity with end-to-end cloud and AI services https://www.missioncloud.com/AssemblyAI - The best way to build Voice AI apps - https://www.assemblyai.com/briefRobots & Pencils - Cloud-native AI solutions that power results https://robotsandpencils.com/The AI Daily Brief helps you understand the most important news and discussions in AI. Subscribe to the podcast version of The AI Daily Brief wherever you listen: https://pod.link/1680633614Our Newsletter is BACK: https://aidailybrief.beehiiv.com/Interested in sponsoring the show? sponsors@aidailybrief.ai
Mixing Music with Dee Kei | Audio Production, Technical Tips, & Mindset
A listener-requested episode is back by popular demand! Dee Kei and Lu tackle one of the most common questions in the show's history — what computer should you actually buy for mixing, producing, recording, and video editing in 2026?They make the case for why Mac has become the default choice for anyone working in a collaborative environment, breaking down file compatibility issues, AirDrop convenience, and why the cost gap with PCs no longer justifies the headache. From there, they go deep on the Mac mini's rise as one of the best value computers ever made, why it became so hard to find, and how the M4 and M5 chip generations have made fan noise, overheating, and lag basically non-issues for audio work.The conversation also covers RAM and storage decisions that actually matter (versus the ones that don't), why external Thunderbolt drives are now just as fast and far more affordable than maxing out internal storage, and surprising GPU improvements that are changing video editing speeds for producers who dabble in content creation. They wrap with an unexpected but relevant detour into Apple's privacy stance and what that means for engineers handling sensitive, high-value sessions.If you've been putting off a computer upgrade or just don't know what specs actually matter anymore, this episode cuts through the noise.SUBSCRIBE TO OUR PATREON FOR EXCLUSIVE CONTENT!SUBSCRIBE TO YOUTUBEJoin the ‘Mixing Music Podcast' Discord!HIRE DEE KEIHIRE LUHIRE JAMESFind Dee Kei and Lu on Social Media:Instagram: @DeeKeiMixes @MasteredbyLu @JamesParrishMixesTwitter: @DeeKeiMixes @MasteredbyLuThe Mixing Music Podcast is sponsored by Izotope, Antares (Auto Tune), Sweetwater, Plugin Boutique, Lauten Audio, Filepass, & CanvaThe Mixing Music Podcast is a video and audio series on the art of music production and post-production. Dee Kei, Lu, and James are professionals in the Los Angeles music industry having worked with names like Odetari, 6arelyhuman, Trey Songz, Keyshia Cole, Benny the Butcher, carolesdaughter, Crying City, Daphne Loves Derby, Natalie Jane, charlieonnafriday, bludnymph, Lay Bankz, Rico Nasty, Ayesha Erotica, ATEEZ, Dizzy Wright, Kanye West, Blackway, The Game, Dylan Espeseth, Tara Yummy, Asteria, Kets4eki, Shaquille O'Neal, Republic Records, Interscope Records, Arista Records, Position Music, Capital Records, Mercury Records, Universal Music Group, apg, Hive Music, Sony Music, and many others.This podcast is meant to be used for educational purposes only. This show is filmed and recorded at Dee Kei's private studio in North Hollywood, California. If you would like to sponsor the show, please email us at deekeimixes@gmail.com.Advertising Inquiries: https://redcircle.com/brandsPrivacy & Opt-Out: https://redcircle.com/privacy
Kaltura's Ruthie Eisenberg and Yair Neumann reveal how the video giant is reinventing itself as an agentic digital experience company built on AI avatars and hyper-personalized content.Topics Include:Kaltura founded 2006, went public on NASDAQ in 2021.Kaltura reinventing itself from video company to agentic digital experience company.Shift from static content delivery to hyper-personalized conversational experiences.Partners and customers now demand intelligence, not just video infrastructure.Kaltura's mission: powering agentic experiences across customer and learner journeys.AWS co-sell motion strengthened as Kaltura runs on AWS AI infrastructure.Camille used a Kaltura avatar to scale her own presentations.Most enterprise websites bury content behind thousands of static links.Kaltura builds personalised web pages on the fly, in real time.Over 80% of content users see is surfaced for the very first time.Acquisitions of eSelf.ai and PassFactory complete Kaltura's agentic content flywheel.PassFactory answers: what should this specific person see next?eSelf.ai enables multimodal conversational avatars that guide users emotionally.20 years of behavioral data underpins Kaltura's content intelligence advantage.GPU scarcity and compute costs shape every AI architecture decision Kaltura makes.Kaltura optimises model tiers — strongest for planning, lighter models for execution.Fidelity, speed, and cost form a constant triangle in every AI product decision.Go-to-market and product teams now work closer together than ever before.Pricing shifting from seat-based SaaS to consumption and outcome-based models.Kaltura co-creating pricing frameworks with customers across different verticals.Internal product agent now handles research, stories, and data analysis autonomously.Small two-to-three person squads move fastest in the current AI environment.Yair's advice: fail at least once a week, succeed once a quarter.Kaltura scaled its CEO via avatar for a live investor earnings call.Ruthie's advice: keep the customer at the centre of every single decision.Participants:Ruthie Eisenberg – Vice President, Strategic Partnerships, KalturaYair Neumann – Senior Vice President of Product, KalturaKamil Davidov – Sales Leader Israel ISV-BizApps, Amazon Web ServicesJohan Broman – EMEA ISV Head of Solutions Architecture, Amazon Web ServicesSee how Amazon Web Services gives you the freedom to migrate, innovate, and scale your software company at https://aws.amazon.com/isv/
Gerard and Laurent welcome Michel Boutouil, co-founder and CEO of Polarise, a leading European AI infrastructure provider and NVIDIA Cloud Partner based in Berlin. After discussing about what happens outside of a datacenter, it is time to dive inside one. Polarise is one of the few genuinely European NeoCloud companies — essentially a European counterpart to CoreWeave — specializing in GPU infrastructure for AI inference. Through its partnership with NVIDIA, Polarise designs its datacenters around the GPU rack itself, using liquid cooling from the outset rather than starting with a traditional real estate-first approach. The company has already developed AI factories in Germany, Norway and the UK. In the conversation, we explore the growing commoditization of large language models and why the real long-term value may lie in AI factories — facilities that are fundamentally different from conventional datacenters. Given Europe's notoriously long grid-connection timelines, Polarise focuses on refurbishing brownfield sites with under 50MW of grid access instead of pursuing massive gigawatt-scale campuses. It's a pragmatic “pod” strategy: adapt to the grid's constraints rather than try to reshape the entire energy system. We also tackle the thorny issue of digital sovereignty. With the U.S. CLOUD Act allowing U.S. authorities access to data managed by American tech companies, it is fair to ask what hyperscalers are doing with European data — and whether Europe needs its own sovereign AI infrastructure. Polarise has secured €1 billion in backing from Swiss investor SWI Stoneweg Icona, but even that is modest compared with the hyperscalers' spending power. For comparison, SpaceX has reportedly invested around $40 billion in Colossus 1 and 2 alone. So, what does the future of the European AI ecosystem look like? Michel's answer is clear: Europe should not try to outspend China or the United States head-on. Instead, it should play to its strengths — smart execution, agility, flexibility, and the ability to learn quickly from the mistakes being made elsewhere. “Today's show is supported by the BMW Foundation Herbert Quandt. The BMW Foundation unites leaders across sectors to develop solutions that foster an innovative economy and a future-proof society. A key focus is "Energy Transition & Climate Change," where the Foundation drives "International collaboration to accelerate the energy transition." With rising energy demands from AI and data centres, new partnerships, effective collaboration, and the exchange of science-based solutions and strategies are essential.”
This AI:AM highlights cut brings together Cameron Berg, David Duvenaud, Michiel Bakker, Shawn “swyx” Wang, and Bing Xu to examine what we understand about frontier AI systems and what happens as more decisions move into their hands. Berg grounds model-consciousness debates in experiments on architecture, agency, valence, and welfare, while Duvenaud argues that even well-aligned AI could gradually disempower humans through ordinary economic choices. Bakker frames Europe's AI challenge as a sovereignty problem, and swyx turns to practitioner stakes around agents, evals, maintainable code, and who owns the system of record. Xu closes the loop at the infrastructure layer, arguing that self-improving compute and GPU-kernel automation may deepen rather than weaken the CUDA moat. For full show notes, links, and references, read the episode page:https://www.cognitiverevolution.ai/ai-am-4-cameron-on-model-consciousness-duvenaud-s-gradual-disempowerment-swyx-s-ai-eng-alpha/ Mercury: Command is Mercury's new conversational interface, giving you natural-language access to your finances and helping you take actions within your existing permissions and approval policies. Visit https://mercury.com to learn more and apply online in minutes. Sponsor: Claude: Claude by Anthropic is an AI collaborator that understands your workflow and helps you tackle research, writing, coding, and organization with deep context. Get started with Claude and explore Claude Pro at https://claude.ai/tcr CHAPTERS: (00:00) About the Episode (00:38) Special Sponsor (02:26) Model consciousness indicators (10:47) Valence inside models (Part 1) (17:09) Sponsor: Claude (19:01) Valence inside models (Part 2) (19:01) Misalignment and uncertainty (25:16) Gradual disempowerment threat (35:10) Slow zones and successors (47:41) Europe's AI bind (55:13) Frontier code benchmarks (01:01:59) Routing and memory (01:10:42) Agent infrastructure strain (01:16:25) Self improving infrastructure (01:27:56) Routing compute costs (01:35:39) Sovereign AI financing (01:42:24) Judging AI judges (01:47:26) Building AI DNA (01:52:38) Episode Outro (01:55:09) Outro PRODUCED BY: https://aipodcast.ing SOCIAL LINKS: Website: https://www.cognitiverevolution.ai Twitter (Podcast): https://x.com/cogrev_podcast Twitter (Nathan): https://x.com/labenz LinkedIn: https://linkedin.com/in/nathanlabenz/ Youtube: https://youtube.com/@CognitiveRevolutionPodcast Apple: https://podcasts.apple.com/de/podcast/the-cognitive-revolution-ai-builders-researchers-and/id1669813431 Spotify: https://open.spotify.com/show/6yHyok3M3BjqzR0VB5MSyk
In Episode 320 of The Block Runner Podcast, hosts William, Iman, and TJ are joined by SuperFan and Dr. Mien of BigNoodle, the AI native art platform behind the Heroes collection that launched on the block pad. The guys trace BigNoodle's path from a Bitcoin Amsterdam hackathon and DMT inspired generative art to its next chapter: a decentralized AI compute network. The core thesis: Bitcoin turns energy into value, and Big Noodle wants to turn that same energy into intelligence. They dig into decentralized physical AI infrastructure, GPU and CPU boxes that aim to cut the cost of inference by as much as ninety percent, and compute as a brand new asset class in a world of data center power shortages and GPU scarcity. They also cover whether you can plug a frontier scale model into a distributed network, the mining style reward mechanism behind it, and their yield product called Bullion, where epoch based profit share lets you contribute a unit to a compute pool the way you would add to a DeFi pool. Plus why censorship resistant AI matters as the frontier labs start drawing political lines, and the sci fi units they minted for the NAT.fun hackathon. Disclosure: The hosts are founders of NAT.fun and hold positions in assets discussed. Nothing in this episode is financial advice. Watch the full episode on YouTube and subscribe to the newsletter at TheBlockRunner.com.
Hyperscalers are spending more than $750 billion a year on AI infrastructure, and much of it is for physical hardware that needs financing. That's creating a compelling opportunity for asset-based lenders who can underwrite real collateral and contracts rather than picking technology winners. On this episode of Disruptive Forces, host Anu Rajakumar speaks with Sean Hinze of Neuberger's Specialty Finance team. Together, they discuss: Why hyperscalers prefer off-balance-sheet financing and what that means for private credit How to underwrite GPU deals when chip technology evolves every two years Why power is the new bottleneck — and what a 68-gigawatt US shortfall means for lenders Where the strongest relative value sits today across chips, power equipment, and fiber How to separate hype from opportunity in a crowded space This communication is provided for informational and educational purposes only and nothing herein constitutes investment, legal, accounting or tax advice, or a recommendation to buy, sell or hold a security. Information is obtained from sources deemed reliable, but there is no representation or warranty as to its accuracy, completeness or reliability. This communication is not directed at any investor or category of investors and should not be regarded as investment advice or a suggestion to engage in or refrain from any investment-related course of action. Neuberger is not providing this material in a fiduciary capacity and has a financial interest in the sale of its products and services. Investment decisions should be made based on an investor's individual objectives and circumstances and in consultation with his or her advisors. All information is current as of the date of this material and is subject to change without notice. Any views or opinions expressed may not reflect those of the firm as a whole. Neuberger products and services may not be available in all jurisdictions or to all client types. This material is not intended as a formal research report and should not be relied upon as a basis for making an investment decision. The firm, its employees and advisory accounts may hold positions of any companies discussed. This material may include estimates, outlooks, projections and other "forward-looking statements." Due to a variety of factors, actual events or market behavior may differ significantly from any views expressed. Investing entails risks, including possible loss of principal. Indexes are unmanaged and are not available for direct investment. Past performance is no guarantee of future results. Use of Artificial Intelligence Tools. Neuberger may utilize AI tools in its business operations to improve operational efficiency and for assistance in research and analyzing data among other uses. AI tools are dependent on historical data, consequently, if the content or analyses that AI applications assist Neuberger in producing are or are alleged to be deficient, inaccurate, or biased, a client account may be adversely affected. Additionally, AI tools used by Neuberger may produce inaccurate, misleading or incomplete responses that could lead to errors in Neuberger's and its employees' judgement, decision-making, investment research or other business activities, which could have a negative impact on the performance of a client account. The application of AI in investment processes, research, or analysis is evolving and subject to limitations, including data quality, algorithmic biases, and interpretive errors. AI outputs should not be relied upon as the sole basis for investment decisions. No assurance is given regarding the accuracy, completeness, or timeliness of information generated by AI. This material is being issued on a limited basis through various global subsidiaries and affiliates of Neuberger Berman Group LLC. Please visit www.nb.com/disclosure-global-communications for the specific entities and jurisdictional limitations and restrictions. The "Neuberger" name and logo are service marks of Neuberger Berman Group LLC. © 2026 Neuberger Berman Group LLC. All rights reserved. M-003297
Is the open model GLM-5.2 really Opus 4.8 level?
Brad's out of town this week, so Will welcomes Expedition: Handheld and The Full Nerd's Adam Patrick Murray to run down the current state of the handheld gaming console market. We talk about Intel's new GPU-first handheld processor, the current state of x86 emulation on ARM handhelds, the pros and cons of the Analog Pocket, and a bunch more! Make sure you check out Adam's work on Expedition: Handheld and The Full Nerd! Support the Pod! Contribute to the Tech Pod Patreon and get access to our booming Discord, a monthly bonus episode, your name in the credits, and other great benefits! You can support the show at: https://patreon.com/techpod
Anthropic pulled the plug on its Mythos / Fable 5 model after the U.S. government raised concerns, and IREN has completed its acquisition of Nostrum for 490 MW of capacity in Spain. Welcome back to The Blockspace Podcast! Anthropic and Uncle Sam are trading blows again, with the frontier LLM company pulling its recently released Mythos / Fable 5 model after whistleblowers said the model's guardrails were bypassed. Lygos Finance's CEO Jay Patel joins us for his reaction to the news and the market rally with a reported, imminent peace deal coming for the Iran War this week. For other news, we cover IREN's closing its acquisition of Nostrum, which will give it a 490 MW foothold in Spain for AI data center development, and the EPA's stance that it won't regulate AI data centers. Check out Dimetrics, the AI industry's Bloomberg terminal. Track financial metrics and news for AI stocks, GPU rental prices, state-by-state data center pushback, and more with the compute industry's most powerful dashboard. Subscribe to our newsletter to receive updates for all of our shows and content.