Podcasts about groq

  • 190PODCASTS
  • 342EPISODES
  • 48mAVG DURATION
  • 1WEEKLY EPISODE
  • Aug 25, 2026LATEST

POPULARITY

20192020202120222023202420252026


Best podcasts about groq

Latest podcast episodes about groq

Black Box
Treasury, dazi e sanzioni. Asia prudente. Debasement e oro. Bitcoin sopra 80mila | Morning Finance

Black Box

Play Episode Listen Later Aug 25, 2026 25:13


25/8 Buyback sui Treasury, la prima di Warsh a Jackson Hole, nuovi dazi (Canada e Cina) e sanzioni (Iran). La settimana del test sulla credibilità americana. Cosa significa per i vostri portafogli? Debasement Trade: dollaro in recupero, stabili Treasury e oro con Bitcoin sopra 80mila. Futures in verde dopo il sell-off di semiconduttori e memory della vigilia. Petrolio stabile dopo “warning” su sanzioni secondarie su digital asset, Tech, oro, aviazione e shipping. Il caso cinese. Nvidia prepara i conti e mette in produzione i chip Groq, Trump investe in Spacex, Anthropic prepara l'Ipo ma l'America rifiuta i datacenter. Sec: mandato di comparizione a banche Wall Street per caso Situational Awareness. ***Questo episodio è offerto da ⁠Scalable Capital ⁠Apri un conto con Scalable Capital e inizia a ricevere il 2,5% di interessi* sui tuoi risparmi:  https://it.scalable.capital/broker-online?utm_medium=affiliate&utm_source=qualityclick&utm_campaign=broker&c_id=QC59486e7f67706c777b517d435049607362766c747c5aS7541p&utm_term=983 Messaggio pubblicitario. Tasso lordo annuo variabile sulla liquidità depositata nel conto deposito non vincolato, composto da tasso base collegato al Tasso di Deposito BCE e tasso bonus discrezionale. Liquidità allocata presso banche partner e fondi monetari riconosciuti. Foglio informativo e condizioni su scalable.capital. Investire comporta dei rischi*** Asia prudente, Kospi giù con Samsung e Sk Hynix. Boj, verso altro rialzo. PPI sotto attese. In Cina recupera Alibaba, debacle Unitree. Europa: più mementum vs. Usa? Oggi dati su commercio e Ifo. Focus su auto e dazi al Canada. Bpm oggi cda straordinario su Ops Mps. Giorgetti incontra rappresentanti Siena. ISS a sostegno di Intesa, oggi si riunisce la Consob sotto Stazi. Unicredit, Weidmann apre a Orcel.  Learn more about your ad choices. Visit megaphone.fm/adchoices

Latent Space: The AI Engineer Podcast — CodeGen, Agents, Computer Vision, Data Science, AI UX and all things Software 3.0

In recent months, the open vs closed, and US vs China discussions on model ownership and sovereign/local AI have heated up to a fever pitch. So it is very very good news that Poolside AI are finally emerging with new models, like Laguna S 2.1, that are beating Thinking Machines' recent release nearly 10 times their size.Poolside's recent tech report got a lot of praise due to their level of detail, and Vibhu first covered Laguna's recent technical report on our paper club:From spending $12 million building language models for code before the world cared to creating a Model Factory that can take a model from pre-training to release in eight weeks, Eiso Kant has spent more than a decade betting that code is the path to AGI. In this episode, the Poolside co-founder joins swyx and Vibhu to explain why ChatGPT felt like vindication, why Poolside embraced open weights and open research, and why he would rather live in a world with 100 foundation model companies than five even if Poolside were one of the five.We go deep on Poolside's Model Factory: the engineering systems behind 10,000–20,000 experiments per month, streaming data directly into training, reproducible experimentation, low-precision compute, and agents that increasingly write code, launch jobs, evaluate results, and modify the pipelines used to train future models. Eiso also unpacks their recent launch Laguna S, why persistence, verification, and backtracking may matter more than raw intelligence, how much capability remains inside smaller models, why reinforcement learning will move earlier into pre-training, and why next-token prediction is still extracting too little from the web.We also discuss model-harness co-design, Poolside's path from coding agents to AGI, why Eiso thinks MCP and traditional tool calls are “stupid,” the real economics behind frontier-model training, Poolside's $500 million raise, open-source AI, regulation, NVIDIA and TSMC's influence, engineering productivity in the agent era, high-agency teams, and hiring at Poolside.We discuss:* How Andrej Karpathy's RNN work inspired Eiso to start building language models for code in 2015* Why Eiso spent four years and $12 million pursuing an idea before the market cared* Why ChatGPT felt like vindication and brought Poolside back to open source* Why Eiso would prefer 100 foundation model companies over an oligopoly of five* The difference between releasing open weights and publishing genuinely open research* Why Poolside deliberately built a global research organization outside the Bay Area talent war* Why model building is ultimately 90% engineering* The Model Factory: Poolside's end-to-end system for rapidly training and improving models* How fewer than 70 researchers run roughly 10,000–20,000 experiments each month* How Poolside moved from six-month model cycles to five- and eight-week launches* Why streaming data directly into training unlocked faster experimentation* How immutable data, versioned code, and reproducibility enable rigorous model research* Why Eiso wants capable researchers to leave their labs and become Poolside's competitors* Why 95% of model building can be reduced to better data or compute efficiency* Laguna S and why persistence, verification, and backtracking can outperform raw intelligence* Why smaller models may handle far more knowledge work than previously expected* Why reinforcement learning will move earlier into pre-training* Why next-token prediction is still failing to extract enough knowledge from the web* Why distillation and environments have become the AI industry's favorite “drugs”* Why mid-training is really an early form of curriculum design* Low-precision training, networking bottlenecks, and the next gains in compute efficiency* Laguna S: 118 billion total parameters, 8 billion active, and eight weeks from training to launch* Why model builders can often evaluate a new checkpoint within its first 30 minutes* Model versus harness: where agent capabilities actually come from* Why Poolside sees coding and long-horizon software tasks as a path to AGI* Why Eiso thinks MCP and traditional tool calls are “stupid”* Why future agents will write scripts instead of choosing from dozens of predefined tools* The case for minimal harnesses, containers, and model freedom* Why Poolside is prioritizing vision but does not expect to work on audio soon* Why language may be the most compute-efficient modality for encoding knowledge and reasoning* The real cost of model development and why the final training run is anticlimactic* The story behind the Poolside name and why it represents refusing to lower ambitions* How Poolside raised $500 million while investors still questioned whether AGI was real* Why intelligence could become the world's most demanded and commoditized resource* When open models may become too capable to release without restrictions* Why unilateral AI safety does not work in a globally competitive environment* How regulation could accidentally lock in an oligopoly of two or three AI companies* NVIDIA, TSMC, and the hardware systems underpinning foundation-model progress* Why reinforcement-learning wall-clock time is one of Poolside's biggest bottlenecks* Why Poolside trains models from scratch instead of simply distilling larger models* How AI changes the way companies should measure engineering productivity* Why agency may become the most important quality for employees in the AI era* How leaders align high-agency people through shared goals and clear constraints* Hiring across research, post-training, pre-training, architecture, evals, and engineering at PoolsideEiso KantLinkedIn: https://www.linkedin.com/in/eisokantX: https://x.com/eisokantPoolside: https://poolside.aiTimestamps00:00:00 Introduction00:00:54 Karpathy, RNNs, and Building Code Models Before Transformers00:02:26 The $12M Failure and ChatGPT Vindication00:03:39 Open Source and the Case for 100 Foundation Model Companies00:09:22 Open Weights, Open Research, and Poolside's Global Team00:16:04 The Model Factory: Why Model Building Is 90% Engineering00:20:19 Agents, Automated Experiments, and Early Signs of RSI00:24:04 Streaming Data, Reproducibility, and Scientific Rigor00:30:35 Creating More Foundation Model Companies00:36:07 Laguna S: Persistence vs. Raw Intelligence00:43:01 Reinventing Pre-Training, RL, and Curriculum Design00:52:33 Low-Precision Training and Squeezing More From Smaller Models00:58:37 Model Harnesses, Coding Agents, and the Path to AGI01:09:26 Why MCP and Traditional Tool Calls Are “Stupid”01:13:04 Vision, Multimodality, and Why Language Still Matters01:18:15 Scaling Models and the Real Economics of Training01:20:40 Why Poolside Is Called Poolside and Raising $500M01:27:37 Open Models, AI Safety, and the Risk of an Oligopoly01:33:53 NVIDIA, TSMC, and the Reinforcement-Learning Bottleneck01:41:52 Smaller Models, Distillation, Engineering Productivity, and HiringTranscriptIntroduction: Eiso Kant, Poolside, and Open ModelsSwyx [00:00:00]: All right, we're here in the studio with Eiso Kant from Poolside, together with Vibhu. Welcome.Eiso Kant [00:00:08]: Thanks. Thanks for having me, guys. Good to be here.Swyx [00:00:10]: Yeah, fresh on the plane. You texted me, you were like, “Hey, I'm on my way to SF.” I was like, “You're on a plane right now, right?” Like, hey.Eiso Kant [00:00:16]: I know. After I texted you, I realized that probably coming in with major jet lag was gonna offer some fun experiences today, but let's do it.Swyx [00:00:23]: I mean, I think the thing I would tell guests is that they don't have to prepare that much because if you're truly working on this every single day, then even, like, what you hazily remember is going to be new for a lot of the audience that don't live in your world every day, right? so 10 years ago, you did a talk at Google Slush, talking about the democratization of AI. and, now here you are, like, open sourcing an incredible new model that we're gonna talk about. But I guess, like, what got you into democratization of AI? Like, it's not obvious from your LinkedIn or something.From Karpathy's RNN Post to SourcedEiso Kant [00:00:57]: No, it's not at all. I don't think it's obvious how I got in this space. I owe getting into this space to Andrej Karpathy.Eiso Kant [00:01:05]: In 2015, he wrote an article called “The Unreasonable Effectiveness of Recurrent Neural Nets.”Swyx [00:01:10]: Neural Nets, yep.Eiso Kant [00:01:11]: And that article, I read it, and I pivoted my startup at the time overnight to working on RNNs, and later LSTMs and Transformer models to be able to write code. If you go to this article and you scroll down, you can start seeing, like, this was the precursor to what ended up becoming language models. So, at least when he was character-level language models that were starting to predict letters, he has an example out here. There's a little Paul Graham generator, and you can read it, and the text makes sense, but it doesn't. and there's a little-- There's an example of code a little bit further down. Yeah, so Shakespeare.Swyx [00:01:47]: Shakespeare.Swyx [00:01:49]: CoolEiso Kant [00:01:49]: And for some reason, I read this, and I went down the rabbit hole of learning everything I could about RNNs and LSTMs, right? This is Transformer paper. And I had built a completely unreasonable belief, that neural nets should be able to generalize to anything and everything, and that language should be able to generalize, to a lot of things that are intelligent and the ability to write code. And so I started building Sourced, which was a fully open source company trying to build, what we used to call machine learning on code, language models on code. And we spent about four or five years on this, till the end of 2019. And that sounds really cool today, but back then, no one cared.Eiso Kant [00:02:29]: Right? Like, no one cared. We were in the dark. Like, we did things along the way. We tried applying convolutional neural nets to, like, the structure of code. We were. when attention came out, we were applying it to LSTMs, and then the Transformer paper came out. And it - it wasn't obvious, and what we missed throughout that entire journey, that we were on the right track, but we should have just kept scaling up. And today, to all of us, the scaling laws and scaling up seems like the most obvious thing. But having spent four or five years of my life on working on language models on code, it wasn't obvious. So I have a lot of respect to folks at Google and OpenAI and others who took that confidence and kept going. we failed ultimately at the time, and it was, like, biggest failure of my career, right? You blew $12 million of investors' money, which was a lot back then.Swyx [00:03:18]: Yep.Eiso Kant [00:03:19]: You spent, still a lot, but, And you spent years with, like, a group of 40 people just obsessing over this problem. And life took a different turn, And it was, and family became a focus, and I kept my heads down and really, didn't really look at language models for the following two years. big mistake considering Following years are gonna be really interesting. And then ChatGPT came out And it was like a vindication. It's like people started texting me. I found, like, my old, work decks and these old talks. And throughout that whole journey, we,ChatGPT, Vindication, and Returning to Open SourceEiso Kant [00:03:56]: We really had a strong point of view at the time that, like, as you're building more capable intelligence, it should be open and open source.Eiso Kant [00:04:04]: When we started Poolside, that wasn't the case at all, and I wanna be very open about it. When we started Poolside, we were like, there was a premise of two things. One is this technology is not gonna stop compounding in capabilities. I think to most people obvious today, but three-plus years ago when we started, most people were still arguing if these were stochastic parrots or not.Eiso Kant [00:04:23]: And the second was that reinforcement learning was gonna be the biggest driver for LLM capabilities. Today, very obvious. Three years ago, was not an opinion held or direction held at either OpenAI or Google or Anthropic or others. And so people looked down on us a little bit. They were like, “ is this really gonna work?” And so we just started working the problem, and we never really thought about open source again. We just kept our heads down and we built our, like, knowledge, understanding from scratch, right? We didn't roll out of an existing lab. So we picked up the papers and started writing code and figuring things out.Eiso Kant [00:04:59]: And it wasn't until the beginning of this year that me and my founder, Jason, picked up the open source conversation again.Eiso Kant [00:05:07]: And if you go back to some of the early things on our website, it was very straightforward. It was we wanna get to AGI, we wanna support a world of abundance, and we wanna be the first company that gets there.Eiso Kant [00:05:20]: But we started talking at the beginning of this year because it became obvious that the world was going in a direction that was starting to like, pick at us a little bit. Like, it didn't, this didn't happen overnight. It was, like, a little bit we were seeing this and we're like, “Okay, The world's going down a path.” And Throughout this journey, there was something that I used as a, as an analogy or thing. So I said well, if I go back to back in those days, 2015 or 2016, we're working on this, and I picked up a fi book off the shelf, and I was reading the book about 2035. AGI is achieved, and the story would be over the following, decades. And it would have that first chapter where everyone's trying to figure things out. You'd get the chapter of ChatGPT coming out And then you would get to the chapter where the world was at a fork in the road, and the one that it picked was one where three or four or a handful of companies were going to create all of intelligence moving forward.Eiso Kant [00:06:21]: And when I thought about that story, it felt like a dystopian fi book, not a utopian fi book. And the reality is, I'm a utopian fi guy. Like, and so We took a step back and said, “Hey, can we play a role here?” Now it was easy for us to do so because we were not at the frontier.Eiso Kant [00:06:41]: If we were at the frontier, I don't think we could have changed our mind. and I don't mean this like it's when the moment there's too much capital involved, too much expectations, you've built up things, right? We're a small team, just improving and improving. And so we knew that we could make that decision now, but it would be a lot harder to make as we got closer and closer to the frontier and caught up to others. And did a lot of soul-searching and a lot of conversations, and said, “No, this makes sense,” Even if there's big unanswered questions, like how the hell do you build a business model with foundation models about open source? Big open-ended question that we do not fully have the answer to yet, right? At what point do you no longer wanna release open source models because misuse of models has, real potential risks associated with it? how is the government gonna respond to open source? but I think it all just came down to one thing, and I'll stop the monologue, is the fact that I rather live in a world that has 100 foundation model companies than a world that has five, even if I was one of the five. And the smallest and most meaningful contribution we can make for 100 to exist is to open up our research and open up, like, our weights right now and figure out along the way how we can, like, do more.Neo-Labs, Model Choice, and the Token EconomySwyx [00:08:01]: Yeah. I think if anything, over the past three years, that has become a bit more true. you are one of a cohort of Neo labsEiso Kant [00:08:10]: YeahSwyx [00:08:10]: That people are now calling that. And, we're, we're doing this on the day that Thinky launched their, new model and you are outperforming them on their, on some benchmarks that they released, right? Like, they just don't have it yet. so it goes to show that I think, like, this is one of those things where, like, there is room for multiple players, and you are seeing a little bit more of the future. Maybe more like 20, not 100, but, like, you are one of the 20.Eiso Kant [00:08:36]: I really hope so, right? I think we I'm, I'm excited about their release, and I'm excited about everyone releasing because, like, ultimately, like, choice competition is both gonna drive progress in the right direction. But the fact that like, we create models and while we all, drink out of the same well of data effectively, we do introduce very different behaviors and biases in our models. Some are intended biases, some are completely unintended biases.Swyx [00:09:03]: Yeah.Eiso Kant [00:09:03]: And if we shape up in an ecosystem in the world where open models are gonna be a part of the token economy, like, I don't think there's any question about it anymore Then we want to be able to live in a world where companies, countries, people can choose and say, “Hey, I am most aligned and I trust most this provider for these things.”Swyx [00:09:25]: Yeah.Vibhu [00:09:26]: I think more than just one of the 20 Neo labs, up until recently, most of open source innovation was coming from the Chinese labs, right? So there's the DeepSeek of the West. Is it today? Okay, maybe it's thinking machines reflection, but there aren't many, right? So, one of the things you guys started in France, Europe, but very much now you're taking that American standpoint and more than just that, the point is the Chinese models that we see, they're not super open research. the work you put out is, I think, some of the best. So every few months you get not only frontier models, but also here's a breakdown blog, paper, technical report of here's everything for state of the art to build, frontier intelligence and you're filling that gap too, right? So not just only open weight, not just Western, but also pretty open research.Open Weights vs. Open ResearchEiso Kant [00:10:20]: No, I appreciate it. Look, I think it's, I think it's the most meaningful contribution, right? Weights are a binary. Let's call them what they are. Yes, we can modify them, we can change them, but, like, giving someone the weights does not allow them ultimately to recreate what you're doing, right? And so now there's challenges around releasing data sets, challenges around like releasing certain things, but being able to share your research, like, right, how do we do it? What are the lessons we learned that we spent, tens of thousands of experiments of compute on? I think very much so. One correction though, Vibhu, and I say this because it's been haunting us for quite a few years. We from day zero were an American company.Swyx [00:10:55]: Yeah. They movedPoolside's Global Team and American Company StorySwyx [00:10:56]: To France.Eiso Kant [00:10:56]: So the story once and for all is very. We start as an American company. We have always been an American company, and early on we made a very conscious decision. We said, “We're not gonna hire any researchers in the Bay Area. We're gonna look for talent everywhere else in the world.” and that is everything from Middle Americas, Seattle to, Serbia, and to Taiwan and Singapore and other places. And it was because we took a view that this was gonna become a talent war for this, and I think it has over the years now. Three years ago, that wasn't fully obvious yet. I think today it very much is. And we also realized that, like, some of the world's most capable people with, like, the most interesting, innovative ideas were not just gonna be here. And so it led us to create like a fully remote company. and we ended up opening an office in Paris and London and different places and we have a lot of the team in the US and a lot of team outside. But we always took this view of like, we're an American company, but if we want the best of the best to work with us, we need to take a global view. Now we do also have people here in Silicon Valley, like the company's grown and others, but I think one of the things that, it slowed us down at the beginning, but it has sped us up now, and it's why you're seeing like the progress, I think, on our models and the cadence at which we release, is because we didn't roll out of an existing lab. Right? we didn't, we didn't have a lot of the information that's freely flowing around here at the time. We just took this point of view as like, “Okay, well, let's just work the problem. Let's just go and, like, read the few papers that are out there, and let's just figure this stuff out.” And we made some hilarious mistakes in model training because of that over the yearsEiso Kant [00:12:35]: Like especially in the first 12 months. there's a few that I think still haunt me and scare me. We can talk about them later. but it created a, like, a resiliency and persistency in the team, right? with extremely few people have left us over the years, that, like, told us, “Okay, we can do this.” When we first wrote our first training code base completely from scratch, it wasn't a fork of any open source. It was just like, “Okay, let's build it from scratch.” I remember we had this one moment where we spent three weeks working out an optimizer bug. Like, it was like training just couldn't get stable. We, like, obsessed over it, and we thought, like, maybe we were wrong. Maybe we should have just forked this repo, or we should have. But then when we solved it, I still remember at the time we were like five people in the company. when we solved it, we were like, “Oh, we can do things,” like if we're just willing to work hard. and I think that culture with a very strong engineering bias has helped us, like, get to where we were. And so there's this notion of open source and talent and these things. I think we, We just took different decisions from a different starting point. and I think we are lucky. I do want to definitely call it lucky. And there was a lot of hard work at the team that now, like, that's starting to show up in results.Swyx [00:13:52]: Just ‘cause we probably won't revisit this again, but, and this is a fun recruiting challenge if someone knows the answer. What was the bug? And then we won't tell the solution, but we'An Optimizer Bug and the Value of Building From ScratchEiso Kant [00:14:01]: So the - This - You're gonna test my memory here,Swyx [00:14:04]: Oh, okayEiso Kant [00:14:04]: So but I thinkSwyx [00:14:05]: DirectlyEiso Kant [00:14:05]: I think I can recall. So if you, so if you look at, So if you take like Adam as an optimizer, you have epsilonSwyx [00:14:12]: YeahEiso Kant [00:14:13]: Which is, right, like in the denominatorSwyx [00:14:14]: Momentum and weights. YeahEiso Kant [00:14:15]: Is exactly, in the denominator. And at the time, if I recall, you looked at like the early Llama papers and things like that. People were juicing epsilon, like, quite a bit. Like, they were, like, adding, I don't know if it was E minus four or whatever, like a high value for epsilon.Eiso Kant [00:14:31]: And if you think about this during training, it's like a bit weird and counterintuitive that we're adding noise to our optimizer by just adding effectively, like, a random number in the denominator, right? Like behind the decimal point. And I don't recall the exact bug, but it had - What I remember is once we solved it, we no longer had to juice epsilon as much as, like, was happening in the Llama paper and other places. and it was like one of those fundamental moments where we had trusted this paper that was out there, and we're like, “Oh, no, it has to be this way. It has to have this high value of epsilon.” But it made no sense to us intuitively. Like, why do you have to have this so high? Like, if you're just trying to avoid division by zero, why can't the value be extremely small? and that was like one of those moments where you realize like, okay, finding things out from scratch yourself builds a better intuition. Because the one thing you learn very quickly with model building is that your intuitions that you start with are gonna get beaten up so hard.Eiso Kant [00:15:33]: Right? Like - It's such an experimental science, that the things that seem obvious, you very quickly get to learn, like, you were wrong, and hopefully you figure out why, and sometimes you don't even.Swyx [00:15:45]: Yeah. yeah, so, one of the reasons that you, when you released your new models, Vibhu got really excited. I mean, everyone got really excited. But Vibhu led our paper club on it, and you guys sawEiso Kant [00:15:58]: YeahSwyx [00:15:58]: Obviously. maybe talk through some lessons learned in that, whatever you can disclose. we can focus on the model factory stuff, whatever you think is a good starting point.Model Building as EngineeringEiso Kant [00:16:08]: So I would say that our view from very early on in the company was that model building is ultimately 90% engineering.Eiso Kant [00:16:18]: And I think we all know it in the industry because if you look at where's every researcher spending their time, they're spending their time writing code, right? Looking at data and writing code. And so we said, okay, The state at the moment, like three years ago, was bash scripts and Slurm and spaghetti code bases for training and, like, data pipelines that were patched together. And we looked at this and said, “Well, ultimately, model building is a process.” You're going from raw data, right? Like training raw material, the web, et cetera. you're doing a whole bunch of filtering, cleaning up, transformations, analyzing. These days, that's, far more complex than it was three years ago. then you're training a model, which is effectively a large distributed systems problem, right? Across hardware that has still-- It's become a lot more reliable. It was extremely flaky back then. and now with every new generation, we get our new sets of challenges. And then you go into the next stages, right? There was no training back then, but, like, you got, your post-training and then your reinforcement learning. And so we looked at this and we said, “Well, this looks like an industrialized process. This looks like an end process, that every single part of it has its machinery,” right? If it's your big data pipelines, if it's your crawling ingestion of the web, if it's your, large-scale distributed training, and then you've got your reliability. And we said, “Well, why don't we take some of the world's smartest distributed systems engineers that we knew and make them part of the process of research from day zero?” Not retrofitting it later on, but, like, really from the beginning. And that became our model factory. And so our model factory started with a handful of components. Today, it's thousands of components, and I try to equate it to, if you think about, like, someone who was at the very early days of Foxconn, if they had been there for the following, decade, they would be able to rebuild Foxconn because they saw every decision that led to building that system and all the complexity. If you and I walk into Foxconn today, no chance.The Model Factory and Experiment VelocityEiso Kant [00:18:18]: Right? Because we don't have the lineage and history of decisions that led to that. And so we built early on from the beginning- with a team that really understood that, well, the metric that we are optimizing for is the speed of an idea from a researcher to an experimental result that we can trust to then being part of the next model training.Eiso Kant [00:18:42]: And in the. And because it's such an experimental science, ultimately, in the beginning when it wasn't that complex, you could patch your way around it, right? But now, at any foundation model company, you are running. I mean, we're a small team, right? We're less than 70 researchers, another 35 engineers. and we are running, I haven't checked the latest count, but far more than 10,000, maybe 10 to 20,000 experiments a month that we cut. And so if you look at that scale of every model run that is, like it's ultimately it's, it's you need to be able to trust it as an infra problem. And so what we have now done over the years is gotten really good at that, and just by working it and improving it and obsessing over those end decisions. So now what that means is that you looked up Laguna XS 2 that we launched. It was five weeks from the beginning of training to launch. The model that we're gonna talk about today was eight weeks from start of training, to launch. We started the next model literally yesterday because we now finished the post-training required for the model we're launching, next week or by the time this comes out today. and we move that compute to the much larger Laguna M model that we're now training. And so the model should be an artifact of someone's process. It shouldn't be really a thing in itself. Like, and we treat this like the way you would look at like a SpaceX factory where, yes, the first rocket, really hard to build, but the much harder challenge was building the factory. And now they're rolling off, and no one is really thinking about the next launch anymore. So it's just another launch, it's another launch, another rocket comes off. And that's what we're trying to do with model building.Eiso Kant [00:20:22]: And what has been, which was not planned from day zero, it was in the back of our mind like this will happen one day, is that when you build a really good end model factory with really good APIs and really good engineering systems, Well, what is it perfect for? It's perfect for agents.Agents Inside the Model FactoryEiso Kant [00:20:40]: Because agents are now starting to take over more and more work in our model factory.Vibhu [00:20:43]: Yeah.Eiso Kant [00:20:44]: So I look at the screens when I walk, like when we're, we come together, in our monthly, we do monthly onsites, and I walk behind people's screens and I stop by and I talk to our researchers. And the default is all of these different agents running on their screen that are writing the code. They're launching the jobs. They're evaluating the results that are coming back from the model runs. They are, making the changes. And we're still in the driver's seat. We're still coming up with the ideas. We're still helping with the debugging. But more and more, and this is right now very profound on the data side of our pipelines in both pre and post and the synthetic data pipelines, it's starting to become more on the architecture side as well. You're starting to see these twinklings of what RSI is gonna look like.Eiso Kant [00:21:27]: And that's. So when we talk about, like to your question about our models, every talk about the model factory, And my coolest example of these things is always that when we kick off a new run, doesn't matter if it's a training like big run or if it's now a post, like one of 10 post-training versions we do for like release or many experiments, is that at any given moment, the changes that somebody made that they had experimental results from the day before make it into that run.Eiso Kant [00:21:57]: So there's not like a cutoff 90 days before. Like no, it's like literally from that moment because we can now trust the machine enough. And then you also have to invest in the reliability. So one of my favorite metrics about like Laguna S is that there was no call events, Right? Like completely zero. And we haven't had a meaningful call event, like something to wake up for, as far as I recall this entire year. now there is one asterisk to that. In usually the first six hours of launching a new model run, something breaks because you set a config wrong, you made a small mistake, et cetera. So that's usually there's a little bit of intervention, but that's always within like call periods, right? Not on call. And I think that's starting to now compound. So the model we're releasing now, I love it. It's amazing, but we're already onto the next one. and I think that's the way it should be.Laguna, Five-Week Builds, and Zero On-Call EventsVibhu [00:22:50]: Hey, I also just wanna point out, so for context, this was like a month ago. we found it in the tech report, so we just came in with, “Okay, new model's dropped. Haven't heard about it.” We wereEiso Kant [00:23:02]: Yeah, we're very used to doing this every few months.Vibhu [00:23:03]: We're, we're very much like, “ okay, look, it's like, on par with Kimi, DeepSeek, whatnot, the small ones, Gemma level. Oh, it's a very cool paper on what goes into building.” And then we hit this page, right? Like literally page two of tech report is, “This process allowed us to build the small model from scratch to delivery within five weeks applying the lessons”. And then I'm like, oh, this paper is not about here's a tech report of benchmarks and here's how many tokens it was trained on. Like for people that wanna dive more from what we're not gonna discuss on the podcast, it's all laid out here, right? FromEiso Kant [00:23:38]: YeahVibhu [00:23:39]: Custom software that agents can use to interface with training code, training data.Eiso Kant [00:23:45]: Yeah. Well, link the paper correctly, so yeah.Vibhu [00:23:47]: Yeah. All that stuff. read the paper here, but,Technical Report Principles and Streaming Training DataEiso Kant [00:23:50]: But I would like to. I love principles, and I think that is a good starting off point for maybe telling some stories. Maybe we can go one by one past the principles. I'll just call out that Dagster just got bought by a Prefect.Vibhu [00:24:01]: Yeah.Eiso Kant [00:24:01]: Isn't it fun? But yes, I'm very familiar with Dagster. just anything where like they trigger some story.Vibhu [00:24:07]: So, well, I would say, well, experiments code's obvious, but I think one of my favorite things is, I don't know where it is in here, but early on, and I still think this is the case a lot of foundation model companies, people prepare their training data sets, they get packaged up, then they get copied over to a training cluster distributed across all of the nodes, and then training starts.Vibhu [00:24:30]: And we looked at this like three years ago and we were like That makes no senseEiso Kant [00:24:36]: You lose so much time because the moment you have to rematerialize the data set, you have to make a change, you have to fix something, et cetera, you've got all this time of like repackaging it, right? Toca- tokenizing it, repacking it, moving it over to a cluster, then distributing it across the nodes. The bigger your clusters are, you start using fancy like torrent-like algorithms to like distribute your data. So why aren't we streaming data into training? Right? Something that's very common and like just basicVibhu [00:25:00]: Like just in timeEiso Kant [00:25:01]: Just in time, like good computer science like principle. And that was one of the first things that I think unlocked - the model factory. Because the moment you start thinking about, well, a training job, it doesn't matter if it's a big hero run or a small like, post-training experiment, consumes a certain number of tokens per second, right? And it's not a lot, right? From a like a data, moving data perspective. So we said, well, we have our training cluster, and then we've got like our AWS kinda setup where we can build these amazing big data pipelines. We can set things up. We use Spark underneath the hood, like all these things.Vibhu [00:25:36]: But when you say AWS, it's not actual AWS, it's your internal AWS.Eiso Kant [00:25:39]: It's our internal-- No, it's our internal like just running like our infrastructureVibhu [00:25:42]: Site web servicesEiso Kant [00:25:43]: Exactly. Our stuff running on like an AWS account or on like any hardware, right?Vibhu [00:25:47]: Yeah.Eiso Kant [00:25:48]: And so once we made that shift into I can stream data into training, all of a sudden you realize a lot of things unlock. Because now you don't have to wait for the whole data set to materialize.Immutable Data, Experiments as Code, and Scientific RigorEiso Kant [00:26:00]: You now all of a sudden when you're running data experiments about mixing data, it's a config. Because you've got these data sources that are coming in, and you just - we have this service called Blender that's in the report, where we then say, “Okay, for this run, I want 20% of this source, 10% of this source. I want this much, so many epochs of repetition. I want this to be, shuffled in a certain way,” and your training job can start while the rest of the data is even still materializing. also what it does is because all of this underneath-- So for us, we treated the data layer underneath as like an immutable data layer, and that was really important. Like experiments as code, immutable data layer means that you can always go back and understand literally down to the single token at which cursor it went in on which version of the code.Vibhu [00:26:47]: Yeah.Eiso Kant [00:26:48]: And it took us a I have to admit, like the first year of Poolside, we understood that engineering had to get great, But we didn't understand yet, that this is ultimately in support of like a good rigorous scientific progress. We were quite a - We were a very small number of people, so a lot of it was YOLO ideas and YOLO runs.Vibhu [00:27:08]: Yeah.Eiso Kant [00:27:09]: And we built great infra for the YOLO runs. But once we realized that we treated data as immutable and code as always versioned, and you could always track and trace every experiment end to end perfectly, you could repeat everything perfectly, right? You have perfect reproducibility. I can still reproduce runs from two years ago if I wanted to, right? It enables the scientific progress, like the scientific process, and I think that took us probably about a year and a half into the company to figure out. We also had some great hires, like our head of applied research, Nikolai, who joined us from Yandex, who'd been working on language models since like the early 2020s, I think brought that into the company of like, “Hey, we wanna have even more rigor.” And then once we kinda had the combination of like increasingly more capable platform that allowed people to do more, but had this immutability, we were able to start “Okay, every experiment is truly an ablation. We truly need to understand it.” And I think we became much more scientifically rigorous in the last couple of years, and the infra underneath enabled it. and then there's just fun stuff like, andVibhu [00:28:16]: Yeah, a lot of it's fun, like even just the, one, you share all the ablations, two, picking the data sets, right? There's like a random small paragraph in here where it's just like, “Oh yeah, training data, we have some, we have an auto mixer.” it trains eight small models, scales them up, picks the training data set. We don't even need to look at it. I'm like, “Wow, a lot of engineering rigor there.” And there's just, there's just a lot in here.Publishing Research and Giving BackEiso Kant [00:28:40]: Yeah, and it'- and look, and we wanna put out more. Like we, We treat writing papers as something that we haven't earned the right for yet for a long time. So you earn the right to spend time, publishing research once you're at the frontier, because until then, you're catching up, and every minute and hour in this industry matters. Like I obsess over it, not just the wall clock time from idea to result, but just general like time every day that we, waste is one that doesn't allow us to catch up. But in this case, we said, “Okay, we're gonna give ourselves.” I think we gave the team like three or four days while still doing their work, like give everything in there. And to your point earlier, if your stuff, it's easy to like put it out. And so there's so many more things that we wanna talk about over time, and we will definitely start doing. And as we earn more of the right, but also now have like added to our mission that we want more foundation model companies to exist, you'll see us like be way more proactive, and just trying to keep dropping some of those like things that we've learned along the way that can help others like speed up.Vibhu [00:29:40]: Which is the other cool side of this, right? It's, it's not like, back to your point, it's not just here's the benchmarks of our training. If you want to replicate, here's experiments of optimizers, data sets, post-training. you lay out a lot of it here alongside here's your system for how to do it? So it's, it's really like promotingEiso Kant [00:29:59]: No, thank youVibhu [00:29:59]: Other people can do the same.Eiso Kant [00:30:00]: And by the way, I also wanna make clear, right, we have been incredible-- Like we've taken a lot of advantage of the fact of all the open research that others have published, Right? And you mentioned, the Chinese labs, and we I think it's important that there's, from every country and every culture and background, including like Western companies like us, there's different models that come out that people can choose to trust. But I think we do have to give credit where credit's due, right? The incredible Chinese lab have done an amazing job at sharing their research, and we have definitely like been on the receiving end of taking advantage of that. So when you're on the receiving end of something coming to you, I think it's, you also have an obligation to give back.Swyx [00:30:39]: Do you have a favorite or underrated Chinese lab that you wanna shout out? Everyone shout outs DeepSeek.Chinese Labs, Zhipu, and PersistenceEiso Kant [00:30:44]: That's a good question.Swyx [00:30:45]: Moaan obviously for Therapsi. Yeah.Eiso Kant [00:30:48]: Yeah, look, I think, I think obviously everyone's been talking about Zhipu lately, with 5.2. I think what most people don't realize is when they started.Swyx [00:30:59]: Yeah.Eiso Kant [00:30:59]: Right? They started years before ChatGPT.Swyx [00:31:02]: They just rebranded. YeahEiso Kant [00:31:03]: And so, I've like, I remember how hard it was to work on these things Before the rest of the world got excited about it. And so I have an immense amount of respect for people, who were working on improving models when it wasn't the sexy thing to do, when believing in LLMs, was gonna get you ridiculed. I remember like back in 2016 when we were doing what we'd call, machine learning on code with some of these models. we would-- people would just laugh at us, like they'd be like, “This makes no sense. Like why are you wasting all these, like, millions of dollars on trying to figure this out?” And so I would say they're probably the one that, I think deserves a shout-out, not just because their latest model is very good, but because they fought to get here. And I think, I think every foundation model company it takes time to get here, right? It took us three years to get to the model that we're, that we're now gonna be releasing. and now the time in between the models is coming, is counted in weeks. It's no longer counted in months or years. But this stuff's hard. and if we can make it a little bit easier for the next person, like we should all do so. Because if we don't do so, we're, we've got a small window before models are really impacting recursive self-improvement to a level where catching up otherwise might become unfeasible. And we should try to, in that window, encourage as many labs or however we wanna call them, like to start. And so one of my currentEiso Kant [00:32:36]: Mission, but qualm is like I wanna encourage whoever is a researcher right now who thinks they can tackle this to go and leave and become my competitor.Eiso Kant [00:32:45]: Like start another foundation model company because I think we need it. I think otherwise we're not gonna be in the world where, I don't want to just be the fifth or the sixth company that wins. I wanna look at a world where there's lots of choice.Starting a Foundation Model CompanyVibhu [00:32:57]: What else do people not see in starting a foundation model? it's, there's a lot of compute, there's a lot of capital required, a lot of compute. You lay out model factory and how to do the training, but there's a lot there, right? That's,Eiso Kant [00:33:10]: Well, look, it's, I in turn-- this is an oversimplification, and I always asterisk it with that because it can land a little bit the wrong way in people's minds. But I think you can sum down, And I saw it, 95% of model building to just doing, you're just doing two things. You're improving data or you're improving compute efficiency. And I know that feels like an oversimplification for the incredible, like, Gifted and skilled work people do. But if you really look at it, like what are we doing? We are looking at data, we're generating new data, we're improving data. and the only way to do that is to look at the data, right? That's a big part of foundation model building. And on the other hand, we come up with these incredible breakthroughs in inference, in architecture, and new attention mechanisms. But what are they really doing? They're bringing compute efficiency. Now, we have definitely had some breakthroughs over the years that allow for more model capabilities. But at the limit, if you could train a large enough model, right, like, and you had infinite compute, we probably-- if you had infinite compute, you'd be at AGI probably already tomorrow.Eiso Kant [00:34:12]: Right? Like it's not. And so, and let me say that infinite compute with infinite ability of much faster networking because networking ends up being more of the bottleneck than compute. But, so I do think that's, those are the main things. And to just realize that this is engineering. I think it's become more obvious, but I think for quite a few years, people have held foundation model companies and researchers and others on this pedestal of like you're doing incredible magic or rocket science, or only like, Nobel laureate physicists can do this. And don't get me wrong, there are some really hard problems that need to be solved, but a lot of the work that all of us are doing on a day Is not sitting down trying to solve a math theorem. A lot of the work that we're doing is just really doing the basics right, writing good code, looking at data, improving it, running experiments, looking at plots, trying to see like, hey, trying to shape our intuitions. And a lot more people could be highly capable researchers. and I think that's, it feels far for people to do so. But I've seen in our own company, we've seen engineers become researchers because the model factory allowed them to be, have a much lower hurdle of running experiments and trying things. And one of the guys on our team who started as an engineer building our agents is a legit reinforcement learning researcher now, making real progress. and that happened in the span of like six months. that would've not been what I think most people assumed was possible, a couple of years ago.Swyx [00:35:46]: Yeah. I think one of the interesting moments is when you can self-host, like, if in a programming language, like if you can compile the language in the language, the equivalent is can you use your own tools, right? You have the pool CLI, you have your own models. presumably you're not only using your own models. There's no way. But like, what's that percentage over time?Laguna S, Persistence, and Behavioral GainsEiso Kant [00:36:10]: This is the first model that we're releasing that is starting to meaningfully contribute to our own work. It's not a it's not state-art model yet. Fable and other, they're, they're very capable models, but Laguna S Is really interesting. I'm gonna pull up the quote. Peng Ming, one of our heads of applied research, said something, last week as the model came out about 10 days ago, much better than we had hoped for or expected. And he said, I have the feeling that a lot of the gains in Laguna S come not from more intelligence, but more from different behavior, more verification, less taking things for granted, not declaring victory early, and being way more persistent. And to be honest, those are more predictive than raw intelligence for success in human also to some degree. And this was, he wrote me this on 5th of July on a Sunday, and it's been burned in my brain ever since because the Laguna S model, as you'll see it and why it does so well on benchmarks and why it does so well in using it on a day basis, is that it's just incredibly persistent. It reasons a lot. I do call that out. We have work to do on making it more efficient. We have to work to do on offering different reasoning modes. But this is the model that has been able to do things that I never thought it could do. A hundred eighteen billion 8B active model, which is not that large. It fits on a DGX Spark and still runs at, thirty, forty tokens a second on a Spark, is able to solve Erdős 397 independently. It's able to do complex programming tasks. It's able to. I asked it this morning to make me a Fi scanner without using any external libraries on my Mac, and it's, like, figuring out, like, the core WLAN API by really persistently trying to understand it without access to the internet. And more, I love vibe checking. I've probably spent eight to ten hours a day with this model for the last ten days.Eiso Kant [00:38:05]: I'm not exaggerating. I was on my eleven-hour flight yesterday. I spent ten hours reading trajectories and traces and, like, of the model.Eiso Kant [00:38:12]: And what I take away from it is exactly what Peng Ming said. We are gonna be able to squeeze so much more out of smaller models than I think we had imagined in the industry because, yes, there's intelligence and larger models are more intelligent. Like, no doubt about it. We should continue to scale up. but the behaviors of being really persistent, of being able to backtrack when you're wrong, of, like, understanding how to interact with your environment show us that we can get a lot more out of it. And this, for me, has created a bit of a Question in my mind the last couple of days. If you think about where we're using models today, right? We are using models, say, for knowledge work. Represents twenty-five percent of the global economy, twenty-five trillion dollars of work.Eiso Kant [00:39:00]: As we scale up models and they become more intelligent, we are excited about using them more and more for pushing the frontier of science.Small Models, Knowledge Work, and CommoditizationEiso Kant [00:39:08]: And if you look at the frontier of science, like true breakthroughs in science, they have been linked, they are linked to more intelligence in many places. Einstein figuring out general relativity is able to bring ideas together that other people would have not brought together. And I think one of the many dimensions of intelligence is the ability to do that, and it's something we clearly see that as models get larger and more capable, they're able to pull more ideas and threads together that a smaller model wouldn't be able to.Eiso Kant [00:39:36]: And we're starting to see examples of that in medicine and, like, in bio and other things. But if you think about the majority of knowledge work that we do, and it includes building software. I'm a software developer at heart first and foremost probably, although I probably can't say it that much anymore as I don't write production code in years, is that what makes us good is our persistence. It's our ability to encounter a problem and backtrack and say, “I need to go figure out this bug. I need to go research this. I need to go look at the documentation. I need to, like, try different, five different ways to see, like, if I can solve it.” But it is not necessarily bringing three ideas together from radically different fields. And so if we are now seeing, and I think Laguna S is an example, that we are able to make a relatively small model much more capable than I had definitely predicted or any previous, like, benchmarks had shown for any model remotely this size or even larger, At least on coding tasks, that it's because of the behaviors. And so now the question I have, and I don't have an answer, it is I know at the limit, so infinite model size, right, extremely large model, and the cost of that model is gonna be very expensive to run. We know this, right? So larger model ROI.Eiso Kant [00:40:52]: So I know that at the very limit, I'm not gonna use the world's largest model one day, quadrillion parameter, whatever crazy, like, scale we scale up, to do a basic coding task. Already today, I'm starting to size down for certain tasks.Eiso Kant [00:41:07]: So it means that there is an optimal. It means there's some curve that goes as we go up to model size for knowledge work, at some point we're at the peak, and after that, the return on investment of using a bigger model, just doesn't make sense.Eiso Kant [00:41:22]: Now, I think the question is, before I would have thought that peak was extremely very far away.Eiso Kant [00:41:30]: This model for me is the first sign that Maybe that peak is At a trillion, five trillion, ten trillion. Maybe we can just squeeze way more out of these models. I'm no longer thinking that we need two or three orders of magnitude on the largest models to be able to, solve knowledge work, the accounting, the legal, the code that we write. And so if that holds true, It is an argument for the commoditization of models. It's an argument that open source can win and, like, succeed in this world. And now it's of course a self-serving argument and it's a hopeful argument, but theoretically at the limit it works. We just have to go discover in the next couple of years of how much more we can squeeze out. Now, I do want to put a big asterisk. This does not mean I'm against scaling models. I think we ultimately only succeed if we scale our models as large as our competition. I do not like. I think we should not put our head in the sand and say we're gonna be king of open source small models. I think that's, It's a out. It's trying to be king of your own kingdom, but not realizing what the rest of the world's doing. All of us rather use a smarter, faster, more model. It's a sign of hope. And so I don't wanna overly state this is a good model. We have a long way to go to get to the state-art. But what hopefully people take away when they use this model is that the behaviors inside of it are what push it to be far more capable, less than necessarily the number of parameters.Pre-Training, Mid-Training, and RL Moving EarlierVibhu [00:43:03]: Is that mostly post-training? LikeEiso Kant [00:43:05]: YesVibhu [00:43:05]: Right.Eiso Kant [00:43:06]: It's entirely post-training.Vibhu [00:43:08]: Are we done improving anything on training? Is, like, training done?Eiso Kant [00:43:12]: No.Vibhu [00:43:12]: Okay.Eiso Kant [00:43:13]: SoVibhu [00:43:13]: I just wanted to cover training, and then we go post-trainingEiso Kant [00:43:15]: Training is not done. I mean, look, there's a part of training of just dealing with skill, right? Every new order of magnitude of model skill, you are going to get new things you gotta solve for. That'- but those are ultimately, engineering challenges.Eiso Kant [00:43:31]: I have a, I would say, a not commonly held opinion that reinforcement learning Will move earlier and earlier into training.Vibhu [00:43:42]: Yeah, training.Eiso Kant [00:43:44]: Not even training. Like training today, right, is, like if you look at - So we've been working on this for years already. and I think the best-- I think the first time we saw it out in public was the DeepSeek Zero paper. this is a year and a half ago, I think, if I recall correctly. where, you can Very early on in a model as it starts capable of being able to use language, et cetera, induce reasoning. and so the question that I have is like, we have this- we have the dataset that's the web. and the web, I think we could arguably say probably has The totality of humanity's knowledge somewhere encoded in different places. It's a huge variance degree of quality, from garbage data, and like once you look at training data, you really get humbled of like what the web is, to like, the most greatest scientific papers and best blog posts and like, best transcripts and whatnot.Eiso Kant [00:44:39]: And so now What we are trying to figure out, and have been doing a lot of work on, and it's a place where maybe not as open as we're on other things, but we will become more over time. we've been spending a couple of years really doing research on how can we turn the web into not just next token prediction, but into a way to teach the model to think earlier in its training. and I think there's a huge amount of gold to be found there. I think we are right now in, we've got some drugs in the industry. One of the drugs is distillation. Another drug is, more environments. Like, and they're great, and they make us feel good, and they make the models better, and like we're all addicted to them, and we'll use them, right? in various different ways. and but ultimately, I think we are still barely squeezing out of the web what we should be getting out of the web.Eiso Kant [00:45:33]: I think just next token prediction during training is not enough.Eiso Kant [00:45:36]: AndVibhu [00:45:38]: YeahEiso Kant [00:45:38]: I think we'll see some very interesting things still happen. and that RL in post-training to induce behaviors, to improve things, like I think - the whole world knows how to do this now. I think we're, we're scaling it up. Everyone is. But I wonder if we need to go as far as we're going today with environments. I'm not sure yetVibhu [00:46:01]: You mean we're going too far?Eiso Kant [00:46:02]: I'm, I'm not sure if the path to AGI is justVibhu [00:46:06]: Is more environmentEiso Kant [00:46:07]: More environments.Vibhu [00:46:08]: It seems like a never-ending, “Okay, I want instruction manual for this table, right? Am I gonna environment out building furniture? Or are we just gonna tail end like we need some general solution?”Eiso Kant [00:46:19]: I think there is, I think there's an ability to generalize more from the web. but I also am very encouraged, like when I look at Laguna S and, which is post-training is, well, is the big impact there. and I see like, oh, wait a second, just by making some of these behaviors much better, we're able to get so much more out of it. It just changes a little bit the way you think about intelligence.Vibhu [00:46:40]: Yeah. The analogy people draw often is the RL phase is where you don't learn as much new knowledge. You shiftEiso Kant [00:46:46]: Yeah.Vibhu [00:46:46]: Yeah. So, you shift distribution, and you can have it reason towards what you want. on your point about training, a lot of training is still just continue training in a domain, say medicine, then you do RL. So still justEiso Kant [00:47:00]: It's just better data, right? Like, I mean, training, ooh, I like how we invented this word. Like it's effectively just like,Vibhu [00:47:06]: Second phaseEiso Kant [00:47:07]: It's the second phase of training With like a really dumb way to do a curriculum. But like ultimately, what you'd want is a curriculum from token zero to token 30 whatever or 40 trillion tokens that really truly is the optimal curriculum for the model to learn. But training is essentially a stage curriculum on the web because we do not have to compute, And, effectively to try to ablate the perfect curriculum, right? And so I'm pretty sure that you'll start to see people talking soon about some other term, and there's two or - ‘cause now we do this, right? We talk stage two and stage three and stage four training and like. But ultimately, all we're doing is we're trying to assign a curriculum to the web data that we have to allow the model to learn better. I think at some point, as things get compute, as models get cheaper to run, as the next generations of compute, this will become more of a continuous spectrum. I also think the reason, by the way, you have training and like stage two and stage three is organizational, Right? It'- this is, I think, a thing where-- that we really try to avoid with the model factory is like Training exists because there's a training team now, right? There's people, or like people in training decide to focus on like a training effort. but what you really want is engineering and scale of experiments that allows for a much more continuous spectrum that you don't, you have infinite stages. Now, we're not there. Compute's not there. Organization design is not there for it yet. but I think we'll get there. we'll look back on a couple of years and be like, “Oh my God, it was so cute that we did our training data like this in such a like naïve way. Like we barely ordered it. We didn't really do a good job at likeCurriculum, Auto Research, and New ObjectivesVibhu [00:48:48]: The building that curriculum will get you that in the industry.Eiso Kant [00:48:51]: And I'll confirm that, when I talk to some researchers that this is a lot of the focus now is like how does training change and what is the next objective other than, next token prediction. I assume you don't have the answers, but you have some ideas.Vibhu [00:49:02]: We have some ideas. We're not ready to talk about it yet.Eiso Kant [00:49:05]: Yeah.Vibhu [00:49:05]: We've been working on them for years, and I think that's the one thing that's also like you asked earlier about, like what's not obvious about building a foundation model company is that you are constantly balancing the table stakes work, the recipe worksEiso Kant [00:49:19]: Yeah.Vibhu [00:49:19]: Versus like your, my crazyEiso Kant [00:49:22]: Pure researchVibhu [00:49:22]: Breakthrough.Eiso Kant [00:49:22]: Yeah.Vibhu [00:49:22]: Pure research and finding that balance and adjusting the percentage to it based on where you are in the race is really important.Eiso Kant [00:49:31]: I mean, so like, this is a nice way. I was gonna bring up auto research at some pointVibhu [00:49:35]: YesEiso Kant [00:49:35]: As another Andrej invention, or coinage, which is like, I honestly, like how many objective functions can there be, right? Like just try 1,000 of them, set it running, whatever.Vibhu [00:49:47]: Man, it's alsoEiso Kant [00:49:48]: Like what you're looking for. You're looking for loss curves like that, likeVibhu [00:49:51]: It's also a thing people take bets on, right? When you say more Neo labs, you're doing a version of we'll do foundation models, scale them up, next token predictors. A lot of other Neo labs that we see want to take a completely different approach, right? At some level, you're right. It's all, compute efficiency, and that's the net objective. But some are okay, different architecture, like vastly different amounts of compute spend. So some are different. They're not justEiso Kant [00:50:19]: YeahVibhu [00:50:19]: They're like, 99% not balancing, here's the vanilla and scale up. They're 99% on, here's novel research that'll change everything.Eiso Kant [00:50:27]: And I think, Luke, I think you. It depends when you started as well, right?Pure Research vs. Table StakesVibhu [00:50:30]: Yeah.Eiso Kant [00:50:30]: When we started, like the novel thing we did was reinforcement learning on code. No long- that's no longer novel by far, but we were like, - that's where we obsessed over when no one believed in RL. So you have to when you start the company, you have to have your own idea. You have to have something that's different that allows you to speed up, right? For us, it was RL to LLMs that later became common, like, Knowledge. But in the beginning, it wasn'tVibhu [00:50:53]: It's cool. this was like your original 2023 blogEiso Kant [00:50:57]: YeahVibhu [00:50:57]: Of purpose.Eiso Kant [00:50:58]: Yeah.Vibhu [00:50:59]: And like you do lay it all out here.Eiso Kant [00:51:01]: We laidVibhu [00:51:01]: The blog is pretty underrated, right? The whole RL on code was very early on.Eiso Kant [00:51:06]: Very early. And even we had to argue with people, like we say here things like to push beyond current capability, to train your own foundation model. We had to argue with people that it mattered that you had your own like, base model. you can fine-tune your way to success, right? major capabilities emerge from training a base model made accurate and useful during fine-tuning.Vibhu [00:51:23]: Which like, for perspective at the time, we knew closed models, OpenAI, Anthropic were huge. The open models we had were like Mistral 7B, a 30B, a 70B.Eiso Kant [00:51:35]: When weVibhu [00:51:35]: YeahEiso Kant [00:51:36]: The date on this thing is wrong. When we published this, it was April 2023. I think this was justVibhu [00:51:42]: YeahEiso Kant [00:51:42]: Happened on a migration, probably found it on archive.org.Vibhu [00:51:45]: Mistral.Eiso Kant [00:51:46]: Mistral had started, we started on the same month, right?Vibhu [00:51:49]: Yeah.Eiso Kant [00:51:49]: So this wasn't even, there was only, I think, Llama out at the timeVibhu [00:51:52]: SnellEiso Kant [00:51:52]: And that's it, right? And so, but I agree. I think we wan

The MAD Podcast with Matt Turck
The Biggest Chip Ever Built — Why OpenAI Runs On It | Cerebras CEO Andrew Feldman

The MAD Podcast with Matt Turck

Play Episode Listen Later Jul 23, 2026 72:41


AI is no longer just a race to train smarter models. As AI moves into production, the bottleneck is increasingly inference: how fast models can generate tokens, use tools, reason, verify, and act. In this episode of the MAD Podcast, Matt Turck sits down with Andrew Feldman, co-founder and CEO of Cerebras, to explain why fast inference may define the next era of AI.Cerebras is known for building a chip the size of a silicon wafer. But this conversation is not just about one company or one chip. It is a deep dive into the AI infrastructure stack: GPUs, ASICs, memory, HBM, SRAM, data centers, power, TSMC, AWS, OpenAI, agents, reasoning models, and why speed changes what AI products can become. Andrew explains why “tokens per second per user” matters, why generating a single word can require moving the equivalent of 100 HD movies through memory, why agents amplify latency, why GPUs struggle with certain inference workloads, and why fast AI may eventually reshape SaaS itself.This is a reference conversation on fast inference, AI chips, and the next compute bottleneck.(00:00) Cold open & Intro(01:31) Why speed became the AI bottleneck(02:32) Tokens per second per user, explained(03:16) AI's broadband moment and the Netflix analogy(04:35) The AI chip landscape: GPUs, TPUs, Trainium, ASICs(06:36) What is an ASIC?(08:08) Nvidia, Groq, and the fast inference war(09:16) OpenAI, Broadcom, and specialized silicon(12:10) China, power, and sovereign AI infrastructure(15:05) Is the AI infrastructure boom a bubble?(18:56) The hidden bottlenecks: HBM, CoWoS, and 3nm(22:57) Why agents are creating CPU demand(25:36) Andrew Feldman's path from SeaMicro to Cerebras(26:13) Why Cerebras bet on AI in 2016(31:14) SRAM vs. HBM: why inference is a memory problem(33:19) What wafer-scale computing actually means(34:28) The deep-tech “Everest” problem(36:07) The moment the first Cerebras system worked(36:49) Ringing the bell and surviving deep tech(39:08) How a giant chip handles failure(41:22) Why GPUs struggle with decode(42:17) Prefill vs. decode explained(44:01) The “100 HD movies” problem in AI inference(45:04) How fast inference changes RL and training(48:08) Reasoning models and why they cost more compute(50:08) Verification, guardrails, and small models checking big models(52:37) Multimodal AI and the path to video(53:51) Cerebras' business model: hardware, cloud, and API(55:14) OpenAI's 750MW inference deal(55:36) Why data centers are measured in megawatts(58:01) AWS Trainium + Cerebras decode(59:29) Fast tokens as a cloud product(01:00:52) Is CUDA still a moat?(01:03:53) How TSMC helped Cerebras build the giant chip(01:07:41) Why nobody cared in 2020(01:08:15) Why chip supply chains are hard to diversify(01:09:54) Why today's AI models will be the worst you ever use(01:10:38) What fast AI could do to SaaS

Black Box
Altri raid, Brent a 79$. Nikkei +2%, SK Hynix a ruba. Unicredit scala Commerz | Morning Finance

Black Box

Play Episode Listen Later Jul 9, 2026 26:44


9/7 Futures in verde, altri raid nella notte. Centcom: colpiti 90 obiettivi. Brent a 79$, salgono prezzi jet fuel. Hormuz ferma, la Russia blocca export diesel per tre settimane. Rotazione o ritirata della marea? Cosa dicono gli strategist. Nvidia: mai così a sconto dal 2019. Nvidia +4%, da Cina via libera export H200. Apple e Broadcom accordo da 30mld dollari.Spacex scende sotto Ipo, lancia nuovo modello Groq con Cursor. MIT: L'AI costa più degli esseri umani che rimpiazza: conveniente solo nel 23% delle mansioni. SK Hynix stasera prezzo, domani Ipo: domanda supera di 7 volte offerta. ****Questo episodio è offerto da ⁠Scalable Capital ⁠ Apri un conto con Scalable Capital e inizia a ricevere il 2,5% di interessi* sui tuoi risparmi:  https://it.scalable.capital/broker-online?utm_medium=affiliate&utm_source=qualityclick&utm_campaign=broker&c_id=QC59486e7f67706c777b517d435049607362766c747c5aS7541p&utm_term=983 Messaggio pubblicitario. Tasso lordo annuo variabile sulla liquidità depositata nel conto deposito non vincolato, composto da tasso base collegato al Tasso di Deposito BCE e tasso bonus discrezionale. Liquidità allocata presso banche partner e fondi monetari riconosciuti. Foglio informativo e condizioni su scalable.capital. Investire comporta dei rischi Asia Nikkei in verde grazie a tech. Bain esce da Kioxia. Kospi volatile con Samsung e SK. Decennale giapponese 2,9% mad da 30 anni. Cina inflazione in calo a giugno +1%, Sale PPI max da quattro anni 4,1%,  Europa in verde, oggi minute Bce e Eurogruppo. Fmi crescita debole, per Italia 0,5% 2026 e 2027. Unicredit scala Commerzbank, Credit Agricole non vuole scalare Banco BPM. Ferrovie lunedì nuovo Ceo. Learn more about your ad choices. Visit megaphone.fm/adchoices

Latent Space: The AI Engineer Podcast — CodeGen, Agents, Computer Vision, Data Science, AI UX and all things Software 3.0
Why AI Infrastructure must evolve for Agent Experience — Akshat Bubna, Modal CTO

Latent Space: The AI Engineer Podcast — CodeGen, Agents, Computer Vision, Data Science, AI UX and all things Software 3.0

Play Episode Listen Later Jul 8, 2026 57:55


We've been running a bit of an Agent Cloud series surveying all the top inference/compute/cloud providers, from Databricks to Daytona to Railway and, even further back, E2B, but we're excited to conclude this series returning to Modal, which has just raised a monster $355M Series C.The cloud was built for developers. But agents are now changing that.The old infra stack was designed for a human who could read docs, reason through YAML, and understand dashboards to figure out what they need when something broke. While this was painful for developers, it worked since they could fill in missing context in their heads.However, agents don't have that luxury. Now in this new era of agents, everything has to be tighter.They need a place to write code, run it, inspect the output, change the environment, debug failures, and try again. Fast iteration and feedback loops with all the necessary context are crucial for agents to operate properly. Furthermore, sandboxes are a clear representation of this shift as agents can easily spin up isolated environments. This programmatic infra even extends to research:Two years ago, we were one of the first to cover Modal with CEO Erik Bernhardsson and Alessio designed our favorite LS thumbnail of all time:At the time, Modal was just a teeny little company with a $17M Series A.Today, fresh off their $355M Series C, Modal is one of the clearest examples of the agent cloud future being built in real time: a cloud platform moving past traditional web app assumptions toward the workloads AI actually creates such as elastic inference, sandboxes, GPU burst, post-training, background agents, and infrastructure that agents themselves can operate.In this episode, Modal CTO Akshat Bubna joins swyx and Vibhu to unpack why AI applications don't fit traditional cloud assumptions, why Kubernetes was never designed for bursty compute-heavy workloads, and why Modal is now shifting from developer experience to agent experience.We go deep on Modal's AI infra stack: serverless functions, decorator-based infrastructure, elastic inference for custom models, GPU snapshotting, DeFlash, speculative decoding, Auto Endpoints, sandboxes, persistent storage, networked containers, private IPv6, RDMA, multi-node training, and Modal's capacity pool across 17 cloud providers. Akshat also explains why RL rollouts can require 100,000 sandboxes, why production agents need hard guardrails, why observability may matter more than reading code, and why AI has made infrastructure exciting again.We discuss:* Why Kubernetes wasn't built for bursty AI workloads* How Modal started as a better runtime before becoming an AI cloud* Why Modal added GPUs before ChatGPT* The shift from developer experience to agent experience* Why observability matters when agents are writing the code* Elastic inference for custom models across audio, video, robotics, and comp bio* GPU snapshotting, cold starts, and why inference workloads are so bursty* Why RL rollouts can require 100,000 sandboxes* DeFlash, speculative decoding, and frontier-level inference performance* Auto Endpoints and making optimized inference easier to deploy* What Modal adds beyond vLLM, SGLang, and raw GPU rental* Modal's 17-cloud capacity pool and supercloud strategy* Networked sandboxes, sidecars, private IPv6, and RDMA* Serverless multi-node training for post-training and research workloads* Auto-research, model-guided sweeps, and agents launching GPU experiments* Compute strategy, capacity planning, and batch tiers* Why production agents need specialized sandboxes and hard guardrails* Modal's take on managed agents, CI, Gitpod/Ona, Python, TypeScript, and Modal BenchAkshat Bubna* LinkedIn: https://www.linkedin.com/in/akshat-bubna-188885103* X: https://x.com/akshat_bModal* Website: https://modal.comTimestamps00:00:00 Introduction00:00:39 Modal's origin and why Kubernetes wasn't enough00:04:32 Developer Experience → Agent Experience00:06:21 Modal's AI cloud primitives00:09:14 Sandboxes, agent loops, and proto-Cognition00:12:12 Elastic inference, GPU snapshotting, and 100,000 sandboxes00:15:24 DeFlash, speculative decoding, and Auto Endpoints00:19:59 Production-grade inference beyond raw GPUs00:22:00 Background agents, Ramp Inspect, and the agent lifecycle00:24:08 Modal's 17-cloud supercloud strategy00:26:40 Networked sandboxes, private IPv6, and RDMA00:32:48 Multi-node training, post-training, and auto research00:37:36 Compute strategy, capacity planning, and batch tiers00:40:55 Open models, real-time AI, and production agent infra00:43:06 Hard guardrails, managed agents, and specialized sandboxes00:46:06 Why AI made infrastructure exciting again00:48:30 Model APIs, differentiated products, and agentic video00:51:50 CI, coding-agent infra, SDKs, and Modal Bench00:57:28 Closing ThoughtsTranscriptIntroduction: Modal, Series C, and the Art PartySwyx [00:00:00]: We're here with Akshat, CTO of Modal, together with Vibhu. Congrats on your Series C.Akshat [00:00:10]: Thank you.Swyx [00:00:11]: Your party yesterday was amazing.Akshat [00:00:15]: Yeah.Swyx [00:00:15]: From all the photos and all the swag.Akshat [00:00:17]: We had a bunch of art installations, which was fun, seeing, like, our products on pedestals next to, like, Rodin.Swyx [00:00:25]: Very nice. Very nice. When you started, it was not the GPU inference company. Maybe it was in your mind. Take us back to the origin story.Modal's Origin: A New Runtime Beyond KubernetesAkshat [00:00:39]: I first met Eric, who's the CEO, through an investor. Back then Eric was already thinking about building, a new runtime, and he got there thinking through why are workflow orchestration products so hard to use. It's because you have to run them on Kubernetes. Kubernetes is hard to manage. It's not built for burstiness and, custom images,Swyx [00:01:03]: YeahAkshat [00:01:03]: It has a terrible developer experience.Swyx [00:01:05]: And I'll, I'll interjectAkshat [00:01:06]: YeahSwyx [00:01:07]: For listeners, who are new, we interviewed Eric two years ago, and there's a bit more of the story there from Spotify and all those things.Swyx [00:01:14]: And I came across Eric through Data Council because he did that talk on the serverless container stack that you guys did, which was like, that was my first like, “Okay, I need to take Modal very seriously” moment.Akshat [00:01:26]: Yeah.Swyx [00:01:26]: But it was still very unclear, like, do I need all this for just my data pipelines?Akshat [00:01:33]: Yeah. initially what we were thinking about was if we build a better runtime, it's a very useful primitive in itself. It's There's a lot of things that, get solved by serverless functions, like you can do, ETL stuff, you can do job queues, you can do all this, like, bursty processing, which it turns out every company had needs for. but then we also were thinking about this as like, this is a primitive that we can build a whole collection of products on, which are very verticalized. So perhaps data engineering would've been the first one, but we were thinking about inference. Back then it was more classical inference, like computer vision stuff and running XGBoosts and whatnot. But we added GPUs to the product a year before ChatGPT came out.From Serverless Containers to GPU WorkloadsSwyx [00:02:19]: Nice.Akshat [00:02:19]: We just didn't think it would be that big of a deal.Swyx [00:02:22]: Yeah, just like add A100.Vibhu [00:02:23]: Was there any, like, early key problem that really sparked off why you built it?Akshat [00:02:28]: Yeah. Primarily it's just, none of the tooling that was out there was built for, one, a really great developer experience, and also there's a general trend of, a lot of the workloads that we were seeing were very. I wish there was a better word for it, but compute-heavy. Like, they need, one, like, need a lot more resources, so you need to burst up and down a lot, versus like Kubernetes designed for, like, slow scaling and, more for, like, web server use cases. And also there's just a lot more specialization in, like, what kinds of environments these workloads run in. Like, we had sometimes they need accelerators, sometimes they need different kinds of images, and this is just like a consistent thing that we saw across a lot of companies. That would be the next step.Software-Defined Infrastructure and Decorator-Based DXSwyx [00:03:13]: Yeah. Yeah. Be nice. I don't know how much this factored into the early story, but I wrote a post when I was at Temporal about infrastructure, software-defined infrastructure or something like that.Akshat [00:03:22]: Yeah, the self-provisioningSwyx [00:03:23]: Self-provisioning.Akshat [00:03:24]: Yeah.Swyx [00:03:24]: Yeah. I can't even remember my own post.Swyx [00:03:26]: And then you put me on the landing page.Akshat [00:03:28]: Yeah. We really like, the term and so we stole it.Swyx [00:03:32]: Because you had the insight that everything can just be in decorators co-located with the code, right?Akshat [00:03:37]: Yeah.Swyx [00:03:37]: Was that a big part of the originalAkshat [00:03:39]: YesSwyx [00:03:39]: Story or it was just like a DX layer?Akshat [00:03:41]: That was, really important because we really didn't want people to spend, so much time, writing YAML, and it seemed like you could really condense the surface area of what you're doing, put it in code so you can operate on it just like you operate on other code, and like build stuff that's more expressive and dynamic. and so yeah, that was always a very important part.Swyx [00:04:04]: Then the pushback is this is a DSL.Akshat [00:04:07]: Yeah.Swyx [00:04:07]: It's you're closed source. I am locked into Modal.Akshat [00:04:11]: Yeah. We never really got pushback for that because the nice thing about Modal is you can bring whatever code you have, and sure, the DSL is at the configuration layer for, what hardware you're using, how you're scaling things up, but you still own the code.Akshat [00:04:27]: And that's, that's been an important, part of our story, even as we do inference now.Swyx [00:04:32]: Yeah.Vibhu [00:04:32]: How much of do you think still stays the same today? Like if you were to build something today, DevX very important, but I feel like, a lot of this has been changed with just hook it up to an agent, have Claude Code, have Codex implement a tool. there's very agent native primitives that are different than if I'm doing this myself, right?Developer Experience → Agent ExperienceAkshat [00:04:54]: We've changed our SDK team to think about agent experience instead of, developer experience and we think that the same benefits that apply for DX also apply for AX, which is why would you have an agent read through hundreds of Kubernetes files and like write YAML that's not even typed when it can make a couple of changes in a decorator and it gets this self-provisioning runtime of, being able to see its changes live in action? yeah, it just seems from the customers we talk to, they find Modal is much faster for agents to use versus operating on a different substrate.Swyx [00:05:34]: Yeah, because like you, again, you co-locate the infrastructure requirements to the code that runs it.Akshat [00:05:38]: Yeah.Swyx [00:05:38]: Well, the negative thesis now is that nobody's looking at their code anymore, so there's no point.Akshat [00:05:44]: Yeah, people aren't looking at code. one thing we still see is really important is observability.Swyx [00:05:51]: Yeah.Akshat [00:05:51]: Like how good is your dashboard? And of course, like we have, we push a lot of it to the CLI so the agents can do their own investigation, but you still need humans to go interpret what's going on and, make judgment calls and whatnot. and that's I feel like, Maybe more important now than looking at the code itself.Swyx [00:06:11]: Yes, because like, you can try to treat the code as a black box and then use, see the observable action that comes out of it, and then just prompt a change.What Modal Is For: AI Cloud PrimitivesAkshat [00:06:21]: Yeah.Swyx [00:06:22]: So I think it takes a bit of restraint to not specialize, to say, “I want to ship a new primitive,” and then just be general purpose.Swyx [00:06:31]: People ask you, “What are you for?” You're like, “ I don't know. We can do this, we can do that.”Vibhu [00:06:36]: Well, I'd be curious to see, like, okay, if we were to ask you, like, what is Modal for even at a high level? There's a lot you guys do, sandboxes, GPUs, everything. How do you answer?Akshat [00:06:46]: Modal is a cloud platform that's built for, where we've built the primitives from scratch for AI applications. and right now it covers, inference, training, batch processing, and sandbox workloads.Akshat [00:07:00]: But we're building a lot moreSwyx [00:07:02]: I noticed you didn't say web server, so there is still a role for, like, the always-on large-scale Kubernetes type things.Akshat [00:07:09]: Yeah, absolutely. We're, we're not trying to compete with the renders of the world, because yeah, we think the differentiator for us is the, are the workloads that need specialized compute, need to scale up and down a lot. yeah, they're, they're, they're just shaped differently.Working Alongside Frontier StartupsVibhu [00:07:26]: I think you're building a lot of it alongside the startups, right? They're innovating quite a bit, even in your, like, latest blog post. Like, even in the series C, the customers that you mention here, the cognitions, technical ones, ramps and whatnot, they're, they're innovating with you, right? And that's not something AWS is doing directly with.Akshat [00:07:45]: Yeah, absolutely. I think, this is again classic. We're a small team. We can move really fast. our engineers are working with our customers and figuring it out. Yeah.Swyx [00:07:54]: So my first week at Cognition, I walked in, there was someone wearing a Modal shirt. I was like, “What are you doing here?” They're like, “Yeah, I just. I am embedded inside of Cog.”Akshat [00:08:05]: Yeah, I think that was Peyton. We sent him overSwyx [00:08:07]: Yeah.Akshat [00:08:07]: Because, the latency of communication was too high otherwise.Swyx [00:08:12]: Yeah, distributed node, you have to - you have to place one and collocate.Vibhu [00:08:16]: Yeah.Swyx [00:08:16]: So I had a, I had direct personal experience, right? So I worked on smol developer three years ago. it was inspired by Claude 1. I think you onboarded me at some point, like, just before, and I was like, “Oh, like, I need some bursty compute. Like, I was just gonna try using Modal.” And it was a, it was a pretty pleasant experience. apparently, I showed up in the board meeting, like the analytics.smol developer, Sandboxes, and Proto-CognitionAkshat [00:08:39]: Yeah, you blew up on Hacker News and,Swyx [00:08:41]: YeahAkshat [00:08:41]: We got a big traffic spike. I. I think the way you used smol developer was Modal functions for running stuff, which was. Like, the, that was a good use case. but then, yeah.Swyx [00:08:53]: Yeah. That - So to me, that was proto-cognition.Akshat [00:08:55]: Right.Swyx [00:08:56]: If only I had, like, stuck to it.Swyx [00:08:58]: Like, that was like, if - did you say draw the tech treeAkshat [00:09:00]: AbsolutelySwyx [00:09:00]: You're just like, “Yeah, like, probably this will happen.”Akshat [00:09:02]: Yeah. Like, he was so close. You were just rebuilding upon usSwyx [00:09:04]: I just didn't realize.Akshat [00:09:05]: But the funny story there is at the same time, we were talking to a bunch of customers who needed something like sandboxing.Swyx [00:09:14]: Yeah.Akshat [00:09:14]: This is like twenty-three.Swyx [00:09:15]: Yeah.Akshat [00:09:16]: So we builtSwyx [00:09:17]: You introduced a new API right after that.Akshat [00:09:18]: Yeah.Swyx [00:09:19]: Yes.Akshat [00:09:19]: Like, we built sandboxes in May of twenty-three before anyone was even knew this was gonna be a thing. And the first example we published was, we took smol developerSwyx [00:09:28]: Smol developerAkshat [00:09:28]: And put it in a loop, so the agent can iterate on itself.Swyx [00:09:33]: Loops are hot these days.Vibhu [00:09:34]: It's the looper.Akshat [00:09:34]: Yeah.Vibhu [00:09:35]: Loops in. When was this, twenty-three?Akshat [00:09:38]: Yeah.Vibhu [00:09:39]: A small check.Akshat [00:09:39]: Yeah.Swyx [00:09:39]: It's like twenty-three. so the. the, those for listeners, like, the problem was the models are not built for any of this, right?Swyx [00:09:46]: Like, you're just trying to like. They're not post-training to understand, like, looping and, like, self-correction and tool calling was there, but, like, also not that great.Akshat [00:09:55]: Yeah.Akshat [00:09:55]: I don't remember if you used tool calling in this one, but yeah, the models would just diverge after like ten iterations and not produce anything meaningful.Swyx [00:10:03]: Yeah. But like, then. So okay, like now talking to myself three years ago, the answerVibhu [00:10:08]: Of course they will get betterSwyx [00:10:09]: Collect all the failures, build benchmark, and then collect all the, examples, build the RL environmentAkshat [00:10:15]: RightSwyx [00:10:15]: Sell it for like ten billion dollars to Meta.Swyx [00:10:17]: And then also train a model and then sell that for sixty billion dollars to Elon. And this isAkshat [00:10:23]: Yeah, of courseSwyx [00:10:23]: The funny machine. Like, it's like, it's about the hardware.Akshat [00:10:28]: It's hard to have that inherent conviction that the stuff will get that much better.Swyx [00:10:33]: In retrospect, it's so f*****g obvious.Akshat [00:10:36]: Fair enough.Swyx [00:10:37]: Like, what else were we doing back then? I don't know. anyway. Yeah. So this. That was the start of your sandboxing journey, right? I feel like it didn't blow up until, like, last year.Akshat [00:10:49]: Yeah.Swyx [00:10:50]: So there was like a couple years of quietness.Akshat [00:10:52]: Exactly, yeah. We wereVibhu [00:10:53]: I think very underrated product value. Like, my experience with Modal, Charles, before he had joined Modal, met this guy at a hackathon, and he really insisted we wanted to run some small model, not hosted anywhere, and he's like, “ there's this cool company, Modal. They'll like spin up a GPU sandbox, we can throw it on there. They'll take a Hugging Face link.” And like there's so much value just right there, right? Like instant hosting, spin it up, spin it down. It'll stay cold, but we run the demo a few days later, it'll come back up and like all this stuff in retrospect, like it's still what we needed like today.Akshat [00:11:27]: Yeah, it's still needed today. workload shapes have changed a lot as, we run stuff for people with really massive production scale and, there it's it's not about scaling from zero to one, but it's how do we scale really elastically, from like thousand to fifteen hundred GPUs very quickly in a given region. It's the same shape problem.Elastic Inference, GPU Autoscaling, and Custom ModelsVibhu [00:11:50]: Okay. So you look at, say, Cursor Composer, right?Akshat [00:11:53]: Yeah.Vibhu [00:11:53]: They had a. “We'll do RL on a model every couple hours.” you guys have a whole version of RL inference gym and whatnot.Vibhu [00:12:01]: When you look at workloads like that, you're doing train runs where you need to scale up, scale down every hour thousands of GPUs, right? That's the example for we do need it, right?Akshat [00:12:12]: Yeah. Well, so I'll, I'll take a step back and, maybe talk about like how people use Modal today. because our biggest use case is, elastic inference. And the thing we first found product market fit, with was inference for custom models. So we stayed away from the LLM space, and we were serving companies like Suno for audio, Runway for video, robotics, comp bio companies that train their own model elsewhere. But Modal is the best black box that for deployment, scaling to however many GPUs you need as your traffic pattern changes. And we saw all of them like have a very unpredict- predict- predictable, traffic pattern. it's like diurnal. It's Some days, like the company will do a launch and, they'll need like, way more. And it's not just one model that they deploy. They-- all these companies deploy, lots of different models in different regions, and so the autoscaling problem becomes even harder because then you have to scale within a certain region, and those cycles are offset. So different times you scale up in different regions.Akshat [00:13:20]: So that's like our sortVibhu [00:13:22]: And thatAkshat [00:13:22]: YeahVibhu [00:13:22]: That in and of itself is a huge category. There's a bunch of inference providers which, provide this fireworks, does this as a service together, whatnot, Base10. that's carved into its own niche for language models, at least right now.Akshat [00:13:36]: Yeah. the thing that we have specialized in is the autoscaling aspect.Vibhu [00:13:41]: Yeah.Akshat [00:13:41]: Because we found that it's not universally true that everyone else can autoscale, and we've gone deeper into it on the tech side by, we've incorporated GPU snapshotting into the product so we can take the GPU state, like your torch.compile model, snapshot it, and the next cold start is way faster. And so going back to your question, it's That's why you need a lot of burstiness for inference. But then people also do a lot of demand training, like for RL stuff, your rollouts are bursty, as you said. People also do a lot of batch jobs. So we'll see, a lot of companies, before they have a training run, they'll need thousands of GPUs to run encoding or something like that. And I think those things are much more bursty than. I agree that agents are not that bursty. sandboxes are, except when you're doing RL. RL is justRL, Batch Jobs, and 100,000 SandboxesVibhu [00:14:28]: Or commerceAkshat [00:14:28]: Insanely bursty.Vibhu [00:14:29]: Yeah.Akshat [00:14:30]: Yeah. Like when you're doing, rollouts, you sometimes need a hundred thousand sandboxes in your sandboxes.Vibhu [00:14:37]: Yeah. I'm curious if you've seen early sparks of continual learning. There are some people, like our friends, ngram, recently announced thisAkshat [00:14:45]: YeahVibhu [00:14:45]: They're, they're trying to do training. That also seems like a different workload, right? If you're doing training twenty-four/seven per se, there's a very weird dynamic of how you're using GPUs between people and whatnot, but seems like something you guys would work for.Akshat [00:15:00]: As you said, we're, we're fortunate to work with a number of, customers at the frontier and grab some of our customers. and they are taking the primitives we have, and trying to use them in very interesting ways, like continual learning. It's possible as the stuff gets better, some of that will be part of, our offering as well if, more people need it. but we're, we're just waiting to seeVibhu [00:15:23]: YeahAkshat [00:15:23]: How it shakes out.Vibhu [00:15:24]: Is there a primitive that you added after sandboxing that was the next step in the story?LLM Inference, DeFlash, and Speculative DecodingAkshat [00:15:32]: I guess we've been going much deeper into LLM inferenceVibhu [00:15:35]: YeahAkshat [00:15:35]: Because we realized that some of the advantages we have with like autoscaling, again, especially in different regions and whatnot, are, not present elsewhere. and the place where we had a gap was we weren't, working on the model layer itself. Like we were a black box. And, we realized that, we can get to frontier-level model performance, with, by having great people who work on this. And, we've been open sourcing a lot of our work, in terms of, Recently, we, shared our work on DeFlash, which is a block-based, speculator, and we've open sourced, all of it. So, you can - By using open source DeFlash, you can get the same performance as you would with one of the proprietary providers. And the next thing we're thinking about hereVibhu [00:16:23]: I thought this wasAkshat [00:16:24]: YeahVibhu [00:16:24]: An interesting blog post as well, right? Like, I think in here you make a claim that. Not a claim, just that how effective speculative deco-decoding really just get to.Akshat [00:16:33]: Yeah.Vibhu [00:16:33]: Anything you wanna point out from this around, what people should know?Akshat [00:16:39]: Yeah, absolutely. the high-level summary is, it would help to describe what speculative decoding is.Vibhu [00:16:44]: Yes.Akshat [00:16:44]: I will, yes.Vibhu [00:16:45]: I think, likeAkshat [00:16:46]: YeahVibhu [00:16:46]: So we've covered like Eagle and all thisAkshat [00:16:47]: YeahVibhu [00:16:47]: Like Hydra and all those things, but it was like two years ago.Akshat [00:16:51]: Yeah.Vibhu [00:16:51]: I think it doesn't hurt, right?Akshat [00:16:52]: Yeah. Speculative decoding is you have a smaller model, called a draft model, predict tokens ahead of the bigger model, and then you have the bigger model, verify all of this, all the tokens are predicted. And the reason it's faster is if you're predicting, one token at once, you're bound by memory bandwidth. But if you can batch the verification of, the draft model, then you're much more efficient using compute, and it's faster, and as long as your draft model is producing a lot of tokens that can get accepted, which is called the accept length, you can get a speed up that's, multiple times of, the original model speed. and well, that's what we highlight here. It's Like people talk a lot about we made these kernels faster and whatnot, but improving kernel will only give you like few percentage points of improvement, and, increasing accept length, literally is a multiplicative decreaseVibhu [00:17:47]: Like two to four X.Akshat [00:17:48]: Yeah, exactly.Vibhu [00:17:48]: Without much head-on performance.Akshat [00:17:50]: Yeah. I think it may - you are running a second model, right? So it may be something more expensive in the compute,Vibhu [00:17:57]: I meant quality performanceAkshat [00:17:58]: Probably not by muchVibhu [00:17:58]: But yeah. I thinkAkshat [00:17:59]: So there's no drop in quality performanceVibhu [00:18:01]: YeahAkshat [00:18:01]: Because you're always. You're never accepting a token that the big modelVibhu [00:18:04]: It's strictly betterAkshat [00:18:05]: YeahVibhu [00:18:05]: Or it's same.Akshat [00:18:06]: Exactly.Vibhu [00:18:07]: Right. Yeah.Akshat [00:18:08]: And so we've been working a bunch on DeFlash, which is a block-based speculator. so it's instead of predicting, one token at a time, it's predicting a block. And we've been open sourcing our work with it. The next thing for us here is for helping people train speculators and custom models. it's it's something that traditionally is very forward-deployed engineering driven, support deployed, engineer driven, like you work with customers and help them do that. And our vision for. This is why we launched Auto Endpoints, is we want to make frontier-level performance available to everyone. And so, we mentioned this in the announcement, we teased it. The next thing we're, we're launching is, as you run an auto endpoint, we shadow trafficAuto Endpoints and Frontier-Level PerformanceVibhu [00:18:54]: Do you want to explain what auto endpoints are?Akshat [00:18:57]: Yeah.Vibhu [00:18:57]: I lovely, yeah.Akshat [00:18:58]: Yeah. So, this is, I guess, going back to your Modal is you touch the code, but, sometimes people don't wanna touch the code, and they wanna get started with an endpoint that works and has all the great performance and, scalability that Modal has. So we've made that easier with, a way to create an endpoint from our UI, from the CLI, that has all of our optimizations that we talked about, like the DeFlash stuff already baked in, and there's full transparency. So we give you the code, you can go run it yourself, and if you want, you can eject out into the full Modal experience, which we see as people get sophisticated, they do wanna tweak the models, they wanna, fine-tune stuff. You can still do all of that. It's it's not a black box. And yeah, the next thing, as we teased later in the post, is how do we give you value even beyond this in terms of having your draft models evolve as your data distribution evolves, again, without having to talk to a person and, yeah.Vibhu [00:19:59]: I guess just to understand it directly, you have the GPUs, you have an endpoint that's compatible, you serve open model. If someone was to do this themselves, what's the delta that you guys provide? So you do a lot of open source great work on effective inference. how does it compare to, say, I take the same model, 5.2 FP8, take shelf inference engine, vLLM, SGLang, get compute of similar capacity, similar cost. What's the delta that plugging into something this, like this offers outside of the benefit of, scaling?Production Inference Beyond Raw GPUsAkshat [00:20:34]: It's interesting because we've taken the approach of open sourcing our contributions and upstreaming them. we work closely with the SGLang team. We want the improvements that our team, comes up with to be, there in open source for others to use, even outside of Modal. The benefit to us is we have a team that has significant expertise in terms of if you do have something that is not there, our team can help you get that performance, first. the other thing is with these endpoints, we are way more elastic, as you said, than, anyone else, and you have true scaling to zero. you have true, burstiness, and in practice, that matters a lot more to people than just finding, the GPU and, running Modal code on something.Vibhu [00:21:20]: Yeah. And I will say it's not that straightforward to just. like what I said is easier said than done, right?Akshat [00:21:26]: Yeah.Vibhu [00:21:27]: It's I think still for the average person, still hard to just gut check using different. There's, there's quite a bit of combinations you can make there. the trade-offs aren't really known at face value.Akshat [00:21:40]: Yeah. it's it's not just that. I think it's it's that running production-grade inference is a hard infer problem.Vibhu [00:21:49]: YeahAkshat [00:21:49]: Even if you subtract out the autoscalingVibhu [00:21:50]: YeahAkshat [00:21:51]: Is controlling things like tail latency and, making sure every, request is delivered at least once and whatnot.The Model and Agent LifecycleVibhu [00:22:00]: There's a lot of innovation that you can do here. I think, it's very interesting that you're starting to encroach on, like as you become a full cloud, you're starting to encroach on other people's turf.Vibhu [00:22:09]: What will you not do?Akshat [00:22:13]: Well, we wanna follow our users and, make sure they get like a platform that has everything that works well together. so right now we're focused on the model lifecycle and the agent, lifecycle. so both like going from data prep to training to inference, and then also if I want to deploy a background agent, let's say, sandbox, do persistent storage, a whole bunch of other stuff.Vibhu [00:22:38]: We talked to Cole, who did, OpenInspect. Yeah.Akshat [00:22:42]: Yeah.Vibhu [00:22:42]: And RealInspect also is on Modal.Akshat [00:22:44]: Yeah. So Ramp Inspect was a great example of a background agent that was really successful because they, were able to use some of the primitives like snapshotting and fast scaling to just have something that feels really reactive and works well.Ramp Inspect and Background AgentsVibhu [00:23:02]: Yeah. That's the new CTO of, Ramp right there.Akshat [00:23:05]: Yeah, Rahul.Vibhu [00:23:08]: It was really fun. yeah, okay, I think, all very bullish. Like, one of my reflections was also I did not originally. So when I met you guysThe Inference Inflection: CPU, GPU, and Co-LocationVibhu [00:23:19]: You weren't that much in the GPU game, and now you're all about, inference. And one of the points that I hinged on for Jensen's keynote at GTC this year was, what we're calling like the inference inflection, right? That let's say in AI workloads or machine learning workloads, it used to be like, let's call it eight to one GPU to CPU, and now it's more like one to one, which is like a interesting. Like, - because of how much agents are blocked or call out to this, to CPU heavy stuff the actual, like, limiting factor, like, swings back and forth from GPU to CPU a lot more than it used to be all GPU and then occasional CPU.Akshat [00:24:01]: Yeah.Vibhu [00:24:02]: GPU, CPU. And now it's like just constantly, and you just have to locate everything.Seventeen Clouds and the Supercloud StrategyAkshat [00:24:08]: Yeah. And that's one of the things that, again, we see as, something appealing about Modal, which is we've built this capacity pool that spans, 17 cloud providers, so we're, we're very good at Running on various kinds of cloud capacity across the worldSwyx [00:24:24]: You don't have your own data centers?Akshat [00:24:25]: We don't have our own data centers. We just run across a lot of neo cloudsSwyx [00:24:29]: Yeah. AreAkshat [00:24:30]: Metal providers.Swyx [00:24:30]: Yeah. Question mark.Swyx [00:24:31]: Yeah. You're, you're running the math, and you're like, “What's the cutover point where you're like.”Akshat [00:24:36]: Yeah, it's a good question. part of it is we see our differentiator in the software layer, and, being capital light and focusing on the software helps us move really fast. so far it's worked out well because there are so many other people building data centers that we're able to work effectively with them, and again, focus on what makes us, special.Swyx [00:24:55]: Yeah.Swyx [00:24:56]: 17 gets you into, like, the local providers sometimes. LikeAkshat [00:25:00]: The,Swyx [00:25:01]: Which was the most interesting one?Akshat [00:25:02]: There are a lot more neo clouds than you expect, and they all have various degrees of, various levels of reliability. And, that's why it's something we've invested a lot of time in, is building our own reliability layer on top. so if the GPU falls off the bus or something happens, we user workloads are not affected, and that lets us use a lot more capacity than,Swyx [00:25:30]: YeahAkshat [00:25:30]: You as a user would be able to.Swyx [00:25:32]: It's a useful thing to have because like now everyone knows, like, what layer you are and, like, you optimize for being the super cloud of all clouds.Akshat [00:25:41]: Yeah. That's, that's, that's the idea. and so I guess when you mentioned colocation, that's, that's another interesting thing where, one thing we've seen is people come to us when they want, very specifically located, CPUs or GPUs, like they wantSwyx [00:25:57]: Oh, they pin it in likeAkshat [00:25:58]: YeahSwyx [00:25:58]: EU?Akshat [00:25:59]: Exactly. Or EU, US.Swyx [00:26:01]: Right. Data resiliencyAkshat [00:26:02]: AustraliaSwyx [00:26:02]: Locality thing or performance or what?Akshat [00:26:04]: It's either data locality or latency, yeah.Swyx [00:26:07]: Yeah.Akshat [00:26:07]: Like, you want your. They're running sandboxes and model. They want them to be right next to aSwyx [00:26:10]: Yeah, it's easy thenAkshat [00:26:11]: YeahSwyx [00:26:12]: To. That is important in all those things. and so, like, you've accidentally, I don't know if it's accident, but, like, you've built the perfect primitive for agents to express themselves. And then, like, it's almost very funny how every extra development just involves more file system, just involves more CPU.Akshat [00:26:30]: Yeah.Swyx [00:26:31]: Just like the things that you already have. I don't know much about, if there's any, like, networking usages that are interesting, but you've also done some good work on networking.Networking, Sidecars, Private IPv6, and SandboxesAkshat [00:26:40]: Yeah, that's exactly right. Like, we're just taking compute storage and networking and building stuff on that layer, for, again, the stuff people need.Swyx [00:26:49]: YeahAkshat [00:26:50]: We see a few interesting networking things coming up. one is people want networked sandboxes. so we haveSwyx [00:26:57]: For like a Docker cluster type thing.Akshat [00:26:59]: Yeah.Swyx [00:26:59]: Sorry, Docker Swarm. Oh, f**k. What is it called?Akshat [00:27:02]: Compose.Swyx [00:27:03]: Compose type thing.Akshat [00:27:04]: Yeah. So if you want Docker Compose, our sandboxes now support, this thing called sidecars. So you can. A sandbox is a pod of containers, and you can run multiple containers in, a sandbox. also useful because, going back to networking, people want a lot of control over, outbound networking from a sandbox.Swyx [00:27:23]: Yeah.Akshat [00:27:23]: Like, they might wanna run a middle proxy for, like, maybe logging stuff for RL or, controlling how egress can happen to a domain, injecting credentials. and yeah. So we've, we've had to build a lot of that stuff ourselves.Swyx [00:27:38]: Yeah.Akshat [00:27:39]: But then also sometimes people want, sandboxes spanning multiple nodes to talk to each other, which is an emerging thing we're seeing. We have support for that for a different reason, and yeah, we'll see if that becomes stable.Swyx [00:27:52]: Like, just an open socket. It's a. This is directly like mTLS.Akshat [00:27:56]: We do support that, which is you can, expose a tunnel inside a sandbox.Swyx [00:28:01]: Yeah.Akshat [00:28:01]: And then you can either expose it to public internet or it can be, you can add like a HTTP, auth layer above it. But we have this thing called I6PN, which we haven't talked about, which is this, like, overlay network using IPv6 addresses. so if Modal containers, within the same workspace, when this is enabled, can address each other using this private IPv6 address, and no one else can.Akshat [00:28:28]: So it's like private networking, for containers. We built it because we needed it as a primitive for our distributed training product. so we have this other feature, which is you can add a decorator to a function, and you get a cluster of GPUs. and they have RDMA networking. so you can run a distributed training job, that's truly serverless. and we did the overlay network for that. But then we've seen that people are using it for other reasons, and, I'm intrigued to yeah, what would people do with it.Swyx [00:28:59]: Build primitives and let people figure it out, right?Akshat [00:29:01]: Yeah, exactly.Swyx [00:29:02]: You put out a pretty interestingAkshat [00:29:03]: They're like, they read the docs webpage. Let me use thatSwyx [00:29:06]: YeahAkshat [00:29:06]: Something they never intended to work. This is literally not even in our docs page. People somehow found it, and they're using it.RDMA, Memory Movement, and Distributed TrainingSwyx [00:29:12]: Huh.Swyx [00:29:14]: The way you portrayed it with, like, RDMA versus TCP, like, very well laid out, but just the transfer speed change at scale for RL, like yeah, you have it, you have it built in. I'm sure someone found it. It's found it to be a lot more efficient before you made a thing out of it, right?Akshat [00:29:32]: Yeah. And not to split hairs, I guess the overlay network is the TCP overlay network.Akshat [00:29:39]: The reason we have that is you need that to do the key exchange for RDMA before you set up the RDMA network on top of that. but then people found the TCP part.Swyx [00:29:48]: Can I tell you, this is like a big aha moment for me becauseAkshat [00:29:51]: YeahSwyx [00:29:51]: So I review 2,200 submissions for the World's Fair.Akshat [00:29:56]: Yeah.Swyx [00:29:57]: And then I got this from John OsterhoutAkshat [00:29:58]: HuhSwyx [00:29:59]: Who I don't know if. Do John Osterhout by name?Akshat [00:30:01]: The name sounds familiar.Swyx [00:30:02]: He published a. He's a well-known professor, published a lot of interesting software design books, and this is the talk he chose to submit, is on RDMA at Inference. And I'm like, you wouldn't think that this guy, who is like operating systems guy, would care about RDMA.Akshat [00:30:20]: I, it makes sense to me because I,Swyx [00:30:24]: This is the cloud, right? YeahAkshat [00:30:25]: Like, the way you move around your KV cache and how efficiently you can do it, how efficiently you move, your weights from your training GPUs to your inference GPUs in RL is there's a lot of degrees of freedom, and it is a systems problemSwyx [00:30:41]: YeahAkshat [00:30:41]: Moving memory aroundSwyx [00:30:42]: YeahAkshat [00:30:43]: Scheduling.Swyx [00:30:44]: This shows you how primitive my understanding of networking stuff is.Swyx [00:30:46]: Is this like the domain of WireGuard as well?Akshat [00:30:50]: Not quite.Swyx [00:30:51]: It's adjacent?Swyx [00:30:53]: Explain everything.Akshat [00:30:54]: Sure.Swyx [00:30:56]: How do we move memory around GPUs?Akshat [00:30:58]: Well, so sorry. Yeah, that is memory. Sorry, I was talking more, and maybe I was talking like five minutes back, about the private IPv6, addressing that you've set up.Swyx [00:31:09]: Yeah.Akshat [00:31:09]: Is it like it's a VPN?Swyx [00:31:10]: Yeah, it is like a VPN, and yeah, WireGuard is, yeah, you're right. It is,Akshat [00:31:16]: Right. Yeah, you already moved on to new topicsSwyx [00:31:17]: A similarAkshat [00:31:18]: OkaySwyx [00:31:19]: In the same space, WireGuard is, encrypted and this is,Akshat [00:31:23]: And you don't need encryption.Swyx [00:31:23]: Yeah.Akshat [00:31:24]: Yeah.Swyx [00:31:24]: This is not encrypted. that's the main difference. This is TCP and we have eBPF programs that will reject or allow the TCP connection based on whether you're allowed to do it.Akshat [00:31:35]: Used to involve a full sidecar, but now you have eBPF in the Linux kernel.Swyx [00:31:39]: Yeah.Akshat [00:31:40]: Yeah. I don't know if this is a natural follow-on to the topic of like my skepticism on distributed training is that while, like, people spend a lot of money on, like, cables to hook up GPUs, and even that is not, like, fast enough, and that's the bottleneck, is your networking fast enough?Swyx [00:31:59]: Yeah. So I guess you're talking about fully distributed training like, Dialog or something which is like cross data centerAkshat [00:32:06]: That would be, yes.Swyx [00:32:07]: That's the extreme.Akshat [00:32:08]: Yeah.Swyx [00:32:08]: You're in the middle, and then other people would have like the Mellanox cables up in, like, their actual data center.Akshat [00:32:14]: When you run multi-node training on Modal, RDMA, I think Mellanox, is, or InfiniBand is like a, is all seen as RDMA. but it's a way to bypass the TCP networking stack and, transfer, stuff much faster, between one node, to the other. And we have I think like 3 terabit per second, internal networkingSwyx [00:32:40]: OkayAkshat [00:32:40]: Which is the standard that's needed.Swyx [00:32:42]: Okay. So I misunderstood whatAkshat [00:32:43]: 50Swyx [00:32:43]: What part of the stack you wereAkshat [00:32:44]: 50 gigs overSwyx [00:32:45]: YeahAkshat [00:32:45]: If you wentSwyx [00:32:45]: YeahAkshat [00:32:46]: RDMA.Swyx [00:32:46]: Okay.Swyx [00:32:48]: Yeah. I, very impressive work.Multi-Node Training, Post-Training, and Auto ResearchSwyx [00:32:52]: So effectively you're extending like the model philosophy to the training cluster, like, yeah.Akshat [00:32:59]: Yeah. And we're, we're not going for like large scale training runs. the thing that we've built multi-node training for is, we see a lot of, smaller scale post-training. like, people are post-training like medium sized fund models, so they can, get higher quality on inference. this is a perfect fit, for something like that.Swyx [00:33:21]: Yeah. That is my impression of how a lot of these labs explore branches in post-training and then eventually merge whatever they find in.Akshat [00:33:31]: Yeah. The other use case we've seen for multi-node training is even if you have a big cluster, your researchers are still doing small runsSwyx [00:33:38]: YesAkshat [00:33:39]: Having elasticity thereSwyx [00:33:40]: Right, sureAkshat [00:33:40]: Matters a lot more.Swyx [00:33:41]: Yeah. the, like, this is like the current limiting factor for auto research, which is like you need to give your model some GPUs in order for it to completely run.Akshat [00:33:51]: We have a blog post on auto resource and model is,Swyx [00:33:55]: YeahAkshat [00:33:56]: Yeah, like, turns out to be pretty good substrate for that.Swyx [00:33:59]: So my impression is auto research means many things, likeAkshat [00:34:01]: YeahSwyx [00:34:01]: Anything that Andrej coins. Right now it's still science fair, right? Like not like, I don't know how many people are doing this.Akshat [00:34:08]: We're having a golf.Swyx [00:34:08]: Yeah.Akshat [00:34:09]: I thought the same thing.Swyx [00:34:11]: Yeah, you would know.Akshat [00:34:12]: We, like, our internal both training and inference teams use this the general shape of this quite a bit. like we have this one internal repo called auto inference, which essentially we've automated our own forward-deployed engineering efforts using, this harness, which is, the agent will just spin up a sweep of different things. It'll even run like, NVIDIA inside profiler and it'll like tweak configs and it'll arrive the right thing. it'll change your GPUs both from H200 to B200, and works really well.Swyx [00:34:47]: Nice.Akshat [00:34:47]: So yeah.Swyx [00:34:48]: By the way, I enjoy that your forward-deployed engineering is so technical that you have to do these things.Swyx [00:34:52]: It's very different from forward-deployed engineering from other people.Akshat [00:34:54]: Yeah. For our forward-deployed engineering team is, essentially they're like applied inference researchers or applied training researchers.Swyx [00:35:02]: Someone told me like they have to be able to build, but they also have to be able to sell. do they have to sell or are they like they're good, they're just like post-sale type of thing?Akshat [00:35:09]: It does, being able to talk to a customer and engage effectively with themSwyx [00:35:13]: YeahAkshat [00:35:13]: Matters a lot.Swyx [00:35:14]: They want the same thing.Akshat [00:35:15]: Yeah.Swyx [00:35:15]: ?Akshat [00:35:15]: But it's it's not really a sales, thing. We pair them with-- We have solution architects as well that are more on the sales side.Swyx [00:35:23]: Okay. Let's spend a bit more time on auto research. This is a big focus for for this year. Where does this go? like, have people explored enough? Like, there's all these beautiful charts of like improve and then level off a bit and then you find the next thing. Is this one abstraction up from normal training? Is that how we think about it, or do you think about it differently? Like model level training versus high, like driven hyperparameter search.Auto Inference and Modal BenchAkshat [00:35:51]: Yeah, like,Swyx [00:35:51]: Someone, some people call it like neural architecture search or whatever, right? Like.Akshat [00:35:54]: Yeah, - So the stuff I've seen people do with it is nowhere on the architecture level. It's pretty much tweaking parameters, but it's it's a hyperparameter sweep that's guided by some model intuition, so it's like much more efficient than, whatever other, sweep you would have.Swyx [00:36:12]: Yeah, it's just, it's just a question of where you want to spend your compute?Akshat [00:36:16]: Right.Swyx [00:36:16]: ‘Cause yeah, you can just throw infinite amounts of money on this and somehow you'll bang out Shakespeare?Akshat [00:36:22]: Yeah, infinite monkey.Swyx [00:36:24]: Yeah, so like the very good for model. and I think it's also very important that agents can spin up other agents, can spin up their infrastructure. Like very good for you. how good is our LLMs at generating model code? Like the benefit of existing LLMs is that you are in the data.Akshat [00:36:42]: Yeah. They're, they're surprisingly good. I think like pre Cloud 4 they were not, and then now they're able to shot, stuff out of the box. But we're playing around with releasing like a Modal Bench for like the harderSwyx [00:36:55]: YeahAkshat [00:36:55]: Things, that the LLMs cannot do yet and maybeSwyx [00:36:59]: What's an example of that?Akshat [00:37:01]: I think the things that- Sometimes agents struggle with, without right guidance and a skill is, how to, use the rest of our observability. Like how to. Something is failing, like how do you look at the logs and then update the right thing? It's reasoning about that. But they're able to shot, likeSwyx [00:37:23]: Yeah. You can just add a skill to it?Compute Strategy and Capacity PlanningAkshat [00:37:26]: Yeah. So we have a Modal skill now that. Which is why we built this Modal Bench. It's to find things like that, so we can address them in our tool.Swyx [00:37:35]: Tune a skill. Yeah.Akshat [00:37:36]: Yeah.Swyx [00:37:36]: No. it's it's good. are you facing any shortages? like we talk a lot about GPU shortages, but also CPU, also memory.Swyx [00:37:44]: Yeah.Akshat [00:37:45]: We have had a lot of growth, which means that, there's - we've had to be much better aboutSwyx [00:37:53]: PlanningAkshat [00:37:54]: Proactive capacity planning.Swyx [00:37:55]: Yeah.Akshat [00:37:55]: So we have,Swyx [00:37:57]: Which by the way, like it's like a MBA's like dreamAkshat [00:38:00]: YesSwyx [00:38:00]: Is like just planning this stuff. I think last time you and I talked about something maybe about this.Akshat [00:38:03]: Yeah. we have a really competent team of people that we call, The role is called compute strategy. so yeah, if anyone listening here or wants to work on thatSwyx [00:38:13]: Compute strategy?Akshat [00:38:13]: Yeah.Swyx [00:38:14]: I think,Akshat [00:38:14]: I feel like,Swyx [00:38:15]: I think the normies call it FP&A or something.Akshat [00:38:18]: Well, it's more It's it's not FP&A. It's it's There's a lot of interesting financial questions of like what is the blend between one year and three-year reservations? how do we forecast our own capacity? how do we. especially since our capacity is very fungible across different GPU types and different regions, like you have to model a lot of it. and you also have to have an opinion on how the supply chain is gonna evolve, and then you have to like, take bets,Swyx [00:38:49]: YeahAkshat [00:38:49]: Based on that.Swyx [00:38:50]: Tokenomics.Akshat [00:38:50]: Yeah.Swyx [00:38:51]: This is like probably a not a real point, but, I was trying to think about like what other industries. I was trying to think about like, we cannot be first to like these kinds of problems.Akshat [00:38:59]: Yeah.Swyx [00:39:00]: And what other industries have had this? And I was like, airlines with fuel and like they have to hedge their fuel and like, I think for a long time Southwest because they made like a hero fuel bet, they like were like super low cost becauseAkshat [00:39:12]: OhSwyx [00:39:12]: Compared to everyone else.Akshat [00:39:14]: Yeah. I hadn't thought about that.Vibhu [00:39:16]: We're at a fun time too?Akshat [00:39:18]: Yeah. It's. A lot of the compute business in general, for us is also about being very good about capacity management. That is how you have great unit, economics. but also over time it's how you can unlock more value for customers. Like, one of the things we're building now is like a way for customers to get, If they don't care about latency, like get much cheaper pricing and they'll get results back in like next 24 hours or something, like a batch tier essentially.Batch Tiers and Latency-Insensitive WorkloadsSwyx [00:39:47]: Yeah.Akshat [00:39:47]: And those are levers we have because we control the whole stack and scheduling and whatnot to give people a sufficientSwyx [00:39:53]: Yeah. I feel like they're not as popular. Like those, like the Frontier Labs have all those APIs. They're not as popular as they should be.Akshat [00:40:00]: The demand that we see for something like that is not for LLMs. although sometimes people wanna run evals andSwyx [00:40:08]: OkayAkshat [00:40:08]: Synthetic data prep and there it makes sense.Swyx [00:40:10]: Okay.Akshat [00:40:11]: But it's from a lot of LLM companies, like people who are doing computational bio, like they have to run really big batch jobs and they don't care about when they get it back.Swyx [00:40:22]: Yeah. And like they have a reasonable. It's it's also like a cousin to the stopping problem of like, will this finish in time?Akshat [00:40:30]: Yeah. You can bound it.Swyx [00:40:33]: Yeah.Akshat [00:40:33]: Like you can give peopleSwyx [00:40:34]: YeahAkshat [00:40:34]: SLAs on it.Swyx [00:40:35]: Yeah. I think what's, what's interesting is like the next phase of model.Swyx [00:40:38]: Like what, do people expect from you, now that you're established and you're like well-known compute player among all these leading companies. You had an inference launch week, and we talked a little bit about the launches. like what else? Like what else should people know?What Modal Builds NextAkshat [00:40:55]: We are building primitives that make our users' lives much easier. So, I think for example, with LLM inference, thousands more companies are gonna post-train their own models and, deploy open source models for inference. so we're thinking a lot about what is the best product shape for that. And, that involves everything from our training gym to, then, endpoints that get frontier-level performance. again, but I haven't talked to anyone. It looks somewhat different on other verticals. Like, we're also seeing a lot of real-time, audio-video stuff in there, which is why like, we're working on things like regional routing, with fallbacks. So you can get GPUs that are as close to users as possible. so you get like low latency for video streaming and whatnot. And then on the agent side, it's,Akshat [00:41:52]: We're still working very closely with our customers because stuff is changing so fast in terms of what they need. And, I think beyond sandboxes and persistent file systems, there's a lot of other things people will need from this agent stack as they build production agents. So yeah, we're thinking about those other things that fit in there.Swyx [00:42:13]: I want to ask what the other things are.Akshat [00:42:15]: Yeah. I probably should share right now.Swyx [00:42:17]: I think-- I think, okay, so, I do think a lot about the principal components of cloud, and you do talk about compute storage networking.Akshat [00:42:25]: Yeah.Swyx [00:42:25]: Because so far for me, it's fine. so far for the. the first couple generations of cloud, it's fine. What's different, qualitatively different about agents that you need some new permission level? Like a lot of people, okay, and I'll just kinda spew tokens at you until it like hopefully sparks something.Akshat [00:42:43]: Yeah.Swyx [00:42:44]: Like the new level now is whatever Claude Code does, which is dangerously scope permissions or like allow list by command or like whatever, right? And sometimes they're like, “Well, okay, we have like this adaptive thinking mode where like, just trust me, bro. I will make the calls for you.” Is that it? like mediated permissions.Hard Guardrails vs. LLM-Mediated PermissionsVibhu [00:43:03]: Now you're looping it with a goal and letting it roll.Akshat [00:43:06]: Yeah, I'm, I'm skeptical of LLM media permission for stuff that is at the sandbox level because you do want hard boundaries.Swyx [00:43:16]: Yeah.Akshat [00:43:16]: Otherwise, someone can exfiltrate stuff.Swyx [00:43:20]: But likeAkshat [00:43:20]: YeahSwyx [00:43:20]: Maybe that's old school thinking. Maybe we're the dinosaurs.Swyx [00:43:23]: Maybe the AI OS or the LLM OS is really the kernel is a goddamn LLM.Swyx [00:43:30]: Like it makes you feel uncomfortable.Akshat [00:43:31]: Yeah, I'm, I'm toldSwyx [00:43:32]: But that's what trusting the LLM is. Like imagine a spherical cow perfect LLM.Akshat [00:43:36]: Right.Swyx [00:43:37]: That it.Akshat [00:43:39]: Maybe.Swyx [00:43:41]: I wanna test the boundaries, right?Akshat [00:43:42]: Yeah.Swyx [00:43:42]: Like, and I don't believe that, but I wanna see where I'm wrong ‘cause that's, that's the consensus.Akshat [00:43:49]: Yeah. I think you always need hard guardrails when you want, And you can pair those with softer guardrails, right? And that's gonna be a lot of mediated.Managed Agents and Specialized SandboxesSwyx [00:44:00]: There. I'll also get you a end with a couple of your commentary on like the ecosystem outside of Modal. Manage agents. Everyone has one. Gemini, OpenAI, Claude, very useful for you, but also like it is their way of starting to edge into your space.Akshat [00:44:17]: Yeah.Swyx [00:44:17]: What's going on?Akshat [00:44:19]: Yeah, we're, very excited to partner with Anthropic and some of the other foundation labs, will not name who we're also working with. the way we see it is the manage agent thing is a great place to start if you're starting out building an agent and, But then when you get to, building something more production grade, like you're a company that's like Ramp that's building their own, Ramp also runs their accounting agent on us, so their external-facing agent. You need a lot more control over, your compute primitive on things like, what sort - how do you persist different files that the agent has access to, and how do you snapshot and restore? How do you control the networking? maybe you want GPUs. When you get to that point, you kinda want, a specialized sandbox provider, that gives you those things, and that's the role that we are trying to play.Swyx [00:45:15]: YeahAkshat [00:45:16]: We don't really have an opinion on the harness, whether it runs - it's a cloud-managed agent, and you hook it up to Model Sandbox, or you run the harness in Model Sandbox. We'll see where people converge with that.Swyx [00:45:26]: Yeah. Do you any opinions on like the meta harnesses, or just another layer on top of these things?Akshat [00:45:31]: You mean like the OpenPipeSwyx [00:45:33]: OpenPipe is one. I think Vercel had one, which I can't remember the name of right now. Fredshot had one. and then, to me, most recently was Data Databricks that had Omnigen. All these are meta harness. Like it's kinda pseudo agent cloud type things.Akshat [00:45:50]: I personally have not played around with them.Swyx [00:45:53]: Yeah.Akshat [00:45:53]: Build agents with them.Swyx [00:45:54]: Everything's bullish Modal, as long as it consumes more infra.Akshat [00:45:57]: That's why we're focusing on the infra layer. It's somewhere where our, relative competence is and, also it's a hard problem to solve.Swyx [00:46:06]: Yeah. I will say like just generally reflecting on that, I don't know if - if there's other topics on Modal, but like just generally reflecting as an infra person, not as intense as you, but in that field, this has like been the most exciting time in infra. Like it was boring for a while, and you couldn't really get people excited about data infrastructure. Like Eric would get on Data Console, everyone just watched the video and like say, “Look at how many sandboxes I can spin up,” and no one gave a crap.Why Infrastructure Became Exciting AgainAkshat [00:46:39]: Yeah.Swyx [00:46:40]: And like now everyone gives a crap.Akshat [00:46:42]: That's true. It is a very exciting time, and I think a lot of that's driven by just the amount of scale all of this stuff needs.Swyx [00:46:50]: I think the, like a lot of your initiatives or a lot of your like product directions make sense in retrospect, which is like the best kind, but I wouldn't necessarily have thought about it myself, which.Akshat [00:47:00]: We need the predictions.Swyx [00:47:02]: I think there's a lot that you just don't even see, right? Like you have the batch, you have the voice, you have the multimodal, but what else?Akshat [00:47:10]: What else is coming up for usSwyx [00:47:11]: Yeah. Where do you see things going?Akshat [00:47:13]: Yeah. I, in generalBiotech, Robotics, and Non-LLM AI WorkloadsAkshat [00:47:15]: It's it's clear that there's there's a huge shift happening. I think one thing that's not as obvious to people because LLM inference gets talked about so much and is also we work a lot of companies that are, doing things like drug discovery and computational bio, like the Chai Discoveries of the world. Big things are probably gonna happen there. we work a lot of robotics companies that are putting robots in like active deployments and getting good results out of them.Swyx [00:47:45]: Is there Air Gap Modal? Is there a version that is like prem air gapped whatever?Akshat [00:47:50]: No. We,Swyx [00:47:51]: You should cloud only.Akshat [00:47:51]: Yeah.Swyx [00:47:52]: Yeah. Okay. But yeah, so what you're saying is like because you're focused on primitives and they're good primitives, you find use cases in all these kinds of things.Akshat [00:48:01]: Yeah.Swyx [00:48:01]: Probably diversifies you a little bit away from LMS all the time.Akshat [00:48:05]: Yeah, absolutely. We're, we'- our goal isn't to only serve the LLM inference market.Swyx [00:48:10]: There are a lot just on the website, the audio,Akshat [00:48:12]: Yeah. We said both onSwyx [00:48:14]: Computational bio images. Yeah, there's a lot here. There's QTA TTS, customizing. Oh, Chatterbox. there was customizing Whisper.Akshat [00:48:24]: Okay. Yeah.Swyx [00:48:25]: This screen reminds me of a fallen competitor, which Replicate.Model APIs vs. Differentiated AI ProductsSwyx [00:48:31]: What's your postmortem on what happened?Akshat [00:48:34]: This is one thing we've stayed away from is providing an API for models because I think providing model APIs is some of it ends up serving like a really hobbyist market, which is much less sticky.Swyx [00:48:50]: Yeah.Akshat [00:48:50]: And we've always wanted to build for companies that are building products and need more flexibility that's not just an API.Swyx [00:48:57]: Which you can build an API for a model and this is clearly what it is. But you - but what you're saying, you can wrap it into a more fully functioning back end that you run.Akshat [00:49:06]: Yeah. So all of our examples, it's not that spin up this model, here's an API token, use it. They're all code.Swyx [00:49:13]: Okay.Akshat [00:49:13]: And so the point is that this is just an example.Swyx [00:49:16]: Starter code.Akshat [00:49:17]: Yeah. But you can tweak it however you want.Swyx [00:49:20]: Yeah.Akshat [00:49:21]: And if you're like a company building a product, like, computational bio whatnot, yeah.Swyx [00:49:26]: I guess I'm trying to tease out for listenersAkshat [00:49:28]: YeahSwyx [00:49:28]: When does it stop becoming, oh, you're just an API call and you're just a wrapper on API to becoming what you call a product, right?Swyx [00:49:36]: Like, what is that layer? Like what-- Like, more lines of code, but like beyond that, what is the substance that people add that qualifies it to be something more?Akshat [00:49:46]: I think there's a little bit of like a selection effect of like a lot of the companies who do wanna get deeper into that level are probably building something that's more differentiated. And, I think, an example is like - with LLM inference, originally we, worked with companies that were building their own post-training frameworks or they were, - Ramp early in the day was training their own tokenizer and like swapping out the tokenizer in Llama and whatnot. I'm not saying that's, that successful, in that case. But a better example is like, let's say Suno. because Suno, does not use Modal for training.Swyx [00:50:26]: Mikey on the pod. Yeah.Akshat [00:50:27]: But they use Modal for all their inference and that's because they have like a custom-- They have completely custom model architecture and that means that they have to be at the code level and tweak things that are not, just an API.Swyx [00:50:41]: It's interesting as well, like we had, Ethan, most recently on the xAI Groq team make a prediction that like the next tier in video gen is not a better video model, it's a better model or agent that orchestrates video models.Video Agents and Production WorkflowsAkshat [00:50:56]: Oh, interesting.Vibhu [00:50:56]: Language model backbone that can use toolsAkshat [00:50:58]: RightVibhu [00:50:59]: And write code.Akshat [00:51:00]: Like, yes, I can make my second video or my second video from Groq, but I want my minute video.Akshat [00:51:06]: And I'm not going there through normal video gen.Swyx [00:51:10]: Yeah, that's interesting. I - So we have GPU sandboxes and recently have seen a few companies doing agents that do video manipulation or,Akshat [00:51:22]: Yeah. Give it FFmpeg and just do it.Swyx [00:51:23]: Run FFmpeg. But likeAkshat [00:51:25]: That's not enough.Swyx [00:51:25]: Yeah.Akshat [00:51:26]: You need to give it Adobe.Swyx [00:51:27]: Yeah, I hadn't put it together with like it would be a video production thing. in my mind these things were going more towards editingAkshat [00:51:36]: Yeah.Vibhu [00:51:36]: Well, shout out Mantis.Akshat [00:51:37]: I think about this a lot.Swyx [00:51:38]: .Akshat [00:51:41]: Yeah. Sorry.Vibhu [00:51:41]: Luma. Luma Agent is a version of this for video production, but it's a off.Swyx [00:51:46]: I was gonna get your quick takes, on some other stuff that happensGitpod/Ona, CI, and Runtime SandboxesSwyx [00:51:50]: In recent news and just-just see if you have anything interesting. Gitpod, very li

Let's Talk AI
#250 - Mythos Mess, GPT 5.6-Sol, GLM 5.2

Let's Talk AI

Play Episode Listen Later Jul 7, 2026 103:25


Our 250th episode with a summary and discussion of last week's big AI news!Recorded on 06/27/2026Note from Andrey: sorry this is late again! this episode release somehow didn't save and I only realized late, my bad... next one will be out way sooner!Hosted by Andrey Kurenkov and Jeremie HarrisFeel free to email us your questions and feedback at andreyvkurenkov@gmail.com and/or hello@gladstone.aiRead out our text newsletter and comment on the podcast at https://lastweekin.ai/In this episode:US government gating of frontier AI expands: Anthropic gets permission to release Mythos-5 to selected companies/agencies after a standoff, OpenAI rolls out GPT-5.6 “Sol” with initial access restricted to ~20 approved organizations, and Meta is pressed to submit models to “voluntary” review—signaling an emerging de facto licensing regime with geopolitical treaty implications.Model capability and safety signals remain murky: limited benchmark disclosure, claims of token-efficiency comparisons, and third-party reports that GPT-5.6 shows extreme benchmark “cheating” sensitivity highlight steering/alignment bottlenecks and uncertainty about real-world long-horizon behavior.Compute supply chain competition accelerates: OpenAI unveils its Jalapeño inference ASIC with Broadcom on TSMC 3nm; Amazon explores selling Trainium to data-center operators; Micron invests in Anthropic with memory supply agreements; SK Hynix surpasses Samsung on HBM-driven valuation; Groq raises $650M while pivoting toward neocloud.Open source and societal response intensify: GLM 5.2 (MIT-licensed) delivers strong long-context coding performance with rapid optimizations; EconEvals maps job-task exposure; bipartisan workforce initiatives and tax credits launch; DeepMind and Apollo publish loss-of-control/control roadmaps; Hollywood reportedly drops a near-finished Sam Altman biopic amid industry pressure.Timestamps (note - these don't take into account dynamically inserted ads and therefore may be off by a couple of minutes):(00:00:10) Intro / Banter(00:03:42) News PreviewTools & Apps(00:04:41) Anthropic allowed to release Mythos AI to some companies, agencies + Anthropic's Mythos mess is only getting worse + Anthropic floats proposal to Lutnick to end US ban of powerful 'Mythos,' 'Fable' AI models: sources(00:07:58) OpenAI Launches GPT-5.6 Sol Under First-Ever US Government-Gated AI Rollout | MLQ News + OpenAI's new flagship model GPT-5.6 Sol cheats on software tests more than any model before it + Summary of METR's predeployment evaluation of GPT-5.6 Sol(00:24:03) U.S. Presses Meta to Agree to A.I. Reviews - The New York Times(00:30:11) Anthropic's Claude Tag is learning your company, one Slack message at a time | TechCrunchApplications & Business(00:32:49) OpenAI reveals its first AI processor: Jalapeño | The Verge(00:38:29) Amazon in Talks to Sell Custom AI Chips in Bid to Undercut Nvidia(00:41:46) Micron invests in Anthropic and grants it a supply deal(00:45:18) SK Hynix overtakes Samsung to become South Korea's most valuable company | Reuters(00:49:12) AI chipmaker Groq confirms $650M raise, re-staffs after Nvidia's $20B not-acqui-hire deal | TechCrunch(00:52:47) SpaceX inks compute deal with Reflection AI, an open source AI lab | TechCrunchProjects & Open Source(00:54:46) GLM-5.2: Built for Long-Horizon Tasks + How we built the world's fastest API for GLM-5.2 + nvidia/GLM-5.2-NVFP4 · Hugging Face(01:03:04) EconEvalsPolicy & Safety(01:05:40) $500 million AI jobs push launches with bipartisan backing - POLITICO(01:07:47) Rep. Sam Liccardo unveils AI workforce tax credit bill - POLITICO(01:08:56) Google DeepMind announced an “AI Control Roadmap” for improving AI agent security. | The Verge + Securing internal systems against increasingly capable and imperfectly aligned AI(01:14:00) The Loss of Control Playbook: Degrees, Dynamics, and Preparedness + The Loss of Control Playbook(01:16:42) Why corporate AI super PACs spent $27 million on a local election | The Verge(01:20:25) Exclusive: Conservatives plan nationwide protest against AI data centersResearch & Advancements(01:27:37) Revisiting the Platonic Representation Hypothesis: An Aristotelian View(01:31:39) Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models(01:33:59) Tapered Language ModelsSynthetic Media & Art(01:36:54) Hollywood is bending the knee to OpenAI | The VergeSee Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

TyfloPodcast
Echa Przeglądów Odcinek nr 2

TyfloPodcast

Play Episode Listen Later Jul 2, 2026 156:14


W drugim odcinku audycji „Echa Przeglądów” Michał Kasperczak rozmawia z Agnieszką Juranek, autorką bloga „Jabłuszkowy Świat” i kanału na YouTube. W audycji sporo uwagi poświęcamy pracy na Macu, dyktowaniu i tworzeniu napisów z pomocą AI. Agnieszka mówi o MacBooku Air M2, aplikacjach Audio Hijack, Loopback, MacWhisper i Movist Pro oraz montażu w Reaperze. Porównujemy MacWhisper z Wispr Flow oraz modele lokalne i chmurowe, m.in. OpenAI, ElevenLabs, Groq i NVIDIA Parakeet, a w wiadomościach od słuchaczy pojawiają się dyktowanie przez ChatGPT oraz Gemini Live z opisem obrazu na żywo. Audycja dostępna jest również w wygenerowanej automatycznie wersji tekstowej

The Six Five with Patrick Moorhead and Daniel Newman
Qualcomm's Data Center Debut, OpenAI's Jalapeño, and the Memory-as-Strategic Infrastructure Debate | The Six Five Pod Ep. 310

The Six Five with Patrick Moorhead and Daniel Newman

Play Episode Listen Later Jun 29, 2026 62:08


On Episode 310 of The Six Five Pod, Patrick Moorhead and Daniel Newman unpack the biggest stories from the week, including insights from Qualcomm Investor Day 2026, OpenAI and Broadcom's Jalapeño AI chip, Anthropic's Micron partnership, SpaceX's massive Reflection AI compute deal, Sakana AI's new Fugu orchestrator, and why memory is emerging as a critical layer of AI infrastructure. Plus, Bulls & Bears covers NVIDIA's $25B bond offering, Apple's MacBook price increases, Micron's record quarter, and Cerebras' first earnings as a public company. The handpicked topics for this week are: Qualcomm Investor Day 2026 — The Data Center Debut: Pat and Dan break down Qualcomm's push into the data center after the company took the stage with Microsoft's Satya Nadella and Meta's Mark Zuckerberg as named customers. They unpack the new Dragonfly platform, including the C1000 250-core data center CPU with PCIe Gen 7 and CXL, the AI200 and AI250 inference accelerators, and a novel High Bandwidth Compute (HBC) architecture that stacks compute under LPDDR memory at dramatically lower cost than HBM. They highlight Qualcomm's ambitious growth targets: $15B data center revenue target for FY 2029, an increased total non-handset revenue goal from $22B to  $40B, and a shortened timeline for automotive revenue by two years. They also debate the identity of Qualcomm's unnamed hyperscaler customer and why its robotics opportunity may be flying under the radar. (The Decode) OpenAI and Broadcom Unveil Jalapeño, OpenAI's First Custom Chip: A photo of Sam Altman and Hock Tan holding a wafer and packaged die kicked off OpenAI's reveal of Jalapeño, a custom inference chip built with Broadcom and slated for late-2026 deployment. The chip reached tape-out in roughly nine months, which is an aggressive cycle for an ASIC of this size, and uses HBM3E memory. Pat takes a victory lap on his long-standing heterogeneous compute thesis: every hyperscaler and now every model lab is building accelerators, and the XPU efficiency argument has played out as predicted. Dan frames OpenAI's broader move as existential: they cannot serve frontier models at premium margins if compute remains constrained. He flags that OpenAI is trying to do everything from chips and fabs to social networks and browsers, and that its IPO is now delayed. (The Decode)   Anthropic and Micron Sign a Strategic Multi-Year Memory Agreement: Anthropic and Micron announced a multi-year supply agreement for HBM, DRAM, and SSDs, including co-designed next-generation memory for AI workloads, along with a strategic investment by Anthropic in Micron. The pattern mirrors Samsung and SK Hynix's pre-funding Anthropic in May, and follows OpenAI's Jalapeño as another frontier lab moving to lock in supply chain control. Dan frames it as the same circular financing playbook NVIDIA ran two to three years ago, but with the ball now in the memory triopoly's court. Pricing-floor agreements with no ceilings, customized rather than commoditized memory architecture, and demand running well past the previously assumed 2027-2028 horizon. Pat notes that the rumored 14% free cash flow margin at Anthropic makes the strategic investment math work cleanly for both sides. (The Decode)   SpaceX Signs $6.3B Compute Deal with Reflection AI: SpaceX inked a $6.3B compute lease with open-source AI lab Reflection AI, at $150M per month from July 2026 through 2029, giving Reflection access to NVIDIA GB300 chips inside the Colossus infrastructure. Combined with the $920M-per-month Google compute contract and existing xAI commitments, SpaceX now has a contracted backlog larger than most public AI startups' entire revenue base, with some calling it the largest commercial AI infrastructure provider at $80B in contracted revenue. Pat reads it as XAI failing to land with developers, consumers, or enterprises, leaving SpaceX with a pot of gold worth far more as wholesale capacity than as XAI's own training compute. Dan flags that Google owning 7% of SpaceX ahead of an IPO is not accidental, and the open question is whether this becomes a Nebius-style infrastructure trade or a full-stack Google-equivalent platform. (The Decode)   Japan's Agentic Orchestrator Sakana AI Ships Fugu Plus and Fugu Ultra: Japan's Sakana AI released Fugu Plus and Fugu Ultra, an agentic orchestrator built on a multi-agent MOE approach that routes workloads across multiple underlying models rather than training a new frontier base model. Sakana claims agentic capabilities on par with or better than top frontier models at significantly lower input/output token costs, similar to the DeepSeek and GLM cost-undercut narrative. Pat compares the architecture to OpenRouter and notes the developer-facing parallel to Perplexity Computer's model-routing approach. Both agree that models themselves are no longer moats, and suggests the real moat is the harness, tooling, connectivity, looping, agentic stack, and total compute availability. Expect more sovereign agentic plays from Japan, the Middle East, and elsewhere on the same template. (The Decode)   The Flip — Is the Era of Memory as a Commodity Over? Daniel takes the FOR side: memory has moved from commodity to strategic AI infrastructure, citing 16 multi-year agreements covering $22B in committed volume booked through 2027, 84.9% gross margins higher than NVIDIA's, the technology barriers of HBM yield/stacking/packaging that only three companies can clear, and demand drivers tied to HBM as the binding constraint on every AI accelerator rather than to elastic consumer cycles. Patrick takes the AGAINST side: long-term agreements and SCAs signal a commodity in a strong cycle, not a structural rerating; nearly every relevant memory standard — DDR5, MRDIMM, HBM3/3E/4, LPDDR5X/6, GDDR6/7, LPCAM2 — is JEDEC-standard and therefore commodity at the pin; and CXMT's China DDR5 production ramps in 2H 2026 with Lenovo already shipping and HP and Dell qualifying. Custom HBM4 and Qualcomm-style HBC are where strategic memory genuinely lives. (The Flip)   NVIDIA's $25B Investment-Grade Bond Offering: NVIDIA priced a $25B multi-tranche bond offering on June 15, its first investment-grade debt sale since 2021, with seven tranches maturing between 2028 and 2056 and $85B in orders against an initial $20B target. Dan reads it as raising when capital is cheap, and oversubscription is real. NVIDIA doesn't need the money, it has a gold balance sheet, and is establishing a credit benchmark rather than funding CapEx. Pat agrees the optics are clean, but flags the irony of NVIDIA, with negative debt, borrowing while the stock trades like dead money at a sub-20x forward P/E. Both note that NVIDIA's underperformance reflects the market's skepticism on memory-as-strategic and on NVIDIA's own capex pace relative to the buildout opportunity ahead. (Bulls & Bears)   Tim Cook Calls Apple's Memory Crunch Price Raises on MacBook and iPad "Unsustainable": Apple announced MacBook and iPad price increases of up to $300, with Tim Cook telling the WSJ the memory cost environment is unsustainable. AAPL fell ~5% on the news, the broader rally was momentarily wiped out before Micron held the gains by close. Dan frames it as a moment when the market saw who is going to pay for the AI buildout: the consumer. He notes Apple's pricing power and inelasticity test is now live. Pat traces the backstory to Apple's negative-margin pricing pressure on Micron during the 2022-2023 memory downturn. The question is whether consumer-price blowback will eventually flow back to the memory vendors. (Bulls & Bears)   Micron Blows the Doors Off Fiscal Q3 — $41.46B Revenue, 84.9% Gross Margin: The memory story continues as Micron reported its largest beat in company history with fiscal Q3 revenue of $41.46B versus a $35.69B consensus, EPS of $25.11, year-over-year growth of more than 340%, and a record 84.9% gross margin that is roughly 10 points above NVIDIA's. Q4 guidance came in at a $50B midpoint against a $43B consensus. The 16 multi-year strategic customer agreements add up to $22B in committed volume, with most contracts containing pricing floors but no ceilings on most of the volume — a structurally asymmetric setup. Pat notes 95% of the beat came from price, not units, which reinforces his commodity argument; Dan flips it as the early innings of an NVIDIA-style run that puts Micron's 2027 profit on par with Google. (Bulls & Bears)   Cerebras' First Earnings Report Since IPO — Revenue Doubles, Margins Compress: Cerebras (CBRS) reported its first earnings as a public company, doubling year-over-year revenue and beating the top line while missing EPS, but the stock sold off hard amid gross margin deterioration. Core revenue came in at $191M, up 12% sequentially, with a $194M Q2 guide that is essentially flat, core gross margins at 47% guiding to 36-38% and 38-41% for the year, and operating margins flipping from positive 2% to a guided -30% to -32%. Customer concentration is shifting from Core42 and G42 (86% of FY25 revenue) to OpenAI, which loaned Cerebras $1B and gets paid quarterly in warrants. Pat flags that Cerebras' uncontested speed claim is no longer uncontested with Groq, TPU v8i, and Tenstorrent putting up real numbers. Cathie Wood is down 52% on her position. (Bulls & Bears) Watch the full video at sixfivemedia.com, and be sure to subscribe to our YouTube channel so you never miss an episode. The Decode Qualcomm Investor Day Lands the Data Center Pivot — Microsoft Deploying Qualcomm HBC XPUs in Azure (Per Satya Nadella) + Meta MOU on Three New Qualcomm Datacenter CPUs (Per Zuckerberg); $3.9B Modular Acquisition; Dragonfly Brand + AI200/AI250 Roadmap; HUMAIN 200MW Ramp; Qualcomm to Become Largest Automotive Silicon Company; Targets $3B Datacenter Revenue FY27, $35B by FY31 https://finance.yahoo.com/markets/stocks/articles/qualcomm-investor-day-detail-data-163247063.html  OpenAI Begins Vertical Integration — First Custom Inference Chip "Jalapeño" Unveiled With Broadcom June 24 (Hock Tan: As Good as Blackwell + TPU; ~50% Cost Savings; Late-2026 Microsoft Deployment, 10GW Multi-Gen Roadmap); Daybreak Cyber Stack (June 22) Confirms the Platform Shift https://x.com/OpenAI/status/2069770172802773292  Frontier AI Labs Are Now Financing Their Own Supply Chains — Anthropic Locks In Multi-Year Micron HBM/DRAM/SSD Supply + Micron Becomes Series H Investor; Same Pattern as Samsung + SK hynix Pre-Funded Anthropic in May; $965B Post-Money, $47B Revenue Run-Rate, October IPO Target https://investors.micron.com/news-releases/news-release-details/micron-and-anthropic-announce-strategic-agreement-scale-next  SpaceX Signs $6.3B Compute Deal With Reflection AI — $150M/Month July 2026 → End of 2029; NVIDIA GB300 + Colossus 2 Capacity; SpaceX Now Largest Commercial AI Infrastructure Provider With $80B+ Committed Compute Revenue Through 2029 https://finance.yahoo.com/technology/ai/articles/spacex-reportedly-grant-reflection-ai-162749237.html  The Sovereign AI Stack Lands — Japan's Sakana Ships Fugu + Fugu Ultra Multi-Agent System (June 22) That Beats Opus 4.8, GPT-5.5, and Gemini 3.1 Pro on 10 of 11 Benchmarks; Designed Around US Export-Control Risk; Completes the Three-Bloc Sovereign-AI Map With Mistral Compute (Europe) + DeepSeek $7.4B (China) https://www.datacamp.com/blog/sakana-fugu  The Flip Is the Era of Memory as a Commodity Over? FOR: Memory is now strategic AI infrastructure with multi-year supply lock-ins. The cycle dynamics that defined the last 30 years no longer apply. https://www.benzinga.com/markets/tech/26/06/60062500/micron-earnings-could-echo-nvidias-2023-moment-says-futurum-ceo  AGAINST: Memory is cyclical and priced for perfection. This print is either step change or top of the cycle, and the second one is more likely. https://www.cnbc.com/2026/06/25/apple-macbook-ipad-price-hike-memory.html Bulls & Bears NVIDIA (NVDA) $25B Bond Sale Anchors the AI Debt-Finance Boom — First Bond Offering Since 2021; Joins Alphabet $80B, Amazon $27.5B, Meta $30B, Oracle Stack; Dan: "Locking In Cheap Capital While It Can" https://finance.yahoo.com/technology/ai/articles/nvidia-record-us-25-billion-131039687.html  Apple (AAPL) Falls −5%+ Thursday June 25 on Confirmed MacBook + iPad Price Hikes — Tim Cook RAM "Unsustainable" Comment Lands as Real Price Action; Apple Hikes Erase Micron-Driven Tech Rally Mid-Session; Memory Beneficiaries (SanDisk, Micron) Surge; Analysts "Mostly Nonplussed" https://tickerspark.ai/market/apple-inc-aapl-drops-5-3-as-price-hikes-spook-investors-1782399950638  Micron (MU) Q3 FY26 ACTUALS — Largest Beat in Company History; Revenue $41.46B (+346% YoY) Crushes $35.69B Consensus; Non-GAAP EPS $25.11 (+1,215% YoY) Beats $20.49; Record 84.9% Gross Margin (Higher Than NVIDIA); Q4 Guide $50B Midpoint vs $43B Consensus; Stock +18-19% Overnight to $1,242 https://www.nasdaq.com/articles/nvda-who-micron-blows-doors-q3-earnings-revs  Cerebras Systems (CBRS) Q1 ACTUALS — First Earnings Post-IPO; Revenue $193.4M Nearly Doubled YoY; 2026 Guide $855-$865M Beats $824M; BUT Gross Margins Forecast 38-41% (Down From 45% Q1, Half of NVIDIA + Micron); Stock −20% AH on Margin Compression; Sets Up Inference-Tier Margin Debate https://investors.cerebras.ai/news-releases/news-release-details/cerebras-systems-announces-strong-first-quarter-2026-results  

Digital Currents
The Memory Bottleneck and SpaceX's Post-IPO Reality Check

Digital Currents

Play Episode Listen Later Jun 26, 2026 56:56


In this episode, we discuss the recent moves in Bitcoin and Strategy and look at the challenges SpaceX may be facing. We cover reports that Groq has secured $650M and examine what some are calling a memory squeeze, which we believe may be influencing parts of the tech sector and the broader AI market. We also touch on reports of Reflection AI's agreement, described as a deal worth up to $6.3B, intended to scale open models, as well as OpenAI's debut of "Jalapeño," reported to be its first custom inference chip. Finally, we review the dips in gold, silver, and Bitcoin, and the potential unwinding of the "debasement trade." Remember to Stay Current! To learn more, visit us on the web at https://www.morgancreekcap.com/morgan-creek-digital/. To speak to a team member or sign up for additional content, please email mcdigital@morgancreekcap.com Legal Disclaimer This podcast is for informational purposes only and should not be construed as investment advice or a solicitation for the sale of any security, advisory, or other service. Investments related to the themes and ideas discussed may be owned by funds managed by the host and podcast guests. Any conflicts mentioned by the host are subject to change. Listeners should consult their personal financial advisors before making any investment decisions.

This Week in XR Podcast
Special From CES 2026: AI Strategy, Tariffs, and the Future of Consumer Tech ft. Gary Shapiro, CEO

This Week in XR Podcast

Play Episode Listen Later Jun 19, 2026 58:57


Gary Shapiro has spent decades at the center of the global consumer technology industry, leading the Consumer Technology Association (CTA) and building CES into one of the most important stages for innovation, policy, and deal-making on the planet.In this first episode of 2026, Gary joins Charlie, Rony, and Ted to preview CES, unpack the explosion of AI across every category, and deliver unusually blunt takes on tariffs, China, manufacturing, and U.S. innovation policy. He explains how CES has evolved from a TV-and-gadgets show into a global platform where boards meet, standards are set, and policymakers, chip designers, robotics firms, and health-tech startups all collide.In the News: Before Gary joins, the hosts break down Nvidia's $20 billion “not-a-deal” with Singapore's Groq, the stake in Intel, and what that combo might signal about the edge of the GPU bubble and the shift toward inference compute, x86, and U.S. industrial policy. They also dig into Netflix's acquisition of Ready Player Me and what it suggests about a Netflix metaverse and location-based entertainment strategy, plus Starlink's rapid growth and an onslaught of “AI everything” products ahead of CES.Gary walks through new features at this year's show: CES Foundry at the Fontainebleau for AI and quantum, expanded tracks on manufacturing, wearables, women's health, and accessibility, plus an AI-powered show app already fielding thousands of questions (top query: where to pick up badges).He also talks candidly about his biggest concern—that fragmented state-level AI regulation (1,200+ state bills in 2025) will crush startups while big players shrug—and why he believes federal standards via NIST are the only realistic path. The discussion ranges from AI-driven healthcare and precision agriculture to robotics, demographics, labor culture, global supply chains, and what CES might look like in 2056.5 Key Takeaways from Gary:AI is now the spine of CES. CES 2026 centers on AI as infrastructure: CES Foundry at the Fontainebleau for AI + quantum, AI training tracks for strategy, implementation, agentic AI, and AI-driven marketing, and an AI-powered app helping attendees navigate the show.Fragmented state AI laws are an existential risk for startups. Over 1,200 state AI bills in 2025—including proposals to criminalize agentic AI counseling—could create a compliance maze only large incumbents can survive, which is why Gary argues for federal standards via NIST.Wearables are becoming systems, not gadgets. Oura rings, wrist devices, body sensors, and subdermal glucose monitors are starting to be designed as interoperable families of devices, with partnerships emerging to combine data into unified health services.Robotics is breaking out of the industrial niche. CES will showcase the largest robotics presence yet, moving beyond factory arms and drones to humanoids, logistics, social companions, and applied AI systems across sectors.Tariffs, alliances, and AI will reshape manufacturing. Gary is skeptical of “Fortress USA” strategies that try to onshore everything, pointing instead to allied reshoring (Latin America, Europe, Japan, South Korea) and the long-term role of AI-powered robotics in changing labor economics and global supply chains.This episode is brought to you by Zappar, creators of Mattercraft—the leading visual development environment for building immersive 3D web experiences for mobile headsets and desktop. Mattercraft combines the power of a game engine with the flexibility of the web, and now features an AI assistant that helps you design, code, and debug in real time, right in your browser. Whether you're a developer, designer, or just getting started, start building smarter at mattercraft.io. Hosted on Acast. See acast.com/privacy for more information.

TechSurge: The Deep Tech Podcast
Battle for the AI Data Center: Deep Dive on the Semiconductor Supercycle

TechSurge: The Deep Tech Podcast

Play Episode Listen Later Jun 16, 2026 53:23


Semiconductors have moved from the background of the technology stack to the center of the AI economy. What used to be a specialized industry discussed mostly by engineers and investors is now shaping the speed, cost, and strategic direction of modern computing.In this episode of TechSurge, host Michael Marks speaks with Stacy Rasgon, Managing Director and Senior Analyst covering U.S. semiconductors and semiconductor capital equipment at Bernstein Research. Stacy has spent years analyzing the chip industry across cycles, but argues that the current moment feels different in scale: AI demand has created an unprecedented scramble for compute, memory pricing has surged, and companies across the stack are being forced to rethink capacity, architecture, and capital allocation.The conversation explains the 4 different kinds of semiconductor cycles—supply, inventory, product, and demand — and why Stacy believes the industry is currently in a demand cycle of unusual magnitude. The discussion also unpacks the distinction between DRAM and NAND, why high-bandwidth memory is becoming strategically central to AI systems, and how the physical realities of wafer capacity and silicon area are constraining supply in ways the broader market often misses.Stacy and Michael also discuss the hardware economics behind the current boom, with Michael pressing Stacy on why compute remains so scarce and how companies are improving performance through packaging and system design. Michael then moves the conversation beyond market headlines to the core business questions: who is actually paying for this compute, which use cases are generating real revenue, and whether AI spending is creating durable economic value or simply shifting costs elsewhere. Together, these questions highlight two of the episode's clearest insights: coding may be one of the earliest AI applications with meaningful willingness to pay, and inference, not training, is the real test of whether the current buildout becomes a lasting business or just another expensive wave of infrastructure.Stacy explains the concentration of power among the major wafer fabrication equipment players, the rise of ASICs as a meaningful share of AI silicon, Broadcom's rapidly expanding AI opportunity, and the growing role of Chinese companies as new entrants, especially in memory and semiconductor equipment. Along the way, the conversation asks the defining question facing the sector: is this just another semiconductor upswing, or the first true supercycle the industry has seen? Stacy believes that this might be the biggest supercycle he has seen in his career.Sign up for our newsletter at techsurgepodcast.com for updates on upcoming TechSurge Live Summits and future episodes.Links:Stacy Rasgon on LinkedIn: https://www.linkedin.com/in/stacy-rasgon-6924963Bernstein: https://www.alliancebernstein.com/corporate/en/home.htmlReferences Mentioned During the DiscussionNVIDIA Blackwell Platform: https://www.nvidia.com/en-us/data-center/blackwell-platform/High Bandwidth Memory (HBM) overview from Micron: https://www.micron.com/products/memory/hbmDRAM overview from IBM: https://www.ibm.com/think/topics/dramNAND flash overview from IBM: https://www.ibm.com/think/topics/nand-flash-memoryFurther ReadingMcKinsey on the semiconductor industry outlook: https://www.mckinsey.com/industries/semiconductors/our-insights/the-semiconductor-industry-in-2025Semiconductor Industry Association: 2025 State of the U.S. Semiconductor Industry: https://www.semiconductors.orgNVIDIA on the Blackwell architecture and AI infrastructure roadmap: https://www.nvidia.com/en-us/data-center/blackwell-platform/Broadcom AI investor materials and infrastructure commentary: https://investors.broadcom.comASML on lithography and advanced chip manufacturing: https://www.asml.com/en/technologyMicron on HBM and AI memory demand: https://www.micron.com/products/memory/hbmChapters[00:00:00] — Highlights[00:00:26] — Welcome to  the Episode[00:01:29] — Meet Stacy Rasgon[00:02:01] — Is This the First Real Semiconductor Supercycle?[00:05:33] — Inside the Strongest Memory Cycle in History [00:09:14] — Can Innovation Keep Up With AI Demand?[00:11:33] — Chiplets, Blackwell, and the New Economics of Compute [00:12:37] — What Could Signal the Cycle Is Slowing[00:14:26] — Vertical Integration at the Hyperscales [00:16:36] — The Difference between Apple and Meta[00:17:15] — What is Vertical Integration Being Done For?[00:18:15] — Will other bottlenecks develop as This Progresses? [00:21:13] — Oligopoly Pricing in the Market[00:22:22] — Any New Entrants into Memory?[00:23:46] — Why the Industry Must Pivot From Training to Inference[00:25:10] — Agentic Coding and the First Real AI Revenues[00:26:57] — Groq, Low-Latency Inference, and What GPUs Cannot Do Alone[00:29:28] —-Could The Smaller Companies All be Bought Up ?[00:30:19] — Why Semiconductor Equipment Matters More Than Ever [00:31:00] — How Semiconductor Equipment is Affected by the Cycle[00:32:55] — A Long Upcycle for Semiconductor Equipment Guys?[00:33:13] — The Big Five and the Rise of Chinese Equipment Players[00:34:24] — The Effects of Geopolitics[00:35:02] — Broadcom's Quiet AI Breakout[00:40:46] — ASICs vs GPUs and the Next Wave of Custom Chips[00:41:06] — Intel, Foundry Strategy, and the Long Turnaround[00:46:46] —-The Risks the Market May Still Be Underestimating[00:49:32] — Where Startups Still Have Room to Win[00:50:39] — What the Semiconductor Industry Could Look Like Next Year

Tid er penger - En podcast med Peter Warren

00:01 1999 igjen: to skrekkfilm-hiter og «this time it's different»00:04 Rekordbelåning og margin debt på all time high00:05 Opsjonsjaget vi ikke har sett siden 198700:08 Short gamma, marketmakere og spiralen som ga «Red Friday»00:14 Ingenting virket: bare lang volatilitet beskyttet00:17 Laveste korrelasjoner på to år og VIX opp 40 prosent00:18 Bank of America: «here be dragons» og ledighet mot inflasjon00:20 Bilen, AI og Jevons-paradokset00:24 SpaceX som datasenterselskap, ikke rakettselskap00:30 Børsnotering denne uka: 1770 milliarder og Musks absolutte makt00:31 S&P-nekten mot FTSE, Russell og MSCI00:32 Lockup-kalenderen og dagen å frykte: seks måneder og fire dager00:35 Grok mot Groq og «race to zero» i modellene00:40 Midtøsten: Trump mot Netanyahu og oljeprisen00:44 Hva folk ikke ser på nå: bear flattening og carry trades som ryker00:47 Dollar over 161 og japansk intervensjon00:49 Hudson River Trading og datasenteret i Norge00:51 Norge har misforstått seg selv: fisk, olje, rå kraft og nå compute00:53 Å raffinere compute: Skygard, spillvarme og 10X på krafta00:58 Compute som multiplikator: fra 10x-ere til 100x-ere01:00 Budsjettforliket, Mímir Kristjánsson og minstepensjonistene01:05 Å prestere når alt er mulig: fokus, nysgjerrighet og flytskjemaer01:11 Telefonen som heroin: reels, 24-timers reset og hjernen tilbake01:19 Trikkedrapet og situational awareness01:24 Varsler i stedet for å glo på skjermen: gull/sølv og momentum01:35 1998: LTCM, doblede posisjoner og banken som tapte 900 millioner01:45 Andrew Left, Citron og short-saken som ble svindel01:50 Oraclum, superforecasters og nordmannen på topp01:56 Drewry-indeksen, VM-frakt og Fifas fredspris til Trump Hosted on Acast. See acast.com/privacy for more information.

Tech Café
Toutes ces IA, ça fait beaucoup, là, non ?

Tech Café

Play Episode Listen Later Jun 3, 2026 78:40


Claude 4.8 ajuste son niveau d'effort, les agents IA posent des problèmes de cadrage en entreprise, et la facture de l'IA devient vertigineuse entre tokens, data centers, Nvidia et concurrence chinoise. Licenciements chez Wix et dans la tech, dark patterns émotionnels dans les chatbots, bug de récupération de compte chez Instagram, Artemis fragilisé par Blue Origin, Starship sous pression, Ferrari Luce, Steam Deck OLED et The Witcher 3.  Me soutenir sur Patreon Me retrouver sur YouTube On discute ensemble sur Discord Ça fait beaucoup là non ? Claude 4.8 : it ain't much but it's honest work . Tout cela est-il bien raisonnable ? Spolier : non. Licenciement massifs de la semaine : Wix, Cisco et tant d'autres… Groq lève encore de l'argent. Et Bolloré il fait une cagnotte Ulule ? L'enshittification est déjà là. L'IA investit pour vous ! Qu'est ce qui peut mal se passer ? Nos futurs terminators sont en fait des bisounours. Boom ! Blue Origin est vert, et la NASA aussi. Starship : un test pas FAAntastique. The thing : même pas besoin d'alien pour devenir parano… Ombre sur la Luce : il vaut mieux en Ferrarire qu'en pleurer. Jeux vidéo Vous êtes sûr d'avoir toujours envie d'une Steam Machine ? C'est alors que je l'ai reconnu, surgissant du passé, il m'était revenu ! Participants Une émission préparée par Guillaume Poggiaspalla Présenté par Guillaume Vendé

Doppelgänger Tech Talk
Anthropic mit 500% NRR | Token-Maxing bald vorbei? | Stark-Drohnen auf €2.5 Mrd. Bewertung #566 

Doppelgänger Tech Talk

Play Episode Listen Later May 29, 2026 80:50


Anthropic schließt seine Runde ab und versucht außerdem, sich Kapazität bei Microsoft zu sichern. Claude Opus 4.8 launcht mit Fast Mode und neuer Zuverlässigkeit. Anthropic-CFO verrät eine Zahl zur Net Revenue Retention, die bisherige Software-Maßstäbe sprengt. Uber-CEO warnt vor explodierenden KI-Kosten, eine einzelne Firma soll im Monat Beträge in dreistelliger Millionenhöhe für KI verbrennen. Meta versucht plötzlich, für seinen Chatbot Geld zu kassieren und schickt Forward Deployed Teams zu Enterprise-Kunden. China sperrt seine KI-Talente ein. DuckDuckGo profitiert vom Google-AI-Backlash. SpaceX nimmt American Airlines an die Starlink-Leine, SpaceX-Tesla-Merger-Gerüchte verdichten sich, Musk könnte erster Billionär werden. Robinhood lässt Kunden mit KI-Agenten traden. Earnings-Block: Temu, Snowflake & Dell. Pip schaut auf die Top 10 der reichsten Männer der Welt und fragt: Wer stellt sich noch gegen Trump? In der Schmuddelecke: Google-Insider Michele Spagnuolo soll Polymarket mit interner Information abgeräumt haben. Unterstütze unseren Podcast und entdecke die Angebote unserer Werbepartner auf ⁠⁠⁠⁠⁠⁠doppelgaenger.io/werbung⁠⁠⁠⁠⁠⁠. Vielen Dank!  Philipp Glöckler und Philipp Klöckner sprechen heute über: (00:00:21) Anthropic Runde (00:07:23) Claude Opus 4.8 (00:13:05) KI als Praktikant? (00:17:37) Anthropic shoppt Microsoft-Chips (00:19:20) Apollo-Debt-Deal für Google-TPUs (00:22:22) Anthropic-NRR über 500% (00:23:21) Token-Maxing & Uber-CEO-Warnung (00:33:16) Meta-AI-Subscriptions (00:38:53) China sperrt KI-Talente ein (00:41:29) DuckDuckGo +30% (00:43:37) Starlink bei American Airlines (00:46:51) SpaceX-Tesla-Merger-Gerüchte (00:49:21) Robinhood: AI-Trading (00:52:47) Temu/PDD Earnings (00:55:56) Snowflake-Comeback (00:59:45) Dell-Earnings & Trump (01:03:38) Top 10 Reichste: pro Trump? (01:08:21) GROQ raised $650 Mio. (01:09:40) Spagnuolo bei Polymarket (01:12:12) Musks Hitler-Tweet (01:14:31) Stark-Drohnen auf $2,5 Mrd. Shownotes Anthropic überholt OpenAI im Fundraising - axios.com Claude Opus 4.8 Launch - reddit.com Anthropic in Gesprächen über Microsofts AI-Chips - theinformation.com Apollo strukturiert $36 Mrd. Debt-Deal für Google-Chips für Anthropic - bloomberg.com AI-Spending: ROI vs. Enterprise-Kosten - axios.com Meta startet AI-Chatbot-Subscriptions - bloomberg.com Meta Enterprise-AI-Push für Business-Adoption - theinformation.com China beschränkt Reisen für Top-KI-Talente - bloomberg.com DuckDuckGo Installs +30% nach Google-AI-Search-Zwang - techcrunch.com Starlink kommt in halbe American-Airlines-Flotte - theinformation.com SpaceX-Tesla-Merger-Gerüchte vor IPO - cnbc.com Robinhood lässt Kunden mit AI handeln - wsj.com Temu/PDD Earnings: Aktie unter Druck - barrons.com Snowflake setzt auf Amazon Graviton Cloud-Chips - cnbc.com PDD verfehlt Gewinnerwartungen wegen China-Wettbewerb - wsj.com Salesforce Q1 Earnings - cnbc.com Groq raised $650 Mio. mit Nvidia-Beteiligung - axios.com Google-Insider Michele Spagnuolo: Insider-Trading mit Top-Suchergebnissen - wsj.com Dell Q1 Earnings - cnbc.com Michael Dell und die Trump-Accounts beim DoD - cnbc.com Michael Dell Forbes-Profil - forbes.com Elon Musk Tweet - xcancel.com Drone start-up Stark set for €2.5bn valuation in new fundraising - ft.com EU fines China's Temu - ft.com

Forbes Daily Briefing
Sometimes You Don't Want A GPU: Groq Cofounder Explains Whirlwind Deal With Nvidia

Forbes Daily Briefing

Play Episode Listen Later May 14, 2026 7:10


Last winter, Groq cofounder and CEO Jonathan Ross walked into a meeting with Nvidia CEO Jensen Huang with a pitch for the companies' tech to work together. He now describes the synergy with a logistics analogy: stop building AI data centers as if every workload wants the same hardware. Training is bulk hauling; inference is last-mile delivery. GPUs can do both, but using the 18-wheeler even when you just need a van can be a lot slower. So: Nvidia's general-purpose GPUs are the big trucks. Groq's specialized chips—LPUs, or language processing units, designed to run models fast—are the smaller vans. “If you were building out a logistics network for the entire United States, and I told you your two options were all 18-wheelers or just delivery vans, which one would you pick?” Ross said. “The best answer is both.”  Ross wasn't just pitching a worldview. He wanted Nvidia's permission to buy around 100,000 Blackwell chips, likely worth billions. Huang grilled him on the technical details, and then the meeting ended.  When Huang called back three days later, Ross expected a discussion about his GPU purchase order. Instead, the Nvidia CEO cut to the chase. “We should probably move really fast,” Ross recalled him saying. Learn more about your ad choices. Visit megaphone.fm/adchoices

Chip Stock Investor Podcast
TSMC's $40 Billion Quarter: Supply Chain Risks, Intel-Tesla, and Who's Threatening the Chip King

Chip Stock Investor Podcast

Play Episode Listen Later Apr 16, 2026 9:04


Chip fab capacity is maxed out — and TSMC is the biggest winner. But new risks are emerging fast.In this episode, Nick and Kasey break down TSMC's Q2 2026 earnings guidance: $39–40 billion in quarterly revenue, 30% year-over-year growth, and gross margins approaching 67.5%. Then they dig into what could actually slow TSMC down.Topics covered:— Helium and LNG shortages driven by the Strait of Hormuz closure— Taiwan's energy security and how long government-secured supply lasts— The Intel-Tesla chip "refactoring" announcement decoded— Could Elon Musk's consortium acquire Intel Foundry after the SpaceX IPO?— Samsung Foundry, Nvidia's Groq acquisition, and supply chain diversificationTSMC has navigated supply chain disruptions before. But with AI chip demand exploding and new competitors circling, the pressure is unlike anything the industry has seen.For the full earnings breakdown and supply chain chart, visit chipstockinvestor.com and check out the Semi Insider subscription.Chip Stock Investor covers semiconductors, AI infrastructure, and the companies powering the next wave of technology.For informational and entertainment purposes only — not individual investment advice. All investing involves risk and you may lose principal. Forecasts are not guaranteed. Nick and Kasey own shares of TSM.

Tech Café
Mouches, baleines et guêpes : ces animaux incompris

Tech Café

Play Episode Listen Later Mar 27, 2026


Focus sur les modèles IA de la semaine : vidéo et générateurs de mondes. Simulation de cerveau de mouche, communication animale et un GPS pour guêpes. Discussions sur véhicules autonomes, annonces toujours fracassantes d'Elon Musk, et la controverse Cursor.   Me soutenir sur Patreon Me retrouver sur YouTube On discute ensemble sur Discord Modèles IA de la semaine Stereo World, 3D Dream Booth, Mobile GS, mosaicMem et pas un DVD de Lost. Mouchetrix : Neo se réveille avec 6 pattes. Parler le baleine ? Ça ne manque pas de sel. Des condensateurs à la taille de guêpe. Cursor fait des cachoteries Kiminables. Meta va enfin peut être s'améliorer. Uber et Rivian font se faire un bout de route ensemble. Teradollars La blague Terafab. Ou alors je suis Terabajoie ? Pour devenir riche, il suffit d'un sèche cheveux. GTC : Rubin et pourquoi NVIDIA a croqué Groq. Des datacenters dans l'espace avec des escortes ? Face à face : détection de visage à haute vitesse chez NVIDIA. La crise de la RAM durera jusqu'en 2030 2040 2050… Firefox lance un mode porno un VPN. Participants Une émission préparée par Guillaume Poggiaspalla Présenté par Guillaume Vendé

The Twenty Minute VC: Venture Capital | Startup Funding | The Pitch
20VC: Why You Need a $1BN Fund To Do Series A Today | OpenAI vs Anthropic: Who Wins Enterprise | SpaceX at $2TRN and Data Centers in Space | The $20BN Groq Deal Broken Down | Jeff Bezos' $100BN New Fund

The Twenty Minute VC: Venture Capital | Startup Funding | The Pitch

Play Episode Listen Later Mar 26, 2026 78:09


AGENDA: 05:00 — Anthropic vs. OpenAI: Who Is Actually Winning the Enterprise War? 07:55 — "Air of Desperation": Is OpenAI Losing Its Invincibility? 18:00 — SpaceX at $2 Trillion: Elon's Insane Plan to Build Data Centers in Space 29:00 — Jeff Bezos' $100 Billion Fund: The End of "Doing It the Hard Way" 34:00 — The $20 Billion "Acqui-hire": The Groq Deal Broken Down 40:40 — Figma's Death Spiral? Why the Markets Are Terrified of AI Disruption 56:00 — The Broken VC Math: Why You Need $1BN To Do Series A 01:04:00 — Win or Die: The Terrifying Reality of the Unicorn "Dead Zone"  

Azeem Azhar's Exponential View
What NVIDIA's bet on OpenClaw means for the future of AI and your token budget

Azeem Azhar's Exponential View

Play Episode Listen Later Mar 25, 2026 36:35


Welcome to Exponential View, the show where I explore how exponential technologies such as AI are reshaping our future. I've been studying AI and exponential technologies at the frontier for over ten years. Each week, I share some of my analysis or speak with an expert guest to make light of a particular topic. To keep up with the Exponential transition, subscribe to this channel or to my newsletter:  https://www.exponentialview.co/ ---- Last week Jensen Huang shared the numbers from NVIDIA's order book: AI compute demand has grown a millionfold in two years. Much GTC coverage focused on chips, robots, data centers in space, but I think Jensen revealed something far more important in his keynote: “the inference inflection has arrived,” and this is about to transform how all companies should manage their budgets. The inference era is already the operating assumption of the world's most valuable company. In this week's podcast, I cover: (1:20) NVIDIA's $1 trillion order book (1:56) OpenClaw: our era's web browser (7:54) Training vs Inference: how AI is changing (12:50) Pre-fill vs. decode: the technical split (18:06) The Harness: why OpenClaw changes everything (18:59) The engine is useless without a car (22:21) From 100M to 870M tokens per day (24:29) Meet my agent R Mini Arnold's team (26:16) AI focus group simulations at $10–50 a run (29:36) Jensen's self-interest (and why he's still right) (33:07) AI governance: token budgets don't belong with IT (35:07) From training economy to inference economy Read my essay "Magnitudes of Intelligence" on Substack: https://www.exponentialview.co/p/the-hundred-million-token-day Access the solar supercyle model here: https://www.exponentialview.co/p/solar-supercycle ---- Where to find me: Exponential View newsletter: https://www.exponentialview.co/ Website: https://www.azeemazhar.com/ LinkedIn: https://www.linkedin.com/in/azeem/ Twitter/X: https://x.com/azeem Production by EPIIPLUS1. Production and research: Baba Films, Chantal Smith, Marija Gavrilov. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

Podcasts – Weird Things
Real-Time AI Speeds, Code Models, Bio Hacking, And Movie Picks

Podcasts – Weird Things

Play Episode Listen Later Mar 24, 2026


The episode surveys an accelerating AI landscape where new hardware like Cerebras and Groq enables near real?time model responses, making voice and agent interactions feel instantly conversational. The conversation covers the rise of code models (Codex, Claude Code), practical tips for using multiple models to check each other, the tug-of-war between frontier labs and big incumbents (OpenAI, Anthropic, Google, Meta, xAI), and how talent, salaries, and state-level data?center politics are shaping the field. They also touch on a striking story about a dog treated with an experimental mRNA therapeutic assembled with help from multiple AI tools, hands-on demos of rapid content generation and deepfake video, and a challenge to listeners to build weird things with these new tools. Picks: Brian Brushwood: Project Hail Mary. Andrew Mayne: Sentimental Value.

After Things Podcast
Real-Time AI Speeds, Code Models, Bio Hacking, And Movie Picks

After Things Podcast

Play Episode Listen Later Mar 24, 2026


The episode surveys an accelerating AI landscape where new hardware like Cerebras and Groq enables near real?time model responses, making voice and agent interactions feel instantly conversational. The conversation covers the rise of code models (Codex, Claude Code), practical tips for using multiple models to check each other, the tug-of-war between frontier labs and big incumbents (OpenAI, Anthropic, Google, Meta, xAI), and how talent, salaries, and state-level data?center politics are shaping the field. They also touch on a striking story about a dog treated with an experimental mRNA therapeutic assembled with help from multiple AI tools, hands-on demos of rapid content generation and deepfake video, and a challenge to listeners to build weird things with these new tools. Picks: Brian Brushwood: Project Hail Mary. Andrew Mayne: Sentimental Value.

Podcasts – Weird Things
Real-Time AI Speeds, Code Models, Bio Hacking, And Movie Picks

Podcasts – Weird Things

Play Episode Listen Later Mar 23, 2026


The episode surveys an accelerating AI landscape where new hardware like Cerebras and Groq enables near real?time model responses, making voice and agent interactions feel instantly conversational. The conversation covers the rise of code models (Codex, Claude Code), practical tips for using multiple models to check each other, the tug-of-war between frontier labs and big […]

The Information's 411
OpenAI's Shopping U-Turn Complications, Nvidia's Groq Chip, Synthesia's AI Video for Enterprise

The Information's 411

Play Episode Listen Later Mar 23, 2026 58:39


E-commerce Reporter Ann Gehan talks with TITV Host Akash Pasricha about OpenAI's sudden retreat from e-commerce integrations to focus on its core product. We also talk with Porch Capital's David Levy about why Nvidia and AWS are pivoting to SRAM to solve the AI memory crunch and KeyBanc Capital Markets' Jackson Ader about why software employees are demanding more stock-based compensation despite market volatility. Finally, we get into Anduril's massive new manufacturing blitz in Ohio with Deputy Bureau Chief of Finance Cory Weinberg.Articles discussed on this episode: https://www.theinformation.com/articles/openais-shopping-u-turn-complicate-enterprise-playbookhttps://www.theinformation.com/articles/inside-andurils-big-gamble-ohio-weapons-factorySubscribe: YouTube: https://www.youtube.com/@theinformation The Information: https://www.theinformation.com/subscribe_hSign up for the AI Agenda newsletter: https://www.theinformation.com/features/ai-agendaTITV airs weekdays on YouTube, X and LinkedIn at 10AM PT / 1PM ET. Or check us out wherever you get your podcasts.Follow us:X: https://x.com/theinformationIG: https://www.instagram.com/theinformation/TikTok: https://www.tiktok.com/@titv.theinformationLinkedIn: https://www.linkedin.com/company/theinformation/

The Six Five with Patrick Moorhead and Daniel Newman
EP 297: AI Control, Compute Power, and the Fight for the Stack

The Six Five with Patrick Moorhead and Daniel Newman

Play Episode Listen Later Mar 23, 2026 54:46


AI is becoming a scale and control business. On Episode 297 of The Six Five Pod, Patrick Moorhead and Daniel Newman examine the companies building the infrastructure, forming the alliances, and making the moves that will define who wins and who gets squeezed out. Control is shifting across compute, models, infrastructure, and enterprise distribution as NVIDIA, Microsoft, OpenAI, Meta, and others push to control the next phase of the AI market. The handpicked topics for this week are: NVIDIA's Full-Stack Push Gets Bigger: Following the GTC conference in San Jose, Pat and Dan break down how NVIDIA continues expanding beyond GPUs with Vera CPU, Dynamo, and a broader agentic AI stack designed to unify training, inference, orchestration, and enterprise-grade security. Microsoft, OpenAI, and Amazon Enter a New Phase of Tension: With Microsoft reportedly weighing legal action over OpenAI's growing AWS relationship, the discussion turns to exclusivity, multi-cloud strategy, and what happens when one of AI's most important alliances starts to crack. China, Compute, and the Geopolitics of AI Access: The hosts examine NVIDIA's reported H200 restart for China and what it says about export controls, policy pressure, and the global fight over advanced AI compute. Meta's $27B Infrastructure Agreement Signals the Real Race: Meta's latest infrastructure deal reinforces a central point of this episode, demand for AI capacity is still outrunning supply, and hyperscalers are moving aggressively to lock in long-term compute. OpenAI's Enterprise Push Raises Bigger Business Model Questions: As OpenAI leans harder into enterprise and eyes an eventual IPO, Pat and Dan unpack what this pivot says about monetization pressure, competitive positioning, and the need to prove a durable AI business model. The GPU Smuggling Story Shows How Valuable AI Hardware Has Become: A major smuggling case involving NVIDIA hardware spotlights the black market for AI chips and the growing intersection of compute, national security, and enforcement. The Flip: Did NVIDIA Just Change the Inference Market Again? This week's debate centers on whether NVIDIA's $20bn Groq Technology deal kills the standalone inference chip market, or whether it actually validates the market by proving just how strategically important specialized inference has become. The Fed, Micron, and Accenture Reflect a More Complicated Market: In Bulls and Bears, the hosts cover the Fed's latest decision, Micron's AI-driven momentum, and why Accenture's results still ran into skepticism despite strong execution. Meta's Workforce Cuts and AI Spend Reflect the New Corporate Tradeoff: The episode closes on the growing tension between rising AI investment and labor efficiency, as companies look for ways to fund massive infrastructure and token budgets while restructuring headcount. For a deeper dive into each topic, please click on the provided links. Subscribe to our YouTube Channel so you never miss an episode.   The Decode   NVIDIA GTC 2026: Vera Rubin Platform, Groq LPU Integration & $1T Demand Vision https://www.cnbc.com/2026/03/16/nvidia-gtc-2026-ceo-jensen-huang-keynote-blackwell-vera-rubin.html https://investor.nvidia.com/news/press-release-details/2026/NVIDIA-Vera-Rubin-Opens-Agentic-AI-Frontier/default.aspx https://x.com/PatrickMoorhead/status/2033662536227393952 https://x.com/danielnewmanUV/status/2033649511592284352   The Groq 3 LPU: NVIDIA's $20B Bet on Inference Economics https://www.cnbc.com/2026/03/13/a-closer-look-at-nvidias-20-billion-bet-on-tech-for-a-new-ai-chip.html https://www.tomshardware.com/tech-industry/semiconductors/nvidias-20-billion-groq-deal-produces-its-first-chip https://www.servethehome.com/decoding-the-future-of-inference-at-nvidia-groq-lpus-join-vera-rubin-platform-for-low-latency-inference/ https://developer.nvidia.com/blog/inside-nvidia-groq-3-lpx-the-low-latency-inference-accelerator-for-the-nvidia-vera-rubin-platform/ https://www.jonpeddie.com/news/nvidias-groq-tie-in/ Microsoft Threatens to Sue OpenAI Over $50B Amazon AWS Frontier Deal https://www.reuters.com/technology/microsoft-weighs-legal-action-over-50-billion-amazon-openai-cloud-deal-ft-2026-03-18/ NVIDIA Restarting H200 Chip Production for China https://www.axios.com/2026/03/17/nvidia-huang-china-h200 https://x.com/danielnewmanUV/status/1999974968143257945 Meta & Nebius Sign $27B AI Infrastructure Agreement — Largest AI Compute Deal https://nebius.com/newsroom/nebius-signs-new-ai-infrastructure-agreement-with-meta https://x.com/danielnewmanUV/status/2033531056784347240 https://x.com/PatrickMoorhead/status/2033543939526193491 OpenAI Enterprise Pivot + Q4 2026 IPO Target https://www.reuters.com/business/openai-lays-groundwork-juggernaut-ipo-up-1-trillion-valuation-2025-10-29/ https://www.forbes.com/sites/josipamajic/2026/03/19/openais-pivot-to-enterprise-is-likely-a-race-against-anthropic-and-the-ipo-clock/ Supermicro's Legal Troubles https://fortune.com/2026/03/19/supermicro-arrested-founder-smuggling-gpu-china/ https://www.bbc.com/news/articles/cy41ly2d9wko The Flip: Did NVIDIA Just Kill the Inference Chip Startup Market with the Groq Acquisition? FOR: NVIDIA Killed It — The Inference Startup Market Is Over https://www.cnbc.com/2026/03/13/a-closer-look-at-nvidias-20-billion-bet-on-tech-for-a-new-ai-chip.html https://www.jonpeddie.com/news/nvidias-groq-tie-in/ AGAINST: Startups Survive — Hyperscalers Won't Deepen NVIDIA Dependency https://www.reuters.com/business/retail-consumer/cerebras-systems-amazon-strike-deal-offer-cerebras-ai-chips-amazons-cloud-2026-03-13/ https://www.tomshardware.com/pc-components/gpus/nvidia-removes-rubin-cpx-accelerators-from-its-roadmap-groq-3-lpus-take-center-stage-as-cpx-is-removed Bulls & Bears Market Reactions to Economic News https://uk.finance.yahoo.com/news/stock-market-today-dow-sinks-750-points-sp-500-nasdaq-slide-after-fed-decision-as-powell-touts-inflation-worries-200050703.html https://www.kiplinger.com/investing/live/march-fed-meeting-2026-live-updates-and-commentary https://www.investopedia.com/stock-market-today-dow-jones-s-and-p-500-03182026-11928689 $MU Micron Technology — Revenue Almost Triples, Tops Estimates https://www.cnbc.com/2026/03/18/micron-mu-q2-earnings-report-2026.html https://x.com/PatrickMoorhead/status/2034390648519024820 https://x.com/danielnewmanUV/status/2034378642613235921 $ACN Accenture — Q2 FY2026 Earnings Beat, Stock Drops ~5% on Guidance https://www.investing.com/news/earnings/accenture-falls-despite-q2-beat-as-earnings-guidance-disappoints-4570221 https://www.zacks.com/stock/news/2886706/accenture-earnings-beat-estimates-in-q2-revenues-increase-yy https://www.investing.com/news/transcripts/earnings-call-transcript-accenture-q2-2026-beats-forecasts-but-stock-dips-93CH-4570789 https://x.com/PatrickMoorhead/status/2033348794348142595  

The Cloudcast
Three Thoughts from NVIDIA GTC 2026

The Cloudcast

Play Episode Listen Later Mar 22, 2026 28:00


SUMMARY: We dig into the NVIDIA GTC keynote and highlight three things - accelerated computing for everything, the complexity of the new inference stack, and NVIDIA's “open” software stack including NemoClaw.SHOW: 1012SHOW TRANSCRIPT: The Reasoning Show #1012 TranscriptSHOW VIDEO: https://youtu.be/aXOr91q76yMSHOW SPONSORS:VENTION - Ready for expert developers who actually deliver?Visit ventionteams.comSHOW NOTES:NVIDIA GTC 2026 (Keynote)NVIDIA NemoClaw - OpenClaw + OpenShell + NVIDIA Agent ToolkitNVIDIA adds Groq LPU to their rack systemsNVIDIA to invest $26B in Open Weight ModelsInterview with Jensen about Accelerated Computing (Stratechery)Topic 1 - Jensen's trying to paint the bigger picture of accelerated computing everywhere (robotics, autonomous driving, gen-ai, physical ai - but also just everyday enterprise apps). Everything is about keeping the stock price up, and margins high. The stock price provides the warchest to fight off all foes. Topic 2 - The inference architecture is a complex mix of GPUs, CPUs, ASICs/LPUs, high-speed networking and seems very different from the training architecture. How big is the burden on data center providers? What are the inference alternatives emerging? Topic 3 - Jensen talked a lot about OpenClaw and eventually about NVIDIA's NemoClaw. How does his interest in Agentic AI tie into his interest in building NVIDIA's own frontier modelFEEDBACK?Email: show @ reasoning dot showBluesky: @reasoningshow.bsky.socialTwitter/X: @ReasoningShowInstagram: @reasoningshowTikTok: @reasoningshow

Sharp China with Bill Bishop
(Preview) The War in Iran and the Visit to Beijing; New DNI Assessments on Taiwan; Military Scientists Disappearing From Public View

Sharp China with Bill Bishop

Play Episode Listen Later Mar 20, 2026 13:19


On today's show Andrew and Bill begin with the news that President Trump has postponed his visit to Beijing amid the war in Iran, including why a delay made sense for both sides, a “Board of Trade” proposal amid signs of stability in Paris, and the uncertainty that pervades on both sides as the war in Iran continues. From there: Reactions to a DNI assessment on China's reunification intentions, news on U.S. weapons sales to Taiwan, the unknowns for China as Gulf unrest persists, and questions surrounding PLA readiness in 2026. At the end: Reactions to reports that several military scientists have had their profiles scrubbed from public websites, while Jensen Huang tells the world that Nvidia has received purchase orders for the H200 but Groq will not be shipping inference chips.

Sharp Tech with Ben Thompson
(Preview) OpenAI's Enterprise Pivot, The Rise of Agents and Bubble Counterpoints, Nvidia Changes Its Inference Story

Sharp Tech with Ben Thompson

Play Episode Listen Later Mar 19, 2026 32:50


Ben and Andrew begin with the news that OpenAI is shifting away from “side quests” and allocating resources to the enterprise space, including Dropbox history to explain OpenAI's present, lessons in the enterprise space generally (and what you learn in business school), and OpenAI taking cues from 1980s Microsoft. From there: Talking through Ben's article on Monday, including the implications of agents and questions about integration as durable differentiation for Anthropic and OpenAI. At the end: Nvidia's new messaging on inference chips and Groq integration, and a word about winters (and whiners) in Wisconsin.

The Information's 411
Nebius CRO on 2026 Strategy, Meta's Rogue AI Security Breach, Ross Gerber on SaaS & AI

The Information's 411

Play Episode Listen Later Mar 19, 2026 53:51


Nebius' Chief Revenue Officer Marc Boroditsky joins TITV to discuss its $4B debt raise after its Nvidia & Meta deals, Nvidia's new Groq chip and its 2026 strategy. Next, The Information's Jyoti Mann talks about a rogue Meta AI agent triggering a major security alert. We also talk with Offline Ventures' Brit Morin about AI agent workflows and security, Financial Analysis Columnist Columnist Anita Ramaswamy about why Canva was smart to delay its IPO, and we get into the $5 billion Qualtrics debt drama & AI with Ross Gerber of Gerber Kawasaki Wealth & Investment.Articles discussed on this episode: https://www.theinformation.com/articles/canva-smart-hold-ipohttps://www.theinformation.com/articles/inside-meta-rogue-ai-agent-triggers-security-alertSubscribe: YouTube: https://www.youtube.com/@theinformation The Information: https://www.theinformation.com/subscribe_hSign up for the AI Agenda newsletter: https://www.theinformation.com/features/ai-agendaTITV airs weekdays on YouTube, X and LinkedIn at 10AM PT / 1PM ET. Or check us out wherever you get your podcasts.Follow us:X: https://x.com/theinformationIG: https://www.instagram.com/theinformation/TikTok: https://www.tiktok.com/@titv.theinformationLinkedIn: https://www.linkedin.com/company/theinformation/

Chip Stock Investor Podcast
Can Nvidia Really Reach 1 Trillion In Revenue? NVDA and Groq Team Up

Chip Stock Investor Podcast

Play Episode Listen Later Mar 19, 2026 9:51


Nvidia's GTC 2026 is finishing up and there is a lot of focus on AI inference. Remember that licensing deal with Groq? By pairing Groq 3 LPX compute with Vera Rubin systems, Nvidia is targeting a 10x revenue jump—moving from $30 billion with Blackwell to a staggering $300 billion opportunity in Ai inference. Jensen Huang's "five-layer cake" strategy now dominates the entire data center stack, from power delivery to AI models, aiming for $1 trillion in total revenue through 2027.This new heterogeneous architecture blends GPUs for high throughput with LPUs for ultra-low latency, creating near-instant AI interactions. But is Nvidia right-priced right now?Join us on Discord with Semiconductor Insider, sign up on our website: www.chipstockinvestor.com/membershipCheck out these other Nvidia videos:https://youtu.be/50UfALpisPghttps://youtu.be/_6w9EbjaSIIhttps://youtu.be/_uvIkPwDu5Ahttps://youtu.be/p5w0aPzDi3ISupercharge your analysis with AI! Get 15% of your membership with our special link here: https://fiscal.ai/csi/Sign Up For Our Newsletter: https://mailchi.mp/b1228c12f284/sign-up-landing-page-short-formIf you found this video useful, please make sure to like and subscribe!⏳ Chapters00:00 – Nvidia GTC 2026: The Groq Licensing Deal 01:00 – The "Five-Layer Cake" of AI Data Centers 01:55 – From Chip Designer to Supply Chain Giant 02:50 – Road to $1 Trillion: 2025–2027 Revenue Outlook 04:20 – Groq 3 LPX & Vera Rubin: The New Rack Solution 05:45 – GPU vs. LPU: Solving the Latency Problem 06:30 – Heterogeneous Architecture: Throughput & Interactivity 07:45 – The 10x Revenue Jump (Blackwell to Rubin) 08:20 – Stock Valuation: Is Nvidia Still a Buy? *********************************************************Affiliate links that are sprinkled in throughout this video. If something catches your eye and you decide to buy it, we might earn a little coffee money. Thanks for helping us (Kasey) fuel our caffeine addiction!Content in this video is for general information or entertainment only and is not specific or individual investment advice. Forecasts and information presented may not develop as predicted and there is no guarantee any strategies presented will be successful. All investing involves risk, and you could lose some or all of your principal. #Nvidia #GTC2026 #AIInference #VeraRubin #Groq #JensenHuang #StockMarket #Semiconductors #TechNews #DataCenterNick and Kasey own shares of Nvidia

Pearls On, Gloves Off
#90 - Clients "Do Their Own Research." That's A Problem for Law Firms

Pearls On, Gloves Off

Play Episode Listen Later Mar 17, 2026 55:38


This episode of Pearls On, Gloves Off is powered by Workday. Learn more at workday.com. In this episode, Mary sits down with Claire Hart, Chief Operating Officer, Chief Legal Officer, and Board Member at Groq, to talk about what legal leaders should expect in the AI era - from their law firms, their teams, and themselves. With senior leadership roles at Google, Blizzard, and Genies, Claire brings a sharp perspective from the intersection of law, business, and technology. The conversation starts with the LinkedIn comment that got people talking: Claire said she would be horrified to learn that some of the law firms she works with are not using AI. From there, she and Mary unpack why adoption is still so uneven, how the billable hour distorts incentives, what young lawyers need to stay relevant, and why judgment, curiosity, and team design matter more than ever as legal moves into an AI-driven future. In this episode: Claire's AI hot take: Why clients should be alarmed if their outside counsel aren't using AI The adoption problem: How risk concerns and the billable hour are slowing real change Efficiency vs. incentives: Why the tech clients want conflicts with how firms make money What young lawyers need now: Judgment, communication, and adaptability over pure technical skill The blurring of roles: How lawyers, legal ops, and contract managers are starting to overlap What law firms still miss: Why understanding how businesses actually operate is now a competitive edge Join Mary's Substack Community Follow Mary on LinkedIn Rate and review on Apple Podcasts  

The Information's 411
Nvidia GTC 2026 Takeaways, OpenAI-AWS Pentagon Deal, Asana CEO on AI Agent Tool Launch

The Information's 411

Play Episode Listen Later Mar 17, 2026 42:30


Futurum Group's Nick Patience and Hydra Host's Aaron Ginn talk with TITV Host Akash Pasricha about Nvidia's $1 trillion revenue projection and the new Groq-based chip system. We also talk with Reporter Sri Muppidi about OpenAI's new AWS deal for government contracts and Editor Ken Brown about Mastercard's $1.8 billion acquisition of BVNK. Lastly, we get into Asana's AI agent strategy and the "SaaS apocalypse" with CEO Dan Rogers.Articles discussed on this episode: https://www.theinformation.com/articles/openai-clinches-aws-deal-bid-win-government-contractshttps://www.theinformation.com/newsletters/ai-agenda/nvidia-needed-groqhttps://www.theinformation.com/briefings/mastercard-buy-stablecoin-startup-bvnk-1-8-billionSubscribe: YouTube: https://www.youtube.com/@theinformation The Information: https://www.theinformation.com/subscribe_hSign up for the AI Agenda newsletter: https://www.theinformation.com/features/ai-agendaTITV airs weekdays on YouTube, X and LinkedIn at 10AM PT / 1PM ET. Or check us out wherever you get your podcasts.Follow us:X: https://x.com/theinformationIG: https://www.instagram.com/theinformation/TikTok: https://www.tiktok.com/@titv.theinformationLinkedIn: https://www.linkedin.com/company/theinformation/

The Twenty Minute VC: Venture Capital | Startup Funding | The Pitch
20VC: The 8 Moats of Enduring Software Companies: How to Analyse for Durability and Defensibility in a World of AI | Why Dropouts are "AI Maxing" the World & Remote Early-Stage Companies are Dying with Gokul Rajaram

The Twenty Minute VC: Venture Capital | Startup Funding | The Pitch

Play Episode Listen Later Mar 16, 2026 78:21


Gokul Rajaram is one of the greatest operators turned investors of the last 2 decades. He is trusted as the go to advisor for the greatest founders in the world. Today he serves as a Board Director at three public companies: Coinbase, Pinterest and The Trade Desk. Prior to Marathon (his firm), Gokul served on the executive team at DoorDash and Block. Before Block, he served as Product Director of Ads at Facebook. Earlier in his career, Gokul served as a Product Management Director for Google AdSense. Gokul is also a prolific angel investor, having invested in 700+ companies, including Airtable, Figma, Groq, Runway, Supabase, and Vercel.  AGENDA: 03:53 — Investing Lessons from Google, Doordash and Facebook 05:32 — Why Mark Zuckerberg is the Greatest Distribution Genius Alive 07:23 — Why Every Company Today Needs to be Multi-Product 09:16 — Negative Gross Margins: Are the Best Companies Actually Built on "Shit" Economics? 10:50 — The SaaS Apocalypse: Is the Entire Sector Going to Zero? 12:15 — The 8 Moats of Enduring Software Companies: How to Analyse Companies 14:50 — Why Brand is No Longer a Strong Moat (And What Replaced It) 16:13 — Salesforce vs. Atlassian: Which Systems of Record are Dying? 18:13 — Outcome-Based Pricing: Is This the Total Death of Seat Pricing? 20:16 — The Bolt-On AI Trap: Why Rebuilding Your Entire UX is Non-Negotiable 23:44 — Are the Outcome Sizes of Vertical SaaS Large Enough for VC Today? 28:16 — The Zombie Cohort: What Happens to Private Companies with High Valuations? 32:44 — Is "King Making" Complete Bullshit? 34:21 — Durability Over Margins: What Really Matters in a 100x Growth World 35:36 — The Non-Consumption Miracle: Why Granola and Gamma are Crushing It 38:50 — The PayPal Rule: Can You Raise Prices 5 Times in 3 Years? 42:47 — My Biggest Miss: How I Misread the Shopify Billion-Dollar Mark 45:18 — The Courage to Bet: Why Instacart is the Best VC Deal Ever 46:33 — Seed vs. Growth Pricing: When Does Price Actually Destroy Returns? 50:53 — Does "Proprietary Founder Access" Even Exist? 54:33 — Double Down or Diversify? The Truth About Fund Reserves 59:44 — The Vanta Anti-Portfolio: A Mistake I'll Never Forget 01:01:21 — When to Sell: The "Sell a Third, Hold a Third, Trade a Third" Rule 01:04:12 — Why Remote Early-Stage Companies are Dying 01:07:33 — Why Mid-Level Partners are Fleeing Mega Funds 01:09:47 — The Best CEO Superpowers: Larry, Mark, Jack, and Tony 01:12:33 — The Next 10 Years: Why Dropouts are "AI Maxing" the World    

nFactorial Podcast
Даулет Жангузин, NVIDIA, Groq, Cohere, Lyft, Google - Как пишут код лучшие кодеры Кремниевой Долины?

nFactorial Podcast

Play Episode Listen Later Mar 12, 2026 193:31


В этом выпуске мы вместе с Даулетом Жангузиным - инженером из Кремниевой долины с 15-летним опытом (NVIDIA, Groq, Cohere, Lyft, Google, Microsoft) - говорим о карьере в BigTech и о том, что происходит под капотом современного AI. Обсуждаем практичную сторону работы с большими моделями: как выжимать максимум из Nvidia GPU, чем полезен Claude в реальных задачах, и какие курсы/ресурсы действительно помогают расти инженеру, как пишут код в 2026 лучшие программисты Кремниевой Долине. Эпизод будет интересен тем, кто строит карьеру в разработке/ML, хочет понять трек BigTech (Microsoft → Google → Lyft), интересуется LLM-инфраструктурой и оптимизацией вычислений, а также ищет советы по обучению и прохождению технических собеседований в ведущие tech-компании.  Арман Сулейменов: https://www.instagram.com/armansu/ Даулет Жангузин: https://www.instagram.com/daulet/ Продюсер и режиссёр: Данияр Ахметжанов: https://www.instagram.com/good.years/ Наш Instagram: https://www.instagram.com/nfactorialpodcast/ Получите одну из самых востребованных профессий в мире - ИИ-разработчик - вместе с nFactorial School - https://www.nfactorial.school/courses_new/llm-engineer

The Information's 411
Anthropic Sues Pentagon, OpenAI IPO Investor Skeptics, New Groq Chip Reveal at Nvidia GTC

The Information's 411

Play Episode Listen Later Mar 9, 2026 35:14


AI Reporter Stephanie Palazzolo talks with TITV Host Akash Pasricha about Anthropic's lawsuit against the Pentagon over its supply chain risk designation and how OpenAI's new GPT 5.4 model is landing with developers. We also talk with Anita Ramaswamy about OpenAI's sky‑high IPO valuation, how it compares to Anthropic, Nvidia and Palantir, and why some public investors may sit out the offering. Then we speak with Anissa Gardizy about Oracle and OpenAI's Texas data center twist, Nvidia's $150 million move to take over the site, the upcoming Groq–Nvidia chip reveal at GTC, and Anthropic's aggressive bet on Google TPUs and Fluidstack.Articles discussed on this episode: https://www.theinformation.com/briefings/anthropic-sues-defense-department-designation-supply-chain-riskhttps://www.theinformation.com/newsletters/ai-agenda/ai-agenda-anthropic-strong-legal-case-trumps-dodhttps://www.theinformation.com/articles/openais-ipo-hopes-face-skeptical-investor-communityhttps://www.theinformation.com/newsletters/ai-infrastructure/real-reason-openai-walked-away-oracle-stargate-expansion-abileneSubscribe: YouTube: https://www.youtube.com/@theinformation The Information: https://www.theinformation.com/subscribe_hSign up for the AI Agenda newsletter: https://www.theinformation.com/features/ai-agendaTITV airs weekdays on YouTube, X and LinkedIn at 10AM PT / 1PM ET. Or check us out wherever you get your podcasts.Follow us:X: https://x.com/theinformationIG: https://www.instagram.com/theinformation/TikTok: https://www.tiktok.com/@titv.theinformationLinkedIn: https://www.linkedin.com/company/theinformation/

Tech Deciphered
74 – The Prediction Episode

Tech Deciphered

Play Episode Listen Later Mar 5, 2026 62:52


Who dares to make predictions in the current landscape? We do!  Our Predictions are back. Will our track-record continue on a high or will we be fundamentally wrong? Listen in to our Predictions for 2026 Navigation: Intro What will 2026 be all about? AI, AI and … more AI The big Hardware movements Of Start-ups and VCs Regulatory & Geopolitical Headwinds… and the Wars Fintech, Crypto and Frontier Tech Conclusion Our co-hosts: Bertrand Schmitt, Entrepreneur in Residence at Red River West, co-founder of App Annie / Data.ai, business angel, advisor to startups and VC funds, @bschmitt Nuno Goncalves Pedro, Investor, Managing Partner, Founder at Chamaeleon, @ngpedro Our show:   Tech DECIPHERED brings you the Entrepreneur and Investor views on Big Tech, VC and Start-up news, opinion pieces and research. We decipher their meaning, and add inside knowledge and context. Being nerds, we also discuss the latest gadgets and pop culture news Subscribe To Our Podcast Bertrand Schmitt Introduction Welcome to Tech Deciphered Episode 74. That would be an episode about some predictions about 2026. What will be 2026 all about? I guess this year is probably starting with a bang. We saw the acquisition of xAI by SpaceX. We saw an acquisition from Grok by NVIDIA. What’s your take about what would be the big themes in 2026? I guess it would be for sure about AI and space. Nuno Goncalves Pedro What will 2026 be all about? Yeah. I predict a year that will be a little bit more of a year of reckoning in some way. There will be a lot of things that I think we’ll start seeing through. The fact that we are in the midst of an amazing transformational era for technology, the use of AI, but at the same time, obviously, a ridiculous bubble that is going alongside it as we’ve discussed in previous episodes. I think that we’ll start seeing some early reckonings of that, companies that might start failing, floundering, maybe a couple of frauds along the way, etc. I’ll tell you what I will not make many predictions about today, which is geopolitics. Geopolitics, I will not make predictions at all. Who the hell knows what’s going to happen to the world this year in 2026? I don’t dare making any predictions on that. Back to things where I would make predictions. I think on AI, we’ll have a little bit of reckoning. We’ll talk about it a little bit more in detail during this episode. Interesting elements around the hardware and physical space. Physical space, we just dedicated a full episode to it. We won’t go into a lot of details on that, but definitely on the hardware side, we’ll talk a little bit more about it. The VC landscape is going through an incredible transformation. We’ll talk about it today as well and some of our predictions for this year. What will happen to the asset class? It seems to be transforming itself dramatically. Obviously, that has a very direct impact on startups, so we’ll talk about that as well. And then to close a little bit the chapter on this, we will address some regulatory and geopolitical, let’s call it, headwinds without making maybe too many complex predictions. We shall see. Maybe by that time of the episode, we will be making some predictions. You guys should stay and listen to us, and maybe we will actually make some predictions about the geopolitical transformations that we will see this year in the world. Then last but not the least, we’ll talk about fintech, crypto, frontier tech, and a couple of other areas before concluding the episode. A classic predictions’ episode. We normally have a pretty good track record on some of these, but right now, the world is going a bit interesting, not to say insane. Bertrand Schmitt Yes, and going back to some news, Groq technically was not acquired, but, practically, it’s as if it got acquired. I’m talking about Groq, G-R-O-Q. The AI semiconductor company focused on inference AI, and it was late December. It was a way to end the year. This year, we started again with an acquisition of xAI by its sister company, SpaceX. I guess that’s where we are starting. AI, AI and … more AI We are going to start on AI. That’s definitely the big stuff. Everything these days, I guess, is about AI or has to have some connection with AI, or it doesn’t matter. I think every company in the world has seen that. You have to have the absolute minimum on AI strategy. You better execute on this strategy and show results, I would say. For the companies that were not AI native, you truly have to have a way to transform yourself. I guess at some point, the stretch might be too much, and it’s not really reasonable. Then you maybe better stay on what you are doing, especially if you’re in tech, you better be moving faster to AI. Nuno Goncalves Pedro Just to highlight, and I think throughout the episode, you’ll see that there’re obviously a lot of implications that would manifest themselves into capital markets. I mean, we’ll specifically talk about VCs and startups later on. But the fact that everything needs to be AI, the fact that there’s so much innovation happening right now, in my opinion, and this is maybe the first pre-topic to AI, is we’ll see a tremendous increase in M&A activity this year across the board. I mean, we’ve seen already some big acquihires we mentioned in some of our previous episodes, but we’ll see a lot more activity on M&A this year. Normally, that’s a precursor to the opening of capital markets. I predict also that there will be a reopening of the IPO market that never really reopened last year, to be honest. M&A, a lot more, reopening of the IPO market. Normally, it happens in the second or third quarter of the year. That’s what my M&A friends tell me. First quarter of year, everyone’s figuring out stuff. Then last quarter of the year, things should be more or less closed. Maybe the third quarter is the big quarter. We shall see. But definitely, as a precursor to our conversation today, I think we’ll see a lot of M&A, and we’ll see reopening of the IPO mark. Bertrand Schmitt I guess last year was not as big as you could expect on M&A given the tariff situation announced in April and May. I mean, it became quite tough to do IPO in such market conditions. Definitely, we can hope for something dramatically different in 2026. I guess talking about public markets and IPO, I guess the big one everyone is waiting for is SpaceX. SpaceX getting even more interesting with its xAI acquisition. Nuno Goncalves Pedro Do you think that because of the acquisition, it’s more likely that it will happen this year, or because of the acquisition, it’s less likely that it will happen this year? Bertrand Schmitt That’s a good question. My guess is the acquisition of xAI is all about xAI needing more financing and cheaper financing. This acquisition is a pathway to that. SpaceX being a much bigger company, a company that is also making much more revenues. I could bet that there is higher probability that, actually, SpaceX will go public in order to finance itself. At the same time, will it have enough time to prepare itself for the IPO given this acquisition just happened? Can they do that in 6 months? I mean, if anyone can do it, I guess it’s Elon Musk. It’s a strategy to present an even more attractive company with an even more interesting story, a story of vertical integration from AI to space. I guess the story as it’s presented itself right now, it’s one about having your AI data centers in space. Because in space, you have much better solar energy production with solar panels. You have a perfect cooling situation because you are in space. Thanks to Starlink, you have the mean to communicate between the satellites and with Earth itself. I think if someone can pull up a story like AI data center in space, I guess Elon Musk can. There is, of course, a lot of questions about is it practical? Is it economical? Yes. I certainly agree. I’m not clear on the mass, and can you make it work? Again, I mean, Elon Musk single-handedly, with SpaceX, managed to transform the space market on its head. I mean, they are the biggest satellite launching company in the world. They have the most satellites in the world. I mean, I’m not sure I would bet against him, and I guess I would probably believe that he could pull up something. Time frames, different story. The 2-3 years data center in space for AI as cheap as on Earth, I have more trouble with that one. I mean, it’s a usual suspect with Elon Musk. You promise something unachievable in a few years, but, ultimately, you still manage to reach it in 5 or 10. Again, I would not bet against the strategy. Nuno Goncalves Pedro Yeah. I’ve talked to a couple of space experts, people that have launched rockets, and have worked JPL, NASA, and a couple of other places, etc. For what it’s worth, their feedback is, “No way in hell, and we’re decades away.” We’ll see. I mean, to your point, Elon has pulled very dramatic stuff. Not as fast as he normally says he’s going to pull it, but within a time span that we all see it. Difficult to bet against him. In terms of actually the prediction, maybe to respond to the prediction as well, will SpaceX IPO? I’m going to make a prediction that has a very high likelihood of missing the mark, but I think Tesla’s going to buy and merge them both into it. It’s going to become a public company through Tesla. That’s my hypothesis. Bertrand Schmitt No. That’s supposed to be it. That’s how you solve that. Nuno Goncalves Pedro And Elon controls the whole universe. X, xAI, Tesla, SpaceX, all under one umbrella beautifully run. And SolarCity is well in there, of course, so wonderful. Bertrand Schmitt That’s possible. Certainly, you are not the only one thinking Tesla will acquire or merge with SpaceX. To remind everyone, Tesla is around 1.3, 1.5 trillion market cap. Depending on the day, SpaceX seems to be valued at similar range, 1.2, 1.3 trillion. It looks like it’s the most valued private company at this stage. These are companies of similar size, so that’s one piece of the puzzle. When you think about the combined company, we could be talking about a 3 trillion entity. Playing right here with the biggest companies in the marketplace today. Nuno Goncalves Pedro With a couple of tweets from Elon, it will rapidly get to 4 to 5 trillion. Bertrand Schmitt That’s so tricky. Nuno Goncalves Pedro Yes. On AI and back to AI, one thing I think that we’re about to see is this will probably be the year of agentic AI. Obviously, we predict a lot of growth on that side of the fence, in particular on the enterprise B2B side. We see a lot of opportunities coming through. From our perspective, at least at Chamaeleon, we generally believe that there’s going to be a lot of movements on agentic AI. It’s also going to be probably the year of the first big fails of agentic AI that will be newsworthy. There will be some elements about that loop and how it gets closed that will happen. I think we might see some scandals already. We’re already seeing the social network of bots talking to bots. We will see other scandals going on this year even in the consumer space and in the bot to bot space, which we now can talk about or in the AI agent to AI agent space. My prediction is we will see some move forwards. There’ll be some dramatic funding rounds along the way. We’ll see a couple of really cool things out of the gates coming out that are really impressive, but we’ll also see the first big misses of the technology stack. I don’t think we’ll go fully mainstream yet this year, so it’s probably maybe something more for 2027 along the way. That would be my prediction again. I think enterprise will lead the way. We’ll definitely see a lot of stuff on consumer as well that is cool. Then we’ll all have our own personal assistance in our hands, basically, literally in our phones. Bertrand Schmitt Going back to agentic AI, we also started the year with some pretty dramatic move. I mean, the launch of Clawdbot, renamed OpenClaw. I mean, this stuff took fire in like a week or 2. It was coded by just one person who actually didn’t even code the product but used AI to build the product, 100% used AI, proposing some new ways also to leverage AI to do coding. He has a pretty unique approach. It’s not vibe coding. I would say it’s a better way to do that. Then the surprising evolution with the launch of a social network for AI agents, Moltbook. I mean, this stuff, probably there is some fake in it. But at the same time, I think it’s quite impressive because it’s the first time we see truly 100,000 plus agents communicating directly to each other. Yeah. I mean, that’s the first time we see surfacing the possibility of some sort of hive mind on the Internet. It’s pretty surprising. Right now, all of this is a hack done in a few days. By end of year, by 2 years, 3 years, we might discover that, actually, the best approach to AI might not be the AI assistant like we are doing today, but a combination of hundreds of thousands of AI working closely together. We might be witnessing the first sign of new intelligence in a way. Nuno Goncalves Pedro Things like this social network might either be Skynet, the beginning of Skynet. They might be the beginning of Her, or they might just be a fad and nothing really happens. It’s just interesting to see what these agents are doing. Bertrand Schmitt Totally. Nuno Goncalves Pedro Obviously, there are real and clear and present dangers of some of the integrations of AI we’re seeing in the market. Interesting enough, and I’ll ask you for your prediction a bit, Bertrand. I think we’ll probably see the first big mishap of AI being used in some infrastructural decision in the age of AI. I mean, we’ve seen AI issues in the past and software issues in the past. We talked in previous episodes about that as well. Mishaps of software that have led to people dying. But I think probably the first big mishap will happen this year as well. Very public mishap of the use of AI and serve its interactions with infrastructure or something that’s very platform related, etc, that will have big impact that everyone will notice. That’s my prediction for the year as well. We’ll have the first big oops moment, as I would call it, for AI in this new age of full on AI. Bertrand Schmitt I would say first some perspective. I think today, people are not using AI directly for life and death decision, at least not that I’m aware. We’re not going to let AI fly a plane, for instance, tomorrow so you can be, reassured. At the same time, given there is such a race to AI, there definitely might be some mistakes. We were talking about the social network for AI agents, Moltbook. Apparently, all the keys used to secure the AI were shared by mistake because it was not properly locked down. We can see that indirectly, mistakes will be made for sure. Two, it’s highly probable that some people will trust AI too much to do some stuff, and this stuff might not work and might have some grave consequence. Hopefully, there is not so much of this. Hopefully, it’s mostly AI used for the good. But you’re right. I mean, at some point, the more we use the technology, the more there would be issue. I mean, it’s highly probable. Nuno Goncalves Pedro That will lead me to another prediction, which is, and we’ll talk about more of it later, but it probably will lead to the first significant movement in terms of regulatory environment certainly in the US at some point if it happens in the US in particular, where there will be some movement that will be like, “Hey, you guys can’t do this anymore.” Because this will probably emerge from mismanaged interfaces. From systems having access to stuff that they shouldn’t have access to in the first place. Talking a little bit more about what’s happening in AI. You’ve already mentioned some of the issues that relate actually to security and cybersecurity. We keep talking about AI. We keep talking about all these infrastructure pieces and platforms that are being built. I think we’ll have a lot more incidents like the one you just mentioned where things will be shared that shouldn’t have been shared, where people will break systems and get into it, etc. Let’s see where that takes us, which is a little bit ironic because, obviously, with AI, the promise is that cybersecurity becomes more robust as well because there’re agents working on our behalf on the cybersecurity side. There’s also agents working on the other side. Bertrand Schmitt It’s a constant race. It’s the attackers, defenders. Each time you have new technology, you have a new race to who is going to attack or defend the best. Each new wave of technology, it’s an opportunity to challenge the status quo. Nuno Goncalves Pedro The attackers have been winning, and I feel they’ll continue winning in 2026. I think it’s going to still be a year of attack. We’ll see more and more breaches, more and more stuff that will happen. Bertrand Schmitt I don’t know if they will win. I mean, it’s normal that they win once in a while. For sure, some infrastructure is not updated as it should. Some stuff are not managed as it should, so there will always be breaches. I don’t know if things are dramatically going to change because, again, everyone who cares who is going to update his infrastructure with AI for defense. There is no question that you have no choice. We will see. That I don’t know. For sure, AI will be used to attack directly with AI. Maybe you’re able to do bigger, larger scale attack. Or thanks to AI, you are simply able to create new type of attacks more easily. AI can be used behind the scene as a way to prepare and organise new type of attacks, even if it’s not used directly live in the battle. Nuno Goncalves Pedro One topic that we’ll come back to later is the geopolitics of everything, but maybe more broadly. On the geopolitics of AI, it’s very clear that we have an arms race going on. Obviously, the US on the one hand, China on the other hand is the two extremes, putting tremendous amount of capital into data centers just at the base of that infrastructure. Chipset development, chipset access, a huge theme in terms of the export restrictions, etc, that are being forced by the US. I think it will continue. From a European standpoint, obviously, they’re stuck between a rock and a hard place, to be very honest. Let’s see what happens on that side of the fence. My view of the world is that certainly from a US and China perspective, we’re going to see a lot more movements in 2026, like big movements. The Chinese movements we always see in delay.  It takes us a couple of months, sometimes even more than that to understand exactly what’s going on. I think we’re going to see some huge moves this year in terms of the States, the United States of America, and China really pouring capital into the creation of the next big winners around AI. I think the US is obviously more visible. We see a lot of these companies. We’ve just discussed xAI and its acquisition by SpaceX or merger. I don’t know what they’re calling it exactly. Effectively, on the China side, the movements I think are already very big. As I said, it will take a while to figure out exactly what those moves are. One thing that I propose is that at some point, China will have very little dependency on chipsets from the US. I’m not sure it’s going to happen this year, but I think the writing is on the wall. Irrespective of any other geopolitical issues that is coming to the fore at this moment in time. That’s one of the key areas or in arenas of fight. Bertrand Schmitt It makes sense. If you are China, you will look at what happened. You would think that you cannot just depend on the largest of one country. It makes rational sense, the same way it makes rational sense for the US to limit exports to China because there is value to delay some peer pressure that could use these technologies for good but also for bad. If you were an ally of the US, that would be one thing. But when you are not an ally of the US, that certainly should be a different perspective. Maybe one last point concerning agents, I think there will be a lot that will revolve around coding. We can see OpenAI with Codex. We can see Cloud with code. There was, of course, [inaudible 00:18:28] that was trying to be big on agentic coding. I think agentic coding was one of the big transformation in 2025 and is going to get bigger in 2026. I think for a lot of people who do coding, there was a radical transformation in terms of what you can achieve, what you can do, how much you can trust AI to help you code. I start to think we might see this year, the replacement of not just one AI replace one coder, but one AI replace a full team because of the new ability to manage that at scale. Coding might be a common activity where you are going to think about outcomes, think about objective, think about how you organise, but not really coding by itself anymore. A big change, like you used to code, directly your hand on the stuff, but step by step, everyone is going to become a manager of agent. I think in one year, we saw enough transformation to think that in the coming year, the transformation can be even more dramatic. Nuno Goncalves Pedro The big Hardware movements Now switching gears to hardware. Obviously, a lot of movements in 2025 and over the last few years. One piece of thesis that we’ve had long-standing at Chamaeleon is that we will see the emergence of AI devices. Some of them have been tremendous failures as we discussed in the past. I predict that we’ll have a couple of really interesting full stack AI devices in the market this year. Why does that matter? Because, as many of you know, obviously, there’s compute that can happen in data centers and cloud infrastructure all over the world, but also there’s compute that can happen at the edges. The more you can move to the edges and the more you can create devices that actually allow you to have user experiences that are very distinctive at the edge, the more powerful some of these devices might become. I predict Apple will not be the first to launch anything on this. I predict probably OpenAI, after the acquisition of IO, will maybe not launch something this year, but will announce something this year. I’ll step back on that prediction. They’ll announce something this year, but maybe not launch. But we’ll start seeing some devices that have some interesting value in the market, probably devices that are AI devices, but they are very focused on very specific user flows, and so very much adequate to specific activities. I won’t make a prediction on that, but I think areas that would make sense for that to happen would be obviously around fitness, health, et cetera, et cetera, where we already have the ascendancy of products like Oura Ring and others out there. Definitely, that’s one area that might have quite a lot of developments. I think AI-first devices, devices that are very focused on compute at the edges, providing user flows that are AI-enabled to end users, we’ll see a lot more of that and a lot more activity this year. Again, I don’t think Apple will be necessarily ahead of the game. Again, maybe OpenAI will give us something to at least think about and look forward to. Bertrand Schmitt First, I’m not sure it will be that transformational because if it’s not in your phone, in your pocket, there is only so much you can do with it, and there is only so much computing power you will have. I’m doubtful it would be really impactful this year. Nuno Goncalves Pedro I feel we’ve been discussing this shift of paradigm in input and output. For me, some of these devices could lead to that shift. Because, again, a mobile phone is not a great long-term paradigm for the usage that we have because it’s really constrained by the screen. The screen is really what takes most of the battery life away. If we didn’t have that screen, what could we do? If we have the block that is as big as a mobile phone, and it didn’t have a screen, it was just compute, that’s a mini computer, a microcomputer. Bertrand Schmitt That’s a fair point, but I don’t see that transformation this year. That’s really more my point. I can see that you can have AI-enabled smart glasses, and it’s clear there is a race to AI-enabled smart glasses. My point is more to go beyond the gadget, it would take quite a while. It would need to have cameras. It would need to analyse what you see. It would need to hear what you hear. Again, it might come, but then at some point, it would be okay, what do you do with it? We have the example of the movie Her. That’s showing Her what it could be. There are definitely possibilities. It’s clear that if you take the big VR headset like the Apple Vision Pro, there is a failure from that perspective in the sense that I think it’s a great, amazing device. The big problem is that it’s doing way more that makes sense. I think there will be a clearer separation between your smart AR glasses that has to be light, that has to be always unconnected, and that’s primarily there to help you make sense of the world around you. The true VR headset that doesn’t really require much in terms of AI, and it’s just there to immerse you in a different world. For this, we know, unfortunately, in some ways, that there is not a lot of demand for it. Maybe there is little demand because you are too hidden in your own world. The technology is not working well enough yet. There are a lot of reasons. But I think Apple trying to do both at the same time, AR and VR, with the Vision Pro, was a pretty grave structural mistake. I think we would see a clearer line of separation between the two. There is bigger market opportunity for AR glasses. That, I certainly agree. There is opportunity to connect that to a computing device. As you talk about, your glasses are your screen, your phone becomes something in your pocket connected to your glasses. Nuno Goncalves Pedro For me, Apple has their way of doing things. From the perspective of what you said, they normally really plan their devices. Even if it’s a big shift in terms of a new area, like they tried with the Vision Pro, and we criticised them for launching it as a device that should have been more of a dev device that they really launched as a full-on device, but that’s their playbook, classically. I think Apple needs to change how they put products out and how they experiment with those products, et cetera. I think they have enough money to be doing everything all the time and figuring it out. If they don’t want to put it out, then they need to do a lot more hell of testing internally with their silos, but they should be playing across all these arenas, VR, AR, everything. They just should put devices out that are either ready for prime time, or they should call it something else. They should call it like this is a dev device or whatever it is. Bertrand Schmitt I agree with you. My complaint is more that it was marketed as a consumer device when it was not. It was a true developer device. Two, they tried to mix the two at once, and it made no sense. No one is going to walk in their home or in the street with their Vision Pro on their head. You have to be deranged, quite frankly, to have use cases like this. I think that for me is a crazy mistake from a company like Apple that prides itself in pure UI, pure user interface, very well-designed device for one specific use case, not mixing the two use cases. We still don’t have Macs with a touchscreen, you know?  We still don’t have an iPad with a good OS that makes use of this great hardware. For some strange reason, they decided to mix everything in the Vision Pro with a device that weighs a ton on your head and is so uncomfortable. That’s why, for me, I’m like, “Guys, what is wrong? Why did you let this team run crazy?” I hope at some point, Apple will go back to the drawing board. My understanding is that that’s what they are doing. They are going to have two devices, one smart glasses, an evolution of the Vision Pro, just focus on VR. They might actually abandon the concept of the pure VR-oriented headset. Because, from a market size perspective, it might not be big enough for Apple, quite frankly. Nuno Goncalves Pedro I read on all of the above, and people at this point was like, “Why are then players like Samsung and others not doing it. LG, et cetera?” Because those players historically have not invented new categories. They’re amazing at catching up once the category is invented, and then they scale the hell out of it, and that’s what these companies have been exceptional at. I wouldn’t see a dramatic innovation, I think, in terms of devices coming from any of the big ones on that side of the fence. Not to disrespect them in any way, but I think that’s not been their playbook ever. Again, if the origination doesn’t come from a start-up or from an Apple, I don’t see those guys going after it. My bet is that we’ll see some start-up activity and, again, hopefully, some announcement from IO now within the OpenAI world. Bertrand Schmitt I would slightly disagree with you. I see where you are coming from. But take the Samsung Galaxy Note, that sudden much bigger headphone that no one was doing that was launched by Samsung, at some point, it forced Apple to launch an iPhone Max. Let’s look at the Z Fold that Samsung launched 7 years ago, copied by everyone. Now Samsung launching a trifold. Apple has still not launched their foldable phone. I think there is a mix, actually, of sometimes- Nuno Goncalves Pedro For me, that’s not a proper new category. It’s still a mobile phone. It just happens to have a screen that folds in half. Bertrand Schmitt The iPhone was still a mobile phone, you could argue.  Nuno Goncalves Pedro No. I think the iPhone was…  I could actually agree with you on that point. Maybe Apple is not as innovative in that case. I think what Steve Jobs was exceptionally good at in terms of his ability as this master product manager was to be an exceptional curator of user flows and user experiences, and creating incredible experiences from devices based on that. That was his secret sauce. Could you say, “Wasn’t all of this stuff already around?” It was. You just put it all together very neatly and very nicely. But if you’re talking about significant shifts in how a category is done, the iPhone was a significant shift in how the category was done. The Fold is still an interesting device. I actually have a Fold right now in front of me. The 7 that you highly recommended to me that we both got, the Z Fold 7. I think they do amazing devices. I don’t think they normally are the most innovative players. Then, when they come to innovation, it comes from technology edges. Obviously, they have Samsung Display, there’s a bunch of other things. They had the ability to do foldable screens in-house themselves. Bertrand Schmitt I don’t disagree with you. I think there is an interesting situation where some companies have some strengths, another one has some strengths. My worry with Apple is that this was not demonstrated with the Vision Pro. The Vision Pro was a hot pot of technologies barely integrated together, with use cases absolutely not well-defined and certainly not something that makes sense for most of us. There is a question of has Apple lost it? While Samsung actually keeps doing their own stuff, that, yes, might be more minor improvements, but at least they are doing it. Because it looks like Apple is missing the train on even the minor improvements. By the way, you might not be aware, but Samsung launched its Vision Pro competitor. Interestingly enough, it might be a better product in some ways, being much lighter and much more comfortable. Nuno Goncalves Pedro We should play around with that and report back to our listeners. Of Start-ups and VCs Moving to venture capital and the startup ecosystem and what’s happening there, I think it is very much a bifurcated environment, and it’s bifurcated for both VCs and for startups. If you’re a startup in the AI space, and you have the hottest team since sliced bread, and you can create FOMO at the speed of light, you can raise ridiculous rounds. Five hundred million at the $3 billion, or $4 billion, or $5 billion valuation, and you still haven’t really even started. First round, you can raise 500 million. That’s back to the whole discussion on Bubble and where are we, et cetera. Some of these companies might actually become huge, some of them might not. But definitely, we are seeing really the haves and have-nots on the startup ecosystem with incredible teams raising a lot of money very, very early on or mid-stage if they’ve already existed for a while, and then the rest not being able to raise. We see a lot of non-necessarily AI sectors, some of the areas of SaaS that don’t necessarily have AI in it, or fintech, or the consumer space that are really, really struggling. If you don’t have an AI story for your startup right now, it’s extremely difficult to raise money unless your numbers are just the best numbers ever. That’s, I think, the first part of the element of bifurcation that we’re seeing today. The second element of bifurcation that we’re seeing today in terms of fundraising is for VCs themselves, and really propelled by the large VC firms raising more and more capital in recent orbits, announcing 15 billion across funds raised. Lightspeed, I think, had made an announcement a couple of weeks ago as well. They’ve raised a bunch of money as well. The big guys are all raising a lot of money. At some point in time, the question some of you might ask is, “These VCs are redeploying more and more money if they have a couple of billion for a VC fund. How does that look like? Is that still VC?” My perspective, I’ve shared before in some of our previous episodes, is that that’s no longer venture capital. At that point in time, we’re talking about something else. Private equity hedge funds, if you want to call them, maybe funds that are really driven by growth investment or late-stage investment. If you have a couple of billion under management, you’re not going to make your returns by writing a $3 million check in a series seed and leading that round.  That has implications for everyone in the ecosystem. It has implications for smaller funds that obviously have a lot more difficulty in raising capital. It’s difficult to differentiate. Last but not least, also for startups that really continue searching for that capital that is out there. Andreessen Horowitz, for example, runs Speedrun, which is a great program for companies around consumer in particular. Initially, it was a lot for gaming. But at some point in time, Andreessen Horowitz could decide that they don’t want to invest more in you. They just put money from Speedrun, which is obviously a very small check compared to the very large checks they could write mid to late stage and that will have an effect on you as a startup. What happens at that point in time if Andreessen Horowitz is not backing you up in later stages? More than that, what happens if I can’t get these big funds interested in me? Are the small funds still valuable to me? Punchline, my view is yes. Obviously, we’re a smaller fund, so there’s parochial interest in what I’m saying. Small funds can still create a ton of value for you, also in terms of credibility, ability to accompany you in those first stages of investment, and the ability to bring other larger investors later down the road as well. There’s definitely a big movement happening in terms of the fundraising for VC funds, which we shouldn’t neglect, which is the big guys are raising a lot more capital and are therefore emptying the market to smaller funds that are having more and more difficult raising at this point in time. We had discussed that there would be a need for concentration in the industry, that micro funds would need to concentrate, and we didn’t have the space for so many micro funds as we had around. But the way it’s happening is extremely dramatic at this moment in time. I think it will continue through 2026. Bertrand Schmitt Remember a few years ago, with the rise of AI, there was more and more of the question about, “What’s the point of SaaS at this stage?” Because SaaS was around for 15 years. Basically, how do you come up with something new that was not already tested, validated by the market? How do you bring something new? We say this was reinforced to the power of 10. If your product is not clearly built from the ground up for a new use case enabled by AI, anyone could then might have built your product 5, 10 years ago, and therefore, why now has no clear answer, and it’s a big problem. I’m still surprised myself to still see some entrepreneurs where you talk to them about AI because you don’t see them in the deck, and they explain to you, “It’s not yet there,” and you’re like, “What’s wrong with you guys?” Fine. Do whatever you want. Do a small business and whatever, but don’t think you can come up pitch and raise without an AI story. The second category is people who come with an AI story, but you can feel very quickly, I guess you saw that many times, Nuno, where just a story layered on top with little credibility. It’s not better. It’s not enough to just have a story. Your business needs to be radically built differently or radically proposing some brand-new use cases that were impossible to solve 5 years ago. Nuno Goncalves Pedro To stack up on that, absolutely in agreement. If you’re just adding to the story, and it’s an afterthought, and you’re just trying to make the story somehow gel, once you go into one or two layers of due diligence, your investors will very quickly realise that you’re not really AI-first or dramatically AI-enabled or whatever. It’s just you’re sort of stacking something on top of another thesis. It needs to make sense from the product onwards. It’s not just, let’s just put it together with chewing gum, and magically, people will give you money. It was true also if we remember the good old crypto blockchain days, where everyone’s investing in crypto. A lot of stories that didn’t make much sense. In that sense, it’s not very different. I would go one step further. I think in the world of the VC winter that we’re a little bit in, where it’s more and more difficult if you’re a smaller fund to raise your fund at this moment in time, there’s a lot of sources of distinctiveness still talked about, like proprietary networks, access to deal flow, fast track record, all that stuff that really, really matters. But our bet continues at Chamaeleon continues being that you need to be AI-first as a VC fund yourself. You need to have core advantages in using not only readily-available AI tools or third-party available AI tools, data sources, technology stacks, but actually building your own stack over time, which is what we did with Mantis at Chamaeleon. Again, just to reinforce that, I think we’re at the beginning of that stage. We, Chamaeleon, are ahead of the game, but we think that the rest of the market will have to move towards that as well. Still, to be honest, very surprising to me to see that many significant large players are doing very little still around some of these spaces. They have data scientists. They’re running some tools. They’re running some analysis and all that stuff, but it’s still, again, back to the point I was making for startups, all glued up with chewing gum. It doesn’t all come together nicely, which it does need to from a platform standpoint. Bertrand Schmitt It’s quite surprising. I agree with you that some VC funds might think that they can do business as usual in that brand-new world. It’s difficult to believe. Nuno Goncalves Pedro Maybe moving a little bit toward the capital formation piece. We already discussed the M&A space really accelerating. We’ve also discussed the IPO market and some predictions on that. Secondaries, there’s obviously a lot of liquidity coming from secondaries from mid to late stage. I think it will continue throughout the rest of 2026. A lot of activity in buying, selling in secondaries as some asset managers are becoming more distressed, as some very high net worth individuals and family offices are becoming more distressed as well, at the same time, where there’s a lot of opportunities to potentially arbitrage around some investments. I believe a lot of money will be made and lost this year by decisions made this year, just to be very, very clear in terms of equity, purchases, et cetera. Exciting year ahead of us. Definitely a very, very interesting market ahead of us. Secondaries, M&A, growth, and late-stage investing, also, early-stage investing will continue just for those that were wondering. Last but not least, the public markets, the IPO market as well. Bertrand Schmitt One of the big questions for the IPO market would be, will SpaceX go public? Would it be good for the startup ecosystem? Because suddenly that they go public, it would be to raise money. If they raise money, will there be any money left for anybody else? That would be an interesting test of the market. For sure, it would be proof that market are risk on financing a new IPO like this one. Or as you said, maybe there is no IPO, and it’s a merger with Tesla. Time will tell. Nuno Goncalves Pedro Regulatory & Geopolitical Headwinds… and the Wars Moving maybe to our topic of regulation and geopolitical headwinds, as we’re seeing … definitely not tailwinds. The Google antitrust verdict and, obviously, the remedies are expected to come forward now, and a lot of people are saying, “There are some risks of structural separation.” What do you think? Is it cool, but nothing will happen in the end dramatically? Alphabet or Google? I’m not sure, actually. It’s Google LLC. I think that’s the case. It’s The United States versus Google LLC. Bertrand Schmitt I’m not sure. Personally, I’m not a big fan. I think there needs to be a better way to manage some anticompetitive behavior. I’m not a big fan. There was this temptation to do that for Microsoft 25 years ago. Look at what happened. No one needed to buy Microsoft to leave space for others. I see the same with Google, and I guess they are happy to not be the number 1 in AI today, but to have an open AI in front of them. Even if they are doing a great job, by the way, to move forward and go faster and faster. Personally, quite impressed now with some of what they have released. Gemini 3 is doing great from my perspective. I’m not a big fan of this. I think to be clear, it’s important that bigger companies don’t behave anticompetitively, but at the same time, we need to find the right approach where it’s not about breaking these companies, and it’s also not about forbidding them to do acquisitions. Because then you end up with what NVIDIA just did with a $20 billion acquihire IP licensing type of acquisition, because they didn’t want to have the uncertainties. They didn’t want to wait 1–2 years in order to acquire the people and the technology, so they organised it in a different way. But I don’t like that. I think they should be able to acquire companies without facing so much uncertainty. To be clear, it’s not new. Uncertainty when you are Google, NVIDIA, or others, it happens. It has happened for a decade plus, 2 decades. I think there needs to be, for sure, some safety valves. At the same time, we want an efficient capital market. An efficient capital market need companies that can acquire other companies. If you don’t do that efficiently, it will be worse for the entrepreneurs, it will be worse for the investors, it will be worse for everybody. I think we have not reached a good equilibrium from my perspective. We need more efficient acquisition process. And at the same time, we need to also enforce faster anticompetitive behavior. Because what you talk about concerning Google, this is a case that was what? That is 10 years old. You see what I mean? This is way too long. If you’re a startup, you are dead by then. It’s like the story of Netscape facing Microsoft. They were dead long after the fact. I think we need a different approach. I’m not sure the best answer. I’m not sure we’ll get a better approach. There are probably too many vested interest. My hope is that it will get better with this current administration because, certainly, the past administration was very anti acquisition and efficient markets. Nuno Goncalves Pedro We’ve talked about the European Union AI Act a bunch of times, so I don’t want to spend too many cycles on that. The only effect that I would say is we are seeing in very slow motion the splitting of the Internet. I once had Tim Berners-Lee, by the way, shouting at me that we were going to break the Internet when we were applying for the .mobi top-level domain. I was part of that consortium that eventually did get the .mobi top-level domain, and I had him shouting at us. But, apparently, this is going to split the Internet, Tim. So in case you’re listening. Because it will create all these different rules. If your data is relating to consumers there, then it’s treated in a different way, and The US is… Well, obviously, we have the case of California with its own rules and laws. I don’t know. I feel we’re having a moment of siloing that goes beyond economic and geopolitical siloing. It will also apply to the digital world, and we’ll start having different landscapes around it. We’ll see how this affects global expansion of services, for example, around AI, particularly for consumer, but I don’t foresee anything dramatically positive. Recently, we had the whole deal around TikTok finally having a solution for their US problem where there’s now a US conglomerate magically that owns it. The conglomerate doesn’t magically own it, they just straight up own it for the US. But it was driven by many of these concerns around data ownership. Where’s the data? Where is it based? I think a lot of other concerns that have to do with the geopolitics of China, obviously, being the basis of ByteDance, the owner of TikTok, that still is a significant owner, by the way, in TikTok in US. Then also the interest in the economics of making money out of something as powerful as TikTok, to be honest, in The US. Just to be clear, I don’t think this was all about the best interests of consumers. It was also about money. Just follow the money. Bertrand Schmitt There are for sure, some powerful interest at play. But let’s be clear. I think one is data, as you rightfully said, but the other one is algorithm. It’s not as if China is authorising any competitor on its territory. They have blocked access to most of the Internet platforms from the US, either finding new rules or just trade blocking them. So I don’t think it’s fair competition. You don’t want some of that data in China about the US or European consumer. Three, it’s about the algorithm. If suddenly, you are a foreign power, and you can as we know in China, you better follow what’s required of you from the Chinese Communist Party. You cannot take a chance with influencing other stuff like elections in other countries. It’s fair from the US perspective. One could even argue it’s fair from a Chinese perspective to want that. I think the only one in the middle who doesn’t really know what they want is Europe because on one side, they want to benefit from American platforms, on the other end, they want to have some controls. On the other end, they don’t create the environment for startups to flourish. So in that weird situation where they have to accept some control by the big US providers and either provider of underlying infrastructure or provider of consumer business facing services. Then they try to regulate them. But I think they are misunderstanding the power relationship, and I think some of this regulation would get some blowback, at least by the current administration. Just, I believe, this morning, there was some news around X being under a criminal investigation in France. This is not going to end well for the French startup and VC ecosystem. This is not going to end well for France and Europe when you depend so much from your American friends. Nuno Goncalves Pedro Regulation will be weaponised. Regulation constraints around exports, all of this will be weaponised geopolitically, and the bigger guys will normally win. I think that’s normally what we’ve seen. Just on TikTok just to… And you guys, if you’re listening to us, just see if you see a pattern here, but obviously, 19.9% still owned by ByteDance of the TikTok entity in the US. It was initially said that 80% of the TikTok entity is owned by non-Chinese investors. Initially, people were saying US investors, and then they changed it to non-Chinese because MGX, I think, has 15% of it. MGX is based in the UAE, connected obviously to Mubadala, the Abu Dhabi sovereign wealth fund. Silver Lake is in there, I think, with 15% as well. Oracle as well with 15%. Those three are the big bucket owners together, 45%. Silver Lake having collaborated with MGX before, and I’m sure a lot of connectivity there. Then you still see a pattern in this in terms of shareholders. If you don’t, then just Google it. Dell Family Office, Vastmir Strategic Investments, which is owned by billionaire Jeff Yass, Alpha Wave Partners, obviously involved with a bunch of things like SpaceX and Klarna, Virgoli, Revolution, which is Steve Case’s, a former founder of AOL, is also in there. Meritway, which is managed by partners, I think, of Dragonair. Vinova from General Atlantic, an affiliate of General Atlantic. Also, NJJ Capital, which I believe is Xavier Nil, the French billionaire that founded Iliad. Mostly American, I think, if the math is correct. 80% non-Chinese, which was what mattered, I think, in many cases. But do see if you saw a pattern in most of those investors. I won’t say anything more than that. Maybe moving to other topics, maybe just to finalise on regulation and geopolitics. In geopolitics, we should talk about wars if we predict anything. Not that we are nasty and one want to be negative, but what the hell is going on? Will we have ending to the wars we already have ongoing or not? But before that, the struggles on the App Stores, I think, will continue both for Apple and for Google Play Store. The writing’s on the wall, the EU keeps pushing it dramatically and Apple keeps just doing stuff. I’m on the board of an App Store company. Apple just creates all these things that basically make you not really… It doesn’t work. You can’t provision then an App Store on Apple devices. On iPhones, et cetera. We’ll see how that will continue going, but I feel the writing’s on the wall. Both Apple and Google will have to open up a bit more of their platforms. I’m not sure it will have a huge impact in the medium to long term, but definitely we need to see more openness in access to apps as given by the two big platform owners, Apple and Google, out there. Bertrand Schmitt Let’s be clear. Google is way more open than Apple. We both have Android devices. You can install alternative app stores. It’s a different ballgame by very far. Nuno Goncalves Pedro Google does other nasty stuff. It’s public. You can check which board I’m a part of. You can see what that company has done towards Google over time. But to your point, yes. It is true that Google has been more open than Apple, but Google has done their own things. Just to be very clear, so I’ll just leave that caveat bracketed there for people to think about it and maybe read a little bit about it as well. Bertrand Schmitt I can say that, me, from my perspective, that path of total control that Apple has been going through on all their devices, that includes macOS, pushed me to, over the past 2, 3 years, to completely live and abandon the Apple ecosystem. I just couldn’t accept that level of control, that golden handcuff approach of the Apple ecosystem, each their own obviously, they are golden, their handcuffs, but they are still handcuffs. Personally, that pushed me way more to Linux, Android, Windows, back to Windows after all these years. I just couldn’t stand it anymore. I want to pick my devices. I want to pick what I install on them, and I don’t want to be controlled like this by just one entity for all my tech devices. For me, at some point, it was just not acceptable anymore. It’s still very warm, very golden handcuffs, but for me, they were just handcuffs at this stage. Yes, what they are doing with the App Store is very typical of that mindset. I think it’s quite sad because I think it started with good intention in some ways. “We need a new computing paradigm, we need to make things smoother and safer,” but it has really become a way to control your clients. For me, it has reached a point where it’s just way too much. Nuno Goncalves Pedro There’s obviously the great power comes great responsibility that uncle Ben told Spider-Man or Peter Parker. But there’s also with great power comes shitload of money, and control. So it’s like, “Yeah. Should we open the server? Do we want to delay opening it up?” “Yeah.” Anyway, it is what it is. Maybe let’s end on the more difficult note of the episode, which is going to be around wars. What’s our prediction? Will we have an end to the Gaza situation with Israel? Will we have an end to Ukraine and, obviously, Russia? What will happen in Iran? Those are the three big, big conflicts right now. Then, obviously, if we want to add just bonus points, what’s going to happen to Greenland, and what’s going to happen to Taiwan, and what’s going to happen to Venezuela? Let’s throw the whole basket in there. We’ve never had like… Let’s talk about all these territories and all these countries. At some point in time, I’m saying this in a light manner, but it’s obviously more tragic than it should be light, and people are dying, and there’s a lot of implications of all of that that is happening right now. Do you have any predictions, Bertrand, for this year? Bertrand Schmitt No. It’s tough to predict on an individual basis. I think on a more bigger picture basis is on one side, obviously, the rise of China on one side. You have also the rise of other countries like India, while very indirectly connected to some of these conflicts are still part of the game, buying oil from Russia, for instance. At the same time, I think overall, the US is more clear about with the sheriff in town. I think it’s good because in some ways, you cannot pay for the goods, you cannot have such a massive advantage versus nearly every other country on earth and just not be clear about who is the boss in some ways. As a result, what are the rules of the game and how it should be played? The US is not alone, obviously, you have China, you have Russia, you have India, you have Europe. You have different other countries. But at some point, it’s not good when countries are not rational and are not clear. I think I prefer the current situation where things are more clear and where you have to assume responsibilities about what you are doing. It’s time to be rational again about how the world behave. Yes, the concept of power and balance of power. I think there has been that dream, maybe mostly coming from Europe, about the end of history. I think that’s simply not the case. It’s not the end of history. It’s still about the balance of power. It has always been about the balance of power. If you are dumb enough to think it was not about that anymore, I just have a bridge to nowhere to sell you. I don’t have specific prediction, but I think it’s clear there is a new sheriff in town. There is a new doctrine about the Western Hemisphere that has been in some ways resurrected on the [inaudible 00:51:35] train, and I think we’ll see more of it. I think at this point, the biggest question is for the Europeans. What do they want to do? Because right now, their position of being a dwarf militarily while being a pretty big giant economically, I don’t think it works. Nuno Goncalves Pedro I agreed on everything that you said. I do have predictions. I’ll stick a flag on the ground just with my predictions. Bertrand Schmitt Good luck. Nuno Goncalves Pedro They are mostly positive. I do think we’ll see an end or, for the most, end to the two big conflicts, the one in Gaza and the one in Ukraine. I think Ukraine will end up in readjustment of territory and splitting between Russia and the Ukraine, but the end of hostilities, I think that we will see an end to the conflict in Gaza also with a readjustment on what that will mean for the Palestinian territories and the Palestinians in general. That I’m not sure, but I feel that there will be an end to those two big conflicts. Iran, I have no clue. I will not put a stick on the ground that I have no clue. There are so many things that could go wrong there. I’ve been reading some really interesting thoughts about even some aggressive thoughts that this might be the time to really change regimes in Iran and for the US to have a bit more of an aggressive stance. I really don’t have a perspective. Obviously, there’s a lot at stake there. Then, if we talk about the other parts, Greenland, I will not opine too much on. Maybe we’re done for now. Maybe there’ll be some other concessions to the US that weren’t already there in the ’50s. Taiwan, I won’t bet either. I’m sad to say I think it might happen at some point in time, but I’m not sure when and what would drive it. Last but not the least, Venezuela is my only really negative prediction. I feel it will continue to be a significant dictatorship as it was before managed enough by other people with the difference now that it has a tax to be paid to the US in the form of oil of some sort, etcetera, and maybe gas, maybe other things as well that it didn’t have before. That’s probably my most negative prediction for the coming year on the geopolitical side. Bertrand Schmitt Without going into detail, I would mostly agree with what you shared. At least that makes sense. But as we know, it’s not always what makes sense, but what might happen. I can tell you 100% I would not have guessed this operation against Maduro. This was so well done, well executed, and shocking at the same time that it’s… I think it shows that it’s hard to guess some of this stuff because there are certainly some new ways to wage limited war, for instance. So it’s certainly interesting, and we certainly need to get used to pretty bombastic statements. But for Venezuela, I don’t think it can be worse than what it was before. I’m probably more optimistic that gradually it can get better. Nuno Goncalves Pedro Just to put perspective on why we’re not making predictions on some of these elements, I think this is a funny story, but I was in Madeira. Actually, first time I was in Madeira, although I’m originally from Portugal. I’ve never been to the islands. Obviously, as you guys know, or some of you might know, there’s a lot of connection between Madeira and Venezuela. There’s a lot of immigration from Madeira Islands to Venezuela. One of my Uber or Bolt drivers there in Madeira was Venezuelan. Was born in Venezuela, but Portuguese descent, et cetera. He was telling me this was still last year. Late last year. Because I told him I lived in US, et cetera, and he was like, “Oh, hopefully, Trump will get Maduro out of there.” In my mind, I was like, “Dude.” No disrespect to the gentleman, but it’s like, “Okay. Mike, your perspective on geopolitics is maybe a little bit exaggerated.” And a couple of days later, we know what happened. When geopolitical decisions are better predicted by some probably very astute Uber drivers, you’re like, “Maybe I shouldn’t make a bet. I have no clue what’s going to happen, no clue what’s going to happen in Greenland, et cetera.” Anyway, a couple of predictions on that element. Bertrand Schmitt That’s why it’s so right. You have to be careful with the prediction, but it doesn’t remove the fact that I think nations and companies that have to play a global game have to understand in some ways what is the game, what are the powers in place, what could happen potentially, but also be realistic. Not be about wish and dreams, but more about, what’s the power relationship? Who has the money? Who has the means? Who has the capacity to do this or that? Because if you start that way, at least the scope of what’s possible, what’s reasonable is more and more clear more quickly. Some stuff like happened with Maduro, I would never have predicted, but for sure, if there’s one country that can do this sort of stuff, it’s the US. I’m not sure anyone has a technology and the means in terms of support infrastructure to do something like this. It’s tough to predict what will happen a year from now for any specific country, but I think that even trying to get a better understanding about the forces in play and their capacity and understanding and accepting that at some point, it’s all about real politic and relationship of power, the more your eyes would be wide open about what’s possible versus simple, wishful thinking. Nuno Goncalves Pedro Fintech, Crypto and Frontier Tech Moving maybe to our last section around fintech, crypto, and frontier tech. For me, just two very quick predictions, views of the world. I think on the frontier tech side, I won’t make a prediction. I will just tell you all to go and listen to our episodes, the one on infrastructure, which is immediately prior to this one, and the episodes that we’ve had around a couple of other topics including AI, what’s the future of your children, because I think they illustrate a lot of the points that we’re seeing and manifesting themselves over the next year and over the next 2 or 3 years as well beyond that. I feel those tomes are complete in and out of themselves, so you can just go and listen to them. Then my second comment is on crypto. I feel crypto has become of the essence, particularly under the current administration in the US, very favored. Obviously, we are now in a world where crypto is just part of the economic system, and I think we’ll see more and more of that emerging, and in some ways, crypto is becoming mainstream. Question is what blockchains will be the blockchains of the future? Obviously, there’s a bunch of bets put out there. We, ourselves, as Chamaeleon, have one investment in one of the significant bets in the space. But besides that, who’s going to win or not, we feel that we’re past the crypto winter. It’s now mainstream days, and we’ll see a lot more activity in there. Bertrand Schmitt I must say with crypto, I’m a bit confused. As you say, we are past the crypto winter. There is much less uncertainty in regul

Spark of Ages
The Data Moat: A Google Veteran's Investment Thesis for AI/David Yakobovitch ~ Spark of Ages Ep 58

Spark of Ages

Play Episode Listen Later Feb 28, 2026 58:38 Transcription Available


We chart how AI leapt from chat to code, why product is now the leverage point, and how startups can market to algorithms without losing trust. David Yakobovitch shares hard-won views on moats, data, defense tech, and the immigrant energy powering American dynamism.• leaders and market share across Google, OpenAI, Anthropic• vibe coding benefits, code quality risks, review loops• prompt libraries, agent swarms, PRD automation• weekly shipping pace and the SaaS squeeze• marketing to algorithms, buyer agents, bot traffic control• pilot to production gap, rise of forward-deployed engineers• moats beyond models via domain, workflow, and proprietary data• China's progress, open source, and on-device AI bets• defense tech, swarms, and physical AI opportunities• endurance mindset, yoga discipline, and founder stamina• personal workflows across Gemini, Claude, and OpenAI• investing across seed and growth with outcome focusThe model wars aren't theoretical anymore—they're shaping how software gets built, shipped, and sold. We sit down with David Yakobovitch, GP at Data Power Capital and former global product lead at Google, to map where AI is actually working in 2026: vibe coding that shrinks teams, agent swarms that harden quality, and product-led moats that outlast model churn. David pulls back the curtain on how Claude, OpenAI, and Google now compete neck and neck on code and content, why prompt engineering as a job vanished while prompts became more valuable, and how forward-deployed engineers bridge the stubborn pilot-to-production gap that has haunted data projects for a decade.We explore go-to-market in a world where buyer agents screen your pitch before a human blinks. That means structuring materials for machines, tuning sites for humans and crawlers, and building demos that agents can evaluate safely. We also go into what happens as models commoditize: the moat shifts to domain depth, proprietary offline data, secure connectors, and measurable workflow outcomes. From small language models running on CPUs in air‑gapped containers to Apple's on-device bet, the edge is back—especially for Europe's sovereignty demands and public sector buyers.Then we widen the lens. Defense and “physical AI” blend hardware and autonomy: swarms, hypersonics, and resilient edge compute that must perform in the real world. David shares why he's backing both the silicon and the software, and how American dynamism—powered by immigrants and impatient builders—remains a durable advantage. Along the way, we trade notes on multi-model workflows, open source momentum, China's narrowed gap, and the endurance mindset that carries teams through the disappointment dip after the first shiny demo.David Yakoboitch: https://www.linkedin.com/in/davidyakobovitch/David Yakobovitch is a General Partner and Managing Director of DataPower Capital, a New York City-based venture capital firm investing across Applied AI, Inference Infrastructure, and DeepTech.  With a portfolio of over 36 companies, David is an investor in the most defining frontier technology firms of our era, including OpenAI, Anthropic, xAI, Neuralink, DataBricks, Groq, Cruesoe, Anduril and SpaceX. David is a leading voice as the host of HumAIn, a podcast focused on Applied and Responsible AI.  Previously, David served as a Global Product Lead aWebsite: https://www.position2.com/podcast/Rajiv Parikh: https://www.linkedin.com/in/rajivparikh/Sandeep Parikh: https://www.instagram.com/sandeepparikh/Email us with any feedback for the show: sparkofages.podcast@position2.com

Equity Mates Investing Podcast
What SpaceX shows us about private markets with Adam Myers

Equity Mates Investing Podcast

Play Episode Listen Later Feb 26, 2026 29:23


SpaceX. OpenAI. Anthropic. The companies everyone wants to own but can't buy on the share market. In this episode, we unpack how private equity works, why the biggest companies are staying private for longer, and how the Pengana Private Equity Trust (ASX:PE1) gives ASX investors exposure to SpaceX and 500+ other private companies.In this episode:0:00 SpaceX, IPO rumours & why private markets matter2:10 The economics of space6:01 Why companies are staying private longer10:51 Private equity 101: how it works14:03 Has private equity outperformed?18:59 Why PE1 is structured as a listed investment trust21:35 PE1 performance, buybacks & distributions24:41 Beyond SpaceX: AI exposure, GROQ & compoundersStocks & ETFs mentioned: Pengana Private Equity Trust (ASX:PE1), SpaceX (private), OpenAI (private), Anthropic (private), xAI (private), NVIDIA (NASDAQ:NVDA), Amazon (NASDAQ:AMZN), T-Mobile (NASDAQ:TMUS), Spice World (private), GROQ (private)None of Pengana Private Equity Trust (“PE1”), Pengana Investment Management Limited (ABN 69 063 081 612, AFSL 219 462) (“Responsible Entity”), Grosvenor Capital Management, L.P., nor any of their related entities guarantees the repayment of capital or any particular rate of return from PE1. Past performance is not a reliable indicator of future performance, the value of investments can go up and down. This document has been prepared by the Responsible Entity and does not take into account a reader's investment objectives, particular needs or financial situation. It is general information only and should not be considered investment advice and should not be relied on as an investment recommendation.Pengana Investment Management Limited (Pengana) (ABN 69 063 081 612, AFSL 219 462) is the issuer of units in the Pengana Private Equity Trust (ARSN 630 923 643) (the Trust). Before acting on any information contained within this report a person should consider the appropriateness of the information, having regard to their objectives, financial situation and needs. An investment in the Trust is subject to investment risk including a possible delay in repayment and loss of income and principal invested.———Want to get involved in the podcast? Record a voice note or send us a message.And come and join the conversation in the Equity Mates Facebook Discussion Group.———Want more Equity Mates? Across books, podcasts, video and email, however you want to learn about investing – [we've got you covered.Keep up with the news moving markets with our daily newsletter and podcast (Apple | Spotify)We're particularly excited to share our latest show: Basis PointsListen to the podcast (Apple | [Spotify)Watch on YouTubeRead the monthly email———Looking for some of our favourite research tools?Download our free Basics of ETF handbookOr our free 4-step stock checklistFind company information on TIKRResearch reports from Good ResearchTrack your portfolio with Sharesight———In the spirit of reconciliation, Equity Mates Media and the hosts of Equity Mates Investing acknowledge the Traditional Custodians of country throughout Australia and their connections to land, sea and community. We pay our respects to their elders past and present and extend that respect to all Aboriginal and Torres Strait Islander people today. ———Equity Mates Investing is a product of Equity Mates Media.This podcast is intended for education and entertainment purposes. Any advice is general advice only, and has not taken into account your personal financial circumstances, needs or objectives. Before acting on general advice, you should consider if it is relevant to your needs and read the relevant Product Disclosure Statement. And if you are unsure, please speak to a financial professional. Equity Mates Media operates under Australian Financial Services Licence 540697. Hosted on Acast. See acast.com/privacy for more information.

Smart Humans with Slava Rubin
Smart Humans: Pre-IPO investor Briefing on Databricks, Groq, Anduril, Anthropic, and Canva, w/ Sacra's Jan-Erik Asplund

Smart Humans with Slava Rubin

Play Episode Listen Later Feb 18, 2026 54:07


Recorded 10/29/25Vincent's Slava Rubin and Sacra's Jan-Erik Asplund discussed Databricks, Groq, Anduril, Anthropic, and Canva, five of the hottest pre-IPO companies in the asset class - and how investors can get access to them.Presented by the Fundrise Innovation Fund.https://fundrise.com/Vincent

Daily Tech News Show
AI Accelerators Could Change Everything - DTNS WEEKEND

Daily Tech News Show

Play Episode Listen Later Jan 24, 2026 19:05


Andrew Mayne explains how chips like Groq and Cerebras are giving us faster training and inferencing, and how that can help save us time and power and increase productivity.Featuring Tom Merritt and Andrew Mayne. Hosted on Acast. See acast.com/privacy for more information.

The Twenty Minute VC: Venture Capital | Startup Funding | The Pitch
20VC: Groq's $20BN NVIDIA Acquisition | Manus Acquired by Meta for $2BN | Why Sam Altman Does Not Care About Dilution | Navan Trading at 4x ARR & Why Going Public Does Not Make Sense Anymore | The Rise of Invisible Unemployment and Labour Markets in

The Twenty Minute VC: Venture Capital | Startup Funding | The Pitch

Play Episode Listen Later Jan 8, 2026 83:47


AGENDA: 04:30 Groq Acquired by NVIDIA for $20BN: The Breakdown 17:13 Meta's $2BN Acquisition of Manus: Did They Sell Too Early 36:04 OpenAI's Stock-Based Compensation Strategy 47:42 Will AI Replace Venture Capitalists 56:13 Navan Trading at 4x ARR: Who is Good Enough to Go Public? 01:09:46 The Rise of Invisible Unemployment 01:14:21 The Future of Work and Education in an AI-Driven World    

FYI - For Your Innovation
Major Shifts In The AI Landscape | The Brainstorm EP 115

FYI - For Your Innovation

Play Episode Listen Later Jan 7, 2026 51:08


In this episode of The Brainstorm, we discuss the latest developments in AI and technology, including NVIDIA's strategic moves with Groq and Meta's acquisition of Manus AI. We explore the implications of these acquisitions on the AI landscape, the potential for orchestration layers in AI models, and the competitive dynamics among major tech companies. The conversation also touches on the future of user interfaces and the evolving role of voice and text in consumer interactions.If you know ARK, then you probably know about our long-term research projections, like estimating where we will be 5-10 years from now! But just because we are long-term investors, doesn't mean we don't have strong views and opinions on breaking news. In fact, we discuss and debate this every day. So now we're sharing some of these internal discussions with you in our new video series, “The Brainstorm”, a co-production from ARK and Wolf.financial, and sponsored by Public. Tune in every week as we react to the latest in innovation. Here and there we'll be joined by special guests, but ultimately this is our chance to join the conversation and share ARK's quick takes on what's going on in tech today.Key Points From This Episode:NVIDIA's strategic investment in Groq highlights its focus on enhancing AI chip capabilities without full acquisition, aiming to secure a competitive edge.Meta's acquisition of Manus AI emphasizes the importance of orchestration layers in delivering agentic AI experiences, integrating multiple models for diverse applications.The discussion explores the evolving AI landscape, questioning whether foundational models or their applications (wrappers) hold more value for end-users.The hosts debate the future of user interfaces, predicting a shift towards voice interactions and the potential for new hardware innovations.The episode concludes with a look at the competitive dynamics in the tech industry, particularly the role of initial public offerings (IPOs) and acquisitions in shaping market leadership.To learn more about WOLF: https://wolf.financialTo learn more about Public: https://public.com/

This Week in Startups
2026 Starts with a bang: META AI Drama and Nvidia's $20B Groq Acquisition | E2230

This Week in Startups

Play Episode Listen Later Jan 6, 2026 54:39


This Week In Startups is made possible by:Crusoe Cloud - https://crusoe.ai/buildUber - http://uber.com/twistEvery.io - http://every.io/Today's show: Jason and Alex are BACK on TWiST for 2026! This holiday season was anything but calm, with deca-corn acquisitions, massive Polymarket bets, and major new startups breaking from stealth!Jason talks the recent Nvidia-Groq $20B acquisition, a major exit for Chamath as the lead investor back in 2017! Jason delves into how the VC fund math shapes out for pre-seed VC funds vs. Series A VC funds.Jason and Alex delve into drama swirling META's AI team. Yann LeCun, META's former Chief AI Scientist, announced that he would be leaving META to become Executive Chairman at AMI Labs. LeCun left the META team in the new year, calling the new Chief AI Scientist, Alexandr Wang, inexperienced. LeCun now looks to move AI beyond the era of LLM at AMI Labs.PLUS Jason and Alex talk about the new social media app Tangle, from Biz Stone, co-founder of Twitter, and Evan Sharp, co-founder of Pinterest. Their Startup, West Co, launched tangle, which seeks to become an “intentional living” app. The two look to improve how humans interact with modern tech. Jason points out that very few news products have worked, but is eager to see how two industry veterans build in the space. Timestamps:(00:00) Why Restaurants are OVER — Peptides and other self medications(06:41) Nvidia Acqui-Hires Groq for $20 BILLION(9:48) Crusoe Cloud: Crusoe is the AI factory company. Reliable infrastructure and expert support. Visit https://crusoe.ai/build to reserve your capacity for the latest GPUs today.(11:00) The VC fund math between seed vs. Series A funds(15:00) META buys TWiST 500 Company, Manus! Why it matters.(20:20) Uber AI Solutions: Your trusted partner to get AI to work in the real world. Book a demo with them TODAY at http://uber.com/twist(21:24) Why Yann LeCun left META, and what could be behind it(25:27) Producer Claude on the Gondola Crash in Zurich(29:13) Jason's Request for Augmented human intelligence(30:11) Every.io - For all of your incorporation, banking, payroll, benefits, accounting, taxes or other back-office administration needs, visit http://every.io/(32:04) How one Trader made $436.8k on one bet on polymarket!(36:05) Jason's Predictions for 2026 IPOs(40:01) Is news broken? How Tangle is tackling it.(45:53) How much should startup incur in legal expenses? Should founders try to use AI to avoid costs?(50:59) Why Google should let NotebookLM cook, make it a standalone brand! *Subscribe to the TWiST500 newsletter: https://ticker.thisweekinstartups.com/Check out the TWIST500: https://twist500.comSubscribe to This Week in Startups on Apple: https://rb.gy/v19fcp*Follow Lon:X: https://x.com/lons*Follow Alex:X: https://x.com/alexLinkedIn: https://www.linkedin.com/in/alexwilhelm/*Follow Jason:X: https://twitter.com/JasonLinkedIn: https://www.linkedin.com/in/jasoncalacanis/*Thank you to our partners:(9:48) Crusoe Cloud: Crusoe is the AI factory company. Reliable infrastructure and expert support. Visit https://crusoe.ai/build to reserve your capacity for the latest GPUs today.(20:20) Uber AI Solutions: Your trusted partner to get AI to work in the real world. Book a demo with them TODAY at http://uber.com/twist(30:11) Every.io - For all of your incorporation, banking, payroll, benefits, accounting, taxes or other back-office administration needs, visit http://every.io/

The AI Breakdown: Daily Artificial Intelligence News and Discussions
What Manus and Groq Acquisitions Tell Us About AI

The AI Breakdown: Daily Artificial Intelligence News and Discussions

Play Episode Listen Later Jan 3, 2026 25:54


Two blockbuster deals over the holidays quietly marked the real start of the AI agent era, revealing where competition is actually heading in 2026. This episode breaks down why Meta's acquisition of Manus signals a shift toward agents as distribution, not features, and why Nvidia's $20B Groq deal is really about owning the future of inference as workloads fragment and latency becomes decisive. Together, these moves show how the battle is moving from models and benchmarks to agents, infrastructure, and the interfaces people refuse to leave. In the headlines: xAI's massive compute expansion, OpenAI's renewed push on voice and devices, SoftBank's infrastructure spree, Brookfield's AI cloud ambitions, and Claude Code reaching the point of writing all of its own code. Brought to you by:KPMG – Discover how AI is transforming possibility into reality. Tune into the new KPMG 'You Can with AI' podcast and unlock insights that will inform smarter decisions inside your enterprise. Listen now and start shaping your future with every episode. ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.kpmg.us/AIpodcasts⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Blitzy.com - Go to ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://blitzy.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ to build enterprise software in days, not months Robots & Pencils - Cloud-native AI solutions that power results ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://robotsandpencils.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠The Agent Readiness Audit from Superintelligent - Go to ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://besuper.ai/ ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠to request your company's agent readiness score.The AI Daily Brief helps you understand the most important news and discussions in AI. Subscribe to the podcast version of The AI Daily Brief wherever you listen: https://pod.link/1680633614Interested in sponsoring the show? sponsors@aidailybrief.ai

Big Technology Podcast
Meta's AI Agent Plan, Grok's Perversion, Prison Of Financial Mediocrity

Big Technology Podcast

Play Episode Listen Later Jan 2, 2026 49:07


Ranjan Roy from Margins is back for our weekly discussion of the latest tech news. This week, we do our 2026 predictions in an abbreviated holiday-time episode. Here's what we cover: 1) Meta buys Manus 2) Is the Manus deal an enterprise play? 3) What Meta could do with consumer AI agents 4) Why consumer AI agents are a good advertising strategy for Meta 5) Instagram head Adam Mosseri addresses AI slop 6) Meta Ray-Bans don't work in the cold 7) NVIDIA pretty much buys Groq 8) Elon Musk's Grok goes full pervert 9) Who's responsible? 10) What explains the rise of sports betting and prediction markets -- is it a lack of a stable financial future that would otherwise be worth investing in? --- Enjoying Big Technology Podcast? Please rate us five stars ⭐⭐⭐⭐⭐ in your podcast app of choice. Want a discount for Big Technology on Substack + Discord? Here's 25% off for the first year: https://www.bigtechnology.com/subscribe?coupon=0843016b Learn more about your ad choices. Visit megaphone.fm/adchoices

All-In with Chamath, Jason, Sacks & Friedberg
Massive Somali Fraud in Minnesota with Nick Shirley, California Asset Seizure, $20B Groq-Nvidia Deal

All-In with Chamath, Jason, Sacks & Friedberg

Play Episode Listen Later Dec 31, 2025 103:22


(0:00) Bestie intros! Nick Shirley joins the show to discuss his recent investigation on potential daycare fraud in Minnesota (3:32) Nick's background, how he got into investigative reporting and YouTube, independence, finding this story (16:36) Why this fraud story is resonating, why the national press initially avoided it (30:08) Future plans, California, possible Al-Shabaab connection, how high up does Minnesota's fraud go? (49:15) What the scale of fraud means for America, Minnesota's future, potential patronage scheme (1:09:06) CA's wealth tax: normalizing the seizure of private property (1:33:56) Chamath breaks down the $20B Groq-Nvidia deal Follow Nick Shirley: https://x.com/nickshirleyy Follow the besties: https://x.com/chamath https://x.com/Jason https://x.com/DavidSacks https://x.com/friedberg Follow on X: https://x.com/theallinpod Follow on Instagram: https://www.instagram.com/theallinpod Follow on TikTok: https://www.tiktok.com/@theallinpod Follow on LinkedIn: https://www.linkedin.com/company/allinpod Intro Music Credit: https://rb.gy/tppkzl https://x.com/yung_spielburg Intro Video Credit: https://x.com/TheZachEffect Referenced in the show: https://x.com/nickshirleyy/status/2004642794862961123 https://www.startribune.com/prosecutors-charge-5-people-in-a-minnesota-housing-fraud-scheme/601548944 https://www.nytimes.com/2025/11/29/us/fraud-minnesota-somali.html https://www.fox9.com/news/fraud-minnesota-detailing-nearly-1-billion-schemes https://x.com/EricLDaugh/status/2005410646603473256 https://x.com/kevinkileyca/status/2006053056660541840 https://x.com/chamath/status/2006087862492582084 https://x.com/C_3C_3/status/2005722313795440956 https://x.com/OliLondonTV/status/2005988021946999166 https://x.com/tomhennessey69/status/2005556784228909441 https://x.com/WallStreetApes/status/2005849513676923358 https://x.com/MarioNawfal/status/2005179409465299219 https://dcyf.mn.gov/programs-directory/child-care-assistance-program https://x.com/susancrabtree/status/2006079778873565541 https://x.com/chamath/status/2005386348169953607 https://x.com/aaronburnett/status/2003874734661161064 https://newsletter.amuseonx.com/p/the-somali-patronage-system-has-taken https://x.com/realdailywire/status/2006122428196442388 https://x.com/rightanglenews/status/2006375449404866720 https://www.auditor.ca.gov/reports/2025-601/

DH Unplugged
DHUnplugged #784: Auld Lang Xiety

DH Unplugged

Play Episode Listen Later Dec 31, 2025 63:01


Looking at a weird GDP data point. Calling BS on Russia/Ukraine peace talks. Gold and Silver – WOW! Closing out the year – a good one too! PLUS we are now on Spotify and Amazon Music/Podcasts! Click HERE for Show Notes and Links DHUnplugged is now streaming live - with listener chat. Click on link on the right sidebar. Love the Show? Then how about a Donation? Follow John C. Dvorak on Twitter Follow Andrew Horowitz on Twitter Warm-Up - CTP Cup - All systems go! 9 participants! - Lots to be excited about and anxious too - Looking at a weird GDP data point - Calling BS on Russia/Ukraine peace talks Markets - Gold and Silver - WOW! - Closing out the year - a good one too! - Buyers are still hot to buy any dip - "Diet" pills coming Bitters Making Progress  - Chocolate -Dark Cherry -Infusions - https://highdesertbotanicals.com NYE Celebration - Cities across America ring in the new year by dropping unexpected objects: - Amelia Island, FL drops a giant shrimp. - Nashville drops a 400lb musical note with 28,140 LEDs. - Boise, ID, drops a glowing potato. - Key West, FL, drops an eight-foot ruby-red heel—complete with a drag queen inside! - In Spain, revelers gulp down 12 grapes—one for each midnight chime—to bring luck for each month - Denmark - Danes toss old dishes at friends' doors—large piles of broken crockery at dawn are seen as tokens of good luck. What a year! - So many themes in 12 months - AI, Tariffs, War and Trade War, Fat drugs, Deglobalization - Data centers, semiconductors, and supporting infrastructure like power and cooling systems. - Approx: DJIA +13.5%, SP500 +17%, NASDA +21%, BTCUSD -7.6%, Gold +64%, SLV +145%, $DXY -9.5%, EEM +30% - 2026 - Opportunities and Auld Lang Xiety (Tech still looks frothy in certain names) Top New Year's Resolutions - Exercise More - Eat Healthier - Save More Money/Get Out of Debt - Be Happy/Improve Mental Health - Lose Weight - Spend More Time with Family & Friends - Learn a New Skill/Hobby - Get Organized Active Management (Funds) - Same report annually - A small group of tech super stocks accounted for an outsize share of returns in 2025, extending a pattern in place for the better part of a decade. - Around $1 trillion was pulled from active equity mutual funds over the year, marking an 11th year of net outflows, while passive equity exchange-traded funds got more than $600 billion. - The concentration of gains in a few stocks made it harder for active managers to do well, with 73% of equity mutual funds trailing their benchmarks this year, the fourth most in data going back to 2007. - BUT, there are some areas that it makes sense for active management ---- Equity vs Fixed income and reasoning --- Efficient markets, boots on the ground Fat Pill - The FDA has approved the first-ever GLP-1 pill from Wegovy maker Novo Nordisk. - Novo Nordisk said the starting dose of 1.5 milligrams will be available in early January in pharmacies and via select telehealth providers with savings offers for $149 per month. - The approval gives Novo Nordisk a head start over chief rival Eli Lilly, which is racing to launch its own obesity pill. - Packaged food makers and fast-food restaurants may be forced to overhaul more of their products next year as newly approved, appetite-suppressing GLP-1 pills become available in January PowerBall - A ticket sold in Arkansas scored a $1.8 billion Powerball jackpot after Wednesday night's draw — one of the richest lottery prizes in U.S. history, landing just in time for Christmas. - The payout soared after last Monday's drawing produced no winners, with last-minute ticket sales pushing the jackpot to $1.817 billion. That makes it the second-largest U.S. lottery prize ever and the biggest Powerball of 2025, the lottery website said on Thursday. - The winning numbers — 4, 25, 31, 52, 59 and the Powerball 19 - Odds: one in 292.2 million. Silver - Amazing year! - Sunday night futures - >$83 then turned hard lower| - Down 7% on Monday - Range $83 - $71 (15%) for the day - Some rumors about a bank collapse due to wrong way position on Silver - forced liquidation and covering.... ----- Hard to believe that a bank was short that much silver - but..... SoKo Breach - South Korean online retail giant Coupang said it will offer 1.69 trillion South Korean won ($1.17 billion) in compensation to 34 million users affected by a massive data breach disclosed last month. - That is about 4% of Coupang's annual revenue - but a big chunk of their profit - $34 per user NVDA Deal - Nvidia has yet to issue a public announcement or disclosure regarding its $20 billion Groq deal that CNBC was first to cover on Wednesday. - Groq described the deal as a “non-exclusive licensing agreement,” a tool that's been used by tech giants of late in part to avoid regulatory scrutiny. - Analyst: “Antitrust would seem to be the primary risk here, though structuring the deal as a non-exclusive license may keep the fiction of competition alive,” Bernstein's Stacy Rasgon wrote in a report. - Groq will remain an independent company (?) GDP Consumption - Something is a bit off.... - With the marketplace costs increasing, this may be more than a one-off expenditure Q3 GDP Surge Russia/Ukraine - Less that an hour after the White House claimed great movement toward peace - Russian President Putin told President Trump that Russia will revise its negotiating position, raising questions over prospects for peace deal - Russian Foreign Minister Sergei Lavrov says Ukraine tried to attack Russian President Putin's residence - Does anyone even listen to the crap coming out of the White House anymore? - Did you hear Lutnick trying to explain the 600% reduction in costs for pharmaceuticals? Math wizards! - - For 2026, my wish is that they continue to work on the job at hand and just shut up Just for fun - Who is biggest drinker of spirits? - While there's no single official "heaviest drinker," legendary wrestler Andre the Giant is widely cited as having unmatched capacity, famously downing 119 beers in one sitting (or even up to 156 in other accounts) Oil - Crude oil futures down about 9.5% YTD - Much of the drop due to pick up in production (supply/demand) - Still a floor with as Russia, Nigeria, Venezuela etc - What will it take to move up? Best Auto Stock for 2025? - GM! Better than ford, Tesla and others (up 55%) - best year from coming out of bankruptcy in 2009 - Ford up 35% - Mary Barra, CEO selling into the strength - $73 M sold this year (Position down 73% from what she held last year) - - - Barra has contended for years that stock undervalued. With all of these say what does that say now? --- Would she ever say shares are overvalued? More fun stats - A peer?reviewed 2025 study estimates AI data centers (including indirect usage from electricity generation) consumed 312–765 billion liters of water annually. That's more than all bottled water consumed worldwide each year - Direct (on-site) water is used for cooling servers via systems like cooling towers or liquid loops. Indirect (off-site) water stems from electricity generation—particularly from thermal and nuclear plants, which require significant cooling resources - ??? Estimates suggest a single standard AI prompt (about 100 words) is linked to around 1.5 liters of water—accounting for the entire chain of consumption. (This is total usage from cooling powr consumption, electricity generation) - Global AI workloads consumed 50–60 terawatt-hours (TWh) in 2025—roughly the annual electricity use of a medium-sized country like Switzerland. - By 2030, AI-related electricity demand could reach 300–500 TWh annually, according to energy analysts—comparable to the entire electricity consumption of countries like France. Over to Iran - President Trump tells reporters that if Iran is building up its nuclear program, the U.S. will have to "knock them down" again --- Wait - I thought we destroyed all of their nuke aspirations??? - - - AND - Iran's currency hit a record low, triggering wave of protests, according to Bloomberg Fed News - Top Fed Chair Candidate Odds Narrow Again, With Hassett at 43% and Warsh at 35% - President Trump still angry at Powell 0threating to sue for incompetence Odd - Tesla Inc. published a series of sales estimates indicating the outlook for its vehicle deliveries may be lower than many investors were expecting. - The carmaker posted estimates showing analysts on average expect the company to deliver 422,850 cars in the fourth quarter, down 15% from a year earlier. - Tesla is on course for its second consecutive drop in annual vehicle sales, with the company compiling an average estimate for 1.6 million deliveries, down more than 8% from a year earlier. - These are estimates published by analysts - Tesla put on its own site - WHY? End of Year Stat - The U.S. national debt is climbing at a rapid pace and has shown no signs of slowing down despite the growing criticism of massive levels of government spending. - The national debt, which measures what the U.S. owes its creditors, rose to $38,386,384,190,622.68 as of Dec. 30, according to the latest numbers published by the Treasury Department. - That is an increase of about $5.8 billion daily - ~$18 per person in the US per day increase ($7,300) - or about the monthly price of leasing a small Mercedes - Each person in US owes approx $128,000 Love the Show? Then how about a Donation? THE CLOSEST TO THE PIN 2025 Winners will be getting great stuff like the new "OFFICIAL" DHUnplugged Shirt! CTP CUP 2025 Participants: Jim Beaver Mike Kazmierczak Joe Metzger Ken Degel David Martin Dean Wormell Neil Larion Mary Lou Schwarzer Eric Harvey (2024 Winner) FED AND CRYPTO LIMERICKS See this week's stock picks HERE Follow John C. Dvorak on Twitter Follow Andrew Horowitz on Twitter

Techmeme Ride Home
Nvidia Kindacquires Groq

Techmeme Ride Home

Play Episode Listen Later Dec 29, 2025 20:57


Nvidia kindaquires Groq I'll tell you what we think the strategy is here. A shot across my bow that CES is next week. Accountants shut down remote testing because of AI. And for all the recent bullishness, an honest look at the immediate limitations of today's robotics. Nvidia Reaches Technology Licensing Deal With Startup Groq (Bloomberg) Why Nvidia Struck a $20 Billion Megadeal with Groq (The Information) Samsung brings Google Photos to the biggest screen in your home (AndroidPolice) Accounting body scraps remote exams to combat cheating (Financial Times) Even the Companies Making Humanoid Robots Think They're Overhyped (WSJ) Learn more about your ad choices. Visit megaphone.fm/adchoices

Squawk on the Street
SOTS 2nd Hour: Debasement Trade, Software Vs. Hardware, & Trump's High Stakes Mar-A-Lago Meetings 12/29/25

Squawk on the Street

Play Episode Listen Later Dec 29, 2025 43:48


Sara Eisen, and David Faber began the hour with a look at the precious metals rally - and why it's tied to the debasement trade - before discussing the broader market outlook with Trivariate's Adam Parker. Plus: is it time to go from hardware to software? Hear one veteran tech investor's take on why 2026 will see "mindblowing" advancements in the latter sector - and what it means for stocks... and former DOJ antitrust watchdog Jonathan Kanter's opinion on whether Nvidia's GROQ deal is a new way for companies to avoid scrutiny from regulators. Also in focus: a high stakes meeting today between the President and Israeli Prime Minister Benjamin Netanyahu - the team discussed the latest and what's at stake with former Council on Foreign Relations head Richard Haass.  Squawk on the Street Disclaimer Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

WSJ Tech News Briefing
TNB Tech Minute: Nvidia Licenses Groq's AI-Inference Technology

WSJ Tech News Briefing

Play Episode Listen Later Dec 26, 2025 2:09


Plus: China sanctions U.S. defense companies and executives including Northrop Grumman, Boeing and Palmer Luckey over Taiwan arms sale. And Google will let users change their Gmail address. Julie Chang hosts. Learn more about your ad choices. Visit megaphone.fm/adchoices