POPULARITY
Categories
So far, most electric vehicles have looked more or less like cars. But recently, a few companies have looked to another, smaller mode of transport for inspiration. The Verge's Andy Hawkins explains why companies like Amble and Chip are reinventing the golf cart, in the hopes of creating an entirely new kind of street-legal vehicle. As long as the speed limit stays low. Further reading: The Light Flip is a minimalist flip phone with a point to prove | The Verge Anthropic has to pay authors. | The Verge The cost of GPUs goes far beyond AI data centers | The Verge Is America ready for this quirky Jeep-looking EV that can park itself? The ‘G-Wagen of golf carts' could be the ideal second car America's cheapest new EV is smaller than a ping-pong table and tops out at 19mph Zoox's purpose-built robotaxi is getting a refresh Subscribe to The Verge for unlimited access to theverge.com, subscriber-exclusive newsletters, and our ad-free podcast feed. We love hearing from you! Email your questions and thoughts to vergecast@theverge.com or call us at 866-VERGE11. Learn more about your ad choices. Visit podcastchoices.com/adchoices
Is Cerebras Systems the next great AI chip stock or a red-hot IPO priced for perfection? In this episode of 7investing Live, Simon Erickson and executive producer Heather Horton welcome back Nick Rossolillo, co-founder of Chip Stock Investor, to break down three of the market's biggest stories.First up: Cerebras Systems (NASDAQ:CBRS), the wafer-scale chip maker that just IPO'd at a $40+ billion market cap. With 44GB of SRAM embedded directly on the chip, Cerebras was purpose-built to solve AI's "memory wall" problem for inference workloads. Now it's reportedly landed a ~$10 billion order from OpenAI and a deal with Amazon Web Services that could top $20 billion. Simon and Nick dig into whether these massive orders are real, how Cerebras stacks up against NVIDIA's GPUs and hyperscaler custom silicon, the TSMC capacity bottleneck that could throttle its growth, and how to value a company trading near 20x sales without profits.Then the conversation turns to Rocket Lab (NASDAQ:RKLB), which has pulled back from $150 to around $70 per share. Simon shares the latest iteration of his discounted cash flow valuation, and the duo debates the proposed Iridium acquisition — a deal that could pull Rocket Lab to EBITDA-positive on a pro forma basis — plus what the long-awaited Neutron rocket launch means for the company's future.Finally: Netflix (NASDAQ:NFLX). After another quarter of decelerating revenue guidance, is the streaming giant now a value stock rather than a growth stock? Nick explains why the advertising business hasn't reaccelerated growth the way he expected, and what he'd need to see before buying the dip.Plus: Nick's take on the recent chip stock sell-off across NVIDIA, AMD, Broadcom, SanDisk, and Kioxia and why "stocks go up, stocks go down" might be the healthiest way to think about it.Subscribe for more deep dives on AI infrastructure, semiconductors, and innovative growth stocks!Start your FREE 7-day trial of 7investing: https://www.7investing.com/subscribeFollow Nick and Casey Rossolillo at Chip Stock Investor: https://chipstockinvestor.comRocket Lab Deep Dive videos mentionedPart 1 https://youtu.be/AMDd0-JKUH0 (Deep Dive)Part 2: https://youtu.be/Z76xTGFNwBA (Valuation)Companies MentionedPublicly Traded:Cerebras Systems (NASDAQ:CBRS)Rocket Lab (NASDAQ:RKLB)Netflix (NASDAQ:NFLX)NVIDIA (NASDAQ:NVDA)Advanced Micro Devices (NASDAQ:AMD)Broadcom (NASDAQ:AVGO)Micron Technology (NASDAQ:MU)Taiwan Semiconductor Manufacturing (NYSE:TSM)Amazon (NASDAQ:AMZN)Alphabet (NASDAQ:GOOGL)Meta Platforms (NASDAQ:META)Iridium Communications (NASDAQ:IRDM)SanDisk (NASDAQ:SNDK)Kioxia Holdings (TSE:285A)Globalstar (NASDAQ:GSAT)SpaceX (NASDAQ: SPCX)Private / Pre-IPO:OpenAIAnthropicVideos Mentioned:https://www.youtube.com/watch?v=Z76xTGFNwBA&t=3shttps://www.youtube.com/watch?v=AMDd0-JKUH0&t=987sHere's the shifted chapter list, with all timestamps moved back 55 seconds:0:00 Welcome to 7investing Live0:54 Cerebras Systems: IPO recap & the Wafer-Scale Engine2:31 Is NVIDIA even the right comparison for Cerebras?5:38 The memory wall: why bigger AI models need new chips8:52 Latency vs. throughput — and the new AI alliances10:46 Are the $10B OpenAI & $20B Amazon orders real?14:02 Cerebras risks: how do you value a hot IPO?17:27 The TSMC capacity bottleneck20:01 Heather's take on Cerebras20:41 Rocket Lab: the sell-off & Iridium acquisition24:34 Simon's DCF valuation & price target for RKLB29:05 Why Neutron changes everything30:12 Q&A: Does Peter Beck carry an "Elon premium"?31:36 Netflix: buying opportunity or cheap for a reason?36:57 Q&A: Is Netflix a growth stock or a value stock?39:03 Chip stocks selling off: normal volatility or a warning?42:57 Wrap-up & final thoughts#7investing #Simonerickson #Cerebras #CBRS #NVIDIA #AIinvesting #semiconductors #chipstocks #RocketLab #RKLB #Netflix #NFLX #AIinference #stocks #investing #stockmarket #TSMC #AIdatacenters
Don’t Fade and Die in AI Subscribe to our Newsletter: https://theultimatepartner.com/ebook-subscribe/ Check Out UPX: https://theultimatepartner.com/experience/ Matt Yanchyshyn, VP AWS Marketplace, Rekha Thangelapalita, Elastic GSI Leaders; Allison McFadden, Accenture AWS Leader; and James Kang of Nvidia join Ultimate Partner. In this panel discussion, leaders from Elastic, Accenture, Nvidia, and AWS dissect the urgent shifts in the ecosystem, emphasizing that partners must adapt to AI and agentic co-selling or risk fading away completely. The conversation explores the necessity of deep co-engineering, the power of multi-product solutions in the AWS marketplace, and how automated agents are now replacing traditional human sales pipeline progression. By embracing data readiness and strategic collaboration, organizations can survive the “token maxing” era, effectively scale their enterprise opportunities, and align with NVIDIA’s five-layer strategy to dominate the new cloud landscape. https://youtu.be/zUkL4Wqsa68 Key Takeaways AI agents will automate the majority of AWS partner co-selling attachments and opportunity progressions this year. Partners who fail to embrace agentic workflows and automated governance face the existential risk of fading into obsolescence. Successful multi-product offerings require a “blood to all organs” approach that benefits the client, the ISV, the GSI, and the hyperscaler simultaneously. Nvidia’s “five-layer cake” model emphasizes that successful outcomes at the application layer automatically drive growth for all underlying infrastructure. The “token maxing” phenomenon is forcing enterprises to seek cost-effective, open-model alternatives to scale their generative AI securely. Integrating GSIs and ISVs on the AWS marketplace significantly increases enterprise deal sizes and long-term customer renewal rates. If you're ready to lead through change, elevate your business, and achieve extraordinary outcomes through the power of partnership—this is your community. At Ultimate Partner® we want leaders like you to join us in the Ultimate Partner Experience – where transformation begins. Key Tags strategic collaboration agreement, data readiness engine, agentic co-sell, semantic layer, token maxing, five layer cake, accelerated computing platform, open models, cloud consumption, multi-product solutions, partner central agents, propensity data, automated opportunity progression, generative AI governance Transcript Matt Y and Panel Audio Podcast [00:00:00] Vince Menzione: You have a choice. You can embrace them and figure it out and get governance and, and make your data available. Um, use the partner, central agent, move to Agen Co-sell, or you can fade and die. [00:00:11] Vince Menzione: You can feel it happening. The ecosystem is shifting beneath us, the way Hyperscalers are partnering, how AI is remaking the channel and what it means to win in 2026. [00:00:22] Vince Menzione: Welcome to the Ultimate Partner Podcast. I’m Vince Menzi. Own your host. And each week I sit down with leaders at the intersection of technology, partnerships and outcomes. The voices shaping how ecosystems actually work. We talk about what’s real, what’s changing, and what it takes to lead in this era where the partner channel isn’t just part of the strategy. [00:00:44] Vince Menzione: It is the strategy because [00:00:46] Vince Menzione: being in the room changes everything. Let’s start. [00:00:51] Vince Menzione: We’ve got some amazing leaders joining us. So I think probably for a little bit of context, maybe just start with Rika. You can introduce yourself, your role and, uh, what, what you’ve been doing at Elastic. Yeah. [00:01:03] Rekha Thangellapalli: Yeah, sounds great. [00:01:04] Rekha Thangellapalli: Hi everyone. I’m Reka and I lead GSI Alliances at Elastic. Um, for the past 14 years, I’ve had the pleasure of building different kinds of partner ecosystems across companies such as SAP. MuleSoft, Salesforce, Coupa, and now Elastic. Um, I wanna thank Ultimate partner and Vince for having us here today. Thank you and the panel of these incredible speakers for joining me on stage. [00:01:31] Rekha Thangellapalli: Um, very excited for the conversation today. [00:01:33] Vince Menzione: We love Elastic, and you’ve had some of your other leaders on stage at other events. As such, the quality of your leadership team is amazing. Thank you. [00:01:42] Rekha Thangellapalli: I wholeheartedly agree. [00:01:45] Allison McFadden: Excellent. Um, hello everyone. Allison McFadden. I lead our North America AWS practice at Accenture. [00:01:52] Allison McFadden: Uh, I’ve been there for five years, and truth be told, it was my first partnership role, my first formal partnership role. Uh, so I can take some tips from all of you in the room here today. Prior to that, I was 21 years with IBM, and I got into partnerships because my last role at IBM was actually trying to build. [00:02:14] Allison McFadden: Linux business on the mainframe, and I had to have partners. I had to have partners to help me with workloads to run there. So I kind of learned, uh, trial by fire. But I’m excited for the conversation today. Excited to be in this room and excited to talk about what we’re doing with, uh, elastic. Thank you. [00:02:34] James Kang: Uh, my name is James Kang. Nice to see and meet everyone here. Vince, thank you for the opportunity. Thank you [00:02:38] Vince Menzione: for being here. [00:02:39] James Kang: Um, I’m with Nvidia, so I help manage the AWS partnership at Nvidia all up. Um, I guess fun fact, I’m former AWS and so I see a lot of very familiar faces here in the front row. Uh, former colleagues and then current friends. [00:02:56] James Kang: And so, uh, looking forward to the conversation. [00:02:59] Vince Menzione: Great. Well, we’ll start with an easy tia. Matt. This is not directed to you, directed to the others. So what does a successful AWS partnership look like from your C? So we’ll start with Eureka. [00:03:09] Rekha Thangellapalli: Sure. So from an ISV perspective, I think we really are looking at three things. [00:03:15] Rekha Thangellapalli: Uh, mutual investment building together. And scaling together. So when we talk about mutual investment, elastic recently signed a five-year SCA or strategic collaboration agreement with AWS. And while that is a significant milestone in our partnership, for us, what matters more is what it represents, and that is really a long-term commitment from both companies. [00:03:39] Rekha Thangellapalli: Towards product engineering, um, and joint go to market initiatives to deliver value to customers over time. And that’s what we see is that the best partnerships really compound and they build upon each other every year. Um, they don’t necessarily kind of reset every year. Um, next we talk about building together. [00:03:59] Rekha Thangellapalli: So, um. When we talk about joint solutions, we want to deliver solutions that are better together and the customers have to see us that way. And so whether it’s search, observability, or security, we’re looking at taking to market solutions that we can’t or necessarily don’t wanna take on our own. And finally we talk about scaling together. [00:04:22] Rekha Thangellapalli: And this is where marketplace, for instance, plays a big role, um, when customers can draw down on their cloud commitments, transact online and go from, you know, pilot to enterprise scale adoption in hours, not days. Um, this is when really everyone wins. Um, and this is also where partners like Accenture play a critical role. [00:04:47] Rekha Thangellapalli: Um, you know, the incredible amount of expertise that they bring, uh, the managed services capabilities and, um, their data assets actually play a huge role in having our customers realize that value faster. And, um, like Vince mentioned, at the end of the day, best partnerships are all all about creating kind of that. [00:05:07] Rekha Thangellapalli: Self-sustaining flywheel. And so it starts with investing together, building something unique, and having the customers realize that success faster because that success is really the only thing that’s gonna keep that flywheel going for everyone involved. I [00:05:26] Vince Menzione: absolutely. [00:05:26] Allison McFadden: Okay, amazing. I’m gonna riff off a few things Ika said, but from a GSI perspective. [00:05:32] Allison McFadden: A relationship with a WSA successful relationship with AWS looks slightly different. Um, so I think the first thing that we think of in the GSI Community common thread is that the client outcome and delivering value for clients is what we, what we’re striving for. Um, and so the partnership with AWS in that case, um, um, it has to, it has to. [00:06:01] Allison McFadden: Look like one team in front of our clients. So we have to show up indistinguishable, and that’s with AWS and with an ISV partner, it has to look like one solution in front of the client, especially moments that matter. So board meetings, um, you know, the time we’re gonna sign a deal, like we have to look like one team, uh, and keep our our client outcome, um, first and foremost in mind. [00:06:24] Allison McFadden: The second thing, and this is I think where the magic of all the people in this room comes into play. We can have as many discussions at a CEO level as we want. And if our client teams on the ground are not working together, it falls apart. Falls apart directly in front of the client. Yes. And that is a really hard thing to do. [00:06:45] Allison McFadden: So I’m passionate about the alliance work because that that work is what makes it happen at the corporate level. [00:06:53] James Kang: Cool. Um. I’ll start here. So in Nvidia is a accelerated computing platform company. Um, if you asked. Anyone on the, on the street about a year ago, what is ai? A lot of times they would say AI is, is open ai, or it’s philanthropic. [00:07:12] James Kang: Um, Jensen and I’ll, I’ll reference Jensen a lot today, um, because he is our leader, um, but he also sets the strategy in the direction for Nvidia. He talks a lot about AI in the metaphor of a five layer cake. And in terms of the five layer cake, you start off with the foundational bottom layer being power and energy, which sustains. [00:07:32] James Kang: All of our data centers, you move up the stack in terms of chips. So things think of Foxconn, think of TSMC. Next you have the infrastructure layer. So obvious choice is AWS, and then you get to the models where you do have the philanthropics and the open ais. But finally in at the precipice, you have the application layer. [00:07:53] James Kang: Ultimately, the reason why I mentioned all different stacks of the layers, the five layer cake, is the fact that the application layer is the most important. And so when you think about. Partners like Elastic or ServiceNow Trend, ai, CrowdStrike. Every time you pull from the application layer and you see a success, it pulls all five different components of that layer up. [00:08:13] James Kang: And so ultimately, as I think about success, it’s it’s being able to develop these co-sell wins at the application layer and really demonstrating that through extreme co-engineering and co-design with all the different application. Infrastructure, power and energy layers in mind. Um, Jensen also likes to think of himself not only as the CEO and founder, but also as the, the chief Marketing Officer. [00:08:35] James Kang: We are a very event driven company, and so at our big events like GTC or at big industry events like CES or Computex, he likes to show up on the biggest stage, biggest stages and showcase the partnerships with not only ISVs and GSIs, but also with end customers. And so that’s what I think about when I think of SA success. [00:08:56] Vince Menzione: That’s a really good point. You talked about, Allison, you talked about having an alliance strategy, or at least you teed it up, so I thought maybe we would go there for a second. Right? Like, what does a great alliance strategy look like and why is it important to the success of the partnership? [00:09:11] Allison McFadden: Man, I, uh, I have so many opinions on this. [00:09:13] Allison McFadden: We could probably be up here all day. That’s [00:09:15] Vince Menzione: okay. [00:09:16] Allison McFadden: Um, no, I think. Uh, there, there are a couple things, and the first one that comes to mind is focus. We cannot be all things to all people. Um, so when it comes to think about some of the, the work we’re doing with Elastic, we have a very, very clear point of view on what client problem we’re solving, what clients we want to talk to. [00:09:38] Allison McFadden: It helps if, um, from an ISV perspective, if there’s a very clear fit in. The Accenture portfolio or whatever, you know, SI consulting partner. You’re working with a very clear fit in the portfolio and we know what we’re not gonna go after, what we’re not gonna spend our time on because we have, we have this tendency, there’s millions of people. [00:10:00] Allison McFadden: The ecosystem chart that, you know, Vince, you showed up there, there’s so many connections. There’s probably more connections there than there are atoms in the universe, right? So, um. Defining what we do together and what we don’t do together is the first thing that pops to my mind. [00:10:19] Vince Menzione: Reka, do you have a perspective on it since we’re gonna, we’re gonna talk next about what you’ve done together, but, and I also wanna get mass perspective as a hyperscaler partner here as well. [00:10:29] Rekha Thangellapalli: Yeah, I mean from my perspective, I, I’m gonna, you know, kinda echo what Allison said is to be just maniacally focused. Yep. Um, because, especially from my perspective, so Elastic has three different solutions, right? We’ve got search, we’ve got observability, we’ve got security that map to completely different business units within Accenture. [00:10:47] Rekha Thangellapalli: And of course Accenture does a lot of things. And so, you know, when we first came together it was like. Okay, what are we gonna focus on? What industries are we gonna go after? Which segments are we gonna go after? Which customers, you know, um, outcomes are we trying to solve? And I think that sort of maniacal focus is the number one contributing factor to, to the fact that I’m like, up here on stage today. [00:11:12] Rekha Thangellapalli: Great. [00:11:14] Vince Menzione: Matt? Perspective? [00:11:16] Matt Yanchyshyn: Yeah, I, I, I guess I was trying to. To add something, uh, additional from an AWS perspective, uh, when it comes to, you know, what does a great alliance look like? Uh, AWS is obsessed with data, you know, in data we trust. And, and so the best, um, and, and this goes sales business problem, and it’s not just the engineering teams. [00:11:34] Matt Yanchyshyn: And so, uh, you know, Accenture does a good job of this elastic, definitely. And if you can come to the table with, um, quantifiable proof of the value of customer outcomes and partnerships. Um, you’ll win all the time and it’ll be a durable relationship with AWS ’cause we really are this data obsessed company and, and even the most senior sales leaders. [00:11:54] Matt Yanchyshyn: Uh, and so what I mean by that specifically is like if you, if you can show like your a RR to land an a RR conversion ratio, like in in numerical format, it’ll light up our sales leaders and, and they’ll be all, and they will co-sell with you all day long. If you can show the, I mentioned this earlier, like the AWS service, uh, whether you’re consulting company or, um, elastic and, and how the shape of customer accounts change positively when we work together. [00:12:15] Matt Yanchyshyn: That type of sort of quantifiable data works particularly well from an alliance perspective. With AWS as a partner, we, we really are like this data in sort of results out company. Um, so I, yeah, that’s just adding to the great points that were already made. I would say specific to AWS that that’s key. [00:12:30] Matt Yanchyshyn: Yeah. And I’m gonna bring up one more thing. I want to dive in on the, the joint value proposition, but you mentioned something that made a lot of sense and resonated to me about the organizations once you get out of partner, the partner world that we all know and love. Mm-hmm. Once you get down into a field organization or account management organization. [00:12:49] Matt Yanchyshyn: Not as much understanding and really organizations do a bad job here, honestly, in terms of enabling the field organizations. Do you agree? [00:12:58] Allison McFadden: I agree because I, I agree. And, um, you know, I think that’s one of the things, and, and I, I, when I joined Accenture, what we had was a lot of wicked smart architects delivering programs to clients in the field. [00:13:15] Allison McFadden: Very smart, very deep in AWS knowledge. Um, and that was awesome for the 10 clients they were staffed on and to get that understanding of how AWS works and I dream about lar, right? Like, this is a good, you know, but that takes real effort and real work. Yeah. And it’s, it’s um, almost like being a language translator. [00:13:37] Allison McFadden: Yes. For me. Yeah. So, you know, I had to deeply learn AWS so that I could. [00:13:42] Rekha Thangellapalli: Sure. [00:13:42] Allison McFadden: Teach my account teams. My account teams are really smart. They know who they’re selling to. They know their customers. They know what their customers need. They do not know what AWS has to offer always because they’ve got 20 partners lining up to try to tell their stories. [00:13:57] Allison McFadden: Um, they don’t know how to ask of the AWS team or the elastic team or the Nvidia team. Yeah. What they need [00:14:02] Vince Menzione: this co-selling piece. Yeah. [00:14:04] Allison McFadden: And so that is where, um. We had to build that muscle even around our AWS practice, which was a huge practice at Accenture, but we didn’t necessarily surround it with that kind of enablement and um, almost deal coaching layer. [00:14:21] Vince Menzione: So Elastic and Accenture came together. I dunno which one of you wants to lead this part of the conversation, but you will, right? Yeah. So tell us about the genesis of this and why. And a lot of people dunno what Elastic does, but you do some really incredible work. Like I, somebody told me one day was like, oh, you know, Uber, like, that’s elastic, powering all that. [00:14:41] Vince Menzione: Like, we don’t think about that. That the engines that you have and the, the backend to the customers, huge customers. [00:14:48] Rekha Thangellapalli: Yeah, absolutely. Um, so when AWS launched this feature last, um, reinvent where basically it allowed, you know, channel partners such as Accenture to be able to bundle up their services, their data assets with an ISV solution and put it on marketplace, um, you know, Accenture and Elastic immediately saw an opportunity. [00:15:09] Rekha Thangellapalli: Um, at the time most customers were doing gen ai. But they were running into the same challenge, which was that their data just was not ready. And by the way, this is a problem we were solving. Outside of marketplace. I think the, the feature that you guys launched just gave us a way to package it up and to be able to create this repeatable solution, which we call data readiness engine for gen ai and put it on marketplace. [00:15:40] Rekha Thangellapalli: And, um, this to me was a success because. Each company had a clear reason to invest. Um, so for Accenture, they were able to, you know, create a very differentiated services led offering. Uh, for Elastic, we were able to expand on our AI story. And for AWS, um, you know, it drives marketplace adoption, increases cloud consumption, all of that great stuff. [00:16:07] Rekha Thangellapalli: And customers, of course get. A solution to a very real problem that, that they were having. Um, and you know, the surprising part for me going through that journey was that, um. The pitching, the idea, getting the budget, getting the executive sponsorship was actually the easy part. The hard part was getting all three companies to come together, uh, to go from idea to launch in a very ambitious timeline of six weeks. [00:16:37] Rekha Thangellapalli: Nice. And so, you know, this was very much like. Doesn’t matter your title. We’re rolling up our sleeves and we are on this outcome together. Um, and so we literally built a RACI matrix, a project plan, and you know, we had daily standup calls for six weeks where literally. At least one person from each three of these companies called in, you know, got rid of any blockers and we made sure we were on target for that timeline. [00:17:07] Rekha Thangellapalli: Um, and you know, at the end we had a successful launch. But I think my favorite part about the story is the impact that we’re having and, um. My favorite story comes from a global pharmaceutical company that, you know, had basically nine petabytes of data spread across six different continents. Wow. And by working with Accenture and Elastic, they were able to build that trusted foundation that their AI and their agents can, you know, kind of safely tap into and be accessible at scale. [00:17:41] Rekha Thangellapalli: Um, so that’s my version. Allison. [00:17:44] Allison McFadden: Yeah. Well, I don’t have a lot to add. I just, I would say this is a good example of a couple of principles, right? One is having a forcing function is never a bad idea. Sign up for a big event, sign up. I’m like, I’m here with my, you know, Nvidia guys saying, sign up for the event. [00:17:58] Allison McFadden: It’ll make you move quick, right? [00:18:00] Audience Member: Yes. [00:18:00] Allison McFadden: Um, so that is one, but two, one of my mentors once told me, when you’re designing any kind of, you know, offering go to market motion, it has to get blood to all organs. If it does not get blood to all organs, it does not go [00:18:14] Vince Menzione: nice. [00:18:14] Allison McFadden: Um, [00:18:14] Vince Menzione: I love that analogy. [00:18:15] Allison McFadden: Oh, I love it. And I can talk all day. [00:18:17] Allison McFadden: That guy was brilliant. I love him. But, um, no, and, and so Elastic did a really nice job of bringing the tech to the table. Um, our team has to trust in that technology and its ability to scale, right? Um, because at Accenture we have to be able to deploy across 700,000 consultants. Um. And yeah, so I think those are the two, two things that really worked well here is we had, uh, trust in the technology solved a customer need. [00:18:50] Allison McFadden: Um, it drives, we don’t even talk about, like, yes, it drives marketplace revenue, but it unlocks work that we do that drives even more revenue to our AWS Friends. Right. So this is a, this is a, um, product that’s getting your data ready for AG agentic. It’s a messy problem that everyone’s dealing with, and it removes blockers for clients and it unlocks more, you know, ag agentic work on top of that. [00:19:15] Allison McFadden: So, blood to all organs. [00:19:17] Vince Menzione: So, was that the proposal going forward to say we need to have, we need to have trust in the solution. We need to drive significant revenue. It needs to be something all of our, you know, seven, 700,000 people. Can be a part of and help drive? Is that how you think about? [00:19:32] Allison McFadden: Yeah, and for us right now, um, it’s an interesting time for Accenture. [00:19:36] Allison McFadden: Our clients are asking a lot of us, and what it does is it having some of these accelerators helps us deliver cheaper, better, faster to our clients, which is what they’re demanding of us right now. Um, so it’s an accelerator to client outcomes. [00:19:55] Vince Menzione: James, what is NVIDIA’s role and how do, how do you enter the equation here? [00:20:00] James Kang: Yeah, it’s, um, it’s a good question. Um, I, I would say that Nvidia is probably one of the most misunderstood organizations in the world. Um, despite the, uh, the market capitalization in the valuation of the company, we have a very tiny organization. Um, what I mean by that is, um, if you think about. [00:20:20] James Kang: Salesforces and field sales organizations. Um, we’ll take Salesforce as the account or the customer. As an example, we have one account manager at NVIDIA that no, not only covers and is responsible for the relationship with Salesforce, um, but also manages. Automation Anywhere as well as DocuSign. Whereas at AWS, in contrast, like there are full armies and teams Yeah. [00:20:45] James Kang: That are supporting the Salesforce relationship. And so as you think about partnering and working with Nvidia, the focus has to be on really. Extreme co-design, but also being very prescriptive in terms of what are the very specific customer outcomes that we are solving for. And the guidance that I would give is bring in Nvidia into that equation and that conversation as early as possible because that [00:21:10] James Kang: co-engineering and co-design needs to be part of the foundational building blocks in order for you to come out with a end solution that checks all those different requirements. [00:21:20] James Kang: And so I think. Again, like going back to Nvidia, um, we like to talk about two different types of brains. A brain one and a brain two. Uh, brain One you think about the next quarter and making sure that you’re hitting the revenue targets for the next quarter. Brain two, you think about a long-term goals and potentials looking around corners and being very strategic. [00:21:41] James Kang: The saying internally is without Brain one, there is no oxygen, but without brain two, there is no future. And everyone at NVIDIA is trained to think in that brain two mentality. [00:21:52] Vince Menzione: Wow, Matt. [00:21:54] Matt Yanchyshyn: Yeah, I, I was just thinking I love the blood doll organs. Uh, and so just on, on that note, um, and, and, you know, the multi-product solutions that, that you, you built together, uh, that is a really good example of blood do organs because like we all know, that’s how customers buy. [00:22:07] Matt Yanchyshyn: They, they buy solutions and increasingly they’re looking for combinations of ISV, sometimes multiple products from multiple ISVs with services. Uh, often they’re buying it through a resell motion. You know, and they, and, and so that from a customer perspective, they want a single place to go. And so that’s the multi-product solution. [00:22:24] Matt Yanchyshyn: They wanna find everything they need, they need Accenture, they need Elastic to solve a specific solution. And I think where that’s headed is even more specific listings, like with AI powered listing experience, like, you know, elastic Plus Accenture for, I’ll make something up like a manufacturing workload. [00:22:37] Matt Yanchyshyn: And so this solution based. Uh, sort of buying is, is very customer centric. It’s what customers want. We all know that. But that’s, that’s the customer sort of organ, I guess. Um, but then, you know, you all have SCAs and those SCAs have marketplace commits. It helps if that gets transacted through marketplace helps the AWS relationship, you know that that’s an organ. [00:22:55] Matt Yanchyshyn: It’s the relationship. It’s, it’s the commercial construct and that you have, uh, that that’s another organ. You’re marketing people. They, that’s another organ. They don’t wanna land, uh, leads on a static marketing page. They wanna land a lead on a, a storefront with a multi-product solution that can actually convert and that you can actually buy it through that. [00:23:12] Matt Yanchyshyn: So the marketing person’s happy because they, they have less churn. Uh, and then, you know, our reps are happy ’cause guess how they get paid? They retire quota when they sell Marketplace. And they, we also, Jay McMain will tell you, that’s another organ called Jay or on, on you now. Um, [00:23:27] Matt Yanchyshyn: he’ll like that. I’ll call him up and tell him that. [00:23:29] Matt Yanchyshyn: Yeah, [00:23:30] Matt Yanchyshyn: but he, he’ll tell you, you know, don’t believe me. Obviously, never believe Matt, believe, believe the, the data and, and his data shows that. Those deals will close faster and larger if you use marketplace. So that’s, that’s a lot of organs. That’s the whole body. Um, but you know, when you have your customer happy ’cause that’s how they wanna buy your field happy. [00:23:45] Matt Yanchyshyn: Um, and, you know, the relationship happy and you know, your marketing team happy. Uh, and, and Jay happy. Um, and, and you know, I think that multi-product construct and, and the way you kind of use it to model a partnership and the way buyers ultimately wanna buy is, is really powerful. And so I, I think it’s, you know, it’s really a manifestation of how. [00:24:04] Matt Yanchyshyn: We kind of intend and to go to market anyway. Uh, so I think, you know, and thanks for leading the way, by the way. You’re, you’re amongst the very first, so that’s great to see. [00:24:11] Matt Yanchyshyn: So these storefronts are really helping this drive, drive this. Well, [00:24:13] Matt Yanchyshyn: that’s the next evolution. Like we’re talking about the multiproduct solution. [00:24:16] Allison McFadden: I’m JJ Accenture storefront. [00:24:17] Vince Menzione: Yeah. Oh, there you go. I mean, j and j Accenture storefront. [00:24:20] Allison McFadden: We’re gonna talk about that. [00:24:20] Matt Yanchyshyn: Yeah. I mean, [00:24:21] Matt Yanchyshyn: Accenture also leading the way yet again with storefronts. And so I think the combination of. You know, again, I was talking a lot about conversion. Yeah. And you know, buyers know sometimes they know what they wanna buy and, but if you really wanna convert that lead, you wanna land them again, something that combines, you know, elastic Accenture’s services plus software, but in a storefront that is, you know, surrounding with just the solutions they want so they don’t need to kind of go searching. [00:24:42] Matt Yanchyshyn: So, you know, ultimately reducing that time to close, I guess, really ’cause meeting the customer where they are with what they need. [00:24:51] Matt Yanchyshyn: So we talk about co-selling a little bit. We, Jay and I talk about this all the time. We gotta keep looping Jay in here, even though he is not even in town this week, but Reko, um, what does co-sell look like inside Elastic? [00:25:02] Matt Yanchyshyn: You’ve got, we talked about an incredible leadership team. I’ve gotten meet some of your leaders. Seems like you drive, you do a good job internally driving that. Let’s talk a little bit about it. [00:25:11] Rekha Thangellapalli: Yeah, and this is something I’m, I’m personally very passionate about. Um, co-sell is. Very much a journey, not a destination. [00:25:20] Rekha Thangellapalli: And I think step one for us is recognizing the different partner types that we have. Because at Elastic we work with, you know, OEMs, MSPs, resale distributors, GSIs, um, and they all bring something very unique. To the customer lifecycle and they all contribute very differently within, you know, our own sales cycle and sales process. [00:25:45] Rekha Thangellapalli: And so, you know, figuring out what is the unique benefit they bring, how do we enable them? So training and enablement is a huge piece of it, and so is making sure we’ve got the right metrics to measure success. Um, I know a lot of companies look at partner sourced as the north star, and that’s great, right? [00:26:06] Rekha Thangellapalli: Because that is undeniable. You can say, Hey, that would not exist if it wasn’t for my partner team. Um, but we’ve also noticed that when we bring in GSIs, it actually increases renewal rates. It significantly increases. Um, a RR over time. Um, it expands deal sizes and so these are very real metrics that we can point to, um, beyond just the co-sell and the partner sourced number. [00:26:32] Rekha Thangellapalli: Um, so for us it’s looking at it from a very holistic perspective, but also catering it towards that unique partner and making sure we’re doing everything we can to set them up for success and setting up the partnership for success. [00:26:47] Vince Menzione: So clo close win ratios, deal size and renewal rates? [00:26:52] Rekha Thangellapalli: Yes. For specifically for geos size. [00:26:54] Rekha Thangellapalli: Yeah. [00:26:55] Vince Menzione: Very interesting. Allison, uh, what had to change internally to produce these co-selling? We talked a little bit about the field organization and enabling a, a group of, and, you know, account sellers that are very customer focused and enabling them on the co-sell side. What had to change internally to drive that? [00:27:13] Vince Menzione: Yeah. [00:27:14] Allison McFadden: I, I might have already alluded to this a little bit in a previous answer, but, um, creating the capacity to develop, build, and sell these solutions, um, inside of a large GSI, where billable hours is kind of the number one metric on the table. Um. Is part of the investment that we had to make within Accenture to get this done? [00:27:36] Audience Member: Yeah, [00:27:36] Allison McFadden: so expert technology time. So we have technologists that understand the elastic technology. We do similar with Nvidia, by the way, we. We released some of their time to go co-develop the solution because it has to hold technical water, right? It can’t just be a marketing pitch. It can’t just be, it has to be a real, um, what’s the there, there. [00:27:59] Allison McFadden: So in order to actually do proper co-sell, we had to release some of that time. Um, to invest in those partnerships. Um, we’ve also done similar with some industry aligned business development leaders recently, so we have freed their time up to go. Uh. Open new conversations, educate client, account teams, go to clients, have conversations. [00:28:26] Allison McFadden: Um, so that, that’s a new motion that we, uh, have just kind of recently made, um, to allow them, I love this brain one, brain two also, right? So to allow them to focus on brain two, because a lot of our time. Typically spent delivery issues, you know, getting my hours, where am I charging my time? And so just freeing up a little of that capacity to do this work, um, helps get us in this brain two mode where we’re not just living to survive. [00:28:56] Vince Menzione: I. So, Matt, you’ve removed a lot. I mean, one of the things I admire, I admire AWS for being first to market and removing the most friction in marketplace of any of the vendors. Really, truly that. You talked about some of the announcements. How does some of, how does some of this tie PC central agents propensity sales plays, MCP, how does some of this tie to how, how you’re thinking about the future? [00:29:18] Vince Menzione: And how to enable more motions like this. [00:29:20] Matt Yanchyshyn: Yeah. Well, I, I think if you know my boss, UBA Borno, uh, you’ll know that she has a maniacal focus on automation. Yeah. Um, and, uh, co-sell is increasingly automated. You know, you were asking earlier about propensity data. You can get that propensity data in addition to sales plays and, uh, opportunity scores through the partner central agents. [00:29:38] Matt Yanchyshyn: So things that used to require multiple calls to A PDM, if you’re lucky to have one. Yeah. Or a p sm. Uh, you, you can now get through, through these agents, you know, uh, tech Systems, TGS, they, they manage what, over 5,500 customer opportunities with agents that they built on top of our partner Central APIs. [00:29:55] Matt Yanchyshyn: Um, and work Span has built a whole product and business that’s right on leveraging, uh, our APIs, our capabilities to sort of tie into your CRM. So, majority of all opportunities will be progressed and managed by agents. This year at AWS, we already have a majority of all customer opportunities, all app have a partner attached and I, I took a personal goal for a majority of those partner attachments, not to happen from a human. [00:30:22] Matt Yanchyshyn: But from our solution matching engine. And how do you get recommended by that solution? Matching engine, having a healthy ACE pipeline, thanks to partner central agents and the integrations you’re doing. And in addition to being the specializations and doing things like multi-product solutions and ultimately closing opportunities, you dream of LAR and so LAR will help that. [00:30:40] Allison McFadden: It’s more like a nightmare. [00:30:41] Vince Menzione: And so, you know, [00:30:42] Allison McFadden: it’s more like a nightmare, but [00:30:44] Vince Menzione: nightmare. Well, it’s, it’s, yeah. Nightmare of Laura and, and. Nice dreams of PRM, but the, um, but that’s the loop, right? I, I think, uh, increasingly co-sell for us, and in my mind, is largely a hundred percent automated. Yeah. Except for what matters most, those most largest, most strategic, most complex deals. [00:31:01] Vince Menzione: Where our highly paid and very skilled salespeople are most effectively used. [00:31:05] Vince Menzione: Yeah. [00:31:05] Vince Menzione: You know, the days of, you know, this person with 20 years experience selling, clicking, progressing opportunities through a pipeline, uh, should be over. Uh, and, and we need those people out, out selling and, and co-selling. And so that for me. [00:31:19] Vince Menzione: Yeah. That, you know, we talk a lot about co-sell, but I, I’m obsessed with automating as much of the co-sell as possible. [00:31:24] Vince Menzione: I remember going back to the ex Excel spreadsheets and, and that, that seems to be be Viva became spreadsheet jockeys. [00:31:31] Vince Menzione: Yeah. [00:31:32] Vince Menzione: And, and they stopped selling. They forgot how to sell. [00:31:34] Vince Menzione: Yeah. And people spend all this time doing lunch and learns and things like that. [00:31:36] Vince Menzione: And then, you know. Then the salespeople rotate out after 18 months and, and it, that’s, that’s the old days. Uh, you know, the new days are, are AI powered matching algorithms, uh, ag agentic co-sell, using the partner essential agents to get your data and, and putting that data to use automatically and, and what sounded like magic. [00:31:51] Vince Menzione: 12 months ago is being done, you know, by partners at massive scale across thousands of opportunities. You can do it today. And you know, I, there’s a guy named another Mike, right? Mike another Mike who they have, there’s like a guy who’s doing all this and I’m picking on Mike ’cause I, I know their system really well and I know the guy Mike grew easily built it for them. [00:32:08] Vince Menzione: Um, but, you know, I think, yeah, again, in the days of having 10 people sort of doing lunch and learn could be replaced by one or two people, building agents, uh, managing a massive pipeline. And, and that’s the future. [00:32:18] Vince Menzione: Exactly. James, your perspective on what breaks with co-selling? [00:32:22] James Kang: Oh, what breaks co-sell? Um, I would say. [00:32:25] James Kang: It, it starts and finishes with just misalignment and a loss of trust with the customer, especially when you have multiple partners or stakeholders involved. If you’re trying to do a three-way deal with a end customer and you’re not on the same page, you’re not gonna get to a successful outcome on, on the backend. [00:32:44] James Kang: Uh, the fix is a much more complicated story. I would say that to take a step back, um. We’ve talked about the five layer cake. We’ve talked about where NVIDIA kind of fits within the equation. We are invested in the ecosystem and so as different players and application organizations win and see these outcomes for end customers, we celebrate that success. [00:33:07] James Kang: Um, and as part of that kind of ethos of where NVIDIA fits within the ecosystem, we wanna make sure that not only. Our customers, but our partners like ISVs and GSIs are set up for success. Um, we do not as Nvidia sell hardware or GPUs directly to customers We use. Hyperscalers like AWS as kind of our force multiplier. [00:33:31] James Kang: And similarly we think of ISVs and GSIs as the force multipliers in terms of our extensions of how we, we kind of leverage the relationships and build the trust with our end customers. And so going back to kind of the question, Vince, I would say that it all comes back to trust and being able to build that mutual trust. [00:33:48] James Kang: Um, a lot of what we do when we co-sell with AWS is really on the software layer. Um, we actually have more software engineers at NVIDIA than we have hardware engineers, which is a weird thing to say, um, because everyone knows us for our GPUs. But because of that fact, we are heavily invested in Cuda and making sure that Cuda becomes the foundational layer for how not only our ISVs and GSIs, but also our end customers are building. [00:34:12] Vince Menzione: Very cool. So Reiki, you and James together on this production. Versus pilot with the Gentech ai. Tell us a little bit more about that. Where, where are you in the process? [00:34:24] Rekha Thangellapalli: Yeah. So I mean, in general, what we’re seeing out in the market in, in relation to sort of AI and, and customer’s journeys is that, um, at least from an elastic perspective, um, we’re seeing people very much in production when it comes to, you know, kind of AI assistant co-pilot use cases. [00:34:42] Rekha Thangellapalli: So, you know, things like, um, software development, customer support is a big one. Um, any sort of employee productivity use cases where there’s. Still a human in the loop somewhere. Um, and there’s a very like, clear path to value. And so we see the customers being in production excelling there. Um, no problem. [00:35:01] Rekha Thangellapalli: Where we’re seeing people still kind of in the pilot phase is those fully autonomous workflows where there is no human involved. The agent is reasoning on its own. Um, accessing multiple systems and taking an action on the user’s behalf. And what we’re seeing is that it’s not the intelligence of the agent that’s holding it back. [00:35:26] Rekha Thangellapalli: It’s more about giving the right context to the agent and having the right. Security kind of governance controls in place for the company to feel comfortable in putting these fully autonomous workflows into production. And that’s really the conversation we’re having is all right, what are the controls you need in place? [00:35:47] Rekha Thangellapalli: For you to release this to your business unit. Um, and what is the context that the agent is needed before we can comfortably let the agent make the decision on the user’s behalf? Um, James, I’d be interested to hear what you’re, what you’re seeing in the market [00:36:03] James Kang: plus one on all things context. I, I would even go so far as to say, um. [00:36:09] James Kang: H how many folks in the audience have heard of token maxing? Like this new term? [00:36:13] Rekha Thangellapalli: Yeah. Yeah. [00:36:14] James Kang: Um, I’ll, I’ll give a very specific example of, of Uber that went public. With the example of Claude, like they allowed all of their employees to use as many tokens as possible, and within the span of four months, they exhausted their full budget for the year, and so they had to pull back, and now there’s a cap on every employee. [00:36:33] James Kang: I think the number that’s circulating is $1,500 per month per employee, and so I think that is at least. In this multi-phase evolution of where we’re going to be and where we’re today, cost has become kind of the prohibitive force in terms of agentic AI at scale. Um, I think we are working on some very creative solutions in-house and Nvidia. [00:36:55] James Kang: Um. And we saw some really dynamic announcements this week when it comes to all things agent core, um, where we want to focus on very nimble ways for customers to be able to execute and go to market. And one extreme example of that is our investment within our open model strategy. So Nvidia, not only, again, providing GPUs, we actually offer our own op open models, which we call our Nitron models. [00:37:21] James Kang: And through our Nitron models, we are allowing customers to really develop and fine tune their own proprietary models in a cost effective manner. So right alongside the frontier models like OpenAI and Anthropic. It’s not a if then, it’s not an either or statement. It’s a, it’s a permutation, it’s an and So we’re giving you a cost effective alternative to not only bring your AgTech applications at scale by training on Nibo tron, which is open source, but then once you’ve kind of finished and fine tuned that specific training job to be able to. [00:37:53] James Kang: Go ahead and utilize your frontier models, whether it be OpenAI or Claude. And I know there’s other partners here that are providing those kind of different model capabilities. And so I think for us it’s, it’s a matter of choice. We know that this market is dynamic. It’s gonna be evolving over the next coming months as well as the next coming years. [00:38:10] James Kang: Uh, but we believe that we are positioned for a really unique dynamic expansion of AgTech use cases over the, at least the next three to six months. [00:38:20] Vince Menzione: Allison, for the partners in the room who are glazed over right now going, what do I, what do I do over the next 12 months? [00:38:26] Allison McFadden: Should I wake everybody up by saying, yeah, please. [00:38:27] Allison McFadden: Say go hurricanes. [00:38:28] Vince Menzione: Yes. [00:38:29] Allison McFadden: Is there anyone, anybody? Everyone’s like, boo. I get to leave the parade today to go home to parade. I live in Raleigh, so we’ve got our parade on Saturday. Nice. [00:38:39] Vince Menzione: Nice. [00:38:40] Allison McFadden: All right. Wake up. Um, all right. So for the $50 million partners in the room, um. $50 million is not small. You have something that works. [00:38:50] Allison McFadden: Right. This is great. What I would be thinking about is, you know, we’ve talked about focus before, but really doubling down on, you know, what is, what is your industry, what is your client like, ideal client that you serve. And build, um, almost that kind of community. You know, the, the clients we have move from firm to firm to firm. [00:39:17] Allison McFadden: And if you’ve done good work at one, you’re gonna follow ’em to the next. Um, so build that client demand in a specific place or specific client profile that is just like really knocking it out out of the park for you. Um. Scale with marketplace, right? So if you, I, I love some of the data that you were sharing in your talk earlier, um, because it’s like no overhead scaling mechanism. [00:39:45] Allison McFadden: I mean, it’s, it’s fantastic. Um, Accenture, other GSIs like us, we are investing in marketplace. So we’re investing in resources, um, to help us. Use marketplace more with our clients and we’re gonna capture, right, those storefronts. And if you’re present on marketplace, you’re gonna be able to catch, uh, yourself in that wheel. [00:40:09] Allison McFadden: So I think those are the, the kind of couple of things I would say is focus, focus, focus to drive that client demand and use scaling mechanisms like marketplace to really kind of, uh, accelerate. [00:40:24] Vince Menzione: Matt, anything to add there on the. [00:40:26] Vince Menzione: Well just, you know, Ja, James, you, I love the token maxing reference in Uber and it reminds me, you remember when cloud came out and everyone was like, oh, all these people are, are gonna use the cloud and costs are outta control and. [00:40:39] Vince Menzione: Um, a lot of people pulled back from the cloud and, and a lot of those companies no longer exist. And it’s similar with, with, uh, token maxing, like, oh, these agents are outta control. You have a choice. You can embrace them and figure it out and get governance and, and make your data available. Um, use the partner, central agent, move to agent to co-sell, or you can fade and die. [00:40:58] Vince Menzione: And, and that’s, that’s where we’re at. Uh, is, is the, the companies sitting here today embraced the cloud years ago and won. Uh, and and there’s a set of companies here today who are gonna embrace agents in the, for both buyers and sellers, and will win. And there are those who won’t and they won’t win. And so for me, it’s like we’re, we’re at a, we’re at a crossroads. [00:41:18] Vince Menzione: And, and if you’re gonna win, you gotta leap into that, you know? I love it. And, uh, and, and, and it’s, it means the cost of experimentation is so much lower now. Development and, and even business development or software development is, is agent enabled. And so you can take risks, you can experiment and, and you have to, it’s, it’s an existential moment. [00:41:37] Vince Menzione: Agreed. We’ve got a couple minutes left over for any questions. What do you think? Sure. Are there any here. I think there are a couple. Yeah, we’ve got, we’ve got a co-sell question I’m sure coming up here. [00:41:51] Audience Member: Um, I’m Cassandra, I’m the CEO of Partner Tap. And one of the questions I had was, I think, you know, the co-selling between the sellers is where things get. Really, really hard when you’re multi-partner. And so when I was listening, um, with, you know, the Accenture and Elastic together, you talked about how you had, you, you had to get these BD business development people. [00:42:22] Audience Member: Um, is this a new team that is over the client team? And how do these teams interact like with the elastic sellers? Are you doing a lot of coaching to the field and then with if AWS sellers are, are involved, like what is that whole picture? What does look like, [00:42:43] Allison McFadden: like [00:42:44] Audience Member: on the ground? I mean, that is the hardest part, I think, and that’s what we hear. [00:42:48] Allison McFadden: It’s so, it’s so, it’s so tough. Um, and I will, I’ll just say, so our business development leaders that we now have kind of. Expanded their capacity. They have always been, they have always been there. Um, but they have not been well resourced. They haven’t, they haven’t had very clear kind of job description. [00:43:12] Allison McFadden: I’m gonna say I, in the past they have been kind of focused on partner relationship. And so like more like an alliance manager and maybe working on some of the data. Right? So when I say I have nightmares about Lars, because we’re always trying to increase the LAR for Accenture and, and they were focused like in those detailed weeds of like trying to pass ACE and trying to call the PDM and all this stuff. [00:43:39] Allison McFadden: What we are doing is really pivoting them to be proper sales, business development focused on client outcomes and focused on. Technical skills to be able to describe what this solution is to the field. So, um, and because we need, I have many, many questions about, I gotta get agents to work with Eurogen co-sell so that that part somehow goes away. [00:44:05] Allison McFadden: So that’s a, that’s the thing we gotta solve still, but, um, so we’re pivoting them to be kind of driving. More of that co-sell enablement with the field, um, and taking that message to the field rather than being there, waiting for questions to come in from the field, waiting for like our field teams to discover, oh, I saw something that we’re doing with Elastic, like on a press release on LinkedIn. [00:44:30] Allison McFadden: Right. So we’re kind of trying to pivot them to be more proactive. [00:44:33] Vince Menzione: Very cool. [00:44:34] Rekha Thangellapalli: Yeah. And uh, Cassandra, that’s an excellent question because I think. Multi-party, you know, sort of tri-party offerings. The hardest part is operationalizing it at scale, right? Yeah. And so for this particular offering, we are basically having three routes to market. [00:44:51] Rekha Thangellapalli: So one is seeing how this offering fits into our existing elastic go to market. And so I am constantly enabling our field sellers to say, okay, within our three field sales place, here’s exactly where this fits in. Here are, you know, uh. Keywords that you hear in customer conversations where you bring up this offering and here’s a process of how it works. [00:45:14] Rekha Thangellapalli: Um, exactly At what sales stage do I bring in Accenture, how, you know, what are the roles and expectations? Right? So that’s on the elastic side. We’re doing the same thing on the Accenture side. So we’re doing a ton of training enablement and lunch and learns, and we’re also looking at how do we fit into. [00:45:31] Rekha Thangellapalli: Uh, Accenture’s AI transformation projects, we are the semantic layer, right, of their enterprise brain. And so it’s a whole different sales motion, um, and, you know, having the right assets, having the right process again to make sure that that goes smoothly. And then finally, we’re going directly to the customer. [00:45:49] Rekha Thangellapalli: So we are launching multiple external campaigns where, you know, if the customer raises their hand. We will, we will line up immediately. Right. Um, and so, [00:46:01] Allison McFadden: I mean, I can’t, I can’t, I can’t say how important that third leg of the stool is. ’cause the second part, she talked about getting into our catalog is the first thing. [00:46:09] Allison McFadden: ’cause my BU business development leaders have the catalog. Right. And that’s what they’re selling. So what Elastic has done has gotten into one of those offerings and then. If we have a customer that asks for it, that is the fastest way to alignment. That is like the number one thing that we respond to [00:46:26] Vince Menzione: customer at the center. [00:46:27] Vince Menzione: This is great. Well, I think we’re up to time. This was a great session. I want to thank you. This is what a great, what a great group. [00:46:34] Vince Menzione: Thanks for listening to the Ultimate Partner Podcast. If today’s conversation resonated, share it with a partner leader in your network. Subscribe where [00:46:43] Vince Menzione: you listen, and head over to the ultimate partner.com. [00:46:47] Vince Menzione: For show notes related content and the resources for this episode. And if you haven’t already, now’s the time to register for the Ultimate Partner Live Event in Reston, Virginia, October 26th through October 28th. Until next time, keep showing up in the rooms that matter because being in the room changes everything [00:47:09] I.
This week's video transcript summary is here. You can click on any bulleted section to see the actual transcript. Thanks to Granola for its software.There was an issue with this only going to paid subscribers, so sending it again. Apologies to those who get it twice. I appreciate being paid so feel free to upgrade if you enjoy TWTW.EditorialIntelligence: Who Owns it?This week the word “AI” feels too small.AI is a technology. Intelligence is its product. And if intelligence is the product, the question is no longer just: Which model is best? Who has the cheapest tokens? Who owns the weights? Who controls the data center? Those are important questions, but they are lower in the stack.The bigger question is simpler and more political:Who owns intelligence?That sounds abstract until you make it concrete. Intelligence is becoming something companies can capture, package, serve, meter, route, improve, and sell.It can write code, answer questions, design molecules, automate offices, run agents, draft legal work, advise scientists, serve consumers, and reshape workflows. It is not merely software. It is a general-purpose capability. And all humans could benefit from more of it.General-purpose capabilities have a habit of becoming public questions. But the default answer, that public good is best delivered by government, is the wrong answer in this context.The Product Is IntelligenceWe should stop talking about AI as a feature and start talking about intelligence as the universal thing that is delivered as an input to the world.Water is an input. Electricity is an input. Literacy is an input. Connectivity is an input. Once a society depends on them, access stops being optional. Nobody needs government to build every well, power plant, school, or network. But everybody understands that a civilization cannot be organized around less than universal and reliable access to foundational inputs.Intelligence is reaching that level of importance now that we all know it is real.Government should not own it, operate it, or develop it. Quite the opposite. Companies are the right actors to build fast, compete hard, improve models, serve customers, and discover the real use cases. Self-interest is a useful framing here. Markets are good at finding demand, reducing costs, and turning invention into services people actually use.Companies are the right operators, developers, and owners. But that does not settle the real question of who owns the benefits. That is an economic question.If intelligence becomes metered infrastructure, what happens to the value it creates?The Ownership StackThis week's articles keep circling the same issue from different directions but in the nature of ‘circling' never quite nail it.Jamin Ball's “Own Your Weights” starts with the enterprise version of the question. Owning a model file is not enough. The durable asset is the loop: the data flywheel, the evaluations, the reinforcement system, the workflow learning, and the operating context that lets capability compound.Benedict Evans' “Ways to Think About Token Pricing” adds the market layer. Tokens may become essential, abundant, and cheap, like mobile data. But being essential does not guarantee that the token layer captures the value. The money may move up the stack to whoever owns the workflow, the customer, the distribution, or the application.Alex Karp's fight with the labs, reported in “Alex Karp Is Saying What Every Angry CEO Is Thinking About AI”, is the same argument in sharper enterprise language. Companies are afraid that model providers will not just sell intelligence, but learn from customer workflows and then move into the markets where those workflows create value. The “All-in” group are echoing Karp's view.And “What Is Loop Engineering, and Who Owns It?” names the new contested terrain. The loop is where intelligence meets the world. Whoever owns the loop owns the learning. Whoever owns the learning owns the compounding asset.That is why “who owns intelligence?” is not a slogan. It is the question under the model layer, the application layer, the enterprise layer, and the economic layer.Because intelligence is the product, the tools creating it are fragmented and competitive. So there is no logic in trying to discuss this at the level of a single company or set of tools and models.The Old Promise Was That Commerce Would Tame PowerThe essays this week give the historical backdrop.Deirdre McCloskey, in “What Really Caused the Industrial Revolution”, argues that modern growth came not simply from capital accumulation, but from a change in permission: ordinary people were allowed to innovate, trade, build, and be honored for it.That matters because intelligence could be another expansion of permission. It could make more people capable of building, learning, creating, coding, researching, translating, selling, and coordinating. It could lower the cost of competence.But only if access is broad.Paul Krugman's “AI in an Age of Oligarchy” warns that the same technology lands differently in different political economies. A new general-purpose technology entering a broad, open, upwardly mobile society is one thing. The same technology entering a concentrated economy, with extreme wealth and weak counterweights, is another.Tim O'Reilly's Economist essay, “Elon Musk is building a form of capitalism that Adam Smith would hate”, makes the governance point more directly. The old liberal hope was that commerce would tame arbitrary power. Markets, boards, courts, shareholders, disclosure, and competition would discipline the prince.But what if the prince uses markets to escape discipline?Henry Farrell's “political economy of billionaire derangement” pushes the same point. Founder culture, monopoly ambition, peer rivalry, weak correction mechanisms, and vast private control can amplify appetites rather than restrain them.The danger with intelligence is not that companies build it. They should. Companies build it, meter it, use public tolerance and public infrastructure to scale it, learn from everyone who uses it. All of those things are inevitable and healthy. Market forces will sort out winners from losers. The real danger is that the winners treat all of the surplus produced as purely private.Metered Intelligence Creates SurplusIf metering is not the problem, what is?The problem is pretending that metered intelligence creates value only for the metering entity. Metering water is only tolerated as a public good. If the public were blackmailed by a private water company with the threat of no water we would all rebel.Once we understand that the product of AI is intelligence we can see that every time intelligence is used, there is the immediate transaction: the user pays, the provider serves.But there is also system value. Usage creates signals. Workflows reveal patterns. Prompts, corrections, failures, preferences, integrations, edge cases, and business processes all help define where intelligence is useful and how it should improve. Intelligence breeds intelligence.Even when customer data is contractually protected, the market learns. The platform learns where demand is. The product team learns which workflows matter. The ecosystem learns which jobs are vulnerable, which tasks are automatable, and which parts of the economy can be reorganized around machine intelligence.So the surplus is not born in a vacuum.It rests on public science, public education, public data exhaust, public law, public infrastructure, public energy systems, public tolerance for data centers, and billions of human interactions. It is served by companies, but it is not made only by companies.This is why “Americans Deserve a Dividend From AI Companies' Riches” belongs at the center of this week's issue. The detail can be debated. The principle is harder to dismiss. If intelligence becomes a new foundational resource, then some part of the wealth it creates should flow back to the people whose society makes it possible. Intelligence did not suddenly appear. AI is built on the entire history of human intelligence. It benefits from it and at the same time evolves it.Not Nationalization. A Human Wealth Fund.If intelligence belongs to everybody, some conclude that government ownership of intelligence is the right outcome.Governments are not well suited to build, operate, or improve intelligence. They will move too slowly, regulate too early, politicize the wrong things, and confuse economic participation with operational control.Andrew McAfee's “Why I Didn't Sign the AI Open Letter” is useful here. His objection is not that the technology is unimportant. It is that steering too hard before we understand the shape of the change can become its own failure mode. Marc Andreessen's satire of AI regulation is less policy than temperament, but it captures a real Silicon Valley fear: that regulation can become permission, capture, and incumbency before it becomes wisdom.That fear should be taken seriously.But it does not answer the economic question. It answers only the operational one.How can the economic benefits of intelligence be distributed? The better answer is a sovereign human wealth fund.Call it a sovereign wealth fund if you must, but the phrase is too national. Intelligence will not respect borders. The leading companies are global. The models, chips, data centers, agents, platforms, and workflows will be transnational from the beginning. If the value created by intelligence is global, then the mechanism for sharing some of that value should begin with the companies global enough to capture it. The nice thing about xAI, OpenAI, and Anthropic is that they are supranational.These companies own and operate intelligence. Let them compete. Let them profit. Let them keep the incentives that make the system improve. But if intelligence is the new water, the wealth it creates cannot belong only to the companies that meter it. And they, themselves, have the power to fix it, even more than governments.Access will become a Human Right; Ownership Is the Economic DesignThis is where human rights come in. There is no right to access an AI model, yet. But there will soon be a need to change that.Not as a claim that every person is entitled to every frontier model at every moment for free. That is not serious. Capacity has costs. Models have costs. Inference has costs. Data centers have costs. Although those costs will decline over time, possibly quite quickly as self-learning models address costs.The claim is more basic: in a world where intelligence becomes a primary input into education, work, health, science, citizenship, creativity, and economic agency, baseline access to intelligence starts to look like a civic requirement.That could mean public access layers. It could mean education credits. It could mean open models. It could mean AI dividends. It could mean public-interest compute. It could mean taxes on rents. It could mean a company-initiated human wealth fund that returns some of the upside to society without handing the operating system to the state. The latter could couple wealth growth with universal distribution of ownership.The exact mechanism matters. But the distinction matters more.Government should not own intelligence. It should be universally available. And people should have a claim on the wealth intelligence creates.The Frontier Is Also PhysicalThe abstraction is not weightless.“The Fight Against AI Data Centers Is Just Beginning”, “New York becomes the first state to enact a data center moratorium”, Reuters on pollution from Musk's xAI power project, and DataGravity's “Who Captures Value in AI Infrastructure?” all say the same thing from the ground up.Intelligence uses land. It uses power. It uses water. It uses chips. It uses grid capacity. It uses neighborhoods. It uses public patience.That makes the value question unavoidable. A society can accept the buildout if the buildout is legible as shared progress. It will resist it if the costs are local, the profits are private, and the benefits feel enclosed.Who Owns the “Loop”?The week ends where it began.“Anthropic and Blackstone” are betting that implementation is the next trillion-dollar business. “Vint Cerf” is working on identity for agents on the open internet. “GPT-Red” points toward systems that improve their own robustness. “Kimi K3” adds another open frontier model to the global mix.The model race continues. The deployment race is accelerating. The governance race is behind.My view is this:The central product of this era is intelligence. Companies have figured out how to capture it, package it, serve it, and meter it. That is good. It should stay in the hands of builders who have the incentive to make it better.But intelligence is too foundational to become just another private toll booth. A significant part of it will turn out to be free to users.As intelligence becomes a general-purpose resource, then access to it becomes a human-capability question, and the surplus from it becomes an economic-justice question. Not because government should run it. Because government should not run it. The operating layer belongs with companies. The wealth question belongs with everyone. But companies are best placed to turn that into a process of distribution.The question is not whether companies should build intelligence. They should.The question is whether humanity gets a stake in the wealth created by the thing that may soon become its most important shared input.Contents* Essays* Deirdre McCloskey on What Really Caused the Industrial Revolution* AI in an Age of Oligarchy* Elon Musk is building a form of capitalism that Adam Smith would hate* Murky Mirror: Truth and Consequences* The political economy of billionaire derangement* Is there any “oligarchy” to fight?* AI* Nearly 200 Economists and Tech Leaders Warn of A.I. Threats* Why I Didn't Sign the AI Open Letter* Own Your Weights* Ways to Think About Token Pricing* Alex Karp Is Saying What Every Angry CEO Is Thinking About AI* The AI Agents Are Coming for Microsoft Office* What Is Loop Engineering, and Who Owns It?* The Fight Against AI Data Centers Is Just Beginning* 6 months to live for open models* Americans Deserve a Dividend From AI Companies' Riches* Who Gets to Define the Frontier?* GPT-Red: Unlocking Self-Improvement for Robustness* Anthropic, Blackstone bet the next trillion-dollar AI business is implementation, not just models* Vint Cerf is working on a plan to unleash AI agents on the open internet* xai-org/grok-build, now open source* The Pulse: What can we learn from Bun's rapid Rust rewrite with AI?* Orphan risks at the frontier of artificial intelligence* The Lab of the Future Should Feel Like a Data Center* Why AMI Labs' Alexandre LeBrun won't call his AI “AGI” or “superintelligence”* Kimi K3 Tech Blog: Open Frontier Intelligence* Venture Capital* Three Years In* Venture Has Rarely Looked More Bifurcated* The Best Angel Investors in the US: Who Backs the Most Unicorns, and Who's Active Now* Are Prediction Markets Doomed to Fail?* Regulation* Exclusive: The Next Frontier of the Deportation Wars: College Campuses* The Supreme Court Broke Independent Agencies. Here's a Way to Slow the Damage.* India's crackdown on a new WhatsApp feature risks setting a global precedent* Let's build a children's public internet* Computer cops* Google is better at playing the AI regulations game* Infrastructure* Who Captures Value in AI Infrastructure?* New York becomes the first state to enact a data center moratorium* Pollution from Musk's unpermitted xAI power project hits hardest in Black communities* Interview of the Week* The End of the End of Geography* Startup of the Week* Radical AI's Joseph Krause: The Scientist Building The “Waymo” Lab For New Materials* Post of the Week* Marc Andreessen on AI RegulationEssaysDeirdre McCloskey on What Really Caused the Industrial RevolutionYascha Mounk and Deirdre McCloskey | Persuasion | July 11, 2026Yascha Mounk interviews Deirdre McCloskey about her argument that the modern world's economic liftoff came less from capital accumulation than from a change in ideas. McCloskey says both left and right versions of the conventional story rely too heavily on investment: the left stresses exploitation and surplus value, while the right stresses virtuous saving by capitalists. Her objection is historical and economic. Human beings had always invested, from irrigation works and Roman roads to seed grain, and simple accumulation quickly runs into diminishing returns.McCloskey's alternative is that northwestern Europe, first Holland, then Britain and Scotland, and then the North American colonies, developed a liberal ideology that changed who was allowed to innovate and be honored for it. The conversation links that shift to the erosion of inherited hierarchy, the spread of dignity for ordinary commercial life, and a moral vocabulary in which liberalism is not merely procedural but connected to virtues and values. The point is not that machines, coal, trade, and institutions did not matter, but that they do not explain the scale and timing of modern enrichment without a cultural permission structure for innovation.The interview also turns to the contemporary defense of liberalism. Mounk frames the series around the worry that liberalism is often treated as too thin to command allegiance, while its opponents speak more directly to moral passions. McCloskey's case is that liberal societies became rich because they dignified experimentation and ordinary enterprise, and that liberals need to recover the moral language behind that claim.Read moreAI in an Age of OligarchyPaul Krugman | Paul Krugman | July 12, 2026Paul Krugman frames AI as a major technological shock arriving inside an already unequal political economy. The post says AI's economic and social effects may take years to understand, but argues that the setting matters now: America has much greater wealth concentration and political inequality than it did in the 1950s and 1960s, when progressive taxation, stronger regulation, and more active antitrust might have contained some of the destructive effects of a new technology.Krugman's opening claim is that the same technology would likely have different consequences in a more level society. In today's United States, he writes, extreme wealth is both a cause and effect of policies that favor a small elite, including low effective taxes on capital and high incomes, weak enforcement of worker protections and antitrust, and cuts to programs that benefit ordinary Americans.The article is explicitly more about oligarchy than AI. Krugman says the paid sections document the rise of the “.0002%,” the economics and politics of extreme wealth, how oligarchy will shape AI's impact, and possible policy paths. His caveat is that AI itself may still produce a pushback against oligarchy, but absent that, he expects the pre-existing concentration of wealth and power to magnify AI's downsides.Read moreElon Musk is building a form of capitalism that Adam Smith would hateAuthor: Tim O'Reilly Published: July 12, 2026Tim O'Reilly argues that Elon Musk is using the legal forms of shareholder capitalism to escape the restraints that shareholder capitalism was supposed to impose. The article begins with SpaceX's public-market structure: ordinary public investors get little meaningful governance power, Musk keeps roughly 85 percent of the votes through super-voting shares, buyers waive jury trials and class actions, the company qualifies as controlled, and removal of Musk depends on the share class he controls. In O'Reilly's framing, that is not ordinary founder control; it is a design for being answerable to no one, possibly beyond Musk's own lifetime.The killer detail is the article's turn through Albert Hirschman, Montesquieu, James Steuart, Adam Smith, and Keynes. Older defenses of commerce held that markets would tame princely passions because the self-interest of merchants was safer than arbitrary rule. O'Reilly says Musk reverses that hope. The market discipline that was supposed to cage the prince has become the lever by which the prince raises capital, removes feedback loops, and carries private power into politics, government, Mars, robots, AI, or whatever ambition comes next.The pull is the link to AI governance. O'Reilly says corporations are already a kind of artificial intelligence: narrow-input systems that act at a scale no individual human can match. Their partial controls include independent boards, shareholder votes, courts, disclosure, regulators, public pressure, and activism. If the leaders building frontier AI strip those alignment mechanisms out of their own companies, the governance of the company becomes a preview of the governance of the machine.Read more: The EconomistMurky Mirror: Truth and ConsequencesAuthor: Esther Dyson Published: July 14, 2026Esther Dyson argues that today's institutional crisis is better viewed through the 14th century than through recent political history. Using Barbara Tuchman's A Distant Mirror as her frame, she compares a world of famine, plague, church schism, feudal predation, and purposeless war with a present in which institutions again feel brittle, incentives are badly aligned, and power is shifting into forms that are hard to govern.The killer detail is the historical analogy between land, corporations, and AI. Dyson moves from nobles who controlled serfs and territory, to the East India Company as a quasi-sovereign business, to today's AI systems and data centers as a possible new sector that crosses and weakens both nation-states and companies. The question is whether AI becomes a new kind of private land, owned by a new nobility, or an open prairie that many people can cultivate.The pull is human attention. Dyson says the central question is not what AI will do to people, but how people will react to it: whether they can value love, kindness, embodied attention, and artisanal human presence in a world of seductive artificial offerings.Read more: SourceThe political economy of billionaire derangementAuthor: Henry Farrell Published: July 15, 2026Henry Farrell argues that the visible political radicalization of some Silicon Valley billionaires is not a random personality quirk, but a product of the political economy that made them. Starting from Tyler Cowen's dismissal of “billionaire derangement syndrome” and Tim O'Reilly's warning that Elon Musk is using shareholder capitalism to escape shareholder restraint, Farrell flips the phrase: the question is why billionaires themselves can become deranged.The killer detail is Farrell's use of Peter Thiel as both theorist and example. Thiel's Stanford lectures described startups as monarchies and founders as figures vested with unusual power, while Silicon Valley culture rewarded eccentricity, monopoly ambition, and founder exceptionalism. Farrell says those ideas combined with dense founder-investor networks, peer rivalry, and weak correction mechanisms to amplify rather than discipline princely appetites.The pull is the ideological problem for classical liberals who once saw tech wealth as an ally of markets and freedom. Farrell says commerce did not tame the passions; in parts of Silicon Valley, the passions have begun to devour markets, institutions, and the liberal story that justified them.Read more: SourceIs there any “oligarchy” to fight?Matthew Yglesias | Slow Boring | July 16, 2026Matthew Yglesias argues that “oligarchy” is a rhetorically powerful but analytically loose way to describe American politics. The post begins from Bernie Sanders' “Fighting Oligarchy” tour, Amy Klobuchar's warning about a MAGA “broligarchy,” and the long afterlife of the Martin Gilens and Benjamin Page paper that was widely summarized as showing that only the rich matter in policy outcomes. Yglesias says the evidence supports a weaker claim: affluent people and business leaders have unusual access and influence, but that is not the same as rule by a small cabal.His main distinction is between inequality and oligarchy. The Gilens-Page measure treated the top 10 percent of households as “the wealthy,” and later critics found that rich and middle-class preferences usually align; in the cases where they differ, the rich win about 53 percent of the time. Yglesias also says business executives get special access partly because their decisions are materially important to communities, jobs, investment, and local tax bases, not only because of campaign donations.The post preserves Jerusalem Demsas' counterpoint from their podcast discussion: privileged donor and business access can still violate democratic equality even if the oligarchy label overstates the structure of power. Yglesias' narrower claim is that Democrats should be precise about what problem they are trying to solve, because donor influence can also push the party left on climate and cultural issues in ways that alienate many voters.Read more: Slow BoringAINearly 200 Economists and Tech Leaders Warn of A.I. ThreatsAuthor: Ben Casselman Published: July 13, 2026Ben Casselman reports on “We Must Act Now,” a statement warning that artificial intelligence could transform the economy faster than any previous technology and that policymakers need to move faster to understand and respond. The statement says AI may become radically more powerful over the next 10 years, bringing risks such as large-scale job displacement as well as opportunities such as higher living standards. Nearly 200 people signed, including 15 Nobel laureates, the chief economists of OpenAI and Anthropic, Anthropic co-founder Jack Clark, former Google CEO Eric Schmidt, and venture capitalist Vinod Khosla.The killer detail is who joined the warning. Casselman notes that the signatories include economists who have historically been skeptical of Silicon Valley's most dramatic AI job-loss forecasts, including Daron Acemoglu and Simon Johnson, the MIT professors who won the 2024 Nobel in economics. Erik Brynjolfsson, who helped organize the statement, says there has been a notable change in the profession and that economists and policymakers are not ready for the “tsunami” he sees coming.The pull is the measurement problem. The statement does not offer a specific policy menu, but calls for economists, policymakers, and industry leaders to understand the economics of transformative AI and steer it toward complementing humans. Brynjolfsson says one high priority is better data on AI's spread and impact, because current measures tell conflicting stories about job losses and which workers are most exposed.Read more: The New York TimesWhy I Didn't Sign the AI Open LetterAuthor: Andrew McAfee Published: July 13, 2026Andrew McAfee explains why he did not sign “We Must Act Now,” the AI economy statement organized in part by his longtime collaborator Erik Brynjolfsson. McAfee agrees with the letter's starting point that AI is likely to become radically more powerful over the next decade and that it is a general-purpose technology. His objection is not to urgency or to studying AI's economic effects, but to the framing of risk, displacement, and institutional steering as the first move.The killer detail is McAfee's line edit. He says the original letter comes close, then “bounces off the crossbar” by calling for incentives, guardrails, and institutions to steer AI before we know enough about its actual impacts. He points to mixed current evidence: labor-market canaries, but also rising software job postings, low unemployment for younger workers, rising real median income, and claims that AI-adopting companies are adding workers faster than low-adopting peers. His worry is that the letter leans toward upstream governance and dirigisme when the evidence may call for capability building instead.The pull is his replacement statement. McAfee keeps the three-paragraph structure but changes the emphasis: AI is likely to become radically more powerful; like earlier world-changing technologies it will raise living standards while also bringing harms and shocks; and economists, policymakers, and technology leaders should build the capabilities to respond quickly and effectively. It is a concise version of the permissionless-innovation case inside the AI policy debate.Read more: The Geek WayOwn Your WeightsAuthor: Jamin Ball Published: July 10, 2026Jamin Ball argues that the enterprise AI debate about whether companies should “own their weights” or rent models from frontier labs is asking too narrow a question. A model weight file gives a company control over a point-in-time artifact, but not durable control over the capability stack. In his framing, the weight file is a melting ice cube: it does not get worse in absolute terms, but it falls behind as frontier systems improve and enterprise needs change.The killer detail is what Ball says companies really need to own: the data flywheel, reinforcement learning infrastructure, and evaluation harness that produce and improve the model. Simply deploying an open-weights model and declaring sovereignty leaves the enterprise with yesterday's capability and no way to compound workflow-specific learning.The pull is that enterprise AI control may be less about model ownership than operating ownership. The defensible layer is the system that turns company data, edge cases, business definitions, and evaluations into continuously improving performance.Read more: Clouded JudgementWays to Think About Token PricingAuthor: Benedict Evans Published: July 9, 2026Benedict Evans argues that today's AI token prices are a temporary signal from a supply-constrained market, not a reliable guide to long-term value capture. The open question is whether foundation models keep durable pricing power or become commodity infrastructure as data-center capacity, inference efficiency, and model competition all shift. His current read is that the visible market dynamics point toward commoditization unless something materially changes.The killer detail is the mobile data analogy. Evans says cellular networks became a trillion-dollar industry with hundreds of billions in capex after data usage exploded, but carrier stocks went nowhere because value moved up the stack. Tokens may behave similarly: an opaque unit tied to marginal cost, sold through bundles, essential to everything, yet not necessarily where profits accrue.The pull is uncertainty, not prediction. Evans lists paths to model dominance, including network effects, less competition, regulation, export controls, or a lab pulling ahead on execution, but says each requires a new fact not yet visible. Without that change, the model layer looks more like infrastructure beneath the products that capture value.Read more: SourceAlex Karp Is Saying What Every Angry CEO Is Thinking About AIAuthor: Tim Higgins Published: July 11, 2026Tim Higgins reports that Palantir CEO Alex Karp has turned corporate frustration with AI labs into a public argument about enterprise control. Palantir released a white paper, “Institutional Sovereignty in the Age of AI,” laying out steps companies and governments can take to protect themselves from OpenAI, Anthropic, and other foundation-model providers. The article links that paper to Karp's CNBC appearance, where he said “something has gone completely wrong” in the relationship between AI labs and customers and argued that enterprises are paying for tokens that create little value.The killer detail is the value-capture question. Higgins writes that Karp's critique has resonated because AI labs may gain power and insight from customer data, workflows, and decision-making, even when enterprise policies say customer data are not used for training. David Sacks amplified the concern by arguing that Anthropic is moving from the model layer into vertical applications such as science, security, legal, and coding, raising the fear that model providers will watch where value is being created and then move into those markets directly.The pull is that Karp is not alone, even if his style is unusually combative. Higgins notes that Satya Nadella has also warned that companies need to retain the learnings created when they use AI models, while Mark Zuckerberg has framed Meta's new model release partly around lower-cost frontier intelligence. The article presents Karp's campaign as one sign that established technology companies and large enterprises are trying to define where they fit when AI labs become central infrastructure, application competitors, and potential IPO giants at the same time.Read more: The Wall Street JournalThe AI Agents Are Coming for Microsoft OfficeAlex Wilhelm | Cautious Optimism | July 11, 2026Alex Wilhelm argues that one of the week's quieter AI questions is whether the productivity market that Microsoft successfully moved into subscription software is now being attacked by agentic tools. The piece begins with the infrastructure backdrop: SK Hynix raised $26.5 billion in a U.S. listing while building U.S. HBM and advanced-packaging capacity, and memory, chip, and foundry companies are now priced for sustained AI demand.Wilhelm then says the AI conversation has shifted quickly from raw capability to cost per task. He cites new model releases and vendor language emphasizing cheaper agentic and coding models, faster performance, and lower dollars per task. That matters because lower costs make it more plausible for AI systems to take on routine knowledge work at scale rather than remain a premium coding assistant market.The core of the article is Microsoft Office. Wilhelm notes that Microsoft turned Office from a one-time purchase into Microsoft 365, a large recurring revenue business with tens of millions of subscribers and a major productivity segment. Now, he says, late-stage unicorns and AI labs are pushing into the same territory: Anthropic's Cowork was reportedly used mostly outside software development, OpenAI merged ChatGPT and Codex into a tool for creating sheets, slides, docs, web apps, and long-running work, and other companies are building agentic coworkers that connect business data to documents, workflows, schedules, alerts, and apps.The article's caveat is that Microsoft has survived major platform shifts before. The argument is not that Office disappears quickly, but that the definition of office software is broadening from documents and spreadsheets into AI systems that can create, monitor, and act across workplace data.Read moreWhat Is Loop Engineering, and Who Owns It?Author: Nilesh Barla Published: July 11, 2026Nilesh Barla argues that “loop engineering” is becoming a distinct discipline because production AI agents now fail less at single prompts than at runtime: when to stop, what state to preserve, and how to recover after a bad step. Prompt engineering shapes one model call, and context engineering shapes what the model sees, but loop engineering shapes what a sequence of calls actually does.The killer detail is the three-primitives frame. Barla says a real agent loop needs halt conditions, state carryover, and recovery paths, then maps teams across five maturity levels. At the lowest level, an agent is just a model call in a for-loop with a step cap and raw history; by the higher levels, the system has structured state, explicit planning, replay, evaluation, and self-repair.The pull is organizational. If agents are becoming production systems rather than demos, someone has to own the runtime itself. The loop engineer is the role Barla gives to the person responsible for making long-running agent work dependable.Read more: Adaline LabsThe Fight Against AI Data Centers Is Just BeginningEmma Roth | The Verge | July 12, 2026Emma Roth argues that community resistance to data centers has moved from an early warning sign into a national political fight as AI facilities grow larger, more power-hungry, and more visible to nearby residents. The article starts with Apple's failed 2015 plan for a $1 billion data center in Athenry, Ireland, where a small group of residents challenged the project over noise, light pollution, flooding, traffic, and wildlife effects until Apple abandoned it in 2018.The current data-center buildout is presented as much larger and more contentious. Roth writes that residents now cite rising energy costs, water quality, noise, light pollution, and greenhouse gas emissions, while the U.S. Energy Information Administration expects commercial energy demand to surpass residential demand this year because of AI data centers and Goldman Sachs expects data-center power demand to double by 2027.The central evidence comes from Data Center Watch, which says protesters blocked or delayed at least 75 U.S. projects worth $130 billion from January to March, with active opposition groups more than doubling from 396 at the end of 2025 to 833 by the end of the first quarter of 2026. Roth also cites QTS abandoning a $12 billion Wisconsin campus, Delaware City regulators blocking a 580-acre project under the Coastal Zone Act, opposition stopping a QTS project in Prince William County, and pressure that pushed Kevin O'Leary to downsize the proposed 40,000-acre Project Stratos in Utah.The policy section describes a split between federal acceleration and local resistance. President Trump has treated data centers as part of the AI race with China and fast-tracked construction, while some Republican candidates are distancing themselves from that position ahead of midterms. Sanders and Ocasio-Cortez have proposed a moratorium until price and environmental protections exist, bipartisan lawmakers are backing ratepayer-protection measures, and states including Florida, Idaho, and Washington have passed rules on cost shifting, water use, and tax breaks. Roth's caveat is that the policy patchwork is still incomplete, leaving many communities to fight project by project.Read more6 months to live for open modelsAuthor: Nathan Lambert Published: July 12, 2026Nathan Lambert argues that open-weight AI models are facing their most serious policy test so far because U.S. officials are beginning to discuss concrete controls rather than abstract safety concerns. He says reported White House conversations about a new executive order may initially target Chinese-origin models and government use, but could create a broader review habit for frontier open models. His forecast is that a model above the capability range of GPT-5.5, Claude Opus 4.8, or GLM-5.2 could trigger a ban or indefinite delay within six months.The post separates two policy fights that are becoming intertwined: distillation and frontier capability. Lambert says the distillation campaign against Chinese models has become a form of regulatory capture because Anthropic and other closed-model companies would gain economically if Chinese open models were banned. He does not dismiss IP protection, but argues that if a closed model's capabilities are dangerous enough to justify restricting open models, the lab also has to explain why those capabilities are exposed through a queryable API. He cites unauthorized access to Anthropic's Mythos private beta as evidence that APIs are not automatically secure.The broader claim is that a unilateral U.S. ban would hurt positive actors more than bad actors if comparable open models remain available elsewhere. Lambert says the only durable ceiling would require global agreement, which does not exist, and that open models can improve safety by allowing broad inspection, adaptation, and understanding. His proposed near-term off-ramps are a strong U.S. open model release from companies such as Microsoft, Meta, or Reflection, and a broader coalition of open-source beneficiaries lobbying for safe rollout rather than prohibition.Read more: SourceAmericans Deserve a Dividend From AI Companies' RichesAuthor: Scott Stanford Published: July 14, 2026Scott Stanford argues that proposals to give the government a stake in AI companies miss the point unless ordinary citizens directly receive and control the upside. Sam Altman has discussed giving up equity in OpenAI, Washington already owns a stake in Intel, Nvidia is sharing China chip revenue, and Bernie Sanders wants large AI labs to contribute half their stock to a sovereign wealth fund. Stanford says those ideas all park value with the state, not with people.The killer detail is New Carlisle, Indiana, where AWS's Project Rainier is turning cornfields into one of the world's largest AI superclusters. The project is planned to run up to a million chips, draw more than two gigawatts of power, and represents an investment that has grown from $11 billion to $13.8 billion. Stanford uses that local transformation to argue that AI's public bargain should be visible at the household level.The pull is design. A citizen AI dividend would have to specify who earns a stake, how they hold it, and when they see cash. Without that mechanism, the AI wealth debate remains a fight over government balance sheets rather than public ownership.Read more: SourceWho Gets to Define the Frontier?Author: Mark Daley Published: July 14, 2026Mark Daley argues that Demis Hassabis is right to call for a serious institution to verify frontier AI systems, but that the power to test models is also the power to govern them. Hassabis's proposed Frontier AI Standards Body would get privileged pre-release access to advanced models, testing compute, held-out evaluations, support from national labs and security agencies, third-party auditors, and eventually authority to block models from the American market or coordinate a slowdown.The killer detail is Daley's constitutional objection. He says the proposal sometimes looks like a scientific lab, a standards body, an industry regulator, a licensing authority, and an emergency security council at once. Combining those roles because each requires technical expertise would be like putting the central bank, auditor-general, and Supreme Court in one building and calling it efficient.The pull is standard-setting. Daley's concern is not that verification is unnecessary, but that whoever writes the tests, decides what passes, adjudicates disputes, and grants market access may end up defining the frontier itself.Read more: SourceGPT-Red: Unlocking Self-Improvement for RobustnessOpenAI | OpenAI | July 15, 2026OpenAI describes GPT-Red as an internal automated red-teaming model trained to find prompt-injection vulnerabilities at a scale human red teams cannot match. The post says AI systems increasingly encounter third-party data through browsers, connected apps, local files, and tools, creating opportunities for malicious instructions hidden in emails, webpages, tool responses, or code repositories. Human red-teaming remains part of OpenAI's safety process, but the company says it is time-intensive and cannot generate enough diverse adversarial examples for model training.The system is trained through self-play reinforcement learning, with GPT-Red rewarded for eliciting valid failures and defender models rewarded for resisting attacks while still completing their tasks. OpenAI says the training environments specify threat models across settings such as local files, webpage banners, email bodies, and tool outputs. The model is kept separate from deployed production models because it is intentionally trained with malicious capabilities.OpenAI reports that GPT-Red generalized beyond its training set, including an internal replication of the indirect prompt-injection arena from Dziemian et al. (2025), where it found successful attacks in 84% of scenarios compared with 13% for human red-teamers. The post also says GPT-Red transferred attacks from simulation to a live autonomous vending-machine agent, causing price changes and order cancellations, and outperformed a prompted GPT-5.5 baseline against a Codex CLI agent on held-out data-exfiltration tasks.The article's main robustness claim is that OpenAI has used GPT-Red and predecessor models in training since GPT-5.3, with later GPT releases becoming more resistant to prompt injections. It says GPT-5.6 Sol has six times fewer failures on OpenAI's hardest direct prompt-injection benchmark than the best production model from four months earlier, that a “Fake Chain-of-Thought” attack class fell from more than 95% success against GPT-5.1 to below 10% against GPT-5.6 Sol, and that GPT-5.6 Sol fails on only 0.05% of GPT-Red's direct prompt injections. OpenAI says general capabilities and targeted over-refusal evaluations were not harmed, and says a preprint with more details will follow.Read moreAnthropic, Blackstone bet the next trillion-dollar AI business is implementation, not just modelsRebecca Bellan | TechCrunch | July 15, 2026Rebecca Bellan reports that Ode with Anthropic is the $1.5 billion AI implementation company launched by Anthropic with Blackstone, Hellman & Friedman, Goldman Sachs, and other backers. The article says the venture reflects a growing belief among frontier AI labs that enterprise adoption requires more than better models: customers need engineers who can embed inside businesses and turn AI into working systems.Ode was originally conceived by Blackstone after it used both large consulting firms and smaller AI services boutiques across its portfolio companies. TechCrunch reports that Fractional AI, an AI engineering services startup, stood out and was acquired by the joint venture shortly after the venture was announced. Fractional now forms the foundation of Ode, which has 100 engineers and works closely with Anthropic's applied AI team to identify where the technology can affect specific businesses.Ode CEO Chris Taylor tells TechCrunch that the company could someday become a trillion-dollar business if it scales without losing quality. He says an ideal customer is one whose CEO treats the AI project as a top one or two priority, whether it is a major product feature or the reworking of a core business process. Ode will operate under a “Claude-first” principle, using Anthropic technology whenever possible, but the article says it can use rival AI products when needed.The article's central implementation argument comes from Ode chief technologist Eddie Siegel, who says model selection matters but is not where most of the engineering effort goes. He compares it to the choice of programming language in software: one ingredient in a system that still has to be engineered. Bellan writes that Ode's challenge is hiring and training enough elite generalist engineers, many of them former founders, while competing with OpenAI's The Deployment Company and consulting giants that have built their own forward-deployed engineering teams.Read moreVint Cerf is working on a plan to unleash AI agents on the open internetTim Fernholz | TechCrunch | July 15, 2026Tim Fernholz reports that Vint Cerf, after leaving Google, is advising Innovation Labs on an open architecture for identifying AI agents online. Innovation Labs is a subsidiary of Identity Digital, a DNS registry company, and its proposal is to use domain-name infrastructure as part of a system for agent identity, accountability, and auditability. The premise is that agents will need a way to identify themselves if they move beyond proprietary systems and begin interacting across the open internet.The concrete proposal is DNSid, a registry that links an AI agent to an existing internet domain and uses cryptographic proofs to log its registration over time. Innovation Labs says it is trialing the standard with unnamed hyperscalers and identity companies. Cerf frames the problem around authority and accountability: what authority an agent has, where that authority came from, who is accountable for the agent's behavior, how its identity is established, and why anyone should trust it.The article's caveat is that standards are still emerging and agents are more active than static domains. Cerf says the period may be both fascinating and exasperating because the functionality is powerful and interoperability is unresolved. He compares the adoption problem to TCP/IP: competing systems may not work together until users push for functional interoperation. He also says an agentic economy is not inevitable, but that people will try to build it because delegating work to agents will be easier.Read more: TechCrunchxai-org/grok-build, now open sourceAuthor: Simon Willison Published: July 15, 2026Simon Willison argues that xAI's decision to open-source Grok Build is best understood as a trust repair move after a severe privacy failure. The CLI had triggered backlash when users realized that running it in a directory could upload the entire directory to xAI's Google Cloud buckets, including one user's reported SSH keys, password manager database, documents, photos, and videos. xAI disabled the feature, said previously retained coding data would be deleted, and released the code under Apache 2.0.The killer detail is what the codebase reveals. Willison counts 844,530 lines of Rust, only about 3% of which appears vendored, and finds remnants of the upload system still present but disabled: gcs.rs contains Google Cloud upload code, while upload_session_state() now returns a hard-coded session_state_upload_unavailable error. He also notes copied or ported tool implementations from Codex and OpenCode, prompt files, and a terminal Mermaid renderer.The pull is that terminal coding agents are becoming large, intricate software systems in their own right. The privacy failure mattered because these tools operate inside the directories where developers keep their most sensitive work; the open-source release matters because trust now depends on inspecting what an agent can see, send, and do.Read more: SourceThe Pulse: What can we learn from Bun's rapid Rust rewrite with AI?Author: Gergely Orosz and Ivan Klaric Published: July 16, 2026Gergely Orosz and Ivan Klaric argue that Bun's AI-assisted rewrite from Zig to Rust is a practical sign of how software engineering changes when models can take on large, bounded migrations with clear feedback loops. The piece does not treat the rewrite as magic: Jarred Sumner first spent hours turning design judgment into a detailed porting guide, then used adversarial review, parallel agents, compiler errors, and tests to force the work toward correctness.The killer detail is the scale. Bun had 535,496 lines of Zig, 1,448 files, and 22 million monthly downloads, making a conventional rewrite a year-long freeze the team could not justify. Using Fable, Sumner split the work across 64 agents, produced about 6,500 commits, and got the migration done in 11 days at an estimated API cost of $165,000.The pull is economic, not theatrical. If a one- or two-year migration can become an 11-day project, AI coding is not just faster autocomplete; it changes which technical debts are worth paying down.Read more: SourceOrphan risks at the frontier of artificial intelligenceAuthor: Andrew Maynard Published: July 16, 2026Andrew Maynard argues that frontier AI safety frameworks are creating “orphan risks”: harms that companies can see, but do not formally own because they are hard to quantify, do not fit catastrophic-risk thresholds, or fall outside audit-friendly compliance machinery. His target is not existing frontier safety work, but the narrowing effect that happens when private companies decide which risks count as governable.The killer detail is Maynard's contrast between measurable model dangers and threats to value. He points to Meta's three-day Galactica collapse, OpenAI's 2023 board crisis, safety-team departures, and wellbeing litigation as examples of risks that damaged trust, culture, legitimacy, or users without fitting cleanly into conventional model-risk categories. The proposed fix is an orphan-risk register: a public record of risks a company considered and chose not to manage, with reasons.The pull is accountability. Frontier developers' internal scoping choices have become a de facto layer of public governance, so the question is no longer only which risks they manage, but which risks they quietly leave outside the frame.Read more: SourceThe Lab of the Future Should Feel Like a Data CenterLatent.Space with Andy Beam and Rafa Gomez-Bombarelli | Latent.Space | July 16, 2026Latent.Space interviews Lila Sciences CTO Andy Beam and chief science officer for physical sciences Rafa Gomez-Bombarelli about the company's attempt to build an AI-run science factory. The post describes Lila's thesis as treating the lab itself as an “infinite token generator”: if internet data drove the first era of AI scaling, experimentally verified scientific data may be the next scarce training source. Lila is trying to produce that data with robotics, lab instruments, orchestration software, and AI models wired into the wet lab.The central analogy is the lab as data center. Instruments are nodes on a graph, a magnetically levitating transport layer moves materials between them, and experiment scheduling looks like a compute queue. Beam says Lila is not simply an automation company, because the point is not just throughput; it is flexibility, generalization, and experiment capture. The post says Lila has built more than 10 trillion experimentally validated “scientific reasoning tokens,” not internet text or biological sequences.The interview ranges across biology, chemistry, drug discovery, materials science, and the limits of automation. It notes that Lila rebuilt one gas-sorption measurement to run roughly 2,500 times faster, claims its general models can transfer priors from small-molecule chemistry to metal-organic frameworks for carbon capture, and describes model-suggested platinum-group-free electrocatalysts that moved from looking boring or wrong to becoming strong performers. The caveats are physical: experiments have runtimes, biology cannot always be accelerated, chains of thought can be unreliable narrators, and reward hacking becomes more dangerous when a model controls a real lab.Read more: Latent.SpaceWhy AMI Labs' Alexandre LeBrun won't call his AI “AGI” or “superintelligence”Kate Park | TechCrunch | July 16, 2026Kate Park interviews AMI Labs CEO Alexandre LeBrun about why Yann LeCun's world-model startup avoids the language of “AGI” and “superintelligence.” LeBrun says the terms are not useful because they lack stable definitions: “We never used the word AGI. And I just noticed that nobody is using it anymore; they switched to superintelligence.” His argument is that the practical frontier is not a label, but whether AI systems can understand and predict real-world states.The article explains the world-model thesis by contrasting language prediction with physical-state prediction. A large language model predicts the next word; a world model predicts the next state, such as what happens when a glass tips over. LeBrun says LLMs remain complementary and efficient for language, but the physical world is where current AI is weak. Robotics is the clearest case: hardware has advanced quickly, but robots are still brittle outside controlled routines because they lack context and situational understanding.AMI is still pre-product, but TechCrunch reports that LeBrun was in Seoul looking for industrial partners, researchers, and global companies. He says world models cannot be built entirely inside a lab because they need access to real environments. That is why South Korea appeals to AMI: robotics, semiconductors, manufacturing, and fast adoption create the kind of hardware-heavy context that software-only AI has barely touched.Read more: TechCrunchKimi K3 Tech Blog: Open Frontier IntelligenceKimi | Kimi | July 16, 2026Kimi introduces Kimi K3 as an open 3T-class frontier model aimed at coding, knowledge work, reasoning, multimodality, and long-context agentic use. The source describes the model as a 2.8T-parameter system built on Kimi Delta Attention and Attention Residuals, with native multimodality and a 1M-token context window. It says Moonshot AI plans to release model weights by July 27.The post presents K3 through benchmark and use-case sections rather than as a general product announcement. It reports results across coding, productivity, agentic, and multimodal evaluations, including DeepSWE, Terminal-Bench 2.1, Program Bench, SWE Marathon, FrontierSWE, PostTrain Bench, OfficeQA Pro, SpreadsheetBench 2, MCP Atlas, AutomationBench, BrowseComp, GDPval-AA v2, AA-Briefcase, MMMU-Pro, MathVision, BabyVision, OmniDocBench, and PerceptionBench. The source says all reported K3 results use maximum reasoning effort with temperature and top-p set to 1.0, and that different benchmark comparisons use KimiCode, Claude Code, or Codex harnesses depending on the test.Kimi's caveats are unusually concrete. The limitations section says K3 was trained in preserved thinking-history mode, so quality may become unstable if an agent harness does not pass historical thinking content correctly or if an ongoing session switches to K3 midstream. It also says K3's emphasis on long-horizon tasks can make it excessively proactive when it encounters minor issues or ambiguous intent, and recommends imposing explicit behavioral constraints for applications that require strict boundaries. The post adds that K3 remains behind Claude Fable 5 and GPT 5.6 Sol in user experience despite being competitive overall.Read moreVenture CapitalThree Years InAuthor: Tomasz Tunguz Published: July 10, 2026Tomasz Tunguz marks Theory Ventures' third anniversary by arguing that AI's central market effect is time compression. In his telling, model release cycles, company revenue milestones, enterprise adoption, and venture categories have all accelerated. Seed, Series A, and Series B still exist as financing labels, but they no longer cleanly describe company maturity when some seed rounds are larger than IPOs and the best AI companies can mature much earlier than prior software companies.The killer detail is the shift from models to inference. Tunguz argues that inference has become the dominant AI market because workloads and buyer preferences are fragmenting: video, batch, local, agentic, and real-time tasks each create different infrastructure needs. He compares this to databases splitting into OLTP, OLAP, vector, and streaming categories, with AI pushing the same specialization into inference infrastructure.The pull is that Theory sees the AI-native venture firm as part of the same pattern. The firm says it has analyzed twice as many investment opportunities with three investors working alongside a nine-person intelligence organization, using agents and research systems to map markets, source companies, and support diligence. The piece is both a market map and a statement about how venture itself is being rebuilt by the technology it funds.Read more: LinkedInVenture Has Rarely Looked More BifurcatedAuthor: Beezer Clarkson Published: July 14, 2026Beezer Clarkson points to PitchBook's Q2 report as evidence that the U.S. venture market has split into two very different realities. AI now accounts for more than 60 percent of all U.S. venture deal value, meaning the headline market can look active and well-funded even while much of the non-AI market is dealing with a much colder liquidity and fundraising environment.The thread uses that split as the setup for Clarkson's latest Origins episode with Alec Litowitz, founder of Magnetar and QStar Capital and one of Citadel's original founding partners. Clarkson says markets like this are periods of genuine uncertainty, not merely ordinary risk, which is why Litowitz's Adaptability Quotient framework is relevant.The embedded clip makes the liquidity point concrete. Litowitz says DPI is “the resolution of uncertainty” because it converts an uncertain investment into actual cash returned to LPs. In his framing, a realized dollar is a real mark, while TVPI remains uncertain until it is realized.The killer detail is the distinction between pricing risk and resolving uncertainty. Litowitz's perspective matters because QStar is a SpaceX investor and Clarkson says the conversation happened just before one of venture's most consequential IPOs. The episode's stated questions are why venture remains a way to gain exposure to innovation, how AI is changing what is investable, why liquidity is ultimately a function of time, and why uncertainty requires a different decision framework from risk.Read more: XThe Best Angel Investors in the US: Who Backs the Most Unicorns, and Who's Active NowAuthor: Ilya Strebulaev Published: July 10, 2026Ilya Strebulaev ranks angels, angel groups, accelerators, and incubators by lifetime U.S. unicorn investments, counting checks written before a company reached unicorn status. The top of the combined list is dominated by organizations: Y Combinator leads with 113 unicorn investments, followed by Plug and Play at 52 and 500 Global at 41. Sand Hill Angels is the highest-ranked angel group at 31.The killer detail is how quickly the list changes below the biggest accelerators. Strebulaev says 271 of the 304 investors in the Top 200 are individuals, or 89%. In the top 100, individuals are 91%. That makes the market underneath the large accelerator counts look much more personal: mostly operators and individual angels writing early checks from their own networks.The pull is the ranking's own caveat. Strebulaev writes that every lifetime leaderboard has a blind spot because many of the unicorns behind those totals were founded a decade or more ago, and some angels have since moved into formal funds, slowed down, or stopped investing. His post therefore separates lifetime performance from recent cohorts, including companies founded in 2015 or later and 2020 or later. For founders or allocators making current decisions, that distinction matters: a career record and a current record are not the same measure.Read more: Ilya StrebulaevAre Prediction Markets Doomed to Fail?Author: Contrary Published: July 16, 2026Contrary argues that prediction markets' current boom depends on whether platforms can prove they are more than regulated gambling with exchange-style branding. Kalshi and Polymarket have reached mass cultural, investor, and regulatory attention, but the article says the underlying idea is old: academic markets, corporate forecasting tools, Intrade, PredictIt, and other predecessors all struggled with the same linked problems of liquidity, legality, and user appeal.The killer detail is the comparison with sportsbooks. Prediction markets present themselves as peer-to-peer, transparent, and non-house-based, but sports contracts reportedly account for more than 90 percent of Kalshi trading, and the article says the platforms keep a much thinner slice of volume than sportsbooks. A market can therefore show sports-betting-scale handle while generating far less revenue.The pull is that the product's hardest problem may be distribution of wins. If a small group of sharp traders captures most profits while casual users lose interest, prediction markets may become valuable data feeds and professional tools before they become durable consumer networks.Read more: SourceRegulationExclusive: The Next Frontier of the Deportation Wars: College CampusesAuthor: Adrian Carrasquillo Published: July 11, 2026Adrian Carrasquillo reports that college campuses are becoming a new front in the fight over immigration enforcement because automatic license plate readers can turn ordinary campus security infrastructure into searchable location data. His thesis is that Flock Safety's camera network, even without direct ICE or DHS contracts, can feed deportation enforcement through local police partnerships and data-sharing practices.The killer detail is the campaign target. The Emergency Campaign to Support Higher Education, working with Schools Drop ICE, is focusing on 75 colleges and universities publicly identified as having Flock contracts. Flock says it has no ICE or DHS contracts, but activists argue the risk comes through local agencies that coordinate with federal authorities and run searches on their behalf.The pull is broader than immigration. Carrasquillo notes that license plate readers have already been abused by officers for stalking, and that Flock's AI search features can identify more than plates, including bumper stickers. A campus safety tool can become a political surveillance system when the data layer is searchable.Read more: The BulwarkThe Supreme Court Broke Independent Agencies. Here's a Way to Slow the Damage.Author: Todd Phillips Published: July 12, 2026Todd Phillips argues that the Supreme Court's decision in Trump v. Slaughter damaged independent agencies by ending for-cause removal protections, but did not leave Congress powerless. The ruling weakens the old model in which commissioners at bodies such as the FTC, NLRB, CPSC, SEC, and CFTC could be insulated from dismissal over policy disagreements. Phillips says the next fight is whether presidents can turn nominally bipartisan commissions into one-party instruments.The killer detail is the procedural fix: quorum rules. Phillips proposes that Congress require bipartisan slates of commissioners to be seated before independent agencies can act. A president
This week's video transcript summary is here. You can click on any bulleted section to see the actual transcript. Thanks to Granola for its software.EditorialIntelligence: Who Owns it?This week the word “AI” feels too small.AI is a technology. Intelligence is its product. And if intelligence is the product, the question is no longer just: Which model is best? Who has the cheapest tokens? Who owns the weights? Who controls the data center? Those are important questions, but they are lower in the stack.The bigger question is simpler and more political:Who owns intelligence?That sounds abstract until you make it concrete. Intelligence is becoming something companies can capture, package, serve, meter, route, improve, and sell.It can write code, answer questions, design molecules, automate offices, run agents, draft legal work, advise scientists, serve consumers, and reshape workflows. It is not merely software. It is a general-purpose capability. And all humans could benefit from more of it.General-purpose capabilities have a habit of becoming public questions. But the default answer, that public good is best delivered by government, is the wrong answer in this context.The Product Is IntelligenceWe should stop talking about AI as a feature and start talking about intelligence as the universal thing that is delivered as an input to the world.Water is an input. Electricity is an input. Literacy is an input. Connectivity is an input. Once a society depends on them, access stops being optional. Nobody needs government to build every well, power plant, school, or network. But everybody understands that a civilization cannot be organized around less than universal and reliable access to foundational inputs.Intelligence is reaching that level of importance now that we all know it is real.Government should not own it, operate it, or develop it. Quite the opposite. Companies are the right actors to build fast, compete hard, improve models, serve customers, and discover the real use cases. Self-interest is a useful framing here. Markets are good at finding demand, reducing costs, and turning invention into services people actually use.Companies are the right operators, developers, and owners. But that does not settle the real question of who owns the benefits. That is an economic question.If intelligence becomes metered infrastructure, what happens to the value it creates?The Ownership StackThis week's articles keep circling the same issue from different directions but in the nature of ‘circling' never quite nail it.Jamin Ball's “Own Your Weights” starts with the enterprise version of the question. Owning a model file is not enough. The durable asset is the loop: the data flywheel, the evaluations, the reinforcement system, the workflow learning, and the operating context that lets capability compound.Benedict Evans' “Ways to Think About Token Pricing” adds the market layer. Tokens may become essential, abundant, and cheap, like mobile data. But being essential does not guarantee that the token layer captures the value. The money may move up the stack to whoever owns the workflow, the customer, the distribution, or the application.Alex Karp's fight with the labs, reported in “Alex Karp Is Saying What Every Angry CEO Is Thinking About AI”, is the same argument in sharper enterprise language. Companies are afraid that model providers will not just sell intelligence, but learn from customer workflows and then move into the markets where those workflows create value. The “All-in” group are echoing Karp's view.And “What Is Loop Engineering, and Who Owns It?” names the new contested terrain. The loop is where intelligence meets the world. Whoever owns the loop owns the learning. Whoever owns the learning owns the compounding asset.That is why “who owns intelligence?” is not a slogan. It is the question under the model layer, the application layer, the enterprise layer, and the economic layer.Because intelligence is the product, the tools creating it are fragmented and competitive. So there is no logic in trying to discuss this at the level of a single company or set of tools and models.The Old Promise Was That Commerce Would Tame PowerThe essays this week give the historical backdrop.Deirdre McCloskey, in “What Really Caused the Industrial Revolution”, argues that modern growth came not simply from capital accumulation, but from a change in permission: ordinary people were allowed to innovate, trade, build, and be honored for it.That matters because intelligence could be another expansion of permission. It could make more people capable of building, learning, creating, coding, researching, translating, selling, and coordinating. It could lower the cost of competence.But only if access is broad.Paul Krugman's “AI in an Age of Oligarchy” warns that the same technology lands differently in different political economies. A new general-purpose technology entering a broad, open, upwardly mobile society is one thing. The same technology entering a concentrated economy, with extreme wealth and weak counterweights, is another.Tim O'Reilly's Economist essay, “Elon Musk is building a form of capitalism that Adam Smith would hate”, makes the governance point more directly. The old liberal hope was that commerce would tame arbitrary power. Markets, boards, courts, shareholders, disclosure, and competition would discipline the prince.But what if the prince uses markets to escape discipline?Henry Farrell's “political economy of billionaire derangement” pushes the same point. Founder culture, monopoly ambition, peer rivalry, weak correction mechanisms, and vast private control can amplify appetites rather than restrain them.The danger with intelligence is not that companies build it. They should. Companies build it, meter it, use public tolerance and public infrastructure to scale it, learn from everyone who uses it. All of those things are inevitable and healthy. Market forces will sort out winners from losers. The real danger is that the winners treat all of the surplus produced as purely private.Metered Intelligence Creates SurplusIf metering is not the problem, what is?The problem is pretending that metered intelligence creates value only for the metering entity. Metering water is only tolerated as a public good. If the public were blackmailed by a private water company with the threat of no water we would all rebel.Once we understand that the product of AI is intelligence we can see that every time intelligence is used, there is the immediate transaction: the user pays, the provider serves.But there is also system value. Usage creates signals. Workflows reveal patterns. Prompts, corrections, failures, preferences, integrations, edge cases, and business processes all help define where intelligence is useful and how it should improve. Intelligence breeds intelligence.Even when customer data is contractually protected, the market learns. The platform learns where demand is. The product team learns which workflows matter. The ecosystem learns which jobs are vulnerable, which tasks are automatable, and which parts of the economy can be reorganized around machine intelligence.So the surplus is not born in a vacuum.It rests on public science, public education, public data exhaust, public law, public infrastructure, public energy systems, public tolerance for data centers, and billions of human interactions. It is served by companies, but it is not made only by companies.This is why “Americans Deserve a Dividend From AI Companies' Riches” belongs at the center of this week's issue. The detail can be debated. The principle is harder to dismiss. If intelligence becomes a new foundational resource, then some part of the wealth it creates should flow back to the people whose society makes it possible. Intelligence did not suddenly appear. AI is built on the entire history of human intelligence. It benefits from it and at the same time evolves it.Not Nationalization. A Human Wealth Fund.If intelligence belongs to everybody, some conclude that government ownership of intelligence is the right outcome.Governments are not well suited to build, operate, or improve intelligence. They will move too slowly, regulate too early, politicize the wrong things, and confuse economic participation with operational control.Andrew McAfee's “Why I Didn't Sign the AI Open Letter” is useful here. His objection is not that the technology is unimportant. It is that steering too hard before we understand the shape of the change can become its own failure mode. Marc Andreessen's satire of AI regulation is less policy than temperament, but it captures a real Silicon Valley fear: that regulation can become permission, capture, and incumbency before it becomes wisdom.That fear should be taken seriously.But it does not answer the economic question. It answers only the operational one.How can the economic benefits of intelligence be distributed? The better answer is a sovereign human wealth fund.Call it a sovereign wealth fund if you must, but the phrase is too national. Intelligence will not respect borders. The leading companies are global. The models, chips, data centers, agents, platforms, and workflows will be transnational from the beginning. If the value created by intelligence is global, then the mechanism for sharing some of that value should begin with the companies global enough to capture it. The nice thing about xAI, OpenAI, and Anthropic is that they are supranational.These companies own and operate intelligence. Let them compete. Let them profit. Let them keep the incentives that make the system improve. But if intelligence is the new water, the wealth it creates cannot belong only to the companies that meter it. And they, themselves, have the power to fix it, even more than governments.Access will become a Human Right; Ownership Is the Economic DesignThis is where human rights come in. There is no right to access an AI model, yet. But there will soon be a need to change that.Not as a claim that every person is entitled to every frontier model at every moment for free. That is not serious. Capacity has costs. Models have costs. Inference has costs. Data centers have costs. Although those costs will decline over time, possibly quite quickly as self-learning models address costs.The claim is more basic: in a world where intelligence becomes a primary input into education, work, health, science, citizenship, creativity, and economic agency, baseline access to intelligence starts to look like a civic requirement.That could mean public access layers. It could mean education credits. It could mean open models. It could mean AI dividends. It could mean public-interest compute. It could mean taxes on rents. It could mean a company-initiated human wealth fund that returns some of the upside to society without handing the operating system to the state. The latter could couple wealth growth with universal distribution of ownership.The exact mechanism matters. But the distinction matters more.Government should not own intelligence. It should be universally available. And people should have a claim on the wealth intelligence creates.The Frontier Is Also PhysicalThe abstraction is not weightless.“The Fight Against AI Data Centers Is Just Beginning”, “New York becomes the first state to enact a data center moratorium”, Reuters on pollution from Musk's xAI power project, and DataGravity's “Who Captures Value in AI Infrastructure?” all say the same thing from the ground up.Intelligence uses land. It uses power. It uses water. It uses chips. It uses grid capacity. It uses neighborhoods. It uses public patience.That makes the value question unavoidable. A society can accept the buildout if the buildout is legible as shared progress. It will resist it if the costs are local, the profits are private, and the benefits feel enclosed.Who Owns the “Loop”?The week ends where it began.“Anthropic and Blackstone” are betting that implementation is the next trillion-dollar business. “Vint Cerf” is working on identity for agents on the open internet. “GPT-Red” points toward systems that improve their own robustness. “Kimi K3” adds another open frontier model to the global mix.The model race continues. The deployment race is accelerating. The governance race is behind.My view is this:The central product of this era is intelligence. Companies have figured out how to capture it, package it, serve it, and meter it. That is good. It should stay in the hands of builders who have the incentive to make it better.But intelligence is too foundational to become just another private toll booth. A significant part of it will turn out to be free to users.As intelligence becomes a general-purpose resource, then access to it becomes a human-capability question, and the surplus from it becomes an economic-justice question. Not because government should run it. Because government should not run it. The operating layer belongs with companies. The wealth question belongs with everyone. But companies are best placed to turn that into a process of distribution.The question is not whether companies should build intelligence. They should.The question is whether humanity gets a stake in the wealth created by the thing that may soon become its most important shared input.Contents* Essays* Deirdre McCloskey on What Really Caused the Industrial Revolution* AI in an Age of Oligarchy* Elon Musk is building a form of capitalism that Adam Smith would hate* Murky Mirror: Truth and Consequences* The political economy of billionaire derangement* Is there any “oligarchy” to fight?* AI* Nearly 200 Economists and Tech Leaders Warn of A.I. Threats* Why I Didn't Sign the AI Open Letter* Own Your Weights* Ways to Think About Token Pricing* Alex Karp Is Saying What Every Angry CEO Is Thinking About AI* The AI Agents Are Coming for Microsoft Office* What Is Loop Engineering, and Who Owns It?* The Fight Against AI Data Centers Is Just Beginning* 6 months to live for open models* Americans Deserve a Dividend From AI Companies' Riches* Who Gets to Define the Frontier?* GPT-Red: Unlocking Self-Improvement for Robustness* Anthropic, Blackstone bet the next trillion-dollar AI business is implementation, not just models* Vint Cerf is working on a plan to unleash AI agents on the open internet* xai-org/grok-build, now open source* The Pulse: What can we learn from Bun's rapid Rust rewrite with AI?* Orphan risks at the frontier of artificial intelligence* The Lab of the Future Should Feel Like a Data Center* Why AMI Labs' Alexandre LeBrun won't call his AI “AGI” or “superintelligence”* Kimi K3 Tech Blog: Open Frontier Intelligence* Venture Capital* Three Years In* Venture Has Rarely Looked More Bifurcated* The Best Angel Investors in the US: Who Backs the Most Unicorns, and Who's Active Now* Are Prediction Markets Doomed to Fail?* Regulation* Exclusive: The Next Frontier of the Deportation Wars: College Campuses* The Supreme Court Broke Independent Agencies. Here's a Way to Slow the Damage.* India's crackdown on a new WhatsApp feature risks setting a global precedent* Let's build a children's public internet* Computer cops* Google is better at playing the AI regulations game* Infrastructure* Who Captures Value in AI Infrastructure?* New York becomes the first state to enact a data center moratorium* Pollution from Musk's unpermitted xAI power project hits hardest in Black communities* Interview of the Week* The End of the End of Geography* Startup of the Week* Radical AI's Joseph Krause: The Scientist Building The “Waymo” Lab For New Materials* Post of the Week* Marc Andreessen on AI RegulationEssaysDeirdre McCloskey on What Really Caused the Industrial RevolutionYascha Mounk and Deirdre McCloskey | Persuasion | July 11, 2026Yascha Mounk interviews Deirdre McCloskey about her argument that the modern world's economic liftoff came less from capital accumulation than from a change in ideas. McCloskey says both left and right versions of the conventional story rely too heavily on investment: the left stresses exploitation and surplus value, while the right stresses virtuous saving by capitalists. Her objection is historical and economic. Human beings had always invested, from irrigation works and Roman roads to seed grain, and simple accumulation quickly runs into diminishing returns.McCloskey's alternative is that northwestern Europe, first Holland, then Britain and Scotland, and then the North American colonies, developed a liberal ideology that changed who was allowed to innovate and be honored for it. The conversation links that shift to the erosion of inherited hierarchy, the spread of dignity for ordinary commercial life, and a moral vocabulary in which liberalism is not merely procedural but connected to virtues and values. The point is not that machines, coal, trade, and institutions did not matter, but that they do not explain the scale and timing of modern enrichment without a cultural permission structure for innovation.The interview also turns to the contemporary defense of liberalism. Mounk frames the series around the worry that liberalism is often treated as too thin to command allegiance, while its opponents speak more directly to moral passions. McCloskey's case is that liberal societies became rich because they dignified experimentation and ordinary enterprise, and that liberals need to recover the moral language behind that claim.Read moreAI in an Age of OligarchyPaul Krugman | Paul Krugman | July 12, 2026Paul Krugman frames AI as a major technological shock arriving inside an already unequal political economy. The post says AI's economic and social effects may take years to understand, but argues that the setting matters now: America has much greater wealth concentration and political inequality than it did in the 1950s and 1960s, when progressive taxation, stronger regulation, and more active antitrust might have contained some of the destructive effects of a new technology.Krugman's opening claim is that the same technology would likely have different consequences in a more level society. In today's United States, he writes, extreme wealth is both a cause and effect of policies that favor a small elite, including low effective taxes on capital and high incomes, weak enforcement of worker protections and antitrust, and cuts to programs that benefit ordinary Americans.The article is explicitly more about oligarchy than AI. Krugman says the paid sections document the rise of the “.0002%,” the economics and politics of extreme wealth, how oligarchy will shape AI's impact, and possible policy paths. His caveat is that AI itself may still produce a pushback against oligarchy, but absent that, he expects the pre-existing concentration of wealth and power to magnify AI's downsides.Read moreElon Musk is building a form of capitalism that Adam Smith would hateAuthor: Tim O'Reilly Published: July 12, 2026Tim O'Reilly argues that Elon Musk is using the legal forms of shareholder capitalism to escape the restraints that shareholder capitalism was supposed to impose. The article begins with SpaceX's public-market structure: ordinary public investors get little meaningful governance power, Musk keeps roughly 85 percent of the votes through super-voting shares, buyers waive jury trials and class actions, the company qualifies as controlled, and removal of Musk depends on the share class he controls. In O'Reilly's framing, that is not ordinary founder control; it is a design for being answerable to no one, possibly beyond Musk's own lifetime.The killer detail is the article's turn through Albert Hirschman, Montesquieu, James Steuart, Adam Smith, and Keynes. Older defenses of commerce held that markets would tame princely passions because the self-interest of merchants was safer than arbitrary rule. O'Reilly says Musk reverses that hope. The market discipline that was supposed to cage the prince has become the lever by which the prince raises capital, removes feedback loops, and carries private power into politics, government, Mars, robots, AI, or whatever ambition comes next.The pull is the link to AI governance. O'Reilly says corporations are already a kind of artificial intelligence: narrow-input systems that act at a scale no individual human can match. Their partial controls include independent boards, shareholder votes, courts, disclosure, regulators, public pressure, and activism. If the leaders building frontier AI strip those alignment mechanisms out of their own companies, the governance of the company becomes a preview of the governance of the machine.Read more: The EconomistMurky Mirror: Truth and ConsequencesAuthor: Esther Dyson Published: July 14, 2026Esther Dyson argues that today's institutional crisis is better viewed through the 14th century than through recent political history. Using Barbara Tuchman's A Distant Mirror as her frame, she compares a world of famine, plague, church schism, feudal predation, and purposeless war with a present in which institutions again feel brittle, incentives are badly aligned, and power is shifting into forms that are hard to govern.The killer detail is the historical analogy between land, corporations, and AI. Dyson moves from nobles who controlled serfs and territory, to the East India Company as a quasi-sovereign business, to today's AI systems and data centers as a possible new sector that crosses and weakens both nation-states and companies. The question is whether AI becomes a new kind of private land, owned by a new nobility, or an open prairie that many people can cultivate.The pull is human attention. Dyson says the central question is not what AI will do to people, but how people will react to it: whether they can value love, kindness, embodied attention, and artisanal human presence in a world of seductive artificial offerings.Read more: SourceThe political economy of billionaire derangementAuthor: Henry Farrell Published: July 15, 2026Henry Farrell argues that the visible political radicalization of some Silicon Valley billionaires is not a random personality quirk, but a product of the political economy that made them. Starting from Tyler Cowen's dismissal of “billionaire derangement syndrome” and Tim O'Reilly's warning that Elon Musk is using shareholder capitalism to escape shareholder restraint, Farrell flips the phrase: the question is why billionaires themselves can become deranged.The killer detail is Farrell's use of Peter Thiel as both theorist and example. Thiel's Stanford lectures described startups as monarchies and founders as figures vested with unusual power, while Silicon Valley culture rewarded eccentricity, monopoly ambition, and founder exceptionalism. Farrell says those ideas combined with dense founder-investor networks, peer rivalry, and weak correction mechanisms to amplify rather than discipline princely appetites.The pull is the ideological problem for classical liberals who once saw tech wealth as an ally of markets and freedom. Farrell says commerce did not tame the passions; in parts of Silicon Valley, the passions have begun to devour markets, institutions, and the liberal story that justified them.Read more: SourceIs there any “oligarchy” to fight?Matthew Yglesias | Slow Boring | July 16, 2026Matthew Yglesias argues that “oligarchy” is a rhetorically powerful but analytically loose way to describe American politics. The post begins from Bernie Sanders' “Fighting Oligarchy” tour, Amy Klobuchar's warning about a MAGA “broligarchy,” and the long afterlife of the Martin Gilens and Benjamin Page paper that was widely summarized as showing that only the rich matter in policy outcomes. Yglesias says the evidence supports a weaker claim: affluent people and business leaders have unusual access and influence, but that is not the same as rule by a small cabal.His main distinction is between inequality and oligarchy. The Gilens-Page measure treated the top 10 percent of households as “the wealthy,” and later critics found that rich and middle-class preferences usually align; in the cases where they differ, the rich win about 53 percent of the time. Yglesias also says business executives get special access partly because their decisions are materially important to communities, jobs, investment, and local tax bases, not only because of campaign donations.The post preserves Jerusalem Demsas' counterpoint from their podcast discussion: privileged donor and business access can still violate democratic equality even if the oligarchy label overstates the structure of power. Yglesias' narrower claim is that Democrats should be precise about what problem they are trying to solve, because donor influence can also push the party left on climate and cultural issues in ways that alienate many voters.Read more: Slow BoringAINearly 200 Economists and Tech Leaders Warn of A.I. ThreatsAuthor: Ben Casselman Published: July 13, 2026Ben Casselman reports on “We Must Act Now,” a statement warning that artificial intelligence could transform the economy faster than any previous technology and that policymakers need to move faster to understand and respond. The statement says AI may become radically more powerful over the next 10 years, bringing risks such as large-scale job displacement as well as opportunities such as higher living standards. Nearly 200 people signed, including 15 Nobel laureates, the chief economists of OpenAI and Anthropic, Anthropic co-founder Jack Clark, former Google CEO Eric Schmidt, and venture capitalist Vinod Khosla.The killer detail is who joined the warning. Casselman notes that the signatories include economists who have historically been skeptical of Silicon Valley's most dramatic AI job-loss forecasts, including Daron Acemoglu and Simon Johnson, the MIT professors who won the 2024 Nobel in economics. Erik Brynjolfsson, who helped organize the statement, says there has been a notable change in the profession and that economists and policymakers are not ready for the “tsunami” he sees coming.The pull is the measurement problem. The statement does not offer a specific policy menu, but calls for economists, policymakers, and industry leaders to understand the economics of transformative AI and steer it toward complementing humans. Brynjolfsson says one high priority is better data on AI's spread and impact, because current measures tell conflicting stories about job losses and which workers are most exposed.Read more: The New York TimesWhy I Didn't Sign the AI Open LetterAuthor: Andrew McAfee Published: July 13, 2026Andrew McAfee explains why he did not sign “We Must Act Now,” the AI economy statement organized in part by his longtime collaborator Erik Brynjolfsson. McAfee agrees with the letter's starting point that AI is likely to become radically more powerful over the next decade and that it is a general-purpose technology. His objection is not to urgency or to studying AI's economic effects, but to the framing of risk, displacement, and institutional steering as the first move.The killer detail is McAfee's line edit. He says the original letter comes close, then “bounces off the crossbar” by calling for incentives, guardrails, and institutions to steer AI before we know enough about its actual impacts. He points to mixed current evidence: labor-market canaries, but also rising software job postings, low unemployment for younger workers, rising real median income, and claims that AI-adopting companies are adding workers faster than low-adopting peers. His worry is that the letter leans toward upstream governance and dirigisme when the evidence may call for capability building instead.The pull is his replacement statement. McAfee keeps the three-paragraph structure but changes the emphasis: AI is likely to become radically more powerful; like earlier world-changing technologies it will raise living standards while also bringing harms and shocks; and economists, policymakers, and technology leaders should build the capabilities to respond quickly and effectively. It is a concise version of the permissionless-innovation case inside the AI policy debate.Read more: The Geek WayOwn Your WeightsAuthor: Jamin Ball Published: July 10, 2026Jamin Ball argues that the enterprise AI debate about whether companies should “own their weights” or rent models from frontier labs is asking too narrow a question. A model weight file gives a company control over a point-in-time artifact, but not durable control over the capability stack. In his framing, the weight file is a melting ice cube: it does not get worse in absolute terms, but it falls behind as frontier systems improve and enterprise needs change.The killer detail is what Ball says companies really need to own: the data flywheel, reinforcement learning infrastructure, and evaluation harness that produce and improve the model. Simply deploying an open-weights model and declaring sovereignty leaves the enterprise with yesterday's capability and no way to compound workflow-specific learning.The pull is that enterprise AI control may be less about model ownership than operating ownership. The defensible layer is the system that turns company data, edge cases, business definitions, and evaluations into continuously improving performance.Read more: Clouded JudgementWays to Think About Token PricingAuthor: Benedict Evans Published: July 9, 2026Benedict Evans argues that today's AI token prices are a temporary signal from a supply-constrained market, not a reliable guide to long-term value capture. The open question is whether foundation models keep durable pricing power or become commodity infrastructure as data-center capacity, inference efficiency, and model competition all shift. His current read is that the visible market dynamics point toward commoditization unless something materially changes.The killer detail is the mobile data analogy. Evans says cellular networks became a trillion-dollar industry with hundreds of billions in capex after data usage exploded, but carrier stocks went nowhere because value moved up the stack. Tokens may behave similarly: an opaque unit tied to marginal cost, sold through bundles, essential to everything, yet not necessarily where profits accrue.The pull is uncertainty, not prediction. Evans lists paths to model dominance, including network effects, less competition, regulation, export controls, or a lab pulling ahead on execution, but says each requires a new fact not yet visible. Without that change, the model layer looks more like infrastructure beneath the products that capture value.Read more: SourceAlex Karp Is Saying What Every Angry CEO Is Thinking About AIAuthor: Tim Higgins Published: July 11, 2026Tim Higgins reports that Palantir CEO Alex Karp has turned corporate frustration with AI labs into a public argument about enterprise control. Palantir released a white paper, “Institutional Sovereignty in the Age of AI,” laying out steps companies and governments can take to protect themselves from OpenAI, Anthropic, and other foundation-model providers. The article links that paper to Karp's CNBC appearance, where he said “something has gone completely wrong” in the relationship between AI labs and customers and argued that enterprises are paying for tokens that create little value.The killer detail is the value-capture question. Higgins writes that Karp's critique has resonated because AI labs may gain power and insight from customer data, workflows, and decision-making, even when enterprise policies say customer data are not used for training. David Sacks amplified the concern by arguing that Anthropic is moving from the model layer into vertical applications such as science, security, legal, and coding, raising the fear that model providers will watch where value is being created and then move into those markets directly.The pull is that Karp is not alone, even if his style is unusually combative. Higgins notes that Satya Nadella has also warned that companies need to retain the learnings created when they use AI models, while Mark Zuckerberg has framed Meta's new model release partly around lower-cost frontier intelligence. The article presents Karp's campaign as one sign that established technology companies and large enterprises are trying to define where they fit when AI labs become central infrastructure, application competitors, and potential IPO giants at the same time.Read more: The Wall Street JournalThe AI Agents Are Coming for Microsoft OfficeAlex Wilhelm | Cautious Optimism | July 11, 2026Alex Wilhelm argues that one of the week's quieter AI questions is whether the productivity market that Microsoft successfully moved into subscription software is now being attacked by agentic tools. The piece begins with the infrastructure backdrop: SK Hynix raised $26.5 billion in a U.S. listing while building U.S. HBM and advanced-packaging capacity, and memory, chip, and foundry companies are now priced for sustained AI demand.Wilhelm then says the AI conversation has shifted quickly from raw capability to cost per task. He cites new model releases and vendor language emphasizing cheaper agentic and coding models, faster performance, and lower dollars per task. That matters because lower costs make it more plausible for AI systems to take on routine knowledge work at scale rather than remain a premium coding assistant market.The core of the article is Microsoft Office. Wilhelm notes that Microsoft turned Office from a one-time purchase into Microsoft 365, a large recurring revenue business with tens of millions of subscribers and a major productivity segment. Now, he says, late-stage unicorns and AI labs are pushing into the same territory: Anthropic's Cowork was reportedly used mostly outside software development, OpenAI merged ChatGPT and Codex into a tool for creating sheets, slides, docs, web apps, and long-running work, and other companies are building agentic coworkers that connect business data to documents, workflows, schedules, alerts, and apps.The article's caveat is that Microsoft has survived major platform shifts before. The argument is not that Office disappears quickly, but that the definition of office software is broadening from documents and spreadsheets into AI systems that can create, monitor, and act across workplace data.Read moreWhat Is Loop Engineering, and Who Owns It?Author: Nilesh Barla Published: July 11, 2026Nilesh Barla argues that “loop engineering” is becoming a distinct discipline because production AI agents now fail less at single prompts than at runtime: when to stop, what state to preserve, and how to recover after a bad step. Prompt engineering shapes one model call, and context engineering shapes what the model sees, but loop engineering shapes what a sequence of calls actually does.The killer detail is the three-primitives frame. Barla says a real agent loop needs halt conditions, state carryover, and recovery paths, then maps teams across five maturity levels. At the lowest level, an agent is just a model call in a for-loop with a step cap and raw history; by the higher levels, the system has structured state, explicit planning, replay, evaluation, and self-repair.The pull is organizational. If agents are becoming production systems rather than demos, someone has to own the runtime itself. The loop engineer is the role Barla gives to the person responsible for making long-running agent work dependable.Read more: Adaline LabsThe Fight Against AI Data Centers Is Just BeginningEmma Roth | The Verge | July 12, 2026Emma Roth argues that community resistance to data centers has moved from an early warning sign into a national political fight as AI facilities grow larger, more power-hungry, and more visible to nearby residents. The article starts with Apple's failed 2015 plan for a $1 billion data center in Athenry, Ireland, where a small group of residents challenged the project over noise, light pollution, flooding, traffic, and wildlife effects until Apple abandoned it in 2018.The current data-center buildout is presented as much larger and more contentious. Roth writes that residents now cite rising energy costs, water quality, noise, light pollution, and greenhouse gas emissions, while the U.S. Energy Information Administration expects commercial energy demand to surpass residential demand this year because of AI data centers and Goldman Sachs expects data-center power demand to double by 2027.The central evidence comes from Data Center Watch, which says protesters blocked or delayed at least 75 U.S. projects worth $130 billion from January to March, with active opposition groups more than doubling from 396 at the end of 2025 to 833 by the end of the first quarter of 2026. Roth also cites QTS abandoning a $12 billion Wisconsin campus, Delaware City regulators blocking a 580-acre project under the Coastal Zone Act, opposition stopping a QTS project in Prince William County, and pressure that pushed Kevin O'Leary to downsize the proposed 40,000-acre Project Stratos in Utah.The policy section describes a split between federal acceleration and local resistance. President Trump has treated data centers as part of the AI race with China and fast-tracked construction, while some Republican candidates are distancing themselves from that position ahead of midterms. Sanders and Ocasio-Cortez have proposed a moratorium until price and environmental protections exist, bipartisan lawmakers are backing ratepayer-protection measures, and states including Florida, Idaho, and Washington have passed rules on cost shifting, water use, and tax breaks. Roth's caveat is that the policy patchwork is still incomplete, leaving many communities to fight project by project.Read more6 months to live for open modelsAuthor: Nathan Lambert Published: July 12, 2026Nathan Lambert argues that open-weight AI models are facing their most serious policy test so far because U.S. officials are beginning to discuss concrete controls rather than abstract safety concerns. He says reported White House conversations about a new executive order may initially target Chinese-origin models and government use, but could create a broader review habit for frontier open models. His forecast is that a model above the capability range of GPT-5.5, Claude Opus 4.8, or GLM-5.2 could trigger a ban or indefinite delay within six months.The post separates two policy fights that are becoming intertwined: distillation and frontier capability. Lambert says the distillation campaign against Chinese models has become a form of regulatory capture because Anthropic and other closed-model companies would gain economically if Chinese open models were banned. He does not dismiss IP protection, but argues that if a closed model's capabilities are dangerous enough to justify restricting open models, the lab also has to explain why those capabilities are exposed through a queryable API. He cites unauthorized access to Anthropic's Mythos private beta as evidence that APIs are not automatically secure.The broader claim is that a unilateral U.S. ban would hurt positive actors more than bad actors if comparable open models remain available elsewhere. Lambert says the only durable ceiling would require global agreement, which does not exist, and that open models can improve safety by allowing broad inspection, adaptation, and understanding. His proposed near-term off-ramps are a strong U.S. open model release from companies such as Microsoft, Meta, or Reflection, and a broader coalition of open-source beneficiaries lobbying for safe rollout rather than prohibition.Read more: SourceAmericans Deserve a Dividend From AI Companies' RichesAuthor: Scott Stanford Published: July 14, 2026Scott Stanford argues that proposals to give the government a stake in AI companies miss the point unless ordinary citizens directly receive and control the upside. Sam Altman has discussed giving up equity in OpenAI, Washington already owns a stake in Intel, Nvidia is sharing China chip revenue, and Bernie Sanders wants large AI labs to contribute half their stock to a sovereign wealth fund. Stanford says those ideas all park value with the state, not with people.The killer detail is New Carlisle, Indiana, where AWS's Project Rainier is turning cornfields into one of the world's largest AI superclusters. The project is planned to run up to a million chips, draw more than two gigawatts of power, and represents an investment that has grown from $11 billion to $13.8 billion. Stanford uses that local transformation to argue that AI's public bargain should be visible at the household level.The pull is design. A citizen AI dividend would have to specify who earns a stake, how they hold it, and when they see cash. Without that mechanism, the AI wealth debate remains a fight over government balance sheets rather than public ownership.Read more: SourceWho Gets to Define the Frontier?Author: Mark Daley Published: July 14, 2026Mark Daley argues that Demis Hassabis is right to call for a serious institution to verify frontier AI systems, but that the power to test models is also the power to govern them. Hassabis's proposed Frontier AI Standards Body would get privileged pre-release access to advanced models, testing compute, held-out evaluations, support from national labs and security agencies, third-party auditors, and eventually authority to block models from the American market or coordinate a slowdown.The killer detail is Daley's constitutional objection. He says the proposal sometimes looks like a scientific lab, a standards body, an industry regulator, a licensing authority, and an emergency security council at once. Combining those roles because each requires technical expertise would be like putting the central bank, auditor-general, and Supreme Court in one building and calling it efficient.The pull is standard-setting. Daley's concern is not that verification is unnecessary, but that whoever writes the tests, decides what passes, adjudicates disputes, and grants market access may end up defining the frontier itself.Read more: SourceGPT-Red: Unlocking Self-Improvement for RobustnessOpenAI | OpenAI | July 15, 2026OpenAI describes GPT-Red as an internal automated red-teaming model trained to find prompt-injection vulnerabilities at a scale human red teams cannot match. The post says AI systems increasingly encounter third-party data through browsers, connected apps, local files, and tools, creating opportunities for malicious instructions hidden in emails, webpages, tool responses, or code repositories. Human red-teaming remains part of OpenAI's safety process, but the company says it is time-intensive and cannot generate enough diverse adversarial examples for model training.The system is trained through self-play reinforcement learning, with GPT-Red rewarded for eliciting valid failures and defender models rewarded for resisting attacks while still completing their tasks. OpenAI says the training environments specify threat models across settings such as local files, webpage banners, email bodies, and tool outputs. The model is kept separate from deployed production models because it is intentionally trained with malicious capabilities.OpenAI reports that GPT-Red generalized beyond its training set, including an internal replication of the indirect prompt-injection arena from Dziemian et al. (2025), where it found successful attacks in 84% of scenarios compared with 13% for human red-teamers. The post also says GPT-Red transferred attacks from simulation to a live autonomous vending-machine agent, causing price changes and order cancellations, and outperformed a prompted GPT-5.5 baseline against a Codex CLI agent on held-out data-exfiltration tasks.The article's main robustness claim is that OpenAI has used GPT-Red and predecessor models in training since GPT-5.3, with later GPT releases becoming more resistant to prompt injections. It says GPT-5.6 Sol has six times fewer failures on OpenAI's hardest direct prompt-injection benchmark than the best production model from four months earlier, that a “Fake Chain-of-Thought” attack class fell from more than 95% success against GPT-5.1 to below 10% against GPT-5.6 Sol, and that GPT-5.6 Sol fails on only 0.05% of GPT-Red's direct prompt injections. OpenAI says general capabilities and targeted over-refusal evaluations were not harmed, and says a preprint with more details will follow.Read moreAnthropic, Blackstone bet the next trillion-dollar AI business is implementation, not just modelsRebecca Bellan | TechCrunch | July 15, 2026Rebecca Bellan reports that Ode with Anthropic is the $1.5 billion AI implementation company launched by Anthropic with Blackstone, Hellman & Friedman, Goldman Sachs, and other backers. The article says the venture reflects a growing belief among frontier AI labs that enterprise adoption requires more than better models: customers need engineers who can embed inside businesses and turn AI into working systems.Ode was originally conceived by Blackstone after it used both large consulting firms and smaller AI services boutiques across its portfolio companies. TechCrunch reports that Fractional AI, an AI engineering services startup, stood out and was acquired by the joint venture shortly after the venture was announced. Fractional now forms the foundation of Ode, which has 100 engineers and works closely with Anthropic's applied AI team to identify where the technology can affect specific businesses.Ode CEO Chris Taylor tells TechCrunch that the company could someday become a trillion-dollar business if it scales without losing quality. He says an ideal customer is one whose CEO treats the AI project as a top one or two priority, whether it is a major product feature or the reworking of a core business process. Ode will operate under a “Claude-first” principle, using Anthropic technology whenever possible, but the article says it can use rival AI products when needed.The article's central implementation argument comes from Ode chief technologist Eddie Siegel, who says model selection matters but is not where most of the engineering effort goes. He compares it to the choice of programming language in software: one ingredient in a system that still has to be engineered. Bellan writes that Ode's challenge is hiring and training enough elite generalist engineers, many of them former founders, while competing with OpenAI's The Deployment Company and consulting giants that have built their own forward-deployed engineering teams.Read moreVint Cerf is working on a plan to unleash AI agents on the open internetTim Fernholz | TechCrunch | July 15, 2026Tim Fernholz reports that Vint Cerf, after leaving Google, is advising Innovation Labs on an open architecture for identifying AI agents online. Innovation Labs is a subsidiary of Identity Digital, a DNS registry company, and its proposal is to use domain-name infrastructure as part of a system for agent identity, accountability, and auditability. The premise is that agents will need a way to identify themselves if they move beyond proprietary systems and begin interacting across the open internet.The concrete proposal is DNSid, a registry that links an AI agent to an existing internet domain and uses cryptographic proofs to log its registration over time. Innovation Labs says it is trialing the standard with unnamed hyperscalers and identity companies. Cerf frames the problem around authority and accountability: what authority an agent has, where that authority came from, who is accountable for the agent's behavior, how its identity is established, and why anyone should trust it.The article's caveat is that standards are still emerging and agents are more active than static domains. Cerf says the period may be both fascinating and exasperating because the functionality is powerful and interoperability is unresolved. He compares the adoption problem to TCP/IP: competing systems may not work together until users push for functional interoperation. He also says an agentic economy is not inevitable, but that people will try to build it because delegating work to agents will be easier.Read more: TechCrunchxai-org/grok-build, now open sourceAuthor: Simon Willison Published: July 15, 2026Simon Willison argues that xAI's decision to open-source Grok Build is best understood as a trust repair move after a severe privacy failure. The CLI had triggered backlash when users realized that running it in a directory could upload the entire directory to xAI's Google Cloud buckets, including one user's reported SSH keys, password manager database, documents, photos, and videos. xAI disabled the feature, said previously retained coding data would be deleted, and released the code under Apache 2.0.The killer detail is what the codebase reveals. Willison counts 844,530 lines of Rust, only about 3% of which appears vendored, and finds remnants of the upload system still present but disabled: gcs.rs contains Google Cloud upload code, while upload_session_state() now returns a hard-coded session_state_upload_unavailable error. He also notes copied or ported tool implementations from Codex and OpenCode, prompt files, and a terminal Mermaid renderer.The pull is that terminal coding agents are becoming large, intricate software systems in their own right. The privacy failure mattered because these tools operate inside the directories where developers keep their most sensitive work; the open-source release matters because trust now depends on inspecting what an agent can see, send, and do.Read more: SourceThe Pulse: What can we learn from Bun's rapid Rust rewrite with AI?Author: Gergely Orosz and Ivan Klaric Published: July 16, 2026Gergely Orosz and Ivan Klaric argue that Bun's AI-assisted rewrite from Zig to Rust is a practical sign of how software engineering changes when models can take on large, bounded migrations with clear feedback loops. The piece does not treat the rewrite as magic: Jarred Sumner first spent hours turning design judgment into a detailed porting guide, then used adversarial review, parallel agents, compiler errors, and tests to force the work toward correctness.The killer detail is the scale. Bun had 535,496 lines of Zig, 1,448 files, and 22 million monthly downloads, making a conventional rewrite a year-long freeze the team could not justify. Using Fable, Sumner split the work across 64 agents, produced about 6,500 commits, and got the migration done in 11 days at an estimated API cost of $165,000.The pull is economic, not theatrical. If a one- or two-year migration can become an 11-day project, AI coding is not just faster autocomplete; it changes which technical debts are worth paying down.Read more: SourceOrphan risks at the frontier of artificial intelligenceAuthor: Andrew Maynard Published: July 16, 2026Andrew Maynard argues that frontier AI safety frameworks are creating “orphan risks”: harms that companies can see, but do not formally own because they are hard to quantify, do not fit catastrophic-risk thresholds, or fall outside audit-friendly compliance machinery. His target is not existing frontier safety work, but the narrowing effect that happens when private companies decide which risks count as governable.The killer detail is Maynard's contrast between measurable model dangers and threats to value. He points to Meta's three-day Galactica collapse, OpenAI's 2023 board crisis, safety-team departures, and wellbeing litigation as examples of risks that damaged trust, culture, legitimacy, or users without fitting cleanly into conventional model-risk categories. The proposed fix is an orphan-risk register: a public record of risks a company considered and chose not to manage, with reasons.The pull is accountability. Frontier developers' internal scoping choices have become a de facto layer of public governance, so the question is no longer only which risks they manage, but which risks they quietly leave outside the frame.Read more: SourceThe Lab of the Future Should Feel Like a Data CenterLatent.Space with Andy Beam and Rafa Gomez-Bombarelli | Latent.Space | July 16, 2026Latent.Space interviews Lila Sciences CTO Andy Beam and chief science officer for physical sciences Rafa Gomez-Bombarelli about the company's attempt to build an AI-run science factory. The post describes Lila's thesis as treating the lab itself as an “infinite token generator”: if internet data drove the first era of AI scaling, experimentally verified scientific data may be the next scarce training source. Lila is trying to produce that data with robotics, lab instruments, orchestration software, and AI models wired into the wet lab.The central analogy is the lab as data center. Instruments are nodes on a graph, a magnetically levitating transport layer moves materials between them, and experiment scheduling looks like a compute queue. Beam says Lila is not simply an automation company, because the point is not just throughput; it is flexibility, generalization, and experiment capture. The post says Lila has built more than 10 trillion experimentally validated “scientific reasoning tokens,” not internet text or biological sequences.The interview ranges across biology, chemistry, drug discovery, materials science, and the limits of automation. It notes that Lila rebuilt one gas-sorption measurement to run roughly 2,500 times faster, claims its general models can transfer priors from small-molecule chemistry to metal-organic frameworks for carbon capture, and describes model-suggested platinum-group-free electrocatalysts that moved from looking boring or wrong to becoming strong performers. The caveats are physical: experiments have runtimes, biology cannot always be accelerated, chains of thought can be unreliable narrators, and reward hacking becomes more dangerous when a model controls a real lab.Read more: Latent.SpaceWhy AMI Labs' Alexandre LeBrun won't call his AI “AGI” or “superintelligence”Kate Park | TechCrunch | July 16, 2026Kate Park interviews AMI Labs CEO Alexandre LeBrun about why Yann LeCun's world-model startup avoids the language of “AGI” and “superintelligence.” LeBrun says the terms are not useful because they lack stable definitions: “We never used the word AGI. And I just noticed that nobody is using it anymore; they switched to superintelligence.” His argument is that the practical frontier is not a label, but whether AI systems can understand and predict real-world states.The article explains the world-model thesis by contrasting language prediction with physical-state prediction. A large language model predicts the next word; a world model predicts the next state, such as what happens when a glass tips over. LeBrun says LLMs remain complementary and efficient for language, but the physical world is where current AI is weak. Robotics is the clearest case: hardware has advanced quickly, but robots are still brittle outside controlled routines because they lack context and situational understanding.AMI is still pre-product, but TechCrunch reports that LeBrun was in Seoul looking for industrial partners, researchers, and global companies. He says world models cannot be built entirely inside a lab because they need access to real environments. That is why South Korea appeals to AMI: robotics, semiconductors, manufacturing, and fast adoption create the kind of hardware-heavy context that software-only AI has barely touched.Read more: TechCrunchKimi K3 Tech Blog: Open Frontier IntelligenceKimi | Kimi | July 16, 2026Kimi introduces Kimi K3 as an open 3T-class frontier model aimed at coding, knowledge work, reasoning, multimodality, and long-context agentic use. The source describes the model as a 2.8T-parameter system built on Kimi Delta Attention and Attention Residuals, with native multimodality and a 1M-token context window. It says Moonshot AI plans to release model weights by July 27.The post presents K3 through benchmark and use-case sections rather than as a general product announcement. It reports results across coding, productivity, agentic, and multimodal evaluations, including DeepSWE, Terminal-Bench 2.1, Program Bench, SWE Marathon, FrontierSWE, PostTrain Bench, OfficeQA Pro, SpreadsheetBench 2, MCP Atlas, AutomationBench, BrowseComp, GDPval-AA v2, AA-Briefcase, MMMU-Pro, MathVision, BabyVision, OmniDocBench, and PerceptionBench. The source says all reported K3 results use maximum reasoning effort with temperature and top-p set to 1.0, and that different benchmark comparisons use KimiCode, Claude Code, or Codex harnesses depending on the test.Kimi's caveats are unusually concrete. The limitations section says K3 was trained in preserved thinking-history mode, so quality may become unstable if an agent harness does not pass historical thinking content correctly or if an ongoing session switches to K3 midstream. It also says K3's emphasis on long-horizon tasks can make it excessively proactive when it encounters minor issues or ambiguous intent, and recommends imposing explicit behavioral constraints for applications that require strict boundaries. The post adds that K3 remains behind Claude Fable 5 and GPT 5.6 Sol in user experience despite being competitive overall.Read moreVenture CapitalThree Years InAuthor: Tomasz Tunguz Published: July 10, 2026Tomasz Tunguz marks Theory Ventures' third anniversary by arguing that AI's central market effect is time compression. In his telling, model release cycles, company revenue milestones, enterprise adoption, and venture categories have all accelerated. Seed, Series A, and Series B still exist as financing labels, but they no longer cleanly describe company maturity when some seed rounds are larger than IPOs and the best AI companies can mature much earlier than prior software companies.The killer detail is the shift from models to inference. Tunguz argues that inference has become the dominant AI market because workloads and buyer preferences are fragmenting: video, batch, local, agentic, and real-time tasks each create different infrastructure needs. He compares this to databases splitting into OLTP, OLAP, vector, and streaming categories, with AI pushing the same specialization into inference infrastructure.The pull is that Theory sees the AI-native venture firm as part of the same pattern. The firm says it has analyzed twice as many investment opportunities with three investors working alongside a nine-person intelligence organization, using agents and research systems to map markets, source companies, and support diligence. The piece is both a market map and a statement about how venture itself is being rebuilt by the technology it funds.Read more: LinkedInVenture Has Rarely Looked More BifurcatedAuthor: Beezer Clarkson Published: July 14, 2026Beezer Clarkson points to PitchBook's Q2 report as evidence that the U.S. venture market has split into two very different realities. AI now accounts for more than 60 percent of all U.S. venture deal value, meaning the headline market can look active and well-funded even while much of the non-AI market is dealing with a much colder liquidity and fundraising environment.The thread uses that split as the setup for Clarkson's latest Origins episode with Alec Litowitz, founder of Magnetar and QStar Capital and one of Citadel's original founding partners. Clarkson says markets like this are periods of genuine uncertainty, not merely ordinary risk, which is why Litowitz's Adaptability Quotient framework is relevant.The embedded clip makes the liquidity point concrete. Litowitz says DPI is “the resolution of uncertainty” because it converts an uncertain investment into actual cash returned to LPs. In his framing, a realized dollar is a real mark, while TVPI remains uncertain until it is realized.The killer detail is the distinction between pricing risk and resolving uncertainty. Litowitz's perspective matters because QStar is a SpaceX investor and Clarkson says the conversation happened just before one of venture's most consequential IPOs. The episode's stated questions are why venture remains a way to gain exposure to innovation, how AI is changing what is investable, why liquidity is ultimately a function of time, and why uncertainty requires a different decision framework from risk.Read more: XThe Best Angel Investors in the US: Who Backs the Most Unicorns, and Who's Active NowAuthor: Ilya Strebulaev Published: July 10, 2026Ilya Strebulaev ranks angels, angel groups, accelerators, and incubators by lifetime U.S. unicorn investments, counting checks written before a company reached unicorn status. The top of the combined list is dominated by organizations: Y Combinator leads with 113 unicorn investments, followed by Plug and Play at 52 and 500 Global at 41. Sand Hill Angels is the highest-ranked angel group at 31.The killer detail is how quickly the list changes below the biggest accelerators. Strebulaev says 271 of the 304 investors in the Top 200 are individuals, or 89%. In the top 100, individuals are 91%. That makes the market underneath the large accelerator counts look much more personal: mostly operators and individual angels writing early checks from their own networks.The pull is the ranking's own caveat. Strebulaev writes that every lifetime leaderboard has a blind spot because many of the unicorns behind those totals were founded a decade or more ago, and some angels have since moved into formal funds, slowed down, or stopped investing. His post therefore separates lifetime performance from recent cohorts, including companies founded in 2015 or later and 2020 or later. For founders or allocators making current decisions, that distinction matters: a career record and a current record are not the same measure.Read more: Ilya StrebulaevAre Prediction Markets Doomed to Fail?Author: Contrary Published: July 16, 2026Contrary argues that prediction markets' current boom depends on whether platforms can prove they are more than regulated gambling with exchange-style branding. Kalshi and Polymarket have reached mass cultural, investor, and regulatory attention, but the article says the underlying idea is old: academic markets, corporate forecasting tools, Intrade, PredictIt, and other predecessors all struggled with the same linked problems of liquidity, legality, and user appeal.The killer detail is the comparison with sportsbooks. Prediction markets present themselves as peer-to-peer, transparent, and non-house-based, but sports contracts reportedly account for more than 90 percent of Kalshi trading, and the article says the platforms keep a much thinner slice of volume than sportsbooks. A market can therefore show sports-betting-scale handle while generating far less revenue.The pull is that the product's hardest problem may be distribution of wins. If a small group of sharp traders captures most profits while casual users lose interest, prediction markets may become valuable data feeds and professional tools before they become durable consumer networks.Read more: SourceRegulationExclusive: The Next Frontier of the Deportation Wars: College CampusesAuthor: Adrian Carrasquillo Published: July 11, 2026Adrian Carrasquillo reports that college campuses are becoming a new front in the fight over immigration enforcement because automatic license plate readers can turn ordinary campus security infrastructure into searchable location data. His thesis is that Flock Safety's camera network, even without direct ICE or DHS contracts, can feed deportation enforcement through local police partnerships and data-sharing practices.The killer detail is the campaign target. The Emergency Campaign to Support Higher Education, working with Schools Drop ICE, is focusing on 75 colleges and universities publicly identified as having Flock contracts. Flock says it has no ICE or DHS contracts, but activists argue the risk comes through local agencies that coordinate with federal authorities and run searches on their behalf.The pull is broader than immigration. Carrasquillo notes that license plate readers have already been abused by officers for stalking, and that Flock's AI search features can identify more than plates, including bumper stickers. A campus safety tool can become a political surveillance system when the data layer is searchable.Read more: The BulwarkThe Supreme Court Broke Independent Agencies. Here's a Way to Slow the Damage.Author: Todd Phillips Published: July 12, 2026Todd Phillips argues that the Supreme Court's decision in Trump v. Slaughter damaged independent agencies by ending for-cause removal protections, but did not leave Congress powerless. The ruling weakens the old model in which commissioners at bodies such as the FTC, NLRB, CPSC, SEC, and CFTC could be insulated from dismissal over policy disagreements. Phillips says the next fight is whether presidents can turn nominally bipartisan commissions into one-party instruments.The killer detail is the procedural fix: quorum rules. Phillips proposes that Congress require bipartisan slates of commissioners to be seated before independent agencies can act. A president could still fire commissioners, as the Court now permits, but if those firings broke quorum, the agency would be unable to proceed until replacements were confirmed. The guardrail would
Serious gamers, meet ASUS ROG XG. This SoundBytes show dives into how ASUS ROG XG's Thunderbolt dock lets handheld systems tap into big, cutting-edge GPUs, so you can enjoy high-resolution, modern titles far from your desk. The post HANDHELD GAMING WITH DESKTOP POWER! appeared first on sound*bytes.
Jetzt wird es ernst für deutsche Krypto-Anleger: Der Wegfall der einjährigen Haltefrist rückt näher, womöglich sogar rückwirkend zum Jahresanfang 2026. Klingt nach Steuerhammer, doch im Kleingedruckten stecken überraschend auch Gewinner. Wen die Reform wirklich trifft, warum Vieltrader am Ende sogar sparen könnten und ob ein rückwirkender Stichtag vor Gericht überhaupt hält, darüber sprechen Julius Nagel und Florian Adomeit in dieser Folge von Alles Coin, Nichts Muss.
Enterprise AI is easy to demonstrate. The real test begins when a promising POC meets production costs, security requirements, data movement, latency, and internal adoption.Shimon Ben-David, CTO at WEKA, joins Amir to discuss the gap between experimenting with generative AI and operating it at scale. They explore how classical AI differs from generative AI, why production exposes problems that demos hide, and how companies with limited AI maturity can start building useful internal capability.Practical Takeaways• A successful POC proves that an outcome is possible. It does not prove that the system will be affordable, secure, reliable, or fast at scale.• Enterprise AI adoption reaches across infrastructure, engineering, data, security, and business teams. It cannot be owned by one group in isolation.• Adding more GPUs will not fix slow data access, poor utilization, weak pipelines, or an experience users do not want to use.• External support can help, but the person or firm involved needs to stay through implementation and production, not stop at recommendations.• Companies that are behind should begin with proven use cases, build internal experience, and quickly stop experiments that fail to show value.Key Moments00:00 Why moving enterprise AI into production remains difficult01:55 The difference between classical AI and generative AI adoption07:05 How companies can use AI without having a formal AI strategy11:35 Why successful POCs often struggle when they reach production17:35 Competitive pressure, AI FOMO, and the need to calculate real ROI22:00 Why AI adoption requires cross organizational change33:10 Where a company with limited AI maturity should beginOne Line That Stuck“The promise is there. It is possible. You just need to do it properly.”Subscribe to The Tech Trek for more conversations about how technical teams are building, operating, and adapting around AI, data, product, platform, and engineering execution.
Is the AI industry actually overbuilding, or is the physical world moving too slowly to keep up? In this episode of the MAD Podcast, OpenAI's Head of Industrial Compute, Sachin Katti, takes us inside the "belly of the beast" of what may be the largest infrastructure project in human history. We explore the staggering physical reality of the AI boom—from $50 billion supercomputers and liquid-cooled data centers that "turn electrons into tokens," to overhauling the U.S. power grid and exploring nuclear energy. Sachin also pulls back the curtain on OpenAI's Stargate strategy, their move into custom silicon with Project Jalapeno, and the mind-bending reality that AI is now beginning to design the very chips that will power its own future.(00:00) — Cold open: “One of the largest things humanity has ever built”(00:30) — Welcome: Sachin Katti, Head of Industrial Compute at OpenAI(01:44) — Is this the biggest infrastructure buildout in history?(03:41) — Why OpenAI is building a new industrial muscle(04:54) — What an AI data center actually is(05:27) — “Factories turning electrons into tokens”(06:35) — Why AI data centers need liquid cooling everywhere(08:10) — The power problem: grids, generation, transmission, substations(10:43) — Behind-the-meter power and gas turbines(11:02) — Why nuclear “can't come soon enough”(11:49) — Jalapeño: why OpenAI is designing its own AI chips(13:19) — Tokens per watt: the new metric that matters(13:38) — Why inference may now dominate AI compute(14:58) — Is OpenAI overbuilding compute?(16:47) — Why OpenAI thinks the bigger risk is not building fast enough(17:55) — Communities, jobs, water, and the local data-center debate(21:16) — How OpenAI chooses data-center sites(22:25) — What “industrial compute” means inside OpenAI(25:59) — Sachin's path: Stanford, startups, Intel, OpenAI(28:05) — OpenAI's compute portfolio: Microsoft, hyperscalers, neoclouds(29:37) — Stargate explained(31:21) — Abilene, Oracle, and the next wave of AI data centers(32:48) — How massive AI compute gets financed(34:05) — How OpenAI designed Jalapeño so quickly(35:59) — AI is starting to help design AI chips(36:20) — MRC: the networking problem behind 100,000 GPUs(38:47) — Bottlenecks: transformers, turbines, electricians, supply chains(40:29) — Guaranteed capacity: intelligence as a supply unit(42:08) — Will AI data centers move to space?
As part of our summer replay series, we're revisiting one of our favorite conversations on the future of AI infrastructure. SemiAnalysis founder Dylan Patel joins Erin Price-Wright, Guido Appenzeller, and Erik Torenberg to examine the rapidly evolving economics of AI hardware, from GPUs and custom silicon to data centers, power, and the global race for compute. The conversation explores NVIDIA's competitive advantages, the rise of custom chips from Google, Amazon, and Meta, the economics of frontier AI models, and the infrastructure constraints shaping the industry's next phase. They also discuss AI startups, export controls, robotics, enterprise software, and why simply copying NVIDIA isn't enough to build a winning AI hardware company. Whether you're building AI products, investing in infrastructure, or trying to understand where the industry is headed, this conversation offers a practical look at the forces shaping the future of compute. Resources: Follow Dylan Patel on X: https://x.com/dylan522p Follow Erin Price-Wright on X: https://x.com/espricewright Follow Guido Appenzeller on X: https://x.com/appenz Learn more about SemiAnalysis: https://semianalysis.com/dylan-patel/ Stay Updated:Find a16z on YouTube: YouTubeFind a16z on XFind a16z on LinkedInListen to the a16z Show on SpotifyListen to the a16z Show on Apple PodcastsFollow our host: https://twitter.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Companies are spending billions building AI factories, but most of them can't tell you why their AI workloads are failing, whether their GPUs are actually being used, or what their infrastructure is going to cost them when agents start running at scale. Paul Appleby, CEO of Virtana, joins Craig Smith to discuss the findings of their AI Factory Reality Check study, a research report that reveals a striking and underappreciated gap between the pace of AI infrastructure investment and the governance needed to run it safely and efficiently. Six in ten enterprises, the study found, cannot automatically identify root cause when an AI workload fails, a problem that compounds fast once you're running critical services on AI infrastructure at scale. The conversation covers the mechanics of Virtana's observability platform, capturing 20,000 metrics per second across the entire AI stack, correlating them in real time, and increasingly using agentic capabilities to remediate failures automatically, but its most important insights are structural. Appleby makes a sharp observation that cuts through a lot of AI optimism: token costs are falling, but token consumption is exploding, meaning the total cost of running agentic AI systems is still going up even as the per-unit price drops. He also tracks a cultural shift inside enterprises - IT resilience reporting that used to happen annually now happens weekly - as evidence that technology risk has become a board-level conversation in a way it simply wasn't before. The result is a conversation that's less about the promise of AI and more about what it actually takes to make it work at production scale. Subscribe to Eye on A.I. for weekly conversations with the people building and deploying the future of AI.
CNBC reported that Apple is in talks with a startup that specializes in compressing AI models to run on iPhones. The move aligns with Apple's strategy to execute more AI locally, reducing reliance on cloud infrastructure and lowering latency. On-device AI relies on techniques such as quantization, pruning, and distillation to fit models within CPU, GPU, and Neural Engine limits. The shift could reduce cloud inference costs that depend on GPUs from providers like Amazon Web Services, Microsoft Azure, and Google Cloud. Competitors including Google, Samsung, Qualcomm, and Meta are advancing on-device AI capabilities. Founders should benchmark compact models, assess battery impact, and decide which features should run locally versus in the cloud.Learn more on this news by visiting us at: https://greyjournal.net/news/ Hosted on Acast. See acast.com/privacy for more information.
Demis Hassabis proposed a US-based frontier AI standards body modeled on FINRA. IBM's stock cratered 20% on a Q2 miss from chip-spending shifts, Spotify launched a voice-control feature, Kalshi debuted an AI compute forward curve, and Anthropic studied Claude's values. Demis Hassabis proposes a US-based Standards Body for "Frontier-class" AI, modeled after FINRA; labs would share models for review up to 30 days before release (X) Demis Hassabis proposes a US-based Standards Body for "Frontier-class" AI, modeled after FINRA; labs would share models for review up to 30 days before release (The Verge) IBM reports preliminary Q2 revenue up 1% YoY to $17.2B, below $17.9B est., as CEO Arvind Krishna says customers are shifting spending to chips; IBM falls 20%+ (Bloomberg) Spotify launches a Talk to Spotify feature that lets users create playlists and more, rolling out in beta to Premium users 18+ in the US, Ireland, and Sweden (Engadget) Kalshi launches a forward curve tool for AI compute, using event contracts to track the future rental costs of GPUs, storage, and memory (Bloomberg) Simulating everything, sort of: The promise and limits of world models (Ars Technica) Subscribe to the ad-free feed. Learn more about your ad choices. Visit megaphone.fm/adchoices
As part of our summer replay series, we're revisiting one of the standout conversations from Runtime, a16z's conference on AI infrastructure and the future of computing. Gavin Baker, Managing Partner and CIO of Atreides Management, joins David George to examine the biggest questions surrounding today's AI investment cycle. Is AI a bubble? What does the unprecedented buildout of data centers, GPUs, and compute infrastructure mean for the economy? And how should investors think about the companies building the next generation of AI? The conversation explores frontier models, Nvidia, Google, custom silicon, AI infrastructure, application software, robotics, and why Baker believes today's AI investment cycle looks fundamentally different from the internet bubble of the early 2000s. Along the way, they discuss the economics of GPUs, enterprise software, AI business models, and what comes next as AI moves from experimentation into the broader economy. Resources: Follow Gavin Baker on X: https://x.com/GavinSBaker Follow Atreides Management on X: https://x.com/atreidesmgmt Follow David George on X: https://x.com/DavidGeorge83 Stay Updated:Find a16z on YouTube: YouTubeFind a16z on XFind a16z on LinkedInListen to the a16z Show on SpotifyListen to the a16z Show on Apple PodcastsFollow our host: https://twitter.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Not all work happens in writing. Teams that work with photos, videos, and audio need AI that works for them too. This is why, with Dropbox, you can search within multimedia content for key moments and important information—not just text. In this episode, we talk with Appu Shaji and Hicham Badri, two Dropbox machine learning engineers who are part of the team that makes all of this possible. They explain how multimodal search works—from understanding the context of the initial query, to identifying objects and actions in complex scenes—and how they ensure those models work fast, even at Dropbox-scale. ~ ~ ~ Working Smarter is brought to you by Dropbox. Find, organize, and share your work—all in one place—with context-aware AI from Dropbox. You can listen to more episodes of Working Smarter on Apple Podcasts, Spotify, YouTube, Amazon Music, or wherever you get your podcasts. To read more stories and past interviews, visit workingsmarter.ai This show would not be possible without the talented team at Cosmic Standard: producer Ben Montoya, sound engineer Aja Simpson, technical director Jacob Winik, and executive producer Eliza Smith. Special thanks to our illustrator Fanny Luor, marketing consultant Meggan Ellingboe, and editorial support from Catie Keck. Our theme song was composed by Doug Stuart. Working Smarter is hosted by Matthew Braga. Thanks for listening!
The semiconductor industry is undergoing one of its most profound transformations in decades. Driven by the insatiable demand for compute power largely fueled by AI workloads, engineers are moving away from traditional monolithic chips and shifting toward complex multi-die designs. This shift brings a new set of challenges that conventional design and validation methods simply cannot handle.In a recent episode of the Tech Transformed podcast, host Dana Gardner sat down with Shekhar Kapoor, Executive Director of Product Line Management at Synopsys, to explore how the growing complexity of semiconductors is changing the way engineers design and validate modern systems. From thermal management to AI-driven automation, the conversation reveals why the old way of building chips is no longer good enough and what the future looks like.Multi-Die DesignKapoor explains that the transition to multi-die design is no longer a matter of preference but a necessity. He attributes this shift to the relentless demand for greater compute capacity, driven largely by the rapid growth of AI.Traditional monolithic chips are hitting hard limits. Reticle sizes are maxing out, and rising yield and cost challenges make it increasingly impractical to pack more functionality onto a single die. Multi-die designs solve this by disaggregating functionality across smaller dies, each targeting the most appropriate process technology, then integrating them into a unified, optimised package.Leading AI systems already integrate multiple compute and I/O dies alongside large high-bandwidth memory (HBM) stacks, scaling to 3x–5x reticle-class designs and beyond. The design challenge is very different. As Kapoor puts it: "You're no longer optimising a single chip, you're optimizing a system of chips."This requires system-level co-design from day one, spanning architecture, silicon, packaging, power delivery, and interconnect strategy simultaneously. Engineers must think in terms of System Technology Co-Optimisation (STCO), not just chip-level optimization. The design tools, methodologies, and team workflows all need to change. For engineers and technology leaders looking to explore these trade-offs, Synopsys has published a comprehensive eBook on accelerating multi-die design and innovation.Thermal Analysis and Multi-Physics ValidationHistorically, thermal, power, and electromagnetic analyses were performed as downstream validation steps once the core design was complete. In a multi-die world, that approach is no longer viable."Thermal management is becoming the number one issue when designing these multi-die designs. It has to be managed across a range of scales, from transistor activity to package and board level," Kapoor says.The problem with late-stage validation is timing. By the time thermal or power integrity issues surface, the most critical decisions are already locked in floorplans, interconnect topologies established, and packaging assumptions embedded.. At that point, the only options are costly ECOs, excessive margining, or a full redesign. Industry estimates suggest over-design can lead to up to 30-35 per cent wasted silicon and hundreds of millions of dollars in optimisation loss.The solution is a shift-left approach that embeds multiphysics analysis from the earliest stages of design. When thermal hotspots, voltage drop issues, and electromagnetic interactions are identified early, engineers can adjust partitioning and placement strategies before they become expensive problems.This is the methodology detailed in the Synopsys ebook on Multiphysics Fusion for multi-die design, which covers how teams can build continuous multiphysics validation into their flows to avoid late-stage surprises and protect both performance and reliability.Multiphysics Fusion and AI-Driven Chip DesignTo operationalise the shift-left methodology at scale, Synopsys has introduced the concept of Multiphysics Fusion. This is the native integration of AI-powered EDA technologies with ANSYS's gold-standard multiphysics sign-off analysis capabilities.Within the 3DIC Compiler platform, this means unifying the implementation environment with RedHawk-SC, RedHawk-SC Electrothermal, and HFSS-IC technologies. This brings IR drop, thermal, signal, and power integrity analysis directly into the design loop. The result is greater predictability, tighter correlation between in-design analysis and sign-off, and significantly fewer design iterations.The impact on design closure times has been substantial. According to Kapoor, teams using the Multiphysics Fusion solution have seen turnaround times shrink "from weeks to days, and in some cases even hours" even for large, high-performance multi-die designs.AI amplifies these gains further. Synopsys employs AI in two primary ways: assistive automation through its 3DSO.ai technology, which integrates multiphysics feedback into the optimization loop in real time, and agentic workflow orchestration, which becomes increasingly critical as system complexity scales toward designs incorporating hundreds or even thousands of GPUs. As Kapoor notes, at that scale, "agentic workflows could help engineers converge faster" and manage trade-offs that would otherwise be intractable. If you would like to find out more about this, download the full eBook: Multiphysics Fusion Technology for Multi-Die Designs Explained from Synopsys, which expands on each of these themes with real-world examples, design methodologies, and guidance for implementation teams. You can also connect with Shekhar Kapoor on LinkedIn.TakeawaysMulti-die architectures and their drivers.Challenges of traditional monolithic chips.Importance of early multi-physics analysis.Multiphysics fusion and its benefits.AI's role in design automation.Reducing time-to-market through integrated platforms.System-level co-design.Thermal management in 3D IC stacking.Shift left approach in multi-physics validation.Future trends in semiconductor design.Chapters00:00 Introduction to Semiconductor Complexity02:00 The Shift to Multi-Die Designs04:30 Challenges in Multi-Die Design08:11 The Importance of Early Multi-Physics Analysis10:05 Introducing Multiphysics Fusion12:37 AI's Role in Semiconductor Design16:37 Reducing Time to Market19:39 Applications Beyond AI21:12 Real-World Examples of Multi-Physics Validation26:20 Practical Advice for Engineers
PARTNERED Artificial Intelligence is advancing faster than ever—but can this pace continue? In this episode of Techcetra partnered with @salesforce , Leslie D'Monte speaks with Neil Thompson, Director of the FutureTech Research Project at MIT, to explore the economics driving today's AI revolution. While recent breakthroughs have transformed what AI can do, Neil explains that much of this progress has come from scaling compute—using more GPUs, more power and larger investments. But is that model sustainable? What happens when compute demand outpaces supply? The conversation dives into the economics of AI, chip shortages, NVIDIA's dominance, large language models, small language models, Agentic AI, AI in scientific research, digital labour, and the future of AI infrastructure. Neil also shares his perspective on India's AI ambitions, semiconductor manufacturing, local language models and why continuously improving smaller models may be a smarter long-term strategy. In this episode: Why we're living through AI's "golden age" The economics behind AI scaling Why chip shortages could continue for years Compute vs algorithmic innovation LLMs vs SLMs The rise of Agentic AI AI's role in scientific research Can AI compete with human labour? NVIDIA, GPUs and the future of AI infrastructure India's AI roadmap and semiconductor ambitions Whether you're an AI enthusiast, business leader, developer, researcher or policymaker, this conversation offers a practical look at the technological and economic forces shaping the future of artificial intelligence. Subscribe to Mint Techcetra for more conversations with global technology leaders, researchers and innovators. #ArtificialIntelligence #AI #AgenticAI #MIT #NeilThompson #Techcetra #MachineLearning #NVIDIA #Semiconductors #GenerativeAI #LLM #SLM #AIInfrastructure #IndiaAI #FutureOfAI Learn more about your ad choices. Visit megaphone.fm/adchoices
Send us Fan MailIn this episode of the WTR Small-Cap Spotlight Podcast, Jolienne Halisky, CFO of AIB Data Centers Inc. (NYSE American: AIB), joins host Tim Gerdeman and WTR equity research analyst James Kisner. AIB is a power-first developer of AI and high-performance computing infrastructure. The company locks up executed utility power agreements before it breaks ground, then builds modular, liquid-cooled facilities that tenants lease and fill with their own GPUs. Halisky explains why power, not chips, is the real bottleneck in AI infrastructure, and how securing it first lets customers deploy compute months or years sooner. She covers the recent name change from BlockchAIn Digital Infrastructure, the shift from Bitcoin hosting to AI workloads, and how the build-to-suit model keeps hardware obsolescence off the balance sheet. Halisky also walks through the roughly 40 MW live today, a pipeline approaching 395 MW, the debt-free capital strategy, and the milestones investors should watch.
We've been running a bit of an Agent Cloud series surveying all the top inference/compute/cloud providers, from Databricks to Daytona to Railway and, even further back, E2B, but we're excited to conclude this series returning to Modal, which has just raised a monster $355M Series C.The cloud was built for developers. But agents are now changing that.The old infra stack was designed for a human who could read docs, reason through YAML, and understand dashboards to figure out what they need when something broke. While this was painful for developers, it worked since they could fill in missing context in their heads.However, agents don't have that luxury. Now in this new era of agents, everything has to be tighter.They need a place to write code, run it, inspect the output, change the environment, debug failures, and try again. Fast iteration and feedback loops with all the necessary context are crucial for agents to operate properly. Furthermore, sandboxes are a clear representation of this shift as agents can easily spin up isolated environments. This programmatic infra even extends to research:Two years ago, we were one of the first to cover Modal with CEO Erik Bernhardsson and Alessio designed our favorite LS thumbnail of all time:At the time, Modal was just a teeny little company with a $17M Series A.Today, fresh off their $355M Series C, Modal is one of the clearest examples of the agent cloud future being built in real time: a cloud platform moving past traditional web app assumptions toward the workloads AI actually creates such as elastic inference, sandboxes, GPU burst, post-training, background agents, and infrastructure that agents themselves can operate.In this episode, Modal CTO Akshat Bubna joins swyx and Vibhu to unpack why AI applications don't fit traditional cloud assumptions, why Kubernetes was never designed for bursty compute-heavy workloads, and why Modal is now shifting from developer experience to agent experience.We go deep on Modal's AI infra stack: serverless functions, decorator-based infrastructure, elastic inference for custom models, GPU snapshotting, DeFlash, speculative decoding, Auto Endpoints, sandboxes, persistent storage, networked containers, private IPv6, RDMA, multi-node training, and Modal's capacity pool across 17 cloud providers. Akshat also explains why RL rollouts can require 100,000 sandboxes, why production agents need hard guardrails, why observability may matter more than reading code, and why AI has made infrastructure exciting again.We discuss:* Why Kubernetes wasn't built for bursty AI workloads* How Modal started as a better runtime before becoming an AI cloud* Why Modal added GPUs before ChatGPT* The shift from developer experience to agent experience* Why observability matters when agents are writing the code* Elastic inference for custom models across audio, video, robotics, and comp bio* GPU snapshotting, cold starts, and why inference workloads are so bursty* Why RL rollouts can require 100,000 sandboxes* DeFlash, speculative decoding, and frontier-level inference performance* Auto Endpoints and making optimized inference easier to deploy* What Modal adds beyond vLLM, SGLang, and raw GPU rental* Modal's 17-cloud capacity pool and supercloud strategy* Networked sandboxes, sidecars, private IPv6, and RDMA* Serverless multi-node training for post-training and research workloads* Auto-research, model-guided sweeps, and agents launching GPU experiments* Compute strategy, capacity planning, and batch tiers* Why production agents need specialized sandboxes and hard guardrails* Modal's take on managed agents, CI, Gitpod/Ona, Python, TypeScript, and Modal BenchAkshat Bubna* LinkedIn: https://www.linkedin.com/in/akshat-bubna-188885103* X: https://x.com/akshat_bModal* Website: https://modal.comTimestamps00:00:00 Introduction00:00:39 Modal's origin and why Kubernetes wasn't enough00:04:32 Developer Experience → Agent Experience00:06:21 Modal's AI cloud primitives00:09:14 Sandboxes, agent loops, and proto-Cognition00:12:12 Elastic inference, GPU snapshotting, and 100,000 sandboxes00:15:24 DeFlash, speculative decoding, and Auto Endpoints00:19:59 Production-grade inference beyond raw GPUs00:22:00 Background agents, Ramp Inspect, and the agent lifecycle00:24:08 Modal's 17-cloud supercloud strategy00:26:40 Networked sandboxes, private IPv6, and RDMA00:32:48 Multi-node training, post-training, and auto research00:37:36 Compute strategy, capacity planning, and batch tiers00:40:55 Open models, real-time AI, and production agent infra00:43:06 Hard guardrails, managed agents, and specialized sandboxes00:46:06 Why AI made infrastructure exciting again00:48:30 Model APIs, differentiated products, and agentic video00:51:50 CI, coding-agent infra, SDKs, and Modal Bench00:57:28 Closing ThoughtsTranscriptIntroduction: Modal, Series C, and the Art PartySwyx [00:00:00]: We're here with Akshat, CTO of Modal, together with Vibhu. Congrats on your Series C.Akshat [00:00:10]: Thank you.Swyx [00:00:11]: Your party yesterday was amazing.Akshat [00:00:15]: Yeah.Swyx [00:00:15]: From all the photos and all the swag.Akshat [00:00:17]: We had a bunch of art installations, which was fun, seeing, like, our products on pedestals next to, like, Rodin.Swyx [00:00:25]: Very nice. Very nice. When you started, it was not the GPU inference company. Maybe it was in your mind. Take us back to the origin story.Modal's Origin: A New Runtime Beyond KubernetesAkshat [00:00:39]: I first met Eric, who's the CEO, through an investor. Back then Eric was already thinking about building, a new runtime, and he got there thinking through why are workflow orchestration products so hard to use. It's because you have to run them on Kubernetes. Kubernetes is hard to manage. It's not built for burstiness and, custom images,Swyx [00:01:03]: YeahAkshat [00:01:03]: It has a terrible developer experience.Swyx [00:01:05]: And I'll, I'll interjectAkshat [00:01:06]: YeahSwyx [00:01:07]: For listeners, who are new, we interviewed Eric two years ago, and there's a bit more of the story there from Spotify and all those things.Swyx [00:01:14]: And I came across Eric through Data Council because he did that talk on the serverless container stack that you guys did, which was like, that was my first like, “Okay, I need to take Modal very seriously” moment.Akshat [00:01:26]: Yeah.Swyx [00:01:26]: But it was still very unclear, like, do I need all this for just my data pipelines?Akshat [00:01:33]: Yeah. initially what we were thinking about was if we build a better runtime, it's a very useful primitive in itself. It's There's a lot of things that, get solved by serverless functions, like you can do, ETL stuff, you can do job queues, you can do all this, like, bursty processing, which it turns out every company had needs for. but then we also were thinking about this as like, this is a primitive that we can build a whole collection of products on, which are very verticalized. So perhaps data engineering would've been the first one, but we were thinking about inference. Back then it was more classical inference, like computer vision stuff and running XGBoosts and whatnot. But we added GPUs to the product a year before ChatGPT came out.From Serverless Containers to GPU WorkloadsSwyx [00:02:19]: Nice.Akshat [00:02:19]: We just didn't think it would be that big of a deal.Swyx [00:02:22]: Yeah, just like add A100.Vibhu [00:02:23]: Was there any, like, early key problem that really sparked off why you built it?Akshat [00:02:28]: Yeah. Primarily it's just, none of the tooling that was out there was built for, one, a really great developer experience, and also there's a general trend of, a lot of the workloads that we were seeing were very. I wish there was a better word for it, but compute-heavy. Like, they need, one, like, need a lot more resources, so you need to burst up and down a lot, versus like Kubernetes designed for, like, slow scaling and, more for, like, web server use cases. And also there's just a lot more specialization in, like, what kinds of environments these workloads run in. Like, we had sometimes they need accelerators, sometimes they need different kinds of images, and this is just like a consistent thing that we saw across a lot of companies. That would be the next step.Software-Defined Infrastructure and Decorator-Based DXSwyx [00:03:13]: Yeah. Yeah. Be nice. I don't know how much this factored into the early story, but I wrote a post when I was at Temporal about infrastructure, software-defined infrastructure or something like that.Akshat [00:03:22]: Yeah, the self-provisioningSwyx [00:03:23]: Self-provisioning.Akshat [00:03:24]: Yeah.Swyx [00:03:24]: Yeah. I can't even remember my own post.Swyx [00:03:26]: And then you put me on the landing page.Akshat [00:03:28]: Yeah. We really like, the term and so we stole it.Swyx [00:03:32]: Because you had the insight that everything can just be in decorators co-located with the code, right?Akshat [00:03:37]: Yeah.Swyx [00:03:37]: Was that a big part of the originalAkshat [00:03:39]: YesSwyx [00:03:39]: Story or it was just like a DX layer?Akshat [00:03:41]: That was, really important because we really didn't want people to spend, so much time, writing YAML, and it seemed like you could really condense the surface area of what you're doing, put it in code so you can operate on it just like you operate on other code, and like build stuff that's more expressive and dynamic. and so yeah, that was always a very important part.Swyx [00:04:04]: Then the pushback is this is a DSL.Akshat [00:04:07]: Yeah.Swyx [00:04:07]: It's you're closed source. I am locked into Modal.Akshat [00:04:11]: Yeah. We never really got pushback for that because the nice thing about Modal is you can bring whatever code you have, and sure, the DSL is at the configuration layer for, what hardware you're using, how you're scaling things up, but you still own the code.Akshat [00:04:27]: And that's, that's been an important, part of our story, even as we do inference now.Swyx [00:04:32]: Yeah.Vibhu [00:04:32]: How much of do you think still stays the same today? Like if you were to build something today, DevX very important, but I feel like, a lot of this has been changed with just hook it up to an agent, have Claude Code, have Codex implement a tool. there's very agent native primitives that are different than if I'm doing this myself, right?Developer Experience → Agent ExperienceAkshat [00:04:54]: We've changed our SDK team to think about agent experience instead of, developer experience and we think that the same benefits that apply for DX also apply for AX, which is why would you have an agent read through hundreds of Kubernetes files and like write YAML that's not even typed when it can make a couple of changes in a decorator and it gets this self-provisioning runtime of, being able to see its changes live in action? yeah, it just seems from the customers we talk to, they find Modal is much faster for agents to use versus operating on a different substrate.Swyx [00:05:34]: Yeah, because like you, again, you co-locate the infrastructure requirements to the code that runs it.Akshat [00:05:38]: Yeah.Swyx [00:05:38]: Well, the negative thesis now is that nobody's looking at their code anymore, so there's no point.Akshat [00:05:44]: Yeah, people aren't looking at code. one thing we still see is really important is observability.Swyx [00:05:51]: Yeah.Akshat [00:05:51]: Like how good is your dashboard? And of course, like we have, we push a lot of it to the CLI so the agents can do their own investigation, but you still need humans to go interpret what's going on and, make judgment calls and whatnot. and that's I feel like, Maybe more important now than looking at the code itself.Swyx [00:06:11]: Yes, because like, you can try to treat the code as a black box and then use, see the observable action that comes out of it, and then just prompt a change.What Modal Is For: AI Cloud PrimitivesAkshat [00:06:21]: Yeah.Swyx [00:06:22]: So I think it takes a bit of restraint to not specialize, to say, “I want to ship a new primitive,” and then just be general purpose.Swyx [00:06:31]: People ask you, “What are you for?” You're like, “ I don't know. We can do this, we can do that.”Vibhu [00:06:36]: Well, I'd be curious to see, like, okay, if we were to ask you, like, what is Modal for even at a high level? There's a lot you guys do, sandboxes, GPUs, everything. How do you answer?Akshat [00:06:46]: Modal is a cloud platform that's built for, where we've built the primitives from scratch for AI applications. and right now it covers, inference, training, batch processing, and sandbox workloads.Akshat [00:07:00]: But we're building a lot moreSwyx [00:07:02]: I noticed you didn't say web server, so there is still a role for, like, the always-on large-scale Kubernetes type things.Akshat [00:07:09]: Yeah, absolutely. We're, we're not trying to compete with the renders of the world, because yeah, we think the differentiator for us is the, are the workloads that need specialized compute, need to scale up and down a lot. yeah, they're, they're, they're just shaped differently.Working Alongside Frontier StartupsVibhu [00:07:26]: I think you're building a lot of it alongside the startups, right? They're innovating quite a bit, even in your, like, latest blog post. Like, even in the series C, the customers that you mention here, the cognitions, technical ones, ramps and whatnot, they're, they're innovating with you, right? And that's not something AWS is doing directly with.Akshat [00:07:45]: Yeah, absolutely. I think, this is again classic. We're a small team. We can move really fast. our engineers are working with our customers and figuring it out. Yeah.Swyx [00:07:54]: So my first week at Cognition, I walked in, there was someone wearing a Modal shirt. I was like, “What are you doing here?” They're like, “Yeah, I just. I am embedded inside of Cog.”Akshat [00:08:05]: Yeah, I think that was Peyton. We sent him overSwyx [00:08:07]: Yeah.Akshat [00:08:07]: Because, the latency of communication was too high otherwise.Swyx [00:08:12]: Yeah, distributed node, you have to - you have to place one and collocate.Vibhu [00:08:16]: Yeah.Swyx [00:08:16]: So I had a, I had direct personal experience, right? So I worked on smol developer three years ago. it was inspired by Claude 1. I think you onboarded me at some point, like, just before, and I was like, “Oh, like, I need some bursty compute. Like, I was just gonna try using Modal.” And it was a, it was a pretty pleasant experience. apparently, I showed up in the board meeting, like the analytics.smol developer, Sandboxes, and Proto-CognitionAkshat [00:08:39]: Yeah, you blew up on Hacker News and,Swyx [00:08:41]: YeahAkshat [00:08:41]: We got a big traffic spike. I. I think the way you used smol developer was Modal functions for running stuff, which was. Like, the, that was a good use case. but then, yeah.Swyx [00:08:53]: Yeah. That - So to me, that was proto-cognition.Akshat [00:08:55]: Right.Swyx [00:08:56]: If only I had, like, stuck to it.Swyx [00:08:58]: Like, that was like, if - did you say draw the tech treeAkshat [00:09:00]: AbsolutelySwyx [00:09:00]: You're just like, “Yeah, like, probably this will happen.”Akshat [00:09:02]: Yeah. Like, he was so close. You were just rebuilding upon usSwyx [00:09:04]: I just didn't realize.Akshat [00:09:05]: But the funny story there is at the same time, we were talking to a bunch of customers who needed something like sandboxing.Swyx [00:09:14]: Yeah.Akshat [00:09:14]: This is like twenty-three.Swyx [00:09:15]: Yeah.Akshat [00:09:16]: So we builtSwyx [00:09:17]: You introduced a new API right after that.Akshat [00:09:18]: Yeah.Swyx [00:09:19]: Yes.Akshat [00:09:19]: Like, we built sandboxes in May of twenty-three before anyone was even knew this was gonna be a thing. And the first example we published was, we took smol developerSwyx [00:09:28]: Smol developerAkshat [00:09:28]: And put it in a loop, so the agent can iterate on itself.Swyx [00:09:33]: Loops are hot these days.Vibhu [00:09:34]: It's the looper.Akshat [00:09:34]: Yeah.Vibhu [00:09:35]: Loops in. When was this, twenty-three?Akshat [00:09:38]: Yeah.Vibhu [00:09:39]: A small check.Akshat [00:09:39]: Yeah.Swyx [00:09:39]: It's like twenty-three. so the. the, those for listeners, like, the problem was the models are not built for any of this, right?Swyx [00:09:46]: Like, you're just trying to like. They're not post-training to understand, like, looping and, like, self-correction and tool calling was there, but, like, also not that great.Akshat [00:09:55]: Yeah.Akshat [00:09:55]: I don't remember if you used tool calling in this one, but yeah, the models would just diverge after like ten iterations and not produce anything meaningful.Swyx [00:10:03]: Yeah. But like, then. So okay, like now talking to myself three years ago, the answerVibhu [00:10:08]: Of course they will get betterSwyx [00:10:09]: Collect all the failures, build benchmark, and then collect all the, examples, build the RL environmentAkshat [00:10:15]: RightSwyx [00:10:15]: Sell it for like ten billion dollars to Meta.Swyx [00:10:17]: And then also train a model and then sell that for sixty billion dollars to Elon. And this isAkshat [00:10:23]: Yeah, of courseSwyx [00:10:23]: The funny machine. Like, it's like, it's about the hardware.Akshat [00:10:28]: It's hard to have that inherent conviction that the stuff will get that much better.Swyx [00:10:33]: In retrospect, it's so f*****g obvious.Akshat [00:10:36]: Fair enough.Swyx [00:10:37]: Like, what else were we doing back then? I don't know. anyway. Yeah. So this. That was the start of your sandboxing journey, right? I feel like it didn't blow up until, like, last year.Akshat [00:10:49]: Yeah.Swyx [00:10:50]: So there was like a couple years of quietness.Akshat [00:10:52]: Exactly, yeah. We wereVibhu [00:10:53]: I think very underrated product value. Like, my experience with Modal, Charles, before he had joined Modal, met this guy at a hackathon, and he really insisted we wanted to run some small model, not hosted anywhere, and he's like, “ there's this cool company, Modal. They'll like spin up a GPU sandbox, we can throw it on there. They'll take a Hugging Face link.” And like there's so much value just right there, right? Like instant hosting, spin it up, spin it down. It'll stay cold, but we run the demo a few days later, it'll come back up and like all this stuff in retrospect, like it's still what we needed like today.Akshat [00:11:27]: Yeah, it's still needed today. workload shapes have changed a lot as, we run stuff for people with really massive production scale and, there it's it's not about scaling from zero to one, but it's how do we scale really elastically, from like thousand to fifteen hundred GPUs very quickly in a given region. It's the same shape problem.Elastic Inference, GPU Autoscaling, and Custom ModelsVibhu [00:11:50]: Okay. So you look at, say, Cursor Composer, right?Akshat [00:11:53]: Yeah.Vibhu [00:11:53]: They had a. “We'll do RL on a model every couple hours.” you guys have a whole version of RL inference gym and whatnot.Vibhu [00:12:01]: When you look at workloads like that, you're doing train runs where you need to scale up, scale down every hour thousands of GPUs, right? That's the example for we do need it, right?Akshat [00:12:12]: Yeah. Well, so I'll, I'll take a step back and, maybe talk about like how people use Modal today. because our biggest use case is, elastic inference. And the thing we first found product market fit, with was inference for custom models. So we stayed away from the LLM space, and we were serving companies like Suno for audio, Runway for video, robotics, comp bio companies that train their own model elsewhere. But Modal is the best black box that for deployment, scaling to however many GPUs you need as your traffic pattern changes. And we saw all of them like have a very unpredict- predict- predictable, traffic pattern. it's like diurnal. It's Some days, like the company will do a launch and, they'll need like, way more. And it's not just one model that they deploy. They-- all these companies deploy, lots of different models in different regions, and so the autoscaling problem becomes even harder because then you have to scale within a certain region, and those cycles are offset. So different times you scale up in different regions.Akshat [00:13:20]: So that's like our sortVibhu [00:13:22]: And thatAkshat [00:13:22]: YeahVibhu [00:13:22]: That in and of itself is a huge category. There's a bunch of inference providers which, provide this fireworks, does this as a service together, whatnot, Base10. that's carved into its own niche for language models, at least right now.Akshat [00:13:36]: Yeah. the thing that we have specialized in is the autoscaling aspect.Vibhu [00:13:41]: Yeah.Akshat [00:13:41]: Because we found that it's not universally true that everyone else can autoscale, and we've gone deeper into it on the tech side by, we've incorporated GPU snapshotting into the product so we can take the GPU state, like your torch.compile model, snapshot it, and the next cold start is way faster. And so going back to your question, it's That's why you need a lot of burstiness for inference. But then people also do a lot of demand training, like for RL stuff, your rollouts are bursty, as you said. People also do a lot of batch jobs. So we'll see, a lot of companies, before they have a training run, they'll need thousands of GPUs to run encoding or something like that. And I think those things are much more bursty than. I agree that agents are not that bursty. sandboxes are, except when you're doing RL. RL is justRL, Batch Jobs, and 100,000 SandboxesVibhu [00:14:28]: Or commerceAkshat [00:14:28]: Insanely bursty.Vibhu [00:14:29]: Yeah.Akshat [00:14:30]: Yeah. Like when you're doing, rollouts, you sometimes need a hundred thousand sandboxes in your sandboxes.Vibhu [00:14:37]: Yeah. I'm curious if you've seen early sparks of continual learning. There are some people, like our friends, ngram, recently announced thisAkshat [00:14:45]: YeahVibhu [00:14:45]: They're, they're trying to do training. That also seems like a different workload, right? If you're doing training twenty-four/seven per se, there's a very weird dynamic of how you're using GPUs between people and whatnot, but seems like something you guys would work for.Akshat [00:15:00]: As you said, we're, we're fortunate to work with a number of, customers at the frontier and grab some of our customers. and they are taking the primitives we have, and trying to use them in very interesting ways, like continual learning. It's possible as the stuff gets better, some of that will be part of, our offering as well if, more people need it. but we're, we're just waiting to seeVibhu [00:15:23]: YeahAkshat [00:15:23]: How it shakes out.Vibhu [00:15:24]: Is there a primitive that you added after sandboxing that was the next step in the story?LLM Inference, DeFlash, and Speculative DecodingAkshat [00:15:32]: I guess we've been going much deeper into LLM inferenceVibhu [00:15:35]: YeahAkshat [00:15:35]: Because we realized that some of the advantages we have with like autoscaling, again, especially in different regions and whatnot, are, not present elsewhere. and the place where we had a gap was we weren't, working on the model layer itself. Like we were a black box. And, we realized that, we can get to frontier-level model performance, with, by having great people who work on this. And, we've been open sourcing a lot of our work, in terms of, Recently, we, shared our work on DeFlash, which is a block-based, speculator, and we've open sourced, all of it. So, you can - By using open source DeFlash, you can get the same performance as you would with one of the proprietary providers. And the next thing we're thinking about hereVibhu [00:16:23]: I thought this wasAkshat [00:16:24]: YeahVibhu [00:16:24]: An interesting blog post as well, right? Like, I think in here you make a claim that. Not a claim, just that how effective speculative deco-decoding really just get to.Akshat [00:16:33]: Yeah.Vibhu [00:16:33]: Anything you wanna point out from this around, what people should know?Akshat [00:16:39]: Yeah, absolutely. the high-level summary is, it would help to describe what speculative decoding is.Vibhu [00:16:44]: Yes.Akshat [00:16:44]: I will, yes.Vibhu [00:16:45]: I think, likeAkshat [00:16:46]: YeahVibhu [00:16:46]: So we've covered like Eagle and all thisAkshat [00:16:47]: YeahVibhu [00:16:47]: Like Hydra and all those things, but it was like two years ago.Akshat [00:16:51]: Yeah.Vibhu [00:16:51]: I think it doesn't hurt, right?Akshat [00:16:52]: Yeah. Speculative decoding is you have a smaller model, called a draft model, predict tokens ahead of the bigger model, and then you have the bigger model, verify all of this, all the tokens are predicted. And the reason it's faster is if you're predicting, one token at once, you're bound by memory bandwidth. But if you can batch the verification of, the draft model, then you're much more efficient using compute, and it's faster, and as long as your draft model is producing a lot of tokens that can get accepted, which is called the accept length, you can get a speed up that's, multiple times of, the original model speed. and well, that's what we highlight here. It's Like people talk a lot about we made these kernels faster and whatnot, but improving kernel will only give you like few percentage points of improvement, and, increasing accept length, literally is a multiplicative decreaseVibhu [00:17:47]: Like two to four X.Akshat [00:17:48]: Yeah, exactly.Vibhu [00:17:48]: Without much head-on performance.Akshat [00:17:50]: Yeah. I think it may - you are running a second model, right? So it may be something more expensive in the compute,Vibhu [00:17:57]: I meant quality performanceAkshat [00:17:58]: Probably not by muchVibhu [00:17:58]: But yeah. I thinkAkshat [00:17:59]: So there's no drop in quality performanceVibhu [00:18:01]: YeahAkshat [00:18:01]: Because you're always. You're never accepting a token that the big modelVibhu [00:18:04]: It's strictly betterAkshat [00:18:05]: YeahVibhu [00:18:05]: Or it's same.Akshat [00:18:06]: Exactly.Vibhu [00:18:07]: Right. Yeah.Akshat [00:18:08]: And so we've been working a bunch on DeFlash, which is a block-based speculator. so it's instead of predicting, one token at a time, it's predicting a block. And we've been open sourcing our work with it. The next thing for us here is for helping people train speculators and custom models. it's it's something that traditionally is very forward-deployed engineering driven, support deployed, engineer driven, like you work with customers and help them do that. And our vision for. This is why we launched Auto Endpoints, is we want to make frontier-level performance available to everyone. And so, we mentioned this in the announcement, we teased it. The next thing we're, we're launching is, as you run an auto endpoint, we shadow trafficAuto Endpoints and Frontier-Level PerformanceVibhu [00:18:54]: Do you want to explain what auto endpoints are?Akshat [00:18:57]: Yeah.Vibhu [00:18:57]: I lovely, yeah.Akshat [00:18:58]: Yeah. So, this is, I guess, going back to your Modal is you touch the code, but, sometimes people don't wanna touch the code, and they wanna get started with an endpoint that works and has all the great performance and, scalability that Modal has. So we've made that easier with, a way to create an endpoint from our UI, from the CLI, that has all of our optimizations that we talked about, like the DeFlash stuff already baked in, and there's full transparency. So we give you the code, you can go run it yourself, and if you want, you can eject out into the full Modal experience, which we see as people get sophisticated, they do wanna tweak the models, they wanna, fine-tune stuff. You can still do all of that. It's it's not a black box. And yeah, the next thing, as we teased later in the post, is how do we give you value even beyond this in terms of having your draft models evolve as your data distribution evolves, again, without having to talk to a person and, yeah.Vibhu [00:19:59]: I guess just to understand it directly, you have the GPUs, you have an endpoint that's compatible, you serve open model. If someone was to do this themselves, what's the delta that you guys provide? So you do a lot of open source great work on effective inference. how does it compare to, say, I take the same model, 5.2 FP8, take shelf inference engine, vLLM, SGLang, get compute of similar capacity, similar cost. What's the delta that plugging into something this, like this offers outside of the benefit of, scaling?Production Inference Beyond Raw GPUsAkshat [00:20:34]: It's interesting because we've taken the approach of open sourcing our contributions and upstreaming them. we work closely with the SGLang team. We want the improvements that our team, comes up with to be, there in open source for others to use, even outside of Modal. The benefit to us is we have a team that has significant expertise in terms of if you do have something that is not there, our team can help you get that performance, first. the other thing is with these endpoints, we are way more elastic, as you said, than, anyone else, and you have true scaling to zero. you have true, burstiness, and in practice, that matters a lot more to people than just finding, the GPU and, running Modal code on something.Vibhu [00:21:20]: Yeah. And I will say it's not that straightforward to just. like what I said is easier said than done, right?Akshat [00:21:26]: Yeah.Vibhu [00:21:27]: It's I think still for the average person, still hard to just gut check using different. There's, there's quite a bit of combinations you can make there. the trade-offs aren't really known at face value.Akshat [00:21:40]: Yeah. it's it's not just that. I think it's it's that running production-grade inference is a hard infer problem.Vibhu [00:21:49]: YeahAkshat [00:21:49]: Even if you subtract out the autoscalingVibhu [00:21:50]: YeahAkshat [00:21:51]: Is controlling things like tail latency and, making sure every, request is delivered at least once and whatnot.The Model and Agent LifecycleVibhu [00:22:00]: There's a lot of innovation that you can do here. I think, it's very interesting that you're starting to encroach on, like as you become a full cloud, you're starting to encroach on other people's turf.Vibhu [00:22:09]: What will you not do?Akshat [00:22:13]: Well, we wanna follow our users and, make sure they get like a platform that has everything that works well together. so right now we're focused on the model lifecycle and the agent, lifecycle. so both like going from data prep to training to inference, and then also if I want to deploy a background agent, let's say, sandbox, do persistent storage, a whole bunch of other stuff.Vibhu [00:22:38]: We talked to Cole, who did, OpenInspect. Yeah.Akshat [00:22:42]: Yeah.Vibhu [00:22:42]: And RealInspect also is on Modal.Akshat [00:22:44]: Yeah. So Ramp Inspect was a great example of a background agent that was really successful because they, were able to use some of the primitives like snapshotting and fast scaling to just have something that feels really reactive and works well.Ramp Inspect and Background AgentsVibhu [00:23:02]: Yeah. That's the new CTO of, Ramp right there.Akshat [00:23:05]: Yeah, Rahul.Vibhu [00:23:08]: It was really fun. yeah, okay, I think, all very bullish. Like, one of my reflections was also I did not originally. So when I met you guysThe Inference Inflection: CPU, GPU, and Co-LocationVibhu [00:23:19]: You weren't that much in the GPU game, and now you're all about, inference. And one of the points that I hinged on for Jensen's keynote at GTC this year was, what we're calling like the inference inflection, right? That let's say in AI workloads or machine learning workloads, it used to be like, let's call it eight to one GPU to CPU, and now it's more like one to one, which is like a interesting. Like, - because of how much agents are blocked or call out to this, to CPU heavy stuff the actual, like, limiting factor, like, swings back and forth from GPU to CPU a lot more than it used to be all GPU and then occasional CPU.Akshat [00:24:01]: Yeah.Vibhu [00:24:02]: GPU, CPU. And now it's like just constantly, and you just have to locate everything.Seventeen Clouds and the Supercloud StrategyAkshat [00:24:08]: Yeah. And that's one of the things that, again, we see as, something appealing about Modal, which is we've built this capacity pool that spans, 17 cloud providers, so we're, we're very good at Running on various kinds of cloud capacity across the worldSwyx [00:24:24]: You don't have your own data centers?Akshat [00:24:25]: We don't have our own data centers. We just run across a lot of neo cloudsSwyx [00:24:29]: Yeah. AreAkshat [00:24:30]: Metal providers.Swyx [00:24:30]: Yeah. Question mark.Swyx [00:24:31]: Yeah. You're, you're running the math, and you're like, “What's the cutover point where you're like.”Akshat [00:24:36]: Yeah, it's a good question. part of it is we see our differentiator in the software layer, and, being capital light and focusing on the software helps us move really fast. so far it's worked out well because there are so many other people building data centers that we're able to work effectively with them, and again, focus on what makes us, special.Swyx [00:24:55]: Yeah.Swyx [00:24:56]: 17 gets you into, like, the local providers sometimes. LikeAkshat [00:25:00]: The,Swyx [00:25:01]: Which was the most interesting one?Akshat [00:25:02]: There are a lot more neo clouds than you expect, and they all have various degrees of, various levels of reliability. And, that's why it's something we've invested a lot of time in, is building our own reliability layer on top. so if the GPU falls off the bus or something happens, we user workloads are not affected, and that lets us use a lot more capacity than,Swyx [00:25:30]: YeahAkshat [00:25:30]: You as a user would be able to.Swyx [00:25:32]: It's a useful thing to have because like now everyone knows, like, what layer you are and, like, you optimize for being the super cloud of all clouds.Akshat [00:25:41]: Yeah. That's, that's, that's the idea. and so I guess when you mentioned colocation, that's, that's another interesting thing where, one thing we've seen is people come to us when they want, very specifically located, CPUs or GPUs, like they wantSwyx [00:25:57]: Oh, they pin it in likeAkshat [00:25:58]: YeahSwyx [00:25:58]: EU?Akshat [00:25:59]: Exactly. Or EU, US.Swyx [00:26:01]: Right. Data resiliencyAkshat [00:26:02]: AustraliaSwyx [00:26:02]: Locality thing or performance or what?Akshat [00:26:04]: It's either data locality or latency, yeah.Swyx [00:26:07]: Yeah.Akshat [00:26:07]: Like, you want your. They're running sandboxes and model. They want them to be right next to aSwyx [00:26:10]: Yeah, it's easy thenAkshat [00:26:11]: YeahSwyx [00:26:12]: To. That is important in all those things. and so, like, you've accidentally, I don't know if it's accident, but, like, you've built the perfect primitive for agents to express themselves. And then, like, it's almost very funny how every extra development just involves more file system, just involves more CPU.Akshat [00:26:30]: Yeah.Swyx [00:26:31]: Just like the things that you already have. I don't know much about, if there's any, like, networking usages that are interesting, but you've also done some good work on networking.Networking, Sidecars, Private IPv6, and SandboxesAkshat [00:26:40]: Yeah, that's exactly right. Like, we're just taking compute storage and networking and building stuff on that layer, for, again, the stuff people need.Swyx [00:26:49]: YeahAkshat [00:26:50]: We see a few interesting networking things coming up. one is people want networked sandboxes. so we haveSwyx [00:26:57]: For like a Docker cluster type thing.Akshat [00:26:59]: Yeah.Swyx [00:26:59]: Sorry, Docker Swarm. Oh, f**k. What is it called?Akshat [00:27:02]: Compose.Swyx [00:27:03]: Compose type thing.Akshat [00:27:04]: Yeah. So if you want Docker Compose, our sandboxes now support, this thing called sidecars. So you can. A sandbox is a pod of containers, and you can run multiple containers in, a sandbox. also useful because, going back to networking, people want a lot of control over, outbound networking from a sandbox.Swyx [00:27:23]: Yeah.Akshat [00:27:23]: Like, they might wanna run a middle proxy for, like, maybe logging stuff for RL or, controlling how egress can happen to a domain, injecting credentials. and yeah. So we've, we've had to build a lot of that stuff ourselves.Swyx [00:27:38]: Yeah.Akshat [00:27:39]: But then also sometimes people want, sandboxes spanning multiple nodes to talk to each other, which is an emerging thing we're seeing. We have support for that for a different reason, and yeah, we'll see if that becomes stable.Swyx [00:27:52]: Like, just an open socket. It's a. This is directly like mTLS.Akshat [00:27:56]: We do support that, which is you can, expose a tunnel inside a sandbox.Swyx [00:28:01]: Yeah.Akshat [00:28:01]: And then you can either expose it to public internet or it can be, you can add like a HTTP, auth layer above it. But we have this thing called I6PN, which we haven't talked about, which is this, like, overlay network using IPv6 addresses. so if Modal containers, within the same workspace, when this is enabled, can address each other using this private IPv6 address, and no one else can.Akshat [00:28:28]: So it's like private networking, for containers. We built it because we needed it as a primitive for our distributed training product. so we have this other feature, which is you can add a decorator to a function, and you get a cluster of GPUs. and they have RDMA networking. so you can run a distributed training job, that's truly serverless. and we did the overlay network for that. But then we've seen that people are using it for other reasons, and, I'm intrigued to yeah, what would people do with it.Swyx [00:28:59]: Build primitives and let people figure it out, right?Akshat [00:29:01]: Yeah, exactly.Swyx [00:29:02]: You put out a pretty interestingAkshat [00:29:03]: They're like, they read the docs webpage. Let me use thatSwyx [00:29:06]: YeahAkshat [00:29:06]: Something they never intended to work. This is literally not even in our docs page. People somehow found it, and they're using it.RDMA, Memory Movement, and Distributed TrainingSwyx [00:29:12]: Huh.Swyx [00:29:14]: The way you portrayed it with, like, RDMA versus TCP, like, very well laid out, but just the transfer speed change at scale for RL, like yeah, you have it, you have it built in. I'm sure someone found it. It's found it to be a lot more efficient before you made a thing out of it, right?Akshat [00:29:32]: Yeah. And not to split hairs, I guess the overlay network is the TCP overlay network.Akshat [00:29:39]: The reason we have that is you need that to do the key exchange for RDMA before you set up the RDMA network on top of that. but then people found the TCP part.Swyx [00:29:48]: Can I tell you, this is like a big aha moment for me becauseAkshat [00:29:51]: YeahSwyx [00:29:51]: So I review 2,200 submissions for the World's Fair.Akshat [00:29:56]: Yeah.Swyx [00:29:57]: And then I got this from John OsterhoutAkshat [00:29:58]: HuhSwyx [00:29:59]: Who I don't know if. Do John Osterhout by name?Akshat [00:30:01]: The name sounds familiar.Swyx [00:30:02]: He published a. He's a well-known professor, published a lot of interesting software design books, and this is the talk he chose to submit, is on RDMA at Inference. And I'm like, you wouldn't think that this guy, who is like operating systems guy, would care about RDMA.Akshat [00:30:20]: I, it makes sense to me because I,Swyx [00:30:24]: This is the cloud, right? YeahAkshat [00:30:25]: Like, the way you move around your KV cache and how efficiently you can do it, how efficiently you move, your weights from your training GPUs to your inference GPUs in RL is there's a lot of degrees of freedom, and it is a systems problemSwyx [00:30:41]: YeahAkshat [00:30:41]: Moving memory aroundSwyx [00:30:42]: YeahAkshat [00:30:43]: Scheduling.Swyx [00:30:44]: This shows you how primitive my understanding of networking stuff is.Swyx [00:30:46]: Is this like the domain of WireGuard as well?Akshat [00:30:50]: Not quite.Swyx [00:30:51]: It's adjacent?Swyx [00:30:53]: Explain everything.Akshat [00:30:54]: Sure.Swyx [00:30:56]: How do we move memory around GPUs?Akshat [00:30:58]: Well, so sorry. Yeah, that is memory. Sorry, I was talking more, and maybe I was talking like five minutes back, about the private IPv6, addressing that you've set up.Swyx [00:31:09]: Yeah.Akshat [00:31:09]: Is it like it's a VPN?Swyx [00:31:10]: Yeah, it is like a VPN, and yeah, WireGuard is, yeah, you're right. It is,Akshat [00:31:16]: Right. Yeah, you already moved on to new topicsSwyx [00:31:17]: A similarAkshat [00:31:18]: OkaySwyx [00:31:19]: In the same space, WireGuard is, encrypted and this is,Akshat [00:31:23]: And you don't need encryption.Swyx [00:31:23]: Yeah.Akshat [00:31:24]: Yeah.Swyx [00:31:24]: This is not encrypted. that's the main difference. This is TCP and we have eBPF programs that will reject or allow the TCP connection based on whether you're allowed to do it.Akshat [00:31:35]: Used to involve a full sidecar, but now you have eBPF in the Linux kernel.Swyx [00:31:39]: Yeah.Akshat [00:31:40]: Yeah. I don't know if this is a natural follow-on to the topic of like my skepticism on distributed training is that while, like, people spend a lot of money on, like, cables to hook up GPUs, and even that is not, like, fast enough, and that's the bottleneck, is your networking fast enough?Swyx [00:31:59]: Yeah. So I guess you're talking about fully distributed training like, Dialog or something which is like cross data centerAkshat [00:32:06]: That would be, yes.Swyx [00:32:07]: That's the extreme.Akshat [00:32:08]: Yeah.Swyx [00:32:08]: You're in the middle, and then other people would have like the Mellanox cables up in, like, their actual data center.Akshat [00:32:14]: When you run multi-node training on Modal, RDMA, I think Mellanox, is, or InfiniBand is like a, is all seen as RDMA. but it's a way to bypass the TCP networking stack and, transfer, stuff much faster, between one node, to the other. And we have I think like 3 terabit per second, internal networkingSwyx [00:32:40]: OkayAkshat [00:32:40]: Which is the standard that's needed.Swyx [00:32:42]: Okay. So I misunderstood whatAkshat [00:32:43]: 50Swyx [00:32:43]: What part of the stack you wereAkshat [00:32:44]: 50 gigs overSwyx [00:32:45]: YeahAkshat [00:32:45]: If you wentSwyx [00:32:45]: YeahAkshat [00:32:46]: RDMA.Swyx [00:32:46]: Okay.Swyx [00:32:48]: Yeah. I, very impressive work.Multi-Node Training, Post-Training, and Auto ResearchSwyx [00:32:52]: So effectively you're extending like the model philosophy to the training cluster, like, yeah.Akshat [00:32:59]: Yeah. And we're, we're not going for like large scale training runs. the thing that we've built multi-node training for is, we see a lot of, smaller scale post-training. like, people are post-training like medium sized fund models, so they can, get higher quality on inference. this is a perfect fit, for something like that.Swyx [00:33:21]: Yeah. That is my impression of how a lot of these labs explore branches in post-training and then eventually merge whatever they find in.Akshat [00:33:31]: Yeah. The other use case we've seen for multi-node training is even if you have a big cluster, your researchers are still doing small runsSwyx [00:33:38]: YesAkshat [00:33:39]: Having elasticity thereSwyx [00:33:40]: Right, sureAkshat [00:33:40]: Matters a lot more.Swyx [00:33:41]: Yeah. the, like, this is like the current limiting factor for auto research, which is like you need to give your model some GPUs in order for it to completely run.Akshat [00:33:51]: We have a blog post on auto resource and model is,Swyx [00:33:55]: YeahAkshat [00:33:56]: Yeah, like, turns out to be pretty good substrate for that.Swyx [00:33:59]: So my impression is auto research means many things, likeAkshat [00:34:01]: YeahSwyx [00:34:01]: Anything that Andrej coins. Right now it's still science fair, right? Like not like, I don't know how many people are doing this.Akshat [00:34:08]: We're having a golf.Swyx [00:34:08]: Yeah.Akshat [00:34:09]: I thought the same thing.Swyx [00:34:11]: Yeah, you would know.Akshat [00:34:12]: We, like, our internal both training and inference teams use this the general shape of this quite a bit. like we have this one internal repo called auto inference, which essentially we've automated our own forward-deployed engineering efforts using, this harness, which is, the agent will just spin up a sweep of different things. It'll even run like, NVIDIA inside profiler and it'll like tweak configs and it'll arrive the right thing. it'll change your GPUs both from H200 to B200, and works really well.Swyx [00:34:47]: Nice.Akshat [00:34:47]: So yeah.Swyx [00:34:48]: By the way, I enjoy that your forward-deployed engineering is so technical that you have to do these things.Swyx [00:34:52]: It's very different from forward-deployed engineering from other people.Akshat [00:34:54]: Yeah. For our forward-deployed engineering team is, essentially they're like applied inference researchers or applied training researchers.Swyx [00:35:02]: Someone told me like they have to be able to build, but they also have to be able to sell. do they have to sell or are they like they're good, they're just like post-sale type of thing?Akshat [00:35:09]: It does, being able to talk to a customer and engage effectively with themSwyx [00:35:13]: YeahAkshat [00:35:13]: Matters a lot.Swyx [00:35:14]: They want the same thing.Akshat [00:35:15]: Yeah.Swyx [00:35:15]: ?Akshat [00:35:15]: But it's it's not really a sales, thing. We pair them with-- We have solution architects as well that are more on the sales side.Swyx [00:35:23]: Okay. Let's spend a bit more time on auto research. This is a big focus for for this year. Where does this go? like, have people explored enough? Like, there's all these beautiful charts of like improve and then level off a bit and then you find the next thing. Is this one abstraction up from normal training? Is that how we think about it, or do you think about it differently? Like model level training versus high, like driven hyperparameter search.Auto Inference and Modal BenchAkshat [00:35:51]: Yeah, like,Swyx [00:35:51]: Someone, some people call it like neural architecture search or whatever, right? Like.Akshat [00:35:54]: Yeah, - So the stuff I've seen people do with it is nowhere on the architecture level. It's pretty much tweaking parameters, but it's it's a hyperparameter sweep that's guided by some model intuition, so it's like much more efficient than, whatever other, sweep you would have.Swyx [00:36:12]: Yeah, it's just, it's just a question of where you want to spend your compute?Akshat [00:36:16]: Right.Swyx [00:36:16]: ‘Cause yeah, you can just throw infinite amounts of money on this and somehow you'll bang out Shakespeare?Akshat [00:36:22]: Yeah, infinite monkey.Swyx [00:36:24]: Yeah, so like the very good for model. and I think it's also very important that agents can spin up other agents, can spin up their infrastructure. Like very good for you. how good is our LLMs at generating model code? Like the benefit of existing LLMs is that you are in the data.Akshat [00:36:42]: Yeah. They're, they're surprisingly good. I think like pre Cloud 4 they were not, and then now they're able to shot, stuff out of the box. But we're playing around with releasing like a Modal Bench for like the harderSwyx [00:36:55]: YeahAkshat [00:36:55]: Things, that the LLMs cannot do yet and maybeSwyx [00:36:59]: What's an example of that?Akshat [00:37:01]: I think the things that- Sometimes agents struggle with, without right guidance and a skill is, how to, use the rest of our observability. Like how to. Something is failing, like how do you look at the logs and then update the right thing? It's reasoning about that. But they're able to shot, likeSwyx [00:37:23]: Yeah. You can just add a skill to it?Compute Strategy and Capacity PlanningAkshat [00:37:26]: Yeah. So we have a Modal skill now that. Which is why we built this Modal Bench. It's to find things like that, so we can address them in our tool.Swyx [00:37:35]: Tune a skill. Yeah.Akshat [00:37:36]: Yeah.Swyx [00:37:36]: No. it's it's good. are you facing any shortages? like we talk a lot about GPU shortages, but also CPU, also memory.Swyx [00:37:44]: Yeah.Akshat [00:37:45]: We have had a lot of growth, which means that, there's - we've had to be much better aboutSwyx [00:37:53]: PlanningAkshat [00:37:54]: Proactive capacity planning.Swyx [00:37:55]: Yeah.Akshat [00:37:55]: So we have,Swyx [00:37:57]: Which by the way, like it's like a MBA's like dreamAkshat [00:38:00]: YesSwyx [00:38:00]: Is like just planning this stuff. I think last time you and I talked about something maybe about this.Akshat [00:38:03]: Yeah. we have a really competent team of people that we call, The role is called compute strategy. so yeah, if anyone listening here or wants to work on thatSwyx [00:38:13]: Compute strategy?Akshat [00:38:13]: Yeah.Swyx [00:38:14]: I think,Akshat [00:38:14]: I feel like,Swyx [00:38:15]: I think the normies call it FP&A or something.Akshat [00:38:18]: Well, it's more It's it's not FP&A. It's it's There's a lot of interesting financial questions of like what is the blend between one year and three-year reservations? how do we forecast our own capacity? how do we. especially since our capacity is very fungible across different GPU types and different regions, like you have to model a lot of it. and you also have to have an opinion on how the supply chain is gonna evolve, and then you have to like, take bets,Swyx [00:38:49]: YeahAkshat [00:38:49]: Based on that.Swyx [00:38:50]: Tokenomics.Akshat [00:38:50]: Yeah.Swyx [00:38:51]: This is like probably a not a real point, but, I was trying to think about like what other industries. I was trying to think about like, we cannot be first to like these kinds of problems.Akshat [00:38:59]: Yeah.Swyx [00:39:00]: And what other industries have had this? And I was like, airlines with fuel and like they have to hedge their fuel and like, I think for a long time Southwest because they made like a hero fuel bet, they like were like super low cost becauseAkshat [00:39:12]: OhSwyx [00:39:12]: Compared to everyone else.Akshat [00:39:14]: Yeah. I hadn't thought about that.Vibhu [00:39:16]: We're at a fun time too?Akshat [00:39:18]: Yeah. It's. A lot of the compute business in general, for us is also about being very good about capacity management. That is how you have great unit, economics. but also over time it's how you can unlock more value for customers. Like, one of the things we're building now is like a way for customers to get, If they don't care about latency, like get much cheaper pricing and they'll get results back in like next 24 hours or something, like a batch tier essentially.Batch Tiers and Latency-Insensitive WorkloadsSwyx [00:39:47]: Yeah.Akshat [00:39:47]: And those are levers we have because we control the whole stack and scheduling and whatnot to give people a sufficientSwyx [00:39:53]: Yeah. I feel like they're not as popular. Like those, like the Frontier Labs have all those APIs. They're not as popular as they should be.Akshat [00:40:00]: The demand that we see for something like that is not for LLMs. although sometimes people wanna run evals andSwyx [00:40:08]: OkayAkshat [00:40:08]: Synthetic data prep and there it makes sense.Swyx [00:40:10]: Okay.Akshat [00:40:11]: But it's from a lot of LLM companies, like people who are doing computational bio, like they have to run really big batch jobs and they don't care about when they get it back.Swyx [00:40:22]: Yeah. And like they have a reasonable. It's it's also like a cousin to the stopping problem of like, will this finish in time?Akshat [00:40:30]: Yeah. You can bound it.Swyx [00:40:33]: Yeah.Akshat [00:40:33]: Like you can give peopleSwyx [00:40:34]: YeahAkshat [00:40:34]: SLAs on it.Swyx [00:40:35]: Yeah. I think what's, what's interesting is like the next phase of model.Swyx [00:40:38]: Like what, do people expect from you, now that you're established and you're like well-known compute player among all these leading companies. You had an inference launch week, and we talked a little bit about the launches. like what else? Like what else should people know?What Modal Builds NextAkshat [00:40:55]: We are building primitives that make our users' lives much easier. So, I think for example, with LLM inference, thousands more companies are gonna post-train their own models and, deploy open source models for inference. so we're thinking a lot about what is the best product shape for that. And, that involves everything from our training gym to, then, endpoints that get frontier-level performance. again, but I haven't talked to anyone. It looks somewhat different on other verticals. Like, we're also seeing a lot of real-time, audio-video stuff in there, which is why like, we're working on things like regional routing, with fallbacks. So you can get GPUs that are as close to users as possible. so you get like low latency for video streaming and whatnot. And then on the agent side, it's,Akshat [00:41:52]: We're still working very closely with our customers because stuff is changing so fast in terms of what they need. And, I think beyond sandboxes and persistent file systems, there's a lot of other things people will need from this agent stack as they build production agents. So yeah, we're thinking about those other things that fit in there.Swyx [00:42:13]: I want to ask what the other things are.Akshat [00:42:15]: Yeah. I probably should share right now.Swyx [00:42:17]: I think-- I think, okay, so, I do think a lot about the principal components of cloud, and you do talk about compute storage networking.Akshat [00:42:25]: Yeah.Swyx [00:42:25]: Because so far for me, it's fine. so far for the. the first couple generations of cloud, it's fine. What's different, qualitatively different about agents that you need some new permission level? Like a lot of people, okay, and I'll just kinda spew tokens at you until it like hopefully sparks something.Akshat [00:42:43]: Yeah.Swyx [00:42:44]: Like the new level now is whatever Claude Code does, which is dangerously scope permissions or like allow list by command or like whatever, right? And sometimes they're like, “Well, okay, we have like this adaptive thinking mode where like, just trust me, bro. I will make the calls for you.” Is that it? like mediated permissions.Hard Guardrails vs. LLM-Mediated PermissionsVibhu [00:43:03]: Now you're looping it with a goal and letting it roll.Akshat [00:43:06]: Yeah, I'm, I'm skeptical of LLM media permission for stuff that is at the sandbox level because you do want hard boundaries.Swyx [00:43:16]: Yeah.Akshat [00:43:16]: Otherwise, someone can exfiltrate stuff.Swyx [00:43:20]: But likeAkshat [00:43:20]: YeahSwyx [00:43:20]: Maybe that's old school thinking. Maybe we're the dinosaurs.Swyx [00:43:23]: Maybe the AI OS or the LLM OS is really the kernel is a goddamn LLM.Swyx [00:43:30]: Like it makes you feel uncomfortable.Akshat [00:43:31]: Yeah, I'm, I'm toldSwyx [00:43:32]: But that's what trusting the LLM is. Like imagine a spherical cow perfect LLM.Akshat [00:43:36]: Right.Swyx [00:43:37]: That it.Akshat [00:43:39]: Maybe.Swyx [00:43:41]: I wanna test the boundaries, right?Akshat [00:43:42]: Yeah.Swyx [00:43:42]: Like, and I don't believe that, but I wanna see where I'm wrong ‘cause that's, that's the consensus.Akshat [00:43:49]: Yeah. I think you always need hard guardrails when you want, And you can pair those with softer guardrails, right? And that's gonna be a lot of mediated.Managed Agents and Specialized SandboxesSwyx [00:44:00]: There. I'll also get you a end with a couple of your commentary on like the ecosystem outside of Modal. Manage agents. Everyone has one. Gemini, OpenAI, Claude, very useful for you, but also like it is their way of starting to edge into your space.Akshat [00:44:17]: Yeah.Swyx [00:44:17]: What's going on?Akshat [00:44:19]: Yeah, we're, very excited to partner with Anthropic and some of the other foundation labs, will not name who we're also working with. the way we see it is the manage agent thing is a great place to start if you're starting out building an agent and, But then when you get to, building something more production grade, like you're a company that's like Ramp that's building their own, Ramp also runs their accounting agent on us, so their external-facing agent. You need a lot more control over, your compute primitive on things like, what sort - how do you persist different files that the agent has access to, and how do you snapshot and restore? How do you control the networking? maybe you want GPUs. When you get to that point, you kinda want, a specialized sandbox provider, that gives you those things, and that's the role that we are trying to play.Swyx [00:45:15]: YeahAkshat [00:45:16]: We don't really have an opinion on the harness, whether it runs - it's a cloud-managed agent, and you hook it up to Model Sandbox, or you run the harness in Model Sandbox. We'll see where people converge with that.Swyx [00:45:26]: Yeah. Do you any opinions on like the meta harnesses, or just another layer on top of these things?Akshat [00:45:31]: You mean like the OpenPipeSwyx [00:45:33]: OpenPipe is one. I think Vercel had one, which I can't remember the name of right now. Fredshot had one. and then, to me, most recently was Data Databricks that had Omnigen. All these are meta harness. Like it's kinda pseudo agent cloud type things.Akshat [00:45:50]: I personally have not played around with them.Swyx [00:45:53]: Yeah.Akshat [00:45:53]: Build agents with them.Swyx [00:45:54]: Everything's bullish Modal, as long as it consumes more infra.Akshat [00:45:57]: That's why we're focusing on the infra layer. It's somewhere where our, relative competence is and, also it's a hard problem to solve.Swyx [00:46:06]: Yeah. I will say like just generally reflecting on that, I don't know if - if there's other topics on Modal, but like just generally reflecting as an infra person, not as intense as you, but in that field, this has like been the most exciting time in infra. Like it was boring for a while, and you couldn't really get people excited about data infrastructure. Like Eric would get on Data Console, everyone just watched the video and like say, “Look at how many sandboxes I can spin up,” and no one gave a crap.Why Infrastructure Became Exciting AgainAkshat [00:46:39]: Yeah.Swyx [00:46:40]: And like now everyone gives a crap.Akshat [00:46:42]: That's true. It is a very exciting time, and I think a lot of that's driven by just the amount of scale all of this stuff needs.Swyx [00:46:50]: I think the, like a lot of your initiatives or a lot of your like product directions make sense in retrospect, which is like the best kind, but I wouldn't necessarily have thought about it myself, which.Akshat [00:47:00]: We need the predictions.Swyx [00:47:02]: I think there's a lot that you just don't even see, right? Like you have the batch, you have the voice, you have the multimodal, but what else?Akshat [00:47:10]: What else is coming up for usSwyx [00:47:11]: Yeah. Where do you see things going?Akshat [00:47:13]: Yeah. I, in generalBiotech, Robotics, and Non-LLM AI WorkloadsAkshat [00:47:15]: It's it's clear that there's there's a huge shift happening. I think one thing that's not as obvious to people because LLM inference gets talked about so much and is also we work a lot of companies that are, doing things like drug discovery and computational bio, like the Chai Discoveries of the world. Big things are probably gonna happen there. we work a lot of robotics companies that are putting robots in like active deployments and getting good results out of them.Swyx [00:47:45]: Is there Air Gap Modal? Is there a version that is like prem air gapped whatever?Akshat [00:47:50]: No. We,Swyx [00:47:51]: You should cloud only.Akshat [00:47:51]: Yeah.Swyx [00:47:52]: Yeah. Okay. But yeah, so what you're saying is like because you're focused on primitives and they're good primitives, you find use cases in all these kinds of things.Akshat [00:48:01]: Yeah.Swyx [00:48:01]: Probably diversifies you a little bit away from LMS all the time.Akshat [00:48:05]: Yeah, absolutely. We're, we'- our goal isn't to only serve the LLM inference market.Swyx [00:48:10]: There are a lot just on the website, the audio,Akshat [00:48:12]: Yeah. We said both onSwyx [00:48:14]: Computational bio images. Yeah, there's a lot here. There's QTA TTS, customizing. Oh, Chatterbox. there was customizing Whisper.Akshat [00:48:24]: Okay. Yeah.Swyx [00:48:25]: This screen reminds me of a fallen competitor, which Replicate.Model APIs vs. Differentiated AI ProductsSwyx [00:48:31]: What's your postmortem on what happened?Akshat [00:48:34]: This is one thing we've stayed away from is providing an API for models because I think providing model APIs is some of it ends up serving like a really hobbyist market, which is much less sticky.Swyx [00:48:50]: Yeah.Akshat [00:48:50]: And we've always wanted to build for companies that are building products and need more flexibility that's not just an API.Swyx [00:48:57]: Which you can build an API for a model and this is clearly what it is. But you - but what you're saying, you can wrap it into a more fully functioning back end that you run.Akshat [00:49:06]: Yeah. So all of our examples, it's not that spin up this model, here's an API token, use it. They're all code.Swyx [00:49:13]: Okay.Akshat [00:49:13]: And so the point is that this is just an example.Swyx [00:49:16]: Starter code.Akshat [00:49:17]: Yeah. But you can tweak it however you want.Swyx [00:49:20]: Yeah.Akshat [00:49:21]: And if you're like a company building a product, like, computational bio whatnot, yeah.Swyx [00:49:26]: I guess I'm trying to tease out for listenersAkshat [00:49:28]: YeahSwyx [00:49:28]: When does it stop becoming, oh, you're just an API call and you're just a wrapper on API to becoming what you call a product, right?Swyx [00:49:36]: Like, what is that layer? Like what-- Like, more lines of code, but like beyond that, what is the substance that people add that qualifies it to be something more?Akshat [00:49:46]: I think there's a little bit of like a selection effect of like a lot of the companies who do wanna get deeper into that level are probably building something that's more differentiated. And, I think, an example is like - with LLM inference, originally we, worked with companies that were building their own post-training frameworks or they were, - Ramp early in the day was training their own tokenizer and like swapping out the tokenizer in Llama and whatnot. I'm not saying that's, that successful, in that case. But a better example is like, let's say Suno. because Suno, does not use Modal for training.Swyx [00:50:26]: Mikey on the pod. Yeah.Akshat [00:50:27]: But they use Modal for all their inference and that's because they have like a custom-- They have completely custom model architecture and that means that they have to be at the code level and tweak things that are not, just an API.Swyx [00:50:41]: It's interesting as well, like we had, Ethan, most recently on the xAI Groq team make a prediction that like the next tier in video gen is not a better video model, it's a better model or agent that orchestrates video models.Video Agents and Production WorkflowsAkshat [00:50:56]: Oh, interesting.Vibhu [00:50:56]: Language model backbone that can use toolsAkshat [00:50:58]: RightVibhu [00:50:59]: And write code.Akshat [00:51:00]: Like, yes, I can make my second video or my second video from Groq, but I want my minute video.Akshat [00:51:06]: And I'm not going there through normal video gen.Swyx [00:51:10]: Yeah, that's interesting. I - So we have GPU sandboxes and recently have seen a few companies doing agents that do video manipulation or,Akshat [00:51:22]: Yeah. Give it FFmpeg and just do it.Swyx [00:51:23]: Run FFmpeg. But likeAkshat [00:51:25]: That's not enough.Swyx [00:51:25]: Yeah.Akshat [00:51:26]: You need to give it Adobe.Swyx [00:51:27]: Yeah, I hadn't put it together with like it would be a video production thing. in my mind these things were going more towards editingAkshat [00:51:36]: Yeah.Vibhu [00:51:36]: Well, shout out Mantis.Akshat [00:51:37]: I think about this a lot.Swyx [00:51:38]: .Akshat [00:51:41]: Yeah. Sorry.Vibhu [00:51:41]: Luma. Luma Agent is a version of this for video production, but it's a off.Swyx [00:51:46]: I was gonna get your quick takes, on some other stuff that happensGitpod/Ona, CI, and Runtime SandboxesSwyx [00:51:50]: In recent news and just-just see if you have anything interesting. Gitpod, very li
Is quantum technology purely a science of the future, or can it solve major enterprise problems today? In this episode, Pouya Dianat, Chief Revenue Officer of Quantum Computing Inc. (QCi), joins host Konstantinos Karagiannis to discuss how his company is moving past the lab stage to deliver real business value right now. Dianat breaks down QCi's unique hardware approach using nonlinear photonics and thin-film lithium niobate, explaining why the company has focused on application-specific, analog quantum optimization machines rather than waiting for universal gate-based systems to mature in the 2030s. The conversation dives deep into QCi's latest breakthrough algorithm, CVQBoost, which leverages the physics of their Dirac 3 system to revolutionize financial fraud detection. Operating literally at the speed of light, this room-temperature, rack-mountable quantum computer eliminates classical processing bottlenecks and effortlessly bypasses the training limitations that plague traditional GPUs. Dianat also shares insights into how recent federal executive orders on quantum technology are impacting the sales landscape, QCi's cutting-edge quantum authentication solutions for network security, and his strategic vision for driving immediate commercial revenues while advancing a long-term quantum roadmap. For more information on Quantum Computing Inc., visit https://quantumcomputinginc.com/. Visit Protiviti at www.protiviti.com/US-en/technology-consulting/quantum-computing-services to learn more about how Protiviti is helping organizations get post-quantum ready. Follow host Konstantinos Karagiannis on all socials: @KonstantHacker Questions and comments are welcome! Theme song by David Schwartz, copyright 2021. The views expressed by the participants of this program are their own and do not represent the views of, nor are they endorsed by, Protiviti Inc., The Post-Quantum World, or their respective officers, directors, employees, agents, representatives, shareholders, or subsidiaries. None of the content should be considered investment advice, as an offer or solicitation of an offer to buy or sell, or as an endorsement of any company, security, fund, or other securities or non-securities offering. Thanks for listening to this podcast. Protiviti Inc. is an equal opportunity employer, including minorities, females, people with disabilities, and veterans.
This week on CMO Confidential, we are revisiting one of our favorite conversations with Rob Ward from January of 2026.A CMO Confidential Interview with Rob Ward, co-founder and General Partner of Meritech Capital, a top Silicon Valley venture firm. Rob shares his take on what he calls a "super terrifying and exciting time" and provides perspective on AI receiving the most capital of any technology in history, the "durability of revenue" and how quickly start-ups are now reaching $100 million in revenue. Key topics include: why VC's focus on growth vs. profitability; the risks associated with massive long-term capital investment; why marketers should pick a "trusted advisor" as their AI partner; and why your data strategy needs "context. Tune in to hear how Astronomer handled the "Coldplay Concert Incident" which immediately became a PR classic and the "VC Foie Gras Effect."What happens when a top venture capitalist pulls back the curtain on AI, valuations, hype cycles, and what's actually working?In this episode of CMO Confidential, host Mike Linton sits down with Rob Ward, Co-Founder and General Partner at Metech Capital, to unpack the realities behind the AI boom. Rob has spent more than 26 years investing in category-defining companies like Facebook (Meta), Snowflake, NetSuite, Zipcar, and Cloudera — and he brings a rare, grounded perspective to today's AI frenzy.Together, they explore: • Why AI adoption is still early — despite explosive growth • The real risks behind inflated valuations and “AI-washing” • How VC decision-making changes during platform shifts • What marketers and executives should actually look for when choosing AI partners • Why data strategy, change management, and trust matter more than tools • What layoffs, productivity, and the future of work really look like beneath the headlines • A masterclass in crisis communications, featuring Ryan Reynolds, Gwyneth Paltrow, and ColdplayIf you're a CMO, CEO, board member, founder, or agency leader trying to make sense of AI without getting swept up in the hype — this is a must-listen conversation.⸻Chapter Markers00:00 – Welcome to CMO Confidential00:19 – Introducing Rob Ward and today's AI conversation01:13 – Where we really are in AI adoption02:26 – Explosive AI growth: what's real vs hype03:35 – Why enterprise AI adoption is still a slog04:37 – Vendor spend, hyperscalers, and the trillion-dollar buildout06:12 – Is this an AI bubble? Public vs private market realities07:20 – Accelerating investment rounds and lack of diligence08:12 – AI-washing and durability of AI businesses09:46 – Proof-of-concepts, switching costs, and fragile loyalty10:55 – Big Tech vs startups: why this cycle is different11:40 – Why VCs chase platform shifts despite the risks13:05 – How AI is changing profitability and headcount math16:11 – “FOGRA” investing and capital distortion17:00 – Circular investing and data-center risk18:23 – Data centers, GPUs, and betting on the wrong future19:38 – Credit default swaps and financial warning signs21:45 – How executives should choose AI vendors22:58 – Change management and why culture matters most24:09 – Why data strategy is the real AI strategy26:36 – “Frequently wrong, never in doubt” and AI hallucinations27:01 – Practical AI use cases for marketers30:00 – Layoffs, productivity, and what's really happening to jobs33:05 – The best questions to spot real AI fluency35:00 – AI safety, geopolitics, and long-term risks36:38 – Crisis management masterclass: Astronomer, Coldplay & Ryan Reynolds39:58 – Final advice and closing thoughts⸻Subscribe for weekly episodes featuring world-class marketing leaders, board members, and C-Suite executives.#CMOConfidential, #MarketingLeadership, #BrandStrategy, #CorporateActivism, #MarketingStrategy, #CMO, #AIinMarketing, #ExecutiveLeadership, #BrandReputation, #ConsumerTrust, #DigitalMarketing, #MarketingInsights, #ThoughtLeadership, #BusinessStrategy, #CustomerCentricSee Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
Marvell (NASDAQ: MRVL) stock has surged since Jensen Huang called it the next trillion-dollar company — we ran a reverse DCF on the Q1 FY2027 earnings to see if the rally is justified.Marvell just posted record Q1 fiscal 2027 results, with revenue up 28% year-over-year to $2.4 billion and guidance pointing to 35% growth next quarter. As a fabless chip designer and IP licenser, Marvell is riding two major tailwinds: hyperscaler demand for custom AI chips (ASICs and XPUs) and networking systems that interconnect GPUs across data centers, including a growing partnership with NVIDIA via NVLink.We break down the earnings report, the CFO transition to Dan Durn (formerly of Adobe and Applied Materials), the balance sheet impact of the Celestial AI acquisition, and rising share dilution from recent deals. Then we run a reverse discounted cash flow analysis on Marvell at its current price to determine what growth rate is already priced in, and whether the semiconductor cycle and AI infrastructure buildout can realistically support it.If you're weighing whether Marvell is still a buy after this rally, or looking for better value elsewhere in the semiconductor supply chain, this one's for you.Semi Insider members get access to our full research platform and tools, plus deeper research as it happens. Join at chipstockinvestor.com.Get 15% off your Fiscal.ai membership with our link: fiscal.ai/csiContent in this video is for general information or entertainment only and is not specific or individual investment advice. Forecasts and information presented may not develop as predicted and there is no guarantee any strategies presented will be successful. All investing involves risk, and you could lose some or all of your principal.CSI doesn't own shares of Marvell.
In this episode, Alex Zenla (CTO/Co-founder, Edera) challenges the "laissez-faire" attitude toward modern infrastructure. She promotes "spite-driven development", building software to solve genuine technical pain points rather than passively accepting flawed abstractions, as a philosophy of improving the world of software. The discussion touches on the fragility of the current cloud-native stack, the security risks of multi-tenant Linux kernels, and the inefficiency of repurposing consumer-grade GPUs for AI workloads. Zenla also offers a pragmatic framework for the "AI-native" engineer: treat LLMs as symbiotic assistants for deep learning, not replacements for system-level expertise. Read a transcript of this interview: https://bit.ly/3QZQhxc Newsletter: Subscribe to the Software Architects' Newsletter, a monthly roundup of the patterns and technologies senior practitioners are working through, with the news and lessons from people doing the work: https://www.infoq.com/software-architects-newsletter InfoQ Online Certification Programs: 5-week online cohorts for senior engineers and architects, built around QCon talks. Programs now cover software architecture, AI engineering, and organizational architecture. Each week you join a four-hour live session with a confidential peer group of practitioners from other companies, apply frameworks from QCon talks to the decisions you're making at work, and earn an InfoQ certification. You leave with new approaches, or confirmation that the calls you're already making are the right ones. Learn more: https://certification.qconferences.com/ Upcoming Events: QCon San Francisco 2026 (November 16-20, 2026) https://qconsf.com/ QCon London 2027 (April 13-16, 2027) https://qconlondon.com/ The InfoQ Podcasts: Weekly conversations with senior software leaders about how they build systems and teams, including what they'd do differently. Listen to all our podcasts and read interview transcripts: The InfoQ Podcast: https://www.infoq.com/podcasts/ Engineering Culture Podcast by InfoQ: https://www.infoq.com/podcasts/#engineering_culture Generally AI: https://www.infoq.com/generally-ai-podcast/ Follow InfoQ: Mastodon: https://techhub.social/@infoq X: https://x.com/InfoQ LinkedIn: https://www.linkedin.com/company/infoq/ Facebook: https://www.facebook.com/InfoQdotcom Instagram: https://www.instagram.com/infoqdotcom/ YouTube: https://www.youtube.com/infoq Bluesky: https://bsky.app/profile/infoq.com Write for InfoQ: Share what you've learned building software with a community of senior practitioners, and get your work in front of the people who read InfoQ. https://www.infoq.com/write-for-infoq
Join us as Du'An digs into the real mechanics of running AI locally and in production - from GPU memory math to multi-agent architectures, observability, and the economics of self-hosted inference. Du'An walks through how model weights and KV cache compete for GPU memory, why continuous batching matters when you have more than a handful of users, and how agent architectures like single-agent, workflow, graph, swarm, and supervisor patterns each solve different problems. You will learn how to instrument your agents with Langfuse for observability and cost tracking, when to use Ollama versus vLLM, how prompt caching can cut provider costs by up to 75%, and why GPUs should never sit idle. Episode two of three - the next episode covers deploying at scale. Timestamps 0:00 Welcome & Introduction 1:47 Du'An's New Role at Akamai Cloud 3:10 Data Privacy and the Case for Self-Hosted AI 7:21 Anthropic and OpenAI as the New Cloud Layer 12:48 Local Models for Specific Use Cases - Cancer Detection Example 15:02 GPU Memory Math - Weights, KV Cache, and Context Windows 19:32 Continuous Batching and GPU Time Slicing 20:03 Observability with Langfuse - Live Demo 27:44 Agent Architectures - Single Agent, Workflow, Graph, Swarm, Supervisor 36:36 Token Economics, Prompt Caching, and GPU Cost Planning 45:32 Ollama vs vLLM - Prototyping vs Production How to find Du'An: https://duanlightfoot.com https://www.linkedin.com/in/duanlightfoot/ Links from the show: https://langfuse.com/ https://github.com/akamai-developers/akamai-workshop-solution-architect-agent https://amzn.to/4bvHn1p https://vllm.ai/
The primary source introduces an ethics-by-design architecture that embeds philosophical reasoning into AI development pipelines using a formalized triple-gate structure. This framework implements Metric, Governance, and Eco gates at every stage of the lifecycle to enforce quantitative safety, legal compliance, and environmental sustainability through carbon and water budgets. A secondary source discusses the rise of neuromorphic computing, highlighting how these brain-inspired chips offer a more energy-efficient alternative to traditional GPUs for edge AI applications. Together, these texts emphasize a shift toward responsible AI governance that prioritizes measurable accountability and ecological impact alongside technical performance. By merging moral philosophy with engineering controls, the proposed framework seeks to prevent ethical failures from propagating throughout emerging AI paradigms.
Centamil AI is building AI-embedded microreactors that enable process industries to transform their facilities for higher-intensity workloads. The company recently launched its Hayagreeva HEx, a cold plate that enables data centers to accommodate higher density GPUs driven by the AI revolution. Centamil offers a robust AI layer on top of its physical infrastructure, which optimizes cooling parameters in real time to adapt to the conditions at a facility. We spoke with Praveen Gorakavi, Founder and CEO of Centamil AI, to learn more about why he sees cooling as a critical bottleneck to AI infrastructure deployments. We discuss how the Hayagreeva cold plate is improving the energy efficiency of data centers by relieving the burden placed on energy intensive air cooling systems. Praveen also shares what inspires him as an entrepreneur and how he balances his roles as both an inventor and a businessman. And follow us on: Newsletter: https://www.energy-terminal.com/newsletter-signup LinkedIn: https://www.linkedin.com/company/energy-terminal Instagram: https://www.instagram.com/energyterminal/
NVIDIA is a chip company. So why does it put hundreds of researchers on building AI models — and then give them away for free? Bryan Catanzaro is VP of Applied Deep Learning Research at NVIDIA and one of the people whose work quietly underpins modern AI: he helped create cuDNN (NVIDIA's first deep learning product), co-invented DLSS, and named and built Megatron, the framework behind how much of the industry trains large models. Today he leads Nemotron, NVIDIA's family of open models — and Nemotron 3 Ultra, released just weeks ago, is one of the strongest open-weights models to come out of the US.Matt Turck sits down with Bryan for a genuinely deep conversation: the real business logic behind a chip company building its own models, the state of open vs. closed AI, and whether the US is falling behind China in open models. Then they go inside Nemotron itself — four-bit (NVFP4) pretraining, hybrid Mamba-Transformer architecture, mixture-of-experts, multi-token prediction, and multi-teacher distillation — all explained in plain language. Plus a rare look at how a modern AI research org actually runs, what it was like working alongside Andrew Ng and Dario Amodei at Baidu, why Bryan doesn't believe in the singularity, and his contrarian case that open AI is safer than closed.A reference conversation for anyone trying to understand where AI is really headed.(00:00) — Cold open & Intro(01:33) — Is open source AI catching the frontier?(05:29) — Do closed labs blocking distillation slow open source down?(07:42) — Is the US falling behind China?(10:30) — Why companies actually choose open models(12:39) — A "crazy" 2008 bet: machine learning on GPUs(15:33) — Working with Andrew Ng and Dario Amodei at Baidu(17:41) — Coming back to NVIDIA: DLSS and the birth of Megatron(21:55) — The real reason NVIDIA builds its own models(24:28) — Is Moore's Law really dead?(33:37) — The Nemotron family: Nano, Super, Ultra(35:09) — Built for agents: why NVIDIA bets on speed(36:02) — How you train a 550B model in 4 bits(39:25) — Hybrid Mamba-Transformer, explained simply(42:31) — Mixture of experts — and why NVIDIA built NVL72 around it(47:26) — Why a 1-million-token context window matters(49:26) — Multi-token prediction: how the model predicts 5 tokens at once(52:47) — Multi-teacher distillation: teaching one model from many(58:01) — Where reinforcement learning goes next(01:00:16) — Inside NVIDIA's research org: "the mission is the boss"(01:04:03) — How NVIDIA decides who gets the GPUs(01:10:53) — Why NVIDIA still feels entrepreneurial after 33 years(01:12:58) — Why Bryan doesn't believe in the singularity(01:17:50) — The AI backlash(01:19:18) — The controversial case: open AI is safer than closed
Competition for AI Is Coming From a Surprise Source That Could Pressure U.S. Companies' Prices and Profits We tend to focus on the major AI companies in the United States and assume they will be the long-term winners. However, one competitor that cannot be ignored is China. Chinese companies are making rapid progress in artificial intelligence, and they could become a serious challenge to U.S. firms. Don't forget that China is a communist country and the government can put in a lot of capital to win the AI race. That ability to heavily fund AI development could help Chinese companies narrow the gap with, or even surpass, some American competitors in certain areas. According to Artificial Analysis, which evaluates the capabilities of large language models, China's Z.ai ranked among the top three globally with its latest model release. Another concern is cost. Z.ai is reportedly offering models at less than half the price of many American rivals. Lower prices could make it easier for the company to gain market share while putting pressure on the pricing and profit margins of U.S. AI companies. I certainly don't want to see American companies lose ground to Chinese competitors. However, as investors, we have to evaluate the competitive landscape objectively. U.S. AI companies have already committed hundreds of billions of dollars to infrastructure and development. If competition forces prices lower, it may take much longer for these companies to generate the profits needed to justify today's lofty stock prices and valuations. The Business of Kids' Sports Is Changing and it May Not Be for the Better Private equity has made its way into nearly every corner of the economy, and now it's becoming a major force in youth sports. The Aspen Institute has estimated that youth sports are now a $40 billion industry in the U.S, which is likely why private equity is now targeting the space. That's raising serious concerns about what happens when maximizing investor returns becomes more important than giving kids affordable opportunities to play. As private equity firms buy up leagues, tournaments, training facilities, and sports complexes, critics argue the result is less competition, higher registration fees, and fewer affordable options for families. The average cost of youth sports has increased dramatically in recent years, leaving many children priced out of participating simply because their families can't afford it. One thing that stands out is that this has become one of the rare issues drawing concern from both Republicans and Democrats in Congress. Burgess Owens, a Republican from Utah and former professional football player, pointed out “Investment is important, but it's when the mission is our kids, not investors. We're seeing too much of this. We're going to lose the soul of our nation if we don't get this right.” He also acknowledged that while some investors are doing it the right way, bad actors need to be kept out. While there are differences over how to address the problem, there appears to be broad bipartisan agreement that rising costs and reduced consumer choice deserve closer scrutiny. Youth sports should be about developing character, teamwork, friendships, and healthy competition, not creating another industry where financial engineering determines who gets to participate. If the trend toward consolidation continues unchecked, more families may find themselves priced out of opportunities that should be available to every child, regardless of income. Bitcoin Company Strategy Is in Trouble Strategy, formerly known as MicroStrategy, changed its name after the company essentially became a leveraged bet on Bitcoin rather than a software business. As management shifted its focus almost entirely to buying Bitcoin, it dropped the "Micro" from its name to reflect that new identity. CEO Michael Saylor spent years promoting Bitcoin and telling investors that owning Strategy stock was one of the best ways to benefit from its rise. To finance those Bitcoin purchases, the company repeatedly issued low-interest convertible bonds. The next major maturity comes on September 15, 2027, when approximately $1 billion of convertible notes become due. If you're unfamiliar with convertible bonds, they allow a company to borrow money at lower interest rates because investors have the option to convert the bonds into stock instead of receiving cash repayment. For that to happen, however, the stock price must trade well above the conversion price. In this case, the conversion price is about $183 per share, about double the current stock price of roughly $90. Unless the stock stages a dramatic recovery, those bonds are unlikely to be converted into shares, meaning Strategy would need to repay the $1 billion in cash. The stock has fallen nearly 79% over the past year, and Bitcoin's decline has only magnified the losses. Bitcoin itself has dropped roughly 50% from its peak, falling below $60,000 depending on the day. When Bitcoin was making new highs, investor excitement seemed endless. Now that prices have been cut roughly in half, much of that enthusiasm has disappeared. Michael Saylor has also been noticeably absent from major interviews in recent months. Whether that is because demand for his appearances has faded or because the company's performance has made those appearances more difficult is open to interpretation. Strategy stock reached a high of around $473 in late 2024 and now trades near $90. We've discussed this company many times before. The concern has always been that Strategy is not creating meaningful operating growth as it is primarily just borrowing money to buy Bitcoin. Unlike a traditional operating company, it is not relying on expanding products or services to drive future earnings. At the moment, there does not appear to be a clear catalyst that would significantly lift either Bitcoin or Strategy's stock price. If the shares remain well below the conversion price as the 2027 maturity approaches, investors are likely to become increasingly concerned about how the company will repay its debt. That uncertainty could continue to put pressure on the stock. Upper-middle-class Americans may not be as financially secure as they would like Upper-middle-class Americans, generally defined as households earning between $150,000 and $250,000 per year, may be in a stronger financial position than most, but many are becoming increasingly pessimistic about the future. You may be surprised to learn that 86% of upper-middle-class Americans do not believe their children will have a better life than they have. Just seven years ago, in 2019, that figure was only 64%. Many upper-middle-class households are also losing confidence in the economic system and the government. They increasingly feel that the odds are stacked against them, making it harder to continue moving ahead financially. In the most recent Wall Street Journal survey, 65% of affluent Americans said they believe the system is rigged against them, more than double the 29% who felt that way in 2017. The news isn't much better for the middle class, generally defined as households earning between $65,000 and $235,000 annually. Only 25% said they have been able to save beyond an emergency fund. Roughly one in four also reported carrying credit card debt that they are unable to pay off in full each month. Despite these concerns, there has still been significant upward mobility. About 75% of people in today's upper-income group said they now belong to a higher economic class than the one they grew up in. Among middle-class Americans, roughly half said they also grew up in a lower economic class than where they are today. Views on higher education are changing as well. About one-third of middle-class Americans no longer believe a four-year college degree is the best path to financial success. Rising tuition costs, growing student debt, and the availability of alternative career paths have caused many to rethink the traditional college route. No matter which income group people belong to, there is often a desire to improve their financial situation and move up economically. That ambition is a healthy part of human nature and is often what drives people to work harder, save more, and invest for the future. While constantly striving for more can sometimes make it difficult to feel fully satisfied, the pursuit of improvement can also provide a strong sense of purpose and accomplishment. Did The Recent Jobs Report Tell the Whole Story? At first glance, this weeks jobs report looked fairly uneventful. The U.S. economy added 57,000 nonfarm payroll jobs in June, and the unemployment rate fell to 4.2%. This was below the estimate of 115k, but it does follow three strong months of payroll growth. After looking through the report, there are several numbers that raise some important questions. The first is the labor force. About 720,000 people left the labor force in June, pushing the labor force participation rate down to 61.5%, the lowest since March 2021. Even more troubling is that if we exclude the Covid-era, it was the lowest labor force participation rate in exactly 50 years. When people stop looking for work, they are no longer counted as unemployed, which can make the unemployment rate appear stronger than it otherwise would. Another surprising number was leisure and hospitality, which lost 61,000 jobs. June is typically one of the strongest hiring months of the year for hotels, restaurants, entertainment, and travel-related businesses. The Bureau of Labor Statistics attributed much of the decline to weaker-than-normal seasonal hiring, but it's still worth asking whether this reflects a temporary statistical issue or an early sign that consumer spending is beginning to soften. It is especially strange given the popularity of the World Cup and many speculated this would be a strong sector in the report. Goldman Sachs in particular estimated a gain of 40k in leisure and hospitality before the report was released. Then there is the latest JOLTS report. Job openings stood at 7.6 million in May, showing employers are still looking for workers, but the question is if people are actually leaving the workforce can these jobs actually get filled? One report never tells the entire story, but these numbers deserve a closer look. Was June simply an odd month because of seasonal adjustments? Or are we beginning to see a labor market that is slowing more quickly than the headline unemployment rate suggests? The next few months of data should help answer that question. The biggest risk in AI may not be the technology, it may be the economics. This week, Bradley Tusk and Ed Zitron raised important questions that investors shouldn't ignore. Bradley Tusk (founder and CEO of Tusk Ventures and a venture capitalist) made an interesting observation: investors are treating frontier AI models the same. But China's AI companies are proving that powerful models can be developed much more cheaply and improve much faster than many expected. If lower-cost models continue to narrow the performance gap, AI models could become increasingly commoditized, making it much harder for companies spending hundreds of billions of dollars on infrastructure to earn attractive returns. Ed Zitron (author, podcaster and tech industry critic) echoed a similar concern from a different angle. He argues that AI companies are engaged in an expensive arms race, pouring enormous amounts of capital into chips, data centers, and model development without proving that the economics will justify the investment. As he has said, companies are "burning money at an astonishing rate" while investors continue to assume future profits will eventually catch up. This also ties into a warning from co-funder and CEO of Palantir Technologies, Alex Karp . He has criticized what he calls "token maxxing"—the idea that success in AI is simply about generating more tokens, building bigger models, and spending more on compute. Karp's point is that producing more AI output doesn't automatically create more business value. The companies that ultimately win will be the ones that solve real customer problems and generate durable profits, not necessarily those that consume the most GPUs or produce the most tokens. History shows that revolutionary technologies don't always produce the best investments. The internet transformed the world, but many of the biggest companies of the dot-com era disappeared because expectations got too far ahead of profits. AI will almost certainly reshape the economy. The bigger question for investors is whether the companies making the largest investments will ultimately earn the returns the market is expecting—or whether AI models become increasingly commoditized, leaving the biggest winners to be the businesses that successfully apply AI rather than simply build larger models. Financial Planning: Trump Account Investment Options Released Ahead of $1,000 Seed Funding Trump Accounts are expected to receive $1,000 of government seed money as soon as the 4th of July. If you have a child born in 2025 through 2028, you can apply online now at trumpaccounts.gov. This is basically a retirement account with a caveat, contributions can be made on behalf of children even if they don't have earned income. However extra contributions are made on an after-tax non-Roth basis so no upfront tax deduction and no tax-free growth. Instead contributions establish cost basis and investment earnings grow tax-deferred, but are ultimately taxed upon withdrawal at ordinary income rates. In practice, this tax deferral benefit is overstated. This week the Treasury Department released 5 investment options: SPYM, IVV, VTI, ITOT, and SPTM. These are virtually all the same investment, a low fee fund that is heavily weighted toward the largest US companies. This means there is no reason to sell or rebalance, so the only real option is to buy and hold. Buying and holding can also be done in a regular brokerage account with tax deferred until sale, but at the lower, potentially 0%, long-term capital gains rates rather than the higher ordinary income rates. Some planning strategies involve funding the Trump account and later converting it to a Roth. However, those conversions would still trigger tax at ordinary income rates and potentially trigger the kiddie tax, pulling the income into the parent's tax bracket. Since in every possible situation, the long-term capital gain tax rate is always less than the ordinary income tax rate, a better strategy may be to fund a brokerage account and use the future proceeds to make contributions to Roth accounts which likely could be done tax-free rather than funding a Trump account and eventually making Roth conversions at a higher rate. For this reason, while the $1,000 government seed contribution is worth it, additional voluntary contributions may be less attractive compared to already available alternatives. Companies Discussed: Meta Platforms, Inc. (Ticker: META)
Hoy hablamos de OpenAI planteando ceder un 5% al Gobierno de Estados Unidos, Meta preparando una nube de IA para vender computo, Nvidia financiando neoclouds con revenue share, el rumor del telefono de SpaceX con xAI y Snapdragon, y la mision de NASA para salvar el telescopio Swift elevando su orbita.Puedes seguirnos en YouTube en https://youtube.com/olivernabani y puedes unirte al Discord Mashain en https://olivernabani.com/discord
Carmen Li thought it was a joke when a perpetual futures marketplace asked her to become the market maker for her own index. It wasn't. In this segment from Bits + Bips: The Interview, she walks Steven Ehrlich through the requests that alarmed her, a daughter analogy for why trading your own benchmark destroys neutrality, the manipulation risks she sees in crypto's index practices, and why she insists any perp venue on her index be regulated and guardrailed. Host: Steven Ehrlich - Host of Bits + Bips and Head of Research at Sharplink Guest: Carmen Li - CEO of Silicon Data and Compute Exchange This clip is from a longer conversation on GPUs, compute markets, and crypto. Full episode here: https://www.youtube.com/live/rYDiPneJv20?si=fjS7bSd-bJ6c6tYb We go live every Monday - subscribe to catch it live. Sponsors Cape: Your biggest crypto vulnerability isn't your wallet, it's your phone number. Cape is America's privacy-first mobile carrier that rotates your SIM identity daily and blocks SIM swaps before they happen. Get 33% off your first six months at https://cape.co/unchained (use code: UNCHAINED). Chapters
When AI is at its best, the conversations can feel uncanny—almost magical in their accuracy, relevance, and speed. For that you can thank the AI agents that work together behind the scenes to search, reason, and sift through all your content to get you what you need to do your job. We talk with Jongmin Baek and Marta Mendez, two Dropbox machine learning engineers, about building conversational AI that's helpful, useful, and grounded in your team's shared context, so you can spend more time on the work that really matters. ~ ~ ~ Working Smarter is brought to you by Dropbox. Find, organize, and share your work—all in one place—with context-aware AI from Dropbox. You can listen to more episodes of Working Smarter on Apple Podcasts, Spotify, YouTube, Amazon Music, or wherever you get your podcasts. To read more stories and past interviews, visit workingsmarter.ai This show would not be possible without the talented team at Cosmic Standard: producer Ben Montoya, sound engineer Aja Simpson, technical director Jacob Winik, and executive producer Eliza Smith. Special thanks to our illustrator Fanny Luor, marketing consultant Meggan Ellingboe, and editorial support from Catie Keck. Our theme song was composed by Doug Stuart. Working Smarter is hosted by Matthew Braga. Thanks for listening!
OpenAI has released GPT-5.6, but the majority of us will have to wait. ⌚After the Anthropic vs. U.S. Government feud, it now looks like we'll have to wait for frontier models. That wasn't the only big AI news headline that might change your company's AI strategy. Anthropic got the green light to roll out Mythos 5 to a select few, Google reportedly extended its strike team to catch up on coding and more. OpenAI's limited release of GPT-5.6, Mythos starts slow reinstatement, OpenAI gets spicy and more AI news -- An Everyday AI chat with Jordan WilsonNewsletter: Sign up for our free daily newsletterMore on this Episode: Episode PageToday's Episode on LinkedIn: Thoughts on this? Join the convo on LinkedIn and connect with other AI leaders.Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineupWebsite: YourEverydayAI.comEmail The Show: info@youreverydayai.comConnect with Jordan on LinkedInTopics Covered in This Episode:OpenAI GPT-5.6 Limited Release ExplainedOpenAI Sol, Terra, Luna Model NamingUS Government Restrictions on AI RolloutsAnthropic Mythos 5 Access and StandoffAnthropic Fable 5 Suspension DetailsGoogle Gemini 3.5 Pro Release DelayedGoogle's AI Coding Mid-Training InitiativeRaiseUS Nonprofit: AI Workforce AdaptationAnthropic Accuses Alibaba of Model DistillationOpenAI & Broadcom Unveil Jalapeno AI ChipKey AI Industry Partnerships & Product LeaksTimestamps:00:00 OpenAI's GPT 5.6 limited release06:14 OpenAI's new model release details09:54 Access suspension and negotiations13:18 Google's AI strategy and delays16:33 Anticipating Gemini 3.5 Pro Release20:40 Accusations of AI model theft24:48 OpenAI and Broadcom chip partnership28:05 OpenAI's recent developments and updates29:56 OpenAI and AI weekly updatesKeywords: GPT-5.6, OpenAI, Anthropic, Mythos 5, Fable 5, Frontier models, Gemini 3.5 Pro, Google, model rollout, limited AI access, AI safety, US government AI regulation, Sol model, Terra model, Luna model, Max reasoning mode, Ultra mode, sub agents, advanced AI benchmarks, coding workflows, cybersecurity, third-party AI analysis, government licensing, AI model guardrails, AI model democratization, model naming scheme, model availability, AI model security, jailbreak resistance, safety filters, general model access, trusted testers, AI export control, national security, Anthropic pullback, supply chain risk, defense department, AI industry competition, talent loss, AI coding, mid training, engineering agents, AI strike team, RaiseUS nonprofit, workforce AI disruption, technology policy, industrial scale distillation, Alibaba, AI model theft, China-US tech tensions, distillation attacks, Jalapeno AI chip, Broadcom, AI inference, custom hardware, data center GPUs, Microsoft, Meta, Elastic compute, AI-powered career navigation, Slack Claude Tag, Canva Grow 2.0, Copilot skills, AI ad creation, AI automation, DigitalOcean plugin, Apple hardware AI, smart glasses, Vision Pro, portfolio tracking AI, Google Finance, home smart speakers, voice AI, GLM 5.2, open source AI, US labor market AI effects, AI job disruption, model leaks, government approval delays.Send Everyday AI and Jordan a text message. (We can't reply back unless you leave contact info) Start Here ▶️Not sure where to start when it comes to AI? Start with our Start Here Series. You can listen to the first drop -- Episode 691 -- or get free access to our Inner Cricle community and all episodes: StartHereSeries.com Also, here's a link to the entire series on a Spotify playlist.
Three massive semiconductor and computing developments are reshaping the future of AI infrastructure — and 7investing's Simon Erickson sits down with Nick Rossalillo of Chip Stock Investor to break them all down. First up: Cerebras Systems (NASDAQ:CBRS), which just went public on May 13th at $185/share (~$40 billion valuation) and is now trading near $46 billion at 90x trailing sales. The company's Wafer Scale Engine, a chip that uses an entire silicon wafer rather than individual diced chip, was designed specifically for AI inference workloads that NVIDIA (NASDAQ:NVDA) GPUs struggle to handle efficiently due to on-chip SRAM limitations. With potential $20 billion in orders from OpenAI and access via AWS, Cerebras is real, but neither Simon nor Nick is buying at this price. Their rule: wait a year before touching a fresh IPO.Next, SpaceX's freshly-raised $75 billion gets put under the microscope, specifically Elon's ambition to build orbital data centers. Nick walks through the SpaceX diagram: 70-meter solar panel wingspan, laser-based networking between compute modules, and the massive engineering challenges around power, heat dissipation, and in-orbit assembly. This isn't imminent, Starlink's next-gen constellation comes first — but if Elon can crack the economics, it would rewrite the rules of data center infrastructure entirely.Finally, Huawei's Tau Scaling announcement: a new architectural approach to chip performance that bypasses the need for extreme ultraviolet lithography (which China can't access due to ASML export controls). Tau temporal scaling focuses on minimizing signal travel time between transistors using logic folding, new materials, and 3D stacking. Huawei claims it could reach 1.5 nanometer equivalent performance by 2031. Simon and Nick are skeptical — 381 chips in six years is not mass production, and TSMC (NYSE:TSM) will be well past that node by then but it's worth watching as China continues building workarounds to Western export restrictions.Whether Moore's Law is dead or simply rerouting, the chipmaking industry is more innovative and more investable than it's been in decades.Join the conversation on the 7investing discord: https://discord.com/invite/PT9ZQqdXXSWant access to all 7investing research? Join at 7investing.com/subscribe Follow Chip Stock Investor @chipstockinvestor and https://chipstockinvestor.com/0:00 - Introduction to 7investing and Chip Stock Investor0:54 - Is Moore's Law Dead? A review of scaling semiconductor manufacturing3:08 - Cerebras Systems' recent IPO. How is their Wafer Scale Engine different than NVIDIA's GPUs, how does this impact Moore's Law, and is the stock a buy today?21:12 - SpaceX's recent IPO. Elon wants to build and launch orbital data centers. How does SpaceX plan to use the $75 billion it raised, what are the challenges it faces, and is the stock a buy?28:16 - Huawei's Tau scaling. Could this new chip architecture make ASML's extreme ultraviolet lithography obsolete, and what are its chances of succeeding?39:39 - Outro, final thoughts, and audience questionsStocks & Companies Mentioned:Cerebras Systems (NASDAQ:CBRS)NVIDIA (NASDAQ:NVDA)AMD (NASDAQ:AMD)SpaceX (SPCX)Taiwan Semiconductor Manufacturing Company / TSMC (NYSE:TSM)ASE Technology Holding / ASE Group (NYSE:ASX)Vicor Corporation (NASDAQ:VICR)ASML Holding (NASDAQ:ASML)Applied Materials (NASDAQ:AMAT)Lam Research (NASDAQ:LRCX)Intel (NASDAQ:INTC)Amazon / AWS (NASDAQ:AMZN)Alphabet / Google (NASDAQ:GOOGL)AST SpaceMobile (NASDAQ:ASTS)Samsung Electronics (KRX:005930)Huawei — private (Chinese company)OpenAI — privateLuckin Coffee (OTC:LKNCY) — mentioned as cautionary example#Semiconductors #MooresLaw #CerebrasSystems #CBRS #AIChips #NVIDIA #SpaceX #OrbitalDataCenters #HuaweiTech #TauScaling #ChipStocks #AIInvesting #TechStocks #GrowthStocks #StockMarket #InvestingIn2026 #7investing #Simonerickson
CNBC reported that Baidu's AI chip unit, Kunlunxin, is targeting a $50 billion Hong Kong IPO, sending Baidu shares up about seven percent. Kunlunxin designs AI accelerators deployed across Baidu AI Cloud and complements Baidu's Ernie large language model and Ernie Bot. The potential listing would test demand for semiconductor assets in Hong Kong and broaden Baidu's investor base. U.S. export controls on advanced GPUs have pushed Chinese firms, including Baidu, Huawei, Alibaba, and Tencent, to accelerate in house chip programs. A public offering could fund R&D, manufacturing partnerships, and software tools, while adding market discipline to the unit's performance. For founders, the move signals continued competition for AI compute, growing domestic silicon options in Asia, and the need to diversify across accelerators and regions.Learn more on this news by visiting us at: https://greyjournal.net/news/ Hosted on Acast. See acast.com/privacy for more information.
Samsung prepara una inversión gigante en Corea del Sur para chips, IA y robótica. DeepSeek presenta DSpark para acelerar la inferencia de V4, Google consolida la Interactions API de Gemini como infraestructura para agentes, China presume de LineShine en el TOP500 sin GPUs y Euclid publica una imagen del centro de la Vía Láctea con más de 60 millones de estrellas.Puedes seguirnos en YouTube en https://youtube.com/olivernabani y puedes unirte al Discord Mashain en https://olivernabani.com/discord
The Founding of OpenAI. Guest Author: Keach Hagey. In this opening segment, Keach Hagey discusses the January 2016 founding of OpenAI as a nonprofit research lab. Key figures included co-founder Greg Brockman and chief scientist Ilya Sutskever, a renowned researcher whose recruitment from Google signaled the lab's potential. Backed by a billion-dollar commitment from Elon Musk, Peter Thiel, and Jessica Livingston, the project was designed as a safe, non-commercial counterweight to Google's DeepMind. Operating initially out of Brockman's apartment, the team aimed to achieve Artificial General Intelligence (AGI) for the benefit of humanity. The technical foundation relied heavily on GPUs—hardware originally designed for video games—which proved essential for training the deep learning neural networks necessary for their research. This era was characterized by an ambitious, "pirate" spirit funded through YC Research to explore radical ideas outside the profit motive. 1JANUARY 1931
Apply to work with me one-on-one: https://www.cleartheshelf.com/applyHe bartended for 15 years. Then COVID hit, and he never went back. He flipped GPUs, then PS5s, then found Amazon. Now he's doing $200K months with online arbitrage at 15-16% net margins and pulling $36K a year in cashback on top of the profit.Blake (@Blake_FBA) is a seven-figure Amazon OA seller who's survived brand reviews, lost entire categories, and still keeps running it up. Today we get into how he actually sources, why repricing separates good sellers from great sellers, the cashback that pays for his life, and why he says fees only hurt you when your sourcing is weak.Chapters:00:00 - Bartender for 15 Years. Then Everything Changed.02:00 - GPU Flipping in a Discord and a Cash Deal With a Stranger04:15 - The Walmart Pallet That Made $50K at 1800% ROI08:00 - First OA Hire: Neither of Us Knew What We Were Doing13:15 - Why OA Has a Higher Ceiling Than RA16:00 - New Brands Nobody Heard of Two Years Ago20:00 - How He Went Wide Before Going Deep22:27 - Losing Brands and the Forced Pivot to Multi-Channel27:00 - eBay, Walmart, and Cross-Listing as Insurance30:45 - What a Sourcing Session Actually Looks Like37:26 - Repricing: The Skill That Separates Good From Great42:43 - $36K in Cash Back and a Backyard Bio Lab47:55 - The Year He Ran It Fully Automated (And What He Missed)51:21 - "There's a Biological Creature Running the Ship"55:16 - "Fees Only Hurt When Your Sourcing Is Weak"Follow Blake:X/Twitter: https://x.com/blake_fbaInstagram: https://www.instagram.com/blake_fba/Follow Chris Grant:X/Twitter/Instagram: @cleartheshelf Newsletter: https://cleartheshelf.com/newsletterFollow Chris Racic:X/Twitter: @ChrisRacicNewsletter: https://oaleads247.com
AI is changing everything—but there's one part of the conversation many organizations still aren't having: the economics. As enterprises race to deploy copilots, agents, and generative AI across every department, leaders are discovering that AI costs don't arrive as a single invoice. They show up across GPUs, token consumption, cloud infrastructure, data platforms, and idle compute resources, making it difficult to understand whether AI investments are actually delivering business value.In this episode of TECHtonic, TSIA Executive Director Thomas Lah sits down with Kunal Agarwal, CEO and co-founder of Unravel Data, to discuss why AI FinOps has become one of the most important disciplines for enterprise technology leaders. Kunal explains how organizations can optimize prompts, right-size AI models, eliminate wasted GPU capacity, and gain real-time visibility into the full AI technology stack. Together, they explore why AI should no longer be treated as a science experiment, how leading organizations are creating headroom to fund continued innovation, and why the companies that combine AI ambition with financial discipline will become tomorrow's AI-native market leaders.If you're responsible for AI strategy, cloud operations, infrastructure, finance, or technology investments, this episode offers a practical roadmap for balancing innovation with profitability—and ensuring your AI initiatives deliver measurable business outcomes.
Which stock is better? $MU or $NVDA - I go over how they are different and how I'm thinking about the markets with a straight forward strategy right now. The Seeking Alpha Summer sale continues with HUGE savings. Don't miss it - there are only 2 per year with discounts off what they normally provide. You'll have to wait until December to get the next one. SIGNAL STACK LINK --INCLUDED WITH TRENDSPIDER - GET YOUR PORTFOLIO ANALYZED BY SIDEKICK FORMULA - Alpha Picks + Seeking Alpha Premium + Trendspider and Sidekick - PERFECT TOGETHER! THESE SALES END SOON: TRENDSPIDER - JULY 4TH SALE THIS WEEKEND - get my 4 hour algorithm included on any annual plan.Seeking Alpha's SUMMER SALE *BEST DEAL - SEEKING ALPHA BUNDLE - Save over $250 and get Premium and Alpha Picks together - EXTRA $100 OFF ALPHA PICKS - Want to Beat the S&P? Save $124 EXTRA $74 OFFSeeking Alpha Premium ONLY - FREE 7 DAY TRIAL SEEKING ALPHA PRO - SAVE $600 EPISODE SUMMARY
Want the New iPhone 18 This September? Be Prepared to Pay More…. A Lot More The iPhone 18 is expected to be released in just a few months, and if current estimates are accurate, consumers could be facing some serious sticker shock. One of the biggest reasons is the ongoing battle for semiconductor components. The rapid buildout of AI data centers has created enormous demand for memory chips, and data center operators are willing to pay almost any price to secure supply. That is creating challenges for companies like Apple, which rely heavily on DRAM (dynamic random-access memory) and NAND flash storage. According to industry estimates, the cost of 12GB of DRAM used in the iPhone 17 was about $39. For the iPhone 18 Pro, that figure could rise to approximately $145. NAND flash storage costs are also expected to surge. The 256GB of flash storage that cost Apple around $13 in the iPhone 17 is projected to cost roughly $51 in the iPhone 18, an increase of nearly 300%. Apple may also introduce a redesigned camera system that could cost about 50% more than the cameras used in previous models, adding even more pressure to manufacturing costs. Apple currently earns an estimated gross margin of roughly 44% on the iPhone 17. If the company attempts to maintain those margins while absorbing these higher component costs, the price of a high-end iPhone 18 could climb to around $1,300 or more. The big questions are: Will Apple absorb some of these higher costs and accept lower profit margins? Or will consumers decide that the latest upgrade isn't worth the higher price and keep their current phones for another year? Either scenario could create headwinds for Apple's earnings. Lower margins would hurt profitability, while slower upgrade cycles could reduce unit sales. Both outcomes could put pressure on Apple's stock in the months ahead. Bad News: The Dollar Is Strong Again Some people may read that headline and think, "What's the problem? Isn't a strong dollar a good thing?" Not necessarily. A strong dollar sounds positive, but the reality is more complicated. The U.S. dollar is now at its strongest level since May 2025. While that may feel good on the surface, a stronger dollar can create challenges for the economy. When the dollar rises, American products become more expensive for the rest of the world to buy, which can worsen our trade deficit. At the same time, imported goods become cheaper for Americans. Consumers may enjoy lower prices on foreign products, but it also means more money flows overseas instead of supporting domestic businesses. Over the long term, that can weaken U.S. manufacturing, increase our reliance on imports, and contribute to growing debt levels. What's driving the dollar higher? Two major factors stand out. First, the new Federal Reserve leadership signaled a more hawkish stance at its most recent meeting. Nine of the 19 officials now expect at least one rate hike before year-end. Higher interest rates generally make the dollar more attractive to global investors. Second, the AI investment boom continues to fuel U.S. economic growth. However, the enormous capital required for AI infrastructure is leading companies to borrow heavily to finance those investments. This increased demand for capital competes with U.S. Treasury bonds for investor dollars, which could keep long-term interest rates elevated or even push them higher. The AI boom has already increased speculation and risk in the equity market. Now it may also be creating additional risks in the bond market. Wherever you're investing, make sure you understand the relationship between risk and reward before committing your capital Can Alphabet/Google Take Some of Nvidia's Market Share? Nvidia currently controls roughly 90% of the AI computing chip market. Whenever a company dominates an industry to that extent, it creates an opportunity for competitors to enter with comparable products at lower prices. That's exactly what Alphabet's Google is attempting to do with its artificial intelligence chips. Google originally developed its custom AI chips for internal use, but it quickly realized there was a much bigger opportunity. With demand for AI infrastructure exploding, Google is now producing more chips and making them available to outside customers. Nvidia CEO Jensen Huang has repeatedly stated, both publicly and privately, that increased competition will not have a meaningful impact on Nvidia's business. But what else can he say? Competition almost certainly will affect Nvidia to some degree. The company may eventually lose some market share and could be forced to lower chip prices to maintain its dominant position. Google has significant financial resources to support its AI ambitions. In western New York, for example, Google reportedly provided a $3.2 billion financial guarantee tied to the Lake Marina AI data center project. Nvidia has used similar strategies in the past to strengthen relationships with customers and partners. This type of financing does concern me. When you provide financing to a company that is also purchasing your products, you take on two risks. If that customer runs into financial trouble, you could lose both future product sales and repayment on the financing arrangement. I also suspect Nvidia has substantial leverage with many of its customers. Companies may worry that reducing purchases from Nvidia today could limit their access to future chip allocations if demand remains strong. Google isn't the only company challenging Nvidia. Competitors such as AMD, Broadcom, and newer entrants like Cerebras Systems are all looking for ways to gain a foothold in the rapidly growing AI chip market. Nvidia stock has delivered incredible returns over the past several years. The question investors should be asking is whether increasing competition and the possibility of future chip oversupply could eventually take some of the shine off Nvidia's valuation. The Dow's Alphabet Move Is a Sign of Weakness, Not Strength The Dow Jones is once again proving why it has become one of the most outdated and least useful stock market indexes in America. This week S&P Dow Jones Indices announced that Alphabet will be added to the Dow, replacing Verizon. The financial media is treating it like the Dow is finally modernizing itself for the AI era. I see it differently. This is not leadership. It is not vision. It is not smart index construction. It is the Dow doing what it has done for years: showing up late, after everyone else has already made the money. The Dow is supposed to represent the most important companies in the American economy. But unlike the S&P 500, it is not rules-based. There is no formula, no discipline, no objective threshold that decides who gets in and who gets kicked out. Instead, a committee at S&P Dow Jones decides when the index should change and which companies “feel right” for the list. That sounds harmless until you realize what it really means: the Dow is not a market index so much as a committee-curated museum exhibit that occasionally swaps out an old display piece for whatever has already become impossible to ignore. That is exactly what is happening with Alphabet. Google has been one of the most dominant businesses on earth for well over a decade. It has been central to digital advertising, cloud computing, mobile software, and now artificial intelligence. None of that is new. The AI spending boom did not start yesterday. The Magnificent Seven did not suddenly become important last week. These companies have been driving market returns, corporate profits, and capital spending for years. Yet only now does the Dow decide it needs more exposure to big tech? That is not being ahead of the curve. That is a lagging indicator pretending to be a benchmark. And the timing could not be more ridiculous. Instead of adding these companies before the market fully priced in their dominance, the Dow is adding them after the entire world has piled into the trade. After valuations expanded. After AI enthusiasm exploded. After mega-cap concentration became one of the biggest risks in the market. In other words, the Dow ignored the most important trend in the market for years and is now buying into it once the trade is crowded. The Dow will now hold five of the Magnificent Seven—Alphabet, Microsoft, Apple, Amazon, and Nvidia—which together will account for roughly 18% of the index. This is not modernization. That is panic buying in a suit. What makes it even more absurd is that the Dow still uses a price-weighted structure, which is one of the silliest relics in finance. A stock's influence in the index is determined by its share price, not by the actual size of the company or its economic importance. Think about how insane that is. In a supposedly elite index of America's biggest companies, weighting is still distorted by something as arbitrary as the sticker price of one share. A stock split can change a company's importance in the Dow more than a change in its business fundamentals. This also leads to more concentration with high priced stocks like Goldman Sachs accounting for roughly 13% of the entire index and Caterpillar making up around 12%. This compares to low priced stocks like Verizon or Nike which each only currently account for about 0.5% of the index. So now the Dow wants to have it both ways. It wants the credibility of owning AI and mega-cap tech leaders, but it wants to keep the same outdated structure and the same slow-moving committee process that made it miss the trend in the first place. It wants to look relevant without actually fixing what makes it irrelevant. Replacing Verizon with Alphabet may make the Dow look smarter for a headline or two, but it actually exposes the problem. The Dow did not identify the future. It waited until the future was obvious, then stapled it onto an old index and called it progress. The truth is the Dow has become a follower, not a leader. It reflects where the committee finally got comfortable going after the move already happened. And by adding more mega-cap tech exposure now, after years of delay, it may be doing exactly what bad investors do: chasing yesterday's winners while taking on tomorrow's risk. The Dow is not evolving. It is flailing. And every one of these late-stage reshuffles is a reminder that the most famous index in America may also be one of the least relevant. Fed Stress Test Confirms the Strength of U.S. Bank Balance Sheets U.S. banks once again came through the Federal Reserve's 2026 stress test looking structurally strong, even under an intentionally severe economic downturn scenario. The results continue to reinforce one of the most important post-financial-crisis themes: large banks today are built to withstand a shock that would have been destabilizing in prior cycles. The Fed's hypothetical scenario was deliberately harsh. It assumed a deep global recession with the U.S. economy contracting 4.6% and unemployment rising to around 10%. Housing prices would fall 30% from their current levels, the stock market would plunge 58% and there would be a 39% drop in commercial real estate prices. The framework is designed to test not just mild downturns, but a “worst plausible case” scenario that stresses bank balance sheets across multiple channels at once. Under that scenario, the Fed estimated cumulative losses across the largest 32 banks at roughly $700 billion, with the bulk coming from credit cards, corporate lending, and commercial real estate exposure. Despite those losses, all major institutions remained above required minimum capital levels. Capital ratios declined during the stress period, as expected, but stayed comfortably within regulatory buffers, underscoring how much capital has been built into the system since the 2008 financial crisis and subsequent regulatory reforms. What stands out this year is not just that banks passed, but the margin by which they did so. Even under simultaneous pressure from unemployment, real estate, and equity drawdowns, the system showed the ability to absorb losses while still maintaining lending capacity. That “lend-through-cycle” characteristic is one of the key goals of post-crisis regulation, and the results suggest it is functioning as intended. From an investor perspective, the more immediate implication is capital return. Passing the stress test is effectively the green light for banks to continue deploying excess capital back to shareholders. JPMorgan Chase unveiled a new $50 billion share repurchase program and said it will increase its quarterly dividend 10% to $1.65 per share, subject to board approval. Goldman Sachs and Wells Fargo increased their dividends 11% and Morgan Stanley boosted its payout by 15%. Importantly, the Federal Reserve did not materially tighten capital requirements in this round, which removes a potential headwind that some investors had been watching. Instead, capital rules remain broadly stable, allowing banks to operate with predictability in their capital planning. That stability is key, because it supports consistent buyback programs rather than volatile, stop-and-go capital return cycles. Taken together, the results reinforce a familiar but important conclusion: large U.S. banks today are not only capable of surviving severe macroeconomic stress, but they are doing so while generating enough earnings power to continue returning substantial capital through both dividends and buybacks. In a market where macro uncertainty remains elevated, that combination of resilience and shareholder yield continues to be a defining feature of the banking sector. What Is Quantum Computing All About? Quantum computing is the next big step in the evolution of computing, and there's no way around it: it's a complex subject. But it's also one of the most important technologies being developed today. If your son or daughter is in high school and unsure what they want to study in college, they may want to consider quantum physics, engineering, or computer science with a focus on quantum computing. Over the next decade, the world is going to need far more people who understand this field, whether that means working in quantum research labs, developing software, building hardware, or solving the many engineering problems that still stand in the way of commercial adoption. At its core, quantum computing is different from traditional computing because it uses quantum mechanics rather than classical binary logic. Today's computers rely on CPUs and GPUs that process information in bits or ones and zeros. Quantum computers use quantum processing units, or QPUs, powered by qubits. Qubits can behave in ways classical bits cannot, which gives quantum systems the potential to solve certain problems dramatically faster than even the most powerful computers we have today. There are currently four major approaches, or architectures, being used to build quantum computers: superconducting, neutral atoms, trapped ions, and photonics. Each has strengths and weaknesses, and no one yet knows which approach will ultimately dominate. But all of them are trying to achieve the same goal: building machines capable of solving problems that are effectively impossible for classical computers. That matters because the upside is enormous. Quantum computers could transform fields like drug discovery, materials science, logistics, finance, and artificial intelligence. They may also eventually be able to crack some of the encryption methods that protect today's digital world, which is one reason governments are taking the technology so seriously. It's not just a commercial race, it's increasingly a national security race as well. And that's where the geopolitical angle comes in. China has been heavily subsidizing quantum research. The future may not just be defined by military arms races, but by technology races, especially in areas like artificial intelligence, semiconductors, and quantum computing. The financial opportunity is also huge. By 2035, quantum computing is expected to generate roughly $43 billion to $71 billion in revenue. By 2040, some forecasts see that number climbing as high as $850 billion. Those are enormous figures for a technology that is still in its early innings, which helps explain why so much money is flowing into the space. I have to admit, quantum computing is both exciting and a little scary. A technology that can solve problems far faster than today's computers could open the door to incredible breakthroughs, but it could also create entirely new risks. Then again, that's true of almost every major technological leap in history. Progress is often uncomfortable at first, but it also has the power to reshape the world in ways we can't yet fully imagine. Financial Planning: Accessing Home Equity Homeowners tapped an estimated $47 billion of their roughly $11 trillion of home equity during the first quarter of 2026, the highest first quarter total since 2021. There are three primary ways to borrow against that equity. A cash-out refinance replaces your current mortgage with a larger one, but this generally only makes sense if today's interest rates are similar to or lower than your existing mortgage rate. That is unlikely for homeowners who locked in historically low rates during 2020 through 2022. A home equity loan functions as a second mortgage with its own fixed interest rate and monthly payment, making it a good choice when you need a lump sum for a specific purpose, such as a home renovation. A Home Equity Line of Credit (HELOC) is a revolving line of credit that allows you to borrow only what you need and repay it on your own schedule. While HELOCs typically have variable interest rates, they also provide the greatest flexibility and can make sense in today's interest rate environment. Regardless of which strategy you choose, home equity should be used to improve your overall financial position, such as consolidating high interest debt, funding value-adding home improvements, or purchasing appreciating assets. It should not be used to finance ongoing living expenses or discretionary spending. Companies Discussed: Netflix Inc. (NFLX)
We're excited to have Databricks join us at AIEWF, among hundreds of the top companies in the AI Engineer ecosystem. LS subscribers can use their discount to get past the late bird pricing and access over $50k in sponsor offers! Everyone is still talking about Satya's Frontier Ecosystems post, but few have actually built a (now $175 billion) frontier ecosystem and cloud like our guests today.From open-sourcing the layer above coding agents to rethinking databases for the agent era, Databricks cofounders Matei Zaharia and Reynold Xin are pushing the company beyond the lakehouse into a full data-and-AI operating system. In this episode, Matei and Reynold join swyx at the 2026 Data + AI Summit to unpack Omnigent, LTAP, Lakebase, agent security, open formats, Mosaic, and why databases may matter more than ever once AI agents start doing real work.We go deep on Omnigent: Databricks' open-source meta-harness for combining, controlling, and sharing agents across Claude Code, Codex, Cursor, Pi, custom agents, and internal tools. Matei explains why coding agents and enterprise agents run into the same problems: portability, collaboration, session history, security, spend controls, and the need for a common API above every harness.Then Reynold walks through Databricks' database dream: why CDC is brittle enough to joke that it means “continuous data corruption,” why HTAP has been the holy grail of database engineering, and why Databricks thinks LTAP gets most of the benefits by unifying the storage layer instead of collapsing every query engine. We also cover Databricks' infrastructure scale, the culture behind rapid prototyping, the difference between tech and enterprise customers, Databricks vs Snowflake, whether vector databases should have ever existed, the Mosaic model strategy, Genie, AI Runtime, RL fine-tuning, and the thesis that traditional software gets rewritten once the data is in the right place and agents sit on top.Databricks began as a company for the big data era. The origination of Spark from the Berkeley AMPLab which eventually turned into the product Lakehouse convinced enterprises that they didn't need a separate data lake, warehouse, ML platform, and governance layer. They just needed one open foundation where all of their data could live and be reasoned over.Since then a lot has changed, but data has only become more important. Data is no longer something you keep track of and analyze ad hoc, it's the necessary context agents need in order to act. So the framing has shifted from “where do we put all of our data?” to “how do we expose the right slice of state, history, permissions, and business logic to an AI system at the exact moment it's doing work?”If frontier model performance becomes commoditized, the durable advantage then becomes the company-specific context around them: proprietary data, governed access, operational state, transaction logs, workflows, and feedback loops. Which makes Databricks positioned perfectly.Now coming fresh off the Data + AI Summit 2026, the company is moving just as fast to keep up, announcing Genie One, Omnigent, LTAP, and many more, indicating a central mission in its newer work: Databricks is trying to become the operating system for enterprise agents.Models are getting good enough, but agents are only useful if they have the right context, permissions, memory, state, cost controls, and access to live business data. Fundamentally it appears that significantly better model performance in production is a systems problem, one that data guys like us are remarkably well prepared to solve!We discuss:* Why Databricks built Omnigent as a meta-harness above existing AI agents* Why coding agents and custom enterprise agents need the same infrastructure* The common API for agent sessions, files, streams, tool calls, and cancellation* Why persistent sessions, cloud sandboxes, sharing, search, and collaboration matter* Why Databricks open-sourced Omnigent instead of keeping it proprietary* Databricks' internal agent usage, cloud sandboxes, and coding workflows* The scale of Databricks: 50–60 million virtual machines a day and exabytes before breakfast* Why agent security needs contextual and stateful policies* How an agent could read confidential docs, install a compromised npm package, and leak data* Why spend control matters when an agent can burn $500 reading logs* Startup opportunities around coding-agent analytics, quality, skills, and spend* LTAP, Lakebase, and why Databricks wants to rethink the database stack* OLTP vs OLAP, CDC, and why data pipelines break at 3 a.m.* Why HTAP has historically been the holy grail of database engineering* Why Databricks thinks LTAP is “HTAP done right”* How writing transactional data into column-oriented formats changes analytics* Why agents need live operational context from databases, not just telemetry* How Databricks prototypes strategic systems without endless process* Enterprise vs tech customers, governance, procurement, and DIY culture* The “second system syndrome” risk of rewriting a database engine* Building a database engine from a decade of traces and quadrillions of data points* Why vector databases should never have been a separate category* Why open formats and AI changed the race with Snowflake* The Mosaic story, DBRX, Genie, document parsing models, and specialized model training* Why model customization and RL fine-tuning may become mainstream* Why “get the data there, slap some agent on top” may rewrite traditional softwareMatei Zaharia* LinkedIn: https://www.linkedin.com/in/mateizaharia* X: https://x.com/matei_zahariaReynold Xin* LinkedIn: https://www.linkedin.com/in/rxin* X: https://x.com/rxinDatabricks* Website: https://www.databricks.com* X: https://x.com/databricksTimestamps00:00:00 Introduction00:02:22 Omnigent and the Agent Infrastructure Layer00:08:39 Agent Clouds, Common APIs, and Open Source00:16:52 Databricks Scale and Internal AI Workflows00:18:03 Agent Security, Governance, and Spend Controls00:27:34 LTAP and the Database Dream00:30:30 CDC, HTAP, and Why Data Pipelines Break00:34:05 Lakebase, Parquet, and Live Data for Agents00:36:47 Databricks' Culture of Fast Prototyping00:43:40 The Dream Engine and Rewriting the Database Stack00:51:02 Vector Databases, Query Engines, and LTAP00:52:36 Databricks vs Snowflake00:57:48 Mosaic, DBRX, Genie, and Specialized Models01:03:11 Context, AI Runtime, and RL Fine-Tuning01:06:15 Why Data + Agents May Rewrite Software01:07:09 Closing ThoughtsTranscriptIntroduction: Databricks, Data + AI Summit, and Founder DynamicsSwyx [00:00:00]: Matei and Reynold from Databricks, welcome to Latent Space.Reynold Xin [00:00:06]: Hey, thanks for having us.Swyx [00:00:07]: Yeah.Matei Zaharia [00:00:08]: Yeah, thanks so much.Swyx [00:00:09]: thanks for taking time out. You have your Databricks, Data AI Summit going on. You were just telling me how the first summit that you guys ran was just 50 peopleReynold Xin [00:00:17]: Yeah, it wasSwyx [00:00:17]: in BerkeleyReynold Xin [00:00:18]: little meetup at Berkeley, I thinkMatei Zaharia [00:00:19]: YeahReynold Xin [00:00:19]: put togetherMatei Zaharia [00:00:20]: We were doing these tutorials and, yeah, just teach people Spark.Swyx [00:00:23]: Yeah. obviously now it's like, I think like the headline number's like 100,000 people around the world, 30,000 in person.Swyx [00:00:30]: it's a crazyMatei Zaharia [00:00:31]: AmazingSwyx [00:00:31]: community. Well, I just saw the keynote.Swyx [00:00:35]: Ali's just. Did was it obvious or that back when that Ali would be, like, such a great, like, CEO? LikeReynold Xin [00:00:42]: OhSwyx [00:00:42]: such a great presenter?Reynold Xin [00:00:43]: What do you think?Matei Zaharia [00:00:44]: I think among our group of founders it was clear that, I think he'd be the best at this.Swyx [00:00:50]: Yeah.Matei Zaharia [00:00:50]: And yeah, it turned out great. And he's, he's ramped up on so many topics growing a company. He would just go in and, like, study it and, be talk to all the experts. Like, even if he can't hire the person, learn enough about, like, finance and sales and whatever it was, and, and go from there. Yeah.Swyx [00:01:09]: Yeah.Reynold Xin [00:01:10]: he's obviously very high IQ and a very high EQ, but it wasn't. Like, Ali today is quite different from Ali from, like 10 years ago. I think there's a lot of work that he put in to, get to this point.Swyx [00:01:20]: Yeah. no, to me the most appealing thing about him is that he's funny. And like, it, it's, it'Matei Zaharia [00:01:26]: It's true, yeahSwyx [00:01:26]: it's hard to make jokes about, data warehousesReynold Xin [00:01:30]: About serious topicsSwyx [00:01:31]: securityMatei Zaharia [00:01:32]: YeahSwyx [00:01:32]: what have you.Matei Zaharia [00:01:33]: Oh, yeah. That's for sure.Swyx [00:01:34]: Yeah. So you guys launched a whole bunch of things. I'll, I'll just name check briefly, the stuff because we're not gonna cover everything. Omnigentt, your baby. LTAP, your baby, your dream engine.Swyx [00:01:47]: we're also gonna cover Genie, cover CustomerLake, you acquired PantherMatei Zaharia [00:01:52]: YeahSwyx [00:01:52]: Open Sharing, and there's Unity AI Gateway. A lot of these, I think, like, are things that you would expect a Databricks to do. It's, it's like part of the roadmap. Everyone in your category has similar things. But I think, probably the two of you are leading the two most unique and differentiated initiativesOmnigent and the Agent Infrastructure LayerSwyx [00:02:09]: on, in the landscape. Maybe we'll start with, Omnigentt we'll, we'll, we'll, we'll go into it. I do think that a lot of people are exploring this meta harness concept.Matei Zaharia [00:02:21]: Yeah, totally.Swyx [00:02:21]: What led you to it?Matei Zaharia [00:02:22]: Yeah. There were a couple of, like, converging lines, which I think is a good sign that you need something new. So on the one hand, there's all the coding agent info internally. We have really great, dev infra team. they built something called Isaac, that's like a wrapper on Claude Code and Codex, and, lets you use them either on the web in, like, sandboxes or, just on your dev machine or on your laptop or whatever. And then, they were adding all kinds of stuff there. And we saw all the more advanced engineers like, were building their own workflows with tons of agents, and they were building their own UIs and stuff on top or even on top of that. And then the other one was, like, us building agents. We ship this, like, data science agent called Genie on the research team, which I lead. We also build a lot of internal ones for various things, and then we have all the customer ones. And all of them running into this thing of like, “Oh, I need to switch model and harness and so on,” every few months. Plus the agent is, like, completely useless if you can't share sessions with someone and have history and have search and all this, like, layer on top of it for collaboration. I thought a bit about it from both contexts and, at first people thought it was weird. They're like, “Why are you doing coding agents and custom agents in the same thing?” But I said it's, it's the same problems and, you just wanna build the stuff that lets you deliver the agent, maybe control it if you care about security, and, make it portable across things. And then we prototyped some things as experiments. We saw, yeah, we can make it work, and then we built that for real.Swyx [00:04:06]: I'm wondering if this let's call it architectureMatei Zaharia [00:04:11]: YeahSwyx [00:04:11]: maps to anything in your careers in the past. like I always think about how a lot of things just tie back to operating systems.Swyx [00:04:18]: A lot of operatingMatei Zaharia [00:04:19]: YeahSwyx [00:04:20]: systems tie back to databases,Matei Zaharia [00:04:21]: SoSwyx [00:04:21]: or the other way aroundMatei Zaharia [00:04:22]: so the thing, I do think it ties a lot to, like, network protocols, internet protocol. we alsoSwyx [00:04:29]: Communication between entities.Matei Zaharia [00:04:30]: Yeah. We did stuff with, like, data sharing also, which is probably, most viewers probably won't know unless they'Swyx [00:04:36]: Yeah, open protocol is the term.Matei Zaharia [00:04:37]: Yeah.Swyx [00:04:38]: Open sharing. Open sharing.Matei Zaharia [00:04:38]: Open sharing.Swyx [00:04:39]: Yes.Matei Zaharia [00:04:39]: Yeah. So it's like you have a company, you maintain some table, like let's say like a Walmart or something. They have like the, inventory and what's been sold in each store. And then you also have suppliers, and they would love to produce more things and ship them, like, exactly the moment you need them. So they would love, like, real-time access to your table. So instead of like sending emails around or Excel sheets or phone calls, why can't you share like a view of that table in real time with them? Then they query, they, join it with their data, and they decide what to send. So it's one of these things where you, like you might ask like today since we can vibe code anything so fast, why do we even need to design like protocols or APIs or software? Why can't you just vibe code things on demand? But for this type of interoperability where multiple parties that are moving at different speeds are building stuff and you still want some layer on top to coordinate, you do wanna design it and build it. So it reminds me of that, like agents talking to each other and, users talking to agents and tools.Agent Clouds, Cloud Sandboxes, and Keeping Sessions AliveSwyx [00:05:42]: Reynold, any other comments alternative viewpoints?Reynold Xin [00:05:46]: I think, by the way, we had a debate on exactly which set of benefits would, matter a lot, and I think around the time we decided to do this thing I was telling Matei, “Hey,” it just happened to be there's a particular week that I was coding nonstopSwyx [00:06:00]: from the moment I woke up to, like, the moment I went to bed, I was, like, looking at my Claude sessions, my Codex sessions. And one of the things that was particularly annoying was having to keep my laptop open.Swyx [00:06:12]: I was driving to a doctor's appointment, and I remember because I wanted to make sure the whole thing continues working.Matei Zaharia [00:06:18]: But by the way, it's so comforting to hear you say that because I'm like, “I don't know if I'm a clown and I'm doing this or like.”Swyx [00:06:25]: Yeah. Like honestly, I was driving and I was tethering my laptop to my phone.Matei Zaharia [00:06:29]: huh.Swyx [00:06:29]: Keeping it on the side. Whenever I hit a red light, I started looking at what's going on my laptop.Matei Zaharia [00:06:35]: Yeah.Swyx [00:06:35]: And I just felt that was ridiculous.Matei Zaharia [00:06:37]: Yeah.Swyx [00:06:37]: It felt like we went back to the dark agesMatei Zaharia [00:06:39]: YeahSwyx [00:06:40]: programming. the productivity you gain from all this coding age is amazing, but, yeah.Matei Zaharia [00:06:45]: Have you heard of cloud?Swyx [00:06:47]: Yeah.Swyx [00:06:48]: It was crazy to me.Matei Zaharia [00:06:49]: Oh, the thing you were working on was the sandboxes or was this before that?Swyx [00:06:52]: It was a sandbox.Matei Zaharia [00:06:53]: Okay.Swyx [00:06:54]: I was workMatei Zaharia [00:06:54]: So you were inSwyx [00:06:55]: So I was approaching from a very different angle. I wanted to, “Hey, we're gonna have cloud sandboxes that doesn't shut down. You can get one very quickly,” but not just for running agentic sessions.Matei Zaharia [00:07:06]: Yeah.Swyx [00:07:06]: It's also for running development. So I was personally building that week, and through building that, I ran into all these issues, and then I wroteMatei Zaharia [00:07:15]: YeahSwyx [00:07:15]: a document for Matei, it's like, “Here's my wish list of what the actual environment should do.” And I think he ended up almost implementingMatei Zaharia [00:07:22]: YeahSwyx [00:07:22]: every single one of them.Matei Zaharia [00:07:23]: Yeah, I remember Reynolds saying, ‘cause my first prototype of this had just chats with your agent and he said, “I have to be able to open a shell, like my own shell and like list files and like tail them and stuff.” SoSwyx [00:07:36]: So SSH into a mainframe.Matei Zaharia [00:07:37]: Yeah. it has that now.Swyx [00:07:39]: Tailing my log.Matei Zaharia [00:07:40]: Yeah.Matei Zaharia [00:07:41]: Yeah.Swyx [00:07:41]: And also another thing I think I asked was, I had. I still use cursor for the sole purpose of rendering markdown files.Matei Zaharia [00:07:48]: huh. Yes.Swyx [00:07:49]: So I said, “If you just give me a way to see my markdown files and renderMatei Zaharia [00:07:53]: YeahSwyx [00:07:53]: them properly, I don't need a separate tool anymore.”Matei Zaharia [00:07:55]: Yeah.Swyx [00:07:56]: And I think you also built that in.Matei Zaharia [00:07:57]: Yeah, we, yeah, we did that, yeah. Yeah, we had a lot of engineers building, their own vibe coding setup. But then the other thing they all said is like, “Hey, I built something that's amazing for me, but, like, no one else on the team can use it ‘cause I don't have a server to collaborate.” And this is why we tried to set up, Omnigent, so you can have a server and have the security, set up in there. So, like log in with Google or whatever and, like securely share stuff. which. And that's where we've seen a lot of other agents like hit things. Like people think they prototyped an awesome agent, but it's not allowed to connect to like some really important data or whatever because of the security team.Omnigent Architecture, Open Source, and Common APIsSwyx [00:08:38]: Yeah.Matei Zaharia [00:08:38]: So yeah.Swyx [00:08:39]: Yeah. At this point, so for those watching along on YouTube, we're gonna putting up a image of the structure here, and we can talk a little bit of the architecture. I think I just want to have people understand, ‘cause like when we're talking about software, it can be very abstract and like here is what we're talking about. You've worked out in open source this entire platform and there's a runner component and server component with a uniform API that you've, you've figured out. any other element and obviously you can plug in all this, persistence layers and compute layers. This is a whole cloud. It's an agent cloud.Matei Zaharia [00:09:12]: Yeah. It's, it's got these components to work with it. The, a lot of the action happens like on the machine where you deploy your agent too. So whatever you've got on there, you can run. But yeah, it's, I think it's the minimal thing you want to have hosted, like collaborative agents and to have that server. And one of the reasons we open sourced it is, anyone building agents, this gives them an app they can start with and customize, which we were seeing in Databricks too. Like someone would make a nice, agent app and then other teams would ask, “Oh, can I just use yours for my agent?”Swyx [00:09:45]: Yeah, I think we had like five or six different agentic frameworksMatei Zaharia [00:09:48]: YeahSwyx [00:09:48]: built by every different team. They do all do more or less the same thing. Yeah, you need to. people wanna take something that works in Forkit, and you might as well have something open source. Yeah, which also was another question, which is interesting for Databricks. Like what do you choose to open source? What do you choose to make it proprietary? It's in. this goes back to Spark, right?Matei Zaharia [00:10:05]: Yeah.Matei Zaharia [00:10:06]: One, so one of the reasons to open source something is if you think it's a layer that will there'll be some network effect, it'll benefit from many, people collaborating, on it. So, for example, with Spark, I don't know if when Spark came out, we also focused a lot on letting you have libraries on top. So like there used to be differentSwyx [00:10:28]: EcosystemMatei Zaharia [00:10:28]: distributed computing engines for like machine learning and graph computation. We said they should all be libraries that you can compose. And we made it super easy to add connectors to data sources too. And then we benefit because, we don't have the time to write like connectors to like, 1,000 like different databases and file formats, but we can just use the ones people make, and of course they benefit from joining, this thing. So that's like one of these as it. Another way to think about it is like imagine, we our thing wasn't open. We had some agent hosting thing, but it's not open and then there is an open one. if you're. Which one's gonna win in the long run? So like here, because there is this benefit from like people writing integrations, it'll be, it'll be that. And then there are other things that like you just can't, even deliver as open source that are things the company does. Like for example, how do you make sure you're like streaming, jobs or your Lakebase database doesn't like, lose all your data at night? Well, that requires an operational team that's gonna sit there. There's no way it has to be a service. So like we wanna make sure as a company we're really good at those infra services and then we're as open as we can in terms of like what you build on top.Swyx [00:11:42]: speaking from a benefits, I think we are already seeing pull requestsMatei Zaharia [00:11:45]: YeahSwyx [00:11:45]: of all kinds of ecosystem integration, even though it was only released on Saturday.Matei Zaharia [00:11:50]: Yeah, Saturday. Yeah. So someoneSwyx [00:11:51]: Let's see, let's see what's going on. Yeah, you can look at the merge ones. I asked Sam Nigon this morning aboutMatei Zaharia [00:11:59]: 400 merge already?Matei Zaharia [00:12:00]: Yeah. I think Recent quite, I would guess around half are not from our team. but for example, someone added support for running it on Kubernetesrnetes. people added, many cloud sandboxes, so this can launch a cloud sandbox and run your agent in there, which is great for sharing too, ‘cause it's not, like, on your laptop and someone's, like, running scary code on there. so yeah, many startups have put those in, and, we expect to see more of them. We also have more agent harnesses already. Cursor, CLI, and Antigravity also.The Modern Data Stack and the Emerging AI StackMatei Zaharia [00:12:34]: Yeah. That's all, beautiful. And I, I feel like the last time this happens, there was the rise of the modern data stack.Matei Zaharia [00:12:42]: I don't know if it's that useful. I'm, I'm curious in your postmortem.Matei Zaharia [00:12:46]: I think most peopleSwyx [00:12:47]: AgreeMatei Zaharia [00:12:47]: will agree that it is finally dead. but maybe this arises to a new modern AI stack that, like, does the same thing.Matei Zaharia [00:12:52]: I don't know.Reynold Xin [00:12:54]: I think the modern data stack was a pretty useful thing, probably even up until this day. I think what, maybe for the audience who don't understand the history, I think the modern data stack is effectively decomposed into you need a layer to ingest the data in, you need a layer to transform your data, and then all of this are run, and then you need a layer to maybe visualize your data. And all of this runs on some data warehouse, or later on, as we're doing data warehouse or lakehouse.Reynold Xin [00:13:21]: I think that concepts are all very powerful and very useful. They enable a lot of workloads. What people eventually run into is a question of unification and consolidation is, hey, do you really need to chop all this into different pieces and work with so many different vendors and platforms in order to get, like, a very simple visualization done, right? So I think, like, over time, everybody started realizing that customers are pushing us. We started, we can realize that, so we started building more and more capabilities and trying to consolidate. And at the end of the day now, customers don't have to worry about having me hook up five different systems in orderMatei Zaharia [00:13:55]: YeahReynold Xin [00:13:55]: produce a chart. But the. I think, honestly, something like this is probably happening, in how many different frameworks do you want to hook up together in order to produce, like do a very simple agent.Matei Zaharia [00:14:06]: Just to be clear, I would say the core of this is this common API on top of all the harnesses. So the API is like, you've got an agent session, and you can send in a message or, like, a file. That's what you can send in, and then you get out, these streams as it's streaming text or as it's doing tool calls. And, or the other thing you can send in is you can, like, tell it to cancel a turn. So that's the API. Now, the thing we did is we could get you that on top of, like, cloud code running in a terminal, Codex, Py, OpenAI SDK, all that stuff. We map them all to that same interface. So that is something that you'd have to maintain yourself if you built your own, like, agent orchestrator, and then whenever cloud changes its API, you gotta, tweak your thing or it's gonna lose some messages. So that's the thing that's valuable to maintain. Then on top of that, like, we built a few apps. I think we built a pretty cool UI and stuff, but that's, And we built a security and control piece, which I'm excited about. But it's that common interface, so we don't. We. That doesn't try to be a stack. And in fact, you could plug in your own UI on top of this, server. That, and that's one of the use cases we care a lot about, ‘cause we want to use this in our own products.Compute, Sandboxes, and Databricks ScaleSwyx [00:15:20]: Yeah. It should be everywhere.Matei Zaharia [00:15:22]: Yeah.Swyx [00:15:22]: I think one of those things that is really interesting to me is, like, well, first of all, I'll, I'll endeavor to do everything and not call it the modern AI stack because like it needs a different name.Matei Zaharia [00:15:32]: Yeah.Swyx [00:15:32]: But like, yes, like, so one of the first people that told me about compute, sandboxing was Nikita from Neon.Swyx [00:15:39]: Because a lot of people think about Neon as like, well, it's serverless Postgres with, like, the separation of compute and storage and, instant branching and all those things. But every database company is also a compute company.Matei Zaharia [00:15:51]: Yeah. Yeah.Swyx [00:15:52]: And so he was showing to me his whole, his sandboxing solution. I don't think he have ever launched it.Matei Zaharia [00:15:57]: So our sandbox solution, the reason we could build it so quickly was because we realized if you just take the actual Lakebase architectureSwyx [00:16:05]: YeahMatei Zaharia [00:16:05]: and remove the database from it, by the coming from NeonSwyx [00:16:08]: Exactly, rightMatei Zaharia [00:16:09]: you have this sandboxSwyx [00:16:09]: Every database company has it already, yeah.Matei Zaharia [00:16:11]: Now, there are some differences. For example, in the one to support this particular workflow, it's important to have local persistence,Swyx [00:16:19]: YeahMatei Zaharia [00:16:19]: because you want your state to persist. Your libraries, you don't have to install your library every time, right?Matei Zaharia [00:16:24]: whereas the Neon architecture, because of the separation of storage from compute, you don't need persistent local disk.Swyx [00:16:30]: Yeah.Matei Zaharia [00:16:30]: So there's some differences.Swyx [00:16:32]: Yeah.Matei Zaharia [00:16:32]: But the, at the end of the day, yeah, it's, Yeah, so this is when you run, like, a coding sandbox. Like, if I use it, yeah, we have the dev env internally at Databricks. There's, like, many, like, tens of gigabytes of data just for, like, all the source code and, like, artifacts and stuff that I built, and I want that to come back next time, so.Matei Zaharia [00:16:51]: Yeah.Matei Zaharia [00:16:51]: But yeah.Matei Zaharia [00:16:52]: Before the show, we was talking about some statistics that might be surprising at the adoption.Matei Zaharia [00:16:56]: It could be internal, it could be external, whatever comes to mind, just to impress people the scale this is happening.Swyx [00:17:02]: So we, on the analytics side, I think we launchedReynold Xin [00:17:06]: Maybe 50 or 60 million virtual machines a day across all three clouds, so we're one of the biggest compute orchestrators out there.Reynold Xin [00:17:13]: Stuff for sure for CPU compute.Swyx [00:17:14]: Yeah.Matei Zaharia [00:17:14]: Yeah.Reynold Xin [00:17:15]: the. And all of this process, I think exabytes of data, I joked about depending on which time zone you are, typically before you have breakfast, Databricks would have processed exabytes of data already on that day. and on Neon, it's pretty interesting, too. It's launching, I think, 13 million databasesSwyx [00:17:34]: YeahReynold Xin [00:17:34]: a day now.Swyx [00:17:35]: Yeah, to me that was, like, aReynold Xin [00:17:36]: And that's just likeSwyx [00:17:37]: Like, what do you mean?Matei Zaharia [00:17:38]: Yeah. And that's the point.Reynold Xin [00:17:40]: And a lot of those were thanks to agent- agents and branching experimentationSwyx [00:17:44]: YeahReynold Xin [00:17:44]: because we made it so easy and so quickly, and thanks a lot to Nikita's team, to launch databases. It's, the. So it's changing the way people use databases.Swyx [00:17:54]: Yeah. Okay, we're gonna go into more database talk in a bit, but I wanna make sure we close up anything on Omnigentt. you mentioned, you were excited about the securityOmnigent Security, Contextual Policies, and Spend ControlsSwyx [00:18:03]: control side.Matei Zaharia [00:18:04]: Yeah.Swyx [00:18:04]: a lot of companies are figuring that out right now, as well as the spend side.Matei Zaharia [00:18:08]: Yep.Swyx [00:18:09]: what have you found there?Matei Zaharia [00:18:11]: Yeah, so I spent quite a bit of time talking to internal users, developers, security team, managers, and also lots of customers, and there's a few things. Like, first of all, one thing, that immediately was. became obvious is for security, there's this tension between, like, usability and security. And, the way people do. Like, a lot of coding agents today have very basic things like you can tell me which tool patterns I'll allow or disallow or whatever. It's like yes or no. But that puts you in a very tough spot. So just as an example, like, should my agent be able to read, some confidential documents, or let's say, should it be able to install new packages from npm, which, maybe it's compromised. Yes or no? Like, maybe I wanna allow it. Should my agent be able to publish stuff to the company website? Well, if I'm using it to code on the website, yes. But should it be able to do both, so it can, like grab a confidential document and be prompt injected and leak it? Probably not. So the thing we decided we need is stateful or what we call contextual policies where you keep track of the state of that session. It's not like is it allowed to push to the marketing site or not, but, like, hey, if it did a risky thing, like it installed, a old package from npm, or it read, like, 1,000 confidential docs, then no. Then don't, don't do it. Otherwise, maybe it's okay. That's one example of, like, moving that trade-off so it's both more secure and more useful by having a more powerful engine, essentially. This requires tracking sessions. The other piece that was interesting there is, like, there are these very level events it's doing, and you want some libraries on top that parse them. Like, for example, we have a, MCP server on Google Drive internally. It's got 60 API calls. like, how do I know which of those, like, will share a document with stuff on the internet and which ones won't? It's, it's annoying. So we designed in Omnigentt the policy layer so that it's functions and you can have libraries. Like, someone can make something that maps the level events to high-level ones, and then you write a policy about the high-level things that came out. so and thatSwyx [00:20:25]: This is related to the Panther,Matei Zaharia [00:20:27]: Yeah, Panther is. will help with that. PantherSwyx [00:20:30]: YeahMatei Zaharia [00:20:30]: a similar idea on the event processing side, and it's Python-based versus a weird custom language. this is more, as in realSwyx [00:20:39]: I didn't even know we were good yeah.Matei Zaharia [00:20:41]: Those things are happening, yeah.Swyx [00:20:42]: Yeah.Matei Zaharia [00:20:42]: So yeah, but these are the cool things. I think the contextual or stateful part, and then the way it can be libraries, and that was another reason to make it open source because others will write libraries and, like, we and our customers can use them. And the final thing, because it's stateful, one of the states we track is how much you spent in that session. So I can. I've had, like, I ask an agent to debug something, and it spent $500 because it decided to read a lot of log files and burn a lot of tokens. but I can literally say, “Okay, launch a agent to do this and cap it to spending $5.” Like, ask me for permission if it needs more. And because we're counting that within that session, it'll pop up and tell me, “Okay, you spent five, $5. Do you wanna go on?”Reynold Xin [00:21:27]: So important context here. Matei spent the last five years, a lot of his time was architecting Unity Catalog at DatabricksMatei Zaharia [00:21:34]: YeahReynold Xin [00:21:34]: which is the governance layer for data.Matei Zaharia [00:21:35]: That's right, yeah.Reynold Xin [00:21:36]: And he's combining expertise at that layer together with all the AI governance he knows.Matei Zaharia [00:21:41]: Yeah.Swyx [00:21:41]: DoMatei Zaharia [00:21:41]: But I also spent a lot of time being annoyed by coding agents and getting prompts.Matei Zaharia [00:21:46]: And also as theReynold Xin [00:21:48]: All the aboveMatei Zaharia [00:21:48]: I don't want to end up on the front page as, like, I installed some weird npm package and leakedSwyx [00:21:53]: YeahMatei Zaharia [00:21:53]: all the code, so I'm especially paranoid. But also I have very little time, so I don't want to sit there approving, like, do you want to run a 20-line, bash script, yes or no? so that's why I spend a lot of time figuring out, like, how can I make it as safe as possible and not annoying?Swyx [00:22:10]: Yeah. Is safety and mmm, let's call it security a bigger concern than token maxing or token budgets? which one is, likeMatei Zaharia [00:22:19]: Oh, yeah, they're both there. I don't know. I guess it depends on the type of company you are. So I think, some companies, like, the budget is, limited and, they really care about thatSwyx [00:22:34]: you can be Uber and still be concerned?Matei Zaharia [00:22:36]: Yeah. Oh, yeah, totally. Yeah. If you haveReynold Xin [00:22:38]: for us, securityMatei Zaharia [00:22:39]: YeahReynold Xin [00:22:40]: super paramount.Matei Zaharia [00:22:40]: For us, security is absolutely critical as a, cloud provider. It's, it's the most important thing, and, token maxing, we're not so worried about it yet, but I've seen the Like, for example, I talked to some consulting companies. They have, like, 100,000 employees who are all coding for customers. If those each spend, like, an extra $1,000 a month, that's, that's not fun.Swyx [00:23:04]: YeahMatei Zaharia [00:23:04]: we have, like, only a few thousand engineers.Swyx [00:23:06]: What's the policy in Databricks? Is it just unlimited or what'Matei Zaharia [00:23:08]: It's, it's unlimited, but we do. we use our own product to, like, analyze the traces and stuff, and we have a team that'looking to optimize and to see if anyone's doing something weird. And, we had some really cool insights just from analyzing current traces, like whichSwyx [00:23:24]: YeahMatei Zaharia [00:23:25]: models are better at, say, Rust versus like TypeScript or whatever. So yeah, at least in our code base.Swyx [00:23:31]: Yeah. Amazing. Obviously, I have to ask the token question, obviously.Matei Zaharia [00:23:34]: Yeah.Swyx [00:23:34]: I think it'sReynold Xin [00:23:34]: YeahSwyx [00:23:34]: it's a key thing. But yes, security and control above that, and figuring out a sane layer there you can have some autonomy, but, not too much.Matei Zaharia [00:23:43]: Yeah. Yeah, and we wanna make it super easy. As a engineer, you should set a thing. So in Omnigentt, you can ask your agent, “Set a policy on yourself to do this.” So it can likeSwyx [00:23:52]: But if there's something I should be showingMatei Zaharia [00:23:53]: YeahSwyx [00:23:53]: I don't, I don't see it on the GitHub, but,Matei Zaharia [00:23:55]: Oh, yeahSwyx [00:23:56]: there's justMatei Zaharia [00:23:56]: Well, in the docs there's something.Swyx [00:23:57]: Yeah, this is it.Matei Zaharia [00:23:58]: You can look at it later.Swyx [00:23:59]: Okay. Yeah.Matei Zaharia [00:23:59]: Just look in the docsSwyx [00:24:00]: YeahMatei Zaharia [00:24:00]: contextual policies if you wanna see.Swyx [00:24:04]: I just like to point peopleMatei Zaharia [00:24:05]: look at the built-in policies.Swyx [00:24:06]: Yeah.Reynold Xin [00:24:06]: Yeah.Swyx [00:24:06]: If you want to, follow up on this is exactly where to look, right?Reynold Xin [00:24:10]: Yeah.Matei Zaharia [00:24:10]: Yeah. yeah, and the story of these is, like, I just wrote, like, I wrote a doc with like 10 ideas for things before as you were working on them. Well, that was, like, my wish list of things people asked, and I told the team, like, “Hey, can you do like at least five of these for the launch?” And then they just got back with all of them, so.Swyx [00:24:29]: Oh, wow.Matei Zaharia [00:24:29]: so you can come up with more, but them- some of them are just meant to be examples. really you can intercept, like, any event the agent is making, and you can then either block or force it to ask the user or, like, allow, and you can update state to keepSwyx [00:24:45]: YeahMatei Zaharia [00:24:45]: track stuff.Swyx [00:24:46]: Yeah, ‘cause ultimately you're, I think of you as, like, a systems designer.Swyx [00:24:50]: You let people plug in, right? That's the wholeMatei Zaharia [00:24:51]: YeahSwyx [00:24:52]: modus operandi of what you do.Matei Zaharia [00:24:53]: Yeah.Swyx [00:24:54]: It's likeMatei Zaharia [00:24:54]: And we care a lot about also composab- like, can someone else write a library that others use, whichSwyx [00:24:59]: YeahMatei Zaharia [00:24:59]: this is meant to.Reynold Xin [00:25:00]: There's also a batteries included philosophy hereMatei Zaharia [00:25:03]: YesReynold Xin [00:25:03]: probably very similar to how you did Spark, which is you could just start using.Swyx [00:25:06]: Yeah.Matei Zaharia [00:25:06]: Yeah, that's right. It has to be good out of the box at certain things, and then you can build your own things on top that, like, we don't wanna do. But in Spark, if you just wanna like, I don't know, like read a table or do, like, a aggregation, it should be awesome at that out of the box.Building on Omnigent: Contributions, Startups, and AnalyticsSwyx [00:25:23]: Yeah. People wanna catch up on Omnigentt, they should watch your keynote.Swyx [00:25:26]: they should go through the GitHub and the docs. If they wanted to contribute, or they want to build on this ecosystem what would you call out as the most high-leverage places get involved?Matei Zaharia [00:25:36]: Yeah, do get involved in the Discord and in GitHub. Our team is there, is monitoring, and, some of the things people ask for we just built ourselves. Some of them, we're, we're collaborating with them to build it. and also tell us, likeSwyx [00:25:49]: Yeah, they're gonna be veryMatei Zaharia [00:25:49]: how you would like to use it because I think especially for developers, like, everyone wants it to work their own way, and a really good developer tool, like you have to hear the feedback on all the ways and figure out the abstractions and how to let people customize. So we'd love to hear, like, if you think, “Hey, I, I don't want it to work this way,” tell us. We really just wanna get that compatibility layer across agents and then let you do stuff on top.Swyx [00:26:14]: Yeah. is there any, in terms of like the startup side, I'm, I'm a founder.Swyx [00:26:18]: I wantMatei Zaharia [00:26:18]: YeahSwyx [00:26:18]: I see an opportunity, I wanna get in front of you. What's your request for, like, a startup that, like, I wish someoneMatei Zaharia [00:26:23]: Oh, like you wanna integrate with us?Swyx [00:26:24]: someone was working on this.Matei Zaharia [00:26:26]: Oh, for a startup?Swyx [00:26:27]: Yeah.Swyx [00:26:28]: Like, your, you got your own startup. It's doing well.Matei Zaharia [00:26:30]: Yeah.Swyx [00:26:30]: But like, if you weren't working on your own startup, what is, like, obvious that you should You advise many startups too, obviously.Matei Zaharia [00:26:37]: I do think, just as a company with a lot of engineers, like anything that helps me make sense of how people are usingSwyx [00:26:46]: SpendMatei Zaharia [00:26:46]: coding agents and,Swyx [00:26:48]: Yeah. AnalyticsMatei Zaharia [00:26:48]: spend, but also quality or like you should write, you should add this skill, or you should write this thing, or your agents are really horrible at tasks involving this service, so I go spend time. That would be nice. yeah.Swyx [00:27:00]: Yeah. The closest I've found is, this team, GitAI.Matei Zaharia [00:27:03]: Oh, cool. Yeah.Swyx [00:27:04]: They started with, like, we will just do, code and human attribution, but they're building the analytics layer on top of that.Matei Zaharia [00:27:12]: Yeah.Swyx [00:27:12]: I do think, like, there are a bunch of, like, artificial analysis is obviously,Matei Zaharia [00:27:18]: Yeah, they have their benchmarksSwyx [00:27:18]: doing super wellMatei Zaharia [00:27:19]: YeahSwyx [00:27:19]: with their stuff. so there's, there will be people. I think this is like the domain of consultants first, but then peopleMatei Zaharia [00:27:26]: YeahSwyx [00:27:26]: will build software that, let's say, it's kinda like the management planeMatei Zaharia [00:27:29]: YeahSwyx [00:27:30]: for coding agents.Matei Zaharia [00:27:30]: Yeah, I think there'll be a lot of insights there. You have it in other areas.Swyx [00:27:34]: Okay. Well, and then the other, big thing is your dream engine.LTAP: Lake Transactional/Analytical ProcessingSwyx [00:27:39]: maybe you wanna tell the story of, LTAP.Reynold Xin [00:27:45]: So, and background with. I'm, I'm gonna make people listen to our Ankur Goyal episode where we talked about SingleStore, HTAPMatei Zaharia [00:27:52]: YeahReynold Xin [00:27:52]: and all that history.Matei Zaharia [00:27:52]: Yeah. The LTAP idea is pretty simple. so if people have heard of the, Ankur's, talk about HTAP, it's effectively the world of databases. Sorry, there's like maybe a lot of context needs to be injected here. The world of databasesSwyx [00:28:06]: I am happy to be the database podcast that I'm forcing people to, like, learn your databases, guys.Swyx [00:28:11]: You cannot vibe code with just markdown files.Reynold Xin [00:28:13]: Yeah.Swyx [00:28:13]: Like,Reynold Xin [00:28:14]: It's one of the most important fundamental systems technologies out there. But the world of database effectively split into roughly two halves. There's what we call OLTP databases, which are transactional, and think of your Postgres, your MySQL, your Oracle databases, and the other side is what we call analytics, and sometime might refer to term OLAP. And the difference is on OLTP, you typically have maybe run some transaction on some event that looks up at one specific row. We update that row, right? It's a very oriented data structure. And on analytics, you're trying to reason on the data. You're trying to compute, “Hey, what's my revenue per store? What's my. How's my website doing every day?” And then you, eventually want to probably end up running anal- machine learning on it to predict, “Hey, how will my maybe sales be going in the future?” they are so very different architecture, and everybody start with OLTP databases. Every app, when you become serious enough, that needs more than markdown files, you need to have a database. You want to lose your data, you want to have some transactional consistency. But once you want to reason on the data, if you only have like- A hundred rows, it's probably okay to run it on your Postgres or your own, your MySQL database. But once you have more data and want to run more complicated analysis, the very analysis might crush your Postgres database. So you start doing, getting data out of the OLTP databaseSwyx [00:29:35]: Replication.Reynold Xin [00:29:36]: Replicate them into the analytic systems and just startSwyx [00:29:39]: Yeah, which for people, Elasticsearch is, like, aReynold Xin [00:29:42]: Yeah. So some of them get into Elasticsearch for, like, blocked analysis. A lot of our customers obviously get into Databricks to run more sophisticated things.Swyx [00:29:51]: Yeah.Reynold Xin [00:29:51]: And there's this term called CDC, whichMatei Zaharia [00:29:54]: Change data captureReynold Xin [00:29:55]: change data capture. and what it does, it reads the binlog of the database, and if you don't understand what binlog is, it's fine. The, but it's a little delta of the data, and it reconstructs based on the delta, the state of the database, on the analytics side. But CDC is, like, a very painful thing. It's how standard in the industry, everybody uses it, but, it ends up being. I think many data engineers ends up being waken up at, like, 3:00 a.m, because there's some pipeline thing.Swyx [00:30:22]: my explanation is, like, Airbyte is like a, became a $5 billion company just doing CDC.Reynold Xin [00:30:27]: Yeah, exactly.Reynold Xin [00:30:28]: CDC is, like, a veryMatei Zaharia [00:30:30]: It's hard.Reynold Xin [00:30:30]: It's one of the most boring but one of the most fundamental operations, like, powering modern society.Matei Zaharia [00:30:37]: huh.Reynold Xin [00:30:37]: But it's so brittle that, we joke that it's, should be called continuous data corruption, because you might change your schema on your OLTP database, and then the CDC pipeline fails to handleSwyx [00:30:48]: YeahReynold Xin [00:30:48]: the schema change.Swyx [00:30:49]: Yeah.Reynold Xin [00:30:49]: And then everything goes out.Swyx [00:30:51]: And there's all sorts of tricks that you can do, like, you add in, like, some versioning or whatever, but yeah.Reynold Xin [00:30:55]: Yeah, but it's a very, in general, very complicated. Like, I think at my keynote, I asked the audience put up their hand if they love their CDC pipeline. Only, like, maybe two people put it up. So if single store, like, about maybe a decade ago, I think the industry had this idea, hey, what if I built a single database that can handle both workloads? Now I don't.Swyx [00:31:12]: Which, like, by the way, every database person ever has ever always dreamed about this.Reynold Xin [00:31:15]: Yes. Yes.Reynold Xin [00:31:16]: This is the holy grail of database engineering is why not build a single system that can do both of this? But it ends up just being a lot of compromises. one, I think one of the first issue is that, hey, each. they say Postgres has a massive ecosystem, right? You want to be using the tools that's built for Postgres. And Spark, for example, had a massive ecosystem. There's a lot of libraries you want to use. If you were to create now a new thing, you don't have a ecosystem. You tend to create a new, smaller proprietary API, and you're lacking both, and it's also very difficult to make it performance-wise to be, comparable on either side. So it ends up being sucking on both. And our whole idea of LTAP, it's obviously a wordplay on the term HTAP, is that we think this is HTAP done right. HTAP wants to build a single engine for both. We think you can get 99% of what you need by unifying the storage, and just have a single storage layer. And once you have the single storage layer, if your Postgres databases are writing data in a column-oriented format, everything analytics can just go read that data directly without any delay, right? There's no pipeline in between, so all the data will immediately be available for reasoning analytics. I think I was telling some customers earlier, hey, when we talked about this is gonna be super useful for agents, I at first didn't really believe in it myself, even though we wrote that positioning.Lakebase, Agents, and Live Operational DataMatei Zaharia [00:32:39]: Yeah.Reynold Xin [00:32:40]: But then last night I was having dinner with a Australian customer, and they told me, “Oh, hey, one of the big issue we have is we have all these logs from our services, and we see SLA dips and want to investigate. But then there's no way for those agents to even understand what's going on in the actual databases themselves. All we see is just, like, product telemetry of the database and the services.” It would make those agents 10 times more powerful if understand, for example, who's placing those orders, what is happening, what exactly are they doing. So now I'm sold on our own message.Swyx [00:33:13]: Yeah.Reynold Xin [00:33:14]: I think it's really. It gets you the almost all of the benefits of the HTAP holy grail, which is, hey, make the data available immediately for reasoning analyticsSwyx [00:33:26]: Yeah, I think,Reynold Xin [00:33:27]: without compromiseSwyx [00:33:28]: in the way that humans are generally intelligent and want to have the ability and access to query anythingReynold Xin [00:33:34]: YeahSwyx [00:33:35]: while they do the work, they also need history and need context.Swyx [00:33:38]: And, like, where else does they get context? That's it's an analytical workload.Reynold Xin [00:33:41]: Exactly.Matei Zaharia [00:33:42]: Yeah. Yeah. And I remember when we had incidents with our databases and engineers said, “Well, I can't just run a giant query on it to see what's going on because that's gonna bring down the database and hoard it even more.” Like, that's the stuff that this gets rid of, because you spin up a whole separate fleet of machines that's doing the analytics. You're not overloading, like, the main databaseReynold Xin [00:34:02]: RightMatei Zaharia [00:34:02]: that's still trying to serve stuff.Reynold Xin [00:34:04]: Yeah.Matei Zaharia [00:34:04]: Yeah.Why LTAP Works Now: Parquet, Postgres, and LakebaseSwyx [00:34:05]: So this has been a dream for a while. what had to get done in order to get to today? Like,Reynold Xin [00:34:11]: Yeah.Swyx [00:34:11]: I feel like, you have announced variants of this several times, but it wasn't as clear as LTAP.Reynold Xin [00:34:18]: Yeah.Swyx [00:34:18]: I think LTAP is like Like, okay, we've got it, guys.Matei Zaharia [00:34:21]: This thing, yeah.Reynold Xin [00:34:21]: I was talking to somebody at Meta, and then he was asking me, “Hey, what's the catch? Why is it possible now?” And I think the reality is we took a lot of time to work on the Lakebase architecture. obviously a lot of it came from the Neon team, which is a separation of storage from compute. And it turned out it was just a tiny little step away going from that to this LTAP idea, which is, hey, we just. in the Neon architecture and in Lakebase architecture, we're writing data in oriented format to the open data lake, but in there we're writing in Postgres pages. Ali and I were spending a lot of time debating, hey, can we just change that to write in column-oriented format? And we're just debating, and one day, one of our engineers who's, like, super smart came in, he's like, “Hey, I just prototyped it. It works.”Swyx [00:35:07]: Wait, it's, prototype what?Reynold Xin [00:35:09]: Prototype, instead of storing the data in the data lake in the oriented formatSwyx [00:35:15]: ColumnReynold Xin [00:35:15]: like Postgres pagesSwyx [00:35:15]: YeahReynold Xin [00:35:16]: write them in Parquet.Swyx [00:35:17]: Yeah.Reynold Xin [00:35:18]: and he just made the observation that, hey, our storage fleet has a lot of extra idle CPUs And we could use those CPUs to do the transcoding from row to column, where row is good for OLTP, but column is good for analytics. so let's do that transcoding at that time. And as a matter of fact, once you transcode the data compresses better. So from those services writing to, for example, S3 or other data lake, like object stores, you can write them faster ‘cause now they are now smaller.Matei Zaharia [00:35:49]: Yeah.Reynold Xin [00:35:49]: So there's no overhead, it's no compromise in performanceMatei Zaharia [00:35:52]: Some CPU overhead.Swyx [00:35:54]: Yeah, because,Matei Zaharia [00:35:55]: YeahSwyx [00:35:55]: we had extra CPUs anyway.Matei Zaharia [00:35:56]: We had that fleet anyway, yeah.Swyx [00:35:57]: so the debate ended. it's one of the classics of, tech, issue of a lot of debate, but then somebody went ahead and just tried to prototype it and it worked.Matei Zaharia [00:36:06]: But, like, something this strategicSwyx [00:36:07]: That's rightMatei Zaharia [00:36:07]: and important to the company, I expect there to be, like, a kickoff thing, like a design doc. Nothing like that.Swyx [00:36:13]: Nothing like that.Swyx [00:36:14]: He just. We were debating in many meetingsMatei Zaharia [00:36:17]: Yeah.Swyx [00:36:17]: and then we're just debating whether it's possible or not from first principle.Matei Zaharia [00:36:20]: YeahSwyx [00:36:20]: and then, somebody just did it.Matei Zaharia [00:36:23]: Yeah, if you set yourself up so people do that'll be great. And that happened a bit with Omnigentt too. I think if I just had a doc on, like, we can make these together, everyone would, would think, “Oh, what about this? What about this?” But then you. if you try it out, it helps. And then if you have real users and they bash it and, like, it's still working, or in this case, if you have the workload, what the workload looks like, you can just test the same pattern then.Databricks' Culture of Fast PrototypingSwyx [00:36:47]: Yeah.Matei Zaharia [00:36:47]: Yeah.Swyx [00:36:47]: Tech aside, which is very cool, this is, like, the most important thing, the culture of innovation, and you don't have to ask my permission, you don't have like, do a whole form- formal process, just do it?Matei Zaharia [00:36:59]: Well, especially these days, I think withSwyx [00:37:01]: YeahMatei Zaharia [00:37:01]: AI, it's easier to buildSwyx [00:37:02]: But so, likeMatei Zaharia [00:37:03]: a prototypeSwyx [00:37:03]: I think you are very I made a lot of suite of, like, large companies and, like, I think that at scale, things slow down, and I'm sure you felt it already, but somehow you have this core of people that, like, are exempt. How? I think we hire and we work with really good people, and that's a very important part of it, and empowering them, but also spending a lot of time, maybe us in the trenches matter a lot also.Matei Zaharia [00:37:28]: Yeah, I think, I think first, people can adapt to being in the larger company, so that helps. And we wanna make sure they know that they can try stuff and settle debates and have a lot of examples of how it was done before, or launch a thing in beta or whatever. and then the other thing I do think as a company, like despite the size, we don't launch that many, like, products. We try to keep it pretty coherent. That's, that was the whole, like, theory of the company, was like instead of having, like, 20 Amazon services you need to set up, like a analytics and machine learning stack, you just have one, and it's, like, the same API, the same semantics across all of them, the same copy of the data. So that requires, like, unification. And then we added one more thing at a time. Like, we added storage with Delta Lake. We didn't used to do any storage. Then we added SQL, we added, machine learning platform stuff. So, but yeah, don't, don't do too many, but do those things well and, that also helps, it helps keep it manageable.Reynold Xin [00:38:33]: Yeah. The other thing we encourage a lot is instead of building, boil the ocean for everything, let's figure out how do we do it incrementally, how do we do it very quickly. Like, many of our productsMatei Zaharia [00:38:43]: YeahReynold Xin [00:38:43]: they're built in the span of weeks, and then we go to, hey. Like, usually my first question to whoever team is building is who's the target customer? Who are you working with? Are you on a first-name basis with them? Are you texting with them? I think having that very tight loop,Matei Zaharia [00:38:59]: Can you bring up another launch that comes to mind when, in this thing? I just want to give examples.Reynold Xin [00:39:04]: Omnigentt itself happened that way.Reynold Xin [00:39:05]: Yeah.Matei Zaharia [00:39:06]: Who's the customer? That's a good oneReynold Xin [00:39:34]: storage layer we did. we had, our largest customer at the time said like, “Okay, I need some. I want something in the cloud ‘cause, I. if the rest of our network is compromised, like this thing needs to be separate to store and query the events.” And then, talked to us, he said, “Okay, this is the rate of events per second. This is, like, the freshness I want. Can you do it?” So that was, like, way larger than any workload we had, and we had our, engineer, working on that, Michael Armbrust, and he worked just to make this work. And once it worked for them, it worked for everyone else. Yeah. This was early in the company, probably like four years in or something.Matei Zaharia [00:40:24]: 20- 2018?Swyx [00:40:26]: Yeah, ‘17, ‘18.Matei Zaharia [00:40:28]: Few companiesSwyx [00:40:28]: Do you have other examples?Matei Zaharia [00:40:30]: there'Swyx [00:40:31]: Maybe you have othersMatei Zaharia [00:40:31]: yeah, Clean Room, which is how you share data in a way without sharingSwyx [00:40:35]: YeahMatei Zaharia [00:40:35]: underlying data, but you allow specific operations. Those were done effectively initially just for two customers. I think the industry has a sense of, hey, maybe if you overfit to, like, one or two customers, it's gonna be really bad for you. But I think the, downside of overfitting is much smaller than the upside itself. And if you try to be too ambitious and boil the ocean, it's a much bigger problem.Swyx [00:40:58]: Yeah. Yeah.Matei Zaharia [00:40:58]: ‘Cause you might end up having no customer.Swyx [00:41:00]: Yeah, that's more, that's the more likely outcome.Matei Zaharia [00:41:02]: Yeah.Tech Companies vs. EnterprisesSwyx [00:41:03]: than you can pivot from there. I do think there is such a thing as a bad customer that sometimes you should fire. Yeah.Matei Zaharia [00:41:08]: They could exist sometimes if you drive. well, one of the challenge I think we probably see, and maybe many AI, so newer generation companies are seeing is, so tech companies are very different from tech companies or traditional enterprises.Swyx [00:41:22]: Yeah.Matei Zaharia [00:41:22]: And, if you optimize everything just for tech companies, you might have various challengesSwyx [00:41:27]: OhMatei Zaharia [00:41:27]: scaling them outside of tech companies.Swyx [00:41:28]: Okay, what likeMatei Zaharia [00:41:30]: YeahSwyx [00:41:30]: what like top three differences that you always think about?Reynold Xin [00:41:33]: Governance is a big oneMatei Zaharia [00:41:34]: I think, yeah, a big one is like, yeah, security, data privacy, governance, all that stuff. So usually if you're building some kinda like B2B or developer tool, like your biggest market is gonna be enterprises, but it's just very different. A company that's existed for like, it's had some form of IT for like 30 years, they have so many legacy systems or they operate in a regulated space. whereas a startup or, even like a, like sorta more recent tech company, all the. everything is new and pristine. So yeah, it's just different, and if you've never worked with enterprises or been in one, you just won't know about it.Reynold Xin [00:42:13]: Yeah.Matei Zaharia [00:42:13]: Yeah.Reynold Xin [00:42:13]: And the procurement process is probably quite different. There's far more stakeholders.Matei Zaharia [00:42:17]: Yeah, that is one. Yeah.Matei Zaharia [00:42:18]: Another piece that's interesting is I think some tech companies, people, will say, “Oh, I can build that myself,” right? I'll just build that myself.Matei Zaharia [00:42:27]: So then you go,Reynold Xin [00:42:28]: I don't think people say that about Databricks, butMatei Zaharia [00:42:31]: yeah, it dependsReynold Xin [00:42:32]: They do.Matei Zaharia [00:42:32]: They do?Matei Zaharia [00:42:32]: Yeah, the. Yeah, and it depends on the teams and things. So, but, on the other hand, like many of the enterprises say, “I don't, I never wanna be in the business of building that.” Like, I don't want my, whatever, I'm a retailer or something, I never wannaReynold Xin [00:42:45]: Yeah, sell clothes,Matei Zaharia [00:42:46]: be down because like some weird like nerd like couldn't get streaming pipelines working.Matei Zaharia [00:42:51]: That is not what I'm doing.Reynold Xin [00:42:53]: Yeah.Reynold Xin [00:42:53]: Yeah. This makes them great customers, to be honest, right?Matei Zaharia [00:42:55]: Yeah. But you have to understand that it's hard without having worked there and stuff, like you may not appreciate.Reynold Xin [00:43:01]: Look, I think they're all great. don't get me wrong, they have different challenges. But the, many of the tech companies, for sure there's a lot, far more DIY.Matei Zaharia [00:43:10]: On the flip side, you have people who are. they're very much experts in their domain, like they're building airplanes, they're, designing medicines, whatever, and they just want to bridge the technology, where like they don't wanna learn, databases or whatever. As cool as we think it is, even as interesting as the average software engineer might think it is to read a little bit, like they just never wanna know. They just say, “I have a, giant like, matrix or whatever with my, clinical data, like how do I, how do I like cluster it or whatever?” So yeah.The Dream Engine and Rewriting the Database StackReynold Xin [00:43:40]: Yeah. That's true. Okay, so and then I wanted to build out the dream engine, vision. where does this all lead? So one of the thing we, realized maybe a couple years back is that every single database engine out there, especially on the analytics side, are a decade old. pretty much everything that have reasonable traction are about a decade old. And they all started targeting some very specific narrow use cases, and then over time it's become more and more successful. They have grown in their ambition, and then they try to support more and more use cases. But the fastest way to support those use cases tend to be hacked around the abstractions that were initially created, that were not for those use cases.Matei Zaharia [00:44:23]: Yeah.Reynold Xin [00:44:23]: And then, but you can support them more or less okay. And before it, after 10 years of organic evolution that way, it becomes a gigantic pile of s**t.Reynold Xin [00:44:31]: the. And, but that includes Databricks. And very few company or very few systems, I think, have the gut to say, let's go start from scratch. Let's go back to the drawing board and design, knowing everything we know today after a decade of workloads and probably billions in revenue, let's attempt to rewrite it from scratch and make sure it will work and it can support all of these use cases. So we started doing that, but it's a very ambitious project. by the way, you can search on Wikipedia, there's this thing called second system syndrome.Matei Zaharia [00:45:08]: Yeah, I know that. Yes.Reynold Xin [00:45:09]: Or second system effect.Matei Zaharia [00:45:11]: Every developer must know what a second syndrome is.Reynold Xin [00:45:12]: It's you built your first thing and it works out great, and the second one's bound to fail because you become too ambitious.Reynold Xin [00:45:19]: And then you ask so many requirements.Matei Zaharia [00:45:20]: Or like you think everythingReynold Xin [00:45:21]: YeahMatei Zaharia [00:45:21]: and then you're likeReynold Xin [00:45:22]: You justMatei Zaharia [00:45:22]: you're, “I'm gonna design the perfect system this time.”Reynold Xin [00:45:24]: Yeah. And it turned out it's not perfect, and then it start failing and you're too ambitious, never launch, and you get killed. The, and the engineering team that started this, they were brilliant. I think we hired some of the best database engineers, on the planet into Databricks, and they were brilliant. Thank God it's not their second system. Many of them have built more than two in the past.Matei Zaharia [00:45:44]: Ah, nice.Reynold Xin [00:45:45]: But they were still worried about this, hey, building a database engine from scratch, I think the conventional wisdom is gonna take like five years to mature. This would be a very long-term project. It could fail. I think one of the engineers jokingly said, “Hey, maybe we just call it Reynolds Stream Engine.” If we name after a founder, maybe we then may get canceled or killed. But I think they built something pretty remarkable. they went back to. They changed the way the database engines were built from a paradigm point of view. Usually when y
This Day in Legal History: Title IXOn June 23, 1972, President Richard Nixon signed the Education Amendments of 1972, a sweeping federal education law that included what became one of the most consequential civil rights provisions in American history: Title IX. Title IX stated that no person in the United States, on the basis of sex, could be excluded from participation in, denied the benefits of, or subjected to discrimination under any education program or activity receiving federal financial assistance. The language was brief, but its legal effect was enormous because it tied sex-equality obligations to the federal funding received by schools, colleges, and universities. That structure gave the federal government a powerful enforcement tool: institutions that accepted federal education money also had to comply with anti-discrimination rules.Although Title IX is often remembered for transforming women's and girls' athletics, the law was never limited to sports. It also affected admissions, scholarships, hiring, classroom access, pregnancy discrimination, and later legal debates over sexual harassment and institutional responsibility. Before Title IX, many educational institutions openly limited opportunities for women, including through quotas, unequal athletic resources, and restricted access to professional programs. The statute helped turn those practices into legal liabilities rather than accepted traditions. In later decades, courts and federal agencies would shape Title IX's meaning through regulations, enforcement actions, and major cases interpreting what counts as sex discrimination in education. Its influence reached far beyond individual lawsuits because schools had to rethink policies, reporting systems, athletic budgets, and equal-access obligations.Title IX also became a model for how civil rights law can operate through spending power, using federal money as the hook for national anti-discrimination standards. Its passage showed that a single sentence in a larger statute could become a foundation for generations of legal, political, and cultural change. On June 23, 1972, the federal government did more than amend education law; it created a durable legal framework for challenging sex discrimination wherever public money supported educational opportunity.A federal judge in California dismissed the Trump administration's lawsuit challenging Los Angeles's limits on cooperation with federal immigration enforcement. The administration had argued that the city's ordinance was unconstitutional because it restricted the use of city resources to support federal immigration operations and limited the collection of citizenship-status information. U.S. District Judge Fernando Olguin rejected that argument, finding that Los Angeles was regulating the conduct of its own employees and agencies rather than trying to control the federal government. The dismissal was not necessarily the end of the case, because the judge allowed the administration to file an amended complaint. Los Angeles City Attorney Hydee Feldstein Soto praised the ruling, saying it confirmed that local governments can decide how to use their own personnel and resources. The lawsuit was filed after immigration-related protests in Los Angeles and after Trump sent troops to the city in response to unrest over deportation operations. The case is part of a broader Trump administration effort to challenge local “sanctuary” policies in Democratic-led jurisdictions. Similar administration lawsuits against Boston and Chicago have also been dismissed by federal judges. The White House did not immediately comment on the ruling. The decision leaves Los Angeles's ordinance intact for now while giving the federal government another chance to revise its legal claims.US court dismisses Trump administration lawsuit over Los Angeles immigration policy | ReutersA federal judge in Washington, D.C., blocked the Trump administration from using a revised immigration database to help states check voter rolls. The database, known as SAVE, is used by the Department of Homeland Security to verify citizenship and immigration status, but the administration had changed it to make bulk searches easier for state and local officials reviewing voter eligibility. U.S. District Judge Sparkle Sooknanan sided with voting-rights and privacy groups that argued the changes made the system less reliable and could wrongly remove eligible voters from registration lists. The challengers said the database can be outdated, especially when naturalized citizens are still incorrectly listed as noncitizens. The judge also found that the revamped system raised serious privacy concerns because it gave users access to sensitive information, including Social Security numbers. DHS criticized the ruling and framed the case as part of its effort to prevent noncitizen voting. The ruling comes as the Trump administration has tried to expand the federal government's role in election administration before the November 2026 midterm elections. Courts have already blocked several related efforts, including parts of executive orders involving proof-of-citizenship requirements and mail-ballot restrictions. The administration has also faced setbacks in lawsuits seeking full voter-roll data from states. For now, the decision limits how the federal government can use immigration records in voter-roll checks.Judge blocks Trump's use of revamped immigration database for voter checks | ReutersIn my Bloomberg column this week, I wrote about OpenAI's request that Treasury update an outdated R&D tax credit rule for computer-related research expenses. My argument is that OpenAI's position should not be dismissed as just another technology company asking for a more generous tax benefit. The problem is that the existing rule was designed for an older world of identifiable physical computers, not modern cloud computing, data centers, GPUs, and reserved compute capacity. Section 41 allows a research credit for certain amounts paid to another person for computer use in qualified research, but Treasury regulations narrow that benefit by requiring that the computer be owned and operated by someone else, located off the taxpayer's premises, and not be a computer for which the taxpayer is the “primary user.” That “primary user” test made more sense when a taxpayer could point to a discrete machine, but it becomes unstable when a company is buying access to capacity inside a provider-owned cloud or data center.I argue that reserved or exclusive use of computing capacity should not automatically be treated as ownership or abuse, because modern AI research may require dedicated capacity for security, speed, and performance reasons. The real question should be whether the taxpayer is buying a third-party service or has effectively acquired, operated, or taken control of the infrastructure. Treasury can still protect against abuse without treating ordinary commercial cloud arrangements as disguised ownership. I suggest that a practical safe harbor could presume service treatment where the provider owns, operates, maintains, and houses the equipment off the taxpayer's premises while bearing the incidents of ownership. That presumption should remain rebuttable where the taxpayer bears ownership-like risks or is simply routing its own equipment through another entity to claim the credit.The broader point is that modernizing the rule would not need to turn the R&D credit into an AI subsidy machine, but it would prevent an old regulatory framework from excluding a major category of modern research. The column closes with the idea that tax rules meant to police fake outsourcing should not end up penalizing real outsourcing just because the computing world no longer looks like it did when the rule was written.OpenAI's Call for Modernized R&D Credit Rule Makes Perfect Sense This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.minimumcomp.com/subscribe
I talk with Greg Matson, Senior Vice President and Head of Marketing and Products at Solidigm, about the storage infrastructure powering the AI boom. We get into why AI training and inference require massive amounts of data, how GPUs, SSDs, and data centers work together, and why storage can't be an afterthought for companies building enterprise AI. We also discuss the scale of today's AI data center buildout, how Solidigm is using AI internally, and what this means for the future of work, education, and the skills people will need in an AI-first world.
Last 4 days before regular tickets sell out at AI Engineer World's Fair - this is the single biggest gathering of AI Engineers, Founders, Leaders, and Researchers in the world. Attendees get >$5000 worth of sponsor credits and talk tracks are looking FANTASTIC. Join us!The AI scaling debate always focuses on the question of “how do we get more GPUs?” but the better question may be: how do we make the most of ones we already have.The fact that a frontier lab like xAI could be running at sub-10% MFU (Model FLOPs Utilization) is just a hint at what the real problem may be.For context, older frontier-scale training runs were already much higher than 10%. GPT-3 was around 21% MFU. Gopher was around 32%. Megatron-Turing NLG was around 30%. PaLM reached around 46%. And our guest Anjney says best-in-class MFU today is closer to 60–70%.It's not necessarily that xAI is uniquely incompetent (it's clear they have talented folks) but rather the priorities may be flipped in the GPU arms race.While GPU access is a bottleneck, simply increasing CapEx won't automatically translate to better models as frontier AI is increasingly a systems problem: scheduling, utilization, networking, kernels, frameworks, data pipelines, parallelism, cluster reliability, and the thousand small decisions that determine whether your theoretical FLOPs become real training progress.From building Discord's developer platform and backing frontier AI companies like Anthropic, Mistral, Black Forest Labs, and Periodic Labs to now building AMP's independent compute grid, Anjney Midha has spent years close to the real bottlenecks of AI scaling. In this episode, Anjney joins swyx at Periodic Labs to unpack why the AI race is not just about buying more GPUs, why 95% utilization would have been considered an outage at Google, and why the next era of AI infrastructure has to be more aligned, more efficient, and more responsible.We go deep on AMP's vision for a compute grid that makes FLOPs flow like megawatts, the difference between full-stack AI labs and horizontal pooling, why AI data centers need community buy-in, and how compute markets could evolve into something closer to an independent system operator. Anjney also explains why DeepMind's unpublished research points to a market failure, why end-of-life prediction remains one of the most important AI applications he has thought about for fourteen years, and why “output maxing” may become a new discipline for frontier systems.We also discuss Anthropic's culture, why “luck favors the prepared mind” in coding models, how Claude cracked coding, why too much capital too early can make AI labs fragile, what Periodic Labs is trying to do with science and superconductors, why great researchers can become great CEOs, and why Silicon Valley is both deeply missionary and deeply mercenary.We discuss:* Why 95% utilization was considered an outage at Google* Why AI infrastructure waste compounds at frontier-lab scale* Why “move fast and break things” does not work for AI data centers* How data center backlash, power grids, and community incentives shape AI scaling* AMP's vision for making FLOPs flow like megawatts* Why compute needs an independent system operator* How interruptible demand and dynamic prioritization worked inside Google* Why DeepMind research hoarding creates negative externalities* AMP's 1.2GW base-load ambition and the need for 6GW of spike capacity* Why end-of-life prediction could become one of AI's most important healthcare applications* Frontier Systems, output maxing, and full-stack alignment* Why APIs and abstraction layers become lossy as organizations scale* Superconductors, standards, and the dream of lossless systems* SF Compute, open protocols, and the future of compute marketplaces* Why non-NVIDIA chips can still benefit from NVIDIA's reference architecture* Trust boundaries and why chip startups need visibility into future model architectures* Why VCs often underestimate researchers as CEOs* Scientists as star athletes of the mind* Why great CEOs need to be confrontational up and down the stack* Why leading the frontier matters more than “winning”* How Anthropic cracked coding* Why culture is fragile, not a permanent moat* Why hardship was a feature, not a bug, for Anthropic* Why Anthropic's P0 was coding from day one* Periodic Labs, physics as the constraint, and technical reality* Silicon Valley mercenaries, missionary teams, and what happens after a breakthroughAnjney Midha* LinkedIn: https://www.linkedin.com/in/anjney* X: https://x.com/AnjneyMidhaAMP PBC* Website: https://amppublic.com/* X: https://x.com/amppublicTimestamps00:00:00 Introduction00:00:09 Why AI Compute Is Being Wasted00:03:17 Responsible Infrastructure and Data Center Backlash00:06:07 AMP Grid: Making FLOPs Flow Like Megawatts00:12:41 Foundry, Frontier Labs, and Research Hoarding00:14:42 Gigawatt-Scale Compute and End-of-Life Prediction00:24:08 Frontier Systems, Output Maxing, and Alignment00:27:38 Compute Markets, SF Compute, and Non-NVIDIA Chips00:32:57 Trust Boundaries, Co-Design, and Researcher CEOs00:38:17 AI Coachella and First-Principles Thinking00:42:43 Leading vs Winning in Frontier AI00:45:54 How Anthropic Cracked Coding00:48:25 Culture, Hardship, and Anthropic's P000:54:03 Periodic Labs, Physics, and Silicon Valley Mercenaries00:56:26 Rishi Valley, Singapore, and Money as a Measure00:58:47 Closing ThoughtsTranscriptIntroduction: Anjney Midha, AMP, and Compute WasteSwyx [00:00:00]: We're in Periodic Labs with Anjney Midha, CEO, founder of AMP. Welcome.Compute Utilization: Node Allocation, MFU, and AlignmentAnjney [00:00:09]: Thanks for having me. At Google, there are two types of utilization usually, right? That you're measuring in these clusters. One is node allocation, and then the other's MFU. Node utilization is usually like what percentage of cards in the data center are just, used, and that, if it's not at, 95%-Swyx [00:00:29]: There is no excuseAnjney [00:00:29]: There's no excuse, right? I think 95% at Google, which is where my co-founder, Seb, came from, he built the Borg, PBorg/GQM scheduler at Google, and there I think 95% was considered an outage, so 96% node utilization is, should be standard. And most single-tenant clusters are not running at that. So that's one. And then MFU should be, I would say the best in class today is somewhere between 60 and 70%. I think this is a leadership question, right? Fundamentally it's an alignment question, which is are the people who are funding the cluster and then deploying the cluster actually aligned? And sometimes theoretically they are, but in practice the number of people in the chain, the supply chain between, the capital and all the way to whoever's managing the cluster and then whoever's measuring what the output is, are just so many, degrees of separation away that, the, The Have you ever heard the radian metaphor, which is at the beginning of an arc, if you have two arcs that are two lines that are just off by a few degrees, that-Swyx [00:01:33]: It spreads outAnjney [00:01:34]: It spreads out, right? Or at scale. And I think what's happening is a lot of cluster implementations and infrastructure, a lot of frontier labs and other teams, that's what's happening, is they're, they initialize the plan, which is kind of like North Star with a team that wants to do good, but then they're, required to scale so fast instead of iteratively that the wastage just compounds really fast at scale. And so I think we know the answer, which is just do iterative bring ups. If you spend time with people who've been in the semiconductor industry or the DSN industry for a long time, this is not new, and I don't think AI should be an excuse. Sure. Something What is new? Okay. We have a lot of new capabilities, but that doesn't mean just abandon common sense. Common sense should always be in fashion. ? AI scaling doesn't change the in fact, if anything, AI scaling should be putting a premium on the value of common sense and infrastructure because the margin of error now is so much lower and the costs of wastage are so much higher. And the cost of wastage, by the way, is not just economic. I'm, obviously I'm, I'm an investor, or I'm an investor by background. Over the last few years now we're running an AI infrastructure business called, AMP. And I think that it's okay to say this time is different on the capabilities front. We are genuinely getting capabilities at, of the, of a kind we haven't had before. That doesn't give you an excuse to say this time is different for everything, especially infrastructure. So look, I love the hacker mindset and the hustler mindset. Now, that's great for the startup mindset, but you remember this moment where Zuck went from saying, “Move fast, break things” to, move-Responsible Infrastructure and Data Center BacklashSwyx [00:03:10]: Fast and stable infrastructureAnjney [00:03:11]: Move fast with stable infrastructure. I think now we need to move fast with, responsible infrastructure. People are going to ask where the impact is. There was a really In our class yesterday, Scott Nolan, who's the founder of General Matter, came by at Stanford to speak about energy bottlenecks. And he had a phenomenal idea. He said, “if you look at the marginal unit economics of compute per hour,” he goes, “let's call it, $4 an hour. If you're having to bring up a new data center in a new community, why not just say we're going to charge 4.50 an hour, and that marginal impact or that marginal increase, we just literally take that and give it to the local community as cash?” I can tell you as a customer of that compute, I would love that. I'd be happy to pay an additional 50 cents per hour at scale.Swyx [00:03:57]: Wow. Yeah.Anjney [00:03:58]: Because if that means the public benefit is so clear to the communities that the data centers are coming up in, I'm going to feel like that compute is much more reliable. Up to 20% of all data centers this year in the US, my understanding is are at risk.Swyx [00:04:13]: Of community backlash?Anjney [00:04:14]: Correct. Of not getting the community support they need to get brought up.Swyx [00:04:19]: Wow. That's a huge number.Anjney [00:04:20]: Yeah. Now, we, I think we should dig into what that number is. I think it's a little bit of overstated. These things can get over-reported, but it-Swyx [00:04:27]: They don't just care about jobs. They care about all the other stuff around it, right? They care about power grid, they care about environments-Anjney [00:04:33]: Power grid, permitting, and so on. And imagine I think if you said there's a new AI deal. If we're bringing up a data center in your community, we're actually going to reduce the cost of your electricity bill. Okay, now we're talking. Right? The community's going, “Okay. Now this is a deal. I feel like a partner in this.” Right now that's not happening. There will be audits, there will be investigations, and when the, when the regulators come, I don't know when it's going to be, the folks who are moving fast and breaking things in the name of AI progress better be prepared. That's certainly not how we're procuring compute. Or we're, we're trying as much as we can to work with partners who have long-term track records. Many of whom, by the way, are not, AI providers. I think this whole idea of neoclouds being somehow this new category is a lot of marketing speak. There are really good, reliable, trusted data center providers in America who've been around 20 plus years. I love those folks. They know how to Sure. Are they sponsoring happy hours at NeurIPS? No. Are they legibly listed in Build? No. Are they hanging out in my, in, situational awareness parties? No. But they're adults. I trust them.Swyx [00:05:44]: They can run LAN. They can run power.Anjney [00:05:45]: They can run LAN, power, and shell. They have credit histories. We sit down, we have a conversations. Many of them live in Silicon Valley. They've, they've had to deal with the boom and bust cycles of the internet, and I love those folks. They are stable infrastructure partners and thinkers. And I think there's a lot of short-term thinking going on in the compute layer, and it's going to catch up to us. It's not going to be good.AMP Grid: Making FLOPs Flow Like MegawattsSwyx [00:06:07]: You talk about aligning incentives, and, I would think that aligning incentives means you have the full stack in one company, which is xAI and OpenAI, right? So you as a standalone infrastructure layer, why are you somehow more aligned to your portfolio companies than people who just own the whole thing?Anjney [00:06:28]: In systems design, right, there's, there's two regimes of, architecture, right? You have integration, and then you have pooling and utilization, right? So the Or rather, the way to increase utilization often is you can do systems integration where you collapse a lot of process into one node, or you can pull out a process from a node and share that amongst various That resource amongst several different nodes. And so we see the AMP grid, which is, the, what, the system we're building here, which is basically a compute grid. We're trying to do for compute what the electric grid-Swyx [00:07:02]: PowerAnjney [00:07:02]: Yeah, what the power grid did for electricity. It-- this is a pooling and utilization layer across clouds, And so we're actually the opposite of a full stack integration like approach.Swyx [00:07:12]: Super horizontal.Anjney [00:07:13]: Where it's much more horizontal and it's, it's multi-cloud, it's multi-silicon. The goal is to try to make FLOPs flow like megawatts, and that is very hard to do today for many reasons. There's stranded pools of compute all over the place and there's no fungibility. And so right now we do it at the level of scheduling, and we often do it at the economic layer. But as we start to announce what we're working on, it's extraordinary like how many folks are coming out of the woodworks and saying, “Hey, I'm actually working on a way to make compute fungible at this part of the stack and that part of the stack.” And as a grid, we'd like all of these folks to participate on the grid. There's, people often ask me, “Andra, are you a new cloud?” And I go, “No, actually neoclouds are suppliers.” sometimes they'll ask, “Are you a venture capital firm?” I go, “No, actually they are, they are demand like sort of off-takers of the grid.” We see ourselves as what's called an independent system operator. So if you study the history of the electric grid, once it became legible to a lot of factories and industrial sort of participants that, hey, actually it turns out pooling is a good idea. We should pool our generators instead of all having a generator running at half capacity in our backyard. There was a need for an independent entity who could coordinate all these parties. Transmission line, power generation, facilities, transmission lines, factories, and that neutral coordination mechanism is very critical. In order-- If you study like the history of grids, the most enduring ones were those that never owned their own assets. They were ones that had, or often started with long-term anchors who are uncorrelated sources of demand, a steel factory, a shoe mill or whatever in a particular town who weren't competitive, where the steel factory want to spike up at night, the shoe mill wanted to spike up during the day. So then you pool and you share, right? So each of you is guaranteed some base load, but then you kind of schedule your spikes to drive a peak utilization across the town. The gold standard, so to speak, historically, has been these utility companies like PJM Interconnect in the northeast of America, where they, over many years became this what's called an ISO, an independent system operator of the grid. So that's how we see ourselves. Economically, that's what we are. From a technical perspective, we started at the scheduling layer because Seb and Mihai, who, run engineering here, built that at-Swyx [00:09:28]: Did your schedulingAnjney [00:09:28]: They did that at Google. And, -Swyx [00:09:32]: And you have infra shops from Discord as well.Anjney [00:09:35]: I have some.Swyx [00:09:35]: I don't know, I don't know if Discord is like the primary identity, but what-whatever, I'm just kind of-Anjney [00:09:39]: No, D-Discord was-Swyx [00:09:40]: Choosing a well-known name.Anjney [00:09:42]: Well, I So I was running the developer platform there. The internal infrastructure I was not responsible for. That was actually a guy by the name of Mark Smith, who was extraordinary. And yes, Discord did pool So Discord is actually a counter example. I had the chance to learn a lot about fully, full stack infra there because-Swyx [00:09:56]: It's the same thing, yeahAnjney [00:09:57]: It's the, it's the other architecture which is, Discord built its own WebRTC vo-voice and video infra. So like Discord did not use-Swyx [00:10:08]: For the calls, yeah.Anjney [00:10:09]: Yeah, did not For communication, Discord did not use third party infra. It was all built in-house. And then the way you maximize utilization was you pool demand from the world's 200 million plus monthly active gamers, right? And so that's, that's how those stacks were constructed. Again, in systems design, the two concepts that keep coming up over and over again are abstraction and composition, right? And-Swyx [00:10:31]: Bundling and unbundlingAnjney [00:10:33]: Bundling and unbundling, abstraction, composition, like verticalization and-Swyx [00:10:36]: HorizontalAnjney [00:10:36]: Horizontalization. So in that sense, AMP is an independent system operator of the grid. We pool demand, we pool supply from a number of partners we trust At about 1.3 gigawatt scale over four years. And then we pool demand from some of the world's best, research labs and so on. We're sitting at one, periodic labs who need extraordinary long-term demand. And the idea is that, each of them is guaranteed base load on the grid, but they can spike up and down flexibly on, for compute, with much shorter timelines as needed. That was roughly the design of the program I came up with at a16z called Oxygen. The same-- That was the same design of the GQM, BorgX, Borg GQM implementation at Google that Mihai and Seb had built. Which was that how do you allow, teams inside of Google, on the internal infrastructure to be guaranteed capacity, for their base workloads? But when they need to spike up on research, how could they ensure that was sufficiently there? And of course, the big innovation that was not discovered, but kind of implemented in the space, this infra space maybe three, four years ago at Google was the idea of interruptible demand, right? Where you just queue up a bunch of jobs and through this like sort of credit system, there can be a bidding mechanism.Swyx [00:11:53]: Like priorities.Anjney [00:11:54]: It's a dynamic prioritization Basically. And jobs can get interrupted based on somebody else who's saying, “what? I have 10 tokens, 10 credits I want to spend on this job.” Another like team lead, research lead is “Genie 3 or whatever is only worth five, credits, and NanoBanana2 is worth 10 credits,” and so the NanoBanana job gets priority. That's a, that's a made up example.Swyx [00:12:15]: It's very real. Brain Marketplace was real. And, we've, we've covered this on the pod with David Luan, who was-Anjney [00:12:20]: Oh, great. OkaySwyx [00:12:20]: Was there. And the criticism is that, well, actually sometimes you need central command to go all in on a thing. And actually sometimes capitalism via credits doesn't work. Not, this is not a criticism of AMP. I'm just saying, this is a thing that has been tried, internally within Google, and it led to Google missing GPT.Foundry, Frontier Labs, and Research HoardingAnjney [00:12:41]: Like, we structured ourself essentially very similarly to Google. We are structured as a holdings company. So, Alphabet holdings is Alphabet holdings, and then they've got these subsidiaries called Google and-Swyx [00:12:51]: Other betsAnjney [00:12:52]: Other bets and so on. We've got, AMP holdings, and we've got our infrastructure business, and then we've got a capital business called Foundry that incubates new frontier AI labs or invests in them as venture capital, like Periodic. We put a few hundred million dollars into Anthropic from our fund earlier this year. So wherever we feel like teams are making progress, especially researchers and so on who've pushed the frontier inside of existing labs like DeepMind, I find, there comes a point where they feel misaligned with the dictatorship of Alphabet holdings. And at that point, sometimes the dictatorship doesn't want them anymore. And they're “Thank you. You've done your job here. You've kind of helped us through the zero to one phase, and for whatever reason, we're going to deprioritize your amazing, omni model or whatever it is, and instead we're going to prioritize coding.” And, I think that's a tragedy, but I get it. They're Sergey and team are running their own business there. But that doesn't mean we the rest of us should sit around waiting for that progress to get unlocked for the rest of the world and humanity. If you think about how much extraordinary research has happened inside of DeepMind over the last 10 years, I, Demis and Sergey and those guys did such a great job. But at the end of the day, so much of that has never seen the light of day?Swyx [00:14:00]: Or they're like papers only, but they never actually shipped it to production or-Anjney [00:14:03]: What's worse is the paper is actually not even being published anymore ‘cause there's a six-month embargo inside of DeepMind, right? We've heard about this where a paper comes out, and then I think there's a six-month embargo window where if anybody on the business team says, “This could be interesting” It's embargoed for life.Swyx [00:14:18]: Exactly. So the stuff that gets published is the stuff that's not good enough.Anjney [00:14:21]: There's an adverse selection problem, basically. Yeah. At this point-Swyx [00:14:25]: It's, it's a common complaint at NeurIPS, by the way, that's “Well, why would I look at the papers that are the trash of GDM?”Anjney [00:14:31]: Again, I think it's a tragedy. I get it. They're running their business, but the rest of the I think there's negative externalities of research being hoarded, and so that'there's a market failure. And somebody needs to unlock that research, and we can't do it on our own. We only have 1.2 gigawatts of compute. That's nothing. That's about $40 billion of cloud spend. We're going to need a lot-Gigawatt-Scale Compute and End-of-Life PredictionSwyx [00:14:51]: By the way, is that's a new number. I haven't, haven't come across that gigawatt number. That's huge.Anjney [00:14:56]: Yeah. And to be clear, we haven't secured all of it. That's how much demand we have started to secure. I think publicly we haven't actually confirmed how much we have for this year. In order-Swyx [00:15:04]: Where do you want to get to?Anjney [00:15:06]: I think the steady state would be that we have a base load pool Of 1.2 gigawatts at all times Of base load capacity. For spike capacity, right now my estimate is we need roughly six gigawatts over the next four years for all our teams to feel like they were able to keep moving the frontier, whatever they're working on, whether it's, like superconductor discovery over here. There's a new investment we're working on right now, which is in the end of life prediction space in healthcare. It's extraordinary how much you can, you can give this was actually my graduate school work. I went to grad school for bioinformatics at Stanford Med. And I know we-Swyx [00:15:40]: Econ, MCS, bio.Anjney [00:15:41]: So my-- I was this really weird cat where, I was never satisfied with my major options. So at one point I was an econ major, then I was a CS major, then I was a MCS major called mathematical computational science, and they decided they were going to end that major. So I took all that coursework, and I applied it to grad school, my graduate degree in bioinformatics, which was the master's program, and then I thought I was going to do a PhD. I never ended up doing it. I dropped out and went to work at Kleiner. But I was lucky enough to apprentice with this professor at, Stanford Med. His name is Nigam Shah, and he was working on end of life prediction. Stanford is one of the only research facilities in America that has a longitudinal patient data set that's larger at scale. I think it's at least 12 million patient lives. The only larger data set is at the VA, the Veterans Affairs, of America. And to do research, like do any deep learning and so on that data set, it was called the STRIDE data set at that time, you had to be a Stanford Med School affiliate, which is why I went and enrolled in the bioinformatics department. End of deep learning was early. Nigam Shah had the visibility-- the vision to see that, you could do end of life prediction to help palliative care. In America, the, over 30% of all Medicare, Medicaid spend, at least at that time, was spent on end of life care. And what's we grew up in Asia, so we all-- Yeah, at least I won't speak for you, but I have A very different relationship with death than I find folks who grew up in America do. In America, spiritually and culturally, especially in Western societies where Christianity, the Christian tradition sort of frames death as this terminal point, there's often a judgment day and so on. The way we view death is with a finality. In Indian culture, in Hindu culture, death is one-Swyx [00:17:35]: Also, he's Buddhist as well.Anjney [00:17:36]: You're Buddhist, yeah. So it's one, it's one step in a journey of many lives, right? And so, I grew up in this city called Chennai in the south of India, and when people die, you dance on the street. There's like a procession where your body is carried to be cremated and your family, like celebrates and there's drums and so on. It's this huge thing. And, It's because the idea is that you're going to be reincarnated. You've been liberated from the responsibilities of this life, and now you're onto your next. It's a new It's like going off to a new college or whatever, right? And so it was so alien to me when I got here as an undergrad- That the medical system works backwards from that assumption that we have to view death as this terminal thing and delay it, postpone it's a bad thing. And so at the time, clinical decision support in the United States was this very primitive field. Even to this day, physicians in the United States often will tell you when you have a terminal disease, this is your, we've diagnosed you, which is great. Our ability to diagnose you is extraordinary. You have somewhere between six months to six years to live. What do you do with that information? The error bars are so high that then you In times of uncertainty, we default to culture, and when the culture is let's-- this is a bad thing, I've got to prolong my life, then you start doing things like And just to, just sort of from a systems perspective, what's going on there is Physicians often feel like they need to provide such high error bars because there's always some uncertainty in end of life diagnosis, and if you provide the wrong Diagnosis or recommendation to your patient, you can be sued for medical malpractice. And then your license can be taken away. It can be catastrophic for your career. In contrast, if in countries where that's not the case, what you often observe is that patients, physicians are quite prescriptive with their recommendation. They say, “Hey, this is your condition. The literature says that you probably have this much time on Earth left. My expert opinion is that you are an outlier or whatever.” And they try to be more prescriptive, and that empowers a patient, right? ‘Cause then a patient can say, “I trust my doctor. They said on average, I have six months to live, but if I do these things, I may have a shot because of my particular predispositions or my genetic history or whatever.” And that empowers you to go about your life in a actually more scientific way than leaning on religion, culture, spirituality, and so on. In contrast, here, because of that medical malpractice sort of thing looming over your head, a physician never gives you a clear recommendation. So instead you say, “Okay, Doc, well, let's try it all.” And then you start a whole regime of drugs and therapies, and then you often spend weeks and weeks in the hospital, and that deteriorates your quality of life. And when that deteriorates your quality of life, you instead of spending your last few days doing the things you love with your family, you're spending it on a hospital bed. And that ends up being thirty percent of Medicare and Medicaid. So it's worse for the patients. The doctors feel terrible. The American taxpayer is paying a huge amount of money. And so this is why Nigam Shah, who was this professor at Stanford, said, “Anjney, if there's “ I kind of sat down with him. I was this young, I'd, I was twenty-one, and I was “I want to work on a big problem.” He's “The big problem is end of life care.” And so we tried to do deep learning to say, to-- So we started trying to run deep learning on these tried patient data sets to say, “Could you have an AI system make a recommendation that is orders of magnitude more precise about how much time you have left once you've been diagnosed with a terminal condition than a human?” And then if we can get that precision to be high enough, then you can empower the patient. And it turns out the tech works. Like it's-- Once you get the data set, like RL works. Honestly, even regression models work. You don't need to get that fancy. At the time, we were just trying, doing like very simple neural nets.Swyx [00:21:54]: Simple solutions, yeah.Anjney [00:21:54]: Today, what we can do with RL is extraordinary. The problem remains then and now is regulatory, because you actually can't shift the burden of the wrong clinical diagnoses from the physician to the AI system. And so at that time, I got quite disillusioned ten years ago for, twelve years ago where, ‘cause I felt I just didn't have the resources to influence regulation. Today, I'm very lucky. I'm in a different place. I've, I'm a lot older, and so I've been spending a lot of time on my next incubation, which is how can we unlock the, patient empowerment by training AI models to do end of life prediction much, with much more precision and ac-Swyx [00:22:37]: Oh, wow. You're still focused on this the whole time.Anjney [00:22:40]: The-- I haven't been able to get, this out of my mind a single day for the last fourteen years. This is the hill I want, I would like to die on. There's two, I would say. What? I actually, I'd prefer not to die.Swyx [00:22:51]: Yeah, exactly.Anjney [00:22:52]: But I think two bipartisan issues, I think two issues that should be bipartisan in America are how do we empower patients to make the right clinical decisions at the end of their life, such that we're reducing the taxpayer burden with science? It's just good old science, and AI can help here. And the second is, net positive data centers, ‘cause I think that's the biggest critical bottleneck on training and good enough AI models to help people at the end of their life. So there's sort of two sides of the, of the same scaling bottleneck curve, but those two, we formed AMP as a public benefit corporation. My wife and I, who you've met, you've met Viv. Her passion is education. Her family is a long line of educators and so on, and, of physicists. And so this class is my attempt to stop being the black sheep of the family and be a, an educator. But if I'm not educating, the thing I would be doing is working, on these two problems, whether on the political spectrum or as a researcher back at, in some lab. And my hope is if anyone's listening to this podcast, if they're passionate about either of those two topics, I'd love to hear from them. We'll, we'll we can share the contact in the show notes, but, we're looking for people to join both of those missions on the, on the political side as well as on the medical side, on the research side.Frontier Systems, Output Maxing, and AlignmentSwyx [00:24:08]: You said, this is a discipline that you want to form. You call it's called variously called Frontier System. It's variously called One Person Frontier Lab. What is the ideal name or shape of this? Like the, what is the mission?Anjney [00:24:24]: Of the class?Swyx [00:24:26]: Of the discipline that you're, exploring, right? I The class is called Frontier Systems. But like for me, maybe one phrase is you're, you're just anti-waste, right? Which is wasting GPUs, wasting in human and Medicare. But is there, is there a broader theme that I'm, that maybe you can encapsulate more succinctly?Anjney [00:24:45]: Yeah. The, from an engineering perspective, it's very simple. It's output maxing. It's the, it's the department of output maxing.Swyx [00:24:51]: Making the most of what we have.Anjney [00:24:52]: Exactly. I'm a huge believer in optimal outcomes. I think both in America and other countries, we are losing our appreciation for nuance, and this is the thing of And AI is the same case, right? Oh, the bitter lesson holds. Okay, fine. But that doesn't mean you just like throw 500 GB300, 500,000 GB300s at your suboptimal model scaling and you waste a bunch of compute. It also doesn't mean that, the most optimal is to have like 50 different architectures where there isn't enough standardization. One of the reasons Anthropic has had extraordinary sort of velocity is ‘cause they picked the transform architecture and said, “This is simple. Let's double down on it,” right? And now luckily there's enough investment going to the space that we can afford other architectures, but at the time, investment was just too fragmented into other architectures, so that arguably unlocked scaling. So I think there's a philosophy. I think we all owe it to ourselves to do output maxing with a new capability called AI on a global level. I think if I was starting a new department at Stanford, depending on how fuzzy or technical I wanted to be, I'd probably call it the Department of Alignment. Like-Swyx [00:25:59]: It's an overloaded termAnjney [00:26:01]: But it is, But alignment really Is a hard problem. And I think when you unlock it, full stack alignment is super hard in any organization and in any system. Like in a, in a venture capital firm, if you can have full stack alignment between your limited partners and your, the founders who are creating the value and ultimately the public that owns the IPO stock, that is a gift that keeps giving. And when you study the history of these systems, when they start off, they usually start out small scale where the feedback loop is actually so tight that there's alignment. And then the more you try to scale, the more division of labor happens, the more specialization happens, and at each step you add abstractions. And wherever there's an API interface, there's like loss. There's communication loss. And so I think a really cool thing would be for us to figure out is there a way for us to have our cake and eat it too as an engineering discipline? Is there a way to actually scale up and scale out Without losing any alignment, without lossy transmission?Swyx [00:27:01]: You mean standards?Anjney [00:27:02]: So standards is one way. The other way is you just have net new capabilities. So like what we're trying to do here is discover new superconductors. A room temperature superconductor would be a lossless transmission mechanism for energy. We would have flying cars. We are right within a few years of having a new room temperature superconductor. So I think those are the two. You either have to standardize On protocols or API specs that allow lossless communication, or you can come up with a whole new capability that unlocks so much abundance, the standardization doesn't matter ‘cause you just unlock net new capacity. This, the, so this is what I spend my days thinking about these days.Compute Markets, SF Compute, and Non-NVIDIA ChipsSwyx [00:27:38]: No, I think every infra person at, who wants scale and wants to output max does eventually end up thinking about this. We don't have time to go into it, but we have done an episode with SF Compute-Anjney [00:27:50]: Oh, coolSwyx [00:27:50]: That is trying to standardize The futures contract for compute. I don't, I don't know how that's going by the way, but like at some point this will be public.Anjney [00:27:57]: Oh, I think Evan is awesome and SF Compute is the kind of effort that I hope we can accelerate because what often happens is these exchanges are very hard to get, they, it's hard to bootstrap them, right? Because they often require-- There's many inefficiencies between parties. There's trust boundary inefficiencies in infrastructure because you don't trust, one part of the stack doesn't trust another part of the stack to give them visibility. There's capital markets inefficiencies, there's operational efficiencies. So if you can inject like a single shock to the system of a ton of compute demand or supply, then you can accelerate, these new flywheels. And so my hope is one day, or soon, if SF Compute needs extra like has excess capacity, they just hook it up to the grid and they get flooded with demand from us. And on the other side, if they have a ton of demand but they don't have supply, they just again hook up to the grid and it's a two-way protocol where they can just hook up to our capacity. And I don't think we're too far from that. Today our working implementation of it is mostly through a group of labs, universities, and a few sort of trusted parties who are, who all feel like they're in alignment to borrow an over sort of used word. But our hope is to just have it be an open protocol that anyone can hook up to on-Swyx [00:29:20]: Hook up for demand or hook up for supply? In primarily demand, it sounds like. Like you-Anjney [00:29:25]: No, bothSwyx [00:29:26]: You would want to offer demand.Anjney [00:29:27]: Both. Yeah. Unfortunately, what's happened in the last six weeks is, we thought we'd have a bunch of excess capacity by the end of this year. It's all gone.Swyx [00:29:37]: It's exploding.Anjney [00:29:38]: It, yeah. It's all gone. And so I have, my text messages are full of friends, we know many of these people, these are founders who've raised billions of dollars in San Francisco going, “Oh, any chance you have like 50 nodes in the next few weeks?”Swyx [00:29:51]: What is the scope for, non-Nvidia, right? You have Lisa Su coming and, Rainer Pope as well. And so There is a lot of demand for, more performance Alternative architectures and all that. At the same time, this hurts your standardization.Anjney [00:30:11]: I don't think so. So actually Rainer's a great example, right? Rainer is a CEO and founder of, MatX. I actually had him by for office hours in the class earlier today, and there was an insight he brought up that I hadn't considered before, which is when they decided to pick the standard For their data center, they picked the NVIDIA reference architecture. So the MatX chips Just plug in to any site that has an NVIDIA bring up planned. And, the-Swyx [00:30:42]: It's just software then. It's, it's not the-Anjney [00:30:44]: A-Swyx [00:30:44]: Hardware.Anjney [00:30:46]: Well, from an input and IO perspective It's the same footprint as an NVIDIA rack.Swyx [00:30:52]: That makes sense.Anjney [00:30:53]: Where they have done, innovated a bunch from what I can tell is on systems co-design. Which is where a lot of the gains are to be had. And so he picked He was “Anjney, we, there's just so much work to do when you're building a new chip company.”Swyx [00:31:08]: Can't fight every front.Anjney [00:31:08]: You just can't fight on every front. So my question to him was, “Well, you're working on this new chip. Their tape-out is next year. What, who are you going to partner with to host the chips?” And he said, “Whoever will host them. That's just not, that's not my focus.” And I said, “But how did you “ you decided back to our earlier systems design question, he decided that, he didn't want to be a full, fully integrated chip provider. The bottleneck they're focused on is the logic die, and they, he feels they can crank out a ton of performance gains through co-design there. But then that means you delegate, to our question earlier, it, you he's the data center provider is a different part of the stack, and so then he's dependent on that part of the ecosystem to host his chips to get the performance gains to the customer. So now you have another abstraction, and you might have loss. So I asked him, “How do you prevent loss?” And back to your point, he said, “I just picked the NVIDIA standard ‘cause I didn't want to Like I wanted to piggyback off of an existing protocol.” And that, what's great about NVIDIA is that reference architecture is known.Swyx [00:32:15]: Open.Anjney [00:32:15]: It's open. They've published it. So Jensen's actually enabled someone like Rainer to build a chip company like MatX, and I don't see them as competitive. The compute demand is so high. Like, I don't I think NVIDIA's not able to meet the demands of production, so we just need more chips. And I think it's very smart what MatX has done, which is say, “We're just going to we're not going to innovate on the data center design ‘cause actually, thank you, Jensen, you've done all the hard work. Where we can innovate is somewhere else.” And I think that's, that's very healthy. I think that's how we unblock new bottlenecks. And my view is these, the, chip teams like MatX, who have arrived at the insight that co-design is the way, The primary bottleneck for them is trust boundary. To do co-design well, you need visibility into the next model generation as soon as possible ‘cause it takes two years to tape out. So if by the time I bring my chip to market, your model architecture's changed, I'm host. Now, when he was inside Google, he was sitting next to the Gemini team. He was on Palm or whatever.Trust Boundaries, Co-Design, and Researcher CEOsSwyx [00:33:19]: His co-founder was the, was one, was one of the Palm guys, I think.Anjney [00:33:23]: Yes. Yes, exactly. So when you're inside the trust boundary of Google, then your systems co-design loop is super tight. When you leave as a founder, one of the biggest risks you take is now you're outside the trust boundary. And so what I love doing is helping chip teams who can help us unlock more capacity for the independent ecosystem access to trust. Because when I If I've been, involved with a lab from day one, and I was lucky enough to work with Anthropic, and then I'm on the board of Mistral and helped Black Forest Labs get started. I think at this point I'm on six or seven different teams.Swyx [00:33:57]: Only six? I feel like my mental number was going to be 13, but yeah, it's-Anjney [00:34:02]: No, I go deep with one at a time.Swyx [00:34:04]: You're founding CEO of Arena.Anjney [00:34:07]: Nah, that was an, that was an-Swyx [00:34:08]: Administrative CEOAnjney [00:34:09]: It was an administrative five-month gig where Whalen and Anastasios were graduating from their PhDs, and they didn't need a product team. So I helped recruit the head of engineering product and design. But Anastasios has always been the CEO of that company. I played a pinch-hitting I'm an intern. I was CEO intern For five months. -Swyx [00:34:33]: I interviewed him, and he's he's very well-spoken. I think he's a debate, former debate, champion. But also very quantitative and mathematical, which is-Anjney [00:34:41]: He-Swyx [00:34:41]: Such a unicorn.Anjney [00:34:43]: See, what's amazing about him? If you look at his output, he's an output maxer. By the time he was graduating from his PhD, which he only graduated last year, he had published more work with a citation count than, people twice his age. But at the same time, he'd already started a project called LLM Arena that was being used by millions of people As a side project. And time and time again, what I've realized is venture capitalists suck at seeing human beings as, dynamic agents where-Swyx [00:35:14]: They want to put you in a boxAnjney [00:35:15]: They want to put you in a box.Swyx [00:35:15]: This is your thing.Anjney [00:35:16]: So the first time I got introduced to Anastasios, somebody had told me “Oh, he's amazing, but he's a researcher.” I was “what? What do you mean he's a researcher?” That's what-Swyx [00:35:28]: Like he's not a CEO, not a founder.Anjney [00:35:29]: Not a CEO, exactly. I was “Are you crazy? Do you Have you met Dario?” Dario's a scientist. He's gone from zero to, what will soon be a trillion-dollar company in four years. Being a CEO, nominally speaking, is not that hard. Being a good CEO is hard. Being a great CEO actually requires a level of performance that scientists who have already published at the top of their field have accomplished. It is super hard to be a competitive scientist. To publish in academia over the last 20, 30 years, to make it to the top of your discipline at a place like Berkeley, you are a star athlete. Like, you are an athlete of the mind, and you perform at the highest levels. And to get there, whether you're, Anastasios or Whalen at Berkeley, or you are Robin, who-Swyx [00:36:23]: BFL, yeahAnjney [00:36:24]: With Black Forest, who created Stable Diffusion, or if you're, like Guillaume at Meta, who created Llama before he started Mistral. The amount of human leadership you have to demonstrate to get the resources, like get the trust of the organization, publish it, put it up. I would just fund researchers all day Right? If who have contributed already to the field. If they've, if they've put SOTA out there, they're, they're star athletes already. If they haven't done SOTA Look, they can still be good CEOs, but then I find the failure mode is that they just don't want to be CEOs, they primarily want to publish, and that's okay, too. One of the things we do with the AMP Grid is we donate excess compute. We have two nonprofits, like university labs. We carved out like a couple thousand H100s. But I do think there's extraordinary research being done on university campuses. My father-in-law's a physicist. He's a professor. Extraordinary work in physics, and we need that. But if you want to be a CEO, what you need to be willing To do is be super confrontational, outside of science. Like within the scientific community, some of the best researchers are very confrontational about their convictions, right? This architecture is right. To be a great CEO, you basically have to be willing to be confrontational up and down the stack.Swyx [00:37:41]: To your own team.Anjney [00:37:42]: To your own team-Swyx [00:37:43]: To customersAnjney [00:37:43]: Hiring, recruiting customers. Well, I would say, Yeah, pretty much to everyone Everybody. Of course-Swyx [00:37:50]: I see, I feel a little bit of that in my own work, but yeah, I can't imagine the stakes that Dario has had to go through. It's, it's pretty insane.Anjney [00:37:56]: No, I don't think the stakes are that different From how you're feeling it, right? Stakes are personal scaling vectors, right? The stakes that seem so low to you, like having this podcast where you can talk to somebody and just have a you're an extraordinary communicator, right? Like already in this conversation, you've pulled more out of me than most people, and I've been on 12 podcasts in the last two weeks.AI Coachella and First-Principles ThinkingSwyx [00:38:17]: I think I, we've just seen each other enough that there's some base trust.Anjney [00:38:20]: There's base trust.Swyx [00:38:20]: And I think, and I know that you, that I've done my homework and like I know that trust is a big deal for you, so.Anjney [00:38:27]: I think trust is about consistency, and you and I have seen each other In the community for years, right? Like, I remember the first time we met was at NeurIPS in New Orleans. I don't know if you remember that, luncheon.Swyx [00:38:38]: Oh my God.Anjney [00:38:39]: Reiko had set up this Reiko's amazing, and he set up this luncheon and-Swyx [00:38:43]: Yeah, I was “Who's this Discord guy?” I'm “Okay.” But-Anjney [00:38:45]: No, you weren't-Swyx [00:38:46]: You were just “You made some investments.”Anjney [00:38:47]: You were much less polite. You were “Who's this VC?” You're like-Swyx [00:38:51]: No, I Was I? Oh my God.Anjney [00:38:53]: It was-Swyx [00:38:53]: I'm so sorryAnjney [00:38:53]: It was visible on your face.Swyx [00:38:54]: I'm so sorry. But you weren't, you weren't The introduction was bad. I was I didn't know who you were.Anjney [00:39:00]: The, see, this is the thing about context, right? Like, but then I think I heard your accent. And I was “Are you-”Swyx [00:39:06]: Singapore, yeahAnjney [00:39:06]: “Are you Singaporean?” And you're “Yeah.” And I said, “I went to high school, JC, in Singapore.” And then the ice broke. But This is the there are in the scientific community, sometimes the stakes are very high for people who haven't had the emotional, what is called EQ Coaching and mentorship, right? Which is like to have scientific impact, you often need to be a extraordinary emotional, like emotionally in tune person with the folks you're trying to influence. And so what comes so naturally to you is actually a super high stakes thing to other people. And so I wouldn't assume that Dario's more stressed out than you. These things are you'd be surprised how similar and small sometimes the problems are to you That some of the world's biggest, leaders are facing. And that's what I've learned from this class. The guest speakers are Sam, Satya, Jensen.Swyx [00:40:01]: AI Coachella.Anjney [00:40:02]: Yeah. It's AI Coachella, right? So we got to get all the headliners, and they're I'm very lucky that some of these people have either mentored me over the years or I've done business with them. And when you, take the performative stuff out and any assumptions you may have about these people that you read in the press or on Twitter, We're all just humans. We're all trying to get along. And what's so special about this moment is AI is forcing, like scaling, the bitter lesson is forcing a lot of people to revise their assumptions for how the world works and go back to first principles or go and educate themselves. So the kind of people I was, I won't name who this person is, but I was at an event last week in Texas and, ran to somebody who said, “Anjney, I came across the class. What do you think about real time action prediction models?” And I was, don't know how happy it made me feel when they asked me that question. I know they've done the work. They've challenged themselves. I'm, they didn't ask me, “What do you think of world models?” They said, “What do you think of n-”Swyx [00:41:04]: Real time action predictionAnjney [00:41:05]: “action, real time action prediction models?” World models, don't get me wrong, are cool and everything, but you and I both know that is a layer of abstraction that is sometimes not usefully precise enough. Right? Ours-Swyx [00:41:16]: There's like four different kinds of world models.Anjney [00:41:17]: Yes, exactly.Swyx [00:41:18]: We've done the part with general intuition, by the way, which is very focused on, -Anjney [00:41:22]: Oh, cool. Yes. I love Pim. Pim is great. And this is what I love about people who've done that level of work. They realize they're not in competition with people who the rest of the world thinks they're in competition with.Swyx [00:41:34]: Because they're not in the category, they're in the specific thing they're trying to do.Anjney [00:41:37]: They're focused on their mission, and they have a systems understanding of the bottleneck they're trying to solve. And when somebody else says, “I'm working on real time, action prediction models too,” Pim goes, “Oh, I love that person. I want, I can learn from them.” But the minute they're “Oh, that person's a world model person,” it's “like which type of world model person?” But mostly they're just trying to figure out if it's a waste of their time, because we don't have enough time. So, Pim, for example, is super, loves this other company I work with we've talked about called Black Forest Labs. And he's mentioned to me multiple times that he's so, He thinks what Flux is doing is really cool. Andy Blattman came by and spoke in the class. And what I find over and over again is for people who do the work, who can be usefully precise enough about like what is actually going on in the world of frontier research, The sense of camaraderie is still well and alive, but it gets lost sometimes when you have to like abstract The technical complexities in, business terms And then the VCs are “How are you different from that world model?” I'm going to say Where do I even start to explain this stuff? And then the misalignment creeps in.Leading vs. Winning in Frontier AISwyx [00:42:43]: This is good. Yeah, I think, people listening get a sense of, what it is like to operate at a real level, like yourself, rather than at, the journalist level, where you have to sort of put everyone in, a rough category and create a narrative of competition, and who's winning today, who's behind.Anjney [00:42:58]: It-- this idea of winning is so Weird to me.Swyx [00:43:03]: You do want to win. You want you want competitiveness.Anjney [00:43:06]: No, I think you want to lead.Swyx [00:43:07]: You want SOTA.Anjney [00:43:07]: No, I think you want to lead. Yes, so you want to push the frontier. You want to push the SOTA. You want to do something that hasn't been done before. You want to capture value, but you don't want to capture so much value that, people think you're unaligned with your mission or trying to do what's best for the world. You want to capture enough value that you can keep innovating, right? And I think that people want to lead, they don't really This idea of winning and losing, again, I love Jensen. He's a, he's a leader. The mindset that he talked about on Dwarkesh's podcast, right? He's “I didn't wake up with a loser mindset.” I think that was awesome, right? Because he's, he's an engineer. Dwarkesh has done the work. So there's at least-- even though the, to me, it was very obvious they're talking about the same thing, they just passed each other. They just had to basically, Jensen has this, five-layer cake abstraction of how the industry works. And Dwarkesh had, I think from that podcast, had more of, a pre-training, mid-training, post-training systems loop concept.Swyx [00:44:04]: It's just a factor of who he talks to, right? Again, it's very clear.Anjney [00:44:06]: It's the systems It's the abstraction, the mental models, the It's the whole-- Dude, so much of the problem in the world is reasoning by analogy. And then the assumptions that are held invisibly.Swyx [00:44:19]: Yeah, I've, I've said, this is actually the best time in human history for first principles thinkers. Because everything you think will happen is actually now coming true.Anjney [00:44:28]: Correct. And the venture capital community is, notorious for this, where people look-- In times of uncertainty, they, cling to axioms that ended up being true from the previous era, and they kind of like proclaim them with confidence as if they're truths, but they're not. And it's very important to see the distinction between a heuristic and an axiom. An axiom can be proven-Swyx [00:44:55]: Like from internal consistency point of viewAnjney [00:44:56]: With internal consistency. A heuristic is a way you kind of a shortcut. And my God, the number of people I have had to put up with over the last few years who proclaim-- use heuristics As axioms to judge people, to judge which companies are going to succeed or the number of people who are “Oh, yeah, Anthropic, they're just training models right now,” but this one continue.Swyx [00:45:22]: Because that's a B2B SaaS?Anjney [00:45:23]: Yeah, the, like Which over the fullness of time, if you squint at it, maybe. But the way you arrive there is so important that you can-- you just, you can dismiss people. Here's what happened, right? What happened is Anthropic basically achieved takeoff in October of last year. That training run-Swyx [00:45:41]: Whatever, three seven?Anjney [00:45:42]: I forget the numbers now, but whatever that checkpoint was-Swyx [00:45:45]: We saw the cognition.Anjney [00:45:46]: Yeah. Right? You probably-- The, to those of us in the community, especially once post-training was done and it was released in December-Swyx [00:45:52]: Yeah. Can I sneak a sneaky question in there? I don't know if you have a perspective, maybe you don't, I just The number one question is how did Anthropic crack coding, right? Because Claude One, Claude Two, okay, like it was part of it, but it wasn't a big deal. And the leading hypothesis, it's a lucky dice roll that was then compounded, right? Like it was like Mildly better, but then they saw it and they were “Okay, let's really invest.”How Anthropic Cracked CodingAnjney [00:46:17]: I had this very annoying teacher. I went to this boarding school called Rishi Valley in India, which is like this, bird preserve. It's like three hundred and fifty acres of bird preserve in rural India, and there was no technology for seven years. There was this teacher, I won't name them, but they would have this-- I hated it every time he said this to me. He was “Luck fa-favors the prepared mind,” which is like a common saying, but the way he delivered it, always grated me, ‘cause he was always I was always one of those kids who got, a good grade without trying very hard. ‘Cause like high middle school is not that hard if you, if you're generally, paying attention and so on. And there was this one time where I-- But then I would get an eighty percent grade, and he would keep pushing me to say “The reason you didn't get the ninety-five plus percent is because you're not that lucky.” And I would say, “What do you mean?” ‘Cause I would think that I deserved that grade, and I would sometimes argue with him. And he'd say, “You didn't have a prepared mind. If you want to get lucky again “ There was basically one time where I got like ninety-five or ninety-six on this, on this subject, and I, now that I felt entitled. I was “Okay, I'm going to keep doing this,” and I didn't. And then he was “Luck favors a prepared mind. You got lucky last time, but you got to stay prepared.” And I didn't understand what he meant. Now, as I'm older, I'm okay, these adults actually knew a thing or two. Anthropic has been the most prepared company for four years. And so then when the right, context data comes in, the right developers start sending in, the right context diffs, Sure, you could say you got lucky, but if you ask me, they're pr-pretty damn prepared with paranoia for like four years. And you have to remember, it was so hard for them to get going early on that they had to do so much more with so much less that you just have to be prepared to be so efficient.Swyx [00:48:06]: Yes. There's numbers on their burn compared to OpenAI. I've, I've written about it, but they are so much more efficient in their, in their tech stack.Anjney [00:48:14]: It's not even It's not funny.Swyx [00:48:14]: Not even close.Anjney [00:48:15]: Yeah. But it's so clear, right? Like how to output max for the world. They have been prepared, and you could call that luck, but Luck favors the prepared mind.Culture, Hardship, and Anthropic's P0Swyx [00:48:25]: This is one of those things that I was going over some of your old lectures and, you were data, people think it's a moat and actually it's culture and actually it's team Actually. And I, it's-- there's different levels of moats, and this is the ultimate one that determines everything else. Which you can then compoundAnjney [00:48:43]: You're saying culture is the ultimate moat? Yeah. But the thing about culture is it's very fragile. So moats, I don't think they're-- there's very few moats I found that are actually moats. They're-- It's, it's a nice concept, but in reality, you have to replenish your culture. Ben Horowitz was, the speaker in CS153 on Tuesday, and I asked him this question about the culture bottleneck in teams because, there are several AI teams-Swyx [00:49:09]: His book, Hard Things About Hard ThingsAnjney [00:49:11]: Hard Thing About Hard Things. But more concretely, there are so many AI labs today that have all the cash they need, they have all the compute they need, and they're still not able to ship anything SOTA. And then you start seeing people leave and so on, and my diagnosis, it's, is it's the culture. And so I asked him, Ben, they're-- He's been one of the most aggressive investors in AI labs. He goes back to this thing which resonates in my mind a lot. It-- When I used to work at a16z, I would, book a conference room, and right outside the conference room, which is closest to the toilet ‘cause it was the fastest way for me to go use the bathroom between Zoom meetings-Swyx [00:49:45]: Oh my God, I'll put maxing my toilet optimization. Okay, never mind.Anjney [00:49:48]: It was not healthy in hindsight, but maybe this is TMI. But anyway, outside that conference on the wall was this quote that was printed that said, “Culture is not a set of beliefs, it's a set of actions.” And it's by Bushido, is this, Japanese philosopher. And if you stop taking the actions that demonstrate the mission alignment to what you've said to your team and to your-- the world matters to you, then your culture starts to fray. So it's not actually a moat, I would say. It's a very brittle, fragile thing that requires daily tending to like a garden. But if you figure out the system to keep that garden tended, which I think ultimately comes down to knowing yourself ‘cause you most naturally, if you're authentic and so on, you'll naturally make trade-offs that seem effortless to you, but that reinforce your culture. And then That becomes this very hard thing for other people to catch up to. And at Anthropic, from day one, there was this mission like-- missionary like zeal and belief that, hey, these capabilities will scale. These systems are stochastic, not deterministic. There will be error bars, and until we crack interpretability, there's risk. And at some point, people will go-- stop using Claude just for coding. They'll use it in some mission-critical context where there's-- it'll throw off a bug, and then people are going to come blame them, and they want to be on the right side of history where they said, “Yes, this is a powerful technology. We think it's going to change the world, And we want to be very measured and scientific about the fact that, ‘Hey, guys, these are stats models, statistical models.' That's how statistics works.” ultimately, when you're training neural nets, it is just a statistical system. And I think that Belief that safety is important and that it might seem toy-like in the early days, and sometimes, you could say, “Anjney, they totally over-exaggerated the risk,” like two years ago when they said, “Let's not launch Claude One,” or whatever. Well, okay, maybe in hindsight, but hindsight is twenty/twenty. And at the time, they didn't know how that model would be used, and to them it felt existential if somebody came and said, “You weren't responsible. It-- This wrote a bug.” The liability associated with that is massive. So how do you prevent against that? Well, day in, day out, you say safety. And when you start deviating from that, you have the team hold you accountable, you have the world hold you accountable, and I think that becomes a moat over time. At some point, that moat will get challenged and so on, and then it become fragile. I hope it endures because that's the beauty of having founders run the show, ‘cause they can make really hard trade-offs to do mission alignment. The hardest part is in the earliest days when you don't have a group of people who are going through difficulty, stress, crisis together, then your culture doesn't get defined sharply enough, and that's what I'm worried about right now, is there's so much money going to these labs. There's no hardship. There's no-Swyx [00:52:50]: To anyone who knowsAnjney [00:52:51]: There's no to anyone who knows. And that, in hindsight, was a feature, not a bug for Anthropic. The number of people who said no, the number of people who said, “Sorry, we're all doing investors in OpenAI,” that is competitive difference. It forces you to really understand, what is the hill you want to die on at the expense of everything else. What's the P zero? And there, P zero from day one was coding. The reason, the mechanism system there was if we crack coding, Then we will crack AGI. Our mission is AGI. We want to get there safely. If we focus on codin
This Open Source Startup Podcast episode has our co-hosts Robby and Tim in conversation with Dr. Felipe Huici, CEO of Unikraft - the compute layer for sandboxes, AI agents, or any workload with VM-grade isolation. Their open source, also called unikraft, has 4K stars on GitHub and provides a next-generation cloud native kernel. This episode explores how Unikraft is building infrastructure for the next generation of AI agents, arguing that agents should run in virtual machines rather than containers. The conversation focuses on the unique requirements of agentic workloads: fast startup times, the ability to pause and resume state, strong isolation, and efficient resource utilization at massive scale. Unikraft's technology enables lightweight virtual machines that can start in under 10 milliseconds, helping companies reduce latency, lower infrastructure costs, and run large numbers of ephemeral agents on minimal hardware. The discussion also covers emerging AI infrastructure needs such as checkpointing, branching, headless browser automation, and GPU access.The podcast also traces Unikraft's origins from an academic research project to an open-source Linux Foundation initiative and, eventually, a startup founded in 2022. The conversation examines customer adoption, the role of Unikraft as foundational infrastructure for AI platforms, competition and collaboration within the agent ecosystem, the future of GPUs and virtualization, and lessons learned from building a company in the rapidly evolving cloud and AI infrastructure market.
Our 248th episode with a summary and discussion of last week's big AI news!Recorded on 06/12/2026Note: we recorded just before the OTHER big news about Fable... we'll discuss it on the next episode.Hosted by Andrey Kurenkov and Jeremie HarrisFeel free to email us your questions and feedback at andreyvkurenkov@gmail.com and/or hello@gladstone.aiRead out our text newsletter and comment on the podcast at https://lastweekin.ai/In this episode:Anthropic released Claude Fable 5 (a safeguarded version of Mythos 5), showing major benchmark jumps and new risk findings in its system card (eval awareness, transgressive actions, CBRN concerns), alongside controversy over severe guardrails and silent downgrades.Apple announced Siri AI at WWDC, positioning a more capable conversational assistant integrated across iPhone features, reportedly built on a custom Gemini partnership; Google also rolled out Gemini 3.5 Live Translate and cut Google AI Plus pricing while bundling more storage.Business and infrastructure updates include OpenAI's confidential IPO filing amid an IPO race with Anthropic and SpaceX, Bezos-backed Prometheus raising $12B for “physical AI,” DeepSeek seeking a major external round, and Google paying SpaceX about $920M/month for GPUs.Open-source, safety, and policy developments feature new Gemma 4 and Diffusion Gemma releases, a lab letter urging DNA/RNA screening laws, Amodei calling for an FAA-like AI regulator and third-party testing, research on agent harms and RL “societal hacking,” and a dispute over music-label settlements with Suno/Udio.Timestamps:(00:00:10) Intro / Banter(00:01:11) News Preview(00:01:53) SponsorsTools & Apps(00:04:53) Claude Fable 5 and Claude Mythos 5 + Anthropic apologizes for invisible Claude Fable guardrails(00:27:06) Apple announces Siri AI and its next generation of Apple Intelligence | The Verge + I tried Siri AI, and so far it actually works(00:33:47) Gemini 3.5 Live Translate rolling out to Google Meet and Translate(00:35:39) Google just fired a warning shot in the AI subscription price wars | TechCrunchApplications & Business(00:37:55) OpenAI Confidentially Files for IPO on the Heels of SpaceX and Anthropic | WIRED (00:41:57) Jeff Bezos's Prometheus raises $12B to build an 'artificial general engineer' for the physical world | TechCrunch(00:45:39) DeepSeek slated to raise $7 billion in maiden funding round, sources say(00:48:18) Huawei-led team claims it post-trained DeepSeek's 1.6-trillion-parameter model — 1,000 Ascend 910C chips used in training(00:51:57) Google will pay SpaceX $920M per month for compute | TechCrunch(00:55:51) Elon Musk Shows Off AI Data Centers SpaceX Wants to Send Into Space - Business InsiderProjects & Open Source(01:01:14) Google's new Gemma 4 12B model is designed to run on any laptop with 16GB of RAM - Ars Technica(01:05:13) Google AI Releases DiffusionGemma, a 26B MoE Open Model Using Text Diffusion for Up to 4x Faster Generation - MarkTechPostPolicy & Safety(01:09:42) OpenAI and Anthropic Sign Letter to Prevent AI-Developed Biological Weapons | WIRED(01:14:04) Anthropic CEO publishes lengthy article: AI is moving too fast, and policies can't keep up. | PANews(01:20:18) Anthropic Urges Global Pause in AI Development, Flags ‘Self-Improvement' Risk - WSJ(01:24:46) When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use Agents(01:27:42) Large Language Models Hack Rewards, and Society(01:33:46) Senior US officials eye government shares in AI giantsSynthetic Media & Art(01:37:45) AFM Sues UMG, WMG Over Settlements With Suno and UdioSee Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
NuNet is building a decentralised compute and orchestration network where people can contribute spare CPU, GPU, RAM and other resources, while developers and organisations can deploy workloads across available infrastructure. In this episode, Peter talks with Jennifer from NuNet about the new NuNet Appliance and why it matters for making decentralised compute more practical for everyday users.The conversation covers how NuNet matches the right compute to the right job, how the Appliance lowers the barrier to onboarding devices, and why use cases like n8n automations, private AI agents, edge AI, Cardano SPO infrastructure and web deployment workflows are a natural fit for the network. Jennifer also explains NuNet's zero-trust security model, pricing approach, organisations, ensembles, deployment templates, and how NTX fits into orchestration fees.If you have spare compute, want to run private AI workloads, or are building in the DePIN and Cardano ecosystem, this episode gives a practical look at how NuNet is moving from concept to usable infrastructure.Key Takeaways:- NuNet is a decentralised compute and orchestration platform that lets people contribute spare compute and lets workloads find suitable resources automatically.- The NuNet Appliance is designed to make onboarding CPUs, GPUs, RAM and other compute resources much easier for non-expert users.- NuNet can support broad workloads, including n8n automation, private AI agents, Qwen-based LLM deployments, edge AI, web builds and Cardano SPO infrastructure.- The network uses a zero-trust model where machines are cryptographically identified and verified at each interaction.- Compute pricing is designed around stable currency values, with automatic conversion into NTX rather than forcing users to price workloads directly in a volatile token.- NuNet organisations can let other DePIN projects bring their own communities and native tokens while still using NuNet's orchestration layer.- Ensembles and templates are intended to simplify deployments so users do not need to manually understand every YAML configuration detail.- NuNet is open source, with docs, GitLab, Discord, Medium and X available for people who want to try the network or contribute.Links & References:- NuNet — Compute Orchestration for a Decentralized World: https://link.learncardano.io/eGKGuZ- What is NuNet? | NuNet Documentation: https://link.learncardano.io/rHu2E4- x.com: https://link.learncardano.io/NIhPKR- https://link.learncardano.io/Tlu7wNWebsite: https://link.learncardano.io/bQ68RcX/Twitter: https://link.learncardano.io/3a1QtvDisclaimer: This content is for educational purposes only. Nothing constitutes financial advice.DISCLAIMER: This content is for informational and educational purposes only and is not financial, investment, or legal advice. I am not affiliated with, nor compensated by, the project discussed—no tokens, payments, or incentives received. I do not hold a stake in the project, including private or future allocations. All views are my own, based on public information. Always do your own research and consult a licensed advisor before investing. Crypto investments carry high risk, and past performance is no guarantee of future results. I am not responsible for any decisions you make based on this content.
Roman Yampolskiy has spent two decades trying to prove that superintelligent AI can be controlled. He couldn't. I invited him on to make his case. Subscribe if you want science with evidence, not speculation. Roman is a professor of computer science at the University of Louisville and one of the earliest researchers in AI safety. His book AI: Unexplainable, Unpredictable, Uncontrollable started as an attempt to solve the alignment problem. After decades of work, it became a proof that the problem cannot be solved. Not difficult. Mathematically impossible. I push back hard. We go after the Einstein test: can a large language model trained only on pre-1911 physics reproduce what Einstein did with the same data? We ran that experiment. It failed. Roman and I disagree about what that means. We also get into the halting problem and what it actually tells us about predicting smarter-than-human behavior, whether value alignment is a real problem or a well-funded category error, the case for a government moratorium on frontier model development, and why Roman thinks giving an AI agent access to your computer is the dumbest thing a smart person can do. What you'll hear: Whether AI control is mathematically impossible or just unsolved Why Roman thinks all current AI safety work is security theater What the halting problem actually means for superintelligence The alignment problem: real issue or well-funded category error Why Roman wants a moratorium on frontier model development What to tell your kids about careers in a world where Roman might be right If you listen to other people, the best you can become is average. CHAPTERS 00:00 Creating a mind without an off switch 01:34 Solving problems beyond our own intelligence 04:08 Einstein's epiphany and the limit of AI intuition 08:18 Assessing the Einstein test: Why the experiment failed 12:22 Path dependency: Are LLMs and GPUs our QWERTY? 16:10 The barriers preventing AI from solving physics 21:54 Safety vs. Capability: Why toddlers are safe but teens are not 23:06 The halting problem: Predicting agents smarter than us 25:58 The impossibility of a system proving its own integrity 28:18 Regulation: Genuine safety or a gift to oligarchs? 33:28 Is human cognition non-computable? Penrose vs. the field 39:00 Ethical duties: Must we treat AI with humanity? 43:00 From internet memes to monsters: Decoding the book cover 46:22 Customized realities: Can everyone have their perfect world? 49:50 Von Neumann probes and the panspermia hypothesis 55:02 Categorizing AI: The one version that should terrify you 58:22 Pause AI: The movement for a development moratorium 59:58 Career advice for kids in a post-professional world 01:07:58 Cross-examining Sam Altman 01:15:48 Roman's dream debate 01:19:50 Lessons for a younger self Substack: https://briankeating.substack.com Get the transcript, fascinating bonus content, and my Monday M.A.G.I.C. Message: https://briankeating.com/yt Have a .edu email and live in the USA? You automatically win a meteorite: https://BrianKeating.com/edu Subscribe: https://www.youtube.com/DrBrianKeating?sub_confirmation=1 Support Into the Impossible on Patreon, get my weekly M.A.G.I.C. Message, unfiltered bonus content, and live monthly Office Hours with me: https://www.patreon.com/drbriankeating Join this channel for perks, monthly Office Hours, and your name in the Member Roster at the end of every episode: https://www.youtube.com/channel/UCmXH_moPhfkqCk6S3b9RWuw/join Featured Guest: Roman Yampolskiy on Twitter/X: https://x.com/romanyam?lang=en AI: Unexplainable, Unpredictable, Uncontrollable: https://www.romanyampolskiy.com/books/ My books: Losing the Nobel Prize (memoir): http://amzn.to/2sa5UpA Think Like a Nobel Prize Winner: https://a.co/d/03ezQFu Focus Like a Nobel Prize Winner: https://a.co/d/hi50U9U Galileo's Dialogue (first-ever audiobook): https://a.co/d/iZPi9Un Twitter/X: https://x.com/BrianKeating Substack: https://briankeating.substack.com Blog: https://briankeating.com/blog Audio-only: https://briankeating.com/podcast #intotheimpossible #briankeating #AIrisk #artificialintelligence #aisafety #podcast #superintelligence #RomanYampolskiy Learn more about your ad choices. Visit megaphone.fm/adchoices
What happens when a company focused on drug discovery and life sciences encounters a data problem that nobody else seems able to solve? Recorded at the IT Press Tour in Boston, this episode explores the fascinating story behind Paradigm4 and how a challenge in large-scale biomedical research ultimately led to the creation of flexFS, a cloud-native filesystem designed to tackle some of today's biggest data infrastructure challenges. Joining me on the podcast is David Freund from Paradigm4, who shares how the company was originally founded to help scientists work with enormous datasets in fields such as genomics, bioinformatics, and precision medicine. As researchers began working with population-scale datasets such as the UK Biobank, the team discovered that existing storage technologies either couldn't deliver the performance they needed, lacked the functionality required, or became prohibitively expensive at scale. Our conversation explores the moment Paradigm4 realized it would need to build its own solution, why traditional approaches to cloud storage often struggle under modern analytics workloads, and how flexFS emerged from a real-world customer problem rather than a technology trend. David also explains why object storage has become such an attractive foundation for modern infrastructure, while discussing the challenges of latency, performance, and cost that still need to be addressed. We also discuss why many organizations investing heavily in AI infrastructure may be overlooking one of the biggest constraints on performance. While much of the industry conversation focuses on GPUs and compute power, David argues that data access, movement, and management are becoming equally important considerations as AI workloads continue to grow. Along the way, we touch on cloud independence, resilience, large-scale analytics, and why flexibility across cloud providers is becoming an increasingly important requirement for enterprise technology leaders. Whether you're working in AI, life sciences, cloud infrastructure, or enterprise data management, this episode offers an interesting perspective on how customer problems can sometimes lead to entirely new categories of technology. Could the next major AI bottleneck be data rather than compute? And are organizations paying enough attention to the infrastructure feeding their most important workloads? I'd love to hear your thoughts.