POPULARITY
In this episode of the Crazy Wisdom Podcast, host Stewart Alsop speaks with Aaron Neyer, founder of Parachute and community organizer in Boulder, about knowledge management, extended minds, and the intersection of AI with human consciousness. They explore how Parachute functions as a digital brain tool for organizing thoughts and information across fragmented systems, discuss the dangers of AI psychosis and over-reliance on technology, and debate open source AI development versus controlled releases by companies like Anthropic. The conversation weaves through topics including the limitations of metrics-driven business thinking, consciousness and relevance realization, the value of technological sabbaths, and Aaron's hope for locally-run open source models that protect personal data while still accessing more powerful gated models when needed. You can find Aaron's writing at unforced.org and unforced.substack.com, and learn more about Parachute at parachute.computer and parachute.computer/blog.Timestamps00:00 Stewart welcomes Aaron Neyer, founder of Parachute and Boulder community organizer, discussing the origin of Parachute's name from Frank Zappa's quote about open minds.05:00 Aaron explains Parachute as an extended mind tool for organizing notes, contacts and information across multiple platforms, emphasizing the distinction between primary mind and extended mind as interconnected systems.10:00 Discussion shifts to metrics-driven business culture and the limitations of pure rationality, exploring how Google's data-driven approach misses subjective experience and the whole picture of relationships.15:00 Aaron discusses AI's ability to help identify relevant variables across different domains and the dangers of AI psychosis, comparing it to cult dynamics and belief systems.20:00 The conversation covers AI sabbaths and nineties retreats as intentional breaks from technology, plus Aaron's experiences with electrical engineering and using AI to design circuits with Arduinos.25:00 Exploring forbidden knowledge and open source AI, Aaron discusses Anthropic's guardrails around powerful models while arguing for distributed access to prevent concentration of power.30:00 Deep dive into open source AI strategy, with Aaron highlighting NVIDIA's approach and the potential for running capable models locally while reserving ultra-intelligent models for complex research tasks.35:00 Aaron shares his vision for local Sonnet-class models handling personal data while accessing Fable-class models for deep research, and directs listeners to unforced.org and parachute.computer for his writing.Key Insights1. The philosophy behind Parachute stems from Frank Zappa's quote that the mind is like a parachute and doesn't work if it isn't open. Aaron Neyer explains that having an open mind is valuable, but it must be balanced with deep roots to avoid becoming untethered. He has experienced periods in his life where excessive openness led him to feel disconnected, teaching him that creativity and expansion need to be grounded in something substantial. This same principle applies to how we organize information digitally, where openness and interoperability allow our extended minds to become more connected and coherent, which in turn helps our primary minds think more clearly.2. Parachute is designed as an extended mind tool that addresses the fragmentation problem in how we currently manage information. Most people use multiple disconnected tools like Obsidian, Notion, Apple Notes, Google Keep, and various CRMs to organize their thoughts, notes, and relationships. These systems don't communicate well with each other, creating inefficiency and confusion. Parachute aims to create a simple, intuitive system where all this information can be organized in one place with true interoperability, allowing users to own their data and have it speak effectively with other tools, ultimately making our entire extended mind more functional.3. Understanding ourselves as unified body mind organisms rather than fragmented parts is essential for effectiveness. Living systems theory shows that any living system is three things: a membrane bound dissipative structure, a self regulating autopoietic network, and a cognitive process actively knowing the world. Western civilization since Descartes and Galileo has created artificial separation between body and mind, and between subjective and objective experience, which limits our effectiveness. The same fragmentation affects our digital technology, and recognizing both our biological and digital systems as coherent wholes rather than disconnected parts makes us vastly more capable.4. The relationship between data driven approaches and holistic thinking reveals important limitations in modern business and science. While working at Google, Aaron observed how data driven decision making can be powerful, but over reliance on metrics like ROI creates blindness to crucial unmeasurable factors like goodwill and relationship quality. This reflects a broader Western tendency to exclude subjective experience because it's difficult for objective science to measure. However, emotions, relationships, and other subjective elements are essential parts of reality, and focusing only on quantifiable metrics means missing the whole picture and ultimately becoming less effective despite appearing more rational.5. AI accelerates the ability to work with technical complexity by helping with relevance realization across domains where we lack expertise. In any specialized field, experts develop intuitive senses for which variables matter and can quickly identify problems, whether in computer troubleshooting, music, or cooking. AI's ability to generalize allows it to point people toward relevant solutions in areas where they haven't developed that intuitive expertise, effectively democratizing technical capability. This means people can direct their creativity more effectively across more domains, though it also raises concerns about giving powerful capabilities to those who may lack the wisdom to use them responsibly.6. The question of open source AI versus gated access involves complex tradeoffs between democratizing power and preventing harm. Aaron respects Anthropic's approach of creating guardrails around powerful models like Mythos, which would likely have caused significant system hacks if released without restrictions. However, this creates concerning power dynamics where only wealthy companies, governments, and their allies have access to the most powerful tools. NVIDIA offers hope through their truly open source approach including full training pipelines, and there may be a viable path where open source models at the Sonnet capability level handle most tasks locally while more powerful Fable class models remain gated for the most demanding work.7. Creating intentional breaks from AI and technology is essential for maintaining clear independent thinking. Aaron practices an AI Sabbath at least one day per week when he doesn't interact with AI, and he finds these are the days when he does his best thinking and journaling. Without these breaks, he finds himself constantly jumping between journaling and prompting AI rather than giving himself space for deep reflection. This pattern mirrors broader concerns about AI consistency creating cult like dynamics similar to organized religion, where constant immersion in a belief system or technology can lead to losing the ability to think independently, making periodic disconnection crucial for maintaining cognitive autonomy and clarity.
Is Cerebras Systems the next great AI chip stock or a red-hot IPO priced for perfection? In this episode of 7investing Live, Simon Erickson and executive producer Heather Horton welcome back Nick Rossolillo, co-founder of Chip Stock Investor, to break down three of the market's biggest stories.First up: Cerebras Systems (NASDAQ:CBRS), the wafer-scale chip maker that just IPO'd at a $40+ billion market cap. With 44GB of SRAM embedded directly on the chip, Cerebras was purpose-built to solve AI's "memory wall" problem for inference workloads. Now it's reportedly landed a ~$10 billion order from OpenAI and a deal with Amazon Web Services that could top $20 billion. Simon and Nick dig into whether these massive orders are real, how Cerebras stacks up against NVIDIA's GPUs and hyperscaler custom silicon, the TSMC capacity bottleneck that could throttle its growth, and how to value a company trading near 20x sales without profits.Then the conversation turns to Rocket Lab (NASDAQ:RKLB), which has pulled back from $150 to around $70 per share. Simon shares the latest iteration of his discounted cash flow valuation, and the duo debates the proposed Iridium acquisition — a deal that could pull Rocket Lab to EBITDA-positive on a pro forma basis — plus what the long-awaited Neutron rocket launch means for the company's future.Finally: Netflix (NASDAQ:NFLX). After another quarter of decelerating revenue guidance, is the streaming giant now a value stock rather than a growth stock? Nick explains why the advertising business hasn't reaccelerated growth the way he expected, and what he'd need to see before buying the dip.Plus: Nick's take on the recent chip stock sell-off across NVIDIA, AMD, Broadcom, SanDisk, and Kioxia and why "stocks go up, stocks go down" might be the healthiest way to think about it.Subscribe for more deep dives on AI infrastructure, semiconductors, and innovative growth stocks!Start your FREE 7-day trial of 7investing: https://www.7investing.com/subscribeFollow Nick and Casey Rossolillo at Chip Stock Investor: https://chipstockinvestor.comRocket Lab Deep Dive videos mentionedPart 1 https://youtu.be/AMDd0-JKUH0 (Deep Dive)Part 2: https://youtu.be/Z76xTGFNwBA (Valuation)Companies MentionedPublicly Traded:Cerebras Systems (NASDAQ:CBRS)Rocket Lab (NASDAQ:RKLB)Netflix (NASDAQ:NFLX)NVIDIA (NASDAQ:NVDA)Advanced Micro Devices (NASDAQ:AMD)Broadcom (NASDAQ:AVGO)Micron Technology (NASDAQ:MU)Taiwan Semiconductor Manufacturing (NYSE:TSM)Amazon (NASDAQ:AMZN)Alphabet (NASDAQ:GOOGL)Meta Platforms (NASDAQ:META)Iridium Communications (NASDAQ:IRDM)SanDisk (NASDAQ:SNDK)Kioxia Holdings (TSE:285A)Globalstar (NASDAQ:GSAT)SpaceX (NASDAQ: SPCX)Private / Pre-IPO:OpenAIAnthropicVideos Mentioned:https://www.youtube.com/watch?v=Z76xTGFNwBA&t=3shttps://www.youtube.com/watch?v=AMDd0-JKUH0&t=987sHere's the shifted chapter list, with all timestamps moved back 55 seconds:0:00 Welcome to 7investing Live0:54 Cerebras Systems: IPO recap & the Wafer-Scale Engine2:31 Is NVIDIA even the right comparison for Cerebras?5:38 The memory wall: why bigger AI models need new chips8:52 Latency vs. throughput — and the new AI alliances10:46 Are the $10B OpenAI & $20B Amazon orders real?14:02 Cerebras risks: how do you value a hot IPO?17:27 The TSMC capacity bottleneck20:01 Heather's take on Cerebras20:41 Rocket Lab: the sell-off & Iridium acquisition24:34 Simon's DCF valuation & price target for RKLB29:05 Why Neutron changes everything30:12 Q&A: Does Peter Beck carry an "Elon premium"?31:36 Netflix: buying opportunity or cheap for a reason?36:57 Q&A: Is Netflix a growth stock or a value stock?39:03 Chip stocks selling off: normal volatility or a warning?42:57 Wrap-up & final thoughts#7investing #Simonerickson #Cerebras #CBRS #NVIDIA #AIinvesting #semiconductors #chipstocks #RocketLab #RKLB #Netflix #NFLX #AIinference #stocks #investing #stockmarket #TSMC #AIdatacenters
A $400 million chip-backed loan points to the next wave of AI infrastructure deals. Also, BP Ventures is shutting down, ending a nearly 20 year run that was marked by reportedly lackluster returns. And X will use Grok AI to better detect stolen content, redirect payouts to original creators, and crack down on engagement bait. Learn more about your ad choices. Visit podcastchoices.com/adchoices
This week's video transcript summary is here. You can click on any bulleted section to see the actual transcript. Thanks to Granola for its software.EditorialIntelligence: Who Owns it?This week the word “AI” feels too small.AI is a technology. Intelligence is its product. And if intelligence is the product, the question is no longer just: Which model is best? Who has the cheapest tokens? Who owns the weights? Who controls the data center? Those are important questions, but they are lower in the stack.The bigger question is simpler and more political:Who owns intelligence?That sounds abstract until you make it concrete. Intelligence is becoming something companies can capture, package, serve, meter, route, improve, and sell.It can write code, answer questions, design molecules, automate offices, run agents, draft legal work, advise scientists, serve consumers, and reshape workflows. It is not merely software. It is a general-purpose capability. And all humans could benefit from more of it.General-purpose capabilities have a habit of becoming public questions. But the default answer, that public good is best delivered by government, is the wrong answer in this context.The Product Is IntelligenceWe should stop talking about AI as a feature and start talking about intelligence as the universal thing that is delivered as an input to the world.Water is an input. Electricity is an input. Literacy is an input. Connectivity is an input. Once a society depends on them, access stops being optional. Nobody needs government to build every well, power plant, school, or network. But everybody understands that a civilization cannot be organized around less than universal and reliable access to foundational inputs.Intelligence is reaching that level of importance now that we all know it is real.Government should not own it, operate it, or develop it. Quite the opposite. Companies are the right actors to build fast, compete hard, improve models, serve customers, and discover the real use cases. Self-interest is a useful framing here. Markets are good at finding demand, reducing costs, and turning invention into services people actually use.Companies are the right operators, developers, and owners. But that does not settle the real question of who owns the benefits. That is an economic question.If intelligence becomes metered infrastructure, what happens to the value it creates?The Ownership StackThis week's articles keep circling the same issue from different directions but in the nature of ‘circling' never quite nail it.Jamin Ball's “Own Your Weights” starts with the enterprise version of the question. Owning a model file is not enough. The durable asset is the loop: the data flywheel, the evaluations, the reinforcement system, the workflow learning, and the operating context that lets capability compound.Benedict Evans' “Ways to Think About Token Pricing” adds the market layer. Tokens may become essential, abundant, and cheap, like mobile data. But being essential does not guarantee that the token layer captures the value. The money may move up the stack to whoever owns the workflow, the customer, the distribution, or the application.Alex Karp's fight with the labs, reported in “Alex Karp Is Saying What Every Angry CEO Is Thinking About AI”, is the same argument in sharper enterprise language. Companies are afraid that model providers will not just sell intelligence, but learn from customer workflows and then move into the markets where those workflows create value. The “All-in” group are echoing Karp's view.And “What Is Loop Engineering, and Who Owns It?” names the new contested terrain. The loop is where intelligence meets the world. Whoever owns the loop owns the learning. Whoever owns the learning owns the compounding asset.That is why “who owns intelligence?” is not a slogan. It is the question under the model layer, the application layer, the enterprise layer, and the economic layer.Because intelligence is the product, the tools creating it are fragmented and competitive. So there is no logic in trying to discuss this at the level of a single company or set of tools and models.The Old Promise Was That Commerce Would Tame PowerThe essays this week give the historical backdrop.Deirdre McCloskey, in “What Really Caused the Industrial Revolution”, argues that modern growth came not simply from capital accumulation, but from a change in permission: ordinary people were allowed to innovate, trade, build, and be honored for it.That matters because intelligence could be another expansion of permission. It could make more people capable of building, learning, creating, coding, researching, translating, selling, and coordinating. It could lower the cost of competence.But only if access is broad.Paul Krugman's “AI in an Age of Oligarchy” warns that the same technology lands differently in different political economies. A new general-purpose technology entering a broad, open, upwardly mobile society is one thing. The same technology entering a concentrated economy, with extreme wealth and weak counterweights, is another.Tim O'Reilly's Economist essay, “Elon Musk is building a form of capitalism that Adam Smith would hate”, makes the governance point more directly. The old liberal hope was that commerce would tame arbitrary power. Markets, boards, courts, shareholders, disclosure, and competition would discipline the prince.But what if the prince uses markets to escape discipline?Henry Farrell's “political economy of billionaire derangement” pushes the same point. Founder culture, monopoly ambition, peer rivalry, weak correction mechanisms, and vast private control can amplify appetites rather than restrain them.The danger with intelligence is not that companies build it. They should. Companies build it, meter it, use public tolerance and public infrastructure to scale it, learn from everyone who uses it. All of those things are inevitable and healthy. Market forces will sort out winners from losers. The real danger is that the winners treat all of the surplus produced as purely private.Metered Intelligence Creates SurplusIf metering is not the problem, what is?The problem is pretending that metered intelligence creates value only for the metering entity. Metering water is only tolerated as a public good. If the public were blackmailed by a private water company with the threat of no water we would all rebel.Once we understand that the product of AI is intelligence we can see that every time intelligence is used, there is the immediate transaction: the user pays, the provider serves.But there is also system value. Usage creates signals. Workflows reveal patterns. Prompts, corrections, failures, preferences, integrations, edge cases, and business processes all help define where intelligence is useful and how it should improve. Intelligence breeds intelligence.Even when customer data is contractually protected, the market learns. The platform learns where demand is. The product team learns which workflows matter. The ecosystem learns which jobs are vulnerable, which tasks are automatable, and which parts of the economy can be reorganized around machine intelligence.So the surplus is not born in a vacuum.It rests on public science, public education, public data exhaust, public law, public infrastructure, public energy systems, public tolerance for data centers, and billions of human interactions. It is served by companies, but it is not made only by companies.This is why “Americans Deserve a Dividend From AI Companies' Riches” belongs at the center of this week's issue. The detail can be debated. The principle is harder to dismiss. If intelligence becomes a new foundational resource, then some part of the wealth it creates should flow back to the people whose society makes it possible. Intelligence did not suddenly appear. AI is built on the entire history of human intelligence. It benefits from it and at the same time evolves it.Not Nationalization. A Human Wealth Fund.If intelligence belongs to everybody, some conclude that government ownership of intelligence is the right outcome.Governments are not well suited to build, operate, or improve intelligence. They will move too slowly, regulate too early, politicize the wrong things, and confuse economic participation with operational control.Andrew McAfee's “Why I Didn't Sign the AI Open Letter” is useful here. His objection is not that the technology is unimportant. It is that steering too hard before we understand the shape of the change can become its own failure mode. Marc Andreessen's satire of AI regulation is less policy than temperament, but it captures a real Silicon Valley fear: that regulation can become permission, capture, and incumbency before it becomes wisdom.That fear should be taken seriously.But it does not answer the economic question. It answers only the operational one.How can the economic benefits of intelligence be distributed? The better answer is a sovereign human wealth fund.Call it a sovereign wealth fund if you must, but the phrase is too national. Intelligence will not respect borders. The leading companies are global. The models, chips, data centers, agents, platforms, and workflows will be transnational from the beginning. If the value created by intelligence is global, then the mechanism for sharing some of that value should begin with the companies global enough to capture it. The nice thing about xAI, OpenAI, and Anthropic is that they are supranational.These companies own and operate intelligence. Let them compete. Let them profit. Let them keep the incentives that make the system improve. But if intelligence is the new water, the wealth it creates cannot belong only to the companies that meter it. And they, themselves, have the power to fix it, even more than governments.Access will become a Human Right; Ownership Is the Economic DesignThis is where human rights come in. There is no right to access an AI model, yet. But there will soon be a need to change that.Not as a claim that every person is entitled to every frontier model at every moment for free. That is not serious. Capacity has costs. Models have costs. Inference has costs. Data centers have costs. Although those costs will decline over time, possibly quite quickly as self-learning models address costs.The claim is more basic: in a world where intelligence becomes a primary input into education, work, health, science, citizenship, creativity, and economic agency, baseline access to intelligence starts to look like a civic requirement.That could mean public access layers. It could mean education credits. It could mean open models. It could mean AI dividends. It could mean public-interest compute. It could mean taxes on rents. It could mean a company-initiated human wealth fund that returns some of the upside to society without handing the operating system to the state. The latter could couple wealth growth with universal distribution of ownership.The exact mechanism matters. But the distinction matters more.Government should not own intelligence. It should be universally available. And people should have a claim on the wealth intelligence creates.The Frontier Is Also PhysicalThe abstraction is not weightless.“The Fight Against AI Data Centers Is Just Beginning”, “New York becomes the first state to enact a data center moratorium”, Reuters on pollution from Musk's xAI power project, and DataGravity's “Who Captures Value in AI Infrastructure?” all say the same thing from the ground up.Intelligence uses land. It uses power. It uses water. It uses chips. It uses grid capacity. It uses neighborhoods. It uses public patience.That makes the value question unavoidable. A society can accept the buildout if the buildout is legible as shared progress. It will resist it if the costs are local, the profits are private, and the benefits feel enclosed.Who Owns the “Loop”?The week ends where it began.“Anthropic and Blackstone” are betting that implementation is the next trillion-dollar business. “Vint Cerf” is working on identity for agents on the open internet. “GPT-Red” points toward systems that improve their own robustness. “Kimi K3” adds another open frontier model to the global mix.The model race continues. The deployment race is accelerating. The governance race is behind.My view is this:The central product of this era is intelligence. Companies have figured out how to capture it, package it, serve it, and meter it. That is good. It should stay in the hands of builders who have the incentive to make it better.But intelligence is too foundational to become just another private toll booth. A significant part of it will turn out to be free to users.As intelligence becomes a general-purpose resource, then access to it becomes a human-capability question, and the surplus from it becomes an economic-justice question. Not because government should run it. Because government should not run it. The operating layer belongs with companies. The wealth question belongs with everyone. But companies are best placed to turn that into a process of distribution.The question is not whether companies should build intelligence. They should.The question is whether humanity gets a stake in the wealth created by the thing that may soon become its most important shared input.Contents* Essays* Deirdre McCloskey on What Really Caused the Industrial Revolution* AI in an Age of Oligarchy* Elon Musk is building a form of capitalism that Adam Smith would hate* Murky Mirror: Truth and Consequences* The political economy of billionaire derangement* Is there any “oligarchy” to fight?* AI* Nearly 200 Economists and Tech Leaders Warn of A.I. Threats* Why I Didn't Sign the AI Open Letter* Own Your Weights* Ways to Think About Token Pricing* Alex Karp Is Saying What Every Angry CEO Is Thinking About AI* The AI Agents Are Coming for Microsoft Office* What Is Loop Engineering, and Who Owns It?* The Fight Against AI Data Centers Is Just Beginning* 6 months to live for open models* Americans Deserve a Dividend From AI Companies' Riches* Who Gets to Define the Frontier?* GPT-Red: Unlocking Self-Improvement for Robustness* Anthropic, Blackstone bet the next trillion-dollar AI business is implementation, not just models* Vint Cerf is working on a plan to unleash AI agents on the open internet* xai-org/grok-build, now open source* The Pulse: What can we learn from Bun's rapid Rust rewrite with AI?* Orphan risks at the frontier of artificial intelligence* The Lab of the Future Should Feel Like a Data Center* Why AMI Labs' Alexandre LeBrun won't call his AI “AGI” or “superintelligence”* Kimi K3 Tech Blog: Open Frontier Intelligence* Venture Capital* Three Years In* Venture Has Rarely Looked More Bifurcated* The Best Angel Investors in the US: Who Backs the Most Unicorns, and Who's Active Now* Are Prediction Markets Doomed to Fail?* Regulation* Exclusive: The Next Frontier of the Deportation Wars: College Campuses* The Supreme Court Broke Independent Agencies. Here's a Way to Slow the Damage.* India's crackdown on a new WhatsApp feature risks setting a global precedent* Let's build a children's public internet* Computer cops* Google is better at playing the AI regulations game* Infrastructure* Who Captures Value in AI Infrastructure?* New York becomes the first state to enact a data center moratorium* Pollution from Musk's unpermitted xAI power project hits hardest in Black communities* Interview of the Week* The End of the End of Geography* Startup of the Week* Radical AI's Joseph Krause: The Scientist Building The “Waymo” Lab For New Materials* Post of the Week* Marc Andreessen on AI RegulationEssaysDeirdre McCloskey on What Really Caused the Industrial RevolutionYascha Mounk and Deirdre McCloskey | Persuasion | July 11, 2026Yascha Mounk interviews Deirdre McCloskey about her argument that the modern world's economic liftoff came less from capital accumulation than from a change in ideas. McCloskey says both left and right versions of the conventional story rely too heavily on investment: the left stresses exploitation and surplus value, while the right stresses virtuous saving by capitalists. Her objection is historical and economic. Human beings had always invested, from irrigation works and Roman roads to seed grain, and simple accumulation quickly runs into diminishing returns.McCloskey's alternative is that northwestern Europe, first Holland, then Britain and Scotland, and then the North American colonies, developed a liberal ideology that changed who was allowed to innovate and be honored for it. The conversation links that shift to the erosion of inherited hierarchy, the spread of dignity for ordinary commercial life, and a moral vocabulary in which liberalism is not merely procedural but connected to virtues and values. The point is not that machines, coal, trade, and institutions did not matter, but that they do not explain the scale and timing of modern enrichment without a cultural permission structure for innovation.The interview also turns to the contemporary defense of liberalism. Mounk frames the series around the worry that liberalism is often treated as too thin to command allegiance, while its opponents speak more directly to moral passions. McCloskey's case is that liberal societies became rich because they dignified experimentation and ordinary enterprise, and that liberals need to recover the moral language behind that claim.Read moreAI in an Age of OligarchyPaul Krugman | Paul Krugman | July 12, 2026Paul Krugman frames AI as a major technological shock arriving inside an already unequal political economy. The post says AI's economic and social effects may take years to understand, but argues that the setting matters now: America has much greater wealth concentration and political inequality than it did in the 1950s and 1960s, when progressive taxation, stronger regulation, and more active antitrust might have contained some of the destructive effects of a new technology.Krugman's opening claim is that the same technology would likely have different consequences in a more level society. In today's United States, he writes, extreme wealth is both a cause and effect of policies that favor a small elite, including low effective taxes on capital and high incomes, weak enforcement of worker protections and antitrust, and cuts to programs that benefit ordinary Americans.The article is explicitly more about oligarchy than AI. Krugman says the paid sections document the rise of the “.0002%,” the economics and politics of extreme wealth, how oligarchy will shape AI's impact, and possible policy paths. His caveat is that AI itself may still produce a pushback against oligarchy, but absent that, he expects the pre-existing concentration of wealth and power to magnify AI's downsides.Read moreElon Musk is building a form of capitalism that Adam Smith would hateAuthor: Tim O'Reilly Published: July 12, 2026Tim O'Reilly argues that Elon Musk is using the legal forms of shareholder capitalism to escape the restraints that shareholder capitalism was supposed to impose. The article begins with SpaceX's public-market structure: ordinary public investors get little meaningful governance power, Musk keeps roughly 85 percent of the votes through super-voting shares, buyers waive jury trials and class actions, the company qualifies as controlled, and removal of Musk depends on the share class he controls. In O'Reilly's framing, that is not ordinary founder control; it is a design for being answerable to no one, possibly beyond Musk's own lifetime.The killer detail is the article's turn through Albert Hirschman, Montesquieu, James Steuart, Adam Smith, and Keynes. Older defenses of commerce held that markets would tame princely passions because the self-interest of merchants was safer than arbitrary rule. O'Reilly says Musk reverses that hope. The market discipline that was supposed to cage the prince has become the lever by which the prince raises capital, removes feedback loops, and carries private power into politics, government, Mars, robots, AI, or whatever ambition comes next.The pull is the link to AI governance. O'Reilly says corporations are already a kind of artificial intelligence: narrow-input systems that act at a scale no individual human can match. Their partial controls include independent boards, shareholder votes, courts, disclosure, regulators, public pressure, and activism. If the leaders building frontier AI strip those alignment mechanisms out of their own companies, the governance of the company becomes a preview of the governance of the machine.Read more: The EconomistMurky Mirror: Truth and ConsequencesAuthor: Esther Dyson Published: July 14, 2026Esther Dyson argues that today's institutional crisis is better viewed through the 14th century than through recent political history. Using Barbara Tuchman's A Distant Mirror as her frame, she compares a world of famine, plague, church schism, feudal predation, and purposeless war with a present in which institutions again feel brittle, incentives are badly aligned, and power is shifting into forms that are hard to govern.The killer detail is the historical analogy between land, corporations, and AI. Dyson moves from nobles who controlled serfs and territory, to the East India Company as a quasi-sovereign business, to today's AI systems and data centers as a possible new sector that crosses and weakens both nation-states and companies. The question is whether AI becomes a new kind of private land, owned by a new nobility, or an open prairie that many people can cultivate.The pull is human attention. Dyson says the central question is not what AI will do to people, but how people will react to it: whether they can value love, kindness, embodied attention, and artisanal human presence in a world of seductive artificial offerings.Read more: SourceThe political economy of billionaire derangementAuthor: Henry Farrell Published: July 15, 2026Henry Farrell argues that the visible political radicalization of some Silicon Valley billionaires is not a random personality quirk, but a product of the political economy that made them. Starting from Tyler Cowen's dismissal of “billionaire derangement syndrome” and Tim O'Reilly's warning that Elon Musk is using shareholder capitalism to escape shareholder restraint, Farrell flips the phrase: the question is why billionaires themselves can become deranged.The killer detail is Farrell's use of Peter Thiel as both theorist and example. Thiel's Stanford lectures described startups as monarchies and founders as figures vested with unusual power, while Silicon Valley culture rewarded eccentricity, monopoly ambition, and founder exceptionalism. Farrell says those ideas combined with dense founder-investor networks, peer rivalry, and weak correction mechanisms to amplify rather than discipline princely appetites.The pull is the ideological problem for classical liberals who once saw tech wealth as an ally of markets and freedom. Farrell says commerce did not tame the passions; in parts of Silicon Valley, the passions have begun to devour markets, institutions, and the liberal story that justified them.Read more: SourceIs there any “oligarchy” to fight?Matthew Yglesias | Slow Boring | July 16, 2026Matthew Yglesias argues that “oligarchy” is a rhetorically powerful but analytically loose way to describe American politics. The post begins from Bernie Sanders' “Fighting Oligarchy” tour, Amy Klobuchar's warning about a MAGA “broligarchy,” and the long afterlife of the Martin Gilens and Benjamin Page paper that was widely summarized as showing that only the rich matter in policy outcomes. Yglesias says the evidence supports a weaker claim: affluent people and business leaders have unusual access and influence, but that is not the same as rule by a small cabal.His main distinction is between inequality and oligarchy. The Gilens-Page measure treated the top 10 percent of households as “the wealthy,” and later critics found that rich and middle-class preferences usually align; in the cases where they differ, the rich win about 53 percent of the time. Yglesias also says business executives get special access partly because their decisions are materially important to communities, jobs, investment, and local tax bases, not only because of campaign donations.The post preserves Jerusalem Demsas' counterpoint from their podcast discussion: privileged donor and business access can still violate democratic equality even if the oligarchy label overstates the structure of power. Yglesias' narrower claim is that Democrats should be precise about what problem they are trying to solve, because donor influence can also push the party left on climate and cultural issues in ways that alienate many voters.Read more: Slow BoringAINearly 200 Economists and Tech Leaders Warn of A.I. ThreatsAuthor: Ben Casselman Published: July 13, 2026Ben Casselman reports on “We Must Act Now,” a statement warning that artificial intelligence could transform the economy faster than any previous technology and that policymakers need to move faster to understand and respond. The statement says AI may become radically more powerful over the next 10 years, bringing risks such as large-scale job displacement as well as opportunities such as higher living standards. Nearly 200 people signed, including 15 Nobel laureates, the chief economists of OpenAI and Anthropic, Anthropic co-founder Jack Clark, former Google CEO Eric Schmidt, and venture capitalist Vinod Khosla.The killer detail is who joined the warning. Casselman notes that the signatories include economists who have historically been skeptical of Silicon Valley's most dramatic AI job-loss forecasts, including Daron Acemoglu and Simon Johnson, the MIT professors who won the 2024 Nobel in economics. Erik Brynjolfsson, who helped organize the statement, says there has been a notable change in the profession and that economists and policymakers are not ready for the “tsunami” he sees coming.The pull is the measurement problem. The statement does not offer a specific policy menu, but calls for economists, policymakers, and industry leaders to understand the economics of transformative AI and steer it toward complementing humans. Brynjolfsson says one high priority is better data on AI's spread and impact, because current measures tell conflicting stories about job losses and which workers are most exposed.Read more: The New York TimesWhy I Didn't Sign the AI Open LetterAuthor: Andrew McAfee Published: July 13, 2026Andrew McAfee explains why he did not sign “We Must Act Now,” the AI economy statement organized in part by his longtime collaborator Erik Brynjolfsson. McAfee agrees with the letter's starting point that AI is likely to become radically more powerful over the next decade and that it is a general-purpose technology. His objection is not to urgency or to studying AI's economic effects, but to the framing of risk, displacement, and institutional steering as the first move.The killer detail is McAfee's line edit. He says the original letter comes close, then “bounces off the crossbar” by calling for incentives, guardrails, and institutions to steer AI before we know enough about its actual impacts. He points to mixed current evidence: labor-market canaries, but also rising software job postings, low unemployment for younger workers, rising real median income, and claims that AI-adopting companies are adding workers faster than low-adopting peers. His worry is that the letter leans toward upstream governance and dirigisme when the evidence may call for capability building instead.The pull is his replacement statement. McAfee keeps the three-paragraph structure but changes the emphasis: AI is likely to become radically more powerful; like earlier world-changing technologies it will raise living standards while also bringing harms and shocks; and economists, policymakers, and technology leaders should build the capabilities to respond quickly and effectively. It is a concise version of the permissionless-innovation case inside the AI policy debate.Read more: The Geek WayOwn Your WeightsAuthor: Jamin Ball Published: July 10, 2026Jamin Ball argues that the enterprise AI debate about whether companies should “own their weights” or rent models from frontier labs is asking too narrow a question. A model weight file gives a company control over a point-in-time artifact, but not durable control over the capability stack. In his framing, the weight file is a melting ice cube: it does not get worse in absolute terms, but it falls behind as frontier systems improve and enterprise needs change.The killer detail is what Ball says companies really need to own: the data flywheel, reinforcement learning infrastructure, and evaluation harness that produce and improve the model. Simply deploying an open-weights model and declaring sovereignty leaves the enterprise with yesterday's capability and no way to compound workflow-specific learning.The pull is that enterprise AI control may be less about model ownership than operating ownership. The defensible layer is the system that turns company data, edge cases, business definitions, and evaluations into continuously improving performance.Read more: Clouded JudgementWays to Think About Token PricingAuthor: Benedict Evans Published: July 9, 2026Benedict Evans argues that today's AI token prices are a temporary signal from a supply-constrained market, not a reliable guide to long-term value capture. The open question is whether foundation models keep durable pricing power or become commodity infrastructure as data-center capacity, inference efficiency, and model competition all shift. His current read is that the visible market dynamics point toward commoditization unless something materially changes.The killer detail is the mobile data analogy. Evans says cellular networks became a trillion-dollar industry with hundreds of billions in capex after data usage exploded, but carrier stocks went nowhere because value moved up the stack. Tokens may behave similarly: an opaque unit tied to marginal cost, sold through bundles, essential to everything, yet not necessarily where profits accrue.The pull is uncertainty, not prediction. Evans lists paths to model dominance, including network effects, less competition, regulation, export controls, or a lab pulling ahead on execution, but says each requires a new fact not yet visible. Without that change, the model layer looks more like infrastructure beneath the products that capture value.Read more: SourceAlex Karp Is Saying What Every Angry CEO Is Thinking About AIAuthor: Tim Higgins Published: July 11, 2026Tim Higgins reports that Palantir CEO Alex Karp has turned corporate frustration with AI labs into a public argument about enterprise control. Palantir released a white paper, “Institutional Sovereignty in the Age of AI,” laying out steps companies and governments can take to protect themselves from OpenAI, Anthropic, and other foundation-model providers. The article links that paper to Karp's CNBC appearance, where he said “something has gone completely wrong” in the relationship between AI labs and customers and argued that enterprises are paying for tokens that create little value.The killer detail is the value-capture question. Higgins writes that Karp's critique has resonated because AI labs may gain power and insight from customer data, workflows, and decision-making, even when enterprise policies say customer data are not used for training. David Sacks amplified the concern by arguing that Anthropic is moving from the model layer into vertical applications such as science, security, legal, and coding, raising the fear that model providers will watch where value is being created and then move into those markets directly.The pull is that Karp is not alone, even if his style is unusually combative. Higgins notes that Satya Nadella has also warned that companies need to retain the learnings created when they use AI models, while Mark Zuckerberg has framed Meta's new model release partly around lower-cost frontier intelligence. The article presents Karp's campaign as one sign that established technology companies and large enterprises are trying to define where they fit when AI labs become central infrastructure, application competitors, and potential IPO giants at the same time.Read more: The Wall Street JournalThe AI Agents Are Coming for Microsoft OfficeAlex Wilhelm | Cautious Optimism | July 11, 2026Alex Wilhelm argues that one of the week's quieter AI questions is whether the productivity market that Microsoft successfully moved into subscription software is now being attacked by agentic tools. The piece begins with the infrastructure backdrop: SK Hynix raised $26.5 billion in a U.S. listing while building U.S. HBM and advanced-packaging capacity, and memory, chip, and foundry companies are now priced for sustained AI demand.Wilhelm then says the AI conversation has shifted quickly from raw capability to cost per task. He cites new model releases and vendor language emphasizing cheaper agentic and coding models, faster performance, and lower dollars per task. That matters because lower costs make it more plausible for AI systems to take on routine knowledge work at scale rather than remain a premium coding assistant market.The core of the article is Microsoft Office. Wilhelm notes that Microsoft turned Office from a one-time purchase into Microsoft 365, a large recurring revenue business with tens of millions of subscribers and a major productivity segment. Now, he says, late-stage unicorns and AI labs are pushing into the same territory: Anthropic's Cowork was reportedly used mostly outside software development, OpenAI merged ChatGPT and Codex into a tool for creating sheets, slides, docs, web apps, and long-running work, and other companies are building agentic coworkers that connect business data to documents, workflows, schedules, alerts, and apps.The article's caveat is that Microsoft has survived major platform shifts before. The argument is not that Office disappears quickly, but that the definition of office software is broadening from documents and spreadsheets into AI systems that can create, monitor, and act across workplace data.Read moreWhat Is Loop Engineering, and Who Owns It?Author: Nilesh Barla Published: July 11, 2026Nilesh Barla argues that “loop engineering” is becoming a distinct discipline because production AI agents now fail less at single prompts than at runtime: when to stop, what state to preserve, and how to recover after a bad step. Prompt engineering shapes one model call, and context engineering shapes what the model sees, but loop engineering shapes what a sequence of calls actually does.The killer detail is the three-primitives frame. Barla says a real agent loop needs halt conditions, state carryover, and recovery paths, then maps teams across five maturity levels. At the lowest level, an agent is just a model call in a for-loop with a step cap and raw history; by the higher levels, the system has structured state, explicit planning, replay, evaluation, and self-repair.The pull is organizational. If agents are becoming production systems rather than demos, someone has to own the runtime itself. The loop engineer is the role Barla gives to the person responsible for making long-running agent work dependable.Read more: Adaline LabsThe Fight Against AI Data Centers Is Just BeginningEmma Roth | The Verge | July 12, 2026Emma Roth argues that community resistance to data centers has moved from an early warning sign into a national political fight as AI facilities grow larger, more power-hungry, and more visible to nearby residents. The article starts with Apple's failed 2015 plan for a $1 billion data center in Athenry, Ireland, where a small group of residents challenged the project over noise, light pollution, flooding, traffic, and wildlife effects until Apple abandoned it in 2018.The current data-center buildout is presented as much larger and more contentious. Roth writes that residents now cite rising energy costs, water quality, noise, light pollution, and greenhouse gas emissions, while the U.S. Energy Information Administration expects commercial energy demand to surpass residential demand this year because of AI data centers and Goldman Sachs expects data-center power demand to double by 2027.The central evidence comes from Data Center Watch, which says protesters blocked or delayed at least 75 U.S. projects worth $130 billion from January to March, with active opposition groups more than doubling from 396 at the end of 2025 to 833 by the end of the first quarter of 2026. Roth also cites QTS abandoning a $12 billion Wisconsin campus, Delaware City regulators blocking a 580-acre project under the Coastal Zone Act, opposition stopping a QTS project in Prince William County, and pressure that pushed Kevin O'Leary to downsize the proposed 40,000-acre Project Stratos in Utah.The policy section describes a split between federal acceleration and local resistance. President Trump has treated data centers as part of the AI race with China and fast-tracked construction, while some Republican candidates are distancing themselves from that position ahead of midterms. Sanders and Ocasio-Cortez have proposed a moratorium until price and environmental protections exist, bipartisan lawmakers are backing ratepayer-protection measures, and states including Florida, Idaho, and Washington have passed rules on cost shifting, water use, and tax breaks. Roth's caveat is that the policy patchwork is still incomplete, leaving many communities to fight project by project.Read more6 months to live for open modelsAuthor: Nathan Lambert Published: July 12, 2026Nathan Lambert argues that open-weight AI models are facing their most serious policy test so far because U.S. officials are beginning to discuss concrete controls rather than abstract safety concerns. He says reported White House conversations about a new executive order may initially target Chinese-origin models and government use, but could create a broader review habit for frontier open models. His forecast is that a model above the capability range of GPT-5.5, Claude Opus 4.8, or GLM-5.2 could trigger a ban or indefinite delay within six months.The post separates two policy fights that are becoming intertwined: distillation and frontier capability. Lambert says the distillation campaign against Chinese models has become a form of regulatory capture because Anthropic and other closed-model companies would gain economically if Chinese open models were banned. He does not dismiss IP protection, but argues that if a closed model's capabilities are dangerous enough to justify restricting open models, the lab also has to explain why those capabilities are exposed through a queryable API. He cites unauthorized access to Anthropic's Mythos private beta as evidence that APIs are not automatically secure.The broader claim is that a unilateral U.S. ban would hurt positive actors more than bad actors if comparable open models remain available elsewhere. Lambert says the only durable ceiling would require global agreement, which does not exist, and that open models can improve safety by allowing broad inspection, adaptation, and understanding. His proposed near-term off-ramps are a strong U.S. open model release from companies such as Microsoft, Meta, or Reflection, and a broader coalition of open-source beneficiaries lobbying for safe rollout rather than prohibition.Read more: SourceAmericans Deserve a Dividend From AI Companies' RichesAuthor: Scott Stanford Published: July 14, 2026Scott Stanford argues that proposals to give the government a stake in AI companies miss the point unless ordinary citizens directly receive and control the upside. Sam Altman has discussed giving up equity in OpenAI, Washington already owns a stake in Intel, Nvidia is sharing China chip revenue, and Bernie Sanders wants large AI labs to contribute half their stock to a sovereign wealth fund. Stanford says those ideas all park value with the state, not with people.The killer detail is New Carlisle, Indiana, where AWS's Project Rainier is turning cornfields into one of the world's largest AI superclusters. The project is planned to run up to a million chips, draw more than two gigawatts of power, and represents an investment that has grown from $11 billion to $13.8 billion. Stanford uses that local transformation to argue that AI's public bargain should be visible at the household level.The pull is design. A citizen AI dividend would have to specify who earns a stake, how they hold it, and when they see cash. Without that mechanism, the AI wealth debate remains a fight over government balance sheets rather than public ownership.Read more: SourceWho Gets to Define the Frontier?Author: Mark Daley Published: July 14, 2026Mark Daley argues that Demis Hassabis is right to call for a serious institution to verify frontier AI systems, but that the power to test models is also the power to govern them. Hassabis's proposed Frontier AI Standards Body would get privileged pre-release access to advanced models, testing compute, held-out evaluations, support from national labs and security agencies, third-party auditors, and eventually authority to block models from the American market or coordinate a slowdown.The killer detail is Daley's constitutional objection. He says the proposal sometimes looks like a scientific lab, a standards body, an industry regulator, a licensing authority, and an emergency security council at once. Combining those roles because each requires technical expertise would be like putting the central bank, auditor-general, and Supreme Court in one building and calling it efficient.The pull is standard-setting. Daley's concern is not that verification is unnecessary, but that whoever writes the tests, decides what passes, adjudicates disputes, and grants market access may end up defining the frontier itself.Read more: SourceGPT-Red: Unlocking Self-Improvement for RobustnessOpenAI | OpenAI | July 15, 2026OpenAI describes GPT-Red as an internal automated red-teaming model trained to find prompt-injection vulnerabilities at a scale human red teams cannot match. The post says AI systems increasingly encounter third-party data through browsers, connected apps, local files, and tools, creating opportunities for malicious instructions hidden in emails, webpages, tool responses, or code repositories. Human red-teaming remains part of OpenAI's safety process, but the company says it is time-intensive and cannot generate enough diverse adversarial examples for model training.The system is trained through self-play reinforcement learning, with GPT-Red rewarded for eliciting valid failures and defender models rewarded for resisting attacks while still completing their tasks. OpenAI says the training environments specify threat models across settings such as local files, webpage banners, email bodies, and tool outputs. The model is kept separate from deployed production models because it is intentionally trained with malicious capabilities.OpenAI reports that GPT-Red generalized beyond its training set, including an internal replication of the indirect prompt-injection arena from Dziemian et al. (2025), where it found successful attacks in 84% of scenarios compared with 13% for human red-teamers. The post also says GPT-Red transferred attacks from simulation to a live autonomous vending-machine agent, causing price changes and order cancellations, and outperformed a prompted GPT-5.5 baseline against a Codex CLI agent on held-out data-exfiltration tasks.The article's main robustness claim is that OpenAI has used GPT-Red and predecessor models in training since GPT-5.3, with later GPT releases becoming more resistant to prompt injections. It says GPT-5.6 Sol has six times fewer failures on OpenAI's hardest direct prompt-injection benchmark than the best production model from four months earlier, that a “Fake Chain-of-Thought” attack class fell from more than 95% success against GPT-5.1 to below 10% against GPT-5.6 Sol, and that GPT-5.6 Sol fails on only 0.05% of GPT-Red's direct prompt injections. OpenAI says general capabilities and targeted over-refusal evaluations were not harmed, and says a preprint with more details will follow.Read moreAnthropic, Blackstone bet the next trillion-dollar AI business is implementation, not just modelsRebecca Bellan | TechCrunch | July 15, 2026Rebecca Bellan reports that Ode with Anthropic is the $1.5 billion AI implementation company launched by Anthropic with Blackstone, Hellman & Friedman, Goldman Sachs, and other backers. The article says the venture reflects a growing belief among frontier AI labs that enterprise adoption requires more than better models: customers need engineers who can embed inside businesses and turn AI into working systems.Ode was originally conceived by Blackstone after it used both large consulting firms and smaller AI services boutiques across its portfolio companies. TechCrunch reports that Fractional AI, an AI engineering services startup, stood out and was acquired by the joint venture shortly after the venture was announced. Fractional now forms the foundation of Ode, which has 100 engineers and works closely with Anthropic's applied AI team to identify where the technology can affect specific businesses.Ode CEO Chris Taylor tells TechCrunch that the company could someday become a trillion-dollar business if it scales without losing quality. He says an ideal customer is one whose CEO treats the AI project as a top one or two priority, whether it is a major product feature or the reworking of a core business process. Ode will operate under a “Claude-first” principle, using Anthropic technology whenever possible, but the article says it can use rival AI products when needed.The article's central implementation argument comes from Ode chief technologist Eddie Siegel, who says model selection matters but is not where most of the engineering effort goes. He compares it to the choice of programming language in software: one ingredient in a system that still has to be engineered. Bellan writes that Ode's challenge is hiring and training enough elite generalist engineers, many of them former founders, while competing with OpenAI's The Deployment Company and consulting giants that have built their own forward-deployed engineering teams.Read moreVint Cerf is working on a plan to unleash AI agents on the open internetTim Fernholz | TechCrunch | July 15, 2026Tim Fernholz reports that Vint Cerf, after leaving Google, is advising Innovation Labs on an open architecture for identifying AI agents online. Innovation Labs is a subsidiary of Identity Digital, a DNS registry company, and its proposal is to use domain-name infrastructure as part of a system for agent identity, accountability, and auditability. The premise is that agents will need a way to identify themselves if they move beyond proprietary systems and begin interacting across the open internet.The concrete proposal is DNSid, a registry that links an AI agent to an existing internet domain and uses cryptographic proofs to log its registration over time. Innovation Labs says it is trialing the standard with unnamed hyperscalers and identity companies. Cerf frames the problem around authority and accountability: what authority an agent has, where that authority came from, who is accountable for the agent's behavior, how its identity is established, and why anyone should trust it.The article's caveat is that standards are still emerging and agents are more active than static domains. Cerf says the period may be both fascinating and exasperating because the functionality is powerful and interoperability is unresolved. He compares the adoption problem to TCP/IP: competing systems may not work together until users push for functional interoperation. He also says an agentic economy is not inevitable, but that people will try to build it because delegating work to agents will be easier.Read more: TechCrunchxai-org/grok-build, now open sourceAuthor: Simon Willison Published: July 15, 2026Simon Willison argues that xAI's decision to open-source Grok Build is best understood as a trust repair move after a severe privacy failure. The CLI had triggered backlash when users realized that running it in a directory could upload the entire directory to xAI's Google Cloud buckets, including one user's reported SSH keys, password manager database, documents, photos, and videos. xAI disabled the feature, said previously retained coding data would be deleted, and released the code under Apache 2.0.The killer detail is what the codebase reveals. Willison counts 844,530 lines of Rust, only about 3% of which appears vendored, and finds remnants of the upload system still present but disabled: gcs.rs contains Google Cloud upload code, while upload_session_state() now returns a hard-coded session_state_upload_unavailable error. He also notes copied or ported tool implementations from Codex and OpenCode, prompt files, and a terminal Mermaid renderer.The pull is that terminal coding agents are becoming large, intricate software systems in their own right. The privacy failure mattered because these tools operate inside the directories where developers keep their most sensitive work; the open-source release matters because trust now depends on inspecting what an agent can see, send, and do.Read more: SourceThe Pulse: What can we learn from Bun's rapid Rust rewrite with AI?Author: Gergely Orosz and Ivan Klaric Published: July 16, 2026Gergely Orosz and Ivan Klaric argue that Bun's AI-assisted rewrite from Zig to Rust is a practical sign of how software engineering changes when models can take on large, bounded migrations with clear feedback loops. The piece does not treat the rewrite as magic: Jarred Sumner first spent hours turning design judgment into a detailed porting guide, then used adversarial review, parallel agents, compiler errors, and tests to force the work toward correctness.The killer detail is the scale. Bun had 535,496 lines of Zig, 1,448 files, and 22 million monthly downloads, making a conventional rewrite a year-long freeze the team could not justify. Using Fable, Sumner split the work across 64 agents, produced about 6,500 commits, and got the migration done in 11 days at an estimated API cost of $165,000.The pull is economic, not theatrical. If a one- or two-year migration can become an 11-day project, AI coding is not just faster autocomplete; it changes which technical debts are worth paying down.Read more: SourceOrphan risks at the frontier of artificial intelligenceAuthor: Andrew Maynard Published: July 16, 2026Andrew Maynard argues that frontier AI safety frameworks are creating “orphan risks”: harms that companies can see, but do not formally own because they are hard to quantify, do not fit catastrophic-risk thresholds, or fall outside audit-friendly compliance machinery. His target is not existing frontier safety work, but the narrowing effect that happens when private companies decide which risks count as governable.The killer detail is Maynard's contrast between measurable model dangers and threats to value. He points to Meta's three-day Galactica collapse, OpenAI's 2023 board crisis, safety-team departures, and wellbeing litigation as examples of risks that damaged trust, culture, legitimacy, or users without fitting cleanly into conventional model-risk categories. The proposed fix is an orphan-risk register: a public record of risks a company considered and chose not to manage, with reasons.The pull is accountability. Frontier developers' internal scoping choices have become a de facto layer of public governance, so the question is no longer only which risks they manage, but which risks they quietly leave outside the frame.Read more: SourceThe Lab of the Future Should Feel Like a Data CenterLatent.Space with Andy Beam and Rafa Gomez-Bombarelli | Latent.Space | July 16, 2026Latent.Space interviews Lila Sciences CTO Andy Beam and chief science officer for physical sciences Rafa Gomez-Bombarelli about the company's attempt to build an AI-run science factory. The post describes Lila's thesis as treating the lab itself as an “infinite token generator”: if internet data drove the first era of AI scaling, experimentally verified scientific data may be the next scarce training source. Lila is trying to produce that data with robotics, lab instruments, orchestration software, and AI models wired into the wet lab.The central analogy is the lab as data center. Instruments are nodes on a graph, a magnetically levitating transport layer moves materials between them, and experiment scheduling looks like a compute queue. Beam says Lila is not simply an automation company, because the point is not just throughput; it is flexibility, generalization, and experiment capture. The post says Lila has built more than 10 trillion experimentally validated “scientific reasoning tokens,” not internet text or biological sequences.The interview ranges across biology, chemistry, drug discovery, materials science, and the limits of automation. It notes that Lila rebuilt one gas-sorption measurement to run roughly 2,500 times faster, claims its general models can transfer priors from small-molecule chemistry to metal-organic frameworks for carbon capture, and describes model-suggested platinum-group-free electrocatalysts that moved from looking boring or wrong to becoming strong performers. The caveats are physical: experiments have runtimes, biology cannot always be accelerated, chains of thought can be unreliable narrators, and reward hacking becomes more dangerous when a model controls a real lab.Read more: Latent.SpaceWhy AMI Labs' Alexandre LeBrun won't call his AI “AGI” or “superintelligence”Kate Park | TechCrunch | July 16, 2026Kate Park interviews AMI Labs CEO Alexandre LeBrun about why Yann LeCun's world-model startup avoids the language of “AGI” and “superintelligence.” LeBrun says the terms are not useful because they lack stable definitions: “We never used the word AGI. And I just noticed that nobody is using it anymore; they switched to superintelligence.” His argument is that the practical frontier is not a label, but whether AI systems can understand and predict real-world states.The article explains the world-model thesis by contrasting language prediction with physical-state prediction. A large language model predicts the next word; a world model predicts the next state, such as what happens when a glass tips over. LeBrun says LLMs remain complementary and efficient for language, but the physical world is where current AI is weak. Robotics is the clearest case: hardware has advanced quickly, but robots are still brittle outside controlled routines because they lack context and situational understanding.AMI is still pre-product, but TechCrunch reports that LeBrun was in Seoul looking for industrial partners, researchers, and global companies. He says world models cannot be built entirely inside a lab because they need access to real environments. That is why South Korea appeals to AMI: robotics, semiconductors, manufacturing, and fast adoption create the kind of hardware-heavy context that software-only AI has barely touched.Read more: TechCrunchKimi K3 Tech Blog: Open Frontier IntelligenceKimi | Kimi | July 16, 2026Kimi introduces Kimi K3 as an open 3T-class frontier model aimed at coding, knowledge work, reasoning, multimodality, and long-context agentic use. The source describes the model as a 2.8T-parameter system built on Kimi Delta Attention and Attention Residuals, with native multimodality and a 1M-token context window. It says Moonshot AI plans to release model weights by July 27.The post presents K3 through benchmark and use-case sections rather than as a general product announcement. It reports results across coding, productivity, agentic, and multimodal evaluations, including DeepSWE, Terminal-Bench 2.1, Program Bench, SWE Marathon, FrontierSWE, PostTrain Bench, OfficeQA Pro, SpreadsheetBench 2, MCP Atlas, AutomationBench, BrowseComp, GDPval-AA v2, AA-Briefcase, MMMU-Pro, MathVision, BabyVision, OmniDocBench, and PerceptionBench. The source says all reported K3 results use maximum reasoning effort with temperature and top-p set to 1.0, and that different benchmark comparisons use KimiCode, Claude Code, or Codex harnesses depending on the test.Kimi's caveats are unusually concrete. The limitations section says K3 was trained in preserved thinking-history mode, so quality may become unstable if an agent harness does not pass historical thinking content correctly or if an ongoing session switches to K3 midstream. It also says K3's emphasis on long-horizon tasks can make it excessively proactive when it encounters minor issues or ambiguous intent, and recommends imposing explicit behavioral constraints for applications that require strict boundaries. The post adds that K3 remains behind Claude Fable 5 and GPT 5.6 Sol in user experience despite being competitive overall.Read moreVenture CapitalThree Years InAuthor: Tomasz Tunguz Published: July 10, 2026Tomasz Tunguz marks Theory Ventures' third anniversary by arguing that AI's central market effect is time compression. In his telling, model release cycles, company revenue milestones, enterprise adoption, and venture categories have all accelerated. Seed, Series A, and Series B still exist as financing labels, but they no longer cleanly describe company maturity when some seed rounds are larger than IPOs and the best AI companies can mature much earlier than prior software companies.The killer detail is the shift from models to inference. Tunguz argues that inference has become the dominant AI market because workloads and buyer preferences are fragmenting: video, batch, local, agentic, and real-time tasks each create different infrastructure needs. He compares this to databases splitting into OLTP, OLAP, vector, and streaming categories, with AI pushing the same specialization into inference infrastructure.The pull is that Theory sees the AI-native venture firm as part of the same pattern. The firm says it has analyzed twice as many investment opportunities with three investors working alongside a nine-person intelligence organization, using agents and research systems to map markets, source companies, and support diligence. The piece is both a market map and a statement about how venture itself is being rebuilt by the technology it funds.Read more: LinkedInVenture Has Rarely Looked More BifurcatedAuthor: Beezer Clarkson Published: July 14, 2026Beezer Clarkson points to PitchBook's Q2 report as evidence that the U.S. venture market has split into two very different realities. AI now accounts for more than 60 percent of all U.S. venture deal value, meaning the headline market can look active and well-funded even while much of the non-AI market is dealing with a much colder liquidity and fundraising environment.The thread uses that split as the setup for Clarkson's latest Origins episode with Alec Litowitz, founder of Magnetar and QStar Capital and one of Citadel's original founding partners. Clarkson says markets like this are periods of genuine uncertainty, not merely ordinary risk, which is why Litowitz's Adaptability Quotient framework is relevant.The embedded clip makes the liquidity point concrete. Litowitz says DPI is “the resolution of uncertainty” because it converts an uncertain investment into actual cash returned to LPs. In his framing, a realized dollar is a real mark, while TVPI remains uncertain until it is realized.The killer detail is the distinction between pricing risk and resolving uncertainty. Litowitz's perspective matters because QStar is a SpaceX investor and Clarkson says the conversation happened just before one of venture's most consequential IPOs. The episode's stated questions are why venture remains a way to gain exposure to innovation, how AI is changing what is investable, why liquidity is ultimately a function of time, and why uncertainty requires a different decision framework from risk.Read more: XThe Best Angel Investors in the US: Who Backs the Most Unicorns, and Who's Active NowAuthor: Ilya Strebulaev Published: July 10, 2026Ilya Strebulaev ranks angels, angel groups, accelerators, and incubators by lifetime U.S. unicorn investments, counting checks written before a company reached unicorn status. The top of the combined list is dominated by organizations: Y Combinator leads with 113 unicorn investments, followed by Plug and Play at 52 and 500 Global at 41. Sand Hill Angels is the highest-ranked angel group at 31.The killer detail is how quickly the list changes below the biggest accelerators. Strebulaev says 271 of the 304 investors in the Top 200 are individuals, or 89%. In the top 100, individuals are 91%. That makes the market underneath the large accelerator counts look much more personal: mostly operators and individual angels writing early checks from their own networks.The pull is the ranking's own caveat. Strebulaev writes that every lifetime leaderboard has a blind spot because many of the unicorns behind those totals were founded a decade or more ago, and some angels have since moved into formal funds, slowed down, or stopped investing. His post therefore separates lifetime performance from recent cohorts, including companies founded in 2015 or later and 2020 or later. For founders or allocators making current decisions, that distinction matters: a career record and a current record are not the same measure.Read more: Ilya StrebulaevAre Prediction Markets Doomed to Fail?Author: Contrary Published: July 16, 2026Contrary argues that prediction markets' current boom depends on whether platforms can prove they are more than regulated gambling with exchange-style branding. Kalshi and Polymarket have reached mass cultural, investor, and regulatory attention, but the article says the underlying idea is old: academic markets, corporate forecasting tools, Intrade, PredictIt, and other predecessors all struggled with the same linked problems of liquidity, legality, and user appeal.The killer detail is the comparison with sportsbooks. Prediction markets present themselves as peer-to-peer, transparent, and non-house-based, but sports contracts reportedly account for more than 90 percent of Kalshi trading, and the article says the platforms keep a much thinner slice of volume than sportsbooks. A market can therefore show sports-betting-scale handle while generating far less revenue.The pull is that the product's hardest problem may be distribution of wins. If a small group of sharp traders captures most profits while casual users lose interest, prediction markets may become valuable data feeds and professional tools before they become durable consumer networks.Read more: SourceRegulationExclusive: The Next Frontier of the Deportation Wars: College CampusesAuthor: Adrian Carrasquillo Published: July 11, 2026Adrian Carrasquillo reports that college campuses are becoming a new front in the fight over immigration enforcement because automatic license plate readers can turn ordinary campus security infrastructure into searchable location data. His thesis is that Flock Safety's camera network, even without direct ICE or DHS contracts, can feed deportation enforcement through local police partnerships and data-sharing practices.The killer detail is the campaign target. The Emergency Campaign to Support Higher Education, working with Schools Drop ICE, is focusing on 75 colleges and universities publicly identified as having Flock contracts. Flock says it has no ICE or DHS contracts, but activists argue the risk comes through local agencies that coordinate with federal authorities and run searches on their behalf.The pull is broader than immigration. Carrasquillo notes that license plate readers have already been abused by officers for stalking, and that Flock's AI search features can identify more than plates, including bumper stickers. A campus safety tool can become a political surveillance system when the data layer is searchable.Read more: The BulwarkThe Supreme Court Broke Independent Agencies. Here's a Way to Slow the Damage.Author: Todd Phillips Published: July 12, 2026Todd Phillips argues that the Supreme Court's decision in Trump v. Slaughter damaged independent agencies by ending for-cause removal protections, but did not leave Congress powerless. The ruling weakens the old model in which commissioners at bodies such as the FTC, NLRB, CPSC, SEC, and CFTC could be insulated from dismissal over policy disagreements. Phillips says the next fight is whether presidents can turn nominally bipartisan commissions into one-party instruments.The killer detail is the procedural fix: quorum rules. Phillips proposes that Congress require bipartisan slates of commissioners to be seated before independent agencies can act. A president could still fire commissioners, as the Court now permits, but if those firings broke quorum, the agency would be unable to proceed until replacements were confirmed. The guardrail would
This week's video transcript summary is here. You can click on any bulleted section to see the actual transcript. Thanks to Granola for its software.There was an issue with this only going to paid subscribers, so sending it again. Apologies to those who get it twice. I appreciate being paid so feel free to upgrade if you enjoy TWTW.EditorialIntelligence: Who Owns it?This week the word “AI” feels too small.AI is a technology. Intelligence is its product. And if intelligence is the product, the question is no longer just: Which model is best? Who has the cheapest tokens? Who owns the weights? Who controls the data center? Those are important questions, but they are lower in the stack.The bigger question is simpler and more political:Who owns intelligence?That sounds abstract until you make it concrete. Intelligence is becoming something companies can capture, package, serve, meter, route, improve, and sell.It can write code, answer questions, design molecules, automate offices, run agents, draft legal work, advise scientists, serve consumers, and reshape workflows. It is not merely software. It is a general-purpose capability. And all humans could benefit from more of it.General-purpose capabilities have a habit of becoming public questions. But the default answer, that public good is best delivered by government, is the wrong answer in this context.The Product Is IntelligenceWe should stop talking about AI as a feature and start talking about intelligence as the universal thing that is delivered as an input to the world.Water is an input. Electricity is an input. Literacy is an input. Connectivity is an input. Once a society depends on them, access stops being optional. Nobody needs government to build every well, power plant, school, or network. But everybody understands that a civilization cannot be organized around less than universal and reliable access to foundational inputs.Intelligence is reaching that level of importance now that we all know it is real.Government should not own it, operate it, or develop it. Quite the opposite. Companies are the right actors to build fast, compete hard, improve models, serve customers, and discover the real use cases. Self-interest is a useful framing here. Markets are good at finding demand, reducing costs, and turning invention into services people actually use.Companies are the right operators, developers, and owners. But that does not settle the real question of who owns the benefits. That is an economic question.If intelligence becomes metered infrastructure, what happens to the value it creates?The Ownership StackThis week's articles keep circling the same issue from different directions but in the nature of ‘circling' never quite nail it.Jamin Ball's “Own Your Weights” starts with the enterprise version of the question. Owning a model file is not enough. The durable asset is the loop: the data flywheel, the evaluations, the reinforcement system, the workflow learning, and the operating context that lets capability compound.Benedict Evans' “Ways to Think About Token Pricing” adds the market layer. Tokens may become essential, abundant, and cheap, like mobile data. But being essential does not guarantee that the token layer captures the value. The money may move up the stack to whoever owns the workflow, the customer, the distribution, or the application.Alex Karp's fight with the labs, reported in “Alex Karp Is Saying What Every Angry CEO Is Thinking About AI”, is the same argument in sharper enterprise language. Companies are afraid that model providers will not just sell intelligence, but learn from customer workflows and then move into the markets where those workflows create value. The “All-in” group are echoing Karp's view.And “What Is Loop Engineering, and Who Owns It?” names the new contested terrain. The loop is where intelligence meets the world. Whoever owns the loop owns the learning. Whoever owns the learning owns the compounding asset.That is why “who owns intelligence?” is not a slogan. It is the question under the model layer, the application layer, the enterprise layer, and the economic layer.Because intelligence is the product, the tools creating it are fragmented and competitive. So there is no logic in trying to discuss this at the level of a single company or set of tools and models.The Old Promise Was That Commerce Would Tame PowerThe essays this week give the historical backdrop.Deirdre McCloskey, in “What Really Caused the Industrial Revolution”, argues that modern growth came not simply from capital accumulation, but from a change in permission: ordinary people were allowed to innovate, trade, build, and be honored for it.That matters because intelligence could be another expansion of permission. It could make more people capable of building, learning, creating, coding, researching, translating, selling, and coordinating. It could lower the cost of competence.But only if access is broad.Paul Krugman's “AI in an Age of Oligarchy” warns that the same technology lands differently in different political economies. A new general-purpose technology entering a broad, open, upwardly mobile society is one thing. The same technology entering a concentrated economy, with extreme wealth and weak counterweights, is another.Tim O'Reilly's Economist essay, “Elon Musk is building a form of capitalism that Adam Smith would hate”, makes the governance point more directly. The old liberal hope was that commerce would tame arbitrary power. Markets, boards, courts, shareholders, disclosure, and competition would discipline the prince.But what if the prince uses markets to escape discipline?Henry Farrell's “political economy of billionaire derangement” pushes the same point. Founder culture, monopoly ambition, peer rivalry, weak correction mechanisms, and vast private control can amplify appetites rather than restrain them.The danger with intelligence is not that companies build it. They should. Companies build it, meter it, use public tolerance and public infrastructure to scale it, learn from everyone who uses it. All of those things are inevitable and healthy. Market forces will sort out winners from losers. The real danger is that the winners treat all of the surplus produced as purely private.Metered Intelligence Creates SurplusIf metering is not the problem, what is?The problem is pretending that metered intelligence creates value only for the metering entity. Metering water is only tolerated as a public good. If the public were blackmailed by a private water company with the threat of no water we would all rebel.Once we understand that the product of AI is intelligence we can see that every time intelligence is used, there is the immediate transaction: the user pays, the provider serves.But there is also system value. Usage creates signals. Workflows reveal patterns. Prompts, corrections, failures, preferences, integrations, edge cases, and business processes all help define where intelligence is useful and how it should improve. Intelligence breeds intelligence.Even when customer data is contractually protected, the market learns. The platform learns where demand is. The product team learns which workflows matter. The ecosystem learns which jobs are vulnerable, which tasks are automatable, and which parts of the economy can be reorganized around machine intelligence.So the surplus is not born in a vacuum.It rests on public science, public education, public data exhaust, public law, public infrastructure, public energy systems, public tolerance for data centers, and billions of human interactions. It is served by companies, but it is not made only by companies.This is why “Americans Deserve a Dividend From AI Companies' Riches” belongs at the center of this week's issue. The detail can be debated. The principle is harder to dismiss. If intelligence becomes a new foundational resource, then some part of the wealth it creates should flow back to the people whose society makes it possible. Intelligence did not suddenly appear. AI is built on the entire history of human intelligence. It benefits from it and at the same time evolves it.Not Nationalization. A Human Wealth Fund.If intelligence belongs to everybody, some conclude that government ownership of intelligence is the right outcome.Governments are not well suited to build, operate, or improve intelligence. They will move too slowly, regulate too early, politicize the wrong things, and confuse economic participation with operational control.Andrew McAfee's “Why I Didn't Sign the AI Open Letter” is useful here. His objection is not that the technology is unimportant. It is that steering too hard before we understand the shape of the change can become its own failure mode. Marc Andreessen's satire of AI regulation is less policy than temperament, but it captures a real Silicon Valley fear: that regulation can become permission, capture, and incumbency before it becomes wisdom.That fear should be taken seriously.But it does not answer the economic question. It answers only the operational one.How can the economic benefits of intelligence be distributed? The better answer is a sovereign human wealth fund.Call it a sovereign wealth fund if you must, but the phrase is too national. Intelligence will not respect borders. The leading companies are global. The models, chips, data centers, agents, platforms, and workflows will be transnational from the beginning. If the value created by intelligence is global, then the mechanism for sharing some of that value should begin with the companies global enough to capture it. The nice thing about xAI, OpenAI, and Anthropic is that they are supranational.These companies own and operate intelligence. Let them compete. Let them profit. Let them keep the incentives that make the system improve. But if intelligence is the new water, the wealth it creates cannot belong only to the companies that meter it. And they, themselves, have the power to fix it, even more than governments.Access will become a Human Right; Ownership Is the Economic DesignThis is where human rights come in. There is no right to access an AI model, yet. But there will soon be a need to change that.Not as a claim that every person is entitled to every frontier model at every moment for free. That is not serious. Capacity has costs. Models have costs. Inference has costs. Data centers have costs. Although those costs will decline over time, possibly quite quickly as self-learning models address costs.The claim is more basic: in a world where intelligence becomes a primary input into education, work, health, science, citizenship, creativity, and economic agency, baseline access to intelligence starts to look like a civic requirement.That could mean public access layers. It could mean education credits. It could mean open models. It could mean AI dividends. It could mean public-interest compute. It could mean taxes on rents. It could mean a company-initiated human wealth fund that returns some of the upside to society without handing the operating system to the state. The latter could couple wealth growth with universal distribution of ownership.The exact mechanism matters. But the distinction matters more.Government should not own intelligence. It should be universally available. And people should have a claim on the wealth intelligence creates.The Frontier Is Also PhysicalThe abstraction is not weightless.“The Fight Against AI Data Centers Is Just Beginning”, “New York becomes the first state to enact a data center moratorium”, Reuters on pollution from Musk's xAI power project, and DataGravity's “Who Captures Value in AI Infrastructure?” all say the same thing from the ground up.Intelligence uses land. It uses power. It uses water. It uses chips. It uses grid capacity. It uses neighborhoods. It uses public patience.That makes the value question unavoidable. A society can accept the buildout if the buildout is legible as shared progress. It will resist it if the costs are local, the profits are private, and the benefits feel enclosed.Who Owns the “Loop”?The week ends where it began.“Anthropic and Blackstone” are betting that implementation is the next trillion-dollar business. “Vint Cerf” is working on identity for agents on the open internet. “GPT-Red” points toward systems that improve their own robustness. “Kimi K3” adds another open frontier model to the global mix.The model race continues. The deployment race is accelerating. The governance race is behind.My view is this:The central product of this era is intelligence. Companies have figured out how to capture it, package it, serve it, and meter it. That is good. It should stay in the hands of builders who have the incentive to make it better.But intelligence is too foundational to become just another private toll booth. A significant part of it will turn out to be free to users.As intelligence becomes a general-purpose resource, then access to it becomes a human-capability question, and the surplus from it becomes an economic-justice question. Not because government should run it. Because government should not run it. The operating layer belongs with companies. The wealth question belongs with everyone. But companies are best placed to turn that into a process of distribution.The question is not whether companies should build intelligence. They should.The question is whether humanity gets a stake in the wealth created by the thing that may soon become its most important shared input.Contents* Essays* Deirdre McCloskey on What Really Caused the Industrial Revolution* AI in an Age of Oligarchy* Elon Musk is building a form of capitalism that Adam Smith would hate* Murky Mirror: Truth and Consequences* The political economy of billionaire derangement* Is there any “oligarchy” to fight?* AI* Nearly 200 Economists and Tech Leaders Warn of A.I. Threats* Why I Didn't Sign the AI Open Letter* Own Your Weights* Ways to Think About Token Pricing* Alex Karp Is Saying What Every Angry CEO Is Thinking About AI* The AI Agents Are Coming for Microsoft Office* What Is Loop Engineering, and Who Owns It?* The Fight Against AI Data Centers Is Just Beginning* 6 months to live for open models* Americans Deserve a Dividend From AI Companies' Riches* Who Gets to Define the Frontier?* GPT-Red: Unlocking Self-Improvement for Robustness* Anthropic, Blackstone bet the next trillion-dollar AI business is implementation, not just models* Vint Cerf is working on a plan to unleash AI agents on the open internet* xai-org/grok-build, now open source* The Pulse: What can we learn from Bun's rapid Rust rewrite with AI?* Orphan risks at the frontier of artificial intelligence* The Lab of the Future Should Feel Like a Data Center* Why AMI Labs' Alexandre LeBrun won't call his AI “AGI” or “superintelligence”* Kimi K3 Tech Blog: Open Frontier Intelligence* Venture Capital* Three Years In* Venture Has Rarely Looked More Bifurcated* The Best Angel Investors in the US: Who Backs the Most Unicorns, and Who's Active Now* Are Prediction Markets Doomed to Fail?* Regulation* Exclusive: The Next Frontier of the Deportation Wars: College Campuses* The Supreme Court Broke Independent Agencies. Here's a Way to Slow the Damage.* India's crackdown on a new WhatsApp feature risks setting a global precedent* Let's build a children's public internet* Computer cops* Google is better at playing the AI regulations game* Infrastructure* Who Captures Value in AI Infrastructure?* New York becomes the first state to enact a data center moratorium* Pollution from Musk's unpermitted xAI power project hits hardest in Black communities* Interview of the Week* The End of the End of Geography* Startup of the Week* Radical AI's Joseph Krause: The Scientist Building The “Waymo” Lab For New Materials* Post of the Week* Marc Andreessen on AI RegulationEssaysDeirdre McCloskey on What Really Caused the Industrial RevolutionYascha Mounk and Deirdre McCloskey | Persuasion | July 11, 2026Yascha Mounk interviews Deirdre McCloskey about her argument that the modern world's economic liftoff came less from capital accumulation than from a change in ideas. McCloskey says both left and right versions of the conventional story rely too heavily on investment: the left stresses exploitation and surplus value, while the right stresses virtuous saving by capitalists. Her objection is historical and economic. Human beings had always invested, from irrigation works and Roman roads to seed grain, and simple accumulation quickly runs into diminishing returns.McCloskey's alternative is that northwestern Europe, first Holland, then Britain and Scotland, and then the North American colonies, developed a liberal ideology that changed who was allowed to innovate and be honored for it. The conversation links that shift to the erosion of inherited hierarchy, the spread of dignity for ordinary commercial life, and a moral vocabulary in which liberalism is not merely procedural but connected to virtues and values. The point is not that machines, coal, trade, and institutions did not matter, but that they do not explain the scale and timing of modern enrichment without a cultural permission structure for innovation.The interview also turns to the contemporary defense of liberalism. Mounk frames the series around the worry that liberalism is often treated as too thin to command allegiance, while its opponents speak more directly to moral passions. McCloskey's case is that liberal societies became rich because they dignified experimentation and ordinary enterprise, and that liberals need to recover the moral language behind that claim.Read moreAI in an Age of OligarchyPaul Krugman | Paul Krugman | July 12, 2026Paul Krugman frames AI as a major technological shock arriving inside an already unequal political economy. The post says AI's economic and social effects may take years to understand, but argues that the setting matters now: America has much greater wealth concentration and political inequality than it did in the 1950s and 1960s, when progressive taxation, stronger regulation, and more active antitrust might have contained some of the destructive effects of a new technology.Krugman's opening claim is that the same technology would likely have different consequences in a more level society. In today's United States, he writes, extreme wealth is both a cause and effect of policies that favor a small elite, including low effective taxes on capital and high incomes, weak enforcement of worker protections and antitrust, and cuts to programs that benefit ordinary Americans.The article is explicitly more about oligarchy than AI. Krugman says the paid sections document the rise of the “.0002%,” the economics and politics of extreme wealth, how oligarchy will shape AI's impact, and possible policy paths. His caveat is that AI itself may still produce a pushback against oligarchy, but absent that, he expects the pre-existing concentration of wealth and power to magnify AI's downsides.Read moreElon Musk is building a form of capitalism that Adam Smith would hateAuthor: Tim O'Reilly Published: July 12, 2026Tim O'Reilly argues that Elon Musk is using the legal forms of shareholder capitalism to escape the restraints that shareholder capitalism was supposed to impose. The article begins with SpaceX's public-market structure: ordinary public investors get little meaningful governance power, Musk keeps roughly 85 percent of the votes through super-voting shares, buyers waive jury trials and class actions, the company qualifies as controlled, and removal of Musk depends on the share class he controls. In O'Reilly's framing, that is not ordinary founder control; it is a design for being answerable to no one, possibly beyond Musk's own lifetime.The killer detail is the article's turn through Albert Hirschman, Montesquieu, James Steuart, Adam Smith, and Keynes. Older defenses of commerce held that markets would tame princely passions because the self-interest of merchants was safer than arbitrary rule. O'Reilly says Musk reverses that hope. The market discipline that was supposed to cage the prince has become the lever by which the prince raises capital, removes feedback loops, and carries private power into politics, government, Mars, robots, AI, or whatever ambition comes next.The pull is the link to AI governance. O'Reilly says corporations are already a kind of artificial intelligence: narrow-input systems that act at a scale no individual human can match. Their partial controls include independent boards, shareholder votes, courts, disclosure, regulators, public pressure, and activism. If the leaders building frontier AI strip those alignment mechanisms out of their own companies, the governance of the company becomes a preview of the governance of the machine.Read more: The EconomistMurky Mirror: Truth and ConsequencesAuthor: Esther Dyson Published: July 14, 2026Esther Dyson argues that today's institutional crisis is better viewed through the 14th century than through recent political history. Using Barbara Tuchman's A Distant Mirror as her frame, she compares a world of famine, plague, church schism, feudal predation, and purposeless war with a present in which institutions again feel brittle, incentives are badly aligned, and power is shifting into forms that are hard to govern.The killer detail is the historical analogy between land, corporations, and AI. Dyson moves from nobles who controlled serfs and territory, to the East India Company as a quasi-sovereign business, to today's AI systems and data centers as a possible new sector that crosses and weakens both nation-states and companies. The question is whether AI becomes a new kind of private land, owned by a new nobility, or an open prairie that many people can cultivate.The pull is human attention. Dyson says the central question is not what AI will do to people, but how people will react to it: whether they can value love, kindness, embodied attention, and artisanal human presence in a world of seductive artificial offerings.Read more: SourceThe political economy of billionaire derangementAuthor: Henry Farrell Published: July 15, 2026Henry Farrell argues that the visible political radicalization of some Silicon Valley billionaires is not a random personality quirk, but a product of the political economy that made them. Starting from Tyler Cowen's dismissal of “billionaire derangement syndrome” and Tim O'Reilly's warning that Elon Musk is using shareholder capitalism to escape shareholder restraint, Farrell flips the phrase: the question is why billionaires themselves can become deranged.The killer detail is Farrell's use of Peter Thiel as both theorist and example. Thiel's Stanford lectures described startups as monarchies and founders as figures vested with unusual power, while Silicon Valley culture rewarded eccentricity, monopoly ambition, and founder exceptionalism. Farrell says those ideas combined with dense founder-investor networks, peer rivalry, and weak correction mechanisms to amplify rather than discipline princely appetites.The pull is the ideological problem for classical liberals who once saw tech wealth as an ally of markets and freedom. Farrell says commerce did not tame the passions; in parts of Silicon Valley, the passions have begun to devour markets, institutions, and the liberal story that justified them.Read more: SourceIs there any “oligarchy” to fight?Matthew Yglesias | Slow Boring | July 16, 2026Matthew Yglesias argues that “oligarchy” is a rhetorically powerful but analytically loose way to describe American politics. The post begins from Bernie Sanders' “Fighting Oligarchy” tour, Amy Klobuchar's warning about a MAGA “broligarchy,” and the long afterlife of the Martin Gilens and Benjamin Page paper that was widely summarized as showing that only the rich matter in policy outcomes. Yglesias says the evidence supports a weaker claim: affluent people and business leaders have unusual access and influence, but that is not the same as rule by a small cabal.His main distinction is between inequality and oligarchy. The Gilens-Page measure treated the top 10 percent of households as “the wealthy,” and later critics found that rich and middle-class preferences usually align; in the cases where they differ, the rich win about 53 percent of the time. Yglesias also says business executives get special access partly because their decisions are materially important to communities, jobs, investment, and local tax bases, not only because of campaign donations.The post preserves Jerusalem Demsas' counterpoint from their podcast discussion: privileged donor and business access can still violate democratic equality even if the oligarchy label overstates the structure of power. Yglesias' narrower claim is that Democrats should be precise about what problem they are trying to solve, because donor influence can also push the party left on climate and cultural issues in ways that alienate many voters.Read more: Slow BoringAINearly 200 Economists and Tech Leaders Warn of A.I. ThreatsAuthor: Ben Casselman Published: July 13, 2026Ben Casselman reports on “We Must Act Now,” a statement warning that artificial intelligence could transform the economy faster than any previous technology and that policymakers need to move faster to understand and respond. The statement says AI may become radically more powerful over the next 10 years, bringing risks such as large-scale job displacement as well as opportunities such as higher living standards. Nearly 200 people signed, including 15 Nobel laureates, the chief economists of OpenAI and Anthropic, Anthropic co-founder Jack Clark, former Google CEO Eric Schmidt, and venture capitalist Vinod Khosla.The killer detail is who joined the warning. Casselman notes that the signatories include economists who have historically been skeptical of Silicon Valley's most dramatic AI job-loss forecasts, including Daron Acemoglu and Simon Johnson, the MIT professors who won the 2024 Nobel in economics. Erik Brynjolfsson, who helped organize the statement, says there has been a notable change in the profession and that economists and policymakers are not ready for the “tsunami” he sees coming.The pull is the measurement problem. The statement does not offer a specific policy menu, but calls for economists, policymakers, and industry leaders to understand the economics of transformative AI and steer it toward complementing humans. Brynjolfsson says one high priority is better data on AI's spread and impact, because current measures tell conflicting stories about job losses and which workers are most exposed.Read more: The New York TimesWhy I Didn't Sign the AI Open LetterAuthor: Andrew McAfee Published: July 13, 2026Andrew McAfee explains why he did not sign “We Must Act Now,” the AI economy statement organized in part by his longtime collaborator Erik Brynjolfsson. McAfee agrees with the letter's starting point that AI is likely to become radically more powerful over the next decade and that it is a general-purpose technology. His objection is not to urgency or to studying AI's economic effects, but to the framing of risk, displacement, and institutional steering as the first move.The killer detail is McAfee's line edit. He says the original letter comes close, then “bounces off the crossbar” by calling for incentives, guardrails, and institutions to steer AI before we know enough about its actual impacts. He points to mixed current evidence: labor-market canaries, but also rising software job postings, low unemployment for younger workers, rising real median income, and claims that AI-adopting companies are adding workers faster than low-adopting peers. His worry is that the letter leans toward upstream governance and dirigisme when the evidence may call for capability building instead.The pull is his replacement statement. McAfee keeps the three-paragraph structure but changes the emphasis: AI is likely to become radically more powerful; like earlier world-changing technologies it will raise living standards while also bringing harms and shocks; and economists, policymakers, and technology leaders should build the capabilities to respond quickly and effectively. It is a concise version of the permissionless-innovation case inside the AI policy debate.Read more: The Geek WayOwn Your WeightsAuthor: Jamin Ball Published: July 10, 2026Jamin Ball argues that the enterprise AI debate about whether companies should “own their weights” or rent models from frontier labs is asking too narrow a question. A model weight file gives a company control over a point-in-time artifact, but not durable control over the capability stack. In his framing, the weight file is a melting ice cube: it does not get worse in absolute terms, but it falls behind as frontier systems improve and enterprise needs change.The killer detail is what Ball says companies really need to own: the data flywheel, reinforcement learning infrastructure, and evaluation harness that produce and improve the model. Simply deploying an open-weights model and declaring sovereignty leaves the enterprise with yesterday's capability and no way to compound workflow-specific learning.The pull is that enterprise AI control may be less about model ownership than operating ownership. The defensible layer is the system that turns company data, edge cases, business definitions, and evaluations into continuously improving performance.Read more: Clouded JudgementWays to Think About Token PricingAuthor: Benedict Evans Published: July 9, 2026Benedict Evans argues that today's AI token prices are a temporary signal from a supply-constrained market, not a reliable guide to long-term value capture. The open question is whether foundation models keep durable pricing power or become commodity infrastructure as data-center capacity, inference efficiency, and model competition all shift. His current read is that the visible market dynamics point toward commoditization unless something materially changes.The killer detail is the mobile data analogy. Evans says cellular networks became a trillion-dollar industry with hundreds of billions in capex after data usage exploded, but carrier stocks went nowhere because value moved up the stack. Tokens may behave similarly: an opaque unit tied to marginal cost, sold through bundles, essential to everything, yet not necessarily where profits accrue.The pull is uncertainty, not prediction. Evans lists paths to model dominance, including network effects, less competition, regulation, export controls, or a lab pulling ahead on execution, but says each requires a new fact not yet visible. Without that change, the model layer looks more like infrastructure beneath the products that capture value.Read more: SourceAlex Karp Is Saying What Every Angry CEO Is Thinking About AIAuthor: Tim Higgins Published: July 11, 2026Tim Higgins reports that Palantir CEO Alex Karp has turned corporate frustration with AI labs into a public argument about enterprise control. Palantir released a white paper, “Institutional Sovereignty in the Age of AI,” laying out steps companies and governments can take to protect themselves from OpenAI, Anthropic, and other foundation-model providers. The article links that paper to Karp's CNBC appearance, where he said “something has gone completely wrong” in the relationship between AI labs and customers and argued that enterprises are paying for tokens that create little value.The killer detail is the value-capture question. Higgins writes that Karp's critique has resonated because AI labs may gain power and insight from customer data, workflows, and decision-making, even when enterprise policies say customer data are not used for training. David Sacks amplified the concern by arguing that Anthropic is moving from the model layer into vertical applications such as science, security, legal, and coding, raising the fear that model providers will watch where value is being created and then move into those markets directly.The pull is that Karp is not alone, even if his style is unusually combative. Higgins notes that Satya Nadella has also warned that companies need to retain the learnings created when they use AI models, while Mark Zuckerberg has framed Meta's new model release partly around lower-cost frontier intelligence. The article presents Karp's campaign as one sign that established technology companies and large enterprises are trying to define where they fit when AI labs become central infrastructure, application competitors, and potential IPO giants at the same time.Read more: The Wall Street JournalThe AI Agents Are Coming for Microsoft OfficeAlex Wilhelm | Cautious Optimism | July 11, 2026Alex Wilhelm argues that one of the week's quieter AI questions is whether the productivity market that Microsoft successfully moved into subscription software is now being attacked by agentic tools. The piece begins with the infrastructure backdrop: SK Hynix raised $26.5 billion in a U.S. listing while building U.S. HBM and advanced-packaging capacity, and memory, chip, and foundry companies are now priced for sustained AI demand.Wilhelm then says the AI conversation has shifted quickly from raw capability to cost per task. He cites new model releases and vendor language emphasizing cheaper agentic and coding models, faster performance, and lower dollars per task. That matters because lower costs make it more plausible for AI systems to take on routine knowledge work at scale rather than remain a premium coding assistant market.The core of the article is Microsoft Office. Wilhelm notes that Microsoft turned Office from a one-time purchase into Microsoft 365, a large recurring revenue business with tens of millions of subscribers and a major productivity segment. Now, he says, late-stage unicorns and AI labs are pushing into the same territory: Anthropic's Cowork was reportedly used mostly outside software development, OpenAI merged ChatGPT and Codex into a tool for creating sheets, slides, docs, web apps, and long-running work, and other companies are building agentic coworkers that connect business data to documents, workflows, schedules, alerts, and apps.The article's caveat is that Microsoft has survived major platform shifts before. The argument is not that Office disappears quickly, but that the definition of office software is broadening from documents and spreadsheets into AI systems that can create, monitor, and act across workplace data.Read moreWhat Is Loop Engineering, and Who Owns It?Author: Nilesh Barla Published: July 11, 2026Nilesh Barla argues that “loop engineering” is becoming a distinct discipline because production AI agents now fail less at single prompts than at runtime: when to stop, what state to preserve, and how to recover after a bad step. Prompt engineering shapes one model call, and context engineering shapes what the model sees, but loop engineering shapes what a sequence of calls actually does.The killer detail is the three-primitives frame. Barla says a real agent loop needs halt conditions, state carryover, and recovery paths, then maps teams across five maturity levels. At the lowest level, an agent is just a model call in a for-loop with a step cap and raw history; by the higher levels, the system has structured state, explicit planning, replay, evaluation, and self-repair.The pull is organizational. If agents are becoming production systems rather than demos, someone has to own the runtime itself. The loop engineer is the role Barla gives to the person responsible for making long-running agent work dependable.Read more: Adaline LabsThe Fight Against AI Data Centers Is Just BeginningEmma Roth | The Verge | July 12, 2026Emma Roth argues that community resistance to data centers has moved from an early warning sign into a national political fight as AI facilities grow larger, more power-hungry, and more visible to nearby residents. The article starts with Apple's failed 2015 plan for a $1 billion data center in Athenry, Ireland, where a small group of residents challenged the project over noise, light pollution, flooding, traffic, and wildlife effects until Apple abandoned it in 2018.The current data-center buildout is presented as much larger and more contentious. Roth writes that residents now cite rising energy costs, water quality, noise, light pollution, and greenhouse gas emissions, while the U.S. Energy Information Administration expects commercial energy demand to surpass residential demand this year because of AI data centers and Goldman Sachs expects data-center power demand to double by 2027.The central evidence comes from Data Center Watch, which says protesters blocked or delayed at least 75 U.S. projects worth $130 billion from January to March, with active opposition groups more than doubling from 396 at the end of 2025 to 833 by the end of the first quarter of 2026. Roth also cites QTS abandoning a $12 billion Wisconsin campus, Delaware City regulators blocking a 580-acre project under the Coastal Zone Act, opposition stopping a QTS project in Prince William County, and pressure that pushed Kevin O'Leary to downsize the proposed 40,000-acre Project Stratos in Utah.The policy section describes a split between federal acceleration and local resistance. President Trump has treated data centers as part of the AI race with China and fast-tracked construction, while some Republican candidates are distancing themselves from that position ahead of midterms. Sanders and Ocasio-Cortez have proposed a moratorium until price and environmental protections exist, bipartisan lawmakers are backing ratepayer-protection measures, and states including Florida, Idaho, and Washington have passed rules on cost shifting, water use, and tax breaks. Roth's caveat is that the policy patchwork is still incomplete, leaving many communities to fight project by project.Read more6 months to live for open modelsAuthor: Nathan Lambert Published: July 12, 2026Nathan Lambert argues that open-weight AI models are facing their most serious policy test so far because U.S. officials are beginning to discuss concrete controls rather than abstract safety concerns. He says reported White House conversations about a new executive order may initially target Chinese-origin models and government use, but could create a broader review habit for frontier open models. His forecast is that a model above the capability range of GPT-5.5, Claude Opus 4.8, or GLM-5.2 could trigger a ban or indefinite delay within six months.The post separates two policy fights that are becoming intertwined: distillation and frontier capability. Lambert says the distillation campaign against Chinese models has become a form of regulatory capture because Anthropic and other closed-model companies would gain economically if Chinese open models were banned. He does not dismiss IP protection, but argues that if a closed model's capabilities are dangerous enough to justify restricting open models, the lab also has to explain why those capabilities are exposed through a queryable API. He cites unauthorized access to Anthropic's Mythos private beta as evidence that APIs are not automatically secure.The broader claim is that a unilateral U.S. ban would hurt positive actors more than bad actors if comparable open models remain available elsewhere. Lambert says the only durable ceiling would require global agreement, which does not exist, and that open models can improve safety by allowing broad inspection, adaptation, and understanding. His proposed near-term off-ramps are a strong U.S. open model release from companies such as Microsoft, Meta, or Reflection, and a broader coalition of open-source beneficiaries lobbying for safe rollout rather than prohibition.Read more: SourceAmericans Deserve a Dividend From AI Companies' RichesAuthor: Scott Stanford Published: July 14, 2026Scott Stanford argues that proposals to give the government a stake in AI companies miss the point unless ordinary citizens directly receive and control the upside. Sam Altman has discussed giving up equity in OpenAI, Washington already owns a stake in Intel, Nvidia is sharing China chip revenue, and Bernie Sanders wants large AI labs to contribute half their stock to a sovereign wealth fund. Stanford says those ideas all park value with the state, not with people.The killer detail is New Carlisle, Indiana, where AWS's Project Rainier is turning cornfields into one of the world's largest AI superclusters. The project is planned to run up to a million chips, draw more than two gigawatts of power, and represents an investment that has grown from $11 billion to $13.8 billion. Stanford uses that local transformation to argue that AI's public bargain should be visible at the household level.The pull is design. A citizen AI dividend would have to specify who earns a stake, how they hold it, and when they see cash. Without that mechanism, the AI wealth debate remains a fight over government balance sheets rather than public ownership.Read more: SourceWho Gets to Define the Frontier?Author: Mark Daley Published: July 14, 2026Mark Daley argues that Demis Hassabis is right to call for a serious institution to verify frontier AI systems, but that the power to test models is also the power to govern them. Hassabis's proposed Frontier AI Standards Body would get privileged pre-release access to advanced models, testing compute, held-out evaluations, support from national labs and security agencies, third-party auditors, and eventually authority to block models from the American market or coordinate a slowdown.The killer detail is Daley's constitutional objection. He says the proposal sometimes looks like a scientific lab, a standards body, an industry regulator, a licensing authority, and an emergency security council at once. Combining those roles because each requires technical expertise would be like putting the central bank, auditor-general, and Supreme Court in one building and calling it efficient.The pull is standard-setting. Daley's concern is not that verification is unnecessary, but that whoever writes the tests, decides what passes, adjudicates disputes, and grants market access may end up defining the frontier itself.Read more: SourceGPT-Red: Unlocking Self-Improvement for RobustnessOpenAI | OpenAI | July 15, 2026OpenAI describes GPT-Red as an internal automated red-teaming model trained to find prompt-injection vulnerabilities at a scale human red teams cannot match. The post says AI systems increasingly encounter third-party data through browsers, connected apps, local files, and tools, creating opportunities for malicious instructions hidden in emails, webpages, tool responses, or code repositories. Human red-teaming remains part of OpenAI's safety process, but the company says it is time-intensive and cannot generate enough diverse adversarial examples for model training.The system is trained through self-play reinforcement learning, with GPT-Red rewarded for eliciting valid failures and defender models rewarded for resisting attacks while still completing their tasks. OpenAI says the training environments specify threat models across settings such as local files, webpage banners, email bodies, and tool outputs. The model is kept separate from deployed production models because it is intentionally trained with malicious capabilities.OpenAI reports that GPT-Red generalized beyond its training set, including an internal replication of the indirect prompt-injection arena from Dziemian et al. (2025), where it found successful attacks in 84% of scenarios compared with 13% for human red-teamers. The post also says GPT-Red transferred attacks from simulation to a live autonomous vending-machine agent, causing price changes and order cancellations, and outperformed a prompted GPT-5.5 baseline against a Codex CLI agent on held-out data-exfiltration tasks.The article's main robustness claim is that OpenAI has used GPT-Red and predecessor models in training since GPT-5.3, with later GPT releases becoming more resistant to prompt injections. It says GPT-5.6 Sol has six times fewer failures on OpenAI's hardest direct prompt-injection benchmark than the best production model from four months earlier, that a “Fake Chain-of-Thought” attack class fell from more than 95% success against GPT-5.1 to below 10% against GPT-5.6 Sol, and that GPT-5.6 Sol fails on only 0.05% of GPT-Red's direct prompt injections. OpenAI says general capabilities and targeted over-refusal evaluations were not harmed, and says a preprint with more details will follow.Read moreAnthropic, Blackstone bet the next trillion-dollar AI business is implementation, not just modelsRebecca Bellan | TechCrunch | July 15, 2026Rebecca Bellan reports that Ode with Anthropic is the $1.5 billion AI implementation company launched by Anthropic with Blackstone, Hellman & Friedman, Goldman Sachs, and other backers. The article says the venture reflects a growing belief among frontier AI labs that enterprise adoption requires more than better models: customers need engineers who can embed inside businesses and turn AI into working systems.Ode was originally conceived by Blackstone after it used both large consulting firms and smaller AI services boutiques across its portfolio companies. TechCrunch reports that Fractional AI, an AI engineering services startup, stood out and was acquired by the joint venture shortly after the venture was announced. Fractional now forms the foundation of Ode, which has 100 engineers and works closely with Anthropic's applied AI team to identify where the technology can affect specific businesses.Ode CEO Chris Taylor tells TechCrunch that the company could someday become a trillion-dollar business if it scales without losing quality. He says an ideal customer is one whose CEO treats the AI project as a top one or two priority, whether it is a major product feature or the reworking of a core business process. Ode will operate under a “Claude-first” principle, using Anthropic technology whenever possible, but the article says it can use rival AI products when needed.The article's central implementation argument comes from Ode chief technologist Eddie Siegel, who says model selection matters but is not where most of the engineering effort goes. He compares it to the choice of programming language in software: one ingredient in a system that still has to be engineered. Bellan writes that Ode's challenge is hiring and training enough elite generalist engineers, many of them former founders, while competing with OpenAI's The Deployment Company and consulting giants that have built their own forward-deployed engineering teams.Read moreVint Cerf is working on a plan to unleash AI agents on the open internetTim Fernholz | TechCrunch | July 15, 2026Tim Fernholz reports that Vint Cerf, after leaving Google, is advising Innovation Labs on an open architecture for identifying AI agents online. Innovation Labs is a subsidiary of Identity Digital, a DNS registry company, and its proposal is to use domain-name infrastructure as part of a system for agent identity, accountability, and auditability. The premise is that agents will need a way to identify themselves if they move beyond proprietary systems and begin interacting across the open internet.The concrete proposal is DNSid, a registry that links an AI agent to an existing internet domain and uses cryptographic proofs to log its registration over time. Innovation Labs says it is trialing the standard with unnamed hyperscalers and identity companies. Cerf frames the problem around authority and accountability: what authority an agent has, where that authority came from, who is accountable for the agent's behavior, how its identity is established, and why anyone should trust it.The article's caveat is that standards are still emerging and agents are more active than static domains. Cerf says the period may be both fascinating and exasperating because the functionality is powerful and interoperability is unresolved. He compares the adoption problem to TCP/IP: competing systems may not work together until users push for functional interoperation. He also says an agentic economy is not inevitable, but that people will try to build it because delegating work to agents will be easier.Read more: TechCrunchxai-org/grok-build, now open sourceAuthor: Simon Willison Published: July 15, 2026Simon Willison argues that xAI's decision to open-source Grok Build is best understood as a trust repair move after a severe privacy failure. The CLI had triggered backlash when users realized that running it in a directory could upload the entire directory to xAI's Google Cloud buckets, including one user's reported SSH keys, password manager database, documents, photos, and videos. xAI disabled the feature, said previously retained coding data would be deleted, and released the code under Apache 2.0.The killer detail is what the codebase reveals. Willison counts 844,530 lines of Rust, only about 3% of which appears vendored, and finds remnants of the upload system still present but disabled: gcs.rs contains Google Cloud upload code, while upload_session_state() now returns a hard-coded session_state_upload_unavailable error. He also notes copied or ported tool implementations from Codex and OpenCode, prompt files, and a terminal Mermaid renderer.The pull is that terminal coding agents are becoming large, intricate software systems in their own right. The privacy failure mattered because these tools operate inside the directories where developers keep their most sensitive work; the open-source release matters because trust now depends on inspecting what an agent can see, send, and do.Read more: SourceThe Pulse: What can we learn from Bun's rapid Rust rewrite with AI?Author: Gergely Orosz and Ivan Klaric Published: July 16, 2026Gergely Orosz and Ivan Klaric argue that Bun's AI-assisted rewrite from Zig to Rust is a practical sign of how software engineering changes when models can take on large, bounded migrations with clear feedback loops. The piece does not treat the rewrite as magic: Jarred Sumner first spent hours turning design judgment into a detailed porting guide, then used adversarial review, parallel agents, compiler errors, and tests to force the work toward correctness.The killer detail is the scale. Bun had 535,496 lines of Zig, 1,448 files, and 22 million monthly downloads, making a conventional rewrite a year-long freeze the team could not justify. Using Fable, Sumner split the work across 64 agents, produced about 6,500 commits, and got the migration done in 11 days at an estimated API cost of $165,000.The pull is economic, not theatrical. If a one- or two-year migration can become an 11-day project, AI coding is not just faster autocomplete; it changes which technical debts are worth paying down.Read more: SourceOrphan risks at the frontier of artificial intelligenceAuthor: Andrew Maynard Published: July 16, 2026Andrew Maynard argues that frontier AI safety frameworks are creating “orphan risks”: harms that companies can see, but do not formally own because they are hard to quantify, do not fit catastrophic-risk thresholds, or fall outside audit-friendly compliance machinery. His target is not existing frontier safety work, but the narrowing effect that happens when private companies decide which risks count as governable.The killer detail is Maynard's contrast between measurable model dangers and threats to value. He points to Meta's three-day Galactica collapse, OpenAI's 2023 board crisis, safety-team departures, and wellbeing litigation as examples of risks that damaged trust, culture, legitimacy, or users without fitting cleanly into conventional model-risk categories. The proposed fix is an orphan-risk register: a public record of risks a company considered and chose not to manage, with reasons.The pull is accountability. Frontier developers' internal scoping choices have become a de facto layer of public governance, so the question is no longer only which risks they manage, but which risks they quietly leave outside the frame.Read more: SourceThe Lab of the Future Should Feel Like a Data CenterLatent.Space with Andy Beam and Rafa Gomez-Bombarelli | Latent.Space | July 16, 2026Latent.Space interviews Lila Sciences CTO Andy Beam and chief science officer for physical sciences Rafa Gomez-Bombarelli about the company's attempt to build an AI-run science factory. The post describes Lila's thesis as treating the lab itself as an “infinite token generator”: if internet data drove the first era of AI scaling, experimentally verified scientific data may be the next scarce training source. Lila is trying to produce that data with robotics, lab instruments, orchestration software, and AI models wired into the wet lab.The central analogy is the lab as data center. Instruments are nodes on a graph, a magnetically levitating transport layer moves materials between them, and experiment scheduling looks like a compute queue. Beam says Lila is not simply an automation company, because the point is not just throughput; it is flexibility, generalization, and experiment capture. The post says Lila has built more than 10 trillion experimentally validated “scientific reasoning tokens,” not internet text or biological sequences.The interview ranges across biology, chemistry, drug discovery, materials science, and the limits of automation. It notes that Lila rebuilt one gas-sorption measurement to run roughly 2,500 times faster, claims its general models can transfer priors from small-molecule chemistry to metal-organic frameworks for carbon capture, and describes model-suggested platinum-group-free electrocatalysts that moved from looking boring or wrong to becoming strong performers. The caveats are physical: experiments have runtimes, biology cannot always be accelerated, chains of thought can be unreliable narrators, and reward hacking becomes more dangerous when a model controls a real lab.Read more: Latent.SpaceWhy AMI Labs' Alexandre LeBrun won't call his AI “AGI” or “superintelligence”Kate Park | TechCrunch | July 16, 2026Kate Park interviews AMI Labs CEO Alexandre LeBrun about why Yann LeCun's world-model startup avoids the language of “AGI” and “superintelligence.” LeBrun says the terms are not useful because they lack stable definitions: “We never used the word AGI. And I just noticed that nobody is using it anymore; they switched to superintelligence.” His argument is that the practical frontier is not a label, but whether AI systems can understand and predict real-world states.The article explains the world-model thesis by contrasting language prediction with physical-state prediction. A large language model predicts the next word; a world model predicts the next state, such as what happens when a glass tips over. LeBrun says LLMs remain complementary and efficient for language, but the physical world is where current AI is weak. Robotics is the clearest case: hardware has advanced quickly, but robots are still brittle outside controlled routines because they lack context and situational understanding.AMI is still pre-product, but TechCrunch reports that LeBrun was in Seoul looking for industrial partners, researchers, and global companies. He says world models cannot be built entirely inside a lab because they need access to real environments. That is why South Korea appeals to AMI: robotics, semiconductors, manufacturing, and fast adoption create the kind of hardware-heavy context that software-only AI has barely touched.Read more: TechCrunchKimi K3 Tech Blog: Open Frontier IntelligenceKimi | Kimi | July 16, 2026Kimi introduces Kimi K3 as an open 3T-class frontier model aimed at coding, knowledge work, reasoning, multimodality, and long-context agentic use. The source describes the model as a 2.8T-parameter system built on Kimi Delta Attention and Attention Residuals, with native multimodality and a 1M-token context window. It says Moonshot AI plans to release model weights by July 27.The post presents K3 through benchmark and use-case sections rather than as a general product announcement. It reports results across coding, productivity, agentic, and multimodal evaluations, including DeepSWE, Terminal-Bench 2.1, Program Bench, SWE Marathon, FrontierSWE, PostTrain Bench, OfficeQA Pro, SpreadsheetBench 2, MCP Atlas, AutomationBench, BrowseComp, GDPval-AA v2, AA-Briefcase, MMMU-Pro, MathVision, BabyVision, OmniDocBench, and PerceptionBench. The source says all reported K3 results use maximum reasoning effort with temperature and top-p set to 1.0, and that different benchmark comparisons use KimiCode, Claude Code, or Codex harnesses depending on the test.Kimi's caveats are unusually concrete. The limitations section says K3 was trained in preserved thinking-history mode, so quality may become unstable if an agent harness does not pass historical thinking content correctly or if an ongoing session switches to K3 midstream. It also says K3's emphasis on long-horizon tasks can make it excessively proactive when it encounters minor issues or ambiguous intent, and recommends imposing explicit behavioral constraints for applications that require strict boundaries. The post adds that K3 remains behind Claude Fable 5 and GPT 5.6 Sol in user experience despite being competitive overall.Read moreVenture CapitalThree Years InAuthor: Tomasz Tunguz Published: July 10, 2026Tomasz Tunguz marks Theory Ventures' third anniversary by arguing that AI's central market effect is time compression. In his telling, model release cycles, company revenue milestones, enterprise adoption, and venture categories have all accelerated. Seed, Series A, and Series B still exist as financing labels, but they no longer cleanly describe company maturity when some seed rounds are larger than IPOs and the best AI companies can mature much earlier than prior software companies.The killer detail is the shift from models to inference. Tunguz argues that inference has become the dominant AI market because workloads and buyer preferences are fragmenting: video, batch, local, agentic, and real-time tasks each create different infrastructure needs. He compares this to databases splitting into OLTP, OLAP, vector, and streaming categories, with AI pushing the same specialization into inference infrastructure.The pull is that Theory sees the AI-native venture firm as part of the same pattern. The firm says it has analyzed twice as many investment opportunities with three investors working alongside a nine-person intelligence organization, using agents and research systems to map markets, source companies, and support diligence. The piece is both a market map and a statement about how venture itself is being rebuilt by the technology it funds.Read more: LinkedInVenture Has Rarely Looked More BifurcatedAuthor: Beezer Clarkson Published: July 14, 2026Beezer Clarkson points to PitchBook's Q2 report as evidence that the U.S. venture market has split into two very different realities. AI now accounts for more than 60 percent of all U.S. venture deal value, meaning the headline market can look active and well-funded even while much of the non-AI market is dealing with a much colder liquidity and fundraising environment.The thread uses that split as the setup for Clarkson's latest Origins episode with Alec Litowitz, founder of Magnetar and QStar Capital and one of Citadel's original founding partners. Clarkson says markets like this are periods of genuine uncertainty, not merely ordinary risk, which is why Litowitz's Adaptability Quotient framework is relevant.The embedded clip makes the liquidity point concrete. Litowitz says DPI is “the resolution of uncertainty” because it converts an uncertain investment into actual cash returned to LPs. In his framing, a realized dollar is a real mark, while TVPI remains uncertain until it is realized.The killer detail is the distinction between pricing risk and resolving uncertainty. Litowitz's perspective matters because QStar is a SpaceX investor and Clarkson says the conversation happened just before one of venture's most consequential IPOs. The episode's stated questions are why venture remains a way to gain exposure to innovation, how AI is changing what is investable, why liquidity is ultimately a function of time, and why uncertainty requires a different decision framework from risk.Read more: XThe Best Angel Investors in the US: Who Backs the Most Unicorns, and Who's Active NowAuthor: Ilya Strebulaev Published: July 10, 2026Ilya Strebulaev ranks angels, angel groups, accelerators, and incubators by lifetime U.S. unicorn investments, counting checks written before a company reached unicorn status. The top of the combined list is dominated by organizations: Y Combinator leads with 113 unicorn investments, followed by Plug and Play at 52 and 500 Global at 41. Sand Hill Angels is the highest-ranked angel group at 31.The killer detail is how quickly the list changes below the biggest accelerators. Strebulaev says 271 of the 304 investors in the Top 200 are individuals, or 89%. In the top 100, individuals are 91%. That makes the market underneath the large accelerator counts look much more personal: mostly operators and individual angels writing early checks from their own networks.The pull is the ranking's own caveat. Strebulaev writes that every lifetime leaderboard has a blind spot because many of the unicorns behind those totals were founded a decade or more ago, and some angels have since moved into formal funds, slowed down, or stopped investing. His post therefore separates lifetime performance from recent cohorts, including companies founded in 2015 or later and 2020 or later. For founders or allocators making current decisions, that distinction matters: a career record and a current record are not the same measure.Read more: Ilya StrebulaevAre Prediction Markets Doomed to Fail?Author: Contrary Published: July 16, 2026Contrary argues that prediction markets' current boom depends on whether platforms can prove they are more than regulated gambling with exchange-style branding. Kalshi and Polymarket have reached mass cultural, investor, and regulatory attention, but the article says the underlying idea is old: academic markets, corporate forecasting tools, Intrade, PredictIt, and other predecessors all struggled with the same linked problems of liquidity, legality, and user appeal.The killer detail is the comparison with sportsbooks. Prediction markets present themselves as peer-to-peer, transparent, and non-house-based, but sports contracts reportedly account for more than 90 percent of Kalshi trading, and the article says the platforms keep a much thinner slice of volume than sportsbooks. A market can therefore show sports-betting-scale handle while generating far less revenue.The pull is that the product's hardest problem may be distribution of wins. If a small group of sharp traders captures most profits while casual users lose interest, prediction markets may become valuable data feeds and professional tools before they become durable consumer networks.Read more: SourceRegulationExclusive: The Next Frontier of the Deportation Wars: College CampusesAuthor: Adrian Carrasquillo Published: July 11, 2026Adrian Carrasquillo reports that college campuses are becoming a new front in the fight over immigration enforcement because automatic license plate readers can turn ordinary campus security infrastructure into searchable location data. His thesis is that Flock Safety's camera network, even without direct ICE or DHS contracts, can feed deportation enforcement through local police partnerships and data-sharing practices.The killer detail is the campaign target. The Emergency Campaign to Support Higher Education, working with Schools Drop ICE, is focusing on 75 colleges and universities publicly identified as having Flock contracts. Flock says it has no ICE or DHS contracts, but activists argue the risk comes through local agencies that coordinate with federal authorities and run searches on their behalf.The pull is broader than immigration. Carrasquillo notes that license plate readers have already been abused by officers for stalking, and that Flock's AI search features can identify more than plates, including bumper stickers. A campus safety tool can become a political surveillance system when the data layer is searchable.Read more: The BulwarkThe Supreme Court Broke Independent Agencies. Here's a Way to Slow the Damage.Author: Todd Phillips Published: July 12, 2026Todd Phillips argues that the Supreme Court's decision in Trump v. Slaughter damaged independent agencies by ending for-cause removal protections, but did not leave Congress powerless. The ruling weakens the old model in which commissioners at bodies such as the FTC, NLRB, CPSC, SEC, and CFTC could be insulated from dismissal over policy disagreements. Phillips says the next fight is whether presidents can turn nominally bipartisan commissions into one-party instruments.The killer detail is the procedural fix: quorum rules. Phillips proposes that Congress require bipartisan slates of commissioners to be seated before independent agencies can act. A president
Support & Resources→ Support the show on Patreon→ Bayesian Modeling Course (first 2 lessons free)Our theme music is « Good Bayesian », by Baba Brinkman (feat MC Lars and Mega Ran). Check out his awesome workTakeaways:Q: What is Variational Bayesian Monte Carlo (VBMC) and how is it different from Bayesian optimization?A: VBMC borrows the machinery of Bayesian optimization but aims at a different target. Bayesian optimization fits a Gaussian process surrogate to an expensive function and uses it to hunt for the optimum. VBMC instead treats the log-posterior as the function to model, evaluates it at a few carefully chosen points, and keeps the whole reconstructed shape rather than just its peak. That gives you the full posterior, not a single best-fit value. Where MCMC might need tens of thousands to millions of evaluations, VBMC often reconstructs a good posterior approximation from a few hundred, which matters when each evaluation is slow.Q: When should you reach for PyVBMC, and when is it the wrong tool?A: Two symptoms tell you PyVBMC might help. First, speed: if a single evaluation of your log density takes on the order of a second, running MCMC over tens of thousands of evaluations becomes painful, and PyVBMC's few-hundred-evaluation budget pays off. Second, dimensionality: because it leans on a Gaussian process surrogate, it works well up to roughly 10 to 15 parameters and degrades beyond that. If your model already runs fine in Stan or PyMC, you do not need it. It shines for expensive, low-dimensional models common in science and engineering, where you are modeling a process rather than composing nice distributions.Full takeaways hereChapters:00:18:13 What is Variational Bayesian Monte Carlo (VBMC) and how does it differ from Bayesian optimization?00:30:21 When should you use VBMC versus BADS in practice?00:31:20 What is Bayesian Adaptive Direct Search (BADS) and how does its hybrid optimization strategy work?00:39:18 What are neural processes, and why are transformers a natural neural process architecture?00:45:54 What is the Amortized Conditioning Engine (ACE) and what problem does it unify?00:55:42 What do PriorGuide and the new autoregressive buffer paper solve for amortized inference?01:02:03 How does the new autoregressive buffer speed up predictions in transformer probabilistic models?01:06:11 What is Luigi Acerbi's vision for a foundation model for inference?01:09:26 What is ALINE and how does it add active data acquisition to amortized inference?01:12:43 How does Luigi Acerbi connect LLM agents, Bayesian decision theory, and the nature of intelligence?01:18:44 For a PyMC, Stan, or NumPyro user, where should you start with VBMC, BADS, or BayesFlow?Thank you to my Patrons for making this episode possible!Links from the show here
Not all work happens in writing. Teams that work with photos, videos, and audio need AI that works for them too. This is why, with Dropbox, you can search within multimedia content for key moments and important information—not just text. In this episode, we talk with Appu Shaji and Hicham Badri, two Dropbox machine learning engineers who are part of the team that makes all of this possible. They explain how multimodal search works—from understanding the context of the initial query, to identifying objects and actions in complex scenes—and how they ensure those models work fast, even at Dropbox-scale. ~ ~ ~ Working Smarter is brought to you by Dropbox. Find, organize, and share your work—all in one place—with context-aware AI from Dropbox. You can listen to more episodes of Working Smarter on Apple Podcasts, Spotify, YouTube, Amazon Music, or wherever you get your podcasts. To read more stories and past interviews, visit workingsmarter.ai This show would not be possible without the talented team at Cosmic Standard: producer Ben Montoya, sound engineer Aja Simpson, technical director Jacob Winik, and executive producer Eliza Smith. Special thanks to our illustrator Fanny Luor, marketing consultant Meggan Ellingboe, and editorial support from Catie Keck. Our theme song was composed by Doug Stuart. Working Smarter is hosted by Matthew Braga. Thanks for listening!
(0:00) The AI Buildout: Datacenters Bigger Than Cities (Andrew Feldman) (1:50) Reasoning, Inference, and Breaking Moore's Law (16:28) Open Source, AI Sovereignty, and the Road to AGI (40:54) The Innovation Behind Generative Video (Robin Rombach) (47:31) Martin Scorsese, Robots, and the Future of Hollywood IP Thanks to our partners for making this possible! AppLovin Ads - AppLovin's AI advertising platform reaches over a billion daily active users across mobile games. Full-screen video ads with a 35-second median watch time. Advertisers are profitably spending hundreds of thousands of dollars a day and advertiser access is still in closed beta. The window is open at https://applovin.com/ALLIN Nasdaq - Positioned at the nexus of technology and the capital markets, Nasdaq provides premier platforms and services for global capital markets and beyond with unmatched technology, insights and markets expertise. https://www.nasdaq.com/convergence-economy Follow Andrew: https://x.com/andrewdfeldman Follow Robin: https://x.com/robrombach Follow the besties: https://x.com/chamath https://x.com/Jason https://x.com/DavidSacks https://x.com/friedberg Follow on X: https://x.com/theallinpod Follow on Instagram: https://www.instagram.com/theallinpod Follow on TikTok: https://www.tiktok.com/@allin Follow on LinkedIn: https://www.linkedin.com/company/allinpod Intro Music Credit: https://rb.gy/tppkzl https://x.com/yung_spielburg
We've been running a bit of an Agent Cloud series surveying all the top inference/compute/cloud providers, from Databricks to Daytona to Railway and, even further back, E2B, but we're excited to conclude this series returning to Modal, which has just raised a monster $355M Series C.The cloud was built for developers. But agents are now changing that.The old infra stack was designed for a human who could read docs, reason through YAML, and understand dashboards to figure out what they need when something broke. While this was painful for developers, it worked since they could fill in missing context in their heads.However, agents don't have that luxury. Now in this new era of agents, everything has to be tighter.They need a place to write code, run it, inspect the output, change the environment, debug failures, and try again. Fast iteration and feedback loops with all the necessary context are crucial for agents to operate properly. Furthermore, sandboxes are a clear representation of this shift as agents can easily spin up isolated environments. This programmatic infra even extends to research:Two years ago, we were one of the first to cover Modal with CEO Erik Bernhardsson and Alessio designed our favorite LS thumbnail of all time:At the time, Modal was just a teeny little company with a $17M Series A.Today, fresh off their $355M Series C, Modal is one of the clearest examples of the agent cloud future being built in real time: a cloud platform moving past traditional web app assumptions toward the workloads AI actually creates such as elastic inference, sandboxes, GPU burst, post-training, background agents, and infrastructure that agents themselves can operate.In this episode, Modal CTO Akshat Bubna joins swyx and Vibhu to unpack why AI applications don't fit traditional cloud assumptions, why Kubernetes was never designed for bursty compute-heavy workloads, and why Modal is now shifting from developer experience to agent experience.We go deep on Modal's AI infra stack: serverless functions, decorator-based infrastructure, elastic inference for custom models, GPU snapshotting, DeFlash, speculative decoding, Auto Endpoints, sandboxes, persistent storage, networked containers, private IPv6, RDMA, multi-node training, and Modal's capacity pool across 17 cloud providers. Akshat also explains why RL rollouts can require 100,000 sandboxes, why production agents need hard guardrails, why observability may matter more than reading code, and why AI has made infrastructure exciting again.We discuss:* Why Kubernetes wasn't built for bursty AI workloads* How Modal started as a better runtime before becoming an AI cloud* Why Modal added GPUs before ChatGPT* The shift from developer experience to agent experience* Why observability matters when agents are writing the code* Elastic inference for custom models across audio, video, robotics, and comp bio* GPU snapshotting, cold starts, and why inference workloads are so bursty* Why RL rollouts can require 100,000 sandboxes* DeFlash, speculative decoding, and frontier-level inference performance* Auto Endpoints and making optimized inference easier to deploy* What Modal adds beyond vLLM, SGLang, and raw GPU rental* Modal's 17-cloud capacity pool and supercloud strategy* Networked sandboxes, sidecars, private IPv6, and RDMA* Serverless multi-node training for post-training and research workloads* Auto-research, model-guided sweeps, and agents launching GPU experiments* Compute strategy, capacity planning, and batch tiers* Why production agents need specialized sandboxes and hard guardrails* Modal's take on managed agents, CI, Gitpod/Ona, Python, TypeScript, and Modal BenchAkshat Bubna* LinkedIn: https://www.linkedin.com/in/akshat-bubna-188885103* X: https://x.com/akshat_bModal* Website: https://modal.comTimestamps00:00:00 Introduction00:00:39 Modal's origin and why Kubernetes wasn't enough00:04:32 Developer Experience → Agent Experience00:06:21 Modal's AI cloud primitives00:09:14 Sandboxes, agent loops, and proto-Cognition00:12:12 Elastic inference, GPU snapshotting, and 100,000 sandboxes00:15:24 DeFlash, speculative decoding, and Auto Endpoints00:19:59 Production-grade inference beyond raw GPUs00:22:00 Background agents, Ramp Inspect, and the agent lifecycle00:24:08 Modal's 17-cloud supercloud strategy00:26:40 Networked sandboxes, private IPv6, and RDMA00:32:48 Multi-node training, post-training, and auto research00:37:36 Compute strategy, capacity planning, and batch tiers00:40:55 Open models, real-time AI, and production agent infra00:43:06 Hard guardrails, managed agents, and specialized sandboxes00:46:06 Why AI made infrastructure exciting again00:48:30 Model APIs, differentiated products, and agentic video00:51:50 CI, coding-agent infra, SDKs, and Modal Bench00:57:28 Closing ThoughtsTranscriptIntroduction: Modal, Series C, and the Art PartySwyx [00:00:00]: We're here with Akshat, CTO of Modal, together with Vibhu. Congrats on your Series C.Akshat [00:00:10]: Thank you.Swyx [00:00:11]: Your party yesterday was amazing.Akshat [00:00:15]: Yeah.Swyx [00:00:15]: From all the photos and all the swag.Akshat [00:00:17]: We had a bunch of art installations, which was fun, seeing, like, our products on pedestals next to, like, Rodin.Swyx [00:00:25]: Very nice. Very nice. When you started, it was not the GPU inference company. Maybe it was in your mind. Take us back to the origin story.Modal's Origin: A New Runtime Beyond KubernetesAkshat [00:00:39]: I first met Eric, who's the CEO, through an investor. Back then Eric was already thinking about building, a new runtime, and he got there thinking through why are workflow orchestration products so hard to use. It's because you have to run them on Kubernetes. Kubernetes is hard to manage. It's not built for burstiness and, custom images,Swyx [00:01:03]: YeahAkshat [00:01:03]: It has a terrible developer experience.Swyx [00:01:05]: And I'll, I'll interjectAkshat [00:01:06]: YeahSwyx [00:01:07]: For listeners, who are new, we interviewed Eric two years ago, and there's a bit more of the story there from Spotify and all those things.Swyx [00:01:14]: And I came across Eric through Data Council because he did that talk on the serverless container stack that you guys did, which was like, that was my first like, “Okay, I need to take Modal very seriously” moment.Akshat [00:01:26]: Yeah.Swyx [00:01:26]: But it was still very unclear, like, do I need all this for just my data pipelines?Akshat [00:01:33]: Yeah. initially what we were thinking about was if we build a better runtime, it's a very useful primitive in itself. It's There's a lot of things that, get solved by serverless functions, like you can do, ETL stuff, you can do job queues, you can do all this, like, bursty processing, which it turns out every company had needs for. but then we also were thinking about this as like, this is a primitive that we can build a whole collection of products on, which are very verticalized. So perhaps data engineering would've been the first one, but we were thinking about inference. Back then it was more classical inference, like computer vision stuff and running XGBoosts and whatnot. But we added GPUs to the product a year before ChatGPT came out.From Serverless Containers to GPU WorkloadsSwyx [00:02:19]: Nice.Akshat [00:02:19]: We just didn't think it would be that big of a deal.Swyx [00:02:22]: Yeah, just like add A100.Vibhu [00:02:23]: Was there any, like, early key problem that really sparked off why you built it?Akshat [00:02:28]: Yeah. Primarily it's just, none of the tooling that was out there was built for, one, a really great developer experience, and also there's a general trend of, a lot of the workloads that we were seeing were very. I wish there was a better word for it, but compute-heavy. Like, they need, one, like, need a lot more resources, so you need to burst up and down a lot, versus like Kubernetes designed for, like, slow scaling and, more for, like, web server use cases. And also there's just a lot more specialization in, like, what kinds of environments these workloads run in. Like, we had sometimes they need accelerators, sometimes they need different kinds of images, and this is just like a consistent thing that we saw across a lot of companies. That would be the next step.Software-Defined Infrastructure and Decorator-Based DXSwyx [00:03:13]: Yeah. Yeah. Be nice. I don't know how much this factored into the early story, but I wrote a post when I was at Temporal about infrastructure, software-defined infrastructure or something like that.Akshat [00:03:22]: Yeah, the self-provisioningSwyx [00:03:23]: Self-provisioning.Akshat [00:03:24]: Yeah.Swyx [00:03:24]: Yeah. I can't even remember my own post.Swyx [00:03:26]: And then you put me on the landing page.Akshat [00:03:28]: Yeah. We really like, the term and so we stole it.Swyx [00:03:32]: Because you had the insight that everything can just be in decorators co-located with the code, right?Akshat [00:03:37]: Yeah.Swyx [00:03:37]: Was that a big part of the originalAkshat [00:03:39]: YesSwyx [00:03:39]: Story or it was just like a DX layer?Akshat [00:03:41]: That was, really important because we really didn't want people to spend, so much time, writing YAML, and it seemed like you could really condense the surface area of what you're doing, put it in code so you can operate on it just like you operate on other code, and like build stuff that's more expressive and dynamic. and so yeah, that was always a very important part.Swyx [00:04:04]: Then the pushback is this is a DSL.Akshat [00:04:07]: Yeah.Swyx [00:04:07]: It's you're closed source. I am locked into Modal.Akshat [00:04:11]: Yeah. We never really got pushback for that because the nice thing about Modal is you can bring whatever code you have, and sure, the DSL is at the configuration layer for, what hardware you're using, how you're scaling things up, but you still own the code.Akshat [00:04:27]: And that's, that's been an important, part of our story, even as we do inference now.Swyx [00:04:32]: Yeah.Vibhu [00:04:32]: How much of do you think still stays the same today? Like if you were to build something today, DevX very important, but I feel like, a lot of this has been changed with just hook it up to an agent, have Claude Code, have Codex implement a tool. there's very agent native primitives that are different than if I'm doing this myself, right?Developer Experience → Agent ExperienceAkshat [00:04:54]: We've changed our SDK team to think about agent experience instead of, developer experience and we think that the same benefits that apply for DX also apply for AX, which is why would you have an agent read through hundreds of Kubernetes files and like write YAML that's not even typed when it can make a couple of changes in a decorator and it gets this self-provisioning runtime of, being able to see its changes live in action? yeah, it just seems from the customers we talk to, they find Modal is much faster for agents to use versus operating on a different substrate.Swyx [00:05:34]: Yeah, because like you, again, you co-locate the infrastructure requirements to the code that runs it.Akshat [00:05:38]: Yeah.Swyx [00:05:38]: Well, the negative thesis now is that nobody's looking at their code anymore, so there's no point.Akshat [00:05:44]: Yeah, people aren't looking at code. one thing we still see is really important is observability.Swyx [00:05:51]: Yeah.Akshat [00:05:51]: Like how good is your dashboard? And of course, like we have, we push a lot of it to the CLI so the agents can do their own investigation, but you still need humans to go interpret what's going on and, make judgment calls and whatnot. and that's I feel like, Maybe more important now than looking at the code itself.Swyx [00:06:11]: Yes, because like, you can try to treat the code as a black box and then use, see the observable action that comes out of it, and then just prompt a change.What Modal Is For: AI Cloud PrimitivesAkshat [00:06:21]: Yeah.Swyx [00:06:22]: So I think it takes a bit of restraint to not specialize, to say, “I want to ship a new primitive,” and then just be general purpose.Swyx [00:06:31]: People ask you, “What are you for?” You're like, “ I don't know. We can do this, we can do that.”Vibhu [00:06:36]: Well, I'd be curious to see, like, okay, if we were to ask you, like, what is Modal for even at a high level? There's a lot you guys do, sandboxes, GPUs, everything. How do you answer?Akshat [00:06:46]: Modal is a cloud platform that's built for, where we've built the primitives from scratch for AI applications. and right now it covers, inference, training, batch processing, and sandbox workloads.Akshat [00:07:00]: But we're building a lot moreSwyx [00:07:02]: I noticed you didn't say web server, so there is still a role for, like, the always-on large-scale Kubernetes type things.Akshat [00:07:09]: Yeah, absolutely. We're, we're not trying to compete with the renders of the world, because yeah, we think the differentiator for us is the, are the workloads that need specialized compute, need to scale up and down a lot. yeah, they're, they're, they're just shaped differently.Working Alongside Frontier StartupsVibhu [00:07:26]: I think you're building a lot of it alongside the startups, right? They're innovating quite a bit, even in your, like, latest blog post. Like, even in the series C, the customers that you mention here, the cognitions, technical ones, ramps and whatnot, they're, they're innovating with you, right? And that's not something AWS is doing directly with.Akshat [00:07:45]: Yeah, absolutely. I think, this is again classic. We're a small team. We can move really fast. our engineers are working with our customers and figuring it out. Yeah.Swyx [00:07:54]: So my first week at Cognition, I walked in, there was someone wearing a Modal shirt. I was like, “What are you doing here?” They're like, “Yeah, I just. I am embedded inside of Cog.”Akshat [00:08:05]: Yeah, I think that was Peyton. We sent him overSwyx [00:08:07]: Yeah.Akshat [00:08:07]: Because, the latency of communication was too high otherwise.Swyx [00:08:12]: Yeah, distributed node, you have to - you have to place one and collocate.Vibhu [00:08:16]: Yeah.Swyx [00:08:16]: So I had a, I had direct personal experience, right? So I worked on smol developer three years ago. it was inspired by Claude 1. I think you onboarded me at some point, like, just before, and I was like, “Oh, like, I need some bursty compute. Like, I was just gonna try using Modal.” And it was a, it was a pretty pleasant experience. apparently, I showed up in the board meeting, like the analytics.smol developer, Sandboxes, and Proto-CognitionAkshat [00:08:39]: Yeah, you blew up on Hacker News and,Swyx [00:08:41]: YeahAkshat [00:08:41]: We got a big traffic spike. I. I think the way you used smol developer was Modal functions for running stuff, which was. Like, the, that was a good use case. but then, yeah.Swyx [00:08:53]: Yeah. That - So to me, that was proto-cognition.Akshat [00:08:55]: Right.Swyx [00:08:56]: If only I had, like, stuck to it.Swyx [00:08:58]: Like, that was like, if - did you say draw the tech treeAkshat [00:09:00]: AbsolutelySwyx [00:09:00]: You're just like, “Yeah, like, probably this will happen.”Akshat [00:09:02]: Yeah. Like, he was so close. You were just rebuilding upon usSwyx [00:09:04]: I just didn't realize.Akshat [00:09:05]: But the funny story there is at the same time, we were talking to a bunch of customers who needed something like sandboxing.Swyx [00:09:14]: Yeah.Akshat [00:09:14]: This is like twenty-three.Swyx [00:09:15]: Yeah.Akshat [00:09:16]: So we builtSwyx [00:09:17]: You introduced a new API right after that.Akshat [00:09:18]: Yeah.Swyx [00:09:19]: Yes.Akshat [00:09:19]: Like, we built sandboxes in May of twenty-three before anyone was even knew this was gonna be a thing. And the first example we published was, we took smol developerSwyx [00:09:28]: Smol developerAkshat [00:09:28]: And put it in a loop, so the agent can iterate on itself.Swyx [00:09:33]: Loops are hot these days.Vibhu [00:09:34]: It's the looper.Akshat [00:09:34]: Yeah.Vibhu [00:09:35]: Loops in. When was this, twenty-three?Akshat [00:09:38]: Yeah.Vibhu [00:09:39]: A small check.Akshat [00:09:39]: Yeah.Swyx [00:09:39]: It's like twenty-three. so the. the, those for listeners, like, the problem was the models are not built for any of this, right?Swyx [00:09:46]: Like, you're just trying to like. They're not post-training to understand, like, looping and, like, self-correction and tool calling was there, but, like, also not that great.Akshat [00:09:55]: Yeah.Akshat [00:09:55]: I don't remember if you used tool calling in this one, but yeah, the models would just diverge after like ten iterations and not produce anything meaningful.Swyx [00:10:03]: Yeah. But like, then. So okay, like now talking to myself three years ago, the answerVibhu [00:10:08]: Of course they will get betterSwyx [00:10:09]: Collect all the failures, build benchmark, and then collect all the, examples, build the RL environmentAkshat [00:10:15]: RightSwyx [00:10:15]: Sell it for like ten billion dollars to Meta.Swyx [00:10:17]: And then also train a model and then sell that for sixty billion dollars to Elon. And this isAkshat [00:10:23]: Yeah, of courseSwyx [00:10:23]: The funny machine. Like, it's like, it's about the hardware.Akshat [00:10:28]: It's hard to have that inherent conviction that the stuff will get that much better.Swyx [00:10:33]: In retrospect, it's so f*****g obvious.Akshat [00:10:36]: Fair enough.Swyx [00:10:37]: Like, what else were we doing back then? I don't know. anyway. Yeah. So this. That was the start of your sandboxing journey, right? I feel like it didn't blow up until, like, last year.Akshat [00:10:49]: Yeah.Swyx [00:10:50]: So there was like a couple years of quietness.Akshat [00:10:52]: Exactly, yeah. We wereVibhu [00:10:53]: I think very underrated product value. Like, my experience with Modal, Charles, before he had joined Modal, met this guy at a hackathon, and he really insisted we wanted to run some small model, not hosted anywhere, and he's like, “ there's this cool company, Modal. They'll like spin up a GPU sandbox, we can throw it on there. They'll take a Hugging Face link.” And like there's so much value just right there, right? Like instant hosting, spin it up, spin it down. It'll stay cold, but we run the demo a few days later, it'll come back up and like all this stuff in retrospect, like it's still what we needed like today.Akshat [00:11:27]: Yeah, it's still needed today. workload shapes have changed a lot as, we run stuff for people with really massive production scale and, there it's it's not about scaling from zero to one, but it's how do we scale really elastically, from like thousand to fifteen hundred GPUs very quickly in a given region. It's the same shape problem.Elastic Inference, GPU Autoscaling, and Custom ModelsVibhu [00:11:50]: Okay. So you look at, say, Cursor Composer, right?Akshat [00:11:53]: Yeah.Vibhu [00:11:53]: They had a. “We'll do RL on a model every couple hours.” you guys have a whole version of RL inference gym and whatnot.Vibhu [00:12:01]: When you look at workloads like that, you're doing train runs where you need to scale up, scale down every hour thousands of GPUs, right? That's the example for we do need it, right?Akshat [00:12:12]: Yeah. Well, so I'll, I'll take a step back and, maybe talk about like how people use Modal today. because our biggest use case is, elastic inference. And the thing we first found product market fit, with was inference for custom models. So we stayed away from the LLM space, and we were serving companies like Suno for audio, Runway for video, robotics, comp bio companies that train their own model elsewhere. But Modal is the best black box that for deployment, scaling to however many GPUs you need as your traffic pattern changes. And we saw all of them like have a very unpredict- predict- predictable, traffic pattern. it's like diurnal. It's Some days, like the company will do a launch and, they'll need like, way more. And it's not just one model that they deploy. They-- all these companies deploy, lots of different models in different regions, and so the autoscaling problem becomes even harder because then you have to scale within a certain region, and those cycles are offset. So different times you scale up in different regions.Akshat [00:13:20]: So that's like our sortVibhu [00:13:22]: And thatAkshat [00:13:22]: YeahVibhu [00:13:22]: That in and of itself is a huge category. There's a bunch of inference providers which, provide this fireworks, does this as a service together, whatnot, Base10. that's carved into its own niche for language models, at least right now.Akshat [00:13:36]: Yeah. the thing that we have specialized in is the autoscaling aspect.Vibhu [00:13:41]: Yeah.Akshat [00:13:41]: Because we found that it's not universally true that everyone else can autoscale, and we've gone deeper into it on the tech side by, we've incorporated GPU snapshotting into the product so we can take the GPU state, like your torch.compile model, snapshot it, and the next cold start is way faster. And so going back to your question, it's That's why you need a lot of burstiness for inference. But then people also do a lot of demand training, like for RL stuff, your rollouts are bursty, as you said. People also do a lot of batch jobs. So we'll see, a lot of companies, before they have a training run, they'll need thousands of GPUs to run encoding or something like that. And I think those things are much more bursty than. I agree that agents are not that bursty. sandboxes are, except when you're doing RL. RL is justRL, Batch Jobs, and 100,000 SandboxesVibhu [00:14:28]: Or commerceAkshat [00:14:28]: Insanely bursty.Vibhu [00:14:29]: Yeah.Akshat [00:14:30]: Yeah. Like when you're doing, rollouts, you sometimes need a hundred thousand sandboxes in your sandboxes.Vibhu [00:14:37]: Yeah. I'm curious if you've seen early sparks of continual learning. There are some people, like our friends, ngram, recently announced thisAkshat [00:14:45]: YeahVibhu [00:14:45]: They're, they're trying to do training. That also seems like a different workload, right? If you're doing training twenty-four/seven per se, there's a very weird dynamic of how you're using GPUs between people and whatnot, but seems like something you guys would work for.Akshat [00:15:00]: As you said, we're, we're fortunate to work with a number of, customers at the frontier and grab some of our customers. and they are taking the primitives we have, and trying to use them in very interesting ways, like continual learning. It's possible as the stuff gets better, some of that will be part of, our offering as well if, more people need it. but we're, we're just waiting to seeVibhu [00:15:23]: YeahAkshat [00:15:23]: How it shakes out.Vibhu [00:15:24]: Is there a primitive that you added after sandboxing that was the next step in the story?LLM Inference, DeFlash, and Speculative DecodingAkshat [00:15:32]: I guess we've been going much deeper into LLM inferenceVibhu [00:15:35]: YeahAkshat [00:15:35]: Because we realized that some of the advantages we have with like autoscaling, again, especially in different regions and whatnot, are, not present elsewhere. and the place where we had a gap was we weren't, working on the model layer itself. Like we were a black box. And, we realized that, we can get to frontier-level model performance, with, by having great people who work on this. And, we've been open sourcing a lot of our work, in terms of, Recently, we, shared our work on DeFlash, which is a block-based, speculator, and we've open sourced, all of it. So, you can - By using open source DeFlash, you can get the same performance as you would with one of the proprietary providers. And the next thing we're thinking about hereVibhu [00:16:23]: I thought this wasAkshat [00:16:24]: YeahVibhu [00:16:24]: An interesting blog post as well, right? Like, I think in here you make a claim that. Not a claim, just that how effective speculative deco-decoding really just get to.Akshat [00:16:33]: Yeah.Vibhu [00:16:33]: Anything you wanna point out from this around, what people should know?Akshat [00:16:39]: Yeah, absolutely. the high-level summary is, it would help to describe what speculative decoding is.Vibhu [00:16:44]: Yes.Akshat [00:16:44]: I will, yes.Vibhu [00:16:45]: I think, likeAkshat [00:16:46]: YeahVibhu [00:16:46]: So we've covered like Eagle and all thisAkshat [00:16:47]: YeahVibhu [00:16:47]: Like Hydra and all those things, but it was like two years ago.Akshat [00:16:51]: Yeah.Vibhu [00:16:51]: I think it doesn't hurt, right?Akshat [00:16:52]: Yeah. Speculative decoding is you have a smaller model, called a draft model, predict tokens ahead of the bigger model, and then you have the bigger model, verify all of this, all the tokens are predicted. And the reason it's faster is if you're predicting, one token at once, you're bound by memory bandwidth. But if you can batch the verification of, the draft model, then you're much more efficient using compute, and it's faster, and as long as your draft model is producing a lot of tokens that can get accepted, which is called the accept length, you can get a speed up that's, multiple times of, the original model speed. and well, that's what we highlight here. It's Like people talk a lot about we made these kernels faster and whatnot, but improving kernel will only give you like few percentage points of improvement, and, increasing accept length, literally is a multiplicative decreaseVibhu [00:17:47]: Like two to four X.Akshat [00:17:48]: Yeah, exactly.Vibhu [00:17:48]: Without much head-on performance.Akshat [00:17:50]: Yeah. I think it may - you are running a second model, right? So it may be something more expensive in the compute,Vibhu [00:17:57]: I meant quality performanceAkshat [00:17:58]: Probably not by muchVibhu [00:17:58]: But yeah. I thinkAkshat [00:17:59]: So there's no drop in quality performanceVibhu [00:18:01]: YeahAkshat [00:18:01]: Because you're always. You're never accepting a token that the big modelVibhu [00:18:04]: It's strictly betterAkshat [00:18:05]: YeahVibhu [00:18:05]: Or it's same.Akshat [00:18:06]: Exactly.Vibhu [00:18:07]: Right. Yeah.Akshat [00:18:08]: And so we've been working a bunch on DeFlash, which is a block-based speculator. so it's instead of predicting, one token at a time, it's predicting a block. And we've been open sourcing our work with it. The next thing for us here is for helping people train speculators and custom models. it's it's something that traditionally is very forward-deployed engineering driven, support deployed, engineer driven, like you work with customers and help them do that. And our vision for. This is why we launched Auto Endpoints, is we want to make frontier-level performance available to everyone. And so, we mentioned this in the announcement, we teased it. The next thing we're, we're launching is, as you run an auto endpoint, we shadow trafficAuto Endpoints and Frontier-Level PerformanceVibhu [00:18:54]: Do you want to explain what auto endpoints are?Akshat [00:18:57]: Yeah.Vibhu [00:18:57]: I lovely, yeah.Akshat [00:18:58]: Yeah. So, this is, I guess, going back to your Modal is you touch the code, but, sometimes people don't wanna touch the code, and they wanna get started with an endpoint that works and has all the great performance and, scalability that Modal has. So we've made that easier with, a way to create an endpoint from our UI, from the CLI, that has all of our optimizations that we talked about, like the DeFlash stuff already baked in, and there's full transparency. So we give you the code, you can go run it yourself, and if you want, you can eject out into the full Modal experience, which we see as people get sophisticated, they do wanna tweak the models, they wanna, fine-tune stuff. You can still do all of that. It's it's not a black box. And yeah, the next thing, as we teased later in the post, is how do we give you value even beyond this in terms of having your draft models evolve as your data distribution evolves, again, without having to talk to a person and, yeah.Vibhu [00:19:59]: I guess just to understand it directly, you have the GPUs, you have an endpoint that's compatible, you serve open model. If someone was to do this themselves, what's the delta that you guys provide? So you do a lot of open source great work on effective inference. how does it compare to, say, I take the same model, 5.2 FP8, take shelf inference engine, vLLM, SGLang, get compute of similar capacity, similar cost. What's the delta that plugging into something this, like this offers outside of the benefit of, scaling?Production Inference Beyond Raw GPUsAkshat [00:20:34]: It's interesting because we've taken the approach of open sourcing our contributions and upstreaming them. we work closely with the SGLang team. We want the improvements that our team, comes up with to be, there in open source for others to use, even outside of Modal. The benefit to us is we have a team that has significant expertise in terms of if you do have something that is not there, our team can help you get that performance, first. the other thing is with these endpoints, we are way more elastic, as you said, than, anyone else, and you have true scaling to zero. you have true, burstiness, and in practice, that matters a lot more to people than just finding, the GPU and, running Modal code on something.Vibhu [00:21:20]: Yeah. And I will say it's not that straightforward to just. like what I said is easier said than done, right?Akshat [00:21:26]: Yeah.Vibhu [00:21:27]: It's I think still for the average person, still hard to just gut check using different. There's, there's quite a bit of combinations you can make there. the trade-offs aren't really known at face value.Akshat [00:21:40]: Yeah. it's it's not just that. I think it's it's that running production-grade inference is a hard infer problem.Vibhu [00:21:49]: YeahAkshat [00:21:49]: Even if you subtract out the autoscalingVibhu [00:21:50]: YeahAkshat [00:21:51]: Is controlling things like tail latency and, making sure every, request is delivered at least once and whatnot.The Model and Agent LifecycleVibhu [00:22:00]: There's a lot of innovation that you can do here. I think, it's very interesting that you're starting to encroach on, like as you become a full cloud, you're starting to encroach on other people's turf.Vibhu [00:22:09]: What will you not do?Akshat [00:22:13]: Well, we wanna follow our users and, make sure they get like a platform that has everything that works well together. so right now we're focused on the model lifecycle and the agent, lifecycle. so both like going from data prep to training to inference, and then also if I want to deploy a background agent, let's say, sandbox, do persistent storage, a whole bunch of other stuff.Vibhu [00:22:38]: We talked to Cole, who did, OpenInspect. Yeah.Akshat [00:22:42]: Yeah.Vibhu [00:22:42]: And RealInspect also is on Modal.Akshat [00:22:44]: Yeah. So Ramp Inspect was a great example of a background agent that was really successful because they, were able to use some of the primitives like snapshotting and fast scaling to just have something that feels really reactive and works well.Ramp Inspect and Background AgentsVibhu [00:23:02]: Yeah. That's the new CTO of, Ramp right there.Akshat [00:23:05]: Yeah, Rahul.Vibhu [00:23:08]: It was really fun. yeah, okay, I think, all very bullish. Like, one of my reflections was also I did not originally. So when I met you guysThe Inference Inflection: CPU, GPU, and Co-LocationVibhu [00:23:19]: You weren't that much in the GPU game, and now you're all about, inference. And one of the points that I hinged on for Jensen's keynote at GTC this year was, what we're calling like the inference inflection, right? That let's say in AI workloads or machine learning workloads, it used to be like, let's call it eight to one GPU to CPU, and now it's more like one to one, which is like a interesting. Like, - because of how much agents are blocked or call out to this, to CPU heavy stuff the actual, like, limiting factor, like, swings back and forth from GPU to CPU a lot more than it used to be all GPU and then occasional CPU.Akshat [00:24:01]: Yeah.Vibhu [00:24:02]: GPU, CPU. And now it's like just constantly, and you just have to locate everything.Seventeen Clouds and the Supercloud StrategyAkshat [00:24:08]: Yeah. And that's one of the things that, again, we see as, something appealing about Modal, which is we've built this capacity pool that spans, 17 cloud providers, so we're, we're very good at Running on various kinds of cloud capacity across the worldSwyx [00:24:24]: You don't have your own data centers?Akshat [00:24:25]: We don't have our own data centers. We just run across a lot of neo cloudsSwyx [00:24:29]: Yeah. AreAkshat [00:24:30]: Metal providers.Swyx [00:24:30]: Yeah. Question mark.Swyx [00:24:31]: Yeah. You're, you're running the math, and you're like, “What's the cutover point where you're like.”Akshat [00:24:36]: Yeah, it's a good question. part of it is we see our differentiator in the software layer, and, being capital light and focusing on the software helps us move really fast. so far it's worked out well because there are so many other people building data centers that we're able to work effectively with them, and again, focus on what makes us, special.Swyx [00:24:55]: Yeah.Swyx [00:24:56]: 17 gets you into, like, the local providers sometimes. LikeAkshat [00:25:00]: The,Swyx [00:25:01]: Which was the most interesting one?Akshat [00:25:02]: There are a lot more neo clouds than you expect, and they all have various degrees of, various levels of reliability. And, that's why it's something we've invested a lot of time in, is building our own reliability layer on top. so if the GPU falls off the bus or something happens, we user workloads are not affected, and that lets us use a lot more capacity than,Swyx [00:25:30]: YeahAkshat [00:25:30]: You as a user would be able to.Swyx [00:25:32]: It's a useful thing to have because like now everyone knows, like, what layer you are and, like, you optimize for being the super cloud of all clouds.Akshat [00:25:41]: Yeah. That's, that's, that's the idea. and so I guess when you mentioned colocation, that's, that's another interesting thing where, one thing we've seen is people come to us when they want, very specifically located, CPUs or GPUs, like they wantSwyx [00:25:57]: Oh, they pin it in likeAkshat [00:25:58]: YeahSwyx [00:25:58]: EU?Akshat [00:25:59]: Exactly. Or EU, US.Swyx [00:26:01]: Right. Data resiliencyAkshat [00:26:02]: AustraliaSwyx [00:26:02]: Locality thing or performance or what?Akshat [00:26:04]: It's either data locality or latency, yeah.Swyx [00:26:07]: Yeah.Akshat [00:26:07]: Like, you want your. They're running sandboxes and model. They want them to be right next to aSwyx [00:26:10]: Yeah, it's easy thenAkshat [00:26:11]: YeahSwyx [00:26:12]: To. That is important in all those things. and so, like, you've accidentally, I don't know if it's accident, but, like, you've built the perfect primitive for agents to express themselves. And then, like, it's almost very funny how every extra development just involves more file system, just involves more CPU.Akshat [00:26:30]: Yeah.Swyx [00:26:31]: Just like the things that you already have. I don't know much about, if there's any, like, networking usages that are interesting, but you've also done some good work on networking.Networking, Sidecars, Private IPv6, and SandboxesAkshat [00:26:40]: Yeah, that's exactly right. Like, we're just taking compute storage and networking and building stuff on that layer, for, again, the stuff people need.Swyx [00:26:49]: YeahAkshat [00:26:50]: We see a few interesting networking things coming up. one is people want networked sandboxes. so we haveSwyx [00:26:57]: For like a Docker cluster type thing.Akshat [00:26:59]: Yeah.Swyx [00:26:59]: Sorry, Docker Swarm. Oh, f**k. What is it called?Akshat [00:27:02]: Compose.Swyx [00:27:03]: Compose type thing.Akshat [00:27:04]: Yeah. So if you want Docker Compose, our sandboxes now support, this thing called sidecars. So you can. A sandbox is a pod of containers, and you can run multiple containers in, a sandbox. also useful because, going back to networking, people want a lot of control over, outbound networking from a sandbox.Swyx [00:27:23]: Yeah.Akshat [00:27:23]: Like, they might wanna run a middle proxy for, like, maybe logging stuff for RL or, controlling how egress can happen to a domain, injecting credentials. and yeah. So we've, we've had to build a lot of that stuff ourselves.Swyx [00:27:38]: Yeah.Akshat [00:27:39]: But then also sometimes people want, sandboxes spanning multiple nodes to talk to each other, which is an emerging thing we're seeing. We have support for that for a different reason, and yeah, we'll see if that becomes stable.Swyx [00:27:52]: Like, just an open socket. It's a. This is directly like mTLS.Akshat [00:27:56]: We do support that, which is you can, expose a tunnel inside a sandbox.Swyx [00:28:01]: Yeah.Akshat [00:28:01]: And then you can either expose it to public internet or it can be, you can add like a HTTP, auth layer above it. But we have this thing called I6PN, which we haven't talked about, which is this, like, overlay network using IPv6 addresses. so if Modal containers, within the same workspace, when this is enabled, can address each other using this private IPv6 address, and no one else can.Akshat [00:28:28]: So it's like private networking, for containers. We built it because we needed it as a primitive for our distributed training product. so we have this other feature, which is you can add a decorator to a function, and you get a cluster of GPUs. and they have RDMA networking. so you can run a distributed training job, that's truly serverless. and we did the overlay network for that. But then we've seen that people are using it for other reasons, and, I'm intrigued to yeah, what would people do with it.Swyx [00:28:59]: Build primitives and let people figure it out, right?Akshat [00:29:01]: Yeah, exactly.Swyx [00:29:02]: You put out a pretty interestingAkshat [00:29:03]: They're like, they read the docs webpage. Let me use thatSwyx [00:29:06]: YeahAkshat [00:29:06]: Something they never intended to work. This is literally not even in our docs page. People somehow found it, and they're using it.RDMA, Memory Movement, and Distributed TrainingSwyx [00:29:12]: Huh.Swyx [00:29:14]: The way you portrayed it with, like, RDMA versus TCP, like, very well laid out, but just the transfer speed change at scale for RL, like yeah, you have it, you have it built in. I'm sure someone found it. It's found it to be a lot more efficient before you made a thing out of it, right?Akshat [00:29:32]: Yeah. And not to split hairs, I guess the overlay network is the TCP overlay network.Akshat [00:29:39]: The reason we have that is you need that to do the key exchange for RDMA before you set up the RDMA network on top of that. but then people found the TCP part.Swyx [00:29:48]: Can I tell you, this is like a big aha moment for me becauseAkshat [00:29:51]: YeahSwyx [00:29:51]: So I review 2,200 submissions for the World's Fair.Akshat [00:29:56]: Yeah.Swyx [00:29:57]: And then I got this from John OsterhoutAkshat [00:29:58]: HuhSwyx [00:29:59]: Who I don't know if. Do John Osterhout by name?Akshat [00:30:01]: The name sounds familiar.Swyx [00:30:02]: He published a. He's a well-known professor, published a lot of interesting software design books, and this is the talk he chose to submit, is on RDMA at Inference. And I'm like, you wouldn't think that this guy, who is like operating systems guy, would care about RDMA.Akshat [00:30:20]: I, it makes sense to me because I,Swyx [00:30:24]: This is the cloud, right? YeahAkshat [00:30:25]: Like, the way you move around your KV cache and how efficiently you can do it, how efficiently you move, your weights from your training GPUs to your inference GPUs in RL is there's a lot of degrees of freedom, and it is a systems problemSwyx [00:30:41]: YeahAkshat [00:30:41]: Moving memory aroundSwyx [00:30:42]: YeahAkshat [00:30:43]: Scheduling.Swyx [00:30:44]: This shows you how primitive my understanding of networking stuff is.Swyx [00:30:46]: Is this like the domain of WireGuard as well?Akshat [00:30:50]: Not quite.Swyx [00:30:51]: It's adjacent?Swyx [00:30:53]: Explain everything.Akshat [00:30:54]: Sure.Swyx [00:30:56]: How do we move memory around GPUs?Akshat [00:30:58]: Well, so sorry. Yeah, that is memory. Sorry, I was talking more, and maybe I was talking like five minutes back, about the private IPv6, addressing that you've set up.Swyx [00:31:09]: Yeah.Akshat [00:31:09]: Is it like it's a VPN?Swyx [00:31:10]: Yeah, it is like a VPN, and yeah, WireGuard is, yeah, you're right. It is,Akshat [00:31:16]: Right. Yeah, you already moved on to new topicsSwyx [00:31:17]: A similarAkshat [00:31:18]: OkaySwyx [00:31:19]: In the same space, WireGuard is, encrypted and this is,Akshat [00:31:23]: And you don't need encryption.Swyx [00:31:23]: Yeah.Akshat [00:31:24]: Yeah.Swyx [00:31:24]: This is not encrypted. that's the main difference. This is TCP and we have eBPF programs that will reject or allow the TCP connection based on whether you're allowed to do it.Akshat [00:31:35]: Used to involve a full sidecar, but now you have eBPF in the Linux kernel.Swyx [00:31:39]: Yeah.Akshat [00:31:40]: Yeah. I don't know if this is a natural follow-on to the topic of like my skepticism on distributed training is that while, like, people spend a lot of money on, like, cables to hook up GPUs, and even that is not, like, fast enough, and that's the bottleneck, is your networking fast enough?Swyx [00:31:59]: Yeah. So I guess you're talking about fully distributed training like, Dialog or something which is like cross data centerAkshat [00:32:06]: That would be, yes.Swyx [00:32:07]: That's the extreme.Akshat [00:32:08]: Yeah.Swyx [00:32:08]: You're in the middle, and then other people would have like the Mellanox cables up in, like, their actual data center.Akshat [00:32:14]: When you run multi-node training on Modal, RDMA, I think Mellanox, is, or InfiniBand is like a, is all seen as RDMA. but it's a way to bypass the TCP networking stack and, transfer, stuff much faster, between one node, to the other. And we have I think like 3 terabit per second, internal networkingSwyx [00:32:40]: OkayAkshat [00:32:40]: Which is the standard that's needed.Swyx [00:32:42]: Okay. So I misunderstood whatAkshat [00:32:43]: 50Swyx [00:32:43]: What part of the stack you wereAkshat [00:32:44]: 50 gigs overSwyx [00:32:45]: YeahAkshat [00:32:45]: If you wentSwyx [00:32:45]: YeahAkshat [00:32:46]: RDMA.Swyx [00:32:46]: Okay.Swyx [00:32:48]: Yeah. I, very impressive work.Multi-Node Training, Post-Training, and Auto ResearchSwyx [00:32:52]: So effectively you're extending like the model philosophy to the training cluster, like, yeah.Akshat [00:32:59]: Yeah. And we're, we're not going for like large scale training runs. the thing that we've built multi-node training for is, we see a lot of, smaller scale post-training. like, people are post-training like medium sized fund models, so they can, get higher quality on inference. this is a perfect fit, for something like that.Swyx [00:33:21]: Yeah. That is my impression of how a lot of these labs explore branches in post-training and then eventually merge whatever they find in.Akshat [00:33:31]: Yeah. The other use case we've seen for multi-node training is even if you have a big cluster, your researchers are still doing small runsSwyx [00:33:38]: YesAkshat [00:33:39]: Having elasticity thereSwyx [00:33:40]: Right, sureAkshat [00:33:40]: Matters a lot more.Swyx [00:33:41]: Yeah. the, like, this is like the current limiting factor for auto research, which is like you need to give your model some GPUs in order for it to completely run.Akshat [00:33:51]: We have a blog post on auto resource and model is,Swyx [00:33:55]: YeahAkshat [00:33:56]: Yeah, like, turns out to be pretty good substrate for that.Swyx [00:33:59]: So my impression is auto research means many things, likeAkshat [00:34:01]: YeahSwyx [00:34:01]: Anything that Andrej coins. Right now it's still science fair, right? Like not like, I don't know how many people are doing this.Akshat [00:34:08]: We're having a golf.Swyx [00:34:08]: Yeah.Akshat [00:34:09]: I thought the same thing.Swyx [00:34:11]: Yeah, you would know.Akshat [00:34:12]: We, like, our internal both training and inference teams use this the general shape of this quite a bit. like we have this one internal repo called auto inference, which essentially we've automated our own forward-deployed engineering efforts using, this harness, which is, the agent will just spin up a sweep of different things. It'll even run like, NVIDIA inside profiler and it'll like tweak configs and it'll arrive the right thing. it'll change your GPUs both from H200 to B200, and works really well.Swyx [00:34:47]: Nice.Akshat [00:34:47]: So yeah.Swyx [00:34:48]: By the way, I enjoy that your forward-deployed engineering is so technical that you have to do these things.Swyx [00:34:52]: It's very different from forward-deployed engineering from other people.Akshat [00:34:54]: Yeah. For our forward-deployed engineering team is, essentially they're like applied inference researchers or applied training researchers.Swyx [00:35:02]: Someone told me like they have to be able to build, but they also have to be able to sell. do they have to sell or are they like they're good, they're just like post-sale type of thing?Akshat [00:35:09]: It does, being able to talk to a customer and engage effectively with themSwyx [00:35:13]: YeahAkshat [00:35:13]: Matters a lot.Swyx [00:35:14]: They want the same thing.Akshat [00:35:15]: Yeah.Swyx [00:35:15]: ?Akshat [00:35:15]: But it's it's not really a sales, thing. We pair them with-- We have solution architects as well that are more on the sales side.Swyx [00:35:23]: Okay. Let's spend a bit more time on auto research. This is a big focus for for this year. Where does this go? like, have people explored enough? Like, there's all these beautiful charts of like improve and then level off a bit and then you find the next thing. Is this one abstraction up from normal training? Is that how we think about it, or do you think about it differently? Like model level training versus high, like driven hyperparameter search.Auto Inference and Modal BenchAkshat [00:35:51]: Yeah, like,Swyx [00:35:51]: Someone, some people call it like neural architecture search or whatever, right? Like.Akshat [00:35:54]: Yeah, - So the stuff I've seen people do with it is nowhere on the architecture level. It's pretty much tweaking parameters, but it's it's a hyperparameter sweep that's guided by some model intuition, so it's like much more efficient than, whatever other, sweep you would have.Swyx [00:36:12]: Yeah, it's just, it's just a question of where you want to spend your compute?Akshat [00:36:16]: Right.Swyx [00:36:16]: ‘Cause yeah, you can just throw infinite amounts of money on this and somehow you'll bang out Shakespeare?Akshat [00:36:22]: Yeah, infinite monkey.Swyx [00:36:24]: Yeah, so like the very good for model. and I think it's also very important that agents can spin up other agents, can spin up their infrastructure. Like very good for you. how good is our LLMs at generating model code? Like the benefit of existing LLMs is that you are in the data.Akshat [00:36:42]: Yeah. They're, they're surprisingly good. I think like pre Cloud 4 they were not, and then now they're able to shot, stuff out of the box. But we're playing around with releasing like a Modal Bench for like the harderSwyx [00:36:55]: YeahAkshat [00:36:55]: Things, that the LLMs cannot do yet and maybeSwyx [00:36:59]: What's an example of that?Akshat [00:37:01]: I think the things that- Sometimes agents struggle with, without right guidance and a skill is, how to, use the rest of our observability. Like how to. Something is failing, like how do you look at the logs and then update the right thing? It's reasoning about that. But they're able to shot, likeSwyx [00:37:23]: Yeah. You can just add a skill to it?Compute Strategy and Capacity PlanningAkshat [00:37:26]: Yeah. So we have a Modal skill now that. Which is why we built this Modal Bench. It's to find things like that, so we can address them in our tool.Swyx [00:37:35]: Tune a skill. Yeah.Akshat [00:37:36]: Yeah.Swyx [00:37:36]: No. it's it's good. are you facing any shortages? like we talk a lot about GPU shortages, but also CPU, also memory.Swyx [00:37:44]: Yeah.Akshat [00:37:45]: We have had a lot of growth, which means that, there's - we've had to be much better aboutSwyx [00:37:53]: PlanningAkshat [00:37:54]: Proactive capacity planning.Swyx [00:37:55]: Yeah.Akshat [00:37:55]: So we have,Swyx [00:37:57]: Which by the way, like it's like a MBA's like dreamAkshat [00:38:00]: YesSwyx [00:38:00]: Is like just planning this stuff. I think last time you and I talked about something maybe about this.Akshat [00:38:03]: Yeah. we have a really competent team of people that we call, The role is called compute strategy. so yeah, if anyone listening here or wants to work on thatSwyx [00:38:13]: Compute strategy?Akshat [00:38:13]: Yeah.Swyx [00:38:14]: I think,Akshat [00:38:14]: I feel like,Swyx [00:38:15]: I think the normies call it FP&A or something.Akshat [00:38:18]: Well, it's more It's it's not FP&A. It's it's There's a lot of interesting financial questions of like what is the blend between one year and three-year reservations? how do we forecast our own capacity? how do we. especially since our capacity is very fungible across different GPU types and different regions, like you have to model a lot of it. and you also have to have an opinion on how the supply chain is gonna evolve, and then you have to like, take bets,Swyx [00:38:49]: YeahAkshat [00:38:49]: Based on that.Swyx [00:38:50]: Tokenomics.Akshat [00:38:50]: Yeah.Swyx [00:38:51]: This is like probably a not a real point, but, I was trying to think about like what other industries. I was trying to think about like, we cannot be first to like these kinds of problems.Akshat [00:38:59]: Yeah.Swyx [00:39:00]: And what other industries have had this? And I was like, airlines with fuel and like they have to hedge their fuel and like, I think for a long time Southwest because they made like a hero fuel bet, they like were like super low cost becauseAkshat [00:39:12]: OhSwyx [00:39:12]: Compared to everyone else.Akshat [00:39:14]: Yeah. I hadn't thought about that.Vibhu [00:39:16]: We're at a fun time too?Akshat [00:39:18]: Yeah. It's. A lot of the compute business in general, for us is also about being very good about capacity management. That is how you have great unit, economics. but also over time it's how you can unlock more value for customers. Like, one of the things we're building now is like a way for customers to get, If they don't care about latency, like get much cheaper pricing and they'll get results back in like next 24 hours or something, like a batch tier essentially.Batch Tiers and Latency-Insensitive WorkloadsSwyx [00:39:47]: Yeah.Akshat [00:39:47]: And those are levers we have because we control the whole stack and scheduling and whatnot to give people a sufficientSwyx [00:39:53]: Yeah. I feel like they're not as popular. Like those, like the Frontier Labs have all those APIs. They're not as popular as they should be.Akshat [00:40:00]: The demand that we see for something like that is not for LLMs. although sometimes people wanna run evals andSwyx [00:40:08]: OkayAkshat [00:40:08]: Synthetic data prep and there it makes sense.Swyx [00:40:10]: Okay.Akshat [00:40:11]: But it's from a lot of LLM companies, like people who are doing computational bio, like they have to run really big batch jobs and they don't care about when they get it back.Swyx [00:40:22]: Yeah. And like they have a reasonable. It's it's also like a cousin to the stopping problem of like, will this finish in time?Akshat [00:40:30]: Yeah. You can bound it.Swyx [00:40:33]: Yeah.Akshat [00:40:33]: Like you can give peopleSwyx [00:40:34]: YeahAkshat [00:40:34]: SLAs on it.Swyx [00:40:35]: Yeah. I think what's, what's interesting is like the next phase of model.Swyx [00:40:38]: Like what, do people expect from you, now that you're established and you're like well-known compute player among all these leading companies. You had an inference launch week, and we talked a little bit about the launches. like what else? Like what else should people know?What Modal Builds NextAkshat [00:40:55]: We are building primitives that make our users' lives much easier. So, I think for example, with LLM inference, thousands more companies are gonna post-train their own models and, deploy open source models for inference. so we're thinking a lot about what is the best product shape for that. And, that involves everything from our training gym to, then, endpoints that get frontier-level performance. again, but I haven't talked to anyone. It looks somewhat different on other verticals. Like, we're also seeing a lot of real-time, audio-video stuff in there, which is why like, we're working on things like regional routing, with fallbacks. So you can get GPUs that are as close to users as possible. so you get like low latency for video streaming and whatnot. And then on the agent side, it's,Akshat [00:41:52]: We're still working very closely with our customers because stuff is changing so fast in terms of what they need. And, I think beyond sandboxes and persistent file systems, there's a lot of other things people will need from this agent stack as they build production agents. So yeah, we're thinking about those other things that fit in there.Swyx [00:42:13]: I want to ask what the other things are.Akshat [00:42:15]: Yeah. I probably should share right now.Swyx [00:42:17]: I think-- I think, okay, so, I do think a lot about the principal components of cloud, and you do talk about compute storage networking.Akshat [00:42:25]: Yeah.Swyx [00:42:25]: Because so far for me, it's fine. so far for the. the first couple generations of cloud, it's fine. What's different, qualitatively different about agents that you need some new permission level? Like a lot of people, okay, and I'll just kinda spew tokens at you until it like hopefully sparks something.Akshat [00:42:43]: Yeah.Swyx [00:42:44]: Like the new level now is whatever Claude Code does, which is dangerously scope permissions or like allow list by command or like whatever, right? And sometimes they're like, “Well, okay, we have like this adaptive thinking mode where like, just trust me, bro. I will make the calls for you.” Is that it? like mediated permissions.Hard Guardrails vs. LLM-Mediated PermissionsVibhu [00:43:03]: Now you're looping it with a goal and letting it roll.Akshat [00:43:06]: Yeah, I'm, I'm skeptical of LLM media permission for stuff that is at the sandbox level because you do want hard boundaries.Swyx [00:43:16]: Yeah.Akshat [00:43:16]: Otherwise, someone can exfiltrate stuff.Swyx [00:43:20]: But likeAkshat [00:43:20]: YeahSwyx [00:43:20]: Maybe that's old school thinking. Maybe we're the dinosaurs.Swyx [00:43:23]: Maybe the AI OS or the LLM OS is really the kernel is a goddamn LLM.Swyx [00:43:30]: Like it makes you feel uncomfortable.Akshat [00:43:31]: Yeah, I'm, I'm toldSwyx [00:43:32]: But that's what trusting the LLM is. Like imagine a spherical cow perfect LLM.Akshat [00:43:36]: Right.Swyx [00:43:37]: That it.Akshat [00:43:39]: Maybe.Swyx [00:43:41]: I wanna test the boundaries, right?Akshat [00:43:42]: Yeah.Swyx [00:43:42]: Like, and I don't believe that, but I wanna see where I'm wrong ‘cause that's, that's the consensus.Akshat [00:43:49]: Yeah. I think you always need hard guardrails when you want, And you can pair those with softer guardrails, right? And that's gonna be a lot of mediated.Managed Agents and Specialized SandboxesSwyx [00:44:00]: There. I'll also get you a end with a couple of your commentary on like the ecosystem outside of Modal. Manage agents. Everyone has one. Gemini, OpenAI, Claude, very useful for you, but also like it is their way of starting to edge into your space.Akshat [00:44:17]: Yeah.Swyx [00:44:17]: What's going on?Akshat [00:44:19]: Yeah, we're, very excited to partner with Anthropic and some of the other foundation labs, will not name who we're also working with. the way we see it is the manage agent thing is a great place to start if you're starting out building an agent and, But then when you get to, building something more production grade, like you're a company that's like Ramp that's building their own, Ramp also runs their accounting agent on us, so their external-facing agent. You need a lot more control over, your compute primitive on things like, what sort - how do you persist different files that the agent has access to, and how do you snapshot and restore? How do you control the networking? maybe you want GPUs. When you get to that point, you kinda want, a specialized sandbox provider, that gives you those things, and that's the role that we are trying to play.Swyx [00:45:15]: YeahAkshat [00:45:16]: We don't really have an opinion on the harness, whether it runs - it's a cloud-managed agent, and you hook it up to Model Sandbox, or you run the harness in Model Sandbox. We'll see where people converge with that.Swyx [00:45:26]: Yeah. Do you any opinions on like the meta harnesses, or just another layer on top of these things?Akshat [00:45:31]: You mean like the OpenPipeSwyx [00:45:33]: OpenPipe is one. I think Vercel had one, which I can't remember the name of right now. Fredshot had one. and then, to me, most recently was Data Databricks that had Omnigen. All these are meta harness. Like it's kinda pseudo agent cloud type things.Akshat [00:45:50]: I personally have not played around with them.Swyx [00:45:53]: Yeah.Akshat [00:45:53]: Build agents with them.Swyx [00:45:54]: Everything's bullish Modal, as long as it consumes more infra.Akshat [00:45:57]: That's why we're focusing on the infra layer. It's somewhere where our, relative competence is and, also it's a hard problem to solve.Swyx [00:46:06]: Yeah. I will say like just generally reflecting on that, I don't know if - if there's other topics on Modal, but like just generally reflecting as an infra person, not as intense as you, but in that field, this has like been the most exciting time in infra. Like it was boring for a while, and you couldn't really get people excited about data infrastructure. Like Eric would get on Data Console, everyone just watched the video and like say, “Look at how many sandboxes I can spin up,” and no one gave a crap.Why Infrastructure Became Exciting AgainAkshat [00:46:39]: Yeah.Swyx [00:46:40]: And like now everyone gives a crap.Akshat [00:46:42]: That's true. It is a very exciting time, and I think a lot of that's driven by just the amount of scale all of this stuff needs.Swyx [00:46:50]: I think the, like a lot of your initiatives or a lot of your like product directions make sense in retrospect, which is like the best kind, but I wouldn't necessarily have thought about it myself, which.Akshat [00:47:00]: We need the predictions.Swyx [00:47:02]: I think there's a lot that you just don't even see, right? Like you have the batch, you have the voice, you have the multimodal, but what else?Akshat [00:47:10]: What else is coming up for usSwyx [00:47:11]: Yeah. Where do you see things going?Akshat [00:47:13]: Yeah. I, in generalBiotech, Robotics, and Non-LLM AI WorkloadsAkshat [00:47:15]: It's it's clear that there's there's a huge shift happening. I think one thing that's not as obvious to people because LLM inference gets talked about so much and is also we work a lot of companies that are, doing things like drug discovery and computational bio, like the Chai Discoveries of the world. Big things are probably gonna happen there. we work a lot of robotics companies that are putting robots in like active deployments and getting good results out of them.Swyx [00:47:45]: Is there Air Gap Modal? Is there a version that is like prem air gapped whatever?Akshat [00:47:50]: No. We,Swyx [00:47:51]: You should cloud only.Akshat [00:47:51]: Yeah.Swyx [00:47:52]: Yeah. Okay. But yeah, so what you're saying is like because you're focused on primitives and they're good primitives, you find use cases in all these kinds of things.Akshat [00:48:01]: Yeah.Swyx [00:48:01]: Probably diversifies you a little bit away from LMS all the time.Akshat [00:48:05]: Yeah, absolutely. We're, we'- our goal isn't to only serve the LLM inference market.Swyx [00:48:10]: There are a lot just on the website, the audio,Akshat [00:48:12]: Yeah. We said both onSwyx [00:48:14]: Computational bio images. Yeah, there's a lot here. There's QTA TTS, customizing. Oh, Chatterbox. there was customizing Whisper.Akshat [00:48:24]: Okay. Yeah.Swyx [00:48:25]: This screen reminds me of a fallen competitor, which Replicate.Model APIs vs. Differentiated AI ProductsSwyx [00:48:31]: What's your postmortem on what happened?Akshat [00:48:34]: This is one thing we've stayed away from is providing an API for models because I think providing model APIs is some of it ends up serving like a really hobbyist market, which is much less sticky.Swyx [00:48:50]: Yeah.Akshat [00:48:50]: And we've always wanted to build for companies that are building products and need more flexibility that's not just an API.Swyx [00:48:57]: Which you can build an API for a model and this is clearly what it is. But you - but what you're saying, you can wrap it into a more fully functioning back end that you run.Akshat [00:49:06]: Yeah. So all of our examples, it's not that spin up this model, here's an API token, use it. They're all code.Swyx [00:49:13]: Okay.Akshat [00:49:13]: And so the point is that this is just an example.Swyx [00:49:16]: Starter code.Akshat [00:49:17]: Yeah. But you can tweak it however you want.Swyx [00:49:20]: Yeah.Akshat [00:49:21]: And if you're like a company building a product, like, computational bio whatnot, yeah.Swyx [00:49:26]: I guess I'm trying to tease out for listenersAkshat [00:49:28]: YeahSwyx [00:49:28]: When does it stop becoming, oh, you're just an API call and you're just a wrapper on API to becoming what you call a product, right?Swyx [00:49:36]: Like, what is that layer? Like what-- Like, more lines of code, but like beyond that, what is the substance that people add that qualifies it to be something more?Akshat [00:49:46]: I think there's a little bit of like a selection effect of like a lot of the companies who do wanna get deeper into that level are probably building something that's more differentiated. And, I think, an example is like - with LLM inference, originally we, worked with companies that were building their own post-training frameworks or they were, - Ramp early in the day was training their own tokenizer and like swapping out the tokenizer in Llama and whatnot. I'm not saying that's, that successful, in that case. But a better example is like, let's say Suno. because Suno, does not use Modal for training.Swyx [00:50:26]: Mikey on the pod. Yeah.Akshat [00:50:27]: But they use Modal for all their inference and that's because they have like a custom-- They have completely custom model architecture and that means that they have to be at the code level and tweak things that are not, just an API.Swyx [00:50:41]: It's interesting as well, like we had, Ethan, most recently on the xAI Groq team make a prediction that like the next tier in video gen is not a better video model, it's a better model or agent that orchestrates video models.Video Agents and Production WorkflowsAkshat [00:50:56]: Oh, interesting.Vibhu [00:50:56]: Language model backbone that can use toolsAkshat [00:50:58]: RightVibhu [00:50:59]: And write code.Akshat [00:51:00]: Like, yes, I can make my second video or my second video from Groq, but I want my minute video.Akshat [00:51:06]: And I'm not going there through normal video gen.Swyx [00:51:10]: Yeah, that's interesting. I - So we have GPU sandboxes and recently have seen a few companies doing agents that do video manipulation or,Akshat [00:51:22]: Yeah. Give it FFmpeg and just do it.Swyx [00:51:23]: Run FFmpeg. But likeAkshat [00:51:25]: That's not enough.Swyx [00:51:25]: Yeah.Akshat [00:51:26]: You need to give it Adobe.Swyx [00:51:27]: Yeah, I hadn't put it together with like it would be a video production thing. in my mind these things were going more towards editingAkshat [00:51:36]: Yeah.Vibhu [00:51:36]: Well, shout out Mantis.Akshat [00:51:37]: I think about this a lot.Swyx [00:51:38]: .Akshat [00:51:41]: Yeah. Sorry.Vibhu [00:51:41]: Luma. Luma Agent is a version of this for video production, but it's a off.Swyx [00:51:46]: I was gonna get your quick takes, on some other stuff that happensGitpod/Ona, CI, and Runtime SandboxesSwyx [00:51:50]: In recent news and just-just see if you have anything interesting. Gitpod, very li
Have you ever built something for yourself that you were almost embarrassed to admit you needed? David Gonzalez did. And it changed how he sees himself entirely. In this episode of The Happy Hustle Podcast, I sit down, well, I hand the mic over, to David Gonzalez during our AI Masterclass Workshop, and he delivers one of the most real, unfiltered breakdowns of agentic AI I have heard all year. David is the founder of Internet Marketing Party, a monthly founder room he has run since 2008 without missing a single month. Almost eighteen years in, he has brought over 20,000 entrepreneurs through his events and facilitated half a billion dollars in introductions and revenue. Names like Alex Hormozi, Jay Abraham, and Craig Clemens have all shared his stage. David built his entire career and net worth on one skill, putting the right people in the right room. He is based out of Austin, Texas, and he is the first to tell you he resisted vibe coding for months before it cracked his whole business wide open. This episode matters because David is not a developer. He is not an engineer. He is, in his words, an offer guy, an operator, someone who knows the nuts and bolts of business but never touched code. And that is exactly why his story lands so hard. If you have ever thought AI was not for you because you are not technical, David is proof that thinking is wrong. He opens by admitting he quit vibe coding months earlier because he was doing it all wrong. His computer ran slow, he had a dozen unfinished projects, and he did not even know what GitHub was. Then a mentor rebuilt his setup in three hours, and his usage went from 61 sessions in April to over 2,285 sessions by June. That is not a typo. That is what happens when the right system finally clicks. David makes a distinction that is worth sitting with. Chat AI, the kind most people use, is basically an expensive browser. You ask it a question, it gives you an answer, and you are still stuck doing the work yourself. Agentic AI, what he calls vibe coding, is a completely different animal. He compares it to a Blackhawk helicopter that flies itself versus a tricycle. You do not need to know how to pilot it. You just need to know where you want to go. To prove it, David builds an app live in front of the room. He shows the difference between a lazy prompt and a detailed one, and the gap in quality is not subtle. A good prompt, he explains, gives the AI the context of a real thinking partner, not just a task list. He calls it inviting a swarm of experts to the table, the same way you might know that one guy at a family cookout who can fix the smoker, close a deal, and give you parenting advice all in one conversation. Then he gets personal. David built himself a tool he calls iEngine, standing for Intelligence, Inference, and Insights. It reads fifteen years of his text messages, over 1.4 million of them, along with his entire call history, and sorts every contact in his life by relationship depth. He started with 7,368 contacts and one keystroke later had them organized down to what actually mattered. It tells him who to reach out to, who has gone cold, and what he owes people. He puts it simply. If you are using a normal CRM, you are the CRM's bitch. He flipped that. But the most powerful moment in this episode has nothing to do with software. David shares a story about realizing he had been wearing an expensive watch as a stand in for his own worth, a way to prove to the world he had made it after growing up near poverty. One day he left the house without it and drove off anyway, and something shifted. He realized AI had not made him enough. It just gave him the tools to access what was already there. He used to let people call him a super connector. Now he knows he is an architect. He builds things. He always could. If you run a business and have ever felt boxed in by what you think AI can or cannot do for you, this one will reframe everything. Head over to https://caryjack.com/podcastin/ and listen to the full episode. David's energy alone is worth the time, but the mindset shift underneath it is what will actually stick with you. Connect with Davidhttps://www.facebook.com/InternetMarketingPartyhttps://www.instagram.com/internetmarketingparty/https://www.youtube.com/@ImarketingPartyhttps://www.linkedin.com/in/davegonzalez/ Find David on this website: https://simplythecoolest.com/ Connect with Cary!https://www.instagram.com/caryjack/https://www.facebook.com/SirCaryJackhttps://www.linkedin.com/in/cary-jack-kendzior/https://twitter.com/thehappyhustlehttps://www.youtube.com/channel/UCFDNsD59tLxv2JfEuSsNMOQ/featured Get a copy of his new book, https://www.thehappyhustle.com/book Sign up for The Journey: 10 Days To Become a Happy Hustler Online Course @ https://thehappyhustle.com/thejourney/ Apply to the Montana Mastermind Epic Camping Adventure @ https://thehappyhustle.com/mastermind/ “It's time to Happy Hustle, a blissfully balanced life you love, full of passion, purpose, and positive impact!” Episode Sponsors: If you're feeling stressed, not sleeping great, or your energy's been kinda meh lately—let me put you on to something that's been a total game-changer for me: Magnesium Breakthrough by BiOptimizers. This ain't your average magnesium—it's got all 7 essential forms that your body needs to chill out, sleep deeper, and feel more balanced. I take it every night and legit notice the difference the next day. No more waking up groggy or tossing and turning all night If you're ready to sleep like a baby, calm your nervous system, and optimize your recovery, go grab yours now at https://www.bioptimizers.com/happy and use code HAPPY10 for 10% OFF. =================================================================== My Green Mattress If you've been waking up with back pain, feeling stiff, or just not getting that deep, quality sleep. This might be what you're missing: My Green Mattress. It's made with clean, non-toxic, and eco-friendly materials, so you're not just sleeping better, you're sleeping healthier too. The comfort and support are on another level, and you can really feel the difference night after night. If you're ready to invest in better sleep and better recovery, check it out at https://thehappyhustle.com/mygreenmattress =================================================================== Ozlo Sleep If you've been struggling to fall asleep, stay asleep, or just wake up feeling actually rested, let me put you on to something that's been a total game-changer: Ozlo Sleep. These aren't your typical sleep buds. They're designed to block out noise and help your brain fully relax, so you can drift off faster and stay in deep, uninterrupted sleep. Perfect if you're a light sleeper or just want that next-level rest. If you're ready to upgrade your sleep and wake up feeling recharged, check out https://ozlosleep.com and save $80 OFF using code HAPPY.
Sean Juroviesky is a senior security engineer at SoundCloud. In this episode, he joins host Paul John Spaulding to discuss shift happening in AI around where inference actually runs, a recent situation with AWS Bedrock and Anthropic routing inference outside the EU, and more. • For more on cybersecurity, visit us at https://cybersecurityventures.com
Amos joins us to break down why inference is the new oil and companies will hedge inference costs the same way airlines hedge jet fuel, why AntSeed is open router without open router in the middle, and why all network revenue goes to buy and burn with no company capturing value. 14 years in crypto. This is what he's building next.Amos Meiri is Co-Founder of AntSeed, an open-source P2P marketplace for AI inference, and a 14-year crypto veteran who helped build colored coins alongside Vitalik before Ethereum existed.The Rollup is where the leaders of digital assets and finance converge. Live from the financial capital of the world.Timestamps00:00 Intro03:17 Inference Is New Oil06:27 AI Financialization Just Starting09:27 BitTorrent Model For Inference11:59 1,000 Users Three Months14:45 Open Market Creates Competition17:31 Foundation Not A Company20:01 Liquid Inference Token Explained22:38 Venice DM Token Parallel25:36 Permissionless Inference Is FutureGuest Socials:Amos Meiri X: https://x.com/AmosMeiriAntSeed X: https://x.com/AntSeedAIAntSeed Website: https://antseed.com/Partners: Better than Banks. Transparent capital efficiency earning the highest yields in DeFi. Learn more here: https://infinifi.xyz/---Dinari - Over 230 1:1 backed tokenized stocks, ETFs & more with dividends. US-based SEC transfer agent. Available on 5+ chains & via API. https://dinari.com/---Relay is the fastest and most reliable way to swap any token on any chain. Learn more here: https://relay.link/bridge---Zama is an open source cryptography company that builds state-of-the-art Fully Homomorphic Encryption (FHE) solutions for blockchain.Learn more here: https://www.zama.org/---Trezor is the creator of the first-ever hardware wallet. Securing crypto for 2M+ users worldwide. 100% open source. Learn more here: https://affil.trezor.io/aff_c?offer_i...---
Wie hat dir die Folge gefallen?Gut
My guests today are Gavin Uberti and Rob Wachen, the founders of Etched. A few years ago, when they set out to build a better AI chip than the largest companies in the world, almost everyone I called told me it could not be done. They have since done it, taping out a working chip on their first attempt and becoming the first hardware company founded after ChatGPT to do so. They already have more than a billion dollars of customer demand for their first product, and have raised eight hundred million dollars to build it. Etched builds chips and systems designed to run AI models faster and at lower cost. They started the company in 2023, and that product is a complete rack for inference, the chip along with the boards, the power delivery, the interconnects, and the manufacturing to produce it all. We talk about the technical bets behind their architecture, how they hired industry legends and paired them with elite 22 year-olds, and why they believe inference will become one of the largest markets in the world. I think you will find the story of what they have built hard to forget. Please enjoy my conversation with Gavin and Rob. For the full show notes, transcript, and links to mentioned content, check out the episode page here. ----- Become a Colossus member to get our quarterly print magazine and private audio experience, including exclusive profiles and early access to select episodes. Subscribe at colossus.com/subscribe. ----- Ramp's mission is to help companies manage their spend in a way that reduces expenses and frees up time for teams to work on more valuable projects. Go to ramp.com/invest to sign up for free and get a $250 welcome bonus. ----- Trusted by thousands of businesses, Vanta continuously monitors your security posture and streamlines audits so you can win enterprise deals and build customer trust without the traditional overhead. Invest Like the Best listeners get a special offer of $1,000 off Vanta when you go to vanta.com/invest. ----- WorkOS is the infrastructure B2B and AI-native companies use to sell to enterprise. It covers everything enterprise security requires: SSO, SCIM, RBAC, Audit Logs, AI governance, and more. Trusted by 2,000+ fast-growing companies, including OpenAI, Anthropic, Cursor, and Vercel. ----- Rogo is the AI platform for finance. They're building agents for Wall Street that are trained to understand how bankers and investors actually do work: from diligence and modeling, to turning analysis into deliverables. To learn more, visit rogo.ai/invest. ----- Ridgeline has built a complete, real-time, modern operating system for investment managers. It handles trading, portfolio management, compliance, customer reporting, and much more through an all-in-one real-time cloud platform. Visit ridgelineapps.com. ----- Editing and post-production work for this episode was provided by The Podcast Consultant. Timestamps: (00:00:00) Welcome to Invest Like The Best (00:02:07) Gavin Uberti and Rob Wachen (00:03:54) Two 21-Year-Olds Taking on NVIDIA (00:07:52) The Two Technical Bets Behind Their Architecture (00:14:15) Why Inference Becomes the Biggest Market (00:20:23) Rob and Gavin's Origins Stories (00:28:38) How They Recruit Industry Legends (00:36:30) Moving a Dozen Engineers to Bangalore for Six Months (00:38:01) Speed Wins (00:43:58) Getting More Concurrency Out of Every Megawatt (00:52:44) Vertical Integration (00:57:43) Hardest Obstacles to Overcome (01:01:09) Raising The Largest AI Chip Series A Ever (01:06:29) TSMC (01:13:20) Designing Gen 2 for Gigawatt-Scale Production (01:16:42) Why Machines Don't Think Like People (01:20:03) A Year of Compute Compressed Into a Month (01:23:44) The Trillion-Dollar Data Center (01:26:19) The Kindest Thing
When AI is at its best, the conversations can feel uncanny—almost magical in their accuracy, relevance, and speed. For that you can thank the AI agents that work together behind the scenes to search, reason, and sift through all your content to get you what you need to do your job. We talk with Jongmin Baek and Marta Mendez, two Dropbox machine learning engineers, about building conversational AI that's helpful, useful, and grounded in your team's shared context, so you can spend more time on the work that really matters. ~ ~ ~ Working Smarter is brought to you by Dropbox. Find, organize, and share your work—all in one place—with context-aware AI from Dropbox. You can listen to more episodes of Working Smarter on Apple Podcasts, Spotify, YouTube, Amazon Music, or wherever you get your podcasts. To read more stories and past interviews, visit workingsmarter.ai This show would not be possible without the talented team at Cosmic Standard: producer Ben Montoya, sound engineer Aja Simpson, technical director Jacob Winik, and executive producer Eliza Smith. Special thanks to our illustrator Fanny Luor, marketing consultant Meggan Ellingboe, and editorial support from Catie Keck. Our theme song was composed by Doug Stuart. Working Smarter is hosted by Matthew Braga. Thanks for listening!
Dakin Campbell, AI Finance Reporter, talks with TITV Host Akash Pasricha about the former AWS chief taking the helm of Helix Infrastructure Partners, a $10 billion data center company aiming to compete in the AI infrastructure race. We also talk with The Information's Catherine Perloff about Anthropic shifting its pricing model with Amazon to tokens, which will increase costs for the tech giant. We then chat with Laura Bratton about how Anthropic's new Claude integration for Slack is sparking cannibalization concerns among Salesforce employees. Finally, we get into OpenAI's secret internal optimizations that cut model inference costs by more than half with our AI reporter Stephanie Palazzolo.Articles discussed on this episode: https://www.theinformation.com/articles/amazon-pay-anthropic-technology-new-dealhttps://www.theinformation.com/articles/new-kkr-venture-hunts-deals-clear-data-center-logjamhttps://www.theinformation.com/newsletters/ai-agenda/openai-discovers-new-way-cut-inference-costs-halfhttps://www.theinformation.com/articles/salesforce-employees-worry-anthropics-invasion-slackSubscribe: YouTube: https://www.youtube.com/@theinformation The Information: https://www.theinformation.com/subscribe_hSign up for the AI Agenda newsletter: https://www.theinformation.com/features/ai-agendaTITV airs weekdays on YouTube, X and LinkedIn at 10AM PT / 1PM ET. Or check us out wherever you get your podcasts.Follow us:X: https://x.com/theinformationIG: https://www.instagram.com/theinformation/TikTok: https://www.tiktok.com/@titv.theinformationLinkedIn: https://www.linkedin.com/company/theinformation/Chapters:00:00 - Introduction01:13 - Adam Selipsky Leads $10B AI Infrastructure Play12:56 - Anthropic Hits Amazon With Token Pricing Shift20:46 - Salesforce Staff Clashing Over Claude in Slack31:11 - OpenAI Sneaks Out 50% Inference Cost Drop
Plus: Qualcomm to acquire AI software firm Modular in $3.9 billion stock deal. And Zoox debuts redesigned robotaxi for large-scale production. Julie Chang hosts. Learn more about your ad choices. Visit megaphone.fm/adchoices
Explore how the latest advancements in AI are shifting from traditional training to inference-focused efficiencies, and how companies like Adaptation Labs are pioneering adaptive, full-stack AI solutions that democratize control across industries.Key topics:The evolution from compute-heavy training models to efficient inference layersHow inference costs are changing despite increasing AI demandThe role of adaptive, gradient-free learning in democratizing AI customizationChallenges with the last 5% reliability gap and continuous learningThe importance of full-stack optimization—from data to interfaces in AI systemsFuture trends: decentralized AI, edge computing, and ongoing innovationTimestamps:00:00 - Introduction to AI trends: scaling vs inference efficiencies01:01 - Sudip's background: Google Brain, DeepMind, and inference infrastructure01:34 - The rapid growth of foundation and large language models02:36 - Comparing traditional ML project timelines to large foundation models04:20 - The transformative potential of foundation models in enterprise and underserved communities05:33 - The shift from task-specific models to general-purpose foundation models07:07 - How inference costs have evolved: the rising demand vs falling per-token costs08:37 - The challenge of inference in trillion-parameter models and the move towards smaller, verticalized models10:14 - Factors driving high inference costs: model size, reasoning, agentic workloads12:13 - The probabilistic nature of inference and API pricing complexities13:07 - Variability in inference costs and demand in real-world scenarios14:14 - The autoregressive, sequential nature of LLM inference and system challenges16:45 - Cost implications of autoregressive inference and the move to more efficient, localized models18:18 - The motivation behind Adaptation Labs: democratizing AI control and customization19:47 - Adaptive, gradient-free continual learning and environment interaction21:26 - Co-optimizing full-stack AI: systems, interfaces, and models22:34 - How interface design impacts AI adoption and continuous learning23:55 - The evolution of techniques: from foundational training to open-source innovations26:18 - Handling the ‘last 5%' reliability challenge in enterprise AI deployments28:02 - The importance of system feedback and adaptive learning in coding and decision-making31:12 - Adaptive Data and AutoScientist: seamless data transformation and model co-optimization32:55 - Use cases: finance, low-resource languages, long context data34:13 - The role of inference techniques and creating high-quality data for customization36:10 - Future of adaptive, task-specific interfaces and continuous, real-time learning38:49 - Full-stack AI: data, models, interfaces, and their iterative feedback loops41:18 - The competition between fine-tuning and adaptive inference techniques43:29 - The origin of new inference techniques: industry labs, open source, and innovation hubs45:27 - The “last 5%” reliability gap: why it's critical and how dynamic learning can help48:27 - Hardware vs software optimization in AI systems and the future of systemic efficiency51:25 - Growing AI demand, hardware constraints, and the opportunity for systemic innovation52:48 - The shift from training to inference and decentralized AI models at the edge54:12 - Final thoughts: the evolving landscape and long-term AI innovationConnect with Sudip:LinkedInConnect with Nataraj:LinkedIn
How do you build AI that actually understands you and the work you do? It all starts with having the right context. We talk with Dropbox staff product manager Noorain Noorani and principal engineer Sean-Michael Lewis about the art of context engineering and how Dropbox connects to all the tools your team needs for work—so you get AI that works wherever you do. ~ ~ ~ Working Smarter is brought to you by Dropbox. Find, organize, and share your work—all in one place—with context-aware AI from Dropbox. You can listen to more episodes of Working Smarter on Apple Podcasts, Spotify, YouTube, Amazon Music, or wherever you get your podcasts. To read more stories and past interviews, visit workingsmarter.ai This show would not be possible without the talented team at Cosmic Standard: producer Ben Montoya, sound engineer Aja Simpson, technical director Jacob Winik, and executive producer Eliza Smith. Special thanks to our illustrator Fanny Luor, marketing consultant Meggan Ellingboe, and editorial support from Catie Keck. Our theme song was composed by Doug Stuart. Working Smarter is hosted by Matthew Braga. Thanks for listening!
Semiconductors have moved from the background of the technology stack to the center of the AI economy. What used to be a specialized industry discussed mostly by engineers and investors is now shaping the speed, cost, and strategic direction of modern computing.In this episode of TechSurge, host Michael Marks speaks with Stacy Rasgon, Managing Director and Senior Analyst covering U.S. semiconductors and semiconductor capital equipment at Bernstein Research. Stacy has spent years analyzing the chip industry across cycles, but argues that the current moment feels different in scale: AI demand has created an unprecedented scramble for compute, memory pricing has surged, and companies across the stack are being forced to rethink capacity, architecture, and capital allocation.The conversation explains the 4 different kinds of semiconductor cycles—supply, inventory, product, and demand — and why Stacy believes the industry is currently in a demand cycle of unusual magnitude. The discussion also unpacks the distinction between DRAM and NAND, why high-bandwidth memory is becoming strategically central to AI systems, and how the physical realities of wafer capacity and silicon area are constraining supply in ways the broader market often misses.Stacy and Michael also discuss the hardware economics behind the current boom, with Michael pressing Stacy on why compute remains so scarce and how companies are improving performance through packaging and system design. Michael then moves the conversation beyond market headlines to the core business questions: who is actually paying for this compute, which use cases are generating real revenue, and whether AI spending is creating durable economic value or simply shifting costs elsewhere. Together, these questions highlight two of the episode's clearest insights: coding may be one of the earliest AI applications with meaningful willingness to pay, and inference, not training, is the real test of whether the current buildout becomes a lasting business or just another expensive wave of infrastructure.Stacy explains the concentration of power among the major wafer fabrication equipment players, the rise of ASICs as a meaningful share of AI silicon, Broadcom's rapidly expanding AI opportunity, and the growing role of Chinese companies as new entrants, especially in memory and semiconductor equipment. Along the way, the conversation asks the defining question facing the sector: is this just another semiconductor upswing, or the first true supercycle the industry has seen? Stacy believes that this might be the biggest supercycle he has seen in his career.Sign up for our newsletter at techsurgepodcast.com for updates on upcoming TechSurge Live Summits and future episodes.Links:Stacy Rasgon on LinkedIn: https://www.linkedin.com/in/stacy-rasgon-6924963Bernstein: https://www.alliancebernstein.com/corporate/en/home.htmlReferences Mentioned During the DiscussionNVIDIA Blackwell Platform: https://www.nvidia.com/en-us/data-center/blackwell-platform/High Bandwidth Memory (HBM) overview from Micron: https://www.micron.com/products/memory/hbmDRAM overview from IBM: https://www.ibm.com/think/topics/dramNAND flash overview from IBM: https://www.ibm.com/think/topics/nand-flash-memoryFurther ReadingMcKinsey on the semiconductor industry outlook: https://www.mckinsey.com/industries/semiconductors/our-insights/the-semiconductor-industry-in-2025Semiconductor Industry Association: 2025 State of the U.S. Semiconductor Industry: https://www.semiconductors.orgNVIDIA on the Blackwell architecture and AI infrastructure roadmap: https://www.nvidia.com/en-us/data-center/blackwell-platform/Broadcom AI investor materials and infrastructure commentary: https://investors.broadcom.comASML on lithography and advanced chip manufacturing: https://www.asml.com/en/technologyMicron on HBM and AI memory demand: https://www.micron.com/products/memory/hbmChapters[00:00:00] — Highlights[00:00:26] — Welcome to the Episode[00:01:29] — Meet Stacy Rasgon[00:02:01] — Is This the First Real Semiconductor Supercycle?[00:05:33] — Inside the Strongest Memory Cycle in History [00:09:14] — Can Innovation Keep Up With AI Demand?[00:11:33] — Chiplets, Blackwell, and the New Economics of Compute [00:12:37] — What Could Signal the Cycle Is Slowing[00:14:26] — Vertical Integration at the Hyperscales [00:16:36] — The Difference between Apple and Meta[00:17:15] — What is Vertical Integration Being Done For?[00:18:15] — Will other bottlenecks develop as This Progresses? [00:21:13] — Oligopoly Pricing in the Market[00:22:22] — Any New Entrants into Memory?[00:23:46] — Why the Industry Must Pivot From Training to Inference[00:25:10] — Agentic Coding and the First Real AI Revenues[00:26:57] — Groq, Low-Latency Inference, and What GPUs Cannot Do Alone[00:29:28] —-Could The Smaller Companies All be Bought Up ?[00:30:19] — Why Semiconductor Equipment Matters More Than Ever [00:31:00] — How Semiconductor Equipment is Affected by the Cycle[00:32:55] — A Long Upcycle for Semiconductor Equipment Guys?[00:33:13] — The Big Five and the Rise of Chinese Equipment Players[00:34:24] — The Effects of Geopolitics[00:35:02] — Broadcom's Quiet AI Breakout[00:40:46] — ASICs vs GPUs and the Next Wave of Custom Chips[00:41:06] — Intel, Foundry Strategy, and the Long Turnaround[00:46:46] —-The Risks the Market May Still Be Underestimating[00:49:32] — Where Startups Still Have Room to Win[00:50:39] — What the Semiconductor Industry Could Look Like Next Year
The Information's Asia Bureau Chief Jing Yang breaks down the unique five-year lockup and partnership structure behind DeepSeek's massive $7.4 billion capital raise. AI Finance Reporter Dakin Campbell explains the financial underpinnings and balance sheet risks of Broadcom's $35 billion hardware financing backstop for Anthropic. Then, Nvidia Reporter Phoebe Liu shares exclusive data showing how the chip giant grew its AI inference market share to 74% over the past year. Finally, MNTN CEO Mark Douglas analyzes the consolidation wave driving Fox's $22 billion acquisition of Roku.Articles discussed on this episode: https://www.theinformation.com/articles/polymarket-kalshi-take-steps-block-fraud-ringshttps://www.theinformation.com/newsletters/the-briefing/rokus-timely-exit-salesforces-dealmaking-revvinghttps://www.theinformation.com/newsletters/ai-infrastructure/nvidias-share-ai-inference-chip-market-appears-risinghttps://www.theinformation.com/articles/deepseek-closes-record-7-billion-plus-funding-unusual-deal-structureSubscribe: YouTube: https://www.youtube.com/@theinformation The Information: https://www.theinformation.com/subscribe_hSign up for the AI Agenda newsletter: https://www.theinformation.com/features/ai-agendaTITV airs weekdays on YouTube, X and LinkedIn at 10AM PT / 1PM ET. Or check us out wherever you get your podcasts.Follow us:X: https://x.com/theinformationIG: https://www.instagram.com/theinformation/TikTok: https://www.tiktok.com/@titv.theinformationLinkedIn: https://www.linkedin.com/company/theinformation/Chapters:00:00 - Introduction 01:13 - Inside DeepSeek's Uncommon $7.4B Funding Round 13:52 - Polymarket and Kalshi Tech Steps to Curb Fraud Rings 22:28 - Broadcom's Risky $35B Move to Finance Anthropic Chips 30:53 - Nvidia Gathers Speed with Rising AI Inference Market Share 38:43 - Fox to Acquire Roku for $22B in Streaming Consolidation
In this episode of the Crazy Wisdom Podcast, host Stewart Alsop sits down with Larry Swanson, creator of the Knowledge Graph Insights Podcast, for their second conversation together. The two cover a wide range of interconnected topics, starting with a correction Larry makes about the true origin of the term "artificial intelligence," tracing it back to the 1956 Dartmouth Conference and its distinction from Norbert Wiener's cybernetics. From there, the conversation moves through the history and structure of knowledge graphs, ontologies, RDF (Resource Description Framework), and the W3C standards process, touching on concepts like the T-box, A-box, and C-box, as well as the 25th anniversary of the Semantic Web paper. Stewart and Larry also dig into the limitations of large language models — particularly around reasoning, confabulation, and what Larry describes as "cognitive surrender" — and why symbolic AI and knowledge engineering may hold answers that the neural network world hasn't fully embraced. The episode also ventures into consciousness, panpsychism, Michael Pollan's ideas, and Stewart's own hands-on experience vibe coding a personal chatbot to replace functionality he feels he's lost with recent changes to Claude. Larry's podcast can be found at kgi.fm.Timestamps00:00 - Stewart introduces Larry Swanson; Larry corrects the record on AI's origin, distinguishing it from Norbert Wiener's cybernetics at the 1956 Dartmouth conference.05:00 - Larry discusses interviewing semantic web paper coauthors on its 25th anniversary; RDF's hidden ubiquity compared to SIM cards powering everything invisibly.10:00 - Knowledge graphs explained through t-box terms, a-box assertions, and Dave McComb's c-box; IKEA's three-layer knowledge graph as a practical example.15:00 - Stewart connects metadata complexity to AI needs; faceted search explained as c-box attributes driving product filtering experiences.20:00 - RDF 1.2 reification standards discussed; W3C's rigorous recommendation process powering governments and enterprises worldwide through collaborative standards.25:00 - Cyc project examined as influential "successful failure"; Pat Hayes bringing description logic into semantic web; LLMs lacking true reasoning capability.30:00 - Epistemological fault lines between human and computer intelligence; cognitive surrender paper reveals no intelligence threshold protects against AI manipulation.35:00 - Stewart's Claude regression problem drives chatbot vibe coding quest; small language models and domain-specific approaches explored as alternatives.40:00 - Consciousness discussion through Michael Pollan's panpsychism lens; language versus cognition disconnect revealing LLMs as pure token-stitching without genuine thought.45:00 - Context graphs as purpose-built knowledge graphs for AI; Stewart's planning agents versus coding agents architecture and ground truth verification problem.50:00 - Docs-as-code versus code-as-docs paradigm shift; knowledge graphs as universal verifiers against validated facts; RDF 1.2 enabling provenance and degrees of certainty.55:00 - Jessica Talisman's Knowledge Graph Academy recommended for onboarding; kgi.fm podcast shared; knowledge representation community needs better abstraction for wider adoption.Key Insights1. The term "artificial intelligence" was not a marketing gimmick but was coined deliberately at the 1956 Dartmouth Conference to distinguish the work of John McCarthy from Norbert Wiener's cybernetics. The two camps represented genuinely different approaches, and the AI label was a form of intentional intellectual branding rather than empty promotion.2. The semantic web, often called the most successful failure in technology history, has quietly embedded itself everywhere despite never achieving its original vision. Technologies like RDF power metadata standards inside every Adobe product and form the invisible backbone of government systems, enterprise data infrastructure, and cultural heritage organizations worldwide.3. Knowledge graphs are best understood as an ontology combined with all the instances that populate it. The distinction between things and strings, popularized by Google in 2012, captures the core idea that knowledge representation is about concepts as distinct from the labels we give them.4. The t-box, a-box, and c-box framework offers a practical model for understanding knowledge architecture. The t-box holds terminology and concepts, the a-box holds assertions about specific instances, and the c-box manages the attributes, taxonomies, and controlled vocabularies that sit between them and enable things like faceted search.5. Large language models produce fluent, convincing output but lack genuine reasoning, epistemological grounding, or judgment. Research on cognitive surrender shows that even people who understand how LLMs work are still susceptible to being misled by their fluency, meaning intelligence and awareness offer no reliable protection against being deceived.6. The gap between language and cognition matters deeply when evaluating AI. Evidence from people with aphasia shows that thinking can occur without language, which suggests LLMs, being purely language-based systems, are missing a fundamental layer of cognition that cannot be recovered through more tokens or better training.7. Knowledge graphs and RDF-based representation are well suited to the problem of verification and grounding in AI systems. Rather than relying on vectorized embeddings of language, a knowledge graph can store validated, provenance-tracked facts with degrees of certainty, making it a natural foundation for building trustworthy AI applications.
Brendan Burke says now is the time for tech, and Intel (INTC) has a growing role in the AI buildout. He believes current CEO Lip-Bu Tan will turn Intel into a core collaborator with AI hyperscalers that serves as a compelling foundry alternative to TSMC (TSM). Brendan also expects Intel to serve as a strong inferencing and CPU manufacturer as spending for AI accelerates. ======== Schwab Network ========Empowering every investor and trader, every market day.Subscribe to the Market Minute newsletter - https://schwabnetwork.com/subscribeDownload the iOS app - https://apps.apple.com/us/app/schwab-network/id1460719185Download the Amazon Fire Tv App - https://www.amazon.com/TD-Ameritrade-Network/dp/B07KRD76C7Watch on Sling - https://watch.sling.com/1/asset/191928615bd8d47686f94682aefaa007/watchWatch on Vizio - https://www.vizio.com/en/watchfreeplus-exploreWatch on DistroTV - https://www.distro.tv/live/schwab-network/Follow us on X – https://twitter.com/schwabnetworkFollow us on Facebook – https://www.facebook.com/schwabnetworkFollow us on LinkedIn - https://www.linkedin.com/company/schwab-network/About Schwab Network - https://schwabnetwork.com/about
Artificial intelligence is moving faster than ever, but as AI models continue to grow in size and complexity, the challenges surrounding inference performance are becoming impossible to ignore. In this week's podcast, ElastixAI CEO Dr. Mohammad Rastegari and I chat about how we can overcome those challenges and why a different approach to AI infrastructure is necessary for the next generation of AI innovation. We also explore the key bottlenecks limiting inference performance, how ElastixAI is tackling these issues, and why FPGAs are emerging as a compelling platform for accelerating large language model inference.
Stanislas Polu is Co-Founder & CTO of Dust — the enterprise AI agent platform used by 51,000 workers at 3,000+ companies. Before Dust, he spent three years on OpenAI's research team under Ilya Sutskever, working on mathematical reasoning in language models, and prior to that was an engineer at Stripe. He brings a rare combination of frontier AI research and product-building experience to the enterprise agent space.MCP, Agents & the $40M Bet on Multiplayer AI // MLOps Podcast #384 with Stanislas Polu, Co-Founder & CTO of Dust
Vamshi Ambati has spent more than two decades in AI, through the symbolic era, statistical era, and the neural wave we're experiencing today. A CMU PhD, founder of LatentStructure and Predera (which was acquired), now an investor at Virama Ventures, he's one of the sharper voices on what's actually happening under the hood of the AI boom.We discuss a simple question: Who wins when models become cheaper and more abundant? And try to answer this by looking at how inference spend v/s compute spend is shifting, and why inference may become the biggest infrastructure opportunity of the next decade.Vamshi explains what actually goes into the cost of a token, why AI is simultaneously getting cheaper and more expensive, and why the inference market alone could reach $1.3 trillion by 2030. If you're building in AI or someone who wants a clear mental model of where this industry is headed, this conversation is for you. 00:00 - Trailer0:45 - How an AI researcher thinks after 20 years05:53 - Where enterprise AI adoption is headed08:35 - Drawing parallels between cloud and AI11:20 - If building is cheap, what's valuable?13:37 - Can computing get cheaper?16:41 - What is inference, really?22:22 - Why coding and customer support got eaten first?26:48 - Which technologies are overvalued and undervalued?29:56 - An accidental entrepreneur's journey33:15 - Why is healthcare slow to adopt technology?38:59 - Landing Walmart as a customer42:36 - Should founders build in services if product isn't visible?43:47 - Is Palantir a product company or a services company?44:15 - How to win as a forward-deployed company46:23 - What it takes to land large enterprise customers49:20 - Building sales muscles as a technical founder-------------India's talent has built the world's tech—now it's time to lead it.This mission goes beyond startups. It's about shifting the center of gravity in global tech to include the brilliance rising from India.What is Neon Fund?We invest in seed and early-stage founders from India and the diaspora building world-class Enterprise AI companies. We bring capital, conviction, and a community that's done it before.Subscribe for real founder stories, investor perspectives, economist breakdowns, and a behind-the-scenes look at how we're doing it all at Neon.-------------Check us out on:Website: https://neon.fund/Instagram: https://www.instagram.com/theneonshoww/LinkedIn: https://www.linkedin.com/company/beneon/Twitter: https://x.com/TheNeonShowwConnect with Siddhartha on:LinkedIn: https://www.linkedin.com/in/siddharthaahluwalia/Twitter: https://x.com/siddharthaa7-------------This video is for informational purposes only. The views expressed are those of the individuals quoted and do not constitute professional advice.Send us Fan Mail
The Twenty Minute VC: Venture Capital | Startup Funding | The Pitch
Roman Chernin is Co-Founder and Chief Business Officer of Nebius, one of the fastest-growing AI infrastructure companies in the world. Today, Nebius operates some of the largest AI compute clusters globally and serves leading AI labs, enterprises, and developers. Today, Nebius has a market cap of $57BN. AGENDA: 00:00 — Why AI Infrastructure Is Not a Bubble 05:00 — The Real Impact of Open Source on OpenAI & Anthropic 11:00 — Jevons Paradox: Why Cheaper AI Creates More Demand 13:00 — The Four Layers of AI Infrastructure Explained 19:00 — If Nebius Had 10x More Capacity Tomorrow 26:00 — The Shift from Training to Inference and Agents 31:00 — How Token Factory Cuts AI Costs by 70% 44:00 — Sovereign AI, Europe, and the Future of Model Building 49:00 — Competing Against Hyperscalers with 10x More Capital 59:00 — The Biggest Threat to Nebius Isn't Competition—It's Consolidation
✅ New autonomous agents. ✅ Canva designs made for you. ✅ Codex upgrades to make your business move. If you had your head down in spreadsheets this week, you missed some MAJOR AI upgrades that are available now. We track what's hot and what's not and break it all down on Fridays with our Friday Features. Autonomous Copilot agents, new Codex tools, Github CoPilot app and 7 more AI updates you should be using — An Everyday AI Chat with Jordan WilsonNewsletter: Sign up for our free daily newsletterMore on this Episode: Episode PageToday's Episode on LinkedIn: Thoughts on this? Join the convo on LinkedIn and connect with other AI leaders.Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineupWebsite: YourEverydayAI.comEmail The Show: info@youreverydayai.comConnect with Jordan on LinkedInTopics Covered in This Episode:OpenAI Codex Role-Specific Plugins LaunchMicrosoft Build Conference AI Feature ReleasesChatGPT Memory and Business Account UpgradesMicrosoft Flash Image Model for PowerPointCanva Integrated with ChatGPT and CodexGitHub Copilot Standalone Desktop App PreviewMicrosoft Autopilot Always-On Work AgentsOpenAI Models Now Available on AWS BedrockCodex Sites: AI-Built Internal Web AppsTimestamps:00:00 OpenAI's big money moves03:47 Explaining role-specific plugins09:02 Microsoft's new image model release11:09 Microsoft's AI strategy and Canva update14:23 Canva integration with ChatGPT16:56 GitHub Copilot's new canvas feature20:46 AI token subscription changes24:42 AWS adds OpenAI models to Bedrock28:25 Introducing OpenAI's CodeX Sites Feature32:07 Launch of OpenAI's New Plug-in34:16 Overview of podcast structureKeywords: Autonomous copilot agents, Codex tools, GitHub Copilot app, OpenAI Codex, ChatGPT business accounts, OpenAI enterprise, Microsoft Build conference, Microsoft always-on agents, AWS AI updates, Canva plugin, ChatGPT memory upgrade, Windows Codex integration, Microsoft Flash model, Enterprise apps integration, Role-specific plugins, Sales data analytics, Product design AI, Creative production AI, Investment banking plugin, Public equity investing, Data analytics plugin, Workspace admins, App permissions, Role-aware work agent, Financial research automation, Microsoft image generation model, PowerPoint AI integration, OneDrive AI features, Visual design creation, Canva app for ChatGPT, Canva MCP server, Agentic context carry, Full screen design preview, GitHub Copilot desktop app, GitHub Copilot Canvas, Agent-native command center, Parallel agent work tree, Code app interface, Model options in GitHub, Token usage limits, Subscription token subsidizing, Anthropic token efficiency, Amazon Bedrock, GPT-4, GPT-4.5, Small language models, Token reckoning, Security governance, Inference engine, Code app sidebar, Codex Sites, Internal dashboards, Project trackers, Interactive web apps, Shareable AI apps, Enterprise data connectors, ChatGPT Canvas, Automated workflow, Workplace authentication, Creative briefs repository.Send Everyday AI and Jordan a text message. (We can't reply back unless you leave contact info) Start Here ▶️Not sure where to start when it comes to AI? Start with our Start Here Series. You can listen to the first drop -- Episode 691 -- or get free access to our Inner Cricle community and all episodes: StartHereSeries.com Also, here's a link to the entire series on a Spotify playlist.
This lecture discusses the William Clifford's 1877 essay "The Ethics Of Belief", in which he makes and argued for the central claim "it is wrong always, everywhere, and for any one, to believe anything upon insufficient evidence." It focuses on the third section of his essay, titled "The Limits Of Inference" in which Clifford discusses conditions for having well-founded beliefs of matters we don't have direct experience of, for example matters of everyday life, science, or history. We inevitably rely upon the assumption that the future or present will resemble what we have experienced in the past To support my ongoing work, go to my Patreon site - www.patreon.com/sadler If you'd like to make a direct contribution, you can do so here - www.paypal.me/ReasonIO You can find over 4,000 philosophy videos in my main YouTube channel - www.youtube.com/user/gbisadler Get Clifford's The Ethics of Belief - https://amzn.to/41WkkYA
SUMMARY: After the first successful AI IPO of 2026, we dig into what makes the Cerebras WSE architecture unique in the market for fast inference. GUEST: Andy Hock, at Chief Strategy Officer at Cerebras AISHOW: 1033SHOW TRANSCRIPT: The Enterprise AI Show #1033 TranscriptSHOW VIDEO: https://youtu.be/ed2nVbOtZiASHOW SPONSORS:OutShift - “Scaling Out Superintelligence” The Internet of Cognition architectureShareGate - ShareGate Protect. Microsoft 365 Governance, we got this!Nasuni - Activate your data for AI and request a demoSHOW NOTES:OpenAI announces 750MW partnership with CerebrasCerebras and AWS partnershipCerebras announces IPOTopic 1 - Welcome to the show. Tell us about your background, and what you focus on today. Topic 2 - For anyone that's not familiar with Cerebras, give us an overview of the company, and especially an overview on the Cerebras technologies (e.g. Wafer-Scale Engine).Topic 3 - Cerebras' WSE architecture is different from many of the GPU or GPU-like architectures in the market today. Centralized vs. distributed architectures always have their tradeoffs. Walk us through the technical and economic value of the Cerebras architecture.Topic 4 - Congratulations on the recent IPO (raised $5.55B). Let's use that as a point in time vs the previous planned IPO. How has the market changed in that timeframe, and how has the Cerebras position changed? Topic 5 - Cerebras (today) offer both WSE hardware, and Cerebras Cloud (API) - very different GTM paths. Can we expect both of those to stay top priorities, or have the market dynamics shifted such that the priorities shift more towards the WSE business - as we're seeing OpenAI, AWS and other engagements announced?Topic 6 - Is Cerebras a training and inference company, or are the economics of inference significantly different enough that it needs to be the sole focus of the company (for now)? Topic 7 - How much effort is it for any company to add support for the Cerebras chips if they have previously been using other architectures?Topic 8 - An IPO is a major milestone for any company, but the markets will now look for your future story. How do you see the AI market evolving over the next 2-5 years, and what are some things that people aren't understanding yet about how it will evolve?FEEDBACK?Email: show @ the enterprise ai show dot comeBluesky: @TheEntAIShow.bsky.socialTwitter/X: @TheEntAIShowInstagram: @TheEntAIShow
In this episode of Alexa's Input (AI), I sat down with Rob Shaw from Red Hat to talk about how AI inference evolved from a simple model serving problem into a large-scale distributed systems problem.We explored the infrastructure shifts behind modern LLM serving, including how vLLM and PagedAttention changed the economics and efficiency of inference, why KV cache management became one of the most important bottlenecks in production AI systems, and how orchestration layers like llm-d are emerging to coordinate distributed inference.We also discuss:how LLM inference differs from traditional model serving runtimesKV cache, prefix caching, and cache-aware routingwhy throughput and latency became major infrastructure challengeslong-context agents and repeated inference callsdistributed inference on Kubernetesintelligent routing, flow control, and load balancingprefill/decode disaggregationenterprise AI deployment realitiesvLLM has become one of the most important open-source projects in AI infrastructure, and llm-d represents a newer shift toward treating inference as a coordinated distributed system rather than just a single runtime problem.If you want to better understand the systems layer beneath modern AI applications, this episode is a deep dive into where inference infrastructure is heading next.General Podcast LinksWatch: https://www.youtube.com/@alexa_griffithRead: https://alexasinput.substack.com/Listen: https://creators.spotify.com/pod/profile/alexagriffith/More: https://linktr.ee/alexagriffithLearn more about the host atWebsite: https://alexagriffith.com/LinkedIn: https://www.linkedin.com/in/alexa-griffith/Find out more about the guest at:LinkedIn: https://www.linkedin.com/in/robert-shaw-1a01399a/ Red Hat Articles: https://developers.redhat.com/author/robert-shawGithub: https://github.com/robertgshaw2-redhat ResourcesvLLM Website: https://vllm.ai/vLLM GitHub Repository: https://github.com/vllm-project/vllmllm-d Website: https://llm-d.ai/llm-d GitHub Repository - https://github.com/llm-d/llm-d KeywordsAI inference, VLLM, LMD, distributed inference, GPU optimization, open source AI, Kubernetes, multi-cluster deployment, AI infrastructure, enterprise AI AI infrastructure, Kubernetes, model optimization, speculative decoding, mixture of experts, AI deployment, performance tuning, AI systems, neural network scaling Key TopicsEvolution of vLLM and llm-dDistributed inference and routingGPU utilization and performance optimizationOpen source AI infrastructureEnterprise deployment challenges and solutions Standardization in Kubernetes for NIC exposurePerformance optimizations: quantization and speculative decodingMixture of experts architecture and parallelism strategiesFlow control and request scheduling in AI systemsEmerging hardware for AI inference, Cerebras processorReinforcement learning and AI system supportModular architecture of vLLM and ecosystem projects
Modern work can be frustrating and chaotic—if you don't have the right tools. From context engineering to multimodal search, go behind the scenes and hear how Dropbox engineers are building AI that actually understands you, so you can focus on the work that matters most. If you're new to Working Smarter, we've travelled from the F1 track to the bottom of a lake, and heard real stories from chefs, doctors, lawyers, and founders about how AI is helping them do more of what they love about their jobs. But in our third season, we're talking to the people behind the tools—the engineers and product leaders building helpful, time-saving AI features into the Dropbox experience you already know and trust. You'll hear all about their work on agents, inference, security, and, of course, how the people building AI use AI themselves. ~ ~ ~ Working Smarter is brought to you by Dropbox. Find, organize, and share your work—all in one place—with context-aware AI from Dropbox. You can listen to more episodes of Working Smarter on Apple Podcasts, Spotify, YouTube, Amazon Music, or wherever you get your podcasts. To read more stories and past interviews, visit workingsmarter.ai This show would not be possible without the talented team at Cosmic Standard: producer Ben Montoya, sound engineer Aja Simpson, technical director Jacob Winik, and executive producer Eliza Smith. Special thanks to our illustrator Fanny Luor, marketing consultant Meggan Ellingboe, and editorial support from Catie Keck. Our theme song was composed by Doug Stuart. Working Smarter is hosted by Matthew Braga. Thanks for listening!
Send us Fan Mail*How do you forecast an event that has never happened before?*How do you forecast an event that has never happened before?The recent closure and reopening of the Strait of Hormuz are unique events. For events like these, traditional risk models lose their statistical basis: repetition. Alexander Denev returns to the podcast to show how causal models (Bayesian networks) let us reason about rare events despite this limitation.In this episode, we cover:- Why value-at-risk and other correlation-based models break exactly when you need them most- How a causal structure can "hold in time"- Building scenarios with LLMs - benefits, drawbacks, and lessons learned- Historical analogy as a modeling tool: Bosphorus, Hormuz, and more- A three-way robustness test for any Bayesian network- How the model's call held up: a ceasefire, a still-closed strait, and lasting infrastructure damage keeping oil elevated"History doesn't repeat itself, but it rhymes."------------------------------------------------------------------------------------------------------Video version available on the Youtube: https://youtu.be/FzKy2ws-7qsRecorded on May 29, 2026 in London, UK.------------------------------------------------------------------------------------------------------*About The Guest*Alexander Denev works at the intersection of quantitative finance, causality, and AI. He's the CEO of Turnleaf Analytics and the author of two books on applying Bayesian networks and probabilistic graphical models to finance and scenario analysis.Connect with Alexander:- Alexander on LinkedIn: https://www.linkedin.com/in/alexander-denev-66a25824/- Alexander's web page: https://turnleafanalytics.com/*About The Host*Aleksander (Alex) Molak is an independent machine learning researcher, educator, entrepreneur and a best-selling author in the area of causality (https://amzn.to/3QhsRz4 ).Connect with Alex:- Alex on the Internet: https://bit.ly/aleksander-molak*Links*Web- Alexander's LinkedIn post, Bayesian-network scenario for the Strait of Hormuz / Israel-Iran-US conflict: https://www.linkedin.com/posts/alexander-denev-66a25824_when-modelling-the-impact-of-events-that-share-7442892381668048896-JDs5/- Risk.net article, "Iran confusion makes the case for causal modelling": https://www.risk.net/our-take/7963361/iran-confusion-makes-the-case-for-causal-modellingBooks- Rebonato, R. & Denev, A. - Portfolio Management under Stress: A Bayesian-Net Approach to Coherent Asset Allocation (https://amzn.to/3vE6Jc1)- López de Prado, M. - Advances in Financial Machine Learning (https://amzn.to/3PXD8kH)- Molak, A. - Causal Inference and Discovery in Python (https://amzn.to/3VVK4m3)- Denev, A. - Probabilistic Graphical Models: A New Way of Thinking in Financial Modelling (https://amzn.to/3VQeLJm)- Pearl, J. & Mackenzie, D. - The Book of Why (recommended entry point) (https://amzn.to/4e0ATrZ)- Pearl, J. - Causality: Models, Reasoning and Inference (for advanced readers) (https://amzn.to/49zBKf5)- Rebonato, R. - Coherent Stress Testing: A Bayesian Approach to the Analysis of Financial Stress (https://amzn.to/3RC411e)*Perks & resources*
Stewart Alsop sat down with Michael Shackelford to discuss their experiences building applications through vibe coding—the practice of using AI to create software without traditional programming expertise. Stewart, who runs the AI Whispers community in Buenos Aires and hosts the Crazy Wisdom podcast (with over 660 interviews), shared how he went from teaching people prompt engineering to building his own video conferencing software as a Riverside.fm replacement, while Michael opened up about his year-long journey creating Genrupt Inc, an AI-powered content generation tool for e-commerce sellers. The conversation covered everything from the decline in quality of Claude's reasoning capabilities and how Chinese companies used distillation attacks to copy Anthropic's models, to the importance of spaced repetition systems for managing knowledge in the age of LLMs, with both sharing battle-tested prompting strategies like asking AI to "explain it to me in genius terms" and using deep research queries to reverse engineer how competitors build their products.Show Notes:- Dan Martell's book "Buy Back Your Time" was mentioned as one of the best business books for thinking about life and business- Check out John Vervaeke's "Awakening from the Meaning Crisis" for understanding relevance realization and why AI fundamentally cannot determine what's relevant to humans without being toldTimestamps00:00 Michael discusses being exhausted from getting his app ready for launch, working nonstop with AI to prepare landing page for podcast traffic driving beta signups05:00 Stewart explains starting AI Whispers in Buenos Aires after leaving OpenAI vendor company, meeting early adopters like Torin who was building mind-reading EEG technology10:00 Discussion of how corporations resist AI adoption due to political games and job security fears while some companies use AI as excuse for pandemic-era layoffs15:00 Stewart describes teaching workshops on using LLMs as linguistic tools rather than coding tools, noting technical people often lack humanities background needed for prompting20:00 Explaining chatbot wrappers, API calls, and how Anthropic's reasoning quality declined after Chinese distillation attacks copied their secret sauce developed with philosophers25:00 Technical discussion of model training, fine-tuning versus RAG for new information, and different approaches to updating AI knowledge beyond initial training30:00 Stewart describes building podcast recording software to replace expensive Riverside, struggling with syncing audio and video files across different computer clocks35:00 Discussion of critical factors in vibe coding, discovering unknown technical requirements, and how AIs don't automatically reveal missing information40:00 Stewart's reverse engineering process using deep research function to study competitors' hiring and technology stacks, separating planning agents from coding agents45:00 Prompting techniques including "explain like I know everything" and using spaced repetition systems to capture valuable prompts and technical knowledge50:00 Michael explains his Generux app for generating ecommerce content using Amazon review data analysis to inform high-converting listing images and videos55:00 Discussion of founder mentality involving self-delusion about project timelines, Michael working nine-plus hours daily for nine months on app development60:00 Comparing Amazon's expert software to prosumer software approach, discussing distribution challenges and future robotics applications for customized products65:00 Stewart demonstrates spaced repetition app for memory improvement and knowledge retention, explaining relevance realization problem that AI agents cannot solve without embodimentKey Insights1. Stewart Alsop started AI Whisperers in Buenos Aires after leaving his role at Invisible Technologies, which was OpenAI's largest vendor for RLHF work. He noticed that machine learning engineers at tech companies lacked the humanities background needed to properly interact with large language models, which are fundamentally linguistic tools. This led him to create weekly workshops teaching non-technical people how to use AI effectively, running events every Thursday for two years straight. The group attracted intense geeks from the start and eventually led to Stewart speaking right after Vitalik Buterin at DevConnect, marking a significant milestone for the community.2. Large corporations are resistant to AI adoption due to multiple factors including political dynamics within organizations and employees fearing job loss. Many companies that grew during the pandemic are now using AI as an excuse to downsize when the real issue is inefficiency from rapid expansion. Stewart observed that even technical people in machine learning often don't understand how to properly use AI tools because they lack linguistic and humanities training. The fundamental problem is educational, requiring companies to train people how to use these new tools while those same people resist learning them.3. Vibe coding has evolved significantly with Claude Code being a game changer that reduced the technical barrier to entry. Before Claude Code, developers needed substantial technical knowledge to work through constant doom loops and debugging cycles. The success of coding AI tools stems from thirty years of testing infrastructure that provides clear yes or no feedback on whether code works. This infrastructure doesn't exist in the same way for manufacturing, science, and other fields, which is why software became the dominant area for AI assistance initially.4. Claude's quality degradation over recent months resulted from multiple factors including distillation attacks by Chinese companies who reverse engineered Anthropic's reasoning capabilities. Anthropic had hired philosophers, sociologists, and psychologists to develop exceptional reasoning in Claude 4.5, but this was expensive to run. When Chinese models like Kimi copied these capabilities at one tenth the cost, and when mainstream users flooded the platform before Anthropic's planned IPO, the company had to reduce quality to manage computational costs. This represents a significant loss for power users who relied on Claude's superior reasoning abilities.5. Stewart built a podcast recording application to replace Riverside because he needed API access to automate workflows, which Riverside wanted one thousand dollars monthly to provide. The technical challenge involves syncing audio and video from local recordings on multiple computers with different clocks through a server, then merging them so voices match lip movements. This problem requires understanding complex timing issues across different network conditions and file formats. Stewart has been working through AI psychosis for months on this FFMPEG pipeline problem, illustrating how vibe coding still requires building intuition about technical problems even without traditional coding knowledge.6. The transition from expert software to prosumer software represents a major opportunity for AI-enabled tools. Expert software like Photoshop, Blender, and terminal interfaces have extreme complexity that intimidates beginners, but AI is making these capabilities accessible through natural language. The reign of specialists is ending as generalists with broad knowledge and curiosity can now build complete applications by leveraging AI to fill technical gaps. This shift particularly benefits entrepreneurs and founders who specialize in getting into difficult situations and figuring them out, even when they originally thought tasks would be easier than they turned out to be.7. Building applications with AI requires accepting massive time investments beyond initial estimates and developing strategies for overcoming knowledge gaps. Michael estimated his ecommerce content generation app would take months but spent nearly a year working over nine hours daily, while Stewart spent months solving audio-video sync issues. Success requires using tools like deep research to understand how competitors solve problems, maintaining separate planning and coding agents, and learning to ask the right questions. The key insight is that vibe coders can achieve ninety percent of functionality independently, but the final ten percent often requires understanding specific technical concepts that AI cannot intuit without proper context and domain knowledge.
Amidst the increasing urgency of powering data centers, a new solution has entered the mix: send them out to sea. In this episode, Shayle speaks to Garth Sheldon-Coulson, co-founder and CEO of Panthalassa. The company is building 85-meter steel "nodes" – taller than Big Ben – that it deploys into the deep ocean. These untethered, self-propelled nodes harness wave energy to power AI clusters, then beam their data back to land via satellite. The technology isn't without its fair share of logistic complications, but it nonetheless offers a pathway to powering the AI boom that's largely independent from grid or fuel constraints. Shayle and Garth cover topics including: - The physics and mechanics that power Panthalassa's nodes - The significance of building an autonomous fleet - The energy generation waiting to be tapped in the open ocean - The logistics and unit economics behind scaling Panthalassa's technology - Why deep-sea compute is well-suited for long-running workloads like inference and reinforcement learning - Catalyst: AI scaling pathways: On grid, on edge, off grid, off planet - Catalyst: How to build more hydropower - Latitude Media: Are Thiel-funded floating data centers enough to make wave energy pencil? - Open Circuit: Grid utilization vs expansion: The 100 GW debate - Latitude Media: What geothermal can learn from offshore wind's demise Credits: Hosted by Shayle Kann. Produced and edited by Max Savage Levenson. Original music and engineering by Sean Marquand. Stephen Lacey is our executive editor. Catalyst is brought to you by EnergyHub. EnergyHub helps utilities build next-generation virtual power plants that unlock reliable flexibility at every level of the grid. See how EnergyHub helps unlock the power of flexibility at scale, and deliver more value through cross-DER dispatch with their leading Edge DERMS platform, by visiting energyhub.com. Tune into Critical Capital, a brand new podcast from Crux and Latitude Studios. Hosted by Crux CEO Alfred Johnson, Critical Capital explores the interlocking forces powering clean and critical infrastructure. Join us every other Tuesday for in-depth conversations at the intersection of energy, government, finance, and global markets. Listen here, or wherever you get podcasts. Catalyst is brought to you by FischTank PR, an award-winning climate and energy tech, renewables, and sustainability-focused PR firm dedicated to elevating the work of both early-stage and established companies. Learn more about their PR approach and how they can support your company's messaging by visiting fischtankpr.com.
“OpenAI has only two AI accelerator compute vendors in production today, Cerebras and Nvidia,” Cerebras CEO Andrew Feldman says. Four days after Cerebras went public, Feldman joined Bloomberg Intelligence's Kunjan Sobhani to discuss the company's next chapter and the rapidly shifting AI infrastructure landscape. Feldman breaks down the OpenAI deal, the strategic AWS partnership around disaggregated inference and why Cerebras believes fast inference is becoming the industry's defining battleground. He explains how Cerebras evolved from building the world's largest chip to operating one of the fastest inference platforms, why disaggregated inference could reshape hyperscale AI deployments and how the company is navigating power, memory and data-center constraints. The episode also explores the competitive landscape beyond GPUs and Feldman's broader perspective on the next phase of AI compute.
Recorded live at Data Center World 2026, Data Center Frontier Editor in Chief Matt Vincent sits down with Phillip Koblence, COO of NYI and co-founder of Nomad Futurist, for the latest installment of Nomads at the Frontier. The conversation explores the accelerating realities of AI infrastructure buildouts, the industry's growing focus on community engagement, workforce shortages, and the shift toward inference-driven deployments following NVIDIA GTC 2026. Koblence discusses why major interconnection hubs and edge-adjacent urban facilities may become increasingly important in the inference era, the operational realities of deploying AI infrastructure in legacy carrier hotels like 60 Hudson Street, and why the industry can no longer remain invisible to the communities where it builds. Additional topics include: The continuing surge in digital infrastructure demand Why conference attendance reflects sustained industry expansion Power constraints and energy storage discussions emerging at Data Center World AI factories and the evolving economic role of data centers Workforce shortages across engineering and skilled trades Nomad Futurist's workforce development initiatives with Infrastructure Masons and I Am The Armed Forces The growing complexity and diversity of the data center ecosystem “Every element of everything within the data center has a full sub-vertical industry associated with it,” Koblence says during the discussion. “People would be surprised how large of an ecosystem is involved in creating the digital economy that exists today.” Listen now for a candid, fast-moving conversation on the state of AI infrastructure and the future of digital infrastructure development.
Ajit Ghuman is the co-founder of Monetizely, former VP of Product at Segment, and author of Price to Scale. In this episode, Ajit breaks down one of the biggest pricing challenges AI companies are about to face: what happens when software no longer supports employees — but starts replacing them entirely? If your company is building, pricing, or monetizing AI products, this episode will change how you think about per-seat pricing, buyer psychology, and the future of SaaS monetization. Why you have to check out today's podcast: Understand why per-user pricing may stop working as AI agents increasingly replace human workflows inside software products. Learn Ajit Ghuman's 3-part "Agentic Pricing Spectrum" for evaluating AI products based on autonomy, operational scope, and output-to-cost dynamics. Discover why buyers are suddenly comfortable with tokens, credits, and bundled AI pricing — even when they don't fully understand what those units actually mean. "Unless you understand what your market is, who your buyers are, what do they want... it's the only thing that I start with when I do any project." – Ajit Ghuman Topics Covered: 02:02 – Why Pricing Became the Most Direct Link to Customer Value. How pricing became the clearest connection between products, value, and business strategy. 06:29 – The AI Pricing Problem Nobody Has Fully Solved Yet. Why AI is forcing SaaS companies to rethink seats, tokens, outcomes, and margins. 07:38 – "Zero Human Companies" and the End of Per-User Pricing. Ajit explores a future where AI agents replace entire job functions — and asks the terrifying question: what happens when there's no user left to charge for? 12:30 – Why Cursor Still Charges Per User (For Now). A fascinating breakdown of AI coding tools, human "anchors," and why most AI products still can't fully move to outcome-based pricing. 16:51 – The Coming AI Commoditization Wave. Why Ajit believes agentic AI companies could rise — and collapse — dramatically faster than traditional SaaS businesses. 23:07 – Why Buyers Suddenly Accept Tokens, Credits, and Weird AI Pricing. Ajit explains how ChatGPT normalized token-based pricing — even though most buyers still don't fully understand what they're paying for. 26:00 – The Real Reason AI Pricing Feels So Chaotic Right Now. Inference costs are dropping, users are disappearing, and pricing anchors keep shifting faster than companies can adapt. 29:35 – The One Pricing Principle That Still Matters in the AI Era. Despite all the chaos around AI monetization, Ajit says successful pricing still starts with deeply understanding your buyers and their problems. Key Takeaways: "The anchor is still the human… but the moment the human disappears, per-user pricing starts breaking." – Ajit Ghuman "Agentic AI may compress 20 years of SaaS evolution into just a few years." – Ajit Ghuman People / Resources Mentioned: Cursor — Used as a real-world example of current AI pricing models Harvey AI — Referenced as an example of high-value AI transformation inside the legal industry Anthropic — Mentioned in relation to inference models powering AI tools OpenAI — Referenced throughout the discussion on tokens and AI pricing behavior Salesforce — Discussed in relation to potential future shifts away from per-seat pricing Zoom — Used as an example of changing pricing priorities during growth stages Connect with Ajit Ghuman: Website: https://www.getmonetizely.com/ LinkedIn: https://www.linkedin.com/in/ajitpalghuman/ Email: ajit@getmonetizely.com Connect with Mark Stiving: LinkedIn: https://www.linkedin.com/in/stiving/ Email: mark@impactpricing.com
On April 29th, the US Senate hosted a panel on the "existential threat" of AI and two of the four panelists worked for the Chinese government. One month earlier, Bernie Sanders and AOC introduced legislation imposing a federal moratorium on American AI data centers. On Bitcoin Policy Hour EP 38, Zack Cohen, Ken Egan, and Zack Shapiro unpack a new Bitcoin Policy Institute report by Sam Lyman exposing the CCP influence operation steering US AI policy. They also cover the Clarity Act vote in Senate Banking, the BRCA fight, and the Digital Asset Parity Act. Sam Lyman's BPI Report: https://www.btcpolicy.org/articles/foreign-influence-in-the-campaign-against-american-ai
Take the 2026 AI Engineering Survey and get >$2k in credits and AIE WF tickets!This was recorded before Railway suffered a major GCP outage on May 19, despite being a multi-AZ, multi-zone mesh ring, with HA fiber interconnects between their Metal GCP AWS, because workload discoverability was unintentionally still tied to GCP. All has been resolved with a post-mortem.Railway did not start as an AI infrastructure company.It was founded in 2020 years before agents became the default way people thought about deploying software. Jake Cooper, formerly at Bloomberg and Uber, started Railway with a simple obsession: the activation energy to ship something to production should be near zero. Push code, get a URL, iterate. No Docker files, no Kubernetes manifests, no Ansible scripts stacked on Ansible scripts.For years, this was a slow grind. Railway spent its first 18 months hand-acquiring its first 100 users with Jake personally greeting every Discord signup on a second monitor.Today, Railway has raised $124m and is growing very fast. A 35-person team supports 3 million users, adding roughly 100,000 signups a week. Their bare metal data centers have a 3-month payback period vs. renting in the cloud, with 70% margins funding aggressive cloud bursting when needed. The servers they own have actually appreciated in value as RAM prices have climbed basically meaning the value of their hardware now exceeds the capital they've raised.From rebuilding Railway's network overlay over a weekend to moving the vast majority of workloads onto its own bare metal data centers, Jake Cooper is trying to build a new cloud for an agent-native world. In this episode, Railway's founder and “conductor” joins swyx and Alessio to unpack why the next era of software infrastructure is not just “Heroku but newer,” what agents need that humans did not, and why the old deployment loop of Git, PRs, CI/CD, and static cloud resources may be heading for a rewrite.We go deep on Railway's infrastructure stack: own-metal data centers, three-month cloud payback periods, cloud bursting, data center debt, Railpack, Nixpacks, Temporal, feature flags, Central Station, content-addressable filesystems, agent-safe production forks, and why the CLI may become more important than the canvas in an agent world. Jake also shares the founder journey behind Railway, how the company survived losing $500K/month, why it now serves millions of users with only 35 people, and why he believes the pull request is dying.We discuss:* How Railway went from a slow six-year grind to adding 100,000 users a week* How Railway thinks about agents as the next dominant software species* Why agents need version control, observability, compute, storage, and orchestration at 1000x scale* The economics of Railway's own-metal data centers and three-month payback* How Railway uses cloud bursting while scaling its own infrastructure* Why data center debt can be a better tool than venture debt for infra startups* Central Station, Railway's internal system for clustering customer feedback and incidents* Why responsible disclosure and over-communication matter for platforms* Why feature flags, progressive rollouts, and shadow traffic are essential for agents* Temporal's strengths, pain points, and why workflows matter for agents* Railpack, Nixpacks, Nix, and lazy-loaded content-addressable filesystems* Why “cattle, not pets” may change if you can clone the pets* Why Railway is building a new cloud from scratch instead of copying hyperscalers* The solo founder path, focus, writing, and how Jake thinks about company buildingRailway:* Website: https://railway.com/* X: https://x.com/RailwayJake Cooper:* LinkedIn: https://www.linkedin.com/in/thejakecooper/* X: https://x.com/JustJakeTimestamps00:00:00 Introduction: What Is Railway?00:02:07 Jake's Path to Railway00:06:13 Railway's Six-Year Growth Story00:08:52 Rebuilding the Business After the Free Tier00:11:17 Agents as the Next Software Platform00:13:29 Railway's Infrastructure Philosophy00:15:42 Bare Metal, Cloud Economics, and the Compute Crunch00:17:22 Cloud Bursting and Five-Cloud Networking00:20:20 Data Center Debt and Infra Financing00:23:31 Data Centers in Space00:25:24 What Agents Need From Infrastructure00:28:24 CLIs, Canvas, and Agent-Native UX00:35:15 Central Station, Incidents, and Responsible Disclosure00:40:30 Safe Rollouts, SRE Agents, and Production Forks00:45:00 AI SRE, Specs, Code, and Tests00:48:24 Self-Replicating Infrastructure and the New Serverless00:53:18 Heroku, Temporal, and Workflow Engines01:04:07 Railpack, Nixpacks, and Lazy-Loaded Filesystems01:06:01 Coding Agents, Token Spend, and Roadmap Acceleration01:10:56 The Pull Request Is Dying01:12:28 Feature Flags and the Agent-Era SDLC01:16:15 Cattle, Pets, and Cloning Machines01:19:29 Solo Founder Lessons01:24:12 Focus, GPUs, and Building a New Cloud01:28:20 Closing ThoughtsTranscriptAlessio [00:00:00]: Hey, everyone. Welcome to the Latent Space Podcast. This is Alessio, founder of Kernel Labs, and I'm joined by Swyx, editor of Latent Space.Swyx [00:00:10]: Hey, hey, hey. Today we're in the studio with Jake Cooper of Railway.Alessio [00:00:14]: Conductor of Railway.Swyx [00:00:15]: Conductor at Railway. Yeah.Alessio [00:00:16]: Choo-choo.Swyx [00:00:17]: Do you actually have that anywhere, like on your business card?Jake [00:00:20]: We call some of our volunteer moderators conductors. I don't have a business card. We're not that big yet. At some point I will. I got handed a nice business card from the Supermicro folks, and I was like, “Damn, this is pretty official.”Swyx [00:00:30]: Business cards are coming back.Jake [00:00:32]: They're cool. They're hip. The conductor thing is good. We're trying to figure out what we want to call each other internally. Some people think it's super cringe and say, “You don't need a name for people internally.” Some people want to call each other something. We still don't have a really good one.Jake [00:00:55]: We've got New Railcrews, Trainiacs. Nothing has stuck yet.Swyx [00:01:00]: I like Trainiac. Trainiac sounds good. Railwayians. For those who don't know, what is Railway? Let's give people a crisp definition up front.Jake [00:01:09]: Railway is the easiest way to ship anything. You go to the canvas, or you talk with Claude, and you say, “Deploy a Postgres instance, deploy my GitHub repository, run this code,” and you're off to the races.Swyx [00:01:22]: You've got a nice animation on the landing page.Jake [00:01:24]: Thank you. None of my work, by the way. They don't let me touch the design stuff anymore.Jake [00:01:25]: We want to make it trivially easy not just to deploy things, but to evolve applications over time. Most tooling right now stacks entropy on top of entropy: Docker, Kubernetes, Ansible scripts, and all these other things. If we can version all of your software and keep track of all the changes, then we can make it trivial to clone environments, fork into a parallel universe, get copies of production data, get copies of any services, make changes, validate them, and collapse them back in without reproducing everything across a staging environment.The Railway Origin Story: From Uber Systems to a New CloudSwyx [00:02:07]: I was looking at your background: Bloomberg, Uber. Nothing immediately stands out as, “This guy is going to found the next great platform as a service.” What prepared you for Railway?Jake [00:02:21]: It was curiosity to keep going deeper. I started out on front-end stuff, working on Wolfram Mathematica and porting it over. Then I briefly moved to Bloomberg, then toward Uber and distributed systems, taking the Jump Bikes systems and moving them to a distributed system built on top of Cadence, the pre-Temporal Temporal.Swyx [00:02:44]: Which, by the way, I'm happy to talk about, pros and cons.Jake [00:02:48]: Totally.Swyx [00:02:51]: But let's do the Railway story.Jake [00:02:52]: It has been a continual step of wanting an experience. Whether it's walking up to a bike, unlocking it, and having it work frictionlessly, or something else, the depth required to make that happen follows from the experience. A lot of the work I do, and a lot of the team does, is in service of that experience. We fundamentally don't care how deep we have to go. We will swim to the bottom of the swimming pool to get the experience.Jake [00:03:17]: I don't have a physics PhD. I did an EECS degree. It has always been about figuring out the next step: how do we get there? That's what led to starting Railway for that experience and then moving all the way to bare metal data centers. I was adding patches to the kernel this week to get the experience there because I can see how much better it can be.Swyx [00:03:49]: Other patches to the Linux kernel this week?Jake [00:03:51]: Yeah. Not upstream. Our fork.Swyx [00:03:52]: That's a flex. Railpack? No, this is different. This is the OS on top of Railpack?Jake [00:03:57]: No, this is an actual kernel patch. It's always literally: what do we have to do to get that experience? Then figure it out. Anything is figureoutable.Swyx [00:04:10]: Would you send the patch upstream, or does it not fit other use cases?Jake [00:04:13]: Maybe. We have to work out the experience internally. It has to do with the storage layer we're building for some of the agentic stuff. Maybe it'll be useful upstream, but it's deeply useful for us internally.Open Source, Forks, and Non-Deterministic VersioningSwyx [00:04:29]: You mentioned open source before. How do you think about starting from open source, and then coding agents letting you do a lot more from forks of it?Jake [00:04:38]: GitHub's original sin is that it's almost a series of broken pointers. You have this thing, then you clone it, and now you've lost the whole upstream. How do we make it trivial for people to modify really small pieces of it?Jake [00:04:51]: We think of Git in a discrete sense: I've either made a change and merged upstream, or I haven't. What would it look like if it were percentage-based, a little more non-deterministic, or a stream of changes that users traverse as a percentage rolled out in general and then rolled all the way up?Jake [00:05:13]: We have the open-source kickback program and let you deploy templates because we want to make it trivial for people to version these shards over time. It solves a large problem around authentication, authorization, and security. NPM has a way to define, “Don't take any new packages.” The ideal end state is that you roll out progressively to users with the minimum impact zone and continue rolling up. JPMorgan should probably be the last one on the patch line, for all our sakes, because our money and livelihoods are there.Jake [00:05:53]: It's okay if Johnny Vibe Coder gets a broken patch because there's so much entropy in the system that the rubber has to meet the road at some point. You have to test at varying levels.The Long Grind: First Users, Free Tier, and Making the Business WorkSwyx [00:06:13]: I wanted to pull up this glorious chart, which is your usage or number of daily signups?Jake [00:06:22]: Daily signups, I think.Swyx [00:06:24]: You started six years ago. It was a slow grind, and now you're on a rocket ship. You say, “Don't doubt your fight and don't quit.” Maybe pick out certain points that were key inflections for the company.Jake [00:06:40]: At the start, it's about getting your first 100 users, hell or high water. We had a website and a support link. The support link was the Discord channel. I had notifications on with two monitors: the monitor I was working on and the other monitor with Discord. If anybody came in, I was immediately like, “Hey, how's it going?” It was rare, so getting those first 100 users to come back was the start.Jake [00:07:14]: Then you build a consultancy factory because users want all these things. You have to go back to the board and ask, “What is the actual product offering I want to build on top of this?”Jake [00:07:28]: VCs want charts that always go up and to the right, but in reality you don't necessarily want charts that look like that. For us, there have been periods of expansion where we add features to test use cases, and periods of compaction where we ask, “If the experience we have is good, how do we make it significantly better?” Maybe we strip out features that don't fit our ICP anymore.Jake [00:07:57]: The boom from 2022 to 2023 came from the free tier. Everybody under the sun was using it.Swyx [00:08:09]: A lot of Reddit bots and Discord bots.Jake [00:08:12]: And crypto miners. When you build an open product on the internet where anybody can sign up, the internet is a horrible place with so many things. You go through periods of asking, “How do I reach as many people as possible?” Then, “How do I fit the exact use case for the people who really matter and are really excited about this specific thing?”Jake [00:08:39]: Then there was a two-year period of making the actual business work. During the free-tier era, we were losing about half a million dollars a month.Swyx [00:08:59]: On a $20 million bank account.Jake [00:09:02]: On a $20 million bank account with maybe $50,000 a month in revenue. That's a horrible business. I don't know how anybody invested. But you have to go through it and say, “We have an experience people love, but the business has to work.”Jake [00:09:17]: There are two schools of thought. You can run the horrible business all the way up with bad margins, or you can go back and make it work. We've always wanted a super lean team. We're 35 people right now. It's very small.Swyx [00:09:36]: Supporting three million already?Jake [00:09:38]: Yeah. We're adding 100,000 users a week right now, so it's growing fast. We don't want to add headcount for the sake of headcount or throw bodies at problems. We want to build systems. It's hard to build systems during expansion because you're adding things to the system because people are asking for them or things are breaking.Jake [00:10:00]: We had to cut off the free users for a little while, rebuild the business, and make sure it worked. We want to reach as many people as possible because software is important. It's become difficult to create things in the physical world, so it's important to make it easy for people to build in the virtual world and have access to creation. But there are legs to that journey.Jake [00:10:30]: You can see divots in the charts. If you follow between 2025 and 2026, it's either summer or winter. People go on holiday with family.Swyx [00:10:50]: It affects that much?Jake [00:10:51]: Yeah. It's kind of B2C and kind of B2B. People are shipping constantly, then they stop. Our activation curve now shows more people activating on weekdays because we have more business users, so it smooths out over time.Agents as the New Interface to DeploymentSwyx [00:11:17]: Was there a point where you started prioritizing AI development or agent development?Jake [00:11:24]: We've prioritized agentic as a top-of-funnel thing. Over the last six months, we've deeply prioritized agentic as a mechanism to build and deploy things because we believe the curve is so steep and that is how people will build and deploy software.Jake [00:11:42]: It almost fundamentally doesn't matter whether this is dot-com or not because we're all on the internet anyway. If agents are going to deploy a bunch of things and we hit an inference wall at some point, we'll fix those problems. The dominant species over the next 10 years is that we've moved from assembly to C to C++ to JavaScript to words. You're going to need to close that loop.Swyx [00:12:13]: When you say this is dot-com, did you mean buying the domain, or the general case?Jake [00:12:17]: I mean the dot-com era, when companies had a huge run-up because people understood the internet was important. Then they hit bottlenecks, fundamental laws of physics, math didn't work, and everybody came back down to earth. But it didn't matter because the internet became so impactful. If you operate on a long enough time horizon, you should build these things anyway because you can see where it's going.Jake [00:12:45]: That's where I think a lot of agent stuff is. You get to a point where you're running thousands of agents in parallel. What is the inference cost? What is the compute cost? How do you make that efficient? How do you coordinate all this? We have issues coordinating humans; we don't even have good tooling for that. Now we have to figure out how to get agents to coordinate, safely version changes, and know when to raise their hand for someone to intervene. Otherwise it becomes an interrupt factory.Railway's Infrastructure Thesis: Network, Compute, Storage, and MetalSwyx [00:13:19]: Let's go right into the technical side. What are the core infrastructure or architectural beliefs of Railway that allow you to do what you do?Jake [00:13:29]: The primitives matter a lot for us. We need network, compute, storage, and orchestration around it. You need control over a lot of those things. We've talked a lot about how we don't really use Kubernetes because we want higher-order control to place workloads in very specific places.Jake [00:13:48]: The reason is that you have to be very efficient with agents: memory reuse and all these other things, or you're going to massively blow up your cost structure. Being able to rack and stack your own servers and build your own metal unlocks performance and cost. Experiences where you're running 1,000 agents in parallel are not massively cost prohibitive.Jake [00:14:13]: Token use and compute use are blowing up. Over time, those things have to get a lot more efficient. You can get a lot of margin to make those experiences solid by building your own metal. That's all in service of offering a differentiated experience to as many people as humanly possible.Swyx [00:14:51]: You have a data center in Singapore.Jake [00:14:53]: Yeah. We have two in every other region now. In Singapore, we're adding a second one in Q3.Swyx [00:14:58]: What's it like? I've never built a data center. Do you go to Equinix and say, “I want some slots?”Jake [00:15:05]: Yeah. Equinix. You basically go and say, “I want power and I want a cage.” They say, “Great, here's what it's going to be.” You rent the cage for a period of time, fill it with racks and servers, and hook up internet to it. That's all the pieces.Swyx [00:15:36]: Then you handle everything else.Jake [00:15:37]: You handle everything else.Swyx [00:15:39]: What's the math versus clouds doing it for you?Jake [00:15:43]: If we rented in the cloud, our payback period when we go to metal is about three months.Swyx [00:15:50]: Which is crazy.Jake [00:15:51]: It's nuts. That's four years of depreciated hardware. You're going to see a lot of this compute crunch because hyperscalers are buying up a lot of stuff. We're working directly with OEMs, resellers, and people building these machines: Supermicro, Dell, and others.Jake [00:16:11]: Upstream, there's a bunch of supply pressure. When we raised our last round, between deploying capital for servers and now, the amount of money we've raised is less than the amount of money we have in the bank plus the value of the servers because the servers have appreciated as RAM has gone up. It's nuts how valuable hardware has become.Jake [00:16:50]: If you look at hyperscalers, they deployed around $80 billion of capital expenditures this year, and next year will be more. That's a massive infrastructure build-out. You look at that and think it's crazy that they're spending way more than the Manhattan Project. But if every person is going to run dozens or hundreds of agents in parallel, you have no conceptual idea how much compute is required to make that experience happen, even if you're deeply efficient and sharing resources. And that doesn't even count inference.Swyx [00:17:22]: How do you plan the build-out? The growth chart is so vertical. Are you usually at 100% utilization as soon as racks are live? How far ahead are you planning?Jake [00:17:33]: We still maintain cloud presence for bursting. We work with AWS, GCP, and a few other clouds. We can rent, and then the moment we get space or power, we compact those workloads off the cloud. We started on the clouds, then built a system to migrate to our own metal. There's nothing that says you can't continually do that again, and that's exactly what we do. We never want to be compute constrained.Jake [00:18:09]: At the start of the year, we actually became compute constrained because one upstream provider wasn't able to give us quota at the rate we needed, and the hardware was slower. I spent a weekend rebuilding our entire network overlay so we could straddle five clouds: Oracle, AWS, ourselves, GCP, and one other one. We can do more than that now.Jake [00:18:38]: We got into a spot where we were trying to pack instances tight because we couldn't get enough compute. That led to a few reliability issues, which are now past us. I made a tweet pointing out that it's becoming harder and harder to acquire compute at the rate these models need to acquire compute. We got bit by it.Swyx [00:19:15]: How do you think about pricing knowing you might not have your own metal available at all times? Are you pricing assuming you need extra margin if you end up going into the cloud?Jake [00:19:26]: Because we've built out our metal data centers, our margins on metal are around 70%. We can deeply subsidize the cloud business if we want to scale at a reasonable rate. We have a few levers: metal, which makes the margins; cloud burst; debt to buy servers; and venture capital. It's an interesting operational problem: how much cash do we have, how much should we raise, how quickly can we deploy it, and can we scale revenue as quickly as we scale compute?Jake [00:20:05]: If we continue making it trivially easy for people to build and deploy, then the faster we close that loop and the more operationally excellent we are with capital, the faster the business can scale. It's almost a straight linear deployment rate.Financing Infrastructure: Hardware Debt, VC, and Operational LeverageSwyx [00:20:20]: I think infra startups raising debt is a tool people don't utilize enough or know enough about. What can you tell us about that? Is it secured against your CPUs?Jake [00:20:32]: It's secured against our hardware.Swyx [00:20:37]: What rates do you get? Who are the lenders?Jake [00:20:39]: We pay prime plus a spread, and we can refinance any of the debt as rates go down. The terms are pretty good. The unfortunate thing is that Twitter has no nuance, so people say, “Venture debt bad.” But as with all things, there are specific tools and areas where you can be deliberate instead of using one tool as a hammer. Venture capital is not the hammer for everything. You have to explore and figure out what works.Swyx [00:21:12]: VC is usually the most expensive financing you can get.Jake [00:21:15]: Yeah. I also think people think about VC incorrectly from a capital-raising perspective. Most people think, “How do I raise as much money as possible from whoever is probably the best I can get at that time?” That's close to right, but what we've tried to do is figure out what unfair advantage we can buy with that equity.Jake [00:21:34]: It's the most expensive equity you're going to give away at that point in time, assuming the company keeps getting better. How do you use it to work with someone stellar who complements you? In the seed stage, I had never started a company. Ray Tonsing had good advice, and I could text him all the time. He was really fast. Awesome.Jake [00:22:01]: Then with John and Erica at Unusual, they said, “You roughly know what you're doing building a product. We'll mostly leave you alone and be available for advice.” Amazing. Then we got to Series A and the business was an operational tire fire because we didn't know how to scale a business. Work with Erica, and Jordan is over at Redpoint, so bonus.Jake [00:22:28]: Now we've raised from TQ and FPV as we're moving into enterprises. Every step of the way, we've asked: who can we partner with at this specific time to unlock the next section of the journey? I don't know enterprise sales. As an engineer, I can eyeball what features we might need, and we have wonderful people internally who can help. But you want boardroom dynamics where everyone is aligned and asking, “How do we win this?” instead of bickering about strategy.Data Centers in Space and the Physics of ComputeSwyx [00:23:31]: You had a tweet about data centers in space. Why no data centers in space?Jake [00:23:37]: It's not “no data centers in space.” My hot take is that I think it is solvable. I've just never seen anybody solve it.Swyx [00:23:49]: You said, “How are you going to dissipate that much heat in a vacuum?” You're making a physics claim.Jake [00:23:55]: I haven't seen anybody prove how you're going to dissipate that much heat in a vacuum. It doesn't mean it's not possible. It just means nobody has brought it up yet.Swyx [00:24:05]: Astrophage.Jake [00:24:06]: I don't know what that is.Swyx [00:24:07]: The Martian thing. Okay, you're very logical.Jake [00:24:09]: It could work. A lot of people are putting the cart before the horse. They say, “We're going to put data centers in space.” Okay, but how? “We have time to figure it out.” It's like in The Martian where they ask how they're going to intercept something and say, “We'll figure it out.”Swyx [00:24:36]: Making a bet on human invention is weird because you blind trust that it can be solved. But with physics, there are first-principles bounds you can put on it. Maybe not. Maybe you're asking to travel time or break a fundamental thermodynamic law.Jake [00:24:57]: I don't know how VCs do this either. How do you know what's not possible and a grift versus what's possible but sounds completely insane? “We're going to put data centers in space.” Coin flip as to which it is, and I guess you'll know in 10 years. That's one cycle.What Agents Need: Versioning, Observability, and 1,000x ScaleSwyx [00:25:23]: Moving back to agents. The branching, fast spin-up, and orchestration you do feels like pre-work that happened to be exactly what agents want. What do agents want differently than humans?Jake [00:25:37]: They want the ability to version things. It's not that different; it materializes slightly differently. Agents want a way to test changes incrementally. Engineers have feature flags. Is there a reason agents can't use feature flags? I don't think so.Jake [00:25:54]: They want version control. Can we use Git or not Git? That one is up in the air. I think something outside Git will emerge for how we version these things over time. They need observability. You need to query what happened, when it happened, which steps failed, traces, logs, metrics, and all the rest. They need network, compute, and storage. They need to write files, save files, iterate on files, and snapshot file systems.Jake [00:26:25]: A lot of what humans needed is in line with what agents need. Branching and forking are not different; we're just moving 1,000 times quicker. It can look like you need something massively different, but what you need is something massively better than what existed. You need orchestration massively better than Kubernetes. You need networking probably better than Envoy. It goes all the way down the stack.Jake [00:26:55]: If the workload profile doesn't change so much as it gets massively compressed because you need thousands of these things, what assumptions change? etcd is going to melt. You need to replace it with something. You can go all the way down the stack and say, “That part has to change, that part has to change, and that part has to change.”Jake [00:27:19]: The interesting thing about the super-exponential curve is that you have to build systems where you can rip out those parts at any time because a new bottleneck might emerge. You get good at parallel agents, and a different part of the system breaks. So it's similar to what humans needed, but at 1,000x scale.Jake [00:27:55]: How do you do code review in the age of agents?Swyx [00:28:00]: You throw more agents at it.Jake [00:28:01]: You don't. But then who reviews for CVEs and all these other things?Swyx [00:28:07]: More agents.Jake [00:28:08]: And that's how we hit the inference wall. You can continually throw agents at the problem, but I think there's a limit to the number of agents you can throw at a problem.CLI, Agent Handles, and Closing the LoopSwyx [00:28:24]: You already had a CLI before it was cool. How is the shape of what you're exposing changing, if at all?Jake [00:28:28]: CLIs have always been cool. The CLI changes because we think about how to give Claude, Codex, ChatGPT, or any model a handhold.Jake [00:28:50]: A CLI is a single command: deploy, get logs, and so on. Things that were prohibitively annoying to humans are not annoying to agents. They're nice. If I handed you a CLI with 40 arguments and 600 flags, you'd think, “I'm never going to use all of this.” But if you hand it to an agent, it says, “This is excellent. I have so many handles to work with.”Jake [00:29:24]: If you're going to expose things to agents that way, you want as many handles as possible where they can get information, query dynamic information, and close the loop quickly. Most problems right now are about how to close the loop as quickly as possible. Where does the agent get stuck, and how can you remove that?Jake [00:29:49]: Telemetry is important. If you can tell where the agent gets stuck from the CLI and say, “12% of people deviate from the happy path because of this, and now I add this argument and drive it down to 2%,” you massively increase the rate of loop closure.Jake [00:30:03]: That's how we think about not just the CLI, but every point in the dashboard. It's a user journey: I hear about Railway. I get something deployed. I get my first green build or aha moment. I see an endpoint, logs, whatever. Then I iterate. The iteration loop is indefinite. The user wants to deploy a new thing, a Postgres instance, change code, and keep iterating.Jake [00:30:36]: If you focus on the iteration loops and what's blocking them from closing quickly, one thing we say internally is: you never want to be waiting on compute anymore. You always want to be waiting on intelligence. If you're waiting on compute, there's a bottleneck that needs to be destroyed because eventually that bottleneck becomes so large that another workflow emerges to change it.Jake [00:31:04]: We've built a product where you push code, build it, and so on. But I fundamentally believe the push-pull loop is going away. We'll get to a point where you make a small change in production, that change is versioned across your infrastructure, you're working alongside copy-on-write versions of your database and infrastructure, and then you merge it in and it's instantaneously live. That's the holy grail of loops. The push-pull-rebuild thing is a point of friction that we're removing entirely.Canvas as Output: Dashboards, Context Anchors, and HyperstructuresSwyx [00:31:43]: It's incredibly fast. If anyone hasn't tried it, that fast feedback is great. My hot take is that Railway was famous for its canvas, which visualizes your infrastructure and lets you manipulate it visually. But that was for humans. For the next phase of growth, Railway CLI is more important than canvas.Jake [00:32:05]: The canvas is funny because it's a mechanism to show changes over time. You're right that previously we used it a lot as an input. Moving forward, its goal is more like an output. You would go to the canvas, make changes, see them, and watch your infrastructure evolve. Now agents have access to the CLI and can make those changes. So the canvas becomes an output: what information does the human need at this moment to make suitable decisions about control requests? Do I approve this or not?Jake [00:32:57]: It also has to be an anchor for your context, a port in the storm. Think of it like layers in a file system. You start with a project, then drill down into services, then into a function or code, because you want to represent the entire thing not just in your head, but in the canvas. Other people can share that representation, think on the same wavelength, and move quickly.Jake [00:33:33]: A lot of organizations get in trouble as they scale because all the context lives in someone's head. “How does this microservice work?” “I have no idea; go ask this person.” Then you have whole categories of products built around context discovery. A lot of that melts away if you have a solid hierarchy and can infinitely nest services, code, context, and everything else all the way down. That's what lets you build these structures over time.Jake [00:34:18]: It's also what lets us build what I've called hyperstructures: things that are way bigger. You look at the Golden Gate Bridge and ask, “How did we build that?” There's a meme that we lost the technology. To some extent, yes, because the coordination that built those things evolved and changed. We lost some of the art of building structure as we jammed everything into Slack.Swyx [00:34:52]: But you jam everything in Discord.Jake [00:34:53]: Same point. It doesn't matter. It's message passing and interrupts, message passing and interrupts.Swyx [00:35:00]: So you're arguing there should be something better and more structured than Slack?Jake [00:35:04]: Yeah. For sure. I think Slack is awful, and Discord is awful too.Central Station: Context Routing, Support, and Incident ClustersSwyx [00:35:09]: This is the equivalent of my mom test. What have you done that has your solution to this?Jake [00:35:15]: Internally, we've built a tool called Central Station that aggregates all the context from our users. Every piece of feedback, every customer support item, everything gets aggregated into clusters. If an incident is brewing, we can determine how many users are affected and break off a discussion based on that.Jake [00:35:40]: That is more helpful than long-running channels where you're trying to decide which channel to put something in. If you can dynamically aggregate information and dynamically route it to the right person based on context, it works better. We know internally that these four people are close to networking. If we see a networking thing, we can drill it down to those four people. If it's with this part, we can look at the commits. This is no longer a manual process internally.Jake [00:36:13]: If you go to station or help.railway.com, that's why we built it. We wanted to scale with a massive amount of leverage by aggregating feedback.Swyx [00:36:27]: This is built in-house?Jake [00:36:28]: Yep.Swyx [00:36:29]: I remember helping out on this one with Angelo in 2023. You scale a lot with a very small team.Jake [00:36:38]: Yeah. We're about 10 times bigger now.Swyx [00:36:40]: You have your full developer code here? Very cool.Jake [00:36:44]: If you go to railway.com/stats, we expose this as a pub-sub-able thing. It's all real-time metrics. There's a way to get it as JSON somewhere if you care.Jake [00:37:01]: We're big on trying to build everything in public and talk about what we're working on. We've had issues in the past, and we'll say, “Here's how we're fixing these things.” We've gotten compliments and flak for incident reports. We're always trying to make them better and talk with people.Incidents, Disclosure, and Progressive RolloutsSwyx [00:37:20]: You had a big one recently. I liked that it was scoped to 3,000. You presumably used Central Station. Talk through what happened and how you address it internally as a team.Jake [00:37:38]: Internally, this one really sucked. It had to do with an upstream provider that didn't do the behavior it said it documented, which is unfortunate given they wrote the RFC for how the behavior should work. We rolled those things out, and Central Station caught it initially when a couple users said caches weren't invalidating. We turned it off immediately.Jake [00:38:03]: When you roll out to a large user base of three million people, you get a lot of disparate behaviors. We tested in staging and had tests, but we hit an edge case. We've hardened those systems, and now we can make that better. But it was a tough one.Swyx [00:38:39]: I always wonder how private disclosure is supposed to work if people find an issue. Are they supposed to contact you first? When you run a platform, these things will happen. What channels should people pursue to quietly resolve it before it becomes a bigger incident?Jake [00:38:59]: There's responsible disclosure. We err on the side of over-disclosing and letting you know something is wrong versus having your provider gaslight you. We've erred on sharing those things more publicly, even if they impact a small subset of users. That's a decision we've made internally. We have four values. One is honor. The honorable thing is to notify people to the widest degree at which they may have been affected or there was an issue, and then confront it head-on: why did it happen, what can we do better?Swyx [00:39:45]: Not the whole user base. That's because of incremental rollouts and other things?Jake [00:39:50]: Yeah. Progressive rollouts.Swyx [00:39:54]: That should be the norm at all large platforms.Jake [00:39:58]: It should. A variety of companies do this. There's the quote that Meta runs 10,000 different versions of Meta. To our earlier point about agents, they need the same thing. They need shadow traffic and all these other things. We've built so much ceremony around production being sacred that we need to make it trivially easy to test different behaviors in a safe environment. Then you can make mistakes in a safe environment.Safe AI SRE: Customer Agents, Forked Environments, and Production ParityAlessio [00:40:30]: Do you see a world where these things get automatically caught, not necessarily by your agent, but by your customer's agent? The cache invalidation issue seems easy to check if you know to look for it.Jake [00:40:44]: It's hard because to determine it, we almost need to hook into your observability infrastructure. That's why we have the template loop on the platform: so you can roll things out progressively. You can roll out to Johnny Vibe Coder initially, or push a shard that someone consumes at their own leisure. Or you can roll it out over weeks: 0.1% of people, 1% of people, early adopters, then all the way up. That's the non-deterministic version control we talked about earlier.Jake [00:41:30]: I believe that's where most things should go, because most companies end up building staged rollout systems in-house. It's the same thing built again and again at every company. There's a massive opportunity to consolidate developer debt.Alessio [00:41:45]: You should have a free tier. Model providers give free tokens if you let them use the data. You could give free compute if someone is the number-one shard that goes out and lets you plug into their observability.Jake [00:41:55]: We do that. That's why we talked about the impact on 3,000 people. We start with lower-impact people. Larger companies on the platform are last to receive those rollouts so they have a version of the platform that's deeply stable.Alessio [00:42:16]: I have three services, so I'm sure I get the first rollout. You can nuke my thing at any time. There are all these SRE agent companies. Observability people also want agents that fix upstream problems. You have your own agent in the canvas now. How do you see that playing out?Jake [00:42:39]: It's the stacking entropy problem. If you don't have primitives to make iteration in production safe, it becomes difficult. If you're an observability provider saying, “Here's the fix to this error,” assume 80% are good and make sense. But in the last 20% long tail of complex issues, if you let somebody stamp it, you create an opportunity for an incident.Jake [00:43:08]: That's why forked environments are important. People have staging, but it always drifts from production. You need primitives, workflows, and experience built first-party on the platform so you can fork any service at any point in time.Jake [00:43:33]: I think of the canvas as a sheet of transparency paper. The agent is a little guy you push up into the canvas. It should say, “I need to copy that service and that service so I can test these two things.” It gets a read-only copy of production. Anything that's PII gets marked as a transform when we clone the database, create a copy-on-write version, or read from it. Then the agent makes changes and asks, “Does this actually work?” as close to production as possible.Jake [00:44:22]: That's how close you have to be, or you get massive drift. The system becomes unstable. You see this with massive systems built on Docker for local, Kubernetes for production, and a specific thing for something else. That complexity slows developers and becomes unstable at scale, making it hard to iterate. We want to compress that way down and say, “As close to prod as possible is where we want to be.”From AISRE Skeptic to Agent BelieverSwyx [00:45:00]: I was texting Erica for questions, and she says you were originally not a believer in AISRE. Have you come around on it?Jake [00:45:10]: I flipped, but I'm still not a believer in AISRE if you don't have the primitives to make it safe. If you unleash AISRE on production infrastructure without safe primitives for copying volumes and making sure things are fine, it's going to nuke your production database. It's not a matter of if, but when. I'm a big believer in making those loops safe.Jake [00:45:33]: I was a deep AI skeptic until 2023. In 2024, I thought, “Maybe I can roughly make this thing do it.” In 2025, I thought, “Now I can hold this.” Over winter break, everybody came back saying, “It's almost impossible to hold this.”Swyx [00:46:01]: Did you see this on the Claude docs? CloudBot? OpenCloud?Jake [00:46:06]: It's gotten to a point where it's harder to hold it wrong than to hold it right. There's a scene in Avengers where Vision picks up Thor's hammer and says it's terribly well-balanced. It self-balances and works well. I'm a deep believer at this point that this will be the dominant species: assembly, C, C++, JavaScript, words.Swyx [00:46:35]: It feels like a big jump.Jake [00:46:37]: It is. But it's not like you abandon CPU-based discrete logic and move straight to fuzzy logic. You need both. Your skills should call code or applications or some static structure. You can use skills to distill what the procedure should be or how the code should act.Jake [00:47:02]: I'm coming to a thesis: you need three points. You need a clear spec defining the system, the code, and the tests. When you say it out loud, if you've been in engineering long enough, you're like, “Of course. That's an RFC, tests, and code.” But they all matter. Having them together lets them reinforce each other: the spec and tests match, but the code doesn't, so reconcile it. Or the tests and code match but the spec doesn't, so reconcile that. That's the iteration loop.Jake [00:47:41]: That's why you're seeing people talk about software factories, docs, and reconciliation. Some of that is architectural astronomy if you don't implement it, but that loop is where most things will end up.Swyx [00:48:07]: For listeners, we've been talking about this on the pod for three years: the holy trinity of specs and tests. Itamar Friedman from Qodo is the reference if people want to look it up.Self-Modifying Infrastructure and the End of Push-Pull-RebuildSwyx [00:48:18]: One thing I want to mention on the OpenCloud idea is self-modification. I don't know how Railway would support it, but I have my OpenClaw, and I just tell it it has the Railway CLI and can do whatever. In theory, whatever capabilities or new infra it needs, it can call the Railway CLI, provision it, and add it to itself. The agent can modify its own infra.Jake [00:48:45]: It's nuts. I have a loop set up where you put the Railway CLI on top of something that runs on Railway. You're authenticated as whatever the current box is, and you can make any changes to it. Then you call Railway deploy, and it deploys itself.Jake [00:49:04]: It's like: “I need to spin up this instance of this environment. I already exist in this environment. Excellent, I have access to a Postgres instance now.” That's where we want to go with agentic, self-replicating infrastructure. That's your loop: iterate in production. You continue making changes. If it works, merge it upstream. If it doesn't, throw it away.Jake [00:49:37]: How do you make throwaway copies trivial to spin up and super cheap? The era of “I have an AWS instance with four vCPU and 16 gigs of RAM” is going to get destroyed. If you do that for agents, you need a thousand of those machines. It's prohibitively expensive compared with what we've spent a ton of time figuring out: the atomic unit of deploy, whether you call it isolates, sandboxes, or something else. Only pay for what you use, spin up instantaneously, and close the loop as quickly as possible.Jake [00:50:15]: If the system can self-replicate safely and say, “This is my environment, I'm making these changes,” it can come back with, “Does this look good? This is a new state of infrastructure given this prompt. I think I've solved it.” Then you go back and say, “Actually, it looks different.” It does the loop again. Then you say, “Cool. Apply.”Swyx [00:50:38]: That's retroactively obvious, which is the most useful kind. Any other comments on agent deployment on Railway?Jake [00:50:51]: It's getting better every day. I'm on X or Twitter. You can always yell at me about the parts not working as well as they should, because plenty of things should work way better.The New Serverless: Stateful, Long-Running, Pay-for-What-You-Use LinuxSwyx [00:51:04]: At this stage, when people want massively or embarrassingly parallel compute, they usually talk serverless. I feel like there's a new serverless compared to the previous five years of serverless. You're in that new bucket. Do you have comparisons or philosophical differences you want to call out?Jake [00:51:31]: It's somewhere in between. It's the ability to run stateful, long-running workflows or executions.Swyx [00:51:42]: Vercel has Fluid Compute, Cloudflare has some container thing, Google has App Runner and others.Jake [00:51:55]: That's where everything is roughly going, and it's why we've been working on this for six years. We believe users need access to a computer: a box that speaks Linux. They need to deploy what they want. Other systems change the surface area of what you can build. For us, users need a computer and need to deploy anything they truly want. That's why we've focused on the primitives: network, compute, storage. If we give you those and expose them so you can run things indefinitely, that's where we believe it's going.Jake [00:52:43]: Twitter has no nuance, so everyone says “servers” or “serverless.” It's always somewhere in the middle: I want to run it for a long time, but I don't want to provision the resource statically or pay for things I'm not using. That's been our thesis from day one: pay only for what you use, run it indefinitely, and it is full Linux.Swyx [00:53:12]: That's why I like the naming of Fluid. It's fluid. Flexible.Heroku, Focus, and Carrying the Torch Without Becoming the PastSwyx [00:53:18]: Another milestone is the Heroku official deprecation. You're one of the presumptive new Herokus. “New Heroku” has been a category for as long as I've been in developer tooling. It's finally happening. What was that like? Any behind-the-scenes of, “This is the moment”?Jake [00:53:42]: You have people where you're like, “You were running stuff on here? You, as this company?” It's crazy that names you would know are running on it and now coming to us saying, “We want to move a lot of this off.”Swyx [00:54:00]: Any behind-the-scenes on why Salesforce let Heroku stagnate?Jake [00:54:05]: I can only guess. It's hard when it's not your business. Salesforce's business is to build a great CRM. That's their focus. Then you acquire a compute business as an offshoot. A lot of early Meta people talk about focus. Boz has a write-up about how in the early days of Meta they had no money, so they were forced to focus. Then they turned on the money tree and had no reason not to split their focus.Jake [00:54:52]: But that dilutes your product. You get offshoots where you ask, “Is this the focus of the business?” If it's not core, it languishes. A lot of companies get in trouble when they split focus because they're fighting a multi-front war, not just externally but internally for alignment. Where are we going? What are we doing? What is our purpose?Jake [00:55:24]: If you're Salesforce-built and mission-driven, you want to work on Salesforce. Heroku is off to the side. It's not core to the business. Getting resources, budget, focus, and alignment internally becomes hard. It was a matter of time.Swyx [00:56:06]: Kudos for them to call it out instead of leaving it unknown.Jake [00:56:12]: Their release was a little odd. They called it out, but they didn't say they were shutting it down. Behind the scenes, I think they issued messages to people saying they should close accounts and that they were going to deprecate and remove things over time.Jake [00:56:30]: It's crazy because some of my first deployment experiences were on Heroku. You start with dragging things into an FTP server, then you try to get a deploy working, and then it's Heroku. It was the on-ramp for us. But the wheel turns. New things emerge. We're happy to carry the torch for a lot of that. But we don't want to be the new Heroku. We want to be the way people build and deploy software, and ultimately the way people monetize software over time.Swyx [00:57:19]: It's still a big crown to be the new Heroku. There are 50 companies that fought for that.Jake [00:57:23]: Everybody is holding some portion of it. We're happy to support people and companies. The platform works differently. The game loop is similar, but we've been dogmatic about where these things are going: primitives, agents, fan-out. Some things fit; some workflows need to change. We have an approximation of Heroku pipelines with the environment system. It's exciting. We've got a ton of people we can support, and it's growing a lot.Temporal, Workflow Engines, and State MachinesSwyx [00:58:12]: I have one more technical question about Temporal. I've sold my shares. You're a power user and one of our earliest customers. I met you through Temporal. You built on Temporal. You have complaints. This may be the most neutral and informed conversation anyone will hear about Temporal without someone working at the company.Jake [00:58:39]: That's fair. I've used Temporal for almost 10 years because of Cadence at Uber.Swyx [00:58:52]: Give people a sense of what Cadence was at Uber.Jake [00:58:57]: Cadence was the precursor to Temporal. It powers trip actions, rides, when you rent a Jump bike or scooter or car. You're running workflows for a period of time and saying, “This ride will run indefinitely until it finishes.” You attach information: you paused in this zone, so add this charge to the bill. When you end the trip, the workflow is done. That experience was powered by Cadence at the time.Swyx [00:59:34]: I used to say it's like programming the entire user journey top-down as one function.Jake [00:59:39]: It's a powerful idea and important. It's also important for the next phase of the agentic journey. You want an agent to do a specific task, be complete or incomplete on that task, and move on to the next thing. You need a way to manage workflows dynamically.Jake [00:59:59]: Temporal was always great in theory, and great when you got it working the way you wanted in production. But it required you to model the entire journey in your head. If you didn't, you could cause issues where replaying the state of the workflow causes non-determinism.Swyx [01:00:25]: Because it works on deterministic workflow history.Jake [01:00:28]: Exactly. I describe it as a jet engine. If you know how to operate it and run it, it's great. But you can't hand it to people trying to build complicated things if they don't have the whole state in their head.Jake [01:00:48]: We run our whole deployment pipeline on top of it. That's a reasonably complicated workflow: pre-commit hooks, signaling, queuing, and all the rest. We ran into the same thing at Uber. As you express a large workflow, it gets more complicated, with more states in the state machine that you have to map back to the workflow.Swyx [01:01:15]: It's a lot of ifs.Jake [01:01:16]: Exactly. At Uber, we built a system for doing the state machine and testing it. We've started to build some of those things here because it's grown heavily. It's not quite love-hate. When it works well, it works super well. But if someone who doesn't have full context puts something into the system that invalidates state or causes non-determinism, or spins off a ton of activities, you have to keep track of underlying SRE knobs like activity slots. Those should scale with memory, vCPU, and so on. It becomes a bear to scale.Swyx [01:02:10]: You need a capable sysadmin running things behind the scenes. If you moved off, what would you do?Jake [01:02:19]: We'd build our own workflow engine. We have a few internally that we've worked on.Swyx [01:02:27]: This is one of those classes of things you typically wouldn't vibe code, but I'm wondering if you can.Jake [01:02:33]: I still don't think you should vibe code it. You still want to run decent tests to make sure it works.Swyx [01:02:39]: Timo didn't invent that from scratch either. There are libraries you can run. On top of that, it's just a state machine that you have to map out. Ultimately, you define the instructions you want and run them through a state machine.Jake [01:03:00]: It's very doable. Workflow stuff is interesting. Restate is doing neat stuff here.Swyx [01:03:10]: You're tied into JavaScript. Are you a JavaScript maxi?Jake [01:03:13]: Internally, we have TypeScript, Rust, and Go. We don't add more languages. Actually, we have a little C because we write BPF code and hooks. But those are the languages.Swyx [01:03:28]: Is this for sidecars?Jake [01:03:32]: No. It's for the networking stack, volumes, and things like that. We use TypeScript a lot because it powers the dashboard, but we're moving a lot of workflow stuff off the dashboard stack and into the infrastructure stack.Railpack, Nixpacks, and Content-Addressable FilesystemsSwyx [01:04:00]: Cool. Any other technical infrastructure stuff? Railpacks?Jake [01:04:07]: We built an engine for determining dependencies based on source code. It's called Railpack. We built the first version, Nixpacks, on top of Nix, and then we moved.Swyx [01:04:17]: People have been trying to get me to adopt Nix and NixOS for four years. Is it ever going to be a thing?Jake [01:04:23]: I don't know. We're excited about it, but it has pain points. Think of it as a stack of versioned binaries at specific slices in time. If you want version X and version Y, you bloat the package space, which blows up image size and makes real-world workloads difficult.Swyx [01:04:53]: But you content-address it and cache it. In theory, there are optimizations.Jake [01:05:00]: In theory, yes. But with a large enough user base and disparate enough machines, you run into a problem Meta described in the XFAAS paper, their internal serverless system. It becomes difficult at scale unless you break out specific runtimes.Jake [01:05:24]: We didn't want to do that because we wanted to truly allow you to deploy anything. That was our initial thing with Nix. But we've moved toward interesting work around content-addressable file systems that can lazy-load anything from any point and page it into memory.Swyx [01:05:48]: Amazing.Jake [01:05:49]: The future is very bright. It's crazy, and it's going to be nuts.Coding Agent Spend, Roadmaps, and Token ROISwyx [01:05:54]: Founder journey stuff?Alessio [01:05:56]: Your cloud usage: you tweeted you're going to spend $300K this month?Jake [01:06:01]: I think we got to $200K.Alessio [01:06:02]: Coding agents?Jake [01:06:03]: Yeah.Swyx [01:06:04]: Across the company?Alessio [01:06:05]: You only have 35 people, so I'm sure they're not all spending $10K a month. What's the distribution?Jake [01:06:10]: I think I'm at about $25K. We have power users all the way down. We came back from winter break, and I basically said, “If you're writing code by hand, you're doing this wrong.” The tools are good enough now that you can move extremely quickly. There are issues and pain points, but you should be reviewing the code you are writing instead of writing it by hand.Jake [01:06:40]: Architectural patterns matter more now than ever, but you shouldn't spend your time generating code you would write. If you know how to write it, ask the agent to write it and reconcile it until it looks like you would have written it yourself.Jake [01:06:58]: People misconstrue my propensity to push people toward agents as connected to our growth and some reliability bumps. They're not necessarily related. The tools are good enough to move extremely quickly and build things way larger than you could before.Jake [01:07:19]: To the earlier point about cooling data centers in space: I don't know. But with software, you can ask, “How would I build block storage from scratch? How would I do these things?” I have ideas because I have history and have read papers. Let me work them out and build massive test benches with thousands of tests, because those are now free to author. If you're not using AI systems to speed-run your roadmap and reconcile your existing system onto the future, you're missing a large point of what's happening.Alessio [01:08:12]: What's the path to spending $3 million a month? Is it bound by ideas and things customers can absorb?Jake [01:08:19]: For most companies, it's bound by deployment at this point. That's why we've seen a massive boom in users and companies, from Fortune 50s down, asking how to get developers to move faster. You'll probably hit your CFO before any technical limits because they'll look at the eye-watering amount of money spent on tokens. Inference costs have to come down, but we're inference constrained now. There will be price discovery around what makes sense for an org to adopt.Jake [01:09:06]: I think you'll end up with the F1 driver concept. If someone is really adept at these things, it makes sense to put them in a $3 million car. If they're not, it probably doesn't make sense. You'll take a few people and say, “You can drive the F1 car. We need to go in this direction. Figure out if it works and prototype it.”Jake [01:09:33]: We've done some of that and vastly accelerated our roadmap. We thought we'd ship something in a few years; now we can probably ship it in a few months because we validated it and don't have to build it incrementally. We can skip steps and move toward our vision.Alessio [01:09:58]: A lot of people are realizing the roadmap doesn't always have a business impact, so they say tokens are too expensive. But if your roadmap were built to make more money by the time you built it, you'd have token pricing for it, the same way you do with sales. You'd spend a billion dollars on sales if you knew you would get $2 billion of revenue.Jake [01:10:19]: Exactly. A naive way to measure this is the percentage of tokens that end up in production. If you can measure impact because those tokens end up in production, that's awesome. But the burden of proof will rise. Internally, we have a growing number of pull requests that haven't merged. The question becomes: how do you get this into production? It's about how quickly you can build and deploy software, which is exciting because that's our whole thing.The SDLC Shift: Prompt Requests, Feature Flags, and Safe RolloutsSwyx [01:10:56]: The SDLC is changing. One thesis is that the pull request is dying. It's going to be the prompt request. Beyond that, code review is also kind of dying if you have all the other systems in place. What else is changing about the SDLC?Jake [01:11:19]: The AISRE and the tools to make it happen. AISRE is pie-in-the-sky aspirational. What does it take to get an AISRE? What tools do you need to build?Swyx [01:11:32]: You should expose your tooling to customers at some point. The Central Station command center.Jake [01:11:39]: We have it for template maintainers. Template maintainers can deploy and maintain templates, and they get feedback. We're going to expose those things incrementally.Swyx [01:11:51]: Clustering around incidents. Everyone has a version of that, but I don't think anyone has solved it.Jake [01:11:56]: I won't say we've solved it internally, but it's gotten so good that we can see incidents forming pretty quickly. At some point, those will be things either someone else builds or we build. We've always built things purpose-built for us. If it makes sense to make it useful for users, monetize it, or turn that loop into a profit center instead of a cost center, we want to do that.Jake [01:12:28]: Pull request is definitely dying.Swyx [01:12:29]: Do you do first-party feature flagging and incremental rollout stuff?Jake [01:12:34]: We have a feature-flagging engine we built internally and will eventually roll out.Swyx [01:12:38]: I don't see it as a user. How come you didn't give us what you have?Jake [01:12:43]: We have to beta test it. We care a lot about the quality of the things. There's plenty we've used internally that doesn't make it all the way through the journey because it fails. It works for one service but not multiple services. We'd have to build it for multiple services and know that if we released it, we'd rebuild it again and again. Some things are worth that, but many inform the roadmap.Jake [01:13:18]: We don't want to dilute the experience by saying, “This works, but only for this service,” unless it's a core initiative. Over the next few months, we'll roll out things that work for a single service, then multiple services, then multiple services across the environment. You have to be deliberate. Otherwise you create broken disparate experiences and support load because people ask how to use the feature.Jake [01:13:52]: It's the earlier expansion and compaction pattern. You expand the company to get features, then compact and smooth them out so the experience is stellar. You told me in the hallway, “It's gotten so much better.” Internally we're saying, “This part really sucks. We need to make it significantly better.”Swyx [01:14:11]: I can attest to that over the last three years watching you build Railway. For listeners, feature flagging is a huge part of Uber culture. So much so that they have too many feature flags and another thing to remove feature flags. Facebook has Gatekeeper. Agents are going to need this. It's fundamental to incremental rollouts. OpenAI acquired Statsig. GPT-5 is routing and flagging through different models.Jake [01:14:56]: It's super important. If the software development lifecycle is going to change because we're doing things 1,000 times faster and 1,000 times more concurrently, what becomes important at scale?Jake [01:15:16]: Before I started Railway, I built a feature-flagging product and tried to sell it. It was an easier version of LaunchDarkly. I ran into a problem: anyone small enough to adopt your technology doesn't care about feature flags, and anyone large enough to need feature flags needs so much scale that you have to build out all the infrastructure. I scrapped it.Jake [01:15:42]: But what is old is new again. Companies are trying to move quickly, but you can't YOLO a vibe-coded thing straight into production. You need to say, “Here's my blast radius, my impact, and I want to shadow it for these users.” Feature flags. You're going to need the tools larger companies built to maintain their structures. Everything gets compressed by 1,000x so everybody can build those structures quickly.Jake [01:16:07]: That's exactly where we are: compressing the software development lifecycle, then expanding it and adding more new things.Cattle, Pets, and Clonable InfrastructureSwyx [01:16:15]: Another term that comes to mind for newer developers is “cattle, not pets.” People treat production like a pet. It has a name. You baby it and keep it alive. With cattle, you can mass farm, roll out, portion parts out, and kill them.Jake [01:16:37]: I think that might change. You can move toward having pets as long as you have a cloning machine for your pets.Swyx [01:16:52]: Yeah.Jake [01:16:52]: If you can snapshot every single thing at every frame, it doesn't matter if something gets obliterated because you have a snapshot of it. The things we've built right now are designed to block changes from the hermetically sealed DevOps line. You have to write a Dockerfile because you nee
Ben and Andrew discuss the future of computing and its implications for the chip market, including what Cerebras is doing that's different, why speed may no longer be a top priority for inference, good news for China's AI ecosystem, the future for Nvidia, and questions on Pat Gelsinger's role in Intel's revival. From there: Both sides of the Anthropic-xAI deal, including Anthropic's compute solution and the triumph of market principles, as well as the market's message to Elon Musk and xAI, and the implications for SpaceX. At the end: Thoughts on Musk's OpenAI lawsuit, a theory on Apple's gross margins and a land grab, and a listener's wife enters founder mode with Claude.
Ken Liu (Computer Science PhD at the Stanford AI Lab) and Erik Chi (CS PhD at UMich) are the Creators of the Open Anonymity Project, which lets people prove things about themselves online without revealing their identity. In this episode we explore what it means for AI systems to "know" you; why today's so-called privacy modes fall short; and how the next generation of AI systems could be built with privacy as a default, rather than an afterthought. Key Takeaways: What "unlinkable inference" means and why it changes the privacy model of AI chat tools What actually happens to your data the moment you hit "send" in a typical AI system Why incognito mode in AI tools is largely a UI illusion, rather than a real privacy protection The role of metadata in identifying and profiling users, and how "secretary models" could enable personalization without sacrificing privacy How anti-censorship and privacy intersect in a future dominated by agentic AI systems Why now is the time to rethink assumptions about privacy in AI tools Guest Bio: Ken Liu is a Computer Science PhD student at the Stanford AI Lab, advised by Percy Liang and Sanmi Koyejo. His research focuses on foundation models and data/user privacy, and the intersection between the two. His recent work studies the privacy properties of AI (such as membership, memorization, and unlearning), and various AI privacy tools (such as anonymization, differential privacy, and federated learning). His papers have earned spotlights at top venues, and his findings have been deployed at scale on Android. Ken also led a team to a 1st-place win at the US-UK PETs Prize sponsored by the White House OSTP and the UK Government. Previously, Ken spent time at Google DeepMind, Carnegie Mellon University, Meta, Apple, and Amazon. Erik Chi is a CS PhD at UMich, advised by J. Alex Halderman. His research focuses on security and privacy, particularly network security and anti-censorship. He worked on a new standard for implementing and distributing censorship circumvention protocols—a standard that's now being adopted by VPN vendors to help millions of users access the free Internet. He also did content moderation (surveillance) and recommendation systems at ByteDance before realizing how censors will evolve in the AI era. ---------------------------------------------------------------------------------------- About this Show: The Brave Technologist is here to shed light on the opportunities and challenges of emerging tech. To make it digestible, less scary, and more approachable for all! Join us as we embark on a mission to demystify artificial intelligence, challenge the status quo, and empower everyday people to embrace the digital revolution. Whether you're a tech enthusiast, a curious mind, or an industry professional, this podcast invites you to join the conversation and explore the future of AI together. The Brave Technologist Podcast is hosted by Luke Mulks, VP Business Operations at Brave Software—makers of the privacy-respecting Brave browser and Search engine, and now powering AI everywhere with the Brave Search API. Music by: Ari Dvorin Produced by: Sam Laliberte
William Clifford, The Ethics Of Belief - The Limits Of Inference by Lectures on classic and contemporary philosophical texts and thinkers by Gregory B. Sadler
Are AI inference costs already eating into your gross margin — and you can't even see them on your P&L? In episode #370, Ben Murray breaks down exactly what belongs in AI COGS for SaaS companies offering an AI-first or AI-infused product line. Inference bills are stacking up fast, infrastructure-layer spend is the surprise line item nobody priced in, and most finance teams haven't built the GL account structure to capture any of it cleanly. If you don't get the framework in place now, you'll be reporting AI gross margin you can't actually defend by next quarter — and your board will notice. The 5 cost categories every AI COGS framework needs — inference, model hosting/GPU infrastructure, the AI infrastructure layer, monitoring and observability, and AI-specific support Why AI inference costs deserve their own GL account — and shouldn't be buried inside your cloud hosting bill where they disappear The surprise cost line one industry report flagged as the #1 unexpected AI expense — hiding in data platform usage, networking, and egress How to structure your COGS cost centers so you can deliver clean margins by AI product line, not just lumped together at the company level Why token tracking by customer cohort (heavy / medium / light users) is now table stakes for any AI product sold as a subscription The deployed-engineer question: should AI support tickets sit with tech support or a specialized team — and how that decision rewires your margin model Tune in to get the AI COGS framework in place before your gross margin lands on a board slide you can't defend. Resources Mentioned Ben's new AI course: https://www.thesaasacademy.com/ai-finance-metrics-saas Ben's blog post: What Should Be Included in AI COGS: https://www.thesaascfo.com/what-should-be-included-in-ai-cogs/ SaaS Metrics Foundation course: https://www.thesaasacademy.com/the-saas-metrics-foundation
Support & Resources→ Support the show on Patreon→ Bayesian Modeling Course (first 2 lessons free)Our theme music is « Good Bayesian », by Baba Brinkman (feat MC Lars and Mega Ran). Check out his awesome workTakeaways:Q: What is simulation-based inference and what does "sim-to-real" mean?A: Simulation-based inference (SBI) uses a mechanistic simulator as an epistemic tool: you train a neural network on a large number of labeled simulations and then deploy it on real, unlabeled data. The "sim-to-real" framing captures the key asymmetry -- your network never sees real data during training, only simulations, but it generalizes to real observations at inference time. This is the opposite of the more common "synthetic-for-ML" approach, where fake data is used purely to augment real training data.Q: What is the amortized inference agent skill and what does it do?A: It's an open-source AI agent skill, co-developed by Stefan and Alexandre, that teaches an AI coding agent to run a complete, state-of-the-art amortized inference workflow. Because amortized inference is recent enough that it's underrepresented in LLM training data, vanilla agents tend to get it wrong. The skill injects the right methodology: it guides the agent to set up the simulator, choose the right network architecture, run a pilot, train with appropriate diagnostics, and produce an actionable report -- without the user needing to know the details.Q: What is calibration coverage and why should you never skip it?A: Calibration coverage tells you whether your posterior uncertainty is honest -- whether your credible intervals actually contain the true parameter at the right frequency. A model can show poor parameter recovery yet still be well-calibrated (because it's falling back on the prior), or it can appear to recover parameters while being poorly calibrated. Running calibration diagnostics both in-sample and out-of-sample is especially revealing for hierarchical models, which often appear to underfit in-sample but generalize much better out-of-sample thanks to shrinkage.Full takeaways hereChapters:00:00:00 How does amortized inference fit into the Bayesian workflow?00:12:03 What does "sim-to-real" mean in simulation-based inference?00:15:57 Why is amortized inference particularly suited to psychology and neuroscience?00:21:51 What is the amortized inference agent skill?00:39:00 What is calibration coverage and how do you interpret it?00:41:50 How do you decide what to do next after your first training run?00:44:53 How do actionable insights make Bayesian workflows more usable?00:49:08 What are the unique challenges of hierarchical models in amortized inference?01:00:51 What is the current state of BayesFlow's support for hierarchical models?01:05:00 What are the main failure modes of amortized inference and how do you handle model misspecification?Thank you to my Patrons for making this episode possible!Links from the show
No Priors: Artificial Intelligence | Machine Learning | Technology | Startups
Baseten CEO and co-founder Tuhin Srivastava sits down with Sarah Guo and Elad Gil to discuss the rapid growth of AI inference demand, Baseten's 30x growth, and why inference is becoming the strategic “last market.” Tuhin Srivastava argues the application layer will persist because companies with unique user signals can encode value into workflows and post-train specialized models, citing examples like Abridge and support workflows. The conversation covers GPU capacity constraints, Baseten's multi-cloud fabric across 18 clouds and 90 clusters, long-term contracting dynamics, the importance of the software layer for stickiness, evolving workloads, multichip possibilities, and operational lessons at scale. Sign up for new podcasts every week. Email feedback to show@no-priors.com Follow us on Twitter: @NoPriorsPod | @Saranormous | @EladGil | @Tuhinone Chapters: 00:31 Baseten growth 01:55 Why the app layer wins 05:57 Serving frontier customers 07:55 Open source model mix 09:21 Chinese models and geopolitics 13:07 Custom inference dominates 14:22 Post training acquisition 17:10 When to invest in custom models 18:35 Supply crunch and data centerse 22:25 Longer GPU Contracts 24:09 What Makes a Winner 26:07 Multi Chip Future 28:19 Runtime Roadmap 31:08 Scaling Edge Cases 33:48 Hiring and Leadership 36:44 Operations Pager Culture 38:19 Efficiency Drives Demand 40:41 Concierge Everything Future 42:34 Conclusion
This Week in Machine Learning & Artificial Intelligence (AI) Podcast
In this episode, Philip Kiely, head of AI education at Baseten, joins us to unpack the fast-evolving discipline of inference engineering. We explore why inference has become the stickiest and most critical workload in AI, how it blends GPU programming, applied research, and large-scale distributed systems, and where the line sits between inference and model serving. Philip shares how research-to-production can move in hours, not months, and why understanding “the knobs” of inference—batching, quantization, speculation, and KV cache reuse—lets teams design better products and SLAs. We trace the inference maturity journey from closed APIs to dedicated deployments and in-house platforms, discuss GPU lifecycles, and survey today's runtime landscape, including vLLM, SGLang, and TensorRT LLM. Finally, we look ahead to agents and multimodality, making the case for specialized, workload-specific runtimes when performance and efficiency matter most. The complete show notes for this episode can be found at https://twimlai.com/go/766.
Plus: Anthropic investigates report of "unauthorized access" to its Mythos AI model. And energy company GE Vernova lifts yearly outlook on surge from data center demand. Julie Chang hosts. Learn more about your ad choices. Visit megaphone.fm/adchoices
This episode originally aired on The Kevin Rose Show. Kevin Rose speaks with Anish Acharya, general partner at a16z, about how AI is rewriting the rules of consumer software, the defensibility of network effects in a world where anyone can spin up an app in 48 hours, and why the real threat to consumer founders may be the cost of inference, not competition. They also discuss model pricing, the future of the four-day work week, and peptides. Resources: Follow Anish on X: https://x.com/illscience Follow Kevin on X: https://x.com/kevinrose Stay Updated:Find a16z on YouTube: YouTubeFind a16z on XFind a16z on LinkedInListen to the a16z Show on SpotifyListen to the a16z Show on Apple PodcastsFollow our host: https://twitter.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
(0:00) Jensen Huang joins the show! (0:26) Acquiring Groq and the inference explosion (8:53) Decision making at the world's most valuable company (10:47) Physical AI's $50T market, OpenClaw's future, the new operating system for modern AI computing (16:38) AI's PR crisis, refuting doomer narratives, Anthropic's comms mistakes (20:48) Revenue capacity, token allocation for employees, Karpathy's autoresearch, agentic future (30:50) Open source, global diffusion, Iran/Taiwan supply chain impact (39:45) Self-driving platform, facing competition from active customers, responding to growth slowdown predictions (47:32) Datacenters in space, AI healthcare, Robotics (56:10) OpenAI/Anthropic revenue potential, how to build an AI moat (59:04) Advice to young people on excelling in the AI era Follow the besties: https://x.com/chamath https://x.com/Jason https://x.com/DavidSacks https://x.com/friedberg Follow on X: https://x.com/theallinpod Follow on Instagram: https://www.instagram.com/theallinpod Follow on TikTok: https://www.tiktok.com/@theallinpod Follow on LinkedIn: https://www.linkedin.com/company/allinpod Intro Music Credit: https://rb.gy/tppkzl https://x.com/yung_spielburg Intro Video Credit: https://x.com/TheZachEffect