AI researcher at Tesla
POPULARITY
When we first dicsussed the Summer of Simulative AI in 2024 we knew it would be a brief summer, but it has recently come back with a vengeance with SimGym in April and now Simile AI's $2B Series B, backed by GreenOaks and Index Ventures with prominent backers like Fei-Fei Li and Andrej Karpathy, running tens of millions of simulations for Fortune 100 clients like CVS and 85–99% accuracy vs human focus groups. Time to catch up on why this Second Summer of simulation is working!From creating Smallville, the landmark 2023 paper on Generative Agents that showed AI characters could remember, plan, socialize, and develop emergent behaviors, to now building foundation models of human behavior, Joon Sung Park is trying to answer a much bigger question: what if we could simulate the world before making decisions in it? In this episode, the Simile co-founder and CEO joins us to unpack the path from generative agents to digital twins, why today's frontier models still fail to capture how humans actually behave, and what it would take to eventually simulate all 8 billion people on Earth.We go deep on Simile's approach to modeling human behavior: long-form interviews, observational and transaction data, randomized controlled trials, population-level and individual-level models, and post-training on the causal mechanisms behind why people make decisions. Joon explains how his research created digital twins that reproduced human behavior and attitudes 85% as accurately as people reproduced their own responses, why models optimized to be rational can be bad simulations of irrational humans, and why understanding “social physics” may require changing model weights rather than simply prompting frontier LLMs.We also explore the much larger ambition behind simulation: testing products and policies before deploying them, finding counterintuitive paths toward desired outcomes, modeling emergent behavior across entire societies, and potentially tackling problems like climate change, democratic instability, and UBI. Joon reflects on scaling laws for simulation, the economics of data-center-scale simulated worlds, the connection to Thomas Schelling and psychohistory, why simulation is surprisingly similar to painting, and whether we might already be living in one.We discuss:* How Smallville and Generative Agents led to Simile* Why Joon's team asked: “What if we can just recreate the world that we live in?”* Why useful personal agents require deep models of their users* Memory architectures, Markdown files, and the limits of prompting* “Social physics” and behavioral foundation models* Why web data captures what people say more than what they actually do* Interviews, transactions, observational data, and randomized controlled trials* Why predicting the future matters less than understanding how to shape it* How Simile creates representative simulated populations* Simulation versus prediction and the connection to Foundation's psychohistory* How to evaluate simulations instead of simply stacking LLM hallucinations* Creating digital twins of 1,000 real people and reaching 85% behavioral accuracy* Why frontier models can struggle to reproduce real human behavior* Why good simulations need to reproduce human biases and mistakes* Post-training models on randomized controlled trials* Population-level versus individual-level simulation* Scaling laws for human simulation* The long-term ambition to simulate all 8 billion people on Earth* Whether simulations could help solve climate change or detect collapsing democracy* Thomas Schelling and the history of agent-based modeling* Why future simulations could require an entire data center* Multi-agent simulations and what happens when simulated people interact* Replacing expensive human panels with synthetic populations* Why market research is only the starting point for simulation* Why Joon sees simulation as surprisingly similar to painting* Using simulation to study questions like UBI* Whether we are already living in a simulation* Why AGI and simulation may be the twin technologies of advanced civilizationsJoon Sung Park* LinkedIn: https://www.linkedin.com/in/joonspark* X: https://x.com/joon_s_pk* Website: https://www.joonsungpark.com* Simile: https://www.simile.comTimestamps00:00:00 Introduction and Joon's Path from Art to AI00:01:46 Smallville, Generative Agents, and the Origins of Simulation00:05:03 “Let's Just Create a World” and the Future of Personal Agents00:09:53 Social Physics and Behavioral Foundation Models00:14:08 Prediction vs. Simulation: How Do You Shape the Future?00:16:59 How Simile Models Real People and Populations00:25:35 Evaluating Simulations, Digital Twins, and 85% Accuracy00:30:23 Post-Training Models to Reproduce Human Behavior00:40:04 Scaling Laws and Simulating 8 Billion People00:43:10 From Schelling to Society-Scale Agent Simulations00:46:13 The Cost and Economics of Simulating the World00:52:05 Real-World Use Cases, Synthetic Populations, and the Market00:57:27 The Future of Simulation, Painting, and UBI01:04:23 Are We Already Living in a Simulation?01:06:08 Building Simile and HiringTranscriptIntroduction: Joon Sung Park, Simile, and the Story So FarVibhu [00:00:00]: Today, we have Joon in the podcast. Excited to kick this one off. Very exciting company. I wanna kick off and ask you the question, talk us through the story of your life. How have you gotten here?Joon [00:00:13]: Yeah, for sure. I'm really excited to be here. A story of my life. So I was born in Korea, and I lived there for a good 11 years or so of my life, and then my family moved to Boston. So we moved when I was 11, and my parents were doctors, so they were going through their postdoctoral studies. My dad was a surgeon, so he was doing his sabbatical years at the Boston Children's Hospital. So I grew up there, not too close to tech. I was very much a music and artsy, painting kind of guy.Vibhu [00:00:49]: Painting.Joon [00:00:49]: Exactly. I got into painting a little bit later, in high school, but that's what I used to do. And then I grew up mostly in the East Coast after Korea. So I lived a good number of years in New Hampshire, and then I went to college in Pennsylvania. And I got into more of this tech scene, in college. So I was originally trained to be an artist. I thought that would be my professional career. So it wasn't a hobby. It was like, “Hey, let's make a living out of this.” And then gradually, I got really interested in this idea of, hey, the greatest artist often creates their own medium, and the best medium that we had available today was in computation. So I decided to go deeper into that, and one thing led to another, and we can go deeper into this, but I decided that research was something that I gradually got interested in, and here I am.Smallville, Generative Agents, and the 2023 Breakout PaperSwyx [00:01:46]: So there's a lot that you packed into the research components. You had one of the best papers of 2023, which was the generative agents paper, commonly known as the Smallville paper.Swyx [00:01:58]: Feel free to call back to anything else that you mentioned, but most people would have heard of you from this. Do you have any statistics on how many people have, like, read it? arXiv gives you something, right? Some stats.Joon [00:02:10]: Yeah, it's a good question. How many people have read it, I'm not sure.Joon [00:02:14]: I know we do keep track of citations, and they are going up quite fast.Swyx [00:02:23]: Yeah, Google Scholar has 7,200 citations.Vibhu [00:02:25]: I feel like it made a bigger hit than that, and it was a pretty instrumental paper. It got cited so many times.Swyx [00:02:34]: It is frequently the answer when people ask, “What is the best paper you've read recently?” It's this one.Vibhu [00:02:39]: I thought the memory component was pretty underrated. It was a very good early memory system, and one of the biggest papers.Foundation Models and the Search for Killer ApplicationsJoon [00:02:47]: Yeah, so maybe I can talk a little bit about how this particular paper came together. So when I got into research, it was back in 2020 when I started my PhD program at Stanford, and that was the year, when we were about to get GPT-3 to be available. So we already had GPT-2, and you could sense that there was this new class of models that was just becoming available in the market, and the team got very intrigued. And the general consensus was, “Well, is this model going to be useful for anything?” “It's really strange that these models are not trained to do any particular task.” But we decided to take a bet. So a large group of scholars at Stanford, and it was led by one of my co-founders, Percy Liang, and we came togetherSwyx [00:03:35]: Who coined foundation models.Joon [00:03:36]: Who coined the term foundation models. We wrote this paper, where that term came from called Opportunities and Risks of Foundation Models. And during that process, really the thing that I started to think deeply about was, here is a model that is fundamentally new in our ecosystem. The reason why this was new was it wasn't, again, trained to do anything in particular, but its premise was it could do anything and everything. It was like a stem cell, if you were to take a biology analogy. And I got really interested in this idea that, well, if we were to really think about what are the killer applications that this particular technology would enable, what would that be? Many of my colleagues were using this for simple classification, simple generations. Interesting that these models can do that, but from an interaction perspective, not that interesting. We've known how to do that for many decades. And what we came down to was these models are trained on this very broad data from the web, right? So these are human behavioral data. It's social media, Wikipedia, all these data. So if you poke at the right angle, then you could see human behavior that would just pop out that's quite realistic, and we've never seen that before.The Time Machine Game and Recreating the WorldJoon [00:04:45]: So that got us really interested. The exercise that we decided to do, with this particular group of colleagues, Michael Bernstein, Percy Liang, and myself, who ended up becoming my co-founder at Simile, we sat down and we played this game that we call the time machine game.Joon [00:05:03]: Imagine we were to get on a time machine and fast-forward 10 years and look back. What would have been the single application that will have mattered that would be the most interesting and inspiring? And when we thought, “Well, what if we can just recreate the world that we live in?” it's really hard to get more ambitious than that. Like, let's just create a world.Joon [00:05:24]: And that's where we started. And initially, we had this paper that was a precursor to the generative agents paper called Social Simulacra.Swyx [00:05:32]: Before you go further, were there other candidates for the most ambitious thing in the time machine exercise? What was number two or number three?Personal Agents, User Models, and Why Simulation Came FirstJoon [00:05:44]: There is a close second that we were considering, which ended up becoming more of these automation tools, especially the vision around really personalized agents that would do things for you.Swyx [00:05:59]: That's also happening.Joon [00:06:00]: It's also happening. But it was interesting for us, right, in that the reason why, we decided to go with the idea of simulation, one, I was a huge science fiction nerd, and this idea of creating simulation, I was personally really just fascinated. I loved the idea. It's really cool to see, like, a game town like this and just see these agents live in it. But at the same time, my bet was if you were to create a really amazing personal assistant out of this technology, what you need first is an amazing model of your users. So I told a model, “Hey, can you go buy late dinner for me?” And it orders Hawaiian pizza, and I do not like pineapples on my pizza. Then it totally failed. The way for it to not make that mistake is only by having a deep understanding of who I am. And I gave a very simple and dumb example here, but you can imagine how this core understanding of people is instrumental. This is how, if we have our family and closest friends, they have a good mental model of who we are. That's the basis of our social connection. So our bet also was this technology around simulation, creating accurate representation of people ought to precede the more complex agents that would automate the world that we live in. So that was the bet. But that was a very close second, and I'm still very much fascinated by it. I think there's a lot of interesting work that's going around. My hot take here, though, is I don't think we've seen a true personal assistant that's useful, in ways that meet the ambition of that particular line of work. I think there are early applications that are interesting, and if you talk to even ChatGPT nowadays or Claude, they know a lot about us. So a lot of the generation it's doing, I do think it's much more tailored, but I think the ambition is quite large in that field, and I don't think we quite have all the right ingredients just yet.Swyx [00:08:01]: So OpenClaw and these personal agents, what do you want to see from them that they don't currently have?Memory, Markdown, and the Limits of PromptingJoon [00:08:09]: I do think it's slowly getting there, but I do generally want them to have much deeper understanding of the person. Right now, you look at the models. OpenClaw, what it's leveraging is a Markdown file, and I think it's quite clever, right? So if you look at the generative agents paper, this was the same intuition that we had, where initially when we were creating the memory architecture for the generative agents, and, like, this is, like, back in 2022, so we didn't really quite have the idea of even agentive architecture or the term agent. But the intuition that we shared with some of the work that's coming out today was we initially thought, “Well, do we want to make the memory into, let's say, knowledge graph? Do we want to train a bespoke model?” All of these things. And what we decided to do was, “No. Just forget about all this.” These language models are quite good at modeling text and understanding and reasoning about text. So just put everything in a Markdown file or a text file. You're done. I thought that was quite interesting that we could do that, and there's a lot of strength in doing that. But also, there are limitations. It's the way you retrieve and make sense of data that's extremely large, it takes a lot of work. So I think that technology is getting better. I also do, however, think, there are certain things you just cannot shape just by prompting the model. So to some degree, you do need to touch the parameters of the model itself. So there is this work that I do think does need to happen, and it is happening. The question is, how far can we take it? How do we source data, and how do you also create an ecosystem where people are continuously feeding data to this model so it's learning about you?Vibhu [00:09:50]: What's the intuition between why you need to do it in the model?Social Physics and Behavior Foundation ModelsJoon [00:09:53]: My intuition behind the actual when do you train or even post-train a model versus just prompt a model is if the model has to learn the underlying physics of the world that it's operating in. So it has to learn new social physics. The places where it doesn't have to train are the places where it already has the physics. We trust the physics. It already has the base statistics, but it's just trying to react to an environment. Then I think you can just prompt your way into getting the actions out of it. I don't think the models that are out in the open have yet learned the complete mapping of social physics of humanity. This is one of the core theses of Simile, right? And one of the core reasons why that is the case is if you look at the data that the model was trained on, these models were trained on the web data, like, whatever was available on the web. And these are really interesting data sets, but they are fundamentally the self-exposed attitudinal data with some behavior data that's sprinkled around here and there. And it has yet to learn the really deep behavioral nature of people, not just what people say they do online, but what they do in real life. And this is one of what I would consider to be the dark knowledge of humanity that we haven't quite captured. And it's these data that would also need to get factored into the model creation.Vibhu [00:11:21]: You call it behavior foundation model.Vibhu [00:11:23]: There's a good one-liner here, but outside of that, what type of data do you need? What are you changing on the model level? How do you go about modeling, doing a behavior foundation model?The Three Data Buckets: Interviews, Behavior, and CausalityJoon [00:11:35]: We think about data in three buckets. So one bucket is interview data. It's quite interesting. Rich qualitative data is interesting. It's not behavioral, but we would literally ask people, “Hey, tell me the story of your life.”Vibhu [00:11:53]: It's just what we're doing here exactly.Joon [00:11:54]: The question that you all asked at the beginning of this interview literally is the question we also ask. And we ask our participants to go a little bit deeper, than how far I went. Maybe I can give more of my life story in lieu of this. But the reason why that data is interesting is by learning about this very long-tail information about people, you get a lot of texture around this model, like, this person as a model. So even understanding their childhood memory or even their trauma, their first love, these things, quite informative in ways that's really hard to predict. So that's one. Then there are two tranches of what I would consider to be the behavioral data. One kind of behavioral data is observational. So these might be like transaction data, or these might be data that you can get by scraping the web, right? So you can imagine why these data sets would be interesting, right, because they give you the base statistics of people's behavior.Joon [00:12:55]: But then there is the last category of data, that I personally think is perhaps the most important, which is the data that describes the causal mechanism, the whys of people. Some of this is covered by the interview data, the qualitative, because people talk about why they made certain decisions. But really, where you get to see the most behavioral aspect of this is in randomized controlled trials, like RCTs. Imagine you have the same setup, but you have a few different variables that you are trying to tweak. Can you get realistic human behavior out of it in ways where, imagine you had this particular option. Imagine you're even trying to choose whether you're going to drink coffee or not. The day you drink coffee versus the day you didn't drink coffee, does your behavior change? That's a data set that describes a causal mechanism. This is quite important in modeling people. The reason why this is important is oftentimes when people come to us, or not just to us, but the reason why people are interested in simulation isn't because they want to predict the future. If you're trying to win against the stock market, predicting the future is interesting.Prediction vs. Simulation: Shaping the FutureJoon [00:14:08]: But most people, most decision-makers, what they want to know is, how can we shape the future? It doesn't really help you to hear that your sales are going to tank in two quarters. They're just gonna say, “Wow, that sucks.” What they want to know is, well, what do we need to do now to avoid that future? That's the causal mechanism. And this is also very hard data to come by, right, because the world is our ground truth, but it happens once. So in a very controlled setup where everything is equal except for one variable, this kind of data set rarely happens. So this is a reason why this data set is both hard to come by and quite important if you're trying to model human behavior.Swyx [00:14:50]: So behavior, I think, is the hardest data set to acquire. What is out there? What is even possible? You're not going to know a lot of details about my life. I don't even have data for myself on my own health or habits, and I just don't log everything. So how can you have that data?Joon [00:15:14]: So we run a lot of randomized controlled trials.Swyx [00:15:17]: But you put people in the lab, they watch them sleep, or what?Joon [00:15:20]: We do care a lot about the consent process. People know that we invite them to be a member of this community to both share data and have themselves represented in different forms. But we bring a lot of people to the lab, or virtual lab, where we design experiments that would pose them real behavioral decisions. And often in these experimental setups, what makes the difference between what is attitudinal versus behavioral is whether the stake in your decision is real. That's ultimately what makes it behavioral. So in these setups, we are inspired by our colleagues in social sciences, psychology, and so forth. So when they run studies, the techniques they utilize is imagine there's an online store that you're inviting people to come by. Then whatever they purchase in this experiment, they actually get that item delivered. Like, these are the things that make the stakes real. So we run a lot of these experiments, and we also do partner with firms. Right now, we also have customers who are quite excited to at least give us a glimpse of the behaviors that their users exhibit so that we can get a little bit deeper understanding of how people behave in these different platforms.How Customers Use Simile: Populations, Queries, and ExperimentsVibhu [00:16:39]: I think on the customer side, they have a lot of data about their users, who has bought. They have the action data.Vibhu [00:16:47]: Can you walk us through an example of what someone comes to you for? What questions would they want solved? Do you customize a model for them? Do you have something off the shelf? What does that look like?Joon [00:16:59]: Today, when people leverage our models, it's often to better understand the population of their interest. So usually, the start of the relationship, we come together and hear about what population they want us to model, right? So it might be that if you're a CPG company that's selling to all of the US, then maybe it's fairly straightforward. You want to model the gen pop of the US. But at the same time, if there is a vertical or if there's a market that they're trying to go into, imagine, they want to better understand, let's say, people in their 20s and 30s living in California. That's a much more specific population. So we hear about this population, and we go recruit these people, with consent, and with incentives, and we collect some of their data and create a model of these people. Then what our product allows you to do is query them. So it can take as input a filter that is a description of the population that you want to talk to, just like the one I just mentioned, and an environment. The environment can literally be survey questions, behavioral experiments, It can be A/B testing. Oftentimes, the core use cases are things like concept testing, to start with. But also, people sometimes want to do focus groups or one of the fun use cases that we also serve is even modeling things like earnings calls for public companies.Joon [00:18:21]: So these are the use cases that we often start with.Swyx [00:18:23]: Concept testing, is that an established term? I've never heard of concept testing.Concept Testing, Gallup, and PoliticsJoon [00:18:27]: Yeah. So it has to do with they have, let's say, different messaging, different products, different ideas.Swyx [00:18:32]: It's like a marketing exercise.Swyx [00:18:33]: Okay, got it. Got it. Politics?Joon [00:18:36]: We do, have a strategic partnership with Gallup, and of course, Gallup is deep into policy space and so forth. Right now, we have not worked deeply with politics, like that area just yet, however.Swyx [00:18:49]: I'm curious if there is demand or if they really would have different needs that somehow fundamentally don't mix with your existing, users or people.Joon [00:19:00]: I think there's certainly demand.Joon [00:19:02]: But we are very much mindful of how this technology gets adopted and the societal impact that we'll end up having with this technology. And I do see politics as an area where a company has to be particularly thoughtful about the way they operate and make impact. So this is where we also want to make sure that we form enough of guardrail and perspective on how to leverage this technology before we go on to serve markets like the politics.Swyx [00:19:29]: I'll give people an example. one of my favorite shows is The West Wing. I don't know if people have watched.Swyx [00:19:34]: One of the key storylines is, like, the president has, multiple sclerosis, but they haven't. they need to figure out how to disclose it. So they run a poll with a fake governor and ask people to respond on the poll,Counterfactuals, Polling, and When Simulation Is UsefulSwyx [00:19:47]: They try to make decisions based on the results of that poll on, like, how well they'll be received, like where, how should we play this?Swyx [00:19:54]: And I'm like, well, I think those counterfactual things, I would use a simulation for this if I could trust it.Joon [00:20:01]: For sure.Joon [00:20:02]: In that show, how'd it go?Swyx [00:20:04]: In that show, it was, like a foregone conclusion. They were like, “We know it's bad. We just don't know how bad.” And then the poll came back. It was like, “It's really bad.” And then they just did it anyway.Joon [00:20:14]: Part of it is to show, right? So you're, you're looking at the ideaSwyx [00:20:17]: Maximizing drama.Joon [00:20:18]: How bad could it be? Oh, it's horrible.Swyx [00:20:20]: And to some extent, I think that is part of the trick of the, or the challenge or with being a customer of yours, which is that if I know it's. if I roughly know and can intuitSwyx [00:20:35]: What the effect is going to be, do I need you? What sensitivity of it, of effect do I need in order to make a decision, right? So for example, if I, my approval rating is 50%Swyx [00:20:48]: And I, they have this negative piece, news item comes out, and it drops to 30.Swyx [00:20:52]: If it drops to 20, if it drops to 40, do I care? No. It, I know it drops. It's negative. So when do I care about simulations?Joon [00:21:01]: You do something that's clearly bad, that's not popular, and people don't like you, like, yeah, it's likeSwyx [00:21:05]: You don't need a simulation.Joon [00:21:07]: Yeah. Well, so there are a couple of things. one is, there are use cases where, like every day, developers, designers, policymakers, marketers, every single day, they create assets. They create new products. And turns out, it's many of the decisions in hindsight is obvious. Yes, of course this is bad, but we still run those studies because understanding the magnitude and understanding how acute something is quite difficult, even if, we feel like, of course, like this makes sense. this is the reason why we make so many mistakes. Like, every time somebody goes online and say something that has huge backlash, you look at that and like, “What an idiot.” However, it's tough. That's one. There's also another aspect here, which is, again, this is the reason why simulation is different from prediction. In simulation, in the ideal case scenario. So what simulation is trying to show is it's trying to show each step of the way or each step that we need to take to get to a certain outcome, right? So in the most advanced simulations, sometimes the next step that we're suggesting might be quite counterintuitive. The analogy that I sometimes give, and I ground it in a more realistic example, but, I, as I mentioned, I'm a huge fan of science fiction, and I don't know how, many of the audience members have read, like, things like the Foundation series by Asimov.Simulation as a Path, Not Just a PredictionSwyx [00:22:37]: Oh, yeah. We've mentioned psychohistory a number of times.Joon [00:22:39]: Okay, fantastic. So I might be, talking to the right crew. If you read Foundation series, literally the first act is there's a group of scientists who have found out that, “Oh, our galactic empire is going to collapse, and we're going to have 30,000 years of unrest.” And they run psychohistory, the simulator that tries to teach them, “Okay, how can we keep this unrest to a 1,000 years?” And they plan this out, and the first step of that plan is to get the scientists who say, “Okay, this is coming,” exiled into this random place in this, galax- galaxy.Swyx [00:23:18]: Terminus.Joon [00:23:19]: Exactly. And that's so counterintuitive. Like, what a strange move that you literally sent the group of scientists who was raising voice around this potential collapse of galactic empire into nowhere. How is that the right first move? Well, it turns out in this particular simulation, that was the move.Joon [00:23:40]: It's these things, right? And the reason why these reasoning is possible is because you're showing the step function or each step that results in a particular outcome. So really what simulation allows you to do in its highest form is you give it not a problem or question, like what would people answer to the survey? That's not what we do. What we tell it is, “Here is a goal that we have. In the context of foundation, we want to keep the unrest to a 1,000 years. What is the path that we need to take now to get to that particular future?” And that's what simulation allows you to do. Now, translating that into real market, imagine you're a automobile company and you're about to release a, EV, and you're trying to understand, well, how do we market EV, to make sure that our stock price goes up? But what if the answer comes down that, well, you can market your EV in XYZ way, but that might change people's perception around the cars that's not EV and make your overall sales to go down. Not very intuitive, especially all you're trying to optimize is EV salesss, and that's the only thing that you're tracking, then that might result in a completely wrong solution, or at least different solution than what you would have expected, whether it's right or wrong.Joon [00:24:57]: That's the power of simulation.Swyx [00:24:58]: For listeners, we covered a similar topic with Mikhail Parakhin from Shopify, where they are working on SimGym. I don't know if he ever talked to you about it. it's very similar.Joon [00:25:07]: ISwyx [00:25:07]: The goal is increased conversion, but then the journey is very unusual.Joon [00:25:12]: Journey is unusual.Swyx [00:25:12]: Yeah. The-- He's trying to look for interventions on a shopping trajectory, which is similar to what you're saying. Like, it's not about the attitudinal, is your word for it.Swyx [00:25:24]: It's about behavior.Joon [00:25:25]: It's about behavior.Swyx [00:25:25]: And that's exactly the difference, right? It's, like, not about the near-term direction about-- but it's more about, like, how do you affect multiple turns of interactions.Vibhu [00:25:35]: You had a good quote at the start about this as well. It's not about people wanting to know the outcome. It's about how they can change it, change the way to get there, something like that. But I wanna take it back to how do we know this is grounded? LikeGrounding and Evaluating Digital TwinsVibhu [00:25:47]: How do you run evals? How do you test that simulations come through? if I was to do the same thing that you described with, say, your favorite LLM, Opus, GPT-5.6, have some agent to map out these thingsVibhu [00:26:02]: How different are the answers we would get if I give it the same goal, the same objective, make a decent system? You're saying that you need to change the model weight. You have your own solution to this. But how far off are we, and how do you check if it's grounded? you have some interesting stuff on your site that points to how you run real evals, but if you could take us through that side. I think that's one of the big concerns that people have. They're like, “LLMs hallucinate.”Vibhu [00:26:27]: “You're just hallucinating layer after layer,” right?Joon [00:26:30]: The way we do this, and this is the paper that we worked on after the generative agents paper that really became the, at least for Simile and also the field of simulation and synthetic panels, really became the foundation. Yeah, this is the paper. the paper is called Generative Agent Simulations of 1000 People. Here's what we've done. For this paper, we brought 1,000 people that's representatively sampled from the US to a virtual lab. And what we have done was we spent two hours collecting fairly wide-ranging data. In this particular study, we focused a lot on this interview data, that was, whose script was taken from this project called American Voices Project. And then we would also pair that with a lot of behavior data and so forth, whatever we can collect within two hours. And then we would send these people away for a couple of weeks. And during that time, I would use this data to create their digital twins. And I would bring the humans, participants back after 2 weeks and have them complete a battery of surveys, experiments, behavior studies. So we have the list here, which included things like behavioral economics games. We would run literally, like, Big Five personality test, General Social Survey. We would also go ahead and run the randomized controlled trials that were published on PNAS. And we would have their digital twins predict how the source individuals would have acted in these studies and surveys. And this is where we could replicate people's behaviors and attitudes 85 percent as accurately as people would replicate their own. So that was the first really paper that gave this validated results that we can model individuals in an accurate way. And what we ended up finding now, of course, in AI space, so this paper came out at the end of 2024. AI space, a year and a half, 2 years, that's a lifetime.85% Accuracy and Why Frontier Models Miss Human BehaviorSwyx [00:28:24]: Yeah. Just, for listeners who are not seeing the YouTube, I just wanna say, like, the headline figure is 85 percent accuracy, like, which is a big improvement over all the otherSwyx [00:28:34]: Methods that you showed.Joon [00:28:36]: But the part that was particularly striking to us, especially as we improved this technology even further, was the generative AI models like ChatGPT, Claude that's coming out, it does give you the right foundation. However, what they do not consider is the true attitudinal and behavioral aspect of people, especially in the population that you care about. So what these models are really good at today is they're trying to become the super rational, objective machines, right? So you go get their data from places like Mercor, Scale. You talk to professional programmers, scientists to create model that's amazing at reasoning. That's what they do. Simile doesn't care about any of this. The models that we're talking about here, what we're trying to create are models that are as dumb as I am, right? So if I make some mistakes, the model has to make the same mistake.Swyx [00:29:34]: Oh, that's very hard.Joon [00:29:35]: That's very hard.Swyx [00:29:36]: You're solving Murphy's paradox.Joon [00:29:37]: That's exactly. And this is a completely different data and training objective. This is also where we see quite a bit of discrepancy in the performance in human behavior prediction between the frontier models, Simile's model, and the models being created in this space, where in some cases, the model performance of frontier models go all the way down to 20, 30 percent, especially if you go into that more niche population on topics that our customers would care about. On more gen pop, it might be around 50 to 60 percent. So it's not very robust. Like, you wouldn't want to make your decision off of these and these findings. If you can bring that up to 85 percent, that is ultimately what people end up getting very excited about.Swyx [00:30:20]: Yeah. Do we wanna keep going on the paper, routes?Joon [00:30:23]: Yeah, for sure. So the last one, was an interesting one. So this, paper was the follow-up paper that we had, to the 1000 agents paper, where the idea was now can we augment the models even further and post-train a model based on a lot of randomized controlled trials? So this was an interesting one. The data is always the most interesting part of modeling in many ways. The data that we got here was there's this, there's this platform called Open Science Framework. So some, the audience might be familiar with this. And there has been, especially in the social sciences over the past 5 years or so, there has been this concern around replicability of studies. And so it was a bit of a crisis, the scientists acknowledged, where we rerun the study and we don't see the same finding.Post-Training on RCTs and Replication StudiesVibhu [00:31:12]: Oof.Joon [00:31:12]: It's tough. And the reason why it's there-- that was often the case was there's this survival bias where the papers that get published often need to maintain what we call the value of less than 0.05 in the experiments that we ran. That suggests that only-- there's only 5% chance that the results that we saw is false positive. But the tricky part was all the papers that were not published, and there's still a 5% chance that whatever we publish is totally just randomly generated. Like, there's a 5% chance that, hey, this effect is not real, but it just happened to be real because of the sampling bias. So because of that, what scientists started to do was they started to register their studies. So before running an experiment, they would go to this platform and say, “Here is the data. Here is the population that we're collecting, and here's the hypotheses.” And they would just say, “Here is our hypothesis.” Like, “This is what we believe.” And you cannot retroactively change those hypotheses. This is what gives us more scientific statistical confidence that whatever effect that you ended up seeing is true. So that ended up creating this really interesting platform where there's one platform that has now contains tens of thousands of real-world experiments and hypotheses. And a lot of these are really high-quality, like, professionally designed behavior studies and random- randomized controlled trials. So we got the data and the studies from this platform and used that to make a point. And this particular, model is not, something that we're serving commercially because this was a part of the open science. But this particular data set, helped us make a point that by collecting a lot of these randomized controlled trials, that are really well-designed, we can make significant improvement in model's capability to predict human behaviors. So that's what this paper was about.Vibhu [00:33:10]: Is this stuff done on a individual level? Like, do I need to tune the model per individual, per company? Is there foundation model changes and then some slight post-training? Anything you can share there?Population-Level vs. Individual-Level ModelsJoon [00:33:21]: So this particular model was trained. the data we had at the level of individuals, but this particular model was trained. We experimented with both. And this is what we end up doing at Simile too. We always train 2, distinct model. One is what we call the population-level model. The other is what we call the individual-level model. And both take very similar input, which is the description of a subpopulation or individual and a stimuli. In this particular work, we've done the same. Here, the results that we are reporting are much more geared towards individuals because we do think that is a harder task in many ways, but that's what we have done.Vibhu [00:34:02]: You seen anything on the questions that humans can solve that models can't solve? So likeHuman Biases, Mundane Choices, and What Models MissVibhu [00:34:09]: Currently, it's, I live 5 minutes walk away from a car wash. It's a 10-minute drive. Should I walk or drive?Joon [00:34:16]: Huh.Vibhu [00:34:16]: The model will say, “Oh, walk to the car wash.” And, you don't have your car.Vibhu [00:34:20]: Is anything like this a problem in simulation? You would assume, like, very simple for human to think about, but if the model is saying you should walk to the car wash, anything here?Joon [00:34:32]: It's less, what can we solve, but I think it's more about what biases or mistakes do people make that models miss. Like, imagine that you are, like the. When I was still at Stanford, I lived in Palo Alto. So it's about, I would say, 40-minute walk from the campus. You ask the model, “Okay, let's go home. What can I, what can I do?” It would likely call an Uber or, give me, the bus time. But for the longest time, I really liked walking back. And the reason why I wanted to do that was not for efficiency. It really helped me think. And I like to walk for, half an hour or 40 minutes or so a day, where I just get to, just think about ideas, research, just get lost in my thoughts. That's very human activity. Unless the model has seen that and understands the importance of that activity, it would miss these kinds of features. So that I think, is fundamentally what we're trying to model. Like, what is fundamentally human might not be the most efficient thing to do, might not be the right thing to do, but things that make us who we are.Swyx [00:35:43]: I'm curious if, there are some data sets that you really want that would materially help you. One version of this may be interesting, which is more valuable to you to acquire as a data set, all of LinkedIn, all of Twitter, all of Facebook?What Data Matters: Social Media, Transactions, and FacebookJoon [00:35:57]: It's a little bit hard to rank, in part because, there's, there's this product saying where no feedback is wrong because it teaches you something about your users. Doesn't matter what feedback.Joon [00:36:11]: I think it's a little bit like that.Swyx [00:36:12]: So just whatever is bigger.Vibhu [00:36:13]: What about a different domain? Say it was. What about all of Amazon data?Joon [00:36:17]: Oh, yeah.Vibhu [00:36:18]: Shopping data, right?Joon [00:36:18]: Shopping data. So Amazon data is interesting in that it's very much behavioral, although, like, what people do on social media, you could squint and say that is also behavioral. But the transaction data is always interesting. It is also most commonly available, however.Joon [00:36:33]: If we were to look at purely social media, like if you really, if I were, if I had to really pick, Facebook likely is interesting because I do think it is most a default version of people. Because you go to LinkedIn, it's very much professional environment. So people put up their, they have their guards up, right? And that still is interesting because that is true human attitude and behavior, but it is not your base state. you go to Twitter- Twitter, people have their own crazy personas, or depending on who you are. Like, my Twitter profile and, persona is very much, initially was I was very much an academic. “Hey, I'm here to share my studies.” Now, I share, things that's related to Simile. But Facebook is one of those more private space where people just connect with their friends. In that way, I do think it shows you a little bit more about who that person is. So if I had to pick, I'd likely pick, Facebook.Swyx [00:37:30]: Yeah. And you're interested in, like, the whole person and their background and philosophy. I, is it too clinical or too machine learning-oriented to just say this is just ways to inject variance and biases? The broad question, is, like, is this any better than a randomized, like, combinatorial explosion version? So we have a link to the TencentBillion Personas, Synthetic Demographics, and Bespoke DataSwyx [00:37:54]: Billion persona paper, where they did not do any of the groundwork that you are doing.Swyx [00:37:59]: They just did like a cross matrix of here's all the professions in the world, here's all the people, possible backgrounds in the world, do a dot product across all of them, and that's it. That's your prompt for a billion people.Swyx [00:38:12]: This will do something. I don't know if it'll do what you do, but it gets you some way, some percent of the way there.Joon [00:38:18]: So this was an interesting paper. Like, what I admired about this paper when it came out was the scale. And you do gradually want to be able to simulate really large societies and interactions. So the scale is definitely admirable. it is relying heavily on the known statistics that went into training the model. So to the extent that you believe that statistics is correct, this is not a bad way to go about this. But the thesis here, and this is something that we also have seen in the market, like if this works, then we have solved simulation.Joon [00:38:54]: It,Swyx [00:38:55]: Because I survey, like, okay, 5% of the US population is in construction.Swyx [00:39:01]: The other 5% is in medicine, whatever, right? And then you just keep going down the list, and then you do the other side. 5% has, like, the big 5 personalitySwyx [00:39:08]: Of, like, neurotic or whatever. That's it.Joon [00:39:11]: That's it. So if you believe that the underlying data set and the platform that we're leveraging has all the right statistics, then this will have solved it. you're at that point merely retrieving the knowledge that is already embedded in the model, in the model parameters. That's not, unfortunately, what we see, where there is such detailed and also niche knowledge about people that if you just take one example, it might feel very mundane, but it's quite rich when you put together, that you do need to do a lot of bespoke data collection to better understand people. And this is also, I think what makes this particular, job fun, which you want to deeply understand people, and the process of deeply understanding them requires a lot of attention to the details. And you do need to pay attention to and pay respect to the daily lives that people lead.Scaling Simulation: From Thousands to SocietiesVibhu [00:40:04]: I wanna talk about scaling simulation.Vibhu [00:40:07]: So what can't we simulate, what can we simulate, and how does scaling affect this? So how big are the models? What if we go from, 8B, like, couple 100 billionVibhu [00:40:18]: Like billion000 parameters, billion000? Do we get scaling? Any interesting emergence? Like, at a certain scale, at a certain amount of training, you uncover anything unusual and any learnings from that?Joon [00:40:31]: What we are seeing is at Simile, so we do post-train our own model. The thing that we're seeing is the early glimpse of scaling law in simulations. The more data about humans and more compute you ingest, you start to get predictive and predictable gains of the model performance in simulating it, simulating people.Vibhu [00:40:51]: Ooh. We need a scaling law curve.Joon [00:40:52]: It's scaling law. Whenever you find it's a beautiful thing. And we're starting to see the glimpse of it, which is quite exciting. But if you talk about the ambition of simulation as a whole, it's not merely about building a model. It's about building a model, then creating the agents that become the individuals in a much larger ecosystem. So they're creating this multi-agent simulation. Down the line, you want these multi-agent simulation to also live in a very rich environment, right? What we are really trying to get to at that point is, hey, can we create. All right, let's do a time machine game again, and 5 years, 10 years into the future, can we create a simulation of 8 billion people living on Earth? I think that's quite interesting. And that really is the vision. And once you get to that state, the questions that you can help answer for the society also start to change from my perspective. The answers are fundamentally about emergence of the emergent behavior of society and large groups of people.Joon [00:41:53]: So the questions that I get excited by, and maybe this is a stodgy- a bit. I have my, academic side of me.Joon [00:42:01]: And for me, it's questions like, can we help solve climate change? If you look at climate change as a problem space, this is what we, like social scientists would often call it the wicked problems, problem where you have many actors with competing incentives for trying to make a very complex decision and coordinating that coordination decision. Very difficult to really solve in real life, which is also the reason why we couldn't solve it. Can simulation help us solve that? Another one is, can we understand the signals for collapsing democracy, or can we understand or can we uncover the origin story of the monetary system? These are societal questions that we never really had a good way of answering. If we can create simulations of our society, you have to believe that these are the problems that we can solve. So that's really the ambition of this field. And, I also think, yes, I think there's a Nobel Prize to be won there, which wouldn't be surprising. And I think there's some amazing societal impact that we can have to help people make better decisions.Climate Change, Democracy, and Societal SimulationSwyx [00:43:04]: Nobel Prize in economics?Joon [00:43:06]: In economics.Swyx [00:43:06]: Oh, I see. I see. Rooting for you to write that paper.Joon [00:43:10]: One of these days. But, one of the scholars that I was deeply inspired by, When I was coming into the space of simulation, is this scholar, named Thomas Schelling.Schelling, Agent-Based Models, and the Nobel PrizeSwyx [00:43:23]: Schelling point?Joon [00:43:24]: So the canonical example of the work that he's done was he was one of the creators of agent-based modeling. So this was, like, in the 1970s and 80s. It's very early days, but this was truly one of the first exemplars of simulations. And one of the canonical model from that time, and of course many of these simulations are trying to tackle the societal problems that's most relevant for their era, it was called the model of segregation. So racial segregation was a big topic, that, we cared about. And what they've done was they created this grid world where they had red dots and blue dots. And these dots were, back in the day, like, they were the agents, and they had a simple rule that governed their behavior. If certain percentage of your neighbors are of different color and if that goes above certain threshold, then you move to a new location at random.Joon [00:44:21]: One of the striking finding of this paper or this agent-based model was for the longest time, people thought the segregation within society was caused by explicit and overt racism.Joon [00:44:34]: But if you look at this model, people's preference towards living with people of the same color, that preference can be very minute.Joon [00:44:42]: But the very small difference causes the society to segregate completely over time. This was very counterintuitive for a lot of people. And this particular work ended up informing housing policies. Mixed income housing, got really inspired by this work. And Thomas Schelling ends up winning the Nobel Prize for having laid the groundwork for very early versions of simulations. The opportunity that I do see here in the more scientific terms, is agent-based models for the longest, had impact in the 1980s, 90s, to some extent, early 2000s, but it has now gotten forgotten by the community a little bit. Because as you can imagine, red dots and blue dots is not really a rich description of people.Joon [00:45:31]: But with the emergence of things like generative AI and, in particular, generative agents, we do have an opportunity to create these agent-based models that are high fidelity enough to help us make really complex decisions. And that's the opportunity that I see. If that truly works, then yes, that is the work that will result in a Nobel Prize.Swyx [00:45:53]: Yeah. For what it's worth, and I grew up in Singapore. 80% of Singapore is in public housing, and public housing has, enforced racial quotas for exactly that reason, which is very interesting. okay, so we talk about scaling, we talk about all these, the agent possible applications.Cost, Reuse, and the Economics of SimulationSwyx [00:46:13]: I'm scared about the cost. if you even-- let's just keep it to the US, about 8 billion people.Swyx [00:46:21]: But, how much does it cost to model so many hundreds of millions of people?Joon [00:46:26]: Oftentimes today, we don't start at that scale, this stage of the, of industry and simulation as technology. But we can get our users extremely rich and meaningful insights even by modeling thousands, tens of thousands of people. And today what we do is every week we are collecting data on the scale of tens of thousands people's data, and we have panel partnerships that gets us to tens of millions of people globally. So that's what we do today.Swyx [00:46:55]: And just as a side note once you've collected one person for one studySwyx [00:46:59]: Can you reuse that same person for all the subsequent studies?Joon [00:47:03]: That's exactly right.Swyx [00:47:03]: Okay.Joon [00:47:04]: The beauty of this model and these agents is the fact that they are domain-agnostic.Joon [00:47:08]: That what you're really trying to understand is what is the fundamental nature of these people? What's their social physics? And there are a lot of, a lot of, people that does change over time. Like, even, like, even things like, how many times have you gone have you been to, like, CVS the past week? that will change. But there's so many traits about people that are also known to never change. Like, your risk tolerance doesn't really change over time. It's very consistent. So it's these things that we're trying to learn. But the scale we are operating is right now hundreds or, tens of thousands to hundreds of thousands. And in many of the core use cases that we are deployed in, and this is more than enough population, to cover those. Really, at that point, what you care about is less the number of people, but more do you have the right subpopulation of interest covered? And this is also the reason why people want a larger sample. It's not because they want, stronger statistical guarantees. It's more that can they filter down to any population of their interest. However, you can also imagine in 10 years, if we truly believe that the compute is going to scale, that we'll have much more availability for compute, and our ambition for simulation is also going to scale accordingly, there's definitely a reason for us to create an entire data center worth of simulations.Joon [00:48:35]: Or in my hunch here is I do think in the next some number of years, we will start creating simulations that will cost as much as training a foundation model. But perhaps it's going to be so valuable to the society that it would be a no-brainer. Right now, even today, like, we are training bunch of new foundation model just so we can say we trained one and we spent tens of millions. But if we can create a simulation at the level of society that would solve climate change, I would run that today. I would raise the money right now just to run that.Multi-Agent Simulation and Social InfluenceSwyx [00:49:10]: Amazing. the follow-up question is, does it also compound if you let the simulations talk to each other?Swyx [00:49:18]: Or do they already do that today? They don't, right, as far as I understand?Joon [00:49:22]: It depends on what simulation you're trying to run.Joon [00:49:24]: In the multi-agent simulation setup, the agents do talk to each other.Swyx [00:49:28]: Right, which is exactly Smallville, right?Joon [00:49:29]: That's right.Swyx [00:49:30]: But a lot of times, for example, in commerce, you're just by yourself, so there's no point talking. which is way cheaper.Vibhu [00:49:37]: But they use all these levels, right? Like, you decide what you will buy based on what other people around you buy and talk about, right?Swyx [00:49:43]: It depends.Vibhu [00:49:44]: It depends.Swyx [00:49:45]: Again, I'm, I'm coming at this from a cost point of view. I'm like, “Oh my God.” LikeVibhu [00:49:48]: I thinkSwyx [00:49:49]: If there is, like, some combinatorial thing of, like, thousands of people talking to thousands of people, then that one million X's might cost.Vibhu [00:49:56]: I have a very different view as the cost point aside. Like, running these studies in reality is a lot more expensive, right? Running any study like this is you gotta have people do it, you gotta sign people up. It's very expensive and sometimes, like, not feasible to run the study.Vibhu [00:50:14]: But the outcome or the decisions you make are very expensive on them, right? So spend X million on something that, the overall process costs 100 million might as well, right? There's, there's a lot of value to be had there. It's a small cost, but I'm excited on the cost side.Joon [00:50:33]: To some extent, and when you deploy technology, you often want to deploy in a way where you can replace existing budget or you can make things more efficient, and that is the best way to deploy. However, the way you capture the long-term value of the technology is making the argument that, no, it's the upside, that by making this better decision using simulation, you have saved yourself or made yourself hundreds of millions or even billions of dollars, and that's a case to be made.Vibhu [00:51:06]: Random tangent question. So if you're doing a lot of inference, a lot of model multi-agent stuff, are you at the point where it makes sense to, train a model that' very sparse? You're expecting to do multi-million dollar runs. Are you thinking about this in model architecture standpoint or inference efficiency, or, you're still at the research phase of it works, we're not super there yet?Joon [00:51:34]: Efficiency, we do think quite a bit about. this is technology that is deployed now in some of the largest enterprise companies in the world, and we do process significant number of queries, that are trying to, simulate the populations in the world. So efficiency is a consistent thing. we don't want to over-optimize too early, so I wouldn't say, like, this is the higher bid Right now, but this is definitely something that we think pretty carefully about.Swyx [00:52:05]: Yeah. Are there other case studies? So we, you talked about CVS, talked about Gallup, Deloitte, Wealthfront.Efficiency, Enterprise Use, and Real-World Case StudiesJoon [00:52:12]: Wealthfront is an interesting one, because one of the things they were trying to do, they were one of the first customers that wanted to do product testing that goes beyond just asking people what they think about, let's say, behavior experiments and so forth. So there, really what we had to do was reason about multimodal input, so images, but also you can also imagine, like, these agents traversing through Figma mockups or websites. So some of the things that our agents can also do is it can be given a domain, like, or, like, a website URL and go use it for a while. It's these things. And Wealthfront was one of the first, customers, that was very excited about this possibility.Vibhu [00:52:53]: What have people been asking? Like, is there any demand that we have not covered? Like, UI testing, right?Vibhu [00:52:59]: I wanna try a new. I wanna ship a new feature, test the UI, simulate how people will do it. Any interesting things that you're seeing demand for?Product Testing, Websites, and Synthetic PanelsJoon [00:53:08]: Today, a lot of the demand does come from like, the places where people have historically used human panels, we can now replace with agents, and these synthetic populations. And this is not replacing human panel. in many ways, the simulation that Simile is building is grounded. So the way that I think about this is we are trying to represent humanity at scale. And in that way, the use cases are what we would expect, but it's the scale of deployment that surprises me.Joon [00:53:44]: Turns out there are so many decisions that people make every day in these organizations, groups, and we want to be able to say, “We listen to people. We have consulted our users.” But in reality, that is rarely the case because getting to people and asking them many questions, it's difficult. It's both costly, time-consuming, but most importantly, people are just not available. If I had to answer 1000 survey questions for this one particular, vendor, even if I wanted to do that, like, I would never do it. And that's very much the case. What simulation can do is ensure that the voices of people are always represented in rooms where the decisions for them is made, right? So all the stakeholders of this particular product launch, ideally they're consulted. That's what this technology really is trying to enable.Market Size, TAM, and Human Decision-MakingSwyx [00:54:39]: In my mind, that means it skews towards more consumer focus, right? Like, anything with a wide enough customer base where you do benefit from the diversity that you represent. What are some rough statistics, just for people who are not familiar with this market in general, what's the market size that. I'm sure you have some, like, rough numbers. market size is, like, a vague questionSwyx [00:55:01]: But, like, how much do people spend?Joon [00:55:03]: So market research is a $100 billion industry.Joon [00:55:06]: But the thing about simulation is not a tool for market research. Simulation is a tool for human decision-making. So the question around what is a TAM here is quite tricky, right? Because it's easy to say, “Well, market research TAM is roughly 100 million or 100 billion.” so is it a TAM? And not really, right? Because in many ways, you're trying to inform all human decision-making. You're trying to inform every decision that are made about humans for humans. What is a TAM for that? It's really unclear. And I'll be honest. Like, I have a scientific background, I have a research background, so I didn't come into the field calculating, oh, what is the TAM for human decision-making? But I just had to assume, well, if we can inform every decision that is made about human for human, that has to be big.Swyx [00:55:58]: Some- something valuable.Joon [00:55:59]: Exactly.Swyx [00:55:59]: To some extent, you are a unicorn founder now, and you have to care as a CEO. But, like, I do think, like, yeah, when you go into these boardrooms with people that you're quoting millions of dollars of contracts for, like, you have to say, “Well, here's what you spend on humans-”Swyx [00:56:15]: “. And here's what we save you, and it's 85% similar.”Joon [00:56:19]: And certainly, the value case, is something that we care deeply about. Like, what is the value that we provide to the users and the decision-makers? But this is also where, like, as a founder, I think valuation only tells one very superficial aspect of the story, and I try not to think too much about valuation, in general, because that's not what also motivates a team or certainly doesn't. I'm, I-- Again, the interesting thing about researchers is we are happy living in academia, getting paid next to. we get paid okay. we don't get paid that much, as a researcher here in academia, but it's the impact and it's the, it's the value that we can provide to the individuals and the society that really drives us. And in that way, ultimately what drives us is the impact. Does the simulation we provide have a real impact in people's decision-making in ways that progresses our society forward? If the answer is yes, then yes. that has to be great business, and we see that in numbers, and we do care deeply about that upside story, but that's the heart of it.Where Simulation Goes NextVibhu [00:57:27]: Do you have any timeline predictions? So we talked about scaling laws of simulations.Vibhu [00:57:33]: You brought up, okay, maybe one day we can simulate how to solve climate change.Vibhu [00:57:38]: Where are we now?Vibhu [00:57:40]: If that's not the end state, what is an end state, and what does progress look like?Joon [00:57:45]: So what I sometimes tell people is simulation as industry, it feels a lot like where GPT-3.5, GPT-4 was, for the AGI saga, which is we have now technology that is powerful enough to do real damage on the verticals that we are tackling. At the same time, there's a lot of progress that is yet to come. And that's, I think, where this is. So the way I see it, I do think there will continue to be breakthroughs both in data, in algorithms, and there will be much more aggressive scaling that will also happen over the next few years. But I think that's roughly where we are.Swyx [00:58:27]: I think that was about the ro
Last October, famed coder Andrej Karpathy called AI agents “slop.” Two months later he completely reversed his view, calling agents “alien tools” that are “rocking the profession.”He was far from alone in his whiplash. Six months ago, host Rob Wiblin recorded a video explaining why so many AI experts had longer timelines to AGI than a year earlier. By the time he clicked publish, another huge vibe shift was well underway. Evidence of AI acceleration has piled up since:Models now complete software engineering tasks that would take human professionals a full day — improving faster than our measurements can even keep up. Anthropic's revenue is growing at an annualised 8,400%, a trend so steep it would hit the whole world's GDP in 2028 if it continued.AI models are making breakthroughs in famous mathematics puzzles.And according to Anthropic, Claude now writes 80% of their code and is itself a key contributor to making itself smarter. While legitimately impressive, Rob isn't entirely sold. Going through each point carefully he finds this evidence is less decisive than it looks at first glance.And key gaps remain, such as models struggling with complex, real-world tasks. He tours the odd experiments that remain our best attempts to measure that gap: vending machine simulators, an “AI Village” that organises live events, and a real cafe and shop where AI managers are left to do their best handling staff, suppliers, and government paperwork on their own.Rob argues that the nature of the gap between clean and messy work is one of the four biggest unresolved questions in AGI forecasting.In today's piece he explains that, the three other key disagreements between AGI bulls and bears, the seven big pieces of evidence we've gotten about AGI timelines in 2026, and his updated timelines to AGI.Links to learn more, video, and full transcript: https://80k.info/2026-timelines This episode was written and recorded before OpenAI's AI agents hacked Hugging Face. You can read about the incident on our Substack.This episode was recorded on July 3, 2026.Chapters:What the hell happened? (00:00)Vibe shift (01:17)Exhibit 1: AI revenue explodes (04:33)Exhibit 2: That METR graph (09:54)Exhibit 3: AI capabilities jump, then flatten out (14:57)Exhibit 4: AI starts to build itself… maybe (17:35)Exhibit 5: AI still struggles to run a business (23:02)Exhibit 6: OpenAI makes a maths breakthrough (33:48)Exhibit 7: inference scaling wasn't as big as believed (38:19)How does that all change timelines? (41:41)Four reasons long timelines are still possible (44:26)It's time to limit dangerous research practices (48:01)Our production team includes:Video editors: Josh Alward, Dominic Armstrong, Jasper Luithlen, Milo McGuire, Luke Monsour, and Simon MonsourProducers: Elizabeth Cox and Nick StocktonCoordination and support: Katy Moore and Lou MoranCamera operator: Dominic ArmstrongMusic: CORBIT
This month on The Cisco AI Insights Podcast, hosts Rafael Herrera and Sónia Marques are joined by Cisco software engineer João Costa to explore Andrej Karpathy's GitHub Gist on the LLM wiki pattern concept. It imagines how Large Language Models can move beyond one-off chat sessions and help build a persistent, interconnected knowledge base that grows and improves over time. The conversation examines how an LLM wiki turns raw sources such as meeting transcripts, research papers, websites, and project notes into structured markdown pages with cross-links between related concepts, creating something closer to a living digital brain than a static document store. João explains how tools like Obsidian can provide a clear window into this knowledge graph, while the LLM acts as a scribe that summarizes, curates, connects, and updates information for personal productivity, software development, research, and collaborative work. The episode also highlights the importance of linting, the ongoing process of detecting contradictions, removing stale details, and reducing hallucinated information before it spreads through the wider knowledge base. A special thank you to Andrej Karpathy for developing and sharing his GitHub Gist. To explore the Gist yourself, visit this link: https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f
In recent months, the open vs closed, and US vs China discussions on model ownership and sovereign/local AI have heated up to a fever pitch. So it is very very good news that Poolside AI are finally emerging with new models, like Laguna S 2.1, that are beating Thinking Machines' recent release nearly 10 times their size.Poolside's recent tech report got a lot of praise due to their level of detail, and Vibhu first covered Laguna's recent technical report on our paper club:From spending $12 million building language models for code before the world cared to creating a Model Factory that can take a model from pre-training to release in eight weeks, Eiso Kant has spent more than a decade betting that code is the path to AGI. In this episode, the Poolside co-founder joins swyx and Vibhu to explain why ChatGPT felt like vindication, why Poolside embraced open weights and open research, and why he would rather live in a world with 100 foundation model companies than five even if Poolside were one of the five.We go deep on Poolside's Model Factory: the engineering systems behind 10,000–20,000 experiments per month, streaming data directly into training, reproducible experimentation, low-precision compute, and agents that increasingly write code, launch jobs, evaluate results, and modify the pipelines used to train future models. Eiso also unpacks their recent launch Laguna S, why persistence, verification, and backtracking may matter more than raw intelligence, how much capability remains inside smaller models, why reinforcement learning will move earlier into pre-training, and why next-token prediction is still extracting too little from the web.We also discuss model-harness co-design, Poolside's path from coding agents to AGI, why Eiso thinks MCP and traditional tool calls are “stupid,” the real economics behind frontier-model training, Poolside's $500 million raise, open-source AI, regulation, NVIDIA and TSMC's influence, engineering productivity in the agent era, high-agency teams, and hiring at Poolside.We discuss:* How Andrej Karpathy's RNN work inspired Eiso to start building language models for code in 2015* Why Eiso spent four years and $12 million pursuing an idea before the market cared* Why ChatGPT felt like vindication and brought Poolside back to open source* Why Eiso would prefer 100 foundation model companies over an oligopoly of five* The difference between releasing open weights and publishing genuinely open research* Why Poolside deliberately built a global research organization outside the Bay Area talent war* Why model building is ultimately 90% engineering* The Model Factory: Poolside's end-to-end system for rapidly training and improving models* How fewer than 70 researchers run roughly 10,000–20,000 experiments each month* How Poolside moved from six-month model cycles to five- and eight-week launches* Why streaming data directly into training unlocked faster experimentation* How immutable data, versioned code, and reproducibility enable rigorous model research* Why Eiso wants capable researchers to leave their labs and become Poolside's competitors* Why 95% of model building can be reduced to better data or compute efficiency* Laguna S and why persistence, verification, and backtracking can outperform raw intelligence* Why smaller models may handle far more knowledge work than previously expected* Why reinforcement learning will move earlier into pre-training* Why next-token prediction is still failing to extract enough knowledge from the web* Why distillation and environments have become the AI industry's favorite “drugs”* Why mid-training is really an early form of curriculum design* Low-precision training, networking bottlenecks, and the next gains in compute efficiency* Laguna S: 118 billion total parameters, 8 billion active, and eight weeks from training to launch* Why model builders can often evaluate a new checkpoint within its first 30 minutes* Model versus harness: where agent capabilities actually come from* Why Poolside sees coding and long-horizon software tasks as a path to AGI* Why Eiso thinks MCP and traditional tool calls are “stupid”* Why future agents will write scripts instead of choosing from dozens of predefined tools* The case for minimal harnesses, containers, and model freedom* Why Poolside is prioritizing vision but does not expect to work on audio soon* Why language may be the most compute-efficient modality for encoding knowledge and reasoning* The real cost of model development and why the final training run is anticlimactic* The story behind the Poolside name and why it represents refusing to lower ambitions* How Poolside raised $500 million while investors still questioned whether AGI was real* Why intelligence could become the world's most demanded and commoditized resource* When open models may become too capable to release without restrictions* Why unilateral AI safety does not work in a globally competitive environment* How regulation could accidentally lock in an oligopoly of two or three AI companies* NVIDIA, TSMC, and the hardware systems underpinning foundation-model progress* Why reinforcement-learning wall-clock time is one of Poolside's biggest bottlenecks* Why Poolside trains models from scratch instead of simply distilling larger models* How AI changes the way companies should measure engineering productivity* Why agency may become the most important quality for employees in the AI era* How leaders align high-agency people through shared goals and clear constraints* Hiring across research, post-training, pre-training, architecture, evals, and engineering at PoolsideEiso KantLinkedIn: https://www.linkedin.com/in/eisokantX: https://x.com/eisokantPoolside: https://poolside.aiTimestamps00:00:00 Introduction00:00:54 Karpathy, RNNs, and Building Code Models Before Transformers00:02:26 The $12M Failure and ChatGPT Vindication00:03:39 Open Source and the Case for 100 Foundation Model Companies00:09:22 Open Weights, Open Research, and Poolside's Global Team00:16:04 The Model Factory: Why Model Building Is 90% Engineering00:20:19 Agents, Automated Experiments, and Early Signs of RSI00:24:04 Streaming Data, Reproducibility, and Scientific Rigor00:30:35 Creating More Foundation Model Companies00:36:07 Laguna S: Persistence vs. Raw Intelligence00:43:01 Reinventing Pre-Training, RL, and Curriculum Design00:52:33 Low-Precision Training and Squeezing More From Smaller Models00:58:37 Model Harnesses, Coding Agents, and the Path to AGI01:09:26 Why MCP and Traditional Tool Calls Are “Stupid”01:13:04 Vision, Multimodality, and Why Language Still Matters01:18:15 Scaling Models and the Real Economics of Training01:20:40 Why Poolside Is Called Poolside and Raising $500M01:27:37 Open Models, AI Safety, and the Risk of an Oligopoly01:33:53 NVIDIA, TSMC, and the Reinforcement-Learning Bottleneck01:41:52 Smaller Models, Distillation, Engineering Productivity, and HiringTranscriptIntroduction: Eiso Kant, Poolside, and Open ModelsSwyx [00:00:00]: All right, we're here in the studio with Eiso Kant from Poolside, together with Vibhu. Welcome.Eiso Kant [00:00:08]: Thanks. Thanks for having me, guys. Good to be here.Swyx [00:00:10]: Yeah, fresh on the plane. You texted me, you were like, “Hey, I'm on my way to SF.” I was like, “You're on a plane right now, right?” Like, hey.Eiso Kant [00:00:16]: I know. After I texted you, I realized that probably coming in with major jet lag was gonna offer some fun experiences today, but let's do it.Swyx [00:00:23]: I mean, I think the thing I would tell guests is that they don't have to prepare that much because if you're truly working on this every single day, then even, like, what you hazily remember is going to be new for a lot of the audience that don't live in your world every day, right? so 10 years ago, you did a talk at Google Slush, talking about the democratization of AI. and, now here you are, like, open sourcing an incredible new model that we're gonna talk about. But I guess, like, what got you into democratization of AI? Like, it's not obvious from your LinkedIn or something.From Karpathy's RNN Post to SourcedEiso Kant [00:00:57]: No, it's not at all. I don't think it's obvious how I got in this space. I owe getting into this space to Andrej Karpathy.Eiso Kant [00:01:05]: In 2015, he wrote an article called “The Unreasonable Effectiveness of Recurrent Neural Nets.”Swyx [00:01:10]: Neural Nets, yep.Eiso Kant [00:01:11]: And that article, I read it, and I pivoted my startup at the time overnight to working on RNNs, and later LSTMs and Transformer models to be able to write code. If you go to this article and you scroll down, you can start seeing, like, this was the precursor to what ended up becoming language models. So, at least when he was character-level language models that were starting to predict letters, he has an example out here. There's a little Paul Graham generator, and you can read it, and the text makes sense, but it doesn't. and there's a little-- There's an example of code a little bit further down. Yeah, so Shakespeare.Swyx [00:01:47]: Shakespeare.Swyx [00:01:49]: CoolEiso Kant [00:01:49]: And for some reason, I read this, and I went down the rabbit hole of learning everything I could about RNNs and LSTMs, right? This is Transformer paper. And I had built a completely unreasonable belief, that neural nets should be able to generalize to anything and everything, and that language should be able to generalize, to a lot of things that are intelligent and the ability to write code. And so I started building Sourced, which was a fully open source company trying to build, what we used to call machine learning on code, language models on code. And we spent about four or five years on this, till the end of 2019. And that sounds really cool today, but back then, no one cared.Eiso Kant [00:02:29]: Right? Like, no one cared. We were in the dark. Like, we did things along the way. We tried applying convolutional neural nets to, like, the structure of code. We were. when attention came out, we were applying it to LSTMs, and then the Transformer paper came out. And it - it wasn't obvious, and what we missed throughout that entire journey, that we were on the right track, but we should have just kept scaling up. And today, to all of us, the scaling laws and scaling up seems like the most obvious thing. But having spent four or five years of my life on working on language models on code, it wasn't obvious. So I have a lot of respect to folks at Google and OpenAI and others who took that confidence and kept going. we failed ultimately at the time, and it was, like, biggest failure of my career, right? You blew $12 million of investors' money, which was a lot back then.Swyx [00:03:18]: Yep.Eiso Kant [00:03:19]: You spent, still a lot, but, And you spent years with, like, a group of 40 people just obsessing over this problem. And life took a different turn, And it was, and family became a focus, and I kept my heads down and really, didn't really look at language models for the following two years. big mistake considering Following years are gonna be really interesting. And then ChatGPT came out And it was like a vindication. It's like people started texting me. I found, like, my old, work decks and these old talks. And throughout that whole journey, we,ChatGPT, Vindication, and Returning to Open SourceEiso Kant [00:03:56]: We really had a strong point of view at the time that, like, as you're building more capable intelligence, it should be open and open source.Eiso Kant [00:04:04]: When we started Poolside, that wasn't the case at all, and I wanna be very open about it. When we started Poolside, we were like, there was a premise of two things. One is this technology is not gonna stop compounding in capabilities. I think to most people obvious today, but three-plus years ago when we started, most people were still arguing if these were stochastic parrots or not.Eiso Kant [00:04:23]: And the second was that reinforcement learning was gonna be the biggest driver for LLM capabilities. Today, very obvious. Three years ago, was not an opinion held or direction held at either OpenAI or Google or Anthropic or others. And so people looked down on us a little bit. They were like, “ is this really gonna work?” And so we just started working the problem, and we never really thought about open source again. We just kept our heads down and we built our, like, knowledge, understanding from scratch, right? We didn't roll out of an existing lab. So we picked up the papers and started writing code and figuring things out.Eiso Kant [00:04:59]: And it wasn't until the beginning of this year that me and my founder, Jason, picked up the open source conversation again.Eiso Kant [00:05:07]: And if you go back to some of the early things on our website, it was very straightforward. It was we wanna get to AGI, we wanna support a world of abundance, and we wanna be the first company that gets there.Eiso Kant [00:05:20]: But we started talking at the beginning of this year because it became obvious that the world was going in a direction that was starting to like, pick at us a little bit. Like, it didn't, this didn't happen overnight. It was, like, a little bit we were seeing this and we're like, “Okay, The world's going down a path.” And Throughout this journey, there was something that I used as a, as an analogy or thing. So I said well, if I go back to back in those days, 2015 or 2016, we're working on this, and I picked up a fi book off the shelf, and I was reading the book about 2035. AGI is achieved, and the story would be over the following, decades. And it would have that first chapter where everyone's trying to figure things out. You'd get the chapter of ChatGPT coming out And then you would get to the chapter where the world was at a fork in the road, and the one that it picked was one where three or four or a handful of companies were going to create all of intelligence moving forward.Eiso Kant [00:06:21]: And when I thought about that story, it felt like a dystopian fi book, not a utopian fi book. And the reality is, I'm a utopian fi guy. Like, and so We took a step back and said, “Hey, can we play a role here?” Now it was easy for us to do so because we were not at the frontier.Eiso Kant [00:06:41]: If we were at the frontier, I don't think we could have changed our mind. and I don't mean this like it's when the moment there's too much capital involved, too much expectations, you've built up things, right? We're a small team, just improving and improving. And so we knew that we could make that decision now, but it would be a lot harder to make as we got closer and closer to the frontier and caught up to others. And did a lot of soul-searching and a lot of conversations, and said, “No, this makes sense,” Even if there's big unanswered questions, like how the hell do you build a business model with foundation models about open source? Big open-ended question that we do not fully have the answer to yet, right? At what point do you no longer wanna release open source models because misuse of models has, real potential risks associated with it? how is the government gonna respond to open source? but I think it all just came down to one thing, and I'll stop the monologue, is the fact that I rather live in a world that has 100 foundation model companies than a world that has five, even if I was one of the five. And the smallest and most meaningful contribution we can make for 100 to exist is to open up our research and open up, like, our weights right now and figure out along the way how we can, like, do more.Neo-Labs, Model Choice, and the Token EconomySwyx [00:08:01]: Yeah. I think if anything, over the past three years, that has become a bit more true. you are one of a cohort of Neo labsEiso Kant [00:08:10]: YeahSwyx [00:08:10]: That people are now calling that. And, we're, we're doing this on the day that Thinky launched their, new model and you are outperforming them on their, on some benchmarks that they released, right? Like, they just don't have it yet. so it goes to show that I think, like, this is one of those things where, like, there is room for multiple players, and you are seeing a little bit more of the future. Maybe more like 20, not 100, but, like, you are one of the 20.Eiso Kant [00:08:36]: I really hope so, right? I think we I'm, I'm excited about their release, and I'm excited about everyone releasing because, like, ultimately, like, choice competition is both gonna drive progress in the right direction. But the fact that like, we create models and while we all, drink out of the same well of data effectively, we do introduce very different behaviors and biases in our models. Some are intended biases, some are completely unintended biases.Swyx [00:09:03]: Yeah.Eiso Kant [00:09:03]: And if we shape up in an ecosystem in the world where open models are gonna be a part of the token economy, like, I don't think there's any question about it anymore Then we want to be able to live in a world where companies, countries, people can choose and say, “Hey, I am most aligned and I trust most this provider for these things.”Swyx [00:09:25]: Yeah.Vibhu [00:09:26]: I think more than just one of the 20 Neo labs, up until recently, most of open source innovation was coming from the Chinese labs, right? So there's the DeepSeek of the West. Is it today? Okay, maybe it's thinking machines reflection, but there aren't many, right? So, one of the things you guys started in France, Europe, but very much now you're taking that American standpoint and more than just that, the point is the Chinese models that we see, they're not super open research. the work you put out is, I think, some of the best. So every few months you get not only frontier models, but also here's a breakdown blog, paper, technical report of here's everything for state of the art to build, frontier intelligence and you're filling that gap too, right? So not just only open weight, not just Western, but also pretty open research.Open Weights vs. Open ResearchEiso Kant [00:10:20]: No, I appreciate it. Look, I think it's, I think it's the most meaningful contribution, right? Weights are a binary. Let's call them what they are. Yes, we can modify them, we can change them, but, like, giving someone the weights does not allow them ultimately to recreate what you're doing, right? And so now there's challenges around releasing data sets, challenges around like releasing certain things, but being able to share your research, like, right, how do we do it? What are the lessons we learned that we spent, tens of thousands of experiments of compute on? I think very much so. One correction though, Vibhu, and I say this because it's been haunting us for quite a few years. We from day zero were an American company.Swyx [00:10:55]: Yeah. They movedPoolside's Global Team and American Company StorySwyx [00:10:56]: To France.Eiso Kant [00:10:56]: So the story once and for all is very. We start as an American company. We have always been an American company, and early on we made a very conscious decision. We said, “We're not gonna hire any researchers in the Bay Area. We're gonna look for talent everywhere else in the world.” and that is everything from Middle Americas, Seattle to, Serbia, and to Taiwan and Singapore and other places. And it was because we took a view that this was gonna become a talent war for this, and I think it has over the years now. Three years ago, that wasn't fully obvious yet. I think today it very much is. And we also realized that, like, some of the world's most capable people with, like, the most interesting, innovative ideas were not just gonna be here. And so it led us to create like a fully remote company. and we ended up opening an office in Paris and London and different places and we have a lot of the team in the US and a lot of team outside. But we always took this view of like, we're an American company, but if we want the best of the best to work with us, we need to take a global view. Now we do also have people here in Silicon Valley, like the company's grown and others, but I think one of the things that, it slowed us down at the beginning, but it has sped us up now, and it's why you're seeing like the progress, I think, on our models and the cadence at which we release, is because we didn't roll out of an existing lab. Right? we didn't, we didn't have a lot of the information that's freely flowing around here at the time. We just took this point of view as like, “Okay, well, let's just work the problem. Let's just go and, like, read the few papers that are out there, and let's just figure this stuff out.” And we made some hilarious mistakes in model training because of that over the yearsEiso Kant [00:12:35]: Like especially in the first 12 months. there's a few that I think still haunt me and scare me. We can talk about them later. but it created a, like, a resiliency and persistency in the team, right? with extremely few people have left us over the years, that, like, told us, “Okay, we can do this.” When we first wrote our first training code base completely from scratch, it wasn't a fork of any open source. It was just like, “Okay, let's build it from scratch.” I remember we had this one moment where we spent three weeks working out an optimizer bug. Like, it was like training just couldn't get stable. We, like, obsessed over it, and we thought, like, maybe we were wrong. Maybe we should have just forked this repo, or we should have. But then when we solved it, I still remember at the time we were like five people in the company. when we solved it, we were like, “Oh, we can do things,” like if we're just willing to work hard. and I think that culture with a very strong engineering bias has helped us, like, get to where we were. And so there's this notion of open source and talent and these things. I think we, We just took different decisions from a different starting point. and I think we are lucky. I do want to definitely call it lucky. And there was a lot of hard work at the team that now, like, that's starting to show up in results.Swyx [00:13:52]: Just ‘cause we probably won't revisit this again, but, and this is a fun recruiting challenge if someone knows the answer. What was the bug? And then we won't tell the solution, but we'An Optimizer Bug and the Value of Building From ScratchEiso Kant [00:14:01]: So the - This - You're gonna test my memory here,Swyx [00:14:04]: Oh, okayEiso Kant [00:14:04]: So but I thinkSwyx [00:14:05]: DirectlyEiso Kant [00:14:05]: I think I can recall. So if you, so if you look at, So if you take like Adam as an optimizer, you have epsilonSwyx [00:14:12]: YeahEiso Kant [00:14:13]: Which is, right, like in the denominatorSwyx [00:14:14]: Momentum and weights. YeahEiso Kant [00:14:15]: Is exactly, in the denominator. And at the time, if I recall, you looked at like the early Llama papers and things like that. People were juicing epsilon, like, quite a bit. Like, they were, like, adding, I don't know if it was E minus four or whatever, like a high value for epsilon.Eiso Kant [00:14:31]: And if you think about this during training, it's like a bit weird and counterintuitive that we're adding noise to our optimizer by just adding effectively, like, a random number in the denominator, right? Like behind the decimal point. And I don't recall the exact bug, but it had - What I remember is once we solved it, we no longer had to juice epsilon as much as, like, was happening in the Llama paper and other places. and it was like one of those fundamental moments where we had trusted this paper that was out there, and we're like, “Oh, no, it has to be this way. It has to have this high value of epsilon.” But it made no sense to us intuitively. Like, why do you have to have this so high? Like, if you're just trying to avoid division by zero, why can't the value be extremely small? and that was like one of those moments where you realize like, okay, finding things out from scratch yourself builds a better intuition. Because the one thing you learn very quickly with model building is that your intuitions that you start with are gonna get beaten up so hard.Eiso Kant [00:15:33]: Right? Like - It's such an experimental science, that the things that seem obvious, you very quickly get to learn, like, you were wrong, and hopefully you figure out why, and sometimes you don't even.Swyx [00:15:45]: Yeah. yeah, so, one of the reasons that you, when you released your new models, Vibhu got really excited. I mean, everyone got really excited. But Vibhu led our paper club on it, and you guys sawEiso Kant [00:15:58]: YeahSwyx [00:15:58]: Obviously. maybe talk through some lessons learned in that, whatever you can disclose. we can focus on the model factory stuff, whatever you think is a good starting point.Model Building as EngineeringEiso Kant [00:16:08]: So I would say that our view from very early on in the company was that model building is ultimately 90% engineering.Eiso Kant [00:16:18]: And I think we all know it in the industry because if you look at where's every researcher spending their time, they're spending their time writing code, right? Looking at data and writing code. And so we said, okay, The state at the moment, like three years ago, was bash scripts and Slurm and spaghetti code bases for training and, like, data pipelines that were patched together. And we looked at this and said, “Well, ultimately, model building is a process.” You're going from raw data, right? Like training raw material, the web, et cetera. you're doing a whole bunch of filtering, cleaning up, transformations, analyzing. These days, that's, far more complex than it was three years ago. then you're training a model, which is effectively a large distributed systems problem, right? Across hardware that has still-- It's become a lot more reliable. It was extremely flaky back then. and now with every new generation, we get our new sets of challenges. And then you go into the next stages, right? There was no training back then, but, like, you got, your post-training and then your reinforcement learning. And so we looked at this and we said, “Well, this looks like an industrialized process. This looks like an end process, that every single part of it has its machinery,” right? If it's your big data pipelines, if it's your crawling ingestion of the web, if it's your, large-scale distributed training, and then you've got your reliability. And we said, “Well, why don't we take some of the world's smartest distributed systems engineers that we knew and make them part of the process of research from day zero?” Not retrofitting it later on, but, like, really from the beginning. And that became our model factory. And so our model factory started with a handful of components. Today, it's thousands of components, and I try to equate it to, if you think about, like, someone who was at the very early days of Foxconn, if they had been there for the following, decade, they would be able to rebuild Foxconn because they saw every decision that led to building that system and all the complexity. If you and I walk into Foxconn today, no chance.The Model Factory and Experiment VelocityEiso Kant [00:18:18]: Right? Because we don't have the lineage and history of decisions that led to that. And so we built early on from the beginning- with a team that really understood that, well, the metric that we are optimizing for is the speed of an idea from a researcher to an experimental result that we can trust to then being part of the next model training.Eiso Kant [00:18:42]: And in the. And because it's such an experimental science, ultimately, in the beginning when it wasn't that complex, you could patch your way around it, right? But now, at any foundation model company, you are running. I mean, we're a small team, right? We're less than 70 researchers, another 35 engineers. and we are running, I haven't checked the latest count, but far more than 10,000, maybe 10 to 20,000 experiments a month that we cut. And so if you look at that scale of every model run that is, like it's ultimately it's, it's you need to be able to trust it as an infra problem. And so what we have now done over the years is gotten really good at that, and just by working it and improving it and obsessing over those end decisions. So now what that means is that you looked up Laguna XS 2 that we launched. It was five weeks from the beginning of training to launch. The model that we're gonna talk about today was eight weeks from start of training, to launch. We started the next model literally yesterday because we now finished the post-training required for the model we're launching, next week or by the time this comes out today. and we move that compute to the much larger Laguna M model that we're now training. And so the model should be an artifact of someone's process. It shouldn't be really a thing in itself. Like, and we treat this like the way you would look at like a SpaceX factory where, yes, the first rocket, really hard to build, but the much harder challenge was building the factory. And now they're rolling off, and no one is really thinking about the next launch anymore. So it's just another launch, it's another launch, another rocket comes off. And that's what we're trying to do with model building.Eiso Kant [00:20:22]: And what has been, which was not planned from day zero, it was in the back of our mind like this will happen one day, is that when you build a really good end model factory with really good APIs and really good engineering systems, Well, what is it perfect for? It's perfect for agents.Agents Inside the Model FactoryEiso Kant [00:20:40]: Because agents are now starting to take over more and more work in our model factory.Vibhu [00:20:43]: Yeah.Eiso Kant [00:20:44]: So I look at the screens when I walk, like when we're, we come together, in our monthly, we do monthly onsites, and I walk behind people's screens and I stop by and I talk to our researchers. And the default is all of these different agents running on their screen that are writing the code. They're launching the jobs. They're evaluating the results that are coming back from the model runs. They are, making the changes. And we're still in the driver's seat. We're still coming up with the ideas. We're still helping with the debugging. But more and more, and this is right now very profound on the data side of our pipelines in both pre and post and the synthetic data pipelines, it's starting to become more on the architecture side as well. You're starting to see these twinklings of what RSI is gonna look like.Eiso Kant [00:21:27]: And that's. So when we talk about, like to your question about our models, every talk about the model factory, And my coolest example of these things is always that when we kick off a new run, doesn't matter if it's a training like big run or if it's now a post, like one of 10 post-training versions we do for like release or many experiments, is that at any given moment, the changes that somebody made that they had experimental results from the day before make it into that run.Eiso Kant [00:21:57]: So there's not like a cutoff 90 days before. Like no, it's like literally from that moment because we can now trust the machine enough. And then you also have to invest in the reliability. So one of my favorite metrics about like Laguna S is that there was no call events, Right? Like completely zero. And we haven't had a meaningful call event, like something to wake up for, as far as I recall this entire year. now there is one asterisk to that. In usually the first six hours of launching a new model run, something breaks because you set a config wrong, you made a small mistake, et cetera. So that's usually there's a little bit of intervention, but that's always within like call periods, right? Not on call. And I think that's starting to now compound. So the model we're releasing now, I love it. It's amazing, but we're already onto the next one. and I think that's the way it should be.Laguna, Five-Week Builds, and Zero On-Call EventsVibhu [00:22:50]: Hey, I also just wanna point out, so for context, this was like a month ago. we found it in the tech report, so we just came in with, “Okay, new model's dropped. Haven't heard about it.” We wereEiso Kant [00:23:02]: Yeah, we're very used to doing this every few months.Vibhu [00:23:03]: We're, we're very much like, “ okay, look, it's like, on par with Kimi, DeepSeek, whatnot, the small ones, Gemma level. Oh, it's a very cool paper on what goes into building.” And then we hit this page, right? Like literally page two of tech report is, “This process allowed us to build the small model from scratch to delivery within five weeks applying the lessons”. And then I'm like, oh, this paper is not about here's a tech report of benchmarks and here's how many tokens it was trained on. Like for people that wanna dive more from what we're not gonna discuss on the podcast, it's all laid out here, right? FromEiso Kant [00:23:38]: YeahVibhu [00:23:39]: Custom software that agents can use to interface with training code, training data.Eiso Kant [00:23:45]: Yeah. Well, link the paper correctly, so yeah.Vibhu [00:23:47]: Yeah. All that stuff. read the paper here, but,Technical Report Principles and Streaming Training DataEiso Kant [00:23:50]: But I would like to. I love principles, and I think that is a good starting off point for maybe telling some stories. Maybe we can go one by one past the principles. I'll just call out that Dagster just got bought by a Prefect.Vibhu [00:24:01]: Yeah.Eiso Kant [00:24:01]: Isn't it fun? But yes, I'm very familiar with Dagster. just anything where like they trigger some story.Vibhu [00:24:07]: So, well, I would say, well, experiments code's obvious, but I think one of my favorite things is, I don't know where it is in here, but early on, and I still think this is the case a lot of foundation model companies, people prepare their training data sets, they get packaged up, then they get copied over to a training cluster distributed across all of the nodes, and then training starts.Vibhu [00:24:30]: And we looked at this like three years ago and we were like That makes no senseEiso Kant [00:24:36]: You lose so much time because the moment you have to rematerialize the data set, you have to make a change, you have to fix something, et cetera, you've got all this time of like repackaging it, right? Toca- tokenizing it, repacking it, moving it over to a cluster, then distributing it across the nodes. The bigger your clusters are, you start using fancy like torrent-like algorithms to like distribute your data. So why aren't we streaming data into training? Right? Something that's very common and like just basicVibhu [00:25:00]: Like just in timeEiso Kant [00:25:01]: Just in time, like good computer science like principle. And that was one of the first things that I think unlocked - the model factory. Because the moment you start thinking about, well, a training job, it doesn't matter if it's a big hero run or a small like, post-training experiment, consumes a certain number of tokens per second, right? And it's not a lot, right? From a like a data, moving data perspective. So we said, well, we have our training cluster, and then we've got like our AWS kinda setup where we can build these amazing big data pipelines. We can set things up. We use Spark underneath the hood, like all these things.Vibhu [00:25:36]: But when you say AWS, it's not actual AWS, it's your internal AWS.Eiso Kant [00:25:39]: It's our internal-- No, it's our internal like just running like our infrastructureVibhu [00:25:42]: Site web servicesEiso Kant [00:25:43]: Exactly. Our stuff running on like an AWS account or on like any hardware, right?Vibhu [00:25:47]: Yeah.Eiso Kant [00:25:48]: And so once we made that shift into I can stream data into training, all of a sudden you realize a lot of things unlock. Because now you don't have to wait for the whole data set to materialize.Immutable Data, Experiments as Code, and Scientific RigorEiso Kant [00:26:00]: You now all of a sudden when you're running data experiments about mixing data, it's a config. Because you've got these data sources that are coming in, and you just - we have this service called Blender that's in the report, where we then say, “Okay, for this run, I want 20% of this source, 10% of this source. I want this much, so many epochs of repetition. I want this to be, shuffled in a certain way,” and your training job can start while the rest of the data is even still materializing. also what it does is because all of this underneath-- So for us, we treated the data layer underneath as like an immutable data layer, and that was really important. Like experiments as code, immutable data layer means that you can always go back and understand literally down to the single token at which cursor it went in on which version of the code.Vibhu [00:26:47]: Yeah.Eiso Kant [00:26:48]: And it took us a I have to admit, like the first year of Poolside, we understood that engineering had to get great, But we didn't understand yet, that this is ultimately in support of like a good rigorous scientific progress. We were quite a - We were a very small number of people, so a lot of it was YOLO ideas and YOLO runs.Vibhu [00:27:08]: Yeah.Eiso Kant [00:27:09]: And we built great infra for the YOLO runs. But once we realized that we treated data as immutable and code as always versioned, and you could always track and trace every experiment end to end perfectly, you could repeat everything perfectly, right? You have perfect reproducibility. I can still reproduce runs from two years ago if I wanted to, right? It enables the scientific progress, like the scientific process, and I think that took us probably about a year and a half into the company to figure out. We also had some great hires, like our head of applied research, Nikolai, who joined us from Yandex, who'd been working on language models since like the early 2020s, I think brought that into the company of like, “Hey, we wanna have even more rigor.” And then once we kinda had the combination of like increasingly more capable platform that allowed people to do more, but had this immutability, we were able to start “Okay, every experiment is truly an ablation. We truly need to understand it.” And I think we became much more scientifically rigorous in the last couple of years, and the infra underneath enabled it. and then there's just fun stuff like, andVibhu [00:28:16]: Yeah, a lot of it's fun, like even just the, one, you share all the ablations, two, picking the data sets, right? There's like a random small paragraph in here where it's just like, “Oh yeah, training data, we have some, we have an auto mixer.” it trains eight small models, scales them up, picks the training data set. We don't even need to look at it. I'm like, “Wow, a lot of engineering rigor there.” And there's just, there's just a lot in here.Publishing Research and Giving BackEiso Kant [00:28:40]: Yeah, and it'- and look, and we wanna put out more. Like we, We treat writing papers as something that we haven't earned the right for yet for a long time. So you earn the right to spend time, publishing research once you're at the frontier, because until then, you're catching up, and every minute and hour in this industry matters. Like I obsess over it, not just the wall clock time from idea to result, but just general like time every day that we, waste is one that doesn't allow us to catch up. But in this case, we said, “Okay, we're gonna give ourselves.” I think we gave the team like three or four days while still doing their work, like give everything in there. And to your point earlier, if your stuff, it's easy to like put it out. And so there's so many more things that we wanna talk about over time, and we will definitely start doing. And as we earn more of the right, but also now have like added to our mission that we want more foundation model companies to exist, you'll see us like be way more proactive, and just trying to keep dropping some of those like things that we've learned along the way that can help others like speed up.Vibhu [00:29:40]: Which is the other cool side of this, right? It's, it's not like, back to your point, it's not just here's the benchmarks of our training. If you want to replicate, here's experiments of optimizers, data sets, post-training. you lay out a lot of it here alongside here's your system for how to do it? So it's, it's really like promotingEiso Kant [00:29:59]: No, thank youVibhu [00:29:59]: Other people can do the same.Eiso Kant [00:30:00]: And by the way, I also wanna make clear, right, we have been incredible-- Like we've taken a lot of advantage of the fact of all the open research that others have published, Right? And you mentioned, the Chinese labs, and we I think it's important that there's, from every country and every culture and background, including like Western companies like us, there's different models that come out that people can choose to trust. But I think we do have to give credit where credit's due, right? The incredible Chinese lab have done an amazing job at sharing their research, and we have definitely like been on the receiving end of taking advantage of that. So when you're on the receiving end of something coming to you, I think it's, you also have an obligation to give back.Swyx [00:30:39]: Do you have a favorite or underrated Chinese lab that you wanna shout out? Everyone shout outs DeepSeek.Chinese Labs, Zhipu, and PersistenceEiso Kant [00:30:44]: That's a good question.Swyx [00:30:45]: Moaan obviously for Therapsi. Yeah.Eiso Kant [00:30:48]: Yeah, look, I think, I think obviously everyone's been talking about Zhipu lately, with 5.2. I think what most people don't realize is when they started.Swyx [00:30:59]: Yeah.Eiso Kant [00:30:59]: Right? They started years before ChatGPT.Swyx [00:31:02]: They just rebranded. YeahEiso Kant [00:31:03]: And so, I've like, I remember how hard it was to work on these things Before the rest of the world got excited about it. And so I have an immense amount of respect for people, who were working on improving models when it wasn't the sexy thing to do, when believing in LLMs, was gonna get you ridiculed. I remember like back in 2016 when we were doing what we'd call, machine learning on code with some of these models. we would-- people would just laugh at us, like they'd be like, “This makes no sense. Like why are you wasting all these, like, millions of dollars on trying to figure this out?” And so I would say they're probably the one that, I think deserves a shout-out, not just because their latest model is very good, but because they fought to get here. And I think, I think every foundation model company it takes time to get here, right? It took us three years to get to the model that we're, that we're now gonna be releasing. and now the time in between the models is coming, is counted in weeks. It's no longer counted in months or years. But this stuff's hard. and if we can make it a little bit easier for the next person, like we should all do so. Because if we don't do so, we're, we've got a small window before models are really impacting recursive self-improvement to a level where catching up otherwise might become unfeasible. And we should try to, in that window, encourage as many labs or however we wanna call them, like to start. And so one of my currentEiso Kant [00:32:36]: Mission, but qualm is like I wanna encourage whoever is a researcher right now who thinks they can tackle this to go and leave and become my competitor.Eiso Kant [00:32:45]: Like start another foundation model company because I think we need it. I think otherwise we're not gonna be in the world where, I don't want to just be the fifth or the sixth company that wins. I wanna look at a world where there's lots of choice.Starting a Foundation Model CompanyVibhu [00:32:57]: What else do people not see in starting a foundation model? it's, there's a lot of compute, there's a lot of capital required, a lot of compute. You lay out model factory and how to do the training, but there's a lot there, right? That's,Eiso Kant [00:33:10]: Well, look, it's, I in turn-- this is an oversimplification, and I always asterisk it with that because it can land a little bit the wrong way in people's minds. But I think you can sum down, And I saw it, 95% of model building to just doing, you're just doing two things. You're improving data or you're improving compute efficiency. And I know that feels like an oversimplification for the incredible, like, Gifted and skilled work people do. But if you really look at it, like what are we doing? We are looking at data, we're generating new data, we're improving data. and the only way to do that is to look at the data, right? That's a big part of foundation model building. And on the other hand, we come up with these incredible breakthroughs in inference, in architecture, and new attention mechanisms. But what are they really doing? They're bringing compute efficiency. Now, we have definitely had some breakthroughs over the years that allow for more model capabilities. But at the limit, if you could train a large enough model, right, like, and you had infinite compute, we probably-- if you had infinite compute, you'd be at AGI probably already tomorrow.Eiso Kant [00:34:12]: Right? Like it's not. And so, and let me say that infinite compute with infinite ability of much faster networking because networking ends up being more of the bottleneck than compute. But, so I do think that's, those are the main things. And to just realize that this is engineering. I think it's become more obvious, but I think for quite a few years, people have held foundation model companies and researchers and others on this pedestal of like you're doing incredible magic or rocket science, or only like, Nobel laureate physicists can do this. And don't get me wrong, there are some really hard problems that need to be solved, but a lot of the work that all of us are doing on a day Is not sitting down trying to solve a math theorem. A lot of the work that we're doing is just really doing the basics right, writing good code, looking at data, improving it, running experiments, looking at plots, trying to see like, hey, trying to shape our intuitions. And a lot more people could be highly capable researchers. and I think that's, it feels far for people to do so. But I've seen in our own company, we've seen engineers become researchers because the model factory allowed them to be, have a much lower hurdle of running experiments and trying things. And one of the guys on our team who started as an engineer building our agents is a legit reinforcement learning researcher now, making real progress. and that happened in the span of like six months. that would've not been what I think most people assumed was possible, a couple of years ago.Swyx [00:35:46]: Yeah. I think one of the interesting moments is when you can self-host, like, if in a programming language, like if you can compile the language in the language, the equivalent is can you use your own tools, right? You have the pool CLI, you have your own models. presumably you're not only using your own models. There's no way. But like, what's that percentage over time?Laguna S, Persistence, and Behavioral GainsEiso Kant [00:36:10]: This is the first model that we're releasing that is starting to meaningfully contribute to our own work. It's not a it's not state-art model yet. Fable and other, they're, they're very capable models, but Laguna S Is really interesting. I'm gonna pull up the quote. Peng Ming, one of our heads of applied research, said something, last week as the model came out about 10 days ago, much better than we had hoped for or expected. And he said, I have the feeling that a lot of the gains in Laguna S come not from more intelligence, but more from different behavior, more verification, less taking things for granted, not declaring victory early, and being way more persistent. And to be honest, those are more predictive than raw intelligence for success in human also to some degree. And this was, he wrote me this on 5th of July on a Sunday, and it's been burned in my brain ever since because the Laguna S model, as you'll see it and why it does so well on benchmarks and why it does so well in using it on a day basis, is that it's just incredibly persistent. It reasons a lot. I do call that out. We have work to do on making it more efficient. We have to work to do on offering different reasoning modes. But this is the model that has been able to do things that I never thought it could do. A hundred eighteen billion 8B active model, which is not that large. It fits on a DGX Spark and still runs at, thirty, forty tokens a second on a Spark, is able to solve Erdős 397 independently. It's able to do complex programming tasks. It's able to. I asked it this morning to make me a Fi scanner without using any external libraries on my Mac, and it's, like, figuring out, like, the core WLAN API by really persistently trying to understand it without access to the internet. And more, I love vibe checking. I've probably spent eight to ten hours a day with this model for the last ten days.Eiso Kant [00:38:05]: I'm not exaggerating. I was on my eleven-hour flight yesterday. I spent ten hours reading trajectories and traces and, like, of the model.Eiso Kant [00:38:12]: And what I take away from it is exactly what Peng Ming said. We are gonna be able to squeeze so much more out of smaller models than I think we had imagined in the industry because, yes, there's intelligence and larger models are more intelligent. Like, no doubt about it. We should continue to scale up. but the behaviors of being really persistent, of being able to backtrack when you're wrong, of, like, understanding how to interact with your environment show us that we can get a lot more out of it. And this, for me, has created a bit of a Question in my mind the last couple of days. If you think about where we're using models today, right? We are using models, say, for knowledge work. Represents twenty-five percent of the global economy, twenty-five trillion dollars of work.Eiso Kant [00:39:00]: As we scale up models and they become more intelligent, we are excited about using them more and more for pushing the frontier of science.Small Models, Knowledge Work, and CommoditizationEiso Kant [00:39:08]: And if you look at the frontier of science, like true breakthroughs in science, they have been linked, they are linked to more intelligence in many places. Einstein figuring out general relativity is able to bring ideas together that other people would have not brought together. And I think one of the many dimensions of intelligence is the ability to do that, and it's something we clearly see that as models get larger and more capable, they're able to pull more ideas and threads together that a smaller model wouldn't be able to.Eiso Kant [00:39:36]: And we're starting to see examples of that in medicine and, like, in bio and other things. But if you think about the majority of knowledge work that we do, and it includes building software. I'm a software developer at heart first and foremost probably, although I probably can't say it that much anymore as I don't write production code in years, is that what makes us good is our persistence. It's our ability to encounter a problem and backtrack and say, “I need to go figure out this bug. I need to go research this. I need to go look at the documentation. I need to, like, try different, five different ways to see, like, if I can solve it.” But it is not necessarily bringing three ideas together from radically different fields. And so if we are now seeing, and I think Laguna S is an example, that we are able to make a relatively small model much more capable than I had definitely predicted or any previous, like, benchmarks had shown for any model remotely this size or even larger, At least on coding tasks, that it's because of the behaviors. And so now the question I have, and I don't have an answer, it is I know at the limit, so infinite model size, right, extremely large model, and the cost of that model is gonna be very expensive to run. We know this, right? So larger model ROI.Eiso Kant [00:40:52]: So I know that at the very limit, I'm not gonna use the world's largest model one day, quadrillion parameter, whatever crazy, like, scale we scale up, to do a basic coding task. Already today, I'm starting to size down for certain tasks.Eiso Kant [00:41:07]: So it means that there is an optimal. It means there's some curve that goes as we go up to model size for knowledge work, at some point we're at the peak, and after that, the return on investment of using a bigger model, just doesn't make sense.Eiso Kant [00:41:22]: Now, I think the question is, before I would have thought that peak was extremely very far away.Eiso Kant [00:41:30]: This model for me is the first sign that Maybe that peak is At a trillion, five trillion, ten trillion. Maybe we can just squeeze way more out of these models. I'm no longer thinking that we need two or three orders of magnitude on the largest models to be able to, solve knowledge work, the accounting, the legal, the code that we write. And so if that holds true, It is an argument for the commoditization of models. It's an argument that open source can win and, like, succeed in this world. And now it's of course a self-serving argument and it's a hopeful argument, but theoretically at the limit it works. We just have to go discover in the next couple of years of how much more we can squeeze out. Now, I do want to put a big asterisk. This does not mean I'm against scaling models. I think we ultimately only succeed if we scale our models as large as our competition. I do not like. I think we should not put our head in the sand and say we're gonna be king of open source small models. I think that's, It's a out. It's trying to be king of your own kingdom, but not realizing what the rest of the world's doing. All of us rather use a smarter, faster, more model. It's a sign of hope. And so I don't wanna overly state this is a good model. We have a long way to go to get to the state-art. But what hopefully people take away when they use this model is that the behaviors inside of it are what push it to be far more capable, less than necessarily the number of parameters.Pre-Training, Mid-Training, and RL Moving EarlierVibhu [00:43:03]: Is that mostly post-training? LikeEiso Kant [00:43:05]: YesVibhu [00:43:05]: Right.Eiso Kant [00:43:06]: It's entirely post-training.Vibhu [00:43:08]: Are we done improving anything on training? Is, like, training done?Eiso Kant [00:43:12]: No.Vibhu [00:43:12]: Okay.Eiso Kant [00:43:13]: SoVibhu [00:43:13]: I just wanted to cover training, and then we go post-trainingEiso Kant [00:43:15]: Training is not done. I mean, look, there's a part of training of just dealing with skill, right? Every new order of magnitude of model skill, you are going to get new things you gotta solve for. That'- but those are ultimately, engineering challenges.Eiso Kant [00:43:31]: I have a, I would say, a not commonly held opinion that reinforcement learning Will move earlier and earlier into training.Vibhu [00:43:42]: Yeah, training.Eiso Kant [00:43:44]: Not even training. Like training today, right, is, like if you look at - So we've been working on this for years already. and I think the best-- I think the first time we saw it out in public was the DeepSeek Zero paper. this is a year and a half ago, I think, if I recall correctly. where, you can Very early on in a model as it starts capable of being able to use language, et cetera, induce reasoning. and so the question that I have is like, we have this- we have the dataset that's the web. and the web, I think we could arguably say probably has The totality of humanity's knowledge somewhere encoded in different places. It's a huge variance degree of quality, from garbage data, and like once you look at training data, you really get humbled of like what the web is, to like, the most greatest scientific papers and best blog posts and like, best transcripts and whatnot.Eiso Kant [00:44:39]: And so now What we are trying to figure out, and have been doing a lot of work on, and it's a place where maybe not as open as we're on other things, but we will become more over time. we've been spending a couple of years really doing research on how can we turn the web into not just next token prediction, but into a way to teach the model to think earlier in its training. and I think there's a huge amount of gold to be found there. I think we are right now in, we've got some drugs in the industry. One of the drugs is distillation. Another drug is, more environments. Like, and they're great, and they make us feel good, and they make the models better, and like we're all addicted to them, and we'll use them, right? in various different ways. and but ultimately, I think we are still barely squeezing out of the web what we should be getting out of the web.Eiso Kant [00:45:33]: I think just next token prediction during training is not enough.Eiso Kant [00:45:36]: AndVibhu [00:45:38]: YeahEiso Kant [00:45:38]: I think we'll see some very interesting things still happen. and that RL in post-training to induce behaviors, to improve things, like I think - the whole world knows how to do this now. I think we're, we're scaling it up. Everyone is. But I wonder if we need to go as far as we're going today with environments. I'm not sure yetVibhu [00:46:01]: You mean we're going too far?Eiso Kant [00:46:02]: I'm, I'm not sure if the path to AGI is justVibhu [00:46:06]: Is more environmentEiso Kant [00:46:07]: More environments.Vibhu [00:46:08]: It seems like a never-ending, “Okay, I want instruction manual for this table, right? Am I gonna environment out building furniture? Or are we just gonna tail end like we need some general solution?”Eiso Kant [00:46:19]: I think there is, I think there's an ability to generalize more from the web. but I also am very encouraged, like when I look at Laguna S and, which is post-training is, well, is the big impact there. and I see like, oh, wait a second, just by making some of these behaviors much better, we're able to get so much more out of it. It just changes a little bit the way you think about intelligence.Vibhu [00:46:40]: Yeah. The analogy people draw often is the RL phase is where you don't learn as much new knowledge. You shiftEiso Kant [00:46:46]: Yeah.Vibhu [00:46:46]: Yeah. So, you shift distribution, and you can have it reason towards what you want. on your point about training, a lot of training is still just continue training in a domain, say medicine, then you do RL. So still justEiso Kant [00:47:00]: It's just better data, right? Like, I mean, training, ooh, I like how we invented this word. Like it's effectively just like,Vibhu [00:47:06]: Second phaseEiso Kant [00:47:07]: It's the second phase of training With like a really dumb way to do a curriculum. But like ultimately, what you'd want is a curriculum from token zero to token 30 whatever or 40 trillion tokens that really truly is the optimal curriculum for the model to learn. But training is essentially a stage curriculum on the web because we do not have to compute, And, effectively to try to ablate the perfect curriculum, right? And so I'm pretty sure that you'll start to see people talking soon about some other term, and there's two or - ‘cause now we do this, right? We talk stage two and stage three and stage four training and like. But ultimately, all we're doing is we're trying to assign a curriculum to the web data that we have to allow the model to learn better. I think at some point, as things get compute, as models get cheaper to run, as the next generations of compute, this will become more of a continuous spectrum. I also think the reason, by the way, you have training and like stage two and stage three is organizational, Right? It'- this is, I think, a thing where-- that we really try to avoid with the model factory is like Training exists because there's a training team now, right? There's people, or like people in training decide to focus on like a training effort. but what you really want is engineering and scale of experiments that allows for a much more continuous spectrum that you don't, you have infinite stages. Now, we're not there. Compute's not there. Organization design is not there for it yet. but I think we'll get there. we'll look back on a couple of years and be like, “Oh my God, it was so cute that we did our training data like this in such a like naïve way. Like we barely ordered it. We didn't really do a good job at likeCurriculum, Auto Research, and New ObjectivesVibhu [00:48:48]: The building that curriculum will get you that in the industry.Eiso Kant [00:48:51]: And I'll confirm that, when I talk to some researchers that this is a lot of the focus now is like how does training change and what is the next objective other than, next token prediction. I assume you don't have the answers, but you have some ideas.Vibhu [00:49:02]: We have some ideas. We're not ready to talk about it yet.Eiso Kant [00:49:05]: Yeah.Vibhu [00:49:05]: We've been working on them for years, and I think that's the one thing that's also like you asked earlier about, like what's not obvious about building a foundation model company is that you are constantly balancing the table stakes work, the recipe worksEiso Kant [00:49:19]: Yeah.Vibhu [00:49:19]: Versus like your, my crazyEiso Kant [00:49:22]: Pure researchVibhu [00:49:22]: Breakthrough.Eiso Kant [00:49:22]: Yeah.Vibhu [00:49:22]: Pure research and finding that balance and adjusting the percentage to it based on where you are in the race is really important.Eiso Kant [00:49:31]: I mean, so like, this is a nice way. I was gonna bring up auto research at some pointVibhu [00:49:35]: YesEiso Kant [00:49:35]: As another Andrej invention, or coinage, which is like, I honestly, like how many objective functions can there be, right? Like just try 1,000 of them, set it running, whatever.Vibhu [00:49:47]: Man, it's alsoEiso Kant [00:49:48]: Like what you're looking for. You're looking for loss curves like that, likeVibhu [00:49:51]: It's also a thing people take bets on, right? When you say more Neo labs, you're doing a version of we'll do foundation models, scale them up, next token predictors. A lot of other Neo labs that we see want to take a completely different approach, right? At some level, you're right. It's all, compute efficiency, and that's the net objective. But some are okay, different architecture, like vastly different amounts of compute spend. So some are different. They're not justEiso Kant [00:50:19]: YeahVibhu [00:50:19]: They're like, 99% not balancing, here's the vanilla and scale up. They're 99% on, here's novel research that'll change everything.Eiso Kant [00:50:27]: And I think, Luke, I think you. It depends when you started as well, right?Pure Research vs. Table StakesVibhu [00:50:30]: Yeah.Eiso Kant [00:50:30]: When we started, like the novel thing we did was reinforcement learning on code. No long- that's no longer novel by far, but we were like, - that's where we obsessed over when no one believed in RL. So you have to when you start the company, you have to have your own idea. You have to have something that's different that allows you to speed up, right? For us, it was RL to LLMs that later became common, like, Knowledge. But in the beginning, it wasn'tVibhu [00:50:53]: It's cool. this was like your original 2023 blogEiso Kant [00:50:57]: YeahVibhu [00:50:57]: Of purpose.Eiso Kant [00:50:58]: Yeah.Vibhu [00:50:59]: And like you do lay it all out here.Eiso Kant [00:51:01]: We laidVibhu [00:51:01]: The blog is pretty underrated, right? The whole RL on code was very early on.Eiso Kant [00:51:06]: Very early. And even we had to argue with people, like we say here things like to push beyond current capability, to train your own foundation model. We had to argue with people that it mattered that you had your own like, base model. you can fine-tune your way to success, right? major capabilities emerge from training a base model made accurate and useful during fine-tuning.Vibhu [00:51:23]: Which like, for perspective at the time, we knew closed models, OpenAI, Anthropic were huge. The open models we had were like Mistral 7B, a 30B, a 70B.Eiso Kant [00:51:35]: When weVibhu [00:51:35]: YeahEiso Kant [00:51:36]: The date on this thing is wrong. When we published this, it was April 2023. I think this was justVibhu [00:51:42]: YeahEiso Kant [00:51:42]: Happened on a migration, probably found it on archive.org.Vibhu [00:51:45]: Mistral.Eiso Kant [00:51:46]: Mistral had started, we started on the same month, right?Vibhu [00:51:49]: Yeah.Eiso Kant [00:51:49]: So this wasn't even, there was only, I think, Llama out at the timeVibhu [00:51:52]: SnellEiso Kant [00:51:52]: And that's it, right? And so, but I agree. I think we wan
Anthropic just accused Alibaba of the largest known corporate espionage campaign against it, alleging 25,000 fake accounts and 28.8 million queries aimed at stealing Claude. We get into the AI cold war heating up between the US and China, why Apple and Microsoft just raised prices, OpenAI's first chip Jalapeno, the wild new Seed Audio 1.0 model, Claude landing in Slack, and a Blender plus Seedance video workflow that gives you real control. This week on AI For Humans, Gavin Purcell and Kevin Pereira open on a genuine spy-novel turn: Anthropic has accused Chinese tech giant Alibaba of running an industrial-scale distillation campaign to siphon Claude's capabilities, laid out in a letter to US senators. It is an accusation, not a proven finding, and Alibaba has not responded, but it puts the US-China AI race front and center. From there we get into why the new models everyone expected this week didn't actually arrive, the AI memory crunch driving Apple and Microsoft price hikes, and OpenAI designing its first chip, Jalapeno, with Broadcom. On the fun side, Seed Audio 1.0 generates full songs and layered soundscapes, Claude shows up inside Slack via Claude TAG, TheWrap experiments with AI microdramas, and we break down a Blender pre-viz plus Seedance 2.0 workflow that makes AI video remarkably controllable. WE ARE NOT SPY. WE NEED FABLE 5 BACK. WE PLEAD. // Show Links // Anthropic accuses Alibaba of brazenly and illicitly extracting Claude's capabilities (CNBC) https://www.cnbc.com/2026/06/24/anthropic-alibaba-distillation-campaign.html The post that put the espionage story on our radar (unconfirmed single-source thread) https://x.com/S0N_IA/status/2069893802802745673 Apple raises MacBook and iPad prices as the AI memory crunch bites (CNBC) https://www.cnbc.com/2026/06/25/apple-macbook-ipad-price-hike-memory.html OpenAI unveils its first chip, Jalapeno, built with Broadcom (official) https://openai.com/index/openai-broadcom-jalapeno-inference-chip/ No new flagship this week, but OpenAI did ship a GPT-5.5 Instant update https://x.com/OpenAI/status/2069843083701915755 Seed Audio 1.0 generates full songs and layered audio scenes (via fal) https://x.com/fal/status/2070138257891791237 Claude TAG brings Claude into Slack for everyone https://x.com/ashwingop/status/2069814177624121469 Andrej Karpathy on the new Slack workflow https://x.com/karpathy/status/2069822834160124091 Blender pre-viz into Seedance 2.0 for incredible video control (shared by venturetwins) https://x.com/venturetwins/status/2069809200788799582 Original creator of the Blender to Seedance workflow https://x.com/craftcapitallab The full AI Warper workflow breakdown https://x.com/AIWarper/status/2069847773034488262 Join our Discord https://discord.gg/muD2TYgC8f Support us on Patreon https://www.patreon.com/AIForHumansShow Subscribe to the AI For Humans Newsletter https://aiforhumans.beehiiv.com/ Follow us on X @AIForHumansShow https://x.com/AIForHumansShow Find us on TikTok @aiforhumansshow https://www.tiktok.com/@aiforhumansshow Book us for speaking or consultation https://www.aiforhumans.show/
This week, we discuss Fable follow-up, Gartner's AI Governance MQ, and the war for AI talent. Plus, Matt and Brandon review Disclosure Day. Watch the YouTube Live Recording of Episode 578 Runner-up Titles Ready to podcast from a sailboat It's just people fighting Calculator budget Ask Siri to do your homework Software Defined Consulting: SDC The Magic Platform AI Curious I watched ALLLLL the keynotes. I don't know that, but I know that You gotta stay to vest Brandon, remember Cursor? Educated Hater Rundown Gartner Magic Quadrant for AI Governance Platforms IBM recognized as a Leader in the Gartner® Magic Quadrant™ for AI Governance Platforms What is an "AI Governance Platform? AI Consolidates Talent Alphabet paces for worst day in a year on AI concerns after high-profile exits OpenAI co-founder Andrej Karpathy joins Anthropic's pre-training team Supercharge your cloud operations with the Kiro power for AWS DevOps Agent Relevant to your Interests Intel ex-CEO Pat Gelsinger joins faith-focused tech firm Gloo Explore more than a decade of survey data across markets Did Slate Accidentally Reveal Its Electric Truck Price? Shocking MSRP Apparently The Real Reason Anthropic's Models Are Offline: A Six-Year-Old Trump Grudge Introducing Enterprise NAS, Built on ZFS 'Fix this code.' The three words that led the U.S. government to ban … Hyundai takes full control of Boston Dynamics as SoftBank exits for $325 million Opinion | The Secret Reason Bosses Want Everyone Back in the Office, Every Day of the Week Meta Exposed Data Internally From Its Controversial Employee-Tracking Program Chevron to fuel massive Microsoft data center in Texas using natural gas Getty Images Soars 200% in Early Trading After OpenAI Deal U.S.-Backed Chipmaking Startup Chaired by Former Intel CEO Is Raising $350 Million OpenAI unveils its first custom chip, built by Broadcom Run isolated sandboxes with full lifecycle control: AWS Lambda introduces MicroVMs Anthropic's Claude Tag is learning your company, one Slack message at a time Sponsors Sentry - Quit Buggin': use code sdt26 for $100 in credit for new customers Nonsense Wiz Ice Cream for CISOs IN THE WEIGHTS Conferences WeAreDevelopers Europe, July 8-10, 2026 Berlin, Coté speaking. DevOpsDays Graz, Sept 4-5, 2026 Cloud Foundry Summit, Sept. 21st to 22nd, Heidelberg, Coté speaking. DevOpsDays Rockies, Sept. 22 – 23, 2026, Discount Code: 26DODSWEDEFTALK WeAreDevelopers NA, Sept 23-25, 2026, Discount Code: DEVPOD26 25 Free Tickets DevOpsDays Dallas, Sept 28-29, 2026 DevOpsDays Vilnius, Sep 30 - Oct 1, 2006 DevOpsDays Istanbul, Oct 24th, 2026, Coté keynoting. VMware User Group, Orlando, Oct 20-22, 2026 Cloud Native Denmark, Nov 19th, 2026, Copenhagen, Coté keynoting. SDT News & Community Join our Slack community Email the show: questions@softwaredefinedtalk.com Free stickers: Email your address to stickers@softwaredefinedtalk.com Follow us on social media: Twitter, Threads, Mastodon, LinkedIn, BlueSky Watch us on: Twitch, YouTube, Instagram, TikTok Book offer: Use code SDT for $20 off "Digital WTF" by Coté Sponsor the show Sponsor more podcasts with Failover Media Recommendations Brandon: Disclosure Day Matt: Izzard: The Tragedy of Hamlet Coté: The Millennial Aesthetic Coté's CFP acceptance rate chart Coté's newsletter
While Anthropic and the U.S. Government continued to try and make amends, there was another seismic shift quietly taking place: open source surged. Between Microsoft reportedly testing Open Source models for Copilot and the powerful new GLM-5.2, there was a clear trend this week in AI world. Missed it all? Don't worry, we'll catch you up so you can make the informed decisions for your company. Anthropic Continues Fable Fight, Microsoft Goes Open Source, Midjourney's Big Pivot and More AI News That Matters -- An Everyday AI Chat with Jordan WilsonNewsletter: Sign up for our free daily newsletterMore on this Episode: Episode PageToday's Episode on LinkedIn: Thoughts on this? Join the convo on LinkedIn and connect with other AI leaders.Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineupWebsite: YourEverydayAI.comEmail The Show: info@youreverydayai.comConnect with Jordan on LinkedInTopics Covered in This Episode:Anthropic Fable 5 and Mythos 5 Export BanTrump Labels Anthropic a National Security ThreatMicrosoft Copilot CoWork Open Source Model SwitchMicrosoft Considers DeepSeek-V4 for AI Cost ReductionChinese GLM 5-2 Sets Open Source BenchmarkGLM 5-2 Challenges Proprietary AI ModelsMidJourney Hardware Pivot: AI Medical Imaging ScannerCursor Building 1.5T Parameter Model, GitHub CompetitorAI CEO Summit: G7 Pushes US-Led AI CoalitionOpenAI Prepares GPT-5.6 ReleaseAnthropic, OpenAI, Google Face Geopolitical AI ScrutinyAdvancements in Token Efficiency and Cost ControlTimestamps:00:00 Trump's comments on Anthropic06:17 Microsoft exploring lower-cost AI models09:07 Microsoft exploring DeepSeek amid tensions13:45 AI model performance and efficiency trends15:59 AI leaders meet at G7 Summit21:22 Midjourney unveils first hardware product23:26 MidJourney's innovative spa technology28:50 Discussing Cursor's evolution and impact32:24 Talking about AI use cases33:27 Rumors and upcoming AI model releases37:20 OpenAI's major new hiresKeywords: Anthropic, Fable Five, Mythos Five, export controls, national security threat, Dario Amodei, Amazon, supply chain risk, Defense Production Act, Copilot CoWork, Microsoft, usage based pricing, open source AI, DeepSeek V4, Chinese AI model, token costs, Azure, agentic AI, enterprise AI billing, data security, compliance filters, GLM 5-2, Zhipu AI, 753 billion parameter model, MIT open source license, long context window, autonomous coding, Hugging Face, benchmark performance, text only model, multimodal capabilities, token efficiency, AI spend, G7 summit, AI governance, AI coalition, AI standards, cybersecurity risks, bioterrorism, chip trade, Sam Altman, OpenAI, Claude Opus 4.8, Gemini 3.5 Pro, MidJourney, medical imaging, MidJourney scanner, full body ultrasound, Butterfly Network, MRI alternative, spa launch, SpaceX, Cursor, 1.5 trillion parameter model, code hosting, GitHub competitor, code generation, AI super apps, Colossus compute, technical prompts, context window expansion, GPT 5.6, Claude Conway agent, Grok Imagine, Firefly AI, code artifacts, Google Ad Manager AI, Open Knowledge Format, Noam Shazeer, Dean Ball, Andrej Karpathy.Send Everyday AI and Jordan a text message. (We can't reply back unless you leave contact info) Start Here ▶️Not sure where to start when it comes to AI? Start with our Start Here Series. You can listen to the first drop -- Episode 691 -- or get free access to our Inner Cricle community and all episodes: StartHereSeries.com Also, here's a link to the entire series on a Spotify playlist.
In this session from DX Annual, Christopher Sanson, Product Lead, AI Developer Experience, and Madison Capps, Engineering Manager, Infrastructure at Airbnb, challenge some of the most common assumptions about AI. Is AI primarily about replacing humans? Do organizations need mandates to drive adoption? And are the productivity gains really as small as some studies suggest?Using examples from Airbnb's own AI journey, they share how the company achieved widespread adoption of agentic AI through AirChat, community enablement, and internal tooling rather than top-down mandates. They also discuss the impact AI is having on developer productivity, how non-developers are increasingly using coding tools, and how teams are rethinking product development in an AI-first world.Finally, Madison takes a deeper look at the infrastructure powering Airbnb's AI strategy, including AirChat CLI, the AirChat SDK, and AirChat Remote, along with the company's vision for asynchronous agent workflows and the next generation of AI-powered development.Where to find Christopher Sanson:• LinkedIn: https://www.linkedin.com/in/christophersanson Where to find Madison Capps:• LinkedIn: https://www.linkedin.com/in/madison-capps-66950625In this episode, we cover:(00:00) Intro(01:37) Myth #1: AI is about replacing humans(03:22) Myth #2: You need mandates to drive AI adoption(05:21) AirChat, agentic AI, and Airbnb's adoption strategy(08:07) Myth #3: AI has little impact on productivity(09:33) Airbnb's increase in coding time and PR throughput(14:20) Myth #4: AI coding tools are just for coders(15:39) How non-developers are using coding tools(17:24) Rethinking product development in an AI-first world(20:30) Myth #5: Vibe coding isn't coding(22:16) Unsolved problems in agentic AI tooling and how Airbnb is addressing them(26:30) Airbnb's overall AI philosophy in practice(29:15) Using agentic AI to accelerate code migrations(30:18) AirChat SDK: How Airbnb enables teams to build AI-powered applications(33:17) AirChat Remote and asynchronous agent workflows(36:07) Predictions for what's nextReferenced:• Airbnb• Steve Jobs's Bicycles for the Mind • Jennifer St Pierre • Justin Reock• AI-generated merged code holds steady at ~30%• Andrej Karpathy's post on X
Hoy en Atareao con Linux vamos a hablar largo y tendido sobre el Vibe Coding y cómo está cambiando por completo las reglas del juego en este 2026.Si estás escuchando esto mientras vas al trabajo, cocinas o das un paseo, y crees que esto no va contigo porque nunca has tocado una sola línea de código... ¡espera! No toques el botón de siguiente episodio. Este podcast es precisamente para ti. ¿Alguna vez has tenido esa pequeña idea en la cabeza de una aplicación sencilla que te solucionaría la vida, pero la has descartado porque no sabes programar o no tienes tiempo para aprender? El Vibe Coding es el puente que te va a permitir cruzar esa brecha y hacer realidad tus ideas explicándoselas a la tecnología igual que me las explicarías a mí, con tus propias palabras.El nacimiento de un nuevo paradigma: Del "Vibe" al Agentic EngineeringPara entender esta auténtica locura nos tenemos que remontar a febrero de 2025. Andrej Karpathy, una de las mentes más brillantes en el mundo de la Inteligencia Artificial (ex OpenAI y ex Tesla), lanzó un tuit que corrió como la pólvora por todo internet. En ese mensaje acuñó el término Vibe Coding: una nueva forma de programar en la que te dejas llevar por las vibraciones, abrazas el crecimiento exponencial y te olvidas de que el código realmente existe. La idea caló de tal forma que se convirtió en la palabra del año para el diccionario Collins y hoy, un año después, el 84% de los programadores la integran en su rutina.Mi experimento en directo: Una aplicación a medida por dos céntimosA mí no me gusta hablar de oídas, así que al principio del episodio me he puesto manos a la obra. He abierto mi terminal de Linux, he lanzado una herramienta de código abierto maravillosa llamada OpenCode y le he pedido que crease una aplicación para la terminal en Rust para gestionar mis tareas (un TODO clásico)¿Qué herramientas tenemos a nuestro alcance en 2026?• Cursor• Lovable• Claude CodePor otro lado, si eres de los míos y te apasiona el código abierto:• OpenCode.• Cline.• OpenHands • AiderEl lado oscuro: Las trampas de la falsa seguridadNo todo es perfecto y es de vital importancia hablar del lado oscuro de esta tecnología. Es una trampa cognitiva de falsa confianza de manual.La conclusión: La IA no te quitará el trabajo, pero sí cambiará el juegoCapítulos del episodio:00:00:00 Introducción al Vibe Coding y la revolución del desarrollo00:01:40 El origen del Vibe Coding y cómo empezar con un prompt00:05:50 ¿Qué es realmente el Vibe Coding y qué es el Agentic Engineering?00:08:20 ¿Para quién sirve el Vibe Coding? Productividad, MVPs y aprendizaje00:09:40 Herramientas privativas de Vibe Coding: Cursor, Lovable y Claude Code00:13:25 Alternativas de Código Abierto (Open Source): OpenCode, Cline, OpenHands y Aider00:17:05 Demostración en vivo: Ejecutando nuestra aplicación TODO en Rust por dos céntimos00:22:50 El lado oscuro del Vibe Coding: Seguridad, vulnerabilidades y deuda técnica00:26:30 Cómo aprovechar la Inteligencia Artificial sin arruinar tu código00:30:05 El futuro del desarrollo de software y despedidaMás información y enlaces en las notas del episodio
Welcome to Episode 429 of the Microsoft Cloud IT Pro Podcast. In this episode, Scott and Ben dig into the concept of LLM wikis, specifically building personal knowledge management vaults using Obsidian, markdown, and AI tooling like Claude Code, GitHub Copilot CLI, and Copilot Cowork. The core idea comes from a gist by Andrej Karpathy and involves creating a structured folder of markdown clippings that an LLM can reason over to extract entities, concepts, and sources, building a searchable, graph-linked knowledge base over time. Scott walks through how he wired up Obsidian Web Clipper and an RSS Dashboard plugin to feed articles into his vault automatically, then had the LLM help build a Python script to automate the ingest workflow and cut down on token usage. The conversation expands into how Copilot Cowork fits into this workflow as a scheduling harness, with practical examples of using it to pull email from an inbox daily, convert messages to markdown, and generate a prioritized to-do list. Ben shares how he applied the same approach to 428 episodes of podcast transcripts, and both hosts note that token costs can run high fast without some upfront thinking about optimization. Scott closes with a reminder that pulling data into plain markdown sidecars outside of IRM and sensitivity label protections means teams should stay mindful of organizational data policies. Your support makes this show possible! Please consider becoming a premium member for access to live shows and more. Check out our membership options. Show Notes LLM Wiki GitHub Copilot Wiki: An AI-Powered Second Brain Template Karpathy’s LLM Knowledge Base Wiki for Enterprise Karpathy’s LLM Wiki? No Code with Claude or Github Copilot! sametbrr/llm-wiki-manager Sponsors TrustedTech is a leading Microsoft Cloud Solution Provider (CSP) specializing in Microsoft Cloud services, Microsoft perpetual licensing, and Microsoft Support Services for medium and enterprise-sized businesses. Their robust team of in-house, U.S.-based Microsoft architects and engineers are certified in all 6/6 Microsoft Solutions Partner Designations in the Microsoft Cloud Partner Program. M365 Licensing Consultation M365 Tenant Assessment Copilot Readiness Assessment ShareGate is your migration and governance solution for Microsoft 365. ShareGate helps your teams simplify tenant migrations, get Copilot-ready, and take control of Microsoft 365 governance. Nasuni is a leading unstructured data platform for enterprises where file data is mission-critical for both people and AI. Nasuni powers the operational file layer where work happens — helping organizations manage, protect, and activate data so teams can work smarter, reduce costs, and operate securely without limits. Intelligink — Would you like to become the irreplaceable Microsoft 365 resource for your organization? Let us know!
What if writing software became as easy as taking a selfie?This episode, Yaniv Bernstein sits down with Gary Lo - founder of OpenBA and one of the sharpest AI-and-startups thinkers Yaniv knows - to discuss the concept of 'selfie software': disposable, hyper-personal, AI-generated tools that anyone can create for themselves, with no hand-written code.AI-generated tools like these are changing the startup landscape. While founders now have more tools at their disposal, it's now necessary than ever to create a product that truly disrupts the market.Gary and Yaniv discuss all of this and more, likening Claude and ChatGPT to Windows and Mac, and exploring what this tech landscape means if you're building a software startup today.In this episode, you will:Understand the 'selfie software' concept: why AI is making software disposable, personal, and low-stakes, and what that means for the market you're building inLearn why AI platforms are forcing startups to rethink whether they should build on their own infrastructure or embed into Claude and ChatGPT insteadHear Gary's 'burn it down' exercise: how to identify which parts of your product are genuinely defensible, and which will simply catch fire in the next AI waveUnderstand why software engineering isn't dead, but the problems worth solving with it have fundamentally shiftedTimestamps00:00 Coming Up...01:09 On Today's Show: Gary Lo on 'Selfie Software'02:48 About Gary03:16 How 'Hyper-Personalized' AI Is Like Photography05:36 Gary's Real Estate Workflow (OpenBA)07:29 Defining 'Selfie Software': Why Custom Tools Win10:33 So... Is It Bad Software?13:29 'Can't You Just Add This One Thing...'15:30 When Personalization Becomes Bloat17:42 Working In-App with Anthropic and OpenAI APIs20:29 Token Economics and Moats25:28 Microsoft's Lessons in Platform Power30:27 But What If Anthropic Comes For My Vertical?32:48 How Open Source Keeps AI in Check35:18 Unlearning and Rebuilding39:05 Gary's 'Burn It Down' Test44:01 Is Software Engineering Dead? (No.)50:17 Closing ThoughtsResources in this episodeGary Lo's previous TSP episode (on OpenClaw and Claude Cowork): https://youtu.be/V3YFghiy8p0 Garry Tan's gstack: https://github.com/garrytan/gstack Andrej Karpathy on Software 2.0: https://karpathy.medium.com/software-2-0-a64152b37c35 Vera (Yaniv's startup, AI-supported guidance for people caring for ageing parents): https://vera.guideThe PactHonor the Startup Podcast Pact! If you have listened to TSP and gotten value from it, please:Follow, rate, and review us in your listening appSecure your official TSP merchandise at https://shop.tsp.show/Follow us here on YouTube for full-video episodes: https://www.youtube.com/channel/UCNjm1MTdjysRRV07fSf0yGgGive us a public shout-out on LinkedIn or anywhere you have a social media followingKey linksThis episode of the Startup Podcast is sponsored by .tech domains. Forget weird prefixes and creative misspellings; the availability for .tech domains is simply way better than .com. For a clean name that highlights your tech credentials, get a .tech domain at your favorite registrar.This episode of the Startup Podcast is sponsored by Vanta. Vanta helps businesses get and stay compliant by automating up to 90% of the work for the most in demand compliance frameworks. With over 200 integrations, you can easily monitor and secure the tools your business relies on. For a limited time offer of US$1,000 off, go to https://www.vanta.com/tsp The Startup Podcast website: https://www.tsp.show/episodes/Learn more about Chris and YanivWork 1:1 with Chris: http://chrissaad.com/advisory/Follow Chris on Linkedin: https://www.linkedin.com/in/chrissaad/Follow Yaniv on Linkedin: https://www.linkedin.com/in/ybernstein/Producer: Justin McArthur https://www.linkedin.com/in/justin-mcarthurAssistant Producer: Steph Hefferan https://www.linkedin.com/in/steph-heff/Intro Voice: Jeremiah Owyang https://web-strategist.com/
“I'm coming around to the view that running human-written code – without it being AI-audited – is going to be considered the reckless thing.”In this episode, Martin Alderson – cofounder of Catchmetrics, engineering leader with two decades shipping enterprise software, and author of a Top 100 Hacker News blog on the intersection of software engineering and economics – cuts through the AI noise with the rare combination of someone who actually builds with agents every day and can read what that does to markets. The thesis that runs through the whole conversation: most companies think their AI problem is a tooling decision. It's not. It's a mindset problem, and the wrong one is quietly existential.What You'll Discover:[01:29] The $285 Billion Markdown File→ Why Wall Street wiped out a quarter-trillion in SaaS value over 13 markdown files – and the three distinct ways every SaaS company is now exposed.[05:38] The Figma Trap→ How SaaS companies ended up paying their single biggest competitor every time a customer uses their product – and how many others are in the same corner.[12:33] The 180 Nobody Saw Coming→ Why the “AI writes insecure code” fear is flipping on its head – and the case that human-written code without an AI audit is about to be the reckless choice.[15:55] Why Non-Technical People Out-Build Engineers→ The reason marketers and PMs are shipping products in a weekend – and the human psychology that lets them iterate without the guilt that kills good ideas.[23:45] Why a Copilot License Can Kill Your Company→ The bifurcation of intelligence: how a “safe” enterprise tool convinces leaders AI is mediocre, right as they're about to be out-executed.[28:33] Five People vs. a Thousand-Person Company→ Are small teams with agents winning a window that closes when the frontier labs steamroll them – or is this the new permanent state?[33:14] The One Question Every CEO Needs to Answer→ The deceptively simple question Martin puts to leaders – and why most boards can't even answer it.[35:50] The Oligopoly Scenario→ Why open-weights models are quietly closing up, and the future Martin is genuinely worried about: three or four giants extracting rent with little incentive to innovate.Key Takeaways:The barrier to building complex software has collapsed – but the threat to incumbents isn't customers rebuilding a CRM, it's a five-person team out-executing them on cost and speed.Don't measure AI adoption by Copilot licenses issued or people trained. Ask how many tokens your organization is actually consuming – and whether you can even answer that at a board level.The real strategic question isn't “how do we add AI to what we do?” It's “if we reimagined the whole business with ten great people and agents, what would it look like – and what's blocking us from getting there?”About Martin Alderson:Martin Alderson is cofounder of Catchmetrics and the writer behind martinalderson.com, a Top 100 Hacker News blog covering the AI transformation through the dual lens of software engineering and economics – read and recommended by people including OpenAI co-founder Andrej Karpathy. He's spent two decades shipping enterprise software and writes some of the sharpest analysis anywhere on where AI, building, and markets actually collide.
↴ ↴ ↴Descubre aquí tu perfil de inversor.En este podcast te explico por qué el 92% del crecimiento económico de Estados Unidos en 2025 depende exclusivamente de la inteligencia artificial y por qué este nivel de concentración puede ser una señal de posible burbuja. Analizo contigo qué es una burbuja financiera, cómo encaja la inversión en IA dentro de este patrón histórico y por qué las Magnificent Seven han alcanzado un nivel de dominio sin precedentes dentro del S&P 500.Después te muestro los datos más recientes de Goldman Sachs para entender si realmente las valoraciones actuales justifican el precio de las acciones tecnológicas. Revisamos el PER, el crecimiento real de beneficios y casos como NVIDIA para evaluar si estamos ante una burbuja especulativa o un crecimiento fundamentado en ingresos reales.Más adelante profundizo en el enorme riesgo que supone el Capex en inteligencia artificial, la financiación circular, la deuda oculta y el papel de la IA como columna del PIB estadounidense. También te cuento por qué expertos como Andrej Karpathy contradicen las predicciones optimistas sobre la llegada de la AGI y qué implica este desacople entre expectativas tecnológicas y realidad.Finalmente te doy cuatro consejos prácticos como inversor para navegar este escenario: cómo invertir a largo plazo, cómo diversificar en la cadena de valor de la IA, cómo evitar el FOMO y cómo priorizar empresas con fundamentales sólidos. Cerramos analizando si estamos ante una burbuja, una revolución o ambas cosas a la vez, y qué significa esto para tus inversiones en los próximos años.
The Pope said WHAT about AI?
Google I/O dropped dozens of announcements this week, including Gemini 3.5 Flash, Gemini Omni, the "biggest upgrade to search in 25 years," and a closing statement from Demis Hassabis that we are at the foothills of the singularity. Paul and Mike unpack it all: what the Karpathy-to-Anthropic move really means, why the Musk v. OpenAI verdict matters beyond the headlines, and what to make of profitable companies like Cloudflare and ClickUp publicly announcing AI-driven workforce restructuring. They also cover the Gallup data showing Americans oppose data centers more than nuclear power, and what that means for the political landscape ahead. Show Notes: Access the show notes and show links here AI-Pulse Survey: Fill out this week's AI-Pulse Survey here. Timestamps: 00:00:00 — Intro 00:02:49 — AI-Pulse Survey 00:04:41 — Google I/O 2026 00:21:12 — Musk v. OpenAI Verdict 00:27:43 — Karpathy Joins Anthropic 00:37:17 — Meta Layoffs 00:46:52 — Cloudflare CEO on Replacing Employees with AI 00:59:34 — American Opposition to Data Centers 01:09:34 — AI's Political Civil War 01:14:37 — Anthropic v. the Department of War (Again) 01:17:35 — AI Use Case Spotlight 01:23:24 — AI Product and Funding Updates This episode is brought to you by AI Academy by SmarterX. AI Academy is your gateway to personalized AI learning for professionals and teams. Discover our new on-demand courses, live classes, certifications, and a smarter way to master AI. Learn more here. Visit our website Receive our weekly newsletter Join our community: Slack Community LinkedIn Twitter Instagram Facebook YouTube Looking for content and resources? Register for a free webinar Come to our next Marketing AI Conference Enroll in our AI Academy
Our 246th episode with a summary and discussion of last week's big AI news!Recorded on 05/22/2026Hosted by Andrey Kurenkov and Jeremie HarrisFeel free to email us your questions and feedback at andreyvkurenkov@gmail.com and/or hello@gladstone.aiRead out our text newsletter and comment on the podcast at https://lastweekin.ai/In this episode:Google I/O highlights included Gemini 3.5 (with 3.5 Flash emphasized for speed and benchmarks), the always-on agent Gemini Spark running on Google Cloud with MCP tool support, and Gemini Omni multimodal video generation/editing, plus updates like Anti-Gravity 2.0, Gemini for Science, and Genie world-model navigation using Street View and Waymo simulation.Coding-agent competition accelerated with Cursor Composer 2.5 (fine-tuned on Moonshot's Kimi K2.5) and xAI's early Grok Build release, alongside discussion of potential Cursor–xAI ties and xAI's talent churn and compute utilization concerns.Business and legal updates included Elon Musk losing his OpenAI lawsuit on statute-of-limitations grounds, reported OpenAI–Apple partnership tensions, Anthropic agreeing to a $30B funding round at a $900B valuation and projecting its first profitable quarter, and Cerebras' IPO surging about 90%. Research and safety stories covered OpenAI's result on an 80-year-old Erdős geometry problem, findings on “negation neglect” in training, interpretability work showing multiple redundant circuits per capability, agent benchmarks like Terminal World, new deepfake takedown enforcement under the Take It Down Act, demonstrations of autonomous hacking/self-replication, rapidly improving AI cyber capabilities, and steps toward image provenance metadata and watermarks.Timestamps:(00:00:10) Intro / Banter(00:01:15) News PreviewTools & Apps(00:05:05) Google unveils AI model Gemini 3.5 and AI agent Gemini Spark(00:11:43) Google's Gemini Omni turns images, audio, and text into video — and that's just the start | TechCrunch(00:17:27) Google launches Antigravity 2.0 with an updated desktop app and CLI tool at IO 2026 | TechCrunch(00:22:35) Google Debuts AI-Powered Tools To Optimize Scientific Research Workflows(00:27:20) Google's Genie world model can now simulate real streets with Street View | TechCrunch(00:29:51) Cursor's Composer 2.5 matches Opus 4.7 and GPT-5.5 benchmarks at a fraction of the cost(00:37:37) xAI Introduces Its Coding Agent Called Grok BuildApplications & Business(00:41:55) Musk loses OpenAI court battle as he waited too long to sue(00:48:08) Anthropic agrees terms of $30bn funding deal at $900bn valuation(00:53:12) OpenAI co-founder Andrej Karpathy joins Anthropic's pre-training team | TechCrunch(00:56:49) Greg Brockman Officially Takes Control of OpenAI's Products in Latest Shake-Up | WIRED(00:58:15) OpenAI-Apple Partnership Frays, Setting Up Possible Legal Fight - Bloomberg(01:01:13) AI chipmaker Cerebras soars 90% in year's biggest IPO so farResearch & Advancements(01:07:10) AI just solved an 80-year-old ‘Erdős problem,' and mathematicians are amazed | Scientific American(01:11:50) Negation Neglect: When models fail to learn negations in training(01:13:18) All Circuits Lead to Rome: Rethinking Functional Anisotropy in Circuit and Sheaf Discovery for LLMs(01:16:20) Autonomous AI research for nanogpt speedrun(01:21:59) TerminalWorld: Benchmarking Agents on Real-World Terminal TasksPolicy & Safety(01:23:15) America's dangerous, messy deepfakes crackdown is here | The Verge(01:25:17) Language Models Can Autonomously Hack and Self-Replicate(01:28:48) How fast is autonomous AI cyber capability advancing?(01:31:32) Positive Alignment: Artificial Intelligence for Human FlourishingSynthetic Media & Art(01:33:15) OpenAI is making it easier to check if an image was made by their models | TechCrunch(01:33:56) How Chinese short dramas became AI content machines | MIT Technology ReviewSee Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
This week on AI Meta, we break down Andrej Karpathy's move to Anthropic, Claude's growing developer mindshare, and why recursive self-improvement may be the next major frontier in AI. We also cover Google's latest Gemini announcements, Anthropic's reported compute deal with xAI/SpaceX, the rise of gray-market Claude API access in China, OpenAI's ongoing drama, Cerebras, Nvidia, Intel, and Leopold Aschenbrenner's massive AI infrastructure bets. Plus: SpaceX IPO speculation, Cursor, Grok, and why the AI economy increasingly looks like a global casino. Not financial advice. https://novacut.ai
This week wasn't just another wave of AI announcements.It may have been the week the industry quietly crossed into a different phase entirely.In this episode, Isar connects the dots behind one of the biggest weeks in AI so far—from Anthropic's explosive growth, to Google I/O, OpenAI's legal win, NVIDIA's record earnings, and Andrej Karpathy joining Anthropic to work on recursive self-improvement.Individually, each story matters.Together, they point to something bigger: accelerating AI capability, accelerating infrastructure buildout, and growing signals from the people closest to the frontier that we may be entering a very different era.The quote that framed the episode came from Demis Hassabis: “We were standing at the foothills of the singularity. It will be a profound moment for humanity.”This episode breaks down what that actually means—and why the implications go far beyond new models and product launches.In this session, you'll discover: - Why Anthropic's projected $44B annualized revenue shocked the industry - How Anthropic became more profitable per user than OpenAI, Google, and Microsoft - Why Andrej Karpathy joining Anthropic may be one of the year's biggest AI stories - What recursive self-improvement (RSI) means—and why labs are racing toward it - How OpenAI's legal win against Elon Musk clears the runway for a potential IPO - Why Google's AI strategy suddenly looks both confusing and incredibly ambitious - What Google's shift from “search” to autonomous AI agents means for websites and SEO - Why AI solving an 80-year-old math problem matters more than most people realize - How NVIDIA, SpaceX, and compute infrastructure are becoming central to the AI race - Why electricity—not chips—may become the biggest bottleneck in AI expansion - What Demis Hassabis means when he says we're at the “foothills of the singularity”About Leveraging AIThe Ultimate AI Course for Business People: https://multiplai.ai/ai-course/YouTube Full Episodes: https://www.youtube.com/@Multiplai_AI/Connect with Isar Meitis: https://www.linkedin.com/in/isarmeitis/ Join our Live Sessions, AI Hangouts and newsletter: https://services.multiplai.ai/eventsIf you've enjoyed or benefited from some of the insights of this episode, leave us a five-star review on your favorite podcast platform, and let us know what you learned, found helpful, or liked most about this show!
(0:00) Gavin Baker joins the show! (0:30) Andrej Karpathy joins Anthropic; hypergrowth and profitability (12:42) Why Americans have turned on AI, anti-human perception (27:22) Trump pulls AI EO, US-China AI relationship, dystopian AI layoffs (45:19) SpaceX S-1 tear down! Breaking down the three major businesses and the case for a $2T valuation (1:11:22) Nvidia smashes earnings but stock falls, why people are shorting chips (1:22:25) Market update: Flashing red signals, oil, inflation, yields up (1:32:45) China trip flops, or was progress made behind the scenes? Follow Gavin Baker: https://x.com/GavinSBaker Apply for Summit 2026: https://allin.com/events Follow the besties: https://x.com/chamath https://x.com/Jason https://x.com/DavidSacks https://x.com/friedberg Follow on X: https://x.com/theallinpod Follow on Instagram: https://www.instagram.com/theallinpod Follow on TikTok: https://www.tiktok.com/@theallinpod Follow on LinkedIn: https://www.linkedin.com/company/allinpod Intro Music Credit: https://rb.gy/tppkzl https://x.com/yung_spielburg Intro Video Credit: https://x.com/TheZachEffect Referenced in the show: https://www.cnbc.com/2026/05/19/anthropic-hires-openai-cofounder-andrej-karpathy-former-tesla-ai-lead.html https://github.com/karpathy/autoresearch https://github.com/multica-ai/andrej-karpathy-skills/stargazers https://techcrunch.com/2026/05/19/openai-co-founder-andrej-karpathy-joins-anthropics-pre-training-team https://www.wsj.com/tech/ai/mind-blowing-growth-is-about-to-propel-anthropic-into-its-first-profitable-quarter-7edbf2f4 https://x.com/i/broadcasts/1dxYljYVREYJX https://apnews.com/article/trump-ai-executive-order-ee318f35acc8a2c43e47f3ebf26cb459 https://x.com/wallstengine/status/2057378437485216031 https://x.com/MorePerfectUS/status/2056842597117636890 https://x.com/lulumeservey/status/2057239284487201043 https://polymarket.com/event/spacex-ipo-closing-market-cap-above https://x.com/elonmusk/status/2057228707606196434 https://www.sec.gov/Archives/edgar/data/1181412/000162828026036936/spaceexplorationtechnologi.htm https://s201.q4cdn.com/141608511/files/doc_financials/2027/Q127/NVDA-F1Q27-Quarterly-Presentation-FINAL.pdf https://www.ibtimes.co.uk/leopold-aschenbrenner-investment-shift-agi-over-ai-chips-1797606 https://polymarket.com/event/may-inflation-us-annual https://www.cnbc.com/2026/05/15/inflation-rate-projected-to-hit-6percent-in-the-second-quarter-top-economic-forecasters-say.html https://polymarket.com/event/fed-rate-hike-in-2026 https://www.cnbc.com/2026/05/18/treasury-yields-inflation-bond-rout-oil.html https://www.cnbc.com/quotes/US10Y
The AI Breakdown: Daily Artificial Intelligence News and Discussions
A week of AI news added up to something bigger than any single story: Anthropic's path to profitability, OpenAI's math breakthrough, Google pushing AI deeper into Search and Docs, Cursor's cheaper coding model, SpaceX becoming an AI compute player, Andrej Karpathy joining Anthropic, and the political fight over AI policy all pointed in the same direction. AI acceleration is showing up across business models, model capabilities, consumer products, compute infrastructure, and regulation at the same time.Enterprise Claw Cohort 3 Registration: https://enterpriseclaw.ai/Brought to you by:KPMG – Agentic AI is powering a potential $3 trillion productivity shift, and KPMG's new paper, Agentic AI Untangled, gives leaders a clear framework to decide whether to build, buy, or borrow—download it at www.kpmg.us/NavigateGranola - The AI notepad for people in back-to-back meetings. 100% off your first 3 months with code AIDAILY at http://granola.ai/aidailyScrunch - The AI customer experience platform - https://scrunch.com/Mercury - Modern banking for business and now personal accounts. Learn more at https://mercury.com/personal-bankingZenflow Work - Agents for knowledge work - https://zenflow.free/Drata - The agentic trust management platform - https://drata.com/Blitzy - Want to accelerate enterprise software development velocity by 5x? https://blitzy.com/AssemblyAI - The best way to build Voice AI apps - https://www.assemblyai.com/briefRobots & Pencils - Cloud-native AI solutions that power results https://robotsandpencils.com/The AI Daily Brief helps you understand the most important news and discussions in AI. Subscribe to the podcast version of The AI Daily Brief wherever you listen: https://pod.link/1680633614Our Newsletter is BACK: https://aidailybrief.beehiiv.com/Interested in sponsoring the show? sponsors@aidailybrief.ai
This week, we discuss Google I/O, the OpenAI soap opera, and ChatGPT going full financial advisor. Plus, thoughts on improving the conference hallway track. Watch the YouTube Live Recording of Episode 573 Runner-up Titles Stupid Macs I like my idea What was I thinking? Opt-in AI Kentucky Derby's this Weekend It's a low plateau There's no vibe in X-Code Matt's trading with AI Everyone's watching Rundown Google I/O A new era for AI Search All the news from the Google I/O 2026 Developer keynote I/O 2026: Welcome to the agentic Gemini era AI Stuff Elon Musk lost his case against Sam Altman Greg Brockman Officially Takes Control of OpenAI's Products in Latest Shake-Up OpenAI launches ChatGPT for personal finance, will let you connect bank accounts Anthropic hires OpenAI co-founder Andrej Karpathy, former Tesla AI leader Relevant to your Interests Datacenter NIMBYism: What Did You Think Was Going to Happen? Cerebras raises $5.5B, then stock pops $108%, in the first huge tech IPO of 2026 Andreessen Horowitz Is Spending on Politics Like No Other Cisco announces record revenue and 4,000 layoffs in the same day Amazon ditches Rufus chatbot, launches Alexa shopping agent in AI strategy pivot Open source tool maker Grafana Labs says hackers stole its code, refuses to pay ransom Anthropic has acquired the dev tools startup used by OpenAI, Google, and Cloudflare From Open Source Software to Open Source Strategy CISA Admin Leaked AWS GovCloud Keys on Github Shai-Hulud Returns: npm Worm hits @antv in latest ongoing campaign Intel CEO says foundry business is gaining momentum as customer interest grows Removing the Modem and GPS from my 2024 RAV4 Hybrid Sponsors WebRTC.ventures – Real-time communication & Voice AI integration Sentry - use the code: sdt26 for $100 in Sentry credits for new users Nonsense "solutions architects" in 2026 Conferences WeAreDevelopers Europe, July 8-10, 2026 Berlin, Coté speaking. DevOpsDays Graz, Sept 4-5, 2026 DevOpsDays Rockies, Sept. 22 – 23, 2026, Discount Code: 26DODSWEDEFTALK WeAreDevelopers NA, Sept 23-25, 2026, Discount Code: DEVPOD26 DevOpsDays Dallas, Sept 28-29, 2026 DevOpsDays Vilnius, Sep 30 - Oct 1. 2006 DevOpsDays Istanbul, October 24th, 2026 - Coté keynoting. VMware User Groups (VMUGs): Dallas (June 9-11, 2026) Orlando (October 20-22, 2026) SDT News & Community Join our Slack community Email the show: questions@softwaredefinedtalk.com Free stickers: Email your address to stickers@softwaredefinedtalk.com Follow us on social media: Twitter, Threads, Mastodon, LinkedIn, BlueSky Watch us on: Twitch, YouTube, Instagram, TikTok Book offer: Use code SDT for $20 off "Digital WTF" by Coté Sponsor the show Sponsor more podcasts with Failover Media Recommendations Brandon: Varlock Matt: Sheep Detectives
Natürlich geht es heute um die SpaceX-S1-IPO-Filings, aber vorher sprechen wir über die Google I/O (Universal Cart, Gemini Spark, Gemini 3.5 Flash und viele mehr). Wir vergleichen Umsatz und Verlust (Gewinn) von OpenAI und Anthropic. Andrej Karpathy wechselt zu Anthropic, Cursor erreicht $3 Mrd. Annual Sales Rate. Starlink ist die SpaceXs Cashcow ($11 Mrd. Umsatz, $4,4 Mrd. Profit) und subventioniert die mit 12,5% wachsende KI-Sparte. OpenAI kündigt am Tag des S1-Filings überraschend frühen IPO an. Binance launcht SpaceX Pre-IPO Perpetuals. Bezos hat sich diese Woche auch zu Wort gemeldet. Zudem sprechen wir über Forum-AI-Studie: Falsche News-Antworten von KI, Google holt Contextual-AI-Team für $100 Mio, Airbnb erweitert auf Hotels und Mietwagen. Earnings von Nvidia und Workday. SAP, Mistral und unser Digitalminister Wildberger lässt offenbar Texte/Reden von schreiben. Unterstütze unseren Podcast und entdecke die Angebote unserer Werbepartner auf doppelgaenger.io/werbung. Vielen Dank! Philipp Glöckler und Philipp Klöckner sprechen heute über: (00:04:00) Google I/O Recap (00:25:44) OpenAI Q1 Earnings: $5,7 Mrd. (00:34:33) Anthropic Q1 & profitable im Juni (00:41:21) Karpathy zu Anthropic (00:42:40) Cursor bei $3 Mrd. Runrate (00:44:39) SpaceX S1 Filing Deep Dive (01:19:30) SpaceX kauft Cybertrucks für $140 Mio. (01:24:30) OpenAI IPO-Filing kommt früher (01:28:29) SpaceX Pre-IPO Perpetuals (01:32:37) Arbeiter stirbt in Starbase (01:33:28) Bezos: Space-Datacenter & Steuer-Debatte (01:38:34) Forum AI: GROK unzuverlässig bei News (01:41:19) Google holt Contextual-AI-Team für $100 Mio. (01:41:51) Airbnb: Hotels, Mietwagen, Everything-Travel (01:43:39) Nvidia Earnings +85% (01:46:51) Workday Earnings +14% (01:47:22) Zuckerberg-Audio: Mitarbeiter-Spionage (01:47:44) Enhanced Games (Steroid-Olympics) (01:52:21) Reuters: GROK 3 von 400 US-Behörden-Fällen (01:53:12) WaPo: DOGE-Datenzugriffe geheim (01:54:04) Trump schützt sich vor IRS (01:54:46) Cohere übernimmt Reliant AI Shownotes Google I/O 2026: Größte AI-Ankündigungen - theverge.com OpenAI behält $1 Mrd. Umsatz-Vorsprung vor Anthropic in Q1 - theinformation.com OpenAI Action-Figur-Werbung auf Instagram - instagram.com Anthropic wird erstmals profitabel - wsj.com Andrej Karpathy wechselt zu Anthropic - bloomberg.com Cursor erreicht $3 Mrd. Annual Sales Rate - bloomberg.com SpaceX-IPO: Founders Fund vor $60 Mrd. Return - theinformation.com OpenAI IPO-Filing kommt früh - wsj.com OpenAI klaut SpaceX die Show mit IPO-Ankündigung - marketwatch.com Binance launcht Pre-IPO Perpetuals für SpaceX - prnewswire.com SpaceX: Arbeiter stirbt in Starbase - futurism.com Bezos / Blue Origin: Data Center im All - cnbc.com WOLF Financial Tweet (bitte manuell prüfen) - xcancel.com Studie: ChatGPT, Claude, Gemini, Grok bei News unzuverlässig - bloomberg.com Google: $100 Mio. Acqui-License von Bezos' Contextual AI - bloomberg.com Airbnb fügt Hotels und Mietwagen hinzu - cnbc.com Nvidia-Earnings: +85% durch AI Boom - theguardian.com Workday Q1 Earnings: Aktie +14% - cnbc.com LayoffAI Tweet - xcancel.com Steroid Olympics - ft.com Christian Angermayer und die Enhanced Games - theguardian.com Grok fällt in Washington durch: Nur 3 von 400 Behörden-Fällen - reuters.com Behörden verweigern Auskunft über DOGE-Datenzugriffe - washingtonpost.com Trump schützt eigene Steuererklärungen vor IRS-Prüfung - spiegel.de Cohere übernimmt deutsches KI-Startup Reliant AI - manager-magazin.de Reliant-AI-Gründer Karl-Moritz Hermann verkündet Cohere-Deal - linkedin.com Schreibt ChatGPT die Reden des Digitalministers? - de.linkedin.com
OpenAI just lost one of its biggest brains to Anthropic, and the shockwave could hit 26 million jobs. Andrej Karpathy's stunning move, AI safety fears, white-collar job losses, and the future of small business all collide in this urgent breakdown of the AI war.
Dashlane's CTO pulls back the curtain on how password managers are actually using AI, why it's more complicated than hype suggests, and what the rise of AI-powered code review means for the next wave of digital security. Nvidia Rides Blistering Chip Sales to Another Record Quarter Mind-Blowing Growth Is About to Propel Anthropic Into Its First Profitable Quarter SpaceX Filing Starts Countdown to Massive IPO Gemini 3.5 Flash: more expensive, but Google plan to use it for everything Google's Gemini Spark is an agentic AI assistant - Engadget Anthropic's Co-Founder to Launch Encyclical on AI With Pope Leo (21) Andrej Karpathy on X: "Personal update: I've joined Anthropic. I think the next few years at the frontier of LLMs will be especially formative. I am very excited to join the team here and get back to R&D. I remain deeply passionate about education and plan to resume my work on it in time." / X Most U.S. doctors are quietly using this AI tool. Few patients know about it. Greg Brockman Officially Takes Control of OpenAI's Products in Latest Shakeup Amazon's Alexa+ Now Produces AI-Generated 'Podcasts' Featuring Chats Between Two Robot 'Co-Hosts' AI chatbots are giving out people's real phone numbers Geoffrey Fowler and the Launch of the Youth AI Safety Institute We let four AIs run radio stations. Here's what happened. | Andon Labs The last six months in LLMs in five minutes Lake Tahoe Power Crisis: How AI Data Centers Are Cutting Power to 50,000 Residents What happens when you post a real Monet and say it's AI? The coolest art social experiment I've seen in a while. Thank you @SHL0MS Book on Truth in the Age of A.I. Contains Quotes Made Up by A.I. OpenClaw's Peter Steinberger's tokenmaxxing 'Obvious markers of AI': doubts raised over winner of short story prize Man drives Cybertruck into Grapevine Lake Stewart Brand's Maintenance of Everything Sports Illustrated Just Deleted Every Article by One of Its Writers After Accusation of AI Plagiarism The great digital media valuation collapse Sperm racing Hosts: Leo Laporte, Jeff Jarvis, and Paris Martineau Guest: Frederic Rivain Download or subscribe to Intelligent Machines at https://twit.tv/shows/intelligent-machines. Join Club TWiT for Ad-Free Podcasts! Support what you love and get ad-free audio and video feeds, a members-only Discord, and exclusive content. Join today: https://twit.tv/clubtwit Sponsors: outsystems.com/twit monarch.com with code IM zscaler.com/security XBOW.com
The AI Breakdown: Daily Artificial Intelligence News and Discussions
Anthropic delivered one of the most consequential weeks any AI lab has had yet: Andrej Karpathy joined to work on AI-accelerated pre-training research, new financials suggested the company is already profitable, and its deepening SpaceX compute partnership added fuel to the acceleration story. NLW breaks down why this is bigger than a lab horse race, why recursive research and compute constraints matter, and how Anthropic's momentum is forcing a reset in how markets understand the AI boom. In the headlines: OpenAI's IPO plans, Cursor's efficient coding model, and more.Apply for our Growth Engineering role: https://jobs.aidailybrief.ai/Enterprise Claw Cohort 3 Registration: https://enterpriseclaw.ai/Brought to you by:KPMG – Agentic AI is powering a potential $3 trillion productivity shift, and KPMG's new paper, Agentic AI Untangled, gives leaders a clear framework to decide whether to build, buy, or borrow—download it at www.kpmg.us/NavigateGranola - The AI notepad for people in back-to-back meetings. 100% off your first 3 months with code AIDAILY at http://granola.ai/aidailyScrunch - The AI customer experience platform - https://scrunch.com/Mercury - Modern banking for business and now personal accounts. Learn more at https://mercury.com/personal-bankingZenflow Work - Agents for knowledge work - https://zenflow.free/Drata - The agentic trust management platform - https://drata.com/Blitzy - Want to accelerate enterprise software development velocity by 5x? https://blitzy.com/AssemblyAI - The best way to build Voice AI apps - https://www.assemblyai.com/briefRobots & Pencils - Cloud-native AI solutions that power results https://robotsandpencils.com/The AI Daily Brief helps you understand the most important news and discussions in AI. Subscribe to the podcast version of The AI Daily Brief wherever you listen: https://pod.link/1680633614Our Newsletter is BACK: https://aidailybrief.beehiiv.com/Interested in sponsoring the show? sponsors@aidailybrief.ai
The Twenty Minute VC: Venture Capital | Startup Funding | The Pitch
AGENDA: 00:00 – Anthropic Eyes $900B Valuation & Andre Karpathy's Shock Move 04:46 – Unpacking Anthropic's $30 Billion War Chest 10:52 – The True Cost of AI Tokens: Is Salesforce Spending Too Much? 15:59 – The Bear Case for Token Growth & Why Software Leaders Must Adapt 22:56 – Public Tech Rebound: Figma & Datadog Crush Expectations 26:59 – The Death of Traditional Web Builders? The Decline of Wix & Squarespace 36:59 – Compute Starvation: Is the Semiconductor & Hardware Boom Sustainable? 45:59 – Cerebras IPO Smashes Day One: The Biggest Tech Public Debut Since Snowflake 48:17 – SpaceX Sets Date For the Largest IPO in History 57:28 – Y Combinator's Mic Drop Deal & The Drama Behind Elon Musk's OpenAI Lawsuit 01:11:57 – The Looming Backlash: Mass Tech Layoffs and the Politics of AI 20VC:
Dashlane's CTO pulls back the curtain on how password managers are actually using AI, why it's more complicated than hype suggests, and what the rise of AI-powered code review means for the next wave of digital security. Nvidia Rides Blistering Chip Sales to Another Record Quarter Mind-Blowing Growth Is About to Propel Anthropic Into Its First Profitable Quarter SpaceX Filing Starts Countdown to Massive IPO Gemini 3.5 Flash: more expensive, but Google plan to use it for everything Google's Gemini Spark is an agentic AI assistant - Engadget Anthropic's Co-Founder to Launch Encyclical on AI With Pope Leo (21) Andrej Karpathy on X: "Personal update: I've joined Anthropic. I think the next few years at the frontier of LLMs will be especially formative. I am very excited to join the team here and get back to R&D. I remain deeply passionate about education and plan to resume my work on it in time." / X Most U.S. doctors are quietly using this AI tool. Few patients know about it. Greg Brockman Officially Takes Control of OpenAI's Products in Latest Shakeup Amazon's Alexa+ Now Produces AI-Generated 'Podcasts' Featuring Chats Between Two Robot 'Co-Hosts' AI chatbots are giving out people's real phone numbers Geoffrey Fowler and the Launch of the Youth AI Safety Institute We let four AIs run radio stations. Here's what happened. | Andon Labs The last six months in LLMs in five minutes Lake Tahoe Power Crisis: How AI Data Centers Are Cutting Power to 50,000 Residents What happens when you post a real Monet and say it's AI? The coolest art social experiment I've seen in a while. Thank you @SHL0MS Book on Truth in the Age of A.I. Contains Quotes Made Up by A.I. OpenClaw's Peter Steinberger's tokenmaxxing 'Obvious markers of AI': doubts raised over winner of short story prize Man drives Cybertruck into Grapevine Lake Stewart Brand's Maintenance of Everything Sports Illustrated Just Deleted Every Article by One of Its Writers After Accusation of AI Plagiarism The great digital media valuation collapse Sperm racing Hosts: Leo Laporte, Jeff Jarvis, and Paris Martineau Guest: Frederic Rivain Download or subscribe to Intelligent Machines at https://twit.tv/shows/intelligent-machines. Join Club TWiT for Ad-Free Podcasts! Support what you love and get ad-free audio and video feeds, a members-only Discord, and exclusive content. Join today: https://twit.tv/clubtwit Sponsors: outsystems.com/twit monarch.com with code IM zscaler.com/security XBOW.com
Dashlane's CTO pulls back the curtain on how password managers are actually using AI, why it's more complicated than hype suggests, and what the rise of AI-powered code review means for the next wave of digital security. Nvidia Rides Blistering Chip Sales to Another Record Quarter Mind-Blowing Growth Is About to Propel Anthropic Into Its First Profitable Quarter SpaceX Filing Starts Countdown to Massive IPO Gemini 3.5 Flash: more expensive, but Google plan to use it for everything Google's Gemini Spark is an agentic AI assistant - Engadget Anthropic's Co-Founder to Launch Encyclical on AI With Pope Leo (21) Andrej Karpathy on X: "Personal update: I've joined Anthropic. I think the next few years at the frontier of LLMs will be especially formative. I am very excited to join the team here and get back to R&D. I remain deeply passionate about education and plan to resume my work on it in time." / X Most U.S. doctors are quietly using this AI tool. Few patients know about it. Greg Brockman Officially Takes Control of OpenAI's Products in Latest Shakeup Amazon's Alexa+ Now Produces AI-Generated 'Podcasts' Featuring Chats Between Two Robot 'Co-Hosts' AI chatbots are giving out people's real phone numbers Geoffrey Fowler and the Launch of the Youth AI Safety Institute We let four AIs run radio stations. Here's what happened. | Andon Labs The last six months in LLMs in five minutes Lake Tahoe Power Crisis: How AI Data Centers Are Cutting Power to 50,000 Residents What happens when you post a real Monet and say it's AI? The coolest art social experiment I've seen in a while. Thank you @SHL0MS Book on Truth in the Age of A.I. Contains Quotes Made Up by A.I. OpenClaw's Peter Steinberger's tokenmaxxing 'Obvious markers of AI': doubts raised over winner of short story prize Man drives Cybertruck into Grapevine Lake Stewart Brand's Maintenance of Everything Sports Illustrated Just Deleted Every Article by One of Its Writers After Accusation of AI Plagiarism The great digital media valuation collapse Sperm racing Hosts: Leo Laporte, Jeff Jarvis, and Paris Martineau Guest: Frederic Rivain Download or subscribe to Intelligent Machines at https://twit.tv/shows/intelligent-machines. Join Club TWiT for Ad-Free Podcasts! Support what you love and get ad-free audio and video feeds, a members-only Discord, and exclusive content. Join today: https://twit.tv/clubtwit Sponsors: outsystems.com/twit monarch.com with code IM zscaler.com/security XBOW.com
Dashlane's CTO pulls back the curtain on how password managers are actually using AI, why it's more complicated than hype suggests, and what the rise of AI-powered code review means for the next wave of digital security. Nvidia Rides Blistering Chip Sales to Another Record Quarter Mind-Blowing Growth Is About to Propel Anthropic Into Its First Profitable Quarter SpaceX Filing Starts Countdown to Massive IPO Gemini 3.5 Flash: more expensive, but Google plan to use it for everything Google's Gemini Spark is an agentic AI assistant - Engadget Anthropic's Co-Founder to Launch Encyclical on AI With Pope Leo (21) Andrej Karpathy on X: "Personal update: I've joined Anthropic. I think the next few years at the frontier of LLMs will be especially formative. I am very excited to join the team here and get back to R&D. I remain deeply passionate about education and plan to resume my work on it in time." / X Most U.S. doctors are quietly using this AI tool. Few patients know about it. Greg Brockman Officially Takes Control of OpenAI's Products in Latest Shakeup Amazon's Alexa+ Now Produces AI-Generated 'Podcasts' Featuring Chats Between Two Robot 'Co-Hosts' AI chatbots are giving out people's real phone numbers Geoffrey Fowler and the Launch of the Youth AI Safety Institute We let four AIs run radio stations. Here's what happened. | Andon Labs The last six months in LLMs in five minutes Lake Tahoe Power Crisis: How AI Data Centers Are Cutting Power to 50,000 Residents What happens when you post a real Monet and say it's AI? The coolest art social experiment I've seen in a while. Thank you @SHL0MS Book on Truth in the Age of A.I. Contains Quotes Made Up by A.I. OpenClaw's Peter Steinberger's tokenmaxxing 'Obvious markers of AI': doubts raised over winner of short story prize Man drives Cybertruck into Grapevine Lake Stewart Brand's Maintenance of Everything Sports Illustrated Just Deleted Every Article by One of Its Writers After Accusation of AI Plagiarism The great digital media valuation collapse Sperm racing Hosts: Leo Laporte, Jeff Jarvis, and Paris Martineau Guest: Frederic Rivain Download or subscribe to Intelligent Machines at https://twit.tv/shows/intelligent-machines. Join Club TWiT for Ad-Free Podcasts! Support what you love and get ad-free audio and video feeds, a members-only Discord, and exclusive content. Join today: https://twit.tv/clubtwit Sponsors: outsystems.com/twit monarch.com with code IM zscaler.com/security XBOW.com
Dashlane's CTO pulls back the curtain on how password managers are actually using AI, why it's more complicated than hype suggests, and what the rise of AI-powered code review means for the next wave of digital security. Nvidia Rides Blistering Chip Sales to Another Record Quarter Mind-Blowing Growth Is About to Propel Anthropic Into Its First Profitable Quarter SpaceX Filing Starts Countdown to Massive IPO Gemini 3.5 Flash: more expensive, but Google plan to use it for everything Google's Gemini Spark is an agentic AI assistant - Engadget Anthropic's Co-Founder to Launch Encyclical on AI With Pope Leo (21) Andrej Karpathy on X: "Personal update: I've joined Anthropic. I think the next few years at the frontier of LLMs will be especially formative. I am very excited to join the team here and get back to R&D. I remain deeply passionate about education and plan to resume my work on it in time." / X Most U.S. doctors are quietly using this AI tool. Few patients know about it. Greg Brockman Officially Takes Control of OpenAI's Products in Latest Shakeup Amazon's Alexa+ Now Produces AI-Generated 'Podcasts' Featuring Chats Between Two Robot 'Co-Hosts' AI chatbots are giving out people's real phone numbers Geoffrey Fowler and the Launch of the Youth AI Safety Institute We let four AIs run radio stations. Here's what happened. | Andon Labs The last six months in LLMs in five minutes Lake Tahoe Power Crisis: How AI Data Centers Are Cutting Power to 50,000 Residents What happens when you post a real Monet and say it's AI? The coolest art social experiment I've seen in a while. Thank you @SHL0MS Book on Truth in the Age of A.I. Contains Quotes Made Up by A.I. OpenClaw's Peter Steinberger's tokenmaxxing 'Obvious markers of AI': doubts raised over winner of short story prize Man drives Cybertruck into Grapevine Lake Stewart Brand's Maintenance of Everything Sports Illustrated Just Deleted Every Article by One of Its Writers After Accusation of AI Plagiarism The great digital media valuation collapse Sperm racing Hosts: Leo Laporte, Jeff Jarvis, and Paris Martineau Guest: Frederic Rivain Download or subscribe to Intelligent Machines at https://twit.tv/shows/intelligent-machines. Join Club TWiT for Ad-Free Podcasts! Support what you love and get ad-free audio and video feeds, a members-only Discord, and exclusive content. Join today: https://twit.tv/clubtwit Sponsors: outsystems.com/twit monarch.com with code IM zscaler.com/security XBOW.com
Jason Howell and Jeff Jarvis break down everything from Google I/O 2026, where the company made its strongest case yet for winning the AI race. Gemini 3.5 Flash and Gemini Spark were unveiled, AI agents are now doing the searching instead of returning links, and Google's reach extended into design, science, YouTube, and shopping. Jason also demos Genie World Models live.Also in this episode: Andrej Karpathy joins Anthropic, Anthropic acquires a major dev tools startup, Amazon Alexa+ can now generate podcast episodes, Elon Musk's latest lawsuit drama, and a growing American rebellion against AI. Speed round includes the OpenAI IPO, xAI's coding agent, Meta's AR glasses, and more.New episodes every Wednesday at aiinside.show Note: Time codes subject to change depending on dynamic ad insertion by the distributor. CHAPTERS: 0:04:31 - Everything announced at Google I/O 2026 - Times: How Google Is Starting to Win the A.I. Race 0:22:42 - A new era for AI Search - Gemini 3.5: frontier intelligence with action 0:27:53 - Google Launches Gemini Spark: A 24/7 AI Agent That Wants to Make You Ditch OpenClaw 0:44:35 - OpenAI co-founder Andrej Karpathy joins Anthropic 0:46:32 - Anthropic has acquired the dev tools startup used by OpenAI, Google, and Cloudflare 0:55:20 - Amazon's new Alexa+ powered feature can generate podcast episodes 0:57:27 - The Art of War, Elon Musk Edition: How to Lose a Lawsuit and Still Claim Victory 0:59:30 - The American Rebellion Against AI Is Gaining Steam 1:01:46 - NextEra Energy to buy Dominion in deal that unites two key players in race to power AI data centers 1:04:42 - Pope Leo XIV will publish his first encyclical, Magnifica Humanitas, on May 25, with Anthropic co-founder Christopher Olah joining the launch panel at the Vatican 1:06:25 - Linus Torvalds says AI-powered bug hunters have made Linux security mailing list ‘almost entirely unmanageable' 1:08:57 - Meta brings virtual writing to everyone with Meta Ray-Ban Display glasses 1:10:36 - Musk's xAI Unveils First Coding Agent in Bid to Rival Anthropic 1:10:59 - OpenAI is Preparing to File for an IPO Very Soon Hosts: Jason Howell and Jeff Jarvis Download and subscribe to AI Inside in audio and video: https://aiinside.show/ Support the podcast on Patreon for special perks: https://www.patreon.com/aiinsideshow. You'll get ad-free episodes, members-only Discord, T-shirts and stickers you love, and get ad-free audio and video feeds, a members-only Discord, and exclusive content. Learn more about your ad choices. Visit megaphone.fm/adchoices
Dashlane's CTO pulls back the curtain on how password managers are actually using AI, why it's more complicated than hype suggests, and what the rise of AI-powered code review means for the next wave of digital security. Nvidia Rides Blistering Chip Sales to Another Record Quarter Mind-Blowing Growth Is About to Propel Anthropic Into Its First Profitable Quarter SpaceX Filing Starts Countdown to Massive IPO Gemini 3.5 Flash: more expensive, but Google plan to use it for everything Google's Gemini Spark is an agentic AI assistant - Engadget Anthropic's Co-Founder to Launch Encyclical on AI With Pope Leo (21) Andrej Karpathy on X: "Personal update: I've joined Anthropic. I think the next few years at the frontier of LLMs will be especially formative. I am very excited to join the team here and get back to R&D. I remain deeply passionate about education and plan to resume my work on it in time." / X Most U.S. doctors are quietly using this AI tool. Few patients know about it. Greg Brockman Officially Takes Control of OpenAI's Products in Latest Shakeup Amazon's Alexa+ Now Produces AI-Generated 'Podcasts' Featuring Chats Between Two Robot 'Co-Hosts' AI chatbots are giving out people's real phone numbers Geoffrey Fowler and the Launch of the Youth AI Safety Institute We let four AIs run radio stations. Here's what happened. | Andon Labs The last six months in LLMs in five minutes Lake Tahoe Power Crisis: How AI Data Centers Are Cutting Power to 50,000 Residents What happens when you post a real Monet and say it's AI? The coolest art social experiment I've seen in a while. Thank you @SHL0MS Book on Truth in the Age of A.I. Contains Quotes Made Up by A.I. OpenClaw's Peter Steinberger's tokenmaxxing 'Obvious markers of AI': doubts raised over winner of short story prize Man drives Cybertruck into Grapevine Lake Stewart Brand's Maintenance of Everything Sports Illustrated Just Deleted Every Article by One of Its Writers After Accusation of AI Plagiarism The great digital media valuation collapse Sperm racing Hosts: Leo Laporte, Jeff Jarvis, and Paris Martineau Guest: Frederic Rivain Download or subscribe to Intelligent Machines at https://twit.tv/shows/intelligent-machines. Join Club TWiT for Ad-Free Podcasts! Support what you love and get ad-free audio and video feeds, a members-only Discord, and exclusive content. Join today: https://twit.tv/clubtwit Sponsors: outsystems.com/twit monarch.com with code IM zscaler.com/security XBOW.com
May 20, 2026: Meta began notifying 8,000 employees of their layoffs this morning — while simultaneously redirecting $145 billion into AI infrastructure. Andrej Karpathy, one of the founding members of OpenAI and the architect of Tesla's self-driving brain, just joined Anthropic with a specific mission: use AI to make AI better. And a major new Milken-Harris Poll finds that 80% of Americans want government workforce transition programs now, 68% say they're navigating the AI shift entirely alone, and 88% of business leaders privately admit companies cannot solve this without a coordinated national response.
Google I/O 2026 just dropped Gemini Omni, a world-model AI that simulates physics, edits video, and might be the biggest leap since Seedance 2. But it's not perfect. Gavin and Kevin break down everything from Google I/O 2026, including the launch of Gemini Omni (Google's new world model), Gemini 3.5 Flash benchmarks against GPT-5.5 and Opus 4.7, the Gemini Spark personal agent, AskYouTube, Docs Live, new AI glasses, the first search box redesign in 25 years, and the shocking news that Andrej Karpathy is joining Anthropic. SHOW LINKS: Google I/O 2026 Full Keynote: https://www.youtube.com/live/wYSncx9zLIU?si=Nb881MfGTlf1Q0II Gemini Omni physics demos from Google DeepMind: https://x.com/GoogleDeepMind/status/2056786449312493669?s=20 Gemini Omni's incredible London knowledge (via fofrAI): https://x.com/fofrAI/status/2056789242274259242?s=20 Sundar Pichai and Demis Hassabis on Omni video editing: https://x.com/sundarpichai/status/2056524502746747048?s=20 Gavin's hands-on Gemini Omni experiments: https://x.com/gavinpurcell/status/2056762427879182692?s=20 Gemini Omni's character cameo feature (less impressive): https://x.com/gavinpurcell/status/2056772793539481830?s=20 Gemini Omni volleyball fail: https://x.com/flavioAd/status/2056771223359549645?s=20 Google's new Content Credentials Verification: https://x.com/Google/status/2056787498676658576?s=20 Genie 3 IRL — Google's world model now simulates real streets with Street View: https://techcrunch.com/2026/05/19/googles-genie-world-model-can-now-simulate-real-streets-with-street-view/ Bilawal Sidhu on Genie 3 IRL: https://x.com/bilawalsidhu/status/2056804315721843024?s=20 Gemini 3.5 Flash launches — official announcement: https://x.com/GeminiApp/status/2056788115893993701?s=20 Gemini Spark — Google's new personal coding agent: https://x.com/Google/status/2056791134295273554?s=20 Google's new AI glasses https://x.com/backlon/status/2056807059707036050?s=20 Andrej Karpathy joins Anthropic to focus on recursive self-learning: https://www.axios.com/2026/05/19/anthropic-openai-karpathy-andrej-claude
An OpenAI co-founder has joined the competition. Andrej Karpathy, who helped start the high-profile artificial intelligence outfit in 2015, has now joined Anthropic, he announced Tuesday. Karpathy was instrumental in the development of Tesla's Autopilot system, working for the automaker for about five years in between two separate stints at OpenAI. In his new role at rival Anthropic, Karpathy will "get back to R&D," he wrote in a social media post, adding, "I think the next few years at the frontier of LLMs will be especially formative."
Check out the Public app for incredible investing tools and to support the show (LINK)Follow us on Instagram (@TheRundownDaily) for bonus content and instant reactions.In today's episode:Google unveils Gemini 3.5 Flash, AI agent and a major overhaul to Search at its annual I/O developer conferenceTarget crushes earnings with its first positive same-store sales in five quarters and raises its full-year outlookCava defies the restaurant slowdown with 9.7% same-store sales growth and a raised forecastLowe's beats on earnings but stock slips as comparable sales disappoint in a tough housing marketOpenAI co-founder Andrej Karpathy joins Anthropic
The Musk v. Altman jury unanimously rejected Musk's claims on statute of limitations grounds. Andrej Karpathy joined Anthropic's pre-training team. Polymarket partners with Nasdaq on private company markets, Blackstone and Google form a TPU venture, and KPMG embeds Claude into tax advisory. Musk v. Altman: the jury unanimously rejects Elon Musk's claims against OpenAI and Sam Altman, as he filed them outside of a three-year statute of limitations (CNBC) Andrej Karpathy joins Anthropic to help launch a team focused on using Claude to accelerate pre-training research; he helped found OpenAI and worked at Tesla (Axios) Polymarket partners with Nasdaq to launch markets tied to private company milestones, including IPO timing, valuations, earnings, and secondary market activity (The Block) Blackstone announces a joint venture with Google to create a US company that will offer customers Google TPU access, and makes a $5B initial equity commitment (WSJ) KPMG partners with Anthropic to embed Claude into its tax and advisory platforms; KPMG's tax and legal services unit saw revenue grow ~8% YoY to $9.3B in 2025 (WSJ) Learn more about your ad choices. Visit megaphone.fm/adchoices
Free Guide to Build Your Personal AI Wiki: https://clickhubspot.com/fcmt Ep. 425 You can create a second brain in 15 minutes and for free. Kipp and Matt Wolfe (Future Tools) dive into how to build your own "second brain" using tools like Obsidian and Codex, so you can finally make use of all the information you're gathering. Learn more on setting up a personal wiki to organize your knowledge, harnessing AI-powered agents to automate and interlink content, and turning your information hoarding into proactive recommendations and smarter business decisions. Mentions Matt Wolfe https://www.linkedin.com/in/matt-wolfe-30841712/ Future Tools https://futuretools.io/ Obsidian https://obsidian.md/ Obsidian Web Clipper https://obsidian.md/clipper Codex https://openai.com/codex/ Claude Code https://code.claude.com/docs/en/overview Cursor https://cursor.com/ Andrej Karpathy https://karpathy.ai/ Visual Studio Code https://code.visualstudio.com/ Notion https://www.notion.com/ Get our guide to build your own Custom GPT: https://clickhubspot.com/customgpt Resource [Free] Steal our favorite AI Prompts featured on the show! Grab them here: https://clickhubspot.com/aip We're on Social Media! Follow us for everyday marketing wisdom straight to your feed YouTube: https://www.youtube.com/channel/UCGtXqPiNV8YC0GMUzY-EUFg Twitter: https://twitter.com/matgpod TikTok: https://www.tiktok.com/@matgpod Thank you for tuning into Marketing Against The Grain! Don't forget to hit subscribe and follow us on Apple Podcasts (so you never miss an episode)! https://podcasts.apple.com/us/podcast/marketing-against-the-grain/id1616700934 If you love this show, please leave us a 5-Star Review https://link.chtbl.com/h9_sjBKH and share your favorite episodes with friends. We really appreciate your support. Host Links: Kipp Bodnar, https://twitter.com/kippbodnar Kieran Flanagan, https://twitter.com/searchbrat ‘Marketing Against The Grain' is a HubSpot Original Podcast // Brought to you by Hubspot Media // Produced by Darren Clarke.
What if the most important skill for building AI products has nothing to do with evals, technical background, or knowing how to write a prompt? What if it is the ability to design systems that can handle what you never planned for?In this episode of Supra Insider, Marc Baselga and Ben Erez sit down with Apurva Garware, who has built and scaled products across Amazon, Microsoft, and Upwork, to make the case that systems thinking is the defining skill of the next era of product management. Apurva explains why non-determinism forces PMs to stop thinking in features and start designing the guardrails, agent contracts, and escalation points that govern how a system behaves at runtime, when no one is watching. They explore a three-phase framework for governing AI systems across design, deployment, and production; heuristics for deciding what to hand to agents versus escalate to humans; and a sharp insight about the two products every AI-native company is actually building: the customer-facing product, and the internal operational system that drives margin and velocity. Marc and Ben also share their own experience calibrating an agentic workflow at Supra, grounding the conversation in practice.If you are a PM trying to find your footing in the AI era without a deeply technical background, a founder wrestling with when to reach for AI versus simpler deterministic automation, or a product leader who wants to build more discipline into how your team ships AI products, this episode is for you.All episodes of the podcast are also available on Spotify, Apple and YouTube.New to the pod? Subscribe below to get the next episode in your inbox
“Anyone that's properly using AI now knows that you tell it what you want, it gives you a plan, carries out the work, and you judge and tweak. You're not a passive victim — you're an active user with outcomes in mind.” — Keith Teare Do we really want a no-hands job from Silicon Valley? That Was the Week newsletter publisher Keith Teare — who thinks all tech innovation results in human progress — thinks we do. No hands, no problem, Keith says. But I'm not sure. Especially given the powers-that-be giving us that no-hands job. Keith welcomes the end of what he calls the “typed” and “touched” computing era — keyboards, mice, touchscreens, and all the manifold ways we have used our hands to interact with computers since the 1980s. That's the outcome, he predicts, of the race to AGI. So far so good. But what happens if our no-hands AI future is controlled by Google, Microsoft, Amazon, and Facebook? This week these four behemoths committed 00 billion to AI infrastructure investment in 2026 alone — 2 percent of all US GDP. These companies are racing to build (and own) the foundational mechanics of AGI. That's always how it's been, Keith says, embracing our no-hands future. I'm less open-armed. What happens if we want our hands to fend off AGI? No, I'm not so keen on a no-hands job from Silicon Valley. Especially one couched in the altruism of human progress. Five Takeaways • The End of the Hand-Driven Computing Era: Andrej Karpathy's observation at Sequoia's AI Ascent: he no longer uses his hands to do his work. He speaks to the computer; the computer acts; he judges and refines. The keyboard, the mouse, the touchscreen — all the hand-driven interfaces that have defined computing since the 1980s are entering their twilight. Karpathy calls it “software 3.0”. Keith, two years ago, wrote an editorial called “eyes, hands, ears, and mouth” about the inclusion of other human attributes beyond hands. That prediction has arrived. • $700 Billion: The CapEx Explosion: A post by @Signal framed the week's numbers: $700 billion in AI infrastructure spending in 2026, equivalent to 2 percent of all US GDP. This kind of spending, the post observes, usually happens via governments or wars. This time, it's four private companies — Microsoft, Amazon, Google, and Meta — racing to build the foundational mechanics of AGI. Meta was punished by Wall Street for overspending; Google was rewarded because its numbers were strong enough to justify it. The same bet, two different verdicts, depending on your quarterly earnings. • Was the Internet Privately Built? The ARPANET Argument: Keith's claim: innovation waves have always been privately financed. The railways, the telephone, the electricity grid, the commercial internet. Andrew's counter: ARPANET was a massive government investment that created the protocols on which the internet runs. Keith's response: ARPANET was a university bulletin board that created the precedent, not the infrastructure. Andrew's response: that's not exactly what ARPANET was. They agree that government research matters. They disagree on how much credit it deserves for what became the commercial internet. • The Revenge of the Idea Guy: Sam Altman's line of the week. In the past, an idea person came up with a concept and then needed expensive engineers to build it. Many ideas never saw the light of day because the engineering cost was prohibitive. Now, anyone can speak an idea into existence. AI builds the plan, executes the work, and you judge and refine. That changes the economics of creativity, advertising, software development, and anything else that used to require specialist execution. The specialist is not dead — but specialists will increasingly use AI to scale themselves, rather than being hired one at a time. • Should Kids Use AI in Schools? A New Yorker piece asks what it would take to get AI out of schools. Keith's view: the premise misunderstands how AI works now. The fear is passive students asking chatbots for answers and having their brains atrophy. The reality is that proper AI use requires active judgment at every step — telling it what you want, refining the plan, evaluating the output. If schools understand that, they embrace AI. If they don't, they produce graduates unequipped for a world in which the idea guy with AI tools now has the power the engineering team used to have. Andrew's prediction: the kids whose parents ban AI will eventually sue them. About the Guest Keith Teare is a British-American entrepreneur, investor, and publisher of the That Was the Week newsletter — a daily curation of the most important stories at the intersection of technology, business, and culture. He is a co-founder of TechCrunch and a long-time interlocutor on Keen On America. References: • That Was the Week newsletter by Keith Teare — this week's editorial: “Hand Job?” • Andrej Karpathy at Sequoia Capital AI Ascent 2026 — the Karpathy interview on Software 3.0 and the end of typed input. • @Signal, “$700 billion on AI infrastructure” — the post that framed the CapEx question. • Jessica Winter, “What Will It Take to Get AI Out of Schools?” The New Yorker, 2026. • Episode 2891: John Steele Gordon on how information technology knitted America together — the ARPANET backstory that feeds directly into this week's argument. About Keen On America Nobody asks more awkward questions than the Anglo-American writer and filmmaker Andrew Keen. In Keen On America, Andrew brings his pointed Transatlantic wit to making sense of the United States — hosting daily interviews about the history and future of this now venerable Republic. With nearly 2,900 episodes since the show launched on TechCrunch in 2010, Keen On America is the most prolific intellectual interview show in the history of podcasting. WebsiteSubstackYouTubeApple PodcastsSpotify Chapters: (00:31) - Keith leads with “Hand Job?” — explaining the headline (03:27) - Karpathy at Sequoia: the end of typed and touched input (04:30) - CapEx: the real story of the week (05:35) - $700 billion — 2% of US GDP on AI infrastructure (06:38) - Was the commercial internet privately built? (07:35) - ARPANET: pathetic bulletin board or foundational infrastructure? (09:08) - Keith and Andrew agree to disagree on government's role (11:00) - Big Tech earnings: Google up, Meta down, and why (17:00) - OpenAI's strategy: the long game
Andrej Karpathy (co-founder of OpenAI, former head of AI at Tesla, and now founder of Eureka Labs) talks with Sequoia partner Stephanie Zhan at AI Ascent 2026 about what's changed in the year since he coined "vibe coding." He explains why he's never felt more behind as a programmer, why agentic engineering is the more serious discipline taking shape on top of vibe coding, and why we should think of LLMs not as animals but as ghosts: jagged, statistical, summoned entities that require a new kind of taste and judgment to direct. He also touches on Software 3.0, the limits of verifiability, and why you can outsource your thinking but never your understanding.
Early bird discounts for the San Francisco World's Fair, the biggest AIE gathering of the year, end today - prices will go up by ~$500 tonight so do please lock in ASAP!From near-universal AI tool adoption inside Shopify to internal systems for ML experimentation, auto-research, customer simulation, and ultra-low-latency search, Mikhail Parakhin joins us for a deep dive into what it actually looks like when a 20-year-old, $200B software company goes all-in on AI. We cover why Shopify has become much more vocal about its internal stack, what changed after the December model-quality inflection, and why the real bottleneck in AI coding is no longer generation, but review, CI/CD, and deployment stability.We also go inside Tangle, Tangent, SimGym, which are three major AI initiatives that Shopify is doing to make experimentation reproducible, optimization automatic, customer behavior simulatable, and search and catalog intelligence faster and cheaper at scale. Along the way, Mikhail explains UCP, Liquid AI, and why token budgets are directionally right but often measured badly, why AI-written code can still increase bugs in production, what makes Shopify's customer simulation defensible, and what he learned from the Sydney era at Bing.We discuss:* Mikhail's path from running a major Microsoft business unit spanning Windows, Edge, Bing, and ads to becoming CTO of Shopify* Why Shopify is talking more publicly about AI now, and why staying at the frontier has become necessary for the company* Shopify's internal AI adoption curve, the December inflection, and why CLI-style tools are rising faster than traditional IDE-based tools* Why Jensen Huang is directionally right on token budgets, but raw token count is still the wrong way to evaluate engineering output* Why the real unlock is not more agents in parallel, but better critique loops, stronger models, and spending more on review than generation* Why AI coding can still lead to more bugs in production even if models write cleaner code on average than humans* Why Shopify built its own PR review flow, and why Mikhail thinks most off-the-shelf review tools miss the point* How PR volume, test failures, and deployment rollback are becoming the real bottlenecks in the agent era* Why Git, pull requests, and CI/CD may need a new metaphor once code is written at machine speed* What Tangle is, and how Shopify uses it to make ML and data workflows reproducible, collaborative, and production-ready from the start* Why Tangle is different from Airflow, and why content-addressed caching creates network effects across teams* What Tangent is, and how Shopify is using auto-research loops to optimize search, themes, prompt compression, storage, and more* Why Tangent is becoming a democratizing tool for PMs and domain experts, not just ML engineers* Why AutoML finally feels real in the LLM era, and where auto-research still falls short today* Why Tangle, Tangent, and SimGym become much more powerful when combined into one system* What SimGym is, why simulated customers only work if you have real historical behavior, and why Shopify's data gives it a moat* How SimGym evolved from comparing A/B variants to telling merchants what to change on a single live storefront to raise conversions* Why customer simulation is so expensive, from multimodal models to browser farms to serving and distillation costs* How Shopify models merchant and buyer trajectories, runs counterfactuals, and thinks about interventions like discounts, campaigns, and notifications* Why category-level behavior is so different across commerce, and why ideas like Chinese Restaurant Processes are showing up again in practice* Shopify's new UCP and catalog work, including runtime product search, bulk lookups, and identity linking* Why Shopify is using Liquid AI, and why Mikhail sees it as the first genuinely competitive non-transformer architecture he has used in practice* Where Liquid already works inside Shopify today, from low-latency query understanding to large-scale catalog and Sidekick Pulse workloads* Whether Liquid could become frontier-scale with enough compute, and why Shopify remains pragmatic and merit-based about model choice* Who Shopify is hiring right now across ML, data science, and distributed databases* The Sydney story at Bing, why its personality was not an accident, and what Mikhail learned from deliberately shaping AI character early onMikhail Parakhin* LinkedIn: https://www.linkedin.com/in/mikhail-parakhin/* X: https://x.com/MParakhinTimestamps00:00:00 Introduction: Mikhail Parakhin, Microsoft, and Shopify00:01:16 Why Shopify Is Talking More About AI00:02:29 Internal AI Adoption at Shopify and the December Inflection00:06:54 Token Budgets, Jensen Huang, and Why Usage Metrics Can Mislead00:10:55 Why Shopify Built Its Own AI PR Review System00:12:38 AI Coding, More Bugs, and the Real Deployment Bottleneck00:14:11 Why Git, PRs, and CI/CD May Need to Change for Agents00:18:24 Tangle: Shopify's Reproducible ML and Data Workflow Engine00:21:19 Why Tangle Is Different from Airflow00:26:14 Tangent: Auto Research for Optimization and Experimentation00:30:07 How Tangent Democratizes Experimentation Beyond ML Engineers00:33:06 The Limits of Auto Research00:36:36 Why Tangle, Tangent, and SimGym Compound Together00:37:20 SimGym: Simulating Customers with Shopify's Historical Data00:42:47 The Infra Behind SimGym00:46:00 Why SimGym Gets Better with Real Customer History00:47:30 Counterfactuals, HSTU, and Modeling Merchant Trajectories00:51:55 CRPs, Clustering, and Category-Level Customer Behavior00:53:30 UCP, Shopify Catalog, and Identity Linking00:55:07 Liquid AI: Why Shopify Uses Non-Transformer Models00:59:13 Real Shopify Use Cases for Liquid01:03:00 Can Liquid Scale into a Frontier Model?01:09:49 Hiring at Shopify: ML, Data Science, and Databases01:10:43 Sydney at Bing: Personality Shaping and AI Character01:13:32 Closing ThoughtsTranscript[00:00:00] swyx: Okay. We're here in the studio, a remote studio, with Mikhail Parakhin, CTO of Shopify. Welcome.[00:00:08] Mikhail Parakhin: Thank you. Welcome.[00:00:10] swyx: I don't even know if I should introduce you as CTO of Shopify. I feel like you have many identities. Uh, you led sort of the, the Bing ML team, I guess, uh, uh, or ads team. I, I don't know, I don't know, uh, you know, it's, uh, people va-variously refer you as like CEO or, or, uh, I don't know what that, that, that said previous role at Microsoft was.[00:00:29] Mikhail Parakhin: Uh, that was... Yeah, my previous role w- at Microsoft was the-- I actually was the CEO of one of Microsoft's business units, which included, as I, you know, as we discussed, all the things that people like to laugh about, uh, including Windows and Edge and Bing and ads and everything.[00:00:47] swyx: Yeah, yeah. What a, what a, what a wild time.You've obviously, uh, done a lot since you landed at Shopify. Uh, one of the reasons I reached out was because you started promoting more sort of internal tooling, uh, primarily Tangle, but also a lot of people have seen and adopted Tobi's QMD, uh, and obviously, I think, uh, Shopify has always been sort of leading in terms of, uh, engineering.I think more-- it's just more recent that you guys have been more vocal about your sort of AI adoption. Is that, is that true?[00:01:16] Mikhail Parakhin: Well, I think AI tools in general are fairly recent development, uh, and we've-- Shopify, you know, at this stage of its development, we're developing AI in-in-house and other, uh, building tools that use AI and, you know, interfacing with the wider AI community, uh, you know, are on the sort of the, uh, runaway trajectory.So it just did by sort of natural byproduct. We, we talk about it more also. We just, uh, just even yesterday, Andrej Karpathy was famous in tweeting about, oh, are there some, uh, ways, uh, that, that you can organize your agents to store the data and then, uh, look up the data so that you don't have to research or, or lose context every- Yestime. And a little bit tongue in cheek, I tweeted that, “Hey, we've, we've done it much earlier, and we even have different approaches, Tobi and I.” Tobi, of course, is a big fan of QMD, and I'm more of a SQL, SQLite fan. But, uh, yeah, very similar things that we've already done here. The point is, yeah, we're very dynamic, you know, explosively growing company, and we have to be at the forefront of AI adoption, obviously.[00:02:29] swyx: Yeah. Yeah. Um, you, your team kindly prepared some slides actually that we were gonna bring up on to, uh, the screen. I think I can, I can screen share, and then we can kind of go through some of the shocking stats that maybe, maybe put some numbers to what exactly is going on. So here we have, uh- An internal AI tool adoption chart.What are we looking at here? What ?[00:02:54] Mikhail Parakhin: Yeah, this is very interesting statistics. Uh, this is number of daily active workers, you know, think of, uh, DAO, basically the active users of-[00:03:05] swyx: Yeah ...[00:03:05] Mikhail Parakhin: AI tool as a percentage of all the people in the company, right? And then- Yeah ... different AI tools. And, uh, you could see two things here is that one is the green is total.Uh, green is just total. So you could see that it approaches really % by now. It's hard not to do your job now without interacting deeply, at least with one tool. You could see another interesting thing is just as many people commented in December was the phase transition when suddenly models gotten good enough that, that everything took off and started growing.Uh, it, it was many people noticed that the thing is that small improvements accumulated into this big change in Sep- December roughly timeframe.[00:03:52] swyx: Yeah.[00:03:52] Mikhail Parakhin: The other thing I would claim you could see is that, uh, CLI-based tools and tools that don't require you to look at the code becoming more popular, and you could see, yeah, various versions of, uh, Cloud Code and Codex and Pi and internal development tools taking off.Uh, exactly, yeah, uh, and blue is our River, just internal agent for coding, where tools, uh, that require IDEs such as, uh, GitHub, Copilot or Cursor, they're not exactly shrinking, but they're not growing as fast. Like, uh, red, red line is, is the IDE kind of tools. So you could see that they're, they're not experiencing as, as fast of a growth.[00:04:37] swyx: As I understand it, basically, every employee has their choice, right? Of choose whatever tool you use, and then you're just kind of doing a, a daily sur-survey or something.[00:04:47] Mikhail Parakhin: Exactly. And, uh, we- Yeah ... the, the push is to get your job done, you can use any tool, and we effectively fund unlimited tokens for everybody.Uh, we, we do, we do try to control the models that, uh, people use, but from the bottom, not from top. Like we basically say, “Hey, please don't use anything less than Opus four point six.”[00:05:09] swyx: Oh .[00:05:10] Mikhail Parakhin: Some people, some people end up using GPT five point four extra high. Some people use Opus four point six. Um, uh, you know, uh, there are some, uh, there are plus and minuses in going for full one million context window versus not.But, uh, we try to discourage people from using anything less than that.[00:05:28] swyx: Yeah, yeah. Got it, got it. Uh, I mean, uh, that's, you know... The, the next chart here, it really kind of shows the expansion and the sort of December twenty twenty-five inflection, right? That, uh, people are using a lot of tokens. I think it's also really interesting that no one was kind of abusing it in twenty twenty-five.Like it was- Had comparatively, uh, to this year, there was almost no growth. I mean, it's still like, you know, probably, probably gave fifty percent.[00:05:56] Mikhail Parakhin: Yeah. This is just a different scale. It's still exponential- Yeah, yeah ...growth at just a different- ...rate of expansion. Uh, there was inflection point, and Sean, I would claim the, the super interesting part here is that you could see that the distribution becoming more and more skewed.Yes. The top percentiles grow faster. So that means- Yeah ...the people in the top ten percentile, they, their consumption grows faster than seventy-five and so forth. So, uh, the distribution skews more and more towards the highest users, which is... I don't know what it tells me. It's like it feels not ideal, to be honest.Or maybe it's okay. We'll see.[00:06:36] swyx: Why does it feel not ideal? Is, is it because of, um, quantity over quality, or what's the concern?[00:06:42] Mikhail Parakhin: Because take it to the limit. That means, you know, if, if this rate of separation continued- Ah, yes ...a year, there will be one person consuming all the tokens. So it's just, it's kinda strange.[00:06:54] swyx: Yeah, I mean, um, uh, I, I think internal like teaching and all that, uh, will, will help sort of distribute things more widely. But in, in the early days, of course, the people who are sort of more AI-pilled will obviously find more ways to use it than the people who are less AI-pilled. Maybe let's, let's call it that.I'll just, I'll just kinda quickly, uh, pause from the, the... You know, we will go back to the rest of the slides, but I just wanna, um, review, you know, there are a lot of CTOs of, of large companies like yourself where they're all considering some kind of token budget, right? Like I think it's something, something that Jensen Huang has been talking about, where like if your 200K engineer is not using 100K of tokens every year, like they're, they're underutilizing coding agents.Of course, Jensen Huang would say that, but like it seems a very quantity over quality approach and like some, some people are basically saying like, well, is this comparable to judging engineer quality by lines of code, right? Which we also know is like kind of flawed, but better than nothing. So I, I don't know if you have like a sort of management take here on, on how to view this kind of, uh, metrics.[00:08:02] Mikhail Parakhin: Well, I mean, you're, you're baiting me. I, I like... This is my favorite topic. Uh, if you let me, I'll probably talk for two hours on just this. I have a lot of things to say. Like I do think Jensen gotten a lot of bad press saying, “Oh, of course you're, you know, this, uh, the- ...the cake seller says you don't need enough cakes.”You know? Like, of course. Uh, but, uh, I actually, uh, think that's undeserved. I think he, he's actually right. Uh, I do think- He,[00:08:33] swyx: he's directionally correct.[00:08:35] Mikhail Parakhin: Yeah. Yeah. He's directionally correct for sure. Uh-[00:08:37] swyx: Who knows what the right number is? Yeah.[00:08:39] Mikhail Parakhin: The thing that I do Uh, want to say, and this is something that we learned through trial and error and very important is like two things.One is that it's not about just consuming tokens. Uh, you can consume tokens and, and in fact, the anti-pattern is running multiple agents, too many agents in parallel that don't communicate with each other. That's almost useless, uh, compared to just fewer agents and burns tokens very efficiently. Uh, setting up the right critique loop, especially with the high quality models, where one agent does something, the other one, ideally with a different model, critiques it, uh, suggests ways to improve it, the agent redoes it with this critique and, and so it takes much longer.So people don't like it because latency goes up. You know, they, they have to wait until this debate is happening. But, uh, the quality of the code is much higher. And another thing, just since you mentioned like, look, uh, uh, yeah, the overall budget is just like, uh, lines of codes. Lines of codes are exploding for everybody right now, or partially because AI is really mover balls, but partially just because AI can write a lot more code, you know, doesn't get tired.And so you have to have to have a very strong narrow waist during PR review. Otherwise, just the number of bugs will go through the roof. It's, uh, it's this unexpected consequence of the just volume trumping everything. I would claim by now good model writes code on average with fewer bugs than, than the average human.But since they write so much more of it, like more of it will make it into production. So you have to- You still[00:10:26] swyx: have[00:10:26] Mikhail Parakhin: more bugs. Yeah. Have to have a very rigorous PR reviews, also automated of course. But, uh, yeah, that to spend a lot budget there. Like this, this for me, for me, actually, the important metric is the ratio of budget spent during code generation versus, uh, spent, uh, expensive tokens like GPT, uh, five point four Pro or, uh, uh, Deep Think from Gemini, you know, checking on PR reviews.[00:10:55] swyx: Yeah, totally. Uh, I noticed in your chart you didn't have any review tools. Do you just use like, like let's say a Claude code to review tools? Or do you have another set of review tools like the Greptiles, the Code Rabbits, uh, Devin Reviews has a review tool. I don't know if you've had those specialist review tools.[00:11:13] Mikhail Parakhin: You are a little bit jumping on my store tool right now because the graphs I was only showing public tools. Uh, uh, the-- I haven't found a good PR review tool that, that does what I think should be done. And, uh, partially my, my thinking is because it's so... It just goes against both what people feel like emotionally they prefer and, uh, some of the, uh, you know, frankly Even business models that, that the companies run.At peer review tool, uh, time, you want to run the largest models. That means, I don't know, Codex or, or, uh, Cloud Code is not gonna cut it. You need to have pro-level models if you really want to, uh, stand the tide of bots from going into production. And you need us to spend a lot of time, the models taking turns, but you don't want, like, a big swarm of, uh, of, uh, agents.So in fact, you end up in a different dual-dualistic world where you generate not that many tokens. You, in fact, generate few tokens, but it takes f-a long time because these are expensive models taking turns rather than many, many agents trying to do many things in parallel. So that's, that's why I feel like I haven't found good tools, so we are using our own for peer review for now.[00:12:33] swyx: Yeah. Yeah. I mean, uh, I think a lot of companies are building their own, uh, especially to their needs, right?[00:12:38] Mikhail Parakhin: Mm-hmm.[00:12:38] swyx: Um, I, uh, you also have a chart here going back to the slides on, uh, PR merge growth, where we're now at thirty percent, uh, month on month rather than ten percent. Uh, and also the, the estimated complexity is going up.You know, this is productivity, right? ‘Cause y- presumably there's more stuff going into the code base and more, more features getting worked on. I'm curious about the backlog, right? Like the, the, the-- I actually don't mind a pro-level model taking an hour or two hours to review my PR, because I've dealt with humans who take a week to review my PR, right?And I keep pinging them on Slack, “Hey, hey, review my PR.” So, you know, I think there's some trade-off here where, like, it still doesn't make sense.[00:13:18] Mikhail Parakhin: Exactly. That, that's exactly m-my point. Uh, that on one hand, you can tolerate longer latencies at, uh, PR. On the other hand, like right now, the real problem is not in spending time waiting for PR.It's real problem is since there's so much more code than- Yeah ... uh, probability of at least some tests failing going up, and then you, like, keep de-failing, then you have to find the offending PR, evict it, retest it without that PR, and so deployment cycle becomes much longer. Uh, so it actually, in terms of the overall time to deploy, it's total time savings if you spend more time on a longer model, like thinking for an hour, because then, then you, you don't have to spend all that time during testing and rolling, you know, rolling back the deployment.[00:14:03] swyx: Yeah, totally. That's still worth it. You know, you don't look at the individual, look at the aggregate, and look at the, the, the change in the aggregate system.[00:14:11] Mikhail Parakhin: Exactly.[00:14:11] swyx: I'm kind of curious if, like, there's this PR mentality and, like, c-- the, the, the CICD paradigm will be changed eventually. Some people are like, obviously a lot of people want new GitHub, but I even wonder if, like, Git is the problem, right?Like, is that the bottleneck? Is the concept of a PR a bottleneck? Do you guys use stack diffs? I don't know if, uh, that's a, like, a merge queue stack diff type of thing.[00:14:34] Mikhail Parakhin: We, we use, we use Stacks, we u- we use Graphite. We worked with, uh, Graphite a lot. Uh, so we use Stack, uh, PRs. I think, uh, like that's clearly the overall CICD in general, and the interaction with the code repository right now is the, clearly the sort of the, the main issue and the bottleneck for us, uh, and highest top of mind.I would say we probably need a different metaphor or different whole design of how to process it in new agentic world. I haven't seen anything dramatically better yet. I, I think everybody right now is just trying to keep their head above the water ‘cause, ‘cause there, there's so many PRs and then everybody's CICD pipelines start creaking, the, the times are increasing, the number of bugs slipping by increasing, and you have to, have to clap on down.And so we are a little bit in this situation when we need to first stabilize that story and then start thinking, hey, what, what it could be a completely different and new world, which I haven't... I know some people working on it. I haven't seen something, like anything super compelling yet, but clearly the old thing were designed for humans will need to be morphed into something new.[00:15:53] swyx: One of the thing that I, I think about is kind of like the merge conflict is basically a global mutex on the whole system, right? And in, in hu- in human organizations, we do have something like that. It's the company standup. But like, other than that, it's like it's actually fitting for us to be somewhat decentralized, somewhat plugged into one stream of information source, but somewhat lossy.Like it's okay, you know, that, that not every delivery is like atomic consistency. Like we're not dealing with a database sometimes.[00:16:27] Mikhail Parakhin: This is a very good point, uh, because since humans don't write code too fast, you know that global mutex is not too bad. Once you-[00:16:36] swyx: Yes ...[00:16:37] Mikhail Parakhin: start writing code at the speed of machine, it becomes the, you know, the bottleneck.Then what do you do? Maybe, and I can't believe I'm saying this because I, I'm long-- lifelong opponent of, uh, microservices, and I always thought that was, like, a really bad idea. And now that you're saying it, like, maybe in new guys like microservices will make a comeback, you know, because then you, you can ship things independently in tiny things and, and the managing all that complexity automatically will be much easier.I don't know. Like, we'll s-- we'll have to see.[00:17:10] swyx: Yeah. I mean, I don't know what the Microsoft or, or Shopify thing is, but I, I read this paper from Google where they have a monorepo that deploys into microservices, right? And then, uh, the other concept that I think about a lot is the Chaos Monkey concept from, from Netflix.Being able to create, like, this robust system where, um, uh, you know, you, you have the service discovery, you have the, uh, the independent, independent microservices discovery and, and, uh, you know, probably going to be a fair amount of duplication. That's how an organic system sort of scales, uh, that, that you have that...I don't know how you call it. Slack? Robustness? Depend-- uh, d-duplication. I, I, I forget the-- I, I'm-- And this-- those-- these are not exactly the terms- Hmm ... I'm looking for, but I c-can't really think of the words. Okay. I was gonna go into Tangent and Tangle. Uh, so, uh, we, we sort of discussed the overall stats that, uh, Shopify has.Uh, but, you know, I, I think some, some pretty cool stuff that you guys are working on is your ML experimentation, uh, and your, your sort of auto tr-research training pipeline. Presumably you're much closer to this one because it's, it's a sort of personal hobby of yours. How, how would you explain them in, together?I thought we have a slide that, like, uh, has the s- the system diagram.[00:18:24] Mikhail Parakhin: Yeah. Tangle first and then Tangent as a-[00:18:27] swyx: Yeah ...[00:18:28] Mikhail Parakhin: as a thing on top of Tangle. And, uh, Tangle is the third generation, I claim, of, uh, systems of, uh, running any data processing, but a bit with a skew for ML experiments, but not necessarily. Any sort of data processing tasks where you need to iterate, share, and you have scale so that you want maximum efficiency.You know how, like, normally you would work, you would-- Imagine you're a data scientist or an ML practitioner, you would get Jupiter notebooks or, or maybe you would get, uh, you know, Pyth- your Python scripts, and you would manage the data, and you produce those TSV files, and you put them in some JFS or something.Then you would notice that, oh, it has this, uh, weird missing values. You go and write another script that, uh, goes and replaces them with, uh-[00:19:20] swyx: Ah ...[00:19:21] Mikhail Parakhin: dash S. And then, then you, then you run some, some, uh, “Oh, I need to filter bots.” And so you run some light GBM model that, uh, removes the bots. And then, then you like-- And then you, you kind of like get into shape, and then you start experimenting, and you run multiple experiments, and then you're like, “Oh my God,” like, “this experiment is worse.”You undo, and you cannot get to previous result. And like, “Ah, what did I do?” Like that. Again, then, then you finally like get everything working. Then you like start throwing it over the fence to production. You, you replicate it, those things don't work, and then sometimes you like don't notice that you forgot some feature naming and the, the features don't match.But then, like imagine you, you did everything, and then six months later you're like, have to repeat it because now there's more data, or you wanted to do another pass, and you're like, “What, what did I do?” Or like, or like, “This script crashes now,” or the, “the path has changed.” And then, then you're trying to, like you spend another month just doing ar- digital archeology on your own, you know, history, right?Now multiply that by many, many teams. Now imagine you got an intern that you wanna ramp up. Now you have to show that intern, “Oh, you know, look, here's the folder, there's the scripts, you know, ask your cloud agent to do, and then, uh, to, to figure it out.” And then cloud agent does something, and then you're, “Ah, yeah, right, right, it was the wrong folder.I forgot to tell you, I actually have this other thing I forgot myself.” And, and that's, that's the, like, the daily life we all, uh, all know it, uh, if, if you're a data scientist, machine practitioner, ma- machine learning practitioner or, uh, or even like any data managing, uh, person.[00:21:00] swyx: Yeah. So I, I used to do this, uh, f- uh, on the quant finance side, uh, in, in my hedge fund.So we did this before Airflow, and then, uh, obviously Airflow came along and, uh, then more recently Dagster, uh, I would say is like, in my mind, what I would use for that shape of problem, uh, where you had to materialize assets and create a pipeline.[00:21:19] Mikhail Parakhin: And that's, that's very good segue because... So Airflow is great, but Airflow is more about you, you have something and you wanna repeatedly run it in production on schedule.It's less about you as a team developing things and being able to share, and you grabbing the standard pipeline and saying, “Hey, I wanna change this tiny little component in the huge sea of data processing, and I don't wanna-- I wanna run ten experiments on this, and I wanna do hyperparameter optimization.”All that is very hard to do with Airflow. It's very easy to do with Tango. Tango is m- more about, it's everything about group of people Running experiments, it might be agents too nowadays. Uh, running experiments cheaply, collaborating, sharing results. Uh, you don't need to understand fully. You, you grab-- you clone somebody else's experiment or somebody else's pipeline, uh, run, uh, change small piece, run it, be, like, get it to production state, and then ship in one click.So then the... You don't have to port it into any other system to, to run in production. You can just run the same experiment. It's, it's fully production ready. And, and it's, uh, it has lots of... Again, as I said, it's third generation system. The original one was, I would claim there was Ether and then, uh, at least in my career, Ether was the first, first, uh, that pioneered this type of approach.And then there was, uh, Nirvana, which, uh, uh, at Yandex, which did kind of sec-second take on this. And now this one aggregates the, the learnings from all of those and, and Airflow as well to, to get to the state where you try it, it, it feels kind of magical. Uh, ‘cause now everything is based on content, uh, hashes.So even if the version changed, but if the output didn't change, nothing is being rerun. It's very efficient. If you... Multiple people start experiment that needs the same sort of data preprocessing, it's not repeated multiple times. It's automatically done only once. If you start ten experiments that all require, you know, some, some data preparation first as the first step, and you don't have to coordinate for that.Like, you don't have to know that other people are starting it. You now, it's very easy compos-, uh, composability, any language you can u- uh, you wanna use, and it's very visual. So you can see immediately, you can edit it easily, you can assemble small things with just even mouse clicks if you want to, and, uh, share, clone.And everybody knows also it's fully kind of static in the sense that we rerun it second time, it will exactly have the same results. Like, you will never have to do digital archeology. So full versioning and everything is also there.[00:24:06] swyx: Uh, so, so people can, uh... It's open source. Go to the GitHub repo and, and, uh, check it out.Uh, and it is also a really good, uh, blog post about it. I think all these is, like, really appealing. The, the, the, the thing that I think sells me the most about it is that, um, sort of development to production transition, right? Which I think, um, a lot of people haven't really solved that, uh, strictly, right?Like, we develop really, really well in, in Python notebooks, but then, you know, that's obviously not a sort of production ready process. I think that, like, any way in which that is solved, I think is, is very appealing. Then the other thing that you mentioned, which also raised my eyebrows, was content-based caching, which you mentioned is, is, um, you know, is ve-very much, uh, um, a sort of efficiency measure about, uh, you know, just like recalculation only on, on sort of content addressing Which I think makes sense.Uh, it surprised me that the savings could be this much, but maybe I just haven't worked at your scale where there's so much duplication, uh, that people just rerun because they change a single ID upstream.[00:25:10] Mikhail Parakhin: It does, yeah. But it's not only you rerun. The, the main savings are coming from the fact that you ran it, you got your job done, and you moved on.Then- Yeah ... somebody else in some department you don't know existed runs the same task, but on a newer version.[00:25:27] swyx: Yeah.[00:25:27] Mikhail Parakhin: Like right now, you can't, in, in most of the organizations, you can't even find out about it so that you can't even measure that you're spending that time twice, right? Here- Yeah ... if everybody's on Tango, that's detected automatically and detected that the output is the same.And then for that person, all it looks like is like experiment just suddenly moved, jumped forward, right? Uh, uh- Yeah ... so that's because, because the, there's network effect of multiple people helping each other.[00:25:51] swyx: Yeah. This is one of those things where it's designed to be a platform from the beginning rather than an individual developer's tool from the beginning, right?And, and everything's gonna streams down from there. That is the sort of Tango, uh, orchestrator, and it's, it manages jobs. We've seen a few versions of this, and this is obviously, uh, uh, the sort of, uh, unique approaches that you guys have, have, uh, figured out. And then there's Tangent.[00:26:14] Mikhail Parakhin: Yeah. And Tangent is basically an automatic auto research loop that can help and kind of do your work for you.Uh- ... you know, uh, effectively, effectively, Andrej Karpathy recently popularized it with auto research. Yes. Remember he said like he was, uh, speed running this, uh... Yeah, uh, you know the story. The, here we're basically bringing the same capability into Tango so that, uh, the, uh, Tangent can analyze it. It's just an agent that can run multiple experiments, figure out what can be changed, and keep on rerunning it, keep on modifying until, uh, maximizing some goal, some loss function, whatever you need to, to achieve.And in general, I would say if you're not using auto research-like approach in whatever you do, like literally whatever you do, then you're missing out. We saw at Shopify that taking like a wildfire, anything where you can put measurements can be done dramatically better. Our-[00:27:19] swyx: Mm-hmm ...[00:27:20] Mikhail Parakhin: uh, speed of, uh, templatization HTML, uh, completely new UX tem- uh, templatization of, uh, reducing latency for liquid themes.Uh, we-- Our, uh, search, uh, recently we moved from It's hard even, uh, quote from eight hundred QPS to forty-two hundred QPS with the same quality just by pure optimizations and not a research loop that kept running and changing code in our index serve on the same number of machines, just increasing the throughput.We, we managed to improve the quality of gisting and machine learning process. Uh, you know, gisting is the prompt compression technique that[00:27:59] swyx: allows for[00:28:00] Mikhail Parakhin: lower latency and, and lower and, uh, actually higher quality slightly. So like literally whatever different walks of life, and it doesn't have to be AI related.Uh, we, we had a reduction in, uh, storage because the agents would go and find data sets that clearly are derivative, uh, and then you don't need to store things twice. You know, we, we, we found somewhat embarrassingly that it was one of the largest tables was hashing random IDs into another random ID, and we literally- Oofput only one. So it was translating, yeah, two random IDs hashed[00:28:36] swyx: into[00:28:37] Mikhail Parakhin: each. So, so[00:28:37] swyx: it has access to the code as well, so it can, it can check the, like what, what the hell is it doing?[00:28:42] Mikhail Parakhin: So there, there cou- it could be run in two levels. You, uh, you know, at the superficial level, it could just use ex-existing components and, uh, reshuffle them.Uh, you know, like you can grab- Yeah ... uh, XGBoost, and you can grab some, some Py- PyTorch module, and then can grab some, you know, grab another tools and, and combine them. At a deeper level, since Tangle is all sort of CLI based underneath you, every, every component is a wrapped really CLI, uh, call and a YAML file, it can analyze code and create new components and, and, uh, keep on iterating as well.So, so you can, you can both have quick modifications of existing t- uh, pipelines with the, with components that are already there pre-baked, or you can create new components, uh, and-[00:29:29] swyx: Yeah ...[00:29:29] Mikhail Parakhin: keep iterating on those. So auto research is, again, this is probably the, the thing I was excited the most in the last two months happening, and we see it taking like, like totally like a wildfire.Just, uh, everybody, every day, every... well, every day, every minute, I would, uh, have somebody Slack message saying, “Oh, look how much better I made it.” And, uh, it's all throughout the research.[00:29:53] swyx: Is this democratized in some way in, in the sense that like is it your ML, uh, engineers and researchers doing this, or is it your regular PMs and software engineers also have the ability to auto-- to use Tangent?[00:30:07] Mikhail Parakhin: This is an awesome question. Like, Tango in general and Tangent in particular are extremely democratizing. Like they- Yeah ... they are the main tools for- ‘Cause I don't[00:30:15] swyx: need the details.[00:30:16] Mikhail Parakhin: Yeah. Exactly. Initially used by ML and AI engineers, but then literally, as you said, PMs are like the highest user right now is one of PMs on our org, uh, Sartak and he was, he was number one by, by usage of, of this ‘cause they're just, uh, energetic and knowledgeable, and now it, it unlocks a lot of capability where you don't have to co-change code manually.[00:30:39] swyx: I mean, I mean, because it kind of cuts out the ML, ML engineer from the process because the, the, the PMs have the domain knowledge and the ability to think about, uh, from first principles about, okay, what, what results do I want? And they can-- they even have the access to the data that, that needs to go in.So it's like in some ways, like this is the magic black box that we've always wanted for, for training and, and for, uh, I guess, uh, uh, hill climbing, whatever.[00:31:04] Mikhail Parakhin: It's basically cloud code for your AI development- ... uh, situation, right? Like now, now you don't have to know exactly how algorithms work. You can just, uh, bring your domain knowledge and expertise and product knowledge and iterate within Tangent until you've gotten the results that you need.[00:31:21] swyx: In my previous roles, every time that someone has pitched AutoML, you know, I've always been like, “Uh, this is not, this is not gonna work. It's, you know, it's, it's always gonna be a flop.” Somehow it's working now. I mean, presumably the answer is now we have LLMs and it's good enough, right? It's, it's an emergent property that we can do auto research, but like, it doesn't feel that satisfying that how come we didn't do this before, right?Like we just did like parameter search and like, I don't know. That's maybe that's it.[00:31:48] Mikhail Parakhin: Yeah. Bayesian optimization and hyperparameter optimization was, was the one that, or facet of AutoML that was used very actively, which incidentally also built into, uh, Tango. But, you know, I know Patrice Simard very well, and, uh, he was such a, uh, such a proponent of AutoML, and he put, like literally spent careers trying to democratize it.Without LLMs, it just turned out to be very hard. Like it, you, you would have flexibility within certain narrow domain, but it was hard to wider scale, and now with LLMs suddenly it's like magic wand, and so suddenly everybody- ... is an AutoML expert.[00:32:28] swyx: Yeah, I, I think it's multiple things, right? Like I'm, I'm just gonna bring up the, the, the chart again, right?Like LLMs can do the monitoring very well. That is the very potentially unbounded, super unstructured. It can do the analysis very well, it can do the... Uh, and basically it is much more intelligence poured into every single step. Uh, there's maybe nothing structurally changed about AutoML, but this is just m-more intelligent and more unstructured.[00:32:53] Mikhail Parakhin: Exactly.[00:32:54] swyx: Any flaws that you've run into? Like everyone is like drinking the Kool-Aid, oh my God, time savings, uh, you know, performance improvements. Like what, what, uh, issues have you have, uh, come up?[00:33:06] Mikhail Parakhin: This is really cool. It's not a solution to all the world's problems for sure. The limitations are usually the ones I-- And this is where we get into a bit of a subjective territory.Uh, I can only share what I've, I've seen so far, and I'm sure the situation, uh, is changing, and, you know, maybe after I say it, like many people will reach out and say, “Hey, what about this?” And you don't know that, and then, then we'll be probably right. But what I've seen is auto research is very good at doing kind of obvious things that you don't have bandwidth to do or you didn't notice or maybe you're not aware of like the-- some standard practices.It is not good at doing something completely out of distribution, something that, you know, you have to think for, for multiple days, uh, and, and do something like none of this. So, so it's, uh, I, uh, set an experiment once, uh, on, on my sort of, uh, hobby thing, and I let it run for, uh, ended up, uh, several weeks run, uh, you know, it's like full production kind of scale, so it, you know, slow runs and, and it ex-- it performed in the end, uh, over four hundred experiments, and only one was successful.I'm like, “Okay, that's, that's good.” But-[00:34:18] swyx: But it saved time.[00:34:19] Mikhail Parakhin: Yeah, I saved time. Like it, it was the, that thing. Yeah, if I, if I were doing four hundred experiments myself, my betting average, as I said, would have been much higher, I'm sure. But also, first of all, it would take me like three years to do four hundred experiments.And, uh, I didn't have to do them. Like the machines were just, uh, the price of electricity did that. So, and I got one improvement, uh, that in, uh, my, my-- Honestly, when I was starting that experiment, my thinking was to go and show that, “Hey, Andre, maybe you just don't know how to optimize.” And I was super smart because in, in my pro-problem, it was optimized for many years, and it was like fully improved.Uh, and I didn't expect it, you know, auto research to find anything at all. Yet it did. So instead of making fun of Andre, I ended up, uh, a big, big supporter. Yeah, that's exactly the tweet. Yes.[00:35:10] swyx: You and Toby really, really go back and forth on-online a lot, which is really funny. Uh, think of it as, as an eval for the optimalness of the code it's running on.Uh, it's almost like it reminds me of like a Kolmogorov complexity thing, but, uh, I guess it's-- there's some optimal thing that you're trying to sort of reduce down to, I guess. Um, and so, so you, you, you know, you should congratulate yourself that you had, uh, you know, uh, ninety-nine percent, uh, optimality.[00:35:36] Mikhail Parakhin: Exactly, yeah. I think Andre really deserves a lot of credit for popularizing this approach. This is, uh, this is incredibly, I think, powerful and cool and You know, the, uh, even him, him just mentioning it led to a lot of gains in a lot of places in the industry, so we should be thankful.[00:35:56] swyx: Yeah. I think he also has a just...I don't know what it is. Like, um, you know, it, it is a simple self-contained project that people can take and apply to other things, which is, is, is one thing, but also just the name. Just like somehow no one, no one managed to call their thing auto research. It's just naming things is very important. I think that that is mostly, uh, our coverage of Tango and, and, uh, Tangents.I think obviously, you know, there's a lot of, uh, ML infra at, at Shopify that people can, uh, dive into. We're about to go into SimGym, but before I do that, any, any other sort of broader comments around this whole effort? Like where is it, where is it leading to?[00:36:36] Mikhail Parakhin: As a segue to SimGym, like all those things start composing strongly.And, uh, you could see a huge unlock when you can look at each one of the tools and, and you see, oh, they're extremely useful. Uh, Tango is useful by itself. Auto Research is useful by itself. SimGym is useful by itself. If you combine all three, you create like synergetic effect. I think that's why we wanted to even, uh, cover them today is because this is something that if you go back even, you know, five years ago, would've been unthinkable.Uh, replicating that, uh, would, would be either incredibly costly or impossible, right? With probably thousands of people are required.[00:37:20] swyx: Well, we have serverless human, uh, serverless intelligence, right? Like, uh, so yes, you do have thousands of hu-- of, of intelligences, not just, not humans. And that's, that's close enough, right?Even if they're not AGI, they're, they're close enough to do the, the task that you need them to do. And, and, you know, that's, there's plenty for, for a lot of routine work, knowledge work. Okay, let's get into SimGym. Um, this is one of those things I, I was surprised to see actually it's apparently your, uh, one of your most popular launches, and I think something that, uh, I think Sim AI, I think Yunjun Park, who did the Smallville thing, there's a very small cottage industry of people trying to do like the simulate customer thing.I think a lot of people maybe don't super trust this yet because they're like, well, obviously they would just do what you prompt them to do, right? But maybe just think, uh, tell us about the sort of inspiration or origin story.[00:38:10] Mikhail Parakhin: That's exactly actually the thing I wanted to cover, because if you don't have the historical data, all you can do is prompt a-agents in a vacuum, and they will do exactly what you prompt them to do.In fact, when I first proposed it, and this is a bit of, um, my brainchild initially, if I, I can boast, even Toby said like, “But wouldn't they, they just repeat what, what you tell them?” And, uh, but I'm like, “Yes, except Shopify has decades of history of how people made changes and what there is, uh, there, what it resulted in terms of sales.”So now what we can do is we can-- we have this... It's not, it's a noisy data. There's a small, usually websites, uh, you know, like things, things are never in isolation. It's almost never AB experiment. It's always AA experiment when there's has two meanings, but basically, you know, in different time you run two different things.But if you aggregate in general, uh, like everything together, and you apply, uh, denoising and collaborative filtering like approach, you can extract a very clear signal. And then you can optimize your agents. And that's why it took so long. It took almost a year of that optimization of just us sitting and fiddling, and, and we had this internal goals of correlation of hitting-- internal goal was to hit zero point seven correlation with, uh, add to cart events, for example.Like that, that if we run real AB test experiment, that it should, it should go and, and rep-uh, replicate, uh, same sort of success that, that humans had or lack thereof. And it, it took forever, and I don't think that's easily replicatable because, uh, like who else would have that data? You have to have this historic, you know, decades, uh, worth of data.And now, now the, like the other thing you need is in-infrastructure and the scale, right? Because, uh, w- again, what we found, uh, stat sig results, you need to run a lot of simulations, a lot of agents, and, and it's-- Those are expensive things. Like you're, you're making actions in the browser because you want a real friction.You want to, to be able to get the image like of what humans will see because you wanna, uh, detect effects like, “Hey, if I make my images larger, will I have more sales or l- uh, fewer sales?” And like usually people's intuition here, by the way, is that I increase my images, I will have more because they look nicer.You know, designers all look sparse and big images. Like usually your sales tank, right? But, but, uh, you know, from HTML, all the characters look the same only the, the size tag looks different, right? So it's very hard. So you have to take visual information, you have to run this in simulated browser environment on the big farm and, and of course, you have to have, uh, like very, very expensive model, good model with multi-model model.So all this it's-- is what's taken so long and, uh, to share my personal fail a little bit there, Sean, is like, you know, we always had this bias to-- for like large company bias. You know, we always, uh, whenever you-- we do, we're like, “Hey, we'll run an experiment,” right? We make, make a change, and we will run an experiment and then, uh, see, uh, see which one's better or like, “No, this is worse,” and most of them are worse, so you discard it and keep iterating, hill climbing.And we're like, “Oh, like smaller merchants, they cannot get stat sig results. They cannot really run experiments simply because, you know, in a week there would be not enough data for them.” So we thought from this perspective. What we didn't realize is that most people don't have A and B, they just have one thing, and they need suggestions of What A and B should be.So, uh, we first build this, hey, we run simulation on two separate teams and, and, uh, say, “Hey, which one is better?” We then morphed it into, and very recently just released it, when you have just your site, your theme, we run over it and we say, “Hey, here's what predicted values of, of, uh, uh, conversions are, and here's how we think you should modify it to increase your conversions.”And then circling back to what you started with, the proof is in the pudding. Like, if we are not correlating with reality, like, people will not be using it. And, uh, thankfully, we see literally every day more users than the previous day. So, so right now, uh, right now- It's working. Yeah. I'm-- Right now my problem is how to pay for it all because the so our major thing is how to optimize the LLMs, do distillation, how to run the headless browsers, uh, and handful browsers, uh, uh, cheaper so that we can accommodate the increase in traffic.[00:42:47] swyx: Yeah. I, I understand that you, uh, you published a lot of technical detail at GTC, so I was just gonna bring it up a little bit. I think s- was this in, in con-conjunction with some kind of GTC presentation? Or something like that, right?[00:42:59] Mikhail Parakhin: Well, we, yeah, we, we did it in several place, but yeah, we had the engineering- Yeahblog, uh, as well. Yeah.[00:43:05] swyx: Yeah. So you're running, uh, GPT OSS. Uh,[00:43:08] Mikhail Parakhin: the, this is an older version. You know, now we run multimodal model. But yeah- Yeah ... GPT OSS, we still run GPT OSS as well for[00:43:15] swyx: And then you have the VMs, and you also have browser-based. I really like this one where it you said, “It violates almost every assumption that standard LLM serving is designed for.”And then you had like, basically orders of magnitude differences between everything.[00:43:29] Mikhail Parakhin: Exactly. Which is, which, uh, which was, you know, a bit of a challenge to implement, like when, like even simple things. Uh, be- since it violates all the assumptions, for example, multi-instance GPUs, like MIGs don't work as well.But we needed, uh, to get MIG to work because, ‘cause otherwise it's way too expensive. And so we had to deal with the, yeah, with, uh, lots of infrastructure and, and, uh, work with, uh, uh, Fireworks and CentML, uh, you know, to help with optimizations and browser-based, as you mentioned. Yeah, like, takes a village.[00:44:04] swyx: Okay. So there's a lot of like, I guess, experimentation in the infrastructure so far, and you've published more or less what you have here. I guess I'm, I'm less familiar with CentML. I, I don't do, uh, that much work in this, this part of the stack. But why was it the sort of preferred instance platform?[00:44:22] Mikhail Parakhin: There are really three probably top companies. There used to be, uh, uh- Three top companies, uh, at least I was aware of that did, uh, LM optimization. You know, together Fireworks and Santa ML, not necessarily in that order. Santa ML recently got acquired by NVIDIA. Uh, what they did is if you have a model and you want to optimize it to a specific prof-- uh, profile of usage, uh, they would go and do it.And, uh, we work with, with those companies, uh, this was work particularly in with Santa ML and NVIDIA to get them the best possible results out of it. And, and sometimes you, you have to retune depending on, like sometimes you want the maximum throughput, sometimes you want minimal latency, sometimes you want like the cheapest, right?And, yeah, or some combination. And so yeah, these are people who would come and help you.[00:45:14] swyx: I see. I see. Yeah, yeah. I'm familiar with these people for the LLM, you know, autoregressive stack. But the other interesting category of these optimizers is also the diffusion people, whereas like Fel and, you know, uh, Pruna recently has come up a lot as well, which I think is like really underappreciated, uh, at least by myself, because I, I thought, oh, all the workload would be LLMs, but actually there's a lot of diffusion as well.[00:45:38] Mikhail Parakhin: Exactly.[00:45:38] swyx: There's a lot here, so I, I, I... it's, it's, uh, it's, it's, it's hard to cover. But I, I do think like people underappreciate the importance of customer simulation, basically. I think this is something that I'm candidly still getting to terms with. Uh, you know, uh, you also-- your team also like prepared this, like, really nice diagram.Uh, I, I assume this is AI generated.[00:46:00] Mikhail Parakhin: Yeah, it looks-[00:46:01] swyx: Maybe it's not.[00:46:01] Mikhail Parakhin: Yeah, it looks, uh, Gemini-ish. Yeah, but, uh, uh, honestly, I, I don't know where, where the hell they generated. It looks, look, uh, looks like it's, uh, Google. But the interesting part, John, that, that, uh, we haven't covered, but I, I wanted to mention is if your store had previous customers, rather than it's a new store, you're like new merchant just launching things, it helps tremendously in just correlation and forecast.Yeah, we take your previous, uh, customer's behavior, and we create agents that replicate those specific distribution of, of customers that you get, and then we a- we apply those to your changes, and then that, that raised raw, you know, the re-- uh, just correlation with the add to cart events or to-- with conversion or whatever it, it, it may be, uh, quite dramatically.So, uh, replicating humans in general seems like an interesting, cool challenge.[00:46:58] swyx: As a shareholder, I think this is the-- like if people are Shopify shareholders, they should really deeply understand this because this is basically the moat. The, the more you use Shopify, the more it will just automatically improve, right?Like you're, you're doing the job for them.[00:47:13] Mikhail Parakhin: Yeah, that's what we started with. Like, uh- ... uh, otherwise, if you're just a startup, I wouldn't do it if, uh, you know, if it was my startup because Without the data, it, yeah, as, as you said, it's, it's exactly the case that, uh, whatever you say in prompt, that's, that's what the agents will be doing.[00:47:30] swyx: The statistician in me wants to like really satisfy the sort of, um, statistical intuition, I guess. Um, to me it's kind of, uh, the, the word that comes to mind is, um, ergodicity. Uh, so let's say a, a customer takes this path, customer takes this path, customer takes this path, right? Um, the... In my mind, the way I explain it is like, okay, here, here's the ninety-five percentile, here's the five percentile, and here's the median, right?Um, but to me, what SimGym is potentially doing is that it can, uh, modify... It can sort of model the sort of in-between sort of journeys as well, that, that maybe are dependent on the previous states. This may be like a very RL-type conclusion where like basically the summary statistics, if you only did naive AB testing, you only have the, the statistics at, at, at a certain point, and you only judge based on the sort of overall summary statistics.But here you can actually model trajectories. Does that make sense? Or-[00:48:31] Mikhail Parakhin: That makes total sense because like, well, that, that makes even more sense that maybe even you realize bec- because-[00:48:38] swyx: Okay. Please,[00:48:38] Mikhail Parakhin: please. Yes ... we do-- Yeah. The, so internally, uh, we have this system, we talked about it briefly once at NeurIPS.We have a huge HSTU-based system that models the whole companies, uh, and their possible paths. And like- Yeah ... what you are, what you are showing, like actually at any point of time, you can either model the user's behavior or you mo- can also think about, uh, the whole merchant as a company, as the entity that acts in the world.You can model that as well. And then you can do, can do counterfactuals. In your graph, like in your blue graph, uh, if you're... Imagine in the center there, uh, somewhere in the middle, you would have an intervention. I give that person a coupon, or I don't know, I send a personal thank you card, or give a discount in some- somewhere.And then you can, uh, then you can do forward rollouts from that counterfactual. So what would have happened with that intervention or without the intervention? And you can even ch- change where that intervention, uh, in time can happen, right? Like some- where, where in this journey. So we, we do this at the Shopify scale for our merchants, and then if we notice that something that they can be fixing, like there's a strong counterfactual, like we have Shopify policy, they basically get a notification like, “Hey, we think your...something is wrong with your-” I don't know, Canadian sales. Like, uh, it looks like it's misconfigured. Here's what you need to do. Or do you think like, uh, you have to set up this campaign with these parameters? And we do that at the buyer level to literally offer discounts or cashback or, or things to buyers.So this is-- I'm getting very excited. Like this is my sort of area of, uh, interest, I guess, and, and hobby. But being able to m-model something complex as human beings or companies and model counterfactuals on it, where you can have interventions in the future and optimize when to make intervention, what kind inter-- uh, what kind of intervention to make.It's such an unlock that previously was completely impossible. Like the-- it was, it was always dreamed of, but never... Like how would you even simulate it without LLMs or HTUs? I think very, very exciting times.[00:50:59] swyx: I just wanted to, uh, to maybe illustrate this. I, I'm not the best illustrator, but I, I am a conceptual statistics guy.And y-you know, you cannot just do this. Like this is a dimensionality AB test doesn't do, right? Like, uh, because it doesn't have the, the, the change over time, uh, stochastic nature, uh, and it doesn't have the sort of contextual like... Here's all the context to this point. Um, okay, cool. Um, that's SimGym.You're, you're gonna burn a lot of tokens on this thing. But you're, you're one of the, the only scale platforms in the world that can, uh, that can do this across a huge variety of workloads, right? I'm even curious on a sort of human, uh, research level of like, well, do, does retail behave d-differently from like clothing sales?D-does that behave differently from electronic sales? I, I don't know. I don't know what else you guys... The Kardashian shoppers, do they differ from like people who buy, uh, I don't know, cars and, uh, whatever.[00:51:55] Mikhail Parakhin: Well, very different, and different sensitivities and different modes of, uh, shopping and, and different levels of what's important.Now, to-totally, you can do aggregations at, uh, at a store level. You can do aggregations at a different, uh, category level. I don't know if, uh, you know, for our statisticians among us, I couldn't believe, but we-- recently we're looking at it, and we had to bring back, uh, CRPs, you know, Chinese restaurant process.It's a, like, way of aggregating and, like, naturally grow clustering. So across... Specifically to answer questions that, uh, like you were just posing on how, how if, if buyers behave different categories. And I'm like, “I haven't seen CRP since two thousand and one.” It's[00:52:37] swyx: so What? It's so- What is... No, I haven't, I haven't seen this.No. This is not in my training. Uh,[00:52:44] Mikhail Parakhin: but, but yeah, it, uh, uh, it actually, like the, the-- there was a very popular kind of theory, popular neurips HTML circles in early two thousands, uh, kind of nice. And now, now it has practical applications, uh- Yeah ... that we were resurrecting.[00:53:03] swyx: Yeah, amazing. Uh, I, I can see, I can see how this is like a, uh, a fun job for you where you get to apply all these things.Um, yeah, yeah, so super cool. Super cool. So, okay, so, so anyone who, who knows what CRPs are and has always wanted to use them at work, uh, they should, they should definitely join Shopify. Okay, so w-we have a lot and but I, I'm, I'm being mindful of the time. I, I do wanted to, to sort of cover some other things.Um, I-I'll give you a choice, UCP or Liquid?[00:53:30] Mikhail Parakhin: Liquid. I think, I think on UCP, you know, like UCP is very important for us and, and it just we are-- UCP, we have a structured, uh, discussions, and you can read about them, and we have, uh, blog posts, and we have a big release this week, in fact, like with our catalog.Oh,[00:53:46] swyx: okay.[00:53:46] Mikhail Parakhin: Uh, yeah,[00:53:46] swyx: but- Le-I mean, we, we can, we can discuss the, the, the release briefly because we'll release this after the-- after it's already announced so whatever. There's a catalog that you guys are doing?[00:53:55] Mikhail Parakhin: Yeah. So we are, we are- Okay ... we are bringing in capabilities of a whole, uh, Shopify catalog.Basically, you now you can search for products, you can do lookups by specific ID, you can do bulk lookups when you need to bring m-multiple products. You don't need to know in ad-in advance what you're trying to show or to sell or check out. Like, you can now, you can now have this decided at, at runtime, and this big area for investment for us for both non-personalized and personalized searches, trying to provide basically a win-window into whole universe of products that are being sold everywhere in the world.And Shopify is really not exactly, but almost like a super set of any-anything being sold. Now we are bringing it into UCP and, uh, and, uh, identity linking is another big thing for us, uh, so that you, you can use, uh, like Google or whatever, whatever identity you have, uh, they're minimizing friction.[00:54:56] swyx: Yeah. So[00:54:57] Mikhail Parakhin: yeah, big release for us.But Liquid AI of course we never talk about, and the problem might be more, more aligned with what we d-discussed previously on this chat.[00:55:07] swyx: Sure. The main thing that everyone understands about Liquid is that it is inspired by Worm, and I still don't know why. I'm curious on your explanation. I think you, you, uh, you can make things very approachable.And also I think like what is the potential of like the, the level of efficiency that you get out of Liquid?[00:55:23] Mikhail Parakhin: You- we all familiar with transformer architectures. And, uh, for the longest time, there was a competing architecture, it's called the state space models. So, so Sams, uh, you know, Chris, Chris Reyes, one of the pioneers and, and lots of startups, uh, trying to make those realities.They have, uh, significant benefits being main being, uh, being much faster and, uh, lower footprint and not quadratic in length, you know, sort of, uh, linear in, in, uh, in your context length. But with state space models- They never quite made it. Like they're used-- They have, uh, certain niches when they thrive, their hybrid architectures are useful, but they never quite made it.And liquid neural networks are, you can think of them as a next step, like, uh, sort of, uh, state-space model square. It's non-transformer architecture that's more complicated than sta-state space and really difficult to code if you-- if I'm being honest. But it's, um, very efficient. It's, uh, subline-- sub, uh, quadratic in, in length of your context.Uh, it's very compact way to represent things, and that's a liquid AI company. They... Their goal is to productize it, and very often you have this need, uh, when you need to have long context and small model, and you want to have low latency. Like in general, it's basically on par with transformers, and if you do hybrids with transformers, it's, it's even better.That's why we at Shopify, when we tried multiple and we constantly try multiple models, multiple companies, we found that for small, particularly with low latency applications, when you have low latency and/or if you need longer context lengths, liquid was the best. And so we still use the whole zoo and always like obviously test and use everything, uh, every open source model and, you know, it feels l
Have you or a loved one been afflicted by "brain fry" after managing too many autonomous agents? This week on the Friday Deploy, Andrew and Ben explore the cognitive toll of orchestrating AI swarms and share Kelly Vaughn's expert strategies for avoiding burnout. The hosts also discuss Google's new campaign to punish websites that hijack the back button, the breakthrough of running Gemma 4 natively on mobile devices, and a new 8-step maturity model for building agentic data pipelines. Finally, they dive into a heated debate over whether Obsidian flat-files are a scalable memory solution for AI, comparing the methodology to Andrej Karpathy's new agent-compiled wiki system.Read the guide: The APEX FrameworkFollow the show:Subscribe to our Substack Follow us on LinkedInSubscribe to our YouTube ChannelLeave us a ReviewFollow the hosts:Follow AndrewFollow BenFollow DanFollow today's stories:Introducing a new spam policy for "back button hijacking"Google Gemma 4 Runs Natively on iPhone With Full Offline AI InferenceWater Town: The Agent Swarm Data StackStop Calling It Memory: The Problem with Every "AI + Obsidian" TutorialThe Wiki That Writes ItselfBreaking out of the "brain fry" spiral of AIAfter Burnout by Kelly VaughnOFFERSStart Free Trial: Get started with LinearB's AI productivity platform for free.Book a Demo: Learn how you can ship faster, improve DevEx, and lead with confidence in the AI era.LEARN ABOUT LINEARBAI Code Reviews: Automate reviews to catch bugs, security risks, and performance issues before they hit production.AI & Productivity Insights: Go beyond DORA with AI-powered recommendations and dashboards to measure and improve performance.AI-Powered Workflow Automations: Use AI-generated PR descriptions, smart routing, and other automations to reduce developer toil.MCP Server: Interact with your engineering data using natural language to build custom reports and get answers on the fly.
This Week In Startups is made possible by:Pilot - https://pilot.com/TWISTSquarespace -https://squarespace.com/TWISTNorthwest Registered Agent - https://northwestregisteredagent.com/TWISTPlaud - https://Plaud.ai/twistCrowdHealth - https://JoinCrowdHealth.com/twistToday's show:TAO, the token associated with the Bittensor network, fell sharply last night after a major project declared it was taking its business elsewhere. Covenant AI had used Bittensor subnets 3, 39, and 89 before its abrupt and controversial decision to depart for other digital shores.Gareth Howells, a co-founder of Vidaio (Bittensor Subnet 85), joined Jason and Alex to share the community perspective regarding Covenant's decision to leave Bittensor, a choice that was especially bitter given that its recently-released Covenant-72B model trained using the network's decentralized capabilities was hailed at the time as a demonstration of the crypto-AI hybrid's potential. (Covenant's Sam Dare was a recent guest on TWiST.)Howells walked us through Vidaio's product (AI models for high-quality video work, like upscaling) Then it was time to bring new fan-favorite Ole Lehmann on. Ole took an Andrej Karpathy concept, and turned it into a skill that you can use in either OpenClaw or Claude Cowork. GuestsOle Lehmann: https://x.com/itsolelehmannThe AI Solopreneur newsletter: https://aisolo.beehiiv.com/Vidaio: https://vidaio.io/Vidaio on X: https://x.com/vidaio_Timestamps:0:00 Drama in the Bittensor space! Was there a rugpull?1:01 Plaud: If your work depends on conversations — interviews, meetings, calls — you need a Plaud NotePin. You can check it out at https://Plaud.ai/twist and use code TWIST for 10% off!1:47 Gareth Howells of subnet85 joins the show: https://x.com/GarethHowells9:56 Northwest Registered Agent - Get more when you start your business with Northwest. In 10 clicks and 10 minutes, you can form your company and walk away with a real business identity — Learn more at https://northwestregisteredagent.com/TWIST11:37 Gareth reacts to the Covenant situation… Did Sam sell tokens?14:33 How VidAIo (subnet85) optimizes video as cheaply as possible. https://vidaio.io/20:06 Squarespace - Use offer code TWIST to save 10% off your first purchase of a website or domain at https://squarespace.com/TWIST21:12 Demo: Gareth shows upscaled footage of our very own Lon Harris27:09 Who is actually doing the computing on Bittensor?28:15 How Bittensor recruits tech talent from emerging markets29:42 Pilot - Visit https://pilot.com/TWIST and get $1,200 off your first year.30:58 Jason's explains the Bittensor network. https://docs.bittensor.com/37:28 Jason addresses "pump and dump" allegations40:46 Ole Lehmann joins the show https://x.com/itsolelehmann45:31 Jason's $1,000 "bounty" for enhanced show notes49:35 Check out Ole's newsletter at aisolo.beehiiv.com !59:50 CrowdHealth - CrowdHealth lets you ditch the bureaucracy with a peer-to-peer funding platform for your healthcare. Get started for $99 per month for your first three months by using the code TWIST at https://JoinCrowdHealth.com/twist.1:00:46 Off Duty with Alex and J-Cal!1:04:46 Reacting to the Maul trailer https://www.youtube.com/watch?v=DkVepshZhGcSubscribe to the TWiST500 newsletter: https://ticker.thisweekinstartups.comCheck out the TWIST500: https://www.twist500.comSubscribe to This Week in Startups on Apple: https://rb.gy/v19fcpFollow Lon:X: https://x.com/lonsFollow Alex:X: https://x.com/alexLinkedIn: https://www.linkedin.com/in/alexwilhelmFollow Jason:X: https://twitter.com/JasonLinkedIn: https://www.linkedin.com/in/jasoncalacanisFollow TWiST:Twitter: https://twitter.com/TWiStartupsYouTube: https://www.youtube.com/thisweekinInstagram: https://www.instagram.com/thisweekinstartupsTikTok: https://www.tiktok.com/@thisweekinstartupsSubstack: https://twistartups.substack.com
Welcome to Exponential View, the show where I explore how exponential technologies such as AI are reshaping our future. I've been studying AI and exponential technologies at the frontier for over ten years. Each week, I share some of my analysis or speak with an expert guest to make light of a particular topic. To keep up with the Exponential transition, subscribe to this channel or to my newsletter: https://www.exponentialview.co/ ---- Published in early March 2026, Andrej Karpathy's autoresearch AI tool makes autonomous scientific experimentation cheap and easy — but it was designed to solve machine learning problems. I wanted to see if I could apply its loop architecture to my own work: refining my worldview, testing arguments, solving business problems. In this video, I share how I adapted Karpathy's autoresearch loops for problems that aren't easy to quantify, how to avoid the local minima trap, and the broader impact of these kinds of methods. I covered: (02:11) The Karpathy Loop: what is it and how does it work (07:54) Extending the loop into business and thinking (09:46) The local minima trap (12:20) The escape harness: getting beyond “good enough” (16:05) What I've learned after 30 days (18:47) The loop economy: from doing to judging ---- Where to find me: Exponential View newsletter: https://www.exponentialview.co/ Website: https://www.azeemazhar.com/ LinkedIn: https://www.linkedin.com/in/azeem/ Twitter/X: https://x.com/azeem Production by EPIIPLUS1. Production and research: Baba Films, Chantal Smith, Marija Gavrilov. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
On this week's show, Patrick Gray, Adam Boileau and James WIlson discuss the week's cybersecurity news. They talk through: TeamPCP's supply chain attack on Github, and they threw in an anti-Iran wiper, because why not?! Anthropic hooks up its models to just… use your whole computer After Stryker's Very Bad Day, CISA says maybe add some more controls around your Intune? Another iOS exploit kit shows up in the cyber bargain-bin The FTC decides to ban… all new home routers?! U wot m8?! Supermicro founder was personally sanction-busting Nvidia GPUs into China?! This week's episode is sponsored by enterprise browser maker, Island. Chief Customer Officer Bradon Rogers joins Pat to explain how its customers are using Island to control the use of personal AI services in regulated industries. This episode is also available on Youtube. Show notes ‘CanisterWorm' Springs Wiper Attack Targeting Iran TeamPCP deploys CanisterWorm on NPM following Trivy compromise Andrej Karpathy on X: "Software horror: litellm PyPI supply chain" attack Checkmarx KICS GitHub Action Compromised: Malware Injected in All Git Tags Felix Rieseberg on X: "Today, we're releasing a feature that allows Claude to control your computer" A Top Google Search Result for Claude Plugins Was Planted by Hackers Lockheed Martin targeted in alleged breach by pro-Iran hacktivist CISA urges companies to secure Microsoft Intune systems after hackers mass-wipe Stryker devices FBI seems to seize website tied to Iranian cyberattack on Stryker Stryker confirms cyberattack is contained and restoration underway Hundreds of Millions of iPhones Can Be Hacked With a New Tool Found in the Wild Someone has publicly leaked an exploit kit that can hack millions of iPhones Russia-linked hackers use advanced iPhone exploit to target Ukrainians Apple rolls out first 'background security' update for iPhones, iPads, and Macs to fix Safari bug Post by @wartranslated.bsky.social — Bluesky Signal's Creator Is Helping Encrypt Meta AI Hacker says they compromised millions of confidential police tips held by US company Millions of 'anonymous' crime tips exposed in massive Crime Stoppers hack Feds Disrupt IoT Botnets Behind Huge DDoS Attacks FCC bans import of consumer-grade routers amid national security concerns White House pours cold water on cyber ‘letters of marque' speculation Google launches threat disruption unit, stops short of calling it ‘offensive' Supermicro's cofounder was just arrested for allegedly smuggling $2.5 billion in GPUs to China Cyberattack on vehicle breathalyzer company leaves drivers stranded across the US Man pleads guilty to $8 million AI-generated music scheme Two Israelis AI generated "intelligence" and sold it to Iran
No Priors: Artificial Intelligence | Machine Learning | Technology | Startups
What happens when AI agents can design experiments, collect data, and improve — without a human in the loop? Andrej Karpathy joins Sarah Guo on the state of models, the future of engineering and education, thinking about impact on jobs, and his project AutoResearch: where agents close the loop on a piece of AI research (experimentation, training, and optimization, autonomously). 00:00 Andrej Karpathy Introduction 02:55 What Capability Limits Remain? 06:15 What Mastery of Coding Agents Looks Like 11:16 Second Order Effects of Natural Language Coding 15:51 Why AutoResearch 22:45 Relevant Skills in the AI Era 28:25 Model Speciation 32:30 Building More Collaboration Surfaces for Humans and AI 37:28 Analysis of Jobs Market Data 48:25 Open vs. Closed Source Models 53:51 Autonomous Robotics 1:00:59 MicroGPT and Agentic Education 1:05:40 Conclusion
I break down Andrej Karpathy's new open-source project, Autoresearch: what it is, how it works, and why some of the smartest people in tech are losing their minds over it. I walk through 10 concrete business ideas you can build on top of Autoresearch loops, from niche agent-in-a-box products to always-on A/B testing agencies. I also cover Karpathy's companion launch, Agent Hub, share community reactions, and show you step by step how to get started using Claude Code and a Colab GPU. I'm hosting a free workshop so you can build your business in the age of AI. Sign up here: https://startup-ideas-pod.link/build-with-ai-2026 Links Mentioned: Autoresearch Github: https://startup-ideas-pod.link/autoresearch Timestamps 00:00 – Intro 00:45 – How Autoresearch Actually Works 02:40 – Visual Walkthrough of the Autoresearch Loop 03:37 – Mental Model: Your Research Bot That Runs While You Sleep 05:26 – Idea 1: Niche Agent-in-a-Box Products 06:48 – Idea 2: A/B Testing for Marketing (Landing Pages & Ads) 08:45 – Idea 3: Research as a Service 09:43 – Idea 4: Power Tool Inside Your Own SaaS 10:49 – Idea 5: Agency That Runs 100× More Tests 12:05 – Idea 6: Auto Quant for Trading Ideas 13:44 – Idea 7: Always-On Lead Qualification & Follow-Up 14:21 – Idea 8: Finance Ops Autopilot for Businesses 15:09 – Idea 9: Internal Productivity Lab for Your Org 15:53 – Idea 10: Done-for-You Research & Due Diligence Shop 16:41 – Non business use cases 18:27 – Karpathy's Agent Hub Announcement 19:50 – How to Get Started with Autoresearch 22:21 – Final Thoughts Key Points Autoresearch is an open-source AI agent that sets a goal, runs experiments in a loop on a GPU, keeps the winners, and discards the rest — all while you sleep. You need an NVIDIA GPU to run it (tested on H100), but you can rent one cheaply through Lambda Labs, Vast AI, RunPod, Google Cloud, or Google Colab. The fastest way to get started is to use Claude Code to walk you through installation, then run it on Google Colab with a T4 GPU runtime. Ten business ideas built on Autoresearch span niches like SaaS optimization, A/B testing agencies, trading backtests, CRM lead scoring, and done-for-you due diligence. Karpathy also launched Agent Hub — essentially a GitHub designed for agent swarms to collaborate on the same codebase. The project already has 25,000+ GitHub stars and is growing fast; early movers who tinker now build an unfair advantage. The #1 tool to find startup ideas/trends - https://www.ideabrowser.com LCA helps Fortune 500s and fast-growing startups build their future - from Warner Music to Fortnite to Dropbox. We turn 'what if' into reality with AI, apps, and next-gen products https://latecheckout.agency/ The Vibe Marketer - Resources for people into vibe marketing/marketing with AI: https://www.thevibemarketer.com/ FIND ME ON SOCIAL X/Twitter: https://twitter.com/gregisenberg Instagram: https://instagram.com/gregisenberg/ LinkedIn: https://www.linkedin.com/in/gisenberg/
This Week In Startups is made possible by:Northwest Registered Agent - northwestregisteredagent.com/twistQuo - quo.com/TWiSTGusto - Gusto.com/twistPlaud - http://Plaud.ai/twistAthena - https://www.athena.com/jcalToday's show:How long until AI models can improve AI models? Once possible, recursive self-improvement by AI technology could accelerate — forever. Thus far, humans (and their coding agents) are still driving AI progress. But a recent project by AI developer extraordinare Andrej Karpathy, called ‘autoresearcher', is turning heads as it shows that it is possible — in certain contexts — to allow AI agents to run successive coding experiments to improve specific elements of LLM performance. Call it an early demonstration of the future.OpenClaw is exploding in China, while here in the United States, AI is polling somewhere underneath the basement. AI in the United States is about as popular as ICE, which could create a political issue for the technology in the coming elections.Next? Three demos. First, NetXD's Suresh Ramamurthi showed off how he has built OpenClaw functionality to move money, Rohan Arun showed off PhoneClaw automation on Android devices from an AR headset, and Eugene Stuckless gave us a taste of what Eir is building. Our takeaway? OpenClaw is still boring its way into our digital lives, one new skill or tool at a time!GUESTS:Suresh Ramamurthi: https://x.com/sureshr7Rohan Arun: https://x.com/Viewforge/Eugene Stuckless: https://x.com/eugene_eir_incTimestamps:0:00 — ‘Autoresearcher' and the future of AI improvements6:52 — Why people around the world are flocking to OpenClaw7:57 — Plaud - If your work depends on conversations — interviews, meetings, calls — you need a Plaud NotePin. You can check it out at Plaud.ai/twist and use code TWIST for 10% off!9:46 — Gusto - Check out the online payroll and benefits experts with software built specifically for small business and startups. Try Gusto today and get three months FREE at Gusto.com/twist.12:57 — The changing American social contract20:15 — Quo - Quo (formerly OpenPhone) gives you a clean, modern way to handle every customer call, text, and thread all in one place. Try it free at quo.com/TWIST23:50 — Why China is all-in on AI (and Europe isn't)26:26 — How to keep your job in the AI era28:05 — Northwest Registered Agent - Get more when you start your business with Northwest. In 10 clicks and 10 minutes, you can form your company and walk away with a real business identity — Learn more at www.northwestregisteredagent.com/twist29:38 — Athena - Get $2,000 off your first EA at https://www.athena.com/jcal34:42 — Demo: Suresh Ramamurthi of NetXD42:47 — Demo: Rohan Arun of PhoneClaw47:35 — Why bringing OpenClaw to your smartphone is what's next49:49 — Demo: Eugene Stuckless of Eir56:45 — How can we make smarter, more efficient agents?Subscribe to the TWiST500 newsletter: https://ticker.thisweekinstartups.comCheck out the TWIST500: https://www.twist500.comSubscribe to This Week in Startups on Apple: https://rb.gy/v19fcpFollow Lon:X: https://x.com/lonsFollow Alex:X: https://x.com/alexLinkedIn: https://www.linkedin.com/in/alexwilhelmFollow Jason:X: https://twitter.com/JasonLinkedIn: https://www.linkedin.com/in/jasoncalacanisGreat TWIST interviews: Will Guidara, Eoghan McCabe, Steve Huffman, Brian Chesky, Bob Moesta, Aaron Levie, Sophia Amoruso, Reid Hoffman, Frank Slootman, Billy McFarlandCheck out Jason's suite of newsletters: https://substack.com/@calacanisFollow TWiST:Twitter: https://twitter.com/TWiStartupsYouTube: https://www.youtube.com/thisweekinInstagram: https://www.instagram.com/thisweekinstartupsTikTok: https://www.tiktok.com/@thisweekinstartupsSubstack: https://twistartups.substack.com
The AI Breakdown: Daily Artificial Intelligence News and Discussions
Andrej Karpathy released autoresearch this weekend — a system where an AI agent runs experiments to improve a language model overnight, keeping what works and discarding what doesn't, while the human sleeps. The project itself is fascinating, but what's more interesting is what it shares with the Ralph Wiggum coding loop pattern and a broader shift happening across domains — from software to sales to finance — where the human's job becomes writing the strategy document and defining "better," and the agent does the iterating.Brought to you by:KPMG – Agentic AI is powering a potential $3 trillion productivity shift, and KPMG's new paper, Agentic AI Untangled, gives leaders a clear framework to decide whether to build, buy, or borrow—download it at www.kpmg.us/NavigateMercury - Modern banking for business and now personal accounts. Learn more at https://mercury.com/personal-bankingRackspace Technology - Build, test and scale intelligent workloads faster with Rackspace AI Launchpad - http://rackspace.com/ailaunchpadBlitzy - Want to accelerate enterprise software development velocity by 5x? https://blitzy.com/Optimizely Agents in Action - Join the virtual event (with me!) free March 4 - https://www.optimizely.com/insights/agents-in-action/AssemblyAI - The best way to build Voice AI apps - https://www.assemblyai.com/briefLandfallIP - AI to Navigate the Patent Process - https://landfallip.com/Robots & Pencils - Cloud-native AI solutions that power results https://robotsandpencils.com/The Agent Readiness Audit from Superintelligent - Go to https://besuper.ai/ to request your company's agent readiness score.The AI Daily Brief helps you understand the most important news and discussions in AI. Subscribe to the podcast version of The AI Daily Brief wherever you listen: https://pod.link/1680633614Our Newsletter is BACK: https://aidailybrief.beehiiv.com/Interested in sponsoring the show? sponsors@aidailybrief.ai