POPULARITY
When we first dicsussed the Summer of Simulative AI in 2024 we knew it would be a brief summer, but it has recently come back with a vengeance with SimGym in April and now Simile AI's $2B Series B, backed by GreenOaks and Index Ventures with prominent backers like Fei-Fei Li and Andrej Karpathy, running tens of millions of simulations for Fortune 100 clients like CVS and 85–99% accuracy vs human focus groups. Time to catch up on why this Second Summer of simulation is working!From creating Smallville, the landmark 2023 paper on Generative Agents that showed AI characters could remember, plan, socialize, and develop emergent behaviors, to now building foundation models of human behavior, Joon Sung Park is trying to answer a much bigger question: what if we could simulate the world before making decisions in it? In this episode, the Simile co-founder and CEO joins us to unpack the path from generative agents to digital twins, why today's frontier models still fail to capture how humans actually behave, and what it would take to eventually simulate all 8 billion people on Earth.We go deep on Simile's approach to modeling human behavior: long-form interviews, observational and transaction data, randomized controlled trials, population-level and individual-level models, and post-training on the causal mechanisms behind why people make decisions. Joon explains how his research created digital twins that reproduced human behavior and attitudes 85% as accurately as people reproduced their own responses, why models optimized to be rational can be bad simulations of irrational humans, and why understanding “social physics” may require changing model weights rather than simply prompting frontier LLMs.We also explore the much larger ambition behind simulation: testing products and policies before deploying them, finding counterintuitive paths toward desired outcomes, modeling emergent behavior across entire societies, and potentially tackling problems like climate change, democratic instability, and UBI. Joon reflects on scaling laws for simulation, the economics of data-center-scale simulated worlds, the connection to Thomas Schelling and psychohistory, why simulation is surprisingly similar to painting, and whether we might already be living in one.We discuss:* How Smallville and Generative Agents led to Simile* Why Joon's team asked: “What if we can just recreate the world that we live in?”* Why useful personal agents require deep models of their users* Memory architectures, Markdown files, and the limits of prompting* “Social physics” and behavioral foundation models* Why web data captures what people say more than what they actually do* Interviews, transactions, observational data, and randomized controlled trials* Why predicting the future matters less than understanding how to shape it* How Simile creates representative simulated populations* Simulation versus prediction and the connection to Foundation's psychohistory* How to evaluate simulations instead of simply stacking LLM hallucinations* Creating digital twins of 1,000 real people and reaching 85% behavioral accuracy* Why frontier models can struggle to reproduce real human behavior* Why good simulations need to reproduce human biases and mistakes* Post-training models on randomized controlled trials* Population-level versus individual-level simulation* Scaling laws for human simulation* The long-term ambition to simulate all 8 billion people on Earth* Whether simulations could help solve climate change or detect collapsing democracy* Thomas Schelling and the history of agent-based modeling* Why future simulations could require an entire data center* Multi-agent simulations and what happens when simulated people interact* Replacing expensive human panels with synthetic populations* Why market research is only the starting point for simulation* Why Joon sees simulation as surprisingly similar to painting* Using simulation to study questions like UBI* Whether we are already living in a simulation* Why AGI and simulation may be the twin technologies of advanced civilizationsJoon Sung Park* LinkedIn: https://www.linkedin.com/in/joonspark* X: https://x.com/joon_s_pk* Website: https://www.joonsungpark.com* Simile: https://www.simile.comTimestamps00:00:00 Introduction and Joon's Path from Art to AI00:01:46 Smallville, Generative Agents, and the Origins of Simulation00:05:03 “Let's Just Create a World” and the Future of Personal Agents00:09:53 Social Physics and Behavioral Foundation Models00:14:08 Prediction vs. Simulation: How Do You Shape the Future?00:16:59 How Simile Models Real People and Populations00:25:35 Evaluating Simulations, Digital Twins, and 85% Accuracy00:30:23 Post-Training Models to Reproduce Human Behavior00:40:04 Scaling Laws and Simulating 8 Billion People00:43:10 From Schelling to Society-Scale Agent Simulations00:46:13 The Cost and Economics of Simulating the World00:52:05 Real-World Use Cases, Synthetic Populations, and the Market00:57:27 The Future of Simulation, Painting, and UBI01:04:23 Are We Already Living in a Simulation?01:06:08 Building Simile and HiringTranscriptIntroduction: Joon Sung Park, Simile, and the Story So FarVibhu [00:00:00]: Today, we have Joon in the podcast. Excited to kick this one off. Very exciting company. I wanna kick off and ask you the question, talk us through the story of your life. How have you gotten here?Joon [00:00:13]: Yeah, for sure. I'm really excited to be here. A story of my life. So I was born in Korea, and I lived there for a good 11 years or so of my life, and then my family moved to Boston. So we moved when I was 11, and my parents were doctors, so they were going through their postdoctoral studies. My dad was a surgeon, so he was doing his sabbatical years at the Boston Children's Hospital. So I grew up there, not too close to tech. I was very much a music and artsy, painting kind of guy.Vibhu [00:00:49]: Painting.Joon [00:00:49]: Exactly. I got into painting a little bit later, in high school, but that's what I used to do. And then I grew up mostly in the East Coast after Korea. So I lived a good number of years in New Hampshire, and then I went to college in Pennsylvania. And I got into more of this tech scene, in college. So I was originally trained to be an artist. I thought that would be my professional career. So it wasn't a hobby. It was like, “Hey, let's make a living out of this.” And then gradually, I got really interested in this idea of, hey, the greatest artist often creates their own medium, and the best medium that we had available today was in computation. So I decided to go deeper into that, and one thing led to another, and we can go deeper into this, but I decided that research was something that I gradually got interested in, and here I am.Smallville, Generative Agents, and the 2023 Breakout PaperSwyx [00:01:46]: So there's a lot that you packed into the research components. You had one of the best papers of 2023, which was the generative agents paper, commonly known as the Smallville paper.Swyx [00:01:58]: Feel free to call back to anything else that you mentioned, but most people would have heard of you from this. Do you have any statistics on how many people have, like, read it? arXiv gives you something, right? Some stats.Joon [00:02:10]: Yeah, it's a good question. How many people have read it, I'm not sure.Joon [00:02:14]: I know we do keep track of citations, and they are going up quite fast.Swyx [00:02:23]: Yeah, Google Scholar has 7,200 citations.Vibhu [00:02:25]: I feel like it made a bigger hit than that, and it was a pretty instrumental paper. It got cited so many times.Swyx [00:02:34]: It is frequently the answer when people ask, “What is the best paper you've read recently?” It's this one.Vibhu [00:02:39]: I thought the memory component was pretty underrated. It was a very good early memory system, and one of the biggest papers.Foundation Models and the Search for Killer ApplicationsJoon [00:02:47]: Yeah, so maybe I can talk a little bit about how this particular paper came together. So when I got into research, it was back in 2020 when I started my PhD program at Stanford, and that was the year, when we were about to get GPT-3 to be available. So we already had GPT-2, and you could sense that there was this new class of models that was just becoming available in the market, and the team got very intrigued. And the general consensus was, “Well, is this model going to be useful for anything?” “It's really strange that these models are not trained to do any particular task.” But we decided to take a bet. So a large group of scholars at Stanford, and it was led by one of my co-founders, Percy Liang, and we came togetherSwyx [00:03:35]: Who coined foundation models.Joon [00:03:36]: Who coined the term foundation models. We wrote this paper, where that term came from called Opportunities and Risks of Foundation Models. And during that process, really the thing that I started to think deeply about was, here is a model that is fundamentally new in our ecosystem. The reason why this was new was it wasn't, again, trained to do anything in particular, but its premise was it could do anything and everything. It was like a stem cell, if you were to take a biology analogy. And I got really interested in this idea that, well, if we were to really think about what are the killer applications that this particular technology would enable, what would that be? Many of my colleagues were using this for simple classification, simple generations. Interesting that these models can do that, but from an interaction perspective, not that interesting. We've known how to do that for many decades. And what we came down to was these models are trained on this very broad data from the web, right? So these are human behavioral data. It's social media, Wikipedia, all these data. So if you poke at the right angle, then you could see human behavior that would just pop out that's quite realistic, and we've never seen that before.The Time Machine Game and Recreating the WorldJoon [00:04:45]: So that got us really interested. The exercise that we decided to do, with this particular group of colleagues, Michael Bernstein, Percy Liang, and myself, who ended up becoming my co-founder at Simile, we sat down and we played this game that we call the time machine game.Joon [00:05:03]: Imagine we were to get on a time machine and fast-forward 10 years and look back. What would have been the single application that will have mattered that would be the most interesting and inspiring? And when we thought, “Well, what if we can just recreate the world that we live in?” it's really hard to get more ambitious than that. Like, let's just create a world.Joon [00:05:24]: And that's where we started. And initially, we had this paper that was a precursor to the generative agents paper called Social Simulacra.Swyx [00:05:32]: Before you go further, were there other candidates for the most ambitious thing in the time machine exercise? What was number two or number three?Personal Agents, User Models, and Why Simulation Came FirstJoon [00:05:44]: There is a close second that we were considering, which ended up becoming more of these automation tools, especially the vision around really personalized agents that would do things for you.Swyx [00:05:59]: That's also happening.Joon [00:06:00]: It's also happening. But it was interesting for us, right, in that the reason why, we decided to go with the idea of simulation, one, I was a huge science fiction nerd, and this idea of creating simulation, I was personally really just fascinated. I loved the idea. It's really cool to see, like, a game town like this and just see these agents live in it. But at the same time, my bet was if you were to create a really amazing personal assistant out of this technology, what you need first is an amazing model of your users. So I told a model, “Hey, can you go buy late dinner for me?” And it orders Hawaiian pizza, and I do not like pineapples on my pizza. Then it totally failed. The way for it to not make that mistake is only by having a deep understanding of who I am. And I gave a very simple and dumb example here, but you can imagine how this core understanding of people is instrumental. This is how, if we have our family and closest friends, they have a good mental model of who we are. That's the basis of our social connection. So our bet also was this technology around simulation, creating accurate representation of people ought to precede the more complex agents that would automate the world that we live in. So that was the bet. But that was a very close second, and I'm still very much fascinated by it. I think there's a lot of interesting work that's going around. My hot take here, though, is I don't think we've seen a true personal assistant that's useful, in ways that meet the ambition of that particular line of work. I think there are early applications that are interesting, and if you talk to even ChatGPT nowadays or Claude, they know a lot about us. So a lot of the generation it's doing, I do think it's much more tailored, but I think the ambition is quite large in that field, and I don't think we quite have all the right ingredients just yet.Swyx [00:08:01]: So OpenClaw and these personal agents, what do you want to see from them that they don't currently have?Memory, Markdown, and the Limits of PromptingJoon [00:08:09]: I do think it's slowly getting there, but I do generally want them to have much deeper understanding of the person. Right now, you look at the models. OpenClaw, what it's leveraging is a Markdown file, and I think it's quite clever, right? So if you look at the generative agents paper, this was the same intuition that we had, where initially when we were creating the memory architecture for the generative agents, and, like, this is, like, back in 2022, so we didn't really quite have the idea of even agentive architecture or the term agent. But the intuition that we shared with some of the work that's coming out today was we initially thought, “Well, do we want to make the memory into, let's say, knowledge graph? Do we want to train a bespoke model?” All of these things. And what we decided to do was, “No. Just forget about all this.” These language models are quite good at modeling text and understanding and reasoning about text. So just put everything in a Markdown file or a text file. You're done. I thought that was quite interesting that we could do that, and there's a lot of strength in doing that. But also, there are limitations. It's the way you retrieve and make sense of data that's extremely large, it takes a lot of work. So I think that technology is getting better. I also do, however, think, there are certain things you just cannot shape just by prompting the model. So to some degree, you do need to touch the parameters of the model itself. So there is this work that I do think does need to happen, and it is happening. The question is, how far can we take it? How do we source data, and how do you also create an ecosystem where people are continuously feeding data to this model so it's learning about you?Vibhu [00:09:50]: What's the intuition between why you need to do it in the model?Social Physics and Behavior Foundation ModelsJoon [00:09:53]: My intuition behind the actual when do you train or even post-train a model versus just prompt a model is if the model has to learn the underlying physics of the world that it's operating in. So it has to learn new social physics. The places where it doesn't have to train are the places where it already has the physics. We trust the physics. It already has the base statistics, but it's just trying to react to an environment. Then I think you can just prompt your way into getting the actions out of it. I don't think the models that are out in the open have yet learned the complete mapping of social physics of humanity. This is one of the core theses of Simile, right? And one of the core reasons why that is the case is if you look at the data that the model was trained on, these models were trained on the web data, like, whatever was available on the web. And these are really interesting data sets, but they are fundamentally the self-exposed attitudinal data with some behavior data that's sprinkled around here and there. And it has yet to learn the really deep behavioral nature of people, not just what people say they do online, but what they do in real life. And this is one of what I would consider to be the dark knowledge of humanity that we haven't quite captured. And it's these data that would also need to get factored into the model creation.Vibhu [00:11:21]: You call it behavior foundation model.Vibhu [00:11:23]: There's a good one-liner here, but outside of that, what type of data do you need? What are you changing on the model level? How do you go about modeling, doing a behavior foundation model?The Three Data Buckets: Interviews, Behavior, and CausalityJoon [00:11:35]: We think about data in three buckets. So one bucket is interview data. It's quite interesting. Rich qualitative data is interesting. It's not behavioral, but we would literally ask people, “Hey, tell me the story of your life.”Vibhu [00:11:53]: It's just what we're doing here exactly.Joon [00:11:54]: The question that you all asked at the beginning of this interview literally is the question we also ask. And we ask our participants to go a little bit deeper, than how far I went. Maybe I can give more of my life story in lieu of this. But the reason why that data is interesting is by learning about this very long-tail information about people, you get a lot of texture around this model, like, this person as a model. So even understanding their childhood memory or even their trauma, their first love, these things, quite informative in ways that's really hard to predict. So that's one. Then there are two tranches of what I would consider to be the behavioral data. One kind of behavioral data is observational. So these might be like transaction data, or these might be data that you can get by scraping the web, right? So you can imagine why these data sets would be interesting, right, because they give you the base statistics of people's behavior.Joon [00:12:55]: But then there is the last category of data, that I personally think is perhaps the most important, which is the data that describes the causal mechanism, the whys of people. Some of this is covered by the interview data, the qualitative, because people talk about why they made certain decisions. But really, where you get to see the most behavioral aspect of this is in randomized controlled trials, like RCTs. Imagine you have the same setup, but you have a few different variables that you are trying to tweak. Can you get realistic human behavior out of it in ways where, imagine you had this particular option. Imagine you're even trying to choose whether you're going to drink coffee or not. The day you drink coffee versus the day you didn't drink coffee, does your behavior change? That's a data set that describes a causal mechanism. This is quite important in modeling people. The reason why this is important is oftentimes when people come to us, or not just to us, but the reason why people are interested in simulation isn't because they want to predict the future. If you're trying to win against the stock market, predicting the future is interesting.Prediction vs. Simulation: Shaping the FutureJoon [00:14:08]: But most people, most decision-makers, what they want to know is, how can we shape the future? It doesn't really help you to hear that your sales are going to tank in two quarters. They're just gonna say, “Wow, that sucks.” What they want to know is, well, what do we need to do now to avoid that future? That's the causal mechanism. And this is also very hard data to come by, right, because the world is our ground truth, but it happens once. So in a very controlled setup where everything is equal except for one variable, this kind of data set rarely happens. So this is a reason why this data set is both hard to come by and quite important if you're trying to model human behavior.Swyx [00:14:50]: So behavior, I think, is the hardest data set to acquire. What is out there? What is even possible? You're not going to know a lot of details about my life. I don't even have data for myself on my own health or habits, and I just don't log everything. So how can you have that data?Joon [00:15:14]: So we run a lot of randomized controlled trials.Swyx [00:15:17]: But you put people in the lab, they watch them sleep, or what?Joon [00:15:20]: We do care a lot about the consent process. People know that we invite them to be a member of this community to both share data and have themselves represented in different forms. But we bring a lot of people to the lab, or virtual lab, where we design experiments that would pose them real behavioral decisions. And often in these experimental setups, what makes the difference between what is attitudinal versus behavioral is whether the stake in your decision is real. That's ultimately what makes it behavioral. So in these setups, we are inspired by our colleagues in social sciences, psychology, and so forth. So when they run studies, the techniques they utilize is imagine there's an online store that you're inviting people to come by. Then whatever they purchase in this experiment, they actually get that item delivered. Like, these are the things that make the stakes real. So we run a lot of these experiments, and we also do partner with firms. Right now, we also have customers who are quite excited to at least give us a glimpse of the behaviors that their users exhibit so that we can get a little bit deeper understanding of how people behave in these different platforms.How Customers Use Simile: Populations, Queries, and ExperimentsVibhu [00:16:39]: I think on the customer side, they have a lot of data about their users, who has bought. They have the action data.Vibhu [00:16:47]: Can you walk us through an example of what someone comes to you for? What questions would they want solved? Do you customize a model for them? Do you have something off the shelf? What does that look like?Joon [00:16:59]: Today, when people leverage our models, it's often to better understand the population of their interest. So usually, the start of the relationship, we come together and hear about what population they want us to model, right? So it might be that if you're a CPG company that's selling to all of the US, then maybe it's fairly straightforward. You want to model the gen pop of the US. But at the same time, if there is a vertical or if there's a market that they're trying to go into, imagine, they want to better understand, let's say, people in their 20s and 30s living in California. That's a much more specific population. So we hear about this population, and we go recruit these people, with consent, and with incentives, and we collect some of their data and create a model of these people. Then what our product allows you to do is query them. So it can take as input a filter that is a description of the population that you want to talk to, just like the one I just mentioned, and an environment. The environment can literally be survey questions, behavioral experiments, It can be A/B testing. Oftentimes, the core use cases are things like concept testing, to start with. But also, people sometimes want to do focus groups or one of the fun use cases that we also serve is even modeling things like earnings calls for public companies.Joon [00:18:21]: So these are the use cases that we often start with.Swyx [00:18:23]: Concept testing, is that an established term? I've never heard of concept testing.Concept Testing, Gallup, and PoliticsJoon [00:18:27]: Yeah. So it has to do with they have, let's say, different messaging, different products, different ideas.Swyx [00:18:32]: It's like a marketing exercise.Swyx [00:18:33]: Okay, got it. Got it. Politics?Joon [00:18:36]: We do, have a strategic partnership with Gallup, and of course, Gallup is deep into policy space and so forth. Right now, we have not worked deeply with politics, like that area just yet, however.Swyx [00:18:49]: I'm curious if there is demand or if they really would have different needs that somehow fundamentally don't mix with your existing, users or people.Joon [00:19:00]: I think there's certainly demand.Joon [00:19:02]: But we are very much mindful of how this technology gets adopted and the societal impact that we'll end up having with this technology. And I do see politics as an area where a company has to be particularly thoughtful about the way they operate and make impact. So this is where we also want to make sure that we form enough of guardrail and perspective on how to leverage this technology before we go on to serve markets like the politics.Swyx [00:19:29]: I'll give people an example. one of my favorite shows is The West Wing. I don't know if people have watched.Swyx [00:19:34]: One of the key storylines is, like, the president has, multiple sclerosis, but they haven't. they need to figure out how to disclose it. So they run a poll with a fake governor and ask people to respond on the poll,Counterfactuals, Polling, and When Simulation Is UsefulSwyx [00:19:47]: They try to make decisions based on the results of that poll on, like, how well they'll be received, like where, how should we play this?Swyx [00:19:54]: And I'm like, well, I think those counterfactual things, I would use a simulation for this if I could trust it.Joon [00:20:01]: For sure.Joon [00:20:02]: In that show, how'd it go?Swyx [00:20:04]: In that show, it was, like a foregone conclusion. They were like, “We know it's bad. We just don't know how bad.” And then the poll came back. It was like, “It's really bad.” And then they just did it anyway.Joon [00:20:14]: Part of it is to show, right? So you're, you're looking at the ideaSwyx [00:20:17]: Maximizing drama.Joon [00:20:18]: How bad could it be? Oh, it's horrible.Swyx [00:20:20]: And to some extent, I think that is part of the trick of the, or the challenge or with being a customer of yours, which is that if I know it's. if I roughly know and can intuitSwyx [00:20:35]: What the effect is going to be, do I need you? What sensitivity of it, of effect do I need in order to make a decision, right? So for example, if I, my approval rating is 50%Swyx [00:20:48]: And I, they have this negative piece, news item comes out, and it drops to 30.Swyx [00:20:52]: If it drops to 20, if it drops to 40, do I care? No. It, I know it drops. It's negative. So when do I care about simulations?Joon [00:21:01]: You do something that's clearly bad, that's not popular, and people don't like you, like, yeah, it's likeSwyx [00:21:05]: You don't need a simulation.Joon [00:21:07]: Yeah. Well, so there are a couple of things. one is, there are use cases where, like every day, developers, designers, policymakers, marketers, every single day, they create assets. They create new products. And turns out, it's many of the decisions in hindsight is obvious. Yes, of course this is bad, but we still run those studies because understanding the magnitude and understanding how acute something is quite difficult, even if, we feel like, of course, like this makes sense. this is the reason why we make so many mistakes. Like, every time somebody goes online and say something that has huge backlash, you look at that and like, “What an idiot.” However, it's tough. That's one. There's also another aspect here, which is, again, this is the reason why simulation is different from prediction. In simulation, in the ideal case scenario. So what simulation is trying to show is it's trying to show each step of the way or each step that we need to take to get to a certain outcome, right? So in the most advanced simulations, sometimes the next step that we're suggesting might be quite counterintuitive. The analogy that I sometimes give, and I ground it in a more realistic example, but, I, as I mentioned, I'm a huge fan of science fiction, and I don't know how, many of the audience members have read, like, things like the Foundation series by Asimov.Simulation as a Path, Not Just a PredictionSwyx [00:22:37]: Oh, yeah. We've mentioned psychohistory a number of times.Joon [00:22:39]: Okay, fantastic. So I might be, talking to the right crew. If you read Foundation series, literally the first act is there's a group of scientists who have found out that, “Oh, our galactic empire is going to collapse, and we're going to have 30,000 years of unrest.” And they run psychohistory, the simulator that tries to teach them, “Okay, how can we keep this unrest to a 1,000 years?” And they plan this out, and the first step of that plan is to get the scientists who say, “Okay, this is coming,” exiled into this random place in this, galax- galaxy.Swyx [00:23:18]: Terminus.Joon [00:23:19]: Exactly. And that's so counterintuitive. Like, what a strange move that you literally sent the group of scientists who was raising voice around this potential collapse of galactic empire into nowhere. How is that the right first move? Well, it turns out in this particular simulation, that was the move.Joon [00:23:40]: It's these things, right? And the reason why these reasoning is possible is because you're showing the step function or each step that results in a particular outcome. So really what simulation allows you to do in its highest form is you give it not a problem or question, like what would people answer to the survey? That's not what we do. What we tell it is, “Here is a goal that we have. In the context of foundation, we want to keep the unrest to a 1,000 years. What is the path that we need to take now to get to that particular future?” And that's what simulation allows you to do. Now, translating that into real market, imagine you're a automobile company and you're about to release a, EV, and you're trying to understand, well, how do we market EV, to make sure that our stock price goes up? But what if the answer comes down that, well, you can market your EV in XYZ way, but that might change people's perception around the cars that's not EV and make your overall sales to go down. Not very intuitive, especially all you're trying to optimize is EV salesss, and that's the only thing that you're tracking, then that might result in a completely wrong solution, or at least different solution than what you would have expected, whether it's right or wrong.Joon [00:24:57]: That's the power of simulation.Swyx [00:24:58]: For listeners, we covered a similar topic with Mikhail Parakhin from Shopify, where they are working on SimGym. I don't know if he ever talked to you about it. it's very similar.Joon [00:25:07]: ISwyx [00:25:07]: The goal is increased conversion, but then the journey is very unusual.Joon [00:25:12]: Journey is unusual.Swyx [00:25:12]: Yeah. The-- He's trying to look for interventions on a shopping trajectory, which is similar to what you're saying. Like, it's not about the attitudinal, is your word for it.Swyx [00:25:24]: It's about behavior.Joon [00:25:25]: It's about behavior.Swyx [00:25:25]: And that's exactly the difference, right? It's, like, not about the near-term direction about-- but it's more about, like, how do you affect multiple turns of interactions.Vibhu [00:25:35]: You had a good quote at the start about this as well. It's not about people wanting to know the outcome. It's about how they can change it, change the way to get there, something like that. But I wanna take it back to how do we know this is grounded? LikeGrounding and Evaluating Digital TwinsVibhu [00:25:47]: How do you run evals? How do you test that simulations come through? if I was to do the same thing that you described with, say, your favorite LLM, Opus, GPT-5.6, have some agent to map out these thingsVibhu [00:26:02]: How different are the answers we would get if I give it the same goal, the same objective, make a decent system? You're saying that you need to change the model weight. You have your own solution to this. But how far off are we, and how do you check if it's grounded? you have some interesting stuff on your site that points to how you run real evals, but if you could take us through that side. I think that's one of the big concerns that people have. They're like, “LLMs hallucinate.”Vibhu [00:26:27]: “You're just hallucinating layer after layer,” right?Joon [00:26:30]: The way we do this, and this is the paper that we worked on after the generative agents paper that really became the, at least for Simile and also the field of simulation and synthetic panels, really became the foundation. Yeah, this is the paper. the paper is called Generative Agent Simulations of 1000 People. Here's what we've done. For this paper, we brought 1,000 people that's representatively sampled from the US to a virtual lab. And what we have done was we spent two hours collecting fairly wide-ranging data. In this particular study, we focused a lot on this interview data, that was, whose script was taken from this project called American Voices Project. And then we would also pair that with a lot of behavior data and so forth, whatever we can collect within two hours. And then we would send these people away for a couple of weeks. And during that time, I would use this data to create their digital twins. And I would bring the humans, participants back after 2 weeks and have them complete a battery of surveys, experiments, behavior studies. So we have the list here, which included things like behavioral economics games. We would run literally, like, Big Five personality test, General Social Survey. We would also go ahead and run the randomized controlled trials that were published on PNAS. And we would have their digital twins predict how the source individuals would have acted in these studies and surveys. And this is where we could replicate people's behaviors and attitudes 85 percent as accurately as people would replicate their own. So that was the first really paper that gave this validated results that we can model individuals in an accurate way. And what we ended up finding now, of course, in AI space, so this paper came out at the end of 2024. AI space, a year and a half, 2 years, that's a lifetime.85% Accuracy and Why Frontier Models Miss Human BehaviorSwyx [00:28:24]: Yeah. Just, for listeners who are not seeing the YouTube, I just wanna say, like, the headline figure is 85 percent accuracy, like, which is a big improvement over all the otherSwyx [00:28:34]: Methods that you showed.Joon [00:28:36]: But the part that was particularly striking to us, especially as we improved this technology even further, was the generative AI models like ChatGPT, Claude that's coming out, it does give you the right foundation. However, what they do not consider is the true attitudinal and behavioral aspect of people, especially in the population that you care about. So what these models are really good at today is they're trying to become the super rational, objective machines, right? So you go get their data from places like Mercor, Scale. You talk to professional programmers, scientists to create model that's amazing at reasoning. That's what they do. Simile doesn't care about any of this. The models that we're talking about here, what we're trying to create are models that are as dumb as I am, right? So if I make some mistakes, the model has to make the same mistake.Swyx [00:29:34]: Oh, that's very hard.Joon [00:29:35]: That's very hard.Swyx [00:29:36]: You're solving Murphy's paradox.Joon [00:29:37]: That's exactly. And this is a completely different data and training objective. This is also where we see quite a bit of discrepancy in the performance in human behavior prediction between the frontier models, Simile's model, and the models being created in this space, where in some cases, the model performance of frontier models go all the way down to 20, 30 percent, especially if you go into that more niche population on topics that our customers would care about. On more gen pop, it might be around 50 to 60 percent. So it's not very robust. Like, you wouldn't want to make your decision off of these and these findings. If you can bring that up to 85 percent, that is ultimately what people end up getting very excited about.Swyx [00:30:20]: Yeah. Do we wanna keep going on the paper, routes?Joon [00:30:23]: Yeah, for sure. So the last one, was an interesting one. So this, paper was the follow-up paper that we had, to the 1000 agents paper, where the idea was now can we augment the models even further and post-train a model based on a lot of randomized controlled trials? So this was an interesting one. The data is always the most interesting part of modeling in many ways. The data that we got here was there's this, there's this platform called Open Science Framework. So some, the audience might be familiar with this. And there has been, especially in the social sciences over the past 5 years or so, there has been this concern around replicability of studies. And so it was a bit of a crisis, the scientists acknowledged, where we rerun the study and we don't see the same finding.Post-Training on RCTs and Replication StudiesVibhu [00:31:12]: Oof.Joon [00:31:12]: It's tough. And the reason why it's there-- that was often the case was there's this survival bias where the papers that get published often need to maintain what we call the value of less than 0.05 in the experiments that we ran. That suggests that only-- there's only 5% chance that the results that we saw is false positive. But the tricky part was all the papers that were not published, and there's still a 5% chance that whatever we publish is totally just randomly generated. Like, there's a 5% chance that, hey, this effect is not real, but it just happened to be real because of the sampling bias. So because of that, what scientists started to do was they started to register their studies. So before running an experiment, they would go to this platform and say, “Here is the data. Here is the population that we're collecting, and here's the hypotheses.” And they would just say, “Here is our hypothesis.” Like, “This is what we believe.” And you cannot retroactively change those hypotheses. This is what gives us more scientific statistical confidence that whatever effect that you ended up seeing is true. So that ended up creating this really interesting platform where there's one platform that has now contains tens of thousands of real-world experiments and hypotheses. And a lot of these are really high-quality, like, professionally designed behavior studies and random- randomized controlled trials. So we got the data and the studies from this platform and used that to make a point. And this particular, model is not, something that we're serving commercially because this was a part of the open science. But this particular data set, helped us make a point that by collecting a lot of these randomized controlled trials, that are really well-designed, we can make significant improvement in model's capability to predict human behaviors. So that's what this paper was about.Vibhu [00:33:10]: Is this stuff done on a individual level? Like, do I need to tune the model per individual, per company? Is there foundation model changes and then some slight post-training? Anything you can share there?Population-Level vs. Individual-Level ModelsJoon [00:33:21]: So this particular model was trained. the data we had at the level of individuals, but this particular model was trained. We experimented with both. And this is what we end up doing at Simile too. We always train 2, distinct model. One is what we call the population-level model. The other is what we call the individual-level model. And both take very similar input, which is the description of a subpopulation or individual and a stimuli. In this particular work, we've done the same. Here, the results that we are reporting are much more geared towards individuals because we do think that is a harder task in many ways, but that's what we have done.Vibhu [00:34:02]: You seen anything on the questions that humans can solve that models can't solve? So likeHuman Biases, Mundane Choices, and What Models MissVibhu [00:34:09]: Currently, it's, I live 5 minutes walk away from a car wash. It's a 10-minute drive. Should I walk or drive?Joon [00:34:16]: Huh.Vibhu [00:34:16]: The model will say, “Oh, walk to the car wash.” And, you don't have your car.Vibhu [00:34:20]: Is anything like this a problem in simulation? You would assume, like, very simple for human to think about, but if the model is saying you should walk to the car wash, anything here?Joon [00:34:32]: It's less, what can we solve, but I think it's more about what biases or mistakes do people make that models miss. Like, imagine that you are, like the. When I was still at Stanford, I lived in Palo Alto. So it's about, I would say, 40-minute walk from the campus. You ask the model, “Okay, let's go home. What can I, what can I do?” It would likely call an Uber or, give me, the bus time. But for the longest time, I really liked walking back. And the reason why I wanted to do that was not for efficiency. It really helped me think. And I like to walk for, half an hour or 40 minutes or so a day, where I just get to, just think about ideas, research, just get lost in my thoughts. That's very human activity. Unless the model has seen that and understands the importance of that activity, it would miss these kinds of features. So that I think, is fundamentally what we're trying to model. Like, what is fundamentally human might not be the most efficient thing to do, might not be the right thing to do, but things that make us who we are.Swyx [00:35:43]: I'm curious if, there are some data sets that you really want that would materially help you. One version of this may be interesting, which is more valuable to you to acquire as a data set, all of LinkedIn, all of Twitter, all of Facebook?What Data Matters: Social Media, Transactions, and FacebookJoon [00:35:57]: It's a little bit hard to rank, in part because, there's, there's this product saying where no feedback is wrong because it teaches you something about your users. Doesn't matter what feedback.Joon [00:36:11]: I think it's a little bit like that.Swyx [00:36:12]: So just whatever is bigger.Vibhu [00:36:13]: What about a different domain? Say it was. What about all of Amazon data?Joon [00:36:17]: Oh, yeah.Vibhu [00:36:18]: Shopping data, right?Joon [00:36:18]: Shopping data. So Amazon data is interesting in that it's very much behavioral, although, like, what people do on social media, you could squint and say that is also behavioral. But the transaction data is always interesting. It is also most commonly available, however.Joon [00:36:33]: If we were to look at purely social media, like if you really, if I were, if I had to really pick, Facebook likely is interesting because I do think it is most a default version of people. Because you go to LinkedIn, it's very much professional environment. So people put up their, they have their guards up, right? And that still is interesting because that is true human attitude and behavior, but it is not your base state. you go to Twitter- Twitter, people have their own crazy personas, or depending on who you are. Like, my Twitter profile and, persona is very much, initially was I was very much an academic. “Hey, I'm here to share my studies.” Now, I share, things that's related to Simile. But Facebook is one of those more private space where people just connect with their friends. In that way, I do think it shows you a little bit more about who that person is. So if I had to pick, I'd likely pick, Facebook.Swyx [00:37:30]: Yeah. And you're interested in, like, the whole person and their background and philosophy. I, is it too clinical or too machine learning-oriented to just say this is just ways to inject variance and biases? The broad question, is, like, is this any better than a randomized, like, combinatorial explosion version? So we have a link to the TencentBillion Personas, Synthetic Demographics, and Bespoke DataSwyx [00:37:54]: Billion persona paper, where they did not do any of the groundwork that you are doing.Swyx [00:37:59]: They just did like a cross matrix of here's all the professions in the world, here's all the people, possible backgrounds in the world, do a dot product across all of them, and that's it. That's your prompt for a billion people.Swyx [00:38:12]: This will do something. I don't know if it'll do what you do, but it gets you some way, some percent of the way there.Joon [00:38:18]: So this was an interesting paper. Like, what I admired about this paper when it came out was the scale. And you do gradually want to be able to simulate really large societies and interactions. So the scale is definitely admirable. it is relying heavily on the known statistics that went into training the model. So to the extent that you believe that statistics is correct, this is not a bad way to go about this. But the thesis here, and this is something that we also have seen in the market, like if this works, then we have solved simulation.Joon [00:38:54]: It,Swyx [00:38:55]: Because I survey, like, okay, 5% of the US population is in construction.Swyx [00:39:01]: The other 5% is in medicine, whatever, right? And then you just keep going down the list, and then you do the other side. 5% has, like, the big 5 personalitySwyx [00:39:08]: Of, like, neurotic or whatever. That's it.Joon [00:39:11]: That's it. So if you believe that the underlying data set and the platform that we're leveraging has all the right statistics, then this will have solved it. you're at that point merely retrieving the knowledge that is already embedded in the model, in the model parameters. That's not, unfortunately, what we see, where there is such detailed and also niche knowledge about people that if you just take one example, it might feel very mundane, but it's quite rich when you put together, that you do need to do a lot of bespoke data collection to better understand people. And this is also, I think what makes this particular, job fun, which you want to deeply understand people, and the process of deeply understanding them requires a lot of attention to the details. And you do need to pay attention to and pay respect to the daily lives that people lead.Scaling Simulation: From Thousands to SocietiesVibhu [00:40:04]: I wanna talk about scaling simulation.Vibhu [00:40:07]: So what can't we simulate, what can we simulate, and how does scaling affect this? So how big are the models? What if we go from, 8B, like, couple 100 billionVibhu [00:40:18]: Like billion000 parameters, billion000? Do we get scaling? Any interesting emergence? Like, at a certain scale, at a certain amount of training, you uncover anything unusual and any learnings from that?Joon [00:40:31]: What we are seeing is at Simile, so we do post-train our own model. The thing that we're seeing is the early glimpse of scaling law in simulations. The more data about humans and more compute you ingest, you start to get predictive and predictable gains of the model performance in simulating it, simulating people.Vibhu [00:40:51]: Ooh. We need a scaling law curve.Joon [00:40:52]: It's scaling law. Whenever you find it's a beautiful thing. And we're starting to see the glimpse of it, which is quite exciting. But if you talk about the ambition of simulation as a whole, it's not merely about building a model. It's about building a model, then creating the agents that become the individuals in a much larger ecosystem. So they're creating this multi-agent simulation. Down the line, you want these multi-agent simulation to also live in a very rich environment, right? What we are really trying to get to at that point is, hey, can we create. All right, let's do a time machine game again, and 5 years, 10 years into the future, can we create a simulation of 8 billion people living on Earth? I think that's quite interesting. And that really is the vision. And once you get to that state, the questions that you can help answer for the society also start to change from my perspective. The answers are fundamentally about emergence of the emergent behavior of society and large groups of people.Joon [00:41:53]: So the questions that I get excited by, and maybe this is a stodgy- a bit. I have my, academic side of me.Joon [00:42:01]: And for me, it's questions like, can we help solve climate change? If you look at climate change as a problem space, this is what we, like social scientists would often call it the wicked problems, problem where you have many actors with competing incentives for trying to make a very complex decision and coordinating that coordination decision. Very difficult to really solve in real life, which is also the reason why we couldn't solve it. Can simulation help us solve that? Another one is, can we understand the signals for collapsing democracy, or can we understand or can we uncover the origin story of the monetary system? These are societal questions that we never really had a good way of answering. If we can create simulations of our society, you have to believe that these are the problems that we can solve. So that's really the ambition of this field. And, I also think, yes, I think there's a Nobel Prize to be won there, which wouldn't be surprising. And I think there's some amazing societal impact that we can have to help people make better decisions.Climate Change, Democracy, and Societal SimulationSwyx [00:43:04]: Nobel Prize in economics?Joon [00:43:06]: In economics.Swyx [00:43:06]: Oh, I see. I see. Rooting for you to write that paper.Joon [00:43:10]: One of these days. But, one of the scholars that I was deeply inspired by, When I was coming into the space of simulation, is this scholar, named Thomas Schelling.Schelling, Agent-Based Models, and the Nobel PrizeSwyx [00:43:23]: Schelling point?Joon [00:43:24]: So the canonical example of the work that he's done was he was one of the creators of agent-based modeling. So this was, like, in the 1970s and 80s. It's very early days, but this was truly one of the first exemplars of simulations. And one of the canonical model from that time, and of course many of these simulations are trying to tackle the societal problems that's most relevant for their era, it was called the model of segregation. So racial segregation was a big topic, that, we cared about. And what they've done was they created this grid world where they had red dots and blue dots. And these dots were, back in the day, like, they were the agents, and they had a simple rule that governed their behavior. If certain percentage of your neighbors are of different color and if that goes above certain threshold, then you move to a new location at random.Joon [00:44:21]: One of the striking finding of this paper or this agent-based model was for the longest time, people thought the segregation within society was caused by explicit and overt racism.Joon [00:44:34]: But if you look at this model, people's preference towards living with people of the same color, that preference can be very minute.Joon [00:44:42]: But the very small difference causes the society to segregate completely over time. This was very counterintuitive for a lot of people. And this particular work ended up informing housing policies. Mixed income housing, got really inspired by this work. And Thomas Schelling ends up winning the Nobel Prize for having laid the groundwork for very early versions of simulations. The opportunity that I do see here in the more scientific terms, is agent-based models for the longest, had impact in the 1980s, 90s, to some extent, early 2000s, but it has now gotten forgotten by the community a little bit. Because as you can imagine, red dots and blue dots is not really a rich description of people.Joon [00:45:31]: But with the emergence of things like generative AI and, in particular, generative agents, we do have an opportunity to create these agent-based models that are high fidelity enough to help us make really complex decisions. And that's the opportunity that I see. If that truly works, then yes, that is the work that will result in a Nobel Prize.Swyx [00:45:53]: Yeah. For what it's worth, and I grew up in Singapore. 80% of Singapore is in public housing, and public housing has, enforced racial quotas for exactly that reason, which is very interesting. okay, so we talk about scaling, we talk about all these, the agent possible applications.Cost, Reuse, and the Economics of SimulationSwyx [00:46:13]: I'm scared about the cost. if you even-- let's just keep it to the US, about 8 billion people.Swyx [00:46:21]: But, how much does it cost to model so many hundreds of millions of people?Joon [00:46:26]: Oftentimes today, we don't start at that scale, this stage of the, of industry and simulation as technology. But we can get our users extremely rich and meaningful insights even by modeling thousands, tens of thousands of people. And today what we do is every week we are collecting data on the scale of tens of thousands people's data, and we have panel partnerships that gets us to tens of millions of people globally. So that's what we do today.Swyx [00:46:55]: And just as a side note once you've collected one person for one studySwyx [00:46:59]: Can you reuse that same person for all the subsequent studies?Joon [00:47:03]: That's exactly right.Swyx [00:47:03]: Okay.Joon [00:47:04]: The beauty of this model and these agents is the fact that they are domain-agnostic.Joon [00:47:08]: That what you're really trying to understand is what is the fundamental nature of these people? What's their social physics? And there are a lot of, a lot of, people that does change over time. Like, even, like, even things like, how many times have you gone have you been to, like, CVS the past week? that will change. But there's so many traits about people that are also known to never change. Like, your risk tolerance doesn't really change over time. It's very consistent. So it's these things that we're trying to learn. But the scale we are operating is right now hundreds or, tens of thousands to hundreds of thousands. And in many of the core use cases that we are deployed in, and this is more than enough population, to cover those. Really, at that point, what you care about is less the number of people, but more do you have the right subpopulation of interest covered? And this is also the reason why people want a larger sample. It's not because they want, stronger statistical guarantees. It's more that can they filter down to any population of their interest. However, you can also imagine in 10 years, if we truly believe that the compute is going to scale, that we'll have much more availability for compute, and our ambition for simulation is also going to scale accordingly, there's definitely a reason for us to create an entire data center worth of simulations.Joon [00:48:35]: Or in my hunch here is I do think in the next some number of years, we will start creating simulations that will cost as much as training a foundation model. But perhaps it's going to be so valuable to the society that it would be a no-brainer. Right now, even today, like, we are training bunch of new foundation model just so we can say we trained one and we spent tens of millions. But if we can create a simulation at the level of society that would solve climate change, I would run that today. I would raise the money right now just to run that.Multi-Agent Simulation and Social InfluenceSwyx [00:49:10]: Amazing. the follow-up question is, does it also compound if you let the simulations talk to each other?Swyx [00:49:18]: Or do they already do that today? They don't, right, as far as I understand?Joon [00:49:22]: It depends on what simulation you're trying to run.Joon [00:49:24]: In the multi-agent simulation setup, the agents do talk to each other.Swyx [00:49:28]: Right, which is exactly Smallville, right?Joon [00:49:29]: That's right.Swyx [00:49:30]: But a lot of times, for example, in commerce, you're just by yourself, so there's no point talking. which is way cheaper.Vibhu [00:49:37]: But they use all these levels, right? Like, you decide what you will buy based on what other people around you buy and talk about, right?Swyx [00:49:43]: It depends.Vibhu [00:49:44]: It depends.Swyx [00:49:45]: Again, I'm, I'm coming at this from a cost point of view. I'm like, “Oh my God.” LikeVibhu [00:49:48]: I thinkSwyx [00:49:49]: If there is, like, some combinatorial thing of, like, thousands of people talking to thousands of people, then that one million X's might cost.Vibhu [00:49:56]: I have a very different view as the cost point aside. Like, running these studies in reality is a lot more expensive, right? Running any study like this is you gotta have people do it, you gotta sign people up. It's very expensive and sometimes, like, not feasible to run the study.Vibhu [00:50:14]: But the outcome or the decisions you make are very expensive on them, right? So spend X million on something that, the overall process costs 100 million might as well, right? There's, there's a lot of value to be had there. It's a small cost, but I'm excited on the cost side.Joon [00:50:33]: To some extent, and when you deploy technology, you often want to deploy in a way where you can replace existing budget or you can make things more efficient, and that is the best way to deploy. However, the way you capture the long-term value of the technology is making the argument that, no, it's the upside, that by making this better decision using simulation, you have saved yourself or made yourself hundreds of millions or even billions of dollars, and that's a case to be made.Vibhu [00:51:06]: Random tangent question. So if you're doing a lot of inference, a lot of model multi-agent stuff, are you at the point where it makes sense to, train a model that' very sparse? You're expecting to do multi-million dollar runs. Are you thinking about this in model architecture standpoint or inference efficiency, or, you're still at the research phase of it works, we're not super there yet?Joon [00:51:34]: Efficiency, we do think quite a bit about. this is technology that is deployed now in some of the largest enterprise companies in the world, and we do process significant number of queries, that are trying to, simulate the populations in the world. So efficiency is a consistent thing. we don't want to over-optimize too early, so I wouldn't say, like, this is the higher bid Right now, but this is definitely something that we think pretty carefully about.Swyx [00:52:05]: Yeah. Are there other case studies? So we, you talked about CVS, talked about Gallup, Deloitte, Wealthfront.Efficiency, Enterprise Use, and Real-World Case StudiesJoon [00:52:12]: Wealthfront is an interesting one, because one of the things they were trying to do, they were one of the first customers that wanted to do product testing that goes beyond just asking people what they think about, let's say, behavior experiments and so forth. So there, really what we had to do was reason about multimodal input, so images, but also you can also imagine, like, these agents traversing through Figma mockups or websites. So some of the things that our agents can also do is it can be given a domain, like, or, like, a website URL and go use it for a while. It's these things. And Wealthfront was one of the first, customers, that was very excited about this possibility.Vibhu [00:52:53]: What have people been asking? Like, is there any demand that we have not covered? Like, UI testing, right?Vibhu [00:52:59]: I wanna try a new. I wanna ship a new feature, test the UI, simulate how people will do it. Any interesting things that you're seeing demand for?Product Testing, Websites, and Synthetic PanelsJoon [00:53:08]: Today, a lot of the demand does come from like, the places where people have historically used human panels, we can now replace with agents, and these synthetic populations. And this is not replacing human panel. in many ways, the simulation that Simile is building is grounded. So the way that I think about this is we are trying to represent humanity at scale. And in that way, the use cases are what we would expect, but it's the scale of deployment that surprises me.Joon [00:53:44]: Turns out there are so many decisions that people make every day in these organizations, groups, and we want to be able to say, “We listen to people. We have consulted our users.” But in reality, that is rarely the case because getting to people and asking them many questions, it's difficult. It's both costly, time-consuming, but most importantly, people are just not available. If I had to answer 1000 survey questions for this one particular, vendor, even if I wanted to do that, like, I would never do it. And that's very much the case. What simulation can do is ensure that the voices of people are always represented in rooms where the decisions for them is made, right? So all the stakeholders of this particular product launch, ideally they're consulted. That's what this technology really is trying to enable.Market Size, TAM, and Human Decision-MakingSwyx [00:54:39]: In my mind, that means it skews towards more consumer focus, right? Like, anything with a wide enough customer base where you do benefit from the diversity that you represent. What are some rough statistics, just for people who are not familiar with this market in general, what's the market size that. I'm sure you have some, like, rough numbers. market size is, like, a vague questionSwyx [00:55:01]: But, like, how much do people spend?Joon [00:55:03]: So market research is a $100 billion industry.Joon [00:55:06]: But the thing about simulation is not a tool for market research. Simulation is a tool for human decision-making. So the question around what is a TAM here is quite tricky, right? Because it's easy to say, “Well, market research TAM is roughly 100 million or 100 billion.” so is it a TAM? And not really, right? Because in many ways, you're trying to inform all human decision-making. You're trying to inform every decision that are made about humans for humans. What is a TAM for that? It's really unclear. And I'll be honest. Like, I have a scientific background, I have a research background, so I didn't come into the field calculating, oh, what is the TAM for human decision-making? But I just had to assume, well, if we can inform every decision that is made about human for human, that has to be big.Swyx [00:55:58]: Some- something valuable.Joon [00:55:59]: Exactly.Swyx [00:55:59]: To some extent, you are a unicorn founder now, and you have to care as a CEO. But, like, I do think, like, yeah, when you go into these boardrooms with people that you're quoting millions of dollars of contracts for, like, you have to say, “Well, here's what you spend on humans-”Swyx [00:56:15]: “. And here's what we save you, and it's 85% similar.”Joon [00:56:19]: And certainly, the value case, is something that we care deeply about. Like, what is the value that we provide to the users and the decision-makers? But this is also where, like, as a founder, I think valuation only tells one very superficial aspect of the story, and I try not to think too much about valuation, in general, because that's not what also motivates a team or certainly doesn't. I'm, I-- Again, the interesting thing about researchers is we are happy living in academia, getting paid next to. we get paid okay. we don't get paid that much, as a researcher here in academia, but it's the impact and it's the, it's the value that we can provide to the individuals and the society that really drives us. And in that way, ultimately what drives us is the impact. Does the simulation we provide have a real impact in people's decision-making in ways that progresses our society forward? If the answer is yes, then yes. that has to be great business, and we see that in numbers, and we do care deeply about that upside story, but that's the heart of it.Where Simulation Goes NextVibhu [00:57:27]: Do you have any timeline predictions? So we talked about scaling laws of simulations.Vibhu [00:57:33]: You brought up, okay, maybe one day we can simulate how to solve climate change.Vibhu [00:57:38]: Where are we now?Vibhu [00:57:40]: If that's not the end state, what is an end state, and what does progress look like?Joon [00:57:45]: So what I sometimes tell people is simulation as industry, it feels a lot like where GPT-3.5, GPT-4 was, for the AGI saga, which is we have now technology that is powerful enough to do real damage on the verticals that we are tackling. At the same time, there's a lot of progress that is yet to come. And that's, I think, where this is. So the way I see it, I do think there will continue to be breakthroughs both in data, in algorithms, and there will be much more aggressive scaling that will also happen over the next few years. But I think that's roughly where we are.Swyx [00:58:27]: I think that was about the ro
Ein Jahr nach dem letzten Gespräch zieht Michél gemeinsam mit Kai Panitzki Bilanz: Was hat sich bei KI, Robotik und Start-ups wirklich verändert? Im Fokus stehen Physical AI, Europas Chancen im internationalen Wettbewerb und die spannendsten Entwicklungen aus Construction Tech und Venture Capital. Außerdem sprechen die beiden darüber, warum Robotik gerade jetzt zum Milliardenmarkt wird und welche Innovationen bereits heute auf Baustellen Einzug halten. _________________________________________________ Die BLACK BOX auf der BAU 2027 ist das neue Innovationsformat der Messe München und DIGITALWERK – mit Livedemos, Diskussionen, Networking und digitalem Content. Das geht nur mit Unternehmen, die ihre Innovationen nicht nur ausstellen, sondern sichtbar machen. Mehr Infos und Kontaktmöglichkeiten gibt's unter https://www.digitalwerk.io/dw-events/blackbox _________________________________________________ VESTIGAS ist die All-in-One-Lösung für die digitale Baulieferkette – vom digitalen Lieferschein über KI-gestützte Rechnungsprüfung bis zum smarten Rechnungsworkflow. Weniger Papierkram, rechtssichere Prozesse und mehr Zeit für die wirklich wichtigen Aufgaben auf der Baustelle und im Büro. Mehr erfahrt ihr unter www.vestigas.com _________________________________________________ 00:00 – Darum geht's in der Folge 03:04 – Robotik wird zur nächsten KI-Revolution 08:17 – Warum große Investments ins Ausland abwandern 11:30 – Foundation Models einfach erklärt 17:46 – Deutschlands Chancen bei Physical AI 26:37 – Robotik auf der Baustelle: Was heute schon möglich ist 31:36 – Wann Innovation wirklich wirtschaftlich wird 36:32 – Construction Tech 2026: Wo die Branche heute steht 44:52 – Welche Technologien Investoren jetzt suchen 55:06 – KI-Lieblinge, Zukunftsausblick und Europas Chancen
Differentiable physics, neural emulators and foundation models for PDEs are the focus of this conversation with Professor Nils Thuerey, head of the Physics-based Simulation group at TUM. Neil and Nils discuss PhiFlow, PICT, Tadpole, scalable 3D transformers, online synthetic data, open datasets, world models and agents that call physics simulators.Full episode, corrected transcript and resources:https://neilashton.co.uk/podcasts/s4-e5-prof-nils-thuerey-on-differentiable-physics-and-foundation-models/TopicsDifferentiable physics and physics-based deep learningPhiFlow and differentiable simulation across ML frameworksWhen neural emulators can outperform their training dataFoundation models for PDEs and synthetic online trainingScalable 3D transformers and high-resolution simulationsLES, temporal data and correlated CFD datasetsOpen-source tools, startups and physics-aware world modelsAI agents that call physics simulatorsPapersNeural Emulator Superiority: When Machine Learning for PDEs Surpasses its Training Datahttps://arxiv.org/abs/2510.23111Tadpole: Autoencoders as Foundation Models for 3D PDEs with Online Learninghttps://arxiv.org/abs/2605.15284P3D: Scalable Neural Surrogates for High-Resolution 3D Physics Simulations with Global Contexthttps://arxiv.org/abs/2509.10186PICT — A Differentiable, GPU-Accelerated Multi-Block PISO Solver for Simulation-Coupled Learning Tasks in Fluid Dynamicshttps://arxiv.org/abs/2505.16992PhiFlow: Differentiable Simulations for PyTorch, TensorFlow and JAXhttps://proceedings.mlr.press/v235/holl24a.htmlPhysics-based Deep Learninghttps://arxiv.org/abs/2109.05237Learning to Control PDEs with Differentiable Physicshttps://arxiv.org/abs/2001.07457Solver-in-the-Loop: Learning from Differentiable Physics to Interact with Iterative PDE-Solvershttps://arxiv.org/abs/2007.00016tempoGAN: A Temporally Coherent, Volumetric GAN for Super-resolution Fluid Flowhttps://arxiv.org/abs/1801.09710Deep Learning Methods for Reynolds-Averaged Navier-Stokes Simulations of Airfoil Flowshttps://arxiv.org/abs/1810.08217WeatherBench: A Benchmark Dataset for Data-Driven Weather Forecastinghttps://arxiv.org/abs/2002.00469SuperWing: A Comprehensive Transonic Wing Dataset for Data-Driven Aerodynamic Designhttps://arxiv.org/abs/2512.14397LinksNils Thuerey and the Physics-based Simulation grouphttps://ge.in.tum.de/about/n-thuerey/Chapters00:00 Podcast intro00:39 Introducing Prof. Nils Thuerey04:13 Conversation begins05:13 From Computational Numerics to Graphics and Visual Effects07:17 Physics-Based Deep Learning Before ChatGPT10:01 CNNs, Graphics and the Move into Engineering Applications12:37 PhiFlow and Differentiable Physics14:13 Can Neural Emulators Surpass Their Training Data?18:00 The Promise and Limits of Foundation Models for PDEs20:43 Tadpole and Synthetic Online Pre-Training24:07 From Canonical PDEs to Navier-Stokes and Industrial CFD26:35 What Do Foundation Models Actually Learn?28:36 PDE Pre-Training vs. Millions of CFD Simulations33:08 Scaling 3D Transformers and Training Infrastructure35:58 Generating and Training on Data in Real Time38:00 LES, Temporal Data and Turbulence42:15 Overfitting and Correlated Simulation Data44:27 Bringing Differentiable Solvers Back into the Loop45:31 WeatherBench, APEBench and the Value of Benchmarks47:09 SuperWing, Open Datasets and Commercial Data51:31 Open Source, Commercial Models and a Technical Oscar56:17 Academia, Startups and Industry01:00:55 What Will Change Over the Next Five Years?01:02:07 World Models and the Need for Physics01:08:19 Agents, Tool Use and Calling Physics Simulators01:11:22 Career Advice for AI and Simulation01:13:54 Closing Thoughts
Today's clip is from episode 161, featuring Luigi Acerbi. In this conversation, Luigi explains one of the biggest engineering bottlenecks facing transformer-based probabilistic models—and how his group found a way around it.The core challenge is that many inference models treat data as an unordered set, making them naturally permutation invariant. That's statistically elegant, but computationally painful: every time a new data point arrives, the model has to recompute attention over the entire dataset from scratch, preventing the kind of KV caching that makes modern language models so efficient.Luigi walks through his team's solution: a hybrid architecture that keeps the original context fully set-based while introducing a causal-attention buffer for newly arriving data. The result is dramatically faster inference- up to 100× faster in some settings - opening the door to applications like reinforcement learning, active data acquisition, and, ultimately, Luigi's long-term vision of a foundation model for Bayesian inference.Get the full discussion hereSupport & Resources→ Support the show on Patreon→ Bayesian Modeling Course (first 2 lessons free)Our theme music is « Good Bayesian », by Baba Brinkman (feat MC Lars and Mega Ran). Check out his awesome work
For thirty years, healthcare has met every new technology with the same four words, we've seen this before. Most of the time the doubt loses. The internet was called a fad, and by 2009 a majority of US adults were looking up health information online, up from a quarter in 2000. So the reflex is earned. This episode asks the harder question, whether being right about the internet tells you anything at all about AI. Chris Boyer and Reed Smith open with why the reflex is so strong. The 1999 objections to the web read almost word for word like the 2024 objections to AI. Patients won't trust it. It makes mistakes. It can't replace the clinician. Pew found 60% of US adults uncomfortable with a provider relying on AI for their own care, which rhymes with the discomfort that greeted the web. There is even a formal model for the pattern, the Gartner Hype Cycle, though its own critics point out the curve has no data behind it. In the second segment they find where the analogy breaks. The web reached a majority of adults over roughly a decade. ChatGPT reached 100 million users in about two months, the fastest ramp UBS had seen in twenty years. The web moved information you could still judge by its source. A large language model makes the answer in the moment, with no source to check, in the same confident tone whether it's right or wrong. Chris and Reed hand you one question to run in a meeting, does the thing that made the old case turn out fine actually show up here, and one way to sort your AI uses by how much a wrong answer costs. Then the conversation turns to two people who have hosted the room where healthcare argues about this since 1996. Kathy Divis and Mike Schneider of Greystone.Net mark the 30th anniversary of HCIC. They trace digital from a small side function to something so central that a keynote once argued digital marketing was dead, meaning it had stopped needing its own name. Mike makes the point plainly, the technology changes and the questions stay the same, from the internet to CRM to AI. And they share what surfaced across this year's conference proposals, a return to trust and authority as the theme running through the whole organization. In this episode, Chris, Reed and the guests cover: Why we've seen this before is both earned pattern recognition and a way to stop thinking Where the AI story tracks the internet and where it breaks, on speed and on the difference between moving information and making it The one question that separates a real analogy from a comforting one How 30 years of HCIC read the same skepticism across the internet, CRM and AI What it means when a technology stops needing its own name, and why invisible was the goal for the web and the risk for AI Why trust and authority came back as the theme in this year's proposals If someone in your next meeting says we've seen this before, this episode is about the follow-up question that tells you whether they just said something wise or talked the room out of thinking. touchpoint.health Mentions from the Show: Pew Research Center, The Social Life of Health Information, 2009 (25% of adults online for health info in 2000, 61% by 2009): https://www.pewresearch.org/internet/2009/06/11/the-social-life-of-health-information/ Pew Research Center, 60% of Americans Would Be Uncomfortable With a Provider Relying on AI in Their Own Health Care, 2023: https://www.pewresearch.org/science/2023/02/22/60-of-americans-would-be-uncomfortable-with-provider-relying-on-ai-in-their-own-health-care/ Gartner, Hype Cycle Research Methodology: https://www.gartner.com/en/research/methodologies/gartner-hype-cycle Reuters, ChatGPT sets record for fastest-growing user base, 2023 (100M users in about two months): https://finance.yahoo.com/news/1-chatgpt-sets-record-fastest-051929384.html BMJ Open, 2026 chatbot health-information audit (nearly half of answers problematic). CONFIRM primary link and figures before publish. Medical Hallucination in Foundation Models, medRxiv, 2025 (overconfidence and poor calibration): https://www.medrxiv.org/content/10.1101/2025.02.28.25323115v1.full HCIC (Healthcare Interactive Conference), 30th anniversary, produced by Greystone: https://www.hcic.net/ Live Kathy Divis, president, Greystone. https://www.linkedin.com/in/kathyldivis/ Mike Schneider, Greystone. https://www.linkedin.com/in/michaelschneider-greystone/ Reed Smith on LinkedIn: https://www.linkedin.com/in/reedtsmith/ Chris Boyer on LinkedIn: https://www.linkedin.com/in/chrisboyer/ Chris Boyer website: http://www.christopherboyer.com/ Chris Boyer on BlueSky: https://bsky.app/profile/chrisboyer.bsky.social Reed Smith on BlueSky: https://bsky.app/profile/reedsmith.bsky.social Learn more about your ad choices. Visit megaphone.fm/adchoices
In this Tech Doctor podcast, the Tech Doctors summarize the Keynote presentation that Apple made at their 2026 World Wide Developers Conference. Siri AI was the main focus of the Keynote and the Tech Doctors describe how Siri AI will be an integral component of all apps across all of the Apple operating systems. The Tech Doctors debunk the rumor that Apple did not develop its own AI models and they describe how Google was actually involved in the development of Apple’s Foundation Models. If you want to really understand this, the Tech Doctors invite you to watch This Youtube Video This looks to be an exciting year for Apple and the Tech Doctors will be back with more podcasts as things develop.
Liquid AI co-founder and CEO Ramin Hasani joins Nathan to make a technically grounded case against the idea that scale alone defines the future of AI. Drawing on Liquid's path from MIT CSAIL work on liquid time-constant networks to Automated Foundation Model Design, he explains why efficient, hardware-aware architectures can look very different from frontier-scale attention models. The conversation centers on device-native foundation models for phones, laptops, cars, and wearables, including Liquid's open-weight LFM family and production deployments at Shopify and Mercedes-Benz. The stakes are whether useful intelligence can move beyond the data center into local, privacy-sensitive, low-latency applications—and which model and chip companies will own that on-device intelligence layer. For full show notes, links, and references, read the episode page:https://www.cognitiverevolution.ai/intelligence-on-the-edge-liquid-ai-s-ramin-hasani-on-the-search-for-device-native-foundation-models/ Mercury: Command is Mercury's new conversational interface, giving you natural-language access to your finances and helping you take actions within your existing permissions and approval policies. Visit https://mercury.com to learn more and apply online in minutes. Sponsor: Claude: Claude by Anthropic is an AI collaborator that understands your workflow and helps you tackle research, writing, coding, and organization with deep context. Get started with Claude and explore Claude Pro at https://claude.ai/tcr CHAPTERS: (00:00) About the Episode (03:53) Special Sponsor (05:41) Liquid AI origins (22:35) Neurons versus parameters (Part 1) (22:40) Sponsor: Claude (24:32) Neurons versus parameters (Part 2) (30:51) Scaling liquid networks (40:09) Automated model design (52:04) Gating and input dependence (01:01:16) Architecture bias spectrum (01:09:17) Device foundation models (01:18:16) Hardware intelligence layer (01:30:01) Local agent setup (01:36:20) Miniaturizing intelligence limits (01:40:40) Curiosity driven AI future (01:43:45) Episode Outro (01:46:46) Outro PRODUCED BY: https://aipodcast.ing SOCIAL LINKS: Website: https://www.cognitiverevolution.ai Twitter (Podcast): https://x.com/cogrev_podcast Twitter (Nathan): https://x.com/labenz LinkedIn: https://linkedin.com/in/nathanlabenz/ Youtube: https://youtube.com/@CognitiveRevolutionPodcast Apple: https://podcasts.apple.com/de/podcast/the-cognitive-revolution-ai-builders-researchers-and/id1669813431 Spotify: https://open.spotify.com/show/6yHyok3M3BjqzR0VB5MSyk
Nico, Cihat und Jan blicken auf ein Jahr NFC.cool, Smart-Home-Projekte und eine Sherlocking-Überraschung zurück, bevor sie sich ausführlich durch WWDC 2026 arbeiten – von Swift über Foundation Models bis zur Siri-Debatte in der EU.
Ep 286 I asked Apple's senior watchOS team why the latest upgrade isn't coming to so many older models — and they told me "we make power and performance a priority" macOS 27 Golden Gate makes it clear when apps are sneakily running in background iOS 27 Adds Landscape Mode to More Apple Apps Ahead of 'iPhone Ultra' ninxsoft/Mist: A Mac utility that automatically downloads macOS Firmwares / Installers. What's new in Soulver 4? | Soulver for Mac, iPad & iPhone WWDC 2026: Emptying the notebook about AI, bug fixes, and more Stalman: One of the coolest in person demos at WWDC was LM Studio running massive local models on 4 daisy chained Mac Studios with a total of 2TB memory Bringing the latest Gemini models to Apple developers Claude support for Apple's Foundation Models framework | Claude Tim Cook Confirms Apple Will Raise Prices Due to Memory and Storage Costs Apple Just Increased Prices on MacBooks, iPads, and More Zahvalnice Snimano 28.6.2026. Uvodna muzika by Vladimir Tošić, stari sajt je ovde. Logotip by Aleksandra Ilić. Artwork epizode by Saša Montiljo, njegov kutak na Devianartu
In this episode, Professor Ricardo Vinuesa - Associate Chair for Research and Associate Professor of Aerospace Engineering at the University of Michigan - explores with Neil one of the biggest questions in modern fluid mechanics: can AI help us move beyond faster CFD and toward genuine autonomous scientific discovery? Drawing on his work at the intersection of turbulence, machine learning, explainable AI, reduced-order modeling, and flow control, Neil and Prof. Vineusa discusses the promise and limits of foundation models for fluids, why the right latent representations may matter more than simply scaling data, and how agentic AI systems could uncover physical mechanisms that humans might otherwise miss.Agentic Exploration of PDE Spaces using Latent Foundation Models for Parameterized Simulations — Abhijeet Vishwasrao et al.https://arxiv.org/abs/2604.09584The episode's most direct follow-up: multi-agent LLMs and latent foundation models autonomously explore flow physics in a tandem-cylinder setup.Enhancing computational fluid dynamics with machine learning — Ricardo Vinuesa, Steven L. Bruntonhttps://doi.org/10.1038/s43588-022-00264-7A concise roadmap for useful ML in CFD, from faster simulations and turbulence modelling to reduced-order models.Identifying regions of importance in wall-bounded turbulence through explainable deep learning — Andrés Cremades et al.https://doi.org/10.1038/s41467-024-47954-6Uses explainable AI to identify flow structures that matter for prediction and control, not just visually striking turbulence features.β-Variational autoencoders and transformers for reduced-order modelling of fluid flows — Alberto Solera-Rico et al.https://doi.org/10.1038/s41467-024-45578-4Shows how disentangled latent spaces, autoencoders, and transformers can support interpretable reduced-order models of nonlinear flows.Improving turbulence control through explainable deep learning — Miguel Beneitez et al.https://arxiv.org/abs/2504.02354Links explainable AI with deep reinforcement learning to target turbulence-sustaining mechanisms, with relevance for flow control, drag reduction, and energy efficiency.LinksVinuesaLabhttps://www.vinuesalab.com/Ricardo Vinuesa — University of Michigan Aerospace Engineeringhttps://aero.engin.umich.edu/people/ricardo-vinuesa/AI and ML for Fluid Dynamics course — Ricardo Vinuesa & Sergio Hoyashttps://www.flowthermolab.com/courses/ai-ml-for-fluids/VinuesaLab YouTube channelhttps://www.youtube.com/@VinuesaLabAI for Fluid Mechanics, Sustainability & XAI — Ricardo Vinuesahttps://www.youtube.com/watch?v=TOfwf4ffPnURicardo Vinuesa — Modelling and controlling turbulent flows through deep learninghttps://www.youtube.com/watch?v=0AOY_agZ8WMChapters00:00 Podcast Intro03:20 The Evolution of Foundation Models in Fluid Dynamics10:22 Understanding Explainable AI in Fluid Mechanics15:34 Challenges in Data Fidelity for Foundation Models20:29 Machine Learning vs. Reduced Order Modeling24:22 The Shift in Focus: Turbulence Modeling to Surrogate Models29:48 Exploring Agentic Systems for Scientific Discovery37:21 Exploring Latent Representations in Fluid Dynamics40:40 The Role of AI in Autonomous Discovery41:57 Bridging Fluid Mechanics and Computer Science45:28 Data-Driven vs Physics-Driven Models51:34 The Role of Academia in AI and Fluid Mechanics56:27 Optimization and Control in Machine Learning01:00:28 Future of AI in Fluid Dynamics: Beyond ChatGPT
Send us Fan MailAre pathology foundation models actually ready for labs, or are they still stronger on paper than in practice?In this episode of DigiPath Digest #49, I unpack a timely review on pathology foundation models and ask the question that matters most to me: not just what these models can do, but what has to be true before they are genuinely useful in real pathology workflows.I walk through how pathology AI moved from narrow, task-specific models into the era of transformer-based foundation models. That shift matters because pathology is no longer only about looking at H&E in isolation. Today, pathologists are expected to integrate morphology, immunohistochemistry, molecular assays, genomics, and clinical context. That growing complexity is one reason foundation models are getting so much attention.In this discussion, I explain how transformers entered pathology, why image patches are treated like tokens, and how shared embeddings can support classification, regression, segmentation, and multimodal retrieval. I also go through the major pathology foundation models mentioned in the paper, including Virchow/Virchow2, Mayo Clinic Atlas, UNI, CONCH, H-Optimus, GigaPath, and TITAN, and why scale alone is not the full story.A big part of this episode is about the gap between benchmark performance and clinical readiness. I talk about the persistent limitations in training data diversity, the overuse of TCGA, and why public benchmarks can still miss what real pathology practice looks like. I also cover where foundation models still struggle, especially in cytopathology, hematopathology, and underrepresented disease areas, along with the real-world problems of artifacts, domain shift, concept drift, infrastructure burden, regulatory complexity, and workflow disruption.For me, one of the most important themes is this: AI in pathology should augment, not replace, pathologists. The future is not about handing diagnosis to a model. It is about building tools that support pathologists better, fit real workflows, and can be validated in ways that deserve trust.I also spend time on what comes next: explainable AI, counterfactual explanations, conversational interfaces, retrieval-augmented systems, multimodal fusion, and the need for deployment-centric validation rather than paper-only excitement.If you are trying to understand where pathology foundation models really stand today, this episode will help you separate the promise from the practical barriers.Episode Highlights00:01 – Why I chose this paper, what is changing at Digital Pathology Place, and why foundation models are worth paying attention to now.02:15 – The core questions: what pathology foundation models are, where they are, and how difficult they are to apply in pathology.04:50 – Why pathology is becoming more cognitively demanding, and how multimodal complexity is driving interest in scalable AI.07:02 – From narrow AI to transformers: how pathology moved beyond single-task CNN models.10:16 – How transformers work in pathology: image patches as tokens, self-attention, embeddings, and downstream tasks.14:16 – Why multimodality matters, and what kinds of data foundation models may eventually integrate.15:27 – Timeline of key model developments, from “Attention Is All You Need” to gigapixel-scale pathology foundation models.17:13 – The leading models and what scale really looks like: Virchow, Mayo Clinic Atlas, UNI, CONCH, H-Optimus, and GigaPath.19:51 – Why dataset diversity matters more than sheer volume, and why TCGA is not enough.23:17 – Where foundation models still struggle: cytopathology, hematopathology, rare disease, artifacts, scanner shifts, and pen marks.28:06 – Explainability, counterfactual explanations, and why trust in pathology AI needs more than attention maps.30:17 – The real deployment hurdles: regulation, infrastructure, workflow fit, and economics.36:32 – Why AI should augment pathologists, not replace them, and which tedious tasks pathologists would gladly hand over.38:36 – Retrieval-augmented and conversational AI in pathology: where interactive systems may actually help.40:51 – Vision-language models and multimodal fusion with histology, radiology, genomics, and clinical notes.42:16 – The path forward: deployment-centric design, prospective multi-site validation, and human-AI collaboration.44:08 – Closing thoughts on AI literacy, community learning, and what needs to happen next.Resources MentionedMain paper discussed:Pathology Foundation Models: Evolution, Current Landscape, Challenges and Opportunities from a Technical and Clinical Perspectivehttps://doi.org/10.3390/bioengineering13050577Review article / journal landing page:https://doi.org/10.3390/bioengineering13050577Benchmarks mentioned:PathoBench — discussed in the review paper; use the review link here for context until you want to swap in a canonical project page:https://doi.org/10.3390/bioengineering13050577PathBench — public benchmark paper:https://arxiv.org/abs/2505.20202MEDFAIR — benchmark paper:https://arxiv.org/abs/2210.01725MEDFAIR code repository:https://github.com/ys-zong/MEDFAIRModels mentioned:Model overview in the review (Virchow/Virchow2, UNI, CONCH, H-Optimus, GigaPath, TITAN, Mayo Clinic Atlas):https://doi.org/10.3390/bioengineering13050577Virchow:https://arxiv.org/abs/2309.07778UNI:https://arxiv.org/abs/2308.15474CONCH:https://arxiv.org/abs/2307.12914Mayo Clinic Atlas:https://arxiv.org/abs/2501.05409TITAN:https://arxiv.org/abs/2411.19666Dataset mentioned:The Cancer Genome Atlas (TCGA)https://portal.gdc.cancer.gov/Book mentioned:Digital Pathology 101: All You Need to Know to Start and Continue Your Digital Pathology Journeyhttps://digitalpathologyplace.com/Platform:Digital Pathology Placehttps://digitalpathologyplace.com/Support the showGet the "Digital Pathology 101" FREE E-book and join us!
Predictive modeling is a core element in modern systems, and powers capabilities such as fraud detection, loan approvals, and recommendation systems. These systems typically operate on structured, relational data stored in enterprise databases, with rows, columns, and interlinked tables. While computer vision and natural language processing have undergone a neural network revolution, the tabular data The post Foundation Models for Structured Data appeared first on Software Engineering Daily.
Predictive modeling is a core element in modern systems, and powers capabilities such as fraud detection, loan approvals, and recommendation systems. These systems typically operate on structured, relational data stored in enterprise databases, with rows, columns, and interlinked tables. While computer vision and natural language processing have undergone a neural network revolution, the tabular data layer underpinning predictive modeling still largely relies on manual feature engineering and task-specific models. Relational deep learning proposes a new approach. It treats databases as graphs and applies transformer-style attention mechanisms directly over structured relational data. Researchers are now building foundation models for tabular data that aim to generalize across predictive tasks without painstaking feature engineering. Jure Leskovec is a Professor of Computer Science at Stanford University and he previously served as Chief Scientist at Pinterest and was an investigator at the Chan Zuckerberg Biohub. Most recently, he co-founded the machine learning startup, Kumo.AI. In this episode, Jure joins Sean Falconer to discuss the limitations of traditional predictive modeling, why structured enterprise data requires its own modality-specific neural architectures, how graph transformers generalize attention to relational databases, and more. The post Foundation Models for Structured Data appeared first on Software Engineering Daily.
Predictive modeling is a core element in modern systems, and powers capabilities such as fraud detection, loan approvals, and recommendation systems. These systems typically operate on structured, relational data stored in enterprise databases, with rows, columns, and interlinked tables. While computer vision and natural language processing have undergone a neural network revolution, the tabular data layer underpinning predictive modeling still largely relies on manual feature engineering and task-specific models. Relational deep learning proposes a new approach. It treats databases as graphs and applies transformer-style attention mechanisms directly over structured relational data. Researchers are now building foundation models for tabular data that aim to generalize across predictive tasks without painstaking feature engineering. Jure Leskovec is a Professor of Computer Science at Stanford University and he previously served as Chief Scientist at Pinterest and was an investigator at the Chan Zuckerberg Biohub. Most recently, he co-founded the machine learning startup, Kumo.AI. In this episode, Jure joins Sean Falconer to discuss the limitations of traditional predictive modeling, why structured enterprise data requires its own modality-specific neural architectures, how graph transformers generalize attention to relational databases, and more. The post Foundation Models for Structured Data appeared first on Software Engineering Daily.
Predictive modeling is a core element in modern systems, and powers capabilities such as fraud detection, loan approvals, and recommendation systems. These systems typically operate on structured, relational data stored in enterprise databases, with rows, columns, and interlinked tables. While computer vision and natural language processing have undergone a neural network revolution, the tabular data layer underpinning predictive modeling still largely relies on manual feature engineering and task-specific models. Relational deep learning proposes a new approach. It treats databases as graphs and applies transformer-style attention mechanisms directly over structured relational data. Researchers are now building foundation models for tabular data that aim to generalize across predictive tasks without painstaking feature engineering. Jure Leskovec is a Professor of Computer Science at Stanford University and he previously served as Chief Scientist at Pinterest and was an investigator at the Chan Zuckerberg Biohub. Most recently, he co-founded the machine learning startup, Kumo.AI. In this episode, Jure joins Sean Falconer to discuss the limitations of traditional predictive modeling, why structured enterprise data requires its own modality-specific neural architectures, how graph transformers generalize attention to relational databases, and more. The post Foundation Models for Structured Data appeared first on Software Engineering Daily.
Predictive modeling is a core element in modern systems, and powers capabilities such as fraud detection, loan approvals, and recommendation systems. These systems typically operate on structured, relational data stored in enterprise databases, with rows, columns, and interlinked tables. While computer vision and natural language processing have undergone a neural network revolution, the tabular data The post Foundation Models for Structured Data appeared first on Software Engineering Daily.
Cihat Gündüz returns to break down everything from WWDC 2026. We go through Swift 6.4's quality-of-life wins, Apple turning Foundation Models into a full agentic harness, Xcode 27's agent and built-in skills, Device Hub, and where we land on the iPhone Fold.GuestCihat Gündüz (@Jeehut) / XCihat Gündüz (@Jeehut@iosdev.space) - iOS Dev SpaceCihat Gündüz (@jeehut) on ThreadsFlineDevRelated LinksWWDCNotesMy Top 5 AI Wishes for WWDC26 – FlineDevFlineDev/SiteKit: AI-first static site generator written in SwiftQwenLM/Qwen3.6: Qwen3.6 is the large language model series developed by Qwen team, Alibaba Group.Related EpisodesPlatforms State of the Union 2026 with Peter WithamActually Really UsefulSwift Toolkit with Natan RolnikWWDC Notes with Cihat GündüzHacking with Ignite with Paul HudsonChapters(00:00) - What's New in Swift & SwiftUI (09:09) - SwiftData (14:09) - Foundation Models (28:29) - Xcode 27 (38:09) - AI Costs & the "AI Apocalypse" (48:09) - iPhone Fold WatchClick here to watch a video of this episode. TranscriptClick here to view the episode transcript. Support the Show ★ Support this podcast on Patreon ★ Thanks to our supporters: Thanks to our monthly supporters Steven Lipton Welcome new supporters: Social MediaLinkedIn - @leogdionGitHub - @brightdigitGitHub - @leogdionMastodon - @leogdion@c.imYouTube - @brightdigitX - @leogdionX - @brightdigitCreditsMusic from https://filmmusic.io "Blippy Trance" by Kevin MacLeod (https://incompetech.com) License: CC BY (http://creativecommons.org/licenses/by/4.0/)
Wie hat dir die Folge gefallen?Gut
Sende uns Deine NachrichtWas passiert, wenn die mächtigsten Akteure der Welt nicht mehr Regierungen, Zentralbanken oder Militärs sind, sondern Betreiber von Rechenzentren und Entwickler von Künstlicher Intelligenz?In dieser Episode taucht Norman Müller tief in die Frage ein, wie sich globale Machtstrukturen durch KI verändern. Aufbauend auf den Gedanken des Soziologen Prof. Dr. Thomas Druyen diskutieren wir, warum Rechenzentren, Halbleiterfabriken und Foundation Models zur strategischen Infrastruktur des 21. Jahrhunderts werden.Wir sprechen über die Entstehung einer neuen digitalen Elite, die Rolle von KI bei der Steuerung von Informationen und Wahrnehmungen sowie die Frage, ob Europa im globalen KI-Wettbewerb noch eine gestaltende Rolle spielen kann.Am Ende steht eine entscheidende Erkenntnis: Je leistungsfähiger KI wird, desto wichtiger werden jene Fähigkeiten, die Maschinen niemals übernehmen können.Hier geht's zum Artikel:https://ventureaibriefing.substack.com/p/die-neue-weltmacht-denkt-nicht-demokratisch00:00 Ein Gedankenexperiment zur Macht der Zukunft02:20 Infrastruktur versus künstliche Intelligenz04:30 Warum KI eine völlig neue Form von Macht schafft07:00 KI als Filter unserer Wirklichkeit08:30 Wie KI-Systeme trainiert werden10:00 Das Blackbox-Problem erklärt12:00 Wer bestimmt die Werte einer KI?14:00 Der globale Wettlauf um KI-Infrastruktur15:30 Europas Rolle im KI-Wettbewerb18:00 Können Maschinen Menschen ersetzen?18:45 Die unüberwindbare Grenze der KI: VerantwortungSupport the show________________Wenn du uns dabei unterstützen möchtest, diesen Podcast zu einer Allianz von Zukunftsarchitekten der KI-Transformation zu machen, in der wir offen über Chancen, Risiken und reale Erfahrungen mit Künstlicher Intelligenz sprechen, dann abonniere uns auf Substack, YouTube, Spotify oder Apple Podcasts. Dein Abonnement kostet dich nichts, hilft uns aber sehr, noch mehr herausragende Persönlichkeiten für tiefgehende und inspirierende Podcast Gespräche zu gewinnen. Vielen Dank für deinen Support.Vernetze dich mit Norman auf LinkedIn:https://www.linkedin.com/in/muellernorman
Recorded thirty minutes after the WWDC26 State of the Union keynote ended, The Trio delivers a hot-take reaction to everything Apple just announced. Steve makes his boldest claim yet: 2026 is the year the "Universal UI" era begins, anchored by a Siri demo that appeared to actually work in real time. Xcode quietly Sherlocked the Codex app, SwiftUI got reorderable containers and (finally) AsyncImage caching, and Aaron spotted some very suspicious folding phone tea leaves in the new Simulator replacement.## Chapters00:08 Introductions 01:31 Reviewing The Trio's "Universal UI" Concept 02:39 Comparison of "AI" Apps: Siri, Claude, Codex, ChatGPT 05:43 Multimodal Prompts & Private Cloud Compute 07:26 Foundation Model Device Requirements 09:50 Dynamic Profiles and Custom Model Configurations 13:41 Xcode 27 Sherlocked the Codex App 15:42 Xcode and Developer Tool Evolution 21:47 SwiftUI Updates: Reordering and AsyncImage Cache 26:14 A Grab Bag of Random Stuff 28:23 App Actions and Siri Integration 35:04 No Apple Claw? 39:09 Swift Compiler Unable to Type Check Error 40:37 Final Impressions 42:02 Folding Phone Tea Leaves 43:07 Snow Leopard Speed Improvements 43:53 Wrap Up & One More Thing... 45:50 Tag ## Show Notes- Steve declares 2026 the start of the "Universal UI era," with a live Siri demo that actually worked as his primary evidence.- Aaron clocked the demo as mostly staring at a loading spinner; Steve argues Apple had to prove the on-device inference wasn't faked this time.- The Foundation Models framework supports dynamic profiles: configurable system prompts, temperatures, and thinking budgets per scenario within a single app.- Xcode 27 ships an agentic coding UI seemingly inspired by the Codex app, prompting Kotaro to ask point-blank: "Are you saying they Sherlocked Codex?"- SwiftUI finally has a reorderable container, which The Trio immediately wants in Bento Fit after a previous attempt even an "AI" agent couldn't pull off.- AsyncImage gets a built-in cache after years of third-party workarounds; Steve suspects some intern with an unlimited Claude Code budget finally got it done.- App Actions now supports natural language invocation without requiring specific phrases or app name mentions, though exact limits remain fuzzy.- Aaron flags resizable iOS windows (previously iPad-only) and an arbitrary-aspect-ratio Simulator replacement as very suspicious folding phone tea leaves.- Kotaro closes on Snow Leopard-style speed wins across the board, including 80% faster AirDrop, because speed is still a feature worth shipping.## Links**One More Thing**Cleo Family: https://www.cleofamily.app/track**PhillyCocoa:** https://phillycocoa.orgIntro music: "When I Hit the Floor", © 2021 Lorne Behrman. Used with permission of the artist.
In this episode, Neil explores how agents, foundation models, and AI are set to transform the Computer-Aided Engineering (CAE) and Electronic Design Automation (EDA) landscapes. He shares a comprehensive historical perspective and predicts a near-future where AI-driven automation redefines engineering workflows, productivity, and innovation.Main Topics:The evolution of simulation codes from the 1960s to modern commercial softwareThe rise of cloud computing, GPUs, and their impact on CAE and EDA industriesThe integration of AI, surrogate modeling, and foundation models into simulation workflowsThe emergence of agentic AI systems capable of autonomously performing complex engineering tasksThe strategic responses of major software companies to AI and agent technologiesThe potential democratization and automation of engineering design through AI agentsCritical questions on model ownership, transparency, and industry adoptionTimestamps: 00:40 - Introduction: How agents and foundation models will disrupt CAE & EDA01:40 - Historical overview: From code writing in the 60s to commercial software03:10 - Growth of aerospace and automotive industry codes and commercialization04:40 - The impact of HPC, cloud computing, and hardware evolution06:25 - Rise of cloud SaaS models and "sassification" of simulation tools07:40 - Big tech entrance: AWS, Microsoft, and Google in CAE & EDA09:00 - GPU acceleration: Changed landscape in past three to four years09:10 - The role of AI startups offering surrogate models and real-time simulation10:40 - Industry consolidation: Mergers and acquisitions among software giants11:40 - The emergence of foundation models and surrogate systems in simulation13:00 - The significance of agents: Combining AI, models, and automation14:10 - Capabilities of autonomous AI agents in complex engineering workflows15:25 - Practical use cases: Running simulations, setting up experiments, and data analysis16:40 - How agent-driven automation could democratize engineering expertise16:10 - Questions about model ownership, open source codes, and licensing19:40 - The future of AI in engineering: Collaboration, transparency, and scientific rigor21:25 - Final thoughts: Opportunities, challenges, and the transformative potential of AI* Please note that this a personal opinion and not that of NVIDIA
In this episode, Ben Lorica sits down with Doris Xin and Moustafa Abdelbaky, co-founders of Disarray, to discuss why classical machine learning models remain essential despite the rise of foundation models and LLMs. Subscribe to the Gradient Flow Newsletter
Food security expert David Lobell is immersed in the data of agriculture. He uses satellite imagery, yield data, and advanced computational modeling to analyze the roughly 500 million farms worldwide to increase productivity and ensure global food security – now and in the future. Though food is often taken for granted, feeding a hungry world is our greatest environmental challenge, he says. Lobell goes on to explain how data can do much more than increase yields – it also cuts costs, prevents conflicts, reduces emissions and deforestation, and improves nutrition. Smart farming is key to food security and avoiding the problems that stem from hunger, Lobell tells host Russ Altman on this episode of Stanford Engineering's The Future of Everything podcast. Have a question for Russ? Send it our way in writing or via voice memo, and it might be featured on an upcoming episode. Please introduce yourself, let us know where you're listening from, and share your question. You can send questions to thefutureofeverything@stanford.edu. Episode Reference Links: Stanford Profile: David Lobell Connect With Us: Episode Transcripts >>> The Future of Everything Website Connect with Russ >>> Threads / Bluesky / Mastodon Connect with School of Engineering >>> Twitter/X / Instagram / LinkedIn / Facebook Chapters: (00:00:00) Introduction Russ Altman introduces guest David Lobell, a professor of Earth System Science at Stanford University (00:03:01) Path into Food Security How Lobell's interest in math and the environment led him to agriculture. (00:04:31) Understanding Farming Systems How farming differs across smallholder and large-scale operations. (00:06:13) Agriculture's Biggest Challenges Improving productivity in developing regions & reducing agriculture's environmental impact. (00:08:15) Farm Potential How researchers estimate potential outputs & the barriers to better outcomes (00:11:03) Using Satellites to Study Farms How satellites help researchers understand what is happening in agriculture internationally. (00:16:13) What Satellites Can Measure Tracking crops, planting dates, harvest timing, yields, and management practices. (00:18:23) Identifying Crops from Space How seasonal patterns, biomass, and reflectance help distinguish crops. (00:20:01) Why Food Matters How food security connects to political stability, conflict, climate, and the environment. (00:23:58) Cover Crops and Tradeoffs Why a promising sustainability practice can sometimes reduce productivity. (00:26:06) Crop Rotation Insights How different rotations affect yields depending on local conditions. (00:27:35) Personalized Farming The importance of balancing large data with local information and implementation (00:31:47) Future In a Minute Rapid-fire Q&A: smarter farming, food access, and the future. (00:33:01) Conclusion Connect With Us:Episode Transcripts >>> The Future of Everything WebsiteConnect with Russ >>> Threads / Bluesky / MastodonConnect with School of Engineering >>>Twitter/X / Instagram / LinkedIn / Facebook Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
This Week in Machine Learning & Artificial Intelligence (AI) Podcast
In this episode, Jure Leskovec, co-founder and chief scientist at Kumo and professor of computer science at Stanford, joins us to explore two fronts of his work: AI for science and relational deep learning. We begin with AI Virtual Cell, a multiscale effort to learn data-driven representations from proteins to cells to patients using single-cell RNA-seq data, protein language models like ESM, and structure models like AlphaFold—without hand-encoding biology. Jure then dives into relational deep learning, reframing enterprise databases as graphs and training neural networks directly on raw multi-table data. He explains Kumo's Relational Foundation Model (RFM2), which performs in-context learning over subgraphs to make accurate predictions on new databases and tasks with no training, and how this approach benchmarks against RelBench and other multi-table datasets. We also discuss real-world deployments at companies like Reddit, DoorDash, and Coinbase, explainability via attention over tables and columns, integration with agentic systems, deployment options, and practical limitations. The complete show notes for this episode can be found at https://twimlai.com/go/768.
Support & Resources→ Support the show on Patreon→ Bayesian Modeling Course (first 2 lessons free)Our theme music is « Good Bayesian », by Baba Brinkman (feat MC Lars and Mega Ran). Check out his awesome workTakeaways:Q: Why are prior predictive checks so underused in practice, and how do simulations help?A: They're underused because researchers don't always think to run them before seeing data -- but also because doing them rigorously (in the style Michael Betancourt advocates, with prior push-forward checks on interpretable summaries) takes effort. Simulations make it cheap to generate thousands of “what-if world” datasets from your model and check whether they look plausible, catching bad priors before you ever touch real data.Q: How can generative AI help with prior elicitation?A: Rather than forcing a domain expert to choose a distributional family and parameterize it, you can use a generative model to translate their qualitative knowledge directly into a prior. The expert describes what realistic data should look like; the generative model produces synthetic datasets matching that description; those datasets are used to fit a prior distribution. It removes the assumption that experts can think in terms of parameters and replaces it with the more natural question: does this look like your data?Q: What would a foundation model for Bayesian inference actually look like?A: Stefan's bet is that it won't be a fine-tuned general LLM. The right analogy is chess: you don't fine-tune GPT to play chess, you teach it when to call Stockfish. For Bayesian inference, you'd want a semantic layer – an LLM that understands the analysis goal – calling specialized numerical engines (MCMC samplers, amortized inference networks) that do the actual computation. Agent skills are already a step in this direction; the longer-term vision is engines that have been trained from scratch to generalize across large families of models and priors.Full takeaways here.Chapters:00:00 How does amortized inference fit into modern Bayesian workflows?06:01 What role do simulations play across the full Bayesian workflow?12:12 How do you elicit priors from a domain expert who doesn't think in distributions?19:01 What would a foundation model for Bayesian inference actually look like?35:32 What is self-consistency in amortized inference and why does it matter?39:22 How does semi-supervised learning improve simulation-based inference?43:16 Why is sensitivity analysis so important yet so underused in Bayesian practice?47:40 What is multiverse analysis and how does it change how we report Bayesian results?51:32 How does amortized inference make sensitivity and multiverse analysis affordable?01:02:47 How do amortized inference and classical MCMC complement each other?01:10:08 What are the next major directions for BayesFlow and amortized inference research?Thank you to my Patrons for making this episode possible!Links from the show here.
Discover how indie developer Dani Devesa Derksen-Staats created RetroRapid, an accessible retro racing game, and Xarra, a flexible reading and listening app. Learn how thoughtful design, multiple input methods, and community feedback can shape more inclusive experiences across Apple devices. Shaun Preece talks to Dani Devesa Derksen-Staats, an indie developer based in London, about his transition from working at the BBC to building accessibility-focused apps in his spare time. His latest projects highlight how inclusive design can be applied across both gaming and productivity tools. RetroRapid is a retro LCD-style racing game available on iPhone, iPad, Mac, and Apple Watch. Originally a side project, it evolved into a fully released title following feedback at the ARCtic conference. The game was designed with accessibility in mind, supporting multiple input methods including taps, swipes, keyboards, game controllers, and the Apple Watch Digital Crown. Dani explains how audio cues—assigning musical notes to each lane—help players build a mental map of the road, alongside haptic feedback and direct touch controls. Community feedback from AppleVis also played a key role in refining the experience. Dani also introduces Xarra, a newly launched app designed to support reading and focus. It allows users to import text, PDFs, and web content, combining audio playback with synchronised on-screen highlighting at line or word level. Built with accessibility at its core, Xarra supports VoiceOver, Switch Control, Full Keyboard Access, and Voice Control, while preserving image descriptions in audio. It also uses Apple Intelligence and Foundation Models to convert code blocks into natural language when listening to technical content. Relevant Links RetroRapid & Xarra: https://accessibilityupto11.com/apps Accessibility Up to 11: https://accessibilityupto11.com ----Follow on:YouTube: https://www.doubletaponair.com/youtubeX (formerly Twitter): https://www.doubletaponair.com/xInstagram: https://www.doubletaponair.com/instagramTikTok: https://www.doubletaponair.com/tiktokThreads: https://www.doubletaponair.com/threadsFacebook: https://www.doubletaponair.com/facebookLinkedIn: https://www.doubletaponair.com/linkedinSubscribe to the Podcast:Apple: https://www.doubletaponair.com/appleSpotify: https://www.doubletaponair.com/spotifyRSS: https://www.doubletaponair.com/podcastiHeadRadio: https://www.doubletaponair.com/iheartAbout Double TapHosted by the insightful duo, Steven Scott and Shaun Preece, Double Tap is a treasure trove of information for anyone who's blind or partially sighted and has a passion for tech. Steven and Shaun not only demystify tech, but they also regularly feature interviews and welcome guests from the community, fostering an interactive and engaging environment. Tune in every day of the week, and you'll discover how technology can seamlessly integrate into your life, enhancing daily tasks and experiences, even if your sight is limited."Double Tap" is a registered trademark of Double Tap Productions Inc. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Peter Walker brings Carta's proprietary private market data from 60,000 startups and 85% of US unicorns to expose the brutal realities of today's tech landscape. While Q1 saw record capital raised, the money is highly concentrated among foundation models. We review the harsh truth behind the 530,000 open tech jobs in the US and the widening talent divide separating top 10% performers from the rest of the market. The conversation covers why venture capital is squeezing operators, how private equity is no longer a guaranteed exit strategy, and the urgent need to optimize your GTM strategy for AI-native workflows. We also debate the death of long-term product roadmaps, the impact of AI on enterprise pipeline generation, and evaluate whether middle management will survive the next 24 months. Key Takeaways The traditional safety net for slow-growth SaaS companies has disappeared, with Peter Walker noting that "A lot of companies held out this PE route as like, this is my escape hatch if I don't grow that fast. And now they're finding it's like, actually, the PEs don't care about you either." Securing venture capital has never been harder for founders outside of the AI bubble, as Peter Walker states "it's definitely not the easiest time to raise money unless you are already in the legible cohort. And if you're in the legible cohort, you know who you are." Rapid execution is replacing traditional product planning, with Peter Walker emphasizing that "the companies that are moving fast… don't know what they're doing in three weeks... we have no idea what's going to happen in October." Artificial intelligence is fundamentally threatening traditional communication hierarchies, and Sam Jacobs is bearish on the future of people managers because "if the purpose of management is to facilitate decision making at certain executive levels, that is something that AI can do." Connect with the Hosts & Guests Host: Sam Jacobs - https://www.linkedin.com/in/samfjacobs/ Host: AJ Bruno - https://www.linkedin.com/in/ajbruno3/ Host: Asad Zaman - https://www.linkedin.com/in/azaman1/ Guest: Peter Walker - https://www.linkedin.com/in/peterjameswalker/ Topline is more than a YouTube Channel! Subscribe to Topline Newsletter: https://toplinemedia.substack.com/ Tune into Topline Podcast, the #1 podcast for founders, operators, and investors in B2B tech: https://www.joinpavilion.com/topline-podcast Join the free Topline Slack channel to connect with 600+ revenue leaders to keep the conversation going beyond the podcast: https://www.joinpavilion.com/topline-slack Chapters: 00:00 The AI Funding Divide 02:36 Is It Harder to Raise Capital 04:37 VCs Only Care About Growth 06:44 The Death of the PE Exit 10:48 Transitioning to AI Native 14:50 Why Product Roadmaps Are (Kinda) Dead 19:19 Stop Defaulting to VC 29:11 Historic Tech Funding Rounds 31:59 Over 530K Open Tech Jobs 36:55 Hiring Market Concentration 46:29 The Growing Tech Talent Divide 49:11 VC Fund Performance Realities 54:28 Future of Foundation Models 58:29 The End of Middle Management
As the AI landscape evolves, the methods we use to process structured data are undergoing a silent revolution. Join us to explore how Tabular Foundation Models (TFMs) are challenging the decade-long reign of tree-based algorithms, why the traditional "train and predict" workflow is being replaced by "in-context learning," and what this shift means for the future of resilient modeling.To help us, Christoph Molnar, renowned expert in machine learning interpretability and author of the Mindful Modeler newsletter, joins us to share his perspective on the emergence of tabular transformers, the surprising power of synthetic data, and how to maintain model safety in a world without parameter updates.The decline of the "fit and predict" paradigm in tabular dataTransformer architectures vs. traditional models like XGBoost and LightGBMIn-context learning: Predicting without traditional training stepsThe role of Structural Causal Models (SCMs) in generating training dataWhy models trained on "math and probability" succeed on real-world datasetsHardware accessibility and running foundation models on local MacBooksIntegrating SHAP values and conformal prediction for model interpretabilityThe future of the data science workflow: One tool among many or a total shift?This episode is full of technical insights and forward-looking predictions that are sure to change how you approach your next dataset. As we move into a new era of AI, it's the perfect time to explore the fundamentals of the next frontier!What did you think? Let us know.Do you have a question or a discussion topic for the AI Fundamentalists? Connect with them to comment on your favorite topics:LinkedIn - Episode summaries, shares of cited articles, and more.YouTube - Was it something that we said? Good. Share your favorite quotes.Visit our page - see past episodes and submit your feedback! It continues to inspire future episodes.
In this episode of Data in Biotech, host Ross Katz sits down with Kevin Brown, co-founder of Standard BioModel, to explore one of the most ambitious projects in biomedical AI, building a multimodal foundation model that represents the full complexity of a patient across time. Drawing on a career spanning brain-computer interfaces, computer-aided diagnosis at Siemens Healthineers, and oncology data science at Bristol Myers Squibb, Kevin shares the scientific and philosophical journey that led him to a single conviction: a patient is not a document. Rather than reducing a patient to clinical notes, ICD-10 codes, or isolated test results, Standard BioModel's approach maps every available modality - CT imaging, digital pathology, genomics, EKGs, longitudinal EHR data - into a shared latent space, and models how that patient moves through time. The result is a framework designed not just for prediction, but for counterfactual reasoning, clinical trial matching, and personalized intervention, with open-source models already being validated across leading academic medical centers. What you'll learn in this episode: >> Why reducing a patient to text - clinical notes, radiology reports, genomic assay summaries - and how mapping multimodal data into a shared latent embedding space preserves information that never makes it into the written record >> How Standard BioModel's temporal architecture models patients as trajectories through an abstract embedding space rather than static snapshots, enabling counterfactual reasoning about the likely impact of interventions on a patient's future health trajectory >> Why no single foundation model can own every clinical vertical and how building a highly generalizable base model that facilitates downstream fine-tuning is a more defensible and scalable strategy than building narrow, application-specific models >> How the model handles missing modalities in real-world clinical settings, and why the architecture is designed to function effectively even when not every data type is available for every patient >> Why Standard BioModel has chosen to open-source its models and why broad, institution-specific validation across diverse patient populations is not just a scientific priority, but a prerequisite for trustworthy clinical AI Meet our guest: Kevin Brown is the Founder and CEO of Standard Model Biomedicine, where he builds foundation models for biomedicine. He previously led AI work as Director of Artificial Intelligence at SimBioSys, and held data science and applied ML roles at Bristol Myers Squibb and Siemens Healthineers. With a neuroscience research background from New York University, Kevin's work spans generative AI and machine learning for biomedical and medical imaging applications. Connect with Kevin Brown on LinkedIn About the host: Ross Katz is Principal and Data Science Lead at CorrDyn. Ross specializes in building intelligent data systems that empower biotech and healthcare organizations to extract insights and drive innovation. Connect with Ross Katz on LinkedIn Connect with us: Follow the podcast for more insightful discussions on the latest in biotech and data science.Subscribe and leave a review if you enjoyed this episode! Sponsored by… This episode is brought to you by CorrDyn, the leader in data-driven solutions for biotech and healthcare. Discover how CorrDyn is helping organizations turn data into breakthroughs at CorrDyn.
In this episode, we dig deep into the evolving landscape of industrial AI, from April Fool's pranks to real advances in robotics and automation. We break down how the line between hype and reality is blurring, and why it's more challenging than ever to separate fact from fiction in the age of agentic AI. We welcome Jonas Messner from NEURA Robotics to unpack how their 'robot gym' is collecting real-world data, why new forms of memory and multi-modal sensing are critical, and how open platforms are redefining collaboration in physical AI. Join us as we connect industry history, current breakthroughs, and bold visions for the future—where robots learn, adapt, and even monetize their skills in dynamic environments.
These models can partly generalize across species, brain regions and tasks, suggesting that a set of machine-learnable rules govern neural population activity. But will we be able to understand them?
What if the expertise that built foundation models could reshape how you think about AI's future? In this episode, Benjamin sits down with Soumya Batra, founder and CEO of WisePort AI and former safety lead on Llama 2 and Llama 3 at Meta, to explore how foundation models evolved from traditional NLP, why post-training holds the highest leverage for safety and controllability, and what natively agentic AI means for the next frontier of AI development. Whether you're curious about the model training lifecycle or wondering what comes after large language models, this conversation unpacks the technical strategies and vision shaping tomorrow's AI systems.
In this episode, we dive deep into the challenges facing time series AI model leaderboards, from hidden information leakage to the complexities of benchmarking foundation models. I sit down with Marcel Meyer to unpack why traditional approaches fall short and how our new TS Arena leaderboard is setting a new standard for fair, future-proof evaluation. We explore the pitfalls that plague current benchmarks, the surprising ways data contamination can skew results, and the innovative pre-registration protocol we've developed to keep evaluations honest. If you've ever wondered what it takes to build a truly trustworthy AI leaderboard—or why it matters for industry and research alike—this conversation is packed with insights you won't want to miss.
Tete Xiao, VP of Engineering and AI, Bot Auto joined Grayson Brulte on The Road to Autonomy to discuss the fundamental shift from virtual AI to the physical AI required for commercial autonomous trucking.Tete co-authored Segment Anything, the landmark paper that ushered in the era of specific models to an era of foundation models that generalize across large segments of data. This approach which he is implementing at Bot Auto, enables the company to move beyond the limitations of previous technology, treating autonomous trucking as a compute-driven challenge where the system learns to navigate the complex physics of driving a truck.To ensure safety, Bot Auto is utilizing a top-down redundancy architecture that mirrors aviation's triple autopilot systems. Including dual onboard computers and independent software stacks running parallel algorithms with deliberately different logic to prevent a single failure from propagating through the system.This spring, Bot Auto is planning to launch fully autonomous commercial operations with Ryan Transportation on the Houston to Dallas corridor. No safety driver. No safety observer. No human in the cab.Episode Chapters00:00 AUTNMY AI00:25 Segment Anything05:04 Virtual AI to Physical AI09:08 Redundancy and Aviation-Inspired Architecture13:40 Hardware and Software17:00 Launching Fully Autonomous Operations20:00 Foundation Models and Reinforcement Learning27:52 Compute Infrastructure35:22 Staying Ahead42:30 Building a Virtual Driver47:06 AGI48:36 Transportation Company53:59 Future of Bot Auto--------About The Road to AutonomyThe Road to Autonomy is the definitive media brand covering the Autonomy Economy™. Through our podcasts, newsletter, and proprietary market intelligence, we set the narrative for institutional investors, industry executives, and policymakers navigating the convergence of automation, autonomy, and economic growth.Join institutional investors and industry leaders who read This Week in The Autonomy Economy every Sunday. Each edition delivers exclusive insight and commentary on the autonomy economy, helping you stay ahead of what's next.Subscribe today for free: https://www.roadtoautonomy.com/ae/See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
258 | Im 19. Jahrhundert haben Robert Bosch, Werner Siemens, Gottlieb Daimler und viele andere Gründer die Basis für unseren heutigen Wohlstand gelegt. In dieser Solo-Folge besprechen wir, warum die Zeit heute ganz ähnlich ist zu damals.Partner dieser Folge:HolviFinanzen für kleine Unternehmen: Von Chaos zu Klarheit mit Holvi - Das kostenlos Holvi Flex Konto ist perfekt für Solopreneure, Freelancer und Unternehmen, die wachsen wollen. www.holvi.com/podcastMach das 1-minütige Quiz und finde eine Geschäftsidee, die zu dir passt: digitaleoptimisten.de/quiz.Kapitel(00:41) Gottlieb Daimler wird gefeuert und baut in seinem Gartenhaus den ersten Benzinmotor (02:52) Wie eine Handvoll Besessener die Deutschland AG aus dem Nichts baute (05:14) Sam Altman will Intelligenz wie Strom aus der Leitung verkaufen(07:56?) Gazprom, Starlink und die Frage, wem die AI-Infrastruktur gehört(08:46) Europa verliert vier Abhängigkeiten gleichzeitig und das ist eine gute Nachricht (12:34) Hades Mining bohrt mit Lasern nach Europas Rohstoffen der Zukunft (17:40) AI-Arbitrage: Mit Google Maps Scraping und GEO heute ein Business starten (21:23) Warum gerade jetzt der beste Moment seit 150 Jahren ist, zu gründenLearningsWandel früh sehen und handelnDaimler sah eine kommende Wende zu mobilen Motoren und setzte darauf, obwohl Otto skeptisch war. Drei Jahre später lief der erste schnelllaufende Benzinmotor; aus dem Gartenhaus entstand einer der größten Konzerne Deutschlands. Das zeigt, wie frühes Erkennen eines Umbruchs und entschlossenes Handeln disruptives Wachstum ermöglicht.AI wird zur GrundversorgungHypothese: AI wird zur Grundversorgung wie Wasser und Strom. Sam Altman hat gesagt, dass AI als Utility kommt, und OpenAI hat Infrastruktur-Deals über 1,4 Billionen Dollar abgeschlossen sowie rund 30 GW Rechenzentrumskapazität. Europa muss Abhängigkeiten vermeiden, sonst drohen geopolitische Risiken, wie Starlink oder russisches Gas zeigen.AI-Arbitrage als DenkregelHypothese: AI-Arbitrage ermöglicht, Lücke zwischen Machbarkeit und Marktkenntnis zu nutzen. Die Folge nennt Google Maps Scraping und GEO als Beispiele, die sich in realen Geschäften monetarisieren lassen (5.000 bis 10.000 Euro pro Monat). So entsteht eine konkrete Brücke zwischen Technik und Markt, ohne umfangreiche Finanzierung.Anwendungsfokus in KI investieren75 Prozent aller europäischen KI-Investitionen fließen in vertikale Anwendungen, nicht in Foundation Models. Dort liegt die Musik: Anwendungsebene für den deutschen Mittelstand, für Compliance und lokale Märkte. Gründer sollten konkrete vertikale KI-Lösungen entwickeln, statt in generische Modelle zu investieren.KeywordsAI-InfrastrukturAI-ArbitrageGenerative Engine OptimizationEuropäische SouveränitätDeep Tech EuropaWie AI zur Grundversorgung wirdEU Inc europäische DatenplattformAI-Arbitrage Geschäftsmodell für GründerGeothermie Seltene Erden EuropaLaser-Bohrsysteme Hades MiningEuropäische DatensouveränitätMittelstand KI-Lösungen
For the final installment of our LawNext on Location series, Bob heads across the bay, from San Francisco to Oakland, to the headquarters of e-discovery company Everlaw, where he sits down with founder and CEO AJ Shankar for a conversation about technology, AI and being in it for the long game. AJ grew up in Connecticut, came west in 2002 for a computer science PhD at UC Berkeley, and has lived within a few blocks of the Berkeley campus ever since. He stumbled into the legal industry almost by accident — recruited to serve as a technical expert in litigation involving how the internet worked — and quickly realized that the legal world was home to some of the most technically fascinating and underserved problems he'd ever encountered. He never left. AJ had a prior startup, a computer vision company that was acquired, before launching Everlaw in 2011. The company was cloud-native and ML-infused from the start, built on the conviction, AJ says, that there's no single way to find the needle in a discovery haystack, and that building a genuinely useful litigation platform requires solving for collaboration, ease of use and scalability all at once. The bulk of the conversation focuses on generative AI, and how Everlaw has approached it differently than much of the market. Rather than bolting on a chatbot, AJ says, Everlaw embedded AI deliberately throughout the platform — document summarization, coding suggestions, deposition analysis, fact extraction — always grounding responses in the actual documents at hand and citing sources so users can verify the work. The December launch of Deep Dive, which lets litigators pose a question and get a synthesized, cited answer drawn from an entire document corpus in about a minute, is the feature AJ calls a "new era" for discovery — one he genuinely believes represents a categorical shift. As Everlaw continues to grow, it also remains independent, with no private equity and no outside majority owners. As for AJ, he says he is in it for the long game, and has never included an exit slide in a fundraising deck. Thank You To Our Sponsors This episode of LawNext is generously made possible by our sponsors. We appreciate their support and hope you will check them out. Paradigm, home to the practice management platforms PracticePanther, Bill4Time, MerusCase and LollyLaw; the e-payments platform Headnote; and the legal accounting software TrustBooks. Briefpoint, eliminating routine discovery response and request drafting tasks so you can focus on drafting what matters (or just make it home for dinner). Chapters 00:00 Introduction and Setting the Scene 03:23 The Journey to Founding Everlaw 08:36 The Evolution of Everlaw's Technology 11:06 Incorporating Generative AI into Legal Processes 14:04 Deep Dive: A New Era in Discovery 19:17 Transformative Experiences in Legal Discovery 22:27 Previewing Innovations at Legal Week 25:03 Understanding AI's Limitations in Legal Contexts 28:11 Navigating Hype in Legal Technology 30:47 The Impact of Foundation Models on Legal Software 34:36 Future Vision for Everlaw and Legal Tech 38:13 Closing Thoughts and Company Philosophy If you enjoy listening to LawNext, please leave us a review wherever you listen to podcasts.
Send a textIf AI can detect patterns we cannot see, how do we know when its answers are clinically trustworthy?In this episode of DigiPath Digest #39, I explore a big-picture question in digital pathology and medical AI. Many models now match or even exceed human performance in specific diagnostic tasks. But most of that evidence comes from controlled or retrospective datasets. So what happens when we try to bring these tools into real clinical workflows?I review four recent papers that help frame this challenge and point toward the next steps for trustworthy AI in healthcare. You will hear about the role of prospective validation, real-world effectiveness, transparent reporting standards, and multimodal data integration as recurring themes across these studies.Key Highlights00:00 – Introduction What do we do when AI detects signals that humans cannot see? The core challenge is verifying those outputs before trusting them in clinical decision making. 03:32 – AI Across the Healthcare Continuum A narrative review shows AI achieving clinician-level performance in well-defined imaging tasks, including digital pathology. But most evidence comes from retrospective or controlled environments, and prospective validation remains limited. 08:34 – Multi-Omics and AI in Gastric Biopsy Diagnostics Morphology alone cannot fully capture molecular heterogeneity or predict disease progression. Integrating genomics, proteomics, metabolomics, and other omics with AI is shifting gastric pathology toward data-driven precision gastroenterology. 13:38 – Hyperspectral Imaging for Real-Time Surgical Guidance Spectral imaging can analyze tissue composition during surgery without staining, freezing, or contact with the tissue. Studies show promising sensitivity for detecting malignancy and supporting intraoperative decision making. 17:20 – REFINE Reporting Guideline for Foundation Models and LLMs An international consensus guideline introduces a 44-item reporting checklist to standardize how AI studies are described. The goal is transparent, reproducible, and comparable research in medical AI. 22:35 – Big Takeaway AI should be viewed as clinical decision support, not a replacement for clinicians. Real-world validation, ethical governance, and reproducible research standards will determine how these tools enter pathology workflows. References (Articles Discussed)Artificial Intelligence in Healthcare: From Diagnosis to Rehabilitation https://pubmed.ncbi.nlm.nih.gov/41755929/Transforming Gastric Biopsy Diagnostics: Integrating Omics Technologies and Artificial Intelligence https://pubmed.ncbi.nlm.nih.gov/41751306/From Image-Guided Surgery to Computer-Assisted Real-Time Diagnosis with Hyperspectral and Multispectral Imaging https://pubmed.ncbi.nlm.nih.gov/41750768/REFINE Reporting Guideline for Foundation and Large Language Models in Medical Research https://pubmed.ncbi.nlm.nih.gov/41762555/If you enjoy staying current with digital pathology and AI research, this episode will help you connect the dots between promising algorithms and practical clinical adoption.Support the showGet the "Digital Pathology 101" FREE E-book and join us!
Ben Lorica talks with Sudip Roy (Co-founder & CTO, Adaption Labs) about why enterprise AI adoption stalls in the “last 5%” of reliability — and why waiting for the next frontier model release is usually the wrong bet. They unpack “adaptation” as something broader than post-training, including gradient-free, inference-time techniques that can sit above models to route, combine, and continuously improve behavior.Subscribe to the Gradient Flow Newsletter
On AI in Action, IBM researcher Campbell Watson explains how foundation models are accelerating discovery across Earth and space science. Moving beyond traditional numerical methods, his team applies concepts from large language models to multimodal satellite data to build powerful, open-source AI systems. In collaboration with NASA and the European Space Agency, they have developed foundation models for Earth observation, weather and heliophysics. They are using AI for sustainability use cases, such as flood detection, biodiversity monitoring and solar flare forecasting. Designed for hybrid cloud environments and even deployed in orbit, these models point toward a future where AI and quantum computing unlock deeper planetary insights.
Welcome to Exponential View, the show where I explore how exponential technologies such as AI are reshaping our future. I've been studying AI and exponential technologies at the frontier for over ten years. Each week, I share some of my analysis or speak with an expert guest to make light of a particular topic. To keep up with the Exponential transition, subscribe to this channel or to my newsletter: https://www.exponentialview.co/ ----In this episode, I'm joined by Jaime Sevilla, founder of Epoch AI; Hannah Petrovic from my team at Exponential View; and financial journalist Matt Robinson from AI Street. Together we investigate a fundamental question: do the economics of AI companies actually work? We analysed OpenAI's financials from public data to examine whether their revenues can sustain the staggering R&D costs of frontier models. The findings reveal a picture far more precarious than many assume; we also explore where the real infrastructure bottlenecks lie, why compute demand will dwarf energy constraints, and what the rise of long-running agentic workloads means for the entire industry. Read the study here: https://www.exponentialview.co/p/inside-openais-unit-economics-epoch-exponentialviewWe covered: (00:00) Do the economics of frontier AI actually work? (02:48) Piecing together OpenAI's finances from public data (05:24) GPT-5's "rapidly depreciating asset" problem (13:25) Why OpenAI is flirting with ads (17:31) If you were Sam Altman, what would you do differently? (22:54) Energy vs. GPUs; where the real infrastructure bottleneck lies (29:15) What surging compute demand actually looks like (33:12) The most surprising finding from the research (38:02) The race to avoid commoditization (43:35) Agents that outlive their models Where to find me: Exponential View newsletter: https://www.exponentialview.co/ Website: https://www.azeemazhar.com/ LinkedIn: https://www.linkedin.com/in/azhar/ Twitter/X: https://x.com/azeem Where to find Jamie: https://epoch.ai or https://epochai.substack.com Where to find Matt: https://www.ai-street.co Production by supermix.io and EPIIPLUS1 Production and research: Chantal Smith and Marija Gavrilov. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Lawyers have always relied on tools—but AI is different. It doesn't just assist with tasks; it makes decisions, applies judgment, and shapes outcomes. In episode #602 of the Lawyerist Podcast, Stephanie Everett talks with Damien Riehl about what ethical responsibility looks like when AI starts doing legal work on its own. Their conversation examines how AI systems embed values, why verification matters more than transparency, and how lawyers can responsibly use tools they don't fully understand. They also explore what legal expertise looks like in an AI-powered future—and why intuition, trust, and integrity may matter more than ever as machines take over the “widgets” of legal work. Listen to our other episodes on Ethics and Responsibility in AI. EP. 582 Deepfakes, Data, and Duty: Navigating AI Ethics in Law, with Merisa Bowers Apple | Spotify | LTN EP. 543 What Lawyers Need to Know About the Ethics of Using AI, with Hilary Gerzhoy Apple | Spotify | LTN Have thoughts about today's episode? Join the conversation on LinkedIn, Facebook, Instagram, and X! If today's podcast resonates with you and you haven't read The Small Firm Roadmap Revisited yet, get the first chapter right now for free! Looking for help beyond the book? See if our coaching community is right for you. Access more resources from Lawyerist at lawyerist.com. Chapters / Timestamps: 00:00 – Introduction 05:55 – Meet Damien Riehl 08:10 – Why AI Is a Different Kind of Legal Tool 11:05 – When AI Starts Doing Legal Work 14:30 – Ethics, Values, and AI Judgment 18:45 – Foundation Models vs. Legal-Specific AI 21:15 – The “Duck Test” and Trusting AI Output 24:45 – Trust but Verify: Reviewing AI Work 28:40 – What Lawyers Are Underestimating About AI 31:10 – What Still Requires Human Judgment 34:30 – Intuition, Trust, and Integrity in Law 37:40 – What This Means for Billing and the Future 40:40 – Closing Thoughts
Drug development has long been a costly, trial-and-error effort, with nine out of ten clinical programs failing despite major scientific advances. One reason is that biological information remains fragmented in silos, and traditional R&D approaches often rely on narrow, task-specific datasets. Bioptimus aims to change this by using AI to build a foundation model that integrates multimodal, multiscale biological data into a single body of knowledge. The approach has particular promise for rare diseases, where patient numbers and data are scarce, preclinical models are poor, and development economics are challenging. We spoke with Jean-Philippe Vert, co-founder and CEO of Bioptimus, about the inherent messiness of biology, the potential to transform rare disease drug development with a foundation model, and how uncovering similarities between conditions could enable repurposing of existing drugs.
Welcome to Exponential View, the show where I explore how exponential technologies such as AI are reshaping our future. I've been studying AI and exponential technologies at the frontier for over ten years.Each week, I share some of my analysis or speak with an expert guest to make light of a particular topic.To keep up with the Exponential transition, subscribe to this channel or to my newsletter: https://www.exponentialview.co/-----At Davos 2026, the mood was unlike any previous World Economic Forum gathering. With Donald Trump arriving amid escalating geopolitical tensions and European leaders sounding alarms about sovereignty, I recorded live dispatches from the ground. In this special episode, I bring together observations from four days at the annual meeting, tracking the seismic shifts in global order alongside the practical realities of AI adoption in the enterprise.Skip to the best bits:(00:38) Day one at Davos(02:10) Three recurring themes through the week(03:55) Day three at Davos(05:12) Mark Carney's stirring speech(05:52) Why European leaders are sounding the alarm(06:51) Why technological sovereignty just became urgent(09:31) Day four at Davos(12:59) What leaders really have to say on AI adoption(14:07) The case for only using open source modelsWhere to find me:Exponential View newsletter: https://www.exponentialview.co/Website: https://www.azeemazhar.com/LinkedIn: https://www.linkedin.com/in/azhar/Twitter/X: https://x.com/azeemProduction by supermix.io and EPIIPLUS1. Production and research: Chantal Smith and Marija Gavrilov. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
The Stanford PhD who built DSPy thought he was just creating better prompts—until he realized he'd accidentally invented a new paradigm that makes LLMs actually programmable. While everyone obsesses over whether LLMs will get us to AGI, Omar Khattab is solving a more urgent problem: the gap between what you want AI to do and your ability to tell it, the absence of a real programming language for intent. He argues the entire field has been approaching this backwards, treating natural language prompts as the interface when we actually need something between imperative code and pure English, and the implications could determine whether AI systems remain unpredictable black boxes or become the reliable infrastructure layer everyone's betting on. Follow Omar Khattab on X: https://x.com/lateinteractionFollow Martin Casado on X: https://x.com/martin_casadoCheck out everything a16z is doing with artificial intelligence here, including articles, projects, and more podcasts. Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures. Stay Updated:Find a16z on XFind a16z on LinkedInListen to the a16z Show on SpotifyListen to the a16z Show on Apple PodcastsFollow our host: https://twitter.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
AI is changing how companies are built and how venture firms operate, forcing faster decisions, clearer judgment, and new ways of working.In this exclusive conversation, Ben Horowitz shares how Andreessen Horowitz adapts to that shift. He explains why managing GPs is different from running a company, how investors are evaluated at the moment of decision rather than years later, and why verticalized teams help the firm scale without internal politics.Ben also breaks down the current AI cycle, from treating AI as a new computing platform to why application design and model orchestration matter more than raw model size. He discusses the return of M&A and why today's AI market reflects real demand, not just inflated valuations. Resources:Follow Ben on X: https://twitter.com/bhorowitzFollow Jen on X: https://twitter.com/jkhamehl Read Justine's piece ‘There is No God Tier Video Model': https://a16z.com/there-is-no-god-tier-video-model-but-there-is-something-better/ Stay Updated:If you enjoyed this episode, be sure to like, subscribe, and share with your friends!Find a16z on X :https://twitter.com/a16zFind a16z on LinkedIn: https://www.linkedin.com/company/a16zListen to the a16z Podcast on Spotify: https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYXListen to the a16z Podcast on Apple Podcasts: https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures. Stay Updated:Find a16z on XFind a16z on LinkedInListen to the a16z Show on SpotifyListen to the a16z Show on Apple PodcastsFollow our host: https://twitter.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Cofounders Jeremy Wohlwend and Gabriele Corso join the a16z podcast to discuss the launch of Boltz, a public benefit company building AI infrastructure for molecular biology. The conversation explains how breakthroughs following AlphaFold moved the field beyond protein structure prediction into modeling biomolecular interactions and binding strength, why open-source Boltz models saw rapid adoption across pharma and biotech, and how that work is now being productized. They outline the launch of Boltz Lab, a platform that brings protein and small-molecule design agents into scientist workflows, Boltz's decision to operate as an infrastructure company rather than a therapeutics company, and how AI could reduce early drug discovery bottlenecks by improving molecular design and speeding iteration between computation and the lab. Resources: Follow Gabriele on X: https://twitter.com/GabriCorso Follow Jeremy on X: https://twitter.com/jeremyWohlwend Follow Jorge X: https://twitter.com/jorgecondebio Follow Zak on X: https://twitter.com/zakdoric Stay Updated:If you enjoyed this episode, be sure to like, subscribe, and share with your friends!Find a16z on X:https://twitter.com/a16zFind a16z on LinkedIn: https://www.linkedin.com/company/a16zListen to the a16z Podcast on Spotify: https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYXListen to the a16z Podcast on Apple Podcasts: https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711Follow our host: https://twitter.com/eriktorenberg](https://x.com/eriktorenbergPlease note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
No Priors: Artificial Intelligence | Machine Learning | Technology | Startups
Pundits are screaming about the so-called “AI bubble.” But historically slow-to-adopt industries like medicine and law are actually embracing AI at an unprecedented speed. Sarah Guo and Elad Gil look ahead to 2026, breaking down the major trends that will define the next era of AI technologies. They explore the future of AI foundational models, predicting breakthroughs in solving complex scientific problems. They share competing views on the timeline for robotics and self-driving cars, debating whether startups have a chance for survival or if incumbents will dominate. Elad and Sarah also discuss the return of tech IPOs and M&As, forecast a new wave of AI consumer agent software, and explore why consumer product innovation has been slower than expected. Finally, the two offer bold non-AI predictions for the new year, including the acceleration of defense tech startups and the second-order underrated impacts of GLP-1 drugs on biohacking. Plus, stick around to hear predictions on what's next for AI in 2026 from some of tech's biggest names and industry leaders. We hear from Jensen Huang (Founder/CEO NVIDIA), Arvind Jain (Founder/CEO, Glean), Winston Weinberg (Founder/CEO, Harvey), Scott Wu (Founder/CEO, Cognition), Raiza Martin (Founder/CEO Huxe), Zach Ziegler (Founder/CTO, Open Evidence), Aaron Levie (Founder/CEO, Box), Misha Laskin (Founder/CEO, ReflectionAI), Noam Brown (Research Scientist, OpenAI), Joshua Meier (Founder/CEO Chai Discovery), Bryan Johnson (Living Man, Don't Die), Sholto Douglas (Member of the Technical Staff, Anthropic), Ben & Asher Spector (Stanford PhDs) and Dylan Patel (Founder/CEO SemiAnalysis). Sign up for new podcasts every week. Email feedback to show@no-priors.com Follow us on Twitter: @NoPriorsPod | @Saranormous | @EladGil Chapters: 00:00 – Introduction 02:43 – AI Predictions for 2026 04:40 – Adoption of AI in Professional Fields 07:17 – Robotics and Self-Driving Cars 08:25 – Robotics: Incumbents vs. Startups 13:59 – Future of IPOs and M&A in AI 16:42 – Challenges in Consumer AI Innovation 21:08 – Funding of Neo Labs, RL Research 26:28 – Predictions for 2026 Beyond AI 26:44 – The Future of Defense and Technology 28:23 – Biohacking and Peptide Therapies 30:37 – 2026 Prediction from AI Industry Leaders 40:46 – Conclusion
In our season six finale, we dive deeper into how artificial intelligence (AI) is shaping the future of drug discovery and scientific research. With remarkable scale and speed, AI models parse through complex datasets and confirm or generate hypotheses, which can help scientists accelerate R&D. In this episode, co-host Danielle Mandikian welcomes Aviv Regev, Head of gRED, and Jure Leskovec, Professor of Computer Science at Stanford University, to talk about foundation models and autonomous agents. Together, they explore the opportunities and challenges of applying AI in drug discovery, including balancing innovation with scientific rigor and the evolving role of scientists. They also discuss how AI is reshaping the future of research — from building more biologically meaningful models to advancing agent-based systems and lab automation. Read the full text transcript at www.gene.com/stories/foundation-models-and-agents
Chip Huyen is a core developer on Nvidia's Nemo platform, a former AI researcher at Netflix, and taught machine learning at Stanford. She's a two-time founder and the author of two widely read books on AI, including AI Engineering, which has been the most-read book on the O'Reilly platform since its launch. Unlike many AI commentators, Chip has built multiple successful AI products and platforms and works directly with enterprises on their AI strategies, giving her unique visibility into what's actually happening inside companies building AI products.We discuss:1. What people think makes AI apps better vs. what actually makes AI apps better2. What pre-training vs. post-training is, and why fine-tuning should be your last resort3. How RLHF (reinforcement learning from human feedback) actually works4. Why data quality matters more than which vector database you choose5. Why high performers are seeing the most gains from AI coding tools6. Why most AI problems are actually UX issues—Brought to you by:Dscout—The UX platform to capture insights at every stage: from ideation to production: https://www.dscout.com/Justworks—The all-in-one HR solution for managing your small business with confidence: https://ad.doubleclick.net/ddm/trackclk/N9515.5688857LENNYSPODCAST/B33689522.423713855;dc_trk_aid=616485030;dc_trk_cid=237010502;dc_lat=;dc_rdid=;tag_for_child_directed_treatment=;tfua=;gdpr=$Persona—A global leader in digital identity verification: https://withpersona.com/lenny—Where to find Chip Huyen:• X: https://x.com/chipro• LinkedIn: https://www.linkedin.com/in/chiphuyen/• Website: https://huyenchip.com/—Where to find Lenny:• Newsletter: https://www.lennysnewsletter.com• X: https://twitter.com/lennysan• LinkedIn: https://www.linkedin.com/in/lennyrachitsky/—In this episode, we cover:(00:00) Introduction to Chip Huyen(04:28) Chip's viral LinkedIn post(07:05) Understanding AI training: pre-training vs. post-training(08:50) Language modeling explained(13:55) The importance of post-training(15:20) Reinforcement learning and human feedback(22:23) The importance of evals in AI development(31:55) Retrieval augmented generation (RAG) explained(38:50) Challenges in AI tool adoption(43:19) Challenges in measuring productivity(45:20) The three-bucket test(49:10) The future of engineering roles(55:31) ML Engineers vs. AI engineers(57:12) Looking forward: the impact of AI(01:05:48) Model capabilities vs. perceived performance(01:08:23) Lightning round and final thoughts—Referenced:• Chip's LinkedIn post on what actually improves AI apps: https://www.linkedin.com/posts/chiphuyen_aiapplications-aiengineering-activity-7358971409227792384-y0mf/• Prediction and Entropy of Printed English: https://www.princeton.edu/~wbialek/rome/refs/shannon_51.pdf• Why experts writing AI evals is creating the fastest-growing companies in history | Brendan Foody (CEO of Mercor): https://www.lennysnewsletter.com/p/experts-writing-ai-evals-brendan-foody•Inside the expert network training every frontier AI model | Garrett Lord (Handshake CEO): https://www.lennysnewsletter.com/p/inside-handshake-garrett-lord• First interview with Scale AI's CEO: $14B Meta deal, what's working in enterprise AI, and what frontier labs are building next | Jason Droege: https://www.lennysnewsletter.com/p/first-interview-with-scale-ais-ceo-jason-droege• Anthropic's CPO on what comes next | Mike Krieger (co-founder of Instagram): https://www.lennysnewsletter.com/p/anthropics-cpo-heres-what-comes-next• Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar (creators of the #1 eval course): https://www.lennysnewsletter.com/p/why-ai-evals-are-the-hottest-new-skill• The rise of Cursor: The $300M ARR AI tool that engineers can't stop using | Michael Truell (co-founder and CEO): https://www.lennysnewsletter.com/p/the-rise-of-cursor-michael-truell• Stanford webinar—How AI Is Changing Coding and Education, Andrew Ng & Mehran Sahami: https://www.youtube.com/watch?v=J91_npj0Nfw• He saved OpenAI, invented the “Like” button, and built Google Maps: Bret Taylor on the future of careers, coding, agents, and more: https://www.lennysnewsletter.com/p/he-saved-openai-bret-taylor• Anthropic co-founder on quitting OpenAI, AGI predictions, $100M talent wars, 20% unemployment, and the nightmare scenarios keeping him up at night | Ben Mann: https://www.lennysnewsletter.com/p/anthropic-co-founder-benjamin-mann• Lenny's vibe-coded app made on Lovable: https://gdoc-images-grab.lovable.app/• Story of Yanxi Palace: https://www.imdb.com/title/tt8865016/• Steve Jobs's quote: https://www.goodreads.com/quotes/427317-remembering-that-i-ll-be-dead-soon-is-the-most-important—Recommended books:• The Complete Sherlock Holmes: https://www.amazon.com/Complete-Sherlock-Holmes-Volumes/dp/0553328255• AI Engineering: Building Applications with Foundation Models: https://www.amazon.com/AI-Engineering-Building-Applications-Foundation/dp/1098166302• The Selfish Gene: https://www.amazon.com/Selfish-Gene-Anniversary-Introduction/dp/0199291152• From Third World to First: The Singapore Story: 1965-2000: https://www.amazon.com/Third-World-First-Singapore-1965-2000/dp/0060197765—Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email podcast@lennyrachitsky.com.Lenny may be an investor in the companies discussed. To hear more, visit www.lennysnewsletter.com