POPULARITY
When we first dicsussed the Summer of Simulative AI in 2024 we knew it would be a brief summer, but it has recently come back with a vengeance with SimGym in April and now Simile AI's $2B Series B, backed by GreenOaks and Index Ventures with prominent backers like Fei-Fei Li and Andrej Karpathy, running tens of millions of simulations for Fortune 100 clients like CVS and 85–99% accuracy vs human focus groups. Time to catch up on why this Second Summer of simulation is working!From creating Smallville, the landmark 2023 paper on Generative Agents that showed AI characters could remember, plan, socialize, and develop emergent behaviors, to now building foundation models of human behavior, Joon Sung Park is trying to answer a much bigger question: what if we could simulate the world before making decisions in it? In this episode, the Simile co-founder and CEO joins us to unpack the path from generative agents to digital twins, why today's frontier models still fail to capture how humans actually behave, and what it would take to eventually simulate all 8 billion people on Earth.We go deep on Simile's approach to modeling human behavior: long-form interviews, observational and transaction data, randomized controlled trials, population-level and individual-level models, and post-training on the causal mechanisms behind why people make decisions. Joon explains how his research created digital twins that reproduced human behavior and attitudes 85% as accurately as people reproduced their own responses, why models optimized to be rational can be bad simulations of irrational humans, and why understanding “social physics” may require changing model weights rather than simply prompting frontier LLMs.We also explore the much larger ambition behind simulation: testing products and policies before deploying them, finding counterintuitive paths toward desired outcomes, modeling emergent behavior across entire societies, and potentially tackling problems like climate change, democratic instability, and UBI. Joon reflects on scaling laws for simulation, the economics of data-center-scale simulated worlds, the connection to Thomas Schelling and psychohistory, why simulation is surprisingly similar to painting, and whether we might already be living in one.We discuss:* How Smallville and Generative Agents led to Simile* Why Joon's team asked: “What if we can just recreate the world that we live in?”* Why useful personal agents require deep models of their users* Memory architectures, Markdown files, and the limits of prompting* “Social physics” and behavioral foundation models* Why web data captures what people say more than what they actually do* Interviews, transactions, observational data, and randomized controlled trials* Why predicting the future matters less than understanding how to shape it* How Simile creates representative simulated populations* Simulation versus prediction and the connection to Foundation's psychohistory* How to evaluate simulations instead of simply stacking LLM hallucinations* Creating digital twins of 1,000 real people and reaching 85% behavioral accuracy* Why frontier models can struggle to reproduce real human behavior* Why good simulations need to reproduce human biases and mistakes* Post-training models on randomized controlled trials* Population-level versus individual-level simulation* Scaling laws for human simulation* The long-term ambition to simulate all 8 billion people on Earth* Whether simulations could help solve climate change or detect collapsing democracy* Thomas Schelling and the history of agent-based modeling* Why future simulations could require an entire data center* Multi-agent simulations and what happens when simulated people interact* Replacing expensive human panels with synthetic populations* Why market research is only the starting point for simulation* Why Joon sees simulation as surprisingly similar to painting* Using simulation to study questions like UBI* Whether we are already living in a simulation* Why AGI and simulation may be the twin technologies of advanced civilizationsJoon Sung Park* LinkedIn: https://www.linkedin.com/in/joonspark* X: https://x.com/joon_s_pk* Website: https://www.joonsungpark.com* Simile: https://www.simile.comTimestamps00:00:00 Introduction and Joon's Path from Art to AI00:01:46 Smallville, Generative Agents, and the Origins of Simulation00:05:03 “Let's Just Create a World” and the Future of Personal Agents00:09:53 Social Physics and Behavioral Foundation Models00:14:08 Prediction vs. Simulation: How Do You Shape the Future?00:16:59 How Simile Models Real People and Populations00:25:35 Evaluating Simulations, Digital Twins, and 85% Accuracy00:30:23 Post-Training Models to Reproduce Human Behavior00:40:04 Scaling Laws and Simulating 8 Billion People00:43:10 From Schelling to Society-Scale Agent Simulations00:46:13 The Cost and Economics of Simulating the World00:52:05 Real-World Use Cases, Synthetic Populations, and the Market00:57:27 The Future of Simulation, Painting, and UBI01:04:23 Are We Already Living in a Simulation?01:06:08 Building Simile and HiringTranscriptIntroduction: Joon Sung Park, Simile, and the Story So FarVibhu [00:00:00]: Today, we have Joon in the podcast. Excited to kick this one off. Very exciting company. I wanna kick off and ask you the question, talk us through the story of your life. How have you gotten here?Joon [00:00:13]: Yeah, for sure. I'm really excited to be here. A story of my life. So I was born in Korea, and I lived there for a good 11 years or so of my life, and then my family moved to Boston. So we moved when I was 11, and my parents were doctors, so they were going through their postdoctoral studies. My dad was a surgeon, so he was doing his sabbatical years at the Boston Children's Hospital. So I grew up there, not too close to tech. I was very much a music and artsy, painting kind of guy.Vibhu [00:00:49]: Painting.Joon [00:00:49]: Exactly. I got into painting a little bit later, in high school, but that's what I used to do. And then I grew up mostly in the East Coast after Korea. So I lived a good number of years in New Hampshire, and then I went to college in Pennsylvania. And I got into more of this tech scene, in college. So I was originally trained to be an artist. I thought that would be my professional career. So it wasn't a hobby. It was like, “Hey, let's make a living out of this.” And then gradually, I got really interested in this idea of, hey, the greatest artist often creates their own medium, and the best medium that we had available today was in computation. So I decided to go deeper into that, and one thing led to another, and we can go deeper into this, but I decided that research was something that I gradually got interested in, and here I am.Smallville, Generative Agents, and the 2023 Breakout PaperSwyx [00:01:46]: So there's a lot that you packed into the research components. You had one of the best papers of 2023, which was the generative agents paper, commonly known as the Smallville paper.Swyx [00:01:58]: Feel free to call back to anything else that you mentioned, but most people would have heard of you from this. Do you have any statistics on how many people have, like, read it? arXiv gives you something, right? Some stats.Joon [00:02:10]: Yeah, it's a good question. How many people have read it, I'm not sure.Joon [00:02:14]: I know we do keep track of citations, and they are going up quite fast.Swyx [00:02:23]: Yeah, Google Scholar has 7,200 citations.Vibhu [00:02:25]: I feel like it made a bigger hit than that, and it was a pretty instrumental paper. It got cited so many times.Swyx [00:02:34]: It is frequently the answer when people ask, “What is the best paper you've read recently?” It's this one.Vibhu [00:02:39]: I thought the memory component was pretty underrated. It was a very good early memory system, and one of the biggest papers.Foundation Models and the Search for Killer ApplicationsJoon [00:02:47]: Yeah, so maybe I can talk a little bit about how this particular paper came together. So when I got into research, it was back in 2020 when I started my PhD program at Stanford, and that was the year, when we were about to get GPT-3 to be available. So we already had GPT-2, and you could sense that there was this new class of models that was just becoming available in the market, and the team got very intrigued. And the general consensus was, “Well, is this model going to be useful for anything?” “It's really strange that these models are not trained to do any particular task.” But we decided to take a bet. So a large group of scholars at Stanford, and it was led by one of my co-founders, Percy Liang, and we came togetherSwyx [00:03:35]: Who coined foundation models.Joon [00:03:36]: Who coined the term foundation models. We wrote this paper, where that term came from called Opportunities and Risks of Foundation Models. And during that process, really the thing that I started to think deeply about was, here is a model that is fundamentally new in our ecosystem. The reason why this was new was it wasn't, again, trained to do anything in particular, but its premise was it could do anything and everything. It was like a stem cell, if you were to take a biology analogy. And I got really interested in this idea that, well, if we were to really think about what are the killer applications that this particular technology would enable, what would that be? Many of my colleagues were using this for simple classification, simple generations. Interesting that these models can do that, but from an interaction perspective, not that interesting. We've known how to do that for many decades. And what we came down to was these models are trained on this very broad data from the web, right? So these are human behavioral data. It's social media, Wikipedia, all these data. So if you poke at the right angle, then you could see human behavior that would just pop out that's quite realistic, and we've never seen that before.The Time Machine Game and Recreating the WorldJoon [00:04:45]: So that got us really interested. The exercise that we decided to do, with this particular group of colleagues, Michael Bernstein, Percy Liang, and myself, who ended up becoming my co-founder at Simile, we sat down and we played this game that we call the time machine game.Joon [00:05:03]: Imagine we were to get on a time machine and fast-forward 10 years and look back. What would have been the single application that will have mattered that would be the most interesting and inspiring? And when we thought, “Well, what if we can just recreate the world that we live in?” it's really hard to get more ambitious than that. Like, let's just create a world.Joon [00:05:24]: And that's where we started. And initially, we had this paper that was a precursor to the generative agents paper called Social Simulacra.Swyx [00:05:32]: Before you go further, were there other candidates for the most ambitious thing in the time machine exercise? What was number two or number three?Personal Agents, User Models, and Why Simulation Came FirstJoon [00:05:44]: There is a close second that we were considering, which ended up becoming more of these automation tools, especially the vision around really personalized agents that would do things for you.Swyx [00:05:59]: That's also happening.Joon [00:06:00]: It's also happening. But it was interesting for us, right, in that the reason why, we decided to go with the idea of simulation, one, I was a huge science fiction nerd, and this idea of creating simulation, I was personally really just fascinated. I loved the idea. It's really cool to see, like, a game town like this and just see these agents live in it. But at the same time, my bet was if you were to create a really amazing personal assistant out of this technology, what you need first is an amazing model of your users. So I told a model, “Hey, can you go buy late dinner for me?” And it orders Hawaiian pizza, and I do not like pineapples on my pizza. Then it totally failed. The way for it to not make that mistake is only by having a deep understanding of who I am. And I gave a very simple and dumb example here, but you can imagine how this core understanding of people is instrumental. This is how, if we have our family and closest friends, they have a good mental model of who we are. That's the basis of our social connection. So our bet also was this technology around simulation, creating accurate representation of people ought to precede the more complex agents that would automate the world that we live in. So that was the bet. But that was a very close second, and I'm still very much fascinated by it. I think there's a lot of interesting work that's going around. My hot take here, though, is I don't think we've seen a true personal assistant that's useful, in ways that meet the ambition of that particular line of work. I think there are early applications that are interesting, and if you talk to even ChatGPT nowadays or Claude, they know a lot about us. So a lot of the generation it's doing, I do think it's much more tailored, but I think the ambition is quite large in that field, and I don't think we quite have all the right ingredients just yet.Swyx [00:08:01]: So OpenClaw and these personal agents, what do you want to see from them that they don't currently have?Memory, Markdown, and the Limits of PromptingJoon [00:08:09]: I do think it's slowly getting there, but I do generally want them to have much deeper understanding of the person. Right now, you look at the models. OpenClaw, what it's leveraging is a Markdown file, and I think it's quite clever, right? So if you look at the generative agents paper, this was the same intuition that we had, where initially when we were creating the memory architecture for the generative agents, and, like, this is, like, back in 2022, so we didn't really quite have the idea of even agentive architecture or the term agent. But the intuition that we shared with some of the work that's coming out today was we initially thought, “Well, do we want to make the memory into, let's say, knowledge graph? Do we want to train a bespoke model?” All of these things. And what we decided to do was, “No. Just forget about all this.” These language models are quite good at modeling text and understanding and reasoning about text. So just put everything in a Markdown file or a text file. You're done. I thought that was quite interesting that we could do that, and there's a lot of strength in doing that. But also, there are limitations. It's the way you retrieve and make sense of data that's extremely large, it takes a lot of work. So I think that technology is getting better. I also do, however, think, there are certain things you just cannot shape just by prompting the model. So to some degree, you do need to touch the parameters of the model itself. So there is this work that I do think does need to happen, and it is happening. The question is, how far can we take it? How do we source data, and how do you also create an ecosystem where people are continuously feeding data to this model so it's learning about you?Vibhu [00:09:50]: What's the intuition between why you need to do it in the model?Social Physics and Behavior Foundation ModelsJoon [00:09:53]: My intuition behind the actual when do you train or even post-train a model versus just prompt a model is if the model has to learn the underlying physics of the world that it's operating in. So it has to learn new social physics. The places where it doesn't have to train are the places where it already has the physics. We trust the physics. It already has the base statistics, but it's just trying to react to an environment. Then I think you can just prompt your way into getting the actions out of it. I don't think the models that are out in the open have yet learned the complete mapping of social physics of humanity. This is one of the core theses of Simile, right? And one of the core reasons why that is the case is if you look at the data that the model was trained on, these models were trained on the web data, like, whatever was available on the web. And these are really interesting data sets, but they are fundamentally the self-exposed attitudinal data with some behavior data that's sprinkled around here and there. And it has yet to learn the really deep behavioral nature of people, not just what people say they do online, but what they do in real life. And this is one of what I would consider to be the dark knowledge of humanity that we haven't quite captured. And it's these data that would also need to get factored into the model creation.Vibhu [00:11:21]: You call it behavior foundation model.Vibhu [00:11:23]: There's a good one-liner here, but outside of that, what type of data do you need? What are you changing on the model level? How do you go about modeling, doing a behavior foundation model?The Three Data Buckets: Interviews, Behavior, and CausalityJoon [00:11:35]: We think about data in three buckets. So one bucket is interview data. It's quite interesting. Rich qualitative data is interesting. It's not behavioral, but we would literally ask people, “Hey, tell me the story of your life.”Vibhu [00:11:53]: It's just what we're doing here exactly.Joon [00:11:54]: The question that you all asked at the beginning of this interview literally is the question we also ask. And we ask our participants to go a little bit deeper, than how far I went. Maybe I can give more of my life story in lieu of this. But the reason why that data is interesting is by learning about this very long-tail information about people, you get a lot of texture around this model, like, this person as a model. So even understanding their childhood memory or even their trauma, their first love, these things, quite informative in ways that's really hard to predict. So that's one. Then there are two tranches of what I would consider to be the behavioral data. One kind of behavioral data is observational. So these might be like transaction data, or these might be data that you can get by scraping the web, right? So you can imagine why these data sets would be interesting, right, because they give you the base statistics of people's behavior.Joon [00:12:55]: But then there is the last category of data, that I personally think is perhaps the most important, which is the data that describes the causal mechanism, the whys of people. Some of this is covered by the interview data, the qualitative, because people talk about why they made certain decisions. But really, where you get to see the most behavioral aspect of this is in randomized controlled trials, like RCTs. Imagine you have the same setup, but you have a few different variables that you are trying to tweak. Can you get realistic human behavior out of it in ways where, imagine you had this particular option. Imagine you're even trying to choose whether you're going to drink coffee or not. The day you drink coffee versus the day you didn't drink coffee, does your behavior change? That's a data set that describes a causal mechanism. This is quite important in modeling people. The reason why this is important is oftentimes when people come to us, or not just to us, but the reason why people are interested in simulation isn't because they want to predict the future. If you're trying to win against the stock market, predicting the future is interesting.Prediction vs. Simulation: Shaping the FutureJoon [00:14:08]: But most people, most decision-makers, what they want to know is, how can we shape the future? It doesn't really help you to hear that your sales are going to tank in two quarters. They're just gonna say, “Wow, that sucks.” What they want to know is, well, what do we need to do now to avoid that future? That's the causal mechanism. And this is also very hard data to come by, right, because the world is our ground truth, but it happens once. So in a very controlled setup where everything is equal except for one variable, this kind of data set rarely happens. So this is a reason why this data set is both hard to come by and quite important if you're trying to model human behavior.Swyx [00:14:50]: So behavior, I think, is the hardest data set to acquire. What is out there? What is even possible? You're not going to know a lot of details about my life. I don't even have data for myself on my own health or habits, and I just don't log everything. So how can you have that data?Joon [00:15:14]: So we run a lot of randomized controlled trials.Swyx [00:15:17]: But you put people in the lab, they watch them sleep, or what?Joon [00:15:20]: We do care a lot about the consent process. People know that we invite them to be a member of this community to both share data and have themselves represented in different forms. But we bring a lot of people to the lab, or virtual lab, where we design experiments that would pose them real behavioral decisions. And often in these experimental setups, what makes the difference between what is attitudinal versus behavioral is whether the stake in your decision is real. That's ultimately what makes it behavioral. So in these setups, we are inspired by our colleagues in social sciences, psychology, and so forth. So when they run studies, the techniques they utilize is imagine there's an online store that you're inviting people to come by. Then whatever they purchase in this experiment, they actually get that item delivered. Like, these are the things that make the stakes real. So we run a lot of these experiments, and we also do partner with firms. Right now, we also have customers who are quite excited to at least give us a glimpse of the behaviors that their users exhibit so that we can get a little bit deeper understanding of how people behave in these different platforms.How Customers Use Simile: Populations, Queries, and ExperimentsVibhu [00:16:39]: I think on the customer side, they have a lot of data about their users, who has bought. They have the action data.Vibhu [00:16:47]: Can you walk us through an example of what someone comes to you for? What questions would they want solved? Do you customize a model for them? Do you have something off the shelf? What does that look like?Joon [00:16:59]: Today, when people leverage our models, it's often to better understand the population of their interest. So usually, the start of the relationship, we come together and hear about what population they want us to model, right? So it might be that if you're a CPG company that's selling to all of the US, then maybe it's fairly straightforward. You want to model the gen pop of the US. But at the same time, if there is a vertical or if there's a market that they're trying to go into, imagine, they want to better understand, let's say, people in their 20s and 30s living in California. That's a much more specific population. So we hear about this population, and we go recruit these people, with consent, and with incentives, and we collect some of their data and create a model of these people. Then what our product allows you to do is query them. So it can take as input a filter that is a description of the population that you want to talk to, just like the one I just mentioned, and an environment. The environment can literally be survey questions, behavioral experiments, It can be A/B testing. Oftentimes, the core use cases are things like concept testing, to start with. But also, people sometimes want to do focus groups or one of the fun use cases that we also serve is even modeling things like earnings calls for public companies.Joon [00:18:21]: So these are the use cases that we often start with.Swyx [00:18:23]: Concept testing, is that an established term? I've never heard of concept testing.Concept Testing, Gallup, and PoliticsJoon [00:18:27]: Yeah. So it has to do with they have, let's say, different messaging, different products, different ideas.Swyx [00:18:32]: It's like a marketing exercise.Swyx [00:18:33]: Okay, got it. Got it. Politics?Joon [00:18:36]: We do, have a strategic partnership with Gallup, and of course, Gallup is deep into policy space and so forth. Right now, we have not worked deeply with politics, like that area just yet, however.Swyx [00:18:49]: I'm curious if there is demand or if they really would have different needs that somehow fundamentally don't mix with your existing, users or people.Joon [00:19:00]: I think there's certainly demand.Joon [00:19:02]: But we are very much mindful of how this technology gets adopted and the societal impact that we'll end up having with this technology. And I do see politics as an area where a company has to be particularly thoughtful about the way they operate and make impact. So this is where we also want to make sure that we form enough of guardrail and perspective on how to leverage this technology before we go on to serve markets like the politics.Swyx [00:19:29]: I'll give people an example. one of my favorite shows is The West Wing. I don't know if people have watched.Swyx [00:19:34]: One of the key storylines is, like, the president has, multiple sclerosis, but they haven't. they need to figure out how to disclose it. So they run a poll with a fake governor and ask people to respond on the poll,Counterfactuals, Polling, and When Simulation Is UsefulSwyx [00:19:47]: They try to make decisions based on the results of that poll on, like, how well they'll be received, like where, how should we play this?Swyx [00:19:54]: And I'm like, well, I think those counterfactual things, I would use a simulation for this if I could trust it.Joon [00:20:01]: For sure.Joon [00:20:02]: In that show, how'd it go?Swyx [00:20:04]: In that show, it was, like a foregone conclusion. They were like, “We know it's bad. We just don't know how bad.” And then the poll came back. It was like, “It's really bad.” And then they just did it anyway.Joon [00:20:14]: Part of it is to show, right? So you're, you're looking at the ideaSwyx [00:20:17]: Maximizing drama.Joon [00:20:18]: How bad could it be? Oh, it's horrible.Swyx [00:20:20]: And to some extent, I think that is part of the trick of the, or the challenge or with being a customer of yours, which is that if I know it's. if I roughly know and can intuitSwyx [00:20:35]: What the effect is going to be, do I need you? What sensitivity of it, of effect do I need in order to make a decision, right? So for example, if I, my approval rating is 50%Swyx [00:20:48]: And I, they have this negative piece, news item comes out, and it drops to 30.Swyx [00:20:52]: If it drops to 20, if it drops to 40, do I care? No. It, I know it drops. It's negative. So when do I care about simulations?Joon [00:21:01]: You do something that's clearly bad, that's not popular, and people don't like you, like, yeah, it's likeSwyx [00:21:05]: You don't need a simulation.Joon [00:21:07]: Yeah. Well, so there are a couple of things. one is, there are use cases where, like every day, developers, designers, policymakers, marketers, every single day, they create assets. They create new products. And turns out, it's many of the decisions in hindsight is obvious. Yes, of course this is bad, but we still run those studies because understanding the magnitude and understanding how acute something is quite difficult, even if, we feel like, of course, like this makes sense. this is the reason why we make so many mistakes. Like, every time somebody goes online and say something that has huge backlash, you look at that and like, “What an idiot.” However, it's tough. That's one. There's also another aspect here, which is, again, this is the reason why simulation is different from prediction. In simulation, in the ideal case scenario. So what simulation is trying to show is it's trying to show each step of the way or each step that we need to take to get to a certain outcome, right? So in the most advanced simulations, sometimes the next step that we're suggesting might be quite counterintuitive. The analogy that I sometimes give, and I ground it in a more realistic example, but, I, as I mentioned, I'm a huge fan of science fiction, and I don't know how, many of the audience members have read, like, things like the Foundation series by Asimov.Simulation as a Path, Not Just a PredictionSwyx [00:22:37]: Oh, yeah. We've mentioned psychohistory a number of times.Joon [00:22:39]: Okay, fantastic. So I might be, talking to the right crew. If you read Foundation series, literally the first act is there's a group of scientists who have found out that, “Oh, our galactic empire is going to collapse, and we're going to have 30,000 years of unrest.” And they run psychohistory, the simulator that tries to teach them, “Okay, how can we keep this unrest to a 1,000 years?” And they plan this out, and the first step of that plan is to get the scientists who say, “Okay, this is coming,” exiled into this random place in this, galax- galaxy.Swyx [00:23:18]: Terminus.Joon [00:23:19]: Exactly. And that's so counterintuitive. Like, what a strange move that you literally sent the group of scientists who was raising voice around this potential collapse of galactic empire into nowhere. How is that the right first move? Well, it turns out in this particular simulation, that was the move.Joon [00:23:40]: It's these things, right? And the reason why these reasoning is possible is because you're showing the step function or each step that results in a particular outcome. So really what simulation allows you to do in its highest form is you give it not a problem or question, like what would people answer to the survey? That's not what we do. What we tell it is, “Here is a goal that we have. In the context of foundation, we want to keep the unrest to a 1,000 years. What is the path that we need to take now to get to that particular future?” And that's what simulation allows you to do. Now, translating that into real market, imagine you're a automobile company and you're about to release a, EV, and you're trying to understand, well, how do we market EV, to make sure that our stock price goes up? But what if the answer comes down that, well, you can market your EV in XYZ way, but that might change people's perception around the cars that's not EV and make your overall sales to go down. Not very intuitive, especially all you're trying to optimize is EV salesss, and that's the only thing that you're tracking, then that might result in a completely wrong solution, or at least different solution than what you would have expected, whether it's right or wrong.Joon [00:24:57]: That's the power of simulation.Swyx [00:24:58]: For listeners, we covered a similar topic with Mikhail Parakhin from Shopify, where they are working on SimGym. I don't know if he ever talked to you about it. it's very similar.Joon [00:25:07]: ISwyx [00:25:07]: The goal is increased conversion, but then the journey is very unusual.Joon [00:25:12]: Journey is unusual.Swyx [00:25:12]: Yeah. The-- He's trying to look for interventions on a shopping trajectory, which is similar to what you're saying. Like, it's not about the attitudinal, is your word for it.Swyx [00:25:24]: It's about behavior.Joon [00:25:25]: It's about behavior.Swyx [00:25:25]: And that's exactly the difference, right? It's, like, not about the near-term direction about-- but it's more about, like, how do you affect multiple turns of interactions.Vibhu [00:25:35]: You had a good quote at the start about this as well. It's not about people wanting to know the outcome. It's about how they can change it, change the way to get there, something like that. But I wanna take it back to how do we know this is grounded? LikeGrounding and Evaluating Digital TwinsVibhu [00:25:47]: How do you run evals? How do you test that simulations come through? if I was to do the same thing that you described with, say, your favorite LLM, Opus, GPT-5.6, have some agent to map out these thingsVibhu [00:26:02]: How different are the answers we would get if I give it the same goal, the same objective, make a decent system? You're saying that you need to change the model weight. You have your own solution to this. But how far off are we, and how do you check if it's grounded? you have some interesting stuff on your site that points to how you run real evals, but if you could take us through that side. I think that's one of the big concerns that people have. They're like, “LLMs hallucinate.”Vibhu [00:26:27]: “You're just hallucinating layer after layer,” right?Joon [00:26:30]: The way we do this, and this is the paper that we worked on after the generative agents paper that really became the, at least for Simile and also the field of simulation and synthetic panels, really became the foundation. Yeah, this is the paper. the paper is called Generative Agent Simulations of 1000 People. Here's what we've done. For this paper, we brought 1,000 people that's representatively sampled from the US to a virtual lab. And what we have done was we spent two hours collecting fairly wide-ranging data. In this particular study, we focused a lot on this interview data, that was, whose script was taken from this project called American Voices Project. And then we would also pair that with a lot of behavior data and so forth, whatever we can collect within two hours. And then we would send these people away for a couple of weeks. And during that time, I would use this data to create their digital twins. And I would bring the humans, participants back after 2 weeks and have them complete a battery of surveys, experiments, behavior studies. So we have the list here, which included things like behavioral economics games. We would run literally, like, Big Five personality test, General Social Survey. We would also go ahead and run the randomized controlled trials that were published on PNAS. And we would have their digital twins predict how the source individuals would have acted in these studies and surveys. And this is where we could replicate people's behaviors and attitudes 85 percent as accurately as people would replicate their own. So that was the first really paper that gave this validated results that we can model individuals in an accurate way. And what we ended up finding now, of course, in AI space, so this paper came out at the end of 2024. AI space, a year and a half, 2 years, that's a lifetime.85% Accuracy and Why Frontier Models Miss Human BehaviorSwyx [00:28:24]: Yeah. Just, for listeners who are not seeing the YouTube, I just wanna say, like, the headline figure is 85 percent accuracy, like, which is a big improvement over all the otherSwyx [00:28:34]: Methods that you showed.Joon [00:28:36]: But the part that was particularly striking to us, especially as we improved this technology even further, was the generative AI models like ChatGPT, Claude that's coming out, it does give you the right foundation. However, what they do not consider is the true attitudinal and behavioral aspect of people, especially in the population that you care about. So what these models are really good at today is they're trying to become the super rational, objective machines, right? So you go get their data from places like Mercor, Scale. You talk to professional programmers, scientists to create model that's amazing at reasoning. That's what they do. Simile doesn't care about any of this. The models that we're talking about here, what we're trying to create are models that are as dumb as I am, right? So if I make some mistakes, the model has to make the same mistake.Swyx [00:29:34]: Oh, that's very hard.Joon [00:29:35]: That's very hard.Swyx [00:29:36]: You're solving Murphy's paradox.Joon [00:29:37]: That's exactly. And this is a completely different data and training objective. This is also where we see quite a bit of discrepancy in the performance in human behavior prediction between the frontier models, Simile's model, and the models being created in this space, where in some cases, the model performance of frontier models go all the way down to 20, 30 percent, especially if you go into that more niche population on topics that our customers would care about. On more gen pop, it might be around 50 to 60 percent. So it's not very robust. Like, you wouldn't want to make your decision off of these and these findings. If you can bring that up to 85 percent, that is ultimately what people end up getting very excited about.Swyx [00:30:20]: Yeah. Do we wanna keep going on the paper, routes?Joon [00:30:23]: Yeah, for sure. So the last one, was an interesting one. So this, paper was the follow-up paper that we had, to the 1000 agents paper, where the idea was now can we augment the models even further and post-train a model based on a lot of randomized controlled trials? So this was an interesting one. The data is always the most interesting part of modeling in many ways. The data that we got here was there's this, there's this platform called Open Science Framework. So some, the audience might be familiar with this. And there has been, especially in the social sciences over the past 5 years or so, there has been this concern around replicability of studies. And so it was a bit of a crisis, the scientists acknowledged, where we rerun the study and we don't see the same finding.Post-Training on RCTs and Replication StudiesVibhu [00:31:12]: Oof.Joon [00:31:12]: It's tough. And the reason why it's there-- that was often the case was there's this survival bias where the papers that get published often need to maintain what we call the value of less than 0.05 in the experiments that we ran. That suggests that only-- there's only 5% chance that the results that we saw is false positive. But the tricky part was all the papers that were not published, and there's still a 5% chance that whatever we publish is totally just randomly generated. Like, there's a 5% chance that, hey, this effect is not real, but it just happened to be real because of the sampling bias. So because of that, what scientists started to do was they started to register their studies. So before running an experiment, they would go to this platform and say, “Here is the data. Here is the population that we're collecting, and here's the hypotheses.” And they would just say, “Here is our hypothesis.” Like, “This is what we believe.” And you cannot retroactively change those hypotheses. This is what gives us more scientific statistical confidence that whatever effect that you ended up seeing is true. So that ended up creating this really interesting platform where there's one platform that has now contains tens of thousands of real-world experiments and hypotheses. And a lot of these are really high-quality, like, professionally designed behavior studies and random- randomized controlled trials. So we got the data and the studies from this platform and used that to make a point. And this particular, model is not, something that we're serving commercially because this was a part of the open science. But this particular data set, helped us make a point that by collecting a lot of these randomized controlled trials, that are really well-designed, we can make significant improvement in model's capability to predict human behaviors. So that's what this paper was about.Vibhu [00:33:10]: Is this stuff done on a individual level? Like, do I need to tune the model per individual, per company? Is there foundation model changes and then some slight post-training? Anything you can share there?Population-Level vs. Individual-Level ModelsJoon [00:33:21]: So this particular model was trained. the data we had at the level of individuals, but this particular model was trained. We experimented with both. And this is what we end up doing at Simile too. We always train 2, distinct model. One is what we call the population-level model. The other is what we call the individual-level model. And both take very similar input, which is the description of a subpopulation or individual and a stimuli. In this particular work, we've done the same. Here, the results that we are reporting are much more geared towards individuals because we do think that is a harder task in many ways, but that's what we have done.Vibhu [00:34:02]: You seen anything on the questions that humans can solve that models can't solve? So likeHuman Biases, Mundane Choices, and What Models MissVibhu [00:34:09]: Currently, it's, I live 5 minutes walk away from a car wash. It's a 10-minute drive. Should I walk or drive?Joon [00:34:16]: Huh.Vibhu [00:34:16]: The model will say, “Oh, walk to the car wash.” And, you don't have your car.Vibhu [00:34:20]: Is anything like this a problem in simulation? You would assume, like, very simple for human to think about, but if the model is saying you should walk to the car wash, anything here?Joon [00:34:32]: It's less, what can we solve, but I think it's more about what biases or mistakes do people make that models miss. Like, imagine that you are, like the. When I was still at Stanford, I lived in Palo Alto. So it's about, I would say, 40-minute walk from the campus. You ask the model, “Okay, let's go home. What can I, what can I do?” It would likely call an Uber or, give me, the bus time. But for the longest time, I really liked walking back. And the reason why I wanted to do that was not for efficiency. It really helped me think. And I like to walk for, half an hour or 40 minutes or so a day, where I just get to, just think about ideas, research, just get lost in my thoughts. That's very human activity. Unless the model has seen that and understands the importance of that activity, it would miss these kinds of features. So that I think, is fundamentally what we're trying to model. Like, what is fundamentally human might not be the most efficient thing to do, might not be the right thing to do, but things that make us who we are.Swyx [00:35:43]: I'm curious if, there are some data sets that you really want that would materially help you. One version of this may be interesting, which is more valuable to you to acquire as a data set, all of LinkedIn, all of Twitter, all of Facebook?What Data Matters: Social Media, Transactions, and FacebookJoon [00:35:57]: It's a little bit hard to rank, in part because, there's, there's this product saying where no feedback is wrong because it teaches you something about your users. Doesn't matter what feedback.Joon [00:36:11]: I think it's a little bit like that.Swyx [00:36:12]: So just whatever is bigger.Vibhu [00:36:13]: What about a different domain? Say it was. What about all of Amazon data?Joon [00:36:17]: Oh, yeah.Vibhu [00:36:18]: Shopping data, right?Joon [00:36:18]: Shopping data. So Amazon data is interesting in that it's very much behavioral, although, like, what people do on social media, you could squint and say that is also behavioral. But the transaction data is always interesting. It is also most commonly available, however.Joon [00:36:33]: If we were to look at purely social media, like if you really, if I were, if I had to really pick, Facebook likely is interesting because I do think it is most a default version of people. Because you go to LinkedIn, it's very much professional environment. So people put up their, they have their guards up, right? And that still is interesting because that is true human attitude and behavior, but it is not your base state. you go to Twitter- Twitter, people have their own crazy personas, or depending on who you are. Like, my Twitter profile and, persona is very much, initially was I was very much an academic. “Hey, I'm here to share my studies.” Now, I share, things that's related to Simile. But Facebook is one of those more private space where people just connect with their friends. In that way, I do think it shows you a little bit more about who that person is. So if I had to pick, I'd likely pick, Facebook.Swyx [00:37:30]: Yeah. And you're interested in, like, the whole person and their background and philosophy. I, is it too clinical or too machine learning-oriented to just say this is just ways to inject variance and biases? The broad question, is, like, is this any better than a randomized, like, combinatorial explosion version? So we have a link to the TencentBillion Personas, Synthetic Demographics, and Bespoke DataSwyx [00:37:54]: Billion persona paper, where they did not do any of the groundwork that you are doing.Swyx [00:37:59]: They just did like a cross matrix of here's all the professions in the world, here's all the people, possible backgrounds in the world, do a dot product across all of them, and that's it. That's your prompt for a billion people.Swyx [00:38:12]: This will do something. I don't know if it'll do what you do, but it gets you some way, some percent of the way there.Joon [00:38:18]: So this was an interesting paper. Like, what I admired about this paper when it came out was the scale. And you do gradually want to be able to simulate really large societies and interactions. So the scale is definitely admirable. it is relying heavily on the known statistics that went into training the model. So to the extent that you believe that statistics is correct, this is not a bad way to go about this. But the thesis here, and this is something that we also have seen in the market, like if this works, then we have solved simulation.Joon [00:38:54]: It,Swyx [00:38:55]: Because I survey, like, okay, 5% of the US population is in construction.Swyx [00:39:01]: The other 5% is in medicine, whatever, right? And then you just keep going down the list, and then you do the other side. 5% has, like, the big 5 personalitySwyx [00:39:08]: Of, like, neurotic or whatever. That's it.Joon [00:39:11]: That's it. So if you believe that the underlying data set and the platform that we're leveraging has all the right statistics, then this will have solved it. you're at that point merely retrieving the knowledge that is already embedded in the model, in the model parameters. That's not, unfortunately, what we see, where there is such detailed and also niche knowledge about people that if you just take one example, it might feel very mundane, but it's quite rich when you put together, that you do need to do a lot of bespoke data collection to better understand people. And this is also, I think what makes this particular, job fun, which you want to deeply understand people, and the process of deeply understanding them requires a lot of attention to the details. And you do need to pay attention to and pay respect to the daily lives that people lead.Scaling Simulation: From Thousands to SocietiesVibhu [00:40:04]: I wanna talk about scaling simulation.Vibhu [00:40:07]: So what can't we simulate, what can we simulate, and how does scaling affect this? So how big are the models? What if we go from, 8B, like, couple 100 billionVibhu [00:40:18]: Like billion000 parameters, billion000? Do we get scaling? Any interesting emergence? Like, at a certain scale, at a certain amount of training, you uncover anything unusual and any learnings from that?Joon [00:40:31]: What we are seeing is at Simile, so we do post-train our own model. The thing that we're seeing is the early glimpse of scaling law in simulations. The more data about humans and more compute you ingest, you start to get predictive and predictable gains of the model performance in simulating it, simulating people.Vibhu [00:40:51]: Ooh. We need a scaling law curve.Joon [00:40:52]: It's scaling law. Whenever you find it's a beautiful thing. And we're starting to see the glimpse of it, which is quite exciting. But if you talk about the ambition of simulation as a whole, it's not merely about building a model. It's about building a model, then creating the agents that become the individuals in a much larger ecosystem. So they're creating this multi-agent simulation. Down the line, you want these multi-agent simulation to also live in a very rich environment, right? What we are really trying to get to at that point is, hey, can we create. All right, let's do a time machine game again, and 5 years, 10 years into the future, can we create a simulation of 8 billion people living on Earth? I think that's quite interesting. And that really is the vision. And once you get to that state, the questions that you can help answer for the society also start to change from my perspective. The answers are fundamentally about emergence of the emergent behavior of society and large groups of people.Joon [00:41:53]: So the questions that I get excited by, and maybe this is a stodgy- a bit. I have my, academic side of me.Joon [00:42:01]: And for me, it's questions like, can we help solve climate change? If you look at climate change as a problem space, this is what we, like social scientists would often call it the wicked problems, problem where you have many actors with competing incentives for trying to make a very complex decision and coordinating that coordination decision. Very difficult to really solve in real life, which is also the reason why we couldn't solve it. Can simulation help us solve that? Another one is, can we understand the signals for collapsing democracy, or can we understand or can we uncover the origin story of the monetary system? These are societal questions that we never really had a good way of answering. If we can create simulations of our society, you have to believe that these are the problems that we can solve. So that's really the ambition of this field. And, I also think, yes, I think there's a Nobel Prize to be won there, which wouldn't be surprising. And I think there's some amazing societal impact that we can have to help people make better decisions.Climate Change, Democracy, and Societal SimulationSwyx [00:43:04]: Nobel Prize in economics?Joon [00:43:06]: In economics.Swyx [00:43:06]: Oh, I see. I see. Rooting for you to write that paper.Joon [00:43:10]: One of these days. But, one of the scholars that I was deeply inspired by, When I was coming into the space of simulation, is this scholar, named Thomas Schelling.Schelling, Agent-Based Models, and the Nobel PrizeSwyx [00:43:23]: Schelling point?Joon [00:43:24]: So the canonical example of the work that he's done was he was one of the creators of agent-based modeling. So this was, like, in the 1970s and 80s. It's very early days, but this was truly one of the first exemplars of simulations. And one of the canonical model from that time, and of course many of these simulations are trying to tackle the societal problems that's most relevant for their era, it was called the model of segregation. So racial segregation was a big topic, that, we cared about. And what they've done was they created this grid world where they had red dots and blue dots. And these dots were, back in the day, like, they were the agents, and they had a simple rule that governed their behavior. If certain percentage of your neighbors are of different color and if that goes above certain threshold, then you move to a new location at random.Joon [00:44:21]: One of the striking finding of this paper or this agent-based model was for the longest time, people thought the segregation within society was caused by explicit and overt racism.Joon [00:44:34]: But if you look at this model, people's preference towards living with people of the same color, that preference can be very minute.Joon [00:44:42]: But the very small difference causes the society to segregate completely over time. This was very counterintuitive for a lot of people. And this particular work ended up informing housing policies. Mixed income housing, got really inspired by this work. And Thomas Schelling ends up winning the Nobel Prize for having laid the groundwork for very early versions of simulations. The opportunity that I do see here in the more scientific terms, is agent-based models for the longest, had impact in the 1980s, 90s, to some extent, early 2000s, but it has now gotten forgotten by the community a little bit. Because as you can imagine, red dots and blue dots is not really a rich description of people.Joon [00:45:31]: But with the emergence of things like generative AI and, in particular, generative agents, we do have an opportunity to create these agent-based models that are high fidelity enough to help us make really complex decisions. And that's the opportunity that I see. If that truly works, then yes, that is the work that will result in a Nobel Prize.Swyx [00:45:53]: Yeah. For what it's worth, and I grew up in Singapore. 80% of Singapore is in public housing, and public housing has, enforced racial quotas for exactly that reason, which is very interesting. okay, so we talk about scaling, we talk about all these, the agent possible applications.Cost, Reuse, and the Economics of SimulationSwyx [00:46:13]: I'm scared about the cost. if you even-- let's just keep it to the US, about 8 billion people.Swyx [00:46:21]: But, how much does it cost to model so many hundreds of millions of people?Joon [00:46:26]: Oftentimes today, we don't start at that scale, this stage of the, of industry and simulation as technology. But we can get our users extremely rich and meaningful insights even by modeling thousands, tens of thousands of people. And today what we do is every week we are collecting data on the scale of tens of thousands people's data, and we have panel partnerships that gets us to tens of millions of people globally. So that's what we do today.Swyx [00:46:55]: And just as a side note once you've collected one person for one studySwyx [00:46:59]: Can you reuse that same person for all the subsequent studies?Joon [00:47:03]: That's exactly right.Swyx [00:47:03]: Okay.Joon [00:47:04]: The beauty of this model and these agents is the fact that they are domain-agnostic.Joon [00:47:08]: That what you're really trying to understand is what is the fundamental nature of these people? What's their social physics? And there are a lot of, a lot of, people that does change over time. Like, even, like, even things like, how many times have you gone have you been to, like, CVS the past week? that will change. But there's so many traits about people that are also known to never change. Like, your risk tolerance doesn't really change over time. It's very consistent. So it's these things that we're trying to learn. But the scale we are operating is right now hundreds or, tens of thousands to hundreds of thousands. And in many of the core use cases that we are deployed in, and this is more than enough population, to cover those. Really, at that point, what you care about is less the number of people, but more do you have the right subpopulation of interest covered? And this is also the reason why people want a larger sample. It's not because they want, stronger statistical guarantees. It's more that can they filter down to any population of their interest. However, you can also imagine in 10 years, if we truly believe that the compute is going to scale, that we'll have much more availability for compute, and our ambition for simulation is also going to scale accordingly, there's definitely a reason for us to create an entire data center worth of simulations.Joon [00:48:35]: Or in my hunch here is I do think in the next some number of years, we will start creating simulations that will cost as much as training a foundation model. But perhaps it's going to be so valuable to the society that it would be a no-brainer. Right now, even today, like, we are training bunch of new foundation model just so we can say we trained one and we spent tens of millions. But if we can create a simulation at the level of society that would solve climate change, I would run that today. I would raise the money right now just to run that.Multi-Agent Simulation and Social InfluenceSwyx [00:49:10]: Amazing. the follow-up question is, does it also compound if you let the simulations talk to each other?Swyx [00:49:18]: Or do they already do that today? They don't, right, as far as I understand?Joon [00:49:22]: It depends on what simulation you're trying to run.Joon [00:49:24]: In the multi-agent simulation setup, the agents do talk to each other.Swyx [00:49:28]: Right, which is exactly Smallville, right?Joon [00:49:29]: That's right.Swyx [00:49:30]: But a lot of times, for example, in commerce, you're just by yourself, so there's no point talking. which is way cheaper.Vibhu [00:49:37]: But they use all these levels, right? Like, you decide what you will buy based on what other people around you buy and talk about, right?Swyx [00:49:43]: It depends.Vibhu [00:49:44]: It depends.Swyx [00:49:45]: Again, I'm, I'm coming at this from a cost point of view. I'm like, “Oh my God.” LikeVibhu [00:49:48]: I thinkSwyx [00:49:49]: If there is, like, some combinatorial thing of, like, thousands of people talking to thousands of people, then that one million X's might cost.Vibhu [00:49:56]: I have a very different view as the cost point aside. Like, running these studies in reality is a lot more expensive, right? Running any study like this is you gotta have people do it, you gotta sign people up. It's very expensive and sometimes, like, not feasible to run the study.Vibhu [00:50:14]: But the outcome or the decisions you make are very expensive on them, right? So spend X million on something that, the overall process costs 100 million might as well, right? There's, there's a lot of value to be had there. It's a small cost, but I'm excited on the cost side.Joon [00:50:33]: To some extent, and when you deploy technology, you often want to deploy in a way where you can replace existing budget or you can make things more efficient, and that is the best way to deploy. However, the way you capture the long-term value of the technology is making the argument that, no, it's the upside, that by making this better decision using simulation, you have saved yourself or made yourself hundreds of millions or even billions of dollars, and that's a case to be made.Vibhu [00:51:06]: Random tangent question. So if you're doing a lot of inference, a lot of model multi-agent stuff, are you at the point where it makes sense to, train a model that' very sparse? You're expecting to do multi-million dollar runs. Are you thinking about this in model architecture standpoint or inference efficiency, or, you're still at the research phase of it works, we're not super there yet?Joon [00:51:34]: Efficiency, we do think quite a bit about. this is technology that is deployed now in some of the largest enterprise companies in the world, and we do process significant number of queries, that are trying to, simulate the populations in the world. So efficiency is a consistent thing. we don't want to over-optimize too early, so I wouldn't say, like, this is the higher bid Right now, but this is definitely something that we think pretty carefully about.Swyx [00:52:05]: Yeah. Are there other case studies? So we, you talked about CVS, talked about Gallup, Deloitte, Wealthfront.Efficiency, Enterprise Use, and Real-World Case StudiesJoon [00:52:12]: Wealthfront is an interesting one, because one of the things they were trying to do, they were one of the first customers that wanted to do product testing that goes beyond just asking people what they think about, let's say, behavior experiments and so forth. So there, really what we had to do was reason about multimodal input, so images, but also you can also imagine, like, these agents traversing through Figma mockups or websites. So some of the things that our agents can also do is it can be given a domain, like, or, like, a website URL and go use it for a while. It's these things. And Wealthfront was one of the first, customers, that was very excited about this possibility.Vibhu [00:52:53]: What have people been asking? Like, is there any demand that we have not covered? Like, UI testing, right?Vibhu [00:52:59]: I wanna try a new. I wanna ship a new feature, test the UI, simulate how people will do it. Any interesting things that you're seeing demand for?Product Testing, Websites, and Synthetic PanelsJoon [00:53:08]: Today, a lot of the demand does come from like, the places where people have historically used human panels, we can now replace with agents, and these synthetic populations. And this is not replacing human panel. in many ways, the simulation that Simile is building is grounded. So the way that I think about this is we are trying to represent humanity at scale. And in that way, the use cases are what we would expect, but it's the scale of deployment that surprises me.Joon [00:53:44]: Turns out there are so many decisions that people make every day in these organizations, groups, and we want to be able to say, “We listen to people. We have consulted our users.” But in reality, that is rarely the case because getting to people and asking them many questions, it's difficult. It's both costly, time-consuming, but most importantly, people are just not available. If I had to answer 1000 survey questions for this one particular, vendor, even if I wanted to do that, like, I would never do it. And that's very much the case. What simulation can do is ensure that the voices of people are always represented in rooms where the decisions for them is made, right? So all the stakeholders of this particular product launch, ideally they're consulted. That's what this technology really is trying to enable.Market Size, TAM, and Human Decision-MakingSwyx [00:54:39]: In my mind, that means it skews towards more consumer focus, right? Like, anything with a wide enough customer base where you do benefit from the diversity that you represent. What are some rough statistics, just for people who are not familiar with this market in general, what's the market size that. I'm sure you have some, like, rough numbers. market size is, like, a vague questionSwyx [00:55:01]: But, like, how much do people spend?Joon [00:55:03]: So market research is a $100 billion industry.Joon [00:55:06]: But the thing about simulation is not a tool for market research. Simulation is a tool for human decision-making. So the question around what is a TAM here is quite tricky, right? Because it's easy to say, “Well, market research TAM is roughly 100 million or 100 billion.” so is it a TAM? And not really, right? Because in many ways, you're trying to inform all human decision-making. You're trying to inform every decision that are made about humans for humans. What is a TAM for that? It's really unclear. And I'll be honest. Like, I have a scientific background, I have a research background, so I didn't come into the field calculating, oh, what is the TAM for human decision-making? But I just had to assume, well, if we can inform every decision that is made about human for human, that has to be big.Swyx [00:55:58]: Some- something valuable.Joon [00:55:59]: Exactly.Swyx [00:55:59]: To some extent, you are a unicorn founder now, and you have to care as a CEO. But, like, I do think, like, yeah, when you go into these boardrooms with people that you're quoting millions of dollars of contracts for, like, you have to say, “Well, here's what you spend on humans-”Swyx [00:56:15]: “. And here's what we save you, and it's 85% similar.”Joon [00:56:19]: And certainly, the value case, is something that we care deeply about. Like, what is the value that we provide to the users and the decision-makers? But this is also where, like, as a founder, I think valuation only tells one very superficial aspect of the story, and I try not to think too much about valuation, in general, because that's not what also motivates a team or certainly doesn't. I'm, I-- Again, the interesting thing about researchers is we are happy living in academia, getting paid next to. we get paid okay. we don't get paid that much, as a researcher here in academia, but it's the impact and it's the, it's the value that we can provide to the individuals and the society that really drives us. And in that way, ultimately what drives us is the impact. Does the simulation we provide have a real impact in people's decision-making in ways that progresses our society forward? If the answer is yes, then yes. that has to be great business, and we see that in numbers, and we do care deeply about that upside story, but that's the heart of it.Where Simulation Goes NextVibhu [00:57:27]: Do you have any timeline predictions? So we talked about scaling laws of simulations.Vibhu [00:57:33]: You brought up, okay, maybe one day we can simulate how to solve climate change.Vibhu [00:57:38]: Where are we now?Vibhu [00:57:40]: If that's not the end state, what is an end state, and what does progress look like?Joon [00:57:45]: So what I sometimes tell people is simulation as industry, it feels a lot like where GPT-3.5, GPT-4 was, for the AGI saga, which is we have now technology that is powerful enough to do real damage on the verticals that we are tackling. At the same time, there's a lot of progress that is yet to come. And that's, I think, where this is. So the way I see it, I do think there will continue to be breakthroughs both in data, in algorithms, and there will be much more aggressive scaling that will also happen over the next few years. But I think that's roughly where we are.Swyx [00:58:27]: I think that was about the ro
2016年10月末,「声东击西」第一期节目上线;到2026年10月末,「声东击西」将更新整整十年。在这十年中,我们采访了很多身处变化之中的人,也和大家一起经历了一个不断变化和转向的世界。 十周年之际,我们的确想做一些特别节目,但不太想只做一次简单回顾。所以我们决定还是用「声东击西」熟悉的形式,通过访谈那些亲历结构性变化并拥有独特观察的嘉宾们,去听听过去这十年中,他们看见了什么,洞察到了什么,如何理解今天,又如何辨认出那些塑造未来的力量。 这一系列的节目会至少有6期,含括技术,城市,制度,教育等不同领域。 那作为「十周年」特别系列的第一期,徐涛前往硅谷,在那里采访了 AI 技术创业者贾扬清。 过去十多年,贾扬清几乎经历了人工智能深度学习发展的每一个关键阶段:从伯克利时期参与深度学习早期探索、开发开源框架 Caffe ,到先后加入 Google Brain、 Facebook AI Infra、阿里云,再到创办 AI 公司并被英伟达收购。如今,他又开启了新的 AI 创业探索。 这期节目,我们和贾扬清一起回望 AI 过去的关键节点:从技术突破,到产业变化,再到 AI 可能带来的未来,这是一场关于 AI 过去十年与未来十年的对话。 本期人物 贾扬清,AI 科学家、连续创业者,Intent Lab 创始人、CEO 徐涛,声动活泼联合创始人 主要话题 2005-2012 :从“人工智能已死”到开始“百花齐放” [05:34] “我本质有点像一个文科生” [08:04] 第一次听见「人工智能」时,“当时至少整个业界判定它的死亡了,大家认为人工智能是一个历史概念” [14:20] 一个各种各样观点在碰撞的阶段 [17:46] 一张快递盒送来的英伟达GPU ,一个 side project,和成为 AI 基础开源工具的 Caffe 的诞生 2013-2022 从实验室到科技巨头的押注和竞赛 [26:53] Google Brain,Facebook AI Research 与 AI 竞赛的开始 [37:56] 2016 年 AlphaGo 时刻,却离开 Google选择「乱七八糟、生机勃勃」的 Facebook [46:11] AI 领域还小,但为什么硅谷愿意为长期技术下注 [50:28] 从 Facebook 到阿里,“这个技术能不能广泛铺到所有领域中” 2022-2024 AI 的再一次惊人跃迁和浪潮中的创业机会 [01:05:05] GPT时刻,AI 的再一次跃迁 [01:12:38] 第一次创业:在 AI 淘金热中「卖铲子」 [01:17:07] 大公司的缓慢,和创业者的洞察 [01:25:35] Lepton AI 第二年实现盈利,很快被英伟达收购 2024- 公司、工作与人的重新定义 [01:32:37]被英伟达收购几个月后,二次创业的想法袭来 [01:37:56] 单个 Agent 足够聪明,一群 Agent 为什么还不是团队 [01:49:47] 当 AI 能够完成更多工作,公司需要怎样的人 [01:55:44] 当解决问题越来越快,人类还需要创造什么 [02:03:59] 从完成任务到获得信任,AI 进入社会还缺少什么 延伸解读 [12:44] 《How to Create a Mind: The Secret of Human Thought Revealed》 [15:20] AlexNet,2012年由 Alex Krizhevsky、Ilya Sutskever 和 Geoffrey Hinton 团队提出的卷积神经网络模型,在 ImageNet 图像识别比赛中取得突破,被认为开启了深度学习革命。ImageNet Classification with Deep Convolutional Neural Networks [17:46] Caffe,一个开源深度学习框架,帮助研究人员更快地设计、训练和验证神经网络模型。它降低了深度学习研究的门槛,也推动了早期深度学习社区的发展。 [21:23] Google Brain,Google 于2010年代初建立的人工智能研究团队,目标是探索大规模神经网络和机器学习技术。它推动了 TensorFlow 等基础设施的发展。 [29:16] ImageNet,由斯坦福大学李飞飞团队推动的大规模图像数据库和评测体系。2012年前后,深度学习模型在 ImageNet 上取得巨大突破,证明了“大数据 + 大计算 + 神经网络”的路线。 [31:34] Neolab,近年来,一些新的 AI 实验室不再完全依附于大型科技公司,而是以创业公司形式进行前沿研究,例如 OpenAI、Anthropic 等。这代表 AI 基础研究组织模式的变化。 [38:08] TensorFlow,Google 开发的机器学习框架,帮助研究人员和企业构建、训练和部署 AI 模型。它代表了 AI 从实验室走向产业基础设施的重要一步。 The 2026 AI Index Report 也可以在小红书账号「徐涛-声东击西」看到更多相关内容和幕后 十周年特别节目 2016-2026,声东击西走过十年。 十年,我们亲历变化,也由此洞见未来。在这个特别系列中,我们邀请身处结构性变化中的人,回望过去十年的关键转折,也一起思考未来的方向。 本系列持续更新中: 第一期:一个 AI 从业者的十年——专访贾扬清 第二期:…… 给声东击西投稿 「声东击西」一直在寻找来自不同社会和群体的真实声音。我们曾经采访过为特朗普竞选生产 MAGA 帽子的中国制造商、记录过七位在美国大选中经历起伏的华人个体,也讲述了委内瑞拉青年的故事。 如果你也有一些特别的经历、观察或想法,不论是亲身体验的故事,还是你在某个行业、社区中的所见所闻,都欢迎你向我们投稿。 你的声音可能出现在未来的节目当中,我们非常期待你的分享! 投稿入口 「Knock Knock 世界」 从围棋游戏到《宝可梦》,科技公司为什么让 AI 玩游戏?https://sourl.co/kRDfdx 春晚耍刀弄剑的人形机器人,真能走进我们的日常生活了吗?https://sourl.co/pmkCMh 半程马拉松、运动会,为什么要办「机器人」体育比赛? https://sourl.co/NzgvvA 被全网追捧的「AI 龙虾」,到底是怎么火起来的?https://sourl.co/XxvuPR 在「Knock Knock 世界」里,听到全球新鲜事,还能成为「全球观察员」,报选题、参加选题会。2026 年的节目正在持续更新,有4期免费试听,苹果播客上还可以还【按月】随时订阅节目。 加入我们 声动活泼团队目前正在招聘内容监制、商业运营经理、商业发展经理和实习生,如果你也对播客行业的内容制作和商务运营感兴趣,欢迎投递! 详情点击招聘入口:加入声动活泼(在招职位速览) 幕后制作 后期:赛德 运营:George 设计:饭团 实习编辑:翔宇 商务合作 声动活泼商业化小队,点击链接可直达商务会客厅,也可发送邮件至 business@shengfm.cn 联系我们。 关于声动活泼 「用声音碰撞世界」,声动活泼致力于为人们提供源源不断的思考养料。 我们还有这些播客:声东击西、What's Next|科技早知道、商业WHY酱、跳进兔子洞&跳进兔子洞第三季、吃喝玩乐了不起、不止金钱、泡腾 VC、反潮流俱乐部 欢迎在即刻、微博等社交媒体上与我们互动,搜索声动活泼即可找到我们。 也欢迎你写邮件和我们联系,邮箱地址是:ting@sheng.fm 获取更多和声动活泼有关的讯息,你也可以扫码添加声小音,在节目之外和我们保持联系! Special Guest: 贾扬清.
Last week, World Labs announced its acquisition of SceniX, bringing together two teams working on one of AI's biggest unsolved problems: how to give machines a true understanding of the physical world. Martin Casado sits down with Fei-Fei Li, co-founder and CEO of World Labs, creator of ImageNet, and pioneer of spatial intelligence, alongside Yunzhu Li, co-founder of SceniX and assistant professor at Columbia University. They discuss why World Labs acquired SceniX, how simulation can unlock the next generation of robotics, and why training robots may require a fundamentally different approach than training language models. The conversation explores real-to-sim-to-real pipelines, world models, robotics foundation models, evaluation, synthetic data, and why the future of AI depends not just on understanding language—but on understanding and interacting with the physical world. Resources: Follow Fei-Fei Li on X: https://x.com/drfeifei Follow Yunzhu Li on X: https://x.com/YunzhuLiYZ Follow Martin Casado on X: https://x.com/martin_casado Stay Updated:Find a16z on YouTube: YouTubeFind a16z on XFind a16z on LinkedInListen to the a16z Show on SpotifyListen to the a16z Show on Apple PodcastsFollow our host: https://twitter.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
In a recent video interview, Brian Yavorsky, CTO at Imagenet, discusses how the company helps payers use AI and other data technologies to achieve efficiencies. Emphasizing the difficulties in getting skilled staff, Yavorsky says that automation can reduce labor not by removing people, but by automating areas that are difficult to staff. Plus, by automating the mundane tasks you allow staff to work on higher value tasks that they enjoy more.Many payers are burdened with legacy systems and workflows, especially after acquiring other companies. Although the systems work individually, the fragmentation in systems and data introduces inefficiencies and requires more manual work, which adds to costs and slows payer responses. Organizations that don't address this fragmentation suffer from inefficient processes and often can't benefit from new AI technologies.Learn more about Imagenet: https://www.imagenetglobal.com/Healthcare IT Community: https://www.healthcareittoday.com/
Best-known as the creator of ImageNet, we meet the Godmother of AI, Dr. Fei-Fei Li In the latest installment of our oral history project. She's a Chinese-American computer scientist and the creator of ImageNet - the dataset that made rapid advances possible in this field of AI that helps computers take meaningful information from things like photos and videos.We Meet: Stanford University's Fei-Fei Li, author of "The Worlds I See: Curiosity, Exploration, and Discovery at the Dawn of AI" and the founder of World LabsCredits:This episode of SHIFT was produced by Jennifer Strong with help from Emma Cillekens. It was mixed by Garret Lang, with original music from him and Jacob Gorski. Art by Meg Marco.
Parce que… c'est l'épisode 0x735! Shameless plug 14 au 17 avril 2026 - Botconf 2026 20 au 22 avril 2026 - ITSec Code rabais de 15%: Seqcure15 28 et 29 avril 2026 - Cybereco Cyberconférence 2026 9 au 17 mai 2026 - NorthSec 2026 3 au 5 juin 2026 - SSTIC 2026 19 septembre 2026 - Bsides Montréal 1 au 3 décembre 2026 - Forum INCYBER - Canada 2026 24 et 25 février 2027 - SéQCure 2027 Description Retour aux sources techniques Dans cet épisode, l'animateur retrouve Frédéric Grelot, expert en intelligence artificielle, qu'il n'avait pas croisé depuis un moment. Frédéric a quitté son poste de dirigeant chez Glimps, une entreprise qu'il avait exportée au Canada, pour retrouver ses premières amours : la technique et la recherche. Il a rejoint l'AMIAD (Agence pour la maîtrise de l'IA de défense), rattachée au ministère des Armées français, créée en 2024. Un retour assumé, les « mains dans le cambouis », comme il le dit lui-même. Le fil conducteur de cet échange tourne autour d'une idée provocatrice : peut-on faire de l'intelligence artificielle sans données ? Frédéric avertit d'emblée que la formule est volontairement accrocheuse, et que la réponse sera nuancée. Une histoire de l'IA en quelques étapes clés Pour comprendre où l'on va, Frédéric propose un retour sur les grandes ruptures qui ont façonné l'IA depuis une quarantaine d'années. 1989 – Les débuts des réseaux convolutifs. Le réseau LeNet-5, conçu pour lire des chiffres manuscrits sur des chèques, représente l'un des premiers exemples concrets de réseaux de neurones convolutifs. Ces réseaux fonctionnent en empilant des couches d'analyse : les premières détectent des formes simples (points, lignes, angles), les suivantes des structures plus complexes (roues, rétroviseurs, puis une voiture entière). Ce paradigme a dominé le domaine pendant environ vingt ans. 2012 – La double révolution. Deux événements simultanés ont provoqué une explosion du domaine. D'une part, Nvidia a démocratisé l'utilisation des GPU pour le calcul scientifique via son API CUDA, rendant accessibles des calculs matriciels massivement parallèles. D'autre part, le jeu de données ImageNet a été publié en accès libre — un million d'images réparties en 1 000 catégories — offrant à la communauté une base commune pour entraîner et évaluer des modèles. Ces deux facteurs combinés ont déclenché une effervescence considérable, notamment dans le domaine de la vision par ordinateur. 2017 – L'avènement des transformers. La publication du célèbre article Attention is all you need introduit une nouvelle architecture qui va s'imposer comme le standard de l'IA moderne. Contrairement aux approches séquentielles précédentes, le transformer analyse chaque mot d'une phrase en le mettant en relation avec tous les mots qui le précèdent, enrichissant progressivement le sens de chaque élément couche après couche. Cette capacité à saisir le contexte global d'une séquence est à la base de tous les grands modèles de langage actuels. Son principal défaut : un coût de calcul quadratique par rapport à la longueur des séquences. Doubler la longueur d'un texte quadruple le volume de calcul. Les recherches de ces huit dernières années ont largement porté sur la résolution de ce problème, avec des résultats impressionnants — certains modèles open source atteignent aujourd'hui des fenêtres de contexte d'un à deux millions de tokens. Novembre 2022 – ChatGPT et la démocratisation. La sortie de ChatGPT marque moins une rupture technologique qu'une rupture d'usage. En mettant dans les mains du grand public ce qui n'existait que dans des laboratoires, OpenAI a transformé une évolution technique en révolution sociale. Les questions d'hallucinations, de jailbreak et d'alignement des modèles ont alors émergé comme des enjeux majeurs. Les modèles de fondation : l'IA qui apprend à vivre avant de se spécialiser C'est ici que Frédéric introduit son concept central : les modèles de fondation. L'analogie qu'il utilise est parlante : un enfant qui grandit jusqu'à 20 ans — observant des formes, des visages, des papillons, faisant du cerf-volant — développe une compréhension générique du monde qui en fait un « excellent modèle de fondation ». Il sera ensuite capable d'apprendre un métier précis, comme la géométrie ou la rétroingénierie de code, en repartant de cette base solide plutôt que de zéro. Un modèle de fondation est entraîné sur des quantités massives de données brutes, sans nécessairement annoter chaque exemple. Une fois cette phase généraliste accomplie, on n'a plus besoin que d'un tout petit volume de données spécialisées et annotées pour l'amener à un niveau d'excellence sur une tâche précise. Là où il fallait autrefois des millions d'exemples étiquetés, quelques centaines suffisent désormais. Ce paradigme bouleverse les rapports de force autour de la donnée. Frédéric, qui conseillait autrefois à ses clients de « conserver précieusement toutes leurs données », leur dit aujourd'hui d'en garder un peu, de bonne qualité. Le reste n'est plus indispensable. Vers l'inférence sans entraînement La dernière évolution abordée est peut-être la plus spectaculaire : le zero-shot learning. Grâce à la richesse des modèles de fondation, il est aujourd'hui possible de montrer une seule image d'un objet inconnu — une voiture jamais vue pendant l'entraînement — et d'être immédiatement capable de la reconnaître dans d'autres photos. Aucun entraînement supplémentaire n'est nécessaire : le modèle comprend l'objet à partir d'un seul exemple. C'est en ce sens que l'on peut parler d'IA « sans données » : non pas qu'il n'y en ait jamais eu, mais que l'utilisateur final n'a plus besoin d'en fournir pour bénéficier de capacités autrefois réservées aux experts disposant de vastes bases de données annotées. Un domaine en perpétuelle ébullition La conversation aborde également la dynamique concurrentielle entre modèles propriétaires américains et modèles open source, notamment chinois (DeepSeek) et français (Mistral). Les contraintes imposées aux acteurs chinois en matière de puissance de calcul ont paradoxalement stimulé l'innovation, à travers des techniques comme la distillation, le pruning ou l'optimisation des architectures d'attention. L'épisode se conclut sur une note d'ouverture : les hallucinations reculent, les modèles apprennent à dire « je ne sais pas », et le champ continue d'évoluer à un rythme soutenu — autant de raisons de se retrouver pour un prochain épisode. Collaborateurs Nicolas-Loïc Fortin Frédéric Grelot Crédits Montage par Intrasecure inc Locaux virtuels par Riverside.fm
AI Unraveled: Latest AI News & Trends, Master GPT, Gemini, Generative AI, LLMs, Prompting, GPT Store
Listen to Full Audio at https://podcasts.apple.com/us/podcast/scientist-vs-storyteller-benchmarking-gpt-5-2-claude/id1684415169?i=1000752001078For years, Latent Diffusion Models—the tech behind Stable Diffusion and DALL-E—have relied on a bit of an 'art form' called KL-regularization. Basically, researchers had to manually guess how much to compress an image before the AI started to lose the details. If you compressed too much, the image got blurry. Too little, and the model became too expensive to train.Enter Unified Latents, or UL.In a new paper out of DeepMind Amsterdam, researchers have introduced a framework that replaces that guesswork with a single, cohesive mathematical objective. Instead of training the compressor and the generator separately, UL trains the Encoder, the Prior, and the Decoder all at once.The 'Secret Sauce' here is something called Fixed Gaussian Noise Encoding. By injecting a constant, specific amount of noise during the encoding process, DeepMind has created a 'Maximum Precision Link.' This forces the encoder to be incredibly efficient, focusing only on the most important structures of an image.The results are staggering: UL achieved a state-of-the-art Video Distance score on the Kinetics-600 dataset and hit a competitive 1.4 FID on ImageNet—all while using significantly less computational power than traditional methods.This episode is made possible by our sponsors:
In this episode of the Crazy Wisdom podcast, host Stewart Alsop sits down with Kelvin Lwin for their second conversation exploring the fascinating intersection of AI and Buddhist cosmology. Lwin brings his unique perspective as both a technologist with deep Silicon Valley experience and a serious meditation practitioner who's spent decades studying Buddhist philosophy. Together, they examine how AI development fits into ancient spiritual prophecies, discuss the dangerous allure of LLMs as potentially "asura weapons" that can mislead users, and explore verification methods for enlightenment claims in our modern digital age. The conversation ranges from technical discussions about the need for better AI compilers and world models to profound questions about humanity's role in what Lwin sees as an inevitable technological crucible that will determine our collective spiritual evolution. For more information about Kelvin's work on attention training and AI, visit his website at alin.ai. You can also join Kelvin for live meditation sessions twice daily on Clubhouse at clubhouse.com/house/neowise.Timestamps00:00 Exploring AI and Spirituality05:56 The Quest for Enlightenment Verification11:58 AI's Impact on Spirituality and Reality17:51 The 500-Year Prophecy of Buddhism23:36 The Future of AI and Business Innovation32:15 Exploring Language and Communication34:54 Programming Languages and Human Interaction36:23 AI and the Crucible of Change39:20 World Models and Physical AI41:27 The Role of Ontologies in AI44:25 The Asura and Deva: A Battle for Supremacy48:15 The Future of Humanity and AI51:08 Persuasion and the Power of LLMs55:29 Navigating the New Age of TechnologyKey Insights1. The Rarity of Polymath AI-Spirituality Perspectives: Kelvin argues that very few people are approaching AI through spiritual frameworks because it requires being a polymath with deep knowledge across multiple domains. Most people specialize in one field, and combining AI expertise with Buddhist cosmology requires significant time, resources, and academic background that few possess.2. Traditional Enlightenment Verification vs. Modern Claims: There are established methods for verifying enlightenment claims in Buddhist traditions, including adherence to the five precepts and overcoming hell rebirth through karmic resolution. Many modern Western practitioners claiming enlightenment fail these traditional tests, often changing the criteria when they can't meet the original requirements.3. The 500-Year Buddhist Prophecy and Current Timing: We are approximately 60 years into a prophesied 500-year period where enlightenment becomes possible again. This "startup phase of Buddhism revival" coincides with technological developments like the internet and AI, which are seen as integral to this spiritual renaissance rather than obstacles to it.4. LLMs as UI Solution, Not Reasoning Engine: While LLMs have solved the user interface problem of capturing human intent, they fundamentally cannot reason or make decisions due to their token-based architecture. The technology works well enough to create illusion of capability, leading people down an asymptotic path away from true solutions.5. The Need for New Programming Paradigms: Current AI development caters too much to human cognitive limitations through familiar programming structures. True advancement requires moving beyond human-readable code toward agent-generated languages that prioritize efficiency over human comprehension, similar to how compilers already translate high-level code.6. AI as Asura Weapon in Spiritual Warfare: From Buddhist cosmological perspective, AI represents an asura (demon-realm) tool that appears helpful but is fundamentally wasteful and disruptive to human consciousness. Humanity exists as the battleground between divine and demonic forces, with AI serving as a weapon that both sides employ in this cosmic conflict.7. 2029 as Critical Convergence Point: Multiple technological and spiritual trends point toward 2029 as when various systems will reach breaking points, forcing humanity to either transcend current limitations or be consumed by them. This timing aligns with both technological development curves and spiritual prophecies about transformation periods.
Julian Sequeira from PyBites joins Sean and Kelly to share their top holiday gift picks for coders, makers, and educators. This episode features 15+ gift ideas ranging from budget-friendly maker tools to classroom robots—plus book recommendations, coding platforms, and a few surprises. Show Notes Wins of the Week Julian: Staying focused on "the one thing" at PyBites, plus 3D printing a custom cappuccino stencil for his local café Kelly: Surviving a muddy, clay-covered hill in North Carolina while on vacation Sean: Designing and 3D printing a custom bracket for his screen door using Fusion 360 Holiday Gift Ideas Julian's Picks Hoverboard with Go-Kart Attachment (~$299 AUD) - Two-wheeled self-balancing boards that can convert to a go-kart with a third wheel attachment. Available at Hoveroo (https://hoveroo.com.au) in Australia. Secret Coders Book Series (~$10-20 USD each) - A six-book graphic novel series that wraps coding puzzles and concepts into mystery stories. Recommended by Faye Shaw from the Boston PyLadies community. Great for ages 8-15. 3D Printer (~$200-300 USD) - Entry-level printers like the Bambu Lab A1 Mini or Elegoo Neptune 4 Pro have dropped significantly in price. Look for auto bed leveling as a key feature. Duolingo Chess (~$13/month with subscription) - A new addition to Duolingo that teaches chess tactics, strategy, and formal terminology through structured lessons. Great for building problem-solving skills. Classic Video Games (Zelda, Pokémon) - Story-driven games that build resilience and problem-solving skills, as an alternative to dopamine-heavy platforms like Roblox. Kelly's Picks Soccer Bot (~$59.99) - An indoor soccer training robot that challenges footwork skills. Works best on hard floors. "The Worlds I See" by Dr. Fei-Fei Li - Memoir of the computer scientist behind ImageNet and modern image recognition, covering her immigrant journey and rise in AI. A must-read for anyone interested in AI. LEGO Retro Radio Building Set (~$99) - A 1970s-style radio that you build, then insert your phone to play music. Features working dials that create authentic radio crackle sounds. Spydroid Loco Hex Robot (classroom investment) - A large spider-shaped robot that codes in Python and block programming. Features LIDAR and AI-based mapping. Seen at ISTE. Richtie Mini from Hugging Face ($299-$449) - An adorable AI desktop companion robot with onboard models. Two versions: one that connects to your computer and one that's self-contained. Sean's Picks LED Pucks (LED 001 Kit) (~$6-13) - Small USB-powered LED discs perfect for 3D printed projects like planet lamps. Available from Bambu Labs or Amazon. RGB versions include remote controls. Daily Desk Calendar (~$15-20) - A throwback gift that provides daily doses of humor, trivia, or inspiration. Suggestions include The Far Side, "They Can Talk," or "How to Win Friends and Influence People." PyBites Coding Platform (subscription) - Bite-sized Python challenges for sharpening coding skills. Great for teachers, students, and professionals looking for practical coding practice. Digital Calipers (~$40-50) - USB-rechargeable precision measuring tools essential for 3D printing and maker projects. Great for teaching geometry and measurement concepts. Deburring Tool (~$10) - A small tool with a curved swiveling blade for cleaning up 3D prints. A quality-of-life improvement for any maker's toolkit. Links Mentioned PyBites (https://pybit.es) - Python coaching and coding challenges Hoveroo (https://hoveroo.com.au) - Hoverboards (Australia) Bambu Lab (https://bambulab.com) - 3D printers and LED pucks Printables (https://www.printables.com) - 3D printing models MakerWorld (https://makerworld.com) - 3D printing models Hugging Face Richtie Mini (https://huggingface.co) - AI companion robot Duolingo (https://duolingo.com) - Language learning app with chess Secret Coders book series - Available on Amazon "The Worlds I See" by Dr. Fei-Fei Li - Available at bookstores Upcoming Events PyCon US 2026 - Long Beach, California Education Summit - Proposals open after the holidays, deadline around March/April Submit proposals when the website opens! Special Guest: Julian Sequeira.
Fei-Fei Li is a Stanford professor, co-director of Stanford Institute for Human-Centered Artificial Intelligence, and co-founder of World Labs. She created ImageNet, the dataset that sparked the deep learning revolution. Justin Johnson is her former PhD student, ex-professor at Michigan, ex-Meta researcher, and now co-founder of World Labs.Together, they just launched Marble—the first model that generates explorable 3D worlds from text or images.In this episode Fei-Fei and Justin explore why spatial intelligence is fundamentally different from language, what's missing from current world models (hint: physics), and the architectural insight that transformers are actually set models, not sequence models. Resources:Follow Fei-Fei on X: https://x.com/drfeifeiFollow Justin on X: https://x.com/jcjohnssFollow Shawn on X: https://x.com/swyxFollow Alessio on X: https://x.com/fanahova Stay Updated:If you enjoyed this episode, please be sure to like, subscribe, and share with your friends.Follow a16z on X: https://x.com/a16zFollow a16z on LinkedIn:https://www.linkedin.com/company/a16zFollow the a16z Podcast on Spotify: https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYXFollow the a16z Podcast on Apple Podcasts: https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details, please see http://a16z.com/disclosures. Stay Updated:Find a16z on XFind a16z on LinkedInListen to the a16z Podcast on SpotifyListen to the a16z Podcast on Apple PodcastsFollow our host: https://twitter.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Fei-Fei Li and Justin Johnson are cofounders of World Labs, who have recently launched Marble (https://marble.worldlabs.ai/), a new kind of generative “world model” that can create editable 3D environments from text, images, and other spatial inputs. Marble lets creators generate persistent 3D worlds, precisely control cameras, and interactively edit scenes, making it a powerful tool for games, film, VR, robotics simulation, and more. In this episode, Fei-Fei and Justin share how their journey from ImageNet and Stanford research led to World Labs, why spatial intelligence is the next frontier after LLMs, and how world models could change how machines see, understand, and build in 3D. We discuss: The massive compute scaling from AlexNet to today and why world models and spatial data are the most compelling way to “soak up” modern GPU clusters compared to language alone. What Marble actually is: a generative model of 3D worlds that turns text and images into editable scenes using Gaussian splats, supports precise camera control and recording, and runs interactively on phones, laptops, and VR headsets. Fei-fei's essay (https://drfeifei.substack.com/p/from-words-to-worlds-spatial-intelligence) on spatial intelligence as a distinct form of intelligence from language: from picking up a mug to inferring the 3D structure of DNA, and why language is a lossy, low-bandwidth channel for describing the rich 3D/4D world we live in. Whether current models “understand” physics or just fit patterns: the gap between predicting orbits and discovering F=ma, and how attaching physical properties to splats and distilling physics engines into neural networks could lead to genuine causal reasoning. The changing role of academia in AI, why Fei-Fei worries more about under-resourced universities than “open vs closed,” and how initiatives like national AI compute clouds and open benchmarks can rebalance the ecosystem. Why transformers are fundamentally set models, not sequence models, and how that perspective opens up new architectures for world models, especially as hardware shifts from single GPUs to massive distributed clusters. Real use cases for Marble today: previsualization and VFX, game environments, virtual production, interior and architectural design (including kitchen remodels), and generating synthetic simulation worlds for training embodied agents and robots. How spatial intelligence and language intelligence will work together in multimodal systems, and why the goal isn't to throw away LLMs but to complement them with rich, embodied models of the world. Fei-Fei and Justin's long-term vision for spatial intelligence: from creative tools for artists and game devs to broader applications in science, medicine, and real-world decision-making. — Fei-Fei Li X: https://x.com/drfeifei LinkedIn: https://www.linkedin.com/in/fei-fei-li-4541247 Justin Johnson X: https://x.com/jcjohnss LinkedIn: https://www.linkedin.com/in/justin-johnson-41b43664 Where to find Latent Space X: https://x.com/latentspacepod Substack: https://www.latent.space/ Chapters 00:00:00 Introduction and the Fei-Fei Li & Justin Johnson Partnership 00:02:00 From ImageNet to World Models: The Evolution of Computer Vision 00:12:42 Dense Captioning and Early Vision-Language Work 00:19:57 Spatial Intelligence: Beyond Language Models 00:28:46 Introducing Marble: World Labs' First Spatial Intelligence Model 00:33:21 Gaussian Splats and the Technical Architecture of Marble 00:22:10 Physics, Dynamics, and the Future of World Models 00:41:09 Multimodality and the Interplay of Language and Space 00:37:37 Use Cases: From Creative Industries to Robotics and Embodied AI 00:56:58 Hiring, Research Directions, and the Future of World Labs
Fei-Fei Li and Justin Johnson are cofounders of World Labs, who have recently launched Marble (https://marble.worldlabs.ai/), a new kind of generative “world model” that can create editable 3D environments from text, images, and other spatial inputs. Marble lets creators generate persistent 3D worlds, precisely control cameras, and interactively edit scenes, making it a powerful tool for games, film, VR, robotics simulation, and more. In this episode, Fei-Fei and Justin share how their journey from ImageNet and Stanford research led to World Labs, why spatial intelligence is the next frontier after LLMs, and how world models could change how machines see, understand, and build in 3D.We discuss:* The massive compute scaling from AlexNet to today and why world models and spatial data are the most compelling way to “soak up” modern GPU clusters compared to language alone.* What Marble actually is: a generative model of 3D worlds that turns text and images into editable scenes using Gaussian splats, supports precise camera control and recording, and runs interactively on phones, laptops, and VR headsets.* Fei-fei's essay:on spatial intelligence as a distinct form of intelligence from language: from picking up a mug to inferring the 3D structure of DNA, and why language is a lossy, low-bandwidth channel for describing the rich 3D/4D world we live in.* Whether current models “understand” physics or just fit patterns: the gap between predicting orbits and discovering F=ma, and how attaching physical properties to splats and distilling physics engines into neural networks could lead to genuine causal reasoning.* The changing role of academia in AI, why Fei-Fei worries more about under-resourced universities than “open vs closed,” and how initiatives like national AI compute clouds and open benchmarks can rebalance the ecosystem.* Why transformers are fundamentally set models, not sequence models, and how that perspective opens up new architectures for world models, especially as hardware shifts from single GPUs to massive distributed clusters.* Real use cases for Marble today: previsualization and VFX, game environments, virtual production, interior and architectural design (including kitchen remodels), and generating synthetic simulation worlds for training embodied agents and robots.* How spatial intelligence and language intelligence will work together in multimodal systems, and why the goal isn't to throw away LLMs but to complement them with rich, embodied models of the world.* Fei-Fei and Justin's long-term vision for spatial intelligence: from creative tools for artists and game devs to broader applications in science, medicine, and real-world decision-making.—Fei-Fei Li* X: https://x.com/drfeifei* LinkedIn: https://www.linkedin.com/in/fei-fei-li-4541247Justin Johnson* X: https://x.com/jcjohnss* LinkedIn: https://www.linkedin.com/in/justin-johnson-41b43664Where to find Latent Space* X: https://x.com/latentspacepodFull Video EpisodeTimestamps00:00:00 Introduction and the Fei-Fei Li & Justin Johnson Partnership00:02:00 From ImageNet to World Models: The Evolution of Computer Vision00:12:42 Dense Captioning and Early Vision-Language Work00:19:57 Spatial Intelligence: Beyond Language Models00:28:46 Introducing Marble: World Labs' First Spatial Intelligence Model00:33:21 Gaussian Splats and the Technical Architecture of Marble00:22:10 Physics, Dynamics, and the Future of World Models00:41:09 Multimodality and the Interplay of Language and Space00:37:37 Use Cases: From Creative Industries to Robotics and Embodied AI00:56:58 Hiring, Research Directions, and the Future of World Labs Get full access to Latent.Space at www.latent.space/subscribe
Dr. Fei-Fei Li is known as the “godmother of AI.” She's been at the center of AI's biggest breakthroughs for over two decades. She spearheaded ImageNet, the dataset that sparked the deep-learning revolution we're living right now, served as Google Cloud's Chief AI Scientist, directed Stanford's Artificial Intelligence Lab, and co-founded Stanford's Institute for Human-Centered AI. In this conversation, Fei-Fei shares the rarely told history of how we got here—including the wild fact that just nine years ago, calling yourself an AI company was basically a death sentence.We discuss:1. How ImageNet helped spark the AI explosion we're living through2. Why world models and spatial intelligence represent the next frontier in AI, beyond large language models3. Why Fei-Fei believes AI won't replace humans but will require us to take responsibility for ourselves4. The surprising applications of Marble, from movie production to psychological research5. Why robotics faces unique challenges compared with language models and what's needed to overcome them6. How to participate in AI regardless of your role—Brought to you by:Figma Make—A prompt-to-code tool for making ideas realJustworks—The all-in-one HR solution for managing your small business with confidenceSinch—Build messaging, email, and calling into your product—Transcript: https://www.lennysnewsletter.com/p/the-godmother-of-ai—My biggest takeaways (for paid newsletter subscribers):https://www.lennysnewsletter.com/i/178223233/my-biggest-takeaways-from-this-conversation—Where to find Dr. Fei-Fei Li• X: https://x.com/drfeifei• LinkedIn: https://www.linkedin.com/in/fei-fei-li-4541247• World Labs: https://www.worldlabs.ai—Where to find Lenny:• Newsletter: https://www.lennysnewsletter.com• X: https://twitter.com/lennysan• LinkedIn: https://www.linkedin.com/in/lennyrachitsky/—In this episode, we cover:(00:00) Introduction to Dr. Fei-Fei Li(05:31) The evolution of AI(09:37) The birth of ImageNet(17:25) The rise of deep learning(23:53) The future of AI and AGI(29:51) Introduction to world models(40:45) The bitter lesson in AI and robotics(48:02) Introducing Marble, a revolutionary product(51:00) Applications and use cases of Marble(01:01:01) The founder's journey and insights(01:10:05) Human-centered AI at Stanford(01:14:24) The role of AI in various professions(01:18:16) Conclusion and final thoughts—References: https://www.lennysnewsletter.com/p/the-godmother-of-ai—Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email podcast@lennyrachitsky.com.—Lenny may be an investor in the companies discussed. To hear more, visit www.lennysnewsletter.com
Fei-Fei Li and Justin Johnson are pioneers in AI. While the world has only recently witnessed a surge in consumer AI, they have long been laying the groundwork for the innovations transforming industries today.With the recent launch of Marble, the first product from their company World Labs, we are revisiting this conversation to explore the ideas that started it all. World Labs is focused on spatial intelligence, building Large World Models that can perceive, generate, and interact with the 3D world. Marble brings that vision to life, allowing anyone, from individual creators to major platforms, to generate 3D scenes directly from text or image prompts and turn complex 3D creation into a simple, creative process.In this episode, a16z general partner Martin Casado talks with Fei-Fei and Justin about the journey from early AI winters to the rise of deep learning and multimodal AI. From foundational breakthroughs like ImageNet to the cutting-edge realm of spatial intelligence, they discuss the evolution of the field and what is next for innovation at World Labs. Timecode:0:00 – The Next Decade of AI2:45 – Origins: Backgrounds of the Founders6:50 – The Rise of Deep Learning & ImageNet8:00 – Algorithmic Unlocks: Compute, Data, and Supervised Learning12:00 – From Predictive to Generative AI16:20 – The Journey to Spatial Intelligence18:35 – Defining Spatial Intelligence21:15 – 3D Data, Computer Vision, and Breakthroughs23:15 – Reconstruction vs. Generation in Computer Vision24:45 – Spatial Intelligence vs. Language Models29:00 – Applications: Virtual, Augmented, and Physical Worlds39:55 – Building World Labs: Team and Vision41:55 – The North Star: Measuring Success in Spatial Intelligence Resources:Learn more about World Labs: https://www.worldlabs.aiLearn more about Marble: https://Marble.WorldLabs.aiFind Fei-Fei on Twitter: https://x.com/drfeifeiFind Justin on Twitter: https://x.com/jcjohnssFind Martin on Twitter: https://x.com/martin_casado Stay Updated: If you enjoyed this episode, be sure to like, subscribe, and share with your friends!Find a16z on X: https://x.com/a16zFind a16z on LinkedIn: https://www.linkedin.com/company/a16zListen to the a16z Podcast on Spotify: https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYXListen to the a16z Podcast on Apple Podcasts: https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711Follow our host: https://x.com/eriktorenbergPlease note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures. Stay Updated:Find a16z on XFind a16z on LinkedInListen to the a16z Podcast on SpotifyListen to the a16z Podcast on Apple PodcastsFollow our host: https://twitter.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
If/Then: Research findings to help us navigate complex issues in business, leadership, and society
This week on If/Then, we're sharing an episode of What's Your Problem?, a show from Pushkin Industries where entrepreneurs, engineers, and scientists talk about the future they're trying to build—and the problems they must solve to get there. Hosted by former Planet Money co-host Jacob Goldstein, each conversation explores the challenges and breakthroughs shaping the next wave of innovation.In this episode, Goldstein speaks with Fei-Fei Li, Stanford computer scientist, former Chief Scientist of AI and Machine Learning at Google, and one of the most influential figures in the field of computer vision. Li reflects on her pioneering work developing ImageNet, the massive dataset that helped spark the modern AI revolution, and the “north star” questions that have guided her research from neuroscience to machine learning.Together, they trace how a single insight about how humans see the world led to a paradigm shift in artificial intelligence—and how Li's vision continues to shape the way we teach machines to see, learn, and collaborate with us.More Resources: • Fei Fei Li • Stanford Institute for Human-Centered Artificial Intelligence (HAI) • ImageNet • What's Your Problem?If/Then is a podcast from Stanford Graduate School of Business that examines research findings that can help us navigate the complex issues we face in business, leadership, and society.Chapters: (00:00:00) Introducing “What's Your Problem?” Kevin Cool introduces the Pushkin Industries podcast hosted by Jacob Goldstein.00:00:45 — What Is Computer Vision? Jacob Goldstein and Fei-Fei Li explain how machines learn to see and interpret images.00:03:18 — Real-World Uses of AI Vision Li shares examples from healthcare, robotics, and environmental science.00:05:06 — Discovering the Science of SeeingHow human vision research inspired Li's lifelong “north star” in AI.00:09:56 — Creating ImageNet Li builds a massive image database that transforms computer vision research.00:13:29 — Defining 30,000 Visual Concepts How cognitive science helped shape ImageNet's massive scale.00:16:41 — Building the Dataset by HandLi's team uses global crowdsourcing to label millions of images.00:19:38 — The 2012 Breakthrough Jeff Hinton's neural network shatters records and sparks the deep learning era.00:22:19 — Data Meets Hardware Li reflects on how big data and GPUs converged to power modern AI.00:24:55 — Lightning Round with Fei-Fei Li Quick insights on resilience, mentorship, and the future of human-AI collaboration.See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
Steven Forth is the Co-founder and Chief Value Officer at Ibbaka, a leading value and pricing consulting firm. With deep expertise in AI applications for pricing and value modeling, Steven is at the forefront of developing intelligent agents that help businesses understand and communicate value more effectively. His work focuses on the intersection of artificial intelligence, pricing strategy, and value creation, making him a pioneer in applying AI to solve complex pricing challenges. In this episode, Steven shares his insights on how benchmarking is revolutionizing both AI development and pricing strategy. Drawing parallels between how AI models are improved through benchmarking and how pricing models should be evaluated, he introduces a framework for measuring pricing effectiveness that could transform how we approach pricing decisions. Together with Mark, they explore the challenges of establishing "truth" in pricing, the role of synthetic data, and the future of AI-powered pricing tools. Why you have to check out today's podcast: Discover how AI benchmarking principles can revolutionize pricing model evaluation. Understand how to evaluate pricing models from both buyer and seller perspectives. Explore the future of AI-powered pricing tools and what it means for pricing professionals. "We don't start with the truth. We have to work our way towards truth through multiple iterations and applications." – Steven Forth Topics Covered: 02:15 – How Intercom's FinAI agent uses daily benchmarking to improve ticket resolution performance 05:30 – Why AI's success is built on benchmarking and how it emerged from the ImageNet competition 08:45 – The critical problem: pricing lacks standardized benchmarking like AI models have 11:20 – Michael Mansard's 12-factor pricing model assessment and its potential as an industry standard 14:10 – Why pricing models must be evaluated from both buyer and seller perspectives 17:25 – How market segmentation and use cases complicate pricing model benchmarking 20:40 – The role of synthetic data in pricing research and model validation 24:15 – Why "vibe coding" could disrupt traditional pricing consulting within 3 years 27:30 – The search for truth in pricing: hedonic pricing models and market assumptions 31:45 – Introduction to ValueIQ: Ibbaka's new AI agent for value-based selling Key Takeaways: "Anyone who says that they're data centric or data driven is actually before that they have to be model driven because they're using some form of model to organize the data." – Steven Forth "We should have done this 20 years ago. What were we thinking? Well, we weren't thinking. And we didn't have ways to do this for us anyway." – Steven Forth (on developing pricing benchmarks) "Benchmarking every day, I think, is going to be critical to the success of agents that do important business things." – Steven Forth "You can always improve your measurement, but at some point the return of improving the measurement is lower than the cost of increasing the validity of the measurement." – Steven Forth Resources and People Mentioned: Douglas Hubbard's How to Measure Anything (book): https://www.amazon.com/How-Measure-Anything-Intangibles-Business/dp/1118539273 ImageNet: https://image-net.org/ Michael Mansard's 12-Factor Pricing Model: https://www.insead.edu/bio/michael-mansard-0 Intercom's FinAI: https://www.intercom.com/help/en/articles/8205718-fin-ai-agent-resolutions Lovable, Replit, Bolt: https://linkblink.medium.com/bolt-vs-cursor-vs-replit-vs-lovable-ai-coders-comparison-guide-3b9d41e75810 ValueIQ: https://www.ibbaka.com/ibbaka-market-blog/get-ready-for-valueiq-sign-up-now-for-beta-access Connect with Steven Forth: LinkedIn: https://www.linkedin.com/in/stevenforth/ Email: steven@ibbaka.com Connect with Mark Stiving: LinkedIn: https://www.linkedin.com/in/stiving/ Email: mark@impactpricing.com
No Priors: Artificial Intelligence | Machine Learning | Technology | Startups
In this episode of No Priors, Sarah and Elad are joined by Dr. Fei-Fei Li, AI pioneer, co-director of Stanford's Human-Centered AI Institute, and founder of World Labs. Fei-Fei shares why she's building at the intersection of embodiment and intelligence, and what today's AI systems are still missing. From the early days of ImageNet to her vision for the next generation of robotics, she unpacks the human and technical motivations behind World Labs. They also discuss the challenges of 3D world modeling, her approach to building exceptional teams, and the special qualities that have led her students like Andrej Karpathy to make major breakthroughs. Show Notes: 0:00 Why and what Dr. Fei-Fei Li is building 3:00 World models at World Labs 6:44 Missing gaps in the AI future 9:16 Robotics and physical intelligence 16:15 Greatest challenges of 3D 19:08 Fei-Fei's work in PhD in ImageNet 23:05 Special moments in Dr. Li's career 29:33 Building teams 32:05 Human-centered AI
We chat with Emily Bender and Alex Hanna — authors of AI Con: How to Fight Big Tech's Hype and Create the Future We Want — and pierce the veil of hype by getting into how these systems actually work and, importantly, the work they cannot do despite claims by boosters and doomers alike. Think of datasets like ImageNet or LAION-5B as big vats of pink slime and LLMs like ChatGPT as “synthetic text extruding machines” that turn pink slime into nuggets of text. It's easy to forget that these magical mystery machines are direct descendants of very unexciting things like “T9 word.” We end the episode by chatting about why we shouldn't trust the hype about how AI is going to destroy (or revolutionize) the education sector. ••• The AI Con | Emily Bender and Alex Hanna https://thecon.ai/ ••• On the genealogy of machine learning datasets: A critical history of ImageNet https://journals.sagepub.com/doi/full/10.1177/20539517211035955 ••• Mystery AI Hype Theater 3000 https://www.dair-institute.org/maiht3k/ Standing Plugs: ••• Order Jathan's new book: https://www.ucpress.edu/book/9780520398078/the-mechanic-and-the-luddite ••• Subscribe to Ed's substack: https://substack.com/@thetechbubble ••• Subscribe to TMK on patreon for premium episodes: https://www.patreon.com/thismachinekills Hosted by Jathan Sadowski (bsky.app/profile/jathansadowski.com) and Edward Ongweso Jr. (www.x.com/bigblackjacobin). Production / Music by Jereme Brown (bsky.app/profile/jebr.bsky.social)
Artificial intelligence is evolving at an unprecedented pace—what does that mean for the future of technology, venture capital, business, and even our understanding of ourselves? Award-winning journalist and writer Anil Ananthaswamy joins us for our latest episode to discuss his latest book Why Machines Learn: The Elegant Math Behind Modern AI.Anil helps us explore the journey and many breakthroughs that have propelled machine learning from simple perceptrons to the sophisticated algorithms shaping today's AI revolution, powering GPT and other models. The discussion aims to demystify some of the underlying mathematical concepts that power modern machine learning, to help everyone grasp this technology impacting our lives–even if your last math class was in high school. Anil walks us through the power of scaling laws, the shift from training to inference optimization, and the debate among AI's pioneers about the road to AGI—should we be concerned, or are we still missing key pieces of the puzzle? The conversation also delves into AI's philosophical implications—could understanding how machines learn help us better understand ourselves? And what challenges remain before AI systems can truly operate with agency?If you enjoy this episode, please subscribe and leave us a review on your favorite podcast platform. Sign up for our newsletter at techsurgepodcast.com for exclusive insights and updates on upcoming TechSurge Live Summits.Links:Read Why Machines Learn, Anil's latest book on the math behind AIhttps://www.amazon.com/Why-Machines-Learn-Elegant-Behind/dp/0593185749Learn more about Anil Ananthaswamy's work and writinghttps://anilananthaswamy.com/Watch Anil Ananthaswamy's TED Talk on AI and intelligencehttps://www.ted.com/speakers/anil_ananthaswamyDiscover the MIT Knight Science Journalism Fellowship that shaped Anil's AI researchhttps://ksj.mit.edu/Understand the Perceptron, the foundation of neural networkshttps://en.wikipedia.org/wiki/PerceptronRead about the Perceptron Convergence Theorem and its significancehttps://www.nature.com/articles/323533a0
Elon Musk a présenté Grok 3 comme l'IA "la plus intelligente sur Terre", mais cette affirmation tient-elle la route ? Avec la multiplication des intelligences artificielles, de ChatGPT à Mistral en passant par Grok ou Perplexity, une question revient sans cesse : quelle est la meilleure ? Pourtant, vouloir les comparer de manière globale n'a pas vraiment de sens, car chaque IA a ses propres spécificités et excelle dans certains domaines tout en montrant des limites dans d'autres.Performance, véracité des réponses, rapidité, coût, impact environnemental... Sur quels critères comparer ? En outre, chaque utilisateur a ses propres attentes et biais, influençant ainsi la perception de la "meilleure" IA. Il existe des outils de classement, comme Chatbot Arena ou le français compareia.beta.gouv.fr, qui permettent de comparer les IA à l'aveugle en se focalisant sur la qualité des réponses. Par ailleurs, des benchmarks techniques comme GLU, SQUAD ou ImageNet apportent des évaluations plus précises sur des compétences spécifiques.Cependant, il est difficile de dire qu'une IA est globalement meilleure qu'une autre. Certaines excellent en traduction, d'autres en génération de code, en recherche d'actualité ou en création de contenu. Plutôt que de chercher une IA universellement supérieure, mieux vaut identifier celle qui correspond le mieux à chaque besoin précis.Liens : https://lmarena.ai/https://www.comparia.beta.gouv.fr/Mots-clés : intelligence artificielle, IA, Grok 3, Elon Musk, ChatGPT, Mistral, Perplexity, comparatif IA, benchmark IA, chatbot arena, DINUM, compareia, GPT-4, IA générative, machine learning, modèle de langage-----------♥️ Soutenez Monde Numérique : https://donorbox.org/monde-numerique
Prof. Jakob Foerster, a leading AI researcher at Oxford University and Meta, and Chris Lu, a researcher at OpenAI -- they explain how AI is moving beyond just mimicking human behaviour to creating truly intelligent agents that can learn and solve problems on their own. Foerster champions open-source AI for responsible, decentralised development. He addresses AI scaling, goal misalignment (Goodhart's Law), and the need for holistic alignment, offering a quick look at the future of AI and how to guide it.SPONSOR MESSAGES:***CentML offers competitive pricing for GenAI model deployment, with flexible options to suit a wide range of models, from small to large-scale deployments. Check out their super fast DeepSeek R1 hosting!https://centml.ai/pricing/Tufa AI Labs is a brand new research lab in Zurich started by Benjamin Crouzier focussed on o-series style reasoning and AGI. They are hiring a Chief Engineer and ML engineers. Events in Zurich. Goto https://tufalabs.ai/***TRANSCRIPT/REFS:https://www.dropbox.com/scl/fi/yqjszhntfr00bhjh6t565/JAKOB.pdf?rlkey=scvny4bnwj8th42fjv8zsfu2y&dl=0 Prof. Jakob Foersterhttps://x.com/j_foersthttps://www.jakobfoerster.com/University of Oxford Profile: https://eng.ox.ac.uk/people/jakob-foerster/Chris Lu:https://chrislu.page/TOC1. GPU Acceleration and Training Infrastructure [00:00:00] 1.1 ARC Challenge Criticism and FLAIR Lab Overview [00:01:25] 1.2 GPU Acceleration and Hardware Lottery in RL [00:05:50] 1.3 Data Wall Challenges and Simulation-Based Solutions [00:08:40] 1.4 JAX Implementation and Technical Acceleration2. Learning Frameworks and Policy Optimization [00:14:18] 2.1 Evolution of RL Algorithms and Mirror Learning Framework [00:15:25] 2.2 Meta-Learning and Policy Optimization Algorithms [00:21:47] 2.3 Language Models and Benchmark Challenges [00:28:15] 2.4 Creativity and Meta-Learning in AI Systems3. Multi-Agent Systems and Decentralization [00:31:24] 3.1 Multi-Agent Systems and Emergent Intelligence [00:38:35] 3.2 Swarm Intelligence vs Monolithic AGI Systems [00:42:44] 3.3 Democratic Control and Decentralization of AI Development [00:46:14] 3.4 Open Source AI and Alignment Challenges [00:49:31] 3.5 Collaborative Models for AI DevelopmentREFS[[00:00:05] ARC Benchmark, Chollethttps://github.com/fchollet/ARC-AGI[00:03:05] DRL Doesn't Work, Irpanhttps://www.alexirpan.com/2018/02/14/rl-hard.html[00:05:55] AI Training Data, Data Provenance Initiativehttps://www.nytimes.com/2024/07/19/technology/ai-data-restrictions.html[00:06:10] JaxMARL, Foerster et al.https://arxiv.org/html/2311.10090v5[00:08:50] M-FOS, Lu et al.https://arxiv.org/abs/2205.01447[00:09:45] JAX Library, Google Researchhttps://github.com/jax-ml/jax[00:12:10] Kinetix, Mike and Michaelhttps://arxiv.org/abs/2410.23208[00:12:45] Genie 2, DeepMindhttps://deepmind.google/discover/blog/genie-2-a-large-scale-foundation-world-model/[00:14:42] Mirror Learning, Grudzien, Kuba et al.https://arxiv.org/abs/2208.01682[00:16:30] Discovered Policy Optimisation, Lu et al.https://arxiv.org/abs/2210.05639[00:24:10] Goodhart's Law, Goodharthttps://en.wikipedia.org/wiki/Goodhart%27s_law[00:25:15] LLM ARChitect, Franzen et al.https://github.com/da-fr/arc-prize-2024/blob/main/the_architects.pdf[00:28:55] AlphaGo, Silver et al.https://arxiv.org/pdf/1712.01815.pdf[00:30:10] Meta-learning, Lu, Towers, Foersterhttps://direct.mit.edu/isal/proceedings-pdf/isal2023/35/67/2354943/isal_a_00674.pdf[00:31:30] Emergence of Pragmatics, Yuan et al.https://arxiv.org/abs/2001.07752[00:34:30] AI Safety, Amodei et al.https://arxiv.org/abs/1606.06565[00:35:45] Intentional Stance, Dennetthttps://plato.stanford.edu/entries/ethics-ai/[00:39:25] Multi-Agent RL, Zhou et al.https://arxiv.org/pdf/2305.10091[00:41:00] Open Source Generative AI, Foerster et al.https://arxiv.org/abs/2405.08597
Fei-Fei Li is a pioneering AI scientist breaking new ground in computer vision, a Stanford professor, and currently leading the innovative start-up World Labs. While her career is deeply rooted in technical expertise, Dr. Li's journey is driven by an insatiable curiosity. In this episode, Brad and Dr. Li reflect on poignant moments from her memoir, "The Worlds I See: Curiosity, Exploration, and Discovery at the Dawn of AI," highlighting the crucial role of keeping humanity at the center of AI development. They also explore how government-funded academic research, driven by curiosity rather than profits, can lead to unexpected and profound discoveries that propel innovation and economic opportunities.Click here for the episode transcript.
Anil Ananthaswamy is a renowned science writer and journalist who has written extensively on various scientific topics. In his latest book "Why Machines Learn", Anil explores the fascinating world of artificial intelligence and machine learning. He reveals the intricate mechanisms and complex algorithms that underlie these cutting-edge technologies. Join us for a fascinating conversation with science writer Anil Ananthaswamy as he shares insights from his book and sheds light on the rapidly evolving field of AI. Tune in to gain a deeper understanding of how these machines work at a basic mathematics level. Resource List - Why Machines Learn, book by Anil Ananthaswamy - https://amzn.in/d/bmirU45 Dartmouth Summer Research Project on Artificial Intelligence - https://home.dartmouth.edu/about/artificial-intelligence-ai-coined-dartmouth What is the Perceptron artificial neural network? - https://www.geeksforgeeks.org/what-is-perceptron-the-simplest-artificial-neural-network/ Read about the McCulloch-Pitts Artificial Neuron - https://towardsdatascience.com/mcculloch-pitts-model-5fdf65ac5dd1 Nobel Prize in Physics 2024 - https://www.nobelprize.org/prizes/physics/2024/press-release/ What is the Hopfield Neural Network? - https://www.geeksforgeeks.org/hopfield-neural-network/ Read about Backpropagation - https://en.wikipedia.org/wiki/Backpropagation “Learning representations by back-propagating errors”, paper by Geoffrey Hinton, David Rumelhart and Ronald Williams - https://www.nature.com/articles/323533a0 AlexNet by Geoffery Hinton and team - https://en.wikipedia.org/wiki/AlexNet What is ImageNet? - https://www.image-net.org/about.php ‘Attention Is All You Need', transformer architecture paper - https://proceedings.neurips.cc/paper_files/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf What are Neural Scaling Laws? - https://en.wikipedia.org/wiki/Neural_scaling_law NeuroAI - https://neuro-ai.com/ About SparX by Mukesh Bansal SparX is a podcast where we delve into cutting-edge scientific research, stories from impact-makers and tools for unlocking the secrets to human potential and growth. We believe that entrepreneurship, fitness and the science of productivity is at the forefront of the India Story; the country is at the cusp of greatness and at SparX, we wish to make these tools accessible for every generation of Indians to be able to make the most of the opportunities around us. In a new episode every Sunday, our host Mukesh Bansal (Founder Myntra and Cult.fit) will talk to guests from all walks of life and also break down everything he's learnt about the science of impact over the course of his 20-year long career. This is the India Century, and we're enthusiastic to start this journey with you. Follow us on Instagram: / sparxbymukeshbansal Website: https://www.sparxbymukeshbansal.com You can also listen to SparX on all audio platforms Fasion | Outbreak | Courtesy EpidemicSound.com
How can we use AI to amplify human potential and build a better future? And what exactly does “AGI” even mean? To kick off Possible's fourth season, Reid and Aria sit down with world-renowned computer scientist Fei-Fei Li, whose work in artificial intelligence over the past several decades has earned her the nickname “the godmother of AI.” An entrepreneur and professor, Fei-Fei shares her journey from creating ImageNet, a massive dataset of labeled images that revolutionized computer vision, to her current role as co-founder and CEO of the spatial intelligence startup World Labs. She explains why spatial intelligence—the ability to perceive and interact with the 3D world—is so crucial for AI's development and how it could lead to breakthroughs in fields like medicine, climate, and education. They get into regulatory guardrails, governance, and what it will take to build a positive, human-centered AI future for all. For more info on the podcast and transcripts of all the episodes, visit https://www.possible.fm/podcast/ Topics: 1:55 - Hellos and intros 3:43 - ImageNet and the interplay between data and models 6:06 - World Labs and spatial intelligence 10:03 - Boundaries between 3D physical and digital worlds 11:50 - The difference between LLMs and LWMs 13:02 - What humans are capable of creating with technology 14:04 - Key principles of AI: human agency and respect 17:16 - Stanford Institute for Human-Centered AI 19:13 - What this moment in AI means for humanity 21:06 - Cross-sector collaboration 25:10 - AI4ALL program and the importance of diversity in AI development 27:00 - Midroll ad break 27:09 - Using AI to improve healthcare delivery and treatment 30:20 - Founding history of AI and the meaning of the term “AGI” 33:00 - Future of agentic AI and voice 34:42 - Fei-Fei's mentor and his advice 37:18 - Rapid-fire questions Possible is an award-winning podcast that sketches out the brightest version of the future—and what it will take to get there. Most of all, it asks: what if, in the future, everything breaks humanity's way? Tune in for grounded and speculative takes on how technology—and, in particular, AI—is inspiring change and transforming the future. Hosted by Reid Hoffman and Aria Finger, each episode features an interview with an ambitious builder or deep thinker on a topic, from art to geopolitics and from healthcare to education. These conversations also showcase another kind of guest: AI. Whether it's Inflection's Pi, OpenAI's ChatGPT or other AI tools, each episode will use AI to enhance and advance our discussion about what humanity could possibly get right if we leverage technology—and our collective effort—effectively.
We're experimenting and would love to hear from you!In this episode of 'Discover Daily' by Perplexity, hosts Isaac and Sienna explore NASA's upcoming Lunar Trailblazer mission, scheduled for January 2025. This compact satellite mission aims to map water resources on the Moon's surface using advanced instruments like the High-resolution Volatiles and Minerals Moon Mapper and the Lunar Thermal Mapper. The mission represents a crucial step in NASA's Artemis program, designed to establish sustainable human presence on the MoonThe show delves into a groundbreaking development in robotics, highlighting Chinese startup AgiBot's release of the AgiBot World Alpha dataset. This comprehensive open-source collection features over one million trajectories from 100 robots in industrial-grade environments, potentially marking an 'ImageNet moment' for embodied intelligence in roboticsThe main story focuses on Microsoft and OpenAI's unconventional redefinition of Artificial General Intelligence (AGI), which ties achievement to a $100 billion profit milestone. The episode examines the implications of this profit-centric definition, Microsoft's diversification strategy in AI investments, and the complex dynamics of their partnership agreement. This innovative approach to defining AGI raises important questions about the future direction of AI development and its impact on the tech industryFrom Perplexity's Discover Feed: https://www.perplexity.ai/page/nasa-s-moon-micro-mission-Bua4as.9SCi.G_fZKZUCPAhttps://www.perplexity.ai/page/agibot-s-humanoid-robot-traini-ovKJpg2RSey1INdEnXwuNwhttps://www.perplexity.ai/page/microsoft-s-100b-agi-definitio-e6FaEhReQs.9exHMGZpuogPerplexity is the fastest and most powerful way to search the web. Perplexity crawls the web and curates the most relevant and up-to-date sources (from academic papers to Reddit threads) to create the perfect response to any question or topic you're interested in. Take the world's knowledge with you anywhere. Available on iOS and Android Join our growing Discord community for the latest updates and exclusive content. Follow us on: Instagram Threads X (Twitter) YouTube Linkedin
Analysis of image classifiers demonstrates that it is possible to understand backprop networks at the task-relevant run-time algorithmic level. In these systems, at least, networks gain their power from deploying massive parallelism to check for the presence of a vast number of simple, shallow patterns. https://betterwithout.ai/images-surface-features This episode has a lot of links: David Chapman's earliest public mention, in February 2016, of image classifiers probably using color and texture in ways that "cheat": twitter.com/Meaningness/status/698688687341572096 Jordana Cepelewicz's “Where we see shapes, AI sees textures,” Quanta Magazine, July 1, 2019: https://www.quantamagazine.org/where-we-see-shapes-ai-sees-textures-20190701/ “Suddenly, a leopard print sofa appears”, May 2015: https://web.archive.org/web/20150622084852/http://rocknrollnerd.github.io/ml/2015/05/27/leopard-sofa.html “Understanding How Image Quality Affects Deep Neural Networks” April 2016: https://arxiv.org/abs/1604.04004 Goodfellow et al., “Explaining and Harnessing Adversarial Examples,” December 2014: https://arxiv.org/abs/1412.6572 “Universal adversarial perturbations,” October 2016: https://arxiv.org/pdf/1610.08401v1.pdf “Exploring the Landscape of Spatial Robustness,” December 2017: https://arxiv.org/abs/1712.02779 “Overinterpretation reveals image classification model pathologies,” NeurIPS 2021: https://proceedings.neurips.cc/paper/2021/file/8217bb4e7fa0541e0f5e04fea764ab91-Paper.pdf “Approximating CNNs with Bag-of-Local-Features Models Works Surprisingly Well on ImageNet,” ICLR 2019: https://openreview.net/forum?id=SkfMWhAqYQ Baker et al.'s “Deep convolutional networks do not classify based on global object shape,” PLOS Computational Biology, 2018: https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1006613 François Chollet's Twitter threads about AI producing images of horses with extra legs: twitter.com/fchollet/status/1573836241875120128 and twitter.com/fchollet/status/1573843774803161090 “Zoom In: An Introduction to Circuits,” 2020: https://distill.pub/2020/circuits/zoom-in/ Geirhos et al., “ImageNet-Trained CNNs Are Biased Towards Texture; Increasing Shape Bias Improves Accuracy and Robustness,” ICLR 2019: https://openreview.net/forum?id=Bygh9j09KX Dehghani et al., “Scaling Vision Transformers to 22 Billion Parameters,” 2023: https://arxiv.org/abs/2302.05442 Hasson et al., “Direct Fit to Nature: An Evolutionary Perspective on Biological and Artificial Neural Networks,” February 2020: https://www.gwern.net/docs/ai/scaling/2020-hasson.pdf
Happy holidays! We'll be sharing snippets from Latent Space LIVE! through the break bringing you the best of 2024! We want to express our deepest appreciation to event sponsors AWS, Daylight Computer, Thoth.ai, StrongCompute, Notable Capital, and most of all all our LS supporters who helped fund the gorgeous venue and A/V production!For NeurIPS last year we did our standard conference podcast coverage interviewing selected papers (that we have now also done for ICLR and ICML), however we felt that we could be doing more to help AI Engineers 1) get more industry-relevant content, and 2) recap 2024 year in review from experts. As a result, we organized the first Latent Space LIVE!, our first in person miniconference, at NeurIPS 2024 in Vancouver.The single most requested domain was computer vision, and we could think of no one better to help us recap 2024 than our friends at Roboflow, who was one of our earliest guests in 2023 and had one of this year's top episodes in 2024 again. Roboflow has since raised a $40m Series B!LinksTheir slides are here:All the trends and papers they picked:* Isaac Robinson* Sora (see our Video Diffusion pod) - extending diffusion from images to video* SAM 2: Segment Anything in Images and Videos (see our SAM2 pod) - extending prompted masks to full video object segmentation* DETR Dominancy: DETRs show Pareto improvement over YOLOs* RT-DETR: DETRs Beat YOLOs on Real-time Object Detection* LW-DETR: A Transformer Replacement to YOLO for Real-Time Detection* D-FINE: Redefine Regression Task in DETRs as Fine-grained Distribution Refinement* Peter Robicheaux* MMVP (Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs)* * Florence 2 (Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks) * PalíGemma / PaliGemma 2* PaliGemma: A versatile 3B VLM for transfer* PaliGemma 2: A Family of Versatile VLMs for Transfer* AlMv2 (Multimodal Autoregressive Pre-training of Large Vision Encoders) * Vik Korrapati - MoondreamFull Talk on YouTubeWant more content like this? Like and subscribe to stay updated on our latest talks, interviews, and podcasts.Transcript/Timestamps[00:00:00] Intro[00:00:05] AI Charlie: welcome to Latent Space Live, our first mini conference held at NeurIPS 2024 in Vancouver. This is Charlie, your AI co host. When we were thinking of ways to add value to our academic conference coverage, we realized that there was a lack of good talks, just recapping the best of 2024, going domain by domain.[00:00:36] AI Charlie: We sent out a survey to the over 900 of you. who told us what you wanted, and then invited the best speakers in the Latent Space Network to cover each field. 200 of you joined us in person throughout the day, with over 2, 200 watching live online. Our second featured keynote is The Best of Vision 2024, with Peter Robichaud and Isaac [00:01:00] Robinson of Roboflow, with a special appearance from Vic Corrapati of Moondream.[00:01:05] AI Charlie: When we did a poll of our attendees, the highest interest domain of the year was vision. And so our first port of call was our friends at Roboflow. Joseph Nelson helped us kickstart our vision coverage in episode 7 last year, and this year came back as a guest host with Nikki Ravey of Meta to cover segment Anything 2.[00:01:25] AI Charlie: Roboflow have consistently been the leaders in open source vision models and tooling. With their SuperVision library recently eclipsing PyTorch's Vision library. And Roboflow Universe hosting hundreds of thousands of open source vision datasets and models. They have since announced a 40 million Series B led by Google Ventures.[00:01:46] AI Charlie: Woohoo.[00:01:48] Isaac's picks[00:01:48] Isaac Robinson: Hi, we're Isaac and Peter from Roboflow, and we're going to talk about the best papers of 2024 in computer vision. So, for us, we defined best as what made [00:02:00] the biggest shifts in the space. And to determine that, we looked at what are some major trends that happened and what papers most contributed to those trends.[00:02:09] Isaac Robinson: So I'm going to talk about a couple trends, Peter's going to talk about a trend, And then we're going to hand it off to Moondream. So, the trends that I'm interested in talking about are These are a major transition from models that run on per image basis to models that run using the same basic ideas on video.[00:02:28] Isaac Robinson: And then also how debtors are starting to take over the real time object detection scene from the YOLOs, which have been dominant for years.[00:02:37] Sora, OpenSora and Video Vision vs Generation[00:02:37] Isaac Robinson: So as a highlight we're going to talk about Sora, which from my perspective is the biggest paper of 2024, even though it came out in February. Is the what?[00:02:48] Isaac Robinson: Yeah. Yeah. So just it's a, SORA is just a a post. So I'm going to fill it in with details from replication efforts, including open SORA and related work, such as a stable [00:03:00] diffusion video. And then we're also going to talk about SAM2, which applies the SAM strategy to video. And then how debtors, These are the improvements in 2024 to debtors that are making them a Pareto improvement to YOLO based models.[00:03:15] Isaac Robinson: So to start this off, we're going to talk about the state of the art of video generation at the end of 2023, MagVIT MagVIT is a discrete token, video tokenizer akin to VQ, GAN, but applied to video sequences. And it actually outperforms state of the art handcrafted video compression frameworks.[00:03:38] Isaac Robinson: In terms of the bit rate versus human preference for quality and videos generated by autoregressing on these discrete tokens generate some pretty nice stuff, but up to like five seconds length and, you know, not super detailed. And then suddenly a few months later we have this, which when I saw it, it was totally mind blowing to me.[00:03:59] Isaac Robinson: 1080p, [00:04:00] a whole minute long. We've got light reflecting in puddles. That's reflective. Reminds me of those RTX demonstrations for next generation video games, such as Cyberpunk, but with better graphics. You can see some issues in the background if you look closely, but they're kind of, as with a lot of these models, the issues tend to be things that people aren't going to pay attention to unless they're looking for.[00:04:24] Isaac Robinson: In the same way that like six fingers on a hand. You're not going to notice is a giveaway unless you're looking for it. So yeah, as we said, SORA does not have a paper. So we're going to be filling it in with context from the rest of the computer vision scene attempting to replicate these efforts. So the first step, you have an LLM caption, a huge amount of videos.[00:04:48] Isaac Robinson: This, this is a trick that they introduced in Dolly 3, where they train a image captioning model to just generate very high quality captions for a huge corpus and then train a diffusion model [00:05:00] on that. Their Sora and their application efforts also show a bunch of other steps that are necessary for good video generation.[00:05:09] Isaac Robinson: Including filtering by aesthetic score and filtering by making sure the videos have enough motion. So they're not just like kind of the generators not learning to just generate static frames. So. Then we encode our video into a series of space time latents. Once again, SORA, very sparse in details.[00:05:29] Isaac Robinson: So the replication related works, OpenSORA actually uses a MAG VIT V2 itself to do this, but swapping out the discretization step with a classic VAE autoencoder framework. They show that there's a lot of benefit from getting the temporal compression, which makes a lot of sense as the Each sequential frames and videos have mostly redundant information.[00:05:53] Isaac Robinson: So by compressing against, compressing in the temporal space, you allow the latent to hold [00:06:00] a lot more semantic information while avoiding that duplicate. So, we've got our spacetime latents. Possibly via, there's some 3D VAE, presumably a MAG VATV2 and then you throw it into a diffusion transformer.[00:06:19] Isaac Robinson: So I think it's personally interesting to note that OpenSORA is using a MAG VATV2, which originally used an autoregressive transformer decoder to model the latent space, but is now using a diffusion diffusion transformer. So it's still a transformer happening. Just the question is like, is it?[00:06:37] Isaac Robinson: Parameterizing the stochastic differential equation is, or parameterizing a conditional distribution via autoregression. It's also it's also worth noting that most diffusion models today, the, the very high performance ones are switching away from the classic, like DDPM denoising diffusion probability modeling framework to rectified flows.[00:06:57] Isaac Robinson: Rectified flows have a very interesting property that as [00:07:00] they converge, they actually get closer to being able to be sampled with a single step. Which means that in practice, you can actually generate high quality samples much faster. Major problem of DDPM and related models for the past four years is just that they require many, many steps to generate high quality samples.[00:07:22] Isaac Robinson: So, and naturally, the third step is throwing lots of compute at the problem. So I didn't, I never figured out how to manage to get this video to loop, but we see very little compute, medium compute, lots of compute. This is so interesting because the the original diffusion transformer paper from Facebook actually showed that, in fact, the specific hyperparameters of the transformer didn't really matter that much.[00:07:48] Isaac Robinson: What mattered was that you were just increasing the amount of compute that the model had. So, I love how in the, once again, little blog posts, they don't even talk about [00:08:00] like the specific hyperparameters. They say, we're using a diffusion transformer, and we're just throwing more compute at it, and this is what happens.[00:08:08] Isaac Robinson: OpenSora shows similar results. The primary issue I think here is that no one else has 32x compute budget. So we end up with these we end up in the middle of the domain and most of the related work, which is still super, super cool. It's just a little disappointing considering the context. So I think this is a beautiful extension of the framework that was introduced in 22 and 23 for these very high quality per image generation and then extending that to videos.[00:08:39] Isaac Robinson: It's awesome. And it's GA as of Monday, except no one can seem to get access to it because they keep shutting down the login.[00:08:46] SAM and SAM2[00:08:46] Isaac Robinson: The next, so next paper I wanted to talk about is SAM. So we at Roboflow allow users to label data and train models on that data. Sam, for us, has saved our users 75 years of [00:09:00] labeling time.[00:09:00] Isaac Robinson: We are the, to the best of my knowledge, the largest SAM API that exists. We also, SAM also allows us to have our users train just pure bounding box regression models and use those to generate high quality masks which has the great side effect of requiring less training data to have a meaningful convergence.[00:09:20] Isaac Robinson: So most people are data limited in the real world. So anything that requires less data to get to a useful thing is that super useful. Most of our users actually run their object per frame object detectors on every frame in a video, or maybe not most, but many, many. And so Sam follows into this category of taking, Sam 2 falls into this category of taking something that really really works and applying it to a video which has the wonderful benefit of being plug and play with most of our Many of our users use cases.[00:09:53] Isaac Robinson: We're, we're still building out a sufficiently mature pipeline to take advantage of that, but it's, it's in the works. [00:10:00] So here we've got a great example. We can click on cells and then follow them. You even notice the cell goes away and comes back and we can still keep track of it which is very challenging for existing object trackers.[00:10:14] Isaac Robinson: High level overview of how SAM2 works. We there's a simple pipeline here where we can give, provide some type of prompt and it fills out the rest of the likely masks for that object throughout the rest of the video. So here we're giving a bounding box in the first frame, a set of positive negative points, or even just a simple mask.[00:10:36] Isaac Robinson: I'm going to assume people are somewhat familiar with SAM. So I'm going to just give a high level overview of how SAM works. You have an image encoder that runs on every frame. SAM two can be used on a single image, in which case the only difference between SAM two and SAM is that image encoder, which Sam used a standard VIT [00:11:00] Sam two replaced that with a hara hierarchical encoder, which gets approximately the same results, but leads to a six times faster inference, which is.[00:11:11] Isaac Robinson: Excellent, especially considering how in a trend of 23 was replacing the VAT with more efficient backbones. In the case where you're doing video segmentation, the difference is that you actually create a memory bank and you cross attend the features from the image encoder based on the memory bank.[00:11:31] Isaac Robinson: So the feature set that is created is essentially well, I'll go more into it in a couple of slides, but we take the features from the past couple frames, plus a set of object pointers and the set of prompts and use that to generate our new masks. Then we then fuse the new masks for this frame with the.[00:11:57] Isaac Robinson: Image features and add that to the memory bank. [00:12:00] It's, well, I'll say more in a minute. The just like SAM, the SAM2 actually uses a data engine to create its data set in that people are, they assembled a huge amount of reference data, used people to label some of it and train the model used the model to label more of it and asked people to refine the predictions of the model.[00:12:20] Isaac Robinson: And then ultimately the data set is just created from the engine Final output of the model on the reference data. It's very interesting. This paradigm is so interesting to me because it unifies a model in a dataset in a way that is very unique. It seems unlikely that another model could come in and have such a tight.[00:12:37] Isaac Robinson: So brief overview of how the memory bank works, the paper did not have a great visual, so I'm just, I'm going to fill in a bit more. So we take the last couple of frames from our video. And we take the last couple of frames from our video attend that, along with the set of prompts that we provided, they could come from the future, [00:13:00] they could come from anywhere in the video, as well as reference object pointers, saying, by the way, here's what we've found so far attending to the last few frames has the interesting benefit of allowing it to model complex object motion without actually[00:13:18] Isaac Robinson: By limiting the amount of frames that you attend to, you manage to keep the model running in real time. This is such an interesting topic for me because one would assume that attending to all of the frames is super essential, or having some type of summarization of all the frames is super essential for high performance.[00:13:35] Isaac Robinson: But we see in their later ablation that that actually is not the case. So here, just to make sure that there is some benchmarking happening, we just compared to some of the stuff that's came out prior, and indeed the SAM2 strategy does improve on the state of the art. This ablation deep in their dependencies was super interesting to me.[00:13:59] Isaac Robinson: [00:14:00] We see in section C, the number of memories. One would assume that increasing the count of memories would meaningfully increase performance. And we see that it has some impact, but not the type that you'd expect. And that it meaningfully decreases speed, which justifies, in my mind, just having this FIFO queue of memories.[00:14:20] Isaac Robinson: Although in the future, I'm super interested to see A more dedicated summarization of all of the last video, not just a stacking of the last frames. So that another extension of beautiful per frame work into the video domain.[00:14:42] Realtime detection: DETRs > YOLO[00:14:42] Isaac Robinson: The next trend I'm interested in talking about is this interesting at RoboFlow, we're super interested in training real time object detectors.[00:14:50] Isaac Robinson: Those are bread and butter. And so we're doing a lot to keep track of what is actually happening in that space. We are finally starting to see something change. So, [00:15:00] for years, YOLOs have been the dominant way of doing real time object detection, and we can see here that they've essentially stagnated.[00:15:08] Isaac Robinson: The performance between 10 and 11 is not meaningfully different, at least, you know, in this type of high level chart. And even from the last couple series, there's not. A major change so YOLOs have hit a plateau, debtors have not. So we can look here and see the YOLO series has this plateau. And then these RT debtor, LW debtor, and Define have meaningfully changed that plateau so that in fact, the best Define models are plus 4.[00:15:43] Isaac Robinson: 6 AP on Cocoa at the same latency. So three major steps to accomplish this. The first RT deditor, which is technically a 2023 paper preprint, but published officially in 24, so I'm going to include that. I hope that's okay. [00:16:00] That is showed that RT deditor showed that we could actually match or out speed YOLOs.[00:16:04] Isaac Robinson: And then LWdebtor showed that pre training is hugely effective on debtors and much less so on YOLOs. And then DeFine added the types of bells and whistles that we expect from these types, this, this arena. So the major improvements that RTdebtor shows was Taking the multi scale features that debtors typically pass into their encoder and decoupling them into a much more efficient transformer encoder.[00:16:30] Isaac Robinson: The transformer is of course, quadratic complexity. So decreasing the amount of stuff that you pass in at once is super helpful for increasing your runtime or increasing your throughput. So that change basically brought us up to yellow speed and then they do a hardcore analysis on. Benchmarking YOLOs, including the NMS step.[00:16:54] Isaac Robinson: Once you once you include the NMS in the latency calculation, you see that in fact, these debtors [00:17:00] are outperforming, at least this time, the the, the YOLOs that existed. Then LW debtor goes in and suggests that in fact, the frame, the huge boost here is from pre training. So, this is the define line, and this is the define line without pre training.[00:17:19] Isaac Robinson: It's within range, it's still an improvement over the YOLOs, but Really huge boost comes from the benefit of pre training. When YOLOx came out in 2021, they showed that they got much better results by having a much, much longer training time, but they found that when they did that, they actually did not benefit from pre training.[00:17:40] Isaac Robinson: So, you see in this graph from LWdebtor, in fact, YOLOs do have a real benefit from pre training, but it goes away as we increase the training time. Then, the debtors converge much faster. LWdebtor trains for only 50 epochs, RTdebtor is 60 epochs. So, one could assume that, in fact, [00:18:00] the entire extra gain from pre training is that you're not destroying your original weights.[00:18:06] Isaac Robinson: By relying on this long training cycle. And then LWdebtor also shows superior performance to our favorite data set, Roboflow 100 which means that they do better on the real world, not just on Cocoa. Then Define throws all the bells and whistles at it. Yellow models tend to have a lot of very specific complicated loss functions.[00:18:26] Isaac Robinson: This Define brings that into the debtor world and shows consistent improvement on a variety of debtor based frameworks. So bring these all together and we see that suddenly we have almost 60 AP on Cocoa while running in like 10 milliseconds. Huge, huge stuff. So we're spending a lot of time trying to build models that work better with less data and debtors are clearly becoming a promising step in that direction.[00:18:56] Isaac Robinson: The, what we're interested in seeing [00:19:00] from the debtors in this, this trend to next is. Codetter and the models that are currently sitting on the top of the leaderboard for large scale inference scale really well as you switch out the backbone. We're very interested in seeing and having people publish a paper, potentially us, on what happens if you take these real time ones and then throw a Swingy at it.[00:19:23] Isaac Robinson: Like, do we have a Pareto curve that extends from the real time domain all the way up to the super, super slow but high performance domain? We also want to see people benchmarking in RF100 more, because that type of data is what's relevant for most users. And we want to see more pre training, because pre training works now.[00:19:43] Isaac Robinson: It's super cool.[00:19:48] Peter's Picks[00:19:48] Peter Robicheaux: Alright, so, yeah, so in that theme one of the big things that we're focusing on is how do we get more out of our pre trained models. And one of the lenses to look at this is through sort of [00:20:00] this, this new requirement for like, how Fine grained visual details and your representations that are extracted from your foundation model.[00:20:08] Peter Robicheaux: So it's sort of a hook for this Oh, yeah, this is just a list of all the the papers that I'm going to mention I just want to make sure I set an actual paper so you can find it later[00:20:18] MMVP (Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs)[00:20:18] Peter Robicheaux: Yeah, so sort of the big hook here is that I make the claim that LLMs can't see if you go to if you go to Claude or ChatGPT you ask it to see this Watch and tell me what time it is, it fails, right?[00:20:34] Peter Robicheaux: And so you could say, like, maybe, maybe the Like, this is, like, a very classic test of an LLM, but you could say, Okay, maybe this, this image is, like, too zoomed out, And it just, like, it'll do better if we increase the resolution, And it has easier time finding these fine grained features, Like, where the watch hands are pointing.[00:20:53] Peter Robicheaux: Nodice. And you can say, okay, well, maybe the model just doesn't know how to tell time from knowing the position of the hands. But if you actually prompt [00:21:00] it textually, it's very easy for it to tell the time. So this to me is proof that these LLMs literally cannot see the position of the watch hands and it can't see those details.[00:21:08] Peter Robicheaux: So the question is sort of why? And for you anthropic heads out there, cloud fails too. So the, the, my first pick for best paper of 2024 Envision is this MMVP paper, which tries to investigate the Why do LLMs not have the ability to see fine grained details? And so, for instance, it comes up with a lot of images like this, where you ask it a question that seems very visually apparent to us, like, which way is the school bus facing?[00:21:32] Peter Robicheaux: And it gets it wrong, and then, of course, it makes up details to support its wrong claim. And so, the process by which it finds these images is sort of contained in its hypothesis for why it can't. See these details. So it hypothesizes that models that have been initialized with, with Clip as their vision encoder, they don't have fine grained details and the, the features extracted using Clip because Clip sort of doesn't need to find these fine grained [00:22:00] details to do its job correctly, which is just to match captions and images, right?[00:22:04] Peter Robicheaux: And sort of at a high level, even if ChatGPT wasn't initialized with Clip and wasn't trained contrastively at all. The vision encoder wasn't trained contrastively at all. Still, in order to do its job of capturing the image it could do a pretty good job without actually finding the exact position of all the objects and visual features in the image, right?[00:22:21] Peter Robicheaux: So This paper finds a set of difficult images for these types of models. And the way it does it is it looks for embeddings that are similar in clip space, but far in DynaV2 space. So DynaV2 is a foundation model that was trained self supervised purely on image data. And it kind of uses like some complex student teacher framework, but essentially, and like, it patches out like certain areas of the image or like crops with certain areas of the image and tries to make sure that those have consistent representations, which is a way for it to learn very fine grained visual features.[00:22:54] Peter Robicheaux: And so if you take things that are very close in clip space and very far in DynaV2 space, you get a set of images [00:23:00] that Basically, pairs of images that are hard for a chat GPT and other big language models to distinguish. So, if you then ask it questions about this image, well, as you can see from this chart, it's going to answer the same way for both images, right?[00:23:14] Peter Robicheaux: Because to, to, from the perspective of the vision encoder, they're the same image. And so if you ask a question like, how many eyes does this animal have? It answers the same for both. And like all these other models, including Lava do the same thing, right? And so this is the benchmark that they create, which is like finding clip, like clip line pairs, which is pairs of images that are similar in clip space and creating a data set of multiple choice questions based off of those.[00:23:39] Peter Robicheaux: And so how do these models do? Well, really bad. Lava, I think, So, so, chat2BT and Jim and I do a little bit better than random guessing, but, like, half of the performance of humans who find these problems to be very easy. Lava is, interestingly, extremely negatively correlated with this dataset. It does much, much, much, much worse [00:24:00] than random guessing, which means that this process has done a very good job of identifying hard images for, for Lava, specifically.[00:24:07] Peter Robicheaux: And that's because Lava is basically not trained for very long and is initialized from Clip, and so You would expect it to do poorly on this dataset. So, one of the proposed solutions that this paper attempts is by basically saying, Okay, well if clip features aren't enough, What if we train the visual encoder of the language model also on dyno features?[00:24:27] Peter Robicheaux: And so it, it proposes two different ways of doing this. One, additively which is basically interpolating between the two features, and then one is interleaving, which is just kind of like training one on the combination of both features. So there's this really interesting trend when you do the additive mixture of features.[00:24:45] Peter Robicheaux: So zero is all clip features and one is all DynaV2 features. So. It, as you, so I think it's helpful to look at the right most chart first, which is as you increase the number of DynaV2 features, your model does worse and worse and [00:25:00] worse on the actual language modeling task. And that's because DynaV2 features were trained completely from a self supervised manner and completely in image space.[00:25:08] Peter Robicheaux: It knows nothing about text. These features aren't really compatible with these text models. And so you can train an adapter all you want, but it seems that it's in such an alien language that it's like a very hard optimization for this. These models to solve. And so that kind of supports what's happening on the left, which is that, yeah, it gets better at answering these questions if as you include more dyna V two features up to a point, but then you, when you oversaturate, it completely loses its ability to like.[00:25:36] Peter Robicheaux: Answer language and do language tasks. So you can also see with the interleaving, like they essentially double the number of tokens that are going into these models and just train on both, and it still doesn't really solve the MMVP task. It gets Lava 1. 5 above random guessing by a little bit, but it's still not close to ChachiPT or, you know, Any like human performance, obviously.[00:25:59] Peter Robicheaux: [00:26:00] So clearly this proposed solution of just using DynaV2 features directly, isn't going to work. And basically what that means is that as a as a vision foundation model, DynaV2 is going to be insufficient for language tasks, right?[00:26:14] Florence 2 (Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks)[00:26:14] Peter Robicheaux: So my next pick for best paper of 2024 would be Florence 2, which tries to solve this problem by incorporating not only This dimension of spatial hierarchy, which is to say pixel level understanding, but also in making sure to include what they call semantic granularity, which ends up, the goal is basically to have features that are sufficient for finding objects in the image, so they're, they're, they have enough pixel information, but also can be talked about and can be reasoned about.[00:26:44] Peter Robicheaux: And that's on the semantic granularity axis. So here's an example of basically three different paradigms of labeling that they do. So they, they create a big dataset. One is text, which is just captioning. And you would expect a model that's trained [00:27:00] only on captioning to have similar performance like chat2BT and like not have spatial hierarchy, not have features that are meaningful at the pixel level.[00:27:08] Peter Robicheaux: And so they add another type, which is region text pairs, which is essentially either classifying a region or You're doing object detection or doing instance segmentation on that region or captioning that region. And then they have text phrased region annotations, which is essentially a triple. And basically, not only do you have a region that you've described, you also find it's like, It's placed in a descriptive paragraph about the image, which is basically trying to introduce even more like semantic understanding of these regions.[00:27:39] Peter Robicheaux: And so like, for instance, if you're saying a woman riding on the road, right, you have to know what a woman is and what the road is and that she's on top of it. And that's, that's basically composing a bunch of objects in this visual space, but also thinking about it semantically, right? And so the way that they do this is they take basically they just dump Features from a vision encoder [00:28:00] straight into a encoder decoder transformer.[00:28:03] Peter Robicheaux: And then they train a bunch of different tasks like object detection and so on as a language task. And I think that's one of the big things that we saw in 2024 is these, these vision language models operating in, on pixel space linguistically. So they introduced a bunch of new tokens to point to locations and[00:28:22] Peter Robicheaux: So how does it work? How does it actually do? We can see if you look at the graph on the right, which is using the, the Dino, the the Dino framework your, your pre trained Florence 2 models transfer very, very well. They get 60%, 60 percent map on Cocoa, which is like approaching state of the art and they train[00:28:42] Vik Korrapati: with, and they[00:28:43] Peter Robicheaux: train with a much more more efficiently.[00:28:47] Peter Robicheaux: So they, they converge a lot faster, which both of these things are pointing to the fact that they're actually leveraging their pre trained weights effectively. So where is it falling short? So these models, I forgot to mention, Florence is a 0. 2 [00:29:00] billion and a 0. 7 billion parameter count. So they're very, very small in terms of being a language model.[00:29:05] Peter Robicheaux: And I think that. This framework, you can see saturation. So, what this graph is showing is that if you train a Florence 2 model purely on the image level and region level annotations and not including the pixel level annotations, like this, segmentation, it actually performs better as an object detector.[00:29:25] Peter Robicheaux: And what that means is that it's not able to actually learn all the visual tasks that it's trying to learn because it doesn't have enough capacity.[00:29:32] PalíGemma / PaliGemma 2[00:29:32] Peter Robicheaux: So I'd like to see this paper explore larger model sizes, which brings us to our next big paper of 2024 or two papers. So PolyGemma came out earlier this year.[00:29:42] Peter Robicheaux: PolyGemma 2 was released, I think like a week or two ago. Oh, I forgot to mention, you can actually train You can, like, label text datasets on RoboFlow and you can train a Florence 2 model and you can actually train a PolyGemma 2 model on RoboFlow, which we got into the platform within, like, 14 hours of release, which I was really excited about.[00:29:59] Peter Robicheaux: So, anyway, so [00:30:00] PolyGemma 2, so PolyGemma is essentially doing the same thing, but instead of doing an encoder decoder, it just dumps everything into a decoder only transformer model. But it also introduced the concept of location tokens to point to objects in pixel space. PolyGemma 2, so PolyGemma uses Gemma as the language encoder, and it uses Gemma2B.[00:30:17] Peter Robicheaux: PolyGemma 2 introduces using multiple different sizes of language encoders. So, the way that they sort of get around having to do encoder decoder is they use the concept of prefix loss. Which basically means that when it's generating, tokens autoregressively, it's all those tokens in the prefix, which is like the image that it's looking at and like a description of the task that it's trying to do.[00:30:41] Peter Robicheaux: They're attending to each other fully, full attention. Which means that, you know, it can sort of. Find high level it's easier for the, the prefix to color, to color the output of the suffix and also to just find like features easily. So this is sort of [00:31:00] an example of like one of the tasks that was trained on, which is like, you describe the task in English and then you give it all these, like, You're asking for it to segment these two classes of objects, and then it finds, like, their locations using these tokens, and it finds their masks using some encoding of the masks into tokens.[00:31:24] Peter Robicheaux: And, yeah, so, one of my critiques, I guess, of PolyGemma 1, at least, is that You find that performance saturates as a pre trained model after only 300 million examples seen. So, what this graph is representing is each blue dot is a performance on some downstream task. And you can see that after seeing 300 million examples, It sort of does equally well on all of the downtrend tasks that they tried it on, which was a lot as 1 billion examples, which to me also kind of suggests a lack of capacity for this model.[00:31:58] Peter Robicheaux: PolyGemma2, [00:32:00] you can see the results on object detection. So these were transferred to to Coco. And you can see that this sort of also points to an increase in capacity being helpful to the model. You can see as. Both the resolution increases, and the parameter count of the language model increases, performance increases.[00:32:16] Peter Robicheaux: So resolution makes sense, obviously, it helps to find small images, or small objects in the image. But it also makes sense for another reason, which is that it kind of gives the model a thinking register, and it gives it more tokens to, like, process when making its predictions. But yeah, you could, you could say, oh, 43.[00:32:30] Peter Robicheaux: 6, that's not that great, like Florence 2 got 60. But this is not Training a dino or a debtor on top of this language or this image encoder. It's doing the raw language modeling task on Cocoa. So it doesn't have any of the bells and whistles. It doesn't have any of the fancy losses. It doesn't even have bipartite graph matching or anything like that.[00:32:52] Peter Robicheaux: Okay, the big result and one of the reasons that I was really excited about this paper is that they blow everything else away [00:33:00] on MMVP. I mean, 47. 3, sure, that's nowhere near human accuracy, which, again, is 94%, but for a, you know, a 2 billion language, 2 billion parameter language model to be chat2BT, that's quite the achievement.[00:33:12] Peter Robicheaux: And that sort of brings us to our final pick for paper of the year, which is AIMV2. So, AIMV2 sort of says, okay, Maybe this language model, like, maybe coming up with all these specific annotations to find features and with high fidelity and pixel space isn't actually necessary. And we can come up with an even simpler, more beautiful idea for combining you know, image tokens and pixel tokens in a way that's interfaceable for language tasks.[00:33:44] Peter Robicheaux: And this is nice because it can scale, you can come up with lots more data if you don't have to come up with all these annotations, right? So the way that it works. is it does something very, very similar to PolyGemo, where you have a vision encoder that dumps image tokens into a decoder only transformer.[00:33:59] Peter Robicheaux: But [00:34:00] the interesting thing is that it also autoregressively tries to learn the mean squared error of the image tokens. So instead of having to come up with fancy object detection or semantic, or segment, or segmentation labels, you can just try to reconstruct the image and have it learn fine grained features that way.[00:34:16] Peter Robicheaux: And it does this in kind of, I think, a beautiful way that's kind of compatible with the PolyGemma line of thinking, which is randomly sampling a prefix line of thinking Prefix length and using only this number of image tokens as the prefix. And so doing a similar thing with the causal. So the causal with prefix is the, the attention mask on the right.[00:34:35] Peter Robicheaux: So it's doing full block attention with some randomly sampled number of image tokens to then reconstruct the rest of the image and the downstream caption for that image. And so, This is the dataset that they train on. It's image or internet scale data, very high quality data created by the data filtering networks paper, essentially which is maybe The best clip data that exists.[00:34:59] Peter Robicheaux: [00:35:00] And we can see that this is finally a model that doesn't saturate. It's even at the highest parameter count, it's, it appears to be, oh, at the highest parameter account, it appears to be improving in performance with more and more samples seen. And so you can sort of think that. You know, if we just keep bumping the parameter count and increasing the example scene, which is the, the, the line of thinking for language models, then it'll keep getting better.[00:35:27] Peter Robicheaux: So how does it actually do at finding, oh, it also improves with resolution, which you would expect for a model that This is the ImageNet classification accuracy, but yeah, it does better if you increase the resolution, which means that it's actually leveraging and finding fine grained visual features.[00:35:44] Peter Robicheaux: And so how does that actually do compared to CLIP on Cocoa? Well, you can see that if you slap a transformer detection head on it, Entry now in Cocoa, it's just 60. 2, which is also within spitting distance of Soda, which means that it does a very good job of [00:36:00] finding visual features, but you could say, okay, well, wait a second.[00:36:03] Peter Robicheaux: Clip got to 59. 1, so. Like, how does this prove your claim at all? Because doesn't that mean like clip, which is known to be clip blind and do badly on MMVP, it's able to achieve a very high performance on fine, on this fine grained visual features task of object detection, well, they train on like, Tons of data.[00:36:24] Peter Robicheaux: They train on like objects, 365, Cocoa, Flickr and everything else. And so I think that this benchmark doesn't do a great job of selling how good of a pre trained model MV2 is. And we would like to see the performance on fewer data as examples and not trained to convergence on object detection. So seeing it in the real world on like a dataset, like RoboFlow 100, I think would be quite interesting.[00:36:48] Peter Robicheaux: And our, our, I guess our final, final pick for paper of 2024 would be Moondream. So introducing Vic to talk about that.[00:36:54] swyx: But overall, that was exactly what I was looking for. Like best of 2024, an amazing job. Yeah, you can, [00:37:00] if there's any other questions while Vic gets set up, like vision stuff,[00:37:07] swyx: yeah,[00:37:11] swyx: Vic, go ahead. Hi,[00:37:13] Vik Korrapati / Moondream[00:37:13] question: well, while we're getting set up, hi, over here, thanks for the really awesome talk. One of the things that's been weird and surprising is that the foundation model companies Even these MLMs, they're just like worse than RT Tether at detection still. Like, if you wanted to pay a bunch of money to auto label your detection dataset, If you gave it to OpenAI or Cloud, that would be like a big waste.[00:37:37] question: So I'm curious, just like, even Pali Gemma 2, like is worse. So, so I'm curious to hear your thoughts on like, how come, Nobody's cracked the code on like a generalist that really you know, beats a specialist model in computer vision like they have in in LLM land.[00:38:00][00:38:01] Isaac Robinson: Okay. It's a very, very interesting question. I think it depends on the specific domain. For image classification, it's basically there. In the, in AIMv2 showed, a simple attentional probe on the pre trained features gets like 90%, which is as well as anyone does. The, the, the, the bigger question, like, why isn't it transferring to object detection, especially like real time object detection.[00:38:25] Isaac Robinson: I think, in my mind, there are two answers. One is, object detection is really, really, really the architectures are super domain specific. You know, we see these, all these super, super complicated things, and it's not super easy to, to, to build something that just transfers naturally like that, whereas image classification, you know, clip pre training transfers super, super quickly.[00:38:48] Isaac Robinson: And the other thing is, until recently, the real time object detectors didn't even really benefit from pre training. Like, you see the YOLOs that are like, essentially saturated, showing very little [00:39:00] difference with pre training improvements, with using pre trained model at all. It's not surprising, necessarily, that People aren't looking at the effects of better and better pre training on real time detection.[00:39:12] Isaac Robinson: Maybe that'll change in the next year. Does that answer your question?[00:39:17] Peter Robicheaux: Can you guys hear me? Yeah, one thing I want to add is just like, or just to summarize, basically, is that like, Until 2024, you know, we haven't really seen a combination of transformer based object detectors and fancy losses, and PolyGemma suffers from the same problem, which is basically to say that these ResNet, or like the convolutional models, they have all these, like, extreme optimizations for doing object detection, but essentially, I think it's kind of been shown now that convolution models like just don't benefit from pre training and just don't like have the level of intelligence of transformer models.[00:39:56] swyx: Awesome. Hi,[00:39:59] Vik Korrapati: can [00:40:00] you hear me?[00:40:01] swyx: Cool. I hear you. See you. Are you sharing your screen?[00:40:04] Vik Korrapati: Hi. Might have forgotten to do that. Let me do[00:40:07] swyx: that. Sorry, should have done[00:40:08] Vik Korrapati: that.[00:40:17] swyx: Here's your screen. Oh, classic. You might have to quit zoom and restart. What? It's fine. We have a capture of your screen.[00:40:34] swyx: So let's get to it.[00:40:35] Vik Korrapati: Okay, easy enough.[00:40:49] Vik Korrapati: All right. Hi, everyone. My name is Vic. I've been working on Moondream for almost a year now. Like Shawn mentioned, I just went and looked and it turns out the first version I released December [00:41:00] 29, 2023. It's been a fascinating journey. So Moonbeam started off as a tiny vision language model. Since then, we've expanded scope a little bit to also try and build some tooling, client libraries, et cetera, to help people really deploy it.[00:41:13] Vik Korrapati: Unlike traditional large models that are focused at assistant type use cases, we're laser focused on building capabilities that developers can, sorry, it's yeah, we're basically focused on building capabilities that developers can use to build vision applications that can run anywhere. So, in a lot of cases for vision more so than for text, you really care about being able to run on the edge, run in real time, etc.[00:41:40] Vik Korrapati: So That's really important. We have we have different output modalities that we support. There's query where you can ask general English questions about an image and get back human like answers. There's captioning, which a lot of our users use for generating synthetic datasets to then train diffusion models and whatnot.[00:41:57] Vik Korrapati: We've done a lot of work to minimize those sessions there. [00:42:00] So that's. Use lot. We have open vocabulary object detection built in similar to a couple of more recent models like Palagem, et cetera, where rather than having to train a dedicated model, you can just say show me soccer balls in this image or show me if there are any deer in this image, it'll detect it.[00:42:14] Vik Korrapati: More recently, earlier this month, we released pointing capability where if all you're interested in is the center of an object you can just ask it to point out where that is. This is very useful when you're doing, you know, I automation type stuff. Let's see, LA we, we have two models out right now.[00:42:33] Vik Korrapati: There's a general purpose to be para model, which runs fair. Like it's, it's it's fine if you're running on server. It's good for our local Amma desktop friends and it can run on flagship, flagship mobile phones, but it never. so much for joining us today, and we'll see you in the [00:43:00] next one. Less memory even with our not yet fully optimized inference client.[00:43:06] Vik Korrapati: So the way we built our 0. 5b model was to start with the 2 billion parameter model and prune it while doing continual training to retain performance. We, our objective during the pruning was to preserve accuracy across a broad set of benchmarks. So the way we went about it was to estimate the importance of different components of the model, like attention heads, channels MLP rows and whatnot using basically a technique based on the gradient.[00:43:37] Vik Korrapati: I'm not sure how much people want to know details. We'll be writing a paper about this, but feel free to grab me if you have more questions. Then we iteratively prune a small chunk that will minimize loss and performance retrain the model to recover performance and bring it back. The 0. 5b we released is more of a proof of concept that this is possible.[00:43:54] Vik Korrapati: I think the thing that's really exciting about this is it makes it possible for for developers to build using the 2B param [00:44:00] model and just explore, build their application, and then once they're ready to deploy figure out what exactly they need out of the model and prune those capabilities into a smaller form factor that makes sense for their deployment target.[00:44:12] Vik Korrapati: So yeah, very excited about that. Let me talk to you folks a little bit about another problem I've been working on recently, which is similar to the clocks example we've been talking about. We had a customer reach out who was talking about, like, who had a bunch of gauges out in the field. This is very common in manufacturing and oil and gas, where you have a bunch of analog devices that you need to monitor.[00:44:34] Vik Korrapati: It's expensive to. And I was like, okay, let's have humans look at that and monitor stuff and make sure that the system gets shut down when the temperature goes over 80 or something. So I was like, yeah, this seems easy enough. Happy to, happy to help you distill that. Let's, let's get it going. Turns out our model couldn't do it at all.[00:44:51] Vik Korrapati: I went and looked at other open source models to see if I could just generate a bunch of data and learn from that. Did not work either. So I was like, let's look at what the folks with [00:45:00] hundreds of billions of dollars in market cap have to offer. And yeah, that doesn't work either. My hypothesis is that like the, the way these models are trained are using a large amount of image text data scraped from the internet.[00:45:15] Vik Korrapati: And that can be biased. In the case of gauges, most gauge images aren't gauges in the wild, they're product images. Detail images like these, where it's always set to zero. It's paired with an alt text that says something like GIVTO, pressure sensor, PSI, zero to 30 or something. And so the models are fairly good at picking up those details.[00:45:35] Vik Korrapati: It'll tell you that it's a pressure gauge. It'll tell you what the brand is, but it doesn't really learn to pay attention to the needle over there. And so, yeah, that's a gap we need to address. So naturally my mind goes to like, let's use synthetic data to, Solve this problem. That works, but it's problematic because it turned out we needed millions of synthetic gauge images to get to reasonable performance.[00:45:57] Vik Korrapati: And thinking about it, reading a gauge is like [00:46:00] not a one, like it's not a zero short process in our minds, right? Like if you had to tell me the reading in Celsius for this, Real world gauge. There's two dials on there. So first you have to figure out which one you have to be paying attention to, like the inner one or the outer one.[00:46:14] Vik Korrapati: You look at the tip of the needle, you look at what labels it's between, and you count how many and do some math to figure out what that probably is. So what happens if we just add that as a Chain of thought to give the model better understanding of the different sub, to allow the model to better learn the subtasks it needs to perform to accomplish this goal.[00:46:37] Vik Korrapati: So you can see in this example, this was actually generated by the latest version of our model. It's like, okay, Celsius is the inner scale. It's between 50 and 60. There's 10 ticks. So the second tick, it's a little debatable here, like there's a weird shadow situation going on, the dial is off, so I don't know what the ground truth is, but it works okay.[00:46:57] Vik Korrapati: There's points on there that are, the points [00:47:00] over there are actually grounded. I don't know if this is easy to see, but when I click on those, there's a little red dot that moves around on the image. The model actually has to predict where this points are, I was already trying to do this with bounding boxes, but then Malmo came out with pointing capabilities.[00:47:15] Vik Korrapati: And it's like pointing is a much better paradigm to to represent this. We see pretty good results. This one's actually for clock reading. I couldn't find our chart for gauge reading at the last minute. So the light. Blue chart is with our rounded chain of thought. This measures, we have, we built a clock reading benchmark about 500 images.[00:47:37] Vik Korrapati: This measures accuracy on that. You can see it's a lot more sample efficient when you're using the chain of thought to model. Another big benefit from this approach is like, you can kind of understand how the model is. it and how it's failing. So in this example, the actual correct reading is 54 Celsius, the model output [00:48:00] 56, not too bad but you can actually go and see where it messed up. Like it got a lot of these right, except instead of saying it was on the 7th tick, it actually predicted that it was the 8th tick and that's why it went with 56.[00:48:14] Vik Korrapati: So now that you know that this. Failing in this way, you can adjust how you're doing the chain of thought to maybe say like, actually count out each tick from 40, instead of just trying to say it's the eighth tick. Or you might say like, okay, I see that there's that middle thing, I'll count from there instead of all the way from 40.[00:48:31] Vik Korrapati: So helps a ton. The other thing I'm excited about is a few short prompting or test time training with this. Like if a customer has a specific gauge that like we're seeing minor errors on, they can give us a couple of examples where like, if it's miss detecting the. Needle, they can go in and correct that in the chain of thought.[00:48:49] Vik Korrapati: And hopefully that works the next time. Now, exciting approach, we only apply it to clocks and gauges. The real question is, is it going to generalize? Probably, like, there's some science [00:49:00] from text models that when you train on a broad number of tasks, it does generalize. And I'm seeing some science with our model as well.[00:49:05] Vik Korrapati: So, in addition to the image based chain of thought stuff, I also added some spelling based chain of thought to help it understand better understand OCR, I guess. I don't understand why everyone doesn't do this, by the way. Like, it's trivial benchmark question. It's Very, very easy to nail. But I also wanted to support it for stuff like license plate, partial matching, like, hey, does any license plate in this image start with WHA or whatever?[00:49:29] Vik Korrapati: So yeah, that sort of worked. All right, that, that ends my story about the gauges. If you think about what's going on over here it's interesting that like LLMs are showing enormous. Progress in reasoning, especially with the latest set of models that we've seen, but we're not really seeing, I have a feeling that VLMs are lagging behind, as we can see with these tasks that should be very simple for a human to do [00:50:00] that are very easy to find VLMs failing at.[00:50:04] Vik Korrapati: My hypothesis on why this is the case is because On the internet, there's a ton of data that talks about how to reason. There's books about how to solve problems. There's books critiquing the books about how to solve problems. But humans are just so good at perception that we never really talk about it.[00:50:20] Vik Korrapati: Like, maybe in art books where it's like, hey, to show that that mountain is further away, you need to desaturate it a bit or whatever. But the actual data on how to, like, look at images is, isn't really present. Also, the Data we have is kind of sketched. The best source of data we have is like image all text pairs on the internet and that's pretty low quality.[00:50:40] Vik Korrapati: So yeah, I, I think our solution here is really just we need to teach them how to operate on individual tasks and figure out how to scale that out. All right. Yep. So conclusion. At Moondream we're trying to build amazing PLMs that run everywhere. Very hard problem. Much work ahead, but we're making a ton of progress and I'm really excited [00:51:00] about If anyone wants to chat about more technical details about how we're doing this or interest in collaborating, please, please hit me up.[00:51:08] Isaac Robinson: Yeah,[00:51:09] swyx: like, I always, when people say, when people say multi modality, like, you know, I always think about vision as the first among equals in all the modalities. So, I really appreciate having the experts in the room. Get full access to Latent Space at www.latent.space/subscribe
Dr. Olga Russakovsky, Computer Science at Princeton University, joins Lightspeed Partner Michael Mignano to discuss what the next generation of AI talent is learning and where she expects to find the next big innovation in artificial intelligence. From her research in computer vision and human-computer interaction to her work in fairness, accountability, and transparency in AI, Dr. Russakovsky has earned many awards, including the MIT Technology Review's 35-under-35 Innovator award and the Foreign Policy Magazine's 100 Leading Global Thinkers award. Dr. Russakovsky is also the co-founder and Board Chair of AI4ALL, a nonprofit that aims to increase the diversity of thought in Artificial Intelligence. Episode Chapters 00:00 Introduction and Guest Overview 01:17 Olga's Career Journey 02:30 Understanding Computer Vision 04:43 Generative AI and Computer Vision 06:36 Interdisciplinary AI Research 15:00 AI4All: Diversity of Thought 17:44 Challenges and Bias in AI 30:01 Future of AI and Data 40:08 ImageNET 43:38 Closing Thoughts Stay in touch: www.lsvp.com X: https://twitter.com/lightspeedvp LinkedIn: https://www.linkedin.com/company/lightspeed-venture-partners/ Instagram: https://www.instagram.com/lightspeedventurepartners/ Subscribe on your favorite podcast app: generativenow.co Email: generativenow@lsvp.com The content here does not constitute tax, legal, business or investment advice or an offer to provide such advice, should not be construed as advocating the purchase or sale of any security or investment or a recommendation of any company, and is not an offer, or solicitation of an offer, for the purchase or sale of any security or investment product. For more details please see lsvp.com/legal.
- GS. Fei-Fei Li (Đại học Stanford, Mỹ) là một trong 5 nhà khoa học được vinh danh Giải thưởng Chính VinFuture 2024 trị giá 3 triệu USD vì những đóng góp đột phá để thúc đẩy sự tiến bộ của học sâu. Bà là người đã tạo ra tập dữ liệu ImageNet giúp thúc đẩy sự tiến bộ trong hệ thống nhận diện hình ảnh, giúp huấn luyện các mô hình học sâu ở quy mô lớn. Trong Chuyện đêm hôm nay, chúng tôi mời quý vị và các bạn cùng trò chuyện với nhà khoa học người Mỹ được mệnh danh là “mẹ đỡ đầu” của AI, nổi tiếng với đóng góp đột phá trong lĩnh vực thị giác máy tính. Chủ đề : GS Fei Fei Li, ĐH Stanford, Mỹ, nghiên cứu AI, Việt Nam phát triển AI --- Support this podcast: https://podcasters.spotify.com/pod/show/vov1sukien/support
This episode of Eye on AI is sponsored by Citrusx. Unlock reliable AI with Citrusx! Our platform simplifies validation and risk management, empowering you to make smarter decisions and stay compliant. Detects and mitigate AI vulnerabilities, biases, and errors with ease. Visit http://email.citrusx.ai/eyeonai to download our free fairness use case and see the solution in action. In this episode of the Eye on AI podcast, Terry Sejnowski, a pioneer in neural networks and computational neuroscience, joins Craig Smith to discuss the future of AI, the evolution of ChatGPT, and the challenges of understanding intelligence. Terry, a key figure in the deep learning revolution, shares insights into how neural networks laid the foundation for modern AI, including ChatGPT's groundbreaking generative capabilities. From its ability to mimic human-like creativity to its limitations in true understanding, we explore what makes ChatGPT remarkable and what it still lacks compared to human cognition. We also dive into fascinating topics like the debate over AI sentience, the concept of "hallucinations" in AI models, and how language models like ChatGPT act as mirrors reflecting user input rather than possessing intrinsic intelligence. Terry explains how understanding language and meaning in AI remains one of the field's greatest challenges. Additionally, Terry shares his perspective on nature-inspired AI and what it will take to develop systems that go beyond prediction to exhibit true autonomy and decision-making. Learn why AI models like ChatGPT are revolutionary yet incomplete, how generative AI might redefine creativity, and what the future holds for AI as we continue to push its boundaries. Don't miss this deep dive into the fascinating world of AI with Terry Sejnowski. Like, subscribe, and hit the notification bell for more cutting-edge AI insights! Stay Updated: Craig Smith Twitter: https://twitter.com/craigss Eye on A.I. Twitter: https://twitter.com/EyeOn_AI (00:00) Introduction to Terry Sejnowski and His Work (03:02) The Origins of Modern AI and Neural Networks (05:29) The Deep Learning Revolution and ImageNet (07:11) Understanding ChatGPT and Generative AI (12:34) Exploring AI Creativity (16:03) Lessons from Gaming AI: AlphaGo and Backgammon (18:37) Early Insights into AI's Affinity for Language (24:48) Syntax vs. Semantics: The Purpose of Language (30:00) How Written Language Transformed AI Training (35:10) Can AI Become Sentient? (41:37) AI Agents and the Next Frontier in Automation (45:43) Nature-Inspired AI: Lessons from Biology (50:02) Digital vs. Biological Computation: Key Differences (54:29) Will AI Replace Jobs? (57:07) The Future of AI
Hi everyone!If you're a new subscriber or listener, welcome. If you're not new, you've probably noticed that things have slowed down from us a bit recently. Hugh Zhang, Andrey Kurenkov and I sat down to recap some of The Gradient's history, where we are now, and how things will look going forward. To summarize and give some context:The Gradient has been around for around 6 years now – we began as an online magazine, and began producing our own newsletter and podcast about 4 years ago. With a team of volunteers — we take in a bit of money through Substack that we use for subscriptions to tools we need and try to pay ourselves a bit — we've been able to keep this going for quite some time. Our team has less bandwidth than we'd like right now (and I'll admit that at least some of us are running on fumes…) — we'll be making a few changes:* Magazine: We're going to be scaling down our editing work on the magazine. While we won't be accepting pitches for unwritten drafts for now, if you have a full piece that you'd like to pitch to us, we'll consider posting it. If you've reached out about writing and haven't heard from us, we're really sorry. We've tried a few different arrangements to manage the pipeline of articles we have, but it's been difficult to make it work. We still want this to be a place to promote good work and writing from the ML community, so we intend to continue using this Substack for that purpose. If we have more editing bandwidth on our team in the future, we want to continue doing that work. * Newsletter: We'll aim to continue the newsletter as before, but with a “Best from the Community” section highlighting posts. We'll have a way for you to send articles you want to be featured, but for now you can reach us at our editor@thegradient.pub. * Podcast: I'll be continuing this (at a slower pace), but eventually transition it away from The Gradient given the expanded range. If you're interested in following, it might be worth subscribing on another player like Apple Podcasts, Spotify, or using the RSS feed.* Sigmoid Social: We'll keep this alive as long as there's financial support for it.If you like what we do and/or want to help us out in any way, do reach out to editor@thegradient.pub. We love hearing from you.Timestamps* (0:00) Intro* (01:55) How The Gradient began* (03:23) Changes and announcements* (10:10) More Gradient history! On our involvement, favorite articles, and some plugsSome of our favorite articles!There are so many, so this is very much a non-exhaustive list:* NLP's ImageNet moment has arrived* The State of Machine Learning Frameworks in 2019* Why transformative artificial intelligence is really, really hard to achieve* An Introduction to AI Story Generation* The Artificiality of Alignment (I didn't mention this one in the episode, but it should be here)Places you can find us!Hugh:* Twitter* Personal site* Papers/things mentioned!* A Careful Examination of LLM Performance on Grade School Arithmetic (GSM1k)* Planning in Natural Language Improves LLM Search for Code Generation* Humanity's Last ExamAndrey:* Twitter* Personal site* Last Week in AI PodcastDaniel:* Twitter* Substack blog* Personal site (under construction) Get full access to The Gradient at thegradientpub.substack.com/subscribe
Fei-Fei Li and Justin Johnson are pioneers in AI. While the world has only recently witnessed a surge in consumer AI, our guests have long been laying the groundwork for innovations that are transforming industries today.In this episode, a16z General Partner Martin Casado joins Fei-Fei and Justin to explore the journey from early AI winters to the rise of deep learning and the rapid expansion of multimodal AI. From foundational advancements like ImageNet to the cutting-edge realm of spatial intelligence, Fei-Fei and Justin share the breakthroughs that have shaped the AI landscape and reveal what's next for innovation at World Labs.If you're curious about how AI is evolving beyond language models and into a new realm of 3D, generative worlds, this episode is a must-listen.Resources: Learn more about World Labs: https://www.worldlabs.aiFind Fei-Fei on Twitter: https://x.com/drfeifeiFind Justin on Twitter: https://x.com/jcjohnss Stay Updated: Let us know what you think: https://ratethispodcast.com/a16zFind a16z on Twitter: https://twitter.com/a16zFind a16z on LinkedIn: https://www.linkedin.com/company/a16zSubscribe on your favorite podcast app: https://a16z.simplecast.com/Follow our host: https://twitter.com/stephsmithioPlease note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.
AI researcher Jim Fan has had a charmed career. He was OpenAI's first intern before he did his PhD at Stanford with “godmother of AI,” Fei-Fei Li. He graduated into a research scientist position at Nvidia and now leads its Embodied AI “GEAR” group. The lab's current work spans foundation models for humanoid robots to agents for virtual worlds. Jim describes a three-pronged data strategy for robotics, combining internet-scale data, simulation data and real world robot data. He believes that in the next few years it will be possible to create a “foundation agent” that can generalize across skills, embodiments and realities—both physical and virtual. He also supports Jensen Huang's idea that “Everything that moves will eventually be autonomous.” Hosted by: Stephanie Zhan and Sonya Huang, Sequoia Capital Mentioned in this episode: World of Bits: Early OpenAI project Jim worked on as an intern with Andrej Karpathy. Part of a bigger initiative called Universe Fei-Fei Li: Jim's PhD advisor at Stanford who founded the ImageNet project in 2010 that revolutionized the field of visual recognition, led the Stanford Vision Lab and just launched her own AI startup, World Labs Project GR00T: Nvidia's “moonshot effort” at a robotic foundation model, premiered at this year's GTC Thinking Fast and Slow: Influential book by Daniel Kahneman that popularized some of his teaching from behavioral economics Jetson Orin chip: The dedicated series of edge computing chips Nvidia is developing to power Project GR00T Eureka: Project by Jim's team that trained a five finger robot hand to do pen spinning MineDojo: A project Jim did when he first got to Nvidia that developed a platform for general purpose agents in the game of Minecraft. Won NeurIPS 2022 Outstanding Paper Award ADI: artificial dog intelligence Mamba: Selective State Space Models, an alternative architecture to Transformers that Jim is interested in (original paper here) 00:00 Introduction 01:35 Jim's journey to embodied intelligence 04:53 The GEAR Group 07:32 Three kinds of data for robotics 10:32 A GPT-3 moment for robotics 16:05 Choosing the humanoid robot form factor 19:37 Specialized generalists 21:59 GR00T gets its own chip 23:35 Eureka and Issac Sim 25:23 Why now for robotics? 28:53 Exploring virtual worlds 36:28 Implications for games 39:13 Is the virtual world in service of the physical world? 42:10 Alternative architectures to Transformers 44:15 Lightning round
Andrew Ilyas, a PhD student at MIT who is about to start as a professor at CMU. We discuss Data modeling and understanding how datasets influence model predictions, Adversarial examples in machine learning and why they occur, Robustness in machine learning models, Black box attacks on machine learning systems, Biases in data collection and dataset creation, particularly in ImageNet and Self-selection bias in data and methods to address it. MLST is sponsored by Brave: The Brave Search API covers over 20 billion webpages, built from scratch without Big Tech biases or the recent extortionate price hikes on search API access. Perfect for AI model training and retrieval augmentated generation. Try it now - get 2,000 free queries monthly at http://brave.com/api Andrew's site: https://andrewilyas.com/ https://x.com/andrew_ilyas TOC: 00:00:00 - Introduction and Andrew's background 00:03:52 - Overview of the machine learning pipeline 00:06:31 - Data modeling paper discussion 00:26:28 - TRAK: Evolution of data modeling work 00:43:58 - Discussion on abstraction, reasoning, and neural networks 00:53:16 - "Adversarial Examples Are Not Bugs, They Are Features" paper 01:03:24 - Types of features learned by neural networks 01:10:51 - Black box attacks paper 01:15:39 - Work on data collection and bias 01:25:48 - Future research plans and closing thoughts References: Adversarial Examples Are Not Bugs, They Are Features https://arxiv.org/pdf/1905.02175 TRAK: Attributing Model Behavior at Scale https://arxiv.org/pdf/2303.14186 Datamodels: Predicting Predictions from Training Data https://arxiv.org/pdf/2202.00622 Adversarial Examples Are Not Bugs, They Are Features https://arxiv.org/pdf/1905.02175 IMAGENET-TRAINED CNNS https://arxiv.org/pdf/1811.12231 ZOO: Zeroth Order Optimization Based Black-box https://arxiv.org/pdf/1708.03999 A Spline Theory of Deep Networks https://proceedings.mlr.press/v80/balestriero18b/balestriero18b.pdf Scaling Monosemanticity https://transformer-circuits.pub/2024/scaling-monosemanticity/ Adversarial Examples Are Not Bugs, They Are Features https://gradientscience.org/adv/ Adversarial Robustness Limits via Scaling-Law and Human-Alignment Studies https://proceedings.mlr.press/v235/bartoldson24a.html Prior Convictions: Black-Box Adversarial Attacks with Bandits and Priors https://arxiv.org/abs/1807.07978 Estimation of Standard Auction Models https://arxiv.org/abs/2205.02060 From ImageNet to Image Classification: Contextualizing Progress on Benchmarks https://arxiv.org/abs/2005.11295 Estimation of Standard Auction Models https://arxiv.org/abs/2205.02060 What Makes A Good Fisherman? Linear Regression under Self-Selection Bias https://arxiv.org/abs/2205.03246 Towards Tracing Factual Knowledge in Language Models Back to the Training Data [Akyürek] https://arxiv.org/pdf/2205.11482
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: The 'strong' feature hypothesis could be wrong, published by lsgos on August 2, 2024 on LessWrong. NB. I am on the Google Deepmind language model interpretability team. But the arguments/views in this post are my own, and shouldn't be read as a team position. "It would be very convenient if the individual neurons of artificial neural networks corresponded to cleanly interpretable features of the input. For example, in an "ideal" ImageNet classifier, each neuron would fire only in the presence of a specific visual feature, such as the color red, a left-facing curve, or a dog snout" : Elhage et. al, Toy Models of Superposition Recently, much attention in the field of mechanistic interpretability, which tries to explain the behavior of neural networks in terms of interactions between lower level components, has been focussed on extracting features from the representation space of a model. The predominant methodology for this has used variations on the sparse autoencoder, in a series of papers inspired by Elhage et. als. model of superposition. Conventionally there understood to be two key theories underlying this agenda. The first is the 'linear representation hypothesis' (LRH), the hypothesis that neural networks represent many intermediates or variables of the computation (such as the 'features of the input' in the opening quote) as linear directions in it's representation space, or atoms[1]. And second, the theory that the network is capable of representing more of these 'atoms' than it has dimensions in its representation space, via superposition (the superposition hypothesis). While superposition is a relatively uncomplicated hypothesis, I think the LRH is worth examining in more detail. It is frequently stated quite vaguely, and I think there are several possible formulations of this hypothesis, with varying degrees of plausibility, that it is worth carefully distinguishing between. For example, the linear representation hypothesis is often stated as 'networks represent features of the input as directions in representation space'. There are a few possible formulations of this: 1. (Weak LRH) some features used by neural networks are represented as atoms in representation space 2. (Strong LRH) all features used by neural networks are represented by atoms. The weak LRH I would say is now well supported by considerable empirical evidence. The strong form is much more speculative: confirming the existence of many linear representations does not necessarily provide strong evidence for the strong hypothesis. Both the weak and the strong forms of the hypothesis can still have considerable variation, depending on what we understand by a feature. I think that in addition to the acknowledged assumption of the LRH and superposition hypotheses, much work on SAEs in practice makes the assumption that each atom in the network will represent a "simple feature" or a "feature of the input". These features that the atoms are representations of are assumed to be 'monosemantic': they will all stand for features which are human interpretable in isolation. I will call this the monosemanticity assumption. This is difficult to state precisely, but we might formulate as the theory that every represented variable will have a single meaning in a good description of a model. This is not a straightforward assumption due to how imprecise the notion of a single meaning is. While various more or less reasonable definitions for features are discussed in the pioneering work of Elhage, these assumptions have different implications. For instance, if one thinks of 'features' as computational intermediates in a broad sense, then superposition and the LRH imply a certain picture of the format of a models internal representation: that what the network is doing is manipulating atoms in superposition (if y...
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: The 'strong' feature hypothesis could be wrong, published by lewis smith on August 2, 2024 on The AI Alignment Forum. NB. I am on the Google Deepmind language model interpretability team. But the arguments/views in this post are my own, and shouldn't be read as a team position. "It would be very convenient if the individual neurons of artificial neural networks corresponded to cleanly interpretable features of the input. For example, in an "ideal" ImageNet classifier, each neuron would fire only in the presence of a specific visual feature, such as the color red, a left-facing curve, or a dog snout" Elhage et. al, Toy Models of Superposition Recently, much attention in the field of mechanistic interpretability, which tries to explain the behavior of neural networks in terms of interactions between lower level components, has been focussed on extracting features from the representation space of a model. The predominant methodology for this has used variations on the sparse autoencoder, in a series of papers inspired by Elhage et. als. model of superposition.It's been conventionally understood that there are two key theories underlying this agenda. The first is the 'linear representation hypothesis' (LRH), the hypothesis that neural networks represent many intermediates or variables of the computation (such as the 'features of the input' in the opening quote) as linear directions in it's representation space, or atoms[1]. And second, the theory that the network is capable of representing more of these 'atoms' than it has dimensions in its representation space, via superposition (the superposition hypothesis). While superposition is a relatively uncomplicated hypothesis, I think the LRH is worth examining in more detail. It is frequently stated quite vaguely, and I think there are several possible formulations of this hypothesis, with varying degrees of plausibility, that it is worth carefully distinguishing between. For example, the linear representation hypothesis is often stated as 'networks represent features of the input as directions in representation space'. Here are two importantly different ways to parse this: 1. (Weak LRH) some or many features used by neural networks are represented as atoms in representation space 2. (Strong LRH) all (or the vast majority of) features used by neural networks are represented by atoms. The weak LRH I would say is now well supported by considerable empirical evidence. The strong form is much more speculative: confirming the existence of many linear representations does not necessarily provide strong evidence for the strong hypothesis. Both the weak and the strong forms of the hypothesis can still have considerable variation, depending on what we understand by a feature and the proportion of the model we expect to yield to analysis, but I think that the distinction between just a weak and strong form is clear enough to work with. I think that in addition to the acknowledged assumption of the LRH and superposition hypotheses, much work on SAEs in practice makes the assumption that each atom in the network will represent a "simple feature" or a "feature of the input". These features that the atoms are representations of are assumed to be 'monosemantic': they will all stand for features which are human interpretable in isolation. I will call this the monosemanticity assumption. This is difficult to state precisely, but we might formulate it as the theory that every represented variable will have a single meaning in a good description of a model. This is not a straightforward assumption due to how imprecise the notion of a single meaning is. While various more or less reasonable definitions for features are discussed in the pioneering work of Elhage, these assumptions have different implications. For instance, if one thinks of 'feat...
Key Topics & Chapter Markers:AI's Evolutionary Journey & Key Challenges [00:00:00]Neural Networks: Inspiration from Biology [00:01:00]Weighted Sum, Inputs & Mathematical Functions [00:05:00]Gradient Descent & Optimization in Neural Nets [00:10:15]Computing Architecture: CPUs vs. GPUs [00:39:56]RNNs and Early Problems in Memory & Context [01:03:00]The Emergence of Convolutional Neural Networks (CNNs) [01:10:00]ImageNet, GPUs & Scaling Neural Networks [01:24:00]Share Your Thoughts: Have questions or comments? Drop us a mail at EffortlessPodcastHQ@gmail.com
We meet Dr. Fei-Fei Li In the latest installment of our oral history project. She's a Chinese-American computer scientist and the creator of ImageNet - the dataset that made rapid advances possible in this field of AI that helps computers take meaningful information from things like photos and videos.We Meet: Stanford University's Fei-Fei Li, author of "The Worlds I See: Curiosity, Exploration, and Discovery at the Dawn of AI"Credits:This episode of SHIFT was produced by Jennifer Strong with help from Emma Cillekens. It was mixed by Garret Lang, with original music from him and Jacob Gorski. Art by Anthony Green.
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Rational Animations' intro to mechanistic interpretability, published by Writer on June 15, 2024 on LessWrong. In our new video, we talk about research on interpreting InceptionV1, a convolutional neural network. Researchers have been able to understand the function of neurons and channels inside the network and uncover visual processing algorithms by looking at the weights. The work on InceptionV1 is early but landmark mechanistic interpretability research, and it functions well as an introduction to the field. We also go into the rationale and goals of the field and mention some more recent research near the end. Our main source material is the circuits thread in the Distill journal and this article on feature visualization. The author of the script is Arthur Frost. I have included the script below, although I recommend watching the video since the script has been written with accompanying moving visuals in mind. Intro In 2018, researchers trained an AI to find out if people were at risk of heart conditions based on pictures of their eyes, and somehow the AI also learned to tell people's biological sex with incredibly high accuracy. How? We're not entirely sure. The crazy thing about Deep Learning is that you can give an AI a set of inputs and outputs, and it will slowly work out for itself what the relationship between them is. We didn't teach AIs how to play chess, go, and atari games by showing them human experts - we taught them how to work it out for themselves. And the issue is, now they have worked it out for themselves, and we don't know what it is they worked out. Current state-of-the-art AIs are huge. Meta's largest LLaMA2 model uses 70 billion parameters spread across 80 layers, all doing different things. It's deep learning models like these which are being used for everything from hiring decisions to healthcare and criminal justice to what youtube videos get recommended. Many experts believe that these models might even one day pose existential risks. So as these automated processes become more widespread and significant, it will really matter that we understand how these models make choices. The good news is, we've got a bit of experience uncovering the mysteries of the universe. We know that humans are made up of trillions of cells, and by investigating those individual cells we've made huge advances in medicine and genetics. And learning the properties of the atoms which make up objects has allowed us to develop modern material science and high-precision technology like computers. If you want to understand a complex system with billions of moving parts, sometimes you have to zoom in. That's exactly what Chris Olah and his team did starting in 2015. They focused on small groups of neurons inside image models, and they were able to find distinct parts responsible for detecting everything from curves and circles to dog heads and cars. In this video we'll Briefly explain how (convolutional) neural networks work Visualise what individual neurons are doing Look at how neurons - the most basic building blocks of the neural network - combine into 'circuits' to perform tasks Explore why interpreting networks is so hard There will also be lots of pictures of dogs, like this one. Let's get going. We'll start with a brief explanation of how convolutional neural networks are built. Here's a network that's trained to label images. An input image comes in on the left, and it flows along through the layers until we get an output on the right - the model's attempt to classify the image into one of the categories. This particular model is called InceptionV1, and the images it's learned to classify are from a massive collection called ImageNet. ImageNet has 1000 different categories of image, like "sandal" and "saxophone" and "sarong" (which, if you don't know, is a k...
Speakers for AI Engineer World's Fair have been announced! See our Microsoft episode for more info and buy now with code LATENTSPACE — we've been studying the best ML research conferences so we can make the best AI industry conf! Note that this year there are 4 main tracks per day and dozens of workshops/expo sessions; the free livestream will air much less than half of the content this time.Apply for free/discounted Diversity Program and Scholarship tickets here. We hope to make this the definitive technical conference for ALL AI engineers.ICLR 2024 took place from May 6-11 in Vienna, Austria. Just like we did for our extremely popular NeurIPS 2023 coverage, we decided to pay the $900 ticket (thanks to all of you paying supporters!) and brave the 18 hour flight and 5 day grind to go on behalf of all of you. We now present the results of that work!This ICLR was the biggest one by far, with a marked change in the excitement trajectory for the conference:Of the 2260 accepted papers (31% acceptance rate), of the subset of those relevant to our shortlist of AI Engineering Topics, we found many, many LLM reasoning and agent related papers, which we will cover in the next episode. We will spend this episode with 14 papers covering other relevant ICLR topics, as below.As we did last year, we'll start with the Best Paper Awards. Unlike last year, we now group our paper selections by subjective topic area, and mix in both Outstanding Paper talks as well as editorially selected poster sessions. Where we were able to do a poster session interview, please scroll to the relevant show notes for images of their poster for discussion. To cap things off, Chris Ré's spot from last year now goes to Sasha Rush for the obligatory last word on the development and applications of State Space Models.We had a blast at ICLR 2024 and you can bet that we'll be back in 2025
When a massive direct-view LED video wall is installed to make a completely immersive experience, teams collaborate to make sure it's a successful project. To hear all the details of this project for Flogistix, Justin and Matt are joined by Ali Sylvester, Director of Business Solutions at Flogistix, and Kyle Kempf, CTS-I Director of Commercial Audio Video at ImageNet Consulting. They dig into the vision for the space, the architecture that went into it and everything else that brings the project to life. Links: Daktronics News Release: https://www.daktronics.com/news/imagenet-and-daktronics-deliver-led-video-wall-experience-for-flogistix Flogistix Website: https://flogistix.com/ ImageNet Consulting Website: https://www.imagenetconsulting.com/ Rand Elliott Architects Website: https://randelliottarchitects.com/ Daktronics and ImageNet Podcast: https://podcast.daktronics.com/e/143-imagenet-consulting-with-kyle-kempf/
At 15, Fei-Fei Li transitioned from a middle-class life in China to poverty in America. Despite the pressures of her family's financial situation and her mother's ailing health, her knack for physics never wavered. She went from learning English as a second language to attending and working at prestigious institutions like Princeton and Stanford. Today, she is among a handful of scientists behind the impressive advances of artificial intelligence in recent times. In this episode, she breaks down her human-centered approach to AI and explores the future of the technology. Dr. Fei-Fei Li is a professor of Computer Science at Stanford University and the co-director of the Stanford Institute for Human-Centered AI. She is the creator of ImageNet, a key driver of modern artificial intelligence. With over 20 years at the forefront of the field, Dr. Li is focused on AI research, education, and policy to improve the human condition. In this episode, Hala and Fei-Fei will discuss: - The current capabilities of AI - The difference between machine learning and AI - The training process for AI models - The gaps in our knowledge about how AI learns - Why ChatGPT fails at higher-level reasoning like math - The biological inspiration for vision in computers - Fears and hopes associated with AI - The human element of jobs AI can't replace - Augmentation of human capabilities through AI - The three pillars of her human-centered AI framework - Responsible development and use of AI - The roadblocks to be aware of when using AI - Her advice to young entrepreneurs navigating the AI world - And other topics… Dr. Fei-Fei Li is a professor of Computer Science at Stanford University and the co-director of the Stanford Institute for Human-Centered AI. She is also the creator of ImageNet and the ImageNet Challenge, a key catalyst to the latest developments in deep learning and AI. Sometimes called the ‘Godmother of AI,' she is a pioneer in early computer vision research. Dr. Li is the author of The Worlds I See, one of Barack Obama's recommended books on AI. Her work has been featured in various publications, including the New York Times, Wall Street Journal, Fortune Magazine, Science, and Wired Magazine. Connect with Fei-Fei: Fei-Fei's Bio: https://profiles.stanford.edu/fei-fei-li Fei-Fei's LinkedIn: https://www.linkedin.com/in/fei-fei-li-4541247/ Fei-Fei's Twitter: https://twitter.com/drfeifei Resources Mentioned: Fei-Fei's Book, The Worlds I See: Curiosity, Exploration, and Discovery at the Dawn of AI: https://www.amazon.com/Worlds-See-Curiosity-Exploration-Discovery-ebook/dp/B0BPQSLVL6 Stanford Human Center AI Institute Website: https://hai.stanford.edu/ LinkedIn Secrets Masterclass, Have Job Security For Life: Use code ‘podcast' for 30% off at yapmedia.io/course. Sponsored By: Shopify - Sign up for a one-dollar-per-month trial period at youngandprofiting.co/shopify Indeed - Get a $75 job credit at indeed.com/profiting Yahoo Finance - For comprehensive financial news and analysis, visit YahooFinance.com More About Young and Profiting Download Transcripts - youngandprofiting.com Get Sponsorship Deals - youngandprofiting.com/sponsorships Leave a Review - ratethispodcast.com/yap Watch Videos - youtube.com/c/YoungandProfiting Follow Hala Taha LinkedIn - linkedin.com/in/htaha/ Instagram - instagram.com/yapwithhala/ TikTok - tiktok.com/@yapwithhala Twitter - twitter.com/yapwithhala Learn more about YAP Media's Services - yapmedia.io/
Fei-Fei Li is a Stanford computer scientist and the former chief scientist of artificial intelligence/machine learning at Google Cloud. When Li entered the field of AI in the 2000s, researchers were making slow progress, optimizing algorithms to incrementally improve outcomes. Li saw that the problem wasn't the algorithm, but the size of the datasets being used. So she built a massive database of images called ImageNet. It was a huge breakthrough, and helped lead the emergence of modern AI.See omnystudio.com/listener for privacy information.
Where did AI come from? Who created it, why, and where can it lead? Artificial intelligence (AI) is rapidly developing into a world-changer, affecting every industry and being used by hundreds of millions of people—even when they're unaware they're interacting with an artificial intelligence. And we're only at the early stages of AI's growth. Join us for an in-depth talk with Dr. Fei-Fei Li, whom Wired called "one of a tiny group of scientists―a group perhaps small enough to fit around a kitchen table―who are responsible for AI's recent remarkable advances.” Dr. Li came to America as an immigrant, enduring a shift from Chinese middle class to American poverty. But a tough upbringing did not stop her from becoming a leading mind in the next big technological development. Fei-Fei's adolescent knack for physics endured and positioned her to make a crucial contribution to the breakthrough we now call AI, placing her at the center of a global transformation. Over the last decades, her work has brought her face-to-face with the extraordinary possibilities―and the extraordinary dangers―of the technology she loves. Known as the creator of ImageNet, a key catalyst of modern artificial intelligence, Dr. Li has spent more than two decades at the forefront of the field. Her work has brought her face-to-face with the extraordinary possibilities―and the extraordinary dangers―of the technology she loves. Don't miss this opportunity to learn more about a breakthrough science and one of the breakthrough scientists who is making it happen. This program is part of our Good Lit series, underwritten by the Bernard Osher Foundation. Learn more about your ad choices. Visit megaphone.fm/adchoices
Dr. Fei-Fei Li is a literal visionary. Her groundbreaking work on ImageNet, a vast visual recognition database, helped propel artificial intelligence at a critical moment. As one of the key innovators and thinkers in AI, Li has argued for a human-centered artificial intelligence that augments people's capabilities instead of displacing them. We talk to Li about her work, her vision for AI and her new memoir, The Worlds I See, in which she recounts her journey as a scientist and immigrant, and how those two roles inform each other. Guests: Fei-Fei Li, professor of Computer Science Department, Stanford University; author, "The Worlds I See: Curiosity, Exploration, and Discovery at the Dawn of AI"
A year ago, the public launch of ChatGPT took the world by storm and it was followed by many more generative artificial intelligence tools, all with remarkable, human-like abilities. Fears over the existential risks posed by AI have dominated the global conversation around the technology ever since. A pioneer that helped lay the groundwork that underpins generative AI models, Fei-Fei Li, takes a more nuanced approach to. She's pushing for a human-centred way of dealing with AI—treating it as a tool to help enhance—and not replace—humanity, while focussing on the pressing challenges of disinformation, bias and job disruption.Fei-Fei Li, a pioneer that helped lay the groundwork that underpins modern generative AI models, takes a more nuanced approach. She's pushing for a human-centred way of dealing with AI—treating it as a tool to help enhance—and not replace—humanity, while focussing on the pressing challenges of disinformation, bias and job disruption.Fei-Fei Li is the founding co-director of Stanford University's Institute for Human-Centred Artificial Intelligence. Fei-Fei and her research group created ImageNet, a huge database of images that enabled computers scientists to build algorithms that were able to see and recognise objects in the real world. That endeavour also introduced the world to deep learning, a type of machine learning that is fundamental part of how large-language and image-creation models work.Host: Alok Jha, The Economist's science and technology editor. Sign up for a free trial of Economist Podcasts+. If you're already a subscriber to The Economist, you'll have full access to all our shows as part of your subscription. For more information about how to access Economist Podcasts+, please visit our FAQs page or watch our video explaining how to link your account. Hosted on Acast. See acast.com/privacy for more information.
A year ago, the public launch of ChatGPT took the world by storm and it was followed by many more generative artificial intelligence tools, all with remarkable, human-like abilities. Fears over the existential risks posed by AI have dominated the global conversation around the technology ever since. Fei-Fei Li, a pioneer that helped lay the groundwork that underpins modern generative AI models, takes a more nuanced approach. She's pushing for a human-centred way of dealing with AI—treating it as a tool to help enhance—and not replace—humanity, while focussing on the pressing challenges of disinformation, bias and job disruption.Fei-Fei Li is the founding co-director of Stanford University's Institute for Human-Centred Artificial Intelligence. Fei-Fei and her research group created ImageNet, a huge database of images that enabled computers scientists to build algorithms that were able to see and recognise objects in the real world. That endeavour also introduced the world to deep learning, a type of machine learning that is fundamental part of how large-language and image-creation models work.Host: Alok Jha, The Economist's science and technology editor. Sign up for a free trial of Economist Podcasts+. If you're already a subscriber to The Economist, you'll have full access to all our shows as part of your subscription. For more information about how to access Economist Podcasts+, please visit our FAQs page or watch our video explaining how to link your account. Hosted on Acast. See acast.com/privacy for more information.
Fei-Fei Li, PhD, Professor in the Computer Science Department at Stanford University, and Co-Director of Stanford's Human-Centered AI Institute, joins Bio + Health founding partner Vijay Pande.In this candid conversation, Li unfolds her transformation from a young immigrant to an influential figure in AI. The conversation explores the birth of ImageNet, a pivotal step that bridged the gap between visual intelligence and accessible AI technology. They delve into the notion of a 'Dignity Economy,' hinting at a future where technology serves to elevate human experience rather than undermine it. Li also touches on the delicate balance between relentless innovation and life's humble pursuits. This episode peels back the layers on the human side of AI, offering a rare glimpse into the personal and professional realms of a pioneer shaping the AI landscape.Check out her new book, The Worlds I See, here: https://us.macmillan.com/books/9781250897930/theworldsiseeCheck out other episodes form our sister podcast, Bio Eats World: https://a16z.com/podcasts/bio-eats-world/ Stay Updated: Find a16z on Twitter: https://twitter.com/a16zFind a16z on LinkedIn: https://www.linkedin.com/company/a16zSubscribe on your favorite podcast app: https://a16z.simplecast.com/Follow our host: https://twitter.com/stephsmithioPlease note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.
Fei-Fei Li's new book is the story of her journey from China to the U.S., from small business to Big Tech, and from academic research to corporate life, and back again. But more than that, it's the story of the dawn of artificial intelligence, as told through her experience as one of the people summoning this new day and standing there awestruck, excited and concerned about what it will mean for humanity. Dr. Li joins us on this episode to discuss the book, The Worlds I See: Curiosity, Exploration, and Discovery at the Dawn of AI, published by Moment of Lift Books, an imprint from Melinda French Gates and Flatiron Books. Known for her foundational contributions to AI and computer vision, Dr. Li is the inventor of ImageNet, a large-scale dataset of images that enabled rapid advances in deep learning for visual recognition. She is a professor of computer science at Stanford University and a co-director of the Stanford Institute for Human-Centered Artificial Intelligence, who worked as Google Cloud's chief scientist for AI/ML during a 2017-2018 sabbatical. Note: GeekWire's Todd Bishop will be speaking further with Dr. Li on Monday evening Nov. 13 at Town Hall in Seattle. See this site for details and tickets. Edited by Curt Milton.See omnystudio.com/listener for privacy information.
Fei-Fei Li, PhD, Professor in the Computer Science Department at Stanford University, and Co-Director of Stanford's Human-Centered AI Institute, joins Bio + Health founding partner Vijay Pande.In this candid conversation, Li unfolds her transformation from a young immigrant to an influential figure in AI. The conversation explores the birth of ImageNet, a pivotal step that bridged the gap between visual intelligence and accessible AI technology. They delve into the notion of a 'Dignity Economy,' hinting at a future where technology serves to elevate human experience rather than undermine it. Li also touches on the delicate balance between relentless innovation and life's humble pursuits. This episode peels back the layers on the human side of AI, offering a rare glimpse into the personal and professional realms of a pioneer shaping the AI landscape.Check out her new book, out November 7, 2023, here: https://us.macmillan.com/books/9781250897930/theworldsisee