POPULARITY
When we first dicsussed the Summer of Simulative AI in 2024 we knew it would be a brief summer, but it has recently come back with a vengeance with SimGym in April and now Simile AI's $2B Series B, backed by GreenOaks and Index Ventures with prominent backers like Fei-Fei Li and Andrej Karpathy, running tens of millions of simulations for Fortune 100 clients like CVS and 85–99% accuracy vs human focus groups. Time to catch up on why this Second Summer of simulation is working!From creating Smallville, the landmark 2023 paper on Generative Agents that showed AI characters could remember, plan, socialize, and develop emergent behaviors, to now building foundation models of human behavior, Joon Sung Park is trying to answer a much bigger question: what if we could simulate the world before making decisions in it? In this episode, the Simile co-founder and CEO joins us to unpack the path from generative agents to digital twins, why today's frontier models still fail to capture how humans actually behave, and what it would take to eventually simulate all 8 billion people on Earth.We go deep on Simile's approach to modeling human behavior: long-form interviews, observational and transaction data, randomized controlled trials, population-level and individual-level models, and post-training on the causal mechanisms behind why people make decisions. Joon explains how his research created digital twins that reproduced human behavior and attitudes 85% as accurately as people reproduced their own responses, why models optimized to be rational can be bad simulations of irrational humans, and why understanding “social physics” may require changing model weights rather than simply prompting frontier LLMs.We also explore the much larger ambition behind simulation: testing products and policies before deploying them, finding counterintuitive paths toward desired outcomes, modeling emergent behavior across entire societies, and potentially tackling problems like climate change, democratic instability, and UBI. Joon reflects on scaling laws for simulation, the economics of data-center-scale simulated worlds, the connection to Thomas Schelling and psychohistory, why simulation is surprisingly similar to painting, and whether we might already be living in one.We discuss:* How Smallville and Generative Agents led to Simile* Why Joon's team asked: “What if we can just recreate the world that we live in?”* Why useful personal agents require deep models of their users* Memory architectures, Markdown files, and the limits of prompting* “Social physics” and behavioral foundation models* Why web data captures what people say more than what they actually do* Interviews, transactions, observational data, and randomized controlled trials* Why predicting the future matters less than understanding how to shape it* How Simile creates representative simulated populations* Simulation versus prediction and the connection to Foundation's psychohistory* How to evaluate simulations instead of simply stacking LLM hallucinations* Creating digital twins of 1,000 real people and reaching 85% behavioral accuracy* Why frontier models can struggle to reproduce real human behavior* Why good simulations need to reproduce human biases and mistakes* Post-training models on randomized controlled trials* Population-level versus individual-level simulation* Scaling laws for human simulation* The long-term ambition to simulate all 8 billion people on Earth* Whether simulations could help solve climate change or detect collapsing democracy* Thomas Schelling and the history of agent-based modeling* Why future simulations could require an entire data center* Multi-agent simulations and what happens when simulated people interact* Replacing expensive human panels with synthetic populations* Why market research is only the starting point for simulation* Why Joon sees simulation as surprisingly similar to painting* Using simulation to study questions like UBI* Whether we are already living in a simulation* Why AGI and simulation may be the twin technologies of advanced civilizationsJoon Sung Park* LinkedIn: https://www.linkedin.com/in/joonspark* X: https://x.com/joon_s_pk* Website: https://www.joonsungpark.com* Simile: https://www.simile.comTimestamps00:00:00 Introduction and Joon's Path from Art to AI00:01:46 Smallville, Generative Agents, and the Origins of Simulation00:05:03 “Let's Just Create a World” and the Future of Personal Agents00:09:53 Social Physics and Behavioral Foundation Models00:14:08 Prediction vs. Simulation: How Do You Shape the Future?00:16:59 How Simile Models Real People and Populations00:25:35 Evaluating Simulations, Digital Twins, and 85% Accuracy00:30:23 Post-Training Models to Reproduce Human Behavior00:40:04 Scaling Laws and Simulating 8 Billion People00:43:10 From Schelling to Society-Scale Agent Simulations00:46:13 The Cost and Economics of Simulating the World00:52:05 Real-World Use Cases, Synthetic Populations, and the Market00:57:27 The Future of Simulation, Painting, and UBI01:04:23 Are We Already Living in a Simulation?01:06:08 Building Simile and HiringTranscriptIntroduction: Joon Sung Park, Simile, and the Story So FarVibhu [00:00:00]: Today, we have Joon in the podcast. Excited to kick this one off. Very exciting company. I wanna kick off and ask you the question, talk us through the story of your life. How have you gotten here?Joon [00:00:13]: Yeah, for sure. I'm really excited to be here. A story of my life. So I was born in Korea, and I lived there for a good 11 years or so of my life, and then my family moved to Boston. So we moved when I was 11, and my parents were doctors, so they were going through their postdoctoral studies. My dad was a surgeon, so he was doing his sabbatical years at the Boston Children's Hospital. So I grew up there, not too close to tech. I was very much a music and artsy, painting kind of guy.Vibhu [00:00:49]: Painting.Joon [00:00:49]: Exactly. I got into painting a little bit later, in high school, but that's what I used to do. And then I grew up mostly in the East Coast after Korea. So I lived a good number of years in New Hampshire, and then I went to college in Pennsylvania. And I got into more of this tech scene, in college. So I was originally trained to be an artist. I thought that would be my professional career. So it wasn't a hobby. It was like, “Hey, let's make a living out of this.” And then gradually, I got really interested in this idea of, hey, the greatest artist often creates their own medium, and the best medium that we had available today was in computation. So I decided to go deeper into that, and one thing led to another, and we can go deeper into this, but I decided that research was something that I gradually got interested in, and here I am.Smallville, Generative Agents, and the 2023 Breakout PaperSwyx [00:01:46]: So there's a lot that you packed into the research components. You had one of the best papers of 2023, which was the generative agents paper, commonly known as the Smallville paper.Swyx [00:01:58]: Feel free to call back to anything else that you mentioned, but most people would have heard of you from this. Do you have any statistics on how many people have, like, read it? arXiv gives you something, right? Some stats.Joon [00:02:10]: Yeah, it's a good question. How many people have read it, I'm not sure.Joon [00:02:14]: I know we do keep track of citations, and they are going up quite fast.Swyx [00:02:23]: Yeah, Google Scholar has 7,200 citations.Vibhu [00:02:25]: I feel like it made a bigger hit than that, and it was a pretty instrumental paper. It got cited so many times.Swyx [00:02:34]: It is frequently the answer when people ask, “What is the best paper you've read recently?” It's this one.Vibhu [00:02:39]: I thought the memory component was pretty underrated. It was a very good early memory system, and one of the biggest papers.Foundation Models and the Search for Killer ApplicationsJoon [00:02:47]: Yeah, so maybe I can talk a little bit about how this particular paper came together. So when I got into research, it was back in 2020 when I started my PhD program at Stanford, and that was the year, when we were about to get GPT-3 to be available. So we already had GPT-2, and you could sense that there was this new class of models that was just becoming available in the market, and the team got very intrigued. And the general consensus was, “Well, is this model going to be useful for anything?” “It's really strange that these models are not trained to do any particular task.” But we decided to take a bet. So a large group of scholars at Stanford, and it was led by one of my co-founders, Percy Liang, and we came togetherSwyx [00:03:35]: Who coined foundation models.Joon [00:03:36]: Who coined the term foundation models. We wrote this paper, where that term came from called Opportunities and Risks of Foundation Models. And during that process, really the thing that I started to think deeply about was, here is a model that is fundamentally new in our ecosystem. The reason why this was new was it wasn't, again, trained to do anything in particular, but its premise was it could do anything and everything. It was like a stem cell, if you were to take a biology analogy. And I got really interested in this idea that, well, if we were to really think about what are the killer applications that this particular technology would enable, what would that be? Many of my colleagues were using this for simple classification, simple generations. Interesting that these models can do that, but from an interaction perspective, not that interesting. We've known how to do that for many decades. And what we came down to was these models are trained on this very broad data from the web, right? So these are human behavioral data. It's social media, Wikipedia, all these data. So if you poke at the right angle, then you could see human behavior that would just pop out that's quite realistic, and we've never seen that before.The Time Machine Game and Recreating the WorldJoon [00:04:45]: So that got us really interested. The exercise that we decided to do, with this particular group of colleagues, Michael Bernstein, Percy Liang, and myself, who ended up becoming my co-founder at Simile, we sat down and we played this game that we call the time machine game.Joon [00:05:03]: Imagine we were to get on a time machine and fast-forward 10 years and look back. What would have been the single application that will have mattered that would be the most interesting and inspiring? And when we thought, “Well, what if we can just recreate the world that we live in?” it's really hard to get more ambitious than that. Like, let's just create a world.Joon [00:05:24]: And that's where we started. And initially, we had this paper that was a precursor to the generative agents paper called Social Simulacra.Swyx [00:05:32]: Before you go further, were there other candidates for the most ambitious thing in the time machine exercise? What was number two or number three?Personal Agents, User Models, and Why Simulation Came FirstJoon [00:05:44]: There is a close second that we were considering, which ended up becoming more of these automation tools, especially the vision around really personalized agents that would do things for you.Swyx [00:05:59]: That's also happening.Joon [00:06:00]: It's also happening. But it was interesting for us, right, in that the reason why, we decided to go with the idea of simulation, one, I was a huge science fiction nerd, and this idea of creating simulation, I was personally really just fascinated. I loved the idea. It's really cool to see, like, a game town like this and just see these agents live in it. But at the same time, my bet was if you were to create a really amazing personal assistant out of this technology, what you need first is an amazing model of your users. So I told a model, “Hey, can you go buy late dinner for me?” And it orders Hawaiian pizza, and I do not like pineapples on my pizza. Then it totally failed. The way for it to not make that mistake is only by having a deep understanding of who I am. And I gave a very simple and dumb example here, but you can imagine how this core understanding of people is instrumental. This is how, if we have our family and closest friends, they have a good mental model of who we are. That's the basis of our social connection. So our bet also was this technology around simulation, creating accurate representation of people ought to precede the more complex agents that would automate the world that we live in. So that was the bet. But that was a very close second, and I'm still very much fascinated by it. I think there's a lot of interesting work that's going around. My hot take here, though, is I don't think we've seen a true personal assistant that's useful, in ways that meet the ambition of that particular line of work. I think there are early applications that are interesting, and if you talk to even ChatGPT nowadays or Claude, they know a lot about us. So a lot of the generation it's doing, I do think it's much more tailored, but I think the ambition is quite large in that field, and I don't think we quite have all the right ingredients just yet.Swyx [00:08:01]: So OpenClaw and these personal agents, what do you want to see from them that they don't currently have?Memory, Markdown, and the Limits of PromptingJoon [00:08:09]: I do think it's slowly getting there, but I do generally want them to have much deeper understanding of the person. Right now, you look at the models. OpenClaw, what it's leveraging is a Markdown file, and I think it's quite clever, right? So if you look at the generative agents paper, this was the same intuition that we had, where initially when we were creating the memory architecture for the generative agents, and, like, this is, like, back in 2022, so we didn't really quite have the idea of even agentive architecture or the term agent. But the intuition that we shared with some of the work that's coming out today was we initially thought, “Well, do we want to make the memory into, let's say, knowledge graph? Do we want to train a bespoke model?” All of these things. And what we decided to do was, “No. Just forget about all this.” These language models are quite good at modeling text and understanding and reasoning about text. So just put everything in a Markdown file or a text file. You're done. I thought that was quite interesting that we could do that, and there's a lot of strength in doing that. But also, there are limitations. It's the way you retrieve and make sense of data that's extremely large, it takes a lot of work. So I think that technology is getting better. I also do, however, think, there are certain things you just cannot shape just by prompting the model. So to some degree, you do need to touch the parameters of the model itself. So there is this work that I do think does need to happen, and it is happening. The question is, how far can we take it? How do we source data, and how do you also create an ecosystem where people are continuously feeding data to this model so it's learning about you?Vibhu [00:09:50]: What's the intuition between why you need to do it in the model?Social Physics and Behavior Foundation ModelsJoon [00:09:53]: My intuition behind the actual when do you train or even post-train a model versus just prompt a model is if the model has to learn the underlying physics of the world that it's operating in. So it has to learn new social physics. The places where it doesn't have to train are the places where it already has the physics. We trust the physics. It already has the base statistics, but it's just trying to react to an environment. Then I think you can just prompt your way into getting the actions out of it. I don't think the models that are out in the open have yet learned the complete mapping of social physics of humanity. This is one of the core theses of Simile, right? And one of the core reasons why that is the case is if you look at the data that the model was trained on, these models were trained on the web data, like, whatever was available on the web. And these are really interesting data sets, but they are fundamentally the self-exposed attitudinal data with some behavior data that's sprinkled around here and there. And it has yet to learn the really deep behavioral nature of people, not just what people say they do online, but what they do in real life. And this is one of what I would consider to be the dark knowledge of humanity that we haven't quite captured. And it's these data that would also need to get factored into the model creation.Vibhu [00:11:21]: You call it behavior foundation model.Vibhu [00:11:23]: There's a good one-liner here, but outside of that, what type of data do you need? What are you changing on the model level? How do you go about modeling, doing a behavior foundation model?The Three Data Buckets: Interviews, Behavior, and CausalityJoon [00:11:35]: We think about data in three buckets. So one bucket is interview data. It's quite interesting. Rich qualitative data is interesting. It's not behavioral, but we would literally ask people, “Hey, tell me the story of your life.”Vibhu [00:11:53]: It's just what we're doing here exactly.Joon [00:11:54]: The question that you all asked at the beginning of this interview literally is the question we also ask. And we ask our participants to go a little bit deeper, than how far I went. Maybe I can give more of my life story in lieu of this. But the reason why that data is interesting is by learning about this very long-tail information about people, you get a lot of texture around this model, like, this person as a model. So even understanding their childhood memory or even their trauma, their first love, these things, quite informative in ways that's really hard to predict. So that's one. Then there are two tranches of what I would consider to be the behavioral data. One kind of behavioral data is observational. So these might be like transaction data, or these might be data that you can get by scraping the web, right? So you can imagine why these data sets would be interesting, right, because they give you the base statistics of people's behavior.Joon [00:12:55]: But then there is the last category of data, that I personally think is perhaps the most important, which is the data that describes the causal mechanism, the whys of people. Some of this is covered by the interview data, the qualitative, because people talk about why they made certain decisions. But really, where you get to see the most behavioral aspect of this is in randomized controlled trials, like RCTs. Imagine you have the same setup, but you have a few different variables that you are trying to tweak. Can you get realistic human behavior out of it in ways where, imagine you had this particular option. Imagine you're even trying to choose whether you're going to drink coffee or not. The day you drink coffee versus the day you didn't drink coffee, does your behavior change? That's a data set that describes a causal mechanism. This is quite important in modeling people. The reason why this is important is oftentimes when people come to us, or not just to us, but the reason why people are interested in simulation isn't because they want to predict the future. If you're trying to win against the stock market, predicting the future is interesting.Prediction vs. Simulation: Shaping the FutureJoon [00:14:08]: But most people, most decision-makers, what they want to know is, how can we shape the future? It doesn't really help you to hear that your sales are going to tank in two quarters. They're just gonna say, “Wow, that sucks.” What they want to know is, well, what do we need to do now to avoid that future? That's the causal mechanism. And this is also very hard data to come by, right, because the world is our ground truth, but it happens once. So in a very controlled setup where everything is equal except for one variable, this kind of data set rarely happens. So this is a reason why this data set is both hard to come by and quite important if you're trying to model human behavior.Swyx [00:14:50]: So behavior, I think, is the hardest data set to acquire. What is out there? What is even possible? You're not going to know a lot of details about my life. I don't even have data for myself on my own health or habits, and I just don't log everything. So how can you have that data?Joon [00:15:14]: So we run a lot of randomized controlled trials.Swyx [00:15:17]: But you put people in the lab, they watch them sleep, or what?Joon [00:15:20]: We do care a lot about the consent process. People know that we invite them to be a member of this community to both share data and have themselves represented in different forms. But we bring a lot of people to the lab, or virtual lab, where we design experiments that would pose them real behavioral decisions. And often in these experimental setups, what makes the difference between what is attitudinal versus behavioral is whether the stake in your decision is real. That's ultimately what makes it behavioral. So in these setups, we are inspired by our colleagues in social sciences, psychology, and so forth. So when they run studies, the techniques they utilize is imagine there's an online store that you're inviting people to come by. Then whatever they purchase in this experiment, they actually get that item delivered. Like, these are the things that make the stakes real. So we run a lot of these experiments, and we also do partner with firms. Right now, we also have customers who are quite excited to at least give us a glimpse of the behaviors that their users exhibit so that we can get a little bit deeper understanding of how people behave in these different platforms.How Customers Use Simile: Populations, Queries, and ExperimentsVibhu [00:16:39]: I think on the customer side, they have a lot of data about their users, who has bought. They have the action data.Vibhu [00:16:47]: Can you walk us through an example of what someone comes to you for? What questions would they want solved? Do you customize a model for them? Do you have something off the shelf? What does that look like?Joon [00:16:59]: Today, when people leverage our models, it's often to better understand the population of their interest. So usually, the start of the relationship, we come together and hear about what population they want us to model, right? So it might be that if you're a CPG company that's selling to all of the US, then maybe it's fairly straightforward. You want to model the gen pop of the US. But at the same time, if there is a vertical or if there's a market that they're trying to go into, imagine, they want to better understand, let's say, people in their 20s and 30s living in California. That's a much more specific population. So we hear about this population, and we go recruit these people, with consent, and with incentives, and we collect some of their data and create a model of these people. Then what our product allows you to do is query them. So it can take as input a filter that is a description of the population that you want to talk to, just like the one I just mentioned, and an environment. The environment can literally be survey questions, behavioral experiments, It can be A/B testing. Oftentimes, the core use cases are things like concept testing, to start with. But also, people sometimes want to do focus groups or one of the fun use cases that we also serve is even modeling things like earnings calls for public companies.Joon [00:18:21]: So these are the use cases that we often start with.Swyx [00:18:23]: Concept testing, is that an established term? I've never heard of concept testing.Concept Testing, Gallup, and PoliticsJoon [00:18:27]: Yeah. So it has to do with they have, let's say, different messaging, different products, different ideas.Swyx [00:18:32]: It's like a marketing exercise.Swyx [00:18:33]: Okay, got it. Got it. Politics?Joon [00:18:36]: We do, have a strategic partnership with Gallup, and of course, Gallup is deep into policy space and so forth. Right now, we have not worked deeply with politics, like that area just yet, however.Swyx [00:18:49]: I'm curious if there is demand or if they really would have different needs that somehow fundamentally don't mix with your existing, users or people.Joon [00:19:00]: I think there's certainly demand.Joon [00:19:02]: But we are very much mindful of how this technology gets adopted and the societal impact that we'll end up having with this technology. And I do see politics as an area where a company has to be particularly thoughtful about the way they operate and make impact. So this is where we also want to make sure that we form enough of guardrail and perspective on how to leverage this technology before we go on to serve markets like the politics.Swyx [00:19:29]: I'll give people an example. one of my favorite shows is The West Wing. I don't know if people have watched.Swyx [00:19:34]: One of the key storylines is, like, the president has, multiple sclerosis, but they haven't. they need to figure out how to disclose it. So they run a poll with a fake governor and ask people to respond on the poll,Counterfactuals, Polling, and When Simulation Is UsefulSwyx [00:19:47]: They try to make decisions based on the results of that poll on, like, how well they'll be received, like where, how should we play this?Swyx [00:19:54]: And I'm like, well, I think those counterfactual things, I would use a simulation for this if I could trust it.Joon [00:20:01]: For sure.Joon [00:20:02]: In that show, how'd it go?Swyx [00:20:04]: In that show, it was, like a foregone conclusion. They were like, “We know it's bad. We just don't know how bad.” And then the poll came back. It was like, “It's really bad.” And then they just did it anyway.Joon [00:20:14]: Part of it is to show, right? So you're, you're looking at the ideaSwyx [00:20:17]: Maximizing drama.Joon [00:20:18]: How bad could it be? Oh, it's horrible.Swyx [00:20:20]: And to some extent, I think that is part of the trick of the, or the challenge or with being a customer of yours, which is that if I know it's. if I roughly know and can intuitSwyx [00:20:35]: What the effect is going to be, do I need you? What sensitivity of it, of effect do I need in order to make a decision, right? So for example, if I, my approval rating is 50%Swyx [00:20:48]: And I, they have this negative piece, news item comes out, and it drops to 30.Swyx [00:20:52]: If it drops to 20, if it drops to 40, do I care? No. It, I know it drops. It's negative. So when do I care about simulations?Joon [00:21:01]: You do something that's clearly bad, that's not popular, and people don't like you, like, yeah, it's likeSwyx [00:21:05]: You don't need a simulation.Joon [00:21:07]: Yeah. Well, so there are a couple of things. one is, there are use cases where, like every day, developers, designers, policymakers, marketers, every single day, they create assets. They create new products. And turns out, it's many of the decisions in hindsight is obvious. Yes, of course this is bad, but we still run those studies because understanding the magnitude and understanding how acute something is quite difficult, even if, we feel like, of course, like this makes sense. this is the reason why we make so many mistakes. Like, every time somebody goes online and say something that has huge backlash, you look at that and like, “What an idiot.” However, it's tough. That's one. There's also another aspect here, which is, again, this is the reason why simulation is different from prediction. In simulation, in the ideal case scenario. So what simulation is trying to show is it's trying to show each step of the way or each step that we need to take to get to a certain outcome, right? So in the most advanced simulations, sometimes the next step that we're suggesting might be quite counterintuitive. The analogy that I sometimes give, and I ground it in a more realistic example, but, I, as I mentioned, I'm a huge fan of science fiction, and I don't know how, many of the audience members have read, like, things like the Foundation series by Asimov.Simulation as a Path, Not Just a PredictionSwyx [00:22:37]: Oh, yeah. We've mentioned psychohistory a number of times.Joon [00:22:39]: Okay, fantastic. So I might be, talking to the right crew. If you read Foundation series, literally the first act is there's a group of scientists who have found out that, “Oh, our galactic empire is going to collapse, and we're going to have 30,000 years of unrest.” And they run psychohistory, the simulator that tries to teach them, “Okay, how can we keep this unrest to a 1,000 years?” And they plan this out, and the first step of that plan is to get the scientists who say, “Okay, this is coming,” exiled into this random place in this, galax- galaxy.Swyx [00:23:18]: Terminus.Joon [00:23:19]: Exactly. And that's so counterintuitive. Like, what a strange move that you literally sent the group of scientists who was raising voice around this potential collapse of galactic empire into nowhere. How is that the right first move? Well, it turns out in this particular simulation, that was the move.Joon [00:23:40]: It's these things, right? And the reason why these reasoning is possible is because you're showing the step function or each step that results in a particular outcome. So really what simulation allows you to do in its highest form is you give it not a problem or question, like what would people answer to the survey? That's not what we do. What we tell it is, “Here is a goal that we have. In the context of foundation, we want to keep the unrest to a 1,000 years. What is the path that we need to take now to get to that particular future?” And that's what simulation allows you to do. Now, translating that into real market, imagine you're a automobile company and you're about to release a, EV, and you're trying to understand, well, how do we market EV, to make sure that our stock price goes up? But what if the answer comes down that, well, you can market your EV in XYZ way, but that might change people's perception around the cars that's not EV and make your overall sales to go down. Not very intuitive, especially all you're trying to optimize is EV salesss, and that's the only thing that you're tracking, then that might result in a completely wrong solution, or at least different solution than what you would have expected, whether it's right or wrong.Joon [00:24:57]: That's the power of simulation.Swyx [00:24:58]: For listeners, we covered a similar topic with Mikhail Parakhin from Shopify, where they are working on SimGym. I don't know if he ever talked to you about it. it's very similar.Joon [00:25:07]: ISwyx [00:25:07]: The goal is increased conversion, but then the journey is very unusual.Joon [00:25:12]: Journey is unusual.Swyx [00:25:12]: Yeah. The-- He's trying to look for interventions on a shopping trajectory, which is similar to what you're saying. Like, it's not about the attitudinal, is your word for it.Swyx [00:25:24]: It's about behavior.Joon [00:25:25]: It's about behavior.Swyx [00:25:25]: And that's exactly the difference, right? It's, like, not about the near-term direction about-- but it's more about, like, how do you affect multiple turns of interactions.Vibhu [00:25:35]: You had a good quote at the start about this as well. It's not about people wanting to know the outcome. It's about how they can change it, change the way to get there, something like that. But I wanna take it back to how do we know this is grounded? LikeGrounding and Evaluating Digital TwinsVibhu [00:25:47]: How do you run evals? How do you test that simulations come through? if I was to do the same thing that you described with, say, your favorite LLM, Opus, GPT-5.6, have some agent to map out these thingsVibhu [00:26:02]: How different are the answers we would get if I give it the same goal, the same objective, make a decent system? You're saying that you need to change the model weight. You have your own solution to this. But how far off are we, and how do you check if it's grounded? you have some interesting stuff on your site that points to how you run real evals, but if you could take us through that side. I think that's one of the big concerns that people have. They're like, “LLMs hallucinate.”Vibhu [00:26:27]: “You're just hallucinating layer after layer,” right?Joon [00:26:30]: The way we do this, and this is the paper that we worked on after the generative agents paper that really became the, at least for Simile and also the field of simulation and synthetic panels, really became the foundation. Yeah, this is the paper. the paper is called Generative Agent Simulations of 1000 People. Here's what we've done. For this paper, we brought 1,000 people that's representatively sampled from the US to a virtual lab. And what we have done was we spent two hours collecting fairly wide-ranging data. In this particular study, we focused a lot on this interview data, that was, whose script was taken from this project called American Voices Project. And then we would also pair that with a lot of behavior data and so forth, whatever we can collect within two hours. And then we would send these people away for a couple of weeks. And during that time, I would use this data to create their digital twins. And I would bring the humans, participants back after 2 weeks and have them complete a battery of surveys, experiments, behavior studies. So we have the list here, which included things like behavioral economics games. We would run literally, like, Big Five personality test, General Social Survey. We would also go ahead and run the randomized controlled trials that were published on PNAS. And we would have their digital twins predict how the source individuals would have acted in these studies and surveys. And this is where we could replicate people's behaviors and attitudes 85 percent as accurately as people would replicate their own. So that was the first really paper that gave this validated results that we can model individuals in an accurate way. And what we ended up finding now, of course, in AI space, so this paper came out at the end of 2024. AI space, a year and a half, 2 years, that's a lifetime.85% Accuracy and Why Frontier Models Miss Human BehaviorSwyx [00:28:24]: Yeah. Just, for listeners who are not seeing the YouTube, I just wanna say, like, the headline figure is 85 percent accuracy, like, which is a big improvement over all the otherSwyx [00:28:34]: Methods that you showed.Joon [00:28:36]: But the part that was particularly striking to us, especially as we improved this technology even further, was the generative AI models like ChatGPT, Claude that's coming out, it does give you the right foundation. However, what they do not consider is the true attitudinal and behavioral aspect of people, especially in the population that you care about. So what these models are really good at today is they're trying to become the super rational, objective machines, right? So you go get their data from places like Mercor, Scale. You talk to professional programmers, scientists to create model that's amazing at reasoning. That's what they do. Simile doesn't care about any of this. The models that we're talking about here, what we're trying to create are models that are as dumb as I am, right? So if I make some mistakes, the model has to make the same mistake.Swyx [00:29:34]: Oh, that's very hard.Joon [00:29:35]: That's very hard.Swyx [00:29:36]: You're solving Murphy's paradox.Joon [00:29:37]: That's exactly. And this is a completely different data and training objective. This is also where we see quite a bit of discrepancy in the performance in human behavior prediction between the frontier models, Simile's model, and the models being created in this space, where in some cases, the model performance of frontier models go all the way down to 20, 30 percent, especially if you go into that more niche population on topics that our customers would care about. On more gen pop, it might be around 50 to 60 percent. So it's not very robust. Like, you wouldn't want to make your decision off of these and these findings. If you can bring that up to 85 percent, that is ultimately what people end up getting very excited about.Swyx [00:30:20]: Yeah. Do we wanna keep going on the paper, routes?Joon [00:30:23]: Yeah, for sure. So the last one, was an interesting one. So this, paper was the follow-up paper that we had, to the 1000 agents paper, where the idea was now can we augment the models even further and post-train a model based on a lot of randomized controlled trials? So this was an interesting one. The data is always the most interesting part of modeling in many ways. The data that we got here was there's this, there's this platform called Open Science Framework. So some, the audience might be familiar with this. And there has been, especially in the social sciences over the past 5 years or so, there has been this concern around replicability of studies. And so it was a bit of a crisis, the scientists acknowledged, where we rerun the study and we don't see the same finding.Post-Training on RCTs and Replication StudiesVibhu [00:31:12]: Oof.Joon [00:31:12]: It's tough. And the reason why it's there-- that was often the case was there's this survival bias where the papers that get published often need to maintain what we call the value of less than 0.05 in the experiments that we ran. That suggests that only-- there's only 5% chance that the results that we saw is false positive. But the tricky part was all the papers that were not published, and there's still a 5% chance that whatever we publish is totally just randomly generated. Like, there's a 5% chance that, hey, this effect is not real, but it just happened to be real because of the sampling bias. So because of that, what scientists started to do was they started to register their studies. So before running an experiment, they would go to this platform and say, “Here is the data. Here is the population that we're collecting, and here's the hypotheses.” And they would just say, “Here is our hypothesis.” Like, “This is what we believe.” And you cannot retroactively change those hypotheses. This is what gives us more scientific statistical confidence that whatever effect that you ended up seeing is true. So that ended up creating this really interesting platform where there's one platform that has now contains tens of thousands of real-world experiments and hypotheses. And a lot of these are really high-quality, like, professionally designed behavior studies and random- randomized controlled trials. So we got the data and the studies from this platform and used that to make a point. And this particular, model is not, something that we're serving commercially because this was a part of the open science. But this particular data set, helped us make a point that by collecting a lot of these randomized controlled trials, that are really well-designed, we can make significant improvement in model's capability to predict human behaviors. So that's what this paper was about.Vibhu [00:33:10]: Is this stuff done on a individual level? Like, do I need to tune the model per individual, per company? Is there foundation model changes and then some slight post-training? Anything you can share there?Population-Level vs. Individual-Level ModelsJoon [00:33:21]: So this particular model was trained. the data we had at the level of individuals, but this particular model was trained. We experimented with both. And this is what we end up doing at Simile too. We always train 2, distinct model. One is what we call the population-level model. The other is what we call the individual-level model. And both take very similar input, which is the description of a subpopulation or individual and a stimuli. In this particular work, we've done the same. Here, the results that we are reporting are much more geared towards individuals because we do think that is a harder task in many ways, but that's what we have done.Vibhu [00:34:02]: You seen anything on the questions that humans can solve that models can't solve? So likeHuman Biases, Mundane Choices, and What Models MissVibhu [00:34:09]: Currently, it's, I live 5 minutes walk away from a car wash. It's a 10-minute drive. Should I walk or drive?Joon [00:34:16]: Huh.Vibhu [00:34:16]: The model will say, “Oh, walk to the car wash.” And, you don't have your car.Vibhu [00:34:20]: Is anything like this a problem in simulation? You would assume, like, very simple for human to think about, but if the model is saying you should walk to the car wash, anything here?Joon [00:34:32]: It's less, what can we solve, but I think it's more about what biases or mistakes do people make that models miss. Like, imagine that you are, like the. When I was still at Stanford, I lived in Palo Alto. So it's about, I would say, 40-minute walk from the campus. You ask the model, “Okay, let's go home. What can I, what can I do?” It would likely call an Uber or, give me, the bus time. But for the longest time, I really liked walking back. And the reason why I wanted to do that was not for efficiency. It really helped me think. And I like to walk for, half an hour or 40 minutes or so a day, where I just get to, just think about ideas, research, just get lost in my thoughts. That's very human activity. Unless the model has seen that and understands the importance of that activity, it would miss these kinds of features. So that I think, is fundamentally what we're trying to model. Like, what is fundamentally human might not be the most efficient thing to do, might not be the right thing to do, but things that make us who we are.Swyx [00:35:43]: I'm curious if, there are some data sets that you really want that would materially help you. One version of this may be interesting, which is more valuable to you to acquire as a data set, all of LinkedIn, all of Twitter, all of Facebook?What Data Matters: Social Media, Transactions, and FacebookJoon [00:35:57]: It's a little bit hard to rank, in part because, there's, there's this product saying where no feedback is wrong because it teaches you something about your users. Doesn't matter what feedback.Joon [00:36:11]: I think it's a little bit like that.Swyx [00:36:12]: So just whatever is bigger.Vibhu [00:36:13]: What about a different domain? Say it was. What about all of Amazon data?Joon [00:36:17]: Oh, yeah.Vibhu [00:36:18]: Shopping data, right?Joon [00:36:18]: Shopping data. So Amazon data is interesting in that it's very much behavioral, although, like, what people do on social media, you could squint and say that is also behavioral. But the transaction data is always interesting. It is also most commonly available, however.Joon [00:36:33]: If we were to look at purely social media, like if you really, if I were, if I had to really pick, Facebook likely is interesting because I do think it is most a default version of people. Because you go to LinkedIn, it's very much professional environment. So people put up their, they have their guards up, right? And that still is interesting because that is true human attitude and behavior, but it is not your base state. you go to Twitter- Twitter, people have their own crazy personas, or depending on who you are. Like, my Twitter profile and, persona is very much, initially was I was very much an academic. “Hey, I'm here to share my studies.” Now, I share, things that's related to Simile. But Facebook is one of those more private space where people just connect with their friends. In that way, I do think it shows you a little bit more about who that person is. So if I had to pick, I'd likely pick, Facebook.Swyx [00:37:30]: Yeah. And you're interested in, like, the whole person and their background and philosophy. I, is it too clinical or too machine learning-oriented to just say this is just ways to inject variance and biases? The broad question, is, like, is this any better than a randomized, like, combinatorial explosion version? So we have a link to the TencentBillion Personas, Synthetic Demographics, and Bespoke DataSwyx [00:37:54]: Billion persona paper, where they did not do any of the groundwork that you are doing.Swyx [00:37:59]: They just did like a cross matrix of here's all the professions in the world, here's all the people, possible backgrounds in the world, do a dot product across all of them, and that's it. That's your prompt for a billion people.Swyx [00:38:12]: This will do something. I don't know if it'll do what you do, but it gets you some way, some percent of the way there.Joon [00:38:18]: So this was an interesting paper. Like, what I admired about this paper when it came out was the scale. And you do gradually want to be able to simulate really large societies and interactions. So the scale is definitely admirable. it is relying heavily on the known statistics that went into training the model. So to the extent that you believe that statistics is correct, this is not a bad way to go about this. But the thesis here, and this is something that we also have seen in the market, like if this works, then we have solved simulation.Joon [00:38:54]: It,Swyx [00:38:55]: Because I survey, like, okay, 5% of the US population is in construction.Swyx [00:39:01]: The other 5% is in medicine, whatever, right? And then you just keep going down the list, and then you do the other side. 5% has, like, the big 5 personalitySwyx [00:39:08]: Of, like, neurotic or whatever. That's it.Joon [00:39:11]: That's it. So if you believe that the underlying data set and the platform that we're leveraging has all the right statistics, then this will have solved it. you're at that point merely retrieving the knowledge that is already embedded in the model, in the model parameters. That's not, unfortunately, what we see, where there is such detailed and also niche knowledge about people that if you just take one example, it might feel very mundane, but it's quite rich when you put together, that you do need to do a lot of bespoke data collection to better understand people. And this is also, I think what makes this particular, job fun, which you want to deeply understand people, and the process of deeply understanding them requires a lot of attention to the details. And you do need to pay attention to and pay respect to the daily lives that people lead.Scaling Simulation: From Thousands to SocietiesVibhu [00:40:04]: I wanna talk about scaling simulation.Vibhu [00:40:07]: So what can't we simulate, what can we simulate, and how does scaling affect this? So how big are the models? What if we go from, 8B, like, couple 100 billionVibhu [00:40:18]: Like billion000 parameters, billion000? Do we get scaling? Any interesting emergence? Like, at a certain scale, at a certain amount of training, you uncover anything unusual and any learnings from that?Joon [00:40:31]: What we are seeing is at Simile, so we do post-train our own model. The thing that we're seeing is the early glimpse of scaling law in simulations. The more data about humans and more compute you ingest, you start to get predictive and predictable gains of the model performance in simulating it, simulating people.Vibhu [00:40:51]: Ooh. We need a scaling law curve.Joon [00:40:52]: It's scaling law. Whenever you find it's a beautiful thing. And we're starting to see the glimpse of it, which is quite exciting. But if you talk about the ambition of simulation as a whole, it's not merely about building a model. It's about building a model, then creating the agents that become the individuals in a much larger ecosystem. So they're creating this multi-agent simulation. Down the line, you want these multi-agent simulation to also live in a very rich environment, right? What we are really trying to get to at that point is, hey, can we create. All right, let's do a time machine game again, and 5 years, 10 years into the future, can we create a simulation of 8 billion people living on Earth? I think that's quite interesting. And that really is the vision. And once you get to that state, the questions that you can help answer for the society also start to change from my perspective. The answers are fundamentally about emergence of the emergent behavior of society and large groups of people.Joon [00:41:53]: So the questions that I get excited by, and maybe this is a stodgy- a bit. I have my, academic side of me.Joon [00:42:01]: And for me, it's questions like, can we help solve climate change? If you look at climate change as a problem space, this is what we, like social scientists would often call it the wicked problems, problem where you have many actors with competing incentives for trying to make a very complex decision and coordinating that coordination decision. Very difficult to really solve in real life, which is also the reason why we couldn't solve it. Can simulation help us solve that? Another one is, can we understand the signals for collapsing democracy, or can we understand or can we uncover the origin story of the monetary system? These are societal questions that we never really had a good way of answering. If we can create simulations of our society, you have to believe that these are the problems that we can solve. So that's really the ambition of this field. And, I also think, yes, I think there's a Nobel Prize to be won there, which wouldn't be surprising. And I think there's some amazing societal impact that we can have to help people make better decisions.Climate Change, Democracy, and Societal SimulationSwyx [00:43:04]: Nobel Prize in economics?Joon [00:43:06]: In economics.Swyx [00:43:06]: Oh, I see. I see. Rooting for you to write that paper.Joon [00:43:10]: One of these days. But, one of the scholars that I was deeply inspired by, When I was coming into the space of simulation, is this scholar, named Thomas Schelling.Schelling, Agent-Based Models, and the Nobel PrizeSwyx [00:43:23]: Schelling point?Joon [00:43:24]: So the canonical example of the work that he's done was he was one of the creators of agent-based modeling. So this was, like, in the 1970s and 80s. It's very early days, but this was truly one of the first exemplars of simulations. And one of the canonical model from that time, and of course many of these simulations are trying to tackle the societal problems that's most relevant for their era, it was called the model of segregation. So racial segregation was a big topic, that, we cared about. And what they've done was they created this grid world where they had red dots and blue dots. And these dots were, back in the day, like, they were the agents, and they had a simple rule that governed their behavior. If certain percentage of your neighbors are of different color and if that goes above certain threshold, then you move to a new location at random.Joon [00:44:21]: One of the striking finding of this paper or this agent-based model was for the longest time, people thought the segregation within society was caused by explicit and overt racism.Joon [00:44:34]: But if you look at this model, people's preference towards living with people of the same color, that preference can be very minute.Joon [00:44:42]: But the very small difference causes the society to segregate completely over time. This was very counterintuitive for a lot of people. And this particular work ended up informing housing policies. Mixed income housing, got really inspired by this work. And Thomas Schelling ends up winning the Nobel Prize for having laid the groundwork for very early versions of simulations. The opportunity that I do see here in the more scientific terms, is agent-based models for the longest, had impact in the 1980s, 90s, to some extent, early 2000s, but it has now gotten forgotten by the community a little bit. Because as you can imagine, red dots and blue dots is not really a rich description of people.Joon [00:45:31]: But with the emergence of things like generative AI and, in particular, generative agents, we do have an opportunity to create these agent-based models that are high fidelity enough to help us make really complex decisions. And that's the opportunity that I see. If that truly works, then yes, that is the work that will result in a Nobel Prize.Swyx [00:45:53]: Yeah. For what it's worth, and I grew up in Singapore. 80% of Singapore is in public housing, and public housing has, enforced racial quotas for exactly that reason, which is very interesting. okay, so we talk about scaling, we talk about all these, the agent possible applications.Cost, Reuse, and the Economics of SimulationSwyx [00:46:13]: I'm scared about the cost. if you even-- let's just keep it to the US, about 8 billion people.Swyx [00:46:21]: But, how much does it cost to model so many hundreds of millions of people?Joon [00:46:26]: Oftentimes today, we don't start at that scale, this stage of the, of industry and simulation as technology. But we can get our users extremely rich and meaningful insights even by modeling thousands, tens of thousands of people. And today what we do is every week we are collecting data on the scale of tens of thousands people's data, and we have panel partnerships that gets us to tens of millions of people globally. So that's what we do today.Swyx [00:46:55]: And just as a side note once you've collected one person for one studySwyx [00:46:59]: Can you reuse that same person for all the subsequent studies?Joon [00:47:03]: That's exactly right.Swyx [00:47:03]: Okay.Joon [00:47:04]: The beauty of this model and these agents is the fact that they are domain-agnostic.Joon [00:47:08]: That what you're really trying to understand is what is the fundamental nature of these people? What's their social physics? And there are a lot of, a lot of, people that does change over time. Like, even, like, even things like, how many times have you gone have you been to, like, CVS the past week? that will change. But there's so many traits about people that are also known to never change. Like, your risk tolerance doesn't really change over time. It's very consistent. So it's these things that we're trying to learn. But the scale we are operating is right now hundreds or, tens of thousands to hundreds of thousands. And in many of the core use cases that we are deployed in, and this is more than enough population, to cover those. Really, at that point, what you care about is less the number of people, but more do you have the right subpopulation of interest covered? And this is also the reason why people want a larger sample. It's not because they want, stronger statistical guarantees. It's more that can they filter down to any population of their interest. However, you can also imagine in 10 years, if we truly believe that the compute is going to scale, that we'll have much more availability for compute, and our ambition for simulation is also going to scale accordingly, there's definitely a reason for us to create an entire data center worth of simulations.Joon [00:48:35]: Or in my hunch here is I do think in the next some number of years, we will start creating simulations that will cost as much as training a foundation model. But perhaps it's going to be so valuable to the society that it would be a no-brainer. Right now, even today, like, we are training bunch of new foundation model just so we can say we trained one and we spent tens of millions. But if we can create a simulation at the level of society that would solve climate change, I would run that today. I would raise the money right now just to run that.Multi-Agent Simulation and Social InfluenceSwyx [00:49:10]: Amazing. the follow-up question is, does it also compound if you let the simulations talk to each other?Swyx [00:49:18]: Or do they already do that today? They don't, right, as far as I understand?Joon [00:49:22]: It depends on what simulation you're trying to run.Joon [00:49:24]: In the multi-agent simulation setup, the agents do talk to each other.Swyx [00:49:28]: Right, which is exactly Smallville, right?Joon [00:49:29]: That's right.Swyx [00:49:30]: But a lot of times, for example, in commerce, you're just by yourself, so there's no point talking. which is way cheaper.Vibhu [00:49:37]: But they use all these levels, right? Like, you decide what you will buy based on what other people around you buy and talk about, right?Swyx [00:49:43]: It depends.Vibhu [00:49:44]: It depends.Swyx [00:49:45]: Again, I'm, I'm coming at this from a cost point of view. I'm like, “Oh my God.” LikeVibhu [00:49:48]: I thinkSwyx [00:49:49]: If there is, like, some combinatorial thing of, like, thousands of people talking to thousands of people, then that one million X's might cost.Vibhu [00:49:56]: I have a very different view as the cost point aside. Like, running these studies in reality is a lot more expensive, right? Running any study like this is you gotta have people do it, you gotta sign people up. It's very expensive and sometimes, like, not feasible to run the study.Vibhu [00:50:14]: But the outcome or the decisions you make are very expensive on them, right? So spend X million on something that, the overall process costs 100 million might as well, right? There's, there's a lot of value to be had there. It's a small cost, but I'm excited on the cost side.Joon [00:50:33]: To some extent, and when you deploy technology, you often want to deploy in a way where you can replace existing budget or you can make things more efficient, and that is the best way to deploy. However, the way you capture the long-term value of the technology is making the argument that, no, it's the upside, that by making this better decision using simulation, you have saved yourself or made yourself hundreds of millions or even billions of dollars, and that's a case to be made.Vibhu [00:51:06]: Random tangent question. So if you're doing a lot of inference, a lot of model multi-agent stuff, are you at the point where it makes sense to, train a model that' very sparse? You're expecting to do multi-million dollar runs. Are you thinking about this in model architecture standpoint or inference efficiency, or, you're still at the research phase of it works, we're not super there yet?Joon [00:51:34]: Efficiency, we do think quite a bit about. this is technology that is deployed now in some of the largest enterprise companies in the world, and we do process significant number of queries, that are trying to, simulate the populations in the world. So efficiency is a consistent thing. we don't want to over-optimize too early, so I wouldn't say, like, this is the higher bid Right now, but this is definitely something that we think pretty carefully about.Swyx [00:52:05]: Yeah. Are there other case studies? So we, you talked about CVS, talked about Gallup, Deloitte, Wealthfront.Efficiency, Enterprise Use, and Real-World Case StudiesJoon [00:52:12]: Wealthfront is an interesting one, because one of the things they were trying to do, they were one of the first customers that wanted to do product testing that goes beyond just asking people what they think about, let's say, behavior experiments and so forth. So there, really what we had to do was reason about multimodal input, so images, but also you can also imagine, like, these agents traversing through Figma mockups or websites. So some of the things that our agents can also do is it can be given a domain, like, or, like, a website URL and go use it for a while. It's these things. And Wealthfront was one of the first, customers, that was very excited about this possibility.Vibhu [00:52:53]: What have people been asking? Like, is there any demand that we have not covered? Like, UI testing, right?Vibhu [00:52:59]: I wanna try a new. I wanna ship a new feature, test the UI, simulate how people will do it. Any interesting things that you're seeing demand for?Product Testing, Websites, and Synthetic PanelsJoon [00:53:08]: Today, a lot of the demand does come from like, the places where people have historically used human panels, we can now replace with agents, and these synthetic populations. And this is not replacing human panel. in many ways, the simulation that Simile is building is grounded. So the way that I think about this is we are trying to represent humanity at scale. And in that way, the use cases are what we would expect, but it's the scale of deployment that surprises me.Joon [00:53:44]: Turns out there are so many decisions that people make every day in these organizations, groups, and we want to be able to say, “We listen to people. We have consulted our users.” But in reality, that is rarely the case because getting to people and asking them many questions, it's difficult. It's both costly, time-consuming, but most importantly, people are just not available. If I had to answer 1000 survey questions for this one particular, vendor, even if I wanted to do that, like, I would never do it. And that's very much the case. What simulation can do is ensure that the voices of people are always represented in rooms where the decisions for them is made, right? So all the stakeholders of this particular product launch, ideally they're consulted. That's what this technology really is trying to enable.Market Size, TAM, and Human Decision-MakingSwyx [00:54:39]: In my mind, that means it skews towards more consumer focus, right? Like, anything with a wide enough customer base where you do benefit from the diversity that you represent. What are some rough statistics, just for people who are not familiar with this market in general, what's the market size that. I'm sure you have some, like, rough numbers. market size is, like, a vague questionSwyx [00:55:01]: But, like, how much do people spend?Joon [00:55:03]: So market research is a $100 billion industry.Joon [00:55:06]: But the thing about simulation is not a tool for market research. Simulation is a tool for human decision-making. So the question around what is a TAM here is quite tricky, right? Because it's easy to say, “Well, market research TAM is roughly 100 million or 100 billion.” so is it a TAM? And not really, right? Because in many ways, you're trying to inform all human decision-making. You're trying to inform every decision that are made about humans for humans. What is a TAM for that? It's really unclear. And I'll be honest. Like, I have a scientific background, I have a research background, so I didn't come into the field calculating, oh, what is the TAM for human decision-making? But I just had to assume, well, if we can inform every decision that is made about human for human, that has to be big.Swyx [00:55:58]: Some- something valuable.Joon [00:55:59]: Exactly.Swyx [00:55:59]: To some extent, you are a unicorn founder now, and you have to care as a CEO. But, like, I do think, like, yeah, when you go into these boardrooms with people that you're quoting millions of dollars of contracts for, like, you have to say, “Well, here's what you spend on humans-”Swyx [00:56:15]: “. And here's what we save you, and it's 85% similar.”Joon [00:56:19]: And certainly, the value case, is something that we care deeply about. Like, what is the value that we provide to the users and the decision-makers? But this is also where, like, as a founder, I think valuation only tells one very superficial aspect of the story, and I try not to think too much about valuation, in general, because that's not what also motivates a team or certainly doesn't. I'm, I-- Again, the interesting thing about researchers is we are happy living in academia, getting paid next to. we get paid okay. we don't get paid that much, as a researcher here in academia, but it's the impact and it's the, it's the value that we can provide to the individuals and the society that really drives us. And in that way, ultimately what drives us is the impact. Does the simulation we provide have a real impact in people's decision-making in ways that progresses our society forward? If the answer is yes, then yes. that has to be great business, and we see that in numbers, and we do care deeply about that upside story, but that's the heart of it.Where Simulation Goes NextVibhu [00:57:27]: Do you have any timeline predictions? So we talked about scaling laws of simulations.Vibhu [00:57:33]: You brought up, okay, maybe one day we can simulate how to solve climate change.Vibhu [00:57:38]: Where are we now?Vibhu [00:57:40]: If that's not the end state, what is an end state, and what does progress look like?Joon [00:57:45]: So what I sometimes tell people is simulation as industry, it feels a lot like where GPT-3.5, GPT-4 was, for the AGI saga, which is we have now technology that is powerful enough to do real damage on the verticals that we are tackling. At the same time, there's a lot of progress that is yet to come. And that's, I think, where this is. So the way I see it, I do think there will continue to be breakthroughs both in data, in algorithms, and there will be much more aggressive scaling that will also happen over the next few years. But I think that's roughly where we are.Swyx [00:58:27]: I think that was about the ro
Episode NotesEpisode SummaryIntroduction to Open Science – Asier Moneva introduces open science, emphasizing transparency and replicability as essential to modern research.Importance of Transparency – He explains how transparency builds trust, enabling other researchers to assess rigor and replicate findings accurately.Preregistration and Registered Reports – Asier discusses these practices, which require researchers to specify methodologies and hypotheses before data collection to reduce bias.Challenges in Adoption – He notes that implementing open science practices can be challenging due to academic pressures and resource limitations.The “Publish or Perish” Culture – We highlight how the pressure to publish quickly can conflict with the time-intensive requirements of open science.Academic Incentives and Misaligned Goals – We critique the academic reward system that often favors quantity over quality, which can detract from scientific rigor.Advantages for Public Accessibility – Open science also enhances public accessibility, making research available beyond academia and helping inform public policy.Ethical Considerations in Research – Asier emphasizes that open science fosters ethical research practices by reducing questionable practices like p-hacking and selective reporting.Benefits of Open Science for Collaboration – The approach encourages collaboration across disciplines and institutions, providing a more comprehensive understanding of complex issues.Real-World Example of Retraction – He mentions a case where a research paper was retracted due to lack of transparency, illustrating the importance of open science practices.Role of Preprints in Open Science – Asier advocates for preprints as a way to share research and receive feedback before formal publication.Challenges with Platform Fragmentation – He observes that the proliferation of research-sharing platforms can hinder accessibility if findings are scattered across multiple sources.Future of Registered Reports – Asier sees registered reports as a future standard, as they align research design with ethical and rigorous science.Open Science as a Solution to Publication Bias – Open science practices help address publication bias by promoting the dissemination of all research findings, regardless of outcomes.Closing Thoughts on Transparency – Open science is about ensuring reproducibility and holding science accountable, aiming to make research as transparent and accessible as possible.About Our Guest:Asier Monevahttps://asiermoneva.comhttps://nscr.nl/en/medewerker/asier-moneva/https://www.thuas.com/research/research-groups/team-cybercrime-cybersecurityhttps://github.com/amonevahttps://osf.io/7ce24/Resources and References Mentioned in This Episode:The Open Science Framework (OSF)The OSF is an open-source platform supporting transparent and reproducible research across disciplines.The Open Science Framework:https://osf.io/Paper Introducing Registered ReportsThis foundational paper outlines the concept of registered reports, a publishing model aimed at reducing bias and enhancing research rigor.Paper introducing "registered reports":https://psycnet.apa.org/fulltext/2014-20922-001.htmlRetraction Case StudyA recent retraction of a notable article on the replicability of social-behavioral research findings offers insights into challenges within open science practices.RETRACTED ARTICLE: High replicability of newly discovered social-behavioural findings is achievable:https://www.nature.com/articles/s41562-023-01749-9Retraction Note: High replicability of newly discovered social-behavioural findings is achievable:https://www.nature.com/articles/s41562-024-01997-3Podcast episode discussing the retraction in depth:https://open.spotify.com/episode/3rygrbUNocfCEEGd1Byn0V?si=vJDuzQT3S7yJqDEUMycF1w&t=178Other:This episode was recorded in a hotel lobby corner with music playing in the background. If the audio sounds a little unusual at times it is because of the noise removal being used to remove that noise being combined with other ‘sound enhancement' features. I had to go back in and play around with the audio directly before I was even a little happy. The tools work well but they are a little unpredictable. I am increasingly wary of ‘it just works' audio editing tools. I would have left it in, but the bots chasing copyright infringement are ravenous and indiscriminate.
In this Papers Podcast, Dr. Kenny Chiudiscusses his JCPP Advances paper ‘Social anxiety symptoms and their relationship with suicidal ideation and depressive symptoms in adolescents: A prospective study' (https://doi.org/10.1002/jcv2.12249). Kenny is the lead author of the paper. There is an overview of the paper, methodology, key findings, and implications for practice. Discussion points include: Insight into the dataset used, which originated from the Wellcome Trust NSPN (Neuroscience in Psychiatry Network) study. The questionnaire measures used for social anxiety symptoms, generalised anxiety symptoms, depressive symptoms, and suicidal ideation. How the researchers dealt with missing data – a common feature of longitudinal cohort studies due to various reasons – and how they tried to account for this to test their hypothesis. The researcher's experience of pre-registering the analysis on the Open Science Framework. Insight into the analytic models used to analyse the data. Implications of the findings for clinicians and other researchers. In this series, we speak to authors of papers published in one of ACAMH's three journals. These are The Journal of Child Psychology and Psychiatry (JCPP); The Child and Adolescent Mental Health (CAMH) journal; and JCPP Advances. #ListenLearnLike
This episode discusses the principles, practices, and technologies associated with open science and underscores the critical role that various stakeholders, including researchers, funders, publishers, and institutions, play in advancing it. Our guest today is Brian Nosek, the co-founder and Executive Director of the Center for Open Science and a professor at the University of Virginia, who focuses on research credibility, implicit bias, and aligning practices with values. Brian also co-developed the Implicit Association Test and co-founded Project Implicit and the Society for the Improvement of Psychological Science. Additional resources: Center for Open Science: https://www.cos.io/ The Open Science Framework: https://www.cos.io/products/osf FORRT (Framework for Open and Reproducible Research Training): https://forrt.org/ The Turing Way: https://book.the-turing-way.org/ CITI Program's “Preparing for Success in Scholarly Publishing” course: https://about.citiprogram.org/course/preparing-for-success-in-scholarly-publishing/ CITI Program's “Protocol Development and Execution: Beyond a Concept” course: https://about.citiprogram.org/course/protocol-development-execution-beyond-a-concept/ CITI Program's “Technology Transfer” course: https://about.citiprogram.org/course/technology-transfer/
Open Science ist ja schön und gut, aber wie transportiere ich das ganze neue Wissen? Diese Frage stellen sich gerade viele Dozierende und Personen aus der Open Science. Die Gestaltung der Open Science Lehre ist ein maßgeblicher Baustein auf dem Weg die Prinzipien allgemeingültig und zum Standard zu machen. In dieser Folge geht es darum, wie man Open Science am besten lehren und lernen kann. Gemeinsam mit zwei Gästinnen bespricht Luise Tipps, wo man etwas über OS lernen kann und geben einen Ausblick in die Zukunft. Stay positive! Musik: Schlaraffel Schnitt und Post-Production: Helena Mehler und Luise Hönig Moderation und Production: Kai Krautter und Luise Hönig Kooperation: Open Science AG (PsyFako) Gästinnen: Dr. Johanna GerekeDr. Anne-Sophie Waag Quellen: FORRT Lexikon: https://forrt.org/ Wikimedia: https://www.wikimedia.de/ Blog: http://www.hypothesis.org/ Open Science Framework: https://osf.io/ Replication Wiki: https://replication.uni-goettingen.de/wiki/index.php/Main_Page Bitss: https://www.bitss.org/
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Eradicating rodenticides from U.S. pest management is less practical than we thought, published by Holly Elmore on March 24, 2023 on The Effective Altruism Forum. The link goes to the Open Science Framework preprint of the full report. Executive Summary Rodenticide poisons are cruel and reducing their use would likely represent an improvement in wild animal welfare. This report explores the reasons why rodenticides are used, under what circumstances they could be replaced, and whether they are replaceable with currently available alternatives. As summarized in the table below, agricultural use of rodenticides is well-protected by state and federal laws and that seems unlikely to change, but the use of rodenticides in food processing and conservation would likely be reduced if there were an adequate alternative such as solid form rodent birth control. Continued innovation of reactive tools to eliminate rodent infestations should reduce the use cases where rodenticides are the most cost-effective option for residential customers or public health officials, but will not eliminate their availability to handle major infestations. This research is a project of Rethink Priorities. It was written by Holly Elmore. If you're interested in RP's work, you can learn more by visiting our research database. For regular updates, please consider subscribing to our newsletter. Thanks for listening. To help us out with The Nonlinear Library or to learn more, please visit nonlinear.org.
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: How meat-free meal selection varies with menu options: an exploration, published by Sagar K Shah on February 14, 2023 on The Effective Altruism Forum. Summary Increasing consumption of meat-free meals can help reduce demand for factory farmed animal products and anthropogenic greenhouse gas emissions. But relatively little research has been done on how meat-free meal selection is influenced by menu options, such as the availability of meat-analogue options or different types of meat. We conducted a preregistered reanalysis of data from a series of hypothetical discrete choice experiments from Brachem et al. (2019). We explored how meat-free meal selection by 1348 respondents (mostly German students) varied across 26 different menus, depending on the number of meat-free options and whether any options contained fish/poultry meat or meat-analogues. Menus consisted of five options (of which, two or three were meat-free) and were composed using images and descriptions of actual dishes available at restaurants at the University of Göttingen. While our work was motivated by causal hypotheses, our reanalysis was limited to detecting correlations and not causal effects. Specific limitations include: Examining hypotheses that the original study was not designed to evaluate. De facto observational design, despite blinded randomization in the original study. Possible non-random correlations between the presence of poultry/fish or meat-analogue menu options and the appealingness of other dishes. Analysis of self-reported, hypothetical meal preferences, rather than actual behavior. Meat-analogues in menus not reflecting prominent products attracting significant financial investment. Notwithstanding, our reanalysis found meat-free meal selection odds were: higher among menus with an extra meat-free option (odds ratio of 2.3, 90% CI [1.8 to 3.0]). lower among menus featuring poultry or fish options (odds ratio of 0.7, 90% CI [0.6 to 0.9]). not significantly associated with the presence of meat-analogues on a menu (odds ratio of 1.2 (90% CI [0.9 to 1.6])) in our preregistered meat-analogue definition. Estimates varied across analogue definitions, but were never significantly different from 1. Despite the many limitations, these findings might slightly update our beliefs to the extent we believe correlations would be expected if causation were occurring. The poultry/fish option correlation highlights the potential for welfare losses from substitution towards small-bodied animals from menu changes as well as shifts in consumer preferences. Given the study didn't feature very prominent meat analogues, the absence of a correlation in this reanalysis cannot credibly be used to refute a belief that high-quality analogues play an important role in reducing meat consumption. But when coupled with the strong correlation on an additional meat-free option, we think the reanalysis highlights the need for further research on the most effective ways to encourage selection of meat-free meals. It remains an open question whether, at the margin, it would be more cost-effective to advocate for more menu options featuring meat-analogues specifically, or for more meat-free options of any kind. You can read the full post on the Rethink Priorities website, and also see the pre-print and code via the Open Science Framework.. Thanks for listening. To help us out with The Nonlinear Library or to learn more, please visit nonlinear.org.
Dan and James are joined by Brian Nosek (Co-founder and Executive Director of the Center for Open Science) to discuss the recent White House Office of Science Technology & Policy memo ensuring free, immediate, and equitable access to federally funded research. They also cover the implications of this memo for scientific publishing, as well as the mechanics of culture change in science. Open Science Framework hits half a million users (https://www.cos.io/blog/celebrating-a-global-open-science-community) The White house memo (https://www.whitehouse.gov/wp-content/uploads/2022/08/08-2022-OSTP-Public-Access-Memo.pdf) Brian on Twitter (https://twitter.com/BrianNosek) Other links Everything Hertz on social media - Dan on twitter (https://www.twitter.com/dsquintana) - James on twitter (https://www.twitter.com/jamesheathers) - Everything Hertz on twitter (https://www.twitter.com/hertzpodcast) - Everything Hertz on Facebook (https://www.facebook.com/everythinghertzpodcast/) Support us on Patreon (https://www.patreon.com/hertzpodcast) and get bonus stuff! $1 per month: A 20% discount on Everything Hertz merchandise, access to the occasional bonus episode, and the the warm feeling you're supporting the show $5 per month or more: All the stuff you get in the one dollar tier PLUS a bonus episode every month Citation Quintana, D.S., Heathers, J.A.J. (Hosts). (2022, August 31) "161: The memo (with Brian Nosek)", Everything Hertz [Audio podcast], DOI: 10.17605/OSF.IO/A7D86 Special Guest: Brian Nosek.
Only 1% of first year PhD students become research professors. This creates a hypercompetitive environment where scientists will do whatever it takes to get funding, even if that means tweaking statistical analyses to make their findings seem more significant. But decentralized science is working to address this egregious misalignment of incentives and reward science done the right way. Patrick Joyce is the Cofounder and COO of ResearchHub, a Reddit-style forum that allows anyone to share, discuss and curate scientific papers—and earn ERC-20 tokens for doing so. On this episode of Boost VC, Patrick joins us to discuss the misalignment of incentives in science and describe the problem with using bibliometrics to determine who receives funding. Patrick explains why decentralizing science is so important, exploring how DeSci will increase the adoption of open science practices and accelerate innovation in the space. Listen in to understand how DeSci can pressure large academic journals to pay content creators for their work and learn how ResearchHub is leveraging Web3 to make capital available to scientists. Topics Covered How Patrick defines sciencePursuit of knowledgeVerifiable results What's behind the replication crisis in scienceHypercompetitive job marketQuality judged by bibliometrics Why decentralizing science is importantTweak analyses to maximize perceived impactNot honest, transparent or reproducibleOnly most cited scientists receive funding Patrick's take on the first breakthrough for DeSciIncrease adoption of open science practicesNo longer career risk to share work in open What drives Patrick's conviction around DeSciSuccess not based on work ethic, intelligenceLittle innovation in drugs for mental health How ResearchHub rewards quality scienceAllows anyone to share and discuss papersEarn ERC-20 tokens for participation How Patrick builds trust in the scientific communityProvide funding and publishing outletsHelp meet goals of open science What success looks like for decentralized sciencePaywall journals provide open access optionsPay science content creators for work Why the science incentive structure hasn't changedPrevious attempts not fundablePirate organizations like Sci-Hub not legal Patrick's take on what makes publishers the enemyResponsible to own incentive to make moneyCharge scientists to publish, subscription fees Patrick's uniting purpose for the DeSci communityWeb3 offers ROI to invest in scienceMakes capital available to scientists How Patrick defines success for ResearchHubTake power away from bibliometricsCreate rockstar scientists who define culture Connect with Patrick Joyce ResearchHub https://www.researchhub.com/ ResearchHub on Twitter https://twitter.com/researchhub ResearchHub on Discord https://discord.com/invite/ZcCYgcnUp5ResearchHub on Medium https://medium.com/researchhubResearchHub on Reddit https://www.reddit.com/r/ResearchHub/ResearchHub on GitHub https://github.com/ResearchHub Resources Brian Armstrong & Patrick Joyce on The Sheekey Science Show https://www.youtube.com/watch?v=1FmhHLPrUIUClarivate Analytics https://clarivate.com/OpenAlex https://openalex.org/r/medicalstudent https://www.reddit.com/r/medicalstudent/r/ImmunoPsychiatry https://www.reddit.com/r/ImmunoPsychiatry/eLife https://elifesciences.org/Open Science Framework https://www.cos.io/Nature Neuroscience https://www.nature.com/neuro/Chris Hill on Boost VC DeSci EP02 https://www.boost.vc/podcastAuthorea https://www.authorea.com/F1000Research https://f1000research.com/Sci-Hub https://sci-hub.se/Stack Overflow https://stackoverflow.com/Hindawi https://www.hindawi.com/MetaMask https://metamask.io/Fast Grants https://fastgrants.org/Kaggle https://www.kaggle.com/ Connect with Boost VC Boost VC Website https://www.boost.vc/Boost VC on Facebook https://www.facebook.com/boostvc/Boost VC on Twitter https://twitter.com/BoostVCBoost VC on Instagram https://www.instagram.com/boost_vc/
We are delighted to be conversing with Prof Dani Prieto-Alhambra, Professor of Pharmaco and Device Epidemology, NDORMS, University of Oxford, UK, and part-time Professor of Real World Evidence, Erasmus MC, Rotterdam, and Research Coordinator for EHDEN. In this penultimate episode of season 2, Dani returns to the podcast (he also was in episode #2 of season 1) to discuss his perspective on evidence generation and conducting research with real world data today and a vision for tomorrow. We start with exploring the differences (or not) between pharmaco and device epidemiology from Dani's experience, then go on to re-evaluating the research response to the COVID-19 pandemic, including inherent difficulties, and in particular evolving research in areas such as 'Long COVID' and sub-acute COVID that his group is leading. Clearly, there are implications for not only COVID-19, for e.g., vaccine use and safety, but also more widely for medical and health research, and in catching up with much that was challenged due to the pandemic. We cover the upcoming research priorities, inclusive of plans for study-a-thons and evidence-a-thons in EHDEN, based on a call for study proposals within the programme. Following this Dani outlines how he sees research methods and collaborations changing now, for the better, and hopefully permanently, in using federated networks, distributed network analysis across geographies, supported by new platforms and technologies. He goes on to explain the use of study-a-thons and evidence-a-thons in EHDEN and OHDSI, and their emerging role in rapid analysis work to meet the challenge of responding to diverse research needs. In the last third of this episode we discuss the paradigm shift we are seeing in terms of the creative disruption of the open science agenda and OHDSI research framework in EHDEN, in Europe, but also globally, inclusive of the Global South. Specific challenges such as reproducibility and transparency are also coming to the fore with our new methods in being able to be truly open, and with a need to collaborate. We take Dani back to his first exposure to OHDSI, the positive impact on his own career, but also the need to train and support a new generation of researchers where this paradigm shift today will be routine for them tomorrow. Moreover, and with COVID-19 in mind, we have changed science for the better, but we need to reinvigorate faith from certain communities in science globally. The views expressed by the participants are personal and not necessarily reflective of their organisations.
U ovoj epizodi Radio Galaksiji imamo dve gošće sa Odeljenja za psihologiju Filozofskog fakulteta u Beogradu i iz Laboratorije za istraživanje individualnih razlika (LIRA).U gostima su nam bile prof. dr Iris Žeželj i doc. dr Danka Purić, a pričali smo o istraživanjima upitnog i problematičnog zdravstvenog ponašanja, psihološkim osnovama sklonosti ka verovanju i praktikovanju tradicionalne, alternativne i komplementarne medicine i mnogim drugim temama kojima se Iris i Danka sa ostalima iz tima bave kroz projekat Reason4Health. Govorili smo o analizi sadržaja medijskih napisa o tradicionalnoj, komplementarnoj i alternativnoj medicini, rezultatima koje je ovaj tim dobio kroz fokus grupe sa lekarima i fokus grupe sa praktičarima alternativne medicine, ali i budućim temama poput istraživanja zdravstvenog ponašanja pacijenata, njihovog odnosa prema alternativnoj medicini, upitnom ponašanju nepridružavanja zdravstvenim savetima i tretmanima. Naravno, pošto nam je važno da pričamo i o metodologiji i pošto stalno postavljamo pitanje "Kako se istražuje?", pričali smo i o tome na koji način se ovakva pojava u psihologiji istražuje, tj. kakve se sve metode i instrumenti koriste da bi se došlo do merljivih informacija o zdravstvenom ponašanju. Na kraju, razgovarali smo o tome kako iracionalni mentalni sklop (sistem iracionalnih obrazaca mišljenja i iracionalnih uverenja) može da predstavlja važan faktor za sklonost ka upitnim i problematičnim zdravstvenim ponašanjima, kao i svojevrsnu poveznicu između psiholoških dispozicija poput sklonosti prepoznavanja pravila tamo gde ih nema, a konačno i kako se ta saznanja mogu iskoristiti da se to stanje u društvu promeni. Linkovi za dalje istraživanje: projekat Reason4Healthprojekat Reason4Health na Open Science Framework platformi (izveštaji, rezultati, itd)Izveštaj sa fokus grupa sa interesnim grupamaLaboratorija za istraživanje individualnih razlika (LIRA)Support the show
The papers behind the pod:1. https://doi.org/10.7554/eLife.71601 & https://doi.org/10.7554/eLife.679952. https://doi.org/10.3389/fnins.2021.8056793. https://doi.org/10.1038/s41598-021-98356-3It's the 3rd Thursday of January – happy new year! You're listening to 3 Minute 3Rs, your monthly recap of efforts to replace, reduce and refine the use of animals in research. Of course, we focus on those three Rs, but many have suggested adding a fourth R to the list: reproducibility. Designing experiments with reproducibility in mind is a key aspect of reducing unnecessary animal use, as well as being good for advancing science.In 2013 the Center of Open Science and Science Exchange began a collaboration to investigate the reproducibility of 193 experiments from 50 high-impact cancer biology papers. Over eight years of repeated experiments, they found that they could only reproduce 50 experiments from 23 papers, generally due to a lack of detail about the methods used or resources being unavailable. 15 of those 50 repeated experiments used animals, and while just over half of them at least partially confirmed the original results, the repeated results were not always statistically significant. Experimental design was also an issue: only one of the original animal experiments used randomization and none used blinding or calculated a sample size before the study began.Papers describing these results are now available in eLife, with all the relevant data available on the Open Science Framework website and more Replication Studies to come from this collaboration. As the reproducibility crisis continues to rumble on, why not check them out and put designing more robust experiments at the top of your agenda?Next, let's look at how training rats can help make fMRI a less stressful experience. Functional magnetic resonance imaging, or fMRI is a powerful non-invasive procedure that is used to assess brain function and connectivity. However, fMRI research in animals is often confounded due to the physical restraint and loud noises that occur during recordings as these induce stress which can alter information processing and cognition.An article from Frontiers in Neuroscience describes a protocol for habituating rats to fMRI that also avoids the need for surgical head restraint. Rats were gradually trained via 18 sessions over 3 weeks beginning with basic handling phase. After following this protocol, fMRIs in awake rats were successfully conducted without inducing increased stress and still achieving stable images with very low motion artifacts.To learn more about this rat refinement, read the full paper online. Finally, playpens for mice – could they be a viable option for refinement when home cage space is limited? Good environmental enrichment improves the quality of life for laboratory mice by providing increased opportunities to carry out natural behaviours such as running, climbing and burrowing. However, due to space requirements, cost and sanitation constraints many facilities worldwide still use standard housing, which has been associated with potential welfare problems. In their publication in Scientific Reports, Ratuski et al show temporary access to playpens could be an effective method to provide mice housed in standard cages with space and structures to facilitate natural behaviors. In this study, female mice were given access to playpens three times a week for several weeks. Mice in the playpens were more active, compared to mice in conventional cages and over time, the animals entered the playpen more quickly and showed increased anticipatory behaviors before accessing the playpen. All indicating the mice found access to playpens rewarding. Want to learn more? Follow the link in the description. See acast.com/privacy for privacy and opt-out information.
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Evidence from two studies of EA careers advice interventions, published by Jamie_Harris on the AI Alignment Forum. Many thanks to Lauren Mee, David Reinstein, Brenton Mayer, Aaron Gertler, Alex Holness-Tofts, Lynn Tan, Vaidehi Agarwalla, David Moss, and Renee Bell for providing feedback on drafts of this writeup, as well as all who provided feedback on the studies themselves. Summary Animal Advocacy Careers (AAC) ran two longitudinal studies aiming to compare and test the cost-effectiveness of our one-to-one advising calls and our online course. Various forms of these two types of careers advice service have been used by people seeking to build the effective altruism (EA) movement for years, and we expect the results to be informative to EA movement builders, as well as to AAC. We interpret the results as tentative evidence of positive effects from both services, but the effects of each seem to be different. Which is more effective overall depends on your views about which sorts of effects are most important; our guess is that one-to-one calls are slightly more effective per participant, but not by much. One-to-one calls seem substantially more costly per participant, which makes the service harder to scale. There therefore seems to be a tradeoff between costs and apparent effects per participant. We'd guess that the online course was (and will be, once scaled up) slightly more cost-effective, all things considered, but the services might just serve different purposes, especially since the applicants might be different for the different services. Background Animal Advocacy Careers (AAC) ran a longitudinal study testing the effects of our ~1 hour one-to-one careers advising calls, which operated in a similar style to calls given by 80,000 Hours and the organisers of local effective altruism (EA) groups across the world. Over roughly the same time period, we ran a second study using very similar methodology that tested the effects of our ~9 week online course, which taught some core content about effective animal advocacy, effective altruism, and impact-focused career strategy and culminated in support to develop a career plan, either via a group workshop or by redirecting to planning materials by 80,000 Hours. Each study was designed as a randomised controlled trial,[1] and pre-registered on the Open Science Framework (here and here), although a few methodological difficulties mean that we shouldn't interpret the results as giving very conclusive answers. Despite these difficulties, we think that the studies provide useful evidence both for AAC and others focusing on building the effective altruism movement (i.e. the community striving to help others as much as possible using the best evidence available) to help us prioritise our time and resources. We'll be sharing more about the methodological lessons from the studies in a forthcoming post called “EA movement building: Should you run an experiment?” The findings are also written up in the style of a formal academic paper, viewable here. That version provides more detail on the methodology (participants, procedure, and instruments) and contains extensive appendices (predictions, full results, anonymised raw data, R code, and more). In the rest of this post, we summarise some of the key results and takeaways. Which service has larger effects? The ideal evaluation of whether a career advice intervention genuinely increases a participant's expected impact for altruistic causes would be very challenging and expensive.[2] So instead, we designed and collected data on four metrics that we expected to be useful indicators of whether people were making changes in promising directions: “Attitudes,” e.g. views on cause prioritisation, inclination towards effective altruism. “Career plans,” e.g. study plans, internship plans, j...
Facebook announces the Deepfake Detection Challenge, a rolling contest to develop technology to detect deepfakes. The US Senate passes the Deepfake Report Act, bipartisan legislation to understand the risks posed by deepfake videos. And US Representatives Hurd and Kelly announced a new initiative to develop a bipartisan national AI strategy with the Bipartisan Policy Center. In research, AI allows a paralyzed person to “handwrite” using his mind. From the University of Grenoble, a paralyzed man is able to walk using a brain-controlled exoskeleton. From the Moscow Institute of Physics and Technology, researchers use a neural network to reconstruct human thoughts from brain waves in real time using electroencephalography. A report from Elsa Kania and Sam Bendett looks at technology collaborations between Russian and China in A New Sino-Russian High-Tech Partnership. In another response to the National Security Commission on AI, Margarita Konaev publishes With AI, We’ll See Faster Fights, But Longer Wars on the War on the Rocks. James, Witten, Hastie, and Tibshirani release An Introduction to Statistical Learning. Open Science Framework makes THINGS available, an object concept and object image database of nearly 14 GB, over 1800 object concepts and more than 26,000 naturalistic object images. And finally, Janelle Shane explains why the danger of AI is Weirder Than You Think. Click here to visit our website and explore the links mentioned in the episode.
oday with have Brian Nosek on the podcast. Nosek is co-Founder and Executive Director of the Center for Open Science (http://cos.io/) that operates the Open Science Framework (http://osf.io/). The Center for Open Science is enabling open and reproducible research practices worldwide. Brian is also a Professor in the Department of Psychology at the University of Virginia. He received his Ph.D. from Yale University in 2002. He co-founded Project Implicit (http://projectimplicit.net/), a multi-university collaboration for research and education investigating implicit cognition–thoughts and feelings that occur outside of awareness or control. Brian investigates the gap between values and practices, such as when behavior is influenced by factors other than one’s intentions and goals. Research applications of this interest include implicit bias, decision-making, attitudes, ideology, morality, innovation, and barriers to change. Nosek applies this interest to improve the alignment between personal and organizational values and practices. In 2015, he was named one of Nature’s 10 and to the Chronicle for Higher Education Influence list. In this episode we discuss: The genesis of Project Implicit The current state of the field of implicit bias Overuses of the Implicit Association Test (IAT) The common desire people have for simple solutions The potential for misuse of the IAT for real-world selection How hard it is to study human behavior What the IAT is really capturing How the degree to which the IAT is trait or state-like varies by the topic you are investigating Cultural influences on the IAT Brian’s criticism of implicit bias training The latest state of the science on implicit bias How our ideologies creep in even when we are trying to be unbiased The difference between implicit attitudes and conscious attitudes What would an equality of implicit associations look like? Why bias is not necessarily bad The genesis of The Reproducibility Project What are some classic psychological studies that haven’t replicated? The importance of having compassion for the scientist The importance of having the intellectual humility of uncertainty The importance of cultivating the desire to get it right (instead of the desire to be right) What is open science? What is #BroOpenScience? How hostility on social media can cause us to lose the view of the majority The importance of balancing getting it right with being kind to others
Jonathan and Chris interview Brian Nosek, a professor of psychology and the co-founder and director of the Center for Open Science. They discuss problems and solutions in modern scientific research, such as committing scientists… to stick to a protocol. Table of contents. 2:00 The culture of science. 4:18 Publications as currency for career advancement. 7:53 What researchers tell each other at the bar. 10:22 Cynicism. 12:48 The solution to climate change (not really). 18:24 The paper is advertising for the research. 22:16 Weaknesses of the peer review process. 23:58 One data set, many scientists, different conclusions. 27:29 Resistance to sharing. 29:52 The road to the Center for Open Science. 37:49 Signs of success. 44:10 The generational gap in openness. 46:55 Registered reports. LINKS: The Center for Open Science website: http://www.cos.io Project Implicit: https://implicit.harvard.edu/implicit/ For scientists, the Open Science Framework: http://www.osf.io Theme music: "Troll of the Mountain Swing" by the Underscore Orkestra. To contribute to The Body of Evidence, go to our Patreon page at: http://www.patreon.com/thebodyofevidence/.
Joe and Eric talk about Open Science and how it can help reshape the field of the social sciences. If you are interested in your own account on the Open Science Framework or ORCID follow these links: https://osf.io and https://orcid.org
Dan and James discuss how to deal with the problem of scientists who start talking about topics outside their area of expertise. They also discuss what they would do differently if they would do their PhDs again Here's what they cover... The podcast will now be permanently archived on Open Science Framework (https://osf.io/zj7y3/) James did a talk at the Sound Education conference on podcasting for early career researchers. Here's the video (https://www.youtube.com/watch?v=26t6660_f-A) if you want to see him squirm uncomfortably in his chair for 20 minutes and/or hear his thoughts our approach to podcasting The temptation for academics to believe their own press and to have their thoughts reinforced by the praise they get Keeping a handle on what you know and don't know Nassim Nicholas Taleb (https://twitter.com/nntaleb) has FANS The "Pete Evans" effect, James' solution, that we should eat Pete Evans (https://medium.com/@jamesheathers/i-think-i-have-a-solution-i-m-going-to-eat-pete-evans-7e2da6f3967f), pesca-pescaterianism (https://www.youtube.com/watch?v=IC-ZBJ-Kw2E), and the spectacularly bad advice that we should stare into the sun (https://www.sciencealert.com/please-don-t-stare-at-the-sun-even-if-pete-evans-says-it-s-good-for-you) You should follow gynecologist Jennifer Gunter on Twitter (https://twitter.com/DrJenGunter) How much money would you pay for 100,000 engaged twitter followers? Here's the tweet (https://twitter.com/ImHardcory/status/1090213113352372224) James was referring to Should researchers have something like a Hippocratic Oath (https://en.wikipedia.org/wiki/Hippocratic_Oath)? How would we police this? Researchers are not good at admitting they're wrong, do we need to approach retractions differently? Would a bounty system, in which journals offer rewards, for finding errors in their papers, work well? The "Loss of confidence (https://lossofconfidence.com)" project, and Rebecca Willen's CV (https://rmwillen.info/publications/) The "Nobel disease" (http://skepdic.com/nobeldisease.html) Other links - Dan on twitter (www.twitter.com/dsquintana) - James on twitter (www.twitter.com/jamesheathers) - Everything Hertz on twitter (www.twitter.com/hertzpodcast) - Everything Hertz on Facebook (www.facebook.com/everythinghertzpodcast/) Music credits: Lee Rosevere (freemusicarchive.org/music/Lee_Rosevere/) Support us on Patreon (https://www.patreon.com/hertzpodcast) and get bonus stuff! $1 a month or more: Monthly newsletter + Access to behind-the-scenes photos & video via the Patreon app + the the warm feeling you're supporting the show $5 a month or more: All the stuff you get in the $1 tier PLUS a bonus mini episode every month (extras + the bits we couldn't include in our regular episodes) Episode citation and permanent link Quintana, D.S., Heathers, J.A.J. (Hosts). (2019, February 4) "Promiscuous expertise", Everything Hertz [Audio podcast], doi: 10.17605/OSF.IO/VYCAH (https://doi.org/10.17605/OSF.IO/VYCAH)
We had conversations with Christopher Jackson and Jean-Sébastien Caux, two researchers who have started open access publishing platforms. They both told us that academics should be more in charge of the publishing system than they currently are, because publishing is too important for academia to be left at the discretion of the commercial players. Christopher Jackson is a professor of Basin Analysis at Imperial College in London. He has been one of the initiators of EarthArxiv (built on the Open Science Framework. Jean-Sébastien Caux is professor of theoretical Condensed Matter Physics at the University of Amsterdam and a recipient of ERC-advanced grant. He has founded the open-source publishing platform SciPost.org. What role do you think the researchers should play in the publishing industry? What personal initiatives have you taken or are planing to take? How can researchers help each other in promoting academic-lead open-source publishing? Please feel welcome to engage in the discussion on twitter (twitter.com/R2OSpodcast) or on the portal of the Open Science Community Utrecht (openscience-utrecht.com/r2os-episode-3/) where you can also find all the show notes.
Lab Director David Yokum and Executive Director of the Center for Open Science, Brian Nosek, discuss how the core principles of research are not part of daily practice, and they offer some ideas of how we might make them. ****************************** About our guest: Brian Nosek is co-Founder and Executive Director of the Center for Open Science that operates the Open Science Framework. COS is enabling open and reproducible research practices worldwide. Brian is also a Professor in the Department of Psychology at the University of Virginia. He received his Ph.D. from Yale University in 2002. He co-founded Project Implicit, a multi-university collaboration for research and education investigating implicit cognition--thoughts and feelings that occur outside of awareness or control. Brian investigates the gap between values and practices, such as when behavior is influenced by factors other than one's intentions and goals. Research applications of this interest include implicit bias, decision-making, attitudes, ideology, morality, innovation, barriers to change, open science, and reproducibility. In 2015, he was named one of Nature's 10 and to the Chronicle for Higher Education Influence list.
In the final chapter of “Mindware,” Nisbett assures the reader that we’re smarter than we were before started the book, and that we’ll now recognise mistakes in the wild. Are you, dear listener, less likely to make the errors in thinking that we’ve been discussing here? When are you likely to make mistakes? When should you rely on other people’s judgements about a domain? There seems to be an element of politeness when interacting with people who make claims, but is it wrong to, say, ask your doctor how often a diagnosis is wrong? Being sceptical about your own claims and expertise seems to be important in making everyday decisions, so how can we develop this epistemic modesty? Does knowing about experimental methodology help you make better decisions? Does is make you more sceptical? Wouldn’t it be nice if everyone asked to see the evidence before important policy decisions were made? How about an Open Science Framework for public policy? Reading: Mindware by Richard Nisbett, “Keeping It Real” and “The Tools of the Lay Scientist” Guests: Jason Tangen, Rachel Searston, Ruben Laukkonen, Gianni Ribeiro, Jeremy Nash, Brooklyn Corbett, Josephine Echberg, Joshua Adie, Kirsty Kent, Melissa Lane, and Ryan Metcalfe. Learn more at think101.org.
There's a relatively new movement in science called the “Open Science Framework” where researchers put all their cards on the table and make predictions before collecting a single data point. Will it change the way that people conduct experiments? Where do you draw the line between science and mere observation? Carefully controlled experiments trump multiple regression analyses, so why are they often treated equally? Why is the notion of ”wiggling events" so critical in experimentation? Can experimental psychologists calibrate their measurements in the same way that astronomers calibrate their telescopes? Reading: Mindware by Richard Nisbett, “Experiments natural and experiments proper” and “Eekonomics.” Guests: Jason Tangen, Rachel Searston, Ruben Laukkonen, Gianni Ribeiro, Jeremy Nash, and Josephine Echberg. Learn more at think101.org.
How can we make the scientific endeavor more open and transparent? Dr. Brian Nosek is taking on that very question as the executive director of the Center for Open Science, a non-profit technology company developing software that stands to revolutionize the practice of science. The Science Soapbox team chats with Dr. Nosek about the nature of scientific capital-T Truth, the dawn of the “Science Internet,” and why the graduate students of the future should be so jazzed about his Open Science Framework. For show notes, visit sciencesoapbox.org/podcast and subscribe on iTunes or Stitcher.
Wie üblich gibt es auch in dieser Episode wieder Neuigkeiten von der deutschen Science-Crowdfunding-Plattform Sciencestarter.de, dieses Mal allerdings unter dem Blickwinkel scheiternder Projekte. Darüber hinaus gibt es ein paar Neuigkeiten aus dem Bereich Open Access, unter anderem Studien die die Effekte von Open Access Publishing beleuchten. Und zu guter Letzt gibt es ein paar News zu Open Science, besonders mit dem Hinweis auf das mittlerweile in der Betaphase befindliche Open Science Framework.
Rob Wiblin's top recommended EconTalk episodes v0.2 Feb 2020
Brian Nosek of the University of Virginia talks with EconTalk host Russ Roberts about how incentives in academic life create a tension between truth-seeking and professional advancement. Nosek argues that these incentives create a subconscious bias toward making research decisions in favor of novel results that may not be true, particularly in empirical and experimental work in the social sciences. In the second half of the conversation, Nosek details some practical innovations occurring in the field of psychology, to replicate established results and to publicize unpublished results that are not sufficiently exciting to merit publication but that nevertheless advance understanding and knowledge. These include the Open Science Framework and PsychFileDrawer.
Brian Nosek of the University of Virginia talks with EconTalk host Russ Roberts about how incentives in academic life create a tension between truth-seeking and professional advancement. Nosek argues that these incentives create a subconscious bias toward making research decisions in favor of novel results that may not be true, particularly in empirical and experimental work in the social sciences. In the second half of the conversation, Nosek details some practical innovations occurring in the field of psychology, to replicate established results and to publicize unpublished results that are not sufficiently exciting to merit publication but that nevertheless advance understanding and knowledge. These include the Open Science Framework and PsychFileDrawer.
Brian Nosek of the University of Virginia talks with EconTalk host Russ Roberts about how incentives in academic life create a tension between truth-seeking and professional advancement. Nosek argues that these incentives create a subconscious bias toward making research decisions in favor of novel results that may not be true, particularly in empirical and experimental work in the social sciences. In the second half of the conversation, Nosek details some practical innovations occurring in the field of psychology, to replicate established results and to publicize unpublished results that are not sufficiently exciting to merit publication but that nevertheless advance understanding and knowledge. These include the Open Science Framework and PsychFileDrawer.
Brian Nosek of the University of Virginia talks with EconTalk host Russ Roberts about how incentives in academic life create a tension between truth-seeking and professional advancement. Nosek argues that these incentives create a subconscious bias toward making research decisions in favor of novel results that may not be true, particularly in empirical and experimental work in the social sciences. In the second half of the conversation, Nosek details some practical innovations occurring in the field of psychology, to replicate established results and to publicize unpublished results that are not sufficiently exciting to merit publication but that nevertheless advance understanding and knowledge. These include the Open Science Framework and PsychFileDrawer.