Podcasts about Multimodality

Phenomenon of human communication having different forms that combine

  • 108PODCASTS
  • 171EPISODES
  • 38mAVG DURATION
  • 1MONTHLY NEW EPISODE
  • Jul 23, 2026LATEST

POPULARITY

20192020202120222023202420252026


Best podcasts about Multimodality

Latest podcast episodes about Multimodality

Latent Space: The AI Engineer Podcast — CodeGen, Agents, Computer Vision, Data Science, AI UX and all things Software 3.0

In recent months, the open vs closed, and US vs China discussions on model ownership and sovereign/local AI have heated up to a fever pitch. So it is very very good news that Poolside AI are finally emerging with new models, like Laguna S 2.1, that are beating Thinking Machines' recent release nearly 10 times their size.Poolside's recent tech report got a lot of praise due to their level of detail, and Vibhu first covered Laguna's recent technical report on our paper club:From spending $12 million building language models for code before the world cared to creating a Model Factory that can take a model from pre-training to release in eight weeks, Eiso Kant has spent more than a decade betting that code is the path to AGI. In this episode, the Poolside co-founder joins swyx and Vibhu to explain why ChatGPT felt like vindication, why Poolside embraced open weights and open research, and why he would rather live in a world with 100 foundation model companies than five even if Poolside were one of the five.We go deep on Poolside's Model Factory: the engineering systems behind 10,000–20,000 experiments per month, streaming data directly into training, reproducible experimentation, low-precision compute, and agents that increasingly write code, launch jobs, evaluate results, and modify the pipelines used to train future models. Eiso also unpacks their recent launch Laguna S, why persistence, verification, and backtracking may matter more than raw intelligence, how much capability remains inside smaller models, why reinforcement learning will move earlier into pre-training, and why next-token prediction is still extracting too little from the web.We also discuss model-harness co-design, Poolside's path from coding agents to AGI, why Eiso thinks MCP and traditional tool calls are “stupid,” the real economics behind frontier-model training, Poolside's $500 million raise, open-source AI, regulation, NVIDIA and TSMC's influence, engineering productivity in the agent era, high-agency teams, and hiring at Poolside.We discuss:* How Andrej Karpathy's RNN work inspired Eiso to start building language models for code in 2015* Why Eiso spent four years and $12 million pursuing an idea before the market cared* Why ChatGPT felt like vindication and brought Poolside back to open source* Why Eiso would prefer 100 foundation model companies over an oligopoly of five* The difference between releasing open weights and publishing genuinely open research* Why Poolside deliberately built a global research organization outside the Bay Area talent war* Why model building is ultimately 90% engineering* The Model Factory: Poolside's end-to-end system for rapidly training and improving models* How fewer than 70 researchers run roughly 10,000–20,000 experiments each month* How Poolside moved from six-month model cycles to five- and eight-week launches* Why streaming data directly into training unlocked faster experimentation* How immutable data, versioned code, and reproducibility enable rigorous model research* Why Eiso wants capable researchers to leave their labs and become Poolside's competitors* Why 95% of model building can be reduced to better data or compute efficiency* Laguna S and why persistence, verification, and backtracking can outperform raw intelligence* Why smaller models may handle far more knowledge work than previously expected* Why reinforcement learning will move earlier into pre-training* Why next-token prediction is still failing to extract enough knowledge from the web* Why distillation and environments have become the AI industry's favorite “drugs”* Why mid-training is really an early form of curriculum design* Low-precision training, networking bottlenecks, and the next gains in compute efficiency* Laguna S: 118 billion total parameters, 8 billion active, and eight weeks from training to launch* Why model builders can often evaluate a new checkpoint within its first 30 minutes* Model versus harness: where agent capabilities actually come from* Why Poolside sees coding and long-horizon software tasks as a path to AGI* Why Eiso thinks MCP and traditional tool calls are “stupid”* Why future agents will write scripts instead of choosing from dozens of predefined tools* The case for minimal harnesses, containers, and model freedom* Why Poolside is prioritizing vision but does not expect to work on audio soon* Why language may be the most compute-efficient modality for encoding knowledge and reasoning* The real cost of model development and why the final training run is anticlimactic* The story behind the Poolside name and why it represents refusing to lower ambitions* How Poolside raised $500 million while investors still questioned whether AGI was real* Why intelligence could become the world's most demanded and commoditized resource* When open models may become too capable to release without restrictions* Why unilateral AI safety does not work in a globally competitive environment* How regulation could accidentally lock in an oligopoly of two or three AI companies* NVIDIA, TSMC, and the hardware systems underpinning foundation-model progress* Why reinforcement-learning wall-clock time is one of Poolside's biggest bottlenecks* Why Poolside trains models from scratch instead of simply distilling larger models* How AI changes the way companies should measure engineering productivity* Why agency may become the most important quality for employees in the AI era* How leaders align high-agency people through shared goals and clear constraints* Hiring across research, post-training, pre-training, architecture, evals, and engineering at PoolsideEiso KantLinkedIn: https://www.linkedin.com/in/eisokantX: https://x.com/eisokantPoolside: https://poolside.aiTimestamps00:00:00 Introduction00:00:54 Karpathy, RNNs, and Building Code Models Before Transformers00:02:26 The $12M Failure and ChatGPT Vindication00:03:39 Open Source and the Case for 100 Foundation Model Companies00:09:22 Open Weights, Open Research, and Poolside's Global Team00:16:04 The Model Factory: Why Model Building Is 90% Engineering00:20:19 Agents, Automated Experiments, and Early Signs of RSI00:24:04 Streaming Data, Reproducibility, and Scientific Rigor00:30:35 Creating More Foundation Model Companies00:36:07 Laguna S: Persistence vs. Raw Intelligence00:43:01 Reinventing Pre-Training, RL, and Curriculum Design00:52:33 Low-Precision Training and Squeezing More From Smaller Models00:58:37 Model Harnesses, Coding Agents, and the Path to AGI01:09:26 Why MCP and Traditional Tool Calls Are “Stupid”01:13:04 Vision, Multimodality, and Why Language Still Matters01:18:15 Scaling Models and the Real Economics of Training01:20:40 Why Poolside Is Called Poolside and Raising $500M01:27:37 Open Models, AI Safety, and the Risk of an Oligopoly01:33:53 NVIDIA, TSMC, and the Reinforcement-Learning Bottleneck01:41:52 Smaller Models, Distillation, Engineering Productivity, and HiringTranscriptIntroduction: Eiso Kant, Poolside, and Open ModelsSwyx [00:00:00]: All right, we're here in the studio with Eiso Kant from Poolside, together with Vibhu. Welcome.Eiso Kant [00:00:08]: Thanks. Thanks for having me, guys. Good to be here.Swyx [00:00:10]: Yeah, fresh on the plane. You texted me, you were like, “Hey, I'm on my way to SF.” I was like, “You're on a plane right now, right?” Like, hey.Eiso Kant [00:00:16]: I know. After I texted you, I realized that probably coming in with major jet lag was gonna offer some fun experiences today, but let's do it.Swyx [00:00:23]: I mean, I think the thing I would tell guests is that they don't have to prepare that much because if you're truly working on this every single day, then even, like, what you hazily remember is going to be new for a lot of the audience that don't live in your world every day, right? so 10 years ago, you did a talk at Google Slush, talking about the democratization of AI. and, now here you are, like, open sourcing an incredible new model that we're gonna talk about. But I guess, like, what got you into democratization of AI? Like, it's not obvious from your LinkedIn or something.From Karpathy's RNN Post to SourcedEiso Kant [00:00:57]: No, it's not at all. I don't think it's obvious how I got in this space. I owe getting into this space to Andrej Karpathy.Eiso Kant [00:01:05]: In 2015, he wrote an article called “The Unreasonable Effectiveness of Recurrent Neural Nets.”Swyx [00:01:10]: Neural Nets, yep.Eiso Kant [00:01:11]: And that article, I read it, and I pivoted my startup at the time overnight to working on RNNs, and later LSTMs and Transformer models to be able to write code. If you go to this article and you scroll down, you can start seeing, like, this was the precursor to what ended up becoming language models. So, at least when he was character-level language models that were starting to predict letters, he has an example out here. There's a little Paul Graham generator, and you can read it, and the text makes sense, but it doesn't. and there's a little-- There's an example of code a little bit further down. Yeah, so Shakespeare.Swyx [00:01:47]: Shakespeare.Swyx [00:01:49]: CoolEiso Kant [00:01:49]: And for some reason, I read this, and I went down the rabbit hole of learning everything I could about RNNs and LSTMs, right? This is Transformer paper. And I had built a completely unreasonable belief, that neural nets should be able to generalize to anything and everything, and that language should be able to generalize, to a lot of things that are intelligent and the ability to write code. And so I started building Sourced, which was a fully open source company trying to build, what we used to call machine learning on code, language models on code. And we spent about four or five years on this, till the end of 2019. And that sounds really cool today, but back then, no one cared.Eiso Kant [00:02:29]: Right? Like, no one cared. We were in the dark. Like, we did things along the way. We tried applying convolutional neural nets to, like, the structure of code. We were. when attention came out, we were applying it to LSTMs, and then the Transformer paper came out. And it - it wasn't obvious, and what we missed throughout that entire journey, that we were on the right track, but we should have just kept scaling up. And today, to all of us, the scaling laws and scaling up seems like the most obvious thing. But having spent four or five years of my life on working on language models on code, it wasn't obvious. So I have a lot of respect to folks at Google and OpenAI and others who took that confidence and kept going. we failed ultimately at the time, and it was, like, biggest failure of my career, right? You blew $12 million of investors' money, which was a lot back then.Swyx [00:03:18]: Yep.Eiso Kant [00:03:19]: You spent, still a lot, but, And you spent years with, like, a group of 40 people just obsessing over this problem. And life took a different turn, And it was, and family became a focus, and I kept my heads down and really, didn't really look at language models for the following two years. big mistake considering Following years are gonna be really interesting. And then ChatGPT came out And it was like a vindication. It's like people started texting me. I found, like, my old, work decks and these old talks. And throughout that whole journey, we,ChatGPT, Vindication, and Returning to Open SourceEiso Kant [00:03:56]: We really had a strong point of view at the time that, like, as you're building more capable intelligence, it should be open and open source.Eiso Kant [00:04:04]: When we started Poolside, that wasn't the case at all, and I wanna be very open about it. When we started Poolside, we were like, there was a premise of two things. One is this technology is not gonna stop compounding in capabilities. I think to most people obvious today, but three-plus years ago when we started, most people were still arguing if these were stochastic parrots or not.Eiso Kant [00:04:23]: And the second was that reinforcement learning was gonna be the biggest driver for LLM capabilities. Today, very obvious. Three years ago, was not an opinion held or direction held at either OpenAI or Google or Anthropic or others. And so people looked down on us a little bit. They were like, “ is this really gonna work?” And so we just started working the problem, and we never really thought about open source again. We just kept our heads down and we built our, like, knowledge, understanding from scratch, right? We didn't roll out of an existing lab. So we picked up the papers and started writing code and figuring things out.Eiso Kant [00:04:59]: And it wasn't until the beginning of this year that me and my founder, Jason, picked up the open source conversation again.Eiso Kant [00:05:07]: And if you go back to some of the early things on our website, it was very straightforward. It was we wanna get to AGI, we wanna support a world of abundance, and we wanna be the first company that gets there.Eiso Kant [00:05:20]: But we started talking at the beginning of this year because it became obvious that the world was going in a direction that was starting to like, pick at us a little bit. Like, it didn't, this didn't happen overnight. It was, like, a little bit we were seeing this and we're like, “Okay, The world's going down a path.” And Throughout this journey, there was something that I used as a, as an analogy or thing. So I said well, if I go back to back in those days, 2015 or 2016, we're working on this, and I picked up a fi book off the shelf, and I was reading the book about 2035. AGI is achieved, and the story would be over the following, decades. And it would have that first chapter where everyone's trying to figure things out. You'd get the chapter of ChatGPT coming out And then you would get to the chapter where the world was at a fork in the road, and the one that it picked was one where three or four or a handful of companies were going to create all of intelligence moving forward.Eiso Kant [00:06:21]: And when I thought about that story, it felt like a dystopian fi book, not a utopian fi book. And the reality is, I'm a utopian fi guy. Like, and so We took a step back and said, “Hey, can we play a role here?” Now it was easy for us to do so because we were not at the frontier.Eiso Kant [00:06:41]: If we were at the frontier, I don't think we could have changed our mind. and I don't mean this like it's when the moment there's too much capital involved, too much expectations, you've built up things, right? We're a small team, just improving and improving. And so we knew that we could make that decision now, but it would be a lot harder to make as we got closer and closer to the frontier and caught up to others. And did a lot of soul-searching and a lot of conversations, and said, “No, this makes sense,” Even if there's big unanswered questions, like how the hell do you build a business model with foundation models about open source? Big open-ended question that we do not fully have the answer to yet, right? At what point do you no longer wanna release open source models because misuse of models has, real potential risks associated with it? how is the government gonna respond to open source? but I think it all just came down to one thing, and I'll stop the monologue, is the fact that I rather live in a world that has 100 foundation model companies than a world that has five, even if I was one of the five. And the smallest and most meaningful contribution we can make for 100 to exist is to open up our research and open up, like, our weights right now and figure out along the way how we can, like, do more.Neo-Labs, Model Choice, and the Token EconomySwyx [00:08:01]: Yeah. I think if anything, over the past three years, that has become a bit more true. you are one of a cohort of Neo labsEiso Kant [00:08:10]: YeahSwyx [00:08:10]: That people are now calling that. And, we're, we're doing this on the day that Thinky launched their, new model and you are outperforming them on their, on some benchmarks that they released, right? Like, they just don't have it yet. so it goes to show that I think, like, this is one of those things where, like, there is room for multiple players, and you are seeing a little bit more of the future. Maybe more like 20, not 100, but, like, you are one of the 20.Eiso Kant [00:08:36]: I really hope so, right? I think we I'm, I'm excited about their release, and I'm excited about everyone releasing because, like, ultimately, like, choice competition is both gonna drive progress in the right direction. But the fact that like, we create models and while we all, drink out of the same well of data effectively, we do introduce very different behaviors and biases in our models. Some are intended biases, some are completely unintended biases.Swyx [00:09:03]: Yeah.Eiso Kant [00:09:03]: And if we shape up in an ecosystem in the world where open models are gonna be a part of the token economy, like, I don't think there's any question about it anymore Then we want to be able to live in a world where companies, countries, people can choose and say, “Hey, I am most aligned and I trust most this provider for these things.”Swyx [00:09:25]: Yeah.Vibhu [00:09:26]: I think more than just one of the 20 Neo labs, up until recently, most of open source innovation was coming from the Chinese labs, right? So there's the DeepSeek of the West. Is it today? Okay, maybe it's thinking machines reflection, but there aren't many, right? So, one of the things you guys started in France, Europe, but very much now you're taking that American standpoint and more than just that, the point is the Chinese models that we see, they're not super open research. the work you put out is, I think, some of the best. So every few months you get not only frontier models, but also here's a breakdown blog, paper, technical report of here's everything for state of the art to build, frontier intelligence and you're filling that gap too, right? So not just only open weight, not just Western, but also pretty open research.Open Weights vs. Open ResearchEiso Kant [00:10:20]: No, I appreciate it. Look, I think it's, I think it's the most meaningful contribution, right? Weights are a binary. Let's call them what they are. Yes, we can modify them, we can change them, but, like, giving someone the weights does not allow them ultimately to recreate what you're doing, right? And so now there's challenges around releasing data sets, challenges around like releasing certain things, but being able to share your research, like, right, how do we do it? What are the lessons we learned that we spent, tens of thousands of experiments of compute on? I think very much so. One correction though, Vibhu, and I say this because it's been haunting us for quite a few years. We from day zero were an American company.Swyx [00:10:55]: Yeah. They movedPoolside's Global Team and American Company StorySwyx [00:10:56]: To France.Eiso Kant [00:10:56]: So the story once and for all is very. We start as an American company. We have always been an American company, and early on we made a very conscious decision. We said, “We're not gonna hire any researchers in the Bay Area. We're gonna look for talent everywhere else in the world.” and that is everything from Middle Americas, Seattle to, Serbia, and to Taiwan and Singapore and other places. And it was because we took a view that this was gonna become a talent war for this, and I think it has over the years now. Three years ago, that wasn't fully obvious yet. I think today it very much is. And we also realized that, like, some of the world's most capable people with, like, the most interesting, innovative ideas were not just gonna be here. And so it led us to create like a fully remote company. and we ended up opening an office in Paris and London and different places and we have a lot of the team in the US and a lot of team outside. But we always took this view of like, we're an American company, but if we want the best of the best to work with us, we need to take a global view. Now we do also have people here in Silicon Valley, like the company's grown and others, but I think one of the things that, it slowed us down at the beginning, but it has sped us up now, and it's why you're seeing like the progress, I think, on our models and the cadence at which we release, is because we didn't roll out of an existing lab. Right? we didn't, we didn't have a lot of the information that's freely flowing around here at the time. We just took this point of view as like, “Okay, well, let's just work the problem. Let's just go and, like, read the few papers that are out there, and let's just figure this stuff out.” And we made some hilarious mistakes in model training because of that over the yearsEiso Kant [00:12:35]: Like especially in the first 12 months. there's a few that I think still haunt me and scare me. We can talk about them later. but it created a, like, a resiliency and persistency in the team, right? with extremely few people have left us over the years, that, like, told us, “Okay, we can do this.” When we first wrote our first training code base completely from scratch, it wasn't a fork of any open source. It was just like, “Okay, let's build it from scratch.” I remember we had this one moment where we spent three weeks working out an optimizer bug. Like, it was like training just couldn't get stable. We, like, obsessed over it, and we thought, like, maybe we were wrong. Maybe we should have just forked this repo, or we should have. But then when we solved it, I still remember at the time we were like five people in the company. when we solved it, we were like, “Oh, we can do things,” like if we're just willing to work hard. and I think that culture with a very strong engineering bias has helped us, like, get to where we were. And so there's this notion of open source and talent and these things. I think we, We just took different decisions from a different starting point. and I think we are lucky. I do want to definitely call it lucky. And there was a lot of hard work at the team that now, like, that's starting to show up in results.Swyx [00:13:52]: Just ‘cause we probably won't revisit this again, but, and this is a fun recruiting challenge if someone knows the answer. What was the bug? And then we won't tell the solution, but we'An Optimizer Bug and the Value of Building From ScratchEiso Kant [00:14:01]: So the - This - You're gonna test my memory here,Swyx [00:14:04]: Oh, okayEiso Kant [00:14:04]: So but I thinkSwyx [00:14:05]: DirectlyEiso Kant [00:14:05]: I think I can recall. So if you, so if you look at, So if you take like Adam as an optimizer, you have epsilonSwyx [00:14:12]: YeahEiso Kant [00:14:13]: Which is, right, like in the denominatorSwyx [00:14:14]: Momentum and weights. YeahEiso Kant [00:14:15]: Is exactly, in the denominator. And at the time, if I recall, you looked at like the early Llama papers and things like that. People were juicing epsilon, like, quite a bit. Like, they were, like, adding, I don't know if it was E minus four or whatever, like a high value for epsilon.Eiso Kant [00:14:31]: And if you think about this during training, it's like a bit weird and counterintuitive that we're adding noise to our optimizer by just adding effectively, like, a random number in the denominator, right? Like behind the decimal point. And I don't recall the exact bug, but it had - What I remember is once we solved it, we no longer had to juice epsilon as much as, like, was happening in the Llama paper and other places. and it was like one of those fundamental moments where we had trusted this paper that was out there, and we're like, “Oh, no, it has to be this way. It has to have this high value of epsilon.” But it made no sense to us intuitively. Like, why do you have to have this so high? Like, if you're just trying to avoid division by zero, why can't the value be extremely small? and that was like one of those moments where you realize like, okay, finding things out from scratch yourself builds a better intuition. Because the one thing you learn very quickly with model building is that your intuitions that you start with are gonna get beaten up so hard.Eiso Kant [00:15:33]: Right? Like - It's such an experimental science, that the things that seem obvious, you very quickly get to learn, like, you were wrong, and hopefully you figure out why, and sometimes you don't even.Swyx [00:15:45]: Yeah. yeah, so, one of the reasons that you, when you released your new models, Vibhu got really excited. I mean, everyone got really excited. But Vibhu led our paper club on it, and you guys sawEiso Kant [00:15:58]: YeahSwyx [00:15:58]: Obviously. maybe talk through some lessons learned in that, whatever you can disclose. we can focus on the model factory stuff, whatever you think is a good starting point.Model Building as EngineeringEiso Kant [00:16:08]: So I would say that our view from very early on in the company was that model building is ultimately 90% engineering.Eiso Kant [00:16:18]: And I think we all know it in the industry because if you look at where's every researcher spending their time, they're spending their time writing code, right? Looking at data and writing code. And so we said, okay, The state at the moment, like three years ago, was bash scripts and Slurm and spaghetti code bases for training and, like, data pipelines that were patched together. And we looked at this and said, “Well, ultimately, model building is a process.” You're going from raw data, right? Like training raw material, the web, et cetera. you're doing a whole bunch of filtering, cleaning up, transformations, analyzing. These days, that's, far more complex than it was three years ago. then you're training a model, which is effectively a large distributed systems problem, right? Across hardware that has still-- It's become a lot more reliable. It was extremely flaky back then. and now with every new generation, we get our new sets of challenges. And then you go into the next stages, right? There was no training back then, but, like, you got, your post-training and then your reinforcement learning. And so we looked at this and we said, “Well, this looks like an industrialized process. This looks like an end process, that every single part of it has its machinery,” right? If it's your big data pipelines, if it's your crawling ingestion of the web, if it's your, large-scale distributed training, and then you've got your reliability. And we said, “Well, why don't we take some of the world's smartest distributed systems engineers that we knew and make them part of the process of research from day zero?” Not retrofitting it later on, but, like, really from the beginning. And that became our model factory. And so our model factory started with a handful of components. Today, it's thousands of components, and I try to equate it to, if you think about, like, someone who was at the very early days of Foxconn, if they had been there for the following, decade, they would be able to rebuild Foxconn because they saw every decision that led to building that system and all the complexity. If you and I walk into Foxconn today, no chance.The Model Factory and Experiment VelocityEiso Kant [00:18:18]: Right? Because we don't have the lineage and history of decisions that led to that. And so we built early on from the beginning- with a team that really understood that, well, the metric that we are optimizing for is the speed of an idea from a researcher to an experimental result that we can trust to then being part of the next model training.Eiso Kant [00:18:42]: And in the. And because it's such an experimental science, ultimately, in the beginning when it wasn't that complex, you could patch your way around it, right? But now, at any foundation model company, you are running. I mean, we're a small team, right? We're less than 70 researchers, another 35 engineers. and we are running, I haven't checked the latest count, but far more than 10,000, maybe 10 to 20,000 experiments a month that we cut. And so if you look at that scale of every model run that is, like it's ultimately it's, it's you need to be able to trust it as an infra problem. And so what we have now done over the years is gotten really good at that, and just by working it and improving it and obsessing over those end decisions. So now what that means is that you looked up Laguna XS 2 that we launched. It was five weeks from the beginning of training to launch. The model that we're gonna talk about today was eight weeks from start of training, to launch. We started the next model literally yesterday because we now finished the post-training required for the model we're launching, next week or by the time this comes out today. and we move that compute to the much larger Laguna M model that we're now training. And so the model should be an artifact of someone's process. It shouldn't be really a thing in itself. Like, and we treat this like the way you would look at like a SpaceX factory where, yes, the first rocket, really hard to build, but the much harder challenge was building the factory. And now they're rolling off, and no one is really thinking about the next launch anymore. So it's just another launch, it's another launch, another rocket comes off. And that's what we're trying to do with model building.Eiso Kant [00:20:22]: And what has been, which was not planned from day zero, it was in the back of our mind like this will happen one day, is that when you build a really good end model factory with really good APIs and really good engineering systems, Well, what is it perfect for? It's perfect for agents.Agents Inside the Model FactoryEiso Kant [00:20:40]: Because agents are now starting to take over more and more work in our model factory.Vibhu [00:20:43]: Yeah.Eiso Kant [00:20:44]: So I look at the screens when I walk, like when we're, we come together, in our monthly, we do monthly onsites, and I walk behind people's screens and I stop by and I talk to our researchers. And the default is all of these different agents running on their screen that are writing the code. They're launching the jobs. They're evaluating the results that are coming back from the model runs. They are, making the changes. And we're still in the driver's seat. We're still coming up with the ideas. We're still helping with the debugging. But more and more, and this is right now very profound on the data side of our pipelines in both pre and post and the synthetic data pipelines, it's starting to become more on the architecture side as well. You're starting to see these twinklings of what RSI is gonna look like.Eiso Kant [00:21:27]: And that's. So when we talk about, like to your question about our models, every talk about the model factory, And my coolest example of these things is always that when we kick off a new run, doesn't matter if it's a training like big run or if it's now a post, like one of 10 post-training versions we do for like release or many experiments, is that at any given moment, the changes that somebody made that they had experimental results from the day before make it into that run.Eiso Kant [00:21:57]: So there's not like a cutoff 90 days before. Like no, it's like literally from that moment because we can now trust the machine enough. And then you also have to invest in the reliability. So one of my favorite metrics about like Laguna S is that there was no call events, Right? Like completely zero. And we haven't had a meaningful call event, like something to wake up for, as far as I recall this entire year. now there is one asterisk to that. In usually the first six hours of launching a new model run, something breaks because you set a config wrong, you made a small mistake, et cetera. So that's usually there's a little bit of intervention, but that's always within like call periods, right? Not on call. And I think that's starting to now compound. So the model we're releasing now, I love it. It's amazing, but we're already onto the next one. and I think that's the way it should be.Laguna, Five-Week Builds, and Zero On-Call EventsVibhu [00:22:50]: Hey, I also just wanna point out, so for context, this was like a month ago. we found it in the tech report, so we just came in with, “Okay, new model's dropped. Haven't heard about it.” We wereEiso Kant [00:23:02]: Yeah, we're very used to doing this every few months.Vibhu [00:23:03]: We're, we're very much like, “ okay, look, it's like, on par with Kimi, DeepSeek, whatnot, the small ones, Gemma level. Oh, it's a very cool paper on what goes into building.” And then we hit this page, right? Like literally page two of tech report is, “This process allowed us to build the small model from scratch to delivery within five weeks applying the lessons”. And then I'm like, oh, this paper is not about here's a tech report of benchmarks and here's how many tokens it was trained on. Like for people that wanna dive more from what we're not gonna discuss on the podcast, it's all laid out here, right? FromEiso Kant [00:23:38]: YeahVibhu [00:23:39]: Custom software that agents can use to interface with training code, training data.Eiso Kant [00:23:45]: Yeah. Well, link the paper correctly, so yeah.Vibhu [00:23:47]: Yeah. All that stuff. read the paper here, but,Technical Report Principles and Streaming Training DataEiso Kant [00:23:50]: But I would like to. I love principles, and I think that is a good starting off point for maybe telling some stories. Maybe we can go one by one past the principles. I'll just call out that Dagster just got bought by a Prefect.Vibhu [00:24:01]: Yeah.Eiso Kant [00:24:01]: Isn't it fun? But yes, I'm very familiar with Dagster. just anything where like they trigger some story.Vibhu [00:24:07]: So, well, I would say, well, experiments code's obvious, but I think one of my favorite things is, I don't know where it is in here, but early on, and I still think this is the case a lot of foundation model companies, people prepare their training data sets, they get packaged up, then they get copied over to a training cluster distributed across all of the nodes, and then training starts.Vibhu [00:24:30]: And we looked at this like three years ago and we were like That makes no senseEiso Kant [00:24:36]: You lose so much time because the moment you have to rematerialize the data set, you have to make a change, you have to fix something, et cetera, you've got all this time of like repackaging it, right? Toca- tokenizing it, repacking it, moving it over to a cluster, then distributing it across the nodes. The bigger your clusters are, you start using fancy like torrent-like algorithms to like distribute your data. So why aren't we streaming data into training? Right? Something that's very common and like just basicVibhu [00:25:00]: Like just in timeEiso Kant [00:25:01]: Just in time, like good computer science like principle. And that was one of the first things that I think unlocked - the model factory. Because the moment you start thinking about, well, a training job, it doesn't matter if it's a big hero run or a small like, post-training experiment, consumes a certain number of tokens per second, right? And it's not a lot, right? From a like a data, moving data perspective. So we said, well, we have our training cluster, and then we've got like our AWS kinda setup where we can build these amazing big data pipelines. We can set things up. We use Spark underneath the hood, like all these things.Vibhu [00:25:36]: But when you say AWS, it's not actual AWS, it's your internal AWS.Eiso Kant [00:25:39]: It's our internal-- No, it's our internal like just running like our infrastructureVibhu [00:25:42]: Site web servicesEiso Kant [00:25:43]: Exactly. Our stuff running on like an AWS account or on like any hardware, right?Vibhu [00:25:47]: Yeah.Eiso Kant [00:25:48]: And so once we made that shift into I can stream data into training, all of a sudden you realize a lot of things unlock. Because now you don't have to wait for the whole data set to materialize.Immutable Data, Experiments as Code, and Scientific RigorEiso Kant [00:26:00]: You now all of a sudden when you're running data experiments about mixing data, it's a config. Because you've got these data sources that are coming in, and you just - we have this service called Blender that's in the report, where we then say, “Okay, for this run, I want 20% of this source, 10% of this source. I want this much, so many epochs of repetition. I want this to be, shuffled in a certain way,” and your training job can start while the rest of the data is even still materializing. also what it does is because all of this underneath-- So for us, we treated the data layer underneath as like an immutable data layer, and that was really important. Like experiments as code, immutable data layer means that you can always go back and understand literally down to the single token at which cursor it went in on which version of the code.Vibhu [00:26:47]: Yeah.Eiso Kant [00:26:48]: And it took us a I have to admit, like the first year of Poolside, we understood that engineering had to get great, But we didn't understand yet, that this is ultimately in support of like a good rigorous scientific progress. We were quite a - We were a very small number of people, so a lot of it was YOLO ideas and YOLO runs.Vibhu [00:27:08]: Yeah.Eiso Kant [00:27:09]: And we built great infra for the YOLO runs. But once we realized that we treated data as immutable and code as always versioned, and you could always track and trace every experiment end to end perfectly, you could repeat everything perfectly, right? You have perfect reproducibility. I can still reproduce runs from two years ago if I wanted to, right? It enables the scientific progress, like the scientific process, and I think that took us probably about a year and a half into the company to figure out. We also had some great hires, like our head of applied research, Nikolai, who joined us from Yandex, who'd been working on language models since like the early 2020s, I think brought that into the company of like, “Hey, we wanna have even more rigor.” And then once we kinda had the combination of like increasingly more capable platform that allowed people to do more, but had this immutability, we were able to start “Okay, every experiment is truly an ablation. We truly need to understand it.” And I think we became much more scientifically rigorous in the last couple of years, and the infra underneath enabled it. and then there's just fun stuff like, andVibhu [00:28:16]: Yeah, a lot of it's fun, like even just the, one, you share all the ablations, two, picking the data sets, right? There's like a random small paragraph in here where it's just like, “Oh yeah, training data, we have some, we have an auto mixer.” it trains eight small models, scales them up, picks the training data set. We don't even need to look at it. I'm like, “Wow, a lot of engineering rigor there.” And there's just, there's just a lot in here.Publishing Research and Giving BackEiso Kant [00:28:40]: Yeah, and it'- and look, and we wanna put out more. Like we, We treat writing papers as something that we haven't earned the right for yet for a long time. So you earn the right to spend time, publishing research once you're at the frontier, because until then, you're catching up, and every minute and hour in this industry matters. Like I obsess over it, not just the wall clock time from idea to result, but just general like time every day that we, waste is one that doesn't allow us to catch up. But in this case, we said, “Okay, we're gonna give ourselves.” I think we gave the team like three or four days while still doing their work, like give everything in there. And to your point earlier, if your stuff, it's easy to like put it out. And so there's so many more things that we wanna talk about over time, and we will definitely start doing. And as we earn more of the right, but also now have like added to our mission that we want more foundation model companies to exist, you'll see us like be way more proactive, and just trying to keep dropping some of those like things that we've learned along the way that can help others like speed up.Vibhu [00:29:40]: Which is the other cool side of this, right? It's, it's not like, back to your point, it's not just here's the benchmarks of our training. If you want to replicate, here's experiments of optimizers, data sets, post-training. you lay out a lot of it here alongside here's your system for how to do it? So it's, it's really like promotingEiso Kant [00:29:59]: No, thank youVibhu [00:29:59]: Other people can do the same.Eiso Kant [00:30:00]: And by the way, I also wanna make clear, right, we have been incredible-- Like we've taken a lot of advantage of the fact of all the open research that others have published, Right? And you mentioned, the Chinese labs, and we I think it's important that there's, from every country and every culture and background, including like Western companies like us, there's different models that come out that people can choose to trust. But I think we do have to give credit where credit's due, right? The incredible Chinese lab have done an amazing job at sharing their research, and we have definitely like been on the receiving end of taking advantage of that. So when you're on the receiving end of something coming to you, I think it's, you also have an obligation to give back.Swyx [00:30:39]: Do you have a favorite or underrated Chinese lab that you wanna shout out? Everyone shout outs DeepSeek.Chinese Labs, Zhipu, and PersistenceEiso Kant [00:30:44]: That's a good question.Swyx [00:30:45]: Moaan obviously for Therapsi. Yeah.Eiso Kant [00:30:48]: Yeah, look, I think, I think obviously everyone's been talking about Zhipu lately, with 5.2. I think what most people don't realize is when they started.Swyx [00:30:59]: Yeah.Eiso Kant [00:30:59]: Right? They started years before ChatGPT.Swyx [00:31:02]: They just rebranded. YeahEiso Kant [00:31:03]: And so, I've like, I remember how hard it was to work on these things Before the rest of the world got excited about it. And so I have an immense amount of respect for people, who were working on improving models when it wasn't the sexy thing to do, when believing in LLMs, was gonna get you ridiculed. I remember like back in 2016 when we were doing what we'd call, machine learning on code with some of these models. we would-- people would just laugh at us, like they'd be like, “This makes no sense. Like why are you wasting all these, like, millions of dollars on trying to figure this out?” And so I would say they're probably the one that, I think deserves a shout-out, not just because their latest model is very good, but because they fought to get here. And I think, I think every foundation model company it takes time to get here, right? It took us three years to get to the model that we're, that we're now gonna be releasing. and now the time in between the models is coming, is counted in weeks. It's no longer counted in months or years. But this stuff's hard. and if we can make it a little bit easier for the next person, like we should all do so. Because if we don't do so, we're, we've got a small window before models are really impacting recursive self-improvement to a level where catching up otherwise might become unfeasible. And we should try to, in that window, encourage as many labs or however we wanna call them, like to start. And so one of my currentEiso Kant [00:32:36]: Mission, but qualm is like I wanna encourage whoever is a researcher right now who thinks they can tackle this to go and leave and become my competitor.Eiso Kant [00:32:45]: Like start another foundation model company because I think we need it. I think otherwise we're not gonna be in the world where, I don't want to just be the fifth or the sixth company that wins. I wanna look at a world where there's lots of choice.Starting a Foundation Model CompanyVibhu [00:32:57]: What else do people not see in starting a foundation model? it's, there's a lot of compute, there's a lot of capital required, a lot of compute. You lay out model factory and how to do the training, but there's a lot there, right? That's,Eiso Kant [00:33:10]: Well, look, it's, I in turn-- this is an oversimplification, and I always asterisk it with that because it can land a little bit the wrong way in people's minds. But I think you can sum down, And I saw it, 95% of model building to just doing, you're just doing two things. You're improving data or you're improving compute efficiency. And I know that feels like an oversimplification for the incredible, like, Gifted and skilled work people do. But if you really look at it, like what are we doing? We are looking at data, we're generating new data, we're improving data. and the only way to do that is to look at the data, right? That's a big part of foundation model building. And on the other hand, we come up with these incredible breakthroughs in inference, in architecture, and new attention mechanisms. But what are they really doing? They're bringing compute efficiency. Now, we have definitely had some breakthroughs over the years that allow for more model capabilities. But at the limit, if you could train a large enough model, right, like, and you had infinite compute, we probably-- if you had infinite compute, you'd be at AGI probably already tomorrow.Eiso Kant [00:34:12]: Right? Like it's not. And so, and let me say that infinite compute with infinite ability of much faster networking because networking ends up being more of the bottleneck than compute. But, so I do think that's, those are the main things. And to just realize that this is engineering. I think it's become more obvious, but I think for quite a few years, people have held foundation model companies and researchers and others on this pedestal of like you're doing incredible magic or rocket science, or only like, Nobel laureate physicists can do this. And don't get me wrong, there are some really hard problems that need to be solved, but a lot of the work that all of us are doing on a day Is not sitting down trying to solve a math theorem. A lot of the work that we're doing is just really doing the basics right, writing good code, looking at data, improving it, running experiments, looking at plots, trying to see like, hey, trying to shape our intuitions. And a lot more people could be highly capable researchers. and I think that's, it feels far for people to do so. But I've seen in our own company, we've seen engineers become researchers because the model factory allowed them to be, have a much lower hurdle of running experiments and trying things. And one of the guys on our team who started as an engineer building our agents is a legit reinforcement learning researcher now, making real progress. and that happened in the span of like six months. that would've not been what I think most people assumed was possible, a couple of years ago.Swyx [00:35:46]: Yeah. I think one of the interesting moments is when you can self-host, like, if in a programming language, like if you can compile the language in the language, the equivalent is can you use your own tools, right? You have the pool CLI, you have your own models. presumably you're not only using your own models. There's no way. But like, what's that percentage over time?Laguna S, Persistence, and Behavioral GainsEiso Kant [00:36:10]: This is the first model that we're releasing that is starting to meaningfully contribute to our own work. It's not a it's not state-art model yet. Fable and other, they're, they're very capable models, but Laguna S Is really interesting. I'm gonna pull up the quote. Peng Ming, one of our heads of applied research, said something, last week as the model came out about 10 days ago, much better than we had hoped for or expected. And he said, I have the feeling that a lot of the gains in Laguna S come not from more intelligence, but more from different behavior, more verification, less taking things for granted, not declaring victory early, and being way more persistent. And to be honest, those are more predictive than raw intelligence for success in human also to some degree. And this was, he wrote me this on 5th of July on a Sunday, and it's been burned in my brain ever since because the Laguna S model, as you'll see it and why it does so well on benchmarks and why it does so well in using it on a day basis, is that it's just incredibly persistent. It reasons a lot. I do call that out. We have work to do on making it more efficient. We have to work to do on offering different reasoning modes. But this is the model that has been able to do things that I never thought it could do. A hundred eighteen billion 8B active model, which is not that large. It fits on a DGX Spark and still runs at, thirty, forty tokens a second on a Spark, is able to solve Erdős 397 independently. It's able to do complex programming tasks. It's able to. I asked it this morning to make me a Fi scanner without using any external libraries on my Mac, and it's, like, figuring out, like, the core WLAN API by really persistently trying to understand it without access to the internet. And more, I love vibe checking. I've probably spent eight to ten hours a day with this model for the last ten days.Eiso Kant [00:38:05]: I'm not exaggerating. I was on my eleven-hour flight yesterday. I spent ten hours reading trajectories and traces and, like, of the model.Eiso Kant [00:38:12]: And what I take away from it is exactly what Peng Ming said. We are gonna be able to squeeze so much more out of smaller models than I think we had imagined in the industry because, yes, there's intelligence and larger models are more intelligent. Like, no doubt about it. We should continue to scale up. but the behaviors of being really persistent, of being able to backtrack when you're wrong, of, like, understanding how to interact with your environment show us that we can get a lot more out of it. And this, for me, has created a bit of a Question in my mind the last couple of days. If you think about where we're using models today, right? We are using models, say, for knowledge work. Represents twenty-five percent of the global economy, twenty-five trillion dollars of work.Eiso Kant [00:39:00]: As we scale up models and they become more intelligent, we are excited about using them more and more for pushing the frontier of science.Small Models, Knowledge Work, and CommoditizationEiso Kant [00:39:08]: And if you look at the frontier of science, like true breakthroughs in science, they have been linked, they are linked to more intelligence in many places. Einstein figuring out general relativity is able to bring ideas together that other people would have not brought together. And I think one of the many dimensions of intelligence is the ability to do that, and it's something we clearly see that as models get larger and more capable, they're able to pull more ideas and threads together that a smaller model wouldn't be able to.Eiso Kant [00:39:36]: And we're starting to see examples of that in medicine and, like, in bio and other things. But if you think about the majority of knowledge work that we do, and it includes building software. I'm a software developer at heart first and foremost probably, although I probably can't say it that much anymore as I don't write production code in years, is that what makes us good is our persistence. It's our ability to encounter a problem and backtrack and say, “I need to go figure out this bug. I need to go research this. I need to go look at the documentation. I need to, like, try different, five different ways to see, like, if I can solve it.” But it is not necessarily bringing three ideas together from radically different fields. And so if we are now seeing, and I think Laguna S is an example, that we are able to make a relatively small model much more capable than I had definitely predicted or any previous, like, benchmarks had shown for any model remotely this size or even larger, At least on coding tasks, that it's because of the behaviors. And so now the question I have, and I don't have an answer, it is I know at the limit, so infinite model size, right, extremely large model, and the cost of that model is gonna be very expensive to run. We know this, right? So larger model ROI.Eiso Kant [00:40:52]: So I know that at the very limit, I'm not gonna use the world's largest model one day, quadrillion parameter, whatever crazy, like, scale we scale up, to do a basic coding task. Already today, I'm starting to size down for certain tasks.Eiso Kant [00:41:07]: So it means that there is an optimal. It means there's some curve that goes as we go up to model size for knowledge work, at some point we're at the peak, and after that, the return on investment of using a bigger model, just doesn't make sense.Eiso Kant [00:41:22]: Now, I think the question is, before I would have thought that peak was extremely very far away.Eiso Kant [00:41:30]: This model for me is the first sign that Maybe that peak is At a trillion, five trillion, ten trillion. Maybe we can just squeeze way more out of these models. I'm no longer thinking that we need two or three orders of magnitude on the largest models to be able to, solve knowledge work, the accounting, the legal, the code that we write. And so if that holds true, It is an argument for the commoditization of models. It's an argument that open source can win and, like, succeed in this world. And now it's of course a self-serving argument and it's a hopeful argument, but theoretically at the limit it works. We just have to go discover in the next couple of years of how much more we can squeeze out. Now, I do want to put a big asterisk. This does not mean I'm against scaling models. I think we ultimately only succeed if we scale our models as large as our competition. I do not like. I think we should not put our head in the sand and say we're gonna be king of open source small models. I think that's, It's a out. It's trying to be king of your own kingdom, but not realizing what the rest of the world's doing. All of us rather use a smarter, faster, more model. It's a sign of hope. And so I don't wanna overly state this is a good model. We have a long way to go to get to the state-art. But what hopefully people take away when they use this model is that the behaviors inside of it are what push it to be far more capable, less than necessarily the number of parameters.Pre-Training, Mid-Training, and RL Moving EarlierVibhu [00:43:03]: Is that mostly post-training? LikeEiso Kant [00:43:05]: YesVibhu [00:43:05]: Right.Eiso Kant [00:43:06]: It's entirely post-training.Vibhu [00:43:08]: Are we done improving anything on training? Is, like, training done?Eiso Kant [00:43:12]: No.Vibhu [00:43:12]: Okay.Eiso Kant [00:43:13]: SoVibhu [00:43:13]: I just wanted to cover training, and then we go post-trainingEiso Kant [00:43:15]: Training is not done. I mean, look, there's a part of training of just dealing with skill, right? Every new order of magnitude of model skill, you are going to get new things you gotta solve for. That'- but those are ultimately, engineering challenges.Eiso Kant [00:43:31]: I have a, I would say, a not commonly held opinion that reinforcement learning Will move earlier and earlier into training.Vibhu [00:43:42]: Yeah, training.Eiso Kant [00:43:44]: Not even training. Like training today, right, is, like if you look at - So we've been working on this for years already. and I think the best-- I think the first time we saw it out in public was the DeepSeek Zero paper. this is a year and a half ago, I think, if I recall correctly. where, you can Very early on in a model as it starts capable of being able to use language, et cetera, induce reasoning. and so the question that I have is like, we have this- we have the dataset that's the web. and the web, I think we could arguably say probably has The totality of humanity's knowledge somewhere encoded in different places. It's a huge variance degree of quality, from garbage data, and like once you look at training data, you really get humbled of like what the web is, to like, the most greatest scientific papers and best blog posts and like, best transcripts and whatnot.Eiso Kant [00:44:39]: And so now What we are trying to figure out, and have been doing a lot of work on, and it's a place where maybe not as open as we're on other things, but we will become more over time. we've been spending a couple of years really doing research on how can we turn the web into not just next token prediction, but into a way to teach the model to think earlier in its training. and I think there's a huge amount of gold to be found there. I think we are right now in, we've got some drugs in the industry. One of the drugs is distillation. Another drug is, more environments. Like, and they're great, and they make us feel good, and they make the models better, and like we're all addicted to them, and we'll use them, right? in various different ways. and but ultimately, I think we are still barely squeezing out of the web what we should be getting out of the web.Eiso Kant [00:45:33]: I think just next token prediction during training is not enough.Eiso Kant [00:45:36]: AndVibhu [00:45:38]: YeahEiso Kant [00:45:38]: I think we'll see some very interesting things still happen. and that RL in post-training to induce behaviors, to improve things, like I think - the whole world knows how to do this now. I think we're, we're scaling it up. Everyone is. But I wonder if we need to go as far as we're going today with environments. I'm not sure yetVibhu [00:46:01]: You mean we're going too far?Eiso Kant [00:46:02]: I'm, I'm not sure if the path to AGI is justVibhu [00:46:06]: Is more environmentEiso Kant [00:46:07]: More environments.Vibhu [00:46:08]: It seems like a never-ending, “Okay, I want instruction manual for this table, right? Am I gonna environment out building furniture? Or are we just gonna tail end like we need some general solution?”Eiso Kant [00:46:19]: I think there is, I think there's an ability to generalize more from the web. but I also am very encouraged, like when I look at Laguna S and, which is post-training is, well, is the big impact there. and I see like, oh, wait a second, just by making some of these behaviors much better, we're able to get so much more out of it. It just changes a little bit the way you think about intelligence.Vibhu [00:46:40]: Yeah. The analogy people draw often is the RL phase is where you don't learn as much new knowledge. You shiftEiso Kant [00:46:46]: Yeah.Vibhu [00:46:46]: Yeah. So, you shift distribution, and you can have it reason towards what you want. on your point about training, a lot of training is still just continue training in a domain, say medicine, then you do RL. So still justEiso Kant [00:47:00]: It's just better data, right? Like, I mean, training, ooh, I like how we invented this word. Like it's effectively just like,Vibhu [00:47:06]: Second phaseEiso Kant [00:47:07]: It's the second phase of training With like a really dumb way to do a curriculum. But like ultimately, what you'd want is a curriculum from token zero to token 30 whatever or 40 trillion tokens that really truly is the optimal curriculum for the model to learn. But training is essentially a stage curriculum on the web because we do not have to compute, And, effectively to try to ablate the perfect curriculum, right? And so I'm pretty sure that you'll start to see people talking soon about some other term, and there's two or - ‘cause now we do this, right? We talk stage two and stage three and stage four training and like. But ultimately, all we're doing is we're trying to assign a curriculum to the web data that we have to allow the model to learn better. I think at some point, as things get compute, as models get cheaper to run, as the next generations of compute, this will become more of a continuous spectrum. I also think the reason, by the way, you have training and like stage two and stage three is organizational, Right? It'- this is, I think, a thing where-- that we really try to avoid with the model factory is like Training exists because there's a training team now, right? There's people, or like people in training decide to focus on like a training effort. but what you really want is engineering and scale of experiments that allows for a much more continuous spectrum that you don't, you have infinite stages. Now, we're not there. Compute's not there. Organization design is not there for it yet. but I think we'll get there. we'll look back on a couple of years and be like, “Oh my God, it was so cute that we did our training data like this in such a like naïve way. Like we barely ordered it. We didn't really do a good job at likeCurriculum, Auto Research, and New ObjectivesVibhu [00:48:48]: The building that curriculum will get you that in the industry.Eiso Kant [00:48:51]: And I'll confirm that, when I talk to some researchers that this is a lot of the focus now is like how does training change and what is the next objective other than, next token prediction. I assume you don't have the answers, but you have some ideas.Vibhu [00:49:02]: We have some ideas. We're not ready to talk about it yet.Eiso Kant [00:49:05]: Yeah.Vibhu [00:49:05]: We've been working on them for years, and I think that's the one thing that's also like you asked earlier about, like what's not obvious about building a foundation model company is that you are constantly balancing the table stakes work, the recipe worksEiso Kant [00:49:19]: Yeah.Vibhu [00:49:19]: Versus like your, my crazyEiso Kant [00:49:22]: Pure researchVibhu [00:49:22]: Breakthrough.Eiso Kant [00:49:22]: Yeah.Vibhu [00:49:22]: Pure research and finding that balance and adjusting the percentage to it based on where you are in the race is really important.Eiso Kant [00:49:31]: I mean, so like, this is a nice way. I was gonna bring up auto research at some pointVibhu [00:49:35]: YesEiso Kant [00:49:35]: As another Andrej invention, or coinage, which is like, I honestly, like how many objective functions can there be, right? Like just try 1,000 of them, set it running, whatever.Vibhu [00:49:47]: Man, it's alsoEiso Kant [00:49:48]: Like what you're looking for. You're looking for loss curves like that, likeVibhu [00:49:51]: It's also a thing people take bets on, right? When you say more Neo labs, you're doing a version of we'll do foundation models, scale them up, next token predictors. A lot of other Neo labs that we see want to take a completely different approach, right? At some level, you're right. It's all, compute efficiency, and that's the net objective. But some are okay, different architecture, like vastly different amounts of compute spend. So some are different. They're not justEiso Kant [00:50:19]: YeahVibhu [00:50:19]: They're like, 99% not balancing, here's the vanilla and scale up. They're 99% on, here's novel research that'll change everything.Eiso Kant [00:50:27]: And I think, Luke, I think you. It depends when you started as well, right?Pure Research vs. Table StakesVibhu [00:50:30]: Yeah.Eiso Kant [00:50:30]: When we started, like the novel thing we did was reinforcement learning on code. No long- that's no longer novel by far, but we were like, - that's where we obsessed over when no one believed in RL. So you have to when you start the company, you have to have your own idea. You have to have something that's different that allows you to speed up, right? For us, it was RL to LLMs that later became common, like, Knowledge. But in the beginning, it wasn'tVibhu [00:50:53]: It's cool. this was like your original 2023 blogEiso Kant [00:50:57]: YeahVibhu [00:50:57]: Of purpose.Eiso Kant [00:50:58]: Yeah.Vibhu [00:50:59]: And like you do lay it all out here.Eiso Kant [00:51:01]: We laidVibhu [00:51:01]: The blog is pretty underrated, right? The whole RL on code was very early on.Eiso Kant [00:51:06]: Very early. And even we had to argue with people, like we say here things like to push beyond current capability, to train your own foundation model. We had to argue with people that it mattered that you had your own like, base model. you can fine-tune your way to success, right? major capabilities emerge from training a base model made accurate and useful during fine-tuning.Vibhu [00:51:23]: Which like, for perspective at the time, we knew closed models, OpenAI, Anthropic were huge. The open models we had were like Mistral 7B, a 30B, a 70B.Eiso Kant [00:51:35]: When weVibhu [00:51:35]: YeahEiso Kant [00:51:36]: The date on this thing is wrong. When we published this, it was April 2023. I think this was justVibhu [00:51:42]: YeahEiso Kant [00:51:42]: Happened on a migration, probably found it on archive.org.Vibhu [00:51:45]: Mistral.Eiso Kant [00:51:46]: Mistral had started, we started on the same month, right?Vibhu [00:51:49]: Yeah.Eiso Kant [00:51:49]: So this wasn't even, there was only, I think, Llama out at the timeVibhu [00:51:52]: SnellEiso Kant [00:51:52]: And that's it, right? And so, but I agree. I think we wan

Varn Vlog
Solidarity or Silence?: How Leftist Politics Often Overlooks the Disabled with Anthony David Vernon

Varn Vlog

Play Episode Listen Later Jul 2, 2026 89:56 Transcription Available


This podcast and YouTube episode features an in-depth conversation with Anthony David Vernon, a philosopher and educator, exploring the intersection of disability studies, left-wing politics, and the systemic failures of accessibility in a post-pandemic world. The discussion challenges the "normative" framework of society, examining how both civic institutions and political movements often fail to truly incorporate the voices and needs of the disabled community.Key Discussion HighlightsThe ADA and the "Checkmark" Problem: Vernon argues that because the ADA is enforced primarily through personal lawsuits and remains largely unfunded, it often results in "checkmark" compliance rather than true accessibility.The Post-Pandemic Erasure: The conversation explores how the rush to move past COVID-19 safety measures has prioritized "normative desires" over the accessibility needs of high-risk and disabled individuals.Multimodality as Justice: Implementing "ready-made" scaffolding and multiple points of entry into education and digital spaces benefits all learners, not just those with a formal diagnosis. Referenced Works (APA Format)Albers, B. (2022). Able-bodied leftists cannot abandon disabled solidarity to move on from COVID. Truthout. https://truthout.org/articles/abled-bodied-leftists-cannot-abandon-disabled-solidarity-to-move-on-from-covid/Data for Progress. (2023, October 3). Disabled voters do not believe politicians care about disabled Americans. https://www.dataforprogress.org/blog/2023/10/3/disabled-voters-do-not-believe-politicians-care-about-disabled-americansHryhorec, S. (2025, October 26). LET ME IN: Mark Butler's office isn't accessible [Video]. YouTube. https://www.youtube.com/watch?v=DEfbZzCspDkIacoboni, G. (2023). Why politics is failing disabled people and what to do about it. Independent Social Research Foundation (ISRF). https://isrf.org/blog/why-politics-is-failing-disabled-people-and-what-to-do-about-itRotarou, E. S., & Sakellariou, D. (2024). Neoliberalism and disability: The systemic erasure of access. Social Science & Medicine. https://www.sciencedirect.com/science/article/pii/S0277953623007189University of Hawaiʻi at Mānoa. (n.d.). Disability studies and political theory: A framework for inclusion. https://scholarspace.manoa.hawaii.edu/server/api/core/bitstreams/425af050-0220-49dc-b28d-f86d976dcf02/contentVarn, C. D. (2023, October). Multimodal availability for those with learning disabilities. PeerCentered. https://www.peercentered.org/2023/10/multimodal-availability-for-those-with.htmlVernon, A. D. (2023, December). Silence: Non-verbal communication in philosophy. Activated Thinker. https://medium.com/activated-thinker/silence-non-verbal-communication-in-philosophy-d5d148ba8a1dWillies, E. (2025, June 16). Anthony David Vernon advocates for social democracy as a tool of rebellion against fascism [Video]. YouTube. https://www.youtube.com/watch?v=Rs7l6jNGDzwSend us Fan Mail Musis by Bitterlake, Used with Permission, all rights to BitterlakeSupport the showCrew:Host: C. Derick VarnIntro and Outro Music by Bitter Lake.Intro Video Design: Jason MylesArt Design: Corn and C. Derick VarnLinks and Social Media:twitter: @varnvlogblue sky: @varnvlog.bsky.socialYou can find the additional streams on YoutubeCurrent Patreon at the Sponsor Tier: Jordan Sheldon, Mark J. Matthews, Lindsay Kimbrough, RedWolf, DRV, Kenneth McKee, JY Chan, Matthew Monahan, Parzival, Adriel Mixon, Buddy Roark, Daniel Petrovic,Julian, Drea, Free Beer 

The Spark Creativity Teacher Podcast | Education
429: Don't Miss this Essential Lesson on Multimodality

The Spark Creativity Teacher Podcast | Education

Play Episode Listen Later Jun 24, 2026 10:25


Recently I was invited to give a poetry workshop on a reflection day at a local school. They wanted the writing element of the day to help students understand themselves better, so I chose to provide a workshop based on George Ella Lyon's poem, "Where I'm From." You know I love that workshop. Together, we looked at how details bring poetry to life, brainstormed images about their childhood experiences, explored how various creators have interpreted the "I am From" prompt to create videos, paintings, photo essays, poems, and combinations thereof. Then I invited them to work multimodaly as they knit together their images with color and imagery.  But I had never worked with them before, and none of them had heard of multimodal communication, though they're surrounded with it everyday. I realized I had left out a crucial step in the workshop, to help them see that multimodal communication would go well beyond "decorating" their poem or underlining all the lines in color. So how can we introduce this concept to students? How can we help them see that text, images, audio, and video can all convey such very different shades of meaning in communication? This week on the pod, let's talk about introducing multimodality, and showing kids what works and what doesn't. Be sure to grab the free download that goes along with this episode, a slideshow full of examples you can share with your students. You can sign up to have me send it over totally free right here. You'll also be subscribed to my teaching idea emails, though of course you can unsubscribe at any time. OK, let's dive in. Grab your copy of the multimodality introduction slideshow: https://spark-creativity.kit.com/bdde614049 Go Further:  Explore alllll the Episodes of The Spark Creativity Teacher Podcast. Snag three free weeks of community-building attendance question slides Join our community, Creative High School English, on Facebook. Come hang out on Instagram.  Enjoying the podcast? Please consider sharing it with a friend, snagging a screenshot to share on the 'gram, or tapping those ⭐⭐⭐⭐⭐ to help others discover the show. Thank you! 

Crazy Wisdom
Episode #548: The Pixel Path: From Perception to Action, and the Future of Intelligent Robots with Nizar

Crazy Wisdom

Play Episode Listen Later May 25, 2026 56:19


Stewart Alsop interviews Nizar, CEO of Pixel Robotics, on the Crazy Wisdom Podcast to explore the intersection of AI, robotics, and perception. The conversation covers a wide range of technical topics including how transformers enable multimodal representation across text, images, and voice, the role of world models in predicting physical interactions, the advantages of diffusion models over traditional LLMs for certain applications, and the challenges of achieving real-time processing for robotics applications. Nizar explains Pixel Robotics' work on creating accurate 3D meshes from smartphone cameras for companies like L'Oréal, moving away from specialized sensors to make the technology more accessible through sophisticated algorithms, and discusses the future of robotics as closing the perception-action loop to enable robots to perform real tasks beyond simple demonstrations. To find out more visit Pixel Robotics' website.Timestamps00:00 Stewart welcomes Nizar, CEO of Pixel Robotics, discussing what a pixel is as the smallest visual unit on screens composed of red green and blue colors05:00 Discussion of perception systems and how logarithmic laws help compress signals in both human and artificial systems, exploring normalization layers and sigmoid functions in deep learning10:00 Exploring how transformers unified different data modalities including text voice and images, creating common representations through methods like contrastive learning15:00 Nizar explains transformers as brute force learning systems with room for improvement through focused attention mechanisms and knowledge graphs rather than processing everything20:00 Conversation about loss functions local minima versus global minima and how mixture of experts uses specialized small models instead of one massive generalist network25:00 Discussion of deterministic versus probabilistic systems and how explicitly defined task graphs often outperform orchestrator-based approaches in AI systems30:00 Exploring world models as predictive physics-based systems that learn environmental flows and transformations, complementing rather than replacing language models35:00 Nizar discusses real-time processing challenges for robotics requiring millisecond responses with small memory footprints using vision transformers for faster experimentation40:00 Pixel's work creating three d meshes from smartphone cameras for companies like L'Oreal, moving away from specialized sensors toward accessible software-based solutions45:00 Explanation of different three d representations including voxels point clouds and meshes, with meshes being optimal for manipulation and rendering in applications50:00 Future direction involves closing perception-action loops in robotics, moving beyond dancing toy robots toward practical multimodal systems that perform real tasks55:00 Pixel's goal is democratizing high-quality three d scanning through smartphones, making mesh creation accessible to unlock applications in gaming cinema and virtual showroomsKey Insights1. Pixel Robotics derives its name from combining perception and action in robotics, where the pixel represents the digital perception component and robotics represents the physical action component. The pixel serves as a metaphor for how robots must quantize and digitize continuous analog information from the real world into discrete units that computer systems can process, similar to how pixels are the fundamental building blocks of images on a screen. This quantization process is essential because numerical systems cannot work with truly continuous data and must convert reality into tractable digital representations that algorithms can manipulate.2. The transformer architecture has created a fundamental unification in how different types of data can be represented and processed across multiple modalities. Before transformers, researchers working on natural language processing, computer vision, and audio analysis used completely different approaches and methodologies. The breakthrough of transformers was establishing a common representational framework that could handle text, images, voice, and other data types using similar underlying mechanisms. This unification is what enabled the development of truly multimodal AI systems and represents one of the most significant advances beyond just the language modeling capabilities that initially gained public attention.3. Current transformer-based systems represent a brute force approach to learning that will likely be superseded or enhanced by more efficient algorithms. Despite claims that we have exhausted internet text data for training, significant improvements continue to emerge every few months through algorithmic innovations rather than simply adding more data. Future developments will likely involve more specialized attention mechanisms that focus on relevant information rather than correlating everything with everything, mixture of experts architectures with small specialized models, and approaches inspired by biological systems such as logarithmic compression laws and event-based processing that humans use naturally.4. Diffusion-based language models represent a promising alternative to standard next-token prediction that could produce more accurate outputs through an iterative refinement process. Unlike traditional language models that predict one token at a time and cannot revise earlier outputs, diffusion models treat text generation like image denoising, starting with a noisy representation and progressively refining the entire output across multiple steps. This holistic approach allows the model to reconsider and improve all parts of the response simultaneously, potentially leading to higher quality results, though it may be slower than current autoregressive methods. This represents an important direction for overcoming fundamental limitations in how language models currently generate text.5. For robotics applications, real-time performance and small model size are critical constraints that differ significantly from the requirements of large language models deployed in data centers. Vision transformers are being used as a testbed for developing efficient real-time algorithms because they require far fewer computational resources to train and test compared to large language models, making them more practical for rapid experimentation. The goal is to achieve millisecond-level response times with minimal memory footprint so that robots can react quickly to dynamic environments and run on affordable hardware that can be embedded in actual robotic systems rather than requiring expensive server infrastructure.6. Practical robotics implementation requires moving beyond specialized sensors to software solutions that work with ubiquitous devices like smartphones for tasks such as three-dimensional reconstruction. Pixel Robotics evolved from building specialized scanning hardware to focusing on algorithms that can generate high-quality mesh representations of environments using only smartphone cameras, making the technology far more accessible and practical for real-world deployment. This approach enables applications ranging from industrial robotic arm control to virtual showrooms, and more importantly, it allows anyone to capture three-dimensional data without expensive equipment, which can also help generate larger training datasets for future AI development.7. The next frontier in AI and robotics is closing the perception-action loop to enable robots to perform real practical tasks rather than remaining as demonstration systems or toys. While significant progress has been made in cognitive capabilities through language models and in robotic mobility through mechanical engineering advances, the critical challenge is integrating perception with action through systems like Vision-Language-Action models. The fundamental starting point for learning this integration is simple perception-action exercises, such as programming a camera mounted on servo motors to track and center a colored object, which demonstrates the basic principle of using sensory input to drive physical response that underlies all more sophisticated robotic behaviors.

Disruption / Interruption
Disrupting MedTech: Turning the Smartphone into a Life-Saving Medical Device with Gennadi Saiko

Disruption / Interruption

Play Episode Listen Later Feb 19, 2026 30:10


In this episode of Disruption/Interruption, KJ sits down with Gennadi Seko, founder and CEO of Oxilight, who is revolutionizing wound care diagnostics by transforming smartphones into powerful medical imaging devices. Gennadi shares his personal journey from Bay Street finance to medical physics, driven by his grandmother's diabetic foot amputation. He discusses how his company is disrupting the medical device industry by making diagnostic technology portable, affordable, and accessible—moving critical wound care assessments from expensive hospital labs to patients' homes. This conversation explores the intersection of deep tech innovation, healthcare accessibility, and the power of multimodal diagnostics in saving lives and limbs. Four Key Takeaways [26:19] Multimodality is the Game Changer - Instead of multiple expensive single-purpose devices sitting on shelves, combining three technologies (multispectral imaging, fluorescence imaging, and thermal imaging) into one $200 smartphone attachment provides a 360-degree view of wound health and dramatically improves diagnostic specificity. [9:29] The Diabetes Crisis is Escalating - 27% of seniors (65+) in the United States have diabetes, and the disease is now affecting people as young as 25. Diabetic foot complications account for 80% of all non-traumatic amputations, making early detection critical. [21:44] Mobility Saves Lives and Money - Moving diagnostic technology to patients' homes solves the compliance problem and enables early intervention. Preventing one amputation saves healthcare systems 10x in costs while dramatically improving patient quality of life. [14:50] Physiological Imaging Beats Anatomical Measurement - Traditional wound measurement with rulers only tracks size over time, requiring multiple visits. Physiological imaging provides immediate prognostic information from a single snapshot, identifying whether a wound will heal normally or requires intervention. Quote of the Show (23:27): “I don't want to improve hospital healthcare. I want to improve healthcare in general." - Gennadi Seko Join our Anti-PR newsletter where we’re keeping a watchful and clever eye on PR trends, PR fails, and interesting news in tech so you don't have to. You're welcome. Want PR that actually matters? Get 30 minutes of expert advice in a fast-paced, zero-nonsense session from Karla Jo Helms, a veteran Crisis PR and Anti-PR Strategist who knows how to tell your story in the best possible light and get the exposure you need to disrupt your industry. Click here to book your call: https://info.jotopr.com/free-anti-pr-eval Ways to connect with Gennadi Seko: LinkedIn: https://www.linkedin.com/in/gennadisaiko/Company Website: https://oxilight.ca How to get more Disruption/Interruption: Amazon Music - https://music.amazon.com/podcasts/eccda84d-4d5b-4c52-ba54-7fd8af3cbe87/disruption-interruption Apple Podcast - https://podcasts.apple.com/us/podcast/disruption-interruption/id1581985755 Spotify - https://open.spotify.com/show/6yGSwcSp8J354awJkCmJlDSee omnystudio.com/listener for privacy information.

ReachMD CME
From Pixels to Practice: Advancing HCM Care With Multimodality Imaging

ReachMD CME

Play Episode Listen Later Jan 28, 2026 43:15


CME credits: 0.75 Valid until: 28-01-2027 Claim your CME credit at https://reachmd.com/programs/cme/from-pixels-to-practice-advancing-hcm-care-with-multimodality-imaging/39877/ This video series focuses on translating multimodality imaging into practical care for patients with hypertrophic cardiomyopathy (HCM). Expert faculty review echocardiography, cardiac magnetic resonance, and emerging imaging strategies to support diagnosis, guide HCM-specific therapies, including the use of cardiac myosin inhibitors, and monitor treatment response. Case-based discussions highlight imaging patterns that inform prognosis and optimize patient outcomes.=

PeerVoice Clinical Pharmacology Audio
Erwan Donal, MD, PhD / Nina Ajmone Marsan, MD, PhD - At the Cutting Edge of ATTR-CM: How Can We Leverage Advances in Multimodality Cardiac Imaging and Artificial Intelligence to Modernise Diagnosis and Monitoring?

PeerVoice Clinical Pharmacology Audio

Play Episode Listen Later Dec 16, 2025 29:14


Erwan Donal, MD, PhD / Nina Ajmone Marsan, MD, PhD - At the Cutting Edge of ATTR-CM: How Can We Leverage Advances in Multimodality Cardiac Imaging and Artificial Intelligence to Modernise Diagnosis and Monitoring?

PeerVoice Heart & Lung Audio
Erwan Donal, MD, PhD / Nina Ajmone Marsan, MD, PhD - At the Cutting Edge of ATTR-CM: How Can We Leverage Advances in Multimodality Cardiac Imaging and Artificial Intelligence to Modernise Diagnosis and Monitoring?

PeerVoice Heart & Lung Audio

Play Episode Listen Later Dec 16, 2025 29:14


Erwan Donal, MD, PhD / Nina Ajmone Marsan, MD, PhD - At the Cutting Edge of ATTR-CM: How Can We Leverage Advances in Multimodality Cardiac Imaging and Artificial Intelligence to Modernise Diagnosis and Monitoring?

PeerVoice Internal Medicine Audio
Erwan Donal, MD, PhD / Nina Ajmone Marsan, MD, PhD - At the Cutting Edge of ATTR-CM: How Can We Leverage Advances in Multimodality Cardiac Imaging and Artificial Intelligence to Modernise Diagnosis and Monitoring?

PeerVoice Internal Medicine Audio

Play Episode Listen Later Dec 16, 2025 29:14


Erwan Donal, MD, PhD / Nina Ajmone Marsan, MD, PhD - At the Cutting Edge of ATTR-CM: How Can We Leverage Advances in Multimodality Cardiac Imaging and Artificial Intelligence to Modernise Diagnosis and Monitoring?

PeerVoice Clinical Pharmacology Video
Erwan Donal, MD, PhD / Nina Ajmone Marsan, MD, PhD - At the Cutting Edge of ATTR-CM: How Can We Leverage Advances in Multimodality Cardiac Imaging and Artificial Intelligence to Modernise Diagnosis and Monitoring?

PeerVoice Clinical Pharmacology Video

Play Episode Listen Later Dec 16, 2025 29:03


Erwan Donal, MD, PhD / Nina Ajmone Marsan, MD, PhD - At the Cutting Edge of ATTR-CM: How Can We Leverage Advances in Multimodality Cardiac Imaging and Artificial Intelligence to Modernise Diagnosis and Monitoring?

PeerVoice Heart & Lung Video
Erwan Donal, MD, PhD / Nina Ajmone Marsan, MD, PhD - At the Cutting Edge of ATTR-CM: How Can We Leverage Advances in Multimodality Cardiac Imaging and Artificial Intelligence to Modernise Diagnosis and Monitoring?

PeerVoice Heart & Lung Video

Play Episode Listen Later Dec 16, 2025 29:03


Erwan Donal, MD, PhD / Nina Ajmone Marsan, MD, PhD - At the Cutting Edge of ATTR-CM: How Can We Leverage Advances in Multimodality Cardiac Imaging and Artificial Intelligence to Modernise Diagnosis and Monitoring?

PeerVoice Internal Medicine Video
Erwan Donal, MD, PhD / Nina Ajmone Marsan, MD, PhD - At the Cutting Edge of ATTR-CM: How Can We Leverage Advances in Multimodality Cardiac Imaging and Artificial Intelligence to Modernise Diagnosis and Monitoring?

PeerVoice Internal Medicine Video

Play Episode Listen Later Dec 16, 2025 29:03


Erwan Donal, MD, PhD / Nina Ajmone Marsan, MD, PhD - At the Cutting Edge of ATTR-CM: How Can We Leverage Advances in Multimodality Cardiac Imaging and Artificial Intelligence to Modernise Diagnosis and Monitoring?

Latent Space: The AI Engineer Podcast — CodeGen, Agents, Computer Vision, Data Science, AI UX and all things Software 3.0
After LLMs: Spatial Intelligence and World Models — Fei-Fei Li & Justin Johnson, World Labs

Latent Space: The AI Engineer Podcast — CodeGen, Agents, Computer Vision, Data Science, AI UX and all things Software 3.0

Play Episode Listen Later Nov 25, 2025 60:38


Fei-Fei Li and Justin Johnson are cofounders of World Labs, who have recently launched Marble (https://marble.worldlabs.ai/), a new kind of generative “world model” that can create editable 3D environments from text, images, and other spatial inputs. Marble lets creators generate persistent 3D worlds, precisely control cameras, and interactively edit scenes, making it a powerful tool for games, film, VR, robotics simulation, and more. In this episode, Fei-Fei and Justin share how their journey from ImageNet and Stanford research led to World Labs, why spatial intelligence is the next frontier after LLMs, and how world models could change how machines see, understand, and build in 3D.We discuss:* The massive compute scaling from AlexNet to today and why world models and spatial data are the most compelling way to “soak up” modern GPU clusters compared to language alone.* What Marble actually is: a generative model of 3D worlds that turns text and images into editable scenes using Gaussian splats, supports precise camera control and recording, and runs interactively on phones, laptops, and VR headsets.* Fei-fei's essay:on spatial intelligence as a distinct form of intelligence from language: from picking up a mug to inferring the 3D structure of DNA, and why language is a lossy, low-bandwidth channel for describing the rich 3D/4D world we live in.* Whether current models “understand” physics or just fit patterns: the gap between predicting orbits and discovering F=ma, and how attaching physical properties to splats and distilling physics engines into neural networks could lead to genuine causal reasoning.* The changing role of academia in AI, why Fei-Fei worries more about under-resourced universities than “open vs closed,” and how initiatives like national AI compute clouds and open benchmarks can rebalance the ecosystem.* Why transformers are fundamentally set models, not sequence models, and how that perspective opens up new architectures for world models, especially as hardware shifts from single GPUs to massive distributed clusters.* Real use cases for Marble today: previsualization and VFX, game environments, virtual production, interior and architectural design (including kitchen remodels), and generating synthetic simulation worlds for training embodied agents and robots.* How spatial intelligence and language intelligence will work together in multimodal systems, and why the goal isn't to throw away LLMs but to complement them with rich, embodied models of the world.* Fei-Fei and Justin's long-term vision for spatial intelligence: from creative tools for artists and game devs to broader applications in science, medicine, and real-world decision-making.—Fei-Fei Li* X: https://x.com/drfeifei* LinkedIn: https://www.linkedin.com/in/fei-fei-li-4541247Justin Johnson* X: https://x.com/jcjohnss* LinkedIn: https://www.linkedin.com/in/justin-johnson-41b43664Where to find Latent Space* X: https://x.com/latentspacepodFull Video EpisodeTimestamps00:00:00 Introduction and the Fei-Fei Li & Justin Johnson Partnership00:02:00 From ImageNet to World Models: The Evolution of Computer Vision00:12:42 Dense Captioning and Early Vision-Language Work00:19:57 Spatial Intelligence: Beyond Language Models00:28:46 Introducing Marble: World Labs' First Spatial Intelligence Model00:33:21 Gaussian Splats and the Technical Architecture of Marble00:22:10 Physics, Dynamics, and the Future of World Models00:41:09 Multimodality and the Interplay of Language and Space00:37:37 Use Cases: From Creative Industries to Robotics and Embodied AI00:56:58 Hiring, Research Directions, and the Future of World Labs Get full access to Latent.Space at www.latent.space/subscribe

Latent Space: The AI Engineer Podcast — CodeGen, Agents, Computer Vision, Data Science, AI UX and all things Software 3.0
After LLMs: Spatial Intelligence and World Models — Fei-Fei Li & Justin Johnson, World Labs

Latent Space: The AI Engineer Podcast — CodeGen, Agents, Computer Vision, Data Science, AI UX and all things Software 3.0

Play Episode Listen Later Nov 25, 2025


Fei-Fei Li and Justin Johnson are cofounders of World Labs, who have recently launched Marble (https://marble.worldlabs.ai/), a new kind of generative “world model” that can create editable 3D environments from text, images, and other spatial inputs. Marble lets creators generate persistent 3D worlds, precisely control cameras, and interactively edit scenes, making it a powerful tool for games, film, VR, robotics simulation, and more. In this episode, Fei-Fei and Justin share how their journey from ImageNet and Stanford research led to World Labs, why spatial intelligence is the next frontier after LLMs, and how world models could change how machines see, understand, and build in 3D. We discuss: The massive compute scaling from AlexNet to today and why world models and spatial data are the most compelling way to “soak up” modern GPU clusters compared to language alone. What Marble actually is: a generative model of 3D worlds that turns text and images into editable scenes using Gaussian splats, supports precise camera control and recording, and runs interactively on phones, laptops, and VR headsets. Fei-fei's essay (https://drfeifei.substack.com/p/from-words-to-worlds-spatial-intelligence) on spatial intelligence as a distinct form of intelligence from language: from picking up a mug to inferring the 3D structure of DNA, and why language is a lossy, low-bandwidth channel for describing the rich 3D/4D world we live in. Whether current models “understand” physics or just fit patterns: the gap between predicting orbits and discovering F=ma, and how attaching physical properties to splats and distilling physics engines into neural networks could lead to genuine causal reasoning. The changing role of academia in AI, why Fei-Fei worries more about under-resourced universities than “open vs closed,” and how initiatives like national AI compute clouds and open benchmarks can rebalance the ecosystem. Why transformers are fundamentally set models, not sequence models, and how that perspective opens up new architectures for world models, especially as hardware shifts from single GPUs to massive distributed clusters. Real use cases for Marble today: previsualization and VFX, game environments, virtual production, interior and architectural design (including kitchen remodels), and generating synthetic simulation worlds for training embodied agents and robots. How spatial intelligence and language intelligence will work together in multimodal systems, and why the goal isn't to throw away LLMs but to complement them with rich, embodied models of the world. Fei-Fei and Justin's long-term vision for spatial intelligence: from creative tools for artists and game devs to broader applications in science, medicine, and real-world decision-making. — Fei-Fei Li X: https://x.com/drfeifei LinkedIn: https://www.linkedin.com/in/fei-fei-li-4541247 Justin Johnson X: https://x.com/jcjohnss LinkedIn: https://www.linkedin.com/in/justin-johnson-41b43664 Where to find Latent Space X: https://x.com/latentspacepod Substack: https://www.latent.space/ Chapters 00:00:00 Introduction and the Fei-Fei Li & Justin Johnson Partnership 00:02:00 From ImageNet to World Models: The Evolution of Computer Vision 00:12:42 Dense Captioning and Early Vision-Language Work 00:19:57 Spatial Intelligence: Beyond Language Models 00:28:46 Introducing Marble: World Labs' First Spatial Intelligence Model 00:33:21 Gaussian Splats and the Technical Architecture of Marble 00:22:10 Physics, Dynamics, and the Future of World Models 00:41:09 Multimodality and the Interplay of Language and Space 00:37:37 Use Cases: From Creative Industries to Robotics and Embodied AI 00:56:58 Hiring, Research Directions, and the Future of World Labs

Crazy Wisdom
Episode #506: How AI Turns Podcasts into Knowledge Engines

Crazy Wisdom

Play Episode Listen Later Nov 14, 2025 49:38


In this episode of Crazy Wisdom, host Stewart Alsop talks with Kevin Smith, co-founder of Snipd, about how AI is reshaping the way we listen, learn, and interact with podcasts. They explore Snipd's vision of transforming podcasts into living knowledge systems, the evolution of machine learning from finance to large language models, and the broader connection between AI, robotics, and energy as the foundation for the next technological era. Kevin also touches on ideas like the bitter lesson, reinforcement learning, and the growing energy demands of AI. Listeners can try Snipd's premium version free for a month using this promo link.Check out this GPT we trained on the conversationTimestamps00:00 – Stewart Alsop welcomes Kevin Smith, co-founder of Snipd, to discuss AI, podcasting, and curiosity-driven learning.05:00 – Kevin explains Snipd's snipping feature, chatting with episodes, and future plans for voice interaction with podcasts.10:00 – They discuss vector search, embeddings, and context windows, comparing full-episode context to chunked transcripts.15:00 – Kevin shares his background in mathematics and economics, his shift from finance to machine learning, and early startup work in AI.20:00 – They explore early quant models versus modern machine learning, statistical modeling, and data limitations in finance.25:00 – Conversation turns to transformer models, pretraining, and the bitter lesson—how compute-based methods outperform human-crafted systems. 30:00 – Stewart connects this to RLHF, Scale AI, and data scarcity; Kevin reflects on reinforcement learning's future. 35:00 – They pivot to Snipd's podcast ecosystem, hidden gems like Founders Podcast, and how stories shape entrepreneurial insight. 40:00 – ETH Zurich, robotics, and startup culture come up, linking academia to real-world innovation. 45:00 – They close on AI, robotics, and energy as the pillars of the future, debating nuclear and solar power's role in sustaining progress.Key InsightsPodcasts as dynamic knowledge systems: Kevin Smith presents Snipd as an AI-powered tool that transforms podcasts into interactive learning environments. By allowing listeners to “snip” and summarize meaningful moments, Snipd turns passive listening into active knowledge management—bridging curiosity, memory, and technology in a way that reframes podcasts as living knowledge capsules rather than static media.AI transforming how we engage with information: The discussion highlights how AI enables entirely new modes of interaction—chatting directly with podcast episodes, asking follow-up questions, and contextualizing information across an author's full body of work. This evolution points toward a future where knowledge consumption becomes conversational and personalized rather than linear and one-size-fits-all.Vectorization and context windows matter: Kevin explains that Snipd currently avoids heavy use of vector databases, opting instead to feed entire episodes into large models. This choice enhances coherence and comprehension, reflecting how advances in context windows have reshaped how AI understands complex audio content.Machine learning's roots in finance shaped early AI thinking: Kevin's journey from quantitative finance to AI reveals how statistical modeling laid the groundwork for modern learning systems. While finance once relied on rigid, theory-based models, the machine learning paradigm replaced those priors with flexible, data-driven discovery—an essential philosophical shift in how intelligence is approached.The Bitter Lesson and the rise of compute: Together they unpack Richard Sutton's “bitter lesson”—the idea that methods leveraging computation and data inevitably surpass those built from human intuition. This insight serves as a compass for understanding why transformers, pretraining, and scaling have driven recent AI breakthroughs.Reinforcement learning and data scarcity define AI's next phase: Stewart links RLHF and the work of companies like Scale AI and Surge AI to the broader question of data limits. Kevin agrees that the next wave of AI will depend on reinforcement learning and simulated environments that generate new, high-quality data beyond what humans can label.The future hinges on AI, robotics, and energy: Kevin closes with a framework for the next decade: AI provides intelligence, robotics applies it to the physical world, and energy sustains it all. He warns that society must shift from fearing energy use to innovating in production—especially through nuclear and solar power—to meet the demands of an increasingly intelligent, interconnected world.

The Plastic Surgery Revolution
The Future of Aesthetics: The Power of Multimodality (With Taylor Farley, PA-C)

The Plastic Surgery Revolution

Play Episode Listen Later Nov 11, 2025 12:52


In this episode of The Plastic Surgery Revolution, Dr. Steven Davis sits down with Taylor Farley, PA-C to unpack key insights from the 2025 Global Aesthetics Conference in Miami — one of the industry's biggest gatherings for surgical and non-surgical innovation. From the rise of regenerative medicine and peptide therapy to the evolving “multimodality” approach that combines injectables, lasers, and biostimulatory treatments, Dr. Davis and Taylor explore how today's most natural, long-lasting results come from blending science, sequencing, and patient health. They also dive into: How peptides and GLP-1 medications are changing the conversation around functional health and cosmetic outcomes The best timing and sequencing for injectables, radiofrequency, and laser treatments Why collagen regeneration starts with nutrition, protein, and overall wellness The importance of “low and slow” when using GLP-1s like Ozempic and Wegovy How at-home skin care and consistency enhance in-office treatments The takeaway? There's no single magic bullet in modern aesthetics — real transformation comes from a balanced, regenerative approach that treats the whole patient, inside and out. Listen to The Plastic Surgery Revolution wherever you get your podcasts and stay ahead of what's next in beauty, health, and science.

New Books Network
Birgit Abels and Patrick Eisenlohr, "Atmospheric Knowledge: Environmentality, Latency, and Sonic Multimodality" (U California Press, 2025)

New Books Network

Play Episode Listen Later Nov 7, 2025 46:45


How do we know through atmospheres? How can being affected by an atmosphere give rise to knowledge? What role does somatic, nonverbal knowledge play in how we belong to places? Atmospheric Knowledge takes up these questions through detailed analyses of practices that generate atmospheres and in which knowledge emerges through visceral intermingling with atmospheres. From combined musicological and anthropological perspectives, Birgit Abels and Patrick Eisenlohr investigate atmospheres as a compelling alternative to better-known analytics of affect by way of performative and sonic practices across a range of ethnographic settings. With particular focus on oceanic relations and sonic affectedness, Atmospheric Knowledge centers the rich affordances of sonic connections for knowing our environments. A free ebook version of this title is available through Luminos, University of California Press's Open Access publishing program. Visit www.luminosoa.org to learn more. Learn more about your ad choices. Visit megaphone.fm/adchoices Support our show by becoming a premium member! https://newbooksnetwork.supportingcast.fm/new-books-network

New Books in Anthropology
Birgit Abels and Patrick Eisenlohr, "Atmospheric Knowledge: Environmentality, Latency, and Sonic Multimodality" (U California Press, 2025)

New Books in Anthropology

Play Episode Listen Later Nov 7, 2025 46:45


How do we know through atmospheres? How can being affected by an atmosphere give rise to knowledge? What role does somatic, nonverbal knowledge play in how we belong to places? Atmospheric Knowledge takes up these questions through detailed analyses of practices that generate atmospheres and in which knowledge emerges through visceral intermingling with atmospheres. From combined musicological and anthropological perspectives, Birgit Abels and Patrick Eisenlohr investigate atmospheres as a compelling alternative to better-known analytics of affect by way of performative and sonic practices across a range of ethnographic settings. With particular focus on oceanic relations and sonic affectedness, Atmospheric Knowledge centers the rich affordances of sonic connections for knowing our environments. A free ebook version of this title is available through Luminos, University of California Press's Open Access publishing program. Visit www.luminosoa.org to learn more. Learn more about your ad choices. Visit megaphone.fm/adchoices Support our show by becoming a premium member! https://newbooksnetwork.supportingcast.fm/anthropology

New Books in Sociology
Birgit Abels and Patrick Eisenlohr, "Atmospheric Knowledge: Environmentality, Latency, and Sonic Multimodality" (U California Press, 2025)

New Books in Sociology

Play Episode Listen Later Nov 7, 2025 46:45


How do we know through atmospheres? How can being affected by an atmosphere give rise to knowledge? What role does somatic, nonverbal knowledge play in how we belong to places? Atmospheric Knowledge takes up these questions through detailed analyses of practices that generate atmospheres and in which knowledge emerges through visceral intermingling with atmospheres. From combined musicological and anthropological perspectives, Birgit Abels and Patrick Eisenlohr investigate atmospheres as a compelling alternative to better-known analytics of affect by way of performative and sonic practices across a range of ethnographic settings. With particular focus on oceanic relations and sonic affectedness, Atmospheric Knowledge centers the rich affordances of sonic connections for knowing our environments. A free ebook version of this title is available through Luminos, University of California Press's Open Access publishing program. Visit www.luminosoa.org to learn more. Learn more about your ad choices. Visit megaphone.fm/adchoices Support our show by becoming a premium member! https://newbooksnetwork.supportingcast.fm/sociology

New Books in Geography
Birgit Abels and Patrick Eisenlohr, "Atmospheric Knowledge: Environmentality, Latency, and Sonic Multimodality" (U California Press, 2025)

New Books in Geography

Play Episode Listen Later Nov 7, 2025 46:45


How do we know through atmospheres? How can being affected by an atmosphere give rise to knowledge? What role does somatic, nonverbal knowledge play in how we belong to places? Atmospheric Knowledge takes up these questions through detailed analyses of practices that generate atmospheres and in which knowledge emerges through visceral intermingling with atmospheres. From combined musicological and anthropological perspectives, Birgit Abels and Patrick Eisenlohr investigate atmospheres as a compelling alternative to better-known analytics of affect by way of performative and sonic practices across a range of ethnographic settings. With particular focus on oceanic relations and sonic affectedness, Atmospheric Knowledge centers the rich affordances of sonic connections for knowing our environments. A free ebook version of this title is available through Luminos, University of California Press's Open Access publishing program. Visit www.luminosoa.org to learn more. Learn more about your ad choices. Visit megaphone.fm/adchoices Support our show by becoming a premium member! https://newbooksnetwork.supportingcast.fm/geography

New Books in Sound Studies
Birgit Abels and Patrick Eisenlohr, "Atmospheric Knowledge: Environmentality, Latency, and Sonic Multimodality" (U California Press, 2025)

New Books in Sound Studies

Play Episode Listen Later Nov 7, 2025 46:45


How do we know through atmospheres? How can being affected by an atmosphere give rise to knowledge? What role does somatic, nonverbal knowledge play in how we belong to places? Atmospheric Knowledge takes up these questions through detailed analyses of practices that generate atmospheres and in which knowledge emerges through visceral intermingling with atmospheres. From combined musicological and anthropological perspectives, Birgit Abels and Patrick Eisenlohr investigate atmospheres as a compelling alternative to better-known analytics of affect by way of performative and sonic practices across a range of ethnographic settings. With particular focus on oceanic relations and sonic affectedness, Atmospheric Knowledge centers the rich affordances of sonic connections for knowing our environments. A free ebook version of this title is available through Luminos, University of California Press's Open Access publishing program. Visit www.luminosoa.org to learn more. Learn more about your ad choices. Visit megaphone.fm/adchoices Support our show by becoming a premium member! https://newbooksnetwork.supportingcast.fm/sound-studies

Alt Goes Mainstream
Hg's Chris Kindt - AI's transformative role in value creation for private equity

Alt Goes Mainstream

Play Episode Listen Later Oct 1, 2025 57:00


Welcome back to the Alt Goes Mainstream podcast.Today's episode provides a fascinating window into the world of private equity value creation and how AI can help transform both portfolio company and investment firm processes and operations to create value.We sat down in New York with Hg Partner and Head of Value Creation Chris Kindt to dive into how he spearheaded the growth of the firm's in-house value creation efforts and has built a high-performing team to meet the evolving needs of portfolio companies.Chris brings 15 years of experience in architecting and driving value creation to bear, with 11 years of experience at Hg and another 8 years as a consultant at BCG and Parthenon.Chris and his team are responsible for driving value across a portfolio of 55 companies that represents over $180B in enterprise value and would represent the second largest software company in Europe after SAP if it were a conglomerate.Chris and I had a thought-provoking discussion about value creation and the impact of AI on investing and operating companies. We discussed:How to build a high-performing value creation team.Where value creation has the biggest impact.How AI is transforming investment processes.How Hg has become an “AI-first” investment firm.Where AI will have the most impact in a company today and why agentic AI is transforming certain roles and processes within companies.How professionals can generate the most leverage from utilizing AI – without having cognitive decline.Why prompting and prompt engineering is so critical in the age of AI (and how you can use your weekends to perfect prompt engineering).How a management team can buy into AI and why the current time period represents an interesting opportunity for incumbents, particularly for those building mission-critical enterprise software.Thanks Chris for sharing your expertise and wisdom on company building and AI. We hope you enjoy.Note: this episode was filmed in August 2025.A word from AGM podcast sponsor, Ultimus Fund SolutionsThis episode of Alt Goes Mainstream is brought to you by Ultimus Fund Solutions, a leading full-service fund administrator for asset managers in private and public markets. As private markets continue to move into the mainstream, the industry requires infrastructure solutions that help funds and investors keep pace. In an increasingly sophisticated financial marketplace, investment managers must navigate a growing array of challenges: elaborate fund structures, specialized strategies, evolving compliance requirements, a growing need for sophisticated reporting, and intensifying demands for transparency.To assist with these challenging opportunities, more and more fund sponsors and asset managers are turning to Ultimus, a leading service provider that blends high tech and high touch in unique and customized fund administration and middle office solutions for a diverse and growing universe of over 450 clients and 1,800 funds, representing $500 billion assets under administration, all handled by a team of over 1,000 professionals. Ultimus offers a wide range of capabilities across registered funds, private funds and public plans, as well as outsourced middle office services. Delivering operational excellence, Ultimus helps firms manage the ever-changing regulatory environment while meeting the needs of their institutional and retail investors. Ultimus provides comprehensive operational support and fund governance services to help managers successfully launch retail alternative products.Visit www.ultimusfundsolutions.com to learn more about Ultimus' technology enhanced services and solutions or contact Ultimus Executive Vice President of Business Development Gary Harris on email at gharris@ultimusfundsolutions.com.We thank Ultimus for their support of alts going mainstream.Show Notes00:00 Message from our Sponsor, Ultimus01:18 Welcome to the Alt Goes Mainstream Podcast02:10 Guest Introduction: Chris Kindt03:58 Chris Kindt's Background and Journey05:43 Value Creation at Hg06:55 Pre-Investment and Diligence Process07:44 Management Team Dynamics08:40 Common Value Creation Interventions11:14 Building a Cross-Functional Team12:31 Scaling the Value Creation Team15:40 Measuring Value Creation Impact17:49 Investment Philosophy: Inch Wide, Mile Deep19:28 Partnerships with Management Teams20:03 Embedding AI in Value Creation20:24 Internal vs External AI Applications21:47 AI First Culture22:18 Effective AI Utilization24:04 Prompt Engineering and Whispering26:25 Choosing the Right AI Models26:50 AI Models: Strengths and Weaknesses27:42 Transformative Impact of AI28:21 Skills Needed in the AI Era28:26 AI's Role in Investment Firms28:55 Core Insights and Judgements29:04 The Core Skillset and Efficiency29:12 Philosophical Questions on AI and Talent Development30:18 Building Grit in the Age of AI31:12 Maintaining Discipline with AI32:08 AI as a Value Creation Lever32:35 Operational Efficiency and Copilots33:43 Emergence of Reasoning Models and Agentic Frameworks34:31 10x Efficiency in Engineering36:24 Challenges in Implementing AI37:16 AI Immersion Strategy Days39:52 Organizational Agility and AI42:15 AI's Impact on Investment Strategies43:26 AI in Mergers and Acquisitions45:29 The Importance of Proprietary Data47:12 AI Fatigue and Disillusionment48:07 Building AI Products and Agentic Products48:58 Hg's Internal AI Incubator49:59 The Next Wave of AI51:19 Voice and Multimodality in AI51:55 Globalization and Internationalization of AI53:35 Overestimating and Underestimating AI's Impact54:54 The Competitive Landscape of AI55:42 The Future of Value Creation with AI56:23 Conclusion and Final ThoughtsEditing and post-production work for this episode was provided by The Podcast Consultant.

ReachMD CME
Multidisciplinary Collaboration Facilitates Multimodality Therapy

ReachMD CME

Play Episode Listen Later Aug 22, 2025


CME credits: 0.75 Valid until: 22-08-2026 Claim your CME credit at https://reachmd.com/programs/cme/multidisciplinary-collaboration-facilitates-multimodality-therapy/36636/ This online CME activity examines advances in managing resectable locally advanced head and neck squamous cell carcinoma (HNSCC), focusing on the integration of perioperative immune checkpoint inhibitors (ICIs) and multimodal approaches. Faculty review current standards of care and highlight unmet needs that have driven investigation into combining radiation and immunotherapy. Emerging clinical trial data are discussed, including the impact of perioperative ICIs on event-free survival and pathologic response, with attention to patient selection informed by risk stratification and biomarkers. The program also addresses practical considerations for multidisciplinary care, including immune-related adverse event management and strategies to support patient access to these evolving treatment paradigms.

ReachMD CME
Multidisciplinary Collaboration Facilitates Multimodality Therapy

ReachMD CME

Play Episode Listen Later Aug 22, 2025


CME credits: 0.75 Valid until: 22-08-2026 Claim your CME credit at https://reachmd.com/programs/cme/multidisciplinary-collaboration-facilitates-multimodality-therapy/36636/ This online CME activity examines advances in managing resectable locally advanced head and neck squamous cell carcinoma (HNSCC), focusing on the integration of perioperative immune checkpoint inhibitors (ICIs) and multimodal approaches. Faculty review current standards of care and highlight unmet needs that have driven investigation into combining radiation and immunotherapy. Emerging clinical trial data are discussed, including the impact of perioperative ICIs on event-free survival and pathologic response, with attention to patient selection informed by risk stratification and biomarkers. The program also addresses practical considerations for multidisciplinary care, including immune-related adverse event management and strategies to support patient access to these evolving treatment paradigms.

Latent Space: The AI Engineer Podcast — CodeGen, Agents, Computer Vision, Data Science, AI UX and all things Software 3.0

Speak (https://speak.com) may not be very well known to native English speakers, but they have come from a slow start in 2016 to emerge as one of the favorite partners of OpenAI, with their Startup Fund leading and joining their Series B and C as one of the new AI-native unicorns, noting that “Speak has the potential to revolutionize not just language learning, but education broadly”. Today we speak with Speak's CTO, Andrew Hsu, on the journey of building the “3rd generation” of language learning software (with Rosetta Stone being Gen 1, and Duolingo being Gen 2). Speak's premise is that speech and language models can now do what was previously only possible with human tutors—provide fluent, responsive, and adaptive instruction—and this belief has shaped its product and company strategy since its early days. https://www.linkedin.com/in/adhsu/ https://speak.com One of the most interesting strategic decisions discussed in the episode is Speak's early focus on South Korea. While counterintuitive for a San Francisco-based startup, the decision was influenced by a combination of market opportunity and founder proximity via a Korean first employee. South Korea's intense demand for English fluency and a highly competitive education market made it a proving ground for a deeply AI-native product. By succeeding in a market saturated with human-based education solutions, Speak validated its model and built strong product-market fit before expanding to other Asian markets and eventually, globally. The arrival of Whisper and GPT-based LLMs in 2022 marked a turning point for Speak. Suddenly, capabilities that were once theoretical—real-time feedback, semantic understanding, conversational memory—became technically feasible. Speak didn't pivot, but rather evolved into its second phase: from a supplemental practice tool to a full-featured language tutor. This transition required significant engineering work, including building custom ASR models, managing latency, and integrating real-time APIs for interactive lessons. It also unlocked the possibility of developing voice-first, immersive roleplay experiences and a roadmap to real-time conversational fluency. To scale globally and support many languages, Speak is investing heavily in AI-generated curriculum and content. Instead of manually scripting all lessons, they are building agents and pipelines that can scaffold curriculum, generate lesson content, and adapt pedagogically to the learner. This ties into one of Speak's most ambitious goals: creating a knowledge graph that captures what a learner knows and can do in a target language, and then adapting the course path accordingly. This level-adjusting tutor model aims to personalize learning at scale and could eventually be applied beyond language learning to any educational domain. Finally, the conversation touches on the broader implications of AI-powered education and the slow real-world adoption of transformative AI technologies. Despite the capabilities of GPT-4 and others, most people's daily lives haven't changed dramatically. Speak sees itself as part of the generation of startups that will translate AI's raw power into tangible consumer value. The company is also a testament to long-term conviction—founded in 2016, it weathered years of slow growth before AI caught up to its vision. Now, with over $50M ARR, a growing B2B arm, and plans to expand across languages and learning domains, Speak represents what AI-native education could look like in the next decade. Chapters 00:00:00 Introductions & Thiel Fellowship Origins 00:02:13 Genesis of Speak: Early Vision & Market Focus 00:03:44 Building the Product: Iterations and Lessons Learned 00:10:59 AI's Role in Language Learning 00:13:49 Scaling Globally & B2B Expansion 00:16:30 Why Korea? Localizing for Success 00:19:08 Content Creation, The Speak Method, and Engineering Culture 00:23:31 The Impact of Whisper and LLM Advances 00:29:08 AI-Generated Content & Measuring Fluency 00:35:30 Personalization, Dialects, and Pronunciation 00:39:38 Immersive Learning, Multimodality, and Real-Time Voice 00:50:02 Engineering Challenges & Company Culture 00:53:20 Beyond Languages: B2B, Knowledge Graphs, and Broader Learning 00:57:32 Fun Stories, Lessons, and Reflections 01:02:03 Final Thoughts: The Future of AI Learning & Slow Takeoff

Latent Space: The AI Engineer Podcast — CodeGen, Agents, Computer Vision, Data Science, AI UX and all things Software 3.0

Speak (https://speak.com) may not be very well known to native English speakers, but they have come from a slow start in 2016 to emerge as one of the favorite partners of OpenAI, with their Startup Fund leading and joining their Series B and C as one of the new AI-native unicorns, noting that “Speak has the potential to revolutionize not just language learning, but education broadly”.Today we speak with Speak's CTO, Andrew Hsu, on the journey of building the “3rd generation” of language learning software (with Rosetta Stone being Gen 1, and Duolingo being Gen 2). Speak's premise is that speech and language models can now do what was previously only possible with human tutors—provide fluent, responsive, and adaptive instruction—and this belief has shaped its product and company strategy since its early days.https://www.linkedin.com/in/adhsu/https://speak.comOne of the most interesting strategic decisions discussed in the episode is Speak's early focus on South Korea. While counterintuitive for a San Francisco-based startup, the decision was influenced by a combination of market opportunity and founder proximity via a Korean first employee. South Korea's intense demand for English fluency and a highly competitive education market made it a proving ground for a deeply AI-native product. By succeeding in a market saturated with human-based education solutions, Speak validated its model and built strong product-market fit before expanding to other Asian markets and eventually, globally.The arrival of Whisper and GPT-based LLMs in 2022 marked a turning point for Speak. Suddenly, capabilities that were once theoretical—real-time feedback, semantic understanding, conversational memory—became technically feasible. Speak didn't pivot, but rather evolved into its second phase: from a supplemental practice tool to a full-featured language tutor. This transition required significant engineering work, including building custom ASR models, managing latency, and integrating real-time APIs for interactive lessons. It also unlocked the possibility of developing voice-first, immersive roleplay experiences and a roadmap to real-time conversational fluency.To scale globally and support many languages, Speak is investing heavily in AI-generated curriculum and content. Instead of manually scripting all lessons, they are building agents and pipelines that can scaffold curriculum, generate lesson content, and adapt pedagogically to the learner. This ties into one of Speak's most ambitious goals: creating a knowledge graph that captures what a learner knows and can do in a target language, and then adapting the course path accordingly. This level-adjusting tutor model aims to personalize learning at scale and could eventually be applied beyond language learning to any educational domain.Finally, the conversation touches on the broader implications of AI-powered education and the slow real-world adoption of transformative AI technologies. Despite the capabilities of GPT-4 and others, most people's daily lives haven't changed dramatically. Speak sees itself as part of the generation of startups that will translate AI's raw power into tangible consumer value. The company is also a testament to long-term conviction—founded in 2016, it weathered years of slow growth before AI caught up to its vision. Now, with over $50M ARR, a growing B2B arm, and plans to expand across languages and learning domains, Speak represents what AI-native education could look like in the next decade.Full Video EpisodeTimestamps00:00 Introductions & Thiel Fellowship Origins02:13 Genesis of Speak: Early Vision & Market Focus03:44 Building the Product: Iterations and Lessons Learned10:59 AI's Role in Language Learning13:49 Scaling Globally & B2B Expansion16:30 Why Korea? Localizing for Success19:08 Content Creation, The Speak Method, and Engineering Culture23:31 The Impact of Whisper and LLM Advances29:08 AI-Generated Content & Measuring Fluency35:30 Personalization, Dialects, and Pronunciation39:38 Immersive Learning, Multimodality, and Real-Time Voice50:02 Engineering Challenges & Company Culture53:20 Beyond Languages: B2B, Knowledge Graphs, and Broader Learning57:32 Fun Stories, Lessons, and Reflections1:02:03 Final Thoughts: The Future of AI Learning & Slow Takeoff Get full access to Latent.Space at www.latent.space/subscribe

Cardionerds
420. Cardio-Rheumatology: Cardiovascular Multimodality Imaging & Systemic Inflammation with Dr. Monica Mukherjee

Cardionerds

Play Episode Listen Later Jun 20, 2025 17:54


In this episode, CardioNerds Dr. Gurleen Kaur, Dr. Richard Ferraro, and Dr. Jake Roberts are joined by Cardio-Rheumatology expert, Dr. Monica Mukherjee, to discuss the role of utilizing multimodal imaging for cardiovascular disease risk stratification, monitoring, and management in patients with chronic systemic inflammation. The team delves into the contexts for utilizing advanced imaging to assess systemic inflammation with cardiac involvement, as well as the role of imaging in monitoring various specific cardiovascular complications that may develop due to inflammatory diseases. Audio editing by CardioNerds academy intern, Christiana Dangas. CardioNerds Prevention PageCardioNerds Episode PageCardioNerds AcademyCardionerds Healy Honor Roll CardioNerds Journal ClubSubscribe to The Heartbeat Newsletter!Check out CardioNerds SWAG!Become a CardioNerds Patron! Pearls - Cardiovascular Multimodality Imaging & Systemic Inflammation Systemic inflammatory diseases are associated with an elevated CVD risk that has significant implications for early detection, risk stratification, and implementation of therapeutic strategies to address these risks and disease-specific complications. As an example, patients with SLE have a 48-fold increased risk for developing ASCVD compared to the general population. They may also develop disease-specific complications, such as pericarditis, that require focused imaging approaches to detect. In addition to increasing the risk for CAD, systemic inflammatory diseases can also result in cardiac complications, including myocardial, pericardial, and valvular involvement. Assessment of these complications requires the use of different imaging techniques, with the modality and serial studies selected based on the suspected disease process involved. In most contexts, echocardiography remains the starting point for evaluating cardiac involvement in systemic inflammatory diseases and can inform the next steps in terms of diagnostic study selection for the assessment of specific cardiac processes. For example, if echocardiography is completed in an SLE patient and demonstrates potential myocardial or pericardial inflammation, the next steps in evaluation may include completing a cardiac MRI for better characterization. While no current guidelines or standards of care directly guide our selection of advanced imaging studies for screening and management of CVD in patients with systemic inflammatory diseases, our understanding of cardiac involvement in these patients continues to improve and will likely lead to future guideline development. Due to the vast heterogeneity of cardiac involvement both across and within different systemic inflammatory diseases, a personalized approach to caring for each individual patient remains central to CVD evaluation and management in these patients. For example, patients with systemic sclerosis and symptoms of shortness of breath may experience these symptoms due to a range of causes. Echocardiography can be a central guiding tool in assessing these patients for potential concerns related to pulmonary hypertension or diastolic dysfunction. Based on the initial echocardiogram, the next steps in evaluation may involve further ischemic evaluation or right heart catheterization, depending on the pathology of concern. Show notes - Cardiovascular Multimodality Imaging & Systemic Inflammation Episode notes drafted by Dr. Jake Roberts. What are the contexts in which we should consider pursuing multimodal cardiac imaging, and are there certain inflammatory disorders associated with systemic inflammation and higher associated CVD risk for which advanced imaging can help guide early intervention? Systemic inflammatory diseases are associated with elevated CVD risk, which has significant implications for early detection, risk stratification, prognostication, and implementation of therapeutic strategies to address CVD risk and complicat...

Software Defined Talk
Episode 516: Vibe Strategy

Software Defined Talk

Play Episode Listen Later Apr 25, 2025 67:32


This week, we discuss Google being found to be a monopoly, OpenAI's “offer” to buy Chrome, and some hot takes on JSON. Plus, is it better to wait on hold or ask for a callback? Watch the YouTube Live Recording of Episode (https://www.youtube.com/watch?v=EhUxUPJv5g4) 516 (https://www.youtube.com/watch?v=EhUxUPJv5g4) Runner-up Titles Just Fine The SDT “Fine” Scale Callback Asynchronous Friendship I would love to get to know you better…over text Send you Jams to the dry cleaners. JSON Take it xslt-easy! Rundown OpenAI OpenAI in talks to pay about $3 billion to acquire AI coding startup Windsurf (https://www.cnbc.com/2025/04/16/openai-in-talks-to-pay-about-3-billion-to-acquire-startup-windsurf.html) The Cursor Mirage (https://artificialintelligencemadesimple.substack.com/p/the-cursor-mirage) AI is for Tinkerers (https://redmonk.com/kholterhoff/2023/06/27/ai-is-for-tinkerers/) Vibe Coding is for PMs (https://redmonk.com/rstephens/2025/04/18/vibe-coding-is-for-pms/) OpenAI releases new simulated reasoning models with full tool access (https://arstechnica.com/ai/2025/04/openai-releases-new-simulated-reasoning-models-with-full-tool-access/) Clouded Judgement 4.18.25 - The Hidden Value in the AI Application Layer (https://cloudedjudgement.substack.com/p/clouded-judgement-41825-the-hidden?utm_source=post-email-title&publication_id=56878&post_id=161562220&utm_campaign=email-post-title&isFreemail=true&r=2l9&triedRedirect=true&utm_medium=email) OpenAI tells judge it would buy Chrome from Google (https://www.theverge.com/news/653882/openai-chrome-google-us-judge) The Creators of Model Context Protocol (https://www.latent.space/p/mcp?utm_source=substack&utm_medium=email) Judge finds Google holds illegal online ad tech monopolies (https://www.cnbc.com/2025/04/17/judge-finds-google-holds-illegal-online-ad-tech-monopolies.html) Intuit, Owner of TurboTax, Wins Battle Against America's Taxpayers (https://prospect.org/power/2025-04-17-intuit-turbotax-wins-battle-against-taxpayers-irs-direct-file/) Relevant to your Interests Switch 2 Carts Still Taste Bad, Designed Purposefully To Be Spat Out (https://www.gamespot.com/articles/switch-2-carts-still-taste-bad-designed-purposefully-to-be-spat-out/1100-6530649/) CEO Andy Jassy's 2024 Letter to Shareholders (https://www.aboutamazon.com/news/company-news/amazon-ceo-andy-jassy-2024-letter-to-shareholders) Amazon CEO Andy Jassy says AI costs will come down (https://www.cnbc.com/2025/04/10/amazon-ceo-andy-jassys-2025-shareholder-letter.html) Happy 18th Birthday CUDA! (https://www.aboutamazon.com/news/company-news/amazon-ceo-andy-jassy-2024-letter-to-shareholders) Honeycomb Acquires Grit: A Strategic Investment in Pragmatic AI and Customer Value (https://www.honeycomb.io/blog/honeycomb-acquires-grit) Everything Announced at Google Cloud Next in 12 Minutes (https://www.youtube.com/watch?v=2OpHbyN4vEM) GitLab vs GitHub : Key Differences in 2025 (https://spacelift.io/blog/gitlab-vs-github) Old Fashioned Function Keys (https://economistwritingeveryday.com/2025/04/11/old-fashioned-function-keys/) Fake job seekers are flooding U.S. companies that are hiring for remote positions, (https://www.cnbc.com/2025/04/08/fake-job-seekers-use-ai-to-interview-for-remote-jobs-tech-ceos-say.html) NetRise raises $10M to expand software supply chain security platform (https://siliconangle.com/2025/04/15/netrise-raises-10-million-expand-software-supply-chain-security-platform/) Mark Zuckerberg's antitrust testimony aired his wildest ideas from Meta's history (https://www.theverge.com/policy/649520/zuckerberg-meta-ftc-antitrust-testimony-facebook-history) How Much Should I Be Spending On Observability? (https://www.honeycomb.io/blog/how-much-should-i-spend-on-observability-pt1) Did we just make platform engineering much easier by shipping a cloud IDP? (https://seroter.com/2025/04/16/did-we-just-make-platform-engineering-much-easier-by-shipping-a-cloud-idp/) Google Cloud Next 2025: Agentic AI Stack, Multimodality, And Sovereignty (https://www.forrester.com/blogs/google-next-2025-agentic-ai-stack-multimodality-and-sovereignty/) iPhone Shipments Down 9% in China's Q1 Smartphone Boom (https://www.macrumors.com/2025/04/18/iphone-shipments-down-in-china-q1/) Exclusive: Anthropic warns fully AI employees are a year away (https://www.axios.com/2025/04/22/ai-anthropic-virtual-employees-security) Synology requires self-branded drives for some consumer NAS systems, drops full functionality and support for third-party HDDs (https://www.tomshardware.com/pc-components/nas/synology-requires-self-branded-drives-for-some-consumer-nas-systems-drops-full-functionality-and-support-for-third-party-hdds) Porting Tailscale to Plan 9 (https://tailscale.com/blog/plan9-port?ck_subscriber_id=512840665&utm_source=convertkit&utm_medium=email&utm_campaign=[Last%20Week%20in%20AWS]%20Issue%20#418:%20Another%20New%20Capacity%20Dingus%20-%2017270009) CVE Foundation (https://www.thecvefoundation.org/) The Cursor Mirage (https://artificialintelligencemadesimple.substack.com/p/the-cursor-mirage) There's a Lot of Bad Telemetry Out There (https://blog.olly.garden/theres-a-lot-of-bad-telemetry-out-there) Gee Wiz (https://redmonk.com/rstephens/2025/04/04/gee-wiz/?ck_subscriber_id=512840665&utm_source=convertkit&utm_medium=email&utm_campaign=[Last%20Week%20in%20AWS]%20Issue%20#418:%20Another%20New%20Capacity%20Dingus%20-%2017270009) Nonsense Silicon Valley crosswalk buttons hacked to imitate Musk, Zuckerberg's voices (https://techcrunch.com/2025/04/14/silicon-valley-crosswalk-buttons-hacked-to-imitate-musk-zuckerberg-voices/) A Visit to Costco in France (https://davidlebovitz.substack.com/p/a-visit-to-costco-in-france) No sweat: Humanoid robots run a Chinese half-marathon (https://apnews.com/article/china-robot-half-marathon-153c6823bd628625106ed26267874d21) Metre, a consistent measurement of the world (https://mappingignorance.org/2025/04/23/150-years-ago-the-metre-convention-determined-how-we-measure-the-world/) Conferences DevOps Days Atlanta (https://devopsdays.org/events/2025-atlanta/welcome/), April 29th-30th. KCD Texas Austin 2025 (https://community.cncf.io/events/details/cncf-kcd-texas-presents-kcd-texas-austin-2025/), May 15th, Whitney Lee Speaking. Cloud Foundry Day US (https://events.linuxfoundation.org/cloud-foundry-day-north-america/), May 14th, Palo Alto, CA, Coté speaking. Fr (https://vmwarereg.fig-street.com/051325-tanzu-workshop/)ee AI workshop (https://vmwarereg.fig-street.com/051325-tanzu-workshop/), May 13th. day before C (https://events.linuxfoundation.org/cloud-foundry-day-north-america/)loud (https://events.linuxfoundation.org/cloud-foundry-day-north-america/) (https://events.linuxfoundation.org/cloud-foundry-day-north-america/)Foundry (https://events.linuxfoundation.org/cloud-foundry-day-north-america/) Day (https://events.linuxfoundation.org/cloud-foundry-day-north-america/). NDC Oslo (https://ndcoslo.com/), May 21st-23th, Coté speaking. SDT News & Community Join our Slack community (https://softwaredefinedtalk.slack.com/join/shared_invite/zt-1hn55iv5d-UTfN7mVX1D9D5ExRt3ZJYQ#/shared-invite/email) Email the show: questions@softwaredefinedtalk.com (mailto:questions@softwaredefinedtalk.com) Free stickers: Email your address to stickers@softwaredefinedtalk.com (mailto:stickers@softwaredefinedtalk.com) Follow us on social media: Twitter (https://twitter.com/softwaredeftalk), Threads (https://www.threads.net/@softwaredefinedtalk), Mastodon (https://hachyderm.io/@softwaredefinedtalk), LinkedIn (https://www.linkedin.com/company/software-defined-talk/), BlueSky (https://bsky.app/profile/softwaredefinedtalk.com) Watch us on: Twitch (https://www.twitch.tv/sdtpodcast), YouTube (https://www.youtube.com/channel/UCi3OJPV6h9tp-hbsGBLGsDQ/featured), Instagram (https://www.instagram.com/softwaredefinedtalk/), TikTok (https://www.tiktok.com/@softwaredefinedtalk) Book offer: Use code SDT for $20 off "Digital WTF" by Coté (https://leanpub.com/digitalwtf/c/sdt) Sponsor the show (https://www.softwaredefinedtalk.com/ads): ads@softwaredefinedtalk.com (mailto:ads@softwaredefinedtalk.com) Recommendations Brandon: Dope Thief (https://www.rottentomatoes.com/tv/dope_thief) on Apple TV (https://www.rottentomatoes.com/tv/dope_thief) Coté: Check out the recording of the Tanzu Annual update (https://www.youtube.com/watch?v=c1QZXzJcAfQ), all about Tanzu's private AI platform. Next, watch Coté's new MCP for D&D video (#4) figures out something cool to do with MCP Prompts (https://www.youtube.com/watch?v=xEtYBznneFg), they make sense now. And, a regret-a-mmendation: Fields Notes annual subscription (https://fieldnotesbrand.com/limited-editions). Photo Credits Header (https://unsplash.com/photos/a-telephone-sitting-on-top-of-a-wooden-shelf-2XnGRN_caHc)

Pivot
Demis Hassabis on AI, Game Theory, Multimodality, and the Nature of Creativity | Possible

Pivot

Play Episode Listen Later Apr 12, 2025 60:49


How can AI help us understand and master deeply complex systems—from the game Go, which has 10 to the power 170 possible positions a player could pursue, or proteins, which, on average, can fold in 10 to the power 300 possible ways? This week, Reid and Aria are joined by Demis Hassabis. Demis is a British artificial intelligence researcher, co-founder, and CEO of the AI company, DeepMind. Under his leadership, DeepMind developed Alpha Go, the first AI to defeat a human world champion in Go and later created AlphaFold, which solved the 50-year-old protein folding problem. He's considered one of the most influential figures in AI. Demis, Reid, and Aria discuss game theory, medicine, multimodality, and the nature of innovation and creativity. For more info on the podcast and transcripts of all the episodes, visit https://www.possible.fm/podcast/  Listen to more from Possible here. Learn more about your ad choices. Visit podcastchoices.com/adchoices

Possible
Demis Hassabis on AI, game theory, multimodality, and the nature of creativity

Possible

Play Episode Listen Later Apr 9, 2025 56:40


How can AI help us understand and master deeply complex systems—from the game Go, which has 10 to the power 170 possible positions a player could pursue, or proteins, which, on average, can fold in 10 to the power 300 possible ways? This week, Reid and Aria are joined by Demis Hassabis. Demis is a British artificial intelligence researcher, co-founder, and CEO of the AI company, DeepMind. Under his leadership, DeepMind developed Alpha Go, the first AI to defeat a human world champion in Go and later created AlphaFold, which solved the 50-year-old protein folding problem. He's considered one of the most influential figures in AI. Demis, Reid, and Aria discuss game theory, medicine, multimodality, and the nature of innovation and creativity. For more info on the podcast and transcripts of all the episodes, visit https://www.possible.fm/podcast/  Select mentions:  Hitchhiker's Guide to the Galaxy by Douglas Adams AlphaGo documentary: https://www.youtube.com/watch?v=WXuK6gekU1Y Nash equilibrium & US mathematician John Forbes Nash Homo Ludens by Johan Huizinga Veo 2, an advanced, AI-powered video creation platform from Google DeepMind The Culture series by Iain Banks Hartmut Neven, German-American computer scientist Topics: 3:11 - Hellos and intros 5:20 - Brute force vs. self-learning systems 8:24 - How a learning approach helped develop new AI systems 11:29 - AlphaGo's Move 37 16:16 - What will the next Move 37 be? 19:42 - What makes an AI that can play the video game StarCraft impressive 22:32 - The importance of the act of play 26:24 - Data and synthetic data 28:33 - Midroll ad 28:39 - Is it important to have AI embedded in the world? 33:44 - The trade-off between thinking time and output quality 36:03 - Computer languages designed for AI 40:22 - The future of multimodality  43:27 - AI and geographic diversity  48:24 - AlphaFold and the future of medicine 51:18 - Rapid-fire Questions Possible is an award-winning podcast that sketches out the brightest version of the future—and what it will take to get there. Most of all, it asks: what if, in the future, everything breaks humanity's way? Tune in for grounded and speculative takes on how technology—and, in particular, AI—is inspiring change and transforming the future. Hosted by Reid Hoffman and Aria Finger, each episode features an interview with an ambitious builder or deep thinker on a topic, from art to geopolitics and from healthcare to education. These conversations also showcase another kind of guest: AI. Each episode seeks to enhance and advance our discussion about what humanity could possibly get right if we leverage technology—and our collective effort—effectively.

Classroom Caffeine
A Conversation with Raúl Alberto Mora

Classroom Caffeine

Play Episode Listen Later Apr 8, 2025 41:27 Transcription Available


Send us a textIn this episode, Raúl Alberto Mora talks to us about education theory as a driver for innovative teaching, mentoring and supporting one another, and the journey of a career in Education. Raúl is known worldwide for his work in the areas of alternative literacy paradigms in second language education and research, the study of second language literacies in physical and virtual spaces, and the use of sociocritical frameworks in language education. In particular, he studies the applications of alternative literacy paradigms to analyze second-language literacy practices in urban and virtual spaces He works to understand the use of languages a social and semiotic resource. His work has been published in the Journal of Adolescent and Adult Literacy, The ALAN Review, Bilingualism and Bilingual Education, International Journal of Cultural Studies, Social Semiotics, Key Concepts in Intercultural Dialogue, Pedagogies: An International Journal, and other journals. He co-edited The Handbook of Critical Literacies, Translanguaging and Multimodality as Flow, Agency, and a New Sense of Advocacy in and From the Global South, and most recently, Reimagining Literacy in the Age of AI: Theory and Practice. Dr. Raúl Alberto Mora Velez is a researcher at the Educations, Languages, and Learning Environments research group and chairs the award-winning Literacies in Second Languages Project (LSLP) research lab. Raúl is a Research Professor at Universidad Pontificia Bolivariana in Colombia. For more information about our guest, stay tuned to the end of this episode.Links mentioned in this episode:Literacies in Second Languages Project Micro-PapersAmerican Educational Research AssociationLiteracy Research AssociationConnect with Classroom Caffeine at www.classroomcaffeine.com or on Instagram, Facebook, Twitter, and LinkedIn.

Latent Space: The AI Engineer Podcast — CodeGen, Agents, Computer Vision, Data Science, AI UX and all things Software 3.0

Today's episode is with Paul Klein, founder of Browserbase. We talked about building browser infrastructure for AI agents, the future of agent authentication, and their open source framework Stagehand.* [00:00:00] Introductions* [00:04:46] AI-specific challenges in browser infrastructure* [00:07:05] Multimodality in AI-Powered Browsing* [00:12:26] Running headless browsers at scale* [00:18:46] Geolocation when proxying* [00:21:25] CAPTCHAs and Agent Auth* [00:28:21] Building “User take over” functionality* [00:33:43] Stagehand: AI web browsing framework* [00:38:58] OpenAI's Operator and computer use agents* [00:44:44] Surprising use cases of Browserbase* [00:47:18] Future of browser automation and market competition* [00:53:11] Being a solo founderTranscriptAlessio [00:00:04]: Hey everyone, welcome to the Latent Space podcast. This is Alessio, partner and CTO at Decibel Partners, and I'm joined by my co-host Swyx, founder of Smol.ai.swyx [00:00:12]: Hey, and today we are very blessed to have our friends, Paul Klein, for the fourth, the fourth, CEO of Browserbase. Welcome.Paul [00:00:21]: Thanks guys. Yeah, I'm happy to be here. I've been lucky to know both of you for like a couple of years now, I think. So it's just like we're hanging out, you know, with three ginormous microphones in front of our face. It's totally normal hangout.swyx [00:00:34]: Yeah. We've actually mentioned you on the podcast, I think, more often than any other Solaris tenant. Just because like you're one of the, you know, best performing, I think, LLM tool companies that have started up in the last couple of years.Paul [00:00:50]: Yeah, I mean, it's been a whirlwind of a year, like Browserbase is actually pretty close to our first birthday. So we are one years old. And going from, you know, starting a company as a solo founder to... To, you know, having a team of 20 people, you know, a series A, but also being able to support hundreds of AI companies that are building AI applications that go out and automate the web. It's just been like, really cool. It's been happening a little too fast. I think like collectively as an AI industry, let's just take a week off together. I took my first vacation actually two weeks ago, and Operator came out on the first day, and then a week later, DeepSeat came out. And I'm like on vacation trying to chill. I'm like, we got to build with this stuff, right? So it's been a breakneck year. But I'm super happy to be here and like talk more about all the stuff we're seeing. And I'd love to hear kind of what you guys are excited about too, and share with it, you know?swyx [00:01:39]: Where to start? So people, you've done a bunch of podcasts. I think I strongly recommend Jack Bridger's Scaling DevTools, as well as Turner Novak's The Peel. And, you know, I'm sure there's others. So you covered your Twilio story in the past, talked about StreamClub, you got acquired to Mux, and then you left to start Browserbase. So maybe we just start with what is Browserbase? Yeah.Paul [00:02:02]: Browserbase is the web browser for your AI. We're building headless browser infrastructure, which are browsers that run in a server environment that's accessible to developers via APIs and SDKs. It's really hard to run a web browser in the cloud. You guys are probably running Chrome on your computers, and that's using a lot of resources, right? So if you want to run a web browser or thousands of web browsers, you can't just spin up a bunch of lambdas. You actually need to use a secure containerized environment. You have to scale it up and down. It's a stateful system. And that infrastructure is, like, super painful. And I know that firsthand, because at my last company, StreamClub, I was CTO, and I was building our own internal headless browser infrastructure. That's actually why we sold the company, is because Mux really wanted to buy our headless browser infrastructure that we'd built. And it's just a super hard problem. And I actually told my co-founders, I would never start another company unless it was a browser infrastructure company. And it turns out that's really necessary in the age of AI, when AI can actually go out and interact with websites, click on buttons, fill in forms. You need AI to do all of that work in an actual browser running somewhere on a server. And BrowserBase powers that.swyx [00:03:08]: While you're talking about it, it occurred to me, not that you're going to be acquired or anything, but it occurred to me that it would be really funny if you became the Nikita Beer of headless browser companies. You just have one trick, and you make browser companies that get acquired.Paul [00:03:23]: I truly do only have one trick. I'm screwed if it's not for headless browsers. I'm not a Go programmer. You know, I'm in AI grant. You know, browsers is an AI grant. But we were the only company in that AI grant batch that used zero dollars on AI spend. You know, we're purely an infrastructure company. So as much as people want to ask me about reinforcement learning, I might not be the best guy to talk about that. But if you want to ask about headless browser infrastructure at scale, I can talk your ear off. So that's really my area of expertise. And it's a pretty niche thing. Like, nobody has done what we're doing at scale before. So we're happy to be the experts.swyx [00:03:59]: You do have an AI thing, stagehand. We can talk about the sort of core of browser-based first, and then maybe stagehand. Yeah, stagehand is kind of the web browsing framework. Yeah.What is Browserbase? Headless Browser Infrastructure ExplainedAlessio [00:04:10]: Yeah. Yeah. And maybe how you got to browser-based and what problems you saw. So one of the first things I worked on as a software engineer was integration testing. Sauce Labs was kind of like the main thing at the time. And then we had Selenium, we had Playbrite, we had all these different browser things. But it's always been super hard to do. So obviously you've worked on this before. When you started browser-based, what were the challenges? What were the AI-specific challenges that you saw versus, there's kind of like all the usual running browser at scale in the cloud, which has been a problem for years. What are like the AI unique things that you saw that like traditional purchase just didn't cover? Yeah.AI-specific challenges in browser infrastructurePaul [00:04:46]: First and foremost, I think back to like the first thing I did as a developer, like as a kid when I was writing code, I wanted to write code that did stuff for me. You know, I wanted to write code to automate my life. And I do that probably by using curl or beautiful soup to fetch data from a web browser. And I think I still do that now that I'm in the cloud. And the other thing that I think is a huge challenge for me is that you can't just create a web site and parse that data. And we all know that now like, you know, taking HTML and plugging that into an LLM, you can extract insights, you can summarize. So it was very clear that now like dynamic web scraping became very possible with the rise of large language models or a lot easier. And that was like a clear reason why there's been more usage of headless browsers, which are necessary because a lot of modern websites don't expose all of their page content via a simple HTTP request. You know, they actually do require you to run this type of code for a specific time. JavaScript on the page to hydrate this. Airbnb is a great example. You go to airbnb.com. A lot of that content on the page isn't there until after they run the initial hydration. So you can't just scrape it with a curl. You need to have some JavaScript run. And a browser is that JavaScript engine that's going to actually run all those requests on the page. So web data retrieval was definitely one driver of starting BrowserBase and the rise of being able to summarize that within LLM. Also, I was familiar with if I wanted to automate a website, I could write one script and that would work for one website. It was very static and deterministic. But the web is non-deterministic. The web is always changing. And until we had LLMs, there was no way to write scripts that you could write once that would run on any website. That would change with the structure of the website. Click the login button. It could mean something different on many different websites. And LLMs allow us to generate code on the fly to actually control that. So I think that rise of writing the generic automation scripts that can work on many different websites, to me, made it clear that browsers are going to be a lot more useful because now you can automate a lot more things without writing. If you wanted to write a script to book a demo call on 100 websites, previously, you had to write 100 scripts. Now you write one script that uses LLMs to generate that script. That's why we built our web browsing framework, StageHand, which does a lot of that work for you. But those two things, web data collection and then enhanced automation of many different websites, it just felt like big drivers for more browser infrastructure that would be required to power these kinds of features.Alessio [00:07:05]: And was multimodality also a big thing?Paul [00:07:08]: Now you can use the LLMs to look, even though the text in the dome might not be as friendly. Maybe my hot take is I was always kind of like, I didn't think vision would be as big of a driver. For UI automation, I felt like, you know, HTML is structured text and large language models are good with structured text. But it's clear that these computer use models are often vision driven, and they've been really pushing things forward. So definitely being multimodal, like rendering the page is required to take a screenshot to give that to a computer use model to take actions on a website. And it's just another win for browser. But I'll be honest, that wasn't what I was thinking early on. I didn't even think that we'd get here so fast with multimodality. I think we're going to have to get back to multimodal and vision models.swyx [00:07:50]: This is one of those things where I forgot to mention in my intro that I'm an investor in Browserbase. And I remember that when you pitched to me, like a lot of the stuff that we have today, we like wasn't on the original conversation. But I did have my original thesis was something that we've talked about on the podcast before, which is take the GPT store, the custom GPT store, all the every single checkbox and plugin is effectively a startup. And this was the browser one. I think the main hesitation, I think I actually took a while to get back to you. The main hesitation was that there were others. Like you're not the first hit list browser startup. It's not even your first hit list browser startup. There's always a question of like, will you be the category winner in a place where there's a bunch of incumbents, to be honest, that are bigger than you? They're just not targeted at the AI space. They don't have the backing of Nat Friedman. And there's a bunch of like, you're here in Silicon Valley. They're not. I don't know.Paul [00:08:47]: I don't know if that's, that was it, but like, there was a, yeah, I mean, like, I think I tried all the other ones and I was like, really disappointed. Like my background is from working at great developer tools, companies, and nothing had like the Vercel like experience. Um, like our biggest competitor actually is partly owned by private equity and they just jacked up their prices quite a bit. And the dashboard hasn't changed in five years. And I actually used them at my last company and tried them and I was like, oh man, like there really just needs to be something that's like the experience of these great infrastructure companies, like Stripe, like clerk, like Vercel that I use in love, but oriented towards this kind of like more specific category, which is browser infrastructure, which is really technically complex. Like a lot of stuff can go wrong on the internet when you're running a browser. The internet is very vast. There's a lot of different configurations. Like there's still websites that only work with internet explorer out there. How do you handle that when you're running your own browser infrastructure? These are the problems that we have to think about and solve at BrowserBase. And it's, it's certainly a labor of love, but I built this for me, first and foremost, I know it's super cheesy and everyone says that for like their startups, but it really, truly was for me. If you look at like the talks I've done even before BrowserBase, and I'm just like really excited to try and build a category defining infrastructure company. And it's, it's rare to have a new category of infrastructure exists. We're here in the Chroma offices and like, you know, vector databases is a new category of infrastructure. Is it, is it, I mean, we can, we're in their office, so, you know, we can, we can debate that one later. That is one.Multimodality in AI-Powered Browsingswyx [00:10:16]: That's one of the industry debates.Paul [00:10:17]: I guess we go back to the LLMOS talk that Karpathy gave way long ago. And like the browser box was very clearly there and it seemed like the people who were building in this space also agreed that browsers are a core primitive of infrastructure for the LLMOS that's going to exist in the future. And nobody was building something there that I wanted to use. So I had to go build it myself.swyx [00:10:38]: Yeah. I mean, exactly that talk that, that honestly, that diagram, every box is a startup and there's the code box and then there's the. The browser box. I think at some point they will start clashing there. There's always the question of the, are you a point solution or are you the sort of all in one? And I think the point solutions tend to win quickly, but then the only ones have a very tight cohesive experience. Yeah. Let's talk about just the hard problems of browser base you have on your website, which is beautiful. Thank you. Was there an agency that you used for that? Yeah. Herb.paris.Paul [00:11:11]: They're amazing. Herb.paris. Yeah. It's H-E-R-V-E. I highly recommend for developers. Developer tools, founders to work with consumer agencies because they end up building beautiful things and the Parisians know how to build beautiful interfaces. So I got to give prep.swyx [00:11:24]: And chat apps, apparently are, they are very fast. Oh yeah. The Mistral chat. Yeah. Mistral. Yeah.Paul [00:11:31]: Late chat.swyx [00:11:31]: Late chat. And then your videos as well, it was professionally shot, right? The series A video. Yeah.Alessio [00:11:36]: Nico did the videos. He's amazing. Not the initial video that you shot at the new one. First one was Austin.Paul [00:11:41]: Another, another video pretty surprised. But yeah, I mean, like, I think when you think about how you talk about your company. You have to think about the way you present yourself. It's, you know, as a developer, you think you evaluate a company based on like the API reliability and the P 95, but a lot of developers say, is the website good? Is the message clear? Do I like trust this founder? I'm building my whole feature on. So I've tried to nail that as well as like the reliability of the infrastructure. You're right. It's very hard. And there's a lot of kind of foot guns that you run into when running headless browsers at scale. Right.Competing with Existing Headless Browser Solutionsswyx [00:12:10]: So let's pick one. You have eight features here. Seamless integration. Scalability. Fast or speed. Secure. Observable. Stealth. That's interesting. Extensible and developer first. What comes to your mind as like the top two, three hardest ones? Yeah.Running headless browsers at scalePaul [00:12:26]: I think just running headless browsers at scale is like the hardest one. And maybe can I nerd out for a second? Is that okay? I heard this is a technical audience, so I'll talk to the other nerds. Whoa. They were listening. Yeah. They're upset. They're ready. The AGI is angry. Okay. So. So how do you run a browser in the cloud? Let's start with that, right? So let's say you're using a popular browser automation framework like Puppeteer, Playwright, and Selenium. Maybe you've written a code, some code locally on your computer that opens up Google. It finds the search bar and then types in, you know, search for Latent Space and hits the search button. That script works great locally. You can see the little browser open up. You want to take that to production. You want to run the script in a cloud environment. So when your laptop is closed, your browser is doing something. The browser is doing something. Well, I, we use Amazon. You can see the little browser open up. You know, the first thing I'd reach for is probably like some sort of serverless infrastructure. I would probably try and deploy on a Lambda. But Chrome itself is too big to run on a Lambda. It's over 250 megabytes. So you can't easily start it on a Lambda. So you maybe have to use something like Lambda layers to squeeze it in there. Maybe use a different Chromium build that's lighter. And you get it on the Lambda. Great. It works. But it runs super slowly. It's because Lambdas are very like resource limited. They only run like with one vCPU. You can run one process at a time. Remember, Chromium is super beefy. It's barely running on my MacBook Air. I'm still downloading it from a pre-run. Yeah, from the test earlier, right? I'm joking. But it's big, you know? So like Lambda, it just won't work really well. Maybe it'll work, but you need something faster. Your users want something faster. Okay. Well, let's put it on a beefier instance. Let's get an EC2 server running. Let's throw Chromium on there. Great. Okay. I can, that works well with one user. But what if I want to run like 10 Chromium instances, one for each of my users? Okay. Well, I might need two EC2 instances. Maybe 10. All of a sudden, you have multiple EC2 instances. This sounds like a problem for Kubernetes and Docker, right? Now, all of a sudden, you're using ECS or EKS, the Kubernetes or container solutions by Amazon. You're spending up and down containers, and you're spending a whole engineer's time on kind of maintaining this stateful distributed system. Those are some of the worst systems to run because when it's a stateful distributed system, it means that you are bound by the connections to that thing. You have to keep the browser open while someone is working with it, right? That's just a painful architecture to run. And there's all this other little gotchas with Chromium, like Chromium, which is the open source version of Chrome, by the way. You have to install all these fonts. You want emojis working in your browsers because your vision model is looking for the emoji. You need to make sure you have the emoji fonts. You need to make sure you have all the right extensions configured, like, oh, do you want ad blocking? How do you configure that? How do you actually record all these browser sessions? Like it's a headless browser. You can't look at it. So you need to have some sort of observability. Maybe you're recording videos and storing those somewhere. It all kind of adds up to be this just giant monster piece of your project when all you wanted to do was run a lot of browsers in production for this little script to go to google.com and search. And when I see a complex distributed system, I see an opportunity to build a great infrastructure company. And we really abstract that away with Browserbase where our customers can use these existing frameworks, Playwright, Publisher, Selenium, or our own stagehand and connect to our browsers in a serverless-like way. And control them, and then just disconnect when they're done. And they don't have to think about the complex distributed system behind all of that. They just get a browser running anywhere, anytime. Really easy to connect to.swyx [00:15:55]: I'm sure you have questions. My standard question with anything, so essentially you're a serverless browser company, and there's been other serverless things that I'm familiar with in the past, serverless GPUs, serverless website hosting. That's where I come from with Netlify. One question is just like, you promised to spin up thousands of servers. You promised to spin up thousands of browsers in milliseconds. I feel like there's no real solution that does that yet. And I'm just kind of curious how. The only solution I know, which is to kind of keep a kind of warm pool of servers around, which is expensive, but maybe not so expensive because it's just CPUs. So I'm just like, you know. Yeah.Browsers as a Core Primitive in AI InfrastructurePaul [00:16:36]: You nailed it, right? I mean, how do you offer a serverless-like experience with something that is clearly not serverless, right? And the answer is, you need to be able to run... We run many browsers on single nodes. We use Kubernetes at browser base. So we have many pods that are being scheduled. We have to predictably schedule them up or down. Yes, thousands of browsers in milliseconds is the best case scenario. If you hit us with 10,000 requests, you may hit a slower cold start, right? So we've done a lot of work on predictive scaling and being able to kind of route stuff to different regions where we have multiple regions of browser base where we have different pools available. You can also pick the region you want to go to based on like lower latency, round trip, time latency. It's very important with these types of things. There's a lot of requests going over the wire. So for us, like having a VM like Firecracker powering everything under the hood allows us to be super nimble and spin things up or down really quickly with strong multi-tenancy. But in the end, this is like the complex infrastructural challenges that we have to kind of deal with at browser base. And we have a lot more stuff on our roadmap to allow customers to have more levers to pull to exchange, do you want really fast browser startup times or do you want really low costs? And if you're willing to be more flexible on that, we may be able to kind of like work better for your use cases.swyx [00:17:44]: Since you used Firecracker, shouldn't Fargate do that for you or did you have to go lower level than that? We had to go lower level than that.Paul [00:17:51]: I find this a lot with Fargate customers, which is alarming for Fargate. We used to be a giant Fargate customer. Actually, the first version of browser base was ECS and Fargate. And unfortunately, it's a great product. I think we were actually the largest Fargate customer in our region for a little while. No, what? Yeah, seriously. And unfortunately, it's a great product, but I think if you're an infrastructure company, you actually have to have a deeper level of control over these primitives. I think it's the same thing is true with databases. We've used other database providers and I think-swyx [00:18:21]: Yeah, serverless Postgres.Paul [00:18:23]: Shocker. When you're an infrastructure company, you're on the hook if any provider has an outage. And I can't tell my customers like, hey, we went down because so-and-so went down. That's not acceptable. So for us, we've really moved to bringing things internally. It's kind of opposite of what we preach. We tell our customers, don't build this in-house, but then we're like, we build a lot of stuff in-house. But I think it just really depends on what is in the critical path. We try and have deep ownership of that.Alessio [00:18:46]: On the distributed location side, how does that work for the web where you might get sort of different content in different locations, but the customer is expecting, you know, if you're in the US, I'm expecting the US version. But if you're spinning up my browser in France, I might get the French version. Yeah.Paul [00:19:02]: Yeah. That's a good question. Well, generally, like on the localization, there is a thing called locale in the browser. You can set like what your locale is. If you're like in the ENUS browser or not, but some things do IP, IP based routing. And in that case, you may want to have a proxy. Like let's say you're running something in the, in Europe, but you want to make sure you're showing up from the US. You may want to use one of our proxy features so you can turn on proxies to say like, make sure these connections always come from the United States, which is necessary too, because when you're browsing the web, you're coming from like a, you know, data center IP, and that can make things a lot harder to browse web. So we do have kind of like this proxy super network. Yeah. We have a proxy for you based on where you're going, so you can reliably automate the web. But if you get scheduled in Europe, that doesn't happen as much. We try and schedule you as close to, you know, your origin that you're trying to go to. But generally you have control over the regions you can put your browsers in. So you can specify West one or East one or Europe. We only have one region of Europe right now, actually. Yeah.Alessio [00:19:55]: What's harder, the browser or the proxy? I feel like to me, it feels like actually proxying reliably at scale. It's much harder than spending up browsers at scale. I'm curious. It's all hard.Paul [00:20:06]: It's layers of hard, right? Yeah. I think it's different levels of hard. I think the thing with the proxy infrastructure is that we work with many different web proxy providers and some are better than others. Some have good days, some have bad days. And our customers who've built browser infrastructure on their own, they have to go and deal with sketchy actors. Like first they figure out their own browser infrastructure and then they got to go buy a proxy. And it's like you can pay in Bitcoin and it just kind of feels a little sus, right? It's like you're buying drugs when you're trying to get a proxy online. We have like deep relationships with these counterparties. We're able to audit them and say, is this proxy being sourced ethically? Like it's not running on someone's TV somewhere. Is it free range? Yeah. Free range organic proxies, right? Right. We do a level of diligence. We're SOC 2. So we have to understand what is going on here. But then we're able to make sure that like we route around proxy providers not working. There's proxy providers who will just, the proxy will stop working all of a sudden. And then if you don't have redundant proxying on your own browsers, that's hard down for you or you may get some serious impacts there. With us, like we intelligently know, hey, this proxy is not working. Let's go to this one. And you can kind of build a network of multiple providers to really guarantee the best uptime for our customers. Yeah. So you don't own any proxies? We don't own any proxies. You're right. The team has been saying who wants to like take home a little proxy server, but not yet. We're not there yet. You know?swyx [00:21:25]: It's a very mature market. I don't think you should build that yourself. Like you should just be a super customer of them. Yeah. Scraping, I think, is the main use case for that. I guess. Well, that leads us into CAPTCHAs and also off, but let's talk about CAPTCHAs. You had a little spiel that you wanted to talk about CAPTCHA stuff.Challenges of Scaling Browser InfrastructurePaul [00:21:43]: Oh, yeah. I was just, I think a lot of people ask, if you're thinking about proxies, you're thinking about CAPTCHAs too. I think it's the same thing. You can go buy CAPTCHA solvers online, but it's the same buying experience. It's some sketchy website, you have to integrate it. It's not fun to buy these things and you can't really trust that the docs are bad. What Browserbase does is we integrate a bunch of different CAPTCHAs. We do some stuff in-house, but generally we just integrate with a bunch of known vendors and continually monitor and maintain these things and say, is this working or not? Can we route around it or not? These are CAPTCHA solvers. CAPTCHA solvers, yeah. Not CAPTCHA providers, CAPTCHA solvers. Yeah, sorry. CAPTCHA solvers. We really try and make sure all of that works for you. I think as a dev, if I'm buying infrastructure, I want it all to work all the time and it's important for us to provide that experience by making sure everything does work and monitoring it on our own. Yeah. Right now, the world of CAPTCHAs is tricky. I think AI agents in particular are very much ahead of the internet infrastructure. CAPTCHAs are designed to block all types of bots, but there are now good bots and bad bots. I think in the future, CAPTCHAs will be able to identify who a good bot is, hopefully via some sort of KYC. For us, we've been very lucky. We have very little to no known abuse of Browserbase because we really look into who we work with. And for certain types of CAPTCHA solving, we only allow them on certain types of plans because we want to make sure that we can know what people are doing, what their use cases are. And that's really allowed us to try and be an arbiter of good bots, which is our long term goal. I want to build great relationships with people like Cloudflare so we can agree, hey, here are these acceptable bots. We'll identify them for you and make sure we flag when they come to your website. This is a good bot, you know?Alessio [00:23:23]: I see. And Cloudflare said they want to do more of this. So they're going to set by default, if they think you're an AI bot, they're going to reject. I'm curious if you think this is something that is going to be at the browser level or I mean, the DNS level with Cloudflare seems more where it should belong. But I'm curious how you think about it.Paul [00:23:40]: I think the web's going to change. You know, I think that the Internet as we have it right now is going to change. And we all need to just accept that the cat is out of the bag. And instead of kind of like wishing the Internet was like it was in the 2000s, we can have free content line that wouldn't be scraped. It's just it's not going to happen. And instead, we should think about like, one, how can we change? How can we change the models of, you know, information being published online so people can adequately commercialize it? But two, how do we rebuild applications that expect that AI agents are going to log in on their behalf? Those are the things that are going to allow us to kind of like identify good and bad bots. And I think the team at Clerk has been doing a really good job with this on the authentication side. I actually think that auth is the biggest thing that will prevent agents from accessing stuff, not captchas. And I think there will be agent auth in the future. I don't know if it's going to happen from an individual company, but actually authentication providers that have a, you know, hidden login as agent feature, which will then you put in your email, you'll get a push notification, say like, hey, your browser-based agent wants to log into your Airbnb. You can approve that and then the agent can proceed. That really circumvents the need for captchas or logging in as you and sharing your password. I think agent auth is going to be one way we identify good bots going forward. And I think a lot of this captcha solving stuff is really short-term problems as the internet kind of reorients itself around how it's going to work with agents browsing the web, just like people do. Yeah.Managing Distributed Browser Locations and Proxiesswyx [00:24:59]: Stitch recently was on Hacker News for talking about agent experience, AX, which is a thing that Netlify is also trying to clone and coin and talk about. And we've talked about this on our previous episodes before in a sense that I actually think that's like maybe the only part of the tech stack that needs to be kind of reinvented for agents. Everything else can stay the same, CLIs, APIs, whatever. But auth, yeah, we need agent auth. And it's mostly like short-lived, like it should not, it should be a distinct, identity from the human, but paired. I almost think like in the same way that every social network should have your main profile and then your alt accounts or your Finsta, it's almost like, you know, every, every human token should be paired with the agent token and the agent token can go and do stuff on behalf of the human token, but not be presumed to be the human. Yeah.Paul [00:25:48]: It's like, it's, it's actually very similar to OAuth is what I'm thinking. And, you know, Thread from Stitch is an investor, Colin from Clerk, Octaventures, all investors in browser-based because like, I hope they solve this because they'll make browser-based submission more possible. So we don't have to overcome all these hurdles, but I think it will be an OAuth-like flow where an agent will ask to log in as you, you'll approve the scopes. Like it can book an apartment on Airbnb, but it can't like message anybody. And then, you know, the agent will have some sort of like role-based access control within an application. Yeah. I'm excited for that.swyx [00:26:16]: The tricky part is just, there's one, one layer of delegation here, which is like, you're authoring my user's user or something like that. I don't know if that's tricky or not. Does that make sense? Yeah.Paul [00:26:25]: You know, actually at Twilio, I worked on the login identity and access. Management teams, right? So like I built Twilio's login page.swyx [00:26:31]: You were an intern on that team and then you became the lead in two years? Yeah.Paul [00:26:34]: Yeah. I started as an intern in 2016 and then I was the tech lead of that team. How? That's not normal. I didn't have a life. He's not normal. Look at this guy. I didn't have a girlfriend. I just loved my job. I don't know. I applied to 500 internships for my first job and I got rejected from every single one of them except for Twilio and then eventually Amazon. And they took a shot on me and like, I was getting paid money to write code, which was my dream. Yeah. Yeah. I'm very lucky that like this coding thing worked out because I was going to be doing it regardless. And yeah, I was able to kind of spend a lot of time on a team that was growing at a company that was growing. So it informed a lot of this stuff here. I think these are problems that have been solved with like the SAML protocol with SSO. I think it's a really interesting stuff with like WebAuthn, like these different types of authentication, like schemes that you can use to authenticate people. The tooling is all there. It just needs to be tweaked a little bit to work for agents. And I think the fact that there are companies that are already. Providing authentication as a service really sets it up. Well, the thing that's hard is like reinventing the internet for agents. We don't want to rebuild the internet. That's an impossible task. And I think people often say like, well, we'll have this second layer of APIs built for agents. I'm like, we will for the top use cases, but instead of we can just tweak the internet as is, which is on the authentication side, I think we're going to be the dumb ones going forward. Unfortunately, I think AI is going to be able to do a lot of the tasks that we do online, which means that it will be able to go to websites, click buttons on our behalf and log in on our behalf too. So with this kind of like web agent future happening, I think with some small structural changes, like you said, it feels like it could all slot in really nicely with the existing internet.Handling CAPTCHAs and Agent Authenticationswyx [00:28:08]: There's one more thing, which is the, your live view iframe, which lets you take, take control. Yeah. Obviously very key for operator now, but like, was, is there anything interesting technically there or that the people like, well, people always want this.Paul [00:28:21]: It was really hard to build, you know, like, so, okay. Headless browsers, you don't see them, right. They're running. They're running in a cloud somewhere. You can't like look at them. And I just want to really make, it's a weird name. I wish we came up with a better name for this thing, but you can't see them. Right. But customers don't trust AI agents, right. At least the first pass. So what we do with our live view is that, you know, when you use browser base, you can actually embed a live view of the browser running in the cloud for your customer to see it working. And that's what the first reason is the build trust, like, okay, so I have this script. That's going to go automate a website. I can embed it into my web application via an iframe and my customer can watch. I think. And then we added two way communication. So now not only can you watch the browser kind of being operated by AI, if you want to pause and actually click around type within this iframe that's controlling a browser, that's also possible. And this is all thanks to some of the lower level protocol, which is called the Chrome DevTools protocol. It has a API called start screencast, and you can also send mouse clicks and button clicks to a remote browser. And this is all embeddable within iframes. You have a browser within a browser, yo. And then you simulate the screen, the click on the other side. Exactly. And this is really nice often for, like, let's say, a capture that can't be solved. You saw this with Operator, you know, Operator actually uses a different approach. They use VNC. So, you know, you're able to see, like, you're seeing the whole window here. What we're doing is something a little lower level with the Chrome DevTools protocol. It's just PNGs being streamed over the wire. But the same thing is true, right? Like, hey, I'm running a window. Pause. Can you do something in this window? Human. Okay, great. Resume. Like sometimes 2FA tokens. Like if you get that text message, you might need a person to type that in. Web agents need human-in-the-loop type workflows still. You still need a person to interact with the browser. And building a UI to proxy that is kind of hard. You may as well just show them the whole browser and say, hey, can you finish this up for me? And then let the AI proceed on afterwards. Is there a future where I stream my current desktop to browser base? I don't think so. I think we're very much cloud infrastructure. Yeah. You know, but I think a lot of the stuff we're doing, we do want to, like, build tools. Like, you know, we'll talk about the stage and, you know, web agent framework in a second. But, like, there's a case where a lot of people are going desktop first for, you know, consumer use. And I think cloud is doing a lot of this, where I expect to see, you know, MCPs really oriented around the cloud desktop app for a reason, right? Like, I think a lot of these tools are going to run on your computer because it makes... I think it's breaking out. People are putting it on a server. Oh, really? Okay. Well, sweet. We'll see. We'll see that. I was surprised, though, wasn't I? I think that the browser company, too, with Dia Browser, it runs on your machine. You know, it's going to be...swyx [00:30:50]: What is it?Paul [00:30:51]: So, Dia Browser, as far as I understand... I used to use Arc. Yeah. I haven't used Arc. But I'm a big fan of the browser company. I think they're doing a lot of cool stuff in consumer. As far as I understand, it's a browser where you have a sidebar where you can, like, chat with it and it can control the local browser on your machine. So, if you imagine, like, what a consumer web agent is, which it lives alongside your browser, I think Google Chrome has Project Marina, I think. I almost call it Project Marinara for some reason. I don't know why. It's...swyx [00:31:17]: No, I think it's someone really likes the Waterworld. Oh, I see. The classic Kevin Costner. Yeah.Paul [00:31:22]: Okay. Project Marinara is a similar thing to the Dia Browser, in my mind, as far as I understand it. You have a browser that has an AI interface that will take over your mouse and keyboard and control the browser for you. Great for consumer use cases. But if you're building applications that rely on a browser and it's more part of a greater, like, AI app experience, you probably need something that's more like infrastructure, not a consumer app.swyx [00:31:44]: Just because I have explored a little bit in this area, do people want branching? So, I have the state. Of whatever my browser's in. And then I want, like, 100 clones of this state. Do people do that? Or...Paul [00:31:56]: People don't do it currently. Yeah. But it's definitely something we're thinking about. I think the idea of forking a browser is really cool. Technically, kind of hard. We're starting to see this in code execution, where people are, like, forking some, like, code execution, like, processes or forking some tool calls or branching tool calls. Haven't seen it at the browser level yet. But it makes sense. Like, if an AI agent is, like, using a website and it's not sure what path it wants to take to crawl this website. To find the information it's looking for. It would make sense for it to explore both paths in parallel. And that'd be a very, like... A road not taken. Yeah. And hopefully find the right answer. And then say, okay, this was actually the right one. And memorize that. And go there in the future. On the roadmap. For sure. Don't make my roadmap, please. You know?Alessio [00:32:37]: How do you actually do that? Yeah. How do you fork? I feel like the browser is so stateful for so many things.swyx [00:32:42]: Serialize the state. Restore the state. I don't know.Paul [00:32:44]: So, it's one of the reasons why we haven't done it yet. It's hard. You know? Like, to truly fork, it's actually quite difficult. The naive way is to open the same page in a new tab and then, like, hope that it's at the same thing. But if you have a form halfway filled, you may have to, like, take the whole, you know, container. Pause it. All the memory. Duplicate it. Restart it from there. It could be very slow. So, we haven't found a thing. Like, the easy thing to fork is just, like, copy the page object. You know? But I think there needs to be something a little bit more robust there. Yeah.swyx [00:33:12]: So, MorphLabs has this infinite branch thing. Like, wrote a custom fork of Linux or something that let them save the system state and clone it. MorphLabs, hit me up. I'll be a customer. Yeah. That's the only. I think that's the only way to do it. Yeah. Like, unless Chrome has some special API for you. Yeah.Paul [00:33:29]: There's probably something we'll reverse engineer one day. I don't know. Yeah.Alessio [00:33:32]: Let's talk about StageHand, the AI web browsing framework. You have three core components, Observe, Extract, and Act. Pretty clean landing page. What was the idea behind making a framework? Yeah.Stagehand: AI web browsing frameworkPaul [00:33:43]: So, there's three frameworks that are very popular or already exist, right? Puppeteer, Playwright, Selenium. Those are for building hard-coded scripts to control websites. And as soon as I started to play with LLMs plus browsing, I caught myself, you know, code-genning Playwright code to control a website. I would, like, take the DOM. I'd pass it to an LLM. I'd say, can you generate the Playwright code to click the appropriate button here? And it would do that. And I was like, this really should be part of the frameworks themselves. And I became really obsessed with SDKs that take natural language as part of, like, the API input. And that's what StageHand is. StageHand exposes three APIs, and it's a super set of Playwright. So, if you go to a page, you may want to take an action, click on the button, fill in the form, etc. That's what the act command is for. You may want to extract some data. This one takes a natural language, like, extract the winner of the Super Bowl from this page. You can give it a Zod schema, so it returns a structured output. And then maybe you're building an API. You can do an agent loop, and you want to kind of see what actions are possible on this page before taking one. You can do observe. So, you can observe the actions on the page, and it will generate a list of actions. You can guide it, like, give me actions on this page related to buying an item. And you can, like, buy it now, add to cart, view shipping options, and pass that to an LLM, an agent loop, to say, what's the appropriate action given this high-level goal? So, StageHand isn't a web agent. It's a framework for building web agents. And we think that agent loops are actually pretty close to the application layer because every application probably has different goals or different ways it wants to take steps. I don't think I've seen a generic. Maybe you guys are the experts here. I haven't seen, like, a really good AI agent framework here. Everyone kind of has their own special sauce, right? I see a lot of developers building their own agent loops, and they're using tools. And I view StageHand as the browser tool. So, we expose act, extract, observe. Your agent can call these tools. And from that, you don't have to worry about it. You don't have to worry about generating playwright code performantly. You don't have to worry about running it. You can kind of just integrate these three tool calls into your agent loop and reliably automate the web.swyx [00:35:48]: A special shout-out to Anirudh, who I met at your dinner, who I think listens to the pod. Yeah. Hey, Anirudh.Paul [00:35:54]: Anirudh's a man. He's a StageHand guy.swyx [00:35:56]: I mean, the interesting thing about each of these APIs is they're kind of each startup. Like, specifically extract, you know, Firecrawler is extract. There's, like, Expand AI. There's a whole bunch of, like, extract companies. They just focus on extract. I'm curious. Like, I feel like you guys are going to collide at some point. Like, right now, it's friendly. Everyone's in a blue ocean. At some point, it's going to be valuable enough that there's some turf battle here. I don't think you have a dog in a fight. I think you can mock extract to use an external service if they're better at it than you. But it's just an observation that, like, in the same way that I see each option, each checkbox in the side of custom GBTs becoming a startup or each box in the Karpathy chart being a startup. Like, this is also becoming a thing. Yeah.Paul [00:36:41]: I mean, like, so the way StageHand works is that it's MIT-licensed, completely open source. You bring your own API key to your LLM of choice. You could choose your LLM. We don't make any money off of the extract or really. We only really make money if you choose to run it with our browser. You don't have to. You can actually use your own browser, a local browser. You know, StageHand is completely open source for that reason. And, yeah, like, I think if you're building really complex web scraping workflows, I don't know if StageHand is the tool for you. I think it's really more if you're building an AI agent that needs a few general tools or if it's doing a lot of, like, web automation-intensive work. But if you're building a scraping company, StageHand is not your thing. You probably want something that's going to, like, get HTML content, you know, convert that to Markdown, query it. That's not what StageHand does. StageHand is more about reliability. I think we focus a lot on reliability and less so on cost optimization and speed at this point.swyx [00:37:33]: I actually feel like StageHand, so the way that StageHand works, it's like, you know, page.act, click on the quick start. Yeah. It's kind of the integration test for the code that you would have to write anyway, like the Puppeteer code that you have to write anyway. And when the page structure changes, because it always does, then this is still the test. This is still the test that I would have to write. Yeah. So it's kind of like a testing framework that doesn't need implementation detail.Paul [00:37:56]: Well, yeah. I mean, Puppeteer, Playwright, and Slenderman were all designed as testing frameworks, right? Yeah. And now people are, like, hacking them together to automate the web. I would say, and, like, maybe this is, like, me being too specific. But, like, when I write tests, if the page structure changes. Without me knowing, I want that test to fail. So I don't know if, like, AI, like, regenerating that. Like, people are using StageHand for testing. But it's more for, like, usability testing, not, like, testing of, like, does the front end, like, has it changed or not. Okay. But generally where we've seen people, like, really, like, take off is, like, if they're using, you know, something. If they want to build a feature in their application that's kind of like Operator or Deep Research, they're using StageHand to kind of power that tool calling in their own agent loop. Okay. Cool.swyx [00:38:37]: So let's go into Operator, the first big agent launch of the year from OpenAI. Seems like they have a whole bunch scheduled. You were on break and your phone blew up. What's your just general view of computer use agents is what they're calling it. The overall category before we go into Open Operator, just the overall promise of Operator. I will observe that I tried it once. It was okay. And I never tried it again.OpenAI's Operator and computer use agentsPaul [00:38:58]: That tracks with my experience, too. Like, I'm a huge fan of the OpenAI team. Like, I think that I do not view Operator as the company. I'm not a company killer for browser base at all. I think it actually shows people what's possible. I think, like, computer use models make a lot of sense. And I'm actually most excited about computer use models is, like, their ability to, like, really take screenshots and reasoning and output steps. I think that using mouse click or mouse coordinates, I've seen that proved to be less reliable than I would like. And I just wonder if that's the right form factor. What we've done with our framework is anchor it to the DOM itself, anchor it to the actual item. So, like, if it's clicking on something, it's clicking on that thing, you know? Like, it's more accurate. No matter where it is. Yeah, exactly. Because it really ties in nicely. And it can handle, like, the whole viewport in one go, whereas, like, Operator can only handle what it sees. Can you hover? Is hovering a thing that you can do? I don't know if we expose it as a tool directly, but I'm sure there's, like, an API for hovering. Like, move mouse to this position. Yeah, yeah, yeah. I think you can trigger hover, like, via, like, the JavaScript on the DOM itself. But, no, I think, like, when we saw computer use, everyone's eyes lit up because they realized, like, wow, like, AI is going to actually automate work for people. And I think seeing that kind of happen from both of the labs, and I'm sure we're going to see more labs launch computer use models, I'm excited to see all the stuff that people build with it. I think that I'd love to see computer use power, like, controlling a browser on browser base. And I think, like, Open Operator, which was, like, our open source version of OpenAI's Operator, was our first take on, like, how can we integrate these models into browser base? And we handle the infrastructure and let the labs do the models. I don't have a sense that Operator will be released as an API. I don't know. Maybe it will. I'm curious to see how well that works because I think it's going to be really hard for a company like OpenAI to do things like support CAPTCHA solving or, like, have proxies. Like, I think it's hard for them structurally. Imagine this New York Times headline, OpenAI CAPTCHA solving. Like, that would be a pretty bad headline, this New York Times headline. Browser base solves CAPTCHAs. No one cares. No one cares. And, like, our investors are bored. Like, we're all okay with this, you know? We're building this company knowing that the CAPTCHA solving is short-lived until we figure out how to authenticate good bots. I think it's really hard for a company like OpenAI, who has this brand that's so, so good, to balance with, like, the icky parts of web automation, which it can be kind of complex to solve. I'm sure OpenAI knows who to call whenever they need you. Yeah, right. I'm sure they'll have a great partnership.Alessio [00:41:23]: And is Open Operator just, like, a marketing thing for you? Like, how do you think about resource allocation? So, you can spin this up very quickly. And now there's all this, like, open deep research, just open all these things that people are building. We started it, you know. You're the original Open. We're the original Open operator, you know? Is it just, hey, look, this is a demo, but, like, we'll help you build out an actual product for yourself? Like, are you interested in going more of a product route? That's kind of the OpenAI way, right? They started as a model provider and then…Paul [00:41:53]: Yeah, we're not interested in going the product route yet. I view Open Operator as a model provider. It's a reference project, you know? Let's show people how to build these things using the infrastructure and models that are out there. And that's what it is. It's, like, Open Operator is very simple. It's an agent loop. It says, like, take a high-level goal, break it down into steps, use tool calling to accomplish those steps. It takes screenshots and feeds those screenshots into an LLM with the step to generate the right action. It uses stagehand under the hood to actually execute this action. It doesn't use a computer use model. And it, like, has a nice interface using the live view that we talked about, the iframe, to embed that into an application. So I felt like people on launch day wanted to figure out how to build their own version of this. And we turned that around really quickly to show them. And I hope we do that with other things like deep research. We don't have a deep research launch yet. I think David from AOMNI actually has an amazing open deep research that he launched. It has, like, 10K GitHub stars now. So he's crushing that. But I think if people want to build these features natively into their application, they need good reference projects. And I think Open Operator is a good example of that.swyx [00:42:52]: I don't know. Actually, I'm actually pretty bullish on API-driven operator. Because that's the only way that you can sort of, like, once it's reliable enough, obviously. And now we're nowhere near. But, like, give it five years. It'll happen, you know. And then you can sort of spin this up and browsers are working in the background and you don't necessarily have to know. And it just is booking restaurants for you, whatever. I can definitely see that future happening. I had this on the landing page here. This might be a slightly out of order. But, you know, you have, like, sort of three use cases for browser base. Open Operator. Or this is the operator sort of use case. It's kind of like the workflow automation use case. And it completes with UiPath in the sort of RPA category. Would you agree with that? Yeah, I would agree with that. And then there's Agents we talked about already. And web scraping, which I imagine would be the bulk of your workload right now, right?Paul [00:43:40]: No, not at all. I'd say actually, like, the majority is browser automation. We're kind of expensive for web scraping. Like, I think that if you're building a web scraping product, if you need to do occasional web scraping or you have to do web scraping that works every single time, you want to use browser automation. Yeah. You want to use browser-based. But if you're building web scraping workflows, what you should do is have a waterfall. You should have the first request is a curl to the website. See if you can get it without even using a browser. And then the second request may be, like, a scraping-specific API. There's, like, a thousand scraping APIs out there that you can use to try and get data. Scraping B. Scraping B is a great example, right? Yeah. And then, like, if those two don't work, bring out the heavy hitter. Like, browser-based will 100% work, right? It will load the page in a real browser, hydrate it. I see.swyx [00:44:21]: Because a lot of people don't render to JS.swyx [00:44:25]: Yeah, exactly.Paul [00:44:26]: So, I mean, the three big use cases, right? Like, you know, automation, web data collection, and then, you know, if you're building anything agentic that needs, like, a browser tool, you want to use browser-based.Alessio [00:44:35]: Is there any use case that, like, you were super surprised by that people might not even think about? Oh, yeah. Or is it, yeah, anything that you can share? The long tail is crazy. Yeah.Surprising use cases of BrowserbasePaul [00:44:44]: One of the case studies on our website that I think is the most interesting is this company called Benny. So, the way that it works is if you're on food stamps in the United States, you can actually get rebates if you buy certain things. Yeah. You buy some vegetables. You submit your receipt to the government. They'll give you a little rebate back. Say, hey, thanks for buying vegetables. It's good for you. That process of submitting that receipt is very painful. And the way Benny works is you use their app to take a photo of your receipt, and then Benny will go submit that receipt for you and then deposit the money into your account. That's actually using no AI at all. It's all, like, hard-coded scripts. They maintain the scripts. They've been doing a great job. And they build this amazing consumer app. But it's an example of, like, all these, like, tedious workflows that people have to do to kind of go about their business. And they're doing it for the sake of their day-to-day lives. And I had never known about, like, food stamp rebates or the complex forms you have to do to fill them. But the world is powered by millions and millions of tedious forms, visas. You know, Emirate Lighthouse is a customer, right? You know, they do the O1 visa. Millions and millions of forms are taking away humans' time. And I hope that Browserbase can help power software that automates away the web forms that we don't need anymore. Yeah.swyx [00:45:49]: I mean, I'm very supportive of that. I mean, forms. I do think, like, government itself is a big part of it. I think the government itself should embrace AI more to do more sort of human-friendly form filling. Mm-hmm. But I'm not optimistic. I'm not holding my breath. Yeah. We'll see. Okay. I think I'm about to zoom out. I have a little brief thing on computer use, and then we can talk about founder stuff, which is, I tend to think of developer tooling markets in impossible triangles, where everyone starts in a niche, and then they start to branch out. So I already hinted at a little bit of this, right? We mentioned more. We mentioned E2B. We mentioned Firecrawl. And then there's Browserbase. So there's, like, all this stuff of, like, have serverless virtual computer that you give to an agent and let them do stuff with it. And there's various ways of connecting it to the internet. You can just connect to a search API, like SERP API, whatever other, like, EXA is another one. That's what you're searching. You can also have a JSON markdown extractor, which is Firecrawl. Or you can have a virtual browser like Browserbase, or you can have a virtual machine like Morph. And then there's also maybe, like, a virtual sort of code environment, like Code Interpreter. So, like, there's just, like, a bunch of different ways to tackle the problem of give a computer to an agent. And I'm just kind of wondering if you see, like, everyone's just, like, happily coexisting in their respective niches. And as a developer, I just go and pick, like, a shopping basket of one of each. Or do you think that you eventually, people will collide?Future of browser automation and market competitionPaul [00:47:18]: I think that currently it's not a zero-sum market. Like, I think we're talking about... I think we're talking about all of knowledge work that people do that can be automated online. All of these, like, trillions of hours that happen online where people are working. And I think that there's so much software to be built that, like, I tend not to think about how these companies will collide. I just try to solve the problem as best as I can and make this specific piece of infrastructure, which I think is an important primitive, the best I possibly can. And yeah. I think there's players that are actually going to like it. I think there's players that are going to launch, like, over-the-top, you know, platforms, like agent platforms that have all these tools built in, right? Like, who's building the rippling for agent tools that has the search tool, the browser tool, the operating system tool, right? There are some. There are some. There are some, right? And I think in the end, what I have seen as my time as a developer, and I look at all the favorite tools that I have, is that, like, for tools and primitives with sufficient levels of complexity, you need to have a solution that's really bespoke to that primitive, you know? And I am sufficiently convinced that the browser is complex enough to deserve a primitive. Obviously, I have to. I'm the founder of BrowserBase, right? I'm talking my book. But, like, I think maybe I can give you one spicy take against, like, maybe just whole OS running. I think that when I look at computer use when it first came out, I saw that the majority of use cases for computer use were controlling a browser. And do we really need to run an entire operating system just to control a browser? I don't think so. I don't think that's necessary. You know, BrowserBase can run browsers for way cheaper than you can if you're running a full-fledged OS with a GUI, you know, operating system. And I think that's just an advantage of the browser. It is, like, browsers are little OSs, and you can run them very efficiently if you orchestrate it well. And I think that allows us to offer 90% of the, you know, functionality in the platform needed at 10% of the cost of running a full OS. Yeah.Open Operator: Browserbase's Open-Source Alternativeswyx [00:49:16]: I definitely see the logic in that. There's a Mark Andreessen quote. I don't know if you know this one. Where he basically observed that the browser is turning the operating system into a poorly debugged set of device drivers, because most of the apps are moved from the OS to the browser. So you can just run browsers.Paul [00:49:31]: There's a place for OSs, too. Like, I think that there are some applications that only run on Windows operating systems. And Eric from pig.dev in this upcoming YC batch, or last YC batch, like, he's building all run tons of Windows operating systems for you to control with your agent. And like, there's some legacy EHR systems that only run on Internet-controlled systems. Yeah.Paul [00:49:54]: I think that's it. I think, like, there are use cases for specific operating systems for specific legacy software. And like, I'm excited to see what he does with that. I just wanted to give a shout out to the pig.dev website.swyx [00:50:06]: The pigs jump when you click on them. Yeah. That's great.Paul [00:50:08]: Eric, he's the former co-founder of banana.dev, too.swyx [00:50:11]: Oh, that Eric. Yeah. That Eric. Okay. Well, he abandoned bananas for pigs. I hope he doesn't start going around with pigs now.Alessio [00:50:18]: Like he was going around with bananas. A little toy pig. Yeah. Yeah. I love that. What else are we missing? I think we covered a lot of, like, the browser-based product history, but. What do you wish people asked you? Yeah.Paul [00:50:29]: I wish people asked me more about, like, what will the future of software look like? Because I think that's really where I've spent a lot of time about why do browser-based. Like, for me, starting a company is like a means of last resort. Like, you shouldn't start a company unless you absolutely have to. And I remain convinced that the future of software is software that you're going to click a button and it's going to do stuff on your behalf. Right now, software. You click a button and it maybe, like, calls it back an API and, like, computes some numbers. It, like, modifies some text, whatever. But the future of software is software using software. So, I may log into my accounting website for my business, click a button, and it's going to go load up my Gmail, search my emails, find the thing, upload the receipt, and then comment it for me. Right? And it may use it using APIs, maybe a browser. I don't know. I think it's a little bit of both. But that's completely different from how we've built software so far. And that's. I think that future of software has different infrastructure requirements. It's going to require different UIs. It's going to require different pieces of infrastructure. I think the browser infrastructure is one piece that fits into that, along with all the other categories you mentioned. So, I think that it's going to require developers to think differently about how they've built software for, you know

SAGE Clinical Medicine & Research
JHVS: The role of multimodality imaging in aortic valve assessment: A review

SAGE Clinical Medicine & Research

Play Episode Listen Later Feb 21, 2025 2:15


Read the article here: https://journals.sagepub.com/doi/full/10.1177/30494826241296390

Oncology Brothers
Intermediate HCC – the evolving role of Immunotherapy with Multimodality approaches

Oncology Brothers

Play Episode Listen Later Jan 27, 2025 27:12


In this final episode of the four-part series on hepatocellular carcinoma (HCC), hosted by the Oncology Brothers, Drs Rohit and Rahul Gosain, the discussion focuses on the evolving role of immunotherapy (IO) in intermediate HCC. The episode explores multimodal approaches that combine IO and IO-based therapies with loco-regional treatments and highlights the essential role of a multidisciplinary care team. Drs Nina Sanford (radiation oncologist), Mark Yarchoan (medical oncologist), and Ed Kim (interventional radiologist) join the Oncology Brothers to share their insights on: • Current treatment options for intermediate HCC, addressing its heterogeneity and standard treatment pathways • Latest clinical trial data (EMERALD-1, LEAP-012) on combining IO with loco-regional therapies, and the clinical implications • The importance of effective collaboration within the multidisciplinary team for delivering optimal patient care • Combining IO with loco-regional therapy and future perspectives in the field Clinical takeaways • IO and IO-based treatments are moving earlier in the treatment paradigm for patients with intermediate HCC. Earlier integration of these therapies aims to achieve improved systemic control, allowing loco-regional therapy to target oligoprogression, residual lesions or reduce tumour burden • Emerging data supports combining systemic and loco-regional therapies for patients with intermediate HCC. EMERALD-1 and LEAP-012 show promising PFS data using IO-based combination regimens like durvalumab + bevacizumab or pembrolizumab + lenvatinib alongside TACE. Long-term OS data are awaited • Effective communication and coordinated care among specialists, such as medical oncologists, radiation oncologists, hepatologists, and interventional radiologists, are essential to developing optimal treatment strategies for patients with intermediate HCC Follow us on social media: •⁠ ⁠X/Twitter: https://twitter.com/oncbrothers •⁠ ⁠Instagram: https://www.instagram.com/oncbrothers •⁠ YouTube: https://www.youtube.com/channel/UCjfxKlVho5xWH5ltufj4F4A/ Subscribe to our channel for more insights on oncology treatments and patient care!

HODLong 后浪
Ep.47 [EN]: Weekee Tiew: Virtuals is here for years, and let's first grow the agent performance

HODLong 后浪

Play Episode Listen Later Jan 19, 2025 44:15


In this episode I interviewed Wee Kee, the co-founder of Virtuals. From here and there in the conversation you can tell that Virtuals team is a pragmatic and perseverant one. To build a successful product in this fast moving industry, there's no fast track but stay on the table and keep grinding. Shownotes:You guys have been dominating mindshare amongst all the crypto native folks lately. So first of all congratulations! But still would love to hear how you introduce yourself to our audience. There were quite many platforms that empower developers to launch agents. What key factors do you think led you to where you guys are today?Do you think the diversity of the agents came from the no-code requirement?What are some of the agents that grew from Virtuals that left people an impression? Also about Multimodality. What are the best performing categories for agents developed with GAME?How do you decide whether you're going to co-marketing with one of the ecosystem agents?I read on Messari that you guys have integrated Farcaster, but I haven't seen a lot of news that you guys are pushing on that front. Any plans? I was particularly interested to hear your answer on this one, because looks like with a web3 native social network, an agent developed under GAME framework will be capable of doing a lot more

Grow Everything Biotech Podcast
112. The Baked-In Future of Science: Nick Edwards on Potato AI's Quest to Become Your AI Scientist

Grow Everything Biotech Podcast

Play Episode Listen Later Jan 17, 2025 50:08


Erum and Karl have an incredible chat with Nick Edwards, the innovative mind behind Potato AI. Nick goes into the inspiration behind the platform's unique name, rooted in childhood curiosity and scientific wonder, and shares his journey from neuroscience research to building a groundbreaking tool aimed at accelerating scientific discovery. He explains how Potato tackles the overwhelming deluge of scientific literature and the reproducibility crisis using AI to structure and analyze data in transformative ways. With stories from his own career, a stint in consulting, and running his podcast "Once a Scientist," Nick discusses the future of AI, science, and collaboration. This episode has amazing insights into how technology is reshaping the very fabric of research. Grow Everything brings the bioeconomy to life. Hosts Karl Schmieder and Erum Azeez Khan share stories and interview the leaders and influencers changing the world by growing everything. Biology is the oldest technology. And it can be engineered. What are we growing? Learn more at ⁠⁠⁠⁠www.messaginglab.com/groweverything⁠⁠⁠⁠ Chapters: 00:00:00 - Potatoes, Balloons, and a DIY Lab Setup 00:04:56 - Reproducibility Crisis: 70% of Research in Question 00:06:40 - Meet Nick Edwards: The Potato Visionary 00:10:00 - Organizing Scientific Chaos: How AI Helps 00:14:47 - AI Research Assistant vs. AI Scientist: The Journey 00:20:09 - From Neuroscience to Entrepreneurship: A Scientist's Leap 00:25:00 - Biopharma Meets Bioindustrials: Bridging Two Worlds 00:29:27 - Liquid Handlers and Shared Protocols: Optimizing Biotech 00:35:13 - Multimodality and the AI Scientist Dream 00:39:04 - Citizen Scientists and Accessible AI Tools 00:45:00 - The Next Scientific Revolution: AI and Decentralized Science Episode Links: Potato AI Once A Scientist Podcast Merck Digital Science Studio Wiley Elsevier Ginkgo Bioworks Topics Covered: Research, AI research assistant, AI Scientist, Biotech, Lab Automation, Reproducibility Have a question or comment? Message us here: Text or Call (804) 505-5553 ⁠⁠⁠⁠Instagram⁠⁠⁠⁠  / ⁠⁠⁠⁠Twitter⁠⁠⁠⁠ / ⁠⁠⁠⁠LinkedIn⁠⁠⁠⁠ / ⁠⁠⁠⁠Youtube⁠⁠⁠⁠ / ⁠⁠⁠⁠Grow Everything⁠⁠⁠⁠ Email: groweverything@messaginglab.com Music by: Nihilore Production by: Amplafy Media

Scouting Frontiers in AI for Biology: Dynamics, Diffusion, and Design, with Amelie Schreiber

Play Episode Listen Later Dec 14, 2024 107:28


Nathan welcomes back computational biochemist Amelie Schreiber for a fascinating update on AI's revolutionary impact in biology. In this episode of The Cognitive Revolution, we explore recent breakthroughs including AlphaFold3, ESM3, and new diffusion models transforming protein engineering and drug discovery. Join us for an insightful discussion about how AI is reshaping our understanding of molecular biology and making complex protein engineering tasks more accessible than ever before. Help shape our show by taking our quick listener survey at https://bit.ly/TurpentinePulse SPONSORS: Shopify: Shopify is the world's leading e-commerce platform, offering a market-leading checkout system and exclusive AI apps like Quikly. Nobody does selling better than Shopify. Get a $1 per month trial at https://shopify.com/cognitive SelectQuote: Finding the right life insurance shouldn't be another task you put off. SelectQuote compares top-rated policies to get you the best coverage at the right price. Even in our AI-driven world, protecting your family's future remains essential. Get your personalized quote at https://selectquote.com/cognitive Oracle Cloud Infrastructure (OCI): Oracle's next-generation cloud platform delivers blazing-fast AI and ML performance with 50% less for compute and 80% less for outbound networking compared to other cloud providers13. OCI powers industry leaders with secure infrastructure and application development capabilities. New U.S. customers can get their cloud bill cut in half by switching to OCI before December 31, 2024 at https://oracle.com/cognitive Weights & Biases RAG++: Advanced training for building production-ready RAG applications. Learn from experts to overcome LLM challenges, evaluate systematically, and integrate advanced features. Includes free Cohere credits. Visit https://wandb.me/cr to start the RAG++ course today. CHAPTERS: (00:00:00) Teaser (00:00:46) About the Episode (00:04:30) AI for Biology (00:07:14) David Baker's Impact (00:11:49) AlphaFold 3 & ESM3 (00:16:40) Protein Interaction Prediction (Part 1) (00:16:44) Sponsors: Shopify | SelectQuote (00:19:18) Protein Interaction Prediction (Part 2) (00:31:12) MSAs & Embeddings (Part 1) (00:32:32) Sponsors: Oracle Cloud Infrastructure (OCI) | Weights & Biases RAG++ (00:34:49) MSAs & Embeddings (Part 2) (00:35:57) Beyond Structure Prediction (00:51:13) Dynamics vs. Statics (00:57:24) In-Painting & Use Cases (00:59:48) Workflow & Platforms (01:06:45) Design Process & Success Rates (01:13:23) Ambition & Task Definition (01:19:25) New Models: PepFlow & GeoAB (01:28:23) Flow Matching vs. Diffusion (01:30:42) ESM3 & Multimodality (01:37:10) Summary & Future Directions (01:45:34) Outro SOCIAL LINKS: Website: https://www.cognitiverevolution.ai Twitter (Podcast): https://x.com/cogrev_podcast Twitter (Nathan): https://x.com/labenz LinkedIn: https://www.linkedin.com/in/nathanlabenz/ Youtube: https://www.youtube.com/@CognitiveRevolutionPodcast Apple: https://podcasts.apple.com/de/podcast/the-cognitive-revolution-ai-builders-researchers-and/id1669813431

Gemini Update: Search Grounding, JSON Mode, Code Execution, & More – with Google's Logan Kilpatrick and Shrestha Basu Mallick

Play Episode Listen Later Oct 31, 2024 56:59


Nathan interviews Google product managers Shrestha Basu Mallick and Logan Kilpatrick about the Gemini API and AI Studio. They discuss Google's new grounding feature, allowing Gemini models to access real-time web information via Google search. The conversation explores Gemini's rapid growth, its position in the AI landscape, and Google's competitive strategy. Nathan shares insights from integrating Gemini into his own application and ponders the future of large language model capabilities across providers. Tune in for an in-depth look at Google's AI API product strategy and the latest Gemini features. Be notified early when Turpentine's drops new publication: https://www.turpentine.co/exclusiveaccess SPONSORS: Weights & Biases RAG++: Advanced training for building production-ready RAG applications. Learn from experts to overcome LLM challenges, evaluate systematically, and integrate advanced features. Includes free Cohere credits. Visit https://wandb.me/cr to start the RAG++ course today. Shopify: Shopify is the world's leading e-commerce platform, offering a market-leading checkout system and exclusive AI apps like Quikly. Nobody does selling better than Shopify. Get a $1 per month trial at https://shopify.com/cognitive Notion: Notion offers powerful workflow and automation templates, perfect for streamlining processes and laying the groundwork for AI-driven automation. With Notion AI, you can search across thousands of documents from various platforms, generating highly relevant analysis and content tailored just for you - try it for free at https://notion.com/cognitiverevolution LMNT: LMNT is a zero-sugar electrolyte drink mix that's redefining hydration and performance. Ideal for those who fast or anyone looking to optimize their electrolyte intake. Support the show and get a free sample pack with any purchase at https://drinklmnt.com/tcr CHAPTERS: (00:00:00) About the Show (00:00:53) Sponsors: Weights & Biases RAG++ (00:01:28) About the Episode (00:04:15) Gemini API Growth (00:05:26) Intro to AI Studio (00:07:35) Vertex vs. AI Studio (00:09:33) Developer Adoption (00:14:23) Gemini Use Cases (Part 1) (00:17:41) Sponsors: Shopify | Notion (00:20:01) Gemini Use Cases (Part 2) (00:23:08) Multimodality & Flash (00:26:29) Free Tier & Costs (00:31:43) Inference Costs (00:32:55) Fine-tuning & Vision (00:36:59) Sponsors: LMNT (00:38:04) Search Grounding (00:44:42) Grounding Sources (00:46:58) Competitive Landscape (00:50:36) Design Decisions (00:54:54) Outro SOCIAL LINKS: Website: https://www.cognitiverevolution.ai Twitter (Podcast): https://x.com/cogrev_podcast Twitter (Nathan): https://x.com/labenz LinkedIn: https://www.linkedin.com/in/nathanlabenz/ Youtube: https://www.youtube.com/@CognitiveRevolutionPodcast Apple: https://podcasts.apple.com/de/podcast/the-cognitive-revolution-ai-builders-researchers-and/id1669813431 Spotify: https://open.spotify.com/show/6yHyok3M3BjqzR0VB5MSyk

ASTRO Journals
Multimodality Therapy for Locally Advanced Esophageal Cancer: An ASTRO Clinical Practice Guideline. Part 2.

ASTRO Journals

Play Episode Listen Later Oct 25, 2024 30:00


In this podcast, we discuss the The Society of Thoracic Surgeons/American Society for Radiation Oncology Updated Clinical Practice Guidelines on Multimodality Therapy for Locally Advanced Cancer of the Esophagus or Gastroesophageal Junction. Joining in the discussion are Dr. Stephanie Worrell, Associate Professor and Thoracic Section Chief in the Division of Cardiothoracic Surgery at the University of Arizona College of Medicine, and Dr. Karyn Goodman, Professor and Vice Chair for Research and Quality at the Icahn School of Medicine at Mount Sinai, and Associate Director for Clinical Research at The Tisch Cancer Institute, who served as chair and co-chair of the guideline panel, respectively. Together, we cover important updates and recommendations that incorporate surgical aspects into the multi-disciplinary management of this disease along with practical considerations for everyday practice. Additionally, we discuss in depth the recently presented ESOPEC trial presented at the 2024 ASCO annual meeting and how it has impacted the standard of care for esophageal cancers.

ASTRO Journals
Multimodality Therapy for Locally Advanced Esophageal Cancer: An ASTRO Clinical Practice Guideline. Part 2.

ASTRO Journals

Play Episode Listen Later Oct 25, 2024 30:00


In this podcast, we discuss the The Society of Thoracic Surgeons/American Society for Radiation Oncology Updated Clinical Practice Guidelines on Multimodality Therapy for Locally Advanced Cancer of the Esophagus or Gastroesophageal Junction. Joining in the discussion are Dr. Stephanie Worrell, Associate Professor and Thoracic Section Chief in the Division of Cardiothoracic Surgery at the University of Arizona College of Medicine, and Dr. Karyn Goodman, Professor and Vice Chair for Research and Quality at the Icahn School of Medicine at Mount Sinai, and Associate Director for Clinical Research at The Tisch Cancer Institute, who served as chair and co-chair of the guideline panel, respectively. Together, we cover important updates and recommendations that incorporate surgical aspects into the multi-disciplinary management of this disease along with practical considerations for everyday practice. Additionally, we discuss in depth the recently presented ESOPEC trial presented at the 2024 ASCO annual meeting and how it has impacted the standard of care for esophageal cancers.

ASTRO Journals
Multimodality Therapy for Locally Advanced Esophageal Cancer: An ASTRO Clinical Practice Guideline. Part 1.

ASTRO Journals

Play Episode Listen Later Oct 24, 2024 31:56


In this podcast, we discuss the The Society of Thoracic Surgeons/American Society for Radiation Oncology Updated Clinical Practice Guidelines on Multimodality Therapy for Locally Advanced Cancer of the Esophagus or Gastroesophageal Junction. Joining in the discussion are Dr. Stephanie Worrell, Associate Professor and Thoracic Section Chief in the Division of Cardiothoracic Surgery at the University of Arizona College of Medicine, and Dr. Karyn Goodman, Professor and Vice Chair for Research and Quality at the Icahn School of Medicine at Mount Sinai, and Associate Director for Clinical Research at The Tisch Cancer Institute, who served as chair and co-chair of the guideline panel, respectively. Together, we cover important updates and recommendations that incorporate surgical aspects into the multi-disciplinary management of this disease along with practical considerations for everyday practice. Additionally, we discuss in depth the recently presented ESOPEC trial presented at the 2024 ASCO annual meeting and how it has impacted the standard of care for esophageal cancers.

ASTRO Journals
Multimodality Therapy for Locally Advanced Esophageal Cancer: An ASTRO Clinical Practice Guideline. Part 1.

ASTRO Journals

Play Episode Listen Later Oct 24, 2024 31:56


In this podcast, we discuss the The Society of Thoracic Surgeons/American Society for Radiation Oncology Updated Clinical Practice Guidelines on Multimodality Therapy for Locally Advanced Cancer of the Esophagus or Gastroesophageal Junction. Joining in the discussion are Dr. Stephanie Worrell, Associate Professor and Thoracic Section Chief in the Division of Cardiothoracic Surgery at the University of Arizona College of Medicine, and Dr. Karyn Goodman, Professor and Vice Chair for Research and Quality at the Icahn School of Medicine at Mount Sinai, and Associate Director for Clinical Research at The Tisch Cancer Institute, who served as chair and co-chair of the guideline panel, respectively. Together, we cover important updates and recommendations that incorporate surgical aspects into the multi-disciplinary management of this disease along with practical considerations for everyday practice. Additionally, we discuss in depth the recently presented ESOPEC trial presented at the 2024 ASCO annual meeting and how it has impacted the standard of care for esophageal cancers.

Cardionerds
396. Case Report: Unmasking Constrictive Pericarditis Using Multimodality Imaging – University of Nebraska

Cardionerds

Play Episode Listen Later Oct 21, 2024 37:19


CardioNerds (Dr. Dan Ambinder and Dr. Rick Ferraro) join Dr. Mansi Oberoi and Dr. Mohan Gudiwada from the University of Nebraska Medical Center discuss a case of constrictive pericarditis. Expert commentary is provided by Dr. Adam Burdorf, who serves as the Program Director for the Cardiovascular Medicine Fellowship at the University of Nebraska Medical Center. The case discussed involves a 76-year-old woman with a history of monoclonal gammopathy of undetermined significance, chronic obstructive pulmonary disease, type 2 diabetes mellitus, and squamous cell carcinoma was admitted to the hospital for worsening shortness of breath, swelling in lower extremities, hyponatremia, and urinary tract infection. CT chest to evaluate for pulmonary embolism showed incidental pericardial calcifications; the heart failure team was consulted for the management of her decompensated heart failure. Echo images were nondiagnostic. Subsequent invasive hemodynamic monitoring showed elevated right and left-sided filling pressures, diastolic equalization of LV and RV pressures, and positive RV square root sign with ventricular interdependence. Cardiac MRI showed septal flattening on deep inspiration and septal bounce, suggestive of interventricular dependence. After a heart team discussion and with shared-decision making the patient opted for medical management owing to her comorbidities and frailty. Enjoy this 2024 JACC State-of-the-Art Review to learn more about pericardial diseases and best practices for pericardiectomy (Al-Kazac et al., JACC 2024) US Cardiology Review is now the official journal of CardioNerds! Submit your manuscript here. CardioNerds Case Reports PageCardioNerds Episode PageCardioNerds AcademyCardionerds Healy Honor Roll CardioNerds Journal ClubSubscribe to The Heartbeat Newsletter!Check out CardioNerds SWAG!Become a CardioNerds Patron! Case Media - Constrictive Pericarditis Echo: Left Ventricular ejection fraction = 55-60%. Unclear septal motion in the setting of atrial fibrillation MRI: Diastolic septal flattening with deep inspiration as well as a septal bounce suggestive of interventricular dependence and constrictive physiology  References Garcia, M. Constrictive Pericarditis Versus Restrictive Cardiomyopathy. Journal of the American College of Cardiology, vol. 67, no. 17, 2016, pp. 2061–2076. Pathophysiology and Diagnosis of Constrictive Pericarditis. American College of Cardiology, 2017. Geske, J., Anavekar, N., Nishimura, R., et al. Differentiation of Constriction and Restriction: Complex Cardiovascular Hemodynamics. Journal of the American College of Cardiology, vol. 68, no. 21, 2016, pp. 2329–2347. Constrictive Pericarditis. ScienceDirect. Constrictive Pericarditis. Journal of the American College of Cardiology, vol. 83, no. 12, 2024, pp. 1500-1512.

Generative Now | AI Builders on Creating the Future
Andrew Mason: Craft and Control in AI Content Creation with Descript

Generative Now | AI Builders on Creating the Future

Play Episode Listen Later Oct 17, 2024 40:16


Descript CEO and founder Andrew Mason joins Lightspeed Partner and Host Michael Mignano on the podcast to talk about the future of content creation with AI tools. Michael and Andrew talk about the evolution of Descript as an AI product designed for podcast and video creators, navigating a world of synthetic content, and Descript's new features including Descript Rooms and Underlord. Andrew talks about raising a $50 Million Series C led by OpenAI Startup Fund and how seeing an early version of ChatGPT inspired confidence in Descript's foundational vision and goal to simplify media production.   Episode Chapters (00:00) Introduction(00:09) Introducing Descript Rooms(01:10) Descript's Versatility with Media Creation(04:22) Social Clips and Longform Content(07:15) Craft and Control in AI in Content Creation(13:16) Descript AI Tools, OpenAI, and ChatGPT(17:30) Multimodality and Improving Quality(26:17) Trust and Adoption of AI Features(29:29) Detour and Groupon(37:31) Closing Thoughts Stay in touch: ⁠www.lsvp.com⁠ X: ⁠https://twitter.com/lightspeedvp⁠ LinkedIn: ⁠https://www.linkedin.com/company/lightspeed-venture-partners/⁠ Instagram: ⁠https://www.instagram.com/lightspeedventurepartners/⁠ Subscribe on your favorite podcast app: ⁠generativenow.co⁠ Email: generativenow@lsvp.com The content here does not constitute tax, legal, business or investment advice or an offer to provide such advice, should not be construed as advocating the purchase or sale of any security or investment or a recommendation of any company, and is not an offer, or solicitation of an offer, for the purchase or sale of any security or investment product. For more details please see ⁠lsvp.com/legal⁠.

JACC Podcast
ACC/AHA/ASE/ASNC/HFSA/HRS/SCAI/SCCT/SCMR/STS 2024 Appropriate Use Criteria for Multimodality Imaging in Cardiovascular Evaluation of Patients Undergoing Nonemergent, Noncardiac Surgery

JACC Podcast

Play Episode Listen Later Sep 30, 2024 8:33


In this episode, experts discuss a crucial 2024 document outlining appropriate use criteria for multimodality imaging in cardiovascular evaluation before non-emergent non-cardiac surgery, addressing the rising annual surgeries and associated cardiac risks. They delve into balancing the necessity of imaging with cost-effectiveness while exploring the potential of artificial intelligence to enhance future evaluations.

Lab Rats to Unicorns
Rising Stars: Zachi Attia on Multimodality and AI in Medicine_e.001

Lab Rats to Unicorns

Play Episode Listen Later Aug 23, 2024 39:26


Join Steve Lehmann & Jeremy Langsam from Portal's Stargaze team on a bimonthly segment of Lab Rats to Unicorns: Rising Stars. In this episode, they explore the groundbreaking work of Zachi Attia, the Director of Artificial Intelligence at Mayo Clinic. With a rich background in electrical engineering and a Ph.D. in Bioinformatics, Zachi discusses his pivotal role in advancing AI models that predict and screen cardiovascular diseases. From his innovative research to real-world applications that are saving lives, this episode offers an inspiring look into the future of healthcare.

No Priors: Artificial Intelligence | Machine Learning | Technology | Startups
State Space Models and Real-time Intelligence with Karan Goel and Albert Gu from Cartesia

No Priors: Artificial Intelligence | Machine Learning | Technology | Startups

Play Episode Listen Later Jun 27, 2024 34:08


This week on No Priors, Sarah Guo and Elad Gil sit down with Karan Goel and Albert Gu from Cartesia. Karan and Albert first met as Stanford AI Lab PhDs, where their lab invented Space Models or SSMs, a fundamental new primitive for training large-scale foundation models. In 2023, they Founded Cartesia to build real-time intelligence for every device. One year later, Cartesia released Sonic which generates high quality and lifelike speech with a model latency of 135ms—the fastest for a model of this class. Sign up for new podcasts every week. Email feedback to show@no-priors.com Follow us on Twitter: @NoPriorsPod | @Saranormous | @EladGil | @krandiash | @_albertgu Show Notes:  (0:00) Introduction (0:28) Use Cases for Cartesia and Sonic  (1:32) Karan Goel & Albert Gu's professional backgrounds (5:06) Steady State Models (SSMs) versus Transformer Based Architectures  (11:51) Domain Applications for Hybrid Approaches  (13:10) Text to Speech and Voice (17:29) Data, Size of Models and Efficiency  (20:34) Recent Launch of Text to Speech Product (25:01) Multimodality & Building Blocks (25:54) What's Next at Cartesia?  (28:28) Latency in Text to Speech (29:30) Choosing Research Problems Based on Aesthetic  (31:23) Product Demo (32:48) Cartesia Team & Hiring

Autism Weekly
Communicating with Autistic Children through Bilingualism, Multimodality, and Code Switching with Naima Bhana #163

Autism Weekly

Play Episode Listen Later Mar 29, 2024 36:55


This week, we are joined by Dr. Naima Bhana Lopez, an assistant professor of special education and BCBA-D at Niagara University in New York. Dr. Lopez specializes in enhancing social-communication opportunities for children with developmental disabilities. Her work also explores diversity and equity in special education, empowering natural communication partners to improve outcomes for diverse students and families. With expertise in ABA therapy and special education, Dr. Lopez will be discussing the intersection of these fields with diversity. Download to learn more! Resources  https://www.abainternational.org/media/180194/abai_interprofessional_collaboration_resource_document.pdf   https://aac-learning-center-moodle.psu.edu/   https://abavisualized.com/collections/for-parents/products/aba-en-imagenes-una-guia-visual-para-padres-y-maestros (also available on amazon)   https://abavisualized.com/collections/for-parents/products/aba-visualized-guidebook-2nd-edition (also on Amazon) ................................................................ Autism weekly is now found on all of the major listening apps including apple podcasts, google podcasts, stitcher, Spotify, amazon music, and more. Subscribe to be notified when we post a new podcast. Autism weekly is produced by ABS Kids. ABS Kids is proud to provide diagnostic assessments and ABA therapy to children with developmental delays like Autism Spectrum Disorder. You can learn more about ABS Kids and the Autism Weekly podcast by visiting abskids.com.  

Luke's ENGLISH Podcast - Learn British English with Luke Thompson
867. Multimodal Communication (with Nik Peachy)

Luke's ENGLISH Podcast - Learn British English with Luke Thompson

Play Episode Listen Later Feb 6, 2024 86:18


This episode is all about the different modes of communication that we use beyond the 4 linguistic skills of reading, writing, listening speaking. My guest is Nik Peachy who has helped to write a new paper published by OUP called "Multimodality in ELT: Communication Skills for Today's Generation". Listen to Nik and me chatting about the importance of multimodal literacy in our social interactions and in the ways we consume and produce media online.

Cardionerds
349. Case Report: Into the Thick of It – An Unusual Cause of Hypertrophic Cardiomyopathy – Cleveland Clinic

Cardionerds

Play Episode Listen Later Dec 17, 2023 50:05


CardioNerds cofounder Dr. Amit Goyal and cardiology fellows from the Cleveland Clinic (Drs. Alejandro Duran Crane, Gary Parizher, and Simrat Kaur) discuss the following case: A 61-year-old man presented with symptoms of heart failure and left ventricular hypertrophy. He was given a diagnosis of obstructive hypertrophic cardiomyopathy. He eventually underwent septal myectomy, mitral valve replacement, aortic aneurysm repair, and aortic valve replacement with findings of Fabry's disease on surgical pathology. The case discussion focuses on the differential diagnosis for LVH and covers Fabry disease as an HCM mimic. Expert commentary was provided by Dr. Angelika Ewrin. The episode audio was edited by student Dr. Diane Masket. US Cardiology Review is now the official journal of CardioNerds! Submit your manuscript here. CardioNerds Case Reports PageCardioNerds Episode PageCardioNerds AcademyCardionerds Healy Honor Roll CardioNerds Journal ClubSubscribe to The Heartbeat Newsletter!Check out CardioNerds SWAG!Become a CardioNerds Patron! Case Media - An Unusual Cause of Hypertrophic Cardiomyopathy – Cleveland Clinic Pearls - An Unusual Cause of Hypertrophic Cardiomyopathy – Cleveland Clinic Left ventricular hypertrophy is a cardiac manifestation of several different systemic and cardiac processes, and its etiology should be clarified to avoid missed diagnosis and treatment opportunities. Fabry disease is a rare, X-linked inherited disease that can present cardiac and extra-cardiac manifestations, the former of which include hypertrophic cardiomyopathy, conduction defects, coronary artery disease, conduction abnormalities, arrhythmias, and heart failure.  The diagnosis of Fabry disease includes measurement of alpha-galactosidase enzyme activity as well as genetic testing to evaluate for pathogenic variants or variants of unknown significance in the GLA gene. Family members of patients diagnosed with Fabry disease should be screened based on the inheritance pattern.   Multimodality imaging can be helpful in the diagnosis of Fabry disease. Echocardiography can show left ventricular hypertrophy (LVH), reduced global strain, aortic and mitral valve thickening, and aortic root dilation with associated mild to moderate aortic regurgitation. Cardiac MRI can show hypertrophy of papillary muscles, mid-wall late gadolinium enhancement and low-native T1 signal.   The treatment of Fabry disease involves a multi-disciplinary approach with geneticists, nephrologists, cardiologists, nephrologists, and primary care doctors. Enzyme replacement therapy can delay the progression of cardiac disease.    Show Notes - An Unusual Cause of Hypertrophic Cardiomyopathy – Cleveland Clinic What are the causes of left ventricular hypertrophy? LVH is extremely common. It is present in 15-20% of the general population, and is more common in Black individuals, the elderly, obese or hypertensive individuals, with most cases being secondary to hypertension and aortic valve stenosis. In general terms, it is helpful to divide the causes of LVH into three main groups: high afterload states, obstruction to LV ejection, and intrinsic myocardial problems. Increased afterload states include both primary and secondary hypertension and renal artery stenosis. Mechanical obstruction includes aortic stenosis, subaortic stenosis, and coarctation of the aorta. Lastly, several intrinsic problems of the myocardium can cause LV hypertrophy, such as athletic heart with physiological LVH, hypertrophic cardiomyopathy with or without outflow obstruction, and infiltrative or storage diseases such as cardiac amyloidosis, Fabry's disease, or Danon disease, among others.  How does Fabry disease present? Fabry disease is present in all races and is an X-linked lysosomal storage disorder caused by pathogenic variants in the GLA gene that result in reduced alpha-galactosidase enzyme activity,