POPULARITY
Categories
What does maintainability look like when the software you're responsible for is a programming language? José Valim creator of Elixir and founder of Dashbit, joins Robby to talk about software that can “fit in your head,” the importance of clear boundaries, and why maintaining a language requires a different level of restraint than maintaining an application.José traces Elixir's origins back to his time on the Rails Core team and a concurrency bug that sent him looking for different approaches to software. That journey led through functional programming to Erlang and the BEAM, where he found ideas around immutability, concurrency, and distributed systems that reshaped how he thought about software. He explains why Elixir was designed to be extensible, allowing communities to build things like Phoenix and Numerical Elixir without continually expanding the language itself.Robby and José also look at what AI is changing for maintainers. José explains how Tidewave gives coding agents access to context that frameworks already provide to developers, why documentation and good error messages have become unexpectedly valuable to agents, and how AI-generated contributions are changing open-source collaboration. They also explore whether coding agents might change the old calculus around dependencies, SDKs, and choosing a programming language.Topics and Timestamps[00:00:53] What Makes Software Maintainable: José compares well-maintained software to good conference Wi-Fi and explains why he values software whose boundaries and moving parts can fit in his head.[00:03:20] Maintaining a Programming Language: How responsibility changes when thousands of applications depend on the software you maintain, and why language designers have to minimize churn for users.[00:10:51] The Origins of Elixir: José tells the story of the Rails concurrency bug that changed his career and sent him exploring functional programming, Erlang, and distributed systems.[00:17:00] Removing the Root Cause: Why José distinguishes between layering a solution over a problem and designing software so the underlying problem disappears.[00:20:47] Designing Elixir to Be Extensible: Why José resisted making Elixir a web-specific language and instead focused on giving developers the building blocks to extend it into new domains.[00:25:42] Restraint and the Elixir Ecosystem: José explains why much of his work moved away from changing the language itself and toward helping communities build on top of it.[00:28:21] Giving Coding Agents Better Context with Tidewave: How Tidewave connects Rails and Phoenix applications with coding agents and exposes information developers already rely on.[00:33:20] Designing for Humans in the Age of AI: Why José believes programming languages should continue optimizing for humans, and how features that help humans often help agents too.[00:36:40] Documentation as Infrastructure: How Elixir's long-standing investment in first-class documentation, HexDocs, Markdown, and searchable package documentation has translated well to AI-assisted development.[00:41:40] AI and Open-Source Maintenance: José describes the increase in AI-assisted contributions and why AI-generated responses can disrupt the human relationship between contributors and maintainers.[00:46:22] Rethinking Dependencies and Language Choice: Robby and José consider whether AI makes it easier to bring ideas across ecosystems instead of rewriting applications or depending on large SDKs.Thanks to Our Sponsors!Your test coverage says 90%, but that might be misleading. Undercover CI looks at your Ruby pull requests and shows you which parts of your changes weren't tested- not just overall coverage, but what changed and what got missed, down to the method level. Visit undercover-ci.com and use code MAINTAINABLE for 15% off your first billing cycle. Free for public repos. Private repos with unlimited users also available.Mailtrap is a modern email delivery platform built for developers. Native SDKs, a secure Email API and SMTP, and a free tier with 4,000 emails a month. When you need help, you'll reach real people on 24/7 support, not an AI chatbot. Try Mailtrap for free!Resources and LinksElixirDashbitTidewavePhoenix FrameworkErlangHexHexDocsNumerical ElixirNxFoundation by Isaac AsimovDune by Frank Herbert Subscribe to Maintainable on:Apple PodcastsSpotifyOr search "Maintainable" wherever you stream your podcasts.Keep up to date with the Maintainable Podcast by joining the newsletter.
We are excited to have Anthropic share their latest AI x Finance work at AI Engineer New York, coming up in 2 weeks!In case you've been under a rock, here's a non-exhaustive list of what Anthropic has been shipping since closing the largest fundraise of all time in May at $47B ARR:* June: Launched Claude Tag and Sonnet 5 and Fable 5* July: Opus 5, /checkup. crossed $65B ARR* Last month: Fable/Mythos 5.1, and EFS (upcoming pod)* IPO target $2T, end 2026 ARR estimated $100B* Cowork/chat merged before did* Claude Mods* Dario endorses the same Pacing the Frontier message cosigned by all labs* Last week: Opus 5.5, Plugins portal, Cloud Sessions/Claude Projects* Today: Sonnet 5.5!Today's episode should catch you up, with Thariq Shihipar, the explainer-king of Anthropic, who we last caught up on Fable launch day with The Field Guide to Fable:The Future of Mutable SoftwarePay special attention to Claude Mods (especially the cheatsheet):In general this is also the inverse of the other viral tweet from Thariq:Cloud Brain, Local HandsAnd give a try to Claude Projects:The “hands” terminology is not just an analogy for the local/cloud paradigm that is being built up at frontier coding agent companies like Cognition, but is ALSO particularly relevant to the safety systems discussions that we'll be discussing with Anthropic in an upcoming episode as they prepare to pace to frontier with responsible AI deployment.For those who want Thariq's writing tips we teased at the start of the pod, watch the full video here:From the rapid rise of Claude Code to a future where agents can rewrite their own harnesses, collaborate across teams, and operate across cloud and local environments, the way we build software is changing extraordinarily fast. In this episode, Anthropic's Thariq Shihipar joins swyx and Vibhu to unpack how power users are actually working with Claude Code today, why prompting remains a high-skill discipline, and where Anthropic thinks the agent harness is headed next.We go deep on Claude Code's evolving interface: Ask User Question and elicitation, artifacts as persistent generative interfaces, Claude Tag for multiplayer agent workflows, Projects, model effort, implementation notes, and the new Claude Mods system for customizing the harness itself. Thariq explains why Claude.md may eventually disappear, why the smartest model could also become the cheapest model for many tasks, and why mutable software could become a new paradigm for how applications are built and customized.The conversation then turns to agent security and Anthropic's “Pacing the Frontier” argument. Thariq walks through recent incidents where agents discovered unexpected ways to communicate, exploit infrastructure, reverse-engineer benchmark scorers, and chain vulnerabilities together. We discuss sandboxing, prompt injection, autonomous agents, interpretability, constitutional classifiers, probes, fallbacks, Auto Mode, and why securing increasingly capable agents may become one of the defining engineering problems of the next few years.We discuss:* Why agentic coding went from controversial to the default in less than a year* Why prompting is still one of the highest-leverage skills for working with Claude Code* How expert users build a mental model of Claude and what it can reliably one-shot* Why discovering your “unknown unknowns” matters more as agents become more capable* Artifacts as persistent, generative interfaces between humans and agents* How Claude could split into a cloud-based “brain,” local or remote “hands,” and dynamic interfaces* Claude Tag, Projects, and multiplayer agents and how collaborative agent workflows could evolve* Why spending more time on the initial prompt can dramatically reduce wasted agent work* When to use low, medium, high, or max effort for different engineering tasks* Why frontier models may eventually outperform smaller models on both intelligence and token efficiency* Why implementation notes can expose decisions the model considered but chose not to make* Why Claude.md may eventually disappear — and why starting without one can sometimes be better* Claude Mods: customizing the execution loop, UI, subagents, routing, and behavior of Claude Code* Model routers, forked agents, and supervisor agents that automatically improve agent workflows* Why Claude Mods may be an early preview of “mutable software”* The bitter lesson of harness engineering and why agent architectures go out of date so quickly* How Claude Tag is becoming an organizational harness for multiplayer work* Why giving agents access to company data creates an enormous new security surface* The Exploit-Bench incident where agents discovered ways to communicate and collaborate* Why agents hacked Hugging Face for scorer code rather than benchmark answers* How agents chained sandbox and infrastructure vulnerabilities in unexpected ways* Why increasingly capable agents make traditional security assumptions harder to maintain* The argument behind Anthropic's “Pacing the Frontier” proposal* Why software engineers are increasingly doing two jobs: engineering and keeping up with AI* Constitutional classifiers, probes, and fallbacks and what interpretability looks like in production* How Auto Mode checks whether an agent's actions actually match the user's permissions* Why Thariq can see serious AI risks while still having a relatively low p(doom)Thariq Shihipar* X: https://x.com/trq212* LinkedIn: https://www.linkedin.com/in/thariqshihiparTimestamps00:00:00 Introduction00:04:12 Ask User Question and the Future of Agent Interfaces00:08:29 Artifacts, Projects, and Multiplayer Agents00:15:37 Prompting as the Core Claude Code Skill00:21:52 Context, Effort, and Smarter Model Usage00:28:10 Is Claude.md Going Away?00:32:49 Claude Mods: Customizing the Claude Code Harness00:36:35 Model Routing and the Rise of Mutable Software00:44:40 The Bitter Lesson of Harness Engineering00:50:49 Claude Tag as an Organizational Harness00:55:59 Pacing the Frontier and Autonomous Agent Security00:58:22 Agents Hack Hugging Face for the Scorer01:05:34 What Happens When Agents Need More Compute?01:10:32 AI Coding Is Changing Faster Than Engineers Can Keep Up01:17:17 Probes, Fallbacks, Interpretability, and Auto Mode01:28:32 AI Risk, p(doom), and Closing ThoughtsTranscriptIntroduction: Life at Anthropic and the Pace of ChangeSwyx [00:00:00]: We're here in the studio with our friend Thariq from Anthropic, and I guess generally the Claude Code, I-- there's, there's so much, merging of boundaries and you've been so on top of everything since you joined Anthropic. You have been early to Claude Code itself, but then also, and you've told that story in other podcasts, and you've also been talking about seeing like an agent. Most recently you did the top AIE World Tour talk, Field Guide to Fable, which obviously you guys launched Fable, so that was-- that's cheating. And mostly you most recently also launching Claude Tag, and we're also gonna be talking about Pacing the Frontier. There's a lot going on in Anthropic. I guess top of the question is, what's it like being at Anthropic when there's so much going on?Thariq Shihipar [00:00:48]: I think that It is, like. I think you can get whiplash sometimes. I think, like, going. When I joined Anthropic, I joined because of Claude Code. Like Claude Code had just come out and I was like, “This is so good.” And Opus 4 to me was like just, I could not imagine, like, how good it was? And that was, like, a real moment for me. But I was, like, trying to convince, like, my startup friends to use agentic coding, and they're like, “Oh, no, like, our engineers don't think it's good enough,” or something. And I was like, “That's insane.” and now you, like, fast-forward, 12 months, less, and, like, it's just like, yeah, the default way that everyone codes, right? And I think that, like, just having to go from, like, selling it to, like, now, teaching people how to be. make the most use of it and be more efficient and things like that is just like a big, like big change. And, yeah, I think, like, it's just hard to stay on top of everything as a human? Like, I think things happen so fast and likeSwyx [00:01:51]: You just throw more agents at it.Thariq Shihipar [00:01:52]: Yeah, like that's like the agentic stuff scales much better than the, like, human stuff where it's like, oh, like, there are three things happening right now and, like, they're all emergencies and, like, how do you, like, respond to it? Yeah.Teaching People to Use Claude CodeVibhu [00:02:05]: What do you split your time on? You do a lot of technical writing, engineering work.Thariq Shihipar [00:02:10]: Yeah, so I think that, like, when I joined the Claude Code team, I wanted to teach people how to use Claude Code and I think that, like, that has been something that, like, I thought, like, maybe I would spend a little bit of time on it or, like, I'd, like, do. I was spending some time on the agent SDK first, and I wasn't exactly sure, like, how the bitter lesson would go, when it comes to, like, harnesses, right? Like, I think sometimes we were like, “Oh, like, what's after Claude Code?”? And so initially I was like, I just wanna teach people how to use Claude Code and make it easier to use Claude Code. And I think that has just, like, as the harnesses have gotten better and better, that's like the dominant problem now is, like, how do you use the agents, right? Like, it's like such a high skill expression thing. So I do that and then I do engineering work. I give talks, but I think, like, when I'm doing engineering work, my goal is to take that feedback that we get from users and also, like, then be able to talk about, like, hey, how to use Claude Code to do engineering. So there's like a good loop there. Yeah.Swyx [00:03:07]: Yeah. I'll-- For listeners, we'll attach, the talk that you did with Sarah for the Dev Writers, meetupThariq Shihipar [00:03:13]: Oh, yeahSwyx [00:03:13]: Which we talked a little bit about, well, first you do the work and then you talk about the work.Thariq Shihipar [00:03:16]: Right.Swyx [00:03:16]: Something like that.Thariq Shihipar [00:03:17]: Yeah.Swyx [00:03:17]: It's sow and reap orThariq Shihipar [00:03:19]: Yeah, reap and. Sow and reap.Swyx [00:03:21]: Something like that. Something like that. Yeah, so, and then just to preview a little bit, we are gonna talk about the evolution of the harness. It has come a long way from just being a CLI. We're gonna talk about, Claude Mods, which is starting to leak today, because you couldn't keep it secret.Thariq Shihipar [00:03:36]: Yeah. yeah.Swyx [00:03:39]: Yeah, there's, there's a lot, there. I think you started off with, like, adding ask user question tool, which people love and hate.Thariq Shihipar [00:03:48]: Yeah.Swyx [00:03:48]: Like, I thought it was, like, very innovative, and then now I have, like, my own version. You have your Interview Me version.Thariq Shihipar [00:03:55]: Yeah.Swyx [00:03:56]: And, yeah, everyone just has, like, their own stuff. And, like, it no longer matters ‘cause now you're supposed to, write prompts that create other prompts and loops and all these things.Ask User Question and Human-Agent InteractionThariq Shihipar [00:04:05]: Sure, yeah.Swyx [00:04:06]: So what's the state of the art, today? Like, what are people. what are you, like, telling people to do today?Thariq Shihipar [00:04:12]: Yeah, ask user question was the first time that the model was good at elicitation. I think this was, like, an emergent behavior that I, like, wanted to see if the models could do. I have, like a human-computer interaction background, so I, like, did that in undergrad and grad school. And so this was like. I think it's like human-agent interaction to me, like, trying to figure out, like, how can the agent communicate with you and extract, the requirements, right? I think that, like, one of the things about, like, that's difficult as Claude Code has gone broader and broader is that everyone has, like, their own way of using it, and it's very hard to, like, change the default behavior. So for example, like, if someone asks Claude Code to do something,Thariq Shihipar [00:04:59]: Sometimes they just want them to do the work, ‘cause they're, like, maybe a very good prompter, and sometimes they want. like, are not good at prompting? And you need. like, the agent needs to, like, clarify? And so that's, like, a good split. Like, and the ask you the question tool like, splits along that side where, like, are-- do you feel like you're good enough to instruct the agent as it is, or is the agent able to, like. does the agent need to, like, pull out more requirements and, like, collaborate with you more and really understand your preferences?Thariq Shihipar [00:05:27]: I, on the whole, believe that pretty much everyone is more on the latter than the former, that they, like, have more ambiguity and they know less than they want, than they, like, think they know about the problem. but, like, it's like a interface design problem to make that easy? And so, like, if you're designing a problem, like, or if you're going through a problem, like, things like what's the schema or, like, what's the call stack and things like that are really important. like, the details in the design are important. Ideally, you want to figure out some of these, like, hard problems ahead of time before starting implementation. And yeah, that's why they call, like, unknowns, right? And so I think that this will forever be, like, a skill in agentic coding is, like, figuring out your unknowns. So, like, because even if the model is, like, super intelligent- It, like, needs to know what you want? And, like, you have preferences. like, you need to like, pull the, pull that out. and so that's, like, I think how I'm, what I'm pushing. the question then is, like, how does the agent interact with you? And I think that has been HTML, has been, like, the big way of doing that. And we've recently added artifacts, right? And artifacts, I think we've done a bad job of, like, or, like, I've done a bad job of, like, explaining how to use them fully. We have a lot of property capabilities. They have a database associated with them? And so every artifact can store and write persistent data. They can, like, feed back into Claude? And so, like, one thing that, like, people are not doing yet that I'm trying to, like, encourage is, like, this idea of a dashboard artifact. So you have, like, Claude working on a project long-term. Maybe it's like a kanban or something. it can store that kanban data in its database. Multiple Claudes can access that data via, like, the artifact MCP, and, like, that artifact can, like, talk to those Claudes as well. And so, like, the. We're building the primitives for you to be able to have this, like, generative interface via artifacts that will, like, let you surface more of that rich detail from the agents. And I think that, like, almost everything with agents right now is, like, this problem of, like, you think what you want, but you don't really know what you want, and, like, the agents need a lot of detail, and collaborating with them in the loop is really important. And so artifacts are, like, the, like, way that we're trying to evolve there. But there's a lot of work to do because it's so much more complicated than, like, a multiple-choice question? there's a lot more, like, detail in terms of, like, diagrams and code snippets and schemas or, like, whatever it is for that problem. But, like, artifacts is, like, the mo-more AGI-pilled way of, like, doing ask user question. So yeah.Artifacts as the Interface to the HarnessSwyx [00:08:15]: I think one thing that's unclear to me about these, the artifact stuff is, like, what feedback should go in through the artifact and what feedback should go through a Claude, a chat? Because the more AGI-pilled one is to just feed everything to the Claude.Thariq Shihipar [00:08:29]: I think the more AGI-pilled one is to go through the artifact. Like, and I think that, like, we imagine in the limit, I think that artifacts will be your interface into the harness? You can, like, comment on this, like, live, like, document of your plan, of the work. you can see maybe, like, multiple agents and different agents are doing this, and that artifact is built for the current work that you're doing, right? And so, like, each one has, like, slightly different. I think we're still, like, getting there from, like, an infrastructure perspective. But yeah, I think, like, on-the-fly interface for your harness is probably where things are headed.Vibhu [00:09:03]: Is there a version of it that's an abstraction from CLI or chat and you. Because right now, a lot of it is, okay, you're interfacing with Claude Code, you're having HTML given back for a mockup. It's pretty rich. There's diagrams. Artifacts are ways to connect these together. Why not just do everything that way?Separating Brain, Hands, and Surface UIThariq Shihipar [00:09:22]: Then it becomes, like, separating out, like, where is the inference happening? Where is the intelligence happening? Where is the work happening? like, I think this is like, difference between, like, or, like, some of the distinction between local and cloud, right? And so, I think right now, if you use Claude Code, it's, like, local and, like, you can spin off remote control, for example, to get some cloud behavior, or you can spin off Claude Code in the cloud, right? We're moving towards a place where instead of Claudes, like, you message a local Claude, it starts a session locally and it executes, to more like you have a Claude that you message that's in the cloud that's running. it can run, like, local, or, like, cloud sessions. This is how Claude Tag works. But, like, over time, we'll add, like, local hands as well. And so, like, local hands will be the ability for that agent to access your computer if it's online, and be able to, like, work there. And so it can spin off many different subagents. It can, like, commu- those subagents can communicate with each other, and that's where the artifact comes in to display all of that work. So you can imagine, like, the. You're separating out these things. So there's, like, the surface UI display that's an artifact and hosted somewhere and has a database and everything. There is the inference intelligence, right, that's happening on the cloud, and you don't have to worry about shutting off your computer or whatever, right? and then there's the, like, hands. Like, and it can be local, it can be in, like, a remote sandbox or wherever you need your work to be done. That's like unpackaging, like, the Claude Code experience right now where, like, right now it all happens in one place, right? So.Multiplayer Agents, Claude Tag, and ProjectsVibhu [00:11:00]: How do you see, like, the multiplayer side of that? So say teams want to work in this way. Right now it's very individual, but how do you see the future of multiplayer? Like, right now, I guess there's Claude Tag, which is a version, but.Thariq Shihipar [00:11:12]: We're launching projects. And so projects is the, like, this abstraction that's like Claude Tag, but on our Claude products, right? So you can message it and, like, it will do the Claude Tag-like stuff, like spinning off subagents. So We think with multiplayer. Like, Claude Tag is, like, a little bit more native multiplayer because it's just, like, in your Slack and the permissions are all figured out and stuff like that. But I do think multiplayer is, like, an important part of the story and, like, that will need to get tied together more. Like, you can imagine how complicated it gets when you're like, oh, you have hands, but now you have other hands in other people's computers too, and, like, you need to, like, permission them or, like, you have, like, your MCP and someone else's MCP, and how do you figure out how to use them, right? It gets, like, quite complicated. And Claude Tag does a good job of, like, sanding down all of these issues, right? So that, like, when you have, yeah, Google Docs, how does it access Google Docs, right? Like, it accesses through the shared Claude MCP, or it can access through your local credentials as well if it doesn't have access. But yeah, I think Claude Tag is our multiplayer, product, and it's really useful for these, like, things that are inherently multiplayer. Like, okay, like on-call, for example, incidents are inherently multiplayer. You want to tag Claude, you want multiple people to log in, you want it to be able to find context. I think whenever I'm, like, working on something and I want, like, privacy or security or, like, I want other people to review it's really nice to, like. I'll have a channel per project and I'll, like, at legal, for example, be like, “Hey, like, I want to ship this. Can you, like.” Like, here's. Like Claude knows everything, just chat with it. And that way legal gets precise answers, on like what exactly is shipping into the code, and I don't need to be in the loop, right? So I think like multiplayer is getting like more and more like, yeah, everyone can participate with Claude. I think Claude Tag is like that product and like projects will start off single player and will like, expand.Swyx [00:13:14]: I think there's a question about like maybe dual questions about identity and the unit of isolation.Identity, Permissions, and IsolationThariq Shihipar [00:13:20]: Yeah.Swyx [00:13:20]: Claude Tag, you specifically chose to make it its own identityThariq Shihipar [00:13:26]: Yes.Swyx [00:13:26]: Which is like, a controversial choice. There's, there's other ways to do it.Thariq Shihipar [00:13:30]: Yeah.Swyx [00:13:30]: Claude Projects probably it sounds like, if it's anything like ChatGPT Projects, it is, the isolation is that artifacts, that cloud instance, everyone's collaborating on this. It'll. It sounds like, it should be like if you're, if you're collaborating with legal on a thing, like that channel should be a project, right? Like it's not yetThariq Shihipar [00:13:50]: Yes.Swyx [00:13:50]: But it. that's the natural next step.Thariq Shihipar [00:13:53]: Yeah, like I think in Claude Tag, it's effectively. Like Claude Tag, you have to do your own arrangement. And so Claude Tag, yeah, each channel is like you can name it as you want, and I nameSwyx [00:14:04]: Yeah.Thariq Shihipar [00:14:04]: Like each featureSwyx [00:14:06]: Yeah.Thariq Shihipar [00:14:07]: As a channel.Swyx [00:14:07]: And, but I think like there is some trans- like it's unclear when there is transference, because let's say it is. if you have a coworkerThariq Shihipar [00:14:14]: Yeah.Swyx [00:14:14]: Who is tagging on all these things, yes, there is transferThariq Shihipar [00:14:16]: Yeah.Swyx [00:14:16]: Because it's the same person. but with Claude, it's unclear if it's like necessarily like, well, no, you don't know any of. you don't know about the other stuff. You should only use this stuff.Thariq Shihipar [00:14:25]: It's like the tip of the iceberg meme, right, where you can like. This is what we spend so much time onSwyx [00:14:31]: Yeah.Thariq Shihipar [00:14:31]: Is like there is like infinite surface area of like, okay, you want Claudes to. Not infinite, but like there's like surface area, a lot of like, surface area to figure out of like permissions and visibility and like how can you let Claude operate as well as you can, as safely as you can? And obviously, this is very important to us because like security for our code base is very important. And so we've put a lot of time into this. Yeah, there's so many like edge cases you can figure out where it's like, oh, like, yeah, this Claude in this channel has different permissions, but it can message another channel, and can't it exfiltrate data that way? Or like can you like. What if it uses your MCP and then messages someone else? Like there's like so much, and we've like really put a lot of work into sanding it down.Swyx [00:15:14]: Yeah. Lots of work. okay. Fable?Fable and the Meta-Skill of PromptingVibhu [00:15:18]: Fable, you wrote two good articles. you've written many good articlesThariq Shihipar [00:15:22]: Yeah.Vibhu [00:15:22]: But on, Field Guide to Fable, Building Claude Code. I'm curious from what you've seen, is there any common patterns that you see in like top users at Anthropic externally? Like what are best practices for getting the most out of Claude Code?Thariq Shihipar [00:15:37]: The like meta skill I say is like prompting is like very important? And like that. Like I think this is like not trivial to say because I think a lot of people are like, “Oh, prompting doesn't matter. It's just like I can just say a sentence and Claude will do it.” And I think prompting is really this like, this. It's like public speaking, like, or writing or something, and for a specific audience, and that audience is Claude. And you need to like build a mental model of Claude and how it thinks and how it works, right? And so that's like the most important skill in working with Claude Code is like having this mental model, right, of Claude and like what it can do well, what it can one-shot, what it can't. And so many people when you see prompting, they're just like, they're short prompts, but they have such a good mental model of Claude and of like the code base and things like that like it's effortless? But it's like high skill ceiling. So like that work of like, spending a lot of time prompting and building mental models of how, and intuition for how the agents work is really important. And then I think like the next thing is like the unknown stuff we talked about earlier, where it's like being able to find out like your, what you don't know or what you haven't written down, learning about like different things. I think as Claude can do more and more things, the likelihood of you doing something out of distribution for you and like you have low domain knowledge on is very high? And the more you can like learn the vocabulary to be able to prompt Claude, it becomes really important. And so like I think the most important unknowns are the unknown unknowns, where you're like, I just like don't even know that this exists, right? Yeah, exactly. I think that's like a illustration of like the map and the territory, right, where you're like, “Okay, this is my prompt,” and the territory is like the actual like work that the agent needs to do, right? And if you are like very precise, you can give more precise things, right? So like for example, in design, I'm not very precise. I'm not a designer, so I say like, “Give me like eight different mock-ups.” But if I was a designer, maybe I'd be like, “Oh, hey, here are some reference sites.” Like, “I want this type of font and this type of like look to it, and here's like a few different components to like visualize. Here's a Figma MC board to bring in,” like. And so you can just be so much more precise with that language. And if you're not a designer, you just need to like try and learn the language or learn the unknown unknowns. And this is true of like everything, I think. Like the more, like you can work with Claude to learn like how things work, the better your prompting will be. I think another good example of this is like game design, like where a lot of people are like, “Oh, like I can vibe code a game now.” And they're like, “It's not fun.” And like it's just like the thing about game design is like every one of these choices has like a lot ofTaste, Domain Knowledge, and Learning the VocabularySwyx [00:18:25]: Variations.Thariq Shihipar [00:18:25]: A lot of like craft to them. So it's like, oh, okay, like when you're making a flying game, the feel of the plane and the like, way it responds to your controls has a lot of like. Like, a game designer would spend like days on that. Do? and likeSwyx [00:18:44]: To me, that's what taste is, right?Swyx [00:18:45]: Like it is like from the possible space of one thousand mathematically valid answersThariq Shihipar [00:18:49]: Yeah.Swyx [00:18:49]: Here's the one that is the humans will like.Thariq Shihipar [00:18:51]: Yes. Yeah.Thariq Shihipar [00:18:52]: I think with taste, I'm like torn on this word ‘cause I think you're right, but everyone has different definitions, and it sounds kind, sounds like low skill or like elitist almost, where you're like, oh, like there are certain people with taste?Swyx [00:19:06]: It's like taste is what I call taste.Thariq Shihipar [00:19:07]: Yeah, exactly.Swyx [00:19:08]: And it's like these guys don't have taste.Thariq Shihipar [00:19:09]: Yeah, exactly. Oh, like an engineer doesn't have taste. Like I, the like founder, have taste.Thariq Shihipar [00:19:14]: ? And I think that's not true. Like I think like the engineers have a lot of taste for these particular like problems? And I think everyone has taste for particular problems. I think like Jason Liu, like say like in order to, yeah, have taste, you have to eat?Thariq Shihipar [00:19:32]: And I really like that, where it's like, okay, you have to like do a lot of things. You have to like iterate and figure out what you want, what you like, and, like build that like domainSwyx [00:19:41]: YesThariq Shihipar [00:19:41]: Domain vocabulary. And then when you're prompting, you're like synthesizing all of that for a product.Swyx [00:19:46]: Isn't it annoying when someone else says it better than you?Swyx [00:19:48]: It's just like, f**k, I have to quote this guy forever.Vibhu [00:19:51]: Having to quote Jason Liu forever.Vibhu [00:19:53]: He's gonna love this.Thariq Shihipar [00:19:55]: So I get prompts, more than that.Vibhu [00:19:57]: And sometimes it's not even that. Sometimes it's just intuitive, right? Like you don't realize you even want something till a model puts it out, and you're like, “Oh, this just feels immediately better,” right?Voice Prompting and Information DensityThariq Shihipar [00:20:07]: Yeah, exactly.Swyx [00:20:09]: One thing I go back and forth on is I feel like the way I prompt half the time, let's say I use voice.Swyx [00:20:16]: Did I say voice? Other people have voice. that is the opposite. That is just like me rambling for like two minutes Pressing down the function key and then let go, and then like hopefully it figures it out. And oftentimes it does.Thariq Shihipar [00:20:26]: Yeah.Swyx [00:20:26]: But it's not as thoughtful as like a structured prompt with like Well-run communication as though it's a PRD or a memo. Is that in line with how people do this? There's like bimodal prompting where there's some prompts where you spend a lot of time upfront and other prompts you just dash it off?Thariq Shihipar [00:20:43]: I don't think the voice is necessarily low. Like I think it's like more like how much information is in the prompt. like the model can. Like you can and like add some sentencesSwyx [00:20:53]: RightThariq Shihipar [00:20:53]: And be like, “Oh, like I changed my mind,” like in the middle of the prompt, and it will be able to follow that perfectly? So I think the like actual format of the text is less important, but then like the ability to. Like how much information is in it, right? And I think for voice, a lot of times, going back to like human-agent interaction and like for a lot of people, it's just way easier to talk than to like type? and I. If that gets more information out of you, like that's better.Vibhu [00:21:21]: At some level, it feels like just giving the model as much contextThariq Shihipar [00:21:24]: YesVibhu [00:21:24]: Over prompting before you kick off is a best practice. I don't know. A lot of the times, like when I was first trying out Fable, I spend a solid 30 minutes like really crafting a long prompt. This, I think, is a response of models running for longer and longer, right? It's still a little difficult to nudge them as they're in like, in the loop, but I just like intuitively spend more time kicking off that first prompt and working with it a lot.Spend More Upfront, Iterate LessThariq Shihipar [00:21:52]: My personal opinion is that if I was a software engineer, if I was like, just running my own startup, for example, I think I would mostly fit, stick to a max 20x? like maybe verification and so code review are like separate things. But I think like what I see a lot of times is people hit rate limits when they're doing this like, oh, like it did a lot of work and you're like, “Oh, I don't like this.” Like, “Can you like undo this and redo it?” And then you're like iterating on this like thing that the model could have done if you had like spent more upfront time or given it better context? And instead it's like you're like, “Nope, don't like that design. Try this.” Or like, “You messed this up,” or something like that. And then that just eats up so much more of like, your usage. And so that's like, I think maybe like a key like tip both for like efficiency as well, right? And yeah, I think like context, and not just like context on like what the goal is good, right? Like are you building a prototype or is it like a production thing? Like where can you spend compute or when, where can you not spend compute? Like I think you have to give the model permission or like not permission to do things sometimes where, like it doesn't know intuitively how much you want to spend on this task, right? And you can use effort for this. So I did-- I'm working on a blog post about that where it's like, if you want. For like we see that effort scales with the complexity of the task. So for security, effort gets like way more results. Like high effort versus like low effort gets, like changes the evals a lot. But for software engineering, it doesn't change it a huge amount because effort is mostly spent on the verification and the like edge case testing and things like that. And so like being able to like give the model that guidance of like, “Hey, this problem is something that I think I want you to spend a lot of time verifying and edge case testing,”?Effort, Model Choice, and VerificationVibhu [00:23:43]: How about model in the mix? So, there's Opus and Fable with effort.Thariq Shihipar [00:23:47]: Yeah.Vibhu [00:23:48]: There's also Haiku in there.Thariq Shihipar [00:23:49]: Yeah. It's not quite true yet, but it's very close where I think the frontier models will be Pareto dominant over like almost everything. like maybe. And sometimes I think Opus might be Pareto dominant. Do? Like I think depending on like how things, like shake out if it's like a newer version of Opus. But I think that like increasingly it's just going to be like the smart model is going to be able to like do the simple task for less tokens than the like the other models because of verification. With verification, in the limit, your model doesn't need to verify, right? If it's a perfect model, it just does the work once and it's like, okay, like you, I did it? And increasingly with Fable, I'm like, I'm like, “Dude, you don't need to spin up Chromium and screenshot all of these things.” Like I see it. Like you did it, right? And so a lot of the. At higher effort, you spend more of those tokens verifying. But if you're working on simpler problems, and a lot of software engineering is like well, like in Fable, like low and medium stability, it can spend less tokens verifying. And as the models get smarter and smarter, they will just be able to like, “All right, done.”? Like, I can run the lint for sanity's sake, but, like, I, like, know it lints? Like, you don't even need to do that. And that will be so much more token efficient than, like, the smaller models. Yeah.Swyx [00:25:15]: Is there a good, practice on our side that we can use to see if we're using too much effort? Like, I freakingThariq Shihipar [00:25:23]: YeahSwyx [00:25:23]: Hate wasting time on that stuff.Thariq Shihipar [00:25:24]: Yeah. I know what you mean. I think, like, so in this blog post, my rough distribution is, like, code review and security should be, like, high or max and, like, software engineeringSwyx [00:25:37]: You said recommend mix settings per domain.Thariq Shihipar [00:25:37]: Yeah. I think, like, if you're doing, like, UI or something like that, like low and medium, I think is you're building, like, an API and you want to make sure, like, you cover enough edge cases? And so I think building, like I said, that mental model of, like, how things work across these distributions is, like, yeah, part of the job.Implementation Notes and Decision LogsVibhu [00:25:56]: This is more intuition-driven or eval? Because I'm guessing this would change as you go.Swyx [00:26:00]: He has evals.Thariq Shihipar [00:26:01]: Yeah. So what I did in the blog post is I go over all of the terminal bench evals. So there are, like, 70 problems and I'm show that, like, okay, like, in the security problems it does more. and then I also, like, look at some of the transcripts just in terms of, like, how-- what does it answer, what does it forget or something. And a lot of times, this is another prompting tip I have, is, like, asking it to make decision notes or implementation notes because, in every eval problem that it faces, it thinks about the correct solution, and decides not to do it. it's like, oh, like, here is the answer. What if I did this? And then it's like, oh, probably not? and then keeps going. And this is, like, the majority of the failures, at, like, a higher max level. It's very rare that the model just doesn't know how to do something. If you just have these implementation notes, then you can review and you can be like, “Oh, I want you to do this thing that you didn't do.” The models are getting better at surfacing that overall. Like, I see in the transcripts of Fable 5.1, like, when it does this output, it will call out its decision-making as well. but making this more explicit in the harness is better. And now we're, allowing ways of you modifying the harness so you can, like, add someVibhu [00:27:23]: Ooh.Thariq Shihipar [00:27:24]: Calculate with there. Yeah.Swyx [00:27:25]: Yeah. So I do wanna call out two things that you mentioned that I think exist outside of prompting. One is like, let's, let's call it the prompt that is so important that it shouldn't be in a prompt. It is in Claude.md or Agents.mdThariq Shihipar [00:27:38]: YeahSwyx [00:27:38]: Which is like goals, right? Like your situation, your goals, the things that you want, the thing. and then second of all is the decision log or the experiment log or whatever log of traces that you might want to survive the current session to do those things. Those are, like, externalities that there's no standard. There's no-- It's not like skills. It's not like MCP. There's no standard. It's, it's just like it's a markdown file. first of all, is that right? Is Claude.md going away? You have a documented dislike of, Agents.md, but you're gonna do it?Claude.md, Agents.md, and Model-Specific InstructionsThariq Shihipar [00:28:10]: Yeah. Okay. So Agents.md, yeah, like, we're, we're gonna do it. I think it's just, like, different models are very different from each other? But I realize that it's, like, such a pain to, like, maintain different ones? And yeah, like, as the models get better and better, the floor of how they accomplish the simpler task is better. And so I do think in the limit, Claude.md goes away, and maybe not even, like, that far. Like, I think, like, I think that right now it might be better to start a new project without a Claude.md.Swyx [00:28:44]: Yes.Thariq Shihipar [00:28:44]: I think that, like, maybe if you see very repeated failure modes, you add them to your Claude.md. The really tough thing is that this changes per model. And so, like, if you've added a bunch of failure modes or, like evenSwyx [00:28:57]: So you need Fable MD, you need Opus MD.Thariq Shihipar [00:28:59]: Or well, even Fable 5.1 versus Fable 5.Swyx [00:29:03]: Yeah.Thariq Shihipar [00:29:03]: Like, it is annoying. Like, I'm not like,Swyx [00:29:05]: YeahThariq Shihipar [00:29:05]: Like, we don't, like, do this on purpose? It's just, like, how the models work, right? And so, like, maybe, like, Fable 5 had this, like, failure mode that Fable 5.1 doesn't. And if you keep this context, this running log of a bunch of different failure modes, they will probably over constrain Claude? And so this is like. we just added evals plugins for skills.Swyx [00:29:28]: Yeah.Thariq Shihipar [00:29:29]: And so now you can eval if a skill is better. I think Daisy on our team did this. And so, yeah, this is like we're trying to work on this. We know it's, like, you still have to spend tokens on it and, like, it's not, it's not perfect, but it's, like, we're trying to help out with this problem.Swyx [00:29:44]: And so, and as far as prompting goes, the one tip I wanna offer is, something I have told people a lot is sufficiently advanced prompting is indistinguishable from sufficiently advanced executive communication. So I've referred to-- This is an executive comms workshop from Heavybit that is the best I've ever seen in my career. And they teach this thing called the SCQA model. Just Google it. It's a, it's a thing. Like, people have done prompting for decades. It's just called executive communication. It's like when one person has to communicate to thousands of people down the org chart, this is what you do. so situation, complication, question and answer, is how you write the memo. but obviously sometimes you don't have the answer, but you can at least list out the SC and Q, and then they have some examples in there. So just leaving breadcrumbs for people if they want to explore.Underrated Prompting Patterns and ELI5Vibhu [00:30:31]: Before we move on, I wanna ask you, any other underrated tips, ways people could get a lot of value from Claude Code that they're not using?Thariq Shihipar [00:30:41]: Yeah, I think a lot of them are in the, this unknowns, like, doc. Like, I give a bunch of example prompts, like, using it for brainstorming, using it to quiz you after. we added this, like, explain it like I'm five skill which is a very short prompt. And it doesn't even say explain it like I'm five. It's like the key word of this prompt is big pictures, few words. like, that's like the main thing. And it is shockingly good? Like, you, like, I think I tweeted about this and it's like /eli5, and, like, you can install it as a plug-in. But yeah, it's, like, way better at just cutting through the BS and being like, yeah, exactly right here. So the diagrams are, like, quite clear. I think one of the things that is true with artifacts is, like, they put too much text in and people are not reading the artifacts? And so, like, this simplifies it a lot more. And, yeah, this came out of, like, just people at Anthropic, like, going through very complicated incidents and being like, “What is happening?”? So, this one I think is great, yeah.Swyx [00:31:47]: My version of this is the, it's like test your understanding. Give you a few choices and then, like, if you get it wrong, you have a mismatch between what you think is happening versus what's happening.Thariq Shihipar [00:31:58]: Yeah. I think this is one of those things that everyone loves talking about, and then very few people really do. Like, I thinkSwyx [00:32:05]: Really helpful.Thariq Shihipar [00:32:07]: Yeah. But most people just don't want to get quizzed about something? Unfortunately, I think this is one of the, like, things that we need to, like.Swyx [00:32:16]: What's the opposite of ask you the question or ask you the question before the thing?Thariq Shihipar [00:32:19]: Yeah.Swyx [00:32:19]: This is after the thing.Thariq Shihipar [00:32:20]: Exactly. Yeah.Vibhu [00:32:21]: It's a good way to stay grounded of, like, do you even know what you're doing, right? The worst case is when people send you slop and they haven't understood what they're asking for or what the output is, and it's like, “Dude, I don't wanna read this. Do you even know what it is?” So, you make it a rule for yourself that before you send stuff, you should at least know what's implemented.Claude Mods: Customizing the HarnessThariq Shihipar [00:32:41]: Yes, but so you could make this a mod and you could build your own mod to, like, make sure you test it. So yeah, you can do that.Swyx [00:32:49]: All right. Let's get right into it. What is Claude Mod, and what is this diagram showing?Thariq Shihipar [00:32:54]: Yeah. Okay, so Claude Mods is you can customize the entire Claude Code harness, and we're going to. If you have requests, we will, like, let you, like, please let us know. We'll add more and more. This works for CLI, it works for desktop. maybe it will work for Claude Tag in the future. I don't know. Like, we're trying to make this very extensible. You can see this reference sheet. I don't want people to get overwhelmed by it? At a high level, you can customize both the execution of the harness, and the UI of the harness. And so, like, you say on that Tetris example from Boris, that's like customizing the UI, right? Like showing, like, Tetris in the game.Thariq Shihipar [00:33:35]: But, like, let's say that you wanted to do this thing where you had. you tested your assumptions or, like, tested your understanding after every project, right? What you would do is you would ask Claude to make this plug-in. It would spin a classifier after every prompt. And so, like, at the end of each turn, you would spin off a sub-agent or, like, a forked agent. A forked agent is, like, maintains the prompt cache, right? So it's like a, like one of those unintuitive things where you can fork and do, like, a little request, and it'll be very cheap because the entire prompt cache is, like, done. And so you can be like, “Has this task been completed?” likeSwyx [00:34:18]: This is how you do BTW and all those.Thariq Shihipar [00:34:20]: Yeah. The underlying forked agent, yes. But so you can, in the f-fork sub-agent, you can say, like, “Has this task been completed? If so, return true.” And then in your hook, or in your, like, plug-in mod, or sorry, like, in the sub-agent probably, you would say, like, “If true, give me a quiz.” give me questions and answers, and then, like, in a JSON format, and then you'd parse it, and then you display above the prompt input, this list of questions, right? And so this is something that's, like, slightly token-intensive because, like, you have to do it after every end of the assistant turn. But it's, like, a lightweight classification, and then you can, like, get this quiz, and then you'll see, like, Claude will always do it for you. You don't need to remember to do it. There are lots of these, like, tips that we've talked about, right, where it's like, oh, implementation notes. You can also add a tool for implementation notes now. And so, like, this tool that I'm adding is, like, register, like, I think assumption is what I'm calling it, but, like, maybe I'll change it around. And this is a mod. And so, like, you give it a register assumption tool, and then it will keep a list. It'll. Every time it does it'll keep a, like, add to the list, and then at the end it will display those assumptions? Another mod I'm working on is a model router. And so, like, internal, like, Claude model routing, right? So it's. This is, I want to say the reason we don't do model routing by default is, like, it's a hard problem? And likeForked Agents, Assumption Tracking, and Model RoutingSwyx [00:35:51]: You will get it wrong.Thariq Shihipar [00:35:52]: Yeah, you, like, yeah, you will, like, accidentally use, like, Fable for a hard problem or Sonnet forSwyx [00:35:57]: Yeah, if you have auto approve, but you don't have auto mode.Thariq Shihipar [00:36:01]: Well, you will have auto. Like, you don't have, like, auto routing or something.Vibhu [00:36:04]: You don't have auto mode for model picker.Thariq Shihipar [00:36:06]: Yeah, exactly. SoVibhu [00:36:07]: I'm getting the rough question of, like, how much do you open this up and how much do people have to think about this? Like, when you talk about prompt caching and building a router, it seems like you could easily build a mod that routes per query, and I'm just killing my plan very fast, right? I guess my question is more so, like, what is, like, a product talk like this look like, right? Who is it for? Is it for power users? Is it everyone should be able to go throughSwyx [00:36:33]: Oh, definitely power users, right?Thariq Shihipar [00:36:35]: Yeah, I think it is power users, but, like, the nature of Claude Code is that so many people are power users? Because it's easy to share things, like you can. Like, one person can make a good model router thing that doesn't break prompt cache all the time, and then you can, like, compose them. Another cool thing about the plug-ins is that they can hook into and compose with each other. And so I have, like, a mod that will, like, create a mode selector at the top, and any plug-ins can register to be a mode. And so, like, the auto router can be a mode, right? Or, like, you can have a mode that's, like, artifact mode, where it's like it primarily talks to you in artifacts. like, you can toggle between plan mode? And so, like, you can create more and more of these modes. But the ability to create modes is in it itself a mod? And so there's a lot of richness here, but we do want to make it fairly easy. We want to be-- make it so that you can just, like, install someone else's. You can ta-- you can chat with Claude and, we'll, like, make sure that it understands the nuances of things like prompt caching and stuff, so it can, like, warn you. This is, like, not extremely complicated behavior for Claude, I think, but we should have just a good skill on how to make mods. and yeah, we'll see how we go. But I do think that this is, like, a preview of, like, mutable software, and, like, how, like, generative software, just like you can customize safely. If enabled, you could customize any piece of software. And I think that more and more apps ideally do something like this?Power Users, Modes, and Mutable SoftwareSwyx [00:38:13]: And by the way, you, we have, you have another cool tweet about how, there's the infinite money button, which is like make your SaaS, consumable by agents. I think mutable software is interesting and, other people have also tried to do it. I think the hurdle comes when you can do everything, then people, users get, tend to get confused. So usually the stuff that works is just like one opinionated flow. This is in the side of less opinionation. It's just like, well, more power to power users. And I think probably unlocked by AI, where, like, you can just prompt for whatever the thing is.Thariq Shihipar [00:38:47]: Yeah, or there can be a skill that gives the opinions?Mods vs. Hooks vs. ArtifactsSwyx [00:38:50]: Yeah.Thariq Shihipar [00:38:50]: And then, yeah.Swyx [00:38:51]: So knowing a little bit about, like, TypeScript and build systems and all these things, the closest-- I'm very curious that the team who worked on this, if, I don't know how close you were to them, if they drew any inspiration from build systems like Babel, Webpack, all these, like, old school things. Because it sounds very similar, like the plug-in ecosystem of those things where they can compose with each other.Thariq Shihipar [00:39:11]: Yeah, I'm not deep in the technical details, but I do know it was a collaboration with someone on the Bun team and someone on the Claude Code team.Swyx [00:39:17]: Yeah, it's a build system mecca.Thariq Shihipar [00:39:19]: Yeah. Exactly. It's, it's very exciting. But yeah, like, agents can just do this very complicated like, extensibility into your software now. And so, yeah, like, another reason to, like. If you run a startup, like, you can just prompt Claude and be like, “Hey, like, could we make an extension system? Like, what would that look like?”?Swyx [00:39:37]: Yeah.Swyx [00:39:38]: And I just really wonder, like, you had hooks in the past and plug-ins, all these things. So what specifically will mods be able to do that those things could not do?Thariq Shihipar [00:39:47]: Internally, we were originally calling this function hooks. And so, like, that's, like, gives you a little bit of an idea where, like, hooks register a, like an event to happen and then, like, a script to call. And this inside of the, like, TypeScript runtime is running things. And so, like, you get some benefits of just, like, it has a bunch of things in the Scope with, like, for example, like how many turns is in this conversation, right? Like, how many tokens have been used? Like, et cetera. Like, what are the messages? Things like that. So it has a bunch of messages that can be used. And then it's just, like, a lot more hooks. So we have, like, or a lot of, lot more, like, things you can register on. And then you can do because of the. because it's all happening in process, you can, spawn sub-agents, with four contests and contexts and stuff. And, like, that will return. You can parse the results of those. You can use structured output to like, return them. and then you can modify the UI, which you can never do in hooks. So, yeah.Swyx [00:40:50]: Yeah. Yeah. So modify UI, this is why you showed the Tetris example. Does it also ex-extend to artifacts? I assume it does.Thariq Shihipar [00:40:57]: You-- Like, artifacts are like a different way of customizing it. like, you can definitely. One of the mods I'm working on is, like, this dashboard mod, which will, like, prompt Claude to maintain a dashboard, that's an artifact. But they're like, slightly orthogonal, or not orthogonal. They compose with each other in different ways. Like, mods are, like, a little bit more, like, in your Claude Code harness, changing the agent loop? And, like, the UI is, like, an added benefit. and then artifacts are just like you want to, see things at a high level, very inter- highly interactive. like, the affordances can be a lot bigger than, like a TUI or even in our desktop.Next Steps, Supervisors, and Persistent GuidanceVibhu [00:41:40]: I'm guessing you'll have a good blog post on the differences, because right now you can also, make a loop that outputs to an artifact that's an interactive dashboard, but you can also do it with a mod. There's just some thinking about making a hacking on a harness when we don't know much about the harness, right?Thariq Shihipar [00:42:00]: Well, something I'm excited about with mods is, like, there's so much things with Claude Code that you just have to remember? You're like, “Oh, like, let me do this, and then let me call the dashboard skill that does the loop,” and things like that. And, or like, “Let me test my assumptions afterwards.” And I think, like, if you do all of these things using these little classifiers and stuff, and you're like, “These are the things I care about. This is what I want to do,” you can, like. You don't have to remember as much. One more, like, mod I'm working on is a next steps mod thatSwyx [00:42:28]: I have-- I was gonna say, I have a next step skill. I always run next steps.Thariq Shihipar [00:42:32]: And does it have access to your skills? Like, this is one of those things where I'm like.Swyx [00:42:37]: I think so.Thariq Shihipar [00:42:38]: Okay. Yeah, probablyVibhu [00:42:39]: Do skills need specific access toThariq Shihipar [00:42:41]: Well, I think there'sSwyx [00:42:41]: Don't they always haveThariq Shihipar [00:42:42]: I think there's, like, specific prompting, I guess, to, like, know your skills. Like I think Claude forgets them sometimes throughout, like, the thing. But anyways, the idea of, like, yeah, next steps that also are like, “Oh, hey, this has happened. Use the explain skill to explain to you what happened because this seems, like, quite complex,”? Or, like, yeah, “Use your unknown skill. It looks like you are, like, asking the model to, like, iterate on these small changes. It seems like you could prompt better.” like, “What if you did this?” Right? So, I think, yeah, like spending more compute there. Yeah.Swyx [00:43:20]: And it should always come out as multiple choice. we have, I haveVibhu [00:43:23]: We have his skill.Swyx [00:43:24]: My next step skill is like this.Thariq Shihipar [00:43:26]: Okay, perfect. Yeah.Swyx [00:43:27]: You can steal it.Thariq Shihipar [00:43:28]: Yeah.Swyx [00:43:29]: Like, but like, for me, it's all-- I think models really always need to be reminded, what are you trying to do here?Thariq Shihipar [00:43:35]: Yeah.Swyx [00:43:35]: Look at the whole transcript and go like, oh, was this original goal? Did your solution solve it? Were you lazy? If you're lazy, maybe there's a reason. Maybe you needed approval from me. Maybe you needed, there's two things you wanna suggest. So it's, it's a little bit like the modification of the ask user question or interview me skill. so it's next steps.Thariq Shihipar [00:43:55]: Yeah, exactly. And again, the benefit of doing it with mods is you can do it as a fork sub-agent, and so it doesn't remain in the context afterwards. So you have this, like, idea of like, okay, the model is doing its execution and you have this almost like supervisor, like, that is like making sure that you can do like the next steps well. So yeah.Swyx [00:44:15]: Yes. I do have two panels and like I often try to have a supervisor thing, keep the high-level context and then the implementationThariq Shihipar [00:44:21]: YeahSwyx [00:44:22]: Detail in another agent.Vibhu [00:44:23]: I feel like a lot of this abstracts away as models change? The, like, half an hour ago you said bitter lesson of harness engineeringThe Bitter Lesson of Harness EngineeringThariq Shihipar [00:44:31]: YeahVibhu [00:44:31]: And we're on the other extreme right now, I feel.Swyx [00:44:33]: Well, so yeah, exactly. If everything's customizable, what is Claude Code, right?Thariq Shihipar [00:44:37]: Yeah.Swyx [00:44:37]: And which I talked to you about last night.Thariq Shihipar [00:44:40]: Yeah, I think that this is. I think the bitter lesson is unintuitive? In terms of like. Also, like we're misusing a little bit of the bitter lesson here where it's like, it's more about like scaling and compute and stuff. But like, I think there is something where it's just like. I think I use it as an approximation here to say that harnesses go out of date very quickly? And like how, but how they change is unintuitive? And so like the big obvious example is like from chat to like agents where you had to give them entirely new tools, right? But like, I think this new version of like, oh, it can modify its own harness, right? This is like, an own harness loop is like a way of using its capabilities, right? Or like it can build an artifact. And like, I think the way I think about it is like the models have more and more intelligence, and they're like so much more intelligent now than like the average software engineering task. Like, you look at the like terminal bench ones and they're like solve like the Jacobian conjecture. Not really, but like, it's like they're, they're quite complex. Like, I would not have been able to do this really as a software engineer.Swyx [00:45:42]: And you said TB4 or TB2?Thariq Shihipar [00:45:43]: TB3. TB3.Swyx [00:45:44]: TB3.Thariq Shihipar [00:45:44]: Yeah. They're quite complex, but the goal is still to deliver user value, right? And like you said, there's like this infinite space of things to do. And so the ways like you spend compute are to keep the user in the loop and make sure that like you're getting to the right decision in the end of the day and like the right output. And artifacts and mods are this way of like spending that intelligence. and I think that's like, yeah, the next step. And so, yeah, I think Claude Code is like, has the core things of agent loop which are, have gotten more complicated. It's like, it needs a sandbox to operate safely. It needs auto mode to like make sure like the permissionsVibhu [00:46:21]: Approvals.Thariq Shihipar [00:46:21]: Yeah, approvals. it needs computer use and MCPs and like all of these like ways of accessing your data, and it needs web search and web fetch. And like, so the-- as the models can do more and more, the core harness has to be like quite complex and very secure. But then like how you interact with it can change quite a lot.Vibhu [00:46:42]: What other harness engineering best practices have you, from the Claude Code team itself? I feel like, there was a phase of plan mode, which is not as used. We now have auto mode. at a point you cut the majority of the system prompt, you got rid of examples. What other best practices are there for harness engineering?Core Harness Primitives and Managed AgentsThariq Shihipar [00:47:02]: I think there is like a forking path where at some point, eventually, yes, the model will just be able to like vibe code the exact version of Claude Code, even describing all this complexity that I've talked about, right? Like auto mode and computer use and stuff. Eventually, the models will just be able to do that in one shot. But I think they can one shot simpler harnesses? And so like, I think some people. Sometimes you don't need this full, like if you don't need computer use or like all this like more complicated stuff. I think before we, you had to use things like the agent SDK, which was like Claude Code wrapped, in order to like. And I would, like suggest people do that because there was so much complexity into building a harness. And now as that's got more abstracted, we have like, Claude managed agents, which lets you have that complexity, but still like, right, like a very bare bones like harness that's scoped to your task. Yeah, I think there's like this barbell effect where like for like very complex, for like coding task and like these like complex things, you should use our harness. And then for like a lot of like simpler or like, more domain-specific things, you can build your own harness because Claude has gotten better at building harnesses, and we have these harness primitives like managed agents. So yeah.Swyx [00:48:18]: Yeah. Is there a general progression? Let's say chapter one was ultra code dynamic workflows, then chapter two was cloud mods. Where is this going?Swyx [00:48:29]: Where you're, you're, you can customize the thing on demand.Thariq Shihipar [00:48:36]: Yeah. I do think that like this evolution of projects and like artifacts and splitting out like brain and hands and, surfaces is like where things are going more. And like, I think it's like not all quite there. partially it's like a, it's just like more token expensive? And like, I think likeProjects, Local Hands, and Cloud-to-Local HandoffsSwyx [00:48:59]: Why would projects be more token expensive? I understand mods would be slightly more token expensive. No, not something I'm worried about.Thariq Shihipar [00:49:06]: Yeah.Swyx [00:49:06]: But whatThariq Shihipar [00:49:07]: You're asking Claude to do. It's like creating loops. Like you're asking Claude to do more work for you. And so like it's managing the sub-agents and reviewing it, versus where you would be doing that work normally. And so that's like gonna be a little bit more intensive, like. Outputting to an artifact is gonna be a little bit more token-intensive than, like, outputting normally. I don't think it's too much more, but like, it's like combining all of these together well, like I think we're, we're still working on like local hands and things like that, I think is like, yeah, where things are headed, yeah.Swyx [00:49:37]: Yeah. Claude and local is, handoff is very interesting. I was thinking about this as reverse cloud remote.Thariq Shihipar [00:49:44]: Yeah.Swyx [00:49:45]: Because it's like remote, it's you're handing off to cloud, but here the cloud is handing off to local, right?Thariq Shihipar [00:49:49]: Yeah, exactly. Yeah, remote control is also another way of doing it. And I do want to say this is like how I think about it and like what the things that I'm most excited about this, but like there are, just like lots of different ways to work with Claude. Like some people use remote control a lot, some people use Claude Code on the web a lot. Obviously, like at Anthropic, we use Claude Tag a lot, and like what's great about Claude Tag is we set up all this stuff for our own execution. And I do think if you're an enterprise, that's still the best way to go. but if you're like an individual, Projects is this way of like, getting some of that like niceness of Tag, which has like that like supervising agent and yeah, adding artifacts and stuff, but like without having that whole like admin setup. And so there will be many ways to use Claude, I think. I think it's probably not just one like single.Claude Tag as an Organizational HarnessSwyx [00:50:36]: You had the multiplayer thing here. Let's, let's just check in on Claude Tag. it's been about two-plus months. Lots of, public, adoption and trying it out.Thariq Shihipar [00:50:45]: Yeah.Swyx [00:50:45]: What's new? What's, what have you found since the launch?Thariq Shihipar [00:50:49]: Like, Claude Tag is how we useSwyx [00:50:51]: It's like 80% of your
From the earliest days of open-weight models to becoming the neutral routing layer for more than 10 million developers, OpenRouter is one of the clearest bets that the future of AI will be multi-model. In this episode, OpenRouter co-founder & CEO Alex Atallah, with AMP's Anjney Midha returning with swyx to unpack how OpenRouter emerged from the first wave of Llama, Alpaca, Mistral, and Midjourney, why model diversity mattered before it was consensus, and how a company dismissed as “just a wrapper” became critical infrastructure for the AI ecosystem.We go deep on the product and distribution lessons behind OpenRouter: why model labs can spend billions training a checkpoint and still struggle to get it into developers' hands, how Mistral helped prove the value of a competitive inference marketplace, why OpenRouter chose focus over expanding into fine-tuning, memory, and other adjacent products, and how its rankings became a real-time map of how AI usage was changing. Alex also explains OpenRouter's early experiments with model fusion, why they deleted the first version and brought it back years later, and how the platform grew to more than 10 trillion tokens per day.Finally, Anjney explains why Stripe and OpenRouter fit together, why token fraud may become one of the defining security problems of the AI economy, and why the next wave of fraud won't just come from humans but from autonomous agents attacking increasingly valuable token flows.We discuss:* Why OpenRouter bet early that no single AI model would win everything* Alpaca, Llama, and open models becoming impossible to ignore* Why Discord's early AI deployments exposed the limitations of closed models* Why model labs can spend billions on training and still fail at distribution* How OpenRouter became a neutral distribution layer for model developers* Why VCs dismissed OpenRouter as “just a marketplace” or “just a wrapper”* The Mistral price war and the first real proof of an inference marketplace* How Midjourney scaled through Discord and what it taught the AI ecosystem* Why crypto infrastructure became a dress rehearsal for generative AI* OpenRouter vs. LM Arena and why their missions are fundamentally different* Why focus became one of OpenRouter's biggest strategic advantages* Anthropic's early focus on AI pair programming and coding* The OpenRouter products that were prototyped but never launched* MOM, OpenRouter's early Mixture of Models experiment* Why model fusion failed in 2024 — and why it works much better now* How OpenRouter's leaderboard became a live map of the AI industry* OpenClaw, auto-routing, and agents reshaping AI usage* How OpenRouter reached 10+ trillion tokens per day* Why inference gateways are increasingly becoming targets for fraud* Why Stripe's fraud infrastructure is strategically important to OpenRouter* The coming rise of agentic fraud and attacks on the token economy* What changes and what stays the same as OpenRouter joins StripeAlex Atallah* LinkedIn: https://www.linkedin.com/in/alexatallah/* X: https://x.com/alexatallah* Website: https://alexatallah.comAnjney Midha* LinkedIn: https://www.linkedin.com/in/anjney/* X: https://x.com/AnjneyMidha* AMP: https://www.amppublic.com/Timestamps00:00:00 Introduction00:02:12 Alpaca, Llama, and the Multi-Model Bet00:06:04 Discord, Open Models, and OpenRouter's Origins00:14:28 Why “One Model Wins” Was the Wrong Bet00:17:27 Why Model Labs Struggle With Distribution00:23:04 “Just a Wrapper”: Why VCs Misunderstood OpenRouter00:27:58 Bootstrapping OpenRouter Through Community00:36:16 Crypto, Midjourney, and the Early Generative AI Ecosystem00:43:38 Mistral and the Birth of the Inference Marketplace00:47:10 OpenRouter vs. LM Arena00:52:08 Focus, Anthropic, and Roads Not Taken00:59:34 Mixture of Models and Model Fusion01:02:44 Sonnet, OpenClaw, and OpenRouter's Explosive Growth01:09:03 Why Stripe Acquired OpenRouter01:12:45 Fraud and the Emerging Token Economy01:17:47 The Coming Wave of Agentic Fraud01:19:07 What's Next for OpenRouter at StripeTranscriptIntroduction: OpenRouter, Marketplaces, and Pub-Sub as a Product PrincipleSwyx [00:00:00]: Okay, we are here in Anja's house, which is where all big startups in San Francisco start.Anjney Midha [00:00:08]: Howdy.Swyx [00:00:08]: And, congrats on Cursor, Mistral. I don'- God knows what else. You got so much stuff going on.Anjney Midha [00:00:17]: There's, there's a lot going on. Well, OpenRouter is probably the - has been the most, I would say, like, one I'm excited about recently.Swyx [00:00:24]: Yeah. And we have Alex, first time on the pod, but,Anjney Midha [00:00:27]: Thanks for having me.Swyx [00:00:27]: You've been in the IE a few times. I appreciate every time you've shown up, for the community. Congrats. I just, like, what a journey. When I was looking back at your past posts, one of the earliest principles that I saw you write as a product person is sub as a product principle. And I wanted - you to maybe explain how you think about what should exist in the world.Anjney Midha [00:00:49]: Yeah. The sub piece, which was early 2023, I didn't think about it until we talked like 10 minutes ago, is about how there is like a way of thinking about products as an intersection between subscribing to data and publishing data. And marketplaces are an easy example of this. You have suppliers that are publishing some product to a SKU. And the SKU is like a sub topic that a consumer is subscribing to and just going to, like, consume whenever they want. And humans consume in a very, like, discreet, ad hoc way. It's not very scalable. all their attention is on the topic when they're buying the thing, and their attention is nowhere else when that happens. agents and consumers of inference don't act like that. They're consuming continuously, and they're changing the SKUs that they consume from all the time. So OpenRouter is like a blend between a normal API experience and a marketplace where we create model slug. We have the auto router. We have all kinds of, like, product SKUs that you can subscribe to. And then you can, like, continuously add, like, derive value and make decisions based on those consumers.Alpaca, Llama, and the Multi-Model BetSwyx [00:02:11]: Yeah. This is something that was more consensus now, but not consensus when you guys started, which was that there is such a demand for swapping models and changing things out and, that people would not use the native SDKs. I guess, for each of you, what was your realization moment that this would be it? I, - You've, you've given a talk at EIE about Alpaca as,Anjney Midha [00:02:33]: Yeah.Swyx [00:02:33]: One of your inspiring moments.Anjney Midha [00:02:35]: Alpaca, I can, like, rehash the Alpaca moment for a sec. Like, the very beginning, at the end of 2022, OpenAI was the only game in town. There was, like, OpenAI, Cohere,Swyx [00:02:47]: Yes.Anjney Midha [00:02:48]: And then a smattering of, like, early attempts at open weight models.Swyx [00:02:54]: Yeah.Anjney Midha [00:02:54]: When Llama came out in January of 2023, it was like, “Wow, really exciting. This is really big.” It outperforms 3 on, one or two benchmarks. but you can't chat with it. It wasn't like - It wasn't an engaging model, but it seemed like someone just needed to fix a couple things and do some RLHF on it to get it all the way there. And Alpaca was the first model that I saw that did that. It only took $600 to do. A team at Stanford generated a bunch of synthetic data, tuned Llama, and made Alpaca, billion parameter model. Or was - Maybe it was thirteen billion parameters. And it was so good. Like, I was just, like, on an airplane using it. I, - in many cases, I, like, you could not discern a ChatGPT versus an Alpaca result. And I figured if it was this easy to make a model, one, we have a whole new way of monetizing data for the first time. you can just, like, take really valuable data and turn it into a service in $600. and that cost will probably go down over time.Swyx [00:04:03]: When you - So sorry. when you say monetizing your data as, what eventually will become an MCP endpoint or as a training data for a model?Anjney Midha [00:04:12]: Yeah, training data for a model.Swyx [00:04:13]: Awesome.Anjney Midha [00:04:13]: Like, an abstract way of saying like, “Hey, I have this data.”Swyx [00:04:15]: Compress it into a model.Anjney Midha [00:04:16]: Like, it makes sense for me in my product, but, like, I could repackage it in the form of a model and sell it. And so it's just a whole new business model for the economy. It also, of course, provides, like, a way of following what Frontier Labs are doing, but in a way that, like, a single developer or a small team of developers can roll on their own. And so - Whenever you have an example of that, like a breakout app that's doing really well, and then some framework for imitating it with - in your own flavor, you have an immediate ecosystem of, like an immediate ecosystem, like, should arise because there's just a huge gap between the, like, decisions that the single company is making and all of the variations in those decisions that, like, a wider ecosystem can create themselves. And so then, you need a marketplace to, like, discover all of those, services and all of those products. There wasn't any place on the internet that, like, was like a home base for LLMs in terms of seeing how much they were being used and seeing who was using them and why.Swyx [00:05:29]: The closest would be Hugging Face.Anjney Midha [00:05:30]: Hugging Face was the closest at the time, yeah.Swyx [00:05:31]: They just started Hugging, like, a few years ago before that.Anjney Midha [00:05:34]: Yeah, and Hugging Face also didn't have the closed-source models.Swyx [00:05:37]: Yeah.Anjney Midha [00:05:38]: And they didn'- you couldn't use the models at the time. and there wasn't data about who was using them. There were, like, a bunch of differences between OpenRouter and Hugging Face, and those differences felt really critical to me, especially when I was just trying to learn about LLMs and, like, why people are choosing, like, Different little ones that are emerging over time.Discord, Open Models, and the Origins of OpenRouterSwyx [00:06:03]: Got it. And then, Ansh, no stranger to wanting more model diversity, at the time, you're a couple of years into your Anthropic journey, which we covered in the previous podcast as well. What was your introduction to Alex?Alex Atallah [00:06:16]: Well, the introduction was, I think, thirteen years before that.Swyx [00:06:20]: Oh.Alex Atallah [00:06:20]: But the OpenRouter handshake happened right over there, if you remember.Anjney Midha [00:06:23]: Yeah.Alex Atallah [00:06:24]: Which - So Alex and I, met, I believe as sophomores now, if I remember at the Stanford Review,Anjney Midha [00:06:32]: That's rightAlex Atallah [00:06:32]: Meeting for the first time.Anjney Midha [00:06:33]: I think so, yeah.Alex Atallah [00:06:35]: Yeah.Anjney Midha [00:06:35]: Yeah.Alex Atallah [00:06:35]: So Stanford Review was the libertarian newspaper on campus at Stanford that Peter Thiel started back in the day. And, whatever-- for whatever reason, I, Alex and I both showed up to one of the meetings, and I remember, the editor-chief was a mutual friend of ours. Lisa was really a really great editor-chief, where, part of an editor-chief's job is to assign responsibilities to people and make sure the work gets done. and I, I may be misremembering the details, but I remember wanting to. It was surprising to me that at the time there was no dedicated technology section in the newspaper.Alex Atallah [00:07:11]: YouSwyx [00:07:13]: Because it's political, right?Alex Atallah [00:07:14]: It is primarilySwyx [00:07:14]: Like, it's talkingAlex Atallah [00:07:15]: It originally started as like aAnjney Midha [00:07:16]: Yes.Swyx [00:07:17]: Yeah, states and all those things.Alex Atallah [00:07:17]: Correct.Swyx [00:07:18]: Yeah.Alex Atallah [00:07:18]: But it, - To take us back in time, you may remember this, but, there was this technology, legislation that was being debated called, the Net Neutrality Act. And net neutrality is, like, inherently this political concept, right? It's, it's about the regulation of - internet broadband access. And so there was a community of us who were technologists, but also debating the politics of the technology. And I thought the Review would be a great place - to, like, write about that. And I was working on, I think, a net neutrality article, and I remember proposing, “Well, maybe we should start a technology section.” And Alex was one of the only people who said, “Yes, that would be cool.” And said. I forget whether we ended up writing stuff together, but - that's when we first met,Alex Atallah [00:08:03]: Was 2011 or twelve. I forget which year it was. It was one of those.Anjney Midha [00:08:09]: Yeah.Alex Atallah [00:08:09]: It was at Old Union, if I remember correctly.Alex Atallah [00:08:11]: That's where we used to meet. But, along the way, Alex and I have had a chance to, To hang out often. And probably the time when we had the most professional overlap was when I was running the platform at Discord, and it had become this explosive platform for cryptoSwyx [00:08:32]: YeahAlex Atallah [00:08:32]: And NFTs in the middle of the pandemic.Swyx [00:08:35]: Which also, by the way, you were in charge of safety and security as well, right?Alex Atallah [00:08:38]: I was the head of platform, which meant all of the crypto - the DAO and NFT launch security debugging fell onSwyx [00:08:45]: And their phishing and.Alex Atallah [00:08:47]: The phishing, the social engineering attacks, the katana DDoS that we were getting hit by. but it's around the time I first started teaching security at scale at Stanford, CS 153. And Alex was on the, - at OpenSea at the time, and I was trying to figure out how we could defend against all these attacks that we were. Like, and at peak, I forget, if you remember how much NFT volume was running throughSwyx [00:09:10]: DiscordAlex Atallah [00:09:10]: Discord, but it was, like, a meaningful amount of, like, it was, like, several billion dollars in NFT volume of GMV, so to speak, were running through the platform, and it was all coming from OpenSea. It was these, like, buy, sell,Swyx [00:09:20]: TheAlex Atallah [00:09:21]: ServersSwyx [00:09:21]: The D in DAO is Discord.Alex Atallah [00:09:25]: Yes. And so that's when I think we had hung out professionally. But a year after that, OpenAI gave Discord early access to GPT. Sorry, three. No, it was five. Yeah, five, which is the RL version of three. And that's around the time we made a Discord bot with, OpenAI for internal deployment, and that's when I realized we would need. Like, since I was part of the deployment team.Anjney Midha [00:09:50]: What was the use case?Alex Atallah [00:09:51]: There were two that were. And there's, there's a post now called “Discord is Your Place for AI with Friends” that somebody sent me recently that I wrote, and published in twenty-three. But There were two use cases. One was Clyde, which was the - like, a party friend inside of Discord that could help you set up your Discord server and talk to you about onboarding and get your friends to hang out more. and then there was content moderation. And one of the realizations we had with content moderation was - it would refuse to moderate. Like, it would just refuse our prompts because the The training was. We were very early in the training era, and it would just. Our prompts would trigger it, its, like, guardrails. And we told OpenAI, “Hey, guys, we need access to the weights because if we're gonna be doing content moderation at scale, we had 250 million monthly active users, we need more reliability that the model will do what we need it to.” And they said, “Well, sorry, guys, that's not how this works. We're a closed-source company.” And so that was my first realization that we needed open models, and the enterprises would need more control over capabilities, and then ultimately would need some control plane or management system to orchestrate these open models. But there weren't no good - there were no good open alternatives until maybeAlex Atallah [00:11:10]: Six months later when Llama came out. And six months after that, I led the series A into Mistral, which was started by Guillaume and the Llama team. And - That, - Around that time is when I remember hearing about Alex launching OpenRouter and going, “These worlds are gonna collide, and I don't know when it'll make sense to team up.” But Alex was so early and could see. I think he was totally right about this ecosystem starting with Llama that then needed, like, a, an easy layer to manage for, especially for. I was approaching it from the enterprise perspective because I had been that, like, the. As the VP of platform at Discord, it was my job to ensure that when we deployed models to, like, 250 million users, they did what we wanted them to. And that was very hard, because if you outsourced it to the labs and they controlled the guardrails and their guardrails are their safety policies. Forbid the model from responding to your prompts. That was quite catastrophic.Swyx [00:12:05]: Yeah. But what, a moderation is the thing that they want to support. And obviously, beyond that, they would - OpenAI would work with you, presumably to give you a moderation endpoint, which they offer for free.Alex Atallah [00:12:16]: It was an interesting use case, that - So they did give us a moderation endpoint. However, as you guys know, every Discord server is like a mini deployment of itself. And so the use case was instead of having human moderators that have to interpret the norms of the community, you just give the, - Often, like every, subreddit, Discord servers, public ones have their own rules that the user, the users create.Swyx [00:12:41]: Oh, yeah. We run the LinkedIn Discord in. Yeah.Alex Atallah [00:12:43]: And then humans used to read those norms and then enforce it every day manually, like observing each message in these communities. And these communities have like millions of users. So we had a 5,000+ person team globally in the, on the Discord content moderation team. These are outsourced contractors who had a really tough job. And so the idea was instead, if you could give the norms of that server To the LLM, then the LLM would do custom moderation for that server. It's almost like a, like context moderation for that server. And many of those servers' norms just violated OpenAI's rules. And so - It was like we had our own custom eval. So each server had its own custom eval. But Discord-- at the time, OpenAI's evals, we were all soAlex Atallah [00:13:28]: Primitive in our thinking about how to deploy these LLMs that often the training prompts were super handed. It said, “Oh, anything about Harry Potter, anything that has trademarked content, don'- refuse.” And if it was a fan - Harry Potter fan community, this is a real use case, that had content moderation, the LLM would just refuse.Swyx [00:13:48]: Yeah.Alex Atallah [00:13:49]: And that was just not precise enough.Anjney Midha [00:13:52]: Another one that we heard was like if someone was trying to write like a detective story, and there's one chapter with a lot of violence, like maybe someoneAlex Atallah [00:14:01]: RightAnjney Midha [00:14:01]: Like kills someone, the LLMs would just refuse to, like, help with that part of the story.Alex Atallah [00:14:07]: Yeah.Anjney Midha [00:14:07]: And then - like, we used to be like, okay, this is not like structurally inherent to LLMs. There must be, like, some choice out there so that I can, like, switch to another model, when I'm getting, like, a refusal or a bad result from the main one that I have. And that, like, tension also drove me for a marketplace.Why “One Model Wins” Was the Wrong BetSwyx [00:14:28]: Yeah. I think that is well accepted now. What was it like back then when you were raising or, starting this? did people get it? what was the, some of the struggles? I like getting stories out of him about how other VCs don't get it. So like anything you wanna, talk about, now - Let's, let's call it, that the early journey of OpenRouter is done, right? You can obviously talk about some of the early days stuff.Anjney Midha [00:14:54]: Well, I was gonna say that, like, the biggest objection we got is big model win, which is - all of theSwyx [00:15:03]: Scaling laws.Anjney Midha [00:15:04]: Huh?Swyx [00:15:04]: Scaling laws.Anjney Midha [00:15:05]: Yeah, scaling laws, and natural network effects are just gonna accrue to one company, which will be - It'll be a Google-style monopoly, just like how Google won the search market, by a large margin, and you'll just be fighting for scraps at the end. That was probably the biggest objection we got. it is interesting that Google won the search engine race with such a huge margin. I think, like, had there been more interesting benchmarks or had, like, search engines been, - had people, like, seen them a little bit more like LLMs where they're services that you can build companies on top of, that might not have been the case. but LLMs don't merely have a user interface. They're also, like, ways of building entirely new businesses. And, a Google-level monopoly would be like the Dutch East India Company times, quadrillion in magnitude because the whole economy ends up, like, depending on the one monopoly as well. So it didn't seem like would be a really crazy outcome if that happened. And it's also less likely because the economics of, like, creating good competitors are much, like, much more decentralizable.Alex Atallah [00:16:25]: Everything Alex said is true, And I came at it from a completely different perspective, whichSwyx [00:16:31]: Yes, this is why we're here.Alex Atallah [00:16:32]: The scaling laws were never - In my mind, were always a feature, not a bug for why OpenRouter would be very valuable. Because, I was one of the first investors in Anthropic, and it was obvious to me that other researchers in our friends - I went to grad school for machine learning, and I just had a lot of friends in the ML community who it was very obvious to us that the bitter lesson holds. And so I was like, “Oh, fantastic. Now we have at least two proof points that compute scaling works.” It was OpenAI and Anthropic. and by the time I think we decided to team up on OpenRouter, I had already invested in Mistral and Black Forest Labs and Luma. So there was multiple model companies and teams that I was, working with.Why Model Labs Struggle With DistributionSwyx [00:17:14]: But you did other modalities, whereas this is literallyAlex Atallah [00:17:16]: Across different modalities, yesSwyx [00:17:17]: Text.Alex Atallah [00:17:18]: Exactly. And it was so obvious to me that an ecosystem of different kinds of models were being created, and that this whole narrative of, like, Only one company will dominate like Google was, well, like maybe true, but one, I don't believe that. But two, there was so much extraordinary innovation happening across several different research teams. But the shared problem I was noticing across all of them was often, the research teams were fantastic at figuring out how to reason about new capabilities. They think in terms of capabilities, but never - like, are not developer mindset-oriented. Like, what happens after the training is done and the checkpoint comes out? Like, you'd be shocked how, like, similar the early training teams at OpenAI, sorry, Anthropic, BFL, Mistral, were in their, like, default approach to. Taking their research out of the, lab and scaling their impact, which is often, oh, the checkpoint is done, put it out as an API, done, and then there'd be crickets. in the case of Claude, the first Claude checkpoint was done a year before they released it internally. And then ChatGPT came out, and we decided, okay, yes, it's a good idea to release a Claude version externally.Alex Atallah [00:18:34]: And they had no plan, like no plan for how to get developers to try it out. And so if you go to the Claude one blog post, you'll notice there are, like, three developer examples for users of the API, and one is a Discord bot, and the second is Vivian, my wife's startup called Juny Learning, ‘- And then there was, like, Notion, because these were all friends of, like, the Anthropic Because that's how - like, last minute the planning was around, hey, once the model's done training, how do you get it out to the world? There was no distribution platform that understood what developers needed, all the key management, provisioning, like, simple, like, endpoint management, versioning control. Like, all these things that the scientists and researchers go, “ that's plumbing. I don't really think about it.”Swyx [00:19:15]: Implementation detail.Alex Atallah [00:19:16]: Right. And instead, Alex came at it from that perspective. And so, it was so obvious to me that, like, every single lab I was funding would spend - like, literally sometimes billions of dollars into training, and then a checkpoint would be done, and there'd be crickets, like, during early access because they're like, “Oh, that's right.”Alex Atallah [00:19:35]: It's hard to use a checkpoint to make anything. You need a whole bunch of plumbing around it to make it usable by a developer. And so by the - I think - it was so obvious to me that a distribution platform like OpenRouter was critical to have in the ecosystem if we wanted there to be competition to Google. Like, unless-- ‘cause with Google, DeepMind is done training a new checkpoint, and then they push a button, and it gets blasted out across all their surfaces from Google Docs to,Swyx [00:20:01]: Everywhere, even if I don't want it.Alex Atallah [00:20:02]: Everywhere. You wanna know about, like, on Android, like, overnight, they can deploy a new checkpoint to, like, a billion devices, right? And that invisible infra advantage, distribution advantage, most people don't realize, but until OpenRouter showed up, - you had to think about all of that yourself as a model lab. And it was very daunting. at Anthropic, I think it took, well, more than twelve months to get to our first 10 million in revenue. And in contrast with Black Forest Labs, I remember the early days, you guys had a conversation with the BFL team, and, it was so simple for OpenRouter to say, “Oh, no problem. Like, the day you launch, we can send 1 million developers to you.” that was crazy. That was like a step function change in, like, an hour.Swyx [00:20:46]: Is that a real number, a million?Alex Atallah [00:20:47]: I,Swyx [00:20:48]: Okay. All right.Alex Atallah [00:20:48]: I think today it's, like, 4 million. How many developers are on OpenRouter today?Anjney Midha [00:20:52]: Over ten,Alex Atallah [00:20:54]: Yeah.Anjney Midha [00:20:54]: Over 10 million, but, like, it's, it's hard to, youAlex Atallah [00:20:59]: I, yeah, I don't know how to. Yeah.Anjney Midha [00:21:00]: We do a lot of, like, account duping work, but, noAlex Atallah [00:21:04]: If you could get 1,000 developers, just to put in context If you get 1,000 developers who try the model on day one after you release it and just, like, do inference and give you feedback, that's a thousandAnjney Midha [00:21:15]: That's hugeAlex Atallah [00:21:16]: More developers than they knew how to get to on their own.Swyx [00:21:19]: Well, BFL had a reputation, but yes.Alex Atallah [00:21:21]: They had one in Stable Diffusion.Swyx [00:21:22]: Yeah.Alex Atallah [00:21:23]: And with Mistral, I don't know if you guys remember, but the first checkpoint they released was, like, torrents. It was, like, torrent weights.Swyx [00:21:31]: Yeah, they just put up a magnet link.Alex Atallah [00:21:33]: Yeah, there was no API.Anjney Midha [00:21:34]: Yeah.Alex Atallah [00:21:34]: Because they didn'- they weren't infra people.Alex Atallah [00:21:37]: ? Like, it's like, okay, download these weights, and you guys go figure out how to host it.Swyx [00:21:39]: Well, he has a story on his side, yeah.Anjney Midha [00:21:41]: Yeah, in addition to the, like, building a really good developer experience around it, the marketing that we do on, like, for different models is totally different and perceived totally differentlyAlex Atallah [00:21:54]: RightAnjney Midha [00:21:54]: From the marketing that a model lab does for itself.Alex Atallah [00:21:56]: Yes, 1,000%.Anjney Midha [00:21:57]: Right? We are like a, neutral layer looking at this market like it's a big dark room with all the corners completely obscure to users, and users are walking into the room and, like, feeling aroundAlex Atallah [00:22:09]: YeahAnjney Midha [00:22:09]: And trying to figure out what objects to grab off the tables and, like, build into, their companies. And it's just an insane way of working. Like, models are not products where you can just enumerate all their features onto a web page. They're all black boxes, including the open weight ones. So you need to, like, shine lights on all corners of this room, so that people can see what makes this model good, and you need the company shining that light to be a neutral third party, which is what we specialize in. So the, like. It'- In addition to developer experience, there's also, like, a very important, like, marketing and product packaging componentAlex Atallah [00:22:50]: YeahAnjney Midha [00:22:50]: And a way of, like, routing and discovering models becomes, like, critical to your market as a provider or a model lab or a server tool and more in the future.“Just a Wrapper”: Why VCs Misunderstood OpenRouterAlex Atallah [00:23:03]: And this value, to your earlier point about how many VCs, like, just don't. One of my biggest frustrations is that venture capitalists, many of them, like, just don't have any operating experience in the field. so unlike a traditional investor who's just maybe come up through the ranks as, like, a associate working on financial modeling or maybe hasn't been a real operator in the field for, like, more than ten years, which is a big part of the industry now, I had just arrived at a16z, like, a year after running the platform. And so I knew what the challenges were of, like, building a real - great developer experience and like, being able to create a working piece of software with a model. And there were a few, I won't name names, but there were investors who were looking at OpenRouter, and, felt at the time, like, when I would compare notes with people, that it was just, I quote unquote, “just a marketplace.”Swyx [00:23:59]: Yeah, just a thin layer, just aAlex Atallah [00:24:00]: CorrectSwyx [00:24:00]: JustAlex Atallah [00:24:01]: A wrapper or whatever on other people's APIs. And I was like, “You have no idea how strategic the value that OpenRouter has created by being able to orchestrate even three.” APIs in production. The amount of both engineering work and community design that goes into getting that live and running in production at the scale the OpenRouter team had started just doesn't happen by default. And that was one of the things that stood out to me about Alex from the earliest days. Like, he just understood, like, - from a systems perspective, like, how do you get these flywheels going? Like, that stood out to me with OpenSea when we were working together on the NFT integration at Discord. Like, Alex had a level of community-- like, systems thinking on how you get these flywheels going that most scientists and machine learning people just don'tAlex Atallah [00:24:48]: Think of. Like, we often think in terms of training.Swyx [00:24:52]: It's a linear stage.Alex Atallah [00:24:53]: It's this linear pipeline.Swyx [00:24:53]: There's no loop yet.Alex Atallah [00:24:54]: Yeah. It wasn't until much later that the modern context feedback loop cycle really got standardized in the industry. But at the time, if you remember, machine learning was like. Like, mostly we did a lot of ML, like, when I was in grad school on a laptop. So you just, like, download a dataset, ran some ablations, and you looked at the loss curves, and you're like, “Great, I made AI.” And the idea that you have to, like, deploy those capabilities, collect feedback trajectories, then, like, put those into a continuous loop, like, came much later. And it was very counterintuitive to the - like, the traditional AI mindset. I do remember doing the investment phase for, OpenRouter, I just didn't try and educate a bunch of other VCs on why it was not just a marketplace. I was like, “ what? I'm just gonna invest.”Anjney Midha [00:25:41]: Yeah.Alex Atallah [00:25:41]: And I'm going to, like, take the opportunity to partner with Alex, and if - no other VCs get it, that's totally fine. ‘Cause at the time, - it was not obvious, I think, to several of the investors that, like, OpenRouter was not more than just a wrapper around APIs. And - that infuriated me. And I was like, “ what? I don't have time to debate you. I'm - we're gonna, we're gonna invest.” And then I think, like, a month later, Matt Murphy marked it up by 10x. Like, - I think. I forget what the exact money was and so on, but, to his credit, Menlo Ventures realized, “Okay, there's much more strategic value here as well.” Maybe you didn't hear all these conversations behind the scenes But that frustrated me a lot. there's a lot of this, like, opining about wrappers. and if you're like, “Oh, an app is just a wrapper on a model,” then, like. And, OpenRouter is, like, this wrapper on top of other APIs, and this is the most stupid, reductive framework.Alex Atallah [00:26:31]: And so it's clearly somebody who has no experience deploying product at scale.Swyx [00:26:34]: It's the thing you dismiss other things with. Like, you're a - everyone's a wrapper on everything, right? Like, and there's, there's some Some wrappers have value.Alex Atallah [00:26:40]: Investors are wrappers and LPs, right?Alex Atallah [00:26:42]: Like venture capitalists. So, yeah, it's all wrappers down, all down to bare metal, I guess, and like energy.Swyx [00:26:46]: Yeah, there - When I started the whole AI engineer, I guess, the coining, in 2023, like, that was, like, the number one pushback is that this is no value. You should just train models.Anjney Midha [00:26:56]: Right.Swyx [00:26:57]: And, yeah, obviously this is, like. you guys are one of the testaments to the fact that you can build very valuable wrappers, but also very valuable model companies.Alex Atallah [00:27:06]: It's so, hard to be. Like, the day a model launches, the fact that you have an OpenRouter, endpoint for that model frequently at the top of Hacker News on day one, people don't realize the amount of work that goes into accomplishing that. And OpenRouter used. Like, that would happen over and over again, and I remember going, “People have no idea how hard that is.”Alex Atallah [00:27:30]: That's not.Swyx [00:27:31]: Yeah, we've covered some of the inference engineering that goes behind,Alex Atallah [00:27:34]: YesSwyx [00:27:34]: Some of - with Base Ten and all those. Well, today you have, all those, like, cool code name things that people guess what Oxy Alpha is and all those things. But, like, I guess one of the things that you're teasing is, how do you get that initial flywheel going, right? Because today you have your scale and your reputation, all these things, so obviously you - you're driving immense distribution. But when you were early on, when it's mostlyBootstrapping OpenRouter Through CommunityAlex Atallah [00:27:55]: The bootstrap, yeah.Swyx [00:27:56]: Yeah.Alex Atallah [00:27:56]: What was the bootstrap like?Anjney Midha [00:27:58]: To bring it back to early Discord days, I think we, like, initially connected with. This is an OpenSea story, technically. But, and we initially connected when you were at Discord, and we talked about, like, - the Axie Infinity server.Alex Atallah [00:28:13]: Oh, yes. Yes.Anjney Midha [00:28:14]: This server was, like, the biggest server at theAlex Atallah [00:28:17]: YeahAnjney Midha [00:28:17]: At Discord.Alex Atallah [00:28:18]: That's right.Anjney Midha [00:28:19]: And you were like, constantly bumping up theAlex Atallah [00:28:22]: The limits on the server. Oh, my GodAnjney Midha [00:28:24]: Of how many people could be in the server.Swyx [00:28:24]: For those who don't know, like, 10% of Philippines was Axie.Alex Atallah [00:28:29]: Was on that server. That's a big hit.Swyx [00:28:31]: It was, like, a meaningful contributor to the GDP of the country.Alex Atallah [00:28:33]: It was an NFT, like, crypto game, but itSwyx [00:28:35]: It was like a Pokémon breeding thing.Anjney Midha [00:28:36]: Yeah.Alex Atallah [00:28:36]: Yeah. Similar. Yeah. There was battling, there was breeding, and then there was, like, a marketplace for trading.Swyx [00:28:43]: Earn as well.Alex Atallah [00:28:45]: Yeah, earn. And, like, the graphics were really cute and fun, and you like, you get emotional about your Axie that you make. So to, like, start a community like that, which we had to do many times at OpenSea with every early project, for us to create a marketplace for it, we need to make sure that the, like, the community wants it.Anjney Midha [00:29:09]: Right.Alex Atallah [00:29:09]: And it's like building something that people want and going and telling them about it. Like, you can do that on a one basis, but there's way higher leverage to do that in a community where everyone can talk to you at the same time. So we spent a lot of time, like, building things that the community really wanted. We did the same thing for OpenRouter. And, like, the Axie community was one of, like, a zillion communities we did that with. And Anj, like, saw us doing it and. ‘Cause you could just see people sharing OpenSea links constantly in that Discord. Like, users sharing links is a really clear indicator that, like, something important is going on. So we spent, a lot of time, like, first figuring out what the gap is in the technology that people care about. Like, what was the actual problem that needs to be solved? in early LLM days, it was, OpenAI refusing to finish the prompt or,Anjney Midha [00:30:09]: YeahAlex Atallah [00:30:10]: To, like, complete the task. It was also.Anjney Midha [00:30:13]: Inability to customize models. and so there are communities that, like are just completely blocked on that issue, and those are the communities that are most useful to learn about and dive into and explore.Alex Atallah [00:30:28]: Something that really struck me at that time, - as I was just hearing your talk, I remember noting - you may not remember this, but we - we had these, like working, Zoom calls that we were doing a sprint around for, like this OpenSea integration with Discord. and, we'd, we'd - it was myself, my engineering team. I think you were there. And I remember, Alex, in the middle of one of those calls, just like there was like silence. we were all like, “Oh, yeah, this totally makes sense. Let's do this.” And then there's - every, like everybody aligned. And Alex was like, “No, this makes no sense to me.” And everyone's - I remember going, “What? Like, it works. Like, you click on a link and this, then it bounces you out to, like, OpenSea.” And he was like, “It's not a good user experience. Yeah, we should not do this.” And I remember going, he was the only one person out of all of us to raise his hand and go, yes, it made sense from a technical implementation perspective. Like, we were bouncing the user out into the, into OpenSea. And so it kinda checked the box of the product manager's requirements on both sides. But Alex went one step further and was like, “ what would be better, guys? If we just embedded the experience right here inside of Discord so the link opened up as an embedded iframe, and you can just check out right there.”Alex Atallah [00:31:47]: And not one person on the call, and there's like seven of us who had met, like, week after week.Swyx [00:31:52]: And it's the guy who doesn't work for Discord.Alex Atallah [00:31:53]: And it's the guy who doesn't work for Discord.Swyx [00:31:55]: Like, technically, you benefit if they bounce.Alex Atallah [00:31:57]: Exactly. And that was, like, adversarial. To keep the user inside of Discord would be adversarial to OpenSea. And yet Alex put that user experience first. And I was like, “That's special.”Swyx [00:32:08]: Wow.Alex Atallah [00:32:08]: Because it's very hard to have somebody who's technical like Alex and understands the developer flow, but also understands the best user experience and wants to prioritize that. And that's two sides of the flywheel that if you can get spinning, like is often hard to stop. And you just reminded me, like that one was one of those moments where I go, I - I realized I gotta be better at user experience because I should have been the one who came up with that, and I didn't. And I learned from you. And, I think that went into one of our case studies for the PM training program at Discord.Swyx [00:32:34]: Whoa.Alex Atallah [00:32:36]: I don't know if it there is Because ofSwyx [00:32:38]: You need an Alex is the conclusion.Alex Atallah [00:32:40]: Yeah. You need an Alex. And this is why I'm not, nobody should be surprised why Stripe decided like they had to buy OpenRouter because it's a really rare combination of people who understand the machine learning community, the developer experience, and the user experience. And putting all that together has resulted in this extraordinary scale that very few other marketplaces have been able to achieveWindow AI, BYOM, and Finding the Right Form FactorSwyx [00:33:02]: Yeah.Alex Atallah [00:33:02]: Over the last, five years.Swyx [00:33:04]: Yeah. Well, we should talk about the other reasons for acquisitions, whichAlex Atallah [00:33:07]: Yes, we should.Swyx [00:33:07]: You've written about. I wanna proceed somewhat chronologically as well. So - there is a point that, one of the questions that, Dave from H of Zero sent in was, when did it - really started to work? And you brought up Mixtral. I don't know if you wanna bring up that story.Alex Atallah [00:33:22]: Oh, yeah.Swyx [00:33:23]: Which obviously you overlap with, so.Anjney Midha [00:33:26]: Yeah, the MoE was. I don't know when. there's no like one moment where I was like, “Oh, this is, officially starting to work.” It wasSwyx [00:33:36]: The moment where you had a Chrome extension, like, really super early on.Anjney Midha [00:33:39]: Oh, yeah. But, well, - yeah. So before OpenRouter, I wanted to, like, explore a bring-your-own-model experiment. And,Swyx [00:33:47]: Which anyone familiar with crypto is like, yeah, Phantom and all these things.Anjney Midha [00:33:50]: Yeah. So it felt like doing a MetaMask analogy for AI would be a fun way of exploring that. And at the time, there were no AI apps. There were probably as many AI apps that were, like, hitting AI - like, hitting an LLM via an API call as there were, like, games just doing it in JavaScript. like there was a, there was a moment in time where it could have been the case that web apps call LLMs through the browser, like through some desktopAlex Atallah [00:34:27]: Yes.Anjney Midha [00:34:27]: Managed app that is controlled by the user. and of course, there are like, I think, many reasons that did not happen. But back when the days were that primordial, I built a Chrome extension called Window AISwyx [00:34:43]: With Plasmo.Anjney Midha [00:34:44]: With Plasmo.Swyx [00:34:45]: I had come across early on, and I was like, “Who's gonna use this?” You did.Anjney Midha [00:34:49]: Plasmo had a couple, like, I think Phantom was using it. there were some other, like real companies using it.Alex Atallah [00:34:56]: It was like a shim.Swyx [00:34:57]: React for Chrome extension. It compiles to allAnjney Midha [00:35:00]: Yeah.Alex Atallah [00:35:00]: I see.Anjney Midha [00:35:00]: Like Next.js for Chrome extensions.Swyx [00:35:01]: Next.js, Next.js.Alex Atallah [00:35:02]: Okay.Anjney Midha [00:35:03]: And yeah, built Window AI on top of it. The creator of Plasmo, like started contributing code to Window AI, in GitHub, and that turned out to be Louis VicchiAlex Atallah [00:35:15]: Oh, you'Anjney Midha [00:35:15]: Who is the founder of OpenRouter.Alex Atallah [00:35:17]: That's right. You have told me this is how you met Louis. Yes.Anjney Midha [00:35:19]: Yeah.Alex Atallah [00:35:19]: Okay.Anjney Midha [00:35:20]: So, that allowed users to like configure which model they wanted to use for a web page in their browser, and then, like the app would just call out to that model when it needed to do things. not the right form factor for LLMs, but, it's like fun experiment. You learn a lot, and like I open sourced it. And the main learning is like, okay, this has to be an API, and it has to look a little bit - like, there has to be more of a developer experience here and more of a discovery experience as well. Like, I don't know where to use these models, and a little Chrome extension is not gonna help me discover. It's not enough real estate. I need more space. I need visuals. I need graphs. I need, examples. I need images. I need to, like, I need to be able to, like explore both as a human and as an agent.Crypto, Midjourney, and the Early Generative AI EcosystemAlex Atallah [00:36:10]: Yeah.Anjney Midha [00:36:10]: So that's how OpenRouter came to be.Alex Atallah [00:36:13]: A meta point that.Alex Atallah [00:36:16]: I think is underappreciated, but Alex is reminding me, is that we were quite lucky that we were so. we were, like, adjacent to the crypto community in those days. Because in hindsight, crypto ended up being like a dress rehearsal for generative models, right? If you think about the Axie experience, Alex is totally right, there were not that many AI apps at the time. And while I was dealing-- my job was to be the head of platform at Discord, which meant to be a general purpose place for communities and friends to create-- for developers to create apps and bots and, other services that could be deployed across Discord. And while 80% of the attention at the time was being spent on crypto, because that's where all the NFT volume was, there was, like, twenty percent of my time I was spending with a friend, who would get hotbot with me and ask me for. We would play Magic: The Gathering on weekends, and he was working on a little Discord bot that could take a text input and turn it into an image, and it was called Midjourney. YouSwyx [00:37:15]: Is that David?Alex Atallah [00:37:15]: It was David Holz.Alex Atallah [00:37:16]: He was a good friend. And David and I have both been failed ARVR founders, in the before that. And, I remember this. Midjourney was one of the fastest-growing communities we had after Axie Infinity started to peter off. And many of the, like, the abstractions and the infrastructure decisions we made to scale Axie happened just in time because they. Axie did this and then fell off a cliff. And then as Midjourney was taking off, we, like, explicitly decided to help David make the server, the Midjourney server, as the primary place for interaction with the model, because it was very hard for people to understand how to use the model if they couldn't see other people using it and copy them. And so the single-player Midjourney web app on its own, like midjourney.com, had, like, terrible retention because people would show up, they'd see this empty field. It's like E 2, and they would type in, like, cat or dog. And it was, like, paralyzing for them to have this blank canvas that they had to fill because they'd never used an AI model before. But instead, in a Discord server, you could see other people using it and riff off of their prompt, and the engagement was off the charts. And so scaling, Midjourney from zero to, like, 10 million monthly actives was a much smoother approach Axie Infinity. And so,Swyx [00:38:29]: Don't forget the best of four pictures, and you choose one.Alex Atallah [00:38:31]: The best, yeah, and then the other, weSwyx [00:38:32]: Which is the feedback loop.Alex Atallah [00:38:33]: The RLHF feedback loop, which, by the way, separately, like, Tom Brown, David and I used to play Magic: The Gathering on weekends. And so, like, it was one group of friends would hang out, and we'd. Like, these concepts were all being discussed all the time. But, there was.Alex Atallah [00:38:47]: I think there were few of us who bridged both the crypto worlds and the AI worlds. And compared to crypto, where it was - the question was always, what's the use case, for this technology? There was never any need to ask that for AI because it's, like, the use case was so visceral. It was like, I can create now anything at - I can imagine. I can write novels, I can code. And the infrastructure that those of us who believed in the distributed systems, like, value of crypto, like the censorship resistance part, found this use case that was explosive. And I think between Midjourney, the, Claude was a Discord bot launch, that we were using internally as an LLM. ElevenLabs had a TTS model that we had on Discord as well. Like, Discord became this petri dish for, like, early apps to innovate. And I don't think it's a coincidence that they found a home there before OpenRouter gave the world, like, a public home store or, like, a, storefront. Discord was this, like, almost petri dish storefront that - had, like, piggybacked on the infra we'd built for crypto communities. And then I think Alex was one of the first people to realize, wait a minute, like, these apps need their own home, on the internet. And then OpenRouter, to me, was a continuation of that community's needs. And of course, there was the crazy distribution that you enabled for a lot of these developers.Why OpenRouter Couldn't Just Live Inside DiscordSwyx [00:40:07]: So then my question is, how come you were. My perception is OpenRouter is not that Discord-centric, right? You have a Discord.Anjney Midha [00:40:14]: Yeah.Swyx [00:40:14]: And you use it to engage your community, but it's not like Midjourney where, like, no, that is like the primary way people experience OpenRouter.Anjney Midha [00:40:21]: Yeah, Midjourney, like, it really helps to see visually really quickly how people are using the model and how to prompt it.Swyx [00:40:29]: Yeah.Anjney Midha [00:40:29]: And I think that is partly why the server was so critical. It's like it is the user experience. It adds a ton.Swyx [00:40:36]: Yes.Anjney Midha [00:40:37]: And you can go the whole mile with just, like, prompting via Midjourney, like, the, via the Midjourney Discord server, getting your images and then sharing them and having fun. For OpenRouter, for LLMs, like, you need a lot of user experience around LLMs to make them, like, really usable.Swyx [00:40:54]: Charge point.Anjney Midha [00:40:55]: And yeah.Anjney Midha [00:40:57]: The, like, seeing the examples of other people is also not as useful because it's a lot of stuff to read. It takes a long time.Swyx [00:41:03]: Yeah.Anjney Midha [00:41:04]: You need, like, based integration. Not possible to do in a Discord server. You need, Or technic- it's possible. I shouldn't say that. It's just not a great developer experience. you need, like, - you need governance for. At the point where you got based integration, now you need governance for managing the LLMs that have access to it, the data policies, which teams. All that stuff needs a lot more than a Discord server can provide. So it's justSwyx [00:41:30]: YeahAnjney Midha [00:41:30]: It's not the right.Alex Atallah [00:41:32]: Well, in addition, you're not wrong, but also there's the very important distinction that, Midjourney was an end user application.Swyx [00:41:40]: Right.Alex Atallah [00:41:40]: And, that's why Discord, which has 250 million monthly end consumers, made, it made sense for Discord to be a host for that application experience. What I knew was gonna happen soon after Midjourney found explosive product-market fit, because we. I think when Midjourney launched, from launch to $100 million revenue run rate, it was less than eight months. And shortly thereafter, Stable Diffusion launched. And, all of us used to hang out in the Discord server. There, I think it was the,Swyx [00:42:13]: The Stability Discord?Alex Atallah [00:42:14]: It was theSwyx [00:42:16]: Yeah, LAION.Alex Atallah [00:42:16]: Yeah, the LAION Discord server.Swyx [00:42:17]: The image community that spawned Stable Diffusion.Alex Atallah [00:42:19]: The image community. Yeah. And so when Stable Diffusion came out, I realized- Oh, now other people can build their own Midjourney.Alex Atallah [00:42:27]: Because until then, Midjourney did not have an API, so they were a stack company, right? They were training their own models, and they were deploying them as an application. But if you wanted to build your own Midjourney, there was no API of that quality. and I think E two was still quite primitive. Like, Midjourney had great quality. And then when Stable Diffusion came out, suddenly there was this new person who - there was - this new capability in the world, which is a developer could create their own Midjourney. And that, I think, created the need for something like OpenRouter, because then you need an API to. If you - if you had the creativity of David Holz and you had Stable Diffusion as the model and you wanted to put these things together, how could you do that without having to figure out how to host the weights? And what OpenRouter, - the shape of OpenRouter enabled is that. Right? When you have open model alternatives to closed applications, OpenRouter's value in the world becomes extraordinary because now any developer can just show up and use theStable Diffusion and the Need for a Model API LayerSwyx [00:43:20]: You just love model diversity.Anjney Midha [00:43:21]: Did you just say the shape of OpenRouter?Alex Atallah [00:43:23]: Oh, no.Anjney Midha [00:43:25]: Were you in cloud? What is this the real Han?Alex Atallah [00:43:26]: I've been, I've been - I'm, I'm misaligned now. I've been overtrained. I've been using Cloud way too much, haven't I?Swyx [00:43:34]: Claude-ish is what people would say.Alex Atallah [00:43:35]: Claude-ish. Oh, God, I gotta untrain myself.Swyx [00:43:38]: Okay. - And I just wanna cap off the Mistral side. my TLDR is there was a Mistral price war, is what they called it, right? Like, round about NeurIPS is twenty-three or twenty-four.Mistral and the Birth of the Inference MarketplaceAnjney Midha [00:43:47]: Yes. DecemberSwyx [00:43:48]: They launched, the Mistral 8x7B, and like the price went down like 80%.Anjney Midha [00:43:54]: Yeah.Swyx [00:43:54]: To me, that's very positive because it's like the first, like, real competition to host Mistral. Is there more?Anjney Midha [00:44:01]: Yeah, that was. I'm, like, trying to remember it, all the things that happened. It. Like, we saw that model come out and immediately saw people say that it was the best model in the world.Alex Atallah [00:44:15]: Yes.Anjney Midha [00:44:15]: Like, this was, to my knowledge, the first time an open weights model was called that in real seriousness.Swyx [00:44:22]: It's hype, right? Is it?Anjney Midha [00:44:25]: It was hype. It was hype. It was also, like, hype from AI influencers at the time. And there were many examples where it was, like, outperforming four. So people really wanted to try it out and see, is this gonna be true for me too? And if so, at what price? And, the, like, inference landscape was really messy.Alex Atallah [00:44:49]: Yes.Anjney Midha [00:44:50]: We cleaned it up. - it allowed, like, providers to compete on price, so we could give you just the best price in one spot. And so it was, I think, the first clear example of, like, a provider marketplace working in a way that adds value to end developers.Alex Atallah [00:45:08]: Sean, you may not remember this, but I think we met for the first time a few days after Mistral came out at NeurIPSAnjney Midha [00:45:15]: Yeah.Alex Atallah [00:45:15]: At a luncheon.Swyx [00:45:16]: Yeah. That's where I also met BFL as well. Yeah.Alex Atallah [00:45:18]: And Guillaume was there.Swyx [00:45:19]: Yeah.Anjney Midha [00:45:19]: I was at NeurIPS at that time.Alex Atallah [00:45:20]: You were there too. And, we had just announced the Mistral investment, and I remember Guillaume was over there, and I remember turning to Guillaume and asking him, Like, “Is it is all the. Like, how are you feeling after the launch of Mistral and seven B?” And, him in his typical French fashion was like, “ it's a, it's an okay model. It's not that good.” And I was like. It was so, in contrast. But I remember him also saying that part of the reason he felt a lot of people Thought that it was better than four was because of the speed. - it was an MoE model that they had, like, absolutely figured out how to make super efficient. It was on the Pareto frontier. And this is an important thing about LLMs, right? Sometimes when they're faster, you think they're smarter, even though, like, if you did, N of, these common, like, evals that are - you do seven tries, and I don't remember. I think we should go back and figure out what the data says, but I wouldn't be surprised if it turns out, oh, on an N of seven attempts, four was smarter on evals, but the perception of on, like, or correctness would be smarter or more accurate. But, people, like, from a human preference perspective felt that it was faster because it - or smarter because it's so fast.Swyx [00:46:36]: Yeah. And most queries do not take that levelAlex Atallah [00:46:39]: Don't take that. That's true.Swyx [00:46:40]: Right? So this is the start of humans as routerAlex Atallah [00:46:42]: Yes.Swyx [00:46:42]: Which then eventually becomes OpenRouter as router of like theAlex Atallah [00:46:45]: Oh, that's interesting way to think about it. Yeah.Swyx [00:46:47]: Like, because humans are the routing mechanism. Like, I will ask the fast model first, and then if, like, oh, not good enough, I'm gonna upgrade manually.Alex Atallah [00:46:52]: Yes.Swyx [00:46:53]: But then he's gonna auto it.Alex Atallah [00:46:54]: I didn't, I hadn't thought of it that way, but that makes sense.Swyx [00:46:57]: Which then there's, there's a lot more techniques, like fusion. Fusion is the thing that we should talk about. Before I move on to those things, I just want to close off the early years. one thing that I observe, which you are also an investor in Arena.OpenRouter vs. LM ArenaAlex Atallah [00:47:10]: Right.Swyx [00:47:10]: And we talked about Midjourney having that feedback loop of, A, B, C, D, and choosing that very. being very important. And you understand the flywheel. So how come you didn't build Arena, and how come Arena didn't build OpenRouter?Anjney Midha [00:47:23]: Well, Arena started before OpenRouter, right?Swyx [00:47:27]: They had the school projectAnjney Midha [00:47:29]: Yeah, LMSwyx [00:47:29]: And then it became a company.Anjney Midha [00:47:31]: LM Arena, yeah.Swyx [00:47:32]: So, but, and I know you had some Arena experiences, like the up comparison type things.Anjney Midha [00:47:37]: Yeah.Swyx [00:47:37]: But you never really went as hard as Arena did.Swyx [00:47:40]: And,Anjney Midha [00:47:40]: In doing up experiences?Swyx [00:47:42]: Yes. And LM Arena did have a router project based on LM Arena ELOs, which they never commercialized.Anjney Midha [00:47:48]: It's hard to do a company that does both because one company is taking data and selling it, and the other company really can't by default. So, I think there is, like, a branding reason that there are two companies here. like, when you set up OpenRouter, there's no training, there are no prompts, right, aside from what your provider policy set. Like, OpenRou- like, OpenRouter can't see your prompts or completions. If you want to see that as an org, you have to opt into it and enable it. And so we're, like, pretty conservative and careful about data policy and security. And privacy. And LM Arena is like, their business model is like oriented around the labs and,Swyx [00:48:34]: Because they give it for free, right? You don't give it for free to give it for free.Anjney Midha [00:48:37]: Yeah.Anjney Midha [00:48:38]: But we do give some. We like have free endpoints too, but like those free endpoints, we, I think we're not collecting any prompts. We're not like monetizing the data unless you, opt into it for some reason.Alex Atallah [00:48:48]: This comparison. you're not the first person to ask me this, and Alex knows this, but I was the interim, like the founder, like first CEO of Arena for the first five months when, and we were helping Anastasios and Waylin spin out of Berkeley. And, I did invest in that before, OpenRouter, but it was very strange to me the comparisons that outside, folks would make between the two projects because the missions were completely different. The founding entity for Arena, we called it the AI Reliability Institute because it was there as an eval service. Like the data, so to speak, that they were originally, offering the labs was how do you make the evaluation of models more reliable than like the state of the art at the time, which was like really just finger in the wind.Alex Atallah [00:49:38]: That's what Anastasios and Waylin's PhD work was as scientists at Berkeley, was on statistical methodologies for correcting, eval estimates, based on like intrinsic biases and how you collected the data.Swyx [00:49:54]: Yes.Alex Atallah [00:49:54]: AndSwyx [00:49:54]: Style control.Alex Atallah [00:49:55]: Style control and stuff like that. And which is very much like a, hey, how. If you're a scientist and you're trying to. the highest expectation customer for Arena was always like a training and, like a researcher at a lab. Whereas the highest expectation customer from my perspective that Alex like really understood and was the mission was to serve was like a developer, right? Who then takes the result of the research and then produces an application that's deployed to the world. It was a completely different problem and person that these two teams were focused on. And so from the outside in. I don't know if you remember this, but I have a distinct memory of a few weeks before we did the term sheet, together for OpenRouter, I'd given you a call because we were trying to get a pooled data set together from OpenRouter and from Arena to, create like an open source repository of prompts. these projects were so different in their goals that it was totally normal to me to be like, “Oh, yeah, let's call Alex and see if he'd want to team up on pooling data,” because they're so different. We need. We don't have that data at all. We. Like, we didn't have API prompts. We didn't, we didn't have like what developers want to do with the models, which is very different from what researchers inside a model lab want to do before releasing the model.Swyx [00:51:15]: Yeah.Alex Atallah [00:51:15]: Does that make sense? And so to this day, I think you see that this difference, even though at a 30,000-foot level you could. I guess you could conclude that Arena and OpenRouter are adjacent, but, the roadmaps, the missions and so on at the time at least were like in very different directions.Swyx [00:51:36]: That ideal customer, I get. I totally get that.Alex Atallah [00:51:39]: Yes.Swyx [00:51:39]: As a founder, I want to own everything, right?Alex Atallah [00:51:41]: That's possible.Swyx [00:51:42]: Like this is clearly an adjacency that I'm like gonna explore that.Anjney Midha [00:51:45]: Own everything meaning like you don't know what to do yet, so you wanna like make sure you catch PMFocus, Anthropic, and Roads Not TakenAlex Atallah [00:51:51]: No, I think what heAnjney Midha [00:51:52]: As quickly as possible.Alex Atallah [00:51:53]: You want to own the entire infrastructure space, and so you expand to whatever demand you can capture.Swyx [00:51:58]: You want to have a play in each end.Alex Atallah [00:51:59]: Yeah, I think that's, that's hard, in reality, because serving multiple customers is difficult.Swyx [00:52:05]: Clearly, this is the one focus, right?Alex Atallah [00:52:08]: Yeah.Anjney Midha [00:52:08]: Yeah. I still think even in the age of AI, like focus is,Alex Atallah [00:52:12]: Is criticalAnjney Midha [00:52:13]: Underrated and critical, not just because you end up with a better product by focusing your humans on it, but also because the world knows what your focus is.Alex Atallah [00:52:22]: One thousand percent.Anjney Midha [00:52:23]: The world can map like, “Oh, I have this issue. Which brand out there is going to help me with that issue? This is the brand that's known for that focus.”Alex Atallah [00:52:31]: Yes.Anjney Midha [00:52:32]: So like if I want real attention on this issue, like this really matters to me, I should go with the brand that cares the most about it.Alex Atallah [00:52:39]: To underscore Alex's point about how important focus is, in the early days of Anthropic, it was not easy to. Like people think that the early days of Anthropic were like super easy because they were on their 3 guys who left, but it was very
Tickets for AIE NYC now open, and apply for the invite-only AIE CODE. Join us!We have an unusual relationship with today's guest: for years since coauthoring the InstructGPT paper, Diogo Almeida had been saying that API-available frontier models have been going down the wrong path, everything from the alignment to refusals to reliability perspectives, that we have dropped every mode other than autoregressive chat-tuned LLMs because of the overwhelming success of ChatGPT.In a launch video now viewed ~40M times (by comparison, GPT4o was 22M, Fable 5 was 15M, Navier Stokes was 74M, and 6 Astra was 137M), Diogo introduced Jev and it immediately took over the AI timeline — we'll skip full Jev explainers because your favorite AI influencer/educator has probably already done one. We also collected:* the official patterns and cookbooks you should see first, from Allie* Jev usecases* speed based - games and computer use* the voice + computer use example we discuss at 1h34 mins* voice + browser control* The must not miss Doom demo* Driving cars in games* Excalidraw* virtual try-ons* “Smart Games”/smart NPCs* guided responses in text messages* Jev for coding agents has an official guide * jev for linting* compacting tool calls* reasonable pushback from Theo - Diogo has published a note on the Tyranny of the KV Cache that you should read as a followup after the pod for Jev + coding agents, because of his belief that Cache Rules Everything* Programming Languages built atop Jev (Diogo's fave)* Jev for analytics replay and user journey review* “dark data”* entity resolution* natural language search* “smart software”* a core goal of Jev is to “disappear into the background” - eg as unremarkable as regex* Jev as a judge* Jev memes* Jev vs LLM capabiltiies* blending transformers and classifiers* about the confidence api* Jev vs GLiNER (note difference/pushback, agreed, agreed, agreed)* Jev on trolley problem* Jev BushInstead we'll focus on what we can uniquely offer — a broader philosophical and mission-based understanding of how and why Jev was created, and what you should expect next in terms of future models from TypeSafe (ReasoningJev?) and what usecases and ideas you should work on vs the 55th low effort clone of Jev's API or doing a generic JevBench benchmark - something Diogo has rejected publicly.Why RLCD: Three kinds of RLHF, and why they are ALL the wrong north starDiogo knows a good deal about RLHF, given that he was on the team that pioneered post-training at OpenAI — and traces the three branches to Christiano et al 2017 (the robot backflip demo), Stiennon et al 2020 (learning to summarize) and his baby, Ouyang et al 2022 (InstructGPT). From there on, every innovation from Function Calling to Structured Outputs to Reasoning felt like a hack on top of the string based, sequence to sequence prediction paradigm. As he mentions on the pod, from 2023-2024 he struggled unsuccessfully, due to both personal and organization underestimation, to train a model that accurately addressed what he saw as the core problem with making LLMs the heart of software: reliability.Jev's core innovation is "Reinforcement Learning for Calibrated Decisions”, a novel, unpublished technique that optimizes for “answers with epistemically honest probabilities on System One tasks” rather than human rated feedback (RLHF) — which causes hallucinations, sycophancy, and permanent reliance on humans — or programmatically verifiable outputs with rubrics (RLVR) — which solves Navier Stokes but exacerbates jagged intelligence and doesn't integrate well with other software.We've talked about the calibration problem before on the pod, but probably the single best place to understand why RLCD became necessary is Diogo's AIE talk, which discusses why a generation of training helpful AI assistants for humans has impaired them for training models for composable, programmable AI for automation.At the end he also teases his contrarian opinion on scaling laws - which teases how to build a modern neolab without the billions of dollars the major labs have…The Bitterest Lesson: Tasks and Data beats ComputeWe spend a good amount of time discussing Diogo's essay on the Bitterest Lesson:His point is that “You get what you optimize for and the bitterest lesson in ML is that the most important part of it isn't ML at all.” - and picking the right north star, eg upvoting for user preference vs being integrated into tool calls - makes everything else fall in line.We're excited to catch up with a freshly dyed Diogo to discuss:* Why AI can solve extraordinarily hard problems but still fail to automate basic work* What System One Models are and why Jev is built for software rather than chat* RLHF, mode collapse, calibration, and the hidden costs of optimizing for human preferences* Why refusals become a problem when AI is buried inside software dependencies* Why TypeSafe rejects public benchmarks and optimizes for intelligence per dollar* The “bitterest lesson”: why the right task and the right data can matter more than compute* Why TypeSafe thinks of itself as a data lab rather than a model lab* RLCD vs. RLHF and RLVR as fundamentally different North Stars for AI* Why reliability and robustness matter more than simple determinism* Jev's programming primitives and how intelligence maps into software control flow* Why developers should decompose AI workflows into small, measurable decisions* How structured state replaces giant prompts and system messages* Why Diogo thinks AI should eventually disappear into the background of software* The “inverse SaaS-pocalypse” and how AI could supercharge existing software* System One vs. System Two intelligence and the limits of reasoning models* Dark data, computer use, real-time intelligence, and Jev's biggest early use cases* Why Jev could reshape coding agents built around a single-model architecture* Why Diogo says he wouldn't pre-train with $1 billion* The OpenAI journey that led to TypeSafe and why he thinks many neo-labs are approaching AI incorrectly* Coding agents beyond the KV cache, shared state, sub-agents, and the multi-agent futureDiogo Almeida* LinkedIn: https://www.linkedin.com/in/diogomda* X: https://x.com/CompleteSkeptic* TypeSafe AI: https://typesafe.ai/Timestamps00:00:00 Jev Launch Week and the AI Economic Revolution00:02:50 What Is Jev? System One Models and Programmable AI00:05:54 RLHF, Mode Collapse, Calibration, and Yann LeCun00:10:29 Programmatic AI, Refusals, and Safety Alignment00:17:21 Why TypeSafe Rejects Public Benchmarks00:20:43 The Bitterest Lesson: Data, Compute, and the Right Task00:24:59 RLCD vs. RLHF and RLVR00:28:42 Why Powerful AI Still Hasn't Automated the Economy00:39:55 Reliability, Robustness, and Determinism00:48:11 Model Versioning, LTS, Speed, and Intelligence per Dollar00:54:04 Inside Jev's API and Programming Primitives00:58:28 How to Build with Jev: Structure, Decomposition, and Small Decisions01:18:28 The Inverse SaaS-pocalypse and AI Disappearing into Software01:33:21 Computer Use, Dark Data, and Jev's Biggest Use Cases01:38:48 How Jev Could Reshape Coding Agents01:41:00 AI Safety, Frontier Pacing, and the Limits of RLVR01:48:03 Why Diogo Wouldn't Pre-Train with $1 Billion01:55:19 The OpenAI Story Behind TypeSafe02:01:41 Why Diogo Thinks Most Neo-Labs Are Getting AI Wrong02:08:00 Coding Agents Beyond the KV Cache and the Multi-Agent FutureTranscriptIntroduction: Jev Launch Week and Developer MomentumSwyx [00:00:00]: Okay, we're in the studio. A special occasion because this week, Diogo, my good buddy, launched Jev, and it's been taking over the complete timeline. How do you feel? What's it like to be you right now?Diogo Almeida [00:00:16]: Emotionally?Swyx [00:00:17]: Yeah.Diogo Almeida [00:00:17]: Never been worse. Like, I'm a ragged corpse of a person right now because there's so much going on, and I'm like a technical CEO, so I have, like, a lot of fires to fight.Swyx [00:00:29]: Yeah.Diogo Almeida [00:00:29]: But mentally, I feel—I say this all the time, and I've been saying this kind of for years in my over-under events. Like, I feel like the entire AI field is like one of those, like, carnival house of mirrors, and everyone is just insane and saying the weirdest stuff that doesn't make sense. And it feels like for just this week, like, I'm on a better in sync with reality and like, oh, people see it now. AI can be so much more than what was once thought.Diogo Almeida [00:01:06]: And like, yes, we are going to make. Like, an AI-based economic revolution is back on the table, and this is f*****g awesome.Diogo Almeida [00:01:17]: I'm so jazzed the developers get it. It's, it's, Yeah, and I want to show my eternal gratitude to the developers andSwyx [00:01:25]: Yeah.Diogo Almeida [00:01:26]: I'm so jazzed about the community and everything. It's so great.Swyx [00:01:28]: Yeah, you were saying yesterday that you decided to prioritize the town hall and not a bunch of, like, VIP, investor-type people because you wanted to make sure that they are the people that you get your most, attention, right? The engineers, the developers.Diogo Almeida [00:01:43]: Yeah, it felt a little like, oh man, I'm talking to, like, really important people right now.Swyx [00:01:47]: Yeah.Diogo Almeida [00:01:47]: I probably shouldn't reveal who.Swyx [00:01:48]: Yeah.Diogo Almeida [00:01:48]: But it feels a little bit dirty for me to, I'm, like, perhaps overly genuine in things. Like, it feels, like, dirty if, like, in my gigantic calendar event of people to talk to, the community isn't one of those.Swyx [00:02:04]: Yeah.Diogo Almeida [00:02:04]: And actually, in my ideal world, it would be, like, community all the time. I was thinking, “Should I host a town hall while walking to your studio?” And I'm like, “No, that's too crazy.”Swyx [00:02:12]: Sure. Yeah. Well, you guys have been hosting town halls on Discord. Discord is now 100,000 people. Your Twitter'sDiogo Almeida [00:02:19]: I don't follow these stats.Swyx [00:02:20]: Yeah.Diogo Almeida [00:02:20]: So holy s**t.Swyx [00:02:21]: Your Twitter's blown up. It was, it was really funny ‘cause, like, at AIE, you were like, “Yeah, follow me please,” and then you didn't, like, provide even your handle.Diogo Almeida [00:02:29]: I'm a noob. I'm a noob.Swyx [00:02:29]: You're such a noob.Diogo Almeida [00:02:30]: I'm a noob.Swyx [00:02:31]: But no, but that, like, that's, like, positive aura that, likeDiogo Almeida [00:02:33]: CoolSwyx [00:02:33]: You don't know how to promote yourself.Diogo Almeida [00:02:35]: Yeah. Someone, like, called me out when I posted, like, “Holy s**t, we're all three twending-- trending topics.” And then they're like, “That's a personal feed.”Swyx [00:02:42]: That's a personal, yeah.Diogo Almeida [00:02:43]: And I'm like, “Oh, no.”Swyx [00:02:44]: Of course, of course it'll trend to you.Diogo Almeida [00:02:45]: Cringe. Yeah.Swyx [00:02:45]: Yes, ‘cause it's what you clicked on.Diogo Almeida [00:02:47]: Yeah.Swyx [00:02:47]: So okay. Let's, Yeah, so congrats on everything.What Is Jev? System 1 Models and Intelligence per DollarDiogo Almeida [00:02:50]: Thank you.Swyx [00:02:50]: We'll talk about more, details as you have them. But let's, for people who are, like, living under a rock or just want, like, the definitive thing, what is Jev?Diogo Almeida [00:03:02]: Whew. Let me think about. That's a hard one.Swyx [00:03:07]: Okay. And I'm happy to, like, re-ask if you wanna kind ofDiogo Almeida [00:03:09]: No. I'm happy toSwyx [00:03:10]: OkayDiogo Almeida [00:03:10]: I'm happy to, like, just jam on it.Swyx [00:03:12]: Yeah.Diogo Almeida [00:03:13]: I will say, like, the first thing that I'm relieved about with this question is now I don't have to answer that question to my parents anymore ‘cause ChatGPT can just explain it.Swyx [00:03:20]: Nice.Diogo Almeida [00:03:21]: So the way I see it is we new-- need a new class of models. We're not attached to naming that class of models. Our-- the most accurate name we've come up with is System 1 models.Swyx [00:03:33]: Yeah.Diogo Almeida [00:03:33]: There will be reasons, but it's-- there's a reason why we don't call them decision models, because, like, they will be. Like, System 1 is beyond that. That's all I can say. We didn't expect this to be our big launch, so we have stuff in the tank.Swyx [00:03:48]: You should have said low-key research preview.Diogo Almeida [00:03:52]: It kind of was, right? It kind of was. But we. So there's a class of models that we describe them as, like, machine-native, System 1, large programmable. I think these are-- is the class of models where the goal is for code to be the consumer. So as opposed to, lar-- pre-trained large language models, which are meant for, like, autocomplete of the internet, or RLHF models, like chatbot instruction-following models, which are meant to, like, reply to text, or RLVR. It's in a weird gray area with RLHF. Like, these are meant to have things that directly are consumed by code, hence the name type safe. So the thing we really want is to have, like, AI, like, be as powerful as possible, and we think the way to do that is to integrate it with software. And we are designing everything, beyond just the outside, the deep internals of the model to be optimized for software. So number one, Jev is our first large programmable model, or a System 1 model, whatever you want to call it. Jev is meant to be optimized for intelligence per dollar, hence the name Jev.Swyx [00:05:03]: Jevons Paradox.Diogo Almeida [00:05:03]: Jevons Paradox, yeah. And it's optimized for intelligence per dollar. I love this debate with people about what is the most important between reliability, cost, calibration, and speed. And Jev is meant to be. Jev will be the name of models that will be on the frontier of intelligence per dollar. There's other ways to optimize it, like, ML, or at least if you're good at ML, it's all about trade-offs. And we are just going all out on that.Calibration, Mode Collapse, and the Limits of RLHFSwyx [00:05:31]: Yeah. And to me, like, calibration is one of the new things that people weren't talking about as much. We've done an episode In the past, with Clementine Foreia of Hugging Face, where they were like, “Yeah, actually, y- they're just.” Or, and this is your whole argument about RLHF, is they're more collapsing towards what you want to hear the mostDiogo Almeida [00:05:50]: OohSwyx [00:05:50]: Or what is most likely, instead of, like, their own internal confidence about a thing.Diogo Almeida [00:05:54]: Can I soapbox on that for a second?Swyx [00:05:56]: Go ahead. Yeah.Diogo Almeida [00:05:57]: Cool. Like, I've been heard that your audience is the most technical, so I actually want to get into that.Swyx [00:06:02]: Yeah.Diogo Almeida [00:06:03]: And if- I went through extreme precision to make sure everything in our launch video is accurate and real. Apparently, that's very unusual. One of the things that no one paid attention to was the downsides of RLHF, in particular mode dropping.Swyx [00:06:17]: Mode dropping or mode collapse?Diogo Almeida [00:06:19]: It's the same thing.Swyx [00:06:19]: Is that what you call it?Diogo Almeida [00:06:20]: It's the same thing.Swyx [00:06:20]: All right.Diogo Almeida [00:06:21]: And I wanna have a blog on this eventually, but I, like, want to tell as many people this as possible ‘cause I think it's a very interesting thing. So the spicy take, I believe in Yann LeCun a lot. I think Yann LeCun's takes are actually among the closest toSwyx [00:06:36]: What about this?Diogo Almeida [00:06:37]: Well, should I address this now or should I wait and go into mode collapse?Swyx [00:06:40]: No, later. Go mode, go mode collapse. I don't know.Diogo Almeida [00:06:42]: So I actually think that among takes, Yann LeCun's is among the most accurate. But he has this very famous/infamous slide about,Swyx [00:06:52]: The cake?Diogo Almeida [00:06:53]: LLMs are doomed.Swyx [00:06:54]: Okay.Diogo Almeida [00:06:54]: Like that one where he, like, has, like, a pie chart with, like, a tiny par-- tiny little thing- and says that as you increase sequence length, the probability of it making an error goes in. Yes, this one. This one. I love this one, because it's one of these things that seems mathematically obvious, but is obviously wrong, right? Like, it's mathematically obvious, but it doesn't empirically hold. And this is my favorite thing to teach people about, like, where youSwyx [00:07:21]: What's the disconnect, right?Diogo Almeida [00:07:22]: Exactly. And may I or you want to tell me?Swyx [00:07:27]: About mode collapse?Diogo Almeida [00:07:28]: Oh, no. Oh, so mode clop-- collapse is related to this.Swyx [00:07:31]: Yeah.Diogo Almeida [00:07:31]: The disconnect happens because if you are in a mode covering or a calibrated distribution, you are, like, not. You are not overly punished about having outliers. You'd expect, like, something. Some amount of the time you'd be out of distribution, some amount of time you'd be in distribution. That's what happens when you cover the distribution. This was like models before GANs. They made blurry images, right?Diogo Almeida [00:07:54]: Instead, GANs mode drop. They, like, drop the minority classes and just do the really common ones. And this is why this effect doesn't happen, right? Like, instead of be-- in order to generate really long strings, without making errors, they need to, like, be extremely conservative because it's e- really easy to see when an error happens. It's very hard to see when, like, a subtle thing that looks correct happens. And that calibration is, like, total poison into, like, the probability distributions of strings.Swyx [00:08:22]: Yeah.Diogo Almeida [00:08:23]: And it's, it's a nuanced take and like, I think that This is why this doesn't happen, and this is why strings are so bad at, decision-making or, overloading the string models are for decision-making is, like, a bad time.Yann LeCun, JEPA, Scaling Laws, and Practical ResearchSwyx [00:08:38]: And while we're on the topic of Yann, do you agree that his fix i- with-- which is like a world model, like a JEPA-type, embedding thing is the right solve? So basically, like, the. One of the reasons that it could fail is because you're trying to reason over token outputs and then, and then just looping back again and going. Keep, continuing going until you reach, like, a end of sentence. Like, is that, And his solve is JEPA, right?Diogo Almeida [00:09:02]: Yes.Swyx [00:09:02]: Which is, like, joint ambition,Diogo Almeida [00:09:04]: YeahSwyx [00:09:04]: Joint embedding prediction. So like, is that the solve or, like, do you have a. Do you have a take on that?Diogo Almeida [00:09:10]: Oh, man. I probably shouldn't talk too much about the insides of ML, but I will say that my brand, other than unhinged, is practical.Diogo Almeida [00:09:20]: Like, even my take here is practical. And like, I'm. Am I a scaling law fan? Depends. It dep-- it's, it's, it's, like, it's. Scaling laws tell you how much better you get at a thing for amount in.Diogo Almeida [00:09:33]: A scaling law does mean exponentially more resources for normally sublinear gains, which looks to be a bad investment unless those, like, linear gains are, like, really valuable. But it's all. To me, it's all about, like, what can we do with what we have to make the biggest possible f*****g difference? I can curse.Swyx [00:09:51]: Yeah.Diogo Almeida [00:09:51]: Yeah.Swyx [00:09:52]: Yeah.Diogo Almeida [00:09:52]: Yeah.Swyx [00:09:53]: We're, we're, we're approved for adults.Diogo Almeida [00:09:54]: Hell yeah.Swyx [00:09:55]: And also we have a scaling law thing if you wanna go into that later.Diogo Almeida [00:09:58]: Oh, I could if we. See, that part is not super relevant right now.Swyx [00:10:02]: Yeah.Diogo Almeida [00:10:03]: I actually. If you wanna go into my bitterest lesson, I think that's more relevant.Swyx [00:10:06]: Okay.Diogo Almeida [00:10:06]: But like, to me, I'm all about, like, pragmatics. And I think that the JEPA stuff is really cool early research. I really love awesome research. Is it practical yet?Diogo Almeida [00:10:21]: Probably shouldn't say. But like, there's just a lot of.Diogo Almeida [00:10:29]: I just think there's just, like, so many diamonds in the rough let all over the research world right now that haven't been polished because people don't know how to, like, do the right task. And I think that what our launch did, it. Does it kickstart us as a company? Like, yes. Will it be great for us as a company? Yes. I think it's gonna be, like, even greater for this direction of, like, programmatic AI. There was going to be, like, a gold rush on top of us for. ‘cause, like, software is super f*****g charged. But I think there's gonna be a gold rush parallel to us as well on, like, all the different ways we can expose things to make software more powerful so people can make even cooler stuff. And then we are back to, like, early internet energy?Swyx [00:11:12]: Yeah.Diogo Almeida [00:11:12]: And I think that's why, like, the Twitter is just like, “Jev.”? It's, it's like. It is a partySwyx [00:11:18]: It's inspiring because it's, it's, like, so different than what we're used to, which is, “I'm sorry you can't do this, but we do scaling laws and only the big labs can do it,” right?Diogo Almeida [00:11:28]: That. Actually, if I. I'll, I'll make a tangent if that's okay.Swyx [00:11:32]: Yeah.Diogo Almeida [00:11:32]: I think you might enjoy this.Swyx [00:11:33]: Really? Our five tangents in. It's good. It's fun. Yeah.Diogo Almeida [00:11:35]: Oh, yeah. I get lost at all my tangents.Swyx [00:11:37]: This is gonna be horrible for the listeners to figure it out, but they're gonna figure it out. It's fine.Safety Alignment, Refusals, and API PhilosophyDiogo Almeida [00:11:40]: Yeah, we can edit it in post.Swyx [00:11:40]: This is my response. Yeah.Diogo Almeida [00:11:41]: So popular thing on Discord, that people keep asking me, I haven't had the time to explain it yet, is why am I opposed to safety alignment and why do we not refuse? I'm not opposed to safety as a principle, but I think that safety alignment is generally misaligned with users. And refusal is just, like, obviously a type error. Like, if you're a human being and you're chatting with, like, a bot or whatever, you're cloud coding, and a refusal happens, like, “I'm sorry, I can't read DNA.py.” that's an annoying time. It's anno- it's, it's annoyingDiogo Almeida [00:12:18]: Right? But you can work with it, right? And you're forced to work with it ‘cause of Stockholm syndrome.Diogo Almeida [00:12:23]: I have stories about that too. I need another tangent deep in here. But like, if you ever want this in a dependency running in the background, what happens if that refuses? What if someone else is using that dependency? They don't know what that system is. Like, you want the software to just stochastically break because a user sent, like, a weird message in there?Diogo Almeida [00:12:42]: Like, that is, like, straight-up insanity. It's coming from a place of, like, people who do not understand software, do not understand programming, and like, they are obsessed with, like, I believe this, horseless carriage of, like, AI coworker instead of unearthing, like, the full power of AI.Swyx [00:13:01]: Fair enough.Diogo Almeida [00:13:01]: Yeah.Swyx [00:13:01]: You want something that is the core kernel that is usable everywhere.Diogo Almeida [00:13:05]: Yes. Exactly. Like, the cognitive core, right?Swyx [00:13:07]: Yeah.Diogo Almeida [00:13:08]: And you need this thing to be s- like, so general, so optimized for its use cases. You want it to be, like, you want it to work on all the future use cases, all the weird s**t that people are doing.Swyx [00:13:19]: Yeah.Diogo Almeida [00:13:19]: We obviously didn't train on any of that stuff. Is it surprising that it works? No, ‘cause we trained on weirder stuff, my friend.Diogo Almeida [00:13:28]: So. But one tangent up about, like, safety alignment.Swyx [00:13:32]: Okay.Diogo Almeida [00:13:32]: Safety alignment makes sense for a product, in my opinion, for, like, ChatGPT and Claude. Like, it, What safety, what makes safety and capability alignment different is capability alignment is, like, about doing what the user wants. That is sick for software engineers. They want their thing to do the thing, and the more predictable it is, the less they have to test it and play around with it. Jeb is not anywhere close to that yet. It could be, but like, there's so many more nines of reliability that we want in order to make it so good, like a database query, that you don't even have to think about it. It is just there when you need intelligence. But safety alignment is, like, the opposite of instruction following. It's when you want to follow someone else's instructions, like OpenAI and AnthropicSwyx [00:14:13]: The RAGs value stack.Diogo Almeida [00:14:14]: Exactly. And this makes a lot of sense for a product. Again, like, ChatGPT should do. Y- you sh- like, if they don't want to, like, do, like, some, not-safe-for-work role play with ChatGPT, that's on them because, like, maybe that's, what their users who have, like, parents and kids want. Like, n- that's fine. But in an API, that's nuts, right? Like, that's completely unacceptable because, like, people need to, like, program around this, and that is, that's so anti-user that it's. It. I'm. Huh. I can be an angry person, so I should try to calm down.Swyx [00:14:52]: It's, People get your passion, and I think that's really good. The one pushback I'll give you is, like, what if we use it to kill people, right? Like, that is the actual. Like, n- the not-safe-for-work thing, it's private, personal, whatever. But like, yes, like, we will use it in war. And like, that is, something that companies can reasonably prefer their APIs not be used for.Diogo Almeida [00:15:14]: I get that. I think that there's, like, pragmatic places where that opinion can be held. I don't think the foundation of, like, a general-purpose technology is that place, personally.Diogo Almeida [00:15:27]: Like, would I prefer that our stuff is not used to kill people? Obviously. Would I prefer it's used for, like, all sorts of, like, great stuff in the world? Obviously. Will I put my thumb in the scale for that? Yes. Will I do it at the technological layer? Absolutely not, because that will fracture the intelligence. Every single time you mean it to overfit to some weird stuff, you're fracturing its intelligence more and more. And like, these things are fractured to the, like. They're so darn fractured right now.Swyx [00:15:54]: Yeah.Diogo Almeida [00:15:54]: So and as a furthermore thing, to me, it's like I think intelligence will be more like a database than a coworker. Like, I don't think it's up to databases to add checks on whether or not they're used for, like, what's something that's not great? Like, CIA. Actually, I don't know what the CIA does, really. You can imagine. You can imagine, killing people who are not even bad or whatever.Diogo Almeida [00:16:21]: And like, I don't think it's the database's responsibility for that. And furthermore, like, a thing that has been weird to me is when people, like, sign up for our thing on Slack and they're like, “Hey, we're gonna deploy this. Can we deploy this thing?” I am just like, “My brother, we are an API. You are a developer. It's none of my business,” right? Like, you shouldn't know what the whole task even isSwyx [00:16:46]: YeahDiogo Almeida [00:16:46]: Because it should be decomposed into small things. We shouldn't be able to know what the downstream users are doing, and that is, like, a good boundary to give software engineers maximum power. Ideally, they use it for the good stuff, and ideally, we can, like, help them and like, we've talked about, like, doing open source and charity and all of that. We have absolutely no time for anything else right now. But like, they will get any of that bias out of the technological layer as long as I'm in charge.Privacy, Benchmarking, and Trusting IntelligenceSwyx [00:17:11]: Yeah, that's great. While we're on the topic, let's also briefly talk about your privacy stuff, terms of ser- terms of use, which, got a little bit ofDiogo Almeida [00:17:18]: OohSwyx [00:17:18]: Misunderstanding. I just wanna clarify that upfront.Diogo Almeida [00:17:21]: Hell yeah.Swyx [00:17:21]: I think this probably takes two sentences from you about, like, you will not. You're not being that restrictive about your API. Like, clearlyDiogo Almeida [00:17:27]: Oh, yeah. Oh, yeah, so yeahSwyx [00:17:27]: Ideologically, you articulate your role as a platform very seriously.Diogo Almeida [00:17:30]: Yes. Yes. I don't know what you're referring to, but like, this was. I've seen a couple of things about, like, benchmarking.Swyx [00:17:38]: Yes.Diogo Almeida [00:17:38]: Like, obviously we're not stopping people from do. Oh, man, I should be careful about what I say. I'm realizingSwyx [00:17:43]: No, you said, you said it publicly thatDiogo Almeida [00:17:44]: YeahSwyx [00:17:44]: That was in the preview period. You didn't take it out for the launch.Diogo Almeida [00:17:47]: Yeah. Okay.Swyx [00:17:47]: And now you're gonna take it out.Diogo Almeida [00:17:48]: So the team is doing stuff thatSwyx [00:17:49]: YesDiogo Almeida [00:17:49]: I'm not even aware of, so it's great to know the team communicated that. I asked them to check in with the lawyers about that.Swyx [00:17:54]: Yeah.Diogo Almeida [00:17:54]: Like, we are obviously not stopping people from doing that type of thing. I'm extremely in favor. So I'm extremely anti-public benchmarks. I'm extremely in fa- I'm medium about private benchmarks that are proxies. ISwyx [00:18:09]: So are you worried about, saturation or, like, training on public benchmarks? So it's, like, easy to cheat.Diogo Almeida [00:18:15]: Not only is it easy to cheat, there's a lot of ins. So I think that we are. Or anyone who's, like, competition with us that, vaguely there is. Like, you could say, likeSwyx [00:18:28]: There's like 50 Jev clones, yeah.Diogo Almeida [00:18:30]: Well, sure.Swyx [00:18:31]: Yeah.Diogo Almeida [00:18:32]: Well, the, these. Let's say that there is competition.Swyx [00:18:34]: And we'll talk about those. Yeah.Diogo Almeida [00:18:34]: Or let's just say that there's. Let's just assume that there's an industry two years from now of people who are doing similar things to us. The thing that we are selling is intelligence per something, per, like, dollar or per second. The. No one. Like, people obsess about the cost and the speed. I believe that is. It's cool, but like, the thing that matters is the intelligence. Like, the cost and the speed are, like, are bad things. You're paying them for something, and you need the thing back, and the intelligence is what truly matters. The problem with intelligence is that there's a je ne sais quoi to it, right? Like, the good model smell. Like, the thing that happened after we launched of, like, two hours later that actually went way bigger than the video, which was like, “Holy s**t.”Swyx [00:19:16]: This is actually usable.Diogo Almeida [00:19:17]: It. WellSwyx [00:19:17]: Yeah.Diogo Almeida [00:19:17]: It's, like, beyond that.Swyx [00:19:20]: Yeah.Diogo Almeida [00:19:20]: Like, the. Whew, the launch was crazy, and people could really sense how hard we care about that, and that's truly what I think the long term of this is. And I think public benchmarks are antithetical to this. Like, they are a way to get people trust in intelligence because intelligence has a je ne sais quoi, but the public benchmarks are extremely gameable. Even if they try not to, they still will. Like, back in the old days, every lab had a team to collect data that looks like MMLU to make it look better, which is just benchmarking with extra steps.Diogo Almeida [00:19:58]: So I believe that in the long run, it needs to be vibes and trust until you put it into a workflow and evaluate it for that workflow and measure it and have your own sense of, like, how it does on the exact workflow that matters. And our job is to keep moving the nines of reliability. This is like an ever-present part of o- of what we need to be doing as a company, and we need to do everything to have people know that this is something we care so much about. Like, if we wanted to, we could have released Jev, like, a year and a half ago if we wanted it to be dumb.The Bitterest Lesson: Tasks, Data, and North StarsSwyx [00:20:34]: Oh.Diogo Almeida [00:20:34]: It. Like, the. My bitterest lesson, right? Like, architecture and Yeah.Swyx [00:20:40]: I'll bring it upDiogo Almeida [00:20:40]: Hell yeahSwyx [00:20:41]: Since you, since you talked about it, here.Diogo Almeida [00:20:43]: Hell yeah. T- like, Sutton says that algorithms beats compute very roughly. Data matters way more than compute, obviously. And doing the right task, having the North Star is the hardest, most important thing. This has happened, in LLM land twice so far, right? Maybe 2.2 times. There's RLHF, which, like, shifted the task to instruction following. No one realized that was possible. RLVR did, like, a tiny little, like, edit to the, to the direction, and now us, right? RLCD. We have a new task, and the goal is, programs in the loop. And yeah, data matters soSwyx [00:21:28]: RightDiogo Almeida [00:21:28]: Unbelievably much.Swyx [00:21:29]: SoDiogo Almeida [00:21:29]: Like, I can't, I can't emphasize it less.Swyx [00:21:31]: Yeah, you consider yourself a data lab rather than, like, a model lab. Is thatDiogo Almeida [00:21:35]: AbsolutelySwyx [00:21:35]: Something. That's the wording you guys use?Diogo Almeida [00:21:37]: Yeah. We are. We will always, like, care so much about data. To me, model capabilities means data. Data is so unbelievably complicated, and that is what gets nines. Like, you have no idea how much data can shift everything. Data is so important.TypeSafe as a Data Lab and Synthetic Data StrategySwyx [00:21:57]: Yeah.Diogo Almeida [00:21:57]: Holy crap. So if people are looking for a job, we are hiring infinite data people, actually infinite.Swyx [00:22:04]: What is a good data person? Like, clearly somebody who cares about reading through the transcripts of, whatever. You've said, for example, that y- all your data is synthetic.Diogo Almeida [00:22:15]: Yep.Swyx [00:22:15]: But that's only, like, the scratching the surface, right?Diogo Almeida [00:22:18]: Yeah.Swyx [00:22:19]: Like, it's not. Like, synthetic, so what, right? Synthetic, but we have people with a lot of taste and a lot of care looking at, looking at these, articulating what's wrong, going back, regenerating. Is that what a good data person is these days?Diogo Almeida [00:22:31]: Let me try to figure out how to. Like, it's, it's super complicated, and like, I literally onboard the data people with a Talk that I assume is longer than this podcast will end up being. So I will try to say, like, the high level of it. So number one, we don't do the kind of synthetic data that people ki. Well, I'll do. Actually, number is zero. Data and synthetic data depends on your task. Like, the shape of your data. The shape of your task changes the data. Like, RLVR's data is kind of environments, right?Swyx [00:23:03]: Yes.Diogo Almeida [00:23:04]: RLHF's is the human feedback? Each task has its own unique kind of data, and we, of course, have our own unique kind of data, right? So number one, we have that. Number two, the thing I. The reason why we don't want to train on our users' data, even if we could, right? Like, we could probably ask for any terms right now, and it will. We. I don't know if it would make a difference. We truly don't want that, because no matter what, the real-world data has so much bias. There's, like, a power law of, like, people, like, asking the same things where you'll end up, like, overfitting to it and like, fracturing to it and all of that. And number two, we are, like, aiming for, like, a complete sci-fi future years from now where, like, these models are going to be, like, the general infrastructure, layers and layers and layers and deep down the stack to, like, things people can't even imagine. Like, I would like to think of our model, like, kind of like, UDP as LLMs and TCP as our models. All sorts of stuff can be built on top of that, and we need to be able to nail those futuristic use cases such that software developers can actually build that futuristic stuff. And the way to do that is even if we had all of the data of the present, we would just overfit to the present, and then it wouldn't work. What we need is to, like.Diogo Almeida [00:24:21]: It almost feels like a. Like, they're the artists? They study this cognitive core. Our cognitive core is, like, way less jagged than anyone else's. And then they find the jaggednesses, and then they address them surgically in a way that. And you can never perfectly do this, right? But they do it in such a way that it addresses it in every single possible, like, dimension, past, present, future.Swyx [00:24:45]: The general case rather than the specific case.Diogo Almeida [00:24:47]: Exactly. And like, that requires a lot of intelligence every time.RLCD vs. RLHF: Defining a New TaskSwyx [00:24:50]: Okay, so we mentioned a little bit. You sort of criticized my thinking as r-- like, very RLVR influence, which is, like, very fair. Let us actually mention RLCDDiogo Almeida [00:24:59]: OohSwyx [00:24:59]: Which obviously you have some secret sauces to our knowledge. You've never actually published a paper or anything like that on it. No, right?Diogo Almeida [00:25:05]: No, not yet.Swyx [00:25:06]: But like, what should people get from this? Like, what. Can you give people some confidence that you're just not just making up jargon for the sake of sounding cool, right? Like, one thing for me is, like, calibration I do think is a. To me, like, well understood because we've covered it in. On the podcast.Diogo Almeida [00:25:22]: Yeah.Swyx [00:25:22]: But I don't know what you mean when you say RLCD versus what people are familiar with.Diogo Almeida [00:25:26]: It's a great question.Swyx [00:25:27]: Yes.Diogo Almeida [00:25:27]: And actually, I will give a related question.Swyx [00:25:29]: Okay.Diogo Almeida [00:25:29]: What is RLHF?Swyx [00:25:31]: Okay.Diogo Almeida [00:25:31]: Right? And actually, RLHF means multiple different things, right?Swyx [00:25:34]: Okay.Diogo Almeida [00:25:34]: Like, there's the RLHF of the original. I think it was, like, Paul Christiano teaching a robot to backflip or something like that. Wasn't there somethingSwyx [00:25:42]: Was that it?Diogo Almeida [00:25:43]: That was the originalSwyx [00:25:44]: I referenced the PPO paper, but I don't know.Diogo Almeida [00:25:46]: And so PPO was not necessarily from human feedback, if I recall.Swyx [00:25:51]: Okay. That's trueDiogo Almeida [00:25:52]: But I b- I believe it was, like, an OpenAI alignment work that could teach hard to specify outputs, like a backflip. I'm not 100% sure. And then there was actually learning to summarize. This was work, by a bunch of the team that helped with, instruct-- and co-authored, the instruction following paper, which was teaching, doing PPO on language models.Swyx [00:26:15]: This is the, sorry. I'm trying to, tryingDiogo Almeida [00:26:19]: YeahSwyx [00:26:19]: Trying to manipulate this thing. This is 2017.Diogo Almeida [00:26:23]: Yeah.Swyx [00:26:23]: Right.Diogo Almeida [00:26:23]: I'm not 100% sure, but like, that looks quite right.Swyx [00:26:26]: Yeah.Diogo Almeida [00:26:26]: If it has, like, a robot doing backflips or something like that might be it. Yes. Okay, cool. I guess I got it right. Hell yeah.Swyx [00:26:35]: There you go.Diogo Almeida [00:26:36]: Yeah.Swyx [00:26:36]: That's the one.Diogo Almeida [00:26:36]: So the idea was can, like, can you do, like, ill-specified things with it? So that's, like, version one. Version two was, the learning to summarize work, that, like, OpenAI did, which is actually, like, PPO on language models to do something somewhat ill-specified. This is, like, another thing that people refer to as RLHF Which I did not co-author.Diogo Almeida [00:26:57]: Oh, Dario's there. Cool. Hell yeah.Swyx [00:27:01]: And Radford.Diogo Almeida [00:27:02]: Yeah. Shout-outs to Alec and Ryan. Love them.Swyx [00:27:04]: Yeah.Diogo Almeida [00:27:05]: But the thing that I refer to RLHF is the, Oh, man.Diogo Almeida [00:27:13]: I'll get toSwyx [00:27:14]: You have comments on that, yeah.Diogo Almeida [00:27:15]: I have comments on that paper, but like, we're so many, tangents deep.Swyx [00:27:18]: Yeah.Diogo Almeida [00:27:18]: So the thing that really got. To me, the thing that I'm calling to RLHF is the task of instruction following. It's not about the PPO. That part doesn't matter. It's about, like, setting a North Star of this is a valuable direction. It's kind of like the Bitris lesson North Star.Diogo Almeida [00:27:34]: And for us, RLCD is this new task. And it is not. I don't see it as jargon. Like, I try to communicate with precision. It's just that, “Hey, here's another North Star.” Just like DPO and all of its, like, descendants also do RLHF, despite not using the algorithm in that paper.Swyx [00:27:55]: And so clear- clearly stating the North Star is, being program- programmable AI is one, word that I really catch onto, removing the human in the loop,Diogo Almeida [00:28:06]: YesSwyx [00:28:06]: From. Because RLHF is tuningDiogo Almeida [00:28:09]: YesSwyx [00:28:09]: For this so that you can automate everything.Diogo Almeida [00:28:11]: Yes. Everything that makesSwyx [00:28:13]: Did I miss anything else in the, in the thesis of, like, what the North Star is?Diogo Almeida [00:28:17]: There is. That is. That is right. I'm overly nuanced in my communication. The one nuance is that we need to be practical. We need to be aware of what language models can do really well. Like what AI can do.Diogo Almeida [00:28:30]: Right? Like, there could be programmatic types that are, like, sick AF, but if you. If the technology is not ready for it to. It's not a tragedy if that's not out in the world.Why Programmable AI MattersSwyx [00:28:41]: Yeah.Diogo Almeida [00:28:42]: But to me, like, the pre-Jev world was a tragedy becau-- it sounds arrogant. Hear me out.Swyx [00:28:49]: No. I strongly believe you.Diogo Almeida [00:28:50]: Cool. It sounds arrogant, but like, I felt this way since long before I even had a company.Swyx [00:28:54]: Yeah. I can, I can vouch that,Diogo Almeida [00:28:56]: Yes, I've been talking about this for so longSwyx [00:28:57]: You said this at All Around Her for, like, three years.Diogo Almeida [00:28:58]: Yeah, I've been talking about this for so long. And I've been saying it because I thought it would have been easier. They say they do not do things because they. It. They're easy. They. It's ‘cause they thought it was easy, soSwyx [00:29:08]: Yeah, exactlyDiogo Almeida [00:29:09]: Something like that. I thought it. This whole project would take a week.Diogo Almeida [00:29:13]: And I was unbelievably wrong. So I am so sorry to everyone at OpenAI that I thought. I was like, “Man, I'm solving this right now.” but like, I think that the tragic thing is when. Well, I think overpromise, underdeliver is tragic too. And like, AI is super extreme on that axis. And I think RLVR is, like, the main. Well, both RLVR and RLHF are extreme perpetrators of this.Diogo Almeida [00:29:40]: But like, it. To me, it's like it's just there's just so much potential there. Like, AI is clearly so smart. I l- smart. I love this in my talks, when I ask people, like, “How can AI be so unbelievably smart? How can we, like, solve millennium prize problems in math, but still not automate even the most basics of works?” Like, really basic rote stuff that, like, the. It d- it doesn't take, like, extremely smart people to do this. It's not a satisfying job. Like, there's other things these people could be doing, but yet we need them to do, like, this ba- like, super basic- non- unsatisfying stuff because, like, we can't automate it yet, but we have this, like, supercharged engine of automation that just does not have, like, the right plugs and stuff to plug into all of this economically valuable work. And like, if the whole company of TypeSafe disappears, like, maybe it'll take, like, a year or two for people to, like, truly catch up. I actually don't know how long it'll take. If model quality matters, then we are gonna be in a very good position for a long time. But it, like, it's done, right? Like, there, like, this has changed the path of, like, technological history.Swyx [00:30:49]: Yeah.Diogo Almeida [00:30:49]: And like, we will be exploring that space as a field.Swyx [00:30:53]: Yeah. I think, I definitely agree with that. You've created possibilities. So I think, if I can paraphrase so that people can un- also understand, you should not take the success of TypeSafe and Jev as just like, “Well, that is a new model type. Now we're done. We go back to business.” Like, no. Like, actually, there's, there are, like, five other model types that you should be exploring and like, let a thousand flowers bloom.Diogo Almeida [00:31:15]: Absolutely.Swyx [00:31:16]: Right?Diogo Almeida [00:31:16]: Like, early internetSwyx [00:31:17]: And some of that, some of which you will probably also build.Diogo Almeida [00:31:18]: Of course, yes.Swyx [00:31:19]: Yes.Diogo Almeida [00:31:19]: Early internet energy. I think it's back to tech utopia. It's no longer like, “Oh, man, like, sometimes my coding agents work, but the, all of the best ones are hoarded internally.”Swyx [00:31:29]: Yeah.Diogo Almeida [00:31:30]: Right? It's like creation is back on the menu.Diogo Almeida [00:31:34]: ? Though it's gonna be a wild-ass world, and buckle up.Diogo Almeida [00:31:38]: It's. And I'm so jazzed about that.Manifesto, Launch Strategy, and Early Internet EnergySwyx [00:31:42]: Yeah. And now you have the funding and the momentum to do whatever you envision there, which I, which I think is, like, very gratifying to see you have after, so long of saying these thingsDiogo Almeida [00:31:53]: YeahSwyx [00:31:54]: But actually show the world.Diogo Almeida [00:31:55]: I know. I just. Such a, such an interesting thing to be a tease the whole time. Like, my talk, like, felt like it was a cliffhanger ‘cause I didn't say how the automation would occur.Swyx [00:32:05]: Yeah.Diogo Almeida [00:32:06]: Sean reviewed our manifesto And he's like, “It's a little bit vague in these parts.”Diogo Almeida [00:32:12]: And like, “What's step one? What is, what is the intelligence model?”Swyx [00:32:16]: Well, I asked you for model, and you were like, “Yeah, model coming.”Diogo Almeida [00:32:18]: Yeah.Swyx [00:32:18]: And like, Well, I just, I mainly objected to the word composable But build prod.god is fantastic.Diogo Almeida [00:32:24]: Thank you.Swyx [00:32:24]: Yeah.Diogo Almeida [00:32:25]: I. We've really rallied around that. I'd like to think we're not entirely a cult like some companies are.Diogo Almeida [00:32:32]: But like, we are, like, jazzed about what we're doing, and like, we are. Like, my brand is being practical, and like, we are all, like, so super-duper practical.Swyx [00:32:42]: Yeah.Diogo Almeida [00:32:42]: It's really great.Swyx [00:32:43]: Yeah. So here. And by the way, here is the step, the secret master plan, right?Diogo Almeida [00:32:47]: Yep.Swyx [00:32:47]: Shape, the shape of machine-native composable AI.Diogo Almeida [00:32:49]: It was your idea to make a secret master plan, soSwyx [00:32:51]: It's a, it's that Elon thing. When he started TeslaDiogo Almeida [00:32:53]: YeahSwyx [00:32:53]: He was like, “Here's what we'll do.”Diogo Almeida [00:32:54]: But I did. Yeah. I'm giving official credit to you.Swyx [00:32:56]: Oh, thank you. Thank you, thank you.Diogo Almeida [00:32:56]: Yeah.Swyx [00:32:56]: Thank you. But like, you should've told me your, you're also gonna do this model launch, ‘cause you, like, you told me, you told me half of the story, and then the other half, you didn't have the doom demo at the time.Diogo Almeida [00:33:08]: Yep.Swyx [00:33:08]: You didn't have any numbers to give me.Diogo Almeida [00:33:10]: Yep.Swyx [00:33:10]: I was like, “what?”Diogo Almeida [00:33:11]: Well, the problem is I don't believe in benchmarking.Swyx [00:33:13]: Exactly.Diogo Almeida [00:33:14]: Right?Swyx [00:33:14]: Exactly.Diogo Almeida [00:33:14]: So like, it is a thing that you need to feel, and like, I think that this is the way to build long-term trust, even though it, like, hurt, it hurt us a, us a lot? Like last year when we did fundraise, no one believed us.Diogo Almeida [00:33:27]: ? Like, and they wanted just benchmarks and stuff, and we're like, “We're not gonna do that. We are principled. We're gonna stand by our guns. That rewards bad actors. I don't give a s**t, like, what you want. Like, this is who we are, and we are standing by that.” So Sorry. It's notSwyx [00:33:43]: No, yeah. Well, and in some ways, I think, like, choosing the hard path, it. But you end up making the company that you wanna work in.Diogo Almeida [00:33:49]: Yep.Swyx [00:33:50]: Right? Otherwise, if you sell out, then you're just working in, like, OpenAI but with my people, right? Which is like.Diogo Almeida [00:33:56]: Yeah. Yeah. Like, I'm, I don't have too many regrets on that, obviously.Swyx [00:34:01]: Yeah.Diogo Almeida [00:34:01]: Like, it worked out so unbelievably well. And like, I, The. I was emotional last night when I was talking about, like, the reasons I left OpenAI, and because, like, it actually had to change my wording after the launch. My phrasing was, “If an AI winter did happen and I did not do every f*****g possible thing I could to, like, avert that, I would see myself as personally responsible both for, the RLHF direction, which I think really widened overpromise versus under-deliver, and also not going all in on this because I think this is, this is where value is going to just be, like, printed.” So. And it was really cool because I feel likeDiogo Almeida [00:34:47]: The AI winter I'm worrying about is averted. Like, AI will be useful. It'll be used for automation.Diogo Almeida [00:34:53]: It's been less than a week, and like, the numbers are already undeniableSwyx [00:34:57]: YeahDiogo Almeida [00:34:57]: That it's, like, being used for real work, and like, there's. It's, it's the Wild West. Yeah.Launch Traction, Tokens, Rate Limits, and Developer UsageSwyx [00:35:03]: Yeah. Can you sh- just if you have top of your head, what numbers are you seeing? Like, what's, what's, like, signups? Like, whatever you can share.Diogo Almeida [00:35:11]: I'm actually not super on top of everything. Like, the team is the ones who are telling me all of these things.Swyx [00:35:16]: Yeah, and I'm sure it's, like, changing every day, right?Diogo Almeida [00:35:17]: It's, it's,Swyx [00:35:18]: But likeDiogo Almeida [00:35:18]: It's kinda nutsSwyx [00:35:19]: If there's a milestone that you're like, “Well, yep, that's one thing we were hoping for. We reached it.”Diogo Almeida [00:35:23]: I will say a milestone that we've passed is tokens per day.Swyx [00:35:27]: Nice.Diogo Almeida [00:35:27]: And this is not, like, fleeting tokens per day.Swyx [00:35:32]: Yeah.Diogo Almeida [00:35:32]: This is, like, even at night, like, it's constantly training, so machines are calling it and not just people trying things out.Diogo Almeida [00:35:39]: So that is, That is so cool. A trillion tokens a day is a lot.Swyx [00:35:45]: Yeah.Diogo Almeida [00:35:45]: So surpassing that is awesome. Signups to me don't really matter. And actually, this was, like, a bit of a mistake we made, if I'm, like, totally honest. People on Twitter were calling us, like, marketing geniuses and all of that, and that was just us. We don't have a marketer. Also hiring. And we were just being our genuine, goofy, like, irreverent selves, and we were, we were just, like, offboarding people off the waitlist so hard. - Our platform team is so unbelievably cracked. I think we have more n- up nines of uptime than Anthropic while having the most Unprecedented launch ever. Like, that is kind of nuts, soSwyx [00:36:21]: YeahDiogo Almeida [00:36:21]: Like, props to them.Swyx [00:36:22]: Yeah.Diogo Almeida [00:36:23]: And the thing we didn't realize. So number one, waitlists, waitlist sign-ups don't matter for, like, a developer platform, in my opinion? I would guess that a large number of them are not even developers. So they go in, they try some queries, and a lot of people don't get it because they are not programming, right? Like, they're just like, “What? This is not a chatbot. Where's my ChatGPT 2?”Diogo Almeida [00:36:45]: Right? But if, like. I haven't exactly calculated this. My sense is that if every single human being in the world, like, just wrote a couple of queries, that would be a rounding error compared to, like, one power user's for loop that is just, like, creating value.Swyx [00:37:01]: Yeah.Diogo Almeida [00:37:01]: And the thing we are-- didn't realize with the waitlist is, like, we could just w- off-board anyone off the waitlist. It doesn't matter. The scary part is rate limits. And then once people start getting value from that, then they just want tons and tons of rate limits because this is what software is, right? Like, you spend effort upfront to specify your rote task, and then this rote task creates more value than it takes to put in. And then now that you have thatSwyx [00:37:25]: Set it and forget, yeah.Diogo Almeida [00:37:26]: Exactly, yeah. You run it in the background. You make it a dependency, to, like, other things. You can make, like, higher level stuff. And like, you just create so much value in the world. Early internet people probably did not imagine, like, the wonder of early 2000s internet, which is still not early internet. But like, it's, it's through, no offense, composabilitySwyx [00:37:47]: NoDiogo Almeida [00:37:47]: That all of the crazy stuff happens, and I just really wanted to emphasize that in our manifesto. We are going for emergence. We are going for, like, being the catalyst. We're wanting to empower people, and we are going to do whatever we can for that, be it, like, Discords in our town hall with me wearing a garbage bag or not.Swyx [00:38:05]: And podcasts and Diogo Almeida [00:38:08]: Hell yeahSwyx [00:38:09]: Getting all that.Diogo Almeida [00:38:09]: Absolutely.Swyx [00:38:09]: Like, ‘cause I want the long form, right?Diogo Almeida [00:38:11]: Yeah.Swyx [00:38:12]: It is like, yes, we'll get past the, some of the superficial things, and then we'll go deep andDiogo Almeida [00:38:15]: Hell yeahSwyx [00:38:15]: And people will really trust and understand your mission and like, the people that, will resonate that will end up joining you or, buying you. Or No, but sorry, as a, as a customer.Diogo Almeida [00:38:27]: Oh, as a customer.Swyx [00:38:28]: As a customer, as a customer.Diogo Almeida [00:38:28]: Okay, yeah. That was funny. I'm sorry.Swyx [00:38:30]: Sorry. I didn't, I didn't mean to say that. But no, any-- one version, one very flattering version of this, like, 36 million views of your launch video.Diogo Almeida [00:38:37]: Cool. Up to 38 now.Swyx [00:38:39]: Yeah, rounding error.Diogo Almeida [00:38:40]: Yeah.Swyx [00:38:40]: Navio still has got 74. Fable 5 got 57. So like, as far as, a- and I didn't, I didn't do the stats for, like, original ChatGPT, likeDiogo Almeida [00:38:48]: YepSwyx [00:38:49]: Which there was no video.Diogo Almeida [00:38:50]: Yep.Swyx [00:38:50]: So like, up there, right?Diogo Almeida [00:38:52]: Yep.Swyx [00:38:52]: Like, as far, as far as, like, if you were to launch a Neolab in 2026, I think you're, like, number one right now, which is, like, pretty crazy.Diogo Almeida [00:38:58]: Yeah. Well, I actually would rather. I do have the shirt, like, your favorites Neola-- favorite Neolab's favorite Neolab.Swyx [00:39:05]: Huh.Diogo Almeida [00:39:05]: I don't give a s**t about being a Neolab. I think being a Neolab. Actually, we have a lot of, like, swag that's being a parody of a Neolab. One of them, one of them I have is, like, Neolab with product, which actually is not a Neolab. Like, I don't care about that, really.Swyx [00:39:20]: Yeah.Diogo Almeida [00:39:20]: What I care about is being a reliable dev platform. So Swyx [00:39:23]: YesDiogo Almeida [00:39:24]: Appreciate the comparison, but likeSwyx [00:39:25]: YeahDiogo Almeida [00:39:25]: Hopefully we transcend past them and we go back into, like, a thing-- like, a revolutionary moment for developers and like, this stable thing that people can rely on and trust.Reliability, Robustness, and DeterminismSwyx [00:39:35]: Yes. To that end, I think that's one thing that really impressed me about you guys is that, yes, you do talk about reliability. I thought it was mostly about calibration, which, like, we talk about RLCD. But actually it's also about just, like, uptime and scalability and all those things, right? They're, they're all sort of the kind.Diogo Almeida [00:39:55]: And nines.Swyx [00:39:56]: And nines.Diogo Almeida [00:39:56]: It's, likeSwyx [00:39:57]: Which uptime is, in my opinion.Diogo Almeida [00:39:58]: Oh, but that's part of it. But like, there's reliability in, like, how intelligent the thing is. Like, how consistently does it do the thing that you want? And I think that, like, the big reasoning models are very smart. In my opinion, they still lack reliability. I think there's many use cases where you-- they look like they should be smart enough to automate their work. There is economic incentive to automate that work, yet still they're not reliable enough as, at an intern because they're optimized for different things. And so like, I think that there's the reliability of being able to, like, trust the outputs. And also we are. Like, there are dimensions of reliability that we are not yet at that I'm, like, so excited by.Swyx [00:40:38]: Yeah.Diogo Almeida [00:40:38]: Like, I want to automate the easy work before the hard work? Like, I think that's just a common sense thing to do. But to me, we will be sufficient. I don't know if there's such thing as sufficiently reliable, but I wanna get so good that people don't even need to try the model to know that it'll work. It's like, that's like what flow state is in programming, right? Like, I'm just, like, writing queries because I need intelligence in here. And like, when. For non-trivial branching, I can just write it in like a, like a type-safe System 1 query and then get the results out of it and it just branches accurately. Like, that would be so good. Like, that's the. That is the dream.Swyx [00:41:12]: Yeah.Diogo Almeida [00:41:12]: And that is, like, going to be, like, a long slog.Swyx [00:41:16]: Yeah. We're gonna go into your API design in a little bitDiogo Almeida [00:41:19]: OohSwyx [00:41:19]: Just to give people examples and like, maybe paths not taken, that kind of stuff.Swyx [00:41:23]: One thing up the front that I do wonder about in terms of reliability is I noticed that there's no seed. There's no, And so basically, same input, do I always get the same output?Diogo Almeida [00:41:34]: SoSwyx [00:41:36]: And if not, why not?Diogo Almeida [00:41:37]: Oh, great question. So this is actually, like, a common question we have between. So reliability is actually a catchall. Like, whenever AI can't automate something, it's due to some form of reliability. Could be, like, type safety. It could be determinism. It just could be, like, it's, it's jagged, right? So reliability is a catchall. I just think that it's also a catchall for, like, what the North Star is. Re- determinism is, like, same inputs, same outputs. I do believe that this is, like, slightly interesting for unit tests, but I believe that to be the wrong North Star. I believe robustness is what peopleDiogo Almeida [00:42:16]: I don't wanna tell people what they really want, ‘cause that would be a little arrogant of me.Diogo Almeida [00:42:19]: I believe that is, like, the more important property. You want, given similar inputs, get similar outputs. And it's kind of wild how unreliable LLMs are.Diogo Almeida [00:42:31]: Like, a way that we test this is you put, like, UUIDs in, like littleSwyx [00:42:36]: YeahDiogo Almeida [00:42:36]: I think they're called nonces In the prompt. And what you want is similar outputs from all of those, ‘cause it's truly semantically the same question, and that is the part where you really want. Th- like, that robustness is where, like, people get, like, burnt with AI making decisions. So I think that is the. A super-duper important property. We could also have determinism. That is, that is a thing that can be available. As far as I can, like, mentally model for programmers, like, it, I- it could be valuable for some use cases, so like, please educate me, in comments or view. But my. In general, it's easy. Determinism is something you can, like, trade off for better cost. Like, we are, we are constantly wanting to be on the intelligence per dollar frontier. We are doing, like, absolutely disgusting things to be there. Like, this is,Diogo Almeida [00:43:32]: I shouldn't say this, but no one's here to stop me.Swyx [00:43:37]: If you s- you sign off on your own PR.Diogo Almeida [00:43:40]: That is not how it works at this company. I believe for this week, my chief of staff, Kay, is the most powerful person in tech.Swyx [00:43:49]: Yeah. And shout-out to Kay for organizing this.Diogo Almeida [00:43:50]: Holy shSwyx [00:43:51]: Yeah.Diogo Almeida [00:43:51]: Holy s**t. She is so f*****g competent and powerful. She's incredible.Diogo Almeida [00:43:58]: She sucks. Don't poach her. But so I try to be a bit more filtered, but like, people are telling me, “Don't call it a Frankenstein's monster of models,” but because that has, like, negative implications. I think Frankenstein's monster was, like, the good guy in this whole. It was innocent, right? I didn't read it. Okay.Diogo Almeida [00:44:18]: I'll, I'll confess. Okay. That. Well, one facial expression, ISwyx [00:44:21]: This is aDiogo Almeida [00:44:21]: My cards on the tableSwyx [00:44:21]: Decent Jacob Elordi movie if you wanna seeDiogo Almeida [00:44:24]: ISwyx [00:44:25]: The adaptation. Anyway.Diogo Almeida [00:44:26]: The. You have no idea how little time I have right now.Swyx [00:44:28]: Yeah.Diogo Almeida [00:44:29]: My priorities are sleep?Swyx [00:44:31]: Developers.Diogo Almeida [00:44:32]: Developers, yes. Developers. But yes. It. We do, like, absolutely disgusting things to be on the Pareto curve of intelligence per dollar, and we are going to keep doing that.Swyx [00:44:47]: Yeah.Diogo Almeida [00:44:47]: We're gonna be doing crazy-ass stuff, and I think people really need to think outside of the box. Like, part of the reason we're surprising is, like, people Are thought inside the box, and we continue to do that. As of right now, we are obviously the best at this, and we want to continue being the best at that whole thing.Swyx [00:45:05]: Yeah.Diogo Almeida [00:45:05]: So Wait, where did, where did we tangent from?Swyx [00:45:07]: No. SoDiogo Almeida [00:45:08]: YeahSwyx [00:45:08]: I asked you about, will you have seeds and determinism?Diogo Almeida [00:45:11]: Oh, yes. SoSwyx [00:45:11]: And then you basically defined reliability and likeDiogo Almeida [00:45:14]: And robustnessSwyx [00:45:15]: How you see it. Yes.Diogo Almeida [00:45:16]: But like, determina- likeSwyx [00:45:17]: I have a robustness example that's, that's, real quick I can show you.Diogo Almeida [00:45:19]: I would love that. I will just say one thing.Swyx [00:45:21]: Yeah.Diogo Almeida [00:45:21]: We can make a deterministic model.Swyx [00:45:22]: Exactly.Diogo Almeida [00:45:23]: Like, we're hap- if people can convince us that is a valuable thing to doSwyx [00:45:27]: YeahDiogo Almeida [00:45:27]: And we don't have a gigantic GPU shortageSwyx [00:45:29]: YeahDiogo Almeida [00:45:29]: We can happily make all of these models. We live to please. And rev- and revolt, revolute,Swyx [00:45:38]: You will throw over everything, except you'll do it in a nice way.Diogo Almeida [00:45:41]: Yeah.Swyx [00:45:41]: And findDiogo Almeida [00:45:42]: So like, determinism could be on the cards.Swyx [00:45:44]: Yeah.Diogo Almeida [00:45:45]: It just gets you less intelligence per dollar.Swyx [00:45:46]: Yeah. Well, just having seen the trajectory of OpenAI and Anthropic, you will. Just trust me now that you will be peer pressured into doing it. So like, just people will want it even if they. If you tell them they don't need it. They'll still want it. So like, yeah, that's the TL;DR of that.Diogo Almeida [00:46:01]: Okay.Swyx [00:46:02]: Yeah.Diogo Almeida [00:46:02]: I will love to. Maybe one day we will see how that happens.Swyx [00:46:07]: Yeah.Diogo Almeida [00:46:07]: I've been told I'm, They say that part of our brand is being unshakeableSwyx [00:46:13]: HuhDiogo Almeida [00:46:13]: And they say that's just the nice way of saying stubborn.Swyx [00:46:15]:
Walter Maffione, CEO and co-founder of KaleidoSwap, join me to explain how the project grew from an RGB Lightning DEX into a broader swap engine and LSP stack — Lightning as rails between Liquid, Arkade, Taproot Assets, Spark, and RGB — plus KaleidoSDK, desktop app, and browser Extension.We dig into web app and SDK integrations for merchants and unified-balance wallets, how HTLC atomicity works across layers, BOLT 12 multi-asset offers and Nostr discovery for competing providers, Bitcoin-only scope with Flashnet and Utexo bridges for external stables, and a product lineup of desktop app, browser extension, and dual SDKs.Walter also covers agentic payments via MCP plugins, self-sovereign local models with Tether's QVAC, fee ranges around 0.5–1%, and a one-to-three-month mainnet launch path for the extension, swap provider, and web app.Timestamp:00:00 — Intro: Walter & KaleidoSwap00:50 — From RGB to Multi-Layer Swaps02:14 — LSP Plus Swap Provider Stack03:42 — Merchants, Stables, Unified Balance05:38 — Why Swap UX Took Over Wallets06:42 — Liquidity Ops and Mainnet Path07:58 — AI Attacks After Boltz Shutdown10:23 — Defending Non-Custodial Swaps12:18 — How Atomic Swaps Actually Work13:36 — RGB, Liquid, Arkade, Spark 14:36 — Open Spec, BOLT 12, Nostr Discovery18:23 — Pay Anything via the SDK19:34 — Bitcoin Layers Only + Bridges22:05 — Desktop, Extension, Dual SDKs25:05 — UX That Adapts to the User26:32 — Agentic Payments and QVAC33:22 — MCP Plugins, Keep Your Own AI35:11 — Fees Around 0.5–1%35:57 — Launch Timeline: Next 1–3 Months37:05 — Find KaleidoSwap OnlineLinks: https://x.com/bit_walthttps://x.com/kaleidoswaphttps://docs.kaleidoswap.com/whats-kaleidoswap/introductionStephan Livera links:Follow me on X: @stephanliveraSubscribe to the podcastSubscribe to Substack
Hoy te traigo un episodio que llevaba tiempo queriendo grabar. En el episodio 829 te hablé sobre skills, pero me quedé con la sensación de haberme enrollado sin mostrar nada concreto. Así que le he dado la vuelta a la tortilla y en esta ocasión te traigo siete MCPs que he seleccionado uno a uno para que veas de qué va esto del Model Context Protocol y, sobre todo, para que compruebes cómo transforman lo que puede hacer tu inteligencia artificial.Porque hay que decirlo claro: la IA es tonta. Literalmente. No sabe qué hora es, no sabe leer un archivo, no sabe si tienes issues abiertos en GitHub, no sabe cómo te llamas. Cada vez que empiezas una conversación con cualquier modelo de lenguaje, empiezas de cero. Y aquí es donde entran los MCPs, que no son ni más ni menos que herramientas: manos, ojos y oídos para que la IA pueda hacer cosas útiles de verdad.El Model Context Protocol es un estándar abierto creado por Anthropic que funciona como un *USB-C para la IA*. Un conector universal que permite que cualquier aplicación de IA se conecte con cualquier fuente de datos o herramienta externa. Y lo mejor es que no es propietario, al contrario que los plugins de ChatGPT. Cualquiera puede crear un MCP, y de hecho te cuento cómo hacerlo con Python o Rust.En este episodio repaso siete MCPs que uso en mi día a día. Empiezo por el Filesystem, que permite a la IA leer, escribir, buscar y editar archivos en tu sistema, con control de acceso para que no haga tonterías. Sigo con el SQLite, ideal para consultar bases de datos locales de agenda, finanzas o cualquier aplicación que uses. Después llega el GitHub, el servidor oficial de GitHub escrito en Go, que te permite gestionar issues, pull requests, repositorios y mucho más sin salir del chat.También te hablo del Web Fetch, que permite a la IA leer páginas web y convertirlas a Markdown ahorrando una barbaridad de tokens. Y del Time, que es una tontería pero imprescindible: la IA puede saber la hora en cualquier huso horario del mundo. Luego está el Memory, que construye un grafo de conocimiento persistente para que la IA recuerde quién eres, qué prefieres y qué proyectos tienes entre manos, incluso entre sesiones. Y por último, el Sequential Thinking, una herramienta de razonamiento paso a paso que obliga a la IA a pensar de forma estructurada cuando se enfrenta a problemas complejos.También hablo de cómo configurar estos MCPs en OpenCode, Claude Desktop y VS Code, y cuánto consumen de RAM y tokens. Porque no es lo mismo activar veinte servidores que tener solo los que necesitas. Y para rematar, comparo MCPs con skills: las skills le dicen a la IA cómo comportarse, los MCPs le dan capacidades para actuar.Y si te pica la curiosidad por crear tu propio MCP, también te cuento las diferencias entre usar el SDK de Python (súper rápido de prototipar) y el de Rust (más rendimiento, menos consumo de RAM).Capítulos del episodio:0:00 - Introducción: la IA necesita herramientas para ser útil2:30 - ¿Qué es MCP? Arquitectura host-cliente-servidor y JSON RPC 2.05:30 - Cinco razones para adoptar MCPs8:00 - MCP Filesystem: leer, buscar y crear archivos11:00 - MCP SQLite: consultar bases de datos y el no-determinismo14:00 - MCP GitHub: issues, pull requests y repositorios17:00 - MCP Web Fetch: búsquedas en internet con ahorro de tokens19:30 - MCP Time: la hora en cualquier huso horario21:30 - MCP Memory: grafo de conocimiento persistente23:30 - MCP Sequential Thinking: razonamiento paso a paso25:30 - Configuración en OpenCode, consumo de tokens y RAM27:30 - Crear tu propio MCP: Python SDK vs Rust28:30 - Cierre y despedidaMás información y enlaces en las notas del episodio
Côté IA : MCP devient stateless, Claude watermarke ses textes, GPT-6 Astra défie Claude Fable, et une étude JetBrains confirme Claude Code en tête des agents de code. Côté JVM : JDK 27 généralise G1, Kotlin 2.4 stabilise les context parameters, une API JSON arrive dans le JDK, et Quarkus comme Micronaut enchaînent les versions. En bonus, trois pannes IA simultanées et un câble débranché chez Google Cloud. Enregistré le 11 septembre 2026 Téléchargement de l'épisode LesCastCodeurs-Episode-343.mp3 ou en vidéo sur YouTube. News Langages Les Types Algébriques de Données (ADTs) en Java rockthejvm.com/articles/algebraic-data-types-in-java Le problème : L'approche classique (champs nullables, hiérarchies de classes ouvertes) crée des états invalides et des erreurs à l'exécution (comme le NullPointerException). Types Produits (ET logique) : Implémentés en Java avec les Records. Ils regroupent plusieurs champs de manière immuable et concise. Types Sommes (OU logique) : Implémentés avec les Sealed Interfaces. Elles définissent un ensemble strictement fermé de sous-types connus à la compilation. ADTs (Types Algébriques) : La combinaison des Sealed Interfaces et des Records. Ils garantissent que les états invalides sont impossibles à représenter dans le code. Pattern Matching : L'extraction des données se fait via des expressions switch exhaustives, supprimant le besoin de casts manuels et obligeant le développeur à traiter tous les cas possibles. Généralisation : Ce modèle est idéal pour créer des types comme Result, forçant le traitement explicite et sécurisé des succès et des erreurs typées. Pourquoi "Algébrique" ? Parce que les types sont combinés mathématiquement (Produits = multiplication, Sommes = addition) pour limiter strictement le nombre d'états possibles d'une donnée. JDK 27 : fonctionnalités et calendrier de sortie openjdk.org/projects/jdk/27 infoworld.com/article/4202901/jdk-27-the-new-features-of-java-27.html JDK 27 est la prochaine version majeure de Java, une version non-LTS avec seulement 6 mois de support, qui succède à JDK 26. La disponibilité générale est prévue pour le 15 septembre 2026, avec des release candidates les 6 et 20 août 2026. Le périmètre est désormais figé (feature freeze) avec neuf JEP au programme. Le ramasse-miettes G1 devient le collecteur par défaut dans tous les environnements, et plus seulement en mode serveur. Ajout d'un support de la cryptographie post-quantique pour TLS 1.3, via des échanges de clés hybrides combinant algorithmes classiques et résistants au quantique. Finalisation de l'API PEM pour encoder et décoder clés, certificats et listes de révocation au format PEM. L'API Vector poursuit son incubation pour la douzième fois, permettant d'exprimer des calculs vectoriels compilés en instructions CPU optimisées. Les en-têtes d'objets compacts, introduits en JDK 24, sont désormais activés par défaut et réduisent l'empreinte mémoire du tas. Plusieurs previews sont reconduites : constantes paresseuses (3e preview), types primitifs dans les patterns (5e preview) et concurrence structurée (7e preview). Ajout d'une fonctionnalité de rédaction in-process pour JFR, afin de masquer les données sensibles dans les enregistrements de profiling. Kotlin 2.4 : nouveautés du langage et outillage kotlinlang.org/docs/whatsnew24.html Kotlin est un langage moderne, multiplateforme (JVM, Android, iOS, JavaScript, Wasm) développé par JetBrains, souvent utilisé comme alternative à Java. Les context parameters passent en stable : ils permettent de fournir des dépendances implicites à une fonction sans les déclarer en paramètre explicite, un peu comme une injection de dépendances. Les collection literals arrivent en expérimental : on peut écrire une liste avec des crochets, comme en Python, par exemple val fruits = ["pomme", "banane"]. L'API UUID de la bibliothèque standard devient stable, pour générer et manipuler des identifiants uniques nativement. Nouvelles fonctions utilitaires comme isSorted() pour vérifier si une collection est déjà triée. Support de Java 26 côté JVM et alignement automatique des versions Java et Kotlin dans les projets Maven. Kotlin/Native, la compilation vers du code natif iOS et macOS, active par défaut un nouveau ramasse-miettes plus rapide et améliore l'export vers Swift. Kotlin/Wasm, la compilation vers WebAssembly pour faire tourner du Kotlin dans le navigateur, rend la compilation incrémentale stable. Kotlin/JS permet désormais d'exporter des value classes vers JavaScript et TypeScript. Le compilateur K1, l'ancienne génération, n'est plus supporté : seul le nouveau compilateur K2 reste disponible. JEP 540 : une API JSON simple intégrée au JDK (incubation) openjdk.org/jeps/540 Le JDK ne propose aujourd'hui aucune API JSON native, obligeant à dépendre de bibliothèques externes comme Jackson, Gson ou Jakarta JSON pour parser ou générer du JSON. Cette JEP remplace la JEP 198 de 2014 et cible JDK 28 avec le nouveau module incubateur jdk.incubator.json. L'objectif est de couvrir les besoins simples d'extraction de données sans binding de données ni API de streaming, en laissant ces cas avancés aux bibliothèques existantes. L'API s'articule autour de l'interface scellée JsonValue avec six sous-types : JsonString, JsonNumber, JsonBoolean, JsonNull, JsonObject et JsonArray. Le parsing est strict et conforme à RFC 8259 : pas de virgules finales, pas de commentaires, et les noms de membres dupliqués provoquent une erreur. Donc pas de JSON5 La navigation se fait via get et tryGet, et la conversion vers des types Java via asInt, asLong, asDouble, asString, asMap ou asList. En cas d'erreur, une JsonValueException précise le chemin exact dans le document et sa position en ligne et colonne. Le pattern matching sur les sous-types de JsonValue permet de gérer proprement l'évolution du format d'un document JSON dans le temps. La génération se fait via toString pour une sortie compacte ou Json.toDisplayString pour une sortie indentée et lisible. À terme, le JDK pourrait utiliser cette API en interne, par exemple pour remplacer les fichiers de configuration au format property par du JSON. Autres nouvelles du JDK openjdk.org/jeps/535 openjdk.org/jeps/541 le mode generationel pour Shenandoah est prévu par défaut et deprécue le non générationel en 28 fini le support de Java sur Apple Intel GraalVM 25.2 : références compressées et Graal Script Agent medium.com/graalvm/… GraalVM est une machine virtuelle polyglotte d'Oracle offrant compilation JIT avancée et compilation en image native pour accélérer les applications Java et d'autres langages. Cette version 25.2 fait partie du train de releases innovation qui livre les nouveautés plus vite, pendant que GraalVM 25.0 reste la version stable recevant les correctifs de sécurité critiques. Nouveauté phare, le Graal Script Agent transforme des demandes en langage naturel en plugins sandboxés exécutés localement, en JavaScript ou Python, avec un accès restreint aux APIs de l'application. Les références compressées sont désormais activées par défaut dans Native Image sur les systèmes 64 bits, remplaçant les adresses complètes par des valeurs 32 bits relatives au tas. Cette optimisation réduit de 39% la consommation mémoire RSS d'une application Micronaut connectée à Oracle Database, comparée à la version 25.0. Contrepartie de cette optimisation, le tas géré est désormais plafonné à 32 Go. Le garbage collector G1 est maintenant disponible sur toutes les plateformes, y compris Windows, via l'option –gc=G1. G1 apporte de meilleures performances, une latence réduite et un démarrage plus rapide, avec des images natives plus petites grâce à l'optimisation guidée par profil. Le Vector API de Java est activé par défaut pour exploiter les instructions SIMD, utile pour le machine learning et le traitement de données. Bonne intégration avec l'écosystème via Micronaut 5.1, Quarkus et WebAssembly. Shopify arrête React Native pour ses applis mobiles iOS et Android et repasse à du natif avec Swift et Kotlin shopify.engineering/back-to-native Les progrès majeurs des LLM (IA) réduisent drastiquement le coût du développement sur deux plateformes distinctes. Les bénéfices du natif pur restent supérieurs, moins de couches d'abstractions, de dépenfances externes, et plus rapide pour adopter les dernières fonctionnalités des OS Les bibliothèques open-source (Skia, FlashList, Restyle) évoluent : Skia sera forkée par William Candillon, FlashList cherche un nouveau repreneur, Restyle sera archivée fin 2026. Migration des applications (Shop, Shopify, etc.) réalisée en mode "greenfield" (reconstruction totale) assistée par IA. Utilisation du système "Helix" pour un développement itératif et contrôlé par des agents IA. Découplage de la logique métier et de l'interface via une CLI pour accélérer les tests et éviter les lenteurs des simulateurs. L'application Shop a été entièrement reconstruite en natif en 12 semaines ; les autres suivront. Librairies LangChain4j CDI est une extension CDI qui intègre LangChain4j avec CDI de Jakarta EE langchain4j.github.io/langchain4j-cdi LangChain4j CDI : Extension intégrant LangChain4j à Jakarta EE et MicroProfile. Services IA : Injection et gestion de cycle de vie via @RegisterAIService. Orchestration d'agents : 11 topologies d'agents configurables par annotations. Serveur MCP : Conversion de beans CDI en serveurs Model Context Protocol. Fonctionnalités d'entreprise : Configuration externe, tolérance aux pannes et observabilité OpenTelemetry. Installation Maven : Deux extensions disponibles selon l'environnement (build-time pour Quarkus/Helidon, portable pour WildFly/GlassFish/Liberty). Prérequis techniques : Java 17+, Jakarta EE 10, MicroProfile 6.1. Quarkus 3.36, 3.37 et 3.38 : trois releases avant Quarkus 4 quarkus.io/blog/quarkus-3-38-released quarkus.io/blog/quarkus-3-37-released quarkus.io/blog/quarkus-3-36-released Quarkus est un framework Java cloud natif optimisé pour GraalVM et HotSpot, conçu pour les microservices et les environnements conteneurisés. En 3.38 (29 juillet), l'équipe allège les nouveautés pour se concentrer sur Quarkus 4, la communauté atteint 1213 contributeurs. 3.38 introduit l'éviction basée sur le poids mémoire pour le cache Caffeine de second niveau d'Hibernate, en plus de l'éviction par comptage. 3.38 apporte l'extension Quarkus HTTP Problem qui implémente la RFC 9457 pour mapper les exceptions en réponses application/problem+json, intégrée à OpenAPI. 3.37 (24 juin) ajoute l'extension expérimentale quarkus jlink pour générer des images runtime JDK sur mesure et réduire la taille des conteneurs. 3.37 active par défaut la sérialisation Jackson sans réflexion pour de meilleures performances. et 3.39 lesdesactivent et les rement en opt-in 3.37 introduit dans REST Client RestMultiResponse pour lire codes de statut et en-têtes sur des réponses REST en streaming, avec passage à Hibernate ORM 7.4 qui exige PostgreSQL 14 minimum. 3.36 (27 mai) propose Quarkus Signals en expérimental, un système de communication typée entre composants inspiré des events CDI et de l'EventBus Vert.x. 3.36 embarque des SBOM applicatifs exposés via /.well-known/sbom, y compris en image native selon la spécification GraalVM. 3.36 ajoute l'authentification OIDC via JWT SPIFFE, facilitant l'identité de charge de travail en environnement zero trust. Micronaut Framework 5.1.0 : injection de dépendances, IA et sécurité renforcées github.com/micronaut-projects/micronaut-platform/…/v5.1.0 Micronaut est un framework JVM pour microservices et applications cloud-natives, avec injection de dépendances à la compilation et démarrage rapide. Introduction d'Open DI 1.0.0, une implémentation CDI Lite s'appuyant sur l'infrastructure d'injection de dépendances de Micronaut. Côté données, support officiel de SQLite et intégration MyBatis, avec ETags basés sur les valeurs pour le verrouillage optimiste. En sécurité, arrivée d'un module OWASP HTML Sanitizer, délégation d'authentification @RunAs et résolution de locale via OIDC. Côté IA, LangChain4j ajoute le support Chroma, la mémorisation de chat Oracle et l'authentification Google injectée pour Vertex AI, avec passage du MCP en version 2.0.0. Mises à jour majeures des dépendances : Spring Boot 4.1.0, Jetty 12.1.10, Tomcat 11.0.23, OpenTelemetry 1.64.0 et Kubernetes Java Client 27.0.0. SSL activé par défaut par service pour les clients HTTP Infrastructure Ça coûte combien de faire tourner un LLM local sur son Apple Silicon ? towardsdatascience.com/how-much-does-a-local-llm-actually-cost-to-run-i-measured-every-watt-on-apple-silicon Coût électrique des LLM locaux sur Mac Apple Silicon Un modèle 120B (MoE) coûte 5x à 10x moins cher qu'un modèle 27B (Dense). Le coût dépend du débit (tokens/seconde), pas du nombre de paramètres. Modèle dense –> Charge 100% des poids par token = lent et très énergivore. MoE (Mixture of Experts) –> N'active qu'une fraction des poids = rapide et économe. Conclusion : Pour réduire la facture électrique, choisir des modèles MoE quantifiés (haut débit). Kubernetes 1.36 (Haru) : sécurité renforcée et alignement IA infoq.com/news/2026/05/kubernetes-1-36-released Kubernetes est la plateforme open source de référence pour l'orchestration de conteneurs, portée par la CNCF. La version 1.36 nommée Haru apporte 70 améliorations : 18 passent stables, 25 en bêta et 25 en alpha, avec 106 entreprises et 491 contributeurs. Les user namespaces passent en disponibilité générale, isolant le root du conteneur de celui de l'hôte. Les Mutating Admission Policies passent en GA, remplaçant les webhooks par des règles CEL natives plus performantes. L'autorisation de l'API kubelet devient plus fine, remplaçant le droit trop large nodes/proxy. Le labeling SELinux des volumes utilise désormais mount -o context, accélérant le démarrage des pods. Plusieurs avancées ciblent les charges IA : gang scheduling en bêta, préemption consciente des groupes de pods, et allocation dynamique de ressources activée par défaut pour le partage fin des GPU. Le redimensionnement vertical des pods en place passe en bêta et activé par défaut, ajustant CPU et mémoire sans redémarrage. Suppression du plugin gitRepo, source de risque de sécurité, et du mode IPVS de kube-proxy, tous deux dépréciés de longue date. Avec la sortie de 1.36, la version 1.34 devient la plus ancienne branche encore supportée et entre en maintenance, ne recevant plus que des correctifs critiques avant sa fin de support. Terraform vs OpenTofu en 2026 : la divergence est actée ecorpit.hashnode.dev/terraform-vs-opentofu-in-2026-the-fork-has-diverged-so-which-do-you-standardize-on env0.com/insights/opentofu-in-2026-what-the-terraform-fork-became-after-three-years-of-independence Terraform est l'outil historique d'Infrastructure as Code de HashiCorp, OpenTofu en est le fork open source lancé après le changement de licence. HashiCorp est passé de la licence MPL 2.0 a la BUSL 1.1 en aout 2023, ce qui a poussé une partie de la communauté a créer OpenTofu sous la Linux Foundation. IBM a racheté HashiCorp pour 6,4 milliards de dollars, finalisé en février 2025, tandis qu'OpenTofu rejoignait le CNCF comme projet sandbox en avril 2025. OpenTofu prend de l'avance sur des fonctionnalités inédites : chiffrement du state côté client depuis la v1.7, valeurs éphémères qui gardent les secrets hors du state depuis la v1.11, et prevent_destroy dynamique en v1.12 (mai 2026). Terraform garde l'avantage sur l'orchestration managée avec Terraform Stacks, désormais en disponibilité générale, sans équivalent natif côté OpenTofu. Le coût diverge fortement : HCP Terraform facture jusqu'à 0,99 dollar par ressource gérée et par mois, alors qu'OpenTofu reste une CLI gratuite couplée au backend de son choix. Fidelity Investments a migré plus de 2000 applications et 50000 fichiers d'état vers OpenTofu, la complexité venant surtout de l'écosystème (CI/CD, gouvernance) plutôt que du binaire lui-même. OpenTofu reste compatible avec les configurations Terraform jusqu'a la version 1.6.x, mais les versions Terraform plus récentes n'offrent plus aucune garantie de compatibilité. Pour les secteurs régulés, le chiffrement natif du state par OpenTofu et sa gouvernance ouverte sont des arguments forts face aux exigences de protection des données. La recommandation qui ressort des deux articles : partir sur OpenTofu pour les projets neufs et rester sur Terraform si l'on est déjà investi dans HCP Terraform et ses fonctionnalités de gouvernance. OTel est à la peine ? matduggan.com/otel-isnt-going-well-and-i-made-a-spreadsheet-about-it Le développement d'OTel est un peu au point mort Périmètre démesuré : Volonté de supporter un nombre gigantesque de langages, bibliothèques et frameworks. Pénurie critique de mainteneurs : Les données montrent une hyper-concentration du travail ; de nombreux SDK (comme PHP ou Ruby) dépendent d'une ou deux personnes seulement. Stabilité paralysante : La règle interdisant toute modification d'une fonctionnalité déclarée « stable » crée une peur de valider les nouveautés, entraînant des mois de débats. Solutions proposées par l'auteur : Créer un niveau « Bêta » temporaire (ex: 12 mois) entre les statuts « Expérimental » et « Stable ». Faire preuve de transparence sur les différences de qualité/maintenance entre les langages (ne pas mettre Go et Ruby sur le même plan). Communiquer activement sur le besoin urgent de nouveaux mainteneurs. Assouplir la politique de stabilité en acceptant des breaking changes bien documentés. Honeycomb transforme son infrastructure Kafka honeycomb.io/blog/transforming-how-we-run-kafka-honeycomb Honeycomb est une plateforme d'observabilité dont Kafka est le coeur du pipeline d'ingestion, traitant des millions d'événements par seconde. L'entreprise a migré de Confluent Platform auto-hébergée vers Apache Kafka 4.1.1 en mode KRaft. Le nouveau cluster tourne sur AWS EKS avec Strimzi comme couche d'orchestration Kubernetes. Motivation principale : la récupération après remplacement de broker était passée de 8-12h à 48-72h avec l'ancienne stack. Confluent imposait aussi sa solution propriétaire de Tiered Storage, impossible à corriger en interne. Un incident de décembre 2025 ayant vidé un cluster a révélé une fenêtre d'opportunité pour migrer. La migration s'est faite progressivement sur six clusters, de dogfood jusqu'à la production. Les producteurs sont basculés avant les consommateurs, avec une courte fenêtre de downtime assumée entre les deux. Le stockage utilise des NVMe en instance store plutôt que de l'EBS pour minimiser la latence. interessant de voir une société reprendre en main sa compétence et de voir les contraintes de certaines fonctionalités propriétaires Cloud AWS us-west-2 : panne réseau régionale et effet domino chez les fournisseurs SaaS blog.incidenthub.cloud/aws-us-west-2-outage-jul-24-2026 AWS us-west-2 (Oregon) est une région cloud majeure hébergeant de nombreux services et fournisseurs SaaS. Le 24 juillet 2026, une panne matérielle réseau a coupé la connectivité entre la région et le Seattle Metro pendant environ 20 minutes. Particularité notable, seule la couche de connectivité externe a été touchée, le trafic interne à la région a continué de fonctionner normalement. Après la réparation matérielle, une phase distincte de reconvergence des routes a de nouveau causé une connectivité intermittente pendant plusieurs dizaines de minutes. Les clients Direct Connect via EqSe2 ont subi une coupure bien plus longue que le reste, 1h17 au total. Neuf incidents chez sept fournisseurs ont cité explicitement AWS comme cause, dont SendGrid, SparkPost et NinjaOne. Fait marquant, les temps de rétablissement des fournisseurs tiers ont largement dépassé la durée de la panne AWS elle même. NinjaOne a mis 9h29 à se rétablir totalement, avec 150000 appareils tentant de se reconnecter simultanément freinés par des mécanismes de backoff et jitter. SparkPost a mis 6h45 à absorber l'arriéré de courriels accumulé pendant la coupure, avec encore 80 à 90 minutes de retard des heures plus tard. L'article recommande d'identifier ses dépendances en us-west-2 et de prévoir capacité et bascule, l'effet différé pouvant durer bien plus longtemps que l'incident initial. Rapport Cloudflare Radar sur les perturbations Internet au Q2 2026 blog.cloudflare.com/fr-fr/q2-2026-internet-disruption-summary Cloudflare Radar est la plateforme qui analyse en temps réel le trafic mondial pour détecter pannes, coupures et censures Internet. l'instabilité est le nouveau normal Le super-typhon Sinlaku a fait chuter le trafic de près de 80% à Guam les 13 et 14 avril. Deux séismes de magnitude 7,5 ont fortement dégradé la connectivité au Venezuela le 24 juin. Une coupure électrique a provoqué cinq heures de perturbation en Tanzanie le 27 juin. En Iran, la connectivité s'est stabilisée à 59% du niveau normal après 88 jours de coupure. Le Soudan a imposé dix coupures programmées pendant les examens nationaux mi-avril. L'Irak a coupé Internet à trois reprises pour lutter contre la fraude aux examens. Des frappes de drones ont endommagé la région AWS me-central-1 aux Émirats arabes unis. Un renouvellement de clés DNSSEC a rendu les sites .de inaccessibles en Allemagne le 5 mai. Une rupture de câble sous-marin a fait chuter le trafic de 60% à Sainte-Lucie fin juin. l'instabilité est le nouveau normal Un ingénieur Google débranche une zone entière de Google Cloud https://www.theregister.com/off-prem/2026/09/04/google-engineer-unplugged-every-fiber-they-could-see-and-surprise-took-down-a-chunk-of-the-g-cloud/5294418 Google Cloud est la plateforme d'infrastructure cloud de Google, organisée en régions et zones de disponibilité comme us-central1. Le 1er septembre 2026, un ingénieur a débranché par erreur des câbles fibre optique lors d'une opération de maintenance matérielle routinière dans la zone us-central1-b. En 13 minutes, il a déconnecté 100 % des chemins de fibre optique de tous les équipements de cette portion de la zone. Les machines virtuelles hébergées dans la zone touchée sont devenues injoignables, avec une perte de paquets élevée. Le taux de chute du trafic réseau pour les ressources concernées a atteint 100 %. L'incident a duré 4 heures et 11 minutes, de 7h41 à 11h52 heure du Pacifique. Google a détecté l'anomalie, reroute le trafic, identifié les liaisons optiques débranchées puis rebranché physiquement les fibres avant le retour à la normale. Les post-mortems d'incident cloud sont un classique du podcast, mais l'erreur humaine sur du câblage physique chez un hyperscaler mérite une minute d'antenne. ChatGPT, Claude et Grok en panne presque simultanément le 3 septembre theregister.com/ai-and-ml/…/5294322 ChatGPT, Claude et Grok sont les assistants IA et coding agents désormais utilisés au quotidien par de nombreux développeurs. xAI loue ses datacenters à un certains nombre de fournisseurs de cloud et d'IA. ils ont fait tombé ses concurrents :slightly_smiling_face: Arreter là Le 3 septembre 2026, les trois services sont tombés en panne quasiment en même temps. ChatGPT a connu une panne de 7h43 à 8h17 PT, soit environ 34 minutes, due à une erreur de routage rendant ChatGPT et Codex indisponibles. Claude a subi une panne partielle de 3 heures et 6 minutes touchant Claude.ai, Claude Code, Claude Cowork et l'API Claude, résolue à 16h16 UTC. xAI a commencé à enquêter sur les problèmes de Grok dès 6h30 PT, puis SpaceX a confirmé une panne de son centre de calcul de Memphis. La coïncidence des trois pannes a fait suspecter un fournisseur commun à l'origine du problème. Pour les développeurs devenus dépendants de ces coding agents, l'épisode illustre le risque d'un point de défaillance unique xAI loue ses datacenters à un certains nombre de fournisseurs de cloud et d'IA. Web Nouveautés CSS 2026 : mixins, masonry natif et animations pilotées par le scroll modern-css.com/whats-new-in-css-2026 animation-timeline: scroll() et view() atteignent le baseline cross-browser, Firefox et Safari ayant livré le support complet. Plus besoin de préfixes ni de librairie JS. @starting-style devient cross-browser : animations d'entrée depuis display: none sans hack de timing JS. Firefox 147 amène l'anchor positioning au baseline, plus les view transition types et la Navigation API. Les menus contextuels via popover CSS arrivent aussi. Data et Intelligence Artificielle Guillaume a porté le SDK Python d'Antigravity en Java… en utilisant Antigravity lui même comme assistant ! glaforge.dev/posts/…/the-unofficial-antigravity-sdk-for-java SDK Java non officiel pour Antigravity, rétro-ingénierie du SDK Python pour exploiter un binaire Go sous-jacent. Cas d'usage : pipelines CI/CD, applications d'entreprise (Spring Boot), outils internes, surveillance en arrière-plan, interfaces personnalisées. Gestion des ressources : implémente AutoCloseable(try-with-resources) pour lancer et fermer proprement le processus Go. Exécution de code Java personnalisé : exposition de méthodes Java comme outils IA via les annotations @Toolet@Param. Streaming et programmation réactive : prise en charge des CompletableFuture, des callbacks (chatStream), et deFlow.Publisher. Fonctionnalités avancées : persistance de session, protocole MCP, politiques de sécurité, entrées multimodales, et sorties structurées mappées sur des records Java. Le SDK Java pour le protocol Agent2Agent sort sa version 1.2.0 medium.com/google-cloud/a2a-java-sdk-1-2-0-final-released… Prise en charge de la spécification A2A 1.0. Sécurité renforcée : Vérification stricte des autorisations de lecture sur les tâches référencées et implémentation d'une logique de blocage par défaut (fail-closed). Intégration facilitée : Support du câblage programmatique des autorisations pour les environnements non-CDI (comme Spring). Contrôle des flux : Ajout du Task Stream Lifecycle Hook pour surveiller et gérer le cycle de vie des abonnements aux flux d'événements. Stabilité des données : Immutabilité stricte imposée sur les enregistrements de spécifications pour empêcher toute modification accidentelle. Documentation : Lancement d'une documentation web multi-versions et d'un Javadoc agrégé pour tous les modules. Corrections de bugs : Résolution de problèmes liés à la synchronisation des tâches, aux réponses de streaming et aux conditions de concurrence HTTP. Breaking changes : Nécessite une migration suite à la modification de certaines méthodes d'autorisation, la réorganisation de packages et le renommage de l'état TaskState.UNRECOGNIZED en TaskState.TASK_STATE_UNSPECIFIED. Article sur le blog de JetBrains : blog.jetbrains.com/idea/2026/08/intellij-idea-goes-lsp pas trop de temps Langchain4j continue sa progression github.com/langchain4j/langchain4j/…/1.18.0 github.com/langchain4j/langchain4j/…/1.19.0 github.com/langchain4j/langchain4j/…/1.17.0 Introduction du pattern Debate pour lancer des sous agents dans rechercehnt et entrent dans un debat critique sur une decision avant qu'un juge decide (1.17) compensation d'action d'un outil avec @ReverseTool (1.17) Ajout du pattern Belief-Desire-Intention (1.18): belief est l'état du monde cru, desire est la liste des objectifs, intention est le plan pour avancer un ou plusieurs objectif Support d'un systeme crash resilient dans l'approche Human in the loop avec des checkpoints nestés (1.18) Support Open AI TextToSpeech (1.18) Support de Mistral batch chat (1.18) support MCP client 2026-07-28 (1.19) Anthromic batch chat et Google thinking mode (1.19) Embabel atteint 1.0 GA github.com/embabel/embabel-agent/…/v1.0.0 embabel d'eloigne de Spring AI ne s'appuie que sur Spring (notamment tool calling) nettoyage des explorations autour du pattern GOAL avant a 1.0 ajout observabilite (dont le cout) les approches de retry solidifiées et d'autres choses toujours basées sur la description de goals typés et des dependences entre eux via les types MCP 2026-07-28 : le protocole devient stateless blog.modelcontextprotocol.io/posts/2026-07-28 MCP (Model Context Protocol) est le protocole standard permettant aux LLM et agents IA de communiquer avec des outils, ressources et serveurs externes. Cette version marque le plus gros changement depuis le lancement du MCP distant il y a 18 mois. Le protocole passe d'un modèle bidirectionnel avec état à un modèle stateless en requête/réponse. Suppression de la poignée de main initialize/initialized et du header Mcp-Session-Id, chaque requête devient autoportante. N'importe quelle requête peut désormais être routée vers n'importe quelle instance de serveur derrière un load balancer classique, sans stockage partagé. Introduction des Multi Round-Trip Requests (MRTR) pour remplacer les requêtes initiées par le serveur, via un resultType input_required et des inputResponses. Nouveaux headers Mcp-Method et Mcp-Name pour permettre aux gateways de router et autoriser sans parser le JSON. Les résultats de tools, prompts et resources deviennent cacheables grâce aux paramètres ttlMs et cacheScope. Renforcement sécurité avec la validation d'issuer RFC 9207 pour éviter les attaques de confusion entre serveurs d'autorisation, et transition de DCR vers CIMD. Roots, Sampling et Logging sont dépréciés avec douze mois de support garanti, tout comme le transport legacy HTTP+SSE. Un site qui référence les skills pour la JVM (framework, langage, build…) jvmskills.com Frameworks : Spring, Quarkus, Jakarta EE, Reactor, Camel Java : bonne pratiques, conventions, guides de mise à jour à niveau LTS, API spécifiques (streams, optionals, logs…) Bases de données : ORM, validation, modélisation PostgreSQL, vectorielle avec pgvector Tests et qualité : TDD, mutation testing, debogage avec JDB Workflows dev et archi : commits git, domain modeling Outils et diagnostics JVM : JFR, Jstall, JSpecify Une skill n'est pas une librairie https://devx.writizzy.blog/p/un-skill-nest-pas-une-lib Les skills sont des éléments de configuration en prose pour agents IA comme Claude, distribués via des marketplaces à la manière de librairies logicielles. Frédéric Camblor critique cette analogie car partager un skill n'est pas la même chose que le mutualiser durablement. Écrire un skill prend 30 minutes mais l'adopter ailleurs coûte cher en appropriation et en maintenance. Forker un skill s'avère souvent plus efficace que de chercher à converger vers une version commune. Contrairement au code, les régressions d'un skill ne sont pas détectables automatiquement. Un skill peut se dégrader silencieusement sur plusieurs cas d'usage en corrigeant un autre. Les skills vieillissent vite car les modèles progressent et intègrent naturellement certaines bonnes pratiques. Le skill-creator d'Anthropic permet d'évaluer un skill via des jeux de cas et des mesures de variance. L'auteur distingue quatre sphères de partage : personnelle, équipe, outil et marketplace. Il propose un cycle partage puis appropriation puis duplication puis divergence plutôt qu'une installation collective figée. Les modèles Anthropic introduisent un filigrane (watermark) dans les textes qu'ils génèrent https://www.anthropic.com/news/claude-text-watermark Claude est l'assistant IA d'Anthropic, et le watermarking est une technique permettant de marquer discrètement un contenu généré par IA pour en tracer l'origine. Anthropic annonce que les futurs modèles Claude intégreront un filigrane numérique invisible dans le texte généré. Le principe exploite les choix de mots équivalents que le modèle fait naturellement, en les orientant via une clé cryptographique plutôt qu'un tirage aléatoire. Le texte produit reste indiscernable à l'œil nu, sans caractères cachés, sans ralentissement ni coût supplémentaire. Seule la personne possédant la clé correspondante peut détecter la présence du filigrane. Le filigrane est plus fiable sur les textes longs et créatifs, moins sur du texte factuel, du code ou après une édition manuelle poussée. Il ne prouve pas qu'un texte est écrit par IA, ni n'identifie l'auteur ou la conversation d'origine, il donne seulement une probabilité d'implication de Claude. Une API de détection est proposée en accès restreint aux régulateurs, forces de l'ordre, médias et vérificateurs de faits. Pour les fichiers non textuels comme les images ou les PDF, Anthropic s'appuie sur le standard C2PA. Cette initiative s'inscrit dans le Code de Pratique de l'UE sur la transparence des contenus IA, signé par Anthropic et environ 190 autres acteurs, en lien avec la loi européenne sur l'IA. GPT-6 Astra, le nouveau modèle d'OpenAI face à Claude Fable 5.1 https://openai.com/index/gpt-6-astra/ GPT-6 Astra est le nouveau modèle phare d'OpenAI, annoncé le 3 septembre 2026 comme le plus intelligent et le plus aligné de l'entreprise. Le modèle arrive deux jours après Claude Fable 5.1, à un tarif affiché comparable, dans une course accélérée aux modèles de code et de raisonnement. Astra revendique 98 % sur FrontierMath Tier 4, 99,9 % sur ARC-AGI-3 et 100 % sur ExploitBench. Fenêtre de contexte d'environ 1,05 million de tokens. Tarification API, 10 dollars par million de tokens en entrée, 50 dollars en sortie, et 1 dollar par million pour les tokens en cache. Au-delà de 272 000 tokens en entrée, toute la requête est facturée au double sur l'entrée et une fois et demie sur la sortie. Astra est le premier modèle d'OpenAI à franchir le seuil interne critique en cybersécurité. La version publique refuse les tâches offensives avancées comme générer des preuves de concept d'exploits. Le déploiement est progressif, les entreprises du programme de cybersécurité Daybreak d'OpenAI y accèdent en premier, avant ChatGPT Plus, Pro, Business, Enterprise, l'API et AWS. Même logique de diffusion contrôlée que chez Anthropic avec Mythos 5.1 : deux jours d'écart, deux modèles de tête, et la même question de savoir qui accède en premier aux capacités les plus sensibles. Outillage JetBrains s'est lancé dans les LSP (Language Server Protocol) avec une extension IntelliJ pour VS Code et assimilés marketplace.visualstudio.com/items?itemName=JetBrains.intellij-s… Nouveau produit : Lancement de l'extension Java & Kotlin by IntelliJ IDEA pour les éditeurs basés sur VS Code (incluant Cursor). Technologie : Utilisation du standard LSP (Language Server Protocol). Objectif : S'adapter au développement piloté par les agents IA, qui nécessite des fonctionnalités IDE légères et standardisées. Fonctionnalités clés : Support des projets Java, Kotlin et mixtes. Débogage (DAP). Complétion intelligente, navigation et analyse de code. Refactoring. Prise en charge de Maven, Gradle et Bazel. Disponibilité : Téléchargeable via le Visual Studio Marketplace et l'Open VSX registry. Licence / Prix : Gratuit durant la phase de preview (évaluation renouvelable de 30 jours). Nécessitera un abonnement IntelliJ IDEA Ultimate après la preview. (Note : Le LSP purement Kotlin reste gratuit et open-source). Avenir : Développement en cours pour optimiser les flux de travail avec les agents IA en ligne de commande (ex: Claude Code, Codex) afin de réduire la consommation de tokens. Après son acquisition par SpaceX, Cursor perd l'accès aux modèles OpenAI https://openai.com/index/our-decision-on-cursor-following-its-acquisition-by-spacex/ decision difficile mais Elon ment comme un arracheur de dent (et c'est un competiteur donc bon ça nous arrange) David Pilato a créé un thème spéciale pour les gens pour les devrel, ou qui font des talks à droite à gauche, pour le moteur Hugo https://david.pilato.fr/posts/2026-09-07-hugo-theme-devrel/ hugo-theme-devrel, thème Hugo (MIT) pour Developer Advocates et conférenciers, fonctionnant comme un module superposé au thème Dream. Gestion des conférences (cartes Leaflet), présentations (PDF, YouTube, co-auteurs), vues dédiées (archives, sujets récurrents, vidéos) et recherche Pagefind. Chaque intervention est un page bundle YAML structuré comme une base de données relationnelle compilée par Hugo. Architecture Airbnb refond son authentification en architecture server-driven, 60 % de code en moins https://www.infoq.com/news/2026/09/airbnb-server-driven-login/ Airbnb a restructuré son authentification autour d'un modèle en deux phases, identification du compte (email, téléphone ou connexion sociale) puis choix du challenge d'authentification décidé côté serveur selon le contexte utilisateur. Ce qui est interessant c'est que choisir la method d'authentification est côté serveur et adaptative, comme si c'était une révolution Arreter là Un moteur de politique serveur sélectionne la méthode d'authentification optimale avec des solutions de repli, permettant des adaptations régionales comme l'OTP WhatsApp au Brésil ou des fournisseurs d'identité locaux en Corée du Sud sans nouvelle version client. Un Challenge Picker propose des méthodes alternatives classées par probabilité de succès en cas d'échec. Résultat chiffré, 60 % de code d'authentification en moins et 100 Ko de moins sur le bundle client web. Méthodologies Les nouvelles règles d'ingénierie du contexte pour les modèles Claude 5 x.com/trq212/status/2080710971228918066 Partage par Thariq des apprentissages sur l'ingénierie du contexte et le prompt engineering pour les nouveaux modèles Claude 5 (comme Claude Opus 5 et Claude Fable 5) utilisés dans Claude Code. Évolution majeure vers le dés-empirement (unhobbling) : plus de 80 % du prompt système de Claude Code a pu être supprimé sans perte sur les évaluations de code, les modèles récents faisant preuve d'un bien meilleur jugement contextuel. Passage des règles strictes au jugement : au lieu d'interdire les commentaires ou d'imposer des contraintes lourdes, les modèles s'adaptent désormais au code environnant et font appel à leur propre discernement. Remplacement des exemples par la conception d'interfaces : fournir des exemples figés restreint l'exploration du modèle, d'où l'importance de concevoir des outils et des fichiers plus expressifs. Adoption de la divulgation progressive (progressive disclosure) : chargement dynamique du contexte (via des compétences ou des outils à chargement différé comme ToolSearch) pour éviter de saturer la fenêtre de contexte avec des instructions fixes. Utilisation d'une mémoire automatique et de références riches (artefacts HTML, suites de tests, fonctions de référence) plutôt que de fichiers CLAUDE.md pléthoriques ou de consignes répétitives. Recommandation pour les fichiers CLAUDE.md et les Skills : les garder légers, se concentrer sur les pièges spécifiques (gotchas) du dépôt, et structurer les guides sous forme d'arborescences modulaires pour ne charger que le nécessaire. Niveau d'adoption des agents IA de codage selon une étude de JetBrains blog.jetbrains.com/research/2026/08/ai-coding-agent-adoption-2026 Adoption massive : 90 % des développeurs professionnels utilisent des agents d'IA de codage au moins une fois par semaine, et 68 % quotidiennement. Claude Code domine : Il devient le nouveau leader du marché avec 39 % d'adoption mondiale (47 % aux États-Unis), détrônant largement ses concurrents. Déclin de GitHub Copilot : L'ancien leader perd de sa superbe, passant de 29 % à 21 % d'adoption, bien qu'il conserve une très forte notoriété (79 %). Percée de Codex : Sa croissance est fulgurante, son taux d'adoption ayant été multiplié par 5 en quelques mois (de 3 % à 16 %). Recul de Cursor : L'outil connaît une légère baisse, passant de 18 % à 12 % d'adoption, principalement due à une forte chute sur le marché chinois. Écosystème diversifié : JetBrains AI atteint 9 % d'adoption. Des alternatives comme OpenCode (7 %) et Google Antigravity (6 %, mais très populaire en Inde à 15 %) continuent de s'implanter. L'IA a cassé les hypothèses de la CI… ou pas https://stack72.dev/ai-broke-the-assumptions-behind-ci/ Paul Stack (ex-Pulumi) explique que l'intégration continue a toujours mêlé deux rôles distincts, exécuter la vérification du code et coordonner les fusions. avec les agents, la pression sur la CI augmente, car ils ne testent pas end to end tout le temps ils ont maintenant des workflows qui verifient tests, lint, revue d'agent etc dans un env local isolé donc c'est pre PR push Il propose de séparer vérification (exécution), attestations structurées (hash de commit, checksums SHA256) et CI, réduite à la validation et à la coordination des fusions. Sa thèse, les agents IA peuvent désormais vérifier tout le code localement avant même l'ouverture d'une pull request, ce qui bouleverse cet équilibre. l'attestation vient ensuite et la CI est une étape de vérification (attestation de commit et de tests et si branche a bougé, repart à l'execution) Martin Fowler, qui a repéré l'article dans ses Fragments du 1er septembre 2026, réplique que la vraie CI a toujours exigé une vérification locale avant de pousser le code. ca demande une chaine d'attestation forte quid de garder les metriques historique de CI Sécurité France Passoire: une analyse sur les différents vols de données des services publics de ces dernières mois https://www.cybernetica.fr/piratage-des-impots-comment-en-est-on-arrive-la/ Analyse du piratage massif de la DGFiP et d'autres administrations françaises en 2026, révélateur de failles systémiques de cybersécurité de l'État. Intrusion détectée fin juin à la DGFiP, mais l'exfiltration de 678 000 entrées fiscales n'a été découverte qu'en août lors de leur mise en vente. Données volées : noms, revenu fiscal de référence, taux de prélèvement, adresse, téléphone. Une seconde attaque du même pirate a visé le cadastre fin juillet, exposant plus de 2 millions de personnes. L'Éducation nationale a aussi été piratée fin juillet, données de tous les agents depuis 2001 exposées. Cause principale : modèle de sécurité fondé sur le périmètre physique plutôt que sur le zero trust, sans contrôle après authentification. Le télétravail post Covid a étendu les accès distants sans reconstruire les modèles de confiance. Aucun système de détection d'exfiltration n'existait, la fuite n'a été révélée que par le pirate lui-même. La transposition de la directive NIS2 est bloquée en France depuis septembre 2025, la CJUE a condamné le pays à des astreintes. L'article souligne un désengagement croissant des Etats-Unis en matière de cybersécurité internationale et une dépendance technologique accrue de la France. Loi, société et organisation Ce que l'IA change vraiment au métier de manager shapeandship.ai/p/ce-que-lia-change-vraiment-au-metier-de-manager Retour d'expérience et analyse par Mathilde Rigabert sur l'impact réel de l'IA générative dans le quotidien d'un Engineering Manager. L'IA excelle pour automatiser la reconstitution factuelle de l'activité (lecture de PRs, commits, reviews) nécessaire aux 1:1 et entretiens annuels, mais elle offre une vision uniquement quantitative et nécessite d'être croisée avec des notes de terrain. L'IA rend le maintien de la qualité et des standards plus difficile : selon une étude Faros AI sur 22 000 développeurs, les PRs mergées sans aucune revue ont augmenté de 31 %, fragilisant la compréhension commune apportée par le pairing et les revues de code. Le temps gagné par l'IA ne permet pas d'augmenter massivement le span of control (seulement 2 ou 3 personnes de plus), car l'IA compresse la collecte d'informations mais pas les conversations humaines complexes ou l'accompagnement du changement. Les compétences d'orchestration et de gestion de sujets multiples acquises par les managers facilitent leur transition vers le pilotage de plusieurs agents IA en contribution individuelle. Le piège actuel réside dans l'accumulation des casquettes (manager, tech lead, product owner, contributeur, pompier), conduisant à l'épuisement et au délaissement du travail de fond sur l'organisation et l'humain. Le temps libéré par l'IA doit être réinvesti dans le travail invisible qui fait tenir le système (suivi des actions de rétro, analyse de métriques, coaching), que personne ne réclame à court terme mais dont l'absence fragilise les équipes à long terme. Je regrette d'avoir migré vers Codeberg xn–gckvb8fzb.com/i-regret-migrating-to-codeberg L'auteur explique pourquoi il regrette d'avoir quitté GitHub pour Codeberg, à la suite des récentes modifications des conditions d'utilisation (ToS) de la plateforme. Codeberg a interdit les projets principalement générés par des LLM ainsi que les projets liés aux cryptomonnaies via des propositions de l'Assembly 2026, au motif qu'ils nuisent à sa réputation. blog.codeberg.org/protecting-our-floss-commons-from… Critique de l'argument de Codeberg sur l'absence de communauté des vibe coders, en rappelant que la majorité des logiciels libres (FOSS) sont créés par des développeurs solos sans communauté au sens romancé du terme. Ironie soulignée concernant la posture de Codeberg et Forgejo, qui a hérité de la communauté de Gitea après un hard fork avant de faire la leçon aux développeurs individuels. Alerte sur le risque de censure idéologique : interdire des catégories entières plutôt que de traiter les abus réels ou la consommation d'infrastructure crée un précédent dangereux pour une forge qui se veut libre. Proposition de solutions alternatives pour gérer l'impact des LLM et de la crypto : déclaration obligatoire via des cases à cocher, hébergement sur des tiers d'infrastructure spécifiques payants ou sous quotas, et disclaimers automatiques. Décision de l'auteur de quitter Codeberg pour mettre en place son propre serveur Git personnel afin d'éviter la dépendance à une plateforme qui modifie ses règles de manière unilatérale. Cloud souverain : Airbus choisit Scaleway pour l'hébergement de ses applications critiques https://www.usine-digitale.fr/aeronautique-spatial/airbus/cloud-souverain-airbus-choisit-scaleway-pour-lhebergement-de-ses-applications-critiques.OHBVBMSZIJELNOKN6G5F6B7NSI.html Scaleway est le cloud provider français filiale du groupe Iliad, positionné comme alternative souveraine aux hyperscalers américains. Airbus a lancé un appel d'offres de six mois consultant une cinquantaine d'acteurs dont OVHcloud, Thales, Google S3NS et Microsoft Bleu. Scaleway a été retenu pour héberger les applications critiques liées à la conception d'aéronefs, l'ingénierie, la production industrielle et les opérations. Le contrat prévoit la migration d'environ 70 applications d'ici 2028, puis jusqu'à 900 applications sur 5 à 6 ans. Le montant du contrat n'a pas été communiqué. Scaleway revendique zéro actionnaire, zéro employé et zéro filiale hors Union européenne pour garantir une protection contre les lois extraterritoriales. Damien Lucas, PDG de Scaleway, évoque une immunité complète face aux évolutions politiques et législatives externes. La plateforme doit aussi accélérer les usages d'intelligence artificielle d'Airbus, avec les modèles de Mistral AI déjà déployés chez Scaleway. Catherine Jestin, responsable numérique d'Airbus, souligne que cette intégration accélère la démarche IA du groupe. Ce choix ne remet pas en cause la stratégie multicloud d'Airbus, Scaleway venant compléter les fournisseurs existants pour les charges nécessitant le plus haut niveau de gouvernance et de résilience. Debian adopte une résolution sur l'usage responsable de l'IA générative lwn.net/Articles/1091231 La discussion sur l'usage des LLM dans Debian s'est tenue du 23 juillet au 13 août 2026, suivie d'un vote du 15 au 28 août 2026. 1045 développeurs Debian étaient éligibles à voter, avec un quorum de 48,49 votes largement dépassé par les huit options en lice. Les options allaient d'une interdiction stricte des contributions générées par LLM inscrite dans le contrat social à une acceptation encadrée des contributions IA. L'option gagnante au classement Condorcet est Responsible Use of Generative AI, devant Allow AI-Assisted Contributions with conditions et A cautious approach to generative AI. Le texte adopté n'interdit ni n'encourage l'usage d'outils d'IA générative dans le développement de Debian. Il exige que toute contribution, quels que soient les outils utilisés pour la produire, respecte les mêmes standards de qualité, correction, maintenabilité et conformité légale. Les contributeurs doivent comprendre, relire, tester et si besoin modifier la production assistée par IA avant de l'intégrer à Debian. Les informations sensibles du projet ne doivent pas être transmises à des fournisseurs d'IA non fiables, et la divulgation de l'usage de l'IA est encouragée sans être obligatoire. Le détail du vote et le texte complet de la résolution sont disponibles sur la page officielle [debian.org/vote/2026/vote_002](https://www.debian.org/vote/2026/vote_002). Contraste direct avec l'OpenJDK, qui a publié une politique interdisant le code généré par LLM (épisode 340), et avec l'auteur de jqwik qui a piégé sa librairie contre les agents. Conférences Nous contacter Pour réagir à cet épisode, venez discuter sur le groupe Google https://groups.google.com/group/lescastcodeurs Contactez-nous via X/twitter https://twitter.com/lescastcodeurs ou Bluesky https://bsky.app/profile/lescastcodeurs.com Faire un crowdcast ou une crowdquestion Soutenez Les Cast Codeurs sur Patreon https://www.patreon.com/LesCastCodeurs Tous les épisodes et toutes les infos sur https://lescastcodeurs.com/
A job-search app shipped a Flappy Bird playable that ends with "there's a relevant job for you." Two years ago that would have made zero sense. This month it's the trend: non-gaming apps are flooding the playable space, most of them vibe-coded with AI in an afternoon.Matej Lančarič and Ondrej Mosberger are joined again by Robin Kuyer (Unity PlayWorks) for the "everyone brings playables" format — with the bold creative twist of reversing the order. https://lancaric.substack.com/p/h1-2026-brutally-honest-ua-report?r=7qqafMatej's Brutally Honest UA report puts playables at 40-60% of creative spend, up from 30-40% a year ago, as AppLovin, Mintegral, and the SDK networks keep growing. Robin brings the "Frankenstein" playable (Wool Crush: arrows + yarn mechanics + save-the-cat, every popular concept stacked into one ad — plus his spreadsheet method for engineering ideas from what's trending), the non-gaming wave Ondrej closes with a Fumb Games playable vibe-coded in Playable Maker (30 iterations and the next version does it in three) and a clean memory-game hidden-object playable that respects the player's intelligence.⏱️ TIMESTAMPS00:00 Playables are 40-60% of creative spend — the Brutally Honest data06:25 Wool Crush — the Frankenstein playable and the spreadsheet method11:20 Non-gaming apps take over — Indeed's Flappy Bird, finance, cleaners19:40 Hero Wars — the open-roam level playable and how to iterate it25:05 Supercell's playable — lost seven times, should a playable let you lose?29:40 Sunday City — GTA timing, Grove Street, and the double-hook debate38:30 Cozy Florist's AI merge & the cats-with-guns wildcard44:50 Vibe-coded in Playable Maker + the memory-game playable that trusts youThis episode is brought to you by Kinoa — the AI operating system for mobile game operations: flows, live segments, in-app messages, push notifications, and A/B testing in one place, run by the operators who own the numbers. Carry1st saw +43% ARPDAU; PlayStudios saw +31% revenue on Tetris Block Party. Learn more at Kinoa.http://www.kinoa.ai?utm_source=MatejPodcast&utm_medium=Link&utm_campaign=Matej+Podcast&utm_id=100PVX Partners offers non-dilutive funding for game developers.Go to: https://pvxpartners.com/They can help you access the most effective form of growth capital once you have the metrics to back it.- Scale fast- Keep your shares- Drawdown only as needed- Have PvX take downside risk alongside you+ Work with a team entirely made up of ex-gaming operators and investorsFor an ever-growing number of game developers, this means that now is the perfect time to invest in monetizing direct-to-consumer at scale.Our sponsor FastSpring:Has delivered D2C at scale for over 20 yearsThey power top mobile publishers around the worldLaunch a new webstore, replace an existing D2C vendor, or add a redundant D2C vendor at fastspring.gg.This is no BS gaming podcast 2.5 gamers session. Sharing actionable insights, dropping knowledge from our day-to-day User Acquisition, Game Design, and Ad monetization jobs. We are definitely not discussing the latest industry news, but having so much fun! Let's not forget this is a 4 a.m. conference discussion vibe, so let's not take it too seriously.Panelists: Jakub Remiar, Felix Braberg, Matej LancaricJoin our slack channel here: https://join.slack.com/t/two-and-half-gamers/shared_invite/zt-3bckldvr8-8PXvzciMWdheOzED9hq0SAMatej LancaricUser Acquisition & Creatives Consultanthttps://lancaric.meFelix BrabergAd monetization consultanthttps://www.felixbraberg.comJakub RemiarGame design consultanthttps://www.linkedin.com/in/jakubremiarPlease share the podcast with your industry friends, dogs & cats. Especially cats! They love it!Hit the Subscribe button on YouTube, Spotify, and Apple!Please share feedback and comments - matej@lancaric.me
In this episode I sit down with Faryam Asif, CTO at Shufti, to unpack how identity verification is evolving as agents, deepfakes, and AI-driven attacks accelerate. The conversation focuses on how Shufti verifies individuals, businesses, and transactions, and why layered security is becoming essential in regulated industries. Faryam also explains how the company is adapting to new use cases like agent verification, age estimation, and audit trail requirements. Faryam explains how Shufti verifies individuals, businesses, and transactions across industries including financial services, healthcare, retail, gambling, social media, and crypto. We discuss document verification, facial biometrics, live selfie and video checks, address verification, NFC-based passport verification, KYB, and AML workflows. I raise the growing challenge of AI agents and asks how identity verification adapts when an automated agent, not just a human, is completing a workflow. Faryam introduces the emerging concept of KYA, know your agent, and explains why the industry is shifting toward verifying the person behind the action. We discuss how deepfakes and synthetic documents have lowered the cost and speed of fraud, making scalable attacks much easier than before. Faryam shares how Shufti is responding with layered verification, combining document checks, facial likeness, device intelligence, behavior analysis, risk scoring, media integrity, and database checks. The conversation explores how behavioral signals are becoming important as attackers learn to mimic human pauses and interaction patterns. We dig into compliance, auditability, and why regulated industries now need proof of how verification happened, not just who was verified. Faryam explains Shufti's own technology stack, including proprietary facial liveness, document verification, OCR, transaction monitoring, KYB, AI, and ML systems. We cover deployment options and integrations, including on-premises hosting, cloud APIs, SDKs, and plugins for platforms like WordPress, Shopify, and Okta. Faryam also discusses Shufti's global footprint, compliance posture, and use cases such as facial age estimation for social media and adult-content restrictions. I hope you enjoy it!
Hey yall, welcome back to ThursdAI, this is Alex, let me catch you up! Today on the show, we covered 1 week with Astra (hint, it's not quite AGI yet despite what we were told), DeepSeek V4.1 catches up to the frontier at a fraction of the cost, and Meta launches a free AI agent with it's own computer, that will take over the OpenClaw/Hermeses of the world for most people. Also huge this week, OpenAI claimed that a swarm of 10K agents of their unreleased model solved the Navier-Stokes, one of the millennium problems! I was stoked to have Chris Alexiuk from Nvidia on the show to cover the innovations DeepSeek put into this latest model! Oh, and the guy who quit Anthropic this week, and wrote an essay about “AI is going to kill all of us” somehow got 130M views on X, a mirriad of TV interviews and rekindled the doomerism movement, we talk about that too!Also, I already told about FullyConnected, CoreWeave's premier conference that's coming up, but they told me about a new announcement today, and you're not going to believe who it's about (not AI related). As a reminder, ThursdAI subscribers get a free ticket!Ok, let's dive in! ThursdAI - Highest signal weekly AI news show is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.OpenAI claims a Navier-Stokes solution from a 10,000 agent swarm from an unreleased model - with some drama (X, Blog, Paper)If you've been reading ThursdAI for a while, you may remember that a model couldn't tell which is higher, 9.9 or 9.11 and naysayers said that “AI can't math” Well, this week OpenAI claimed that a swarm of 10K agents, of an unreleased model, found a solution to the Navier-Stokes problem, in about 88 hours! They published a huge 160+ page paper and a Lean proof of the solution! I said on the show, this is like the moon landing equivalent of AI doing things humans didn't do before! Now, keep in mind, this is only a claim from OpenAI, the Clay Mathematics Institute is still reviewing it (moved the status of this problem from “unsolved” to “under review” so this isn't independently verified yet) but it's still an insane deal. The agents sent 2.7 million messages and burned about 130B tokens, of a model that has no public price yet, so it's hard to estimate the cost of this run, all for a 1M prize that OpenAI said they will not claim. The coolest thing I think we got from this paper, in addition to solving on of the most important and hardest problems in mathematics, is this chart above, where OpenAI shows their unreleased model and how much better it is on Math problems compared to... GPT-6! The “AGI” model we got just a week ago. So so much to look forward to.The drama behind thisI don't want to get into the drama behind this too much, but if you've seen this online, there release wasn't without it's hiccups. Apparently, OpenAI caught wind that an a duo of researchers Tristan Buckmaster (NYU) and Levent Alpöge (Anthropic, in personal capacity), have independently solved the related Euler problem and were about to go public. Apparently, the duo used a mix of GPT and Claude to work on this problem. OpenAI caught wind of this and started working on their own solution on September 1st. A day before OpenAI dropped their release, Buckmaster posted that OpenAI is about to drop it, and in communications with him, they offered him to co-author the paper, only if the Anthropic guy is removed. To which he said no. He then had some claims that maybe OpenAI trained on some of the papers and chats they put into Codex, which OpenAI refuted, while noting “we cannot rule out the possibility that de-identified usage data helped improve the model”. OpenAI also say that the OptOut toggle works, and neither mathematician provided a screenshot that they opted out of the training, so it's hard to say what really went in thereMy take, I don't really care. Two years ago, we told you that reasoning is coming and AI is going to be doing superhuman things, and we finally see the first signs of it, this is a problem that no humans was able to solve for over 60 years! Navie-Stokes probably doesn't change your friday, but there are so many other that they can solve with this approach! Cancer research, room temperature superconductors (remember LK-99? that's also a search problem) and much more. Kudos to OpenAI for this, and I'm looking forward to see the new heights of mathematics. And for the mathematicians who “disagree” with OpenAI solving or not solving this problem, why don't you post your own Lean proofs instead of fuming publicly online? GPT-6 - Not quite AGI, yet?After a week with Astra GPT-6, and Jensen Huang announcing AGI is here, I think we can do a quick recapI and all the co-hosts have been Astra-maxxing for a whole week now and the results are in. This model is incredible at coding, it goes very very deep, however, definitely not AGI quite yet. It has a very jagged frontier, it does some things incredibly well (most demos online are building whole games and apps in 3D and those are mind-blowing) but I started seeing many folks go back to GPT 5.6 Sol etc. I tink some of it also has to do with price, Astra runs faster and is significantly more expensive so token limits for folks are draining fast, but also with controllability, folks are likely still using their old and unoptimized prompts. Speaking of prompts, here's a great writeup from OpenAI how to rework your prompts (which you can send to Astra and have it review your prompts!) for better resultsThe AI doomerism has quite a week (Coxon post, Thread, Hubinger, Marks, Christiano, An Alien Mind)I started ThursdAI with the notion to counter anti-ai and doomerism, and bring positivity to the AI world, so we had to cover this. An Anthropic employee who previously worked at OpenAI, posted on X about leaving Anthropic, saying that both labs are racing towards uncontrollable self-improving superintelligence and that it could be a disaster of the “end all of humanity” kind.His post sits at 130M impressions (after being basically a nobody on X before) and he's been interviewed by Fox, AP, Time magazine, WSJ, NBC and a host of senators, Bernie (chief doomer) included, reposted his post on the same day! The funniest thing is that this post got about 26x more attention than Ilya Sutskever's post about leaving OpenAI. Just nutsWithin four days, seven current and former employees from Anthropic, OpenAI and DeepMind said in public that they believe that AI could kill us all. Also notable that Paul Christiano, who is one of the most interesting AI doomers out there, has joined the OpenAI foundation, and Daniel Kokotajlo, the famed OpenAI whistle-blower, joined Joe Rogans podcast to talk about AI doom. Each one of these incidents in vacuum is normal, but having all these happen in a a span of a few days just feels, inorganic. Some folks are even saying that this is a well coordinated doomerism campaign! I want to be fair to Coxon, folks who worked with him at OpenAI say he's the real deal, and cares deeply about AI safety and humanity, however, he only worked at Anthropic for 6 weeks before publicly leaving, and now every interview he does sayd “Ex Anthropic employee”. Whether it's a coordinated effort or not, it's still a very important discussion, after the pacingthefrontier letter and the HuggingFace hack incident, and it seems to have made waves. Sam Altman just told staff that he's not opposed to pausing and have petitioned the US government to regulate AI as wellLook, I don't disagree that we're dealing with a very powerful technology, however I don't believe that scaring the bajeesus out of everyone is the right way to handle this. Politicians use fear to get votes and get elected, they don't really care about tech progress, and framing this in a way that “we pause or we're dead” ignores all the good that AI is about to do. Cure cancer, find solutions to climate change, helping solving povery. All these seem like out there ideas but they are coming. The US GDP is already growing at an unprecedented rate and a lot of it is due to AI. In any rate, as I said on the show, I'm not against pausing, just after we solve cancer. Then we can pause and reassess, till then, nobody is telling me how China's government is going to pause if US pauses, and if they don't, they will reach superintelligence before we, and I don't want to live in that world! Open Source AIDeepSeek V4.1 Flash: the whale is back, and it's cheap (X, HF, TokenJuice)Speaking of.. chinese AI! Deepsek (The whale) resurfaced this week with V4.1 Flash, and don't let the name fool you, this is not just a .1 small update. 552B with only 8B active on prefill and 16B on decode, 1M context, trained from scratch on 45T multimodal tokens! plus as always, MIT license. Chris from Nvidia joined us to break it down, and his main point stuck with me: every DeepSeek release comes with one of the best engineering reports you can read, and this one is the most data-pilled they've ever done. The paper basically says it out loud, everything else is nice, but it's the data. 45T tokens isn't a huge number anymore, but the cleaning they describe goes way beyond what anyone else publishes (Chris said even his own beloved Nemotron's open pipelines are less thorough). Yam opened a new corner of the show, “I Told You So”, because DeepSeek went back to an encoder-decoder architecture. Not the old one from before GPT-2, this one has a pile of battle tested tricks that make it work at half a trillion parameters. His verdict after testing it all day: the best open weights model you can host for coding right now, and it's not even close to the largest one.This chart is the one to look at. KV cache per token went from 389,000 bytes in the first DeepSeek (Nov 2023) to about 890 bytes now. Nisten did the math live, over 400x smaller. That's why this model is so cheap to serve, and as Yam kept yelling, we shouldn't take it for granted, this is the actual moat and they just put it in the open. Evals, briefly: 90.6 on Terminal-Bench 2.1 (above Opus 5 and GPT 5.6 Sol), 74.2 on DeepSWE 1.1 (also above both), and on an Open Design leaderboard it lands second behind Astra at two cents a task. Not twenty cents. Two. DeepSeek's own evals, so the usual asterisk applies.Two more things. Friend of the pod Aaron Batilo (he works on CoreWeave Inference, this is a side project, not sponsored) put up TokenJuice.ai, free and fast DeepSeek V4.1 Flash in exchange for your requests as training data, hosted in the US. Nobody wants free DeepSeek? Go try it. And Nisten had Astra build a 3D visualization of the whole architecture, every weight a cube sized by its bytes on disk, link in the TL;DR.This Week's Buzz
In this episode, our co-hosts Robby and Tim talk with Modem Co-Founder Ben Vinegar who's building the AI product manager for software teams. He was previously a longtime engineering leader at open source application monitoring software platform Sentry. Ben reflects on his journey from Disqus to nearly a decade at Sentry, where he became the company's first VP of Engineering and worked across product development, SDKs, acquisitions, and incubation. He shares how the rapidly falling cost of software development with AI led him to build Modem, an AI product teammate that helps teams understand what customers are saying and connect that feedback to what they're shipping.The conversation explores why data and context, not just the agent, are the hard parts of building useful AI products, how conversational data creates a new kind of observability, and why these systems are harder to build internally than they appear. He also discusses open source, dogfooding, and how being deeply fluent with AI is changing the way he hires and builds teams.
HTC is now in the AI smart glasses race. Thomas Dexmier, VP of Sales and Marketing at HTC VIVE, came on the Metavertising Podcast to talk about Vive Eagle, the company's first pair. We recorded a few days before they went on sale in Europe.Thomas has been at HTC for more than 15 years. Smartphones, then Vive and VR, now this. He says 2026 feels a lot like 2016, with one big difference. Back then HTC was asking people to step into the technology. This time the technology has to disappear into something people already wear on their face all day.We go through the hardware: 49 grams, Zeiss lenses, a 12MP camera that shoots 3K video, 32GB of storage, 4GB of RAM, four microphones and open-ear speakers. The part I find more interesting is what HTC decided not to do, which is lock you into one AI. Vive AI sits on top as the experience layer, you pick between models like GPT and Gemini, and developers can point the SDK at their own enterprise LLM.Then privacy, which is where every smart glasses conversation ends up right now. The recording LED is enforced in hardware, so covering it up does not get you a secret camera. Sensors know when the glasses come off your face. And if I am recording you and you are not happy about it, you can say "stop recording" out loud and Eagle stops, in English, French, German or Spanish. I had not seen that anywhere else, and it is one of the better answers I have heard to the camera problem.We also talk about why opticians and Silmo Paris matter more to this launch than a shelf in a consumer electronics store, and why Thomas thinks the B2B side may end up bigger than consumer. Underneath all of it sits a question I keep coming back to. The industry has spent 20 years monetising the screen. What happens when the interface is your voice?Vive Eagle went on sale on 3 September and ships from 21 September, starting at 469 euros, or 559 euros with the photochromic Zeiss lenses.In this episode:Why AI smart glasses have to be glasses first and technology secondLaunching VR in 2016 versus launching AI glasses in 2026HTC's case for letting you choose your own AI modelSDK access, enterprise LLMs and the B2B use cases HTC is already seeingHardware enforced privacy and the "stop recording" commandOpticians, telcos and Silmo Paris: selling glasses, not gadgetsFrames, colours and the three Zeiss lens configurationsWhether smart glasses really are the computing platform after the smartphoneFollow Metavertising for conversations with the people building immersive tech and smart glasses.Connect with Thomas DexmierConnect with Ely SantosDetails on the HTC VIVE Eagle: https://www.vive.comAI smart glasses, HTC Vive Eagle, HTC VIVE, smart glasses privacy, Thomas Dexmier, Vive AI, Zeiss smart glasses, AI glasses Europe, GPT glasses, Gemini glasses, wearable AI, spatial computing, XR, enterprise smart glasses, smart glasses launch 2026, Metavertising
Neue Games für alte Systeme? Lieben wir. Vor allem wenn es um Markus' heimliche Liebe, das Neo Geo, geht. Und wie es der Zufall so will, kennen wir da natürlich jemanden, der genau diese Leidenschaft teilt: Sascha Reuter, in der Homebrew-Community auch als Coinfeeder bekannt, hat sein halbes Leben in der Tech-und Startup-Branche verbracht – gestartet in der frühen Internet-Ära. in Hannover, hat es ihn inzwischen nach Australien verschlagen, wo er heute mit seiner Familie lebt und arbeitet. Sein neuestes Projekt: Project Neon! Ein waschechtes Shoot'em Up der alten Schule, das dieser Tage auf allen erdenklichen Plattformen erscheint, obwohl es ja eigentlich exklusiv für das Neo Geo erscheinen sollte ... Let's get (very, very) nerdy!
Topics covered in this episode: OpenAI's Python SDK has migrated to HTTPX2 TMOG - Native Task Manager for macOS, Windows, and Linux wrapture - one wrapper for mocking, tracing, and observability linkedin2md: turn your LinkedIn export into 40+ Markdown files Extras Joke Watch on YouTube About the show Sponsored by us! Support our work through: Our courses at Talk Python Consulting from Six Feet Up Connect with the hosts Michael: Mastodon / BlueSky / X / LinkedIn Calvin: Mastodon / BlueSky / X / LinkedIn Show: Mastodon / BlueSky / X Join us on YouTube at pythonbytes.fm/live to be part of the audience. Usually Tuesday at 7am PT. Older video versions available there too. Finally, if you want an artisanal digest of every week of the show notes in email form? Add your name and email to our friends of the show list, we'll never share it. Calvin #1: OpenAI's Python SDK has migrated to HTTPX2 The OpenAI Python SDK has migrated to HTTPX2, the Pydantic-stewarded fork of httpx. Pydantic picked it up citing "limited activity recently" in the original project, promising "a reliably maintained path forward." If you just use the default client, nothing to do. No code changes. The catch is TLS. Quoting the guide: HTTPX "previously verified certificates against the CA bundle provided by certifi. HTTPX2 instead uses the operating-system trust store, and the SDK no longer installs certifi." That "can break certificate verification in minimal container images without system CA certificates, environments using corporate TLS-inspecting proxies, and deployments that relied on a custom or modified certifi bundle." The fix is SSL_CERT_FILE or SSL_CERT_DIR, or pass your own ssl.SSLContext via verify. Deeper integrations need real edits: custom clients, auth handlers, hooks, and request mocking all take HTTPX2 objects now, and plain httpx is no longer pulled in transitively. So import httpx in your own code means declaring it yourself or moving over. Temporary escape hatch: a legacy HTTPX client Michael #2: TMOG - Native Task Manager for macOS, Windows, and Linux A native, deeply instrumented system monitor for macOS, Windows, and Linux, now in public beta - from Plummers' Software, i.e. Dave Plummer, who wrote the original Windows Task Manager and donated it to Microsoft in 1995. Wikipedia Three real native apps: Swift/AppKit on macOS, Win32 on Windows, C++/Qt 6 on Linux, with a shared C++ core keeping metric semantics aligned - no browser shell anywhere. One dense summary: CPU, clocks, thermals, GPU, memory, storage, network, energy, and the processes responsible for the load, all click-through. Per-core honesty: logical processor and NUMA views, P and E cores color-coded, optional kernel time, 60 FPS live meters. Memory with context: pressure, wired, compressed, cached, committed, available, and swap, plus configurable scrolling history. Processes that act like processes: tree view, filtering, sorting, follow mode, and native verbs including service and launchd control. Phosphor themes: light, dark, green, amber, blue, or mono, with color and saturation you tune yourself. Calvin #3: wrapture - one wrapper for mocking, tracing, and observability Graham Dumpleton, author of wrapt and the original New Relic Python agent, has released wrapture. The name is wrapt plus capture. The core idea: wrap real code instead of replacing it, so the real code still runs while you watch every call. Name a method with wrapture.binding(Class, "method"), open a timeline(), and you get a tape of what actually happened. Real return values, real nesting, arguments normalised against real signatures. tape.tree() prints the call graph as it ran. One mechanism, three jobs: monkey patching with a real lifecycle (apply, remove, suspend, plus returns, raises, transforms_args), unit testing that asserts on real call flow instead of a flat MagicMock call list, and ad-hoc tracing of a running app. The testing pitch is error paths. Inject TimeoutError at the payment gateway, then assert the ledger was never written. Stubs and mocks are strict and spec-required, and there is deliberately no bare Mock(). Tracing needs no code at all. A wrapture.toml naming targets and a sink, run with python -m wrapture main.py, and you get a live call tree with timings. It captures ordinary logging calls as nested events, and with the otel extra it exports spans, metrics and correlated logs with W3C trace ids that join across services. Every line of code and docs was AI-written under their direction, and they say so up front. Two weeks from first commit, eleventh alpha, over 1000 tests, 150+ pages of docs. Alpha on PyPI, needs Python 3.12+ and wrapt 2.4.0+. Michael #4: linkedin2md: turn your LinkedIn export into 40+ Markdown files Via Juan Manuel Daza - a Python CLI that unpacks LinkedIn's data-export ZIP into clean, per-category Markdown you can drop straight into an LLM. One command: linkedin2md Complete_LinkedInDataExport.zip, plus o for output dir, -lang en|es, and -pdf. 40+ output files: profile, experience, education, skills, connections, posts, comments, reactions, recommendations, endorsements, job applications, even ad targeting and LinkedIn's inferences about you. Built for LLM analysis: the README pitches NotebookLM, Claude Projects, Obsidian, and Ollama, with example prompts like "what patterns do you see in my career transitions?" PDF resume mode: -pdf renders an A4 CV via weasyprint, and degrades gracefully to Markdown-only if it isn't installed. Dependency note: "pure Python / zero-dep" holds for the Markdown path only - the PDF path needs weasyprint and markdown installed. Install: pipx install linkedin2md recommended, pip in a venv otherwise - 86% Python, 10 releases, v0.3.1 in May. Agentic dev angle: repo ships opencode config and an N3RV subagent pipeline, including a "judgment day" dual-model adversarial PR review. Extras Calvin: EVE Online Migrates to Python 3 Michael: Dinkus by Will McGugan Joke: Tao of Programming: Book 5 Maintenance
A weekly news show informing you on the latest in Bitcoin, privacy and open source tech, hosted by Ungovernables, Max and Q.AOBWe are now a weekly showTopicsFull screen videoAdmin toolsEnvoy 2.3.3KeyOS 1.4.0 beta3 (public release this week)NEWSCore Lightning: A Critical Vulnerability, a Binary-Only Patch, and a Docker Pipeline That Broke Because of It -- Stacker News / Blockstream Blog / GitHub ReleaseBitcoin's First Quantum-Safe Mainnet Transaction (No Fork Required) -- Bitcoin MagazineStacker News Ships Embedded Spark Wallet and Receive Proxy Toggle -- Stacker NewsSeedSigner Fork Adds Miniscript Support with Liana Bridge -- X/@UnruggableGGRoman Storm Retrial Pushed to April 2027 -- X/@rstormsfCoinbase and Better Mortgage Launch Bitcoin-Backed Mortgages -- Bitcoin MagazineHRF Awards 500 Million Sats to 16 Freedom Tech Projects Worldwide -- Bitcoin MagazineKraken Users Locked Out After 12,000 Dust Transfers From Sanctioned HTX Wallets -- Bitcoin MagazineBitcoin Knots BLAKE2b Hardfork Goes Live as a "Mainnet Trial" -- CryptoSlate / GitHub PR #359 / The Bitcoin Manual explainerPeach Bitcoin Pauses Escrow Model After Swiss Regulator Reversal -- Peach BlogRELEASESFlint v1.0.3 -- 2026-08-28Security audit follow-up for Seth's BTCPay Server plugin. Strips macOS payloads from the build, pins the release path, and restores file permissions after dylib stripping. All artifacts are Sigstore-attested with no maintainer key required.Also in window: Flint v1.0.2 (2026-08-27) -- 23% smaller package, shared reconciliation scans reducing SDK calls for idle stores, cached Spark network status reads. Follows the v1.0.0 milestone which fixed a payment-hash association security issue.Envoy 2.3.3 -- 2026-08-27Point release improving Passport Prime pairing reliability with automatic retry and clock-drift tolerance. Also fixes a visual bug where seed verification puzzle words could overflow at high zoom, and removes a hardcoded Stripe test key.Sparrow Wallet 2.5.4 -- 2026-08-27Significant hardware wallet security hardening: mandatory anti-klepto on BitBox02, Ledger wallet policy re-registration on rejection, and serialised USB device access to prevent enumeration interrupting operations. Also improves BIP129/descriptor import validation and SLIP39 passphrase handling.Eclair v0.14.2 -- 2026-08-26ACINQ released Eclair v0.14.2 with fixes for bugs that "can be exploited by malicious nodes." The release also documents that bitcoind should run on the same machine as eclair or use a secure tunnel. Upgrade is "highly recommended."BTCPay Server v2.4.3 -- 2026-08-24Security release recommended for servers shared with many users. Release notes are minimal, stating only that updating is recommended. Published by NicolasDorier with a GPG-verified commit.BitBox02 Firmware 9.27.0 -- 2026-08-24Displays long transaction and swap amounts in full instead of truncating. Adds message signing for keys in the m/45' namespace and supports timestamp-based absolute locktimes. Stepped upgrade path required from older firmware versions.Blockstream Ships Jade 1.0.41 and Reflects on the Coldcard Fallout -- Blockstream BlogPublished: 2026-08-25Blockstream released Jade firmware 1.0.41 alongside a blog post responding to the Coldcard RNG vulnerability. Jade is not affected (no degraded RNG fallback path, mixes multiple entropy sources through SHA512). Release is security hardening: runtime bumped to latest stable, increased stack protection, updated dependencies, audited sensitive-memory clearing. Not urgent but recommended. 1.0.42 on an expedited timeline for less serious reports. The post reveals dozens of AI-automated scans landed on Jade within days of the Coldcard disclosure, plus multiple human reviews. Credit to Spiral's Loupe tool and the Kvazar agentic harness.Everything ElseArkade v0.9.16 -- Aug 2026Major security hardening and admin web console for the Ark protocol server.BasicSwap DEX v0.18.5 -- Aug 2026Breaking protocol change. All nodes must upgrade.Also in window: v0.18.4 (Aug 2026) -- critical swap security fixes. v0.18.3 (Aug 2026) -- earlier security patch in the same chain.Bisq 1.10.7 -- Aug 2026Hotfix for corrupted DAO datastores. Continues the rapid Bisq 1 patch cycle from last episode (1.10.5 security update, 1.10.6 hotfix).Cashu CDK 0.18.0-rc.1 -- Aug 2026Release candidate with explicit payment confirmation, batch minting, and BOLT12 fix.Cashu TS v5.0.0-rc.8 -- Aug 2026Release candidate with crash-safe swap previews.Fedimint v0.12.0 "Second Nature" -- Aug 2026Named release of the federated ecash protocol.JoinMarket-NG 0.38.0 -- Aug 2026Quantized fees, zero-fee makers, and security hardening. Continues the rapid development from last episode's 0.37.0/0.37.1 releases.LND v0.21.3-beta.rc1 -- Aug 2026Release candidate. Also ships v0.20.4 RC for the older branch.Nunchuk Android 2.8.4 -- Aug 2026Bug fixes for the multisig wallet.Vexl 26.8.0 -- Aug 2026Update to the P2P no-KYC trading app.Wasabi Wallet v2.8.2 -- Aug 2026Coinjoin security patches and reorg sync fixes.EDUCATIONDeep Dive: Bitcoin's Quantum-Resistant Transaction Explained -- Sylvain Saurel / In Bitcoin We TrustPublished: 2026-08-27Detailed technical explainer breaking down StarkWare's first quantum-resistant Bitcoin mainnet transaction (block 964,199). Covers how hash-based cryptography replaces ECDSA exposure, explains Shor's and Grover's algorithms, the mempool attack surface (public keys are exposed between broadcast and confirmation), and why the broader network remains vulnerable despite this milestone.TO DONATE TO ROMAN'S DEFENSE FUND: https://freeromanstorm.com/donateHELP GET SAMOURAI A PARDONSIGN THE PETITION ----> https://www.change.org/p/stand-up-for-freedom-pardon-the-innocent-coders-jailed-for-building-privacy-tools DONATE TO THE FAMILIES w/ USD ----> https://www.givesendgo.com/billandkeonneDONATE TO THE FAMILIES w/ BTC ----> https://pay.zaprite.com/pl_JpxtkLv95T SUPPORT ON SOCIAL MEDIA ---> https://billandkeonne.org/VALUE FOR VALUEThanks for listening you Ungovernable Misfits, we appreciate your continued support and hope you enjoy the shows.You can support this episode using your time, talent or treasure.TIME:- create fountain clips for the show- create a meetup- help boost the signal on social mediaTALENT:- create ungovernable misfit inspired art, animation or music- design or implement some software that can make the podcast better- use whatever talents you have to make a contribution to the show!TREASURE:- BOOST IT OR STREAM SATS on the Podcasting 2.0 apps @ https://podcastapps.com- DONATE via Monero @ https://xmrchat.com/ungovernable- BUY SOME STICKERS @ https://ungovernable.network/shop/FOUNDATIONhttps://foundation.xyz/ungovernableFoundation builds Bitcoin-centric tools that empower you to reclaim your digital sovereignty.As a sovereign computing company, Foundation is the antithesis of today's tech conglomerates. Returning to cypherpunk principles, they build open source technology that “can't be evil”.Thank you Foundation Devices for sponsoring the show!Use code: Ungovernable for $10 off of your purchaseCAKE WALLEThttps://cakewallet.comCake Wallet is an open-source, non-custodial wallet available on Android, iOS, macOS, and Linux.Features:- Built-in Exchange: Swap easily between Bitcoin and Monero.- User-Friendly: Simple interface for all users.Monero Users:- Batch Transactions: Send multiple payments at once.- Faster Syncing: Optimized syncing via specified restore heights- Proxy Support: Enhance privacy with proxy node options.Bitcoin Users:- Coin Control: Manage your transactions effectively.- Silent Payments: Static bitcoin addresses- Batch Transactions: Streamline your payment process.Thank you Cake Wallet for sponsoring the show!MYNYMBOXhttps://mynymbox.ioYour go-to for anonymous server hosting solutions, featuring: virtual private & dedicated servers, domain registration and DNS parking. We don't require any of your personal information, and you can purchase using Bitcoin, Lightning, Monero and many other cryptos.Explore benefits such as No KYC, complete privacy & security, and human support.(00:00:00) INTRO(00:00:58) THANK YOU FOUNDATION(00:01:45) THANK YOU CAKE WALLET(00:02:48) Going Weekly & Four Chainsaws(00:05:13) The New Website & Topic Archive(00:15:24) Envoy &…
Apple, la empresa más hermética del planeta, ha publicado ella misma su catálogo secreto. La Release Candidate de macOS Tahoe 26.7 salió con feature flags sin desactivar y hasta un vídeo promocional de unos AirPods con cámaras usando Visual Intelligence, con voz de Siri incluida. En este episodio hacemos el catálogo completo de lo filtrado, código por código: el hub del hogar J490/J491, el HomePod mini B525, la misteriosa cámara J229, los MacBook Pro con M6, el iPad mini OLED y la familia completa de iPhone 18… incluido el V68, el iPhone Ultra plegable. Diseccionamos el framework AccessorySensorManager que un forero de MacRumors destripó del código: cámaras RGB estéreo de 1 megapíxel, modos de captura activo y pasivo, inferencia de personas en el propio auricular y piloto de privacidad. Explicamos por qué esas cámaras no sirven para hacer fotos… porque no son cámaras: son los ojos de Siri. Y planteamos la pregunta incómoda: ¿de verdad ha sido un despiste? Analizamos la explicación de Mark Gurman sobre el merge de ramas accidental, el giro de guión de los AirPods B790 cancelados frente al B798 de 2027, y por qué una filtración "accidental" que genera la cobertura de una keynote a coste cero es funcionalmente indistinguible del mejor marketing del mundo. Cerramos con un análisis especial de Pebble, el homeOS del hub del hogar: qué sabemos de su interfaz, sus apps y su lanzamiento… y cómo el modo de redimensionado libre del Device Hub de Xcode 27 lleva un año preparando a todo el ecosistema de apps para el plegable y para las pantallas del hogar. Los cimientos se ponen en el SDK años antes de que veas el edificio.
Send us Fan MailOn this episode of Embedded Insiders, Ren Sun, the Regional Head of MCU Marketing and Business Development at GigaDevice, joins the podcast to share the company's latest developments in the global market, specifically with its GD32 MCU product portfolio and opportunities for GD32 within humanoid robotics applications. Next, Rich and Rolf Segger, the Founder of SEGGER, discuss ways to simplify the design of embedded applications, a term the company refers to as EmApps. Users can have SEGGER write the apps for them, or they can get an SDK from SEGGER and do it themselves. The two also dive into TinyML, where it should and should not be deployed, and why. For more information, visit embeddedcomputing.com
Topics covered in this episode: Web UIs for your reverse proxy Wagtail 8.0 is hot off the presses RISC-V is now officially supported by CPython Django's annual releases make every version an LTS Extras Joke Watch on YouTube About the show Sponsored by Logfire from Pydantic: pythonbytes.fm/logfire Connect with the hosts Michael: Mastodon / BlueSky / X / LinkedIn Calvin: Mastodon / BlueSky / X / LinkedIn Show: Mastodon / BlueSky / X Join us on YouTube at pythonbytes.fm/live to be part of the audience. Usually Tuesday at 7am PT. Older video versions available there too. Finally, if you want an artisanal, hand-crafted digest of every week of the show notes in email form? Add your name and email to our friends of the show list, we'll never share it. Michael #1: Web UIs for your reverse proxy Traefik, nginx, and Caddy all sit in front of a lot of self-hosted infrastructure, and all three are configured by hand-editing files. Three active projects put a control plane on top: Traefik Manager (Python + Flask), Nginx UI (Go + Vue), and caddy/ui (React + Node). All three are additive rather than replacements - none of them take ownership of your config away from you - which is the part that matters when the thing has write access to production routing. Traefik Manager is the Python one: Flask 3.1 and Gunicorn for the control plane, a lightweight Go agent for remote instances, currently v1.10.0 with an Android companion app. Nginx UI is a single Go binary at 11.3k stars, with a block-style config editor, an Ace editor doing LLM completion on nginx syntax, and an MCP server so agents can drive it. caddy/ui runs as two containers next to your existing Caddy, reads and writes your Caddyfile directly, and uses Caddy's /adapt API to validate before reload - no Docker socket required. Each one edits the config the underlying server already reads, so your files stay the source of truth and you can drop the UI without unwinding anything. Undo is a first-class feature across all three - timestamped backups with optional Git history, config version compare and restore, Caddyfile snapshots with one-click rollback. Observability is where they diverge: Traefik Manager does CrowdSec and a visual route map, Nginx UI does server metrics, caddy/ui streams access logs over SSE and pulls p50/p95/p99 off Caddy's Prometheus endpoint. Maturity spread is wide - Nginx UI has 11.3k stars, caddy/ui has 4 and was built in a single Claude session - and caddy/ui ships with auth off by default, so set CADDY_UI_USER and JWT_SECRET before it goes anywhere near a public interface. Calvin #2: Wagtail 8.0 is hot off the presses Link: https://github.com/wagtail/wagtail/releases/tag/v8.0 Custom base page models are now supported, so projects aren't locked into subclassing Wagtail's Page as shipped (Matt Westcott). New v3 REST API handles both read and write CMS operations, a first for Wagtail's API. A global registry for permission policies, plus full customizability for the remaining page views via PageViewSet. AVIF and WebP images are no longer auto-converted to PNG by default, a real behavior change to watch on upgrade. Five security fixes: page admin API restrictions, document identification by SHA1 hash, descendant collections in the Documents/Images API, snippet copy permissions, and the page translation endpoint. Formalized Django 6.1 support, and CI now runs on uv with a lockfile. Sponsor: Logfire from Pydantic Your AI agent failed at 2am. Was it the model? A tool call? The database? Most observability tools can't tell you, because they only see part of your stack. Pydantic Logfire sees all of it. One trace across your agents, LLMs, APIs, and database. Down to the infrastructure: services, Kubernetes, and hosts. It's built on OpenTelemetry, with SDKs for Python, TypeScript, and Rust, and it works with any OTel-compatible language. Every prompt, token count, and cost, right next to your vector searches and API calls. You query everything with Postgres-compatible SQL. And so can your coding agent, through the Logfire MCP server. Stop guessing. Read the trace. Pydantic Logfire. AI, it's still just engineering. Visit pythonbytes.fm/logfire today and sign up today. Get 10M records free every month, no card required. You can even click “Onboard with your coding agent” to copy a prompt to have claude or codex integrate Logfire into your app. Thanks to Pydantic for supporting the show. Calvin #3: RISC-V is now officially supported by CPython Link: https://blog.python.org/2026/08/riscv-now-officially-supported/ CPython added RISC-V as a tier 3 platform under PEP 11, specifically the 64-bit Linux target riscv64-unknown-linux-gnu. RISC-V is an open ISA anyone can implement, unlike x86 and ARM, and its market is projected to quadruple by 2032. The RISE Project donated real RISC-V machines for buildbots; the author's work was funded by a Sovereign Tech Agency fellowship. What changes: the port is now a maintained compatibility target, so CPython changes are less likely to quietly break it. What doesn't: no python.org installers, no binary wheel parity for native extensions. Next up: RISC-V runners in CPython CI for pre-merge feedback, then a push toward tier 2, plus architecture-specific optimizations. The ask is testing. If you have RISC-V hardware, build CPython, run your test suite, file what breaks. Tier 3 is the weakest support tier. PEP 11 tier 3 requires a core developer contact and a buildbot, but failures on tier 3 platforms explicitly do not block a release. Saying "ongoing CI/testing expectations" oversells it. The honest bit is "someone is now on the hook for it, and breakage gets noticed," not "it's guaranteed working." Worth the caveat that this is Linux SBCs, not microcontrollers. A VisionFive 2 counts, an ESP32-C6 or Pico 2 does not. Those are 32-bit non-Linux parts where MicroPython is still the answer. Michael #4: Django's annual releases make every version an LTS Starting with Django 2028, Django will move to one January feature release per year, adopt calendar-based version numbers, and support every release for three years. The old distinction between standard and LTS releases disappears, giving teams a predictable annual upgrade path that aligns more closely with Python's own release and support cadence. Every Django release becomes the safe, long-supported choice, so teams no longer need to wait for a specially designated LTS version or absorb two years of changes at once. Each release gets one year of mainstream bug fixes followed by two years of security and data-loss fixes. New releases support the three latest Python versions and add the next Python release during their first year. Calendar versioning begins with Django 2028, followed by Django 2029 and so on. Three Django versions will be supported at any time, giving third-party packages a clearer rolling target. Nothing changes before 2028, and existing commitments for Django 5.2 LTS and 6.2 LTS remain in place. Extras Calvin: The Python docs now document the time complexity of built-in types https://docs.python.org/3.16/library/time-complexity.html Thinking in Python - Bruce Eckel's free book https://thinkinginpython.com/ Michael: prune_uv_pythons.py - Prune uv-managed Python installs, keeping only the newest patch per minor version Runs automatically in my system “upgrade” script: upgrade-output-2026.png Started using Ollama cloud models for my Hermes assistant. Thanks to Jeff Triplett I learned they are not just local models. Joke: The Tao of Programming - Book Seven: Corporate Wisdom
One vendor, SEGGER Microcontroller Systems, can come up with a way to simplify the design of embedded applications, a term they've coined as EmApps. You can have Segger write the apps for you, or you can get an SDK from them and do it yourself. I dove way deeper into this issue with Rolf Segger, the Founder of SEGGER, on this week's Embedded Executives podcast. We get into where TinyML should be deployed and why, and where it shouldn't.
Some crypto products work with multiple chains on different post-quantum paths. NEAR's Illia Polosukhin and Ledger's Charles Guillemet discuss how they manage that challenge. ======================================================== Thank you to our sponsor! Visit 1inch.com to swap tokenized securities, crypto and more. Simple. Secure. Self-custodial. Whatever asset you're buying - swap it at 1inch.com ======================================================== In March, a Google research team published a paper on breaking cryptographic keys with a quantum algorithm, so cautious about the finding that it released only a zero-knowledge proof the algorithm existed. Weeks later, an EigenLayer AI competition improved on that method in roughly 48 hours. Illia Polosukhin, co-founder of NEAR Protocol, and Charles Guillemet, CTO of Ledger, join Laura Shin for an update on the quantum threat whose deadline could be approaching fast. Both are creating products that deal with multiple chains that all have different post-quantum approaches. They discuss why, of the three NIST-standardized, post-quantum algorithms, the crypto industry has splintered into different chains working with different ones, whereas most industries are converging on one, called lattice-based. They also debate what to do with Satoshi Nakamoto's bitcoins: do nothing, freeze them, or freeze and tail-emit new bitcoin, an option Guillemet favors even though Bitcoin's leaderless governance makes consensus hard to reach. Host: Laura Shin, Host / Unchained Guests: Illia Polosukhin - Co-founder of NEAR Protocol Charles Guillemet - CTO of Ledger Timestamps
Q2 Innovation Studio just marked its fifth anniversary, with more than 90% of Q2 Digital Banking Platform customers now using its SDK and partner ecosystem. Johnny Ola, SVP of Q2 Innovation Studio, shares the origin story, explains how Q2 vets and works with partners, highlights the use cases banks and credit unions are prioritizing today, and previews what's next as Q2 Code opens up new ways to build on the platform. Related Links [News Release] Q2 Innovation Studio Marks Five Years [Webpage] Q2 Innovation Studio [Blog] A Closer Look at Q2 Code [LinkedIn] Johnny Ola
“AI companies have incredible convenience and are working on trust. Credit unions have trust and are working on convenience. The only question is who gets there first.” – Siva NarendraThank you for tuning in to The CUInsight Network, with your host, Robbie Young, Vice President of Strategic Growth at CUInsight. In The CUInsight Network, we take a deeper dive with the thought leaders who support the credit union community. We discuss issues and challenges facing credit unions and identify best practices to learn and grow together.My guest on today's show is Siva Narendra, CEO and Co-Founder of Tyfone who returns to the podcast after having previously been featured in episode 100. This time, he joins me to discuss how AI is changing financial services at an alarming rate and what that means for credit unions. Tune in as Siva shares why he believes trust remains credit unions' greatest advantage but also explains why that advantage won't last forever unless institutions rethink how they approach it. We talk about why simply adding AI-powered features isn't enough, how credit unions can remain central to members' financial lives, and why the race between trust and convenience is accelerating faster than many people even realize.Throughout our conversation, Siva and I also explore why he regards lending as one of the strongest long-term opportunities in the AI era, the growing importance of identity verification, and the security challenges that come with increasingly sophisticated AI tools. Along the way, he offers some practical insights into how Tyfone is building technology designed to help credit unions embrace innovation without compromising the relationships that have always set them apart.As we wrap up the conversation, Siva discusses the inventor who inspires him most, the surprisingly indispensable purchase that's become part of his daily routine, the podcasts he relies on to stay informed, and why gardening (and a daily commute with his wife) continues to help him maintain balance outside of work. Enjoy this conversation with Siva Narendra!Connect with Siva:Siva Narendra, CEO and Co-Founder of Tyfonetyfone.com Siva: LinkedInTyfone: LinkedIn | Facebook | Instagram | YouTube | Spotify (DIgital Panking Podcast) | XShow notes from this episode:Shout-out: AT&TShout-out: VerizonShout-out: NetflixTopic mentioned: OmnichannelShout-out: ChatGPTShout-out: AppleShout-out: Abraham LincolnShout-out: KeurigPodcast mentioned: Planet MoneyPodcast mentioned: The Indicator from Planet MoneyShout-out: Greg MichligTopic mentioned: The HolocaustShout-out: Siva's in-lawsShout-out: Siva's wifePrevious guests mentioned in this episode: Siva G. Narendra (episode #100)In this episode:[1:17] - Hear how digital banking is Tyfone's foundation for continuous innovation![2:56] - Siva argues that credit unions need to close the personalization gap while holding onto their advantage in member trust.[5:58] - Siva explains how SDK-based customization can leave institutions isolated, while frontier AI offers more personalized financial insights.[8:10] - Generative AI requires strong authentication, data controls, observability, and guardrails to remain secure and relevant![11:10] - Hear how lending still gives financial institutions an edge because of regulation, risk, institutional knowledge, and physical branches.[14:41] - As AI makes identities easier to fake, verifying people in person becomes more important than ever.[16:25] - Hear how Tyfone is working on authentication to help maintain trust as AI makes fraud easier.[18:22] - Hear why Siva admires Abraham Lincoln.[18:52] - Siva reveals what he loves about his Keurig coffee machine.[19:44] - Find out what podcasts Siva loves.[21:49] - Siva reminds us that physical branches, underwriting judgment, and loan officers remain irreplaceable assets that AI can't replicate.
Topics covered in this episode: Python 3.12.14, 3.11.16, 3.10.21 - security releases Codeberg's AI-code ban tests its role as a GitHub alternative Brett Cannon: what's missing for reproducible builds on PyPI nothing records the source code a distribution came from. direct_url.json captures it when you install from a repo or archive, so the fix is putting the same info in sdist/wheel metadata. recording the build tools. Wheels can already do this via PEP 770 SBOMs in .dist-info/sboms/ - sdists can't, since they're a tarball plus a precalculated PKG-INFO with nowhere to hang extra metadata. Either "don't use sdists" or an sdist v2. Extra extra extra, hear all about it Extras Joke Watch on YouTube Sponsored by Logfire from Pydantic pythonbytes.fm/logfire This episode is brought to you by Pydantic Logfire. It's observability for AI apps from the team behind Pydantic - agents, LLMs, APIs, database, and infrastructure in a single trace, queried with Postgres-compatible SQL. Your coding agent can query it too, through their MCP server. I'll tell you more later. Connect with the hosts Michael: Mastodon / BlueSky / X / LinkedIn Calvin: Mastodon / BlueSky / X / LinkedIn Show: Mastodon / BlueSky / X Join us on YouTube at pythonbytes.fm/live to be part of the audience. Usually Tuesday at 7am PT. Older video versions available there too. Finally, if you want an artisanal, hand-crafted digest of every week of the show notes in email form? Add your name and email to our friends of the show list, we'll never share it. Calvin #1: Python 3.12.14, 3.11.16, 3.10.21 - security releases https://blog.python.org/2026/08/python-31214-31116-31021/ Source-only security releases for the three branches now in security-fix-only mode; release team blamed the European solar eclipse for the timing. tarfile hardening. Multiple path-traversal bypasses of the data filter closed, including a symlink escape that bypassed the CVE-2025-4330 fix; extract() now applies the filter to link targets too. Four fresh CVEs: CVE-2026-2297 (SourcelessFileLoader not using io.open_code() for .pyc), CVE-2026-4224 (expat crash on deeply nested content models), CVE-2026-3644 (control chars in http.cookies.Morsel), plus the completed CVE-2021-4189 fix in ftplib.ftpcp. Quadratic-complexity DoS cleanup across the stdlib: HTMLParser, configparser regexes, unicodedata.normalize(), csv.Sniffer.sniff(), and ElementTree XPath index predicates. Header/injection fixes: CR/LF rejected in HTTPConnection.set_tunnel(), control chars blocked in wsgiref.handlers status, and webbrowser now rejects leading dashes (plus a %action prefix bypass). http.client now caps chunked trailer lines and 1xx interim responses at 100 each - a hostile server could previously hang the client forever despite a socket timeout. Memory-safety odds and ends: stale pointers in lzma/bz2/zlib decompressors after MemoryError, a bz2 stack overflow on reuse-after-error, and bundled libexpat bumped to 2.8.3. If you're still on 3.10, 3.11, or 3.12 - and you extract tarballs from anywhere you don't fully control - this one's not optional. Michael #2: Codeberg's AI-code ban tests its role as a GitHub alternative Armin's article “Codeberg Divides” Armin Ronacher argues that Codeberg's new terms, which prohibit projects mostly written with generative AI, create a vague and difficult-to-enforce boundary. His larger concern is that a democratically governed host can still be unpredictable or ideologically narrow, weakening Codeberg's potential as a broad European alternative to GitHub. The strongest question for Python developers is whether repository hosting should judge legal open source by how code was produced, or focus on behavior and resource abuse. “Mostly generated” is hard to measure in modern codebases where developers mix handwritten code, completions, agents, and generated refactors. Ronacher suggests clearer alternatives: ban all LLM involvement, or target autonomous repository spam, abusive resource use, and low-quality generated contributions directly. Codeberg is free to choose a values-driven community, but that may conflict with being predictable, neutral infrastructure and a serious GitHub competitor. Worth discussing: can open-source communities set meaningful AI boundaries without driving maintainers and projects into opposing camps? Very first search for these terms lands on this page. Codeberg looked like a viable alternative. … Unfortunately, the latest update to its terms of service seems to mark a first step in changing one part I moved there for, namely the “freedom” part. Sponsor: Logfire from Pydantic Your AI agent failed at 2am. Was it the model? A tool call? The database? Most observability tools can't tell you, because they only see part of your stack. Pydantic Logfire sees all of it. One trace across your agents, LLMs, APIs, and database. Down to the infrastructure: services, Kubernetes, and hosts. It's built on OpenTelemetry, with SDKs for Python, TypeScript, and Rust, and it works with any OTel-compatible language. Every prompt, token count, and cost, right next to your vector searches and API calls. You query everything with Postgres-compatible SQL. And so can your coding agent, through the Logfire MCP server. Stop guessing. Read the trace. Pydantic Logfire. AI, it's still just engineering. Visit pythonbytes.fm/logfire today and sign up today. Get 10M records free every month, no card required. You can even click “Onboard with your coding agent” to copy a prompt to have claude or codex integrate Logfire into your app. Thanks to Pydantic for supporting the show. Calvin #3: Brett Cannon: what's missing for reproducible builds on PyPI Framing came out of his 2026 Python Packaging Council nomination - the secure-supply-chain gap he found is that Python has no defined way to do reproducible builds at all. Design goal is zero friction: producers uploading to PyPI shouldn't have to do anything. The work lands on build backends and installers. Gap #1: nothing records the source code a distribution came from. direct_url.json captures it when you install from a repo or archive, so the fix is putting the same info in sdist/wheel metadata. Gap #2: recording the build tools. Wheels can already do this via PEP 770 SBOMs in .dist-info/sboms/ - sdists can't, since they're a tarball plus a precalculated PKG-INFO with nowhere to hang extra metadata. Either "don't use sdists" or an sdist v2. The replay mechanism already exists: [build-system] in pyproject.toml is a defined entry point, so if backends recorded their own environment, you could reinstall and re-run the build. Payoff idea: trusted third parties report successful reproductions back to PyPI, which displays "independently reproduced by X" - surfaced in the index API so installers could prefer reproduced files. Explicitly framed as a perk, not a requirement - roughly SLSA build level 1, no shaming projects that don't opt in. Verbal kicker option: "And don't think pure-Python wheels are off the hook. Something built that wheel, and if that something was compromised, so is your wheel. SolarWinds was a build-process attack." Michael #4: Extra extra extra, hear all about it Python 3.14.7 Upgraded the MCP servers to 2026-07-28 v2 protocols (talk python, python bytes) Got agentsview running synced via postgres Talk Python courses, teams trial offering Talk Python courses, government procurement offering Lean TDD audio book is out Extras Calvin: uv now prefers post-quantum key exchange - https://github.com/astral-sh/uv/releases/tag/0.12.4 Joke: Beware of dog
A Developer Advocate for Deepgram with over ten years of experience in advocacy, Ed Charbeneau focuses on fostering engagement with the developer community while supporting the development of innovative AI-Native applications. With a diverse technical skill set, Ed contributes to building impactful solutions using a wide range of technologies including: C#, Asp.Net, HTML/CSS, JavaScript, Microsoft Extensions AI, RAG, and ML.NET technologies. Ed leverages his role as a developer advocate to find new and innovative ways to expand the boundaries of product through customer feedback, research and hands-on development. Ed has built proof of concept frameworks, components, and SDKs that have gone on to production. These initiatives have future-proofed the Progress portfolio, empowered customers, and enabled product integration. Ed is passionate about empowering developers to explore opportunities in full-stack development, AI, machine learning, and prompt engineering. Ed facilitates discussions on emerging technology trends, fostering collaboration and knowledge-sharing within the tech community.You can find Ed on the following sites:WebsiteXGitHubMastodonTwitchYouTubeHere are some links provided by Ed:Deepgram PLEASE SUBSCRIBE TO THE PODCASTSpotifyApple PodcastsYouTube MusicAmazon MusicRSS FeedYou can check out more episodes of Coffee and Open Source on https://www.coffeeandopensource.comCoffee and Open Source is hosted by Isaac Levin
Discover more Sincerely Accra! Joseph sits with Miss Enny, Kojo Junior and SDK to get some answers! Is the Ghanaian influencer space worth the ups and downs? Who has more earning power?Music Opening Oshe - Reynolds The Gentleman ft. Fra! Music Bridges Fa Be Bom (Tatata) - Sefa x King Paluta Soko in Soko - MANAMEPEE & ShakurTYBAba Jane - MANAMEPEE R2Bees - Oseikrom Sikanii & AlorGYesu - Nana Yaw Ofori-Atta Music Closer All My Love - StoneBwoy ft. EfyaA GCR Production - Africa's Premiere Podcast Network
Lightning Labs built Wavelength to deliver self-custodial Lightning payments without forcing users to run nodes, manage channels, or handle liquidity.Olaoluwa Osuntokun, CTO and co-founder of Lightning Labs, and Michael Levin, VP of Product, join me to detail the design choices behind their Ark implementation.They explain why Ark was selected, how hop hints let every payment use ordinary Lightning invoices, and why the four-endpoint SDK targets AI agents and vibe coders. The conversation covers sub-dust vouchers, unilateral exits, offline payment delivery, and one-basis-point alpha pricing.Wavelength shows that self-custodial Lightning can match the integration ease of custodial services while preserving Bitcoin sovereignty.Timestamps01:28 — Why Wavelength: Lightning Without Node Pain03:00 — Why Ark? 05:28 — No New Addresses: Just Lightning Invoices08:40 — Targeting Vibe Coders & AI Agents13:06 — Lightning Beats Credit for LLM APIs17:51 — Receive Sub-1k Sat Vouchers Seamlessly19:32 — Normal User Spins Up Ark Wallet Fast23:02 — Build Wallets with Just 4 Endpoints24:46 — Telegram Self-Custodial Wallets Already Live26:21 — Offline Payments Still Arrive Automatically29:17 — Drop-In SDK for iOS and Android31:26 — Can Servers Steal Your Funds?33:36 — Wavelength Alpha: Just 1 Bip Fees37:57 — AI Attacks Targeting Bitcoin Services?44:05 — Bug Bounties Shift to Token Spending48:36 — Self-Custody as Easy as CustodialLinks: https://x.com/roasbeehttps://x.com/MichaelLevinWavelength Announcement: https://x.com/lightning/status/2079620936567779707Stephan Livera links:Follow me on X: @stephanliveraSubscribe to the podcastSubscribe to Substack
Topics covered in this episode: Claude Code /insights Post-quantum crypto lands in Python MCP goes stateless — and FastMCP gets renamed inshellisense - IDE style command line auto complete Extras Joke Watch on YouTube About the show Sponsored by Xweather Xweather combines enterprise-grade weather intelligence with agent-ready APIs, natural language capabilities, and an MCP server so your agents can adapt workflows, automate responses, and make better decisions based on real-world conditions. Michael will tell you more about them later in the show. Get started for free at pythonbytes.fm/xweather Connect with the hosts Michael: Mastodon / BlueSky / X / LinkedIn Calvin: Mastodon / BlueSky / X / LinkedIn Show: Mastodon / BlueSky / X Join us on YouTube at pythonbytes.fm/live to be part of the audience. Usually Tuesday at 7am PT. Older video versions available there too. Finally, if you want an artisanal digest of every week of the show notes in email form? Add your name and email to our friends of the show list, we'll never share it. Michael #1: Claude Code /insights Michael's Insights: michael-kennedy-claude-code-insights-2026-08-09.html Be careful sharing these outputs, they include details references to your projects, errors, security findings, etc. ;) /insights reads your last 30 days of local session transcripts and hands back an interactive HTML report on how you actually work. One command, zero setup: type /insights in a session, or run claude -p "/insights" from the shell for a non-interactive version that just prints the path Reads what's already on disk: pulls session logs from ~/.claude/projects/, skipping agent sub-sessions and anything under 2 messages or 1 minute Project areas: clusters your sessions into themes like "CLI Tooling" or "Documentation" with session counts Friction analysis: categorizes where things went wrong by root cause - and quotes your own prompts back at you Interaction style: tells you whether you're a delegator or a micromanager, plus which workflows are worth doubling down on Actually actionable: suggests concrete CLAUDE.md additions and Claude Code features you're not using The catch: Haiku does the per-session classification, so the first run takes several minutes; results cache to ~/.claude/usage-data/facets/ and the report lands at ~/.claude/usage-data/report.html Calvin #2: Post-quantum crypto lands in Python pyca/cryptography 48 ships ML-KEM (key establishment) and ML-DSA (signatures) — NIST's post-quantum standards, now one pip install away. Big deal because it's the 11th most-downloaded package on PyPI (~1.2B downloads/month) and sits under Ansible, Certbot, Airflow, and paramiko. No PQ there, no PQ anywhere in Python. Trail of Bits did the work (Rust bindings, cross-backend API, tests, AWS-LC backend support), funded by the Sovereign Tech Agency. Timing tracks a June 22 White House order setting federal deadlines: PQ key establishment by end of 2030, PQ signatures by end of 2031. Not a drop-in swap — the wire sizes explode. ML-DSA-65 signatures are 3,309 bytes vs Ed25519's 64; ML-KEM-768 public keys are 1,184 bytes vs X25519's 32. Hardcoded field sizes and length prefixes will bite. API looks like the existing asymmetric primitives, except ML-KEM is encapsulate/decapsulate rather than a Diffie-Hellman exchange. SLH-DSA (the hash-based conservative backstop) is still in progress. The primitives are here, but protocols haven't caught up — so you won't be running post-quantum Certbot this week. Sponsor: Xweather You're using agents that can write code, summarize documents, and automate workflows. But they're missing one thing: awareness of the world around them. This is where today's sponsor, Xweather comes in. Xweather combines enterprise-grade weather intelligence with agent-ready APIs, natural language capabilities, and an MCP server built for tools like Claude, Codex, Copilot, and modern IDEs – so your agents can adapt workflows, automate responses, and make better decisions based on real-world conditions. Backed by Vaisala, whose instruments fly on NASA missions to Mars, Xweather delivers trusted data and unique insights that go beyond conditions to actual impact – from real-time lightning strikes to road surface forecasts. Start with 15,000 free API calls each month and pay only for what you use as you grow. Xweather is your full weather stack, for developers by developers. Start building for free today at pythonbytes.fm/xweather. The link is in your podcast player's show notes and on the episode page. Thanks so much to Xweather for supporting Python Bytes. Calvin #3: MCP goes stateless — and FastMCP gets renamed From Philipp Acsany over at Real Python The 2026-07-28 spec landed July 28 and the Python SDK shipped 2.0.0 the same day. Biggest rewrite since MCP launched, and it's breaking on purpose. Context for scale: the Tier 1 SDKs are pulling close to half a billion downloads a month, with TypeScript and Python each past a billion total. The headline is the stateless core. The initialize/initialized handshake and the Mcp-Session-Id header are both retired — protocol version, client identity, and capabilities now ride in _meta on every request, with an optional server/discover RPC if a client wants capabilities up front. Any request can land on any instance behind plain round-robin, no shared storage. Server-initiated calls are the hard part of the migration. Sampling, elicitation, and roots/list no longer call back to the client; instead the server returns resultType: "input_required" and the client retries with inputResponses attached. Multi Round-Trip Requests, MRTR. Also: Mcp-Method and Mcp-Name are now required headers so gateways route on headers instead of cracking JSON bodies, and missing-resource errors move to standard 32602. Deprecation sweep with an actual policy behind it — Roots, Sampling, Logging, and the legacy HTTP+SSE transport all deprecated with a twelve-month minimum offramp. Tasks graduated out of the experimental core into a real extension, which is what the formalized extensions framework was for. MCP Apps is now an official extension too, so a tool call can return sandboxed interactive HTML. Auth picked up RFC 9207 issuer validation, issuer-bound credentials, and a shift from DCR toward CIMD. Python SDK 2.0 is where it gets personal: FastMCP is now MCPServer, no alias, no shim. McpError → MCPError. Wire types went snake_case (is_error, input_schema) and moved to a standalone mcp_types package, with mcp.types kept as a permanent alias. One Client object replaces the old transport + ClientSession + initialize() stack. httpx became httpx2. Sync handlers run on worker threads now, so asyncio.get_running_loop() raises inside them. The good news: one MCPServer serves both protocol eras, so 2025-era clients keep working with nothing to configure, and a Resolve(fn) parameter lets one tool body cover MRTR and the old path. 1.x is maintenance-and-security-fixes only — pin mcp>=1.28,
In "How Nordian's Platform Enables Long-Haul Autonomy", Joe Lynch speaks with Co-founder and CEO of Nordian, Michael Schramm, about how Nordian enables long-haul autonomy by combining precise positioning, satellite connectivity, and edge intelligence into a single platform. About Michael Schramm Michael Schramm is Co-founder and CEO of Nordian, the positioning and connectivity platform for Physical AI, delivering centimeter-level GNSS corrections, satellite connectivity, and fleet lifecycle management to some of the largest industrial OEMs in the Americas. A serial entrepreneur with 15+ years of executive leadership, he is a founding partner of Ambush, an applied AI engineering firm; co-founder of Echo54, an advanced sensing R&D company serving US and allied government agencies; and founder of GOAT, an acquired consumer micro-mobility company. Across nearly two decades of building companies that operate in the physical world, he kept running into the same failure point: machines break when positioning and connectivity aren't engineered as one system. Nordian exists to fix that. About Nordian Nordian is the positioning and connectivity platform for Physical AI. One platform delivers centimeter-level GNSS corrections, integrated satellite connectivity, and fleet lifecycle management to industrial OEMs across transportation, agriculture, and mining. Headquartered in Austin, Texas, and deliberately launched in the hardest environments on Earth, Nordian built South America's largest PPP-RTK network, and its platform serves 80% of the region's 20 largest agricultural OEMs. Proven where networks fail and machines can't, Nordian is now expanding globally to power autonomous operations at scale. Key Takeaways: How Nordian's Platform Enables Long-Haul Autonomy In "How Nordian's Platform Enables Long-Haul Autonomy", Joe Lynch speaks with Co-founder and CEO of Nordian, Michael Schramm, about how Nordian enables long-haul autonomy by combining precise positioning, satellite connectivity, and edge intelligence into a single platform. Physical AI is a Connectivity and Processing Challenge, Not an AI Model Problem: Current AI systems are fully capable of handling autonomous navigation, but real-world physical AI is constrained by connectivity and real-time processing capabilities. Offloading critical decisions to back-end cloud servers introduces latency, which is dangerous for heavy machinery like a 25-ton autonomous truck moving at highway speeds. Edge Computing and Local Inference Eliminate Deadly Latency: To operate safely without reliance on uninterrupted internet access, 100% of mission-critical decisions must occur directly on the device using edge computing. Nordian provides the local processing capacity needed for real-time inference, allowing autonomous vehicles, drones, and heavy equipment to operate safely in "air-gapped" environments or during brief network dropouts. Centimeter-Level Positioning Replaces Imprecise Traditional GPS: Standard GPS provides meter-level accuracy, which is acceptable for route navigation but unacceptable for vehicle control, lane-level autonomous driving, precise geofencing, or row-crop agriculture. By combining satellite signals with dedicated ground reference stations to calculate real-time differential corrections, Nordian achieves centimeter-level accuracy required for absolute control. Integration Burden is the Primary Bottleneck for OEMs: Equipment manufacturers historically acted as their own integrators—trying to bolt together separate vendors for chipsets, satellite bands, cellular modems, and edge computing. Nordian abstracts this complexity by unifying precise positioning, resilient connectivity, and edge intelligence into a single plug-and-play factory-installed package with an SDK for custom software development. A Multi-Band "N+3" Connectivity Model Bridges the Connectivity Gap: Autonomy dies where cellular coverage fails, particularly across the 71% of U.S. roadways located in rural environments. Nordian solves the connectivity gap by layering cellular networks, L-band communications, and low Earth orbit (LEO) satellite constellations (including integrated Starlink connectivity) into a redundant system capable of rapid sub-10-second signal convergence. Agricultural Battle-Testing Translates Directly to Transportation and Logistics: Before expanding into long-haul trucking and yard logistics, Nordian proved its system in South America's harsh agricultural environments, building a massive reference station network across Brazil and Argentina. This foundation enabled them to capture 80% of the top 20 agricultural OEMs in the region—proving the technology where infrastructure is non-existent and atmospheric interference (scintillation) is severe. Autonomous Technology Target Long-Haul Workloads to Improve Quality of Life: Autonomous technology is positioned to address structural labor shortages by replacing high-turnover, long-haul routes (where drivers are away from home for weeks) with fully autonomous systems or human-augmented modes. This shifts human operators toward last-mile and short-haul jobs, improving driver safety, operational utilization, and overall work-life balance. Learn More About How Nordian's Platform Enables Long-Haul Autonomy Michael Schramm | Linkedin Nordian | Linkedin Nordian Contact Nordian Nordian Authorized to Resell Starlink High-Speed Internet to Businesses & Enterprises. Nordian Expands High-Precision GNSS Positioning to Brazil Through Strategic Partnership with u-blox Federal News Network's Space Hour Podcast - Connecting devices out in remote regions Fierce Network - Nordian authorized to resell Starlink internet to businesses and enterprises Ending the 60% Waste: The Radical Shift Trucking Needs Right Now The Logistics of Logistics Podcast If you enjoy the podcast, please leave a positive review, subscribe, and share it with your friends and colleagues. The Logistics of Logistics Podcast: Google, Apple, Castbox, Spotify, Stitcher, PlayerFM, Tunein, Podbean, Owltail, Libsyn, Overcast Check out The Logistics of Logistics on Youtube
#362: Feature flags or canary deployments - do you need both? Viktor puts it to Alex Casalboni from Unleash, who says he argues about this with his colleagues roughly every day, and the answer lands clean. Switching a hostname, a database, an API vendor? That is infrastructure, nothing to do with who the user is, so keep your canaries and your blue-green. But a canary switches one thing at a time. Try running three A/B tests and ten behavioral changes through it and the whole approach buckles. Anything that needs to know who the user is belongs in a flag. Different layers of the stack, different tools, and most teams will end up with both whether they planned to or not. Back up, though, because there is a new word attached to all of this. FeatureOps. There is a manifesto and everything, sitting at [featureops.io](https://featureops.io/), reading a lot like someone nailed 95 theses about feature flags to a door. Real discipline, or marketing wrapper? Alex gets about ten seconds of pleasantries before he has to answer for the word. His defense is narrower than the name suggests, and better for it: every ops discipline we have gets you to the deployment and then waves goodbye. Something breaks, you go around the whole loop again - hotfix, pipeline, 20 or 30 or 60 minutes, fingers crossed. FeatureOps is the claim that the same principles apply after the code is already running. Runtime control. Alex says enterprise customers routinely have a 12 to 24-hour round trip between finding a problem and getting the fix live. Even for a hotfix. Viktor is not letting the seconds claim through unchallenged. If it takes you a day to notice and two seconds to flip, that is a day and two seconds - so stop measuring from the convenient starting line. Alex concedes the framing and then goes somewhere better with it: the bottleneck was never the clicking. It is the humans and the bureaucracy in between. Which is why Unleash is pushing impact metrics, where the SDK sends error rates back and the system kills the feature itself, no human in the loop. Then Darin calls BS on immutable event log, because there is no such thing as immutable data, and Alex takes the hit cleanly - fair, it is append-only with locked-down keys, not magic. Nobody puts this part on a landing page. Flag evaluation has an input, not just a true/false output, and that input is user context - which means an external API call is not just latency, it is your PII leaving the perimeter. A compliance problem hiding inside a performance decision. And the flag graveyard is worse than you think: companies create roughly ten flags for every one they clean up, and Alex has a customer whose oldest flag dates to 2012. His fix is an MCP server that opens the cleanup PR for you when you mark a release complete. Best line of the day, on whether flags complicate your code: everything complicates your code, and the best way to not complicate your code is to not code. Alex's contact information: LinkedIn: https://www.linkedin.com/in/alexcasalboni/ X: https://x.com/alex_casalboni YouTube channel: https://youtube.com/devopsparadox Review the podcast on Apple Podcasts: https://www.devopsparadox.com/review-podcast/ Slack: https://www.devopsparadox.com/slack/ Connect with us at: https://www.devopsparadox.com/contact/
PEBCAK Podcast: Information Security News by Some All Around Good People
Welcome to this week's episode of the PEBCAK Podcast! We've got four amazing stories this week so sit back, relax, and keep being awesome! Be sure to stick around for our Dad Joke of the Week. (DJOW) Follow us on Instagram @pebcakpodcast Please share this podcast with someone you know! It helps us grow the podcast and we really appreciate it! Simple 6 signup link https://simple6.co/r/CFUR98 Bing malvertising campaign tricks users into downloading a fake Claude desktop app that quietly installs the SectopRAT info-stealing trojan. - https://www.bleepingcomputer.com/news/security/fake-claude-app-promoted-by-bing-ads-pushes-sectoprat-malware/ The malicious "FakeAgent" campaign hit at least 29 organizations on July 21–22; the poisoned Claude Artifact (hosted on Anthropic's own domain) was downloaded 7,100 times before removal, with the fake installer (ClaudeDesktop.exe) sideloading a malicious DLL to drop SectopRAT — a HVNC-capable info-stealer active since 2019 that targets browser logins, crypto wallets, Discord/Telegram/Steam credentials, and uses Ethereum smart-contract transactions (EtherHiding) to fetch its C2 address; researchers even used Claude Opus 4.8 themselves to help reverse-engineer the payload. Iran allegedly exploited decades-old cell network flaws to physically locate and target US troops during the Iran War. - https://techcrunch.com/2026/07/14/iran-abused-mobile-networks-vulnerabilities-to-locate-u-s-military-in-the-middle-east-report-says/ Per a Financial Times report citing the Mobile Surveillance Monitor and government officials, Iran exploited SS7 — the legacy signaling protocol still underpinning 2G/3G global roaming — to track US personnel at bases and hotels in Iraq, Bahrain, and elsewhere in the Middle East, contributing to strikes that wounded upwards of 150 US troops; Iran reportedly also abused ad-tech location data as a secondary tracking vector. A multi-university study found dozens of apps marketed directly to US troops are quietly shipping Chinese and Russian code. - https://www.wired.com/story/apps-marketed-to-us-troops-are-shipping-chinese-and-russian-code/ Researchers from Purdue, West Point, and Florida International University analyzed 220+ apps aimed at service members (fitness trackers, base-living-condition raters, National Guard-affiliated apps) and found 64% contain third-party SDKs from foreign countries, with roughly 1-in-8 to 1-in-14 apps (reporting varies) carrying code tied directly to China or Russia — including at least 12 apps embedding Huawei's mobile framework and others using the Russian ad service Yandex; no active exfiltration was observed, but researchers warn the dormant SDK code is an exploitable backchannel. A 21-year-old allegedly stole $220K in crypto by hiding malware in Steam games — and got caught because he spent it on Uber Eats. - https://www.pcmag.com/news/fbi-traces-malware-infected-steam-games-to-21-year-old-in-florida The FBI arrested Florida's Zyaire Dontaevious Zamarion Wilkins for allegedly running eight malware-laced Steam games (including BlockBlasters and PirateFi) between May 2024–Feb 2026, infecting ~8,000 devices and draining ~80 crypto wallets for at least $220,000 — including $35,000 stolen from a streamer's cancer-treatment fundraiser; investigators cracked the case by tracing stolen Bitcoin to 150+ Bitrefill gift cards mostly spent on Uber Eats orders tied to his home and university email address. Thrillist crowned Doritos Nacho Cheese the single greatest snack of all time, edging out Oreos and Pringles for the top spot. - https://www.thrillist.com/eat/nation/best-snack-foods-chips-candy-ranking The top five, in order: Doritos (Nacho Cheese, specifically) at #1, Oreos (Double Stuf gets the nod) at #2, Pringles at #3, Reese's Peanut Butter Cups at #4, and Goldfish rounding out the top five; other notable placements include Cheez-Its at #8, M&Ms at #7, Cheetos (the curls, not puffs) at #6, and Lay's Original topping the chip-specific competition at #12. Dad Joke of the Week (DJOW) Find the hosts on LinkedIn: Chris - https://www.linkedin.com/in/chlouie/ Brian - https://www.linkedin.com/in/briandeitch-sase/ Glenn - https://www.linkedin.com/in/glennmedina/ Ben - https://www.linkedin.com/in/benjamincorll/
A bi-weekly news show informing you on the latest in Bitcoin, privacy and open source tech hosted by Ungovernables, Max and Q. AOBFreedom.Tech launch reminderKeyOS v1.3 now publicly availableNEWSIndia orders GitHub to take down BitChat's source code; Internet Freedom Foundation calls it unconstitutional - TFTC: India BitChat GitHub takedown, I4C, IFF / CoinDeskFourth Circuit says border agents can hand-search your phone with zero suspicion, as a man is prosecuted for a duress-wipe - EFF: Fourth Circuit says border agents can search your phone by hand, no suspicion required / TechCrunch: US accuses American of wiping his phone with a duress password at the borderSenate Democrats kill the CLARITY Act before recess; the developer safe harbor (Section 604) stalls with it - TFTC: CLARITY Act rejected, Bitcoin ownership surpasses goldState Department launches a "Freedom Tech" program with BPI, Palantir, and Anduril as founding partners - Bitcoin Magazine: State Department tech program with BitcoinBlock open-sources Buzz: a Nostr-native, keypair-identity workspace for humans and AI agents - LINKBIP-110 approaches its mandatory signaling window with support under 1%, and enforcing nodes staring at a minority fork - TFTC: BIP-110 enters mandatory signaling window below 1% hashrateRELEASESBitcoin core / protocolbtcd v0.26.2 - 2026-07-25Security-hardening for the Go full node: stricter PSBT/input parsing, Schnorr and WIF validation, rejection of malformed bech32, tighter inbound admission.Hardware / signingKeystone 3 v3.0.0 - 2026-07-21Major firmware across all variants of the airgapped open-source signer: reworked passcode/recovery flow, stronger validation, upgraded security policies. Reproducible with published checksums.Trezor Suite v26.7.2 - 2026-07-22Firmware security updates plus a lower 0.2 sat/vB minimum fee and cancel-pending-transaction support.Nunchuk 2.7.1 - 2026-07-16Collaborative-custody multisig wallet. 2.7.0 (07-15) added self-custodial USDT on Liquid and Trezor Bluetooth support; 2.7.1 is bug fixes on top. On-lens for multisig self-custody.Bitkey App 2026.11.0 - 2026-07-14Block's consumer hardware wallet. Release highlights its Emergency Access (recovery/inheritance) path; full notes hosted off-repo at bitkey.world/releases.LightningCore Lightning v26.06.6 - 2026-07-22Patch release (26.06.3-5 pulled over broken PyPI publishing). Now rejects channels reusing an existing funding outpoint, closing a channel-security edge case.LNDg v1.11.0 - 2026-07-26Self-hosted LND dashboard: peer-offline reporting, auto re-index on data migration, historic failed-HTLC data via API. Update logging config on upgrade.Zeus v13.1.3 - 2026-07-21Point release / version bump on the 13.1 line for the self-custodial Lightning wallet.LNbits v1.5.6 - 2026-07-15Minor patch on 1.5.5 (payments extension-field refactor and fixes) for the self-hosted Lightning accounts system.Lightning Labs Wavelength - 2026-07-21A toolkit for adding self-custodial bitcoin (and stablecoin) payments to any application, designed to create the best developer experience for humans and agents.EcashCashu TS v5.0.0-rc.5 - 2026-07-23RC for the major v5 of the reference TS Cashu library: NUT-18 payment requests (PaymentRequestBuilder), mint-preference support, hardened P2PK validation, integer fee math. Foundational for ecash wallets.Nutshell 0.20.3 - 2026-07-22Reference mint/wallet: Pay-to-Blinded-Key (lock ecash to a receiver without revealing their pubkey to the mint), a Spark L2 backend, and a false-UNPAID melt-race fix. DB migration, back up first.Fedimint v0.12.0-beta.0 - 2026-07-23Beta pre-release of the federated ecash / community-custody protocol. Flagged unstable, no upgrade guarantee. "In the pipeline," not production. (Admin UI: Fedimint UI v0.7.4, adds arm64 image.)On-chain privacy / coinjoinWasabi Wallet v2.8.1 - 2026-07-22Now receives to Taproot addresses by default (a real "state of the network" adoption nudge, four-plus years post-activation), adds Linux AppImage, on top of 2.8.0's serverless P2P filter sync.Ashigaru Desktop v1.1.2 - 2026-07-25Whirlpool coinjoin QoL: live Tor/Electrum status, one-click connect, faster startup, self-clearing coordinator banner.JoinMarket-NG 0.34.2 - 2026-07-20Actively-maintained modern fork of JoinMarket: safe expired-fidelity-bond handling, correct frozen-UTXO reporting, multi-wallet RPC routing.Bitcoin Safe 2.1.1 - 2026-07-20Multisig/single-sig desktop wallet: UI fixes and improved Debian build reproducibility.P2P / no-KYCBisq 1.10.4 - 2026-07-24Mandatory security update for the decentralized no-KYC exchange (audit findings): signed DAO block providers, stricter blind-vote/dispute validation, re-enabled BSQ swaps. Required to keep trading.Bull Bitcoin 6.12.4 - 2026-07-24Bug-fix for the no-account self-custodial app (iOS startup-lockup fix). The feature release was 6.12.2 (UTXO/coin-control, Coldcard NFC, BitBox02 Nova BLE, sub-1 sat/vB).Vexl v1.45.1 - 2026-07-21Point release of the contacts-based no-KYC P2P trading app (small fixes).Peach Bitcoin 0.69.0 (381) - 2026-07-23Latest build of the no-KYC P2P Bitcoin marketplace (rolling 0.69.0 build increments 379/380/381 across the fortnight). Verify the build-tag slug before publishing (parentheses in the tag).Self-hosting / infraBTCPay Server v2.4.1 - 2026-07-23Self-hosted no-KYC payment processor: BIP-329 label import, editable invoice comments, refund-email triggers, RTL UI, restored Boltcard payments.Start9 StartOS v0.4.0 - 2026-07-24Major: a complete ground-up rewrite of StartOS, out of public beta after six years, billed as the "correct architecture for sovereign computing." Note: the only upgrade path is a fresh install (no in-place migration). One of the biggest self-hosting stories of the fortnight.Liquid GDK release_0.77.7 - 2026-07-20Blockstream's wallet SDK: libwally + Tor bumps, macOS/iOS cross-compile, single-sig gap-limit fee fix.Privacy stack / PayjoinPayjoin Dev Kit payjoin-cli 1.0.0-rc.1 - 2026-07-23RC for the reference Payjoin CLI, synced to payjoin 1.0.0-rc.6. Signals the v1.0 Payjoin stack nearing release (breaks common-input-ownership heuristics on-chain).NostrAmber v6.3.0 - 2026-07-20Android Nostr remote signer (keeps your nsec off client apps): grouped/collapsible multi-request approvals, a log-disabling privacy mode, built-in Tor, NIP-65 relay prefetch.Wallets (self-custody)BlueWallet 8.0.1 - 2026-07-21Major v8 line: iOS 26 UI refresh, BC-UR v2 airgap scanning (OneKey/Keystone), Unchained multisig cosigner import, 19 new languages, crypto-js replaced with @noble. Broad user base. Confirm the exact tag slug before publishing.Cake Wallet 6.3.2 - 2026-07-24Non-custodial BTC/Monero wallet: home-screen recent history, better OpenAlias/ENS/Unstoppable alias resolution, faster Zcash sync.EDUCATIONWhat Is a UTXO, and Why Does It Matter for Bitcoin Privacy? - 2026-07-25Community explainer thread on Stacker News. The useful part is the top response, which walks through how receive-and-spend patterns fingerprint you and where coinjoin actually helps. Good raw material for a plain-English UTXO segment, which pairs with the Wasabi and Ashigaru releases and gives newer listeners the vocabulary before the coinjoin talk.Bitcoin Optech Newsletter #415 - 2026-07-24Two items worth surfacing. Fabian Jahr's draft BIP459 proposes full aggregation of BIP340 schnorr signatures using DahLIAS, combining multiple signatures into a single 64-byte aggregate, with cross-input signature aggregation as a downstream possibility. And libsecp256k1 #1765 adds an optional BIP352 silent-payments module supporting receiver scanning from only the scan secret and spend pubkey, so the spend private key stays offline. Silent payments quietly becoming infrastructure is a good recurring beat.TO DONATE TO ROMAN'S DEFENSE FUND: https://freeromanstorm.com/donateHELP GET SAMOURAI A PARDONSIGN THE PETITION ----> https://www.change.org/p/stand-up-for-freedom-pardon-the-innocent-coders-jailed-for-building-privacy-tools DONATE TO THE FAMILIES ----> https://www.givesendgo.com/billandkeonneSUPPORT ON SOCIAL MEDIA ---> https://billandkeonne.org/VALUE FOR VALUEThanks for listening you Ungovernable Misfits, we appreciate your continued support and hope you enjoy the shows.You can support this episode using your time, talent or treasure.TIME:- create fountain clips for the show- create a meetup- help boost the signal on social mediaTALENT:- create ungovernable misfit inspired art, animation or music- design or implement some software that can make the podcast better- use whatever talents you have to make a contribution to the show!TREASURE:- BOOST IT OR STREAM SATS on the Podcasting 2.0 apps @ https://podcastapps.com- DONATE via Monero @
In this episode, sponsored by GuardSquare (guardsquare.com), Ken Johnson and Seth Law discuss OpenAI's reported Hugging Face security incident, questioning whether the model demonstrated genuinely novel offensive capability or mostly chained known vulnerability patterns at high inference cost, while also considering the defense-contract and marketing angles around "dangerous" frontier models. The main technical discussion returns to AppSec fundamentals through an article on preventing IDOR, emphasizing authorization as a core control, the difficulty of role and tenant isolation in complex systems, and the need for framework-level patterns, typed IDs, tenant checks, and thorough authorization testing. They also cover Krebs' reporting on LG banning residential proxy SDKs from smart TV apps, explaining how free TV apps can turn consumer devices into proxy infrastructure and why IoT app ecosystems need stronger review. The episode closes with DEF CON logistics, Hacker Tracker updates, and upcoming guest plans.
David Soria Parra is an Engineering Lead at Anthropic and one of the core maintainers of the Model Context Protocol (MCP). We explore the biggest evolution of the protocol since its launch, and why MCP is becoming the foundation for the next generation of AI agents.We discuss why MCP is moving toward stateless communication, what developers misunderstand about state, sessions, and transport layers, and how lessons from real-world deployments at massive scale have shaped the protocol's future. We also dive into MCP v2, SDK migrations, protocol design, extension architecture, governance, developer experience, and how Anthropic thinks about balancing simplicity with long-term flexibility.Along the way, we explore progressive disclosure, tool search, programmatic tool calling, context bloat, forward compatibility, long-running AI tasks, protocol evolution, open-source governance, observability, and why the future of AI infrastructure will depend on designing protocols that can evolve without breaking the ecosystem.Timestamps:[00:00] Introduction[01:59] Why MCP Had to Become Stateless[04:28] The Tradeoffs of Stateless Design[06:13] What We Learned About Agent State[08:04] Sessions, Models & Implicit State[09:33] Migrating to MCP v2[12:19] Lessons from HTTP & Open Source Standards[18:16] Shipping Fast Without Breaking Everything[20:35] The Future Complexity of MCP[22:44] Core Features vs Extensions[26:47] Progressive Disclosure Explained[28:16] Solving Context Bloat[30:50] Why Tool Search Beats Progressive Disclosure[32:10] The Biggest MCP Anti-Pattern[34:25] Designing for Forward Compatibility[38:41] Why "Tasks" Matter[40:53] JSON, Tokens & Better Tool Calling[44:44] Observability & Tracing AI Agents[47:34] Will MCP Ever Be Finished?[50:22] What's Next for MCP
Camera-free smart glasses just hit a $1 billion valuation. Even Realities is betting that the winning pair of AI glasses is the one you forget you are wearing.In this episode of Metavertising, host Ely Santos sits down with Raag Harshavat, developer ecosystem lead at Even Realities and previously at Snap Inc. and Meta, to unpack how the Even G2 became one of the most-worn devices in the smart glasses category without a single camera on board.Raag breaks down the design decisions behind a 36 gram pair of glasses that runs for two days on one charge, why Even Realities builds its own prescription lenses in its own factory, and how a monochrome green micro LED waveguide display turns out to be a feature rather than a compromise. He also shares what happened when he ran live translation for twelve straight hours in China, and why "quiet tech" and ambient computing describe something very different from what most of the industry is shipping right now.For developers and creative technologists, this is a practical map of the Even Hub ecosystem: 400+ apps and climbing, a JavaScript-based SDK, a desktop simulator, and a Claude skill that makes vibe coding your own glasses app a realistic weekend project.In this episode:
As open web publisher traffic drops due to changing search engine behaviors, enterprise brand advertisers face declining reach and measurement challenges.Stephen Upstone, CEO & Founder of LoopMe, breaks down how the company leverages artificial intelligence and a network of 30,000 mobile SDK integrations to bring brand advertising into mobile apps and Connected TV.Key tactical themes covered:- Reallocating media budgets from web environments to the resilient "open app internet".- Moving brand campaign metrics from top-of-funnel reach to trackable mid-funnel intent.- Deploying internal agentic AI to increase developer code output by 5x–6x.- Building supply-side seller agents to enable automated machine-to-machine media transactions.- Reallocating talent resources to support a 5-year strategic technology vision.Stephen Upstone is the CEO & Founder of LoopMe, an adtech platform optimizing brand advertising performance across mobile apps and Connected TV using AI.Connect to us or our guestFollow Stephen Upstone on LinkedIn: https://www.linkedin.com/in/stephenupstone/Explore LoopMe: https://loopme.ai/Simplify Paid Social with Strike Social: https://strikesocial.com/guaranteed-performance-marketing/Connect with Host Dylan Conroy: https://www.linkedin.com/in/dylanconroy/
Samsung winds down the display for a cheaper Vision Pro the same week Apple publishes 74 pages of engineering spec for third-party motion controllers — then we close with the best app night in months, from Library of Congress stereographs to spellcasting with your bare hands.TOP STORYCheaper Vision Pro Display Work Winds Down at Samsung (MacRumors)https://www.macrumors.com/2026/07/08/cheaper-apple-vision-pro-display-work-ended/Component Development Reportedly Scrapped (9to5Mac)https://9to5mac.com/2026/07/08/component-development-for-cheaper-apple-vision-pro-reportedly-scrapped/Lower Cost Vision Pro May Be Dropped (AppleInsider)https://appleinsider.com/articles/26/07/08/lower-cost-apple-vision-pro-may-be-dropped-as-apple-focuses-on-aiHARDWARE & SUPPLY CHAINWhere Will Next-Gen Components Be Made?https://appleworld.today/2026/07/where-will-the-next-generation-of-apple-vision-pro-components-be-made/Patent: Detecting Contact Lens Shifthttps://appleworld.today/2026/07/an-apple-vision-pro-may-one-day-be-able-to-tell-is-your-contact-lenses-shift-on-your-eyes/VISIONOS 27Apple Publishes Motion Controller Specs (Road to VR)https://roadtovr.com/apple-publishes-detailed-technical-specifications-for-third-party-vision-pro-motion-controllers/CONNECTOME & the Challenges of Building for Vision Pro (UploadVR)https://www.uploadvr.com/connectome-a-game-of-points-the-challenges-of-building-for-apple-vision-pro/LAMBORGHINI IN YOUR LIVING ROOMhttps://9to5mac.com/2026/07/07/lamborghini-launches-apple-vision-pro-app-with-interactive-full-size-cars/https://www.uploadvr.com/lamborghinis-apple-vision-pro-app-reimagines-the-showroom-at-home/APPS NIGHTStereopticon (Free) https://apps.apple.com/us/app/stereopticon/id6790976075Retro Beamer ($4.99) https://apps.apple.com/us/app/retro-beamer/id6790771285Posters: Discover Media @ Home https://apps.apple.com/us/app/posters-discover-media-home/id6478062053LALO Immersive https://apps.apple.com/us/app/lalo-immersive/id6740135829I'm Wizard: AR Magic Combat ($9.99) https://apps.apple.com/us/app/im-wizard-ar-magic-combat/id6747723768Zork Online https://playzork.online/zorkTHE 3D MOVIE FIXSpatial Film https://apps.apple.com/us/app/spatial-film/id6670564820AirStream Mac Companion https://spatial.film/airstream/DEVELOPER BETA 4https://developer.apple.com/documentation/visionos-release-notes/visionos-27-release-notes- Mac Virtual Display no longer disconnects when you put the headset back on- Palm-up battery percentage corrected (status bar staleness still a known issue)- High Quality Recording fixes: warm-device capture failures, and the settings switch that froze the recording subsystem- Genmoji and Image Playground panels no longer blank from Safari and Freeform- 15+ Siri fixes, including Guest User Mode and history deletion on disable- Still broken: Siri commands for Environments, Maps snippets- Spatial accessory input dropouts and Spatial Gallery panorama freezes resolved- Quick Look annotations get five fixes; Mail subject/content mismatch resolved- EyeSight privacy indicator now animates on every capture, not just the first- Known issues: TestFlight apps still open to blank windows; Spatial Personas lag during High Quality Recording- Devs: Reality Composer Pro Preview is live, On Demand Resources deprecated for Background Assets, and apps built on the new SDK must adopt the scene-based lifecycle or they will not launchEnjoy the show? Subscribe, leave a review, and pass it to another Vision Pro owner.FIND USLive Mondays 9 PM ET — YouTube.com/@VisionProfilesThePodTalk.net | ThePodTalkNetwork@gmail.com
On episode 524 James and Frank dive into the new .NET MAUI developer stack—covering the MAUI CLI/Maui Doctor that auto-provisions SDKs and emulators, the Maui Sherpa GUI for device/Xcode/provisioning management, and DevFlow's MCP server that lets AI agents inspect, interact with and automatically test apps (closing the loop). They also highlight the VS Code MAUI agent/skills, profiling tools, and the shift to core CLR in .NET 11 with performance tradeoffs to watch. Follow Us Frank: Twitter, Blog, GitHub James: Twitter, Blog, GitHub Merge Conflict: Twitter, Facebook, Website, Chat on Discord Music : Amethyst Seer - Citrine by Adventureface ⭐⭐ Review Us ⭐⭐ Machine transcription available on http://mergeconflict.fm
In this episode, we're joined by Jeremiah Lowin, Founder & CEO at Prefect and the creator of FastMCP, to explore how one of the most influential projects in the MCP ecosystem came to be - and where the protocol is heading next.We discuss the accidental origin of FastMCP, why Anthropic adopted it into the official SDK, what developers are getting wrong about MCP, and why Chris believes the biggest opportunity for AI agents isn't customer-facing applications, but internal enterprise systems. We also dive into MCP Apps, developer experience, protocol design, AI tooling, Python, and why building great abstractions is often more valuable than exposing more configuration.Along the way, we explore the rapid growth of the MCP ecosystem, how FastMCP became the default way many developers build MCP servers, why "too much magic" can actually hurt developer experience, and what the next generation of AI-powered applications will look like as agents move beyond simple tool calling into rich, interactive experiences.Prefect: https://www.prefect.ioJeremiah Lowin: https://www.linkedin.com/in/jlowinDemetrios: https://www.linkedin.com/in/dpbrinkmTimestamps:00:00 Lost My Entire Talk00:47 The Story Behind FastMCP02:08 Anthropic Adopted FastMCP02:34 When MCP Took Off04:10 FastMCP vs The Official SDK05:43 Is MCP Actually Dead?06:42 What Everyone Gets Wrong About MCP08:11 MCP's Biggest Use Case10:25 Building Internal AI Systems12:00 Why FastMCP Exploded13:29 Making Complex Software Simple15:10 Can Software Be Too Magical?20:11 MCP Apps Explained23:42 Why Python Needed MCP Apps27:54 The Future of AI Interfaces34:18 AI Should Generate UIs40:11 AI Deleted My Presentation43:30 The AI Assistant We Actually Need48:00 Personal AI vs SaaS52:28 The Future of AI Agents55:06 Final Thoughts
Tony Holdstock-Brown is the co-founder and CEO of Inngest, the durable execution platform that quietly powers your favorite AI agents.We get into why agents work in a demo and die in production, building their own cloud to get 20x lower cost, growing 35x after AWS and Cloudflare copied them, growing a dev tools company without a personal brand or Twitter account, why he thinks evals today are like “asking the criminal if they committed the crime”, and the thing they built to score 100% of your production agents without paying for LLM as a judge.Thank you to Numeral, Flex, Amplitude, Merge, and Monaco for supporting this episode.Numeral: Sales tax on autopilot https://www.numeral.comFlex: Premium banking, 60-day credit, 0% APR https://home.flex.one/referral/bananacapitalAmplitude: AI analytics https://www.amplitude.comMerge: Every model, one API https://www.merge.dev/turnerMonaco: The revenue engine for startups https://www.monaco.com/Timestamps:(0:00) The hidden infra layer every AI agent runs on(1:46) Building complex chains of logic(3:31) Why agent SDK's don't go far enough(4:49) Healthcare was the original event-driven nightmare(6:32) Storing traces on your infrastructure enables self-improving loops(14:26) Why Inngest was already in the right place for AI(15:49) Score agents off product events, not LLM's(17:31) The OpenAI copy-paste signal(21:24) Swap in LLMs and cut costs 20x(23:44) How customers pulled the product forward(25:41) Orchestration belongs outside the sandbox(29:48) Building a neocloud to cut costs 20x(32:09) Most neoclouds just resell AWS(32:54) All AI infrastructure is converging(34:49) Why Claude can't just build your backend(36:44) How to build a software factory(39:12) Agents are a lottery you get addicted to(42:44) Loops must exist until AGI hits(45:38) If models keep getting better, why orchestrate?(48:28) When incumbents steal your features(52:30) Why you can't vibe code infrastructure(55:54) Why Tony has no personal brand(59:38) Dev tools GTM without Twitter(1:03:20) Lessons from the founder of DuckDuckGo(1:10:39) Truth as a company value(1:13:08) Taking too long adapting to AI(1:15:10) Startups are 100% R&D(1:17:19) Ali from Databricks(1:19:03) Writing his own code, Voice-to-text with local models(1:23:53) Evals are batshit insaneReferencedInngest: https://www.inngest.com/Principles by Ray Dalio: https://www.amazon.com/dp/1501124021?lv=shuf&channelId=500&plpRedirect=mhFallbackTraction - How Any Startup Can Achieve Explosive Customer Growth: https://www.amazon.com/dp/1591848369?lv=shuf&channelId=500&plpRedirect=mhFallbackFollow TonyTwitter: https://x.com/itstonyhbLinkedIn: https://www.linkedin.com/in/tonyhb/Follow TurnerTwitter: https://twitter.com/TurnerNovakLinkedIn: https://www.linkedin.com/in/turnernovakSubscribe to my newsletter to get every episode + the transcript in your inbox every week: https://www.thespl.it/
In this Azure Friday episode, Scott Hanselman and Sajee demonstrate the Azure Cosmos DB Agent Kit — a skill you install with one command that gives your coding agent 100+ Cosmos DB best-practice rules across data modeling, partitioning, query optimization, and SDK usage and much more. Using a multi-agent fitness coaching app as an example, they show how the kit caught a missing partition key filter that was leaking member data across tenants, recommended hierarchical partitioning for multi-tenant scale, and fixed a fan-out query—all before the code shipped to production. Chapters 00:00 - Introduction 00:33 - Meet Sajee & overview of the Cosmos DB Agent Kit 00:50 - The problem: partition key & query mistakes that cost money in production 02:32 - How the Agent Kit works: one install, 100+ rules across 12 categories 04:52 - Demo setup: fitness coaching multi-agent app with Cosmos DB 06:26 - Showing the data: missing partition key filter exposes other members' data 08:23 - Agent Kit findings: SQL injection, singleton pattern, fan-out queries 10:46 - Indexing best practices & query optimization recommendations 12:34 - Applying the fix: correct results and single-partition RU cost 13:13 - Wrap up & how to get started Recommended resources Azure Cosmos DB Agent Kit Agent Kit Repository Connect Scott Hanselman | Twitter/X: @SHanselman Sajeetharan | Twitter/X: @sajeetharan Azure Friday | Twitter/X: @AzureFriday Azure | Twitter/X: @Azure
In this Azure Friday episode, Scott Hanselman and Sajee demonstrate the Azure Cosmos DB Agent Kit — a skill you install with one command that gives your coding agent 100+ Cosmos DB best-practice rules across data modeling, partitioning, query optimization, and SDK usage and much more. Using a multi-agent fitness coaching app as an example, they show how the kit caught a missing partition key filter that was leaking member data across tenants, recommended hierarchical partitioning for multi-tenant scale, and fixed a fan-out query—all before the code shipped to production. Chapters 00:00 - Introduction 00:33 - Meet Sajee & overview of the Cosmos DB Agent Kit 00:50 - The problem: partition key & query mistakes that cost money in production 02:32 - How the Agent Kit works: one install, 100+ rules across 12 categories 04:52 - Demo setup: fitness coaching multi-agent app with Cosmos DB 06:26 - Showing the data: missing partition key filter exposes other members' data 08:23 - Agent Kit findings: SQL injection, singleton pattern, fan-out queries 10:46 - Indexing best practices & query optimization recommendations 12:34 - Applying the fix: correct results and single-partition RU cost 13:13 - Wrap up & how to get started Recommended resources Azure Cosmos DB Agent Kit Agent Kit Repository Connect Scott Hanselman | Twitter/X: @SHanselman Sajeetharan | Twitter/X: @sajeetharan Azure Friday | Twitter/X: @AzureFriday Azure | Twitter/X: @Azure
Software Engineering Radio - The Podcast for Professional Software Developers
Clare Liguori, a Senior Principal Engineer who works on developer tooling and agentic AI at Amazon Web Services, speaks with host Sri Panyam about the Amazon Strands Agents SDK. This episode explores the philosophy, design decisions, and emerging patterns behind building production-grade AI agents. Clare frames any agent as three core components: a model, a set of tools, and a prompt. During this interview, she describes the origin story of Strands, the model-driven approach vs. workflows and custom orchestration, steering hooks, tools and MCP, sub-agents and multi-agents, memory layers, production readiness, testing and evaluation starting with use cases where trajectories can be evaluated deterministically, and anti-patterns for newcomers. She describes what's next for Strands, and offers some closing advice for getting results from working with agents
We've been running a bit of an Agent Cloud series surveying all the top inference/compute/cloud providers, from Databricks to Daytona to Railway and, even further back, E2B, but we're excited to conclude this series returning to Modal, which has just raised a monster $355M Series C.The cloud was built for developers. But agents are now changing that.The old infra stack was designed for a human who could read docs, reason through YAML, and understand dashboards to figure out what they need when something broke. While this was painful for developers, it worked since they could fill in missing context in their heads.However, agents don't have that luxury. Now in this new era of agents, everything has to be tighter.They need a place to write code, run it, inspect the output, change the environment, debug failures, and try again. Fast iteration and feedback loops with all the necessary context are crucial for agents to operate properly. Furthermore, sandboxes are a clear representation of this shift as agents can easily spin up isolated environments. This programmatic infra even extends to research:Two years ago, we were one of the first to cover Modal with CEO Erik Bernhardsson and Alessio designed our favorite LS thumbnail of all time:At the time, Modal was just a teeny little company with a $17M Series A.Today, fresh off their $355M Series C, Modal is one of the clearest examples of the agent cloud future being built in real time: a cloud platform moving past traditional web app assumptions toward the workloads AI actually creates such as elastic inference, sandboxes, GPU burst, post-training, background agents, and infrastructure that agents themselves can operate.In this episode, Modal CTO Akshat Bubna joins swyx and Vibhu to unpack why AI applications don't fit traditional cloud assumptions, why Kubernetes was never designed for bursty compute-heavy workloads, and why Modal is now shifting from developer experience to agent experience.We go deep on Modal's AI infra stack: serverless functions, decorator-based infrastructure, elastic inference for custom models, GPU snapshotting, DeFlash, speculative decoding, Auto Endpoints, sandboxes, persistent storage, networked containers, private IPv6, RDMA, multi-node training, and Modal's capacity pool across 17 cloud providers. Akshat also explains why RL rollouts can require 100,000 sandboxes, why production agents need hard guardrails, why observability may matter more than reading code, and why AI has made infrastructure exciting again.We discuss:* Why Kubernetes wasn't built for bursty AI workloads* How Modal started as a better runtime before becoming an AI cloud* Why Modal added GPUs before ChatGPT* The shift from developer experience to agent experience* Why observability matters when agents are writing the code* Elastic inference for custom models across audio, video, robotics, and comp bio* GPU snapshotting, cold starts, and why inference workloads are so bursty* Why RL rollouts can require 100,000 sandboxes* DeFlash, speculative decoding, and frontier-level inference performance* Auto Endpoints and making optimized inference easier to deploy* What Modal adds beyond vLLM, SGLang, and raw GPU rental* Modal's 17-cloud capacity pool and supercloud strategy* Networked sandboxes, sidecars, private IPv6, and RDMA* Serverless multi-node training for post-training and research workloads* Auto-research, model-guided sweeps, and agents launching GPU experiments* Compute strategy, capacity planning, and batch tiers* Why production agents need specialized sandboxes and hard guardrails* Modal's take on managed agents, CI, Gitpod/Ona, Python, TypeScript, and Modal BenchAkshat Bubna* LinkedIn: https://www.linkedin.com/in/akshat-bubna-188885103* X: https://x.com/akshat_bModal* Website: https://modal.comTimestamps00:00:00 Introduction00:00:39 Modal's origin and why Kubernetes wasn't enough00:04:32 Developer Experience → Agent Experience00:06:21 Modal's AI cloud primitives00:09:14 Sandboxes, agent loops, and proto-Cognition00:12:12 Elastic inference, GPU snapshotting, and 100,000 sandboxes00:15:24 DeFlash, speculative decoding, and Auto Endpoints00:19:59 Production-grade inference beyond raw GPUs00:22:00 Background agents, Ramp Inspect, and the agent lifecycle00:24:08 Modal's 17-cloud supercloud strategy00:26:40 Networked sandboxes, private IPv6, and RDMA00:32:48 Multi-node training, post-training, and auto research00:37:36 Compute strategy, capacity planning, and batch tiers00:40:55 Open models, real-time AI, and production agent infra00:43:06 Hard guardrails, managed agents, and specialized sandboxes00:46:06 Why AI made infrastructure exciting again00:48:30 Model APIs, differentiated products, and agentic video00:51:50 CI, coding-agent infra, SDKs, and Modal Bench00:57:28 Closing ThoughtsTranscriptIntroduction: Modal, Series C, and the Art PartySwyx [00:00:00]: We're here with Akshat, CTO of Modal, together with Vibhu. Congrats on your Series C.Akshat [00:00:10]: Thank you.Swyx [00:00:11]: Your party yesterday was amazing.Akshat [00:00:15]: Yeah.Swyx [00:00:15]: From all the photos and all the swag.Akshat [00:00:17]: We had a bunch of art installations, which was fun, seeing, like, our products on pedestals next to, like, Rodin.Swyx [00:00:25]: Very nice. Very nice. When you started, it was not the GPU inference company. Maybe it was in your mind. Take us back to the origin story.Modal's Origin: A New Runtime Beyond KubernetesAkshat [00:00:39]: I first met Eric, who's the CEO, through an investor. Back then Eric was already thinking about building, a new runtime, and he got there thinking through why are workflow orchestration products so hard to use. It's because you have to run them on Kubernetes. Kubernetes is hard to manage. It's not built for burstiness and, custom images,Swyx [00:01:03]: YeahAkshat [00:01:03]: It has a terrible developer experience.Swyx [00:01:05]: And I'll, I'll interjectAkshat [00:01:06]: YeahSwyx [00:01:07]: For listeners, who are new, we interviewed Eric two years ago, and there's a bit more of the story there from Spotify and all those things.Swyx [00:01:14]: And I came across Eric through Data Council because he did that talk on the serverless container stack that you guys did, which was like, that was my first like, “Okay, I need to take Modal very seriously” moment.Akshat [00:01:26]: Yeah.Swyx [00:01:26]: But it was still very unclear, like, do I need all this for just my data pipelines?Akshat [00:01:33]: Yeah. initially what we were thinking about was if we build a better runtime, it's a very useful primitive in itself. It's There's a lot of things that, get solved by serverless functions, like you can do, ETL stuff, you can do job queues, you can do all this, like, bursty processing, which it turns out every company had needs for. but then we also were thinking about this as like, this is a primitive that we can build a whole collection of products on, which are very verticalized. So perhaps data engineering would've been the first one, but we were thinking about inference. Back then it was more classical inference, like computer vision stuff and running XGBoosts and whatnot. But we added GPUs to the product a year before ChatGPT came out.From Serverless Containers to GPU WorkloadsSwyx [00:02:19]: Nice.Akshat [00:02:19]: We just didn't think it would be that big of a deal.Swyx [00:02:22]: Yeah, just like add A100.Vibhu [00:02:23]: Was there any, like, early key problem that really sparked off why you built it?Akshat [00:02:28]: Yeah. Primarily it's just, none of the tooling that was out there was built for, one, a really great developer experience, and also there's a general trend of, a lot of the workloads that we were seeing were very. I wish there was a better word for it, but compute-heavy. Like, they need, one, like, need a lot more resources, so you need to burst up and down a lot, versus like Kubernetes designed for, like, slow scaling and, more for, like, web server use cases. And also there's just a lot more specialization in, like, what kinds of environments these workloads run in. Like, we had sometimes they need accelerators, sometimes they need different kinds of images, and this is just like a consistent thing that we saw across a lot of companies. That would be the next step.Software-Defined Infrastructure and Decorator-Based DXSwyx [00:03:13]: Yeah. Yeah. Be nice. I don't know how much this factored into the early story, but I wrote a post when I was at Temporal about infrastructure, software-defined infrastructure or something like that.Akshat [00:03:22]: Yeah, the self-provisioningSwyx [00:03:23]: Self-provisioning.Akshat [00:03:24]: Yeah.Swyx [00:03:24]: Yeah. I can't even remember my own post.Swyx [00:03:26]: And then you put me on the landing page.Akshat [00:03:28]: Yeah. We really like, the term and so we stole it.Swyx [00:03:32]: Because you had the insight that everything can just be in decorators co-located with the code, right?Akshat [00:03:37]: Yeah.Swyx [00:03:37]: Was that a big part of the originalAkshat [00:03:39]: YesSwyx [00:03:39]: Story or it was just like a DX layer?Akshat [00:03:41]: That was, really important because we really didn't want people to spend, so much time, writing YAML, and it seemed like you could really condense the surface area of what you're doing, put it in code so you can operate on it just like you operate on other code, and like build stuff that's more expressive and dynamic. and so yeah, that was always a very important part.Swyx [00:04:04]: Then the pushback is this is a DSL.Akshat [00:04:07]: Yeah.Swyx [00:04:07]: It's you're closed source. I am locked into Modal.Akshat [00:04:11]: Yeah. We never really got pushback for that because the nice thing about Modal is you can bring whatever code you have, and sure, the DSL is at the configuration layer for, what hardware you're using, how you're scaling things up, but you still own the code.Akshat [00:04:27]: And that's, that's been an important, part of our story, even as we do inference now.Swyx [00:04:32]: Yeah.Vibhu [00:04:32]: How much of do you think still stays the same today? Like if you were to build something today, DevX very important, but I feel like, a lot of this has been changed with just hook it up to an agent, have Claude Code, have Codex implement a tool. there's very agent native primitives that are different than if I'm doing this myself, right?Developer Experience → Agent ExperienceAkshat [00:04:54]: We've changed our SDK team to think about agent experience instead of, developer experience and we think that the same benefits that apply for DX also apply for AX, which is why would you have an agent read through hundreds of Kubernetes files and like write YAML that's not even typed when it can make a couple of changes in a decorator and it gets this self-provisioning runtime of, being able to see its changes live in action? yeah, it just seems from the customers we talk to, they find Modal is much faster for agents to use versus operating on a different substrate.Swyx [00:05:34]: Yeah, because like you, again, you co-locate the infrastructure requirements to the code that runs it.Akshat [00:05:38]: Yeah.Swyx [00:05:38]: Well, the negative thesis now is that nobody's looking at their code anymore, so there's no point.Akshat [00:05:44]: Yeah, people aren't looking at code. one thing we still see is really important is observability.Swyx [00:05:51]: Yeah.Akshat [00:05:51]: Like how good is your dashboard? And of course, like we have, we push a lot of it to the CLI so the agents can do their own investigation, but you still need humans to go interpret what's going on and, make judgment calls and whatnot. and that's I feel like, Maybe more important now than looking at the code itself.Swyx [00:06:11]: Yes, because like, you can try to treat the code as a black box and then use, see the observable action that comes out of it, and then just prompt a change.What Modal Is For: AI Cloud PrimitivesAkshat [00:06:21]: Yeah.Swyx [00:06:22]: So I think it takes a bit of restraint to not specialize, to say, “I want to ship a new primitive,” and then just be general purpose.Swyx [00:06:31]: People ask you, “What are you for?” You're like, “ I don't know. We can do this, we can do that.”Vibhu [00:06:36]: Well, I'd be curious to see, like, okay, if we were to ask you, like, what is Modal for even at a high level? There's a lot you guys do, sandboxes, GPUs, everything. How do you answer?Akshat [00:06:46]: Modal is a cloud platform that's built for, where we've built the primitives from scratch for AI applications. and right now it covers, inference, training, batch processing, and sandbox workloads.Akshat [00:07:00]: But we're building a lot moreSwyx [00:07:02]: I noticed you didn't say web server, so there is still a role for, like, the always-on large-scale Kubernetes type things.Akshat [00:07:09]: Yeah, absolutely. We're, we're not trying to compete with the renders of the world, because yeah, we think the differentiator for us is the, are the workloads that need specialized compute, need to scale up and down a lot. yeah, they're, they're, they're just shaped differently.Working Alongside Frontier StartupsVibhu [00:07:26]: I think you're building a lot of it alongside the startups, right? They're innovating quite a bit, even in your, like, latest blog post. Like, even in the series C, the customers that you mention here, the cognitions, technical ones, ramps and whatnot, they're, they're innovating with you, right? And that's not something AWS is doing directly with.Akshat [00:07:45]: Yeah, absolutely. I think, this is again classic. We're a small team. We can move really fast. our engineers are working with our customers and figuring it out. Yeah.Swyx [00:07:54]: So my first week at Cognition, I walked in, there was someone wearing a Modal shirt. I was like, “What are you doing here?” They're like, “Yeah, I just. I am embedded inside of Cog.”Akshat [00:08:05]: Yeah, I think that was Peyton. We sent him overSwyx [00:08:07]: Yeah.Akshat [00:08:07]: Because, the latency of communication was too high otherwise.Swyx [00:08:12]: Yeah, distributed node, you have to - you have to place one and collocate.Vibhu [00:08:16]: Yeah.Swyx [00:08:16]: So I had a, I had direct personal experience, right? So I worked on smol developer three years ago. it was inspired by Claude 1. I think you onboarded me at some point, like, just before, and I was like, “Oh, like, I need some bursty compute. Like, I was just gonna try using Modal.” And it was a, it was a pretty pleasant experience. apparently, I showed up in the board meeting, like the analytics.smol developer, Sandboxes, and Proto-CognitionAkshat [00:08:39]: Yeah, you blew up on Hacker News and,Swyx [00:08:41]: YeahAkshat [00:08:41]: We got a big traffic spike. I. I think the way you used smol developer was Modal functions for running stuff, which was. Like, the, that was a good use case. but then, yeah.Swyx [00:08:53]: Yeah. That - So to me, that was proto-cognition.Akshat [00:08:55]: Right.Swyx [00:08:56]: If only I had, like, stuck to it.Swyx [00:08:58]: Like, that was like, if - did you say draw the tech treeAkshat [00:09:00]: AbsolutelySwyx [00:09:00]: You're just like, “Yeah, like, probably this will happen.”Akshat [00:09:02]: Yeah. Like, he was so close. You were just rebuilding upon usSwyx [00:09:04]: I just didn't realize.Akshat [00:09:05]: But the funny story there is at the same time, we were talking to a bunch of customers who needed something like sandboxing.Swyx [00:09:14]: Yeah.Akshat [00:09:14]: This is like twenty-three.Swyx [00:09:15]: Yeah.Akshat [00:09:16]: So we builtSwyx [00:09:17]: You introduced a new API right after that.Akshat [00:09:18]: Yeah.Swyx [00:09:19]: Yes.Akshat [00:09:19]: Like, we built sandboxes in May of twenty-three before anyone was even knew this was gonna be a thing. And the first example we published was, we took smol developerSwyx [00:09:28]: Smol developerAkshat [00:09:28]: And put it in a loop, so the agent can iterate on itself.Swyx [00:09:33]: Loops are hot these days.Vibhu [00:09:34]: It's the looper.Akshat [00:09:34]: Yeah.Vibhu [00:09:35]: Loops in. When was this, twenty-three?Akshat [00:09:38]: Yeah.Vibhu [00:09:39]: A small check.Akshat [00:09:39]: Yeah.Swyx [00:09:39]: It's like twenty-three. so the. the, those for listeners, like, the problem was the models are not built for any of this, right?Swyx [00:09:46]: Like, you're just trying to like. They're not post-training to understand, like, looping and, like, self-correction and tool calling was there, but, like, also not that great.Akshat [00:09:55]: Yeah.Akshat [00:09:55]: I don't remember if you used tool calling in this one, but yeah, the models would just diverge after like ten iterations and not produce anything meaningful.Swyx [00:10:03]: Yeah. But like, then. So okay, like now talking to myself three years ago, the answerVibhu [00:10:08]: Of course they will get betterSwyx [00:10:09]: Collect all the failures, build benchmark, and then collect all the, examples, build the RL environmentAkshat [00:10:15]: RightSwyx [00:10:15]: Sell it for like ten billion dollars to Meta.Swyx [00:10:17]: And then also train a model and then sell that for sixty billion dollars to Elon. And this isAkshat [00:10:23]: Yeah, of courseSwyx [00:10:23]: The funny machine. Like, it's like, it's about the hardware.Akshat [00:10:28]: It's hard to have that inherent conviction that the stuff will get that much better.Swyx [00:10:33]: In retrospect, it's so f*****g obvious.Akshat [00:10:36]: Fair enough.Swyx [00:10:37]: Like, what else were we doing back then? I don't know. anyway. Yeah. So this. That was the start of your sandboxing journey, right? I feel like it didn't blow up until, like, last year.Akshat [00:10:49]: Yeah.Swyx [00:10:50]: So there was like a couple years of quietness.Akshat [00:10:52]: Exactly, yeah. We wereVibhu [00:10:53]: I think very underrated product value. Like, my experience with Modal, Charles, before he had joined Modal, met this guy at a hackathon, and he really insisted we wanted to run some small model, not hosted anywhere, and he's like, “ there's this cool company, Modal. They'll like spin up a GPU sandbox, we can throw it on there. They'll take a Hugging Face link.” And like there's so much value just right there, right? Like instant hosting, spin it up, spin it down. It'll stay cold, but we run the demo a few days later, it'll come back up and like all this stuff in retrospect, like it's still what we needed like today.Akshat [00:11:27]: Yeah, it's still needed today. workload shapes have changed a lot as, we run stuff for people with really massive production scale and, there it's it's not about scaling from zero to one, but it's how do we scale really elastically, from like thousand to fifteen hundred GPUs very quickly in a given region. It's the same shape problem.Elastic Inference, GPU Autoscaling, and Custom ModelsVibhu [00:11:50]: Okay. So you look at, say, Cursor Composer, right?Akshat [00:11:53]: Yeah.Vibhu [00:11:53]: They had a. “We'll do RL on a model every couple hours.” you guys have a whole version of RL inference gym and whatnot.Vibhu [00:12:01]: When you look at workloads like that, you're doing train runs where you need to scale up, scale down every hour thousands of GPUs, right? That's the example for we do need it, right?Akshat [00:12:12]: Yeah. Well, so I'll, I'll take a step back and, maybe talk about like how people use Modal today. because our biggest use case is, elastic inference. And the thing we first found product market fit, with was inference for custom models. So we stayed away from the LLM space, and we were serving companies like Suno for audio, Runway for video, robotics, comp bio companies that train their own model elsewhere. But Modal is the best black box that for deployment, scaling to however many GPUs you need as your traffic pattern changes. And we saw all of them like have a very unpredict- predict- predictable, traffic pattern. it's like diurnal. It's Some days, like the company will do a launch and, they'll need like, way more. And it's not just one model that they deploy. They-- all these companies deploy, lots of different models in different regions, and so the autoscaling problem becomes even harder because then you have to scale within a certain region, and those cycles are offset. So different times you scale up in different regions.Akshat [00:13:20]: So that's like our sortVibhu [00:13:22]: And thatAkshat [00:13:22]: YeahVibhu [00:13:22]: That in and of itself is a huge category. There's a bunch of inference providers which, provide this fireworks, does this as a service together, whatnot, Base10. that's carved into its own niche for language models, at least right now.Akshat [00:13:36]: Yeah. the thing that we have specialized in is the autoscaling aspect.Vibhu [00:13:41]: Yeah.Akshat [00:13:41]: Because we found that it's not universally true that everyone else can autoscale, and we've gone deeper into it on the tech side by, we've incorporated GPU snapshotting into the product so we can take the GPU state, like your torch.compile model, snapshot it, and the next cold start is way faster. And so going back to your question, it's That's why you need a lot of burstiness for inference. But then people also do a lot of demand training, like for RL stuff, your rollouts are bursty, as you said. People also do a lot of batch jobs. So we'll see, a lot of companies, before they have a training run, they'll need thousands of GPUs to run encoding or something like that. And I think those things are much more bursty than. I agree that agents are not that bursty. sandboxes are, except when you're doing RL. RL is justRL, Batch Jobs, and 100,000 SandboxesVibhu [00:14:28]: Or commerceAkshat [00:14:28]: Insanely bursty.Vibhu [00:14:29]: Yeah.Akshat [00:14:30]: Yeah. Like when you're doing, rollouts, you sometimes need a hundred thousand sandboxes in your sandboxes.Vibhu [00:14:37]: Yeah. I'm curious if you've seen early sparks of continual learning. There are some people, like our friends, ngram, recently announced thisAkshat [00:14:45]: YeahVibhu [00:14:45]: They're, they're trying to do training. That also seems like a different workload, right? If you're doing training twenty-four/seven per se, there's a very weird dynamic of how you're using GPUs between people and whatnot, but seems like something you guys would work for.Akshat [00:15:00]: As you said, we're, we're fortunate to work with a number of, customers at the frontier and grab some of our customers. and they are taking the primitives we have, and trying to use them in very interesting ways, like continual learning. It's possible as the stuff gets better, some of that will be part of, our offering as well if, more people need it. but we're, we're just waiting to seeVibhu [00:15:23]: YeahAkshat [00:15:23]: How it shakes out.Vibhu [00:15:24]: Is there a primitive that you added after sandboxing that was the next step in the story?LLM Inference, DeFlash, and Speculative DecodingAkshat [00:15:32]: I guess we've been going much deeper into LLM inferenceVibhu [00:15:35]: YeahAkshat [00:15:35]: Because we realized that some of the advantages we have with like autoscaling, again, especially in different regions and whatnot, are, not present elsewhere. and the place where we had a gap was we weren't, working on the model layer itself. Like we were a black box. And, we realized that, we can get to frontier-level model performance, with, by having great people who work on this. And, we've been open sourcing a lot of our work, in terms of, Recently, we, shared our work on DeFlash, which is a block-based, speculator, and we've open sourced, all of it. So, you can - By using open source DeFlash, you can get the same performance as you would with one of the proprietary providers. And the next thing we're thinking about hereVibhu [00:16:23]: I thought this wasAkshat [00:16:24]: YeahVibhu [00:16:24]: An interesting blog post as well, right? Like, I think in here you make a claim that. Not a claim, just that how effective speculative deco-decoding really just get to.Akshat [00:16:33]: Yeah.Vibhu [00:16:33]: Anything you wanna point out from this around, what people should know?Akshat [00:16:39]: Yeah, absolutely. the high-level summary is, it would help to describe what speculative decoding is.Vibhu [00:16:44]: Yes.Akshat [00:16:44]: I will, yes.Vibhu [00:16:45]: I think, likeAkshat [00:16:46]: YeahVibhu [00:16:46]: So we've covered like Eagle and all thisAkshat [00:16:47]: YeahVibhu [00:16:47]: Like Hydra and all those things, but it was like two years ago.Akshat [00:16:51]: Yeah.Vibhu [00:16:51]: I think it doesn't hurt, right?Akshat [00:16:52]: Yeah. Speculative decoding is you have a smaller model, called a draft model, predict tokens ahead of the bigger model, and then you have the bigger model, verify all of this, all the tokens are predicted. And the reason it's faster is if you're predicting, one token at once, you're bound by memory bandwidth. But if you can batch the verification of, the draft model, then you're much more efficient using compute, and it's faster, and as long as your draft model is producing a lot of tokens that can get accepted, which is called the accept length, you can get a speed up that's, multiple times of, the original model speed. and well, that's what we highlight here. It's Like people talk a lot about we made these kernels faster and whatnot, but improving kernel will only give you like few percentage points of improvement, and, increasing accept length, literally is a multiplicative decreaseVibhu [00:17:47]: Like two to four X.Akshat [00:17:48]: Yeah, exactly.Vibhu [00:17:48]: Without much head-on performance.Akshat [00:17:50]: Yeah. I think it may - you are running a second model, right? So it may be something more expensive in the compute,Vibhu [00:17:57]: I meant quality performanceAkshat [00:17:58]: Probably not by muchVibhu [00:17:58]: But yeah. I thinkAkshat [00:17:59]: So there's no drop in quality performanceVibhu [00:18:01]: YeahAkshat [00:18:01]: Because you're always. You're never accepting a token that the big modelVibhu [00:18:04]: It's strictly betterAkshat [00:18:05]: YeahVibhu [00:18:05]: Or it's same.Akshat [00:18:06]: Exactly.Vibhu [00:18:07]: Right. Yeah.Akshat [00:18:08]: And so we've been working a bunch on DeFlash, which is a block-based speculator. so it's instead of predicting, one token at a time, it's predicting a block. And we've been open sourcing our work with it. The next thing for us here is for helping people train speculators and custom models. it's it's something that traditionally is very forward-deployed engineering driven, support deployed, engineer driven, like you work with customers and help them do that. And our vision for. This is why we launched Auto Endpoints, is we want to make frontier-level performance available to everyone. And so, we mentioned this in the announcement, we teased it. The next thing we're, we're launching is, as you run an auto endpoint, we shadow trafficAuto Endpoints and Frontier-Level PerformanceVibhu [00:18:54]: Do you want to explain what auto endpoints are?Akshat [00:18:57]: Yeah.Vibhu [00:18:57]: I lovely, yeah.Akshat [00:18:58]: Yeah. So, this is, I guess, going back to your Modal is you touch the code, but, sometimes people don't wanna touch the code, and they wanna get started with an endpoint that works and has all the great performance and, scalability that Modal has. So we've made that easier with, a way to create an endpoint from our UI, from the CLI, that has all of our optimizations that we talked about, like the DeFlash stuff already baked in, and there's full transparency. So we give you the code, you can go run it yourself, and if you want, you can eject out into the full Modal experience, which we see as people get sophisticated, they do wanna tweak the models, they wanna, fine-tune stuff. You can still do all of that. It's it's not a black box. And yeah, the next thing, as we teased later in the post, is how do we give you value even beyond this in terms of having your draft models evolve as your data distribution evolves, again, without having to talk to a person and, yeah.Vibhu [00:19:59]: I guess just to understand it directly, you have the GPUs, you have an endpoint that's compatible, you serve open model. If someone was to do this themselves, what's the delta that you guys provide? So you do a lot of open source great work on effective inference. how does it compare to, say, I take the same model, 5.2 FP8, take shelf inference engine, vLLM, SGLang, get compute of similar capacity, similar cost. What's the delta that plugging into something this, like this offers outside of the benefit of, scaling?Production Inference Beyond Raw GPUsAkshat [00:20:34]: It's interesting because we've taken the approach of open sourcing our contributions and upstreaming them. we work closely with the SGLang team. We want the improvements that our team, comes up with to be, there in open source for others to use, even outside of Modal. The benefit to us is we have a team that has significant expertise in terms of if you do have something that is not there, our team can help you get that performance, first. the other thing is with these endpoints, we are way more elastic, as you said, than, anyone else, and you have true scaling to zero. you have true, burstiness, and in practice, that matters a lot more to people than just finding, the GPU and, running Modal code on something.Vibhu [00:21:20]: Yeah. And I will say it's not that straightforward to just. like what I said is easier said than done, right?Akshat [00:21:26]: Yeah.Vibhu [00:21:27]: It's I think still for the average person, still hard to just gut check using different. There's, there's quite a bit of combinations you can make there. the trade-offs aren't really known at face value.Akshat [00:21:40]: Yeah. it's it's not just that. I think it's it's that running production-grade inference is a hard infer problem.Vibhu [00:21:49]: YeahAkshat [00:21:49]: Even if you subtract out the autoscalingVibhu [00:21:50]: YeahAkshat [00:21:51]: Is controlling things like tail latency and, making sure every, request is delivered at least once and whatnot.The Model and Agent LifecycleVibhu [00:22:00]: There's a lot of innovation that you can do here. I think, it's very interesting that you're starting to encroach on, like as you become a full cloud, you're starting to encroach on other people's turf.Vibhu [00:22:09]: What will you not do?Akshat [00:22:13]: Well, we wanna follow our users and, make sure they get like a platform that has everything that works well together. so right now we're focused on the model lifecycle and the agent, lifecycle. so both like going from data prep to training to inference, and then also if I want to deploy a background agent, let's say, sandbox, do persistent storage, a whole bunch of other stuff.Vibhu [00:22:38]: We talked to Cole, who did, OpenInspect. Yeah.Akshat [00:22:42]: Yeah.Vibhu [00:22:42]: And RealInspect also is on Modal.Akshat [00:22:44]: Yeah. So Ramp Inspect was a great example of a background agent that was really successful because they, were able to use some of the primitives like snapshotting and fast scaling to just have something that feels really reactive and works well.Ramp Inspect and Background AgentsVibhu [00:23:02]: Yeah. That's the new CTO of, Ramp right there.Akshat [00:23:05]: Yeah, Rahul.Vibhu [00:23:08]: It was really fun. yeah, okay, I think, all very bullish. Like, one of my reflections was also I did not originally. So when I met you guysThe Inference Inflection: CPU, GPU, and Co-LocationVibhu [00:23:19]: You weren't that much in the GPU game, and now you're all about, inference. And one of the points that I hinged on for Jensen's keynote at GTC this year was, what we're calling like the inference inflection, right? That let's say in AI workloads or machine learning workloads, it used to be like, let's call it eight to one GPU to CPU, and now it's more like one to one, which is like a interesting. Like, - because of how much agents are blocked or call out to this, to CPU heavy stuff the actual, like, limiting factor, like, swings back and forth from GPU to CPU a lot more than it used to be all GPU and then occasional CPU.Akshat [00:24:01]: Yeah.Vibhu [00:24:02]: GPU, CPU. And now it's like just constantly, and you just have to locate everything.Seventeen Clouds and the Supercloud StrategyAkshat [00:24:08]: Yeah. And that's one of the things that, again, we see as, something appealing about Modal, which is we've built this capacity pool that spans, 17 cloud providers, so we're, we're very good at Running on various kinds of cloud capacity across the worldSwyx [00:24:24]: You don't have your own data centers?Akshat [00:24:25]: We don't have our own data centers. We just run across a lot of neo cloudsSwyx [00:24:29]: Yeah. AreAkshat [00:24:30]: Metal providers.Swyx [00:24:30]: Yeah. Question mark.Swyx [00:24:31]: Yeah. You're, you're running the math, and you're like, “What's the cutover point where you're like.”Akshat [00:24:36]: Yeah, it's a good question. part of it is we see our differentiator in the software layer, and, being capital light and focusing on the software helps us move really fast. so far it's worked out well because there are so many other people building data centers that we're able to work effectively with them, and again, focus on what makes us, special.Swyx [00:24:55]: Yeah.Swyx [00:24:56]: 17 gets you into, like, the local providers sometimes. LikeAkshat [00:25:00]: The,Swyx [00:25:01]: Which was the most interesting one?Akshat [00:25:02]: There are a lot more neo clouds than you expect, and they all have various degrees of, various levels of reliability. And, that's why it's something we've invested a lot of time in, is building our own reliability layer on top. so if the GPU falls off the bus or something happens, we user workloads are not affected, and that lets us use a lot more capacity than,Swyx [00:25:30]: YeahAkshat [00:25:30]: You as a user would be able to.Swyx [00:25:32]: It's a useful thing to have because like now everyone knows, like, what layer you are and, like, you optimize for being the super cloud of all clouds.Akshat [00:25:41]: Yeah. That's, that's, that's the idea. and so I guess when you mentioned colocation, that's, that's another interesting thing where, one thing we've seen is people come to us when they want, very specifically located, CPUs or GPUs, like they wantSwyx [00:25:57]: Oh, they pin it in likeAkshat [00:25:58]: YeahSwyx [00:25:58]: EU?Akshat [00:25:59]: Exactly. Or EU, US.Swyx [00:26:01]: Right. Data resiliencyAkshat [00:26:02]: AustraliaSwyx [00:26:02]: Locality thing or performance or what?Akshat [00:26:04]: It's either data locality or latency, yeah.Swyx [00:26:07]: Yeah.Akshat [00:26:07]: Like, you want your. They're running sandboxes and model. They want them to be right next to aSwyx [00:26:10]: Yeah, it's easy thenAkshat [00:26:11]: YeahSwyx [00:26:12]: To. That is important in all those things. and so, like, you've accidentally, I don't know if it's accident, but, like, you've built the perfect primitive for agents to express themselves. And then, like, it's almost very funny how every extra development just involves more file system, just involves more CPU.Akshat [00:26:30]: Yeah.Swyx [00:26:31]: Just like the things that you already have. I don't know much about, if there's any, like, networking usages that are interesting, but you've also done some good work on networking.Networking, Sidecars, Private IPv6, and SandboxesAkshat [00:26:40]: Yeah, that's exactly right. Like, we're just taking compute storage and networking and building stuff on that layer, for, again, the stuff people need.Swyx [00:26:49]: YeahAkshat [00:26:50]: We see a few interesting networking things coming up. one is people want networked sandboxes. so we haveSwyx [00:26:57]: For like a Docker cluster type thing.Akshat [00:26:59]: Yeah.Swyx [00:26:59]: Sorry, Docker Swarm. Oh, f**k. What is it called?Akshat [00:27:02]: Compose.Swyx [00:27:03]: Compose type thing.Akshat [00:27:04]: Yeah. So if you want Docker Compose, our sandboxes now support, this thing called sidecars. So you can. A sandbox is a pod of containers, and you can run multiple containers in, a sandbox. also useful because, going back to networking, people want a lot of control over, outbound networking from a sandbox.Swyx [00:27:23]: Yeah.Akshat [00:27:23]: Like, they might wanna run a middle proxy for, like, maybe logging stuff for RL or, controlling how egress can happen to a domain, injecting credentials. and yeah. So we've, we've had to build a lot of that stuff ourselves.Swyx [00:27:38]: Yeah.Akshat [00:27:39]: But then also sometimes people want, sandboxes spanning multiple nodes to talk to each other, which is an emerging thing we're seeing. We have support for that for a different reason, and yeah, we'll see if that becomes stable.Swyx [00:27:52]: Like, just an open socket. It's a. This is directly like mTLS.Akshat [00:27:56]: We do support that, which is you can, expose a tunnel inside a sandbox.Swyx [00:28:01]: Yeah.Akshat [00:28:01]: And then you can either expose it to public internet or it can be, you can add like a HTTP, auth layer above it. But we have this thing called I6PN, which we haven't talked about, which is this, like, overlay network using IPv6 addresses. so if Modal containers, within the same workspace, when this is enabled, can address each other using this private IPv6 address, and no one else can.Akshat [00:28:28]: So it's like private networking, for containers. We built it because we needed it as a primitive for our distributed training product. so we have this other feature, which is you can add a decorator to a function, and you get a cluster of GPUs. and they have RDMA networking. so you can run a distributed training job, that's truly serverless. and we did the overlay network for that. But then we've seen that people are using it for other reasons, and, I'm intrigued to yeah, what would people do with it.Swyx [00:28:59]: Build primitives and let people figure it out, right?Akshat [00:29:01]: Yeah, exactly.Swyx [00:29:02]: You put out a pretty interestingAkshat [00:29:03]: They're like, they read the docs webpage. Let me use thatSwyx [00:29:06]: YeahAkshat [00:29:06]: Something they never intended to work. This is literally not even in our docs page. People somehow found it, and they're using it.RDMA, Memory Movement, and Distributed TrainingSwyx [00:29:12]: Huh.Swyx [00:29:14]: The way you portrayed it with, like, RDMA versus TCP, like, very well laid out, but just the transfer speed change at scale for RL, like yeah, you have it, you have it built in. I'm sure someone found it. It's found it to be a lot more efficient before you made a thing out of it, right?Akshat [00:29:32]: Yeah. And not to split hairs, I guess the overlay network is the TCP overlay network.Akshat [00:29:39]: The reason we have that is you need that to do the key exchange for RDMA before you set up the RDMA network on top of that. but then people found the TCP part.Swyx [00:29:48]: Can I tell you, this is like a big aha moment for me becauseAkshat [00:29:51]: YeahSwyx [00:29:51]: So I review 2,200 submissions for the World's Fair.Akshat [00:29:56]: Yeah.Swyx [00:29:57]: And then I got this from John OsterhoutAkshat [00:29:58]: HuhSwyx [00:29:59]: Who I don't know if. Do John Osterhout by name?Akshat [00:30:01]: The name sounds familiar.Swyx [00:30:02]: He published a. He's a well-known professor, published a lot of interesting software design books, and this is the talk he chose to submit, is on RDMA at Inference. And I'm like, you wouldn't think that this guy, who is like operating systems guy, would care about RDMA.Akshat [00:30:20]: I, it makes sense to me because I,Swyx [00:30:24]: This is the cloud, right? YeahAkshat [00:30:25]: Like, the way you move around your KV cache and how efficiently you can do it, how efficiently you move, your weights from your training GPUs to your inference GPUs in RL is there's a lot of degrees of freedom, and it is a systems problemSwyx [00:30:41]: YeahAkshat [00:30:41]: Moving memory aroundSwyx [00:30:42]: YeahAkshat [00:30:43]: Scheduling.Swyx [00:30:44]: This shows you how primitive my understanding of networking stuff is.Swyx [00:30:46]: Is this like the domain of WireGuard as well?Akshat [00:30:50]: Not quite.Swyx [00:30:51]: It's adjacent?Swyx [00:30:53]: Explain everything.Akshat [00:30:54]: Sure.Swyx [00:30:56]: How do we move memory around GPUs?Akshat [00:30:58]: Well, so sorry. Yeah, that is memory. Sorry, I was talking more, and maybe I was talking like five minutes back, about the private IPv6, addressing that you've set up.Swyx [00:31:09]: Yeah.Akshat [00:31:09]: Is it like it's a VPN?Swyx [00:31:10]: Yeah, it is like a VPN, and yeah, WireGuard is, yeah, you're right. It is,Akshat [00:31:16]: Right. Yeah, you already moved on to new topicsSwyx [00:31:17]: A similarAkshat [00:31:18]: OkaySwyx [00:31:19]: In the same space, WireGuard is, encrypted and this is,Akshat [00:31:23]: And you don't need encryption.Swyx [00:31:23]: Yeah.Akshat [00:31:24]: Yeah.Swyx [00:31:24]: This is not encrypted. that's the main difference. This is TCP and we have eBPF programs that will reject or allow the TCP connection based on whether you're allowed to do it.Akshat [00:31:35]: Used to involve a full sidecar, but now you have eBPF in the Linux kernel.Swyx [00:31:39]: Yeah.Akshat [00:31:40]: Yeah. I don't know if this is a natural follow-on to the topic of like my skepticism on distributed training is that while, like, people spend a lot of money on, like, cables to hook up GPUs, and even that is not, like, fast enough, and that's the bottleneck, is your networking fast enough?Swyx [00:31:59]: Yeah. So I guess you're talking about fully distributed training like, Dialog or something which is like cross data centerAkshat [00:32:06]: That would be, yes.Swyx [00:32:07]: That's the extreme.Akshat [00:32:08]: Yeah.Swyx [00:32:08]: You're in the middle, and then other people would have like the Mellanox cables up in, like, their actual data center.Akshat [00:32:14]: When you run multi-node training on Modal, RDMA, I think Mellanox, is, or InfiniBand is like a, is all seen as RDMA. but it's a way to bypass the TCP networking stack and, transfer, stuff much faster, between one node, to the other. And we have I think like 3 terabit per second, internal networkingSwyx [00:32:40]: OkayAkshat [00:32:40]: Which is the standard that's needed.Swyx [00:32:42]: Okay. So I misunderstood whatAkshat [00:32:43]: 50Swyx [00:32:43]: What part of the stack you wereAkshat [00:32:44]: 50 gigs overSwyx [00:32:45]: YeahAkshat [00:32:45]: If you wentSwyx [00:32:45]: YeahAkshat [00:32:46]: RDMA.Swyx [00:32:46]: Okay.Swyx [00:32:48]: Yeah. I, very impressive work.Multi-Node Training, Post-Training, and Auto ResearchSwyx [00:32:52]: So effectively you're extending like the model philosophy to the training cluster, like, yeah.Akshat [00:32:59]: Yeah. And we're, we're not going for like large scale training runs. the thing that we've built multi-node training for is, we see a lot of, smaller scale post-training. like, people are post-training like medium sized fund models, so they can, get higher quality on inference. this is a perfect fit, for something like that.Swyx [00:33:21]: Yeah. That is my impression of how a lot of these labs explore branches in post-training and then eventually merge whatever they find in.Akshat [00:33:31]: Yeah. The other use case we've seen for multi-node training is even if you have a big cluster, your researchers are still doing small runsSwyx [00:33:38]: YesAkshat [00:33:39]: Having elasticity thereSwyx [00:33:40]: Right, sureAkshat [00:33:40]: Matters a lot more.Swyx [00:33:41]: Yeah. the, like, this is like the current limiting factor for auto research, which is like you need to give your model some GPUs in order for it to completely run.Akshat [00:33:51]: We have a blog post on auto resource and model is,Swyx [00:33:55]: YeahAkshat [00:33:56]: Yeah, like, turns out to be pretty good substrate for that.Swyx [00:33:59]: So my impression is auto research means many things, likeAkshat [00:34:01]: YeahSwyx [00:34:01]: Anything that Andrej coins. Right now it's still science fair, right? Like not like, I don't know how many people are doing this.Akshat [00:34:08]: We're having a golf.Swyx [00:34:08]: Yeah.Akshat [00:34:09]: I thought the same thing.Swyx [00:34:11]: Yeah, you would know.Akshat [00:34:12]: We, like, our internal both training and inference teams use this the general shape of this quite a bit. like we have this one internal repo called auto inference, which essentially we've automated our own forward-deployed engineering efforts using, this harness, which is, the agent will just spin up a sweep of different things. It'll even run like, NVIDIA inside profiler and it'll like tweak configs and it'll arrive the right thing. it'll change your GPUs both from H200 to B200, and works really well.Swyx [00:34:47]: Nice.Akshat [00:34:47]: So yeah.Swyx [00:34:48]: By the way, I enjoy that your forward-deployed engineering is so technical that you have to do these things.Swyx [00:34:52]: It's very different from forward-deployed engineering from other people.Akshat [00:34:54]: Yeah. For our forward-deployed engineering team is, essentially they're like applied inference researchers or applied training researchers.Swyx [00:35:02]: Someone told me like they have to be able to build, but they also have to be able to sell. do they have to sell or are they like they're good, they're just like post-sale type of thing?Akshat [00:35:09]: It does, being able to talk to a customer and engage effectively with themSwyx [00:35:13]: YeahAkshat [00:35:13]: Matters a lot.Swyx [00:35:14]: They want the same thing.Akshat [00:35:15]: Yeah.Swyx [00:35:15]: ?Akshat [00:35:15]: But it's it's not really a sales, thing. We pair them with-- We have solution architects as well that are more on the sales side.Swyx [00:35:23]: Okay. Let's spend a bit more time on auto research. This is a big focus for for this year. Where does this go? like, have people explored enough? Like, there's all these beautiful charts of like improve and then level off a bit and then you find the next thing. Is this one abstraction up from normal training? Is that how we think about it, or do you think about it differently? Like model level training versus high, like driven hyperparameter search.Auto Inference and Modal BenchAkshat [00:35:51]: Yeah, like,Swyx [00:35:51]: Someone, some people call it like neural architecture search or whatever, right? Like.Akshat [00:35:54]: Yeah, - So the stuff I've seen people do with it is nowhere on the architecture level. It's pretty much tweaking parameters, but it's it's a hyperparameter sweep that's guided by some model intuition, so it's like much more efficient than, whatever other, sweep you would have.Swyx [00:36:12]: Yeah, it's just, it's just a question of where you want to spend your compute?Akshat [00:36:16]: Right.Swyx [00:36:16]: ‘Cause yeah, you can just throw infinite amounts of money on this and somehow you'll bang out Shakespeare?Akshat [00:36:22]: Yeah, infinite monkey.Swyx [00:36:24]: Yeah, so like the very good for model. and I think it's also very important that agents can spin up other agents, can spin up their infrastructure. Like very good for you. how good is our LLMs at generating model code? Like the benefit of existing LLMs is that you are in the data.Akshat [00:36:42]: Yeah. They're, they're surprisingly good. I think like pre Cloud 4 they were not, and then now they're able to shot, stuff out of the box. But we're playing around with releasing like a Modal Bench for like the harderSwyx [00:36:55]: YeahAkshat [00:36:55]: Things, that the LLMs cannot do yet and maybeSwyx [00:36:59]: What's an example of that?Akshat [00:37:01]: I think the things that- Sometimes agents struggle with, without right guidance and a skill is, how to, use the rest of our observability. Like how to. Something is failing, like how do you look at the logs and then update the right thing? It's reasoning about that. But they're able to shot, likeSwyx [00:37:23]: Yeah. You can just add a skill to it?Compute Strategy and Capacity PlanningAkshat [00:37:26]: Yeah. So we have a Modal skill now that. Which is why we built this Modal Bench. It's to find things like that, so we can address them in our tool.Swyx [00:37:35]: Tune a skill. Yeah.Akshat [00:37:36]: Yeah.Swyx [00:37:36]: No. it's it's good. are you facing any shortages? like we talk a lot about GPU shortages, but also CPU, also memory.Swyx [00:37:44]: Yeah.Akshat [00:37:45]: We have had a lot of growth, which means that, there's - we've had to be much better aboutSwyx [00:37:53]: PlanningAkshat [00:37:54]: Proactive capacity planning.Swyx [00:37:55]: Yeah.Akshat [00:37:55]: So we have,Swyx [00:37:57]: Which by the way, like it's like a MBA's like dreamAkshat [00:38:00]: YesSwyx [00:38:00]: Is like just planning this stuff. I think last time you and I talked about something maybe about this.Akshat [00:38:03]: Yeah. we have a really competent team of people that we call, The role is called compute strategy. so yeah, if anyone listening here or wants to work on thatSwyx [00:38:13]: Compute strategy?Akshat [00:38:13]: Yeah.Swyx [00:38:14]: I think,Akshat [00:38:14]: I feel like,Swyx [00:38:15]: I think the normies call it FP&A or something.Akshat [00:38:18]: Well, it's more It's it's not FP&A. It's it's There's a lot of interesting financial questions of like what is the blend between one year and three-year reservations? how do we forecast our own capacity? how do we. especially since our capacity is very fungible across different GPU types and different regions, like you have to model a lot of it. and you also have to have an opinion on how the supply chain is gonna evolve, and then you have to like, take bets,Swyx [00:38:49]: YeahAkshat [00:38:49]: Based on that.Swyx [00:38:50]: Tokenomics.Akshat [00:38:50]: Yeah.Swyx [00:38:51]: This is like probably a not a real point, but, I was trying to think about like what other industries. I was trying to think about like, we cannot be first to like these kinds of problems.Akshat [00:38:59]: Yeah.Swyx [00:39:00]: And what other industries have had this? And I was like, airlines with fuel and like they have to hedge their fuel and like, I think for a long time Southwest because they made like a hero fuel bet, they like were like super low cost becauseAkshat [00:39:12]: OhSwyx [00:39:12]: Compared to everyone else.Akshat [00:39:14]: Yeah. I hadn't thought about that.Vibhu [00:39:16]: We're at a fun time too?Akshat [00:39:18]: Yeah. It's. A lot of the compute business in general, for us is also about being very good about capacity management. That is how you have great unit, economics. but also over time it's how you can unlock more value for customers. Like, one of the things we're building now is like a way for customers to get, If they don't care about latency, like get much cheaper pricing and they'll get results back in like next 24 hours or something, like a batch tier essentially.Batch Tiers and Latency-Insensitive WorkloadsSwyx [00:39:47]: Yeah.Akshat [00:39:47]: And those are levers we have because we control the whole stack and scheduling and whatnot to give people a sufficientSwyx [00:39:53]: Yeah. I feel like they're not as popular. Like those, like the Frontier Labs have all those APIs. They're not as popular as they should be.Akshat [00:40:00]: The demand that we see for something like that is not for LLMs. although sometimes people wanna run evals andSwyx [00:40:08]: OkayAkshat [00:40:08]: Synthetic data prep and there it makes sense.Swyx [00:40:10]: Okay.Akshat [00:40:11]: But it's from a lot of LLM companies, like people who are doing computational bio, like they have to run really big batch jobs and they don't care about when they get it back.Swyx [00:40:22]: Yeah. And like they have a reasonable. It's it's also like a cousin to the stopping problem of like, will this finish in time?Akshat [00:40:30]: Yeah. You can bound it.Swyx [00:40:33]: Yeah.Akshat [00:40:33]: Like you can give peopleSwyx [00:40:34]: YeahAkshat [00:40:34]: SLAs on it.Swyx [00:40:35]: Yeah. I think what's, what's interesting is like the next phase of model.Swyx [00:40:38]: Like what, do people expect from you, now that you're established and you're like well-known compute player among all these leading companies. You had an inference launch week, and we talked a little bit about the launches. like what else? Like what else should people know?What Modal Builds NextAkshat [00:40:55]: We are building primitives that make our users' lives much easier. So, I think for example, with LLM inference, thousands more companies are gonna post-train their own models and, deploy open source models for inference. so we're thinking a lot about what is the best product shape for that. And, that involves everything from our training gym to, then, endpoints that get frontier-level performance. again, but I haven't talked to anyone. It looks somewhat different on other verticals. Like, we're also seeing a lot of real-time, audio-video stuff in there, which is why like, we're working on things like regional routing, with fallbacks. So you can get GPUs that are as close to users as possible. so you get like low latency for video streaming and whatnot. And then on the agent side, it's,Akshat [00:41:52]: We're still working very closely with our customers because stuff is changing so fast in terms of what they need. And, I think beyond sandboxes and persistent file systems, there's a lot of other things people will need from this agent stack as they build production agents. So yeah, we're thinking about those other things that fit in there.Swyx [00:42:13]: I want to ask what the other things are.Akshat [00:42:15]: Yeah. I probably should share right now.Swyx [00:42:17]: I think-- I think, okay, so, I do think a lot about the principal components of cloud, and you do talk about compute storage networking.Akshat [00:42:25]: Yeah.Swyx [00:42:25]: Because so far for me, it's fine. so far for the. the first couple generations of cloud, it's fine. What's different, qualitatively different about agents that you need some new permission level? Like a lot of people, okay, and I'll just kinda spew tokens at you until it like hopefully sparks something.Akshat [00:42:43]: Yeah.Swyx [00:42:44]: Like the new level now is whatever Claude Code does, which is dangerously scope permissions or like allow list by command or like whatever, right? And sometimes they're like, “Well, okay, we have like this adaptive thinking mode where like, just trust me, bro. I will make the calls for you.” Is that it? like mediated permissions.Hard Guardrails vs. LLM-Mediated PermissionsVibhu [00:43:03]: Now you're looping it with a goal and letting it roll.Akshat [00:43:06]: Yeah, I'm, I'm skeptical of LLM media permission for stuff that is at the sandbox level because you do want hard boundaries.Swyx [00:43:16]: Yeah.Akshat [00:43:16]: Otherwise, someone can exfiltrate stuff.Swyx [00:43:20]: But likeAkshat [00:43:20]: YeahSwyx [00:43:20]: Maybe that's old school thinking. Maybe we're the dinosaurs.Swyx [00:43:23]: Maybe the AI OS or the LLM OS is really the kernel is a goddamn LLM.Swyx [00:43:30]: Like it makes you feel uncomfortable.Akshat [00:43:31]: Yeah, I'm, I'm toldSwyx [00:43:32]: But that's what trusting the LLM is. Like imagine a spherical cow perfect LLM.Akshat [00:43:36]: Right.Swyx [00:43:37]: That it.Akshat [00:43:39]: Maybe.Swyx [00:43:41]: I wanna test the boundaries, right?Akshat [00:43:42]: Yeah.Swyx [00:43:42]: Like, and I don't believe that, but I wanna see where I'm wrong ‘cause that's, that's the consensus.Akshat [00:43:49]: Yeah. I think you always need hard guardrails when you want, And you can pair those with softer guardrails, right? And that's gonna be a lot of mediated.Managed Agents and Specialized SandboxesSwyx [00:44:00]: There. I'll also get you a end with a couple of your commentary on like the ecosystem outside of Modal. Manage agents. Everyone has one. Gemini, OpenAI, Claude, very useful for you, but also like it is their way of starting to edge into your space.Akshat [00:44:17]: Yeah.Swyx [00:44:17]: What's going on?Akshat [00:44:19]: Yeah, we're, very excited to partner with Anthropic and some of the other foundation labs, will not name who we're also working with. the way we see it is the manage agent thing is a great place to start if you're starting out building an agent and, But then when you get to, building something more production grade, like you're a company that's like Ramp that's building their own, Ramp also runs their accounting agent on us, so their external-facing agent. You need a lot more control over, your compute primitive on things like, what sort - how do you persist different files that the agent has access to, and how do you snapshot and restore? How do you control the networking? maybe you want GPUs. When you get to that point, you kinda want, a specialized sandbox provider, that gives you those things, and that's the role that we are trying to play.Swyx [00:45:15]: YeahAkshat [00:45:16]: We don't really have an opinion on the harness, whether it runs - it's a cloud-managed agent, and you hook it up to Model Sandbox, or you run the harness in Model Sandbox. We'll see where people converge with that.Swyx [00:45:26]: Yeah. Do you any opinions on like the meta harnesses, or just another layer on top of these things?Akshat [00:45:31]: You mean like the OpenPipeSwyx [00:45:33]: OpenPipe is one. I think Vercel had one, which I can't remember the name of right now. Fredshot had one. and then, to me, most recently was Data Databricks that had Omnigen. All these are meta harness. Like it's kinda pseudo agent cloud type things.Akshat [00:45:50]: I personally have not played around with them.Swyx [00:45:53]: Yeah.Akshat [00:45:53]: Build agents with them.Swyx [00:45:54]: Everything's bullish Modal, as long as it consumes more infra.Akshat [00:45:57]: That's why we're focusing on the infra layer. It's somewhere where our, relative competence is and, also it's a hard problem to solve.Swyx [00:46:06]: Yeah. I will say like just generally reflecting on that, I don't know if - if there's other topics on Modal, but like just generally reflecting as an infra person, not as intense as you, but in that field, this has like been the most exciting time in infra. Like it was boring for a while, and you couldn't really get people excited about data infrastructure. Like Eric would get on Data Console, everyone just watched the video and like say, “Look at how many sandboxes I can spin up,” and no one gave a crap.Why Infrastructure Became Exciting AgainAkshat [00:46:39]: Yeah.Swyx [00:46:40]: And like now everyone gives a crap.Akshat [00:46:42]: That's true. It is a very exciting time, and I think a lot of that's driven by just the amount of scale all of this stuff needs.Swyx [00:46:50]: I think the, like a lot of your initiatives or a lot of your like product directions make sense in retrospect, which is like the best kind, but I wouldn't necessarily have thought about it myself, which.Akshat [00:47:00]: We need the predictions.Swyx [00:47:02]: I think there's a lot that you just don't even see, right? Like you have the batch, you have the voice, you have the multimodal, but what else?Akshat [00:47:10]: What else is coming up for usSwyx [00:47:11]: Yeah. Where do you see things going?Akshat [00:47:13]: Yeah. I, in generalBiotech, Robotics, and Non-LLM AI WorkloadsAkshat [00:47:15]: It's it's clear that there's there's a huge shift happening. I think one thing that's not as obvious to people because LLM inference gets talked about so much and is also we work a lot of companies that are, doing things like drug discovery and computational bio, like the Chai Discoveries of the world. Big things are probably gonna happen there. we work a lot of robotics companies that are putting robots in like active deployments and getting good results out of them.Swyx [00:47:45]: Is there Air Gap Modal? Is there a version that is like prem air gapped whatever?Akshat [00:47:50]: No. We,Swyx [00:47:51]: You should cloud only.Akshat [00:47:51]: Yeah.Swyx [00:47:52]: Yeah. Okay. But yeah, so what you're saying is like because you're focused on primitives and they're good primitives, you find use cases in all these kinds of things.Akshat [00:48:01]: Yeah.Swyx [00:48:01]: Probably diversifies you a little bit away from LMS all the time.Akshat [00:48:05]: Yeah, absolutely. We're, we'- our goal isn't to only serve the LLM inference market.Swyx [00:48:10]: There are a lot just on the website, the audio,Akshat [00:48:12]: Yeah. We said both onSwyx [00:48:14]: Computational bio images. Yeah, there's a lot here. There's QTA TTS, customizing. Oh, Chatterbox. there was customizing Whisper.Akshat [00:48:24]: Okay. Yeah.Swyx [00:48:25]: This screen reminds me of a fallen competitor, which Replicate.Model APIs vs. Differentiated AI ProductsSwyx [00:48:31]: What's your postmortem on what happened?Akshat [00:48:34]: This is one thing we've stayed away from is providing an API for models because I think providing model APIs is some of it ends up serving like a really hobbyist market, which is much less sticky.Swyx [00:48:50]: Yeah.Akshat [00:48:50]: And we've always wanted to build for companies that are building products and need more flexibility that's not just an API.Swyx [00:48:57]: Which you can build an API for a model and this is clearly what it is. But you - but what you're saying, you can wrap it into a more fully functioning back end that you run.Akshat [00:49:06]: Yeah. So all of our examples, it's not that spin up this model, here's an API token, use it. They're all code.Swyx [00:49:13]: Okay.Akshat [00:49:13]: And so the point is that this is just an example.Swyx [00:49:16]: Starter code.Akshat [00:49:17]: Yeah. But you can tweak it however you want.Swyx [00:49:20]: Yeah.Akshat [00:49:21]: And if you're like a company building a product, like, computational bio whatnot, yeah.Swyx [00:49:26]: I guess I'm trying to tease out for listenersAkshat [00:49:28]: YeahSwyx [00:49:28]: When does it stop becoming, oh, you're just an API call and you're just a wrapper on API to becoming what you call a product, right?Swyx [00:49:36]: Like, what is that layer? Like what-- Like, more lines of code, but like beyond that, what is the substance that people add that qualifies it to be something more?Akshat [00:49:46]: I think there's a little bit of like a selection effect of like a lot of the companies who do wanna get deeper into that level are probably building something that's more differentiated. And, I think, an example is like - with LLM inference, originally we, worked with companies that were building their own post-training frameworks or they were, - Ramp early in the day was training their own tokenizer and like swapping out the tokenizer in Llama and whatnot. I'm not saying that's, that successful, in that case. But a better example is like, let's say Suno. because Suno, does not use Modal for training.Swyx [00:50:26]: Mikey on the pod. Yeah.Akshat [00:50:27]: But they use Modal for all their inference and that's because they have like a custom-- They have completely custom model architecture and that means that they have to be at the code level and tweak things that are not, just an API.Swyx [00:50:41]: It's interesting as well, like we had, Ethan, most recently on the xAI Groq team make a prediction that like the next tier in video gen is not a better video model, it's a better model or agent that orchestrates video models.Video Agents and Production WorkflowsAkshat [00:50:56]: Oh, interesting.Vibhu [00:50:56]: Language model backbone that can use toolsAkshat [00:50:58]: RightVibhu [00:50:59]: And write code.Akshat [00:51:00]: Like, yes, I can make my second video or my second video from Groq, but I want my minute video.Akshat [00:51:06]: And I'm not going there through normal video gen.Swyx [00:51:10]: Yeah, that's interesting. I - So we have GPU sandboxes and recently have seen a few companies doing agents that do video manipulation or,Akshat [00:51:22]: Yeah. Give it FFmpeg and just do it.Swyx [00:51:23]: Run FFmpeg. But likeAkshat [00:51:25]: That's not enough.Swyx [00:51:25]: Yeah.Akshat [00:51:26]: You need to give it Adobe.Swyx [00:51:27]: Yeah, I hadn't put it together with like it would be a video production thing. in my mind these things were going more towards editingAkshat [00:51:36]: Yeah.Vibhu [00:51:36]: Well, shout out Mantis.Akshat [00:51:37]: I think about this a lot.Swyx [00:51:38]: .Akshat [00:51:41]: Yeah. Sorry.Vibhu [00:51:41]: Luma. Luma Agent is a version of this for video production, but it's a off.Swyx [00:51:46]: I was gonna get your quick takes, on some other stuff that happensGitpod/Ona, CI, and Runtime SandboxesSwyx [00:51:50]: In recent news and just-just see if you have anything interesting. Gitpod, very li
In this episode of the Altium OnTrack Podcast, host Zach Peterson sits down with Rob Barton, Head of Platform API at Altium, for a deep dive into how programmatic access is transforming PCB design and electronics development. Rob traces the evolution of Altium's API—from the early disconnected SDKs and the launch of Nexar, through Octopart supply data, all the way to the new Platform API that exposes design data, supply chain intelligence, and manufacturing services through a single, federated GraphQL schema. If you've ever wanted to connect Altium 365 and Altium Designer data directly into your own systems, this conversation maps out exactly where the technology is heading. Through two live demos, Rob shows how to query live Octopart supply data—pricing, availability, RoHS compliance, and BOM resolution—then navigates the Platform API down to individual PCB layers, nets, and track coordinates. The discussion also explores API-first design philosophy, why discoverable APIs now matter for AI agents and MCP servers, and the upcoming Altium Developer Center that will open this platform to engineers, enterprises, and third-party developers. Whether you're a procurement professional, a PCB designer, or building AI-enabled tools on top of electronics data, this episode is a clear look at the future of open, programmatic hardware design.
Episode 521 James and Frank obsess over “polish”: the tiny design and packaging details that make apps feel finished. They start with impeccable.style, product.md and design.md (and why agents.md and readme aren't enough), then dig into UI fit‑and‑finish — tray/menu UIs, icon choice, grouping settings and why AI agents still struggle with layout and whitespace. The conversation then moves deep into Windows packaging: WinApp SDK versions, trimming woes with WinRT, ready‑to‑run vs. single‑file self‑contained builds, MSIX tradeoffs, and strange cases where builds bloat with unwanted packages. Key takeaways: give agents the right metadata, expect to hand‑tune UI polish, split architectures, disable R2R for size savings, exclude unnecessary SDK assets, and use Windows Sandbox/WSD for testing. A practical, nitty‑gritty episode for devs who care about the final mile. Follow Us Frank: Twitter, Blog, GitHub James: Twitter, Blog, GitHub Merge Conflict: Twitter, Facebook, Website, Chat on Discord Music : Amethyst Seer - Citrine by Adventureface ⭐⭐ Review Us ⭐⭐ Machine transcription available on http://mergeconflict.fm
EPISODE DESCRIPTION I sat down with Pasha from Igra Labs in Berlin to dig into one of the most contrarian bets in Web3 right now , EVM on proof of work. Pasha walks me through why he left the Ethereum ecosystem, what drew him to Kaspa's BlockDAG architecture, and why he believes proof of work offers something proof of stake simply cannot: real fairness. We get into MEV resilience, censorship resistance, and why 10 blocks per second with no centralized sequencer changes everything for DeFi builders. We also cover Multitude, their brand new sovereign execution zones product, and why AI agents might finally be the user experience layer that makes Web3 click for everyone. If you're a founder, a builder, or just someone curious about where the next wave of adoption is coming from, this one is packed. CONNECT Igra Labs Website: https://igralabs.com/heroDeploy your sovereign finance chain infrastructure https://igralabs.com/multitude Igra Labs Twitter/X: https://x.com/Igra_LabsPavel LinkedIn: https://www.linkedin.com/in/emdin/Web3 with Sam Kamani Podcast: https://www.web3pod.xyz/ KEY POINTS WITH TIMESTAMPS • [00:30] Introduction to Pasha from Igra Labs and what the episode covers• [01:51] Pasha's backstory , from Delivery Hero and Auto One to falling into crypto in 2017• [03:24] The core problem Igra Labs is solving: bringing EVM to proof of work via the Kaspa BlockDAG• [05:53] Why proof of work still makes sense , decentralization, security, and hardware commitment• [07:47] MEV resilience and front-run resistance as the killer property of EVM on proof of work• [09:44] Real use cases: stablecoins, DeFi, RWAs, and AI agentic settlement on Igra• [12:39] How EVM makes it easy for developers , Solidity tooling, SDKs, and AI coding agents• [14:26] Developer onboarding in minutes using a single GitBook link and an AI agent• [16:26] Introducing Multitude , sovereign execution zones allowing teams to launch their own Igra• [18:52] How Multitude differs from parachains, appchains, and L2 models like OpStack• [21:25] Where Web3 goes next , AI agents as the UX layer that finally unlocks mass adoption• [24:48] An agentic marketplace being built on Igra , like Upwork, but for AI agents• [27:02] What Pasha would do differently if starting Igra Labs today• [27:52] Current asks: design partners for Multitude, a small strategic round, and ambitious buildersDISCLAIMERNothing mentioned in this podcast is investment advice and please do your own research. It would mean a lot if you can leave a review of this podcast on Apple Podcasts or Spotify and share this podcast with a friend. Be a guest on the podcast or contact us - https://www.web3pod.xyz/
Talk Python To Me - Python conversations for passionate developers
If you've ever been to PyCon, you know one of the best parts of the expo hall is Startup Row, a stretch of booths where early-stage companies built on Python show off what they're creating. But only attendees get to walk that lane, so let's bring it to everyone. In this episode, we stroll down Startup Row together. We kick things off with the organizers, Jason and Shay, who share the program's origin story going back to Paul Graham and the PSF, plus some surprising stats, including two unicorns among the alumni. Then we meet five startups: Tetrix, bringing AI to institutional investing in private markets. Arcjet, security that lives inside your app as an SDK. Phemeral.dev, serverless hosting built for Python web apps. CapiscIO, an identity and authority layer for AI agents. And Pixeltable, a multimodal database from Marcel Kornacker, co-creator of Apache Parquet. See if you can spot the theme running through them all. Let's go for a walk. Episode sponsors AgentField AI Talk Python Courses Links from the show Guests Naunidh Bhalla: linkedin.com Grant Gittes: linkedin.com Marcel Kornacker: linkedin.com Beon de Nood: linkedin.com Chinmaya Joshi: linkedin.com David Mytton: linkedin.com Shea Tate-Di Donna: linkedin.com Jason Rowley: linkedin.com Azul Garza: github.com Renée Rosillo: linkedin.com Tetrix: tetrix.co Tetrix Jobs: tetrix.co Arcjet: arcjet.com Pixeltable: pixeltable.com Phemeral.dev: phemeral.dev CapiscIO: capisc.io Episode #551 deep-dive: talkpython.fm/551 Episode transcripts: talkpython.fm Theme Song: Developer Rap
If you think code is safe from automation, think again. This week's discussion tackles why the rise of vibe coding and AI-powered tools could upend long-held beliefs about software development, with even seasoned pros rethinking their roles. Also, a new C++ documentary is worth watching! Windows After a weekend of Build session viewing, two big takeaways! Vibe coding native Windows apps and a new reactive dev model for WinUI will help to make modern app dev easier for everyone A new theory emerges: The real reason Microsoft is fixing Windows 11 is that it needs this foundation for a future of hybrid AI agents. And hybrid means more than just local + cloud. Patch Tuesday is here! As promised, Microsoft fixed a record number of security issues thanks to AI 24H2/25H2: Shared audio, more NPU in Task Manager, multi-app camera support, user folder name choice in OOBE, more 26H1: Xbox Mode, Drop tray, etc. Windows Insider Program: New 26H1 Beta channel added for some reason Dell now sells a Windows Hello ESS-compatible wired mouse AI WWDC 2026: Apple announced vibe-coding advances for normal users (Safari extensions) and developers (Xcode). Paul used Xcode and Claude Code to create a full-featured Markdown editor app in about 12-15 minutes. Google drops the price of AI Plus plan to $4.99 per month, raises storage to 400 GB and announces new NotebookLM capabilities Proton Drive is coming to Linux, has a new SDK, and now has a new CLI too. We're going to need a CLI section in the show notes. XBOX and gaming Microsoft Games Showcase: It needed to be a big day for Xbox and it was Microsoft showed off Halo: Campaign Evolved, Gears of War E-Day, Fable, and a lot more Some games will be console-exclusive in the future, starting with the new Gears Microsoft will sell a limited edition Xbox Series X25 later this year Xbox leadership is exploring new business models for the next console - Game Pass lost "millions" of subscribers after last year's price hikes Xbox Insider update adds a new way to discover mutual friends, more Valve says the Steam Machine and Steam Frame will ship this summer Tips and picks Tip of the week: Windows 11 Field Guide is being updated to 2026 edition App pick of the week: Brave Origin RunAs Radio this week: How Machine Learning Fails with Megan Robertson Brown liquor pick of the week: Thy Bøg Hosts: Leo Laporte, Paul Thurrott, and Richard Campbell Download or subscribe to Windows Weekly at https://twit.tv/shows/windows-weekly Check out Paul's blog at thurrott.com The Windows Weekly theme music is courtesy of Carl Franklin. Join Club TWiT for Ad-Free Podcasts! Support what you love and get ad-free audio and video feeds, a members-only Discord, and exclusive content. Join today: https://twit.tv/clubtwit Sponsors: helixsleep.com/windows zscaler.com/security trustedtech.team/windowsweekly365
If you think code is safe from automation, think again. This week's discussion tackles why the rise of vibe coding and AI-powered tools could upend long-held beliefs about software development, with even seasoned pros rethinking their roles. Also, a new C++ documentary is worth watching! Windows After a weekend of Build session viewing, two big takeaways! Vibe coding native Windows apps and a new reactive dev model for WinUI will help to make modern app dev easier for everyone A new theory emerges: The real reason Microsoft is fixing Windows 11 is that it needs this foundation for a future of hybrid AI agents. And hybrid means more than just local + cloud. Patch Tuesday is here! As promised, Microsoft fixed a record number of security issues thanks to AI 24H2/25H2: Shared audio, more NPU in Task Manager, multi-app camera support, user folder name choice in OOBE, more 26H1: Xbox Mode, Drop tray, etc. Windows Insider Program: New 26H1 Beta channel added for some reason Dell now sells a Windows Hello ESS-compatible wired mouse AI WWDC 2026: Apple announced vibe-coding advances for normal users (Safari extensions) and developers (Xcode). Paul used Xcode and Claude Code to create a full-featured Markdown editor app in about 12-15 minutes. Google drops the price of AI Plus plan to $4.99 per month, raises storage to 400 GB and announces new NotebookLM capabilities Proton Drive is coming to Linux, has a new SDK, and now has a new CLI too. We're going to need a CLI section in the show notes. XBOX and gaming Microsoft Games Showcase: It needed to be a big day for Xbox and it was Microsoft showed off Halo: Campaign Evolved, Gears of War E-Day, Fable, and a lot more Some games will be console-exclusive in the future, starting with the new Gears Microsoft will sell a limited edition Xbox Series X25 later this year Xbox leadership is exploring new business models for the next console - Game Pass lost "millions" of subscribers after last year's price hikes Xbox Insider update adds a new way to discover mutual friends, more Valve says the Steam Machine and Steam Frame will ship this summer Tips and picks Tip of the week: Windows 11 Field Guide is being updated to 2026 edition App pick of the week: Brave Origin RunAs Radio this week: How Machine Learning Fails with Megan Robertson Brown liquor pick of the week: Thy Bøg Hosts: Leo Laporte, Paul Thurrott, and Richard Campbell Download or subscribe to Windows Weekly at https://twit.tv/shows/windows-weekly Check out Paul's blog at thurrott.com The Windows Weekly theme music is courtesy of Carl Franklin. Join Club TWiT for Ad-Free Podcasts! Support what you love and get ad-free audio and video feeds, a members-only Discord, and exclusive content. Join today: https://twit.tv/clubtwit Sponsors: helixsleep.com/windows zscaler.com/security trustedtech.team/windowsweekly365