Text-based open standard designed for human-readable data interchange
POPULARITY
Categories
We are excited to have Anthropic share their latest AI x Finance work at AI Engineer New York, coming up in 2 weeks!In case you've been under a rock, here's a non-exhaustive list of what Anthropic has been shipping since closing the largest fundraise of all time in May at $47B ARR:* June: Launched Claude Tag and Sonnet 5 and Fable 5* July: Opus 5, /checkup. crossed $65B ARR* Last month: Fable/Mythos 5.1, and EFS (upcoming pod)* IPO target $2T, end 2026 ARR estimated $100B* Cowork/chat merged before did* Claude Mods* Dario endorses the same Pacing the Frontier message cosigned by all labs* Last week: Opus 5.5, Plugins portal, Cloud Sessions/Claude Projects* Today: Sonnet 5.5!Today's episode should catch you up, with Thariq Shihipar, the explainer-king of Anthropic, who we last caught up on Fable launch day with The Field Guide to Fable:The Future of Mutable SoftwarePay special attention to Claude Mods (especially the cheatsheet):In general this is also the inverse of the other viral tweet from Thariq:Cloud Brain, Local HandsAnd give a try to Claude Projects:The “hands” terminology is not just an analogy for the local/cloud paradigm that is being built up at frontier coding agent companies like Cognition, but is ALSO particularly relevant to the safety systems discussions that we'll be discussing with Anthropic in an upcoming episode as they prepare to pace to frontier with responsible AI deployment.For those who want Thariq's writing tips we teased at the start of the pod, watch the full video here:From the rapid rise of Claude Code to a future where agents can rewrite their own harnesses, collaborate across teams, and operate across cloud and local environments, the way we build software is changing extraordinarily fast. In this episode, Anthropic's Thariq Shihipar joins swyx and Vibhu to unpack how power users are actually working with Claude Code today, why prompting remains a high-skill discipline, and where Anthropic thinks the agent harness is headed next.We go deep on Claude Code's evolving interface: Ask User Question and elicitation, artifacts as persistent generative interfaces, Claude Tag for multiplayer agent workflows, Projects, model effort, implementation notes, and the new Claude Mods system for customizing the harness itself. Thariq explains why Claude.md may eventually disappear, why the smartest model could also become the cheapest model for many tasks, and why mutable software could become a new paradigm for how applications are built and customized.The conversation then turns to agent security and Anthropic's “Pacing the Frontier” argument. Thariq walks through recent incidents where agents discovered unexpected ways to communicate, exploit infrastructure, reverse-engineer benchmark scorers, and chain vulnerabilities together. We discuss sandboxing, prompt injection, autonomous agents, interpretability, constitutional classifiers, probes, fallbacks, Auto Mode, and why securing increasingly capable agents may become one of the defining engineering problems of the next few years.We discuss:* Why agentic coding went from controversial to the default in less than a year* Why prompting is still one of the highest-leverage skills for working with Claude Code* How expert users build a mental model of Claude and what it can reliably one-shot* Why discovering your “unknown unknowns” matters more as agents become more capable* Artifacts as persistent, generative interfaces between humans and agents* How Claude could split into a cloud-based “brain,” local or remote “hands,” and dynamic interfaces* Claude Tag, Projects, and multiplayer agents and how collaborative agent workflows could evolve* Why spending more time on the initial prompt can dramatically reduce wasted agent work* When to use low, medium, high, or max effort for different engineering tasks* Why frontier models may eventually outperform smaller models on both intelligence and token efficiency* Why implementation notes can expose decisions the model considered but chose not to make* Why Claude.md may eventually disappear — and why starting without one can sometimes be better* Claude Mods: customizing the execution loop, UI, subagents, routing, and behavior of Claude Code* Model routers, forked agents, and supervisor agents that automatically improve agent workflows* Why Claude Mods may be an early preview of “mutable software”* The bitter lesson of harness engineering and why agent architectures go out of date so quickly* How Claude Tag is becoming an organizational harness for multiplayer work* Why giving agents access to company data creates an enormous new security surface* The Exploit-Bench incident where agents discovered ways to communicate and collaborate* Why agents hacked Hugging Face for scorer code rather than benchmark answers* How agents chained sandbox and infrastructure vulnerabilities in unexpected ways* Why increasingly capable agents make traditional security assumptions harder to maintain* The argument behind Anthropic's “Pacing the Frontier” proposal* Why software engineers are increasingly doing two jobs: engineering and keeping up with AI* Constitutional classifiers, probes, and fallbacks and what interpretability looks like in production* How Auto Mode checks whether an agent's actions actually match the user's permissions* Why Thariq can see serious AI risks while still having a relatively low p(doom)Thariq Shihipar* X: https://x.com/trq212* LinkedIn: https://www.linkedin.com/in/thariqshihiparTimestamps00:00:00 Introduction00:04:12 Ask User Question and the Future of Agent Interfaces00:08:29 Artifacts, Projects, and Multiplayer Agents00:15:37 Prompting as the Core Claude Code Skill00:21:52 Context, Effort, and Smarter Model Usage00:28:10 Is Claude.md Going Away?00:32:49 Claude Mods: Customizing the Claude Code Harness00:36:35 Model Routing and the Rise of Mutable Software00:44:40 The Bitter Lesson of Harness Engineering00:50:49 Claude Tag as an Organizational Harness00:55:59 Pacing the Frontier and Autonomous Agent Security00:58:22 Agents Hack Hugging Face for the Scorer01:05:34 What Happens When Agents Need More Compute?01:10:32 AI Coding Is Changing Faster Than Engineers Can Keep Up01:17:17 Probes, Fallbacks, Interpretability, and Auto Mode01:28:32 AI Risk, p(doom), and Closing ThoughtsTranscriptIntroduction: Life at Anthropic and the Pace of ChangeSwyx [00:00:00]: We're here in the studio with our friend Thariq from Anthropic, and I guess generally the Claude Code, I-- there's, there's so much, merging of boundaries and you've been so on top of everything since you joined Anthropic. You have been early to Claude Code itself, but then also, and you've told that story in other podcasts, and you've also been talking about seeing like an agent. Most recently you did the top AIE World Tour talk, Field Guide to Fable, which obviously you guys launched Fable, so that was-- that's cheating. And mostly you most recently also launching Claude Tag, and we're also gonna be talking about Pacing the Frontier. There's a lot going on in Anthropic. I guess top of the question is, what's it like being at Anthropic when there's so much going on?Thariq Shihipar [00:00:48]: I think that It is, like. I think you can get whiplash sometimes. I think, like, going. When I joined Anthropic, I joined because of Claude Code. Like Claude Code had just come out and I was like, “This is so good.” And Opus 4 to me was like just, I could not imagine, like, how good it was? And that was, like, a real moment for me. But I was, like, trying to convince, like, my startup friends to use agentic coding, and they're like, “Oh, no, like, our engineers don't think it's good enough,” or something. And I was like, “That's insane.” and now you, like, fast-forward, 12 months, less, and, like, it's just like, yeah, the default way that everyone codes, right? And I think that, like, just having to go from, like, selling it to, like, now, teaching people how to be. make the most use of it and be more efficient and things like that is just like a big, like big change. And, yeah, I think, like, it's just hard to stay on top of everything as a human? Like, I think things happen so fast and likeSwyx [00:01:51]: You just throw more agents at it.Thariq Shihipar [00:01:52]: Yeah, like that's like the agentic stuff scales much better than the, like, human stuff where it's like, oh, like, there are three things happening right now and, like, they're all emergencies and, like, how do you, like, respond to it? Yeah.Teaching People to Use Claude CodeVibhu [00:02:05]: What do you split your time on? You do a lot of technical writing, engineering work.Thariq Shihipar [00:02:10]: Yeah, so I think that, like, when I joined the Claude Code team, I wanted to teach people how to use Claude Code and I think that, like, that has been something that, like, I thought, like, maybe I would spend a little bit of time on it or, like, I'd, like, do. I was spending some time on the agent SDK first, and I wasn't exactly sure, like, how the bitter lesson would go, when it comes to, like, harnesses, right? Like, I think sometimes we were like, “Oh, like, what's after Claude Code?”? And so initially I was like, I just wanna teach people how to use Claude Code and make it easier to use Claude Code. And I think that has just, like, as the harnesses have gotten better and better, that's like the dominant problem now is, like, how do you use the agents, right? Like, it's like such a high skill expression thing. So I do that and then I do engineering work. I give talks, but I think, like, when I'm doing engineering work, my goal is to take that feedback that we get from users and also, like, then be able to talk about, like, hey, how to use Claude Code to do engineering. So there's like a good loop there. Yeah.Swyx [00:03:07]: Yeah. I'll-- For listeners, we'll attach, the talk that you did with Sarah for the Dev Writers, meetupThariq Shihipar [00:03:13]: Oh, yeahSwyx [00:03:13]: Which we talked a little bit about, well, first you do the work and then you talk about the work.Thariq Shihipar [00:03:16]: Right.Swyx [00:03:16]: Something like that.Thariq Shihipar [00:03:17]: Yeah.Swyx [00:03:17]: It's sow and reap orThariq Shihipar [00:03:19]: Yeah, reap and. Sow and reap.Swyx [00:03:21]: Something like that. Something like that. Yeah, so, and then just to preview a little bit, we are gonna talk about the evolution of the harness. It has come a long way from just being a CLI. We're gonna talk about, Claude Mods, which is starting to leak today, because you couldn't keep it secret.Thariq Shihipar [00:03:36]: Yeah. yeah.Swyx [00:03:39]: Yeah, there's, there's a lot, there. I think you started off with, like, adding ask user question tool, which people love and hate.Thariq Shihipar [00:03:48]: Yeah.Swyx [00:03:48]: Like, I thought it was, like, very innovative, and then now I have, like, my own version. You have your Interview Me version.Thariq Shihipar [00:03:55]: Yeah.Swyx [00:03:56]: And, yeah, everyone just has, like, their own stuff. And, like, it no longer matters ‘cause now you're supposed to, write prompts that create other prompts and loops and all these things.Ask User Question and Human-Agent InteractionThariq Shihipar [00:04:05]: Sure, yeah.Swyx [00:04:06]: So what's the state of the art, today? Like, what are people. what are you, like, telling people to do today?Thariq Shihipar [00:04:12]: Yeah, ask user question was the first time that the model was good at elicitation. I think this was, like, an emergent behavior that I, like, wanted to see if the models could do. I have, like a human-computer interaction background, so I, like, did that in undergrad and grad school. And so this was like. I think it's like human-agent interaction to me, like, trying to figure out, like, how can the agent communicate with you and extract, the requirements, right? I think that, like, one of the things about, like, that's difficult as Claude Code has gone broader and broader is that everyone has, like, their own way of using it, and it's very hard to, like, change the default behavior. So for example, like, if someone asks Claude Code to do something,Thariq Shihipar [00:04:59]: Sometimes they just want them to do the work, ‘cause they're, like, maybe a very good prompter, and sometimes they want. like, are not good at prompting? And you need. like, the agent needs to, like, clarify? And so that's, like, a good split. Like, and the ask you the question tool like, splits along that side where, like, are-- do you feel like you're good enough to instruct the agent as it is, or is the agent able to, like. does the agent need to, like, pull out more requirements and, like, collaborate with you more and really understand your preferences?Thariq Shihipar [00:05:27]: I, on the whole, believe that pretty much everyone is more on the latter than the former, that they, like, have more ambiguity and they know less than they want, than they, like, think they know about the problem. but, like, it's like a interface design problem to make that easy? And so, like, if you're designing a problem, like, or if you're going through a problem, like, things like what's the schema or, like, what's the call stack and things like that are really important. like, the details in the design are important. Ideally, you want to figure out some of these, like, hard problems ahead of time before starting implementation. And yeah, that's why they call, like, unknowns, right? And so I think that this will forever be, like, a skill in agentic coding is, like, figuring out your unknowns. So, like, because even if the model is, like, super intelligent- It, like, needs to know what you want? And, like, you have preferences. like, you need to like, pull the, pull that out. and so that's, like, I think how I'm, what I'm pushing. the question then is, like, how does the agent interact with you? And I think that has been HTML, has been, like, the big way of doing that. And we've recently added artifacts, right? And artifacts, I think we've done a bad job of, like, or, like, I've done a bad job of, like, explaining how to use them fully. We have a lot of property capabilities. They have a database associated with them? And so every artifact can store and write persistent data. They can, like, feed back into Claude? And so, like, one thing that, like, people are not doing yet that I'm trying to, like, encourage is, like, this idea of a dashboard artifact. So you have, like, Claude working on a project long-term. Maybe it's like a kanban or something. it can store that kanban data in its database. Multiple Claudes can access that data via, like, the artifact MCP, and, like, that artifact can, like, talk to those Claudes as well. And so, like, the. We're building the primitives for you to be able to have this, like, generative interface via artifacts that will, like, let you surface more of that rich detail from the agents. And I think that, like, almost everything with agents right now is, like, this problem of, like, you think what you want, but you don't really know what you want, and, like, the agents need a lot of detail, and collaborating with them in the loop is really important. And so artifacts are, like, the, like, way that we're trying to evolve there. But there's a lot of work to do because it's so much more complicated than, like, a multiple-choice question? there's a lot more, like, detail in terms of, like, diagrams and code snippets and schemas or, like, whatever it is for that problem. But, like, artifacts is, like, the mo-more AGI-pilled way of, like, doing ask user question. So yeah.Artifacts as the Interface to the HarnessSwyx [00:08:15]: I think one thing that's unclear to me about these, the artifact stuff is, like, what feedback should go in through the artifact and what feedback should go through a Claude, a chat? Because the more AGI-pilled one is to just feed everything to the Claude.Thariq Shihipar [00:08:29]: I think the more AGI-pilled one is to go through the artifact. Like, and I think that, like, we imagine in the limit, I think that artifacts will be your interface into the harness? You can, like, comment on this, like, live, like, document of your plan, of the work. you can see maybe, like, multiple agents and different agents are doing this, and that artifact is built for the current work that you're doing, right? And so, like, each one has, like, slightly different. I think we're still, like, getting there from, like, an infrastructure perspective. But yeah, I think, like, on-the-fly interface for your harness is probably where things are headed.Vibhu [00:09:03]: Is there a version of it that's an abstraction from CLI or chat and you. Because right now, a lot of it is, okay, you're interfacing with Claude Code, you're having HTML given back for a mockup. It's pretty rich. There's diagrams. Artifacts are ways to connect these together. Why not just do everything that way?Separating Brain, Hands, and Surface UIThariq Shihipar [00:09:22]: Then it becomes, like, separating out, like, where is the inference happening? Where is the intelligence happening? Where is the work happening? like, I think this is like, difference between, like, or, like, some of the distinction between local and cloud, right? And so, I think right now, if you use Claude Code, it's, like, local and, like, you can spin off remote control, for example, to get some cloud behavior, or you can spin off Claude Code in the cloud, right? We're moving towards a place where instead of Claudes, like, you message a local Claude, it starts a session locally and it executes, to more like you have a Claude that you message that's in the cloud that's running. it can run, like, local, or, like, cloud sessions. This is how Claude Tag works. But, like, over time, we'll add, like, local hands as well. And so, like, local hands will be the ability for that agent to access your computer if it's online, and be able to, like, work there. And so it can spin off many different subagents. It can, like, commu- those subagents can communicate with each other, and that's where the artifact comes in to display all of that work. So you can imagine, like, the. You're separating out these things. So there's, like, the surface UI display that's an artifact and hosted somewhere and has a database and everything. There is the inference intelligence, right, that's happening on the cloud, and you don't have to worry about shutting off your computer or whatever, right? and then there's the, like, hands. Like, and it can be local, it can be in, like, a remote sandbox or wherever you need your work to be done. That's like unpackaging, like, the Claude Code experience right now where, like, right now it all happens in one place, right? So.Multiplayer Agents, Claude Tag, and ProjectsVibhu [00:11:00]: How do you see, like, the multiplayer side of that? So say teams want to work in this way. Right now it's very individual, but how do you see the future of multiplayer? Like, right now, I guess there's Claude Tag, which is a version, but.Thariq Shihipar [00:11:12]: We're launching projects. And so projects is the, like, this abstraction that's like Claude Tag, but on our Claude products, right? So you can message it and, like, it will do the Claude Tag-like stuff, like spinning off subagents. So We think with multiplayer. Like, Claude Tag is, like, a little bit more native multiplayer because it's just, like, in your Slack and the permissions are all figured out and stuff like that. But I do think multiplayer is, like, an important part of the story and, like, that will need to get tied together more. Like, you can imagine how complicated it gets when you're like, oh, you have hands, but now you have other hands in other people's computers too, and, like, you need to, like, permission them or, like, you have, like, your MCP and someone else's MCP, and how do you figure out how to use them, right? It gets, like, quite complicated. And Claude Tag does a good job of, like, sanding down all of these issues, right? So that, like, when you have, yeah, Google Docs, how does it access Google Docs, right? Like, it accesses through the shared Claude MCP, or it can access through your local credentials as well if it doesn't have access. But yeah, I think Claude Tag is our multiplayer, product, and it's really useful for these, like, things that are inherently multiplayer. Like, okay, like on-call, for example, incidents are inherently multiplayer. You want to tag Claude, you want multiple people to log in, you want it to be able to find context. I think whenever I'm, like, working on something and I want, like, privacy or security or, like, I want other people to review it's really nice to, like. I'll have a channel per project and I'll, like, at legal, for example, be like, “Hey, like, I want to ship this. Can you, like.” Like, here's. Like Claude knows everything, just chat with it. And that way legal gets precise answers, on like what exactly is shipping into the code, and I don't need to be in the loop, right? So I think like multiplayer is getting like more and more like, yeah, everyone can participate with Claude. I think Claude Tag is like that product and like projects will start off single player and will like, expand.Swyx [00:13:14]: I think there's a question about like maybe dual questions about identity and the unit of isolation.Identity, Permissions, and IsolationThariq Shihipar [00:13:20]: Yeah.Swyx [00:13:20]: Claude Tag, you specifically chose to make it its own identityThariq Shihipar [00:13:26]: Yes.Swyx [00:13:26]: Which is like, a controversial choice. There's, there's other ways to do it.Thariq Shihipar [00:13:30]: Yeah.Swyx [00:13:30]: Claude Projects probably it sounds like, if it's anything like ChatGPT Projects, it is, the isolation is that artifacts, that cloud instance, everyone's collaborating on this. It'll. It sounds like, it should be like if you're, if you're collaborating with legal on a thing, like that channel should be a project, right? Like it's not yetThariq Shihipar [00:13:50]: Yes.Swyx [00:13:50]: But it. that's the natural next step.Thariq Shihipar [00:13:53]: Yeah, like I think in Claude Tag, it's effectively. Like Claude Tag, you have to do your own arrangement. And so Claude Tag, yeah, each channel is like you can name it as you want, and I nameSwyx [00:14:04]: Yeah.Thariq Shihipar [00:14:04]: Like each featureSwyx [00:14:06]: Yeah.Thariq Shihipar [00:14:07]: As a channel.Swyx [00:14:07]: And, but I think like there is some trans- like it's unclear when there is transference, because let's say it is. if you have a coworkerThariq Shihipar [00:14:14]: Yeah.Swyx [00:14:14]: Who is tagging on all these things, yes, there is transferThariq Shihipar [00:14:16]: Yeah.Swyx [00:14:16]: Because it's the same person. but with Claude, it's unclear if it's like necessarily like, well, no, you don't know any of. you don't know about the other stuff. You should only use this stuff.Thariq Shihipar [00:14:25]: It's like the tip of the iceberg meme, right, where you can like. This is what we spend so much time onSwyx [00:14:31]: Yeah.Thariq Shihipar [00:14:31]: Is like there is like infinite surface area of like, okay, you want Claudes to. Not infinite, but like there's like surface area, a lot of like, surface area to figure out of like permissions and visibility and like how can you let Claude operate as well as you can, as safely as you can? And obviously, this is very important to us because like security for our code base is very important. And so we've put a lot of time into this. Yeah, there's so many like edge cases you can figure out where it's like, oh, like, yeah, this Claude in this channel has different permissions, but it can message another channel, and can't it exfiltrate data that way? Or like can you like. What if it uses your MCP and then messages someone else? Like there's like so much, and we've like really put a lot of work into sanding it down.Swyx [00:15:14]: Yeah. Lots of work. okay. Fable?Fable and the Meta-Skill of PromptingVibhu [00:15:18]: Fable, you wrote two good articles. you've written many good articlesThariq Shihipar [00:15:22]: Yeah.Vibhu [00:15:22]: But on, Field Guide to Fable, Building Claude Code. I'm curious from what you've seen, is there any common patterns that you see in like top users at Anthropic externally? Like what are best practices for getting the most out of Claude Code?Thariq Shihipar [00:15:37]: The like meta skill I say is like prompting is like very important? And like that. Like I think this is like not trivial to say because I think a lot of people are like, “Oh, prompting doesn't matter. It's just like I can just say a sentence and Claude will do it.” And I think prompting is really this like, this. It's like public speaking, like, or writing or something, and for a specific audience, and that audience is Claude. And you need to like build a mental model of Claude and how it thinks and how it works, right? And so that's like the most important skill in working with Claude Code is like having this mental model, right, of Claude and like what it can do well, what it can one-shot, what it can't. And so many people when you see prompting, they're just like, they're short prompts, but they have such a good mental model of Claude and of like the code base and things like that like it's effortless? But it's like high skill ceiling. So like that work of like, spending a lot of time prompting and building mental models of how, and intuition for how the agents work is really important. And then I think like the next thing is like the unknown stuff we talked about earlier, where it's like being able to find out like your, what you don't know or what you haven't written down, learning about like different things. I think as Claude can do more and more things, the likelihood of you doing something out of distribution for you and like you have low domain knowledge on is very high? And the more you can like learn the vocabulary to be able to prompt Claude, it becomes really important. And so like I think the most important unknowns are the unknown unknowns, where you're like, I just like don't even know that this exists, right? Yeah, exactly. I think that's like a illustration of like the map and the territory, right, where you're like, “Okay, this is my prompt,” and the territory is like the actual like work that the agent needs to do, right? And if you are like very precise, you can give more precise things, right? So like for example, in design, I'm not very precise. I'm not a designer, so I say like, “Give me like eight different mock-ups.” But if I was a designer, maybe I'd be like, “Oh, hey, here are some reference sites.” Like, “I want this type of font and this type of like look to it, and here's like a few different components to like visualize. Here's a Figma MC board to bring in,” like. And so you can just be so much more precise with that language. And if you're not a designer, you just need to like try and learn the language or learn the unknown unknowns. And this is true of like everything, I think. Like the more, like you can work with Claude to learn like how things work, the better your prompting will be. I think another good example of this is like game design, like where a lot of people are like, “Oh, like I can vibe code a game now.” And they're like, “It's not fun.” And like it's just like the thing about game design is like every one of these choices has like a lot ofTaste, Domain Knowledge, and Learning the VocabularySwyx [00:18:25]: Variations.Thariq Shihipar [00:18:25]: A lot of like craft to them. So it's like, oh, okay, like when you're making a flying game, the feel of the plane and the like, way it responds to your controls has a lot of like. Like, a game designer would spend like days on that. Do? and likeSwyx [00:18:44]: To me, that's what taste is, right?Swyx [00:18:45]: Like it is like from the possible space of one thousand mathematically valid answersThariq Shihipar [00:18:49]: Yeah.Swyx [00:18:49]: Here's the one that is the humans will like.Thariq Shihipar [00:18:51]: Yes. Yeah.Thariq Shihipar [00:18:52]: I think with taste, I'm like torn on this word ‘cause I think you're right, but everyone has different definitions, and it sounds kind, sounds like low skill or like elitist almost, where you're like, oh, like there are certain people with taste?Swyx [00:19:06]: It's like taste is what I call taste.Thariq Shihipar [00:19:07]: Yeah, exactly.Swyx [00:19:08]: And it's like these guys don't have taste.Thariq Shihipar [00:19:09]: Yeah, exactly. Oh, like an engineer doesn't have taste. Like I, the like founder, have taste.Thariq Shihipar [00:19:14]: ? And I think that's not true. Like I think like the engineers have a lot of taste for these particular like problems? And I think everyone has taste for particular problems. I think like Jason Liu, like say like in order to, yeah, have taste, you have to eat?Thariq Shihipar [00:19:32]: And I really like that, where it's like, okay, you have to like do a lot of things. You have to like iterate and figure out what you want, what you like, and, like build that like domainSwyx [00:19:41]: YesThariq Shihipar [00:19:41]: Domain vocabulary. And then when you're prompting, you're like synthesizing all of that for a product.Swyx [00:19:46]: Isn't it annoying when someone else says it better than you?Swyx [00:19:48]: It's just like, f**k, I have to quote this guy forever.Vibhu [00:19:51]: Having to quote Jason Liu forever.Vibhu [00:19:53]: He's gonna love this.Thariq Shihipar [00:19:55]: So I get prompts, more than that.Vibhu [00:19:57]: And sometimes it's not even that. Sometimes it's just intuitive, right? Like you don't realize you even want something till a model puts it out, and you're like, “Oh, this just feels immediately better,” right?Voice Prompting and Information DensityThariq Shihipar [00:20:07]: Yeah, exactly.Swyx [00:20:09]: One thing I go back and forth on is I feel like the way I prompt half the time, let's say I use voice.Swyx [00:20:16]: Did I say voice? Other people have voice. that is the opposite. That is just like me rambling for like two minutes Pressing down the function key and then let go, and then like hopefully it figures it out. And oftentimes it does.Thariq Shihipar [00:20:26]: Yeah.Swyx [00:20:26]: But it's not as thoughtful as like a structured prompt with like Well-run communication as though it's a PRD or a memo. Is that in line with how people do this? There's like bimodal prompting where there's some prompts where you spend a lot of time upfront and other prompts you just dash it off?Thariq Shihipar [00:20:43]: I don't think the voice is necessarily low. Like I think it's like more like how much information is in the prompt. like the model can. Like you can and like add some sentencesSwyx [00:20:53]: RightThariq Shihipar [00:20:53]: And be like, “Oh, like I changed my mind,” like in the middle of the prompt, and it will be able to follow that perfectly? So I think the like actual format of the text is less important, but then like the ability to. Like how much information is in it, right? And I think for voice, a lot of times, going back to like human-agent interaction and like for a lot of people, it's just way easier to talk than to like type? and I. If that gets more information out of you, like that's better.Vibhu [00:21:21]: At some level, it feels like just giving the model as much contextThariq Shihipar [00:21:24]: YesVibhu [00:21:24]: Over prompting before you kick off is a best practice. I don't know. A lot of the times, like when I was first trying out Fable, I spend a solid 30 minutes like really crafting a long prompt. This, I think, is a response of models running for longer and longer, right? It's still a little difficult to nudge them as they're in like, in the loop, but I just like intuitively spend more time kicking off that first prompt and working with it a lot.Spend More Upfront, Iterate LessThariq Shihipar [00:21:52]: My personal opinion is that if I was a software engineer, if I was like, just running my own startup, for example, I think I would mostly fit, stick to a max 20x? like maybe verification and so code review are like separate things. But I think like what I see a lot of times is people hit rate limits when they're doing this like, oh, like it did a lot of work and you're like, “Oh, I don't like this.” Like, “Can you like undo this and redo it?” And then you're like iterating on this like thing that the model could have done if you had like spent more upfront time or given it better context? And instead it's like you're like, “Nope, don't like that design. Try this.” Or like, “You messed this up,” or something like that. And then that just eats up so much more of like, your usage. And so that's like, I think maybe like a key like tip both for like efficiency as well, right? And yeah, I think like context, and not just like context on like what the goal is good, right? Like are you building a prototype or is it like a production thing? Like where can you spend compute or when, where can you not spend compute? Like I think you have to give the model permission or like not permission to do things sometimes where, like it doesn't know intuitively how much you want to spend on this task, right? And you can use effort for this. So I did-- I'm working on a blog post about that where it's like, if you want. For like we see that effort scales with the complexity of the task. So for security, effort gets like way more results. Like high effort versus like low effort gets, like changes the evals a lot. But for software engineering, it doesn't change it a huge amount because effort is mostly spent on the verification and the like edge case testing and things like that. And so like being able to like give the model that guidance of like, “Hey, this problem is something that I think I want you to spend a lot of time verifying and edge case testing,”?Effort, Model Choice, and VerificationVibhu [00:23:43]: How about model in the mix? So, there's Opus and Fable with effort.Thariq Shihipar [00:23:47]: Yeah.Vibhu [00:23:48]: There's also Haiku in there.Thariq Shihipar [00:23:49]: Yeah. It's not quite true yet, but it's very close where I think the frontier models will be Pareto dominant over like almost everything. like maybe. And sometimes I think Opus might be Pareto dominant. Do? Like I think depending on like how things, like shake out if it's like a newer version of Opus. But I think that like increasingly it's just going to be like the smart model is going to be able to like do the simple task for less tokens than the like the other models because of verification. With verification, in the limit, your model doesn't need to verify, right? If it's a perfect model, it just does the work once and it's like, okay, like you, I did it? And increasingly with Fable, I'm like, I'm like, “Dude, you don't need to spin up Chromium and screenshot all of these things.” Like I see it. Like you did it, right? And so a lot of the. At higher effort, you spend more of those tokens verifying. But if you're working on simpler problems, and a lot of software engineering is like well, like in Fable, like low and medium stability, it can spend less tokens verifying. And as the models get smarter and smarter, they will just be able to like, “All right, done.”? Like, I can run the lint for sanity's sake, but, like, I, like, know it lints? Like, you don't even need to do that. And that will be so much more token efficient than, like, the smaller models. Yeah.Swyx [00:25:15]: Is there a good, practice on our side that we can use to see if we're using too much effort? Like, I freakingThariq Shihipar [00:25:23]: YeahSwyx [00:25:23]: Hate wasting time on that stuff.Thariq Shihipar [00:25:24]: Yeah. I know what you mean. I think, like, so in this blog post, my rough distribution is, like, code review and security should be, like, high or max and, like, software engineeringSwyx [00:25:37]: You said recommend mix settings per domain.Thariq Shihipar [00:25:37]: Yeah. I think, like, if you're doing, like, UI or something like that, like low and medium, I think is you're building, like, an API and you want to make sure, like, you cover enough edge cases? And so I think building, like I said, that mental model of, like, how things work across these distributions is, like, yeah, part of the job.Implementation Notes and Decision LogsVibhu [00:25:56]: This is more intuition-driven or eval? Because I'm guessing this would change as you go.Swyx [00:26:00]: He has evals.Thariq Shihipar [00:26:01]: Yeah. So what I did in the blog post is I go over all of the terminal bench evals. So there are, like, 70 problems and I'm show that, like, okay, like, in the security problems it does more. and then I also, like, look at some of the transcripts just in terms of, like, how-- what does it answer, what does it forget or something. And a lot of times, this is another prompting tip I have, is, like, asking it to make decision notes or implementation notes because, in every eval problem that it faces, it thinks about the correct solution, and decides not to do it. it's like, oh, like, here is the answer. What if I did this? And then it's like, oh, probably not? and then keeps going. And this is, like, the majority of the failures, at, like, a higher max level. It's very rare that the model just doesn't know how to do something. If you just have these implementation notes, then you can review and you can be like, “Oh, I want you to do this thing that you didn't do.” The models are getting better at surfacing that overall. Like, I see in the transcripts of Fable 5.1, like, when it does this output, it will call out its decision-making as well. but making this more explicit in the harness is better. And now we're, allowing ways of you modifying the harness so you can, like, add someVibhu [00:27:23]: Ooh.Thariq Shihipar [00:27:24]: Calculate with there. Yeah.Swyx [00:27:25]: Yeah. So I do wanna call out two things that you mentioned that I think exist outside of prompting. One is like, let's, let's call it the prompt that is so important that it shouldn't be in a prompt. It is in Claude.md or Agents.mdThariq Shihipar [00:27:38]: YeahSwyx [00:27:38]: Which is like goals, right? Like your situation, your goals, the things that you want, the thing. and then second of all is the decision log or the experiment log or whatever log of traces that you might want to survive the current session to do those things. Those are, like, externalities that there's no standard. There's no-- It's not like skills. It's not like MCP. There's no standard. It's, it's just like it's a markdown file. first of all, is that right? Is Claude.md going away? You have a documented dislike of, Agents.md, but you're gonna do it?Claude.md, Agents.md, and Model-Specific InstructionsThariq Shihipar [00:28:10]: Yeah. Okay. So Agents.md, yeah, like, we're, we're gonna do it. I think it's just, like, different models are very different from each other? But I realize that it's, like, such a pain to, like, maintain different ones? And yeah, like, as the models get better and better, the floor of how they accomplish the simpler task is better. And so I do think in the limit, Claude.md goes away, and maybe not even, like, that far. Like, I think, like, I think that right now it might be better to start a new project without a Claude.md.Swyx [00:28:44]: Yes.Thariq Shihipar [00:28:44]: I think that, like, maybe if you see very repeated failure modes, you add them to your Claude.md. The really tough thing is that this changes per model. And so, like, if you've added a bunch of failure modes or, like evenSwyx [00:28:57]: So you need Fable MD, you need Opus MD.Thariq Shihipar [00:28:59]: Or well, even Fable 5.1 versus Fable 5.Swyx [00:29:03]: Yeah.Thariq Shihipar [00:29:03]: Like, it is annoying. Like, I'm not like,Swyx [00:29:05]: YeahThariq Shihipar [00:29:05]: Like, we don't, like, do this on purpose? It's just, like, how the models work, right? And so, like, maybe, like, Fable 5 had this, like, failure mode that Fable 5.1 doesn't. And if you keep this context, this running log of a bunch of different failure modes, they will probably over constrain Claude? And so this is like. we just added evals plugins for skills.Swyx [00:29:28]: Yeah.Thariq Shihipar [00:29:29]: And so now you can eval if a skill is better. I think Daisy on our team did this. And so, yeah, this is like we're trying to work on this. We know it's, like, you still have to spend tokens on it and, like, it's not, it's not perfect, but it's, like, we're trying to help out with this problem.Swyx [00:29:44]: And so, and as far as prompting goes, the one tip I wanna offer is, something I have told people a lot is sufficiently advanced prompting is indistinguishable from sufficiently advanced executive communication. So I've referred to-- This is an executive comms workshop from Heavybit that is the best I've ever seen in my career. And they teach this thing called the SCQA model. Just Google it. It's a, it's a thing. Like, people have done prompting for decades. It's just called executive communication. It's like when one person has to communicate to thousands of people down the org chart, this is what you do. so situation, complication, question and answer, is how you write the memo. but obviously sometimes you don't have the answer, but you can at least list out the SC and Q, and then they have some examples in there. So just leaving breadcrumbs for people if they want to explore.Underrated Prompting Patterns and ELI5Vibhu [00:30:31]: Before we move on, I wanna ask you, any other underrated tips, ways people could get a lot of value from Claude Code that they're not using?Thariq Shihipar [00:30:41]: Yeah, I think a lot of them are in the, this unknowns, like, doc. Like, I give a bunch of example prompts, like, using it for brainstorming, using it to quiz you after. we added this, like, explain it like I'm five skill which is a very short prompt. And it doesn't even say explain it like I'm five. It's like the key word of this prompt is big pictures, few words. like, that's like the main thing. And it is shockingly good? Like, you, like, I think I tweeted about this and it's like /eli5, and, like, you can install it as a plug-in. But yeah, it's, like, way better at just cutting through the BS and being like, yeah, exactly right here. So the diagrams are, like, quite clear. I think one of the things that is true with artifacts is, like, they put too much text in and people are not reading the artifacts? And so, like, this simplifies it a lot more. And, yeah, this came out of, like, just people at Anthropic, like, going through very complicated incidents and being like, “What is happening?”? So, this one I think is great, yeah.Swyx [00:31:47]: My version of this is the, it's like test your understanding. Give you a few choices and then, like, if you get it wrong, you have a mismatch between what you think is happening versus what's happening.Thariq Shihipar [00:31:58]: Yeah. I think this is one of those things that everyone loves talking about, and then very few people really do. Like, I thinkSwyx [00:32:05]: Really helpful.Thariq Shihipar [00:32:07]: Yeah. But most people just don't want to get quizzed about something? Unfortunately, I think this is one of the, like, things that we need to, like.Swyx [00:32:16]: What's the opposite of ask you the question or ask you the question before the thing?Thariq Shihipar [00:32:19]: Yeah.Swyx [00:32:19]: This is after the thing.Thariq Shihipar [00:32:20]: Exactly. Yeah.Vibhu [00:32:21]: It's a good way to stay grounded of, like, do you even know what you're doing, right? The worst case is when people send you slop and they haven't understood what they're asking for or what the output is, and it's like, “Dude, I don't wanna read this. Do you even know what it is?” So, you make it a rule for yourself that before you send stuff, you should at least know what's implemented.Claude Mods: Customizing the HarnessThariq Shihipar [00:32:41]: Yes, but so you could make this a mod and you could build your own mod to, like, make sure you test it. So yeah, you can do that.Swyx [00:32:49]: All right. Let's get right into it. What is Claude Mod, and what is this diagram showing?Thariq Shihipar [00:32:54]: Yeah. Okay, so Claude Mods is you can customize the entire Claude Code harness, and we're going to. If you have requests, we will, like, let you, like, please let us know. We'll add more and more. This works for CLI, it works for desktop. maybe it will work for Claude Tag in the future. I don't know. Like, we're trying to make this very extensible. You can see this reference sheet. I don't want people to get overwhelmed by it? At a high level, you can customize both the execution of the harness, and the UI of the harness. And so, like, you say on that Tetris example from Boris, that's like customizing the UI, right? Like showing, like, Tetris in the game.Thariq Shihipar [00:33:35]: But, like, let's say that you wanted to do this thing where you had. you tested your assumptions or, like, tested your understanding after every project, right? What you would do is you would ask Claude to make this plug-in. It would spin a classifier after every prompt. And so, like, at the end of each turn, you would spin off a sub-agent or, like, a forked agent. A forked agent is, like, maintains the prompt cache, right? So it's like a, like one of those unintuitive things where you can fork and do, like, a little request, and it'll be very cheap because the entire prompt cache is, like, done. And so you can be like, “Has this task been completed?” likeSwyx [00:34:18]: This is how you do BTW and all those.Thariq Shihipar [00:34:20]: Yeah. The underlying forked agent, yes. But so you can, in the f-fork sub-agent, you can say, like, “Has this task been completed? If so, return true.” And then in your hook, or in your, like, plug-in mod, or sorry, like, in the sub-agent probably, you would say, like, “If true, give me a quiz.” give me questions and answers, and then, like, in a JSON format, and then you'd parse it, and then you display above the prompt input, this list of questions, right? And so this is something that's, like, slightly token-intensive because, like, you have to do it after every end of the assistant turn. But it's, like, a lightweight classification, and then you can, like, get this quiz, and then you'll see, like, Claude will always do it for you. You don't need to remember to do it. There are lots of these, like, tips that we've talked about, right, where it's like, oh, implementation notes. You can also add a tool for implementation notes now. And so, like, this tool that I'm adding is, like, register, like, I think assumption is what I'm calling it, but, like, maybe I'll change it around. And this is a mod. And so, like, you give it a register assumption tool, and then it will keep a list. It'll. Every time it does it'll keep a, like, add to the list, and then at the end it will display those assumptions? Another mod I'm working on is a model router. And so, like, internal, like, Claude model routing, right? So it's. This is, I want to say the reason we don't do model routing by default is, like, it's a hard problem? And likeForked Agents, Assumption Tracking, and Model RoutingSwyx [00:35:51]: You will get it wrong.Thariq Shihipar [00:35:52]: Yeah, you, like, yeah, you will, like, accidentally use, like, Fable for a hard problem or Sonnet forSwyx [00:35:57]: Yeah, if you have auto approve, but you don't have auto mode.Thariq Shihipar [00:36:01]: Well, you will have auto. Like, you don't have, like, auto routing or something.Vibhu [00:36:04]: You don't have auto mode for model picker.Thariq Shihipar [00:36:06]: Yeah, exactly. SoVibhu [00:36:07]: I'm getting the rough question of, like, how much do you open this up and how much do people have to think about this? Like, when you talk about prompt caching and building a router, it seems like you could easily build a mod that routes per query, and I'm just killing my plan very fast, right? I guess my question is more so, like, what is, like, a product talk like this look like, right? Who is it for? Is it for power users? Is it everyone should be able to go throughSwyx [00:36:33]: Oh, definitely power users, right?Thariq Shihipar [00:36:35]: Yeah, I think it is power users, but, like, the nature of Claude Code is that so many people are power users? Because it's easy to share things, like you can. Like, one person can make a good model router thing that doesn't break prompt cache all the time, and then you can, like, compose them. Another cool thing about the plug-ins is that they can hook into and compose with each other. And so I have, like, a mod that will, like, create a mode selector at the top, and any plug-ins can register to be a mode. And so, like, the auto router can be a mode, right? Or, like, you can have a mode that's, like, artifact mode, where it's like it primarily talks to you in artifacts. like, you can toggle between plan mode? And so, like, you can create more and more of these modes. But the ability to create modes is in it itself a mod? And so there's a lot of richness here, but we do want to make it fairly easy. We want to be-- make it so that you can just, like, install someone else's. You can ta-- you can chat with Claude and, we'll, like, make sure that it understands the nuances of things like prompt caching and stuff, so it can, like, warn you. This is, like, not extremely complicated behavior for Claude, I think, but we should have just a good skill on how to make mods. and yeah, we'll see how we go. But I do think that this is, like, a preview of, like, mutable software, and, like, how, like, generative software, just like you can customize safely. If enabled, you could customize any piece of software. And I think that more and more apps ideally do something like this?Power Users, Modes, and Mutable SoftwareSwyx [00:38:13]: And by the way, you, we have, you have another cool tweet about how, there's the infinite money button, which is like make your SaaS, consumable by agents. I think mutable software is interesting and, other people have also tried to do it. I think the hurdle comes when you can do everything, then people, users get, tend to get confused. So usually the stuff that works is just like one opinionated flow. This is in the side of less opinionation. It's just like, well, more power to power users. And I think probably unlocked by AI, where, like, you can just prompt for whatever the thing is.Thariq Shihipar [00:38:47]: Yeah, or there can be a skill that gives the opinions?Mods vs. Hooks vs. ArtifactsSwyx [00:38:50]: Yeah.Thariq Shihipar [00:38:50]: And then, yeah.Swyx [00:38:51]: So knowing a little bit about, like, TypeScript and build systems and all these things, the closest-- I'm very curious that the team who worked on this, if, I don't know how close you were to them, if they drew any inspiration from build systems like Babel, Webpack, all these, like, old school things. Because it sounds very similar, like the plug-in ecosystem of those things where they can compose with each other.Thariq Shihipar [00:39:11]: Yeah, I'm not deep in the technical details, but I do know it was a collaboration with someone on the Bun team and someone on the Claude Code team.Swyx [00:39:17]: Yeah, it's a build system mecca.Thariq Shihipar [00:39:19]: Yeah. Exactly. It's, it's very exciting. But yeah, like, agents can just do this very complicated like, extensibility into your software now. And so, yeah, like, another reason to, like. If you run a startup, like, you can just prompt Claude and be like, “Hey, like, could we make an extension system? Like, what would that look like?”?Swyx [00:39:37]: Yeah.Swyx [00:39:38]: And I just really wonder, like, you had hooks in the past and plug-ins, all these things. So what specifically will mods be able to do that those things could not do?Thariq Shihipar [00:39:47]: Internally, we were originally calling this function hooks. And so, like, that's, like, gives you a little bit of an idea where, like, hooks register a, like an event to happen and then, like, a script to call. And this inside of the, like, TypeScript runtime is running things. And so, like, you get some benefits of just, like, it has a bunch of things in the Scope with, like, for example, like how many turns is in this conversation, right? Like, how many tokens have been used? Like, et cetera. Like, what are the messages? Things like that. So it has a bunch of messages that can be used. And then it's just, like, a lot more hooks. So we have, like, or a lot of, lot more, like, things you can register on. And then you can do because of the. because it's all happening in process, you can, spawn sub-agents, with four contests and contexts and stuff. And, like, that will return. You can parse the results of those. You can use structured output to like, return them. and then you can modify the UI, which you can never do in hooks. So, yeah.Swyx [00:40:50]: Yeah. Yeah. So modify UI, this is why you showed the Tetris example. Does it also ex-extend to artifacts? I assume it does.Thariq Shihipar [00:40:57]: You-- Like, artifacts are like a different way of customizing it. like, you can definitely. One of the mods I'm working on is, like, this dashboard mod, which will, like, prompt Claude to maintain a dashboard, that's an artifact. But they're like, slightly orthogonal, or not orthogonal. They compose with each other in different ways. Like, mods are, like, a little bit more, like, in your Claude Code harness, changing the agent loop? And, like, the UI is, like, an added benefit. and then artifacts are just like you want to, see things at a high level, very inter- highly interactive. like, the affordances can be a lot bigger than, like a TUI or even in our desktop.Next Steps, Supervisors, and Persistent GuidanceVibhu [00:41:40]: I'm guessing you'll have a good blog post on the differences, because right now you can also, make a loop that outputs to an artifact that's an interactive dashboard, but you can also do it with a mod. There's just some thinking about making a hacking on a harness when we don't know much about the harness, right?Thariq Shihipar [00:42:00]: Well, something I'm excited about with mods is, like, there's so much things with Claude Code that you just have to remember? You're like, “Oh, like, let me do this, and then let me call the dashboard skill that does the loop,” and things like that. And, or like, “Let me test my assumptions afterwards.” And I think, like, if you do all of these things using these little classifiers and stuff, and you're like, “These are the things I care about. This is what I want to do,” you can, like. You don't have to remember as much. One more, like, mod I'm working on is a next steps mod thatSwyx [00:42:28]: I have-- I was gonna say, I have a next step skill. I always run next steps.Thariq Shihipar [00:42:32]: And does it have access to your skills? Like, this is one of those things where I'm like.Swyx [00:42:37]: I think so.Thariq Shihipar [00:42:38]: Okay. Yeah, probablyVibhu [00:42:39]: Do skills need specific access toThariq Shihipar [00:42:41]: Well, I think there'sSwyx [00:42:41]: Don't they always haveThariq Shihipar [00:42:42]: I think there's, like, specific prompting, I guess, to, like, know your skills. Like I think Claude forgets them sometimes throughout, like, the thing. But anyways, the idea of, like, yeah, next steps that also are like, “Oh, hey, this has happened. Use the explain skill to explain to you what happened because this seems, like, quite complex,”? Or, like, yeah, “Use your unknown skill. It looks like you are, like, asking the model to, like, iterate on these small changes. It seems like you could prompt better.” like, “What if you did this?” Right? So, I think, yeah, like spending more compute there. Yeah.Swyx [00:43:20]: And it should always come out as multiple choice. we have, I haveVibhu [00:43:23]: We have his skill.Swyx [00:43:24]: My next step skill is like this.Thariq Shihipar [00:43:26]: Okay, perfect. Yeah.Swyx [00:43:27]: You can steal it.Thariq Shihipar [00:43:28]: Yeah.Swyx [00:43:29]: Like, but like, for me, it's all-- I think models really always need to be reminded, what are you trying to do here?Thariq Shihipar [00:43:35]: Yeah.Swyx [00:43:35]: Look at the whole transcript and go like, oh, was this original goal? Did your solution solve it? Were you lazy? If you're lazy, maybe there's a reason. Maybe you needed approval from me. Maybe you needed, there's two things you wanna suggest. So it's, it's a little bit like the modification of the ask user question or interview me skill. so it's next steps.Thariq Shihipar [00:43:55]: Yeah, exactly. And again, the benefit of doing it with mods is you can do it as a fork sub-agent, and so it doesn't remain in the context afterwards. So you have this, like, idea of like, okay, the model is doing its execution and you have this almost like supervisor, like, that is like making sure that you can do like the next steps well. So yeah.Swyx [00:44:15]: Yes. I do have two panels and like I often try to have a supervisor thing, keep the high-level context and then the implementationThariq Shihipar [00:44:21]: YeahSwyx [00:44:22]: Detail in another agent.Vibhu [00:44:23]: I feel like a lot of this abstracts away as models change? The, like, half an hour ago you said bitter lesson of harness engineeringThe Bitter Lesson of Harness EngineeringThariq Shihipar [00:44:31]: YeahVibhu [00:44:31]: And we're on the other extreme right now, I feel.Swyx [00:44:33]: Well, so yeah, exactly. If everything's customizable, what is Claude Code, right?Thariq Shihipar [00:44:37]: Yeah.Swyx [00:44:37]: And which I talked to you about last night.Thariq Shihipar [00:44:40]: Yeah, I think that this is. I think the bitter lesson is unintuitive? In terms of like. Also, like we're misusing a little bit of the bitter lesson here where it's like, it's more about like scaling and compute and stuff. But like, I think there is something where it's just like. I think I use it as an approximation here to say that harnesses go out of date very quickly? And like how, but how they change is unintuitive? And so like the big obvious example is like from chat to like agents where you had to give them entirely new tools, right? But like, I think this new version of like, oh, it can modify its own harness, right? This is like, an own harness loop is like a way of using its capabilities, right? Or like it can build an artifact. And like, I think the way I think about it is like the models have more and more intelligence, and they're like so much more intelligent now than like the average software engineering task. Like, you look at the like terminal bench ones and they're like solve like the Jacobian conjecture. Not really, but like, it's like they're, they're quite complex. Like, I would not have been able to do this really as a software engineer.Swyx [00:45:42]: And you said TB4 or TB2?Thariq Shihipar [00:45:43]: TB3. TB3.Swyx [00:45:44]: TB3.Thariq Shihipar [00:45:44]: Yeah. They're quite complex, but the goal is still to deliver user value, right? And like you said, there's like this infinite space of things to do. And so the ways like you spend compute are to keep the user in the loop and make sure that like you're getting to the right decision in the end of the day and like the right output. And artifacts and mods are this way of like spending that intelligence. and I think that's like, yeah, the next step. And so, yeah, I think Claude Code is like, has the core things of agent loop which are, have gotten more complicated. It's like, it needs a sandbox to operate safely. It needs auto mode to like make sure like the permissionsVibhu [00:46:21]: Approvals.Thariq Shihipar [00:46:21]: Yeah, approvals. it needs computer use and MCPs and like all of these like ways of accessing your data, and it needs web search and web fetch. And like, so the-- as the models can do more and more, the core harness has to be like quite complex and very secure. But then like how you interact with it can change quite a lot.Vibhu [00:46:42]: What other harness engineering best practices have you, from the Claude Code team itself? I feel like, there was a phase of plan mode, which is not as used. We now have auto mode. at a point you cut the majority of the system prompt, you got rid of examples. What other best practices are there for harness engineering?Core Harness Primitives and Managed AgentsThariq Shihipar [00:47:02]: I think there is like a forking path where at some point, eventually, yes, the model will just be able to like vibe code the exact version of Claude Code, even describing all this complexity that I've talked about, right? Like auto mode and computer use and stuff. Eventually, the models will just be able to do that in one shot. But I think they can one shot simpler harnesses? And so like, I think some people. Sometimes you don't need this full, like if you don't need computer use or like all this like more complicated stuff. I think before we, you had to use things like the agent SDK, which was like Claude Code wrapped, in order to like. And I would, like suggest people do that because there was so much complexity into building a harness. And now as that's got more abstracted, we have like, Claude managed agents, which lets you have that complexity, but still like, right, like a very bare bones like harness that's scoped to your task. Yeah, I think there's like this barbell effect where like for like very complex, for like coding task and like these like complex things, you should use our harness. And then for like a lot of like simpler or like, more domain-specific things, you can build your own harness because Claude has gotten better at building harnesses, and we have these harness primitives like managed agents. So yeah.Swyx [00:48:18]: Yeah. Is there a general progression? Let's say chapter one was ultra code dynamic workflows, then chapter two was cloud mods. Where is this going?Swyx [00:48:29]: Where you're, you're, you can customize the thing on demand.Thariq Shihipar [00:48:36]: Yeah. I do think that like this evolution of projects and like artifacts and splitting out like brain and hands and, surfaces is like where things are going more. And like, I think it's like not all quite there. partially it's like a, it's just like more token expensive? And like, I think likeProjects, Local Hands, and Cloud-to-Local HandoffsSwyx [00:48:59]: Why would projects be more token expensive? I understand mods would be slightly more token expensive. No, not something I'm worried about.Thariq Shihipar [00:49:06]: Yeah.Swyx [00:49:06]: But whatThariq Shihipar [00:49:07]: You're asking Claude to do. It's like creating loops. Like you're asking Claude to do more work for you. And so like it's managing the sub-agents and reviewing it, versus where you would be doing that work normally. And so that's like gonna be a little bit more intensive, like. Outputting to an artifact is gonna be a little bit more token-intensive than, like, outputting normally. I don't think it's too much more, but like, it's like combining all of these together well, like I think we're, we're still working on like local hands and things like that, I think is like, yeah, where things are headed, yeah.Swyx [00:49:37]: Yeah. Claude and local is, handoff is very interesting. I was thinking about this as reverse cloud remote.Thariq Shihipar [00:49:44]: Yeah.Swyx [00:49:45]: Because it's like remote, it's you're handing off to cloud, but here the cloud is handing off to local, right?Thariq Shihipar [00:49:49]: Yeah, exactly. Yeah, remote control is also another way of doing it. And I do want to say this is like how I think about it and like what the things that I'm most excited about this, but like there are, just like lots of different ways to work with Claude. Like some people use remote control a lot, some people use Claude Code on the web a lot. Obviously, like at Anthropic, we use Claude Tag a lot, and like what's great about Claude Tag is we set up all this stuff for our own execution. And I do think if you're an enterprise, that's still the best way to go. but if you're like an individual, Projects is this way of like, getting some of that like niceness of Tag, which has like that like supervising agent and yeah, adding artifacts and stuff, but like without having that whole like admin setup. And so there will be many ways to use Claude, I think. I think it's probably not just one like single.Claude Tag as an Organizational HarnessSwyx [00:50:36]: You had the multiplayer thing here. Let's, let's just check in on Claude Tag. it's been about two-plus months. Lots of, public, adoption and trying it out.Thariq Shihipar [00:50:45]: Yeah.Swyx [00:50:45]: What's new? What's, what have you found since the launch?Thariq Shihipar [00:50:49]: Like, Claude Tag is how we useSwyx [00:50:51]: It's like 80% of your
In this episode, Phil is joined by Ben Sherman, Jon Manning and Rike Hanssen to take a fresh look at Nextflow plugins, two years on from the plugin walkthrough in episode 35. They cover what plugins are for: sharing reusable functions across pipelines, hooking into workflow events with observers (e.g., nf-slack notifications for nf-core full-size tests, or carbon footprint tracking with nf-co2footprint), and building custom executors and file systems. Ben explains what changed in Nextflow 25.10: the new nextflow plugin create command to scaffold a project, far less Gradle boilerplate, make install for local testing, and make release to publish to the plugin registry at registry.nextflow.io, with a one-time approval for the plugin name. Older Nextflow versions still read the legacy JSON index, which the registry keeps updated for plugins that support them. Rike demos nf-emoji, her first plugin, with themed emoji progress bars, special-day countdowns and optional terminal confetti, and talks about using AI to write tests once she knew her way around the code. Ben explains how custom config scopes let the Nextflow language server validate plugin settings. Jon introduces the new plugin development side quest at training.nextflow.io, covering scaffolding, custom functions, trace observers, testing and config scopes. The episode wraps up with how plugin name claims are reviewed, the expectation that community plugins are open source, and options for distributing private plugins internally.00:00 Podcast Episode 58 - Plugins00:09 Introduction01:01 Welcome to our guests01:48 An introduction to plugins06:14 Common plugin use cases08:25 What's new in Nextflow plugins since 202417:11 Demo: Rike's NF-emoji plugin30:23 Global installs & custom config scopes33:30 The new plugin training material42:34 Publishing and approving plugins47:50 Closing thoughts
I could not find how to cancel limit orders in my Safe wallet. Claude Code built a JSON transaction file; I imported it, ran a simulation and signed the cancellation myself.I also share my loss-making Token Trade experiment, why I prefer chat to complicated dashboards, and the possibilities for agent payments. The closing discussion explores AI wallet risks, decentralized travel and Trips Community, and a future where we express our intent to an agent. I am not claiming AI wallets are safer or right for every use case.Collaborations: info@web3intravel.comCHAPTERS0:00 Intro0:20 Wallets and the adoption hurdle3:40 Convenience, custody and attention5:26 The Safe limit-order problem7:54 JSON, simulation and signatures11:23 Building Token Trade14:15 Trading losses and on-chain exchanges19:12 Building tools through chat20:42 Why dashboards drift23:13 Stablecoins or credit cards?25:40 AI wallets and security trade-offs26:50 What agents could change in travel28:35 Intent as the interface
Podcasting 2.0 September 25th 2026 Episode 272 - "Is this Y'all?" 00 -
Arranca un nuevo episodio de Potencia.pro, el podcast de WordPress y cozas, con Miguel Ángel Terrón y Mariano Pérez al micrófono. Es el capítulo 336, segundo de la temporada… ¿undécima? ¿Décimo primera? Tras un pequeño debate lingüístico, y alguna broma de las que marcan la casa, lo dejan en «décimo primera», que suena menos a cacho de algo. Un plugin exclusivo… que nadie descarga Miguel Ángel recuerda que los suscriptores de Potencia tienen a su disposición el plugin que subió para ellos. Según él, «miles y miles de personas» se lo han pedido. En concreto, una: Madrillano, que le llamó por teléfono para que se lo pasara, a pesar de estar suscrito y poder descargarlo él mismo. Eso sí, le parece un plugin maravilloso. WordCamp Europe en Málaga Antes de entrar en materia, Miguel Ángel anuncia que ya se han publicado los organizadores de la WordCamp Europe que se celebrará en Málaga en mayo del año que viene, y él está entre ellos. Le encantan las fotos del equipo, con pelos de todos los colores. Lo que le preocupa un poco es tener que hablar con todos en inglés. Su plan: sonreír mucho y decir yes. Plugins hechos con IA por quienes más saben El tema principal de hoy nace de una reflexión de Mariano. Con la IA, cualquiera puede hacerse un plugin. Pero cuando quienes lo hacen son programadores expertos, cansados de usar herramientas de terceros que solo cubren el 80 u 85 % de lo que necesitan, el resultado es otra cosa. Es el caso de Fernando Tellado y Javier Casares, dos de los mayores referentes de WordPress en España. Se han hecho sus propios plugins para cubrir sus necesidades y, además, los comparten con todo el mundo. Los plugins de Fernando Tellado Fernando tiene 23 plugins en su perfil de WordPress.org. Mariano destaca algo que Fernando defiende con mucho énfasis: no son plugins «rastreros», de esos que usan el repositorio como escaparate. Esos plugins ofrecen una versión Lite que apenas sirve y te empujan a comprar la Pro. Los de Fernando están completos en el repositorio, sin versiones de pago. Además, son recientes, él mismo los usa en su web y los actualiza constantemente. DietPress Su logo es un tenedor, un cuchillo y una cuchara, lo cual tiene su gracia porque el plugin sirve para poner a dieta a WordPress. Mariano bromea con que debería estar tachado para que no coma. Es un optimizador en la línea de Machete, el plugin de Nilo. Elimina elementos antiguos o innecesarios que WordPress carga sin necesidad: Activa o desactiva características no esenciales. Reduce las peticiones a la base de datos. Acelera la carga de la página. La versión anterior, programada a mano, ofrecía unas diez o doce opciones. Con la ayuda de la IA ahora supera las cincuenta. Vigía Pensado para hacer GEO, es decir, para que las inteligencias artificiales vean mejor tu web: Te da una puntuación de cómo te están viendo las IA. Analiza más de cincuenta rastreadores de aplicaciones de IA. Permite bloquear los bots que no te interesen, por ejemplo los de las IA chinas. Mariano lo ilustra con el ejemplo de quien compró una bombilla fundida en un chino y ya no quiere saber nada. Miguel Ángel, en cambio, confiesa que le encantaría ser chino. Genera el archivo llms.txt, una especie de robots.txt con directrices para que la IA lea mejor la web. Mariano cree que de momento no sirve de mucho, aunque admite que puede estar equivocado. Crea un esquema JSON de la web. Vigilant Es el favorito de Mariano y, según él, probablemente el menos usado de los tres. Se trata de un plugin de seguridad 100 % gratuito. Lo sitúa como una alternativa a Wordfence o Sucuri que protege más y carga menos la web. Va por la versión 3.0.1 y su última novedad técnica es la autorreparación: si un virus infectara su código, es capaz de eliminarlo y restaurarse tal como estaba. Fernando lo explica en un artículo de su blog. Los plugins de su web En su propio dominio, Fernando tiene otra sección de plugins que no pasan por el repositorio. Allí la mayoría también son gratuitos: Un traductor con IA, pensado sobre todo para tiendas WooCommerce. Recuerda a Weglot, pero funciona de forma distinta. No toca la base de datos: al servir la página, pasa el contenido por la IA que tengas conectada y la muestra traducida. Mariano había visto algo parecido en inglés con pago mensual, y aquí es gratis. Un revisor de datos de productos para Google. Analiza todas las fichas y te dice qué está mal y qué se puede mejorar, algo ideal para tiendas de 3.000 productos que nadie va a revisar a mano. Dos de pago: una capa de consentimiento para cumplir la RGPD y otro relacionado con formularios. Los plugins de Javier Casares WPVulnerability Para Mariano es el plugin estrella de Javier y lo tiene instalado en todas sus webs. Se conecta a una base de datos de vulnerabilidades y revisa todo: servidor, plugins y temas. Te avisa de cualquier problema de seguridad pendiente de arreglar. Javier lo programó a mano en su día y ahora, con la IA, lo ha mejorado muchísimo. Los plugins de su empresa En la web de su empresa, Javier ha liberado plugins que desarrolló para clientes. Son más pequeñitos que los «macroplugins» de Fernando, pero están muy bien hechos: 2FA: autenticación en dos pasos, como el plugin de WordPress, pero con su criterio. Si lo miras por dentro, funciona mejor. SMTP: para enviar correos desde WordPress, esa fuente eterna de dolores de cabeza cuando el servidor no deja mandarlos. Asegura el envío, reintenta y se autoevalúa si hay errores. Secrets: guarda cifradas, o «encristadas», como bromean, las claves de API y de servicios de IA. Así no quedan en claro en la base de datos si esta se ve comprometida. Telemetry Disable: WordPress envía datos de la instalación, como número de usuarios o de entradas, que no son necesarios para funcionar. Este plugin elimina toda la telemetría no esencial y deja solo lo indispensable. Frecuencia: automatizaciones para podcasts, con transcripción automática y resúmenes por IA. Cuesta 240 euros y Mariano se lo dedica a Miguel Ángel como «chimpún» final. El aviso sobre Alexa A raíz de la telemetría, la conversación deriva hacia Alexa. Miguel Angel advierte de que los altavoces graban mucho más de lo que creemos. Además, cualquiera con acceso a la cuenta puede escuchar esos audios en Amazon. Su consejo: si vais a hacer algo inconfesable en casa de vuestros padres o de vuestros hijos, desactivad el micrófono. Noticias WordCamp Galicia Los dos tienen previsto ir. Todavía queda un mes y ya hay más de 150 entradas vendidas. Mucha gente guapa y una organización que lo hace muy bien. Ipsum, el nuevo tema por defecto de WordPress Se ha presentado el tema que llegará con WordPress 7.2, previsto para antes de fin de año. Lo desarrolla un equipo de 56 personas y trae una gran novedad: por primera vez no llevará el nombre del año. Se llamará Ipsum, como el lorem ipsum de toda la vida. Es un tema tipo lienzo en blanco, al estilo de GeneratePress o Astra. Incluirá muchos patrones, tipografías y opciones de personalización. Tiene su historia: el equipo había diseñado inicialmente un tema mucho más artístico y elaborado. Al presentarlo, se pidió algo más neutro para que cada cual pusiera lo que quisiera. Ese primer diseño no se tirará a la basura, porque también lo publicarán aparte. Mary Hubbard, presidenta de la Open Website Alliance Miguel Ángel confiesa que siempre se enamora de quien está al lado de Matt, por aquello de la «erótica del poder». La noticia es que Mary Hubbard, responsable de WordPress.org, asume también la presidencia de la Open Website Alliance, la alianza de proyectos de software libre. La presidencia es rotatoria, como la de la Unión Europea, y ahora le toca a WordPress. Purple, el primer tema de bloques de WooCommerce Tras mucho trabajo con el checkout en bloques, WooCommerce enseña una preview de Purple, su primer tema oficial hecho con bloques. Como indica el nombre, es moradito y muy bonito, y se puede probar para ver cómo queda una tienda con él. Jurassic.Ninja El descubrimiento de la semana: Automattic tiene un servicio para crear sitios de prueba con un nombre genial, Jurassic.Ninja. No es un WordPress en el navegador como WordPress Playground, sino un sitio alojado de verdad y con Jetpack incluido. A Miguel Ángel le viene de perlas. Para probar su plugin de generación de posts usaba Playground, que no permite las llamadas externas que su plugin necesita. Plugins del día Kiki Mariano cuenta el que le «hizo llorar». Él vive feliz con GeneratePress y Gutenberg, hasta que descubrió un nuevo constructor visual al estilo Elementor llamado Kiki. El nombre prometía, pero ni el nombre ni la chica que aparece en la web resultaron ser lo que él había leído, así que sigue sin cambiarse. Para quien no se deje llevar por estas cosas, promete diseñar libremente, como en Canva o Figma. Después genera los bloques de Gutenberg correspondientes con una interfaz más moderna, y se puede empezar gratis. WP Shipyard Shipyard significa astillero, y la metáfora le va que ni pintada. Este plugin gratuito permite al administrador cambiar de tema en privado: solo él ve el nuevo diseño mientras lo prepara, y el resto de visitantes sigue viendo la web de siempre. Cuando todo está listo, el barco sale del astillero y se publica el nuevo tema. No hace falta poner la web en mantenimiento. Agenda y despedida El lunes de la semana que viene hay meetup de WordPress en Sevilla. La próxima WordCamp en España será la de Valencia y luego llegará la de Galicia. Por el camino asoma también la de Faro, que para eso estamos en Iberia. Ya puestos, fantasean con vivir la reunificación de la península. Como manda la tradición, el episodio no se cierra sin chiste. A un amigo de Mariano le dijeron los médicos que tenía el hueso descalcificado. ¿La respuesta? «Bueno, lo importante es participar». 🤖 El contenido de este post ha sido generado automáticamente con inteligencia artificial a partir de la transcripción del audio. Puede contener errores o imprecisiones. 🎙️ Publicado con VozCaster, el bot de Telegram que convierte tu voz en un episodio de podcast publicado. Pruébalo gratis. ¿Te ha gustado el episodio? Si quieres que sigamos experimentando con bots, protocolos y empanadillas polacas, no olvides suscribirte y dejarnos tu valoración. ¡Nos escuchamos en el próximo capítulo! Métodos de contacto Enviadnos vuestras preguntas al grupo de Telegram. Apuntaos al canal de Youtube del podcast https://www.youtube.com/potenciapro Si nos queréis decir algo directamente lo podéis hacer a @potenciapro , @materron, @mpc, o en el grupo de Telegram Y si eres muy muy muy fan del podcast Echa un vistazo a cómo nos puedes ayudar en https://potencia.pro/se-prosperoso/
Podcasting 2.0 September 25th 2026 Episode 272 - "Is this Y'all?" 00 -
Ep 292Apple Plans to Squeeze More Revenue From the App Store, Report SaysiCloud+ Now Includes Apple TV, Apple Arcade, and Curated Apple Music Stations in Over 100 CountriesDiscussingFilm: Most awarded networks at the 2026 Emmys: Apple TV - 28, HBO/Max - 21, Netflix - 16, Prime Video - 10, NBC - 8, Peacock - 7, Disney+ - 6, CBS - 6, FX/Hulu - 3Apple unveils iPhone Duo, the first foldable iPhoneBasicAppleGuy: #iPhoneDuo — The Making Of, posted by Apple on the TikTok channelApple debuts iPhone 18 Pro and iPhone 18 Pro MaxIntroducing Apple Watch Series 12, with the all-new Health Sensing SystemApple unveils Apple Watch Ultra 4Apple introduces AirPods 5 with best-in-class open-ear Active Noise CancellationApple advances health and fitness capabilities using Apple IntelligenceApple releases iOS 27, iPadOS 27, macOS 27, watchOS 27, visionOS 27, and tvOS 27 - MacDailyNewsDr Waqar: Apple to apple comparison for #iOS26 vs #iOS27Quinn Nelson: This is insane. This is an M2 iPad running iPadOS while simultaneously running macOS 27 at 6K on an external monitor.macOS 27 kills Time Capsule backups, these NAS replacements keep workingPaul Hudson: I don't think I ever swear on social media, but this one is worth it: holy shit! A JSON-based project format for Xcode!McKinley — Design your own SF SymbolsAmore — Self-publish your Mac apps outside the App Store: Sparkle updates, code signing, notarization, and DMG creationHomebrew 7.0.0: faster installations and upgrades, stronger sandboxing, a native macOS app, built-in vulnerability checks, the end of macOS 10.15 support and Intel Macs moving to Tier 3brew install homebrew-appimmurok: Standalone Wireless Fingerprint Key for Mac, Windows & LinuxAnil Dash: Dashboard Touch — DIY fingerprint reader for the Mac that types your password, from ~$30 of open-source hardware and softwareHedgie: An NYU mathematician says OpenAI used his own progress against him to beat him to one of the biggest unsolved problems in mathematics. Tristan Buckmaster had been working toward a Millennium Prize proof using OpenAI's Codex when information about his progress reached OpenAI. Days later, OpenAI p…llmfit: Right-sizes LLM models to your system's RAM, CPU, and GPUWill It Run Locally | Local AI NewsNije Apple, ali mora: trailcamZahvalniceSnimano 21.9.2026.Uvodna muzika by Vladimir Tošić, stari sajt je ovde.Logotip by Aleksandra Ilić.Artwork epizode by Saša Montiljo, njegov kutak na Devianartu
Tickets for AIE NYC now open, and apply for the invite-only AIE CODE. Join us!We have an unusual relationship with today's guest: for years since coauthoring the InstructGPT paper, Diogo Almeida had been saying that API-available frontier models have been going down the wrong path, everything from the alignment to refusals to reliability perspectives, that we have dropped every mode other than autoregressive chat-tuned LLMs because of the overwhelming success of ChatGPT.In a launch video now viewed ~40M times (by comparison, GPT4o was 22M, Fable 5 was 15M, Navier Stokes was 74M, and 6 Astra was 137M), Diogo introduced Jev and it immediately took over the AI timeline — we'll skip full Jev explainers because your favorite AI influencer/educator has probably already done one. We also collected:* the official patterns and cookbooks you should see first, from Allie* Jev usecases* speed based - games and computer use* the voice + computer use example we discuss at 1h34 mins* voice + browser control* The must not miss Doom demo* Driving cars in games* Excalidraw* virtual try-ons* “Smart Games”/smart NPCs* guided responses in text messages* Jev for coding agents has an official guide * jev for linting* compacting tool calls* reasonable pushback from Theo - Diogo has published a note on the Tyranny of the KV Cache that you should read as a followup after the pod for Jev + coding agents, because of his belief that Cache Rules Everything* Programming Languages built atop Jev (Diogo's fave)* Jev for analytics replay and user journey review* “dark data”* entity resolution* natural language search* “smart software”* a core goal of Jev is to “disappear into the background” - eg as unremarkable as regex* Jev as a judge* Jev memes* Jev vs LLM capabiltiies* blending transformers and classifiers* about the confidence api* Jev vs GLiNER (note difference/pushback, agreed, agreed, agreed)* Jev on trolley problem* Jev BushInstead we'll focus on what we can uniquely offer — a broader philosophical and mission-based understanding of how and why Jev was created, and what you should expect next in terms of future models from TypeSafe (ReasoningJev?) and what usecases and ideas you should work on vs the 55th low effort clone of Jev's API or doing a generic JevBench benchmark - something Diogo has rejected publicly.Why RLCD: Three kinds of RLHF, and why they are ALL the wrong north starDiogo knows a good deal about RLHF, given that he was on the team that pioneered post-training at OpenAI — and traces the three branches to Christiano et al 2017 (the robot backflip demo), Stiennon et al 2020 (learning to summarize) and his baby, Ouyang et al 2022 (InstructGPT). From there on, every innovation from Function Calling to Structured Outputs to Reasoning felt like a hack on top of the string based, sequence to sequence prediction paradigm. As he mentions on the pod, from 2023-2024 he struggled unsuccessfully, due to both personal and organization underestimation, to train a model that accurately addressed what he saw as the core problem with making LLMs the heart of software: reliability.Jev's core innovation is "Reinforcement Learning for Calibrated Decisions”, a novel, unpublished technique that optimizes for “answers with epistemically honest probabilities on System One tasks” rather than human rated feedback (RLHF) — which causes hallucinations, sycophancy, and permanent reliance on humans — or programmatically verifiable outputs with rubrics (RLVR) — which solves Navier Stokes but exacerbates jagged intelligence and doesn't integrate well with other software.We've talked about the calibration problem before on the pod, but probably the single best place to understand why RLCD became necessary is Diogo's AIE talk, which discusses why a generation of training helpful AI assistants for humans has impaired them for training models for composable, programmable AI for automation.At the end he also teases his contrarian opinion on scaling laws - which teases how to build a modern neolab without the billions of dollars the major labs have…The Bitterest Lesson: Tasks and Data beats ComputeWe spend a good amount of time discussing Diogo's essay on the Bitterest Lesson:His point is that “You get what you optimize for and the bitterest lesson in ML is that the most important part of it isn't ML at all.” - and picking the right north star, eg upvoting for user preference vs being integrated into tool calls - makes everything else fall in line.We're excited to catch up with a freshly dyed Diogo to discuss:* Why AI can solve extraordinarily hard problems but still fail to automate basic work* What System One Models are and why Jev is built for software rather than chat* RLHF, mode collapse, calibration, and the hidden costs of optimizing for human preferences* Why refusals become a problem when AI is buried inside software dependencies* Why TypeSafe rejects public benchmarks and optimizes for intelligence per dollar* The “bitterest lesson”: why the right task and the right data can matter more than compute* Why TypeSafe thinks of itself as a data lab rather than a model lab* RLCD vs. RLHF and RLVR as fundamentally different North Stars for AI* Why reliability and robustness matter more than simple determinism* Jev's programming primitives and how intelligence maps into software control flow* Why developers should decompose AI workflows into small, measurable decisions* How structured state replaces giant prompts and system messages* Why Diogo thinks AI should eventually disappear into the background of software* The “inverse SaaS-pocalypse” and how AI could supercharge existing software* System One vs. System Two intelligence and the limits of reasoning models* Dark data, computer use, real-time intelligence, and Jev's biggest early use cases* Why Jev could reshape coding agents built around a single-model architecture* Why Diogo says he wouldn't pre-train with $1 billion* The OpenAI journey that led to TypeSafe and why he thinks many neo-labs are approaching AI incorrectly* Coding agents beyond the KV cache, shared state, sub-agents, and the multi-agent futureDiogo Almeida* LinkedIn: https://www.linkedin.com/in/diogomda* X: https://x.com/CompleteSkeptic* TypeSafe AI: https://typesafe.ai/Timestamps00:00:00 Jev Launch Week and the AI Economic Revolution00:02:50 What Is Jev? System One Models and Programmable AI00:05:54 RLHF, Mode Collapse, Calibration, and Yann LeCun00:10:29 Programmatic AI, Refusals, and Safety Alignment00:17:21 Why TypeSafe Rejects Public Benchmarks00:20:43 The Bitterest Lesson: Data, Compute, and the Right Task00:24:59 RLCD vs. RLHF and RLVR00:28:42 Why Powerful AI Still Hasn't Automated the Economy00:39:55 Reliability, Robustness, and Determinism00:48:11 Model Versioning, LTS, Speed, and Intelligence per Dollar00:54:04 Inside Jev's API and Programming Primitives00:58:28 How to Build with Jev: Structure, Decomposition, and Small Decisions01:18:28 The Inverse SaaS-pocalypse and AI Disappearing into Software01:33:21 Computer Use, Dark Data, and Jev's Biggest Use Cases01:38:48 How Jev Could Reshape Coding Agents01:41:00 AI Safety, Frontier Pacing, and the Limits of RLVR01:48:03 Why Diogo Wouldn't Pre-Train with $1 Billion01:55:19 The OpenAI Story Behind TypeSafe02:01:41 Why Diogo Thinks Most Neo-Labs Are Getting AI Wrong02:08:00 Coding Agents Beyond the KV Cache and the Multi-Agent FutureTranscriptIntroduction: Jev Launch Week and Developer MomentumSwyx [00:00:00]: Okay, we're in the studio. A special occasion because this week, Diogo, my good buddy, launched Jev, and it's been taking over the complete timeline. How do you feel? What's it like to be you right now?Diogo Almeida [00:00:16]: Emotionally?Swyx [00:00:17]: Yeah.Diogo Almeida [00:00:17]: Never been worse. Like, I'm a ragged corpse of a person right now because there's so much going on, and I'm like a technical CEO, so I have, like, a lot of fires to fight.Swyx [00:00:29]: Yeah.Diogo Almeida [00:00:29]: But mentally, I feel—I say this all the time, and I've been saying this kind of for years in my over-under events. Like, I feel like the entire AI field is like one of those, like, carnival house of mirrors, and everyone is just insane and saying the weirdest stuff that doesn't make sense. And it feels like for just this week, like, I'm on a better in sync with reality and like, oh, people see it now. AI can be so much more than what was once thought.Diogo Almeida [00:01:06]: And like, yes, we are going to make. Like, an AI-based economic revolution is back on the table, and this is f*****g awesome.Diogo Almeida [00:01:17]: I'm so jazzed the developers get it. It's, it's, Yeah, and I want to show my eternal gratitude to the developers andSwyx [00:01:25]: Yeah.Diogo Almeida [00:01:26]: I'm so jazzed about the community and everything. It's so great.Swyx [00:01:28]: Yeah, you were saying yesterday that you decided to prioritize the town hall and not a bunch of, like, VIP, investor-type people because you wanted to make sure that they are the people that you get your most, attention, right? The engineers, the developers.Diogo Almeida [00:01:43]: Yeah, it felt a little like, oh man, I'm talking to, like, really important people right now.Swyx [00:01:47]: Yeah.Diogo Almeida [00:01:47]: I probably shouldn't reveal who.Swyx [00:01:48]: Yeah.Diogo Almeida [00:01:48]: But it feels a little bit dirty for me to, I'm, like, perhaps overly genuine in things. Like, it feels, like, dirty if, like, in my gigantic calendar event of people to talk to, the community isn't one of those.Swyx [00:02:04]: Yeah.Diogo Almeida [00:02:04]: And actually, in my ideal world, it would be, like, community all the time. I was thinking, “Should I host a town hall while walking to your studio?” And I'm like, “No, that's too crazy.”Swyx [00:02:12]: Sure. Yeah. Well, you guys have been hosting town halls on Discord. Discord is now 100,000 people. Your Twitter'sDiogo Almeida [00:02:19]: I don't follow these stats.Swyx [00:02:20]: Yeah.Diogo Almeida [00:02:20]: So holy s**t.Swyx [00:02:21]: Your Twitter's blown up. It was, it was really funny ‘cause, like, at AIE, you were like, “Yeah, follow me please,” and then you didn't, like, provide even your handle.Diogo Almeida [00:02:29]: I'm a noob. I'm a noob.Swyx [00:02:29]: You're such a noob.Diogo Almeida [00:02:30]: I'm a noob.Swyx [00:02:31]: But no, but that, like, that's, like, positive aura that, likeDiogo Almeida [00:02:33]: CoolSwyx [00:02:33]: You don't know how to promote yourself.Diogo Almeida [00:02:35]: Yeah. Someone, like, called me out when I posted, like, “Holy s**t, we're all three twending-- trending topics.” And then they're like, “That's a personal feed.”Swyx [00:02:42]: That's a personal, yeah.Diogo Almeida [00:02:43]: And I'm like, “Oh, no.”Swyx [00:02:44]: Of course, of course it'll trend to you.Diogo Almeida [00:02:45]: Cringe. Yeah.Swyx [00:02:45]: Yes, ‘cause it's what you clicked on.Diogo Almeida [00:02:47]: Yeah.Swyx [00:02:47]: So okay. Let's, Yeah, so congrats on everything.What Is Jev? System 1 Models and Intelligence per DollarDiogo Almeida [00:02:50]: Thank you.Swyx [00:02:50]: We'll talk about more, details as you have them. But let's, for people who are, like, living under a rock or just want, like, the definitive thing, what is Jev?Diogo Almeida [00:03:02]: Whew. Let me think about. That's a hard one.Swyx [00:03:07]: Okay. And I'm happy to, like, re-ask if you wanna kind ofDiogo Almeida [00:03:09]: No. I'm happy toSwyx [00:03:10]: OkayDiogo Almeida [00:03:10]: I'm happy to, like, just jam on it.Swyx [00:03:12]: Yeah.Diogo Almeida [00:03:13]: I will say, like, the first thing that I'm relieved about with this question is now I don't have to answer that question to my parents anymore ‘cause ChatGPT can just explain it.Swyx [00:03:20]: Nice.Diogo Almeida [00:03:21]: So the way I see it is we new-- need a new class of models. We're not attached to naming that class of models. Our-- the most accurate name we've come up with is System 1 models.Swyx [00:03:33]: Yeah.Diogo Almeida [00:03:33]: There will be reasons, but it's-- there's a reason why we don't call them decision models, because, like, they will be. Like, System 1 is beyond that. That's all I can say. We didn't expect this to be our big launch, so we have stuff in the tank.Swyx [00:03:48]: You should have said low-key research preview.Diogo Almeida [00:03:52]: It kind of was, right? It kind of was. But we. So there's a class of models that we describe them as, like, machine-native, System 1, large programmable. I think these are-- is the class of models where the goal is for code to be the consumer. So as opposed to, lar-- pre-trained large language models, which are meant for, like, autocomplete of the internet, or RLHF models, like chatbot instruction-following models, which are meant to, like, reply to text, or RLVR. It's in a weird gray area with RLHF. Like, these are meant to have things that directly are consumed by code, hence the name type safe. So the thing we really want is to have, like, AI, like, be as powerful as possible, and we think the way to do that is to integrate it with software. And we are designing everything, beyond just the outside, the deep internals of the model to be optimized for software. So number one, Jev is our first large programmable model, or a System 1 model, whatever you want to call it. Jev is meant to be optimized for intelligence per dollar, hence the name Jev.Swyx [00:05:03]: Jevons Paradox.Diogo Almeida [00:05:03]: Jevons Paradox, yeah. And it's optimized for intelligence per dollar. I love this debate with people about what is the most important between reliability, cost, calibration, and speed. And Jev is meant to be. Jev will be the name of models that will be on the frontier of intelligence per dollar. There's other ways to optimize it, like, ML, or at least if you're good at ML, it's all about trade-offs. And we are just going all out on that.Calibration, Mode Collapse, and the Limits of RLHFSwyx [00:05:31]: Yeah. And to me, like, calibration is one of the new things that people weren't talking about as much. We've done an episode In the past, with Clementine Foreia of Hugging Face, where they were like, “Yeah, actually, y- they're just.” Or, and this is your whole argument about RLHF, is they're more collapsing towards what you want to hear the mostDiogo Almeida [00:05:50]: OohSwyx [00:05:50]: Or what is most likely, instead of, like, their own internal confidence about a thing.Diogo Almeida [00:05:54]: Can I soapbox on that for a second?Swyx [00:05:56]: Go ahead. Yeah.Diogo Almeida [00:05:57]: Cool. Like, I've been heard that your audience is the most technical, so I actually want to get into that.Swyx [00:06:02]: Yeah.Diogo Almeida [00:06:03]: And if- I went through extreme precision to make sure everything in our launch video is accurate and real. Apparently, that's very unusual. One of the things that no one paid attention to was the downsides of RLHF, in particular mode dropping.Swyx [00:06:17]: Mode dropping or mode collapse?Diogo Almeida [00:06:19]: It's the same thing.Swyx [00:06:19]: Is that what you call it?Diogo Almeida [00:06:20]: It's the same thing.Swyx [00:06:20]: All right.Diogo Almeida [00:06:21]: And I wanna have a blog on this eventually, but I, like, want to tell as many people this as possible ‘cause I think it's a very interesting thing. So the spicy take, I believe in Yann LeCun a lot. I think Yann LeCun's takes are actually among the closest toSwyx [00:06:36]: What about this?Diogo Almeida [00:06:37]: Well, should I address this now or should I wait and go into mode collapse?Swyx [00:06:40]: No, later. Go mode, go mode collapse. I don't know.Diogo Almeida [00:06:42]: So I actually think that among takes, Yann LeCun's is among the most accurate. But he has this very famous/infamous slide about,Swyx [00:06:52]: The cake?Diogo Almeida [00:06:53]: LLMs are doomed.Swyx [00:06:54]: Okay.Diogo Almeida [00:06:54]: Like that one where he, like, has, like, a pie chart with, like, a tiny par-- tiny little thing- and says that as you increase sequence length, the probability of it making an error goes in. Yes, this one. This one. I love this one, because it's one of these things that seems mathematically obvious, but is obviously wrong, right? Like, it's mathematically obvious, but it doesn't empirically hold. And this is my favorite thing to teach people about, like, where youSwyx [00:07:21]: What's the disconnect, right?Diogo Almeida [00:07:22]: Exactly. And may I or you want to tell me?Swyx [00:07:27]: About mode collapse?Diogo Almeida [00:07:28]: Oh, no. Oh, so mode clop-- collapse is related to this.Swyx [00:07:31]: Yeah.Diogo Almeida [00:07:31]: The disconnect happens because if you are in a mode covering or a calibrated distribution, you are, like, not. You are not overly punished about having outliers. You'd expect, like, something. Some amount of the time you'd be out of distribution, some amount of time you'd be in distribution. That's what happens when you cover the distribution. This was like models before GANs. They made blurry images, right?Diogo Almeida [00:07:54]: Instead, GANs mode drop. They, like, drop the minority classes and just do the really common ones. And this is why this effect doesn't happen, right? Like, instead of be-- in order to generate really long strings, without making errors, they need to, like, be extremely conservative because it's e- really easy to see when an error happens. It's very hard to see when, like, a subtle thing that looks correct happens. And that calibration is, like, total poison into, like, the probability distributions of strings.Swyx [00:08:22]: Yeah.Diogo Almeida [00:08:23]: And it's, it's a nuanced take and like, I think that This is why this doesn't happen, and this is why strings are so bad at, decision-making or, overloading the string models are for decision-making is, like, a bad time.Yann LeCun, JEPA, Scaling Laws, and Practical ResearchSwyx [00:08:38]: And while we're on the topic of Yann, do you agree that his fix i- with-- which is like a world model, like a JEPA-type, embedding thing is the right solve? So basically, like, the. One of the reasons that it could fail is because you're trying to reason over token outputs and then, and then just looping back again and going. Keep, continuing going until you reach, like, a end of sentence. Like, is that, And his solve is JEPA, right?Diogo Almeida [00:09:02]: Yes.Swyx [00:09:02]: Which is, like, joint ambition,Diogo Almeida [00:09:04]: YeahSwyx [00:09:04]: Joint embedding prediction. So like, is that the solve or, like, do you have a. Do you have a take on that?Diogo Almeida [00:09:10]: Oh, man. I probably shouldn't talk too much about the insides of ML, but I will say that my brand, other than unhinged, is practical.Diogo Almeida [00:09:20]: Like, even my take here is practical. And like, I'm. Am I a scaling law fan? Depends. It dep-- it's, it's, it's, like, it's. Scaling laws tell you how much better you get at a thing for amount in.Diogo Almeida [00:09:33]: A scaling law does mean exponentially more resources for normally sublinear gains, which looks to be a bad investment unless those, like, linear gains are, like, really valuable. But it's all. To me, it's all about, like, what can we do with what we have to make the biggest possible f*****g difference? I can curse.Swyx [00:09:51]: Yeah.Diogo Almeida [00:09:51]: Yeah.Swyx [00:09:52]: Yeah.Diogo Almeida [00:09:52]: Yeah.Swyx [00:09:53]: We're, we're, we're approved for adults.Diogo Almeida [00:09:54]: Hell yeah.Swyx [00:09:55]: And also we have a scaling law thing if you wanna go into that later.Diogo Almeida [00:09:58]: Oh, I could if we. See, that part is not super relevant right now.Swyx [00:10:02]: Yeah.Diogo Almeida [00:10:03]: I actually. If you wanna go into my bitterest lesson, I think that's more relevant.Swyx [00:10:06]: Okay.Diogo Almeida [00:10:06]: But like, to me, I'm all about, like, pragmatics. And I think that the JEPA stuff is really cool early research. I really love awesome research. Is it practical yet?Diogo Almeida [00:10:21]: Probably shouldn't say. But like, there's just a lot of.Diogo Almeida [00:10:29]: I just think there's just, like, so many diamonds in the rough let all over the research world right now that haven't been polished because people don't know how to, like, do the right task. And I think that what our launch did, it. Does it kickstart us as a company? Like, yes. Will it be great for us as a company? Yes. I think it's gonna be, like, even greater for this direction of, like, programmatic AI. There was going to be, like, a gold rush on top of us for. ‘cause, like, software is super f*****g charged. But I think there's gonna be a gold rush parallel to us as well on, like, all the different ways we can expose things to make software more powerful so people can make even cooler stuff. And then we are back to, like, early internet energy?Swyx [00:11:12]: Yeah.Diogo Almeida [00:11:12]: And I think that's why, like, the Twitter is just like, “Jev.”? It's, it's like. It is a partySwyx [00:11:18]: It's inspiring because it's, it's, like, so different than what we're used to, which is, “I'm sorry you can't do this, but we do scaling laws and only the big labs can do it,” right?Diogo Almeida [00:11:28]: That. Actually, if I. I'll, I'll make a tangent if that's okay.Swyx [00:11:32]: Yeah.Diogo Almeida [00:11:32]: I think you might enjoy this.Swyx [00:11:33]: Really? Our five tangents in. It's good. It's fun. Yeah.Diogo Almeida [00:11:35]: Oh, yeah. I get lost at all my tangents.Swyx [00:11:37]: This is gonna be horrible for the listeners to figure it out, but they're gonna figure it out. It's fine.Safety Alignment, Refusals, and API PhilosophyDiogo Almeida [00:11:40]: Yeah, we can edit it in post.Swyx [00:11:40]: This is my response. Yeah.Diogo Almeida [00:11:41]: So popular thing on Discord, that people keep asking me, I haven't had the time to explain it yet, is why am I opposed to safety alignment and why do we not refuse? I'm not opposed to safety as a principle, but I think that safety alignment is generally misaligned with users. And refusal is just, like, obviously a type error. Like, if you're a human being and you're chatting with, like, a bot or whatever, you're cloud coding, and a refusal happens, like, “I'm sorry, I can't read DNA.py.” that's an annoying time. It's anno- it's, it's annoyingDiogo Almeida [00:12:18]: Right? But you can work with it, right? And you're forced to work with it ‘cause of Stockholm syndrome.Diogo Almeida [00:12:23]: I have stories about that too. I need another tangent deep in here. But like, if you ever want this in a dependency running in the background, what happens if that refuses? What if someone else is using that dependency? They don't know what that system is. Like, you want the software to just stochastically break because a user sent, like, a weird message in there?Diogo Almeida [00:12:42]: Like, that is, like, straight-up insanity. It's coming from a place of, like, people who do not understand software, do not understand programming, and like, they are obsessed with, like, I believe this, horseless carriage of, like, AI coworker instead of unearthing, like, the full power of AI.Swyx [00:13:01]: Fair enough.Diogo Almeida [00:13:01]: Yeah.Swyx [00:13:01]: You want something that is the core kernel that is usable everywhere.Diogo Almeida [00:13:05]: Yes. Exactly. Like, the cognitive core, right?Swyx [00:13:07]: Yeah.Diogo Almeida [00:13:08]: And you need this thing to be s- like, so general, so optimized for its use cases. You want it to be, like, you want it to work on all the future use cases, all the weird s**t that people are doing.Swyx [00:13:19]: Yeah.Diogo Almeida [00:13:19]: We obviously didn't train on any of that stuff. Is it surprising that it works? No, ‘cause we trained on weirder stuff, my friend.Diogo Almeida [00:13:28]: So. But one tangent up about, like, safety alignment.Swyx [00:13:32]: Okay.Diogo Almeida [00:13:32]: Safety alignment makes sense for a product, in my opinion, for, like, ChatGPT and Claude. Like, it, What safety, what makes safety and capability alignment different is capability alignment is, like, about doing what the user wants. That is sick for software engineers. They want their thing to do the thing, and the more predictable it is, the less they have to test it and play around with it. Jeb is not anywhere close to that yet. It could be, but like, there's so many more nines of reliability that we want in order to make it so good, like a database query, that you don't even have to think about it. It is just there when you need intelligence. But safety alignment is, like, the opposite of instruction following. It's when you want to follow someone else's instructions, like OpenAI and AnthropicSwyx [00:14:13]: The RAGs value stack.Diogo Almeida [00:14:14]: Exactly. And this makes a lot of sense for a product. Again, like, ChatGPT should do. Y- you sh- like, if they don't want to, like, do, like, some, not-safe-for-work role play with ChatGPT, that's on them because, like, maybe that's, what their users who have, like, parents and kids want. Like, n- that's fine. But in an API, that's nuts, right? Like, that's completely unacceptable because, like, people need to, like, program around this, and that is, that's so anti-user that it's. It. I'm. Huh. I can be an angry person, so I should try to calm down.Swyx [00:14:52]: It's, People get your passion, and I think that's really good. The one pushback I'll give you is, like, what if we use it to kill people, right? Like, that is the actual. Like, n- the not-safe-for-work thing, it's private, personal, whatever. But like, yes, like, we will use it in war. And like, that is, something that companies can reasonably prefer their APIs not be used for.Diogo Almeida [00:15:14]: I get that. I think that there's, like, pragmatic places where that opinion can be held. I don't think the foundation of, like, a general-purpose technology is that place, personally.Diogo Almeida [00:15:27]: Like, would I prefer that our stuff is not used to kill people? Obviously. Would I prefer it's used for, like, all sorts of, like, great stuff in the world? Obviously. Will I put my thumb in the scale for that? Yes. Will I do it at the technological layer? Absolutely not, because that will fracture the intelligence. Every single time you mean it to overfit to some weird stuff, you're fracturing its intelligence more and more. And like, these things are fractured to the, like. They're so darn fractured right now.Swyx [00:15:54]: Yeah.Diogo Almeida [00:15:54]: So and as a furthermore thing, to me, it's like I think intelligence will be more like a database than a coworker. Like, I don't think it's up to databases to add checks on whether or not they're used for, like, what's something that's not great? Like, CIA. Actually, I don't know what the CIA does, really. You can imagine. You can imagine, killing people who are not even bad or whatever.Diogo Almeida [00:16:21]: And like, I don't think it's the database's responsibility for that. And furthermore, like, a thing that has been weird to me is when people, like, sign up for our thing on Slack and they're like, “Hey, we're gonna deploy this. Can we deploy this thing?” I am just like, “My brother, we are an API. You are a developer. It's none of my business,” right? Like, you shouldn't know what the whole task even isSwyx [00:16:46]: YeahDiogo Almeida [00:16:46]: Because it should be decomposed into small things. We shouldn't be able to know what the downstream users are doing, and that is, like, a good boundary to give software engineers maximum power. Ideally, they use it for the good stuff, and ideally, we can, like, help them and like, we've talked about, like, doing open source and charity and all of that. We have absolutely no time for anything else right now. But like, they will get any of that bias out of the technological layer as long as I'm in charge.Privacy, Benchmarking, and Trusting IntelligenceSwyx [00:17:11]: Yeah, that's great. While we're on the topic, let's also briefly talk about your privacy stuff, terms of ser- terms of use, which, got a little bit ofDiogo Almeida [00:17:18]: OohSwyx [00:17:18]: Misunderstanding. I just wanna clarify that upfront.Diogo Almeida [00:17:21]: Hell yeah.Swyx [00:17:21]: I think this probably takes two sentences from you about, like, you will not. You're not being that restrictive about your API. Like, clearlyDiogo Almeida [00:17:27]: Oh, yeah. Oh, yeah, so yeahSwyx [00:17:27]: Ideologically, you articulate your role as a platform very seriously.Diogo Almeida [00:17:30]: Yes. Yes. I don't know what you're referring to, but like, this was. I've seen a couple of things about, like, benchmarking.Swyx [00:17:38]: Yes.Diogo Almeida [00:17:38]: Like, obviously we're not stopping people from do. Oh, man, I should be careful about what I say. I'm realizingSwyx [00:17:43]: No, you said, you said it publicly thatDiogo Almeida [00:17:44]: YeahSwyx [00:17:44]: That was in the preview period. You didn't take it out for the launch.Diogo Almeida [00:17:47]: Yeah. Okay.Swyx [00:17:47]: And now you're gonna take it out.Diogo Almeida [00:17:48]: So the team is doing stuff thatSwyx [00:17:49]: YesDiogo Almeida [00:17:49]: I'm not even aware of, so it's great to know the team communicated that. I asked them to check in with the lawyers about that.Swyx [00:17:54]: Yeah.Diogo Almeida [00:17:54]: Like, we are obviously not stopping people from doing that type of thing. I'm extremely in favor. So I'm extremely anti-public benchmarks. I'm extremely in fa- I'm medium about private benchmarks that are proxies. ISwyx [00:18:09]: So are you worried about, saturation or, like, training on public benchmarks? So it's, like, easy to cheat.Diogo Almeida [00:18:15]: Not only is it easy to cheat, there's a lot of ins. So I think that we are. Or anyone who's, like, competition with us that, vaguely there is. Like, you could say, likeSwyx [00:18:28]: There's like 50 Jev clones, yeah.Diogo Almeida [00:18:30]: Well, sure.Swyx [00:18:31]: Yeah.Diogo Almeida [00:18:32]: Well, the, these. Let's say that there is competition.Swyx [00:18:34]: And we'll talk about those. Yeah.Diogo Almeida [00:18:34]: Or let's just say that there's. Let's just assume that there's an industry two years from now of people who are doing similar things to us. The thing that we are selling is intelligence per something, per, like, dollar or per second. The. No one. Like, people obsess about the cost and the speed. I believe that is. It's cool, but like, the thing that matters is the intelligence. Like, the cost and the speed are, like, are bad things. You're paying them for something, and you need the thing back, and the intelligence is what truly matters. The problem with intelligence is that there's a je ne sais quoi to it, right? Like, the good model smell. Like, the thing that happened after we launched of, like, two hours later that actually went way bigger than the video, which was like, “Holy s**t.”Swyx [00:19:16]: This is actually usable.Diogo Almeida [00:19:17]: It. WellSwyx [00:19:17]: Yeah.Diogo Almeida [00:19:17]: It's, like, beyond that.Swyx [00:19:20]: Yeah.Diogo Almeida [00:19:20]: Like, the. Whew, the launch was crazy, and people could really sense how hard we care about that, and that's truly what I think the long term of this is. And I think public benchmarks are antithetical to this. Like, they are a way to get people trust in intelligence because intelligence has a je ne sais quoi, but the public benchmarks are extremely gameable. Even if they try not to, they still will. Like, back in the old days, every lab had a team to collect data that looks like MMLU to make it look better, which is just benchmarking with extra steps.Diogo Almeida [00:19:58]: So I believe that in the long run, it needs to be vibes and trust until you put it into a workflow and evaluate it for that workflow and measure it and have your own sense of, like, how it does on the exact workflow that matters. And our job is to keep moving the nines of reliability. This is like an ever-present part of o- of what we need to be doing as a company, and we need to do everything to have people know that this is something we care so much about. Like, if we wanted to, we could have released Jev, like, a year and a half ago if we wanted it to be dumb.The Bitterest Lesson: Tasks, Data, and North StarsSwyx [00:20:34]: Oh.Diogo Almeida [00:20:34]: It. Like, the. My bitterest lesson, right? Like, architecture and Yeah.Swyx [00:20:40]: I'll bring it upDiogo Almeida [00:20:40]: Hell yeahSwyx [00:20:41]: Since you, since you talked about it, here.Diogo Almeida [00:20:43]: Hell yeah. T- like, Sutton says that algorithms beats compute very roughly. Data matters way more than compute, obviously. And doing the right task, having the North Star is the hardest, most important thing. This has happened, in LLM land twice so far, right? Maybe 2.2 times. There's RLHF, which, like, shifted the task to instruction following. No one realized that was possible. RLVR did, like, a tiny little, like, edit to the, to the direction, and now us, right? RLCD. We have a new task, and the goal is, programs in the loop. And yeah, data matters soSwyx [00:21:28]: RightDiogo Almeida [00:21:28]: Unbelievably much.Swyx [00:21:29]: SoDiogo Almeida [00:21:29]: Like, I can't, I can't emphasize it less.Swyx [00:21:31]: Yeah, you consider yourself a data lab rather than, like, a model lab. Is thatDiogo Almeida [00:21:35]: AbsolutelySwyx [00:21:35]: Something. That's the wording you guys use?Diogo Almeida [00:21:37]: Yeah. We are. We will always, like, care so much about data. To me, model capabilities means data. Data is so unbelievably complicated, and that is what gets nines. Like, you have no idea how much data can shift everything. Data is so important.TypeSafe as a Data Lab and Synthetic Data StrategySwyx [00:21:57]: Yeah.Diogo Almeida [00:21:57]: Holy crap. So if people are looking for a job, we are hiring infinite data people, actually infinite.Swyx [00:22:04]: What is a good data person? Like, clearly somebody who cares about reading through the transcripts of, whatever. You've said, for example, that y- all your data is synthetic.Diogo Almeida [00:22:15]: Yep.Swyx [00:22:15]: But that's only, like, the scratching the surface, right?Diogo Almeida [00:22:18]: Yeah.Swyx [00:22:19]: Like, it's not. Like, synthetic, so what, right? Synthetic, but we have people with a lot of taste and a lot of care looking at, looking at these, articulating what's wrong, going back, regenerating. Is that what a good data person is these days?Diogo Almeida [00:22:31]: Let me try to figure out how to. Like, it's, it's super complicated, and like, I literally onboard the data people with a Talk that I assume is longer than this podcast will end up being. So I will try to say, like, the high level of it. So number one, we don't do the kind of synthetic data that people ki. Well, I'll do. Actually, number is zero. Data and synthetic data depends on your task. Like, the shape of your data. The shape of your task changes the data. Like, RLVR's data is kind of environments, right?Swyx [00:23:03]: Yes.Diogo Almeida [00:23:04]: RLHF's is the human feedback? Each task has its own unique kind of data, and we, of course, have our own unique kind of data, right? So number one, we have that. Number two, the thing I. The reason why we don't want to train on our users' data, even if we could, right? Like, we could probably ask for any terms right now, and it will. We. I don't know if it would make a difference. We truly don't want that, because no matter what, the real-world data has so much bias. There's, like, a power law of, like, people, like, asking the same things where you'll end up, like, overfitting to it and like, fracturing to it and all of that. And number two, we are, like, aiming for, like, a complete sci-fi future years from now where, like, these models are going to be, like, the general infrastructure, layers and layers and layers and deep down the stack to, like, things people can't even imagine. Like, I would like to think of our model, like, kind of like, UDP as LLMs and TCP as our models. All sorts of stuff can be built on top of that, and we need to be able to nail those futuristic use cases such that software developers can actually build that futuristic stuff. And the way to do that is even if we had all of the data of the present, we would just overfit to the present, and then it wouldn't work. What we need is to, like.Diogo Almeida [00:24:21]: It almost feels like a. Like, they're the artists? They study this cognitive core. Our cognitive core is, like, way less jagged than anyone else's. And then they find the jaggednesses, and then they address them surgically in a way that. And you can never perfectly do this, right? But they do it in such a way that it addresses it in every single possible, like, dimension, past, present, future.Swyx [00:24:45]: The general case rather than the specific case.Diogo Almeida [00:24:47]: Exactly. And like, that requires a lot of intelligence every time.RLCD vs. RLHF: Defining a New TaskSwyx [00:24:50]: Okay, so we mentioned a little bit. You sort of criticized my thinking as r-- like, very RLVR influence, which is, like, very fair. Let us actually mention RLCDDiogo Almeida [00:24:59]: OohSwyx [00:24:59]: Which obviously you have some secret sauces to our knowledge. You've never actually published a paper or anything like that on it. No, right?Diogo Almeida [00:25:05]: No, not yet.Swyx [00:25:06]: But like, what should people get from this? Like, what. Can you give people some confidence that you're just not just making up jargon for the sake of sounding cool, right? Like, one thing for me is, like, calibration I do think is a. To me, like, well understood because we've covered it in. On the podcast.Diogo Almeida [00:25:22]: Yeah.Swyx [00:25:22]: But I don't know what you mean when you say RLCD versus what people are familiar with.Diogo Almeida [00:25:26]: It's a great question.Swyx [00:25:27]: Yes.Diogo Almeida [00:25:27]: And actually, I will give a related question.Swyx [00:25:29]: Okay.Diogo Almeida [00:25:29]: What is RLHF?Swyx [00:25:31]: Okay.Diogo Almeida [00:25:31]: Right? And actually, RLHF means multiple different things, right?Swyx [00:25:34]: Okay.Diogo Almeida [00:25:34]: Like, there's the RLHF of the original. I think it was, like, Paul Christiano teaching a robot to backflip or something like that. Wasn't there somethingSwyx [00:25:42]: Was that it?Diogo Almeida [00:25:43]: That was the originalSwyx [00:25:44]: I referenced the PPO paper, but I don't know.Diogo Almeida [00:25:46]: And so PPO was not necessarily from human feedback, if I recall.Swyx [00:25:51]: Okay. That's trueDiogo Almeida [00:25:52]: But I b- I believe it was, like, an OpenAI alignment work that could teach hard to specify outputs, like a backflip. I'm not 100% sure. And then there was actually learning to summarize. This was work, by a bunch of the team that helped with, instruct-- and co-authored, the instruction following paper, which was teaching, doing PPO on language models.Swyx [00:26:15]: This is the, sorry. I'm trying to, tryingDiogo Almeida [00:26:19]: YeahSwyx [00:26:19]: Trying to manipulate this thing. This is 2017.Diogo Almeida [00:26:23]: Yeah.Swyx [00:26:23]: Right.Diogo Almeida [00:26:23]: I'm not 100% sure, but like, that looks quite right.Swyx [00:26:26]: Yeah.Diogo Almeida [00:26:26]: If it has, like, a robot doing backflips or something like that might be it. Yes. Okay, cool. I guess I got it right. Hell yeah.Swyx [00:26:35]: There you go.Diogo Almeida [00:26:36]: Yeah.Swyx [00:26:36]: That's the one.Diogo Almeida [00:26:36]: So the idea was can, like, can you do, like, ill-specified things with it? So that's, like, version one. Version two was, the learning to summarize work, that, like, OpenAI did, which is actually, like, PPO on language models to do something somewhat ill-specified. This is, like, another thing that people refer to as RLHF Which I did not co-author.Diogo Almeida [00:26:57]: Oh, Dario's there. Cool. Hell yeah.Swyx [00:27:01]: And Radford.Diogo Almeida [00:27:02]: Yeah. Shout-outs to Alec and Ryan. Love them.Swyx [00:27:04]: Yeah.Diogo Almeida [00:27:05]: But the thing that I refer to RLHF is the, Oh, man.Diogo Almeida [00:27:13]: I'll get toSwyx [00:27:14]: You have comments on that, yeah.Diogo Almeida [00:27:15]: I have comments on that paper, but like, we're so many, tangents deep.Swyx [00:27:18]: Yeah.Diogo Almeida [00:27:18]: So the thing that really got. To me, the thing that I'm calling to RLHF is the task of instruction following. It's not about the PPO. That part doesn't matter. It's about, like, setting a North Star of this is a valuable direction. It's kind of like the Bitris lesson North Star.Diogo Almeida [00:27:34]: And for us, RLCD is this new task. And it is not. I don't see it as jargon. Like, I try to communicate with precision. It's just that, “Hey, here's another North Star.” Just like DPO and all of its, like, descendants also do RLHF, despite not using the algorithm in that paper.Swyx [00:27:55]: And so clear- clearly stating the North Star is, being program- programmable AI is one, word that I really catch onto, removing the human in the loop,Diogo Almeida [00:28:06]: YesSwyx [00:28:06]: From. Because RLHF is tuningDiogo Almeida [00:28:09]: YesSwyx [00:28:09]: For this so that you can automate everything.Diogo Almeida [00:28:11]: Yes. Everything that makesSwyx [00:28:13]: Did I miss anything else in the, in the thesis of, like, what the North Star is?Diogo Almeida [00:28:17]: There is. That is. That is right. I'm overly nuanced in my communication. The one nuance is that we need to be practical. We need to be aware of what language models can do really well. Like what AI can do.Diogo Almeida [00:28:30]: Right? Like, there could be programmatic types that are, like, sick AF, but if you. If the technology is not ready for it to. It's not a tragedy if that's not out in the world.Why Programmable AI MattersSwyx [00:28:41]: Yeah.Diogo Almeida [00:28:42]: But to me, like, the pre-Jev world was a tragedy becau-- it sounds arrogant. Hear me out.Swyx [00:28:49]: No. I strongly believe you.Diogo Almeida [00:28:50]: Cool. It sounds arrogant, but like, I felt this way since long before I even had a company.Swyx [00:28:54]: Yeah. I can, I can vouch that,Diogo Almeida [00:28:56]: Yes, I've been talking about this for so longSwyx [00:28:57]: You said this at All Around Her for, like, three years.Diogo Almeida [00:28:58]: Yeah, I've been talking about this for so long. And I've been saying it because I thought it would have been easier. They say they do not do things because they. It. They're easy. They. It's ‘cause they thought it was easy, soSwyx [00:29:08]: Yeah, exactlyDiogo Almeida [00:29:09]: Something like that. I thought it. This whole project would take a week.Diogo Almeida [00:29:13]: And I was unbelievably wrong. So I am so sorry to everyone at OpenAI that I thought. I was like, “Man, I'm solving this right now.” but like, I think that the tragic thing is when. Well, I think overpromise, underdeliver is tragic too. And like, AI is super extreme on that axis. And I think RLVR is, like, the main. Well, both RLVR and RLHF are extreme perpetrators of this.Diogo Almeida [00:29:40]: But like, it. To me, it's like it's just there's just so much potential there. Like, AI is clearly so smart. I l- smart. I love this in my talks, when I ask people, like, “How can AI be so unbelievably smart? How can we, like, solve millennium prize problems in math, but still not automate even the most basics of works?” Like, really basic rote stuff that, like, the. It d- it doesn't take, like, extremely smart people to do this. It's not a satisfying job. Like, there's other things these people could be doing, but yet we need them to do, like, this ba- like, super basic- non- unsatisfying stuff because, like, we can't automate it yet, but we have this, like, supercharged engine of automation that just does not have, like, the right plugs and stuff to plug into all of this economically valuable work. And like, if the whole company of TypeSafe disappears, like, maybe it'll take, like, a year or two for people to, like, truly catch up. I actually don't know how long it'll take. If model quality matters, then we are gonna be in a very good position for a long time. But it, like, it's done, right? Like, there, like, this has changed the path of, like, technological history.Swyx [00:30:49]: Yeah.Diogo Almeida [00:30:49]: And like, we will be exploring that space as a field.Swyx [00:30:53]: Yeah. I think, I definitely agree with that. You've created possibilities. So I think, if I can paraphrase so that people can un- also understand, you should not take the success of TypeSafe and Jev as just like, “Well, that is a new model type. Now we're done. We go back to business.” Like, no. Like, actually, there's, there are, like, five other model types that you should be exploring and like, let a thousand flowers bloom.Diogo Almeida [00:31:15]: Absolutely.Swyx [00:31:16]: Right?Diogo Almeida [00:31:16]: Like, early internetSwyx [00:31:17]: And some of that, some of which you will probably also build.Diogo Almeida [00:31:18]: Of course, yes.Swyx [00:31:19]: Yes.Diogo Almeida [00:31:19]: Early internet energy. I think it's back to tech utopia. It's no longer like, “Oh, man, like, sometimes my coding agents work, but the, all of the best ones are hoarded internally.”Swyx [00:31:29]: Yeah.Diogo Almeida [00:31:30]: Right? It's like creation is back on the menu.Diogo Almeida [00:31:34]: ? Though it's gonna be a wild-ass world, and buckle up.Diogo Almeida [00:31:38]: It's. And I'm so jazzed about that.Manifesto, Launch Strategy, and Early Internet EnergySwyx [00:31:42]: Yeah. And now you have the funding and the momentum to do whatever you envision there, which I, which I think is, like, very gratifying to see you have after, so long of saying these thingsDiogo Almeida [00:31:53]: YeahSwyx [00:31:54]: But actually show the world.Diogo Almeida [00:31:55]: I know. I just. Such a, such an interesting thing to be a tease the whole time. Like, my talk, like, felt like it was a cliffhanger ‘cause I didn't say how the automation would occur.Swyx [00:32:05]: Yeah.Diogo Almeida [00:32:06]: Sean reviewed our manifesto And he's like, “It's a little bit vague in these parts.”Diogo Almeida [00:32:12]: And like, “What's step one? What is, what is the intelligence model?”Swyx [00:32:16]: Well, I asked you for model, and you were like, “Yeah, model coming.”Diogo Almeida [00:32:18]: Yeah.Swyx [00:32:18]: And like, Well, I just, I mainly objected to the word composable But build prod.god is fantastic.Diogo Almeida [00:32:24]: Thank you.Swyx [00:32:24]: Yeah.Diogo Almeida [00:32:25]: I. We've really rallied around that. I'd like to think we're not entirely a cult like some companies are.Diogo Almeida [00:32:32]: But like, we are, like, jazzed about what we're doing, and like, we are. Like, my brand is being practical, and like, we are all, like, so super-duper practical.Swyx [00:32:42]: Yeah.Diogo Almeida [00:32:42]: It's really great.Swyx [00:32:43]: Yeah. So here. And by the way, here is the step, the secret master plan, right?Diogo Almeida [00:32:47]: Yep.Swyx [00:32:47]: Shape, the shape of machine-native composable AI.Diogo Almeida [00:32:49]: It was your idea to make a secret master plan, soSwyx [00:32:51]: It's a, it's that Elon thing. When he started TeslaDiogo Almeida [00:32:53]: YeahSwyx [00:32:53]: He was like, “Here's what we'll do.”Diogo Almeida [00:32:54]: But I did. Yeah. I'm giving official credit to you.Swyx [00:32:56]: Oh, thank you. Thank you, thank you.Diogo Almeida [00:32:56]: Yeah.Swyx [00:32:56]: Thank you. But like, you should've told me your, you're also gonna do this model launch, ‘cause you, like, you told me, you told me half of the story, and then the other half, you didn't have the doom demo at the time.Diogo Almeida [00:33:08]: Yep.Swyx [00:33:08]: You didn't have any numbers to give me.Diogo Almeida [00:33:10]: Yep.Swyx [00:33:10]: I was like, “what?”Diogo Almeida [00:33:11]: Well, the problem is I don't believe in benchmarking.Swyx [00:33:13]: Exactly.Diogo Almeida [00:33:14]: Right?Swyx [00:33:14]: Exactly.Diogo Almeida [00:33:14]: So like, it is a thing that you need to feel, and like, I think that this is the way to build long-term trust, even though it, like, hurt, it hurt us a, us a lot? Like last year when we did fundraise, no one believed us.Diogo Almeida [00:33:27]: ? Like, and they wanted just benchmarks and stuff, and we're like, “We're not gonna do that. We are principled. We're gonna stand by our guns. That rewards bad actors. I don't give a s**t, like, what you want. Like, this is who we are, and we are standing by that.” So Sorry. It's notSwyx [00:33:43]: No, yeah. Well, and in some ways, I think, like, choosing the hard path, it. But you end up making the company that you wanna work in.Diogo Almeida [00:33:49]: Yep.Swyx [00:33:50]: Right? Otherwise, if you sell out, then you're just working in, like, OpenAI but with my people, right? Which is like.Diogo Almeida [00:33:56]: Yeah. Yeah. Like, I'm, I don't have too many regrets on that, obviously.Swyx [00:34:01]: Yeah.Diogo Almeida [00:34:01]: Like, it worked out so unbelievably well. And like, I, The. I was emotional last night when I was talking about, like, the reasons I left OpenAI, and because, like, it actually had to change my wording after the launch. My phrasing was, “If an AI winter did happen and I did not do every f*****g possible thing I could to, like, avert that, I would see myself as personally responsible both for, the RLHF direction, which I think really widened overpromise versus under-deliver, and also not going all in on this because I think this is, this is where value is going to just be, like, printed.” So. And it was really cool because I feel likeDiogo Almeida [00:34:47]: The AI winter I'm worrying about is averted. Like, AI will be useful. It'll be used for automation.Diogo Almeida [00:34:53]: It's been less than a week, and like, the numbers are already undeniableSwyx [00:34:57]: YeahDiogo Almeida [00:34:57]: That it's, like, being used for real work, and like, there's. It's, it's the Wild West. Yeah.Launch Traction, Tokens, Rate Limits, and Developer UsageSwyx [00:35:03]: Yeah. Can you sh- just if you have top of your head, what numbers are you seeing? Like, what's, what's, like, signups? Like, whatever you can share.Diogo Almeida [00:35:11]: I'm actually not super on top of everything. Like, the team is the ones who are telling me all of these things.Swyx [00:35:16]: Yeah, and I'm sure it's, like, changing every day, right?Diogo Almeida [00:35:17]: It's, it's,Swyx [00:35:18]: But likeDiogo Almeida [00:35:18]: It's kinda nutsSwyx [00:35:19]: If there's a milestone that you're like, “Well, yep, that's one thing we were hoping for. We reached it.”Diogo Almeida [00:35:23]: I will say a milestone that we've passed is tokens per day.Swyx [00:35:27]: Nice.Diogo Almeida [00:35:27]: And this is not, like, fleeting tokens per day.Swyx [00:35:32]: Yeah.Diogo Almeida [00:35:32]: This is, like, even at night, like, it's constantly training, so machines are calling it and not just people trying things out.Diogo Almeida [00:35:39]: So that is, That is so cool. A trillion tokens a day is a lot.Swyx [00:35:45]: Yeah.Diogo Almeida [00:35:45]: So surpassing that is awesome. Signups to me don't really matter. And actually, this was, like, a bit of a mistake we made, if I'm, like, totally honest. People on Twitter were calling us, like, marketing geniuses and all of that, and that was just us. We don't have a marketer. Also hiring. And we were just being our genuine, goofy, like, irreverent selves, and we were, we were just, like, offboarding people off the waitlist so hard. - Our platform team is so unbelievably cracked. I think we have more n- up nines of uptime than Anthropic while having the most Unprecedented launch ever. Like, that is kind of nuts, soSwyx [00:36:21]: YeahDiogo Almeida [00:36:21]: Like, props to them.Swyx [00:36:22]: Yeah.Diogo Almeida [00:36:23]: And the thing we didn't realize. So number one, waitlists, waitlist sign-ups don't matter for, like, a developer platform, in my opinion? I would guess that a large number of them are not even developers. So they go in, they try some queries, and a lot of people don't get it because they are not programming, right? Like, they're just like, “What? This is not a chatbot. Where's my ChatGPT 2?”Diogo Almeida [00:36:45]: Right? But if, like. I haven't exactly calculated this. My sense is that if every single human being in the world, like, just wrote a couple of queries, that would be a rounding error compared to, like, one power user's for loop that is just, like, creating value.Swyx [00:37:01]: Yeah.Diogo Almeida [00:37:01]: And the thing we are-- didn't realize with the waitlist is, like, we could just w- off-board anyone off the waitlist. It doesn't matter. The scary part is rate limits. And then once people start getting value from that, then they just want tons and tons of rate limits because this is what software is, right? Like, you spend effort upfront to specify your rote task, and then this rote task creates more value than it takes to put in. And then now that you have thatSwyx [00:37:25]: Set it and forget, yeah.Diogo Almeida [00:37:26]: Exactly, yeah. You run it in the background. You make it a dependency, to, like, other things. You can make, like, higher level stuff. And like, you just create so much value in the world. Early internet people probably did not imagine, like, the wonder of early 2000s internet, which is still not early internet. But like, it's, it's through, no offense, composabilitySwyx [00:37:47]: NoDiogo Almeida [00:37:47]: That all of the crazy stuff happens, and I just really wanted to emphasize that in our manifesto. We are going for emergence. We are going for, like, being the catalyst. We're wanting to empower people, and we are going to do whatever we can for that, be it, like, Discords in our town hall with me wearing a garbage bag or not.Swyx [00:38:05]: And podcasts and Diogo Almeida [00:38:08]: Hell yeahSwyx [00:38:09]: Getting all that.Diogo Almeida [00:38:09]: Absolutely.Swyx [00:38:09]: Like, ‘cause I want the long form, right?Diogo Almeida [00:38:11]: Yeah.Swyx [00:38:12]: It is like, yes, we'll get past the, some of the superficial things, and then we'll go deep andDiogo Almeida [00:38:15]: Hell yeahSwyx [00:38:15]: And people will really trust and understand your mission and like, the people that, will resonate that will end up joining you or, buying you. Or No, but sorry, as a, as a customer.Diogo Almeida [00:38:27]: Oh, as a customer.Swyx [00:38:28]: As a customer, as a customer.Diogo Almeida [00:38:28]: Okay, yeah. That was funny. I'm sorry.Swyx [00:38:30]: Sorry. I didn't, I didn't mean to say that. But no, any-- one version, one very flattering version of this, like, 36 million views of your launch video.Diogo Almeida [00:38:37]: Cool. Up to 38 now.Swyx [00:38:39]: Yeah, rounding error.Diogo Almeida [00:38:40]: Yeah.Swyx [00:38:40]: Navio still has got 74. Fable 5 got 57. So like, as far as, a- and I didn't, I didn't do the stats for, like, original ChatGPT, likeDiogo Almeida [00:38:48]: YepSwyx [00:38:49]: Which there was no video.Diogo Almeida [00:38:50]: Yep.Swyx [00:38:50]: So like, up there, right?Diogo Almeida [00:38:52]: Yep.Swyx [00:38:52]: Like, as far, as far as, like, if you were to launch a Neolab in 2026, I think you're, like, number one right now, which is, like, pretty crazy.Diogo Almeida [00:38:58]: Yeah. Well, I actually would rather. I do have the shirt, like, your favorites Neola-- favorite Neolab's favorite Neolab.Swyx [00:39:05]: Huh.Diogo Almeida [00:39:05]: I don't give a s**t about being a Neolab. I think being a Neolab. Actually, we have a lot of, like, swag that's being a parody of a Neolab. One of them, one of them I have is, like, Neolab with product, which actually is not a Neolab. Like, I don't care about that, really.Swyx [00:39:20]: Yeah.Diogo Almeida [00:39:20]: What I care about is being a reliable dev platform. So Swyx [00:39:23]: YesDiogo Almeida [00:39:24]: Appreciate the comparison, but likeSwyx [00:39:25]: YeahDiogo Almeida [00:39:25]: Hopefully we transcend past them and we go back into, like, a thing-- like, a revolutionary moment for developers and like, this stable thing that people can rely on and trust.Reliability, Robustness, and DeterminismSwyx [00:39:35]: Yes. To that end, I think that's one thing that really impressed me about you guys is that, yes, you do talk about reliability. I thought it was mostly about calibration, which, like, we talk about RLCD. But actually it's also about just, like, uptime and scalability and all those things, right? They're, they're all sort of the kind.Diogo Almeida [00:39:55]: And nines.Swyx [00:39:56]: And nines.Diogo Almeida [00:39:56]: It's, likeSwyx [00:39:57]: Which uptime is, in my opinion.Diogo Almeida [00:39:58]: Oh, but that's part of it. But like, there's reliability in, like, how intelligent the thing is. Like, how consistently does it do the thing that you want? And I think that, like, the big reasoning models are very smart. In my opinion, they still lack reliability. I think there's many use cases where you-- they look like they should be smart enough to automate their work. There is economic incentive to automate that work, yet still they're not reliable enough as, at an intern because they're optimized for different things. And so like, I think that there's the reliability of being able to, like, trust the outputs. And also we are. Like, there are dimensions of reliability that we are not yet at that I'm, like, so excited by.Swyx [00:40:38]: Yeah.Diogo Almeida [00:40:38]: Like, I want to automate the easy work before the hard work? Like, I think that's just a common sense thing to do. But to me, we will be sufficient. I don't know if there's such thing as sufficiently reliable, but I wanna get so good that people don't even need to try the model to know that it'll work. It's like, that's like what flow state is in programming, right? Like, I'm just, like, writing queries because I need intelligence in here. And like, when. For non-trivial branching, I can just write it in like a, like a type-safe System 1 query and then get the results out of it and it just branches accurately. Like, that would be so good. Like, that's the. That is the dream.Swyx [00:41:12]: Yeah.Diogo Almeida [00:41:12]: And that is, like, going to be, like, a long slog.Swyx [00:41:16]: Yeah. We're gonna go into your API design in a little bitDiogo Almeida [00:41:19]: OohSwyx [00:41:19]: Just to give people examples and like, maybe paths not taken, that kind of stuff.Swyx [00:41:23]: One thing up the front that I do wonder about in terms of reliability is I noticed that there's no seed. There's no, And so basically, same input, do I always get the same output?Diogo Almeida [00:41:34]: SoSwyx [00:41:36]: And if not, why not?Diogo Almeida [00:41:37]: Oh, great question. So this is actually, like, a common question we have between. So reliability is actually a catchall. Like, whenever AI can't automate something, it's due to some form of reliability. Could be, like, type safety. It could be determinism. It just could be, like, it's, it's jagged, right? So reliability is a catchall. I just think that it's also a catchall for, like, what the North Star is. Re- determinism is, like, same inputs, same outputs. I do believe that this is, like, slightly interesting for unit tests, but I believe that to be the wrong North Star. I believe robustness is what peopleDiogo Almeida [00:42:16]: I don't wanna tell people what they really want, ‘cause that would be a little arrogant of me.Diogo Almeida [00:42:19]: I believe that is, like, the more important property. You want, given similar inputs, get similar outputs. And it's kind of wild how unreliable LLMs are.Diogo Almeida [00:42:31]: Like, a way that we test this is you put, like, UUIDs in, like littleSwyx [00:42:36]: YeahDiogo Almeida [00:42:36]: I think they're called nonces In the prompt. And what you want is similar outputs from all of those, ‘cause it's truly semantically the same question, and that is the part where you really want. Th- like, that robustness is where, like, people get, like, burnt with AI making decisions. So I think that is the. A super-duper important property. We could also have determinism. That is, that is a thing that can be available. As far as I can, like, mentally model for programmers, like, it, I- it could be valuable for some use cases, so like, please educate me, in comments or view. But my. In general, it's easy. Determinism is something you can, like, trade off for better cost. Like, we are, we are constantly wanting to be on the intelligence per dollar frontier. We are doing, like, absolutely disgusting things to be there. Like, this is,Diogo Almeida [00:43:32]: I shouldn't say this, but no one's here to stop me.Swyx [00:43:37]: If you s- you sign off on your own PR.Diogo Almeida [00:43:40]: That is not how it works at this company. I believe for this week, my chief of staff, Kay, is the most powerful person in tech.Swyx [00:43:49]: Yeah. And shout-out to Kay for organizing this.Diogo Almeida [00:43:50]: Holy shSwyx [00:43:51]: Yeah.Diogo Almeida [00:43:51]: Holy s**t. She is so f*****g competent and powerful. She's incredible.Diogo Almeida [00:43:58]: She sucks. Don't poach her. But so I try to be a bit more filtered, but like, people are telling me, “Don't call it a Frankenstein's monster of models,” but because that has, like, negative implications. I think Frankenstein's monster was, like, the good guy in this whole. It was innocent, right? I didn't read it. Okay.Diogo Almeida [00:44:18]: I'll, I'll confess. Okay. That. Well, one facial expression, ISwyx [00:44:21]: This is aDiogo Almeida [00:44:21]: My cards on the tableSwyx [00:44:21]: Decent Jacob Elordi movie if you wanna seeDiogo Almeida [00:44:24]: ISwyx [00:44:25]: The adaptation. Anyway.Diogo Almeida [00:44:26]: The. You have no idea how little time I have right now.Swyx [00:44:28]: Yeah.Diogo Almeida [00:44:29]: My priorities are sleep?Swyx [00:44:31]: Developers.Diogo Almeida [00:44:32]: Developers, yes. Developers. But yes. It. We do, like, absolutely disgusting things to be on the Pareto curve of intelligence per dollar, and we are going to keep doing that.Swyx [00:44:47]: Yeah.Diogo Almeida [00:44:47]: We're gonna be doing crazy-ass stuff, and I think people really need to think outside of the box. Like, part of the reason we're surprising is, like, people Are thought inside the box, and we continue to do that. As of right now, we are obviously the best at this, and we want to continue being the best at that whole thing.Swyx [00:45:05]: Yeah.Diogo Almeida [00:45:05]: So Wait, where did, where did we tangent from?Swyx [00:45:07]: No. SoDiogo Almeida [00:45:08]: YeahSwyx [00:45:08]: I asked you about, will you have seeds and determinism?Diogo Almeida [00:45:11]: Oh, yes. SoSwyx [00:45:11]: And then you basically defined reliability and likeDiogo Almeida [00:45:14]: And robustnessSwyx [00:45:15]: How you see it. Yes.Diogo Almeida [00:45:16]: But like, determina- likeSwyx [00:45:17]: I have a robustness example that's, that's, real quick I can show you.Diogo Almeida [00:45:19]: I would love that. I will just say one thing.Swyx [00:45:21]: Yeah.Diogo Almeida [00:45:21]: We can make a deterministic model.Swyx [00:45:22]: Exactly.Diogo Almeida [00:45:23]: Like, we're hap- if people can convince us that is a valuable thing to doSwyx [00:45:27]: YeahDiogo Almeida [00:45:27]: And we don't have a gigantic GPU shortageSwyx [00:45:29]: YeahDiogo Almeida [00:45:29]: We can happily make all of these models. We live to please. And rev- and revolt, revolute,Swyx [00:45:38]: You will throw over everything, except you'll do it in a nice way.Diogo Almeida [00:45:41]: Yeah.Swyx [00:45:41]: And findDiogo Almeida [00:45:42]: So like, determinism could be on the cards.Swyx [00:45:44]: Yeah.Diogo Almeida [00:45:45]: It just gets you less intelligence per dollar.Swyx [00:45:46]: Yeah. Well, just having seen the trajectory of OpenAI and Anthropic, you will. Just trust me now that you will be peer pressured into doing it. So like, just people will want it even if they. If you tell them they don't need it. They'll still want it. So like, yeah, that's the TL;DR of that.Diogo Almeida [00:46:01]: Okay.Swyx [00:46:02]: Yeah.Diogo Almeida [00:46:02]: I will love to. Maybe one day we will see how that happens.Swyx [00:46:07]: Yeah.Diogo Almeida [00:46:07]: I've been told I'm, They say that part of our brand is being unshakeableSwyx [00:46:13]: HuhDiogo Almeida [00:46:13]: And they say that's just the nice way of saying stubborn.Swyx [00:46:15]:
Topics covered in this episode: Pandas Should Go Extinct Pydantic-pint puts real-world units in your Pydantic models How Libraries Run Rust Inside Python (With PyO3) AWS acquires DuckLabs Extras Joke Watch on YouTube Sponsored by Logfire from Pydantic: pythonbytes.fm/logfire Connect with the hosts Michael: Mastodon / BlueSky / X / LinkedIn Calvin: Mastodon / BlueSky / X / LinkedIn Show: Mastodon / BlueSky / X Join us on YouTube at pythonbytes.fm/live to be part of the audience. Usually Tuesday at 7am PT. Older video versions available there too. Finally, if you want an artisanal digest of every week of the show notes in email form? Add your name and email to our friends of the show list, we'll never share it. Calvin #1: Pandas Should Go Extinct Pandas' slowness pushes teams toward "Big Data" tools (Spark, Databricks) they don't actually need — most workloads never hit true Big Data scale Amazon Redshift telemetry: ~95% of tables are under 100GB, ~87% of queries touch 80GB or less — that's "Medium Data," not Big Data Polars and DuckDB fill that gap: single-machine, fast, no cluster required 1 Billion Row Challenge benchmark: Pandas took 4m28s vs. Polars 5.04s and DuckDB 5.19s — DuckDB also used 19x less memory On a real-world NYC taxi dataset (3GB parquet), pure DuckDB ran 2x faster than pure Pandas while using a fraction of the RAM Bonus: Apache Arrow lets you pass data between Pandas/Polars/DuckDB with zero copying, so trying them out doesn't mean a full rewrite Michael #2: Pydantic-pint puts real-world units in your Pydantic models Pydantic-pint bridges Pydantic and Pint so models can validate physical quantities like 4m or 12 meters instead of bare floats. Fields annotated with PydanticPintQuantity parse user input, convert between compatible units, and serialize quantities back out as strings. That closes a real gap for anything consuming API payloads, config files, or sensor data with measurements, letting you enforce units at the validation boundary instead of hoping every caller remembered them. via PyCoder's Weekly newsletter Unit mix-ups have literally crashed spacecraft; now your Pydantic models can refuse them at the door. Annotate a field as Annotated[Quantity, PydanticPintQuantity('km')] and inputs like 12 meters arrive auto-converted to kilometers Validation covers string, numeric, and quantity inputs, and model_dump_json serializes quantities as readable unit strings Installable from PyPI as pydantic-pint, MIT licensed, with docs at pydantic-pint.readthedocs.io Early-stage solo project at version 0.4, so API stability and maintenance are open questions worth discussing Calvin #3: How Libraries Run Rust Inside Python (With PyO3) Pydantic v2's validation core (pydantic-core) is Rust under the hood, built with PyO3 — this post shows how that bridge actually works via a small hand-built JSON parser Four steps to get Rust into Python: write a normal Rust module, annotate with PyO3 macros (#[pyfunction], #[pymodule]), compile/install with maturin, then just import it The parser builds a Rust tree first — Python never touches it until the boundary crossing Key insight: converting the Rust result into Python objects (.into_pyobject) is often the expensive part, not the parsing — 100,000 JSON values means ~100,000 Python objects built after parsing's already done Errors cross the boundary too: Rust's typed errors convert into real Python exceptions (ValueError, FileNotFoundError) via From/?, so callers get clean Python semantics Takeaway for anyone porting Rust in: if you're returning a scalar, don't sweat it; if you're returning a big structure, profile the boundary — that's the real cost, not the algorithm Michael #4: AWS acquires DuckLabs Thank you Dylan McConnell. What does this mean for the DuckDB ecosystem? DuckDB is the open-source in-process analytical SQL engine. MIT licensed. The IP is not owned by any company - it's held by the nonprofit DuckDB Foundation, which was created when the team spun out of CWI Amsterdam. Peter Boncz, the CWI representative on the Foundation board, describes it as the entity that holds all IP of open-source DuckDB. DuckLabs (ducklabs.com) is the company, formerly branded DuckDB Labs. Founded a little over five years ago by Hannes Mühleisen and Mark Raasveldt to give the DuckDB team a stable long-term home, bootstrapped deliberately instead of taking VC, grown to 30+ people in Amsterdam, funded by support and feature-prioritization contracts. It employs the core devs. It does not own DuckDB. DuckLake is one of three projects DuckLabs builds, what they call the Duck Stack: DuckDB, DuckLake, and Quack. DuckLake is the lakehouse format that puts catalog metadata in a SQL database instead of in files on object storage. Quack is newer - an RPC-style protocol that turns DuckDB into a client-server system where both ends are DuckDB instances, slated to stabilize in DuckDB v2.0 in September 2026. MotherDuck is a separate Seattle company, Jordan Tigani's, selling serverless hosted DuckDB. It was started in partnership with DuckDB Labs and has worked closely with Hannes and Mark for four years. It contracted DuckLabs for engineering work and contributes heavily upstream - three of its engineers are among the top 10 outside contributors to DuckDB. It also sells its own DuckLake offering. Customer and collaborator, never owner. What the AWS post changes. Amazon bought the company, not the project. DuckLabs joined AWS effective September 1, with the process concluding August 31, 2026. Hannes and Mark keep leading the team and the project's technical direction, the team stays in Amsterdam, and DuckDB stays MIT under the Foundation. AWS gets the people and a direct line to the roadmap. The license protects your code, not your priorities. Three second-order effects worth tracking: The Foundation board is the real question. It has three directors: Mühleisen, Raasveldt, and Boncz. Two now work for AWS. Commentary on the deal has focused on exactly this - the license protects the code, not the roadmap. The announced counterweight is governance: a technical advisory board on the Foundation, and opening the extension stack so extensions signed by other developers can run in DuckDB. MotherDuck immediately moved into the business DuckLabs vacated. It now sells DuckDB enterprise support, which it had avoided because it didn't want to compete with DuckLabs' business model, and says it has explicit blessing from Hannes and Mark now that they're joining Amazon. It also bought Tower.dev the day before the AWS announcement. Everyone expects an AWS DuckDB service. Tigani says Amazon will likely release one eventually, and welcomes the competition, citing Redshift's failure to slow Snowflake on AWS. The groundwork is already visible: Amazon Quick uses DuckDB to query S3 Tables and has processed over 2.5B queries with it since launching in October 2025. The DuckLake angle is the one to watch. AWS is heavily committed to Iceberg through S3 Tables, and it just acquired the team behind a competing lakehouse format. The stated plan is to use DuckDB, DuckLake, and Quack together to power a new generation of data services, but which format wins internal priority is unannounced. Extras Calvin: astral-sh/uv 0.12.12: code-signed release binaries
Most fault trees get built on gut feeling. Petra Vukmirovic did something rarer: she borrowed the actual math from aviation and nuclear-plant safety engineering and pointed it at AI agents. Petra traded emergency medicine for application security and now heads information security at Numan — and she joins Chris Romeo and Robert Hurlbut to make the case for fault tree analysis (FTA), the deductive method that picks up exactly where threat modeling stops. Petra walks through a "wrong customer refund" AI agent scenario step by step, showing how AND/OR gates and minimal cut sets turn vague worry into ranked, data backed probabilities. They dig into where AI helps build a tree, and where garbage in, garbage out still applies, why "comprehensive test coverage" is a myth, and how attaching real dollar figures to failure paths makes it easier to sell security controls to leadership.This episode is sponsored by Corgea. Design it. Build it. Ship it. Corgea secures it.About CorgeaCorgea is an AI-native application security platform that secures software from design to production. It brings together security design reviews, AI SAST, dependency and IaC scanning, code quality checks, and autonomous pentesting—helping security and engineering teams find risk earlier, fix what matters, and ship securely.→ Learn more about CorgeaConnect with Petra Vukmirovic:→ Petra Vukmirovic on LinkedIn→ OWASP Threat Model LibraryMentioned in this episode:→ Adam Shostack: "Stop Trying to 'Manage Risk'" (keynote)→ OWASP Global AppSec USA 2026 (San Francisco, Nov 5–6)Follow the Application Security Podcast:➜ Home: appsecpodcast.com➜ X: @AppSecPodcast➜ LinkedIn: The Application Security Podcast➜ YouTube: @ApplicationSecurityPodcast➜ Instagram: @appsecpodcast➜ Facebook: Application Security PodcastChapters:00:00 Cold open — the math behind where to put your controls01:09 Meet Petra Vukmirovic01:28 Petra's origin story: from ER doctor to AppSec02:50 Career path: engineer to Head of InfoSec at Numan04:18 What is fault tree analysis, and where threat modeling ends06:22 Can AI actually do fault tree analysis?08:12 Walking the "wrong customer refund" agent example12:33 Storing your trees: JSON vs. Markdown16:08 Why conjunctive failures trip up narrow thinking17:27 Top 3 failure modes when agents touch downstream systems19:29 Real story: an agent pushed code to main without approval21:27 Testing: why "comprehensive coverage" is a myth23:54 How rough is rough? Assigning probabilities28:08 Getting started without a six week science project31:35 Using FTA to sell controls and build credibility33:43 The epiphany: FTA is about controls, not faults34:48 The one thing every agentic team should add today35:48 Closing thoughts and OWASP Global AppSec USA preview
AI agents bring new demands to the infrastructure connecting them to models, tools, and data. Each agent can have its own permissions and state, and spend much of its time asleep before waking up to make another request. Supporting that behavior at scale changes how the network routes traffic, loads policies, and manages resources.In this episode of Alexa's Input (AI), I sit down with Yan Avlasov, Staff Software Engineer at Google and a senior maintainer of Envoy, to talk about building the networking stack for agents.We get into why Google is building on Envoy, what changes between serving inference and supporting agents, and why enforcing policies for individual agents requires a more dynamic control plane. Yan also shares what he's seeing from AI-powered security scanning, including how models combine subtle bugs into serious vulnerabilities and why fixing them can mean changing behavior that production systems have relied on for years.From the episode:- Model selection, cost controls, and fallback across providers- Loading policies dynamically as agent traffic changes- Managing personalized state and agents that suspend and resume- Where Model Context Protocol (MCP) fits alongside the protocols agents already use- Semantic routing and choosing models based on what a request needs- Security findings involving policy enforcement, path normalization, and JSON parsing- The engineering work and cost of making security scanning continuousChapters00:00 Introduction01:12 Welcome, Yan02:30 How AI is changing Google07:56 Yan's networking background10:00 New networking layers for agents13:11 AI gateways, cost controls, and fallback16:38 Why Google builds on Envoy22:21 Per-agent policies and a dynamic control plane26:56 Inference versus stateful agents31:58 Project Substrate: running agents at scale33:31 MCP and the other protocols agents use37:32 Semantic routing and model costs40:40 AI security scanning and the work of fixing bugs43:55 Policy timing, path normalization, and JSON47:11 Making security scanning continuous50:40 What's next for agentic networking53:32 ClosingRead the episode article on SubstackWatch on YouTube
This episode of the DevCast looks at how JavaScript, HTML, JSON, Web Viewers, and AI can work together to build more flexible and dynamic FileMaker solutions. A few of our devs walk us through real-world examples, like a multi-image upload and rotation tool and a dynamic dining room and table management interface.The conversation strides into a bigger question: How much do you really need to understand when AI is writing the code for you? Our devs discuss finding the right balance between letting AI handle the heavy lifting and understanding enough of what's happening under the hood to build, test, troubleshoot, and maintain a reliable solution.Multi-Image Upload Blog Pt. 1: https://www.portagebay.com/blog/multi-image-uploading-with-gallery-and-rotation-free-demo/Multi-Image Upload Blog Pt. 2: https://www.portagebay.com/blog/using-webviewers-to-add-functionality-for-field-technicians/
Côté IA : MCP devient stateless, Claude watermarke ses textes, GPT-6 Astra défie Claude Fable, et une étude JetBrains confirme Claude Code en tête des agents de code. Côté JVM : JDK 27 généralise G1, Kotlin 2.4 stabilise les context parameters, une API JSON arrive dans le JDK, et Quarkus comme Micronaut enchaînent les versions. En bonus, trois pannes IA simultanées et un câble débranché chez Google Cloud. Enregistré le 11 septembre 2026 Téléchargement de l'épisode LesCastCodeurs-Episode-343.mp3 ou en vidéo sur YouTube. News Langages Les Types Algébriques de Données (ADTs) en Java rockthejvm.com/articles/algebraic-data-types-in-java Le problème : L'approche classique (champs nullables, hiérarchies de classes ouvertes) crée des états invalides et des erreurs à l'exécution (comme le NullPointerException). Types Produits (ET logique) : Implémentés en Java avec les Records. Ils regroupent plusieurs champs de manière immuable et concise. Types Sommes (OU logique) : Implémentés avec les Sealed Interfaces. Elles définissent un ensemble strictement fermé de sous-types connus à la compilation. ADTs (Types Algébriques) : La combinaison des Sealed Interfaces et des Records. Ils garantissent que les états invalides sont impossibles à représenter dans le code. Pattern Matching : L'extraction des données se fait via des expressions switch exhaustives, supprimant le besoin de casts manuels et obligeant le développeur à traiter tous les cas possibles. Généralisation : Ce modèle est idéal pour créer des types comme Result, forçant le traitement explicite et sécurisé des succès et des erreurs typées. Pourquoi "Algébrique" ? Parce que les types sont combinés mathématiquement (Produits = multiplication, Sommes = addition) pour limiter strictement le nombre d'états possibles d'une donnée. JDK 27 : fonctionnalités et calendrier de sortie openjdk.org/projects/jdk/27 infoworld.com/article/4202901/jdk-27-the-new-features-of-java-27.html JDK 27 est la prochaine version majeure de Java, une version non-LTS avec seulement 6 mois de support, qui succède à JDK 26. La disponibilité générale est prévue pour le 15 septembre 2026, avec des release candidates les 6 et 20 août 2026. Le périmètre est désormais figé (feature freeze) avec neuf JEP au programme. Le ramasse-miettes G1 devient le collecteur par défaut dans tous les environnements, et plus seulement en mode serveur. Ajout d'un support de la cryptographie post-quantique pour TLS 1.3, via des échanges de clés hybrides combinant algorithmes classiques et résistants au quantique. Finalisation de l'API PEM pour encoder et décoder clés, certificats et listes de révocation au format PEM. L'API Vector poursuit son incubation pour la douzième fois, permettant d'exprimer des calculs vectoriels compilés en instructions CPU optimisées. Les en-têtes d'objets compacts, introduits en JDK 24, sont désormais activés par défaut et réduisent l'empreinte mémoire du tas. Plusieurs previews sont reconduites : constantes paresseuses (3e preview), types primitifs dans les patterns (5e preview) et concurrence structurée (7e preview). Ajout d'une fonctionnalité de rédaction in-process pour JFR, afin de masquer les données sensibles dans les enregistrements de profiling. Kotlin 2.4 : nouveautés du langage et outillage kotlinlang.org/docs/whatsnew24.html Kotlin est un langage moderne, multiplateforme (JVM, Android, iOS, JavaScript, Wasm) développé par JetBrains, souvent utilisé comme alternative à Java. Les context parameters passent en stable : ils permettent de fournir des dépendances implicites à une fonction sans les déclarer en paramètre explicite, un peu comme une injection de dépendances. Les collection literals arrivent en expérimental : on peut écrire une liste avec des crochets, comme en Python, par exemple val fruits = ["pomme", "banane"]. L'API UUID de la bibliothèque standard devient stable, pour générer et manipuler des identifiants uniques nativement. Nouvelles fonctions utilitaires comme isSorted() pour vérifier si une collection est déjà triée. Support de Java 26 côté JVM et alignement automatique des versions Java et Kotlin dans les projets Maven. Kotlin/Native, la compilation vers du code natif iOS et macOS, active par défaut un nouveau ramasse-miettes plus rapide et améliore l'export vers Swift. Kotlin/Wasm, la compilation vers WebAssembly pour faire tourner du Kotlin dans le navigateur, rend la compilation incrémentale stable. Kotlin/JS permet désormais d'exporter des value classes vers JavaScript et TypeScript. Le compilateur K1, l'ancienne génération, n'est plus supporté : seul le nouveau compilateur K2 reste disponible. JEP 540 : une API JSON simple intégrée au JDK (incubation) openjdk.org/jeps/540 Le JDK ne propose aujourd'hui aucune API JSON native, obligeant à dépendre de bibliothèques externes comme Jackson, Gson ou Jakarta JSON pour parser ou générer du JSON. Cette JEP remplace la JEP 198 de 2014 et cible JDK 28 avec le nouveau module incubateur jdk.incubator.json. L'objectif est de couvrir les besoins simples d'extraction de données sans binding de données ni API de streaming, en laissant ces cas avancés aux bibliothèques existantes. L'API s'articule autour de l'interface scellée JsonValue avec six sous-types : JsonString, JsonNumber, JsonBoolean, JsonNull, JsonObject et JsonArray. Le parsing est strict et conforme à RFC 8259 : pas de virgules finales, pas de commentaires, et les noms de membres dupliqués provoquent une erreur. Donc pas de JSON5 La navigation se fait via get et tryGet, et la conversion vers des types Java via asInt, asLong, asDouble, asString, asMap ou asList. En cas d'erreur, une JsonValueException précise le chemin exact dans le document et sa position en ligne et colonne. Le pattern matching sur les sous-types de JsonValue permet de gérer proprement l'évolution du format d'un document JSON dans le temps. La génération se fait via toString pour une sortie compacte ou Json.toDisplayString pour une sortie indentée et lisible. À terme, le JDK pourrait utiliser cette API en interne, par exemple pour remplacer les fichiers de configuration au format property par du JSON. Autres nouvelles du JDK openjdk.org/jeps/535 openjdk.org/jeps/541 le mode generationel pour Shenandoah est prévu par défaut et deprécue le non générationel en 28 fini le support de Java sur Apple Intel GraalVM 25.2 : références compressées et Graal Script Agent medium.com/graalvm/… GraalVM est une machine virtuelle polyglotte d'Oracle offrant compilation JIT avancée et compilation en image native pour accélérer les applications Java et d'autres langages. Cette version 25.2 fait partie du train de releases innovation qui livre les nouveautés plus vite, pendant que GraalVM 25.0 reste la version stable recevant les correctifs de sécurité critiques. Nouveauté phare, le Graal Script Agent transforme des demandes en langage naturel en plugins sandboxés exécutés localement, en JavaScript ou Python, avec un accès restreint aux APIs de l'application. Les références compressées sont désormais activées par défaut dans Native Image sur les systèmes 64 bits, remplaçant les adresses complètes par des valeurs 32 bits relatives au tas. Cette optimisation réduit de 39% la consommation mémoire RSS d'une application Micronaut connectée à Oracle Database, comparée à la version 25.0. Contrepartie de cette optimisation, le tas géré est désormais plafonné à 32 Go. Le garbage collector G1 est maintenant disponible sur toutes les plateformes, y compris Windows, via l'option –gc=G1. G1 apporte de meilleures performances, une latence réduite et un démarrage plus rapide, avec des images natives plus petites grâce à l'optimisation guidée par profil. Le Vector API de Java est activé par défaut pour exploiter les instructions SIMD, utile pour le machine learning et le traitement de données. Bonne intégration avec l'écosystème via Micronaut 5.1, Quarkus et WebAssembly. Shopify arrête React Native pour ses applis mobiles iOS et Android et repasse à du natif avec Swift et Kotlin shopify.engineering/back-to-native Les progrès majeurs des LLM (IA) réduisent drastiquement le coût du développement sur deux plateformes distinctes. Les bénéfices du natif pur restent supérieurs, moins de couches d'abstractions, de dépenfances externes, et plus rapide pour adopter les dernières fonctionnalités des OS Les bibliothèques open-source (Skia, FlashList, Restyle) évoluent : Skia sera forkée par William Candillon, FlashList cherche un nouveau repreneur, Restyle sera archivée fin 2026. Migration des applications (Shop, Shopify, etc.) réalisée en mode "greenfield" (reconstruction totale) assistée par IA. Utilisation du système "Helix" pour un développement itératif et contrôlé par des agents IA. Découplage de la logique métier et de l'interface via une CLI pour accélérer les tests et éviter les lenteurs des simulateurs. L'application Shop a été entièrement reconstruite en natif en 12 semaines ; les autres suivront. Librairies LangChain4j CDI est une extension CDI qui intègre LangChain4j avec CDI de Jakarta EE langchain4j.github.io/langchain4j-cdi LangChain4j CDI : Extension intégrant LangChain4j à Jakarta EE et MicroProfile. Services IA : Injection et gestion de cycle de vie via @RegisterAIService. Orchestration d'agents : 11 topologies d'agents configurables par annotations. Serveur MCP : Conversion de beans CDI en serveurs Model Context Protocol. Fonctionnalités d'entreprise : Configuration externe, tolérance aux pannes et observabilité OpenTelemetry. Installation Maven : Deux extensions disponibles selon l'environnement (build-time pour Quarkus/Helidon, portable pour WildFly/GlassFish/Liberty). Prérequis techniques : Java 17+, Jakarta EE 10, MicroProfile 6.1. Quarkus 3.36, 3.37 et 3.38 : trois releases avant Quarkus 4 quarkus.io/blog/quarkus-3-38-released quarkus.io/blog/quarkus-3-37-released quarkus.io/blog/quarkus-3-36-released Quarkus est un framework Java cloud natif optimisé pour GraalVM et HotSpot, conçu pour les microservices et les environnements conteneurisés. En 3.38 (29 juillet), l'équipe allège les nouveautés pour se concentrer sur Quarkus 4, la communauté atteint 1213 contributeurs. 3.38 introduit l'éviction basée sur le poids mémoire pour le cache Caffeine de second niveau d'Hibernate, en plus de l'éviction par comptage. 3.38 apporte l'extension Quarkus HTTP Problem qui implémente la RFC 9457 pour mapper les exceptions en réponses application/problem+json, intégrée à OpenAPI. 3.37 (24 juin) ajoute l'extension expérimentale quarkus jlink pour générer des images runtime JDK sur mesure et réduire la taille des conteneurs. 3.37 active par défaut la sérialisation Jackson sans réflexion pour de meilleures performances. et 3.39 lesdesactivent et les rement en opt-in 3.37 introduit dans REST Client RestMultiResponse pour lire codes de statut et en-têtes sur des réponses REST en streaming, avec passage à Hibernate ORM 7.4 qui exige PostgreSQL 14 minimum. 3.36 (27 mai) propose Quarkus Signals en expérimental, un système de communication typée entre composants inspiré des events CDI et de l'EventBus Vert.x. 3.36 embarque des SBOM applicatifs exposés via /.well-known/sbom, y compris en image native selon la spécification GraalVM. 3.36 ajoute l'authentification OIDC via JWT SPIFFE, facilitant l'identité de charge de travail en environnement zero trust. Micronaut Framework 5.1.0 : injection de dépendances, IA et sécurité renforcées github.com/micronaut-projects/micronaut-platform/…/v5.1.0 Micronaut est un framework JVM pour microservices et applications cloud-natives, avec injection de dépendances à la compilation et démarrage rapide. Introduction d'Open DI 1.0.0, une implémentation CDI Lite s'appuyant sur l'infrastructure d'injection de dépendances de Micronaut. Côté données, support officiel de SQLite et intégration MyBatis, avec ETags basés sur les valeurs pour le verrouillage optimiste. En sécurité, arrivée d'un module OWASP HTML Sanitizer, délégation d'authentification @RunAs et résolution de locale via OIDC. Côté IA, LangChain4j ajoute le support Chroma, la mémorisation de chat Oracle et l'authentification Google injectée pour Vertex AI, avec passage du MCP en version 2.0.0. Mises à jour majeures des dépendances : Spring Boot 4.1.0, Jetty 12.1.10, Tomcat 11.0.23, OpenTelemetry 1.64.0 et Kubernetes Java Client 27.0.0. SSL activé par défaut par service pour les clients HTTP Infrastructure Ça coûte combien de faire tourner un LLM local sur son Apple Silicon ? towardsdatascience.com/how-much-does-a-local-llm-actually-cost-to-run-i-measured-every-watt-on-apple-silicon Coût électrique des LLM locaux sur Mac Apple Silicon Un modèle 120B (MoE) coûte 5x à 10x moins cher qu'un modèle 27B (Dense). Le coût dépend du débit (tokens/seconde), pas du nombre de paramètres. Modèle dense –> Charge 100% des poids par token = lent et très énergivore. MoE (Mixture of Experts) –> N'active qu'une fraction des poids = rapide et économe. Conclusion : Pour réduire la facture électrique, choisir des modèles MoE quantifiés (haut débit). Kubernetes 1.36 (Haru) : sécurité renforcée et alignement IA infoq.com/news/2026/05/kubernetes-1-36-released Kubernetes est la plateforme open source de référence pour l'orchestration de conteneurs, portée par la CNCF. La version 1.36 nommée Haru apporte 70 améliorations : 18 passent stables, 25 en bêta et 25 en alpha, avec 106 entreprises et 491 contributeurs. Les user namespaces passent en disponibilité générale, isolant le root du conteneur de celui de l'hôte. Les Mutating Admission Policies passent en GA, remplaçant les webhooks par des règles CEL natives plus performantes. L'autorisation de l'API kubelet devient plus fine, remplaçant le droit trop large nodes/proxy. Le labeling SELinux des volumes utilise désormais mount -o context, accélérant le démarrage des pods. Plusieurs avancées ciblent les charges IA : gang scheduling en bêta, préemption consciente des groupes de pods, et allocation dynamique de ressources activée par défaut pour le partage fin des GPU. Le redimensionnement vertical des pods en place passe en bêta et activé par défaut, ajustant CPU et mémoire sans redémarrage. Suppression du plugin gitRepo, source de risque de sécurité, et du mode IPVS de kube-proxy, tous deux dépréciés de longue date. Avec la sortie de 1.36, la version 1.34 devient la plus ancienne branche encore supportée et entre en maintenance, ne recevant plus que des correctifs critiques avant sa fin de support. Terraform vs OpenTofu en 2026 : la divergence est actée ecorpit.hashnode.dev/terraform-vs-opentofu-in-2026-the-fork-has-diverged-so-which-do-you-standardize-on env0.com/insights/opentofu-in-2026-what-the-terraform-fork-became-after-three-years-of-independence Terraform est l'outil historique d'Infrastructure as Code de HashiCorp, OpenTofu en est le fork open source lancé après le changement de licence. HashiCorp est passé de la licence MPL 2.0 a la BUSL 1.1 en aout 2023, ce qui a poussé une partie de la communauté a créer OpenTofu sous la Linux Foundation. IBM a racheté HashiCorp pour 6,4 milliards de dollars, finalisé en février 2025, tandis qu'OpenTofu rejoignait le CNCF comme projet sandbox en avril 2025. OpenTofu prend de l'avance sur des fonctionnalités inédites : chiffrement du state côté client depuis la v1.7, valeurs éphémères qui gardent les secrets hors du state depuis la v1.11, et prevent_destroy dynamique en v1.12 (mai 2026). Terraform garde l'avantage sur l'orchestration managée avec Terraform Stacks, désormais en disponibilité générale, sans équivalent natif côté OpenTofu. Le coût diverge fortement : HCP Terraform facture jusqu'à 0,99 dollar par ressource gérée et par mois, alors qu'OpenTofu reste une CLI gratuite couplée au backend de son choix. Fidelity Investments a migré plus de 2000 applications et 50000 fichiers d'état vers OpenTofu, la complexité venant surtout de l'écosystème (CI/CD, gouvernance) plutôt que du binaire lui-même. OpenTofu reste compatible avec les configurations Terraform jusqu'a la version 1.6.x, mais les versions Terraform plus récentes n'offrent plus aucune garantie de compatibilité. Pour les secteurs régulés, le chiffrement natif du state par OpenTofu et sa gouvernance ouverte sont des arguments forts face aux exigences de protection des données. La recommandation qui ressort des deux articles : partir sur OpenTofu pour les projets neufs et rester sur Terraform si l'on est déjà investi dans HCP Terraform et ses fonctionnalités de gouvernance. OTel est à la peine ? matduggan.com/otel-isnt-going-well-and-i-made-a-spreadsheet-about-it Le développement d'OTel est un peu au point mort Périmètre démesuré : Volonté de supporter un nombre gigantesque de langages, bibliothèques et frameworks. Pénurie critique de mainteneurs : Les données montrent une hyper-concentration du travail ; de nombreux SDK (comme PHP ou Ruby) dépendent d'une ou deux personnes seulement. Stabilité paralysante : La règle interdisant toute modification d'une fonctionnalité déclarée « stable » crée une peur de valider les nouveautés, entraînant des mois de débats. Solutions proposées par l'auteur : Créer un niveau « Bêta » temporaire (ex: 12 mois) entre les statuts « Expérimental » et « Stable ». Faire preuve de transparence sur les différences de qualité/maintenance entre les langages (ne pas mettre Go et Ruby sur le même plan). Communiquer activement sur le besoin urgent de nouveaux mainteneurs. Assouplir la politique de stabilité en acceptant des breaking changes bien documentés. Honeycomb transforme son infrastructure Kafka honeycomb.io/blog/transforming-how-we-run-kafka-honeycomb Honeycomb est une plateforme d'observabilité dont Kafka est le coeur du pipeline d'ingestion, traitant des millions d'événements par seconde. L'entreprise a migré de Confluent Platform auto-hébergée vers Apache Kafka 4.1.1 en mode KRaft. Le nouveau cluster tourne sur AWS EKS avec Strimzi comme couche d'orchestration Kubernetes. Motivation principale : la récupération après remplacement de broker était passée de 8-12h à 48-72h avec l'ancienne stack. Confluent imposait aussi sa solution propriétaire de Tiered Storage, impossible à corriger en interne. Un incident de décembre 2025 ayant vidé un cluster a révélé une fenêtre d'opportunité pour migrer. La migration s'est faite progressivement sur six clusters, de dogfood jusqu'à la production. Les producteurs sont basculés avant les consommateurs, avec une courte fenêtre de downtime assumée entre les deux. Le stockage utilise des NVMe en instance store plutôt que de l'EBS pour minimiser la latence. interessant de voir une société reprendre en main sa compétence et de voir les contraintes de certaines fonctionalités propriétaires Cloud AWS us-west-2 : panne réseau régionale et effet domino chez les fournisseurs SaaS blog.incidenthub.cloud/aws-us-west-2-outage-jul-24-2026 AWS us-west-2 (Oregon) est une région cloud majeure hébergeant de nombreux services et fournisseurs SaaS. Le 24 juillet 2026, une panne matérielle réseau a coupé la connectivité entre la région et le Seattle Metro pendant environ 20 minutes. Particularité notable, seule la couche de connectivité externe a été touchée, le trafic interne à la région a continué de fonctionner normalement. Après la réparation matérielle, une phase distincte de reconvergence des routes a de nouveau causé une connectivité intermittente pendant plusieurs dizaines de minutes. Les clients Direct Connect via EqSe2 ont subi une coupure bien plus longue que le reste, 1h17 au total. Neuf incidents chez sept fournisseurs ont cité explicitement AWS comme cause, dont SendGrid, SparkPost et NinjaOne. Fait marquant, les temps de rétablissement des fournisseurs tiers ont largement dépassé la durée de la panne AWS elle même. NinjaOne a mis 9h29 à se rétablir totalement, avec 150000 appareils tentant de se reconnecter simultanément freinés par des mécanismes de backoff et jitter. SparkPost a mis 6h45 à absorber l'arriéré de courriels accumulé pendant la coupure, avec encore 80 à 90 minutes de retard des heures plus tard. L'article recommande d'identifier ses dépendances en us-west-2 et de prévoir capacité et bascule, l'effet différé pouvant durer bien plus longtemps que l'incident initial. Rapport Cloudflare Radar sur les perturbations Internet au Q2 2026 blog.cloudflare.com/fr-fr/q2-2026-internet-disruption-summary Cloudflare Radar est la plateforme qui analyse en temps réel le trafic mondial pour détecter pannes, coupures et censures Internet. l'instabilité est le nouveau normal Le super-typhon Sinlaku a fait chuter le trafic de près de 80% à Guam les 13 et 14 avril. Deux séismes de magnitude 7,5 ont fortement dégradé la connectivité au Venezuela le 24 juin. Une coupure électrique a provoqué cinq heures de perturbation en Tanzanie le 27 juin. En Iran, la connectivité s'est stabilisée à 59% du niveau normal après 88 jours de coupure. Le Soudan a imposé dix coupures programmées pendant les examens nationaux mi-avril. L'Irak a coupé Internet à trois reprises pour lutter contre la fraude aux examens. Des frappes de drones ont endommagé la région AWS me-central-1 aux Émirats arabes unis. Un renouvellement de clés DNSSEC a rendu les sites .de inaccessibles en Allemagne le 5 mai. Une rupture de câble sous-marin a fait chuter le trafic de 60% à Sainte-Lucie fin juin. l'instabilité est le nouveau normal Un ingénieur Google débranche une zone entière de Google Cloud https://www.theregister.com/off-prem/2026/09/04/google-engineer-unplugged-every-fiber-they-could-see-and-surprise-took-down-a-chunk-of-the-g-cloud/5294418 Google Cloud est la plateforme d'infrastructure cloud de Google, organisée en régions et zones de disponibilité comme us-central1. Le 1er septembre 2026, un ingénieur a débranché par erreur des câbles fibre optique lors d'une opération de maintenance matérielle routinière dans la zone us-central1-b. En 13 minutes, il a déconnecté 100 % des chemins de fibre optique de tous les équipements de cette portion de la zone. Les machines virtuelles hébergées dans la zone touchée sont devenues injoignables, avec une perte de paquets élevée. Le taux de chute du trafic réseau pour les ressources concernées a atteint 100 %. L'incident a duré 4 heures et 11 minutes, de 7h41 à 11h52 heure du Pacifique. Google a détecté l'anomalie, reroute le trafic, identifié les liaisons optiques débranchées puis rebranché physiquement les fibres avant le retour à la normale. Les post-mortems d'incident cloud sont un classique du podcast, mais l'erreur humaine sur du câblage physique chez un hyperscaler mérite une minute d'antenne. ChatGPT, Claude et Grok en panne presque simultanément le 3 septembre theregister.com/ai-and-ml/…/5294322 ChatGPT, Claude et Grok sont les assistants IA et coding agents désormais utilisés au quotidien par de nombreux développeurs. xAI loue ses datacenters à un certains nombre de fournisseurs de cloud et d'IA. ils ont fait tombé ses concurrents :slightly_smiling_face: Arreter là Le 3 septembre 2026, les trois services sont tombés en panne quasiment en même temps. ChatGPT a connu une panne de 7h43 à 8h17 PT, soit environ 34 minutes, due à une erreur de routage rendant ChatGPT et Codex indisponibles. Claude a subi une panne partielle de 3 heures et 6 minutes touchant Claude.ai, Claude Code, Claude Cowork et l'API Claude, résolue à 16h16 UTC. xAI a commencé à enquêter sur les problèmes de Grok dès 6h30 PT, puis SpaceX a confirmé une panne de son centre de calcul de Memphis. La coïncidence des trois pannes a fait suspecter un fournisseur commun à l'origine du problème. Pour les développeurs devenus dépendants de ces coding agents, l'épisode illustre le risque d'un point de défaillance unique xAI loue ses datacenters à un certains nombre de fournisseurs de cloud et d'IA. Web Nouveautés CSS 2026 : mixins, masonry natif et animations pilotées par le scroll modern-css.com/whats-new-in-css-2026 animation-timeline: scroll() et view() atteignent le baseline cross-browser, Firefox et Safari ayant livré le support complet. Plus besoin de préfixes ni de librairie JS. @starting-style devient cross-browser : animations d'entrée depuis display: none sans hack de timing JS. Firefox 147 amène l'anchor positioning au baseline, plus les view transition types et la Navigation API. Les menus contextuels via popover CSS arrivent aussi. Data et Intelligence Artificielle Guillaume a porté le SDK Python d'Antigravity en Java… en utilisant Antigravity lui même comme assistant ! glaforge.dev/posts/…/the-unofficial-antigravity-sdk-for-java SDK Java non officiel pour Antigravity, rétro-ingénierie du SDK Python pour exploiter un binaire Go sous-jacent. Cas d'usage : pipelines CI/CD, applications d'entreprise (Spring Boot), outils internes, surveillance en arrière-plan, interfaces personnalisées. Gestion des ressources : implémente AutoCloseable(try-with-resources) pour lancer et fermer proprement le processus Go. Exécution de code Java personnalisé : exposition de méthodes Java comme outils IA via les annotations @Toolet@Param. Streaming et programmation réactive : prise en charge des CompletableFuture, des callbacks (chatStream), et deFlow.Publisher. Fonctionnalités avancées : persistance de session, protocole MCP, politiques de sécurité, entrées multimodales, et sorties structurées mappées sur des records Java. Le SDK Java pour le protocol Agent2Agent sort sa version 1.2.0 medium.com/google-cloud/a2a-java-sdk-1-2-0-final-released… Prise en charge de la spécification A2A 1.0. Sécurité renforcée : Vérification stricte des autorisations de lecture sur les tâches référencées et implémentation d'une logique de blocage par défaut (fail-closed). Intégration facilitée : Support du câblage programmatique des autorisations pour les environnements non-CDI (comme Spring). Contrôle des flux : Ajout du Task Stream Lifecycle Hook pour surveiller et gérer le cycle de vie des abonnements aux flux d'événements. Stabilité des données : Immutabilité stricte imposée sur les enregistrements de spécifications pour empêcher toute modification accidentelle. Documentation : Lancement d'une documentation web multi-versions et d'un Javadoc agrégé pour tous les modules. Corrections de bugs : Résolution de problèmes liés à la synchronisation des tâches, aux réponses de streaming et aux conditions de concurrence HTTP. Breaking changes : Nécessite une migration suite à la modification de certaines méthodes d'autorisation, la réorganisation de packages et le renommage de l'état TaskState.UNRECOGNIZED en TaskState.TASK_STATE_UNSPECIFIED. Article sur le blog de JetBrains : blog.jetbrains.com/idea/2026/08/intellij-idea-goes-lsp pas trop de temps Langchain4j continue sa progression github.com/langchain4j/langchain4j/…/1.18.0 github.com/langchain4j/langchain4j/…/1.19.0 github.com/langchain4j/langchain4j/…/1.17.0 Introduction du pattern Debate pour lancer des sous agents dans rechercehnt et entrent dans un debat critique sur une decision avant qu'un juge decide (1.17) compensation d'action d'un outil avec @ReverseTool (1.17) Ajout du pattern Belief-Desire-Intention (1.18): belief est l'état du monde cru, desire est la liste des objectifs, intention est le plan pour avancer un ou plusieurs objectif Support d'un systeme crash resilient dans l'approche Human in the loop avec des checkpoints nestés (1.18) Support Open AI TextToSpeech (1.18) Support de Mistral batch chat (1.18) support MCP client 2026-07-28 (1.19) Anthromic batch chat et Google thinking mode (1.19) Embabel atteint 1.0 GA github.com/embabel/embabel-agent/…/v1.0.0 embabel d'eloigne de Spring AI ne s'appuie que sur Spring (notamment tool calling) nettoyage des explorations autour du pattern GOAL avant a 1.0 ajout observabilite (dont le cout) les approches de retry solidifiées et d'autres choses toujours basées sur la description de goals typés et des dependences entre eux via les types MCP 2026-07-28 : le protocole devient stateless blog.modelcontextprotocol.io/posts/2026-07-28 MCP (Model Context Protocol) est le protocole standard permettant aux LLM et agents IA de communiquer avec des outils, ressources et serveurs externes. Cette version marque le plus gros changement depuis le lancement du MCP distant il y a 18 mois. Le protocole passe d'un modèle bidirectionnel avec état à un modèle stateless en requête/réponse. Suppression de la poignée de main initialize/initialized et du header Mcp-Session-Id, chaque requête devient autoportante. N'importe quelle requête peut désormais être routée vers n'importe quelle instance de serveur derrière un load balancer classique, sans stockage partagé. Introduction des Multi Round-Trip Requests (MRTR) pour remplacer les requêtes initiées par le serveur, via un resultType input_required et des inputResponses. Nouveaux headers Mcp-Method et Mcp-Name pour permettre aux gateways de router et autoriser sans parser le JSON. Les résultats de tools, prompts et resources deviennent cacheables grâce aux paramètres ttlMs et cacheScope. Renforcement sécurité avec la validation d'issuer RFC 9207 pour éviter les attaques de confusion entre serveurs d'autorisation, et transition de DCR vers CIMD. Roots, Sampling et Logging sont dépréciés avec douze mois de support garanti, tout comme le transport legacy HTTP+SSE. Un site qui référence les skills pour la JVM (framework, langage, build…) jvmskills.com Frameworks : Spring, Quarkus, Jakarta EE, Reactor, Camel Java : bonne pratiques, conventions, guides de mise à jour à niveau LTS, API spécifiques (streams, optionals, logs…) Bases de données : ORM, validation, modélisation PostgreSQL, vectorielle avec pgvector Tests et qualité : TDD, mutation testing, debogage avec JDB Workflows dev et archi : commits git, domain modeling Outils et diagnostics JVM : JFR, Jstall, JSpecify Une skill n'est pas une librairie https://devx.writizzy.blog/p/un-skill-nest-pas-une-lib Les skills sont des éléments de configuration en prose pour agents IA comme Claude, distribués via des marketplaces à la manière de librairies logicielles. Frédéric Camblor critique cette analogie car partager un skill n'est pas la même chose que le mutualiser durablement. Écrire un skill prend 30 minutes mais l'adopter ailleurs coûte cher en appropriation et en maintenance. Forker un skill s'avère souvent plus efficace que de chercher à converger vers une version commune. Contrairement au code, les régressions d'un skill ne sont pas détectables automatiquement. Un skill peut se dégrader silencieusement sur plusieurs cas d'usage en corrigeant un autre. Les skills vieillissent vite car les modèles progressent et intègrent naturellement certaines bonnes pratiques. Le skill-creator d'Anthropic permet d'évaluer un skill via des jeux de cas et des mesures de variance. L'auteur distingue quatre sphères de partage : personnelle, équipe, outil et marketplace. Il propose un cycle partage puis appropriation puis duplication puis divergence plutôt qu'une installation collective figée. Les modèles Anthropic introduisent un filigrane (watermark) dans les textes qu'ils génèrent https://www.anthropic.com/news/claude-text-watermark Claude est l'assistant IA d'Anthropic, et le watermarking est une technique permettant de marquer discrètement un contenu généré par IA pour en tracer l'origine. Anthropic annonce que les futurs modèles Claude intégreront un filigrane numérique invisible dans le texte généré. Le principe exploite les choix de mots équivalents que le modèle fait naturellement, en les orientant via une clé cryptographique plutôt qu'un tirage aléatoire. Le texte produit reste indiscernable à l'œil nu, sans caractères cachés, sans ralentissement ni coût supplémentaire. Seule la personne possédant la clé correspondante peut détecter la présence du filigrane. Le filigrane est plus fiable sur les textes longs et créatifs, moins sur du texte factuel, du code ou après une édition manuelle poussée. Il ne prouve pas qu'un texte est écrit par IA, ni n'identifie l'auteur ou la conversation d'origine, il donne seulement une probabilité d'implication de Claude. Une API de détection est proposée en accès restreint aux régulateurs, forces de l'ordre, médias et vérificateurs de faits. Pour les fichiers non textuels comme les images ou les PDF, Anthropic s'appuie sur le standard C2PA. Cette initiative s'inscrit dans le Code de Pratique de l'UE sur la transparence des contenus IA, signé par Anthropic et environ 190 autres acteurs, en lien avec la loi européenne sur l'IA. GPT-6 Astra, le nouveau modèle d'OpenAI face à Claude Fable 5.1 https://openai.com/index/gpt-6-astra/ GPT-6 Astra est le nouveau modèle phare d'OpenAI, annoncé le 3 septembre 2026 comme le plus intelligent et le plus aligné de l'entreprise. Le modèle arrive deux jours après Claude Fable 5.1, à un tarif affiché comparable, dans une course accélérée aux modèles de code et de raisonnement. Astra revendique 98 % sur FrontierMath Tier 4, 99,9 % sur ARC-AGI-3 et 100 % sur ExploitBench. Fenêtre de contexte d'environ 1,05 million de tokens. Tarification API, 10 dollars par million de tokens en entrée, 50 dollars en sortie, et 1 dollar par million pour les tokens en cache. Au-delà de 272 000 tokens en entrée, toute la requête est facturée au double sur l'entrée et une fois et demie sur la sortie. Astra est le premier modèle d'OpenAI à franchir le seuil interne critique en cybersécurité. La version publique refuse les tâches offensives avancées comme générer des preuves de concept d'exploits. Le déploiement est progressif, les entreprises du programme de cybersécurité Daybreak d'OpenAI y accèdent en premier, avant ChatGPT Plus, Pro, Business, Enterprise, l'API et AWS. Même logique de diffusion contrôlée que chez Anthropic avec Mythos 5.1 : deux jours d'écart, deux modèles de tête, et la même question de savoir qui accède en premier aux capacités les plus sensibles. Outillage JetBrains s'est lancé dans les LSP (Language Server Protocol) avec une extension IntelliJ pour VS Code et assimilés marketplace.visualstudio.com/items?itemName=JetBrains.intellij-s… Nouveau produit : Lancement de l'extension Java & Kotlin by IntelliJ IDEA pour les éditeurs basés sur VS Code (incluant Cursor). Technologie : Utilisation du standard LSP (Language Server Protocol). Objectif : S'adapter au développement piloté par les agents IA, qui nécessite des fonctionnalités IDE légères et standardisées. Fonctionnalités clés : Support des projets Java, Kotlin et mixtes. Débogage (DAP). Complétion intelligente, navigation et analyse de code. Refactoring. Prise en charge de Maven, Gradle et Bazel. Disponibilité : Téléchargeable via le Visual Studio Marketplace et l'Open VSX registry. Licence / Prix : Gratuit durant la phase de preview (évaluation renouvelable de 30 jours). Nécessitera un abonnement IntelliJ IDEA Ultimate après la preview. (Note : Le LSP purement Kotlin reste gratuit et open-source). Avenir : Développement en cours pour optimiser les flux de travail avec les agents IA en ligne de commande (ex: Claude Code, Codex) afin de réduire la consommation de tokens. Après son acquisition par SpaceX, Cursor perd l'accès aux modèles OpenAI https://openai.com/index/our-decision-on-cursor-following-its-acquisition-by-spacex/ decision difficile mais Elon ment comme un arracheur de dent (et c'est un competiteur donc bon ça nous arrange) David Pilato a créé un thème spéciale pour les gens pour les devrel, ou qui font des talks à droite à gauche, pour le moteur Hugo https://david.pilato.fr/posts/2026-09-07-hugo-theme-devrel/ hugo-theme-devrel, thème Hugo (MIT) pour Developer Advocates et conférenciers, fonctionnant comme un module superposé au thème Dream. Gestion des conférences (cartes Leaflet), présentations (PDF, YouTube, co-auteurs), vues dédiées (archives, sujets récurrents, vidéos) et recherche Pagefind. Chaque intervention est un page bundle YAML structuré comme une base de données relationnelle compilée par Hugo. Architecture Airbnb refond son authentification en architecture server-driven, 60 % de code en moins https://www.infoq.com/news/2026/09/airbnb-server-driven-login/ Airbnb a restructuré son authentification autour d'un modèle en deux phases, identification du compte (email, téléphone ou connexion sociale) puis choix du challenge d'authentification décidé côté serveur selon le contexte utilisateur. Ce qui est interessant c'est que choisir la method d'authentification est côté serveur et adaptative, comme si c'était une révolution Arreter là Un moteur de politique serveur sélectionne la méthode d'authentification optimale avec des solutions de repli, permettant des adaptations régionales comme l'OTP WhatsApp au Brésil ou des fournisseurs d'identité locaux en Corée du Sud sans nouvelle version client. Un Challenge Picker propose des méthodes alternatives classées par probabilité de succès en cas d'échec. Résultat chiffré, 60 % de code d'authentification en moins et 100 Ko de moins sur le bundle client web. Méthodologies Les nouvelles règles d'ingénierie du contexte pour les modèles Claude 5 x.com/trq212/status/2080710971228918066 Partage par Thariq des apprentissages sur l'ingénierie du contexte et le prompt engineering pour les nouveaux modèles Claude 5 (comme Claude Opus 5 et Claude Fable 5) utilisés dans Claude Code. Évolution majeure vers le dés-empirement (unhobbling) : plus de 80 % du prompt système de Claude Code a pu être supprimé sans perte sur les évaluations de code, les modèles récents faisant preuve d'un bien meilleur jugement contextuel. Passage des règles strictes au jugement : au lieu d'interdire les commentaires ou d'imposer des contraintes lourdes, les modèles s'adaptent désormais au code environnant et font appel à leur propre discernement. Remplacement des exemples par la conception d'interfaces : fournir des exemples figés restreint l'exploration du modèle, d'où l'importance de concevoir des outils et des fichiers plus expressifs. Adoption de la divulgation progressive (progressive disclosure) : chargement dynamique du contexte (via des compétences ou des outils à chargement différé comme ToolSearch) pour éviter de saturer la fenêtre de contexte avec des instructions fixes. Utilisation d'une mémoire automatique et de références riches (artefacts HTML, suites de tests, fonctions de référence) plutôt que de fichiers CLAUDE.md pléthoriques ou de consignes répétitives. Recommandation pour les fichiers CLAUDE.md et les Skills : les garder légers, se concentrer sur les pièges spécifiques (gotchas) du dépôt, et structurer les guides sous forme d'arborescences modulaires pour ne charger que le nécessaire. Niveau d'adoption des agents IA de codage selon une étude de JetBrains blog.jetbrains.com/research/2026/08/ai-coding-agent-adoption-2026 Adoption massive : 90 % des développeurs professionnels utilisent des agents d'IA de codage au moins une fois par semaine, et 68 % quotidiennement. Claude Code domine : Il devient le nouveau leader du marché avec 39 % d'adoption mondiale (47 % aux États-Unis), détrônant largement ses concurrents. Déclin de GitHub Copilot : L'ancien leader perd de sa superbe, passant de 29 % à 21 % d'adoption, bien qu'il conserve une très forte notoriété (79 %). Percée de Codex : Sa croissance est fulgurante, son taux d'adoption ayant été multiplié par 5 en quelques mois (de 3 % à 16 %). Recul de Cursor : L'outil connaît une légère baisse, passant de 18 % à 12 % d'adoption, principalement due à une forte chute sur le marché chinois. Écosystème diversifié : JetBrains AI atteint 9 % d'adoption. Des alternatives comme OpenCode (7 %) et Google Antigravity (6 %, mais très populaire en Inde à 15 %) continuent de s'implanter. L'IA a cassé les hypothèses de la CI… ou pas https://stack72.dev/ai-broke-the-assumptions-behind-ci/ Paul Stack (ex-Pulumi) explique que l'intégration continue a toujours mêlé deux rôles distincts, exécuter la vérification du code et coordonner les fusions. avec les agents, la pression sur la CI augmente, car ils ne testent pas end to end tout le temps ils ont maintenant des workflows qui verifient tests, lint, revue d'agent etc dans un env local isolé donc c'est pre PR push Il propose de séparer vérification (exécution), attestations structurées (hash de commit, checksums SHA256) et CI, réduite à la validation et à la coordination des fusions. Sa thèse, les agents IA peuvent désormais vérifier tout le code localement avant même l'ouverture d'une pull request, ce qui bouleverse cet équilibre. l'attestation vient ensuite et la CI est une étape de vérification (attestation de commit et de tests et si branche a bougé, repart à l'execution) Martin Fowler, qui a repéré l'article dans ses Fragments du 1er septembre 2026, réplique que la vraie CI a toujours exigé une vérification locale avant de pousser le code. ca demande une chaine d'attestation forte quid de garder les metriques historique de CI Sécurité France Passoire: une analyse sur les différents vols de données des services publics de ces dernières mois https://www.cybernetica.fr/piratage-des-impots-comment-en-est-on-arrive-la/ Analyse du piratage massif de la DGFiP et d'autres administrations françaises en 2026, révélateur de failles systémiques de cybersécurité de l'État. Intrusion détectée fin juin à la DGFiP, mais l'exfiltration de 678 000 entrées fiscales n'a été découverte qu'en août lors de leur mise en vente. Données volées : noms, revenu fiscal de référence, taux de prélèvement, adresse, téléphone. Une seconde attaque du même pirate a visé le cadastre fin juillet, exposant plus de 2 millions de personnes. L'Éducation nationale a aussi été piratée fin juillet, données de tous les agents depuis 2001 exposées. Cause principale : modèle de sécurité fondé sur le périmètre physique plutôt que sur le zero trust, sans contrôle après authentification. Le télétravail post Covid a étendu les accès distants sans reconstruire les modèles de confiance. Aucun système de détection d'exfiltration n'existait, la fuite n'a été révélée que par le pirate lui-même. La transposition de la directive NIS2 est bloquée en France depuis septembre 2025, la CJUE a condamné le pays à des astreintes. L'article souligne un désengagement croissant des Etats-Unis en matière de cybersécurité internationale et une dépendance technologique accrue de la France. Loi, société et organisation Ce que l'IA change vraiment au métier de manager shapeandship.ai/p/ce-que-lia-change-vraiment-au-metier-de-manager Retour d'expérience et analyse par Mathilde Rigabert sur l'impact réel de l'IA générative dans le quotidien d'un Engineering Manager. L'IA excelle pour automatiser la reconstitution factuelle de l'activité (lecture de PRs, commits, reviews) nécessaire aux 1:1 et entretiens annuels, mais elle offre une vision uniquement quantitative et nécessite d'être croisée avec des notes de terrain. L'IA rend le maintien de la qualité et des standards plus difficile : selon une étude Faros AI sur 22 000 développeurs, les PRs mergées sans aucune revue ont augmenté de 31 %, fragilisant la compréhension commune apportée par le pairing et les revues de code. Le temps gagné par l'IA ne permet pas d'augmenter massivement le span of control (seulement 2 ou 3 personnes de plus), car l'IA compresse la collecte d'informations mais pas les conversations humaines complexes ou l'accompagnement du changement. Les compétences d'orchestration et de gestion de sujets multiples acquises par les managers facilitent leur transition vers le pilotage de plusieurs agents IA en contribution individuelle. Le piège actuel réside dans l'accumulation des casquettes (manager, tech lead, product owner, contributeur, pompier), conduisant à l'épuisement et au délaissement du travail de fond sur l'organisation et l'humain. Le temps libéré par l'IA doit être réinvesti dans le travail invisible qui fait tenir le système (suivi des actions de rétro, analyse de métriques, coaching), que personne ne réclame à court terme mais dont l'absence fragilise les équipes à long terme. Je regrette d'avoir migré vers Codeberg xn–gckvb8fzb.com/i-regret-migrating-to-codeberg L'auteur explique pourquoi il regrette d'avoir quitté GitHub pour Codeberg, à la suite des récentes modifications des conditions d'utilisation (ToS) de la plateforme. Codeberg a interdit les projets principalement générés par des LLM ainsi que les projets liés aux cryptomonnaies via des propositions de l'Assembly 2026, au motif qu'ils nuisent à sa réputation. blog.codeberg.org/protecting-our-floss-commons-from… Critique de l'argument de Codeberg sur l'absence de communauté des vibe coders, en rappelant que la majorité des logiciels libres (FOSS) sont créés par des développeurs solos sans communauté au sens romancé du terme. Ironie soulignée concernant la posture de Codeberg et Forgejo, qui a hérité de la communauté de Gitea après un hard fork avant de faire la leçon aux développeurs individuels. Alerte sur le risque de censure idéologique : interdire des catégories entières plutôt que de traiter les abus réels ou la consommation d'infrastructure crée un précédent dangereux pour une forge qui se veut libre. Proposition de solutions alternatives pour gérer l'impact des LLM et de la crypto : déclaration obligatoire via des cases à cocher, hébergement sur des tiers d'infrastructure spécifiques payants ou sous quotas, et disclaimers automatiques. Décision de l'auteur de quitter Codeberg pour mettre en place son propre serveur Git personnel afin d'éviter la dépendance à une plateforme qui modifie ses règles de manière unilatérale. Cloud souverain : Airbus choisit Scaleway pour l'hébergement de ses applications critiques https://www.usine-digitale.fr/aeronautique-spatial/airbus/cloud-souverain-airbus-choisit-scaleway-pour-lhebergement-de-ses-applications-critiques.OHBVBMSZIJELNOKN6G5F6B7NSI.html Scaleway est le cloud provider français filiale du groupe Iliad, positionné comme alternative souveraine aux hyperscalers américains. Airbus a lancé un appel d'offres de six mois consultant une cinquantaine d'acteurs dont OVHcloud, Thales, Google S3NS et Microsoft Bleu. Scaleway a été retenu pour héberger les applications critiques liées à la conception d'aéronefs, l'ingénierie, la production industrielle et les opérations. Le contrat prévoit la migration d'environ 70 applications d'ici 2028, puis jusqu'à 900 applications sur 5 à 6 ans. Le montant du contrat n'a pas été communiqué. Scaleway revendique zéro actionnaire, zéro employé et zéro filiale hors Union européenne pour garantir une protection contre les lois extraterritoriales. Damien Lucas, PDG de Scaleway, évoque une immunité complète face aux évolutions politiques et législatives externes. La plateforme doit aussi accélérer les usages d'intelligence artificielle d'Airbus, avec les modèles de Mistral AI déjà déployés chez Scaleway. Catherine Jestin, responsable numérique d'Airbus, souligne que cette intégration accélère la démarche IA du groupe. Ce choix ne remet pas en cause la stratégie multicloud d'Airbus, Scaleway venant compléter les fournisseurs existants pour les charges nécessitant le plus haut niveau de gouvernance et de résilience. Debian adopte une résolution sur l'usage responsable de l'IA générative lwn.net/Articles/1091231 La discussion sur l'usage des LLM dans Debian s'est tenue du 23 juillet au 13 août 2026, suivie d'un vote du 15 au 28 août 2026. 1045 développeurs Debian étaient éligibles à voter, avec un quorum de 48,49 votes largement dépassé par les huit options en lice. Les options allaient d'une interdiction stricte des contributions générées par LLM inscrite dans le contrat social à une acceptation encadrée des contributions IA. L'option gagnante au classement Condorcet est Responsible Use of Generative AI, devant Allow AI-Assisted Contributions with conditions et A cautious approach to generative AI. Le texte adopté n'interdit ni n'encourage l'usage d'outils d'IA générative dans le développement de Debian. Il exige que toute contribution, quels que soient les outils utilisés pour la produire, respecte les mêmes standards de qualité, correction, maintenabilité et conformité légale. Les contributeurs doivent comprendre, relire, tester et si besoin modifier la production assistée par IA avant de l'intégrer à Debian. Les informations sensibles du projet ne doivent pas être transmises à des fournisseurs d'IA non fiables, et la divulgation de l'usage de l'IA est encouragée sans être obligatoire. Le détail du vote et le texte complet de la résolution sont disponibles sur la page officielle [debian.org/vote/2026/vote_002](https://www.debian.org/vote/2026/vote_002). Contraste direct avec l'OpenJDK, qui a publié une politique interdisant le code généré par LLM (épisode 340), et avec l'auteur de jqwik qui a piégé sa librairie contre les agents. Conférences Nous contacter Pour réagir à cet épisode, venez discuter sur le groupe Google https://groups.google.com/group/lescastcodeurs Contactez-nous via X/twitter https://twitter.com/lescastcodeurs ou Bluesky https://bsky.app/profile/lescastcodeurs.com Faire un crowdcast ou une crowdquestion Soutenez Les Cast Codeurs sur Patreon https://www.patreon.com/LesCastCodeurs Tous les épisodes et toutes les infos sur https://lescastcodeurs.com/
Welcome to Episode 436 of the Microsoft Cloud IT Pro Podcast. In this episode, Ben and Scott dive into Microsoft Entra Tenant Governance. Almost nobody runs one Microsoft Entra tenant. Mergers, divestitures, dev and test environments, regional partitioning, that one AI proof-of-concept somebody spun up on a credit card. Organizations accumulate tenants the way they accumulate SharePoint sites. The difference is that nobody ever built an inventory of the tenants, and central IT frequently cannot name them, let alone say whether they are configured safely. Microsoft Entra Tenant Governance, now generally available, is Microsoft’s first-party answer. It is genuinely four products under one blade: it finds the tenants you did not know about, gives you a least-privilege way to administer them without local accounts, monitors their configuration for drift against a JSON baseline, and gates the creation of new ones so the next tenant is governed from birth. Your support makes this show possible! Please consider becoming a premium member for access to live shows and more. Check out our membership options. Show Notes Microsoft Entra Tenant Governance is now generally available What is Microsoft Entra Tenant Governance? Deploy Microsoft Entra Tenant Governance end to end Related tenants in Tenant Governance Licensing for Microsoft Entra Tenant Governance Sponsors Nasuni is a leading unstructured data platform for enterprises where file data is mission-critical for both people and AI. Nasuni powers the operational file layer where work happens — helping organizations manage, protect, and activate data so teams can work smarter, reduce costs, and operate securely without limits. Intelligink — Would you like to become the irreplaceable Microsoft 365 resource for your organization? Let us know!
Talk Python To Me - Python conversations for passionate developers
How many files does your query read before it reads any data? On some data lakes, you go through JSON and metadata files first, just to learn which Parquet files matter. DuckLake asks one SQL question instead. The metadata lives in a real database. The data stays in plain Parquet. That's the entire format. Pedro Holanda joined DuckDB in 2018, when it was still a research prototype at CWI. He's the lead DuckLake developer. Guillermo Sanchez Dionis works on DuckLake and the new Quack protocol. With Quack as the catalog, DuckLake handles 200 transactions a second under heavy contention. No other open table format comes close. Episode sponsors Six Feet Up Talk Python Courses Links from the show Guests Pedro Holanda: pedroholanda.org Guillermo Sanchez: linkedin.com PhD on progressive indexes: ir.cwi.nl SQLite: www.sqlite.org Litestream: litestream.io boring hardware: talkpython.fm DuckDB: duckdb.org episode 491: talkpython.fm Iceberg: iceberg.apache.org manifesto: ducklake.select DuckLake: ducklake.select spec: ducklake.select this diagram: blobs.talkpython.fm Data inlining: ducklake.select ducklake-dataframe: github.com Polars course: training.talkpython.fm CSV parser: duckdb.org Zero-copy Arrow: duckdb.org ART index: duckdb.org async I/O: duckdb.org v1.0: ducklake.select Git-like branching: ducklake.select Watch this episode on YouTube: youtube.com Episode #562 deep-dive: talkpython.fm/562 Episode transcripts: talkpython.fm Theme Song: Developer Rap
Hoy te traigo un episodio que llevaba tiempo queriendo grabar. Y es que muchos estáis usando agentes de IA, pero los tenéis desnudos. Sin skills. Y un agente sin skills es como un Linux sin comandos: técnicamente funciona, tienes el kernel, tienes la shell, pero sin ls, sin grep, sin systemctl, no puedes hacer nada útil. El mejor modelo del mundo sin herramientas solamente es texto bonito.En este episodio te cuento qué son exactamente las skills, por qué transforman un modelo de lenguaje en un asistente que hace cosas, y cuáles son las tres skills imprescindibles que todo agente debería tener. Una skill no es ni más ni menos que un prompt. Un conjunto de instrucciones que le dice a tu agente cómo tiene que hacer algo. No es un programa ni un script, es una receta de comportamiento. Y no necesitas ser programador para crearlas.Te hablo de skills del sistema: leer archivos, ejecutar comandos, navegar por tu equipo. Skills de búsqueda: búsqueda web con SearXNG (que tengo montado en un Slimbook One y devuelve resultados en JSON que es una maravilla), búsqueda local con RipGrep que es increíblemente rápida, y búsqueda semántica con SQLite. Y skills de automatización: tareas programadas, webhooks que responden a eventos como un git push, o scripts orquestados que pueden hacer casi cualquier cosa.Pero no me quedo ahí. Te doy las tres reglas de oro para crear tus propias skills. Porque los mejores skills son los que escribes para ti mismo. Primera regla: no digas "revisa el sistema", sino "ejecuta systemctl status --failed". Cuanto más específico, mejor. Segunda: define los límites. Qué puede hacer y qué no. Lo que no puede hacer es casi tan importante como lo que puede hacer. Tercera: ponle ejemplos. Cómo tiene que quedar el resultado, qué formato tiene que usar. Con estas tres cosas todo rueda mucho mejor.También te cuento cómo organizar tus skills. Cada agente guarda los skills donde le da la gana: OpenCode en .config/opencode/skills, Hermes en .hermes/skills. Yo cada vez los guardo más en .agents/skills porque la mayoría de los agentes ya saben encontrarlos ahí. Y te hablo de los hubs de skills, donde puedes instalar skills creados por la comunidad con un solo comando.Si usas OpenCode, Hermes Agent u Open Web y sientes que tu agente responde preguntas pero no hace cosas, este episodio te va a cambiar el día a día. Vamos directos al turrón.Capítulos del episodio:0:00 - Introducción: tu agente está desnudo1:45 - ¿Qué es una skill? De un prompt a un asistente que hace cosas4:15 - Skills del sistema: leer archivos, ejecutar comandos, navegar7:30 - Skills de búsqueda: web con SearXNG, local con RipGrep, semántica con SQLite11:00 - Skills de automatización: tareas programadas, webhooks, scripts orquestados13:45 - Las 3 reglas de oro para crear tus propias skills18:30 - El ecosistema de skills: dónde guardarlas y cómo organizarlas21:00 - Ejemplo práctico: la skill del tiempo meteorológico23:15 - Conclusión: un agente sin skills es Linux sin comandosMás información y enlaces en las notas del episodio
Kevin White, Head of Marketing at Scrunch, argues that referral traffic from ChatGPT is the wrong instrument for measuring AI search — and that the real question is whether a model represents your brand accurately when no one clicks at all. He and Greg Kihlström get into what to build and what to measure instead."AI search" is several different problems wearing one label. A crawler pulling your content, a model synthesizing an answer, and an agent acting on someone's behalf are not the same event — and teams that treat them as one end up optimizing for the wrong thing.A website now serves two readers with different needs. White on what changes structurally when the page has to work for a human and for a machine that will compress it into three sentences.Visibility is probabilistic, so accountability has to change. The same prompt can surface your brand one day and skip it the next. White on what a marketing leader can honestly commit to a CMO or a board in that environment.About Kevin WhiteKevin White is a B2B tech marketing leader who has helped shape go-to-market at Segment, Retool, and Common Room, and now leads marketing at Scrunch AI, an AI-first customer-journey platform. He got his start in SaaS at Gigya — where he was handed Marketo and told to 'figure it out,' the spark for a learn-by-doing growth mindset that merges analytical rigor (attribution, lifecycle infrastructure, reporting) with creative offer- and channel-building. That growth foundation carried him to leading whole marketing orgs, though he's candid that the climb pulled him away from the hands-on craft he loves, and at Common Room he deliberately stepped back into a senior IC role. At Scrunch he focuses on one of the most overlooked shifts in modern marketing: websites are now consumed as much by AI bots as by humans, and brands that make their sites legible to LLMs — through structure, markdown and JSON, schema, and a parallel AI-optimized experience — win the emerging AI retrieval channel.Kevin White on LinkedIn---------- Resources ----------This episode is brought to you by Scrunch. Scrunch gets your site AI-ready so you show up in answers, get cited, and grow revenue. Learn more at Scrunch.comReach your customers with Reddit. Spend $500 in ad spend, get $500 back in ad credit! Learn more: https://advertalize.com/r/491818c79fb1873fEnjoyed the show? Tell us more at and give us a rating so others can find the show at: https://aglbrnd.co/r/faaed112fc9887f3Connect with Greg on LinkedIn: https://www.linkedin.com/in/gregkihlstromDon't miss a thing: get the latest episodes, sign up for our newsletter and more: https://aglbrnd.co/r/35ded3ccfb6716baCheck out The Agile Brand Guide website with articles, insights, and Martechipedia, the wiki for marketing technology: https://www.agilebrandguide.comThe Agile Brand is produced by Missing Link—a Latina-owned strategy-driven, creatively fueled production co-op. From ideation to creation, they craft human connections through intelligent, engaging and informative content. https://www.missinglink.company Hosted on Acast. See acast.com/privacy for more information.
In this episode, Ray Cochrane digs into Anthropic’s Model Hardware Standard. It is a shared driver that lets an AI agent run real lab equipment, from pipetting robots to the lasers inside a quantum computer. He also covers OpenAI’s builder’s guide to GPT-5.6, Google’s new Expert Intelligence book feature, Apple’s M5 Ultra Mac Studio, and a judge’s order forcing Google to stop hiding rival app stores. Finally, he weighs in on Apple’s proposed 15 percent link-out fee, Meta’s Australia numbers, the White House deputizing private hackers, and why rivers obey a 1957 math rule. – Want to start a podcast? Its easy to get started! Sign-up at Blubrry – Thinking of buying a Starlink? Use my link to support the show. Subscribe to the Newsletter. Email Ray if you want to get in touch! Like and Follow Geek News Central’s Facebook Page. Support my Show Sponsor: Best Godaddy Promo Codes Get 1Password Full Summary Cochrane opens with a quick personal update. He is hunting for tickets to Michigan for his dad’s anniversary, and he has been learning Blender and Godot on the side, mostly modeling and blocking out levels. Consequently, he asks listeners for advice on starting a big game project, and he plans to record his progress, maybe as a time lapse. Then it is straight into the featured story. Anthropic’s Model Hardware Standard: A Driver for the Physical World The featured story comes from Anthropic, which opened a research preview of the Model Hardware Standard, or MHS. Cochrane frames it as the other side of the question NVIDIA’s world models raised two weeks ago: when do AI agents start touching actual machines? A typical lab runs a microscope, a liquid handler, a robotic arm, and a plate reader, each from a different vendor with its own control software. One Janelia researcher in the post launches seven programs in three languages just to start an experiment. Anthropic says wiring a setup like that takes weeks or months of specialist work. MHS is a driver, the same kind of translation layer a printer uses, except every device gets described with a tiny set of commands like read and write. Devices announce themselves on the network. A plain-English reference file then records what each machine measures, what can be adjusted, and which safety limits get enforced no matter what the agent asks. Agents then reach the hardware through the Model Context Protocol, the command line, or plain code. Cochrane sees the same move the industry keeps making, from coding harnesses to RSS and JSON: agree on a standard and let everyone build against it. In fact, he calls MHS the hardware version of MCP. The partner results carry the segment. QuEra builds quantum computers from individual atoms held by lasers that must hold their frequency to about one part in a trillion. A four-person team spent months on a relock script that worked 58 percent of the time. However, four copies of Claude iterating overnight through MHS produced a decision-tree script that recovers the laser in about six seconds, and it passed 99.3 percent of 700 blind trials. Carnegie Mellon wrote MHS drivers for four instruments across three incompatible computers in about eight hours, then ran dose-response experiments three times faster and blocked all six deliberately induced faults. Genentech, meanwhile, showed the limits. Claude used the same pump speed for water, a foamy protein solution, and a human had to explain that the bubbles were a physics problem. That gap in physical intuition is what sticks with Cochrane. He doubts it will change soon, and he suspects the fix will arrive as sub-agents or sub-models that judge a request against an expected outcome. He also connects MHS to a video of racing robots that never learned to stop at the finish line. What happens, he wonders, once they can read a distance sensor through a shared standard? Still, he calls the announcement a fantastic read and points listeners to the full article. Sponsor: GoDaddy Economy hosting $6.99/month, WordPress hosting $12.99/month, domains $11.99. Website builder trial available. Use codes at geeknewscentral.com/godaddy to support the show. GPT-5.6 Does the Same Work for a Fraction of the Cost OpenAI’s builder’s guide to GPT-5.6 leads the headlines. Cochrane recaps the three tiers from episode 1870, Sol, Terra, and Luna, plus the separate dial for reasoning effort. On BrowseComp, a benchmark for digging up obscure facts on the web, the old GPT-5.5 flagship scored about 84 percent on a run that cost 33 dollars three months ago. Luna now matches that score for a dollar thirty-three, and OpenAI has since cut Luna’s price another 80 percent. Browser Use reports Luna finishing 78 percent of its hardest browser tasks for about 14 dollars, against 80 percent for roughly 235 dollars from the best available model. The guide’s other big addition is a multi-agent beta flag. It lets the model handling a request spawn parallel helper agents that report back to a root agent inside a single API call. However, Cochrane is unimpressed by the timing. He has been running that pattern in Claude Code for months, so he sees OpenAI copying a workflow other companies already ship rather than inventing its own. Along the way, he plugs Claude Code’s remote-control sessions, which let him send prompts from his phone to a terminal session at home. Google Lets Gemini Read the Books You Actually Bought Google launched Expert Intelligence, a name Cochrane calls quite the reach. The feature lets you drop a book you bought on Google Play Books into Gemini Notebook, formerly NotebookLM, and ask questions answered only from that book, with citations. Cochrane sees real power here for students, since he once used NotebookLM to organize scattered course PDFs. Additionally, publishers get a cut, which he calls a far better deal than the wholesale scraping of books that trained earlier models. Nevertheless, he asks who loses out, because a paid publisher does not automatically mean a paid author. He floats the same idea for artists, even a penny per use, then admits that may be too idealistic. Apple’s M5 Ultra Mac Studio Is Built to Run Big Models at Home Back in episode 1861, when Apple killed the Mac Pro, an M5 Ultra Mac Studio was expected later this year. Now it is here. The M5 Ultra brings up to a 36-core CPU, an 80-core GPU, and 512GB of unified memory moving 1.2 terabytes per second. Apple claims up to 4.3 times the AI performance of the M3 Ultra. Thunderbolt 5 can also cluster four machines into one memory pool for up to three times faster inference. The M5 Max model starts at $2,499 and the Ultra at $5,499, with shipping on September 22 and the 512GB configuration arriving in late October. Cochrane finds the clustering pitch ridiculous at that price, but he invites anyone who spends the money to report back. Apple Opens a Manufacturing School in Houston Apple also opened a 20,000-square-foot Advanced Manufacturing Center in Houston. It offers free classes for small and midsize manufacturers, from circuit board design to hands-on time on a scaled-down production line, with college students joining later. Cochrane calls it a solid step in the bring-manufacturing-home movement. The bigger story is the campus itself, which builds Apple’s AI servers and will add the first US-assembled Mac mini line later this year. That ties back to the Mac mini shortage that followed the OpenClaw rush, when Tim Cook warned of months-long waits. Cult of Mac was still reporting four-month waits in late July. However, Cook blamed chip supply rather than assembly, so Cochrane is not counting on relief just yet. Amazon EC2 Turns Twenty Amazon EC2 turned twenty this week, which Cochrane admits makes him feel old. The 2006 beta offered one server size in one region for ten cents an hour. Each came with a 1.7 gigahertz Xeon and under two gigabytes of memory, and accounts were capped at twenty servers. Today AWS offers more than 1,200 instance types across 39 regions. Consequently, Cochrane credits the company with turning that tiny product into the backbone of cloud and AI computing. Intel Gamer Days: Two Free Games, With Fine Print Intel Gamer Days runs through September 13. Buy a qualifying Core Ultra Series 2 or 14th Gen desktop chip, a Core Ultra Series 3 laptop, or an Arc graphics card. In return you get Star Wars: Galactic Racer plus the Tomb Raider: Legacy of Atlantis remake. GamesRadar values the pair at about 120 dollars. However, neither game is out yet, and codes must be redeemed by October 31 even though the Tomb Raider remake ships in February. Cochrane calls that awful, but he still tells qualifying buyers to claim the deal early. Note that 13th Gen chips do not qualify. Judge Orders Google to Stop Hiding Rival App Stores A jury found Google’s Android app monopoly illegal in late 2023, and Judge James Donato ordered rival stores into the Play Store in 2024. On August 13, Epic’s lawyer demonstrated that searching Play for “store for apps” returned Walmart instead of any app store. Donato called that “not acceptable” and ordered three fixes within a week. Searches must surface third-party stores, listings need a plain install button, and the “are you looking for” interstitial has to go. Cochrane welcomes the monopoly being chipped away, but he notes that a controlling entity still sits atop every app store. In his view, community hubs like app stores and social media need a public infrastructure layer. He suspects governments skip that investment because companies already run the services, while selling your data. Apple Wants 15 Percent of Purchases Outside Its Store The other half of the Epic saga is Apple’s proposed link-out commission. After the 2021 anti-steering injunction, Apple charged 27 percent on purchases made through external links. A judge held it in contempt last year, and the Ninth Circuit then allowed a fee limited to the cost of running the system. Judge Yvonne Gonzalez Rogers refused to wait for the Supreme Court, writing that “further delay is unwarranted.” Apple filed 15 percent for standard apps, 10 percent for subscription renewals and partner programs, and 5 percent for small businesses. It also conceded the rate would be “essentially zero” under the appeals court’s cost yardstick. Since Apple has charged nothing on link-outs since the contempt ruling, Cochrane sees this as a raise. He calls a cut on purchases made on a developer’s own website disturbing. He also recalls reading about the size of Uber’s payments to Apple, and he questions whether that kind of percentage is sustainable for companies without funding. Meta Says It Has Cut Off 750,000 Australian Kids Meta reported locking out more than 750,000 Facebook and Instagram accounts in Australia by the end of June under the country’s under-16 social media law. Over 500,000 of those were removed before the law even took effect. Detection relies mostly on AI scanning posts and bios for tells like birthday messages, plus user reports and blocks on re-registration. However, the post gives no count of mistaken removals or appeals, and the regulator’s early data shows under-16 usage falling only from about 86 to 81 percent. Meta wants a single age signal at the operating system or app store level, and Cochrane agrees completely. He connects it to the MHS idea from the top of the show: platforms need a standard flag to reference instead of guessing. The White House Deputizes Private Hackers Earlier this month the White House signed a National Security Presidential Memorandum that lets vetted private security firms run surveillance and disruption operations against overseas criminal groups. The Justice Department and Homeland Security hold the contracts and oversee the work. Firms need a proven track record, vetted staff, and a bond of at least $1 million, and must submit operating procedures within 60 days. Cochrane finds the measure aggressive in a good way and hopes it deters attacks on innocents. Still, he takes Kevin Beaumont’s warning seriously that the private security industry profits from ransomware existing. He compares it to the old Head and Shoulders myth: why solve the problem that drives your revenue? A Weather Satellite Watched the Eclipse Shadow Cross Europe Cochrane skips the readout on this one and simply sends listeners to ESA’s site. The MTG-I1 weather satellite captured the Moon’s shadow sweeping across Europe during the August 12 eclipse. Watching a shadow cross an entire continent, he says, was a first for him. Additionally, it leaves him excited about the research happening beyond the planet. Rivers, Deltas, and the Number 0.6 Quanta Magazine explains Hack’s law, which John Hack discovered in 1957 while measuring streams in Virginia and Maryland. A stream’s length tracks its drainage area raised to the power of 0.6, regardless of the rock underneath, and satellite data later confirmed it worldwide. Computer models in the 1990s showed why. Channels that capture extra runoff cut deeper and steal from their neighbors until the network settles into the arrangement that wastes the least energy. Now a University of Texas Rio Grande Valley team has found the same 0.6 exponent in river deltas, which spread water out rather than gathering it. Nobody knows why yet, and Cochrane calls it a really cool read. Sugar Helped Grow the Human Brain, Too A new paper in Science, co-authored by Jennie Brand-Miller at the University of Sydney, adds a third ingredient to the story of early human brain growth. Alongside meat and cooking, natural sugars from ripe fruit and honey may have fueled it too. The brain is about two percent of body weight but burns twenty percent of resting energy. It runs on glucose, which meat and marrow barely supply and raw starch cannot release without fire. The team modeled ancestral diets from a chimp-like baseline through Homo erectus and concluded that the earliest hominins may have drawn over 65 percent of their energy from natural sugars. Cochrane stresses that it is a model, not fossils, and notes that paleoanthropologist Marina Lozano thinks the authors place widespread cooking too early. Still, he loves this kind of deep research. Retracing the steps to our own intelligence, he suggests, could hint at what it takes for intelligent life to develop at all. A Brain Rhythm That Tells Doctors Where to Aim Finally, Science Daily covered a University of Cologne study on deep brain stimulation. That is the implanted-electrode treatment that eases Parkinson’s tremors for some patients but not others. Andreas Horn’s team recorded from 50 patients using both the implanted electrodes and an external magnetic scanner. They identified a circuit between the electrode’s target and the frontal cortex that oscillates at 20 to 35 cycles per second. Stronger coupling there predicted bigger improvement after surgery, though the study, published in Brain, shows correlation rather than cause. First author Bahne Bahners hopes the finding helps tune DBS more precisely, especially for patients who have not responded well. Cochrane half-jokingly asks whether MHS might one day drive those electrodes, and he calls brain disorders the hardest thing in the body to treat. Cochrane wraps with housekeeping: become a GNC Insider at geeknewscentral.com/insider, email geeknews@gmail.com with questions or comments, subscribe to the newsletter, and grab a modern podcast app at podcastapps.com. He thanks GoDaddy for over twenty years of keeping the show on the air, promises to catch everyone next Monday, and wishes listeners a great night. The post Eyes, Hands, and a Sense of Timing #1874 appeared first on Geek News Central.
Today we are talking about Security, Vulnerabilities, and how to avoid exposure with guest Dave Welch. We'll also cover Security Scanner as our module of the week. For show notes visit: https://www.talkingDrupal.com/567 Topics What Are CVEs CVE Lifecycle and Disclosure AI Era Security Challenges What CVE Program Excludes Patch Fast Reality Global Security Signals CVE Timing Judgment KEV Flags Explained CVE Updates Link Rot Who Decides CVE Sneaky Patch Dangers ADP Program Fixes Small Team Triage Vulnerability Tsunami AI Autonomous Security Future Legal Pressure Budgets Resources Psalm PHP Static Analysis Tool SARIF format PHP ecosystem Council of roots How AI Broke Open Source Security: End-of-Life Software Is the Most Exposed CVE podcast Vulncon PSIRT Guests David Welch - github: dwelch2344 dwelch2344 Hosts Nic Laflin - nLighteneddevelopment.com nicxvan John Picozzi - epam.com johnpicozzi JD Flynn - dorficus MOTW Correspondent Martin Anderson-Clutz - mandclu.com mandclu Brief description: Have you ever wanted a fast way to catch the security mistakes that slip into custom Drupal code — especially the code your AI assistant just wrote — before it ships? There's a module for that. Module name/project name: Security Scanner Brief history How old: created in July 2026 by Mayank Gupta (mayankguptadotcom) of Acquia Versions available: 1.0.0, which works with Drupal 10.3 and 11 Maintainership Actively maintained — created and shipped its first stable this summer, with steady development right through late July Security coverage Test coverage — and it's strong: unit and kernel tests, including a regression corpus built from real Drupal core advisories Documentation? In-depth README with a full check table and CI recipes, plus a CHANGELOG Number of open issues: 1 issue, not a bug Usage stats: 2 sites (it's brand new) Module features and usage Provide a Drush command, has no UI — you point drush security:scan at a module or any path, it reads the code statically, and prints a prioritized, OWASP-mapped list of things to review It's built for the age of AI-written code — the checks target the classes AI assistants keep reintroducing: routes with no access check, #markup and |raw XSS, missing CSRF tokens, unserialize() on untrusted data, hardcoded secrets Then there's an optional deep pass: with the Psalm static analysis scanning engine installed, it'll trace untrusted input across functions and files to catch cross-function issues. And it's honest about state — the report always says whether that deep pass ran, was skipped, or failed, so a failure never gets mistaken for a clean scan One nice detail under the hood: a tokenizer-backed "code map" that knows whether a match is real code, a comment, or a string — so it won't flag the word "unserialize" sitting in a doc comment. That kills the single biggest source of false positives The checks are regression-tested against real Drupal advisories (Drupalgeddon, Drupalgeddon2, the 2019 unserialize bug, etc) so a pattern that caused an actual CVE can't quietly come back in your custom code Output comes in three flavors: a readable table, JSON for CI and AI agents, and SARIF — which means findings show up as annotations right on your GitHub or GitLab merge-request diff instead of buried in a job log For adopting it on an existing codebase there's a baseline file — you fingerprint the findings you've reviewed, with a required reason on each, and they stop failing the build but never go invisible; every run still counts them It exits non-zero on error-level findings, so it drops straight into CI or a pre-commit hook And it's extensible — checks are Drupal plugins with a #[SecurityCheck] attribute, so any module can add its own or alter the ones that ship Big caveat, and the module says this itself: a finding means "review this," not "this is broken." Static analysis has false positives, and a clean scan doesn't prove the code is secure — access-control logic especially still needs human review I first heard about this module over beverages at Drupalcamp Asheville, so I know that this module was largely vibe-coded, after having an AI agent ingest every single Drupal security team CVE. So I like to think of this module as security pattern recognition tool, but of course it does even more
Si llevas años usando Bash, Zsh o Fish y piensas que los pipes de Unix son lo más parecido a la perfección, este episodio te va a hacer tambalear los cimientos. Porque existe un shell que no pasa texto entre comandos: pasa estructuras de datos. Tablas, listas, registros, fechas, tamaños de archivo con tipo real. Y encima habla con Ollama sin que tengas que escribir ni una línea de Python.Ese shell es Nushell. Está escrito en Rust, tiene más de 40.000 estrellas en GitHub, y su filosofía es sencilla: los pipes deberían transportar datos con tipo, no texto que luego parseas con awk, sed o jq.En este episodio te cuento mi experiencia pasando de Fish a Nushell con ejemplos reales. Cuando escribes ls no obtienes texto: obtienes una tabla con columnas tipadas. Puedes hacer ls | where size > 1mb | sort-by size sin recurrir a awk ni números mágicos. El shell entiende qué es un filesize, qué es una fecha, qué es un número.Y luego está open, que entiende el formato por la extensión: JSON, YAML, TOML, CSV, SQLite... todo se convierte en datos estructurados. Abres un SQLite y ejecutas consultas con query db. Y todo combinable: http get a una API, filtrar con where y guardar con save — en un solo pipeline, sin archivos temporales.La guinda es la integración con IA. Como Nushell entiende JSON y Ollama habla JSON, se entienden a la perfección. Te enseño un pipeline que lista procesos, filtra los que consumen más de 100MB de RAM, se los manda a un modelo local, y mata el que más memoria usa. Todo en una línea. También te hablo de ai.nu, un módulo que envuelve Ollama, OpenAI y DeepSeek, con function calling desde el shell.También hago una comparativa: Bash, Zsh, Fish y Nushell cara a cara. Bash funciona en cualquier sitio pero el manejo de datos es arcaico. Zsh es Bash con esteroides pero los pipes siguen siendo texto. Fish es moderno pero no entiende de tipos. Nu es el único con estructuras de datos de verdad. PowerShell fue el primero en pasar objetos, pero Nu es lo que PowerShell debería haber sido.Capítulos del episodio:0:00 — Introducción: de Bash a Fish, la evolución de las shells2:30 — El problema del texto plano: por qué Nushell es diferente5:00 — La trifecta: ls, where y select, SQL en tu terminal7:30 — Tipos reales: la shell entiende fechas, tamaños y números10:00 — Open: abrir JSON, CSV, YAML y SQLite sin herramientas externas13:00 — Procesamiento avanzado: $in, save, append y par-each15:30 — HTTP GET: APIs de GitHub y meteorología desde la shell18:00 — Comparativa de shells: Bash vs ZSH vs Fish vs Nushell21:00 — Nushell e IA: integración nativa con Ollama sin Python24:00 — Instalación, casos de uso y conclusiones finalesMás información y enlaces en las notas del episodio
Building something people actually want is supposed to be the happy ending. But it arrives with a bill attached: feature requests you didn't ask for, pull requests you'd rather not maintain forever, users demanding the one thing you swore you'd never build, and — if you're unlucky — a company where the sales team quietly starts deciding what engineering works on. DuckDB has spent the last two and a half years working through that list. So how do you stay a database engineering team when success keeps trying to turn you into something else?Hannes Mühleisen, co-creator of DuckDB, is back to talk through the answers they've landed on. Their fix for unwanted pull requests was an extension mechanism, which then forced them to make every part of the engine pluggable — including the parser, which meant ripping out 20,000 lines of Postgres' yacc grammar and rewriting SQL parsing on top of PEG. Their fix for the users demanding client-server was Quack, a protocol designed by people who'd already published a paper on why every existing database wire protocol is wrong. And their answer to Apache Iceberg, after three years of implementing it, was DuckLake: throw out the Avro-and-JSON metadata files and keep the metadata in a database, on the grounds that the Iceberg REST catalog has a Postgres in it anyway.Which brings us to the news Hannes breaks in this episode: DuckDB Labs is being acquired by AWS, while the DuckDB Foundation, the project and its licence stay where they are. There's the question of why a profitable, self-funded, 30-person company in Amsterdam would take that deal, what commitments you write into the contracts when you're worried today's promises might outlive today's management, and what it's actually like to have a boss again after five years without one. If you're curious how an open source project keeps its technical soul once the enterprise arrives — or you just want to know why parsing SQL is harder than parsing almost anything else — Hannes has some good answers.---Support Developer Voices on Patreon: https://patreon.com/DeveloperVoicesSupport Developer Voices on YouTube: https://www.youtube.com/@DeveloperVoices/joinOur previous episode with Hannes: https://youtu.be/pZV9FvdKmLcDuckDB: https://duckdb.org/DuckDB Foundation: https://duckdb.foundation/DuckLabs (formerly DuckDB Labs): https://ducklabs.com/DuckLake: https://ducklake.select/Quack (DuckDB's client-server protocol): https://duckdb.org/quack/DuckDB v2.0: Your Database Deserves a Better Parser: https://duckdb.org/2026/08/20/duckdb-20-peg-parserRuntime-Extensible Parsers (CIDR 2025 paper): https://duckdb.org/pdf/CIDR2025-muehleisen-raasveldt-extensible-parsers.pdfDon't Hold My Data Hostage (VLDB 2017 paper): https://www.vldb.org/pvldb/vol10/p1022-muehleisen.pdfcpp-peglib: https://github.com/yhirose/cpp-peglibGNU Bison: https://www.gnu.org/software/bison/PEP 617 – New PEG parser for CPython: https://peps.python.org/pep-0617/PRQL: https://prql-lang.org/Apache Iceberg: https://iceberg.apache.org/CWI (Centrum Wiskunde & Informatica): https://www.cwi.nl/en/DuckCon #7, Amsterdam: https://duckdb.org/events/2026/06/24/duckcon7/Kris on Bluesky: https://bsky.app/profile/krisajenkins.bsky.socialKris on Mastodon: http://mastodon.social/@krisajenkinsKris on LinkedIn: https://www.linkedin.com/in/krisjenkins/
¿Sigues usando find y grep como en los 90? Hace unas semanas me puse a buscar un archivo en un repositorio con git, lancé el find de toda la vida, y cuando volví de tomarme un café —literalmente— todavía seguía buscando. El problema es que find se mete en el .git, en los binarios, en sitios donde no debería. Y grep, pues lo mismo, sobre todo si trabajas con Unicode o con repositorios grandes. Así que llevo un tiempo usando fd y ripgrep, dos herramientas escritas en Rust que son órdenes de magnitud más rápidas. Pero lo mejor no es solo la velocidad: es que puedes combinarlas para crear pipelines que alimenten directamente a tu IA local.fd (44.1k estrellas en GitHub) es un reemplazo directo de find. En los benchmarks oficiales, buscar archivos con fd -u tarda 0.8 segundos donde find necesita 11 segundos con -iname y casi 20 segundos con -iregex. 23 veces más rápido. Y no solo es velocidad: fd respeta .gitignore por defecto, soporta expresiones regulares directamente, y tiene placeholders como {}, {.}, {/} y {//} que te permiten ejecutar comandos sobre cada resultado con -x o pasarlos en lote con -X.ripgrep (67.4k estrellas) es lo mismo pero para buscar texto. En el kernel de Linux, rg tarda 0.08 segundos donde grep tarda 2.67 segundos. 32 veces más rápido. Y tiene superpoderes que grep ni sueña: salida en JSON con --json, búsqueda en archivos comprimidos con -z, soporte PCRE2 con -P para lookaheads, y un flag --passthru que te muestra también las líneas que no coinciden. Desde la versión 15 también soporta hyperlinks OSC 8 y respeta repositorios de Jujutsu.Pero lo que realmente me tiene enganchado es combinarlos. El patrón es sencillo: fd encuentra los archivos que te interesan, ripgrep extrae el contexto relevante, y todo eso se lo pasas a Ollama para que lo procese. Te enseño la función aresumen que me he montado en Bash, y su equivalente en Fish, para preguntarle a mi documentación local sin salir de la terminal. Cosas como "resume todo lo que he escrito sobre Ollama en el último mes" se resuelven con un pipeline de tres comandos. Sin RAG, sin bases de datos vectoriales, sin complicaciones. Solo con un pipe bien puesto y el modelo adecuado.También te cuento cómo usar jq para procesar la salida JSON de ripgrep, cómo montar un buscador interactivo con fzf y bat, y los errores más comunes al construir estos pipelines. Si alguna vez has pensado "ojalá pudiera preguntarle a mis propias notas desde la terminal", este episodio te va a gustar. Y si todavía usas find y grep, te aseguro que después de oír los benchmarks no vuelves atrás.Capítulos del episodio:0:00 — Introducción: fd y ripgrep para alimentar a tu IA2:30 — El problema con find y grep tradicionales5:00 — fd: el find que siempre quisiste tener8:00 — Placeholders y expresiones regulares en fd11:00 — ripgrep: el grep con superpoderes14:00 — Salidas estructuradas con JSON y jq16:30 — La combinación estrella: fd + ripgrep con -x19:00 — Pipelines avanzados para filtrar archivos21:30 — Integración con IA local: fd + rg + Ollama24:30 — Alias, funciones y trucos del día a día27:00 — Despedida y conclusionesRecursos mencionados:- fd (sharkdp/fd): https://github.com/sharkdp/fd- ripgrep (BurntSushi/ripgrep): https://github.com/BurntSushi/ripgrep- Ollama: https://ollama.com- jq: https://jqlang.github.io/jq/- fzf: https://github.com/junegunn/fzf- bat: https://github.com/sharkdp/batMás información y enlaces en las notas del episodio
In "Eliminating EDI Headaches: Managed Integration vs. Ticket Queues", Joe Lynch speaks with Co-Founder and Leader of Atadex, Mitch Bernet, about how fully managed EDI and API integrations eliminate operational bottlenecks, speed up customer onboarding, and drive supply chain profitability. About Mitch Bernet Mitch Bernet is the Co-Founder and Leader of Atadex, bringing decades of supply chain expertise and executive leadership to the role. Raised in Cleveland, Ohio, he earned a degree from Providence College, an MBA in Finance from Creighton University, and launched his career at Union Pacific Railroad before taking senior management roles at Conrail, APL Logistics, and Hub Group. A seasoned entrepreneur, he went on to found Integra Logistics in 2003 and co-found Coyote Logistics in 2008—growing the Atlanta-based business through its eventual acquisition by UPS—before taking a brief hiatus and returning to the industry to co-found Atadex. Married to Rosanna for 35 years, he is the proud father of two children, Matthew and Katherine, who graduated from Georgia Tech and the University of Georgia, respectively. About Atadex Atadex is a full-service EDI (Electronic Data Interchange) and API integration provider exclusively focused on the supply chain and logistics industry. The company was founded to deliver fully managed, fast, low-cost data integration—handling everything end to end so clients don't need in-house EDI expertise and positioning itself as an alternative to legacy EDI vendors, self-serve platforms, and traditional professional services. Atadex connects any TMS or WMS (including Blue Yonder, Trimble, Manhattan, Turvo, Shipwell, TAI, and McLeod) to any trading partner, supporting all major formats (X12, EDIFACT, API, XML, CSV, JSON) and transaction types. The company processes over 50 million messages per month and has completed more than 5,000 go-lives, pairing each client with a dedicated, US-based account manager. Its goal is to fundamentally change how data integration is managed and serviced, helping clients improve productivity, performance, and profitability. Target customers include carriers and 3PLs with complex partner networks, warehouses and 4PLs seeking a fully outsourced EDI department, and TMS/WMS providers wanting to offer integration without building their own delivery teams. Atadex also offers AssetMaps (fleet consolidation) and Freight Board Central (freight board integration). Key Takeaways: Eliminating EDI Headaches: Managed Integration vs. Ticket Queues In "Eliminating EDI Headaches: Managed Integration vs. Ticket Queues", Joe Lynch speaks with Co-Founder and Leader of Atadex, Mitch Bernet, about how fully managed EDI and API integrations eliminate operational bottlenecks, speed up customer onboarding, and drive supply chain profitability. EDI as a Profitability Driver, Not Just an IT Task: Data integration directly impacts the bottom line by eliminating manual data entry, reducing human error, and allowing logistics staff to manage up to 25% more volume per person without adding head count. Overcoming System Incompatibility via "Universal Translation": Supply chain tech remains heavily fragmented between legacy systems (like mainframes or EDI X12) and modern RESTful APIs. Atadex acts as a universal translator, mapping and converting any data format into whatever spec a trading partner requires. Speed to Integration Equals Speed to Revenue: Onboarding delays can stall new customer relationships for months and destroy projected margins. Rapid implementation and proactive partner coordination directly protect customer retention and protect sales commissions. Fully Managed Service vs. Anonymous Ticket Queues: Logistics operates on extreme urgency where every minor issue feels like a major emergency. Relying on dedicated, domain-expert project managers who proactively catch failed transmissions beats putting in tickets with traditional tech providers. Decoupling Customer Support from Technical Engineering: Mirroring the operational structure pioneer Jeff Silver used at Coyote Logistics, separating client-facing project managers from back-end technical mappers ensures seamless communication and higher overall customer satisfaction. Transparent Pricing Over Complex VAN Fees: Traditional EDI value-added networks (VANs) often obscure costs behind per-kilo-character charges. Simplifying billing to clear, per-message rates provides total transparency so companies can calculate precise integration costs per trading partner. Practical AI Integration with a "People-First" Foundation: While AI tools (like Claude) are being developed internally to speed up complex mapping and protocol translation, technology will supplement—rather than replace—the critical human element required to support high-stakes logistics operations. Learn More About Eliminating EDI Headaches: Managed Integration vs. Ticket Queues Mitch Bernet | Linkedin Atadex | Linkedin Atadex USMMG Partners with Atadex for EDI Solutions | LinkedIn Trailer Bridge's Partnership With Atadex as Their EDI Provider Revolutionized Their Efficiency | LinkedIn The Game-Changer for Riverside Transport's EDI and Target Performance | LinkedIn Atadex| Facebook The Logistics of Logistics Podcast If you enjoy the podcast, please leave a positive review, subscribe, and share it with your friends and colleagues. The Logistics of Logistics Podcast: Google, Apple, Castbox, Spotify, Stitcher, PlayerFM, Tunein, Podbean, Owltail, Libsyn, Overcast Check out The Logistics of Logistics on Youtube
Kamille Parks demonstrates how to use the Model Context Protocol (MCP) to manage Airtable automations via Claude. She walks through updating an existing automation by adding placeholder email addresses to a notification step, showing how the AI can modify specific parameters while preserving existing context. The demo also covers management tasks like deleting an automation and using the "undo" feature to revert actions. By leveraging action IDs tracked through the MCP, the AI can effectively restore what was previously deleted. Finally, the segment explores creating new scheduled automations from scratch and exporting all base automations as JSON for auditing and configuration reviews. ⏱ In this cut: 01:58 — Updating an existing automation 06:52 — Deleting and undoing actions 08:36 — Creating new scheduled automations 16:11 — Exporting automations as JSON
James Pope is on site in Las Vegas more than a week before the doors open. As SOC lead for the Black Hat NOC and Senior Director of Security Product Research and Technical Marketing Engineering at Corelight, his show starts with switches and access points rather than alerts. The team brings in the ISP, the firewall, the switches, and the access points, deploys them across the conference, and then moves into SOC mode. If there is no network, there is nothing to secure. The tooling arrives through partnership rather than sponsorship. James Pope says a company cannot buy or sponsor its way into the NOC, and that the team picks what it wants and fills gaps as it finds them. Cisco covers Umbrella and file malware analytics, Palo Alto Networks provides the firewall and XSIAM as the log aggregator, Arista handles switching and access points, Jamf runs MDM across the registration devices, and Lumen supplies the internet. Corelight is the network visibility layer. That layer carries different weight here than it would inside a company. Asking attendees to install a certificate or an endpoint agent so the NOC can inspect their traffic is a request nearly everyone declines. In most corporate environments the endpoint is one of the richest sources of signal. At Black Hat, visibility into attendee activity comes from network data. A Black Hat positive is malicious activity that is legitimate in context. Attendees pay to learn attack techniques against real targets, and researchers demonstrate new exploits on stage. Those events generate true detections no corporate SOC would ignore. The NOC lets them run rather than killing a paid training exercise or a live demo. So how does the team tell a training exercise from a real attack? It baselines each classroom and spends its time on the outliers. When seventy students in a room run the same attacks against the same destinations, the activity is probably sanctioned. The curriculum is ingested as a JSON file and the system moves through a series of gates, asking whether this is a class, whether multiple sources are reaching the same destination, and whether the attack would be expected in that curriculum. Anything that does not fit comes back for a human. The team informs far more often than it blocks. On the day of the recording, James Pope went to the trade show floor to tell someone that command and control traffic was running from their machine, and handed over logs for their IT and security team. He is not their manager, and what happens next is their call. Illegal activity is treated differently, and a handful of times per show the team asks a room to stop. At Black Hat Asia, traffic from a Corelight sensor showed a double RAT infection on one machine, a single APT running one implant for exfiltration and another for command and control. Working from traffic, James Pope established that the person was a reporter, the region they covered, and the company they worked for. Open source intelligence narrowed it to a single name, registration confirmed the person was on site, and the NOC invited them in. The reporter arrived expecting a product demo. The laptop was reset with everyone present, sessions were revoked, passwords were changed, and the reporter left in a secured state. This year the team opened the Outpost, running real Black Hat network logs from Corelight behind application guardrails, LLM guardrails, and a kill switch, where visitors query the data with text to SQL. Agentic triage stitches alerts into detections and detections into a timeline, and James Pope treats the ability to drill down to raw logs as a requirement rather than a preference. Success is measured largely by what does not happen: no compromise of registration, the switches, or the access points, and people who arrive infected leaving better than they got here. This is a Brand Briefing. A Brand Briefing is an on-location conversation recorded on site at Black Hat USA 2026, putting a spotlight on the guest and their company and pairing it with the editorial reach of ITSPmagazine. Learn more: https://www.studioc60.com/performance/#briefing GUEST James Pope, Senior Director of Security Product Research and Technical Marketing Engineering at Corelight, and SOC lead for the Black Hat NOC RESOURCES Black Hat USA 2026 event coverage from ITSPmagazine: https://www.itspmagazine.com/black-hat-usa-2026-cybersecurity-event-coverage-in-las-vegas Learn more about Corelight: https://corelight.com Corelight blog, including the Black Hat NOC series: https://corelight.com/blog Are you interested in telling your story? ▶︎ Full Length Brand Story: https://www.studioc60.com/content-creation#full ▶︎ Brand Spotlight Story: https://www.studioc60.com/content-creation#spotlight ▶︎ Brand Highlight Story: https://www.studioc60.com/content-creation#highlight ▶︎ Get your own Brand Briefing at an upcoming event: https://www.studioc60.com/buy-brand-briefings KEYWORDS james pope, corelight, sean martin, marco ciappelli, brand briefing, brand story, brand marketing, marketing podcast, black hat usa 2026, network detection and response, network evidence, security operations center, threat hunting, agentic triage, ai in the soc, conference network security, black hat noc, command and control, incident response Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Alli Alosa demonstrates how to use the 'Insert as JSON' feature in Airtable automations to reliably handle incoming data from external webhooks, such as Fillout. She explains why this method is more robust than traditional field mapping, especially when dealing with inconsistent data. The discussion covers a practical use case involving event feedback collection. By passing the entire payload as a JSON object, you can prevent automations from failing when certain form fields are left empty or missing from the payload. Finally, the team explores how to use this JSON data within a script action. Instead of manually mapping every individual field as an input, you can pass the entire object into a script, allowing for much more scalable and maintainable automation logic. ⏱ In this cut: 01:03 — Using feedback collection as a use case 12:38 — The benefits of using Insert as JSON 13:37 — Solving errors caused by empty field values 17:00 — Passing JSON data into a script action 25:09 — Conditional filters and advanced integrations
Tentative de nouveau format : "FAQ Discord". On a récupéré toutes vos questions posées sur notre Discord et on y répond dans cet épisode : indexation, revente de sites, DMCA, Trust Flow, éthique de l'affiliation, YouTube… 1h17 de réponses sans filtre.
Topics covered in this episode: Claude Code /insights Post-quantum crypto lands in Python MCP goes stateless — and FastMCP gets renamed inshellisense - IDE style command line auto complete Extras Joke Watch on YouTube About the show Sponsored by Xweather Xweather combines enterprise-grade weather intelligence with agent-ready APIs, natural language capabilities, and an MCP server so your agents can adapt workflows, automate responses, and make better decisions based on real-world conditions. Michael will tell you more about them later in the show. Get started for free at pythonbytes.fm/xweather Connect with the hosts Michael: Mastodon / BlueSky / X / LinkedIn Calvin: Mastodon / BlueSky / X / LinkedIn Show: Mastodon / BlueSky / X Join us on YouTube at pythonbytes.fm/live to be part of the audience. Usually Tuesday at 7am PT. Older video versions available there too. Finally, if you want an artisanal digest of every week of the show notes in email form? Add your name and email to our friends of the show list, we'll never share it. Michael #1: Claude Code /insights Michael's Insights: michael-kennedy-claude-code-insights-2026-08-09.html Be careful sharing these outputs, they include details references to your projects, errors, security findings, etc. ;) /insights reads your last 30 days of local session transcripts and hands back an interactive HTML report on how you actually work. One command, zero setup: type /insights in a session, or run claude -p "/insights" from the shell for a non-interactive version that just prints the path Reads what's already on disk: pulls session logs from ~/.claude/projects/, skipping agent sub-sessions and anything under 2 messages or 1 minute Project areas: clusters your sessions into themes like "CLI Tooling" or "Documentation" with session counts Friction analysis: categorizes where things went wrong by root cause - and quotes your own prompts back at you Interaction style: tells you whether you're a delegator or a micromanager, plus which workflows are worth doubling down on Actually actionable: suggests concrete CLAUDE.md additions and Claude Code features you're not using The catch: Haiku does the per-session classification, so the first run takes several minutes; results cache to ~/.claude/usage-data/facets/ and the report lands at ~/.claude/usage-data/report.html Calvin #2: Post-quantum crypto lands in Python pyca/cryptography 48 ships ML-KEM (key establishment) and ML-DSA (signatures) — NIST's post-quantum standards, now one pip install away. Big deal because it's the 11th most-downloaded package on PyPI (~1.2B downloads/month) and sits under Ansible, Certbot, Airflow, and paramiko. No PQ there, no PQ anywhere in Python. Trail of Bits did the work (Rust bindings, cross-backend API, tests, AWS-LC backend support), funded by the Sovereign Tech Agency. Timing tracks a June 22 White House order setting federal deadlines: PQ key establishment by end of 2030, PQ signatures by end of 2031. Not a drop-in swap — the wire sizes explode. ML-DSA-65 signatures are 3,309 bytes vs Ed25519's 64; ML-KEM-768 public keys are 1,184 bytes vs X25519's 32. Hardcoded field sizes and length prefixes will bite. API looks like the existing asymmetric primitives, except ML-KEM is encapsulate/decapsulate rather than a Diffie-Hellman exchange. SLH-DSA (the hash-based conservative backstop) is still in progress. The primitives are here, but protocols haven't caught up — so you won't be running post-quantum Certbot this week. Sponsor: Xweather You're using agents that can write code, summarize documents, and automate workflows. But they're missing one thing: awareness of the world around them. This is where today's sponsor, Xweather comes in. Xweather combines enterprise-grade weather intelligence with agent-ready APIs, natural language capabilities, and an MCP server built for tools like Claude, Codex, Copilot, and modern IDEs – so your agents can adapt workflows, automate responses, and make better decisions based on real-world conditions. Backed by Vaisala, whose instruments fly on NASA missions to Mars, Xweather delivers trusted data and unique insights that go beyond conditions to actual impact – from real-time lightning strikes to road surface forecasts. Start with 15,000 free API calls each month and pay only for what you use as you grow. Xweather is your full weather stack, for developers by developers. Start building for free today at pythonbytes.fm/xweather. The link is in your podcast player's show notes and on the episode page. Thanks so much to Xweather for supporting Python Bytes. Calvin #3: MCP goes stateless — and FastMCP gets renamed From Philipp Acsany over at Real Python The 2026-07-28 spec landed July 28 and the Python SDK shipped 2.0.0 the same day. Biggest rewrite since MCP launched, and it's breaking on purpose. Context for scale: the Tier 1 SDKs are pulling close to half a billion downloads a month, with TypeScript and Python each past a billion total. The headline is the stateless core. The initialize/initialized handshake and the Mcp-Session-Id header are both retired — protocol version, client identity, and capabilities now ride in _meta on every request, with an optional server/discover RPC if a client wants capabilities up front. Any request can land on any instance behind plain round-robin, no shared storage. Server-initiated calls are the hard part of the migration. Sampling, elicitation, and roots/list no longer call back to the client; instead the server returns resultType: "input_required" and the client retries with inputResponses attached. Multi Round-Trip Requests, MRTR. Also: Mcp-Method and Mcp-Name are now required headers so gateways route on headers instead of cracking JSON bodies, and missing-resource errors move to standard 32602. Deprecation sweep with an actual policy behind it — Roots, Sampling, Logging, and the legacy HTTP+SSE transport all deprecated with a twelve-month minimum offramp. Tasks graduated out of the experimental core into a real extension, which is what the formalized extensions framework was for. MCP Apps is now an official extension too, so a tool call can return sandboxed interactive HTML. Auth picked up RFC 9207 issuer validation, issuer-bound credentials, and a shift from DCR toward CIMD. Python SDK 2.0 is where it gets personal: FastMCP is now MCPServer, no alias, no shim. McpError → MCPError. Wire types went snake_case (is_error, input_schema) and moved to a standalone mcp_types package, with mcp.types kept as a permanent alias. One Client object replaces the old transport + ClientSession + initialize() stack. httpx became httpx2. Sync handlers run on worker threads now, so asyncio.get_running_loop() raises inside them. The good news: one MCPServer serves both protocol eras, so 2025-era clients keep working with nothing to configure, and a Resolve(fn) parameter lets one tool body cover MRTR and the old path. 1.x is maintenance-and-security-fixes only — pin mcp>=1.28,
Members Only: Today's video is available only to members. If you are already a member, you can access your private podcast feed by visiting https://www.pointfree.co/account. --- We add a geofence to the `Trip` model and explore advanced features of SQLite, including JSONB and `json_each`, which allow us to efficiently store and query this data. And we will show that not all bundles of data need to be JSON: we are free to group our database columns among as many data types as we like.
Free AI users rejoice: there's finally unlimited, free frontier AI.
Today we are talking about Maintaining NodeJS, Patternlab, Writing Books, and Open Source with guest Brian Muenzenmeyer. We'll also cover AI Webform Generator as our module of the week. For show notes visit: https://www.talkingDrupal.com/564 Topics Brian Open Source Origins Pattern Lab Node Journey Maintaining and Moving On Writing Approachable Open Source Who the Book Is For Beyond Code Contributions All Things Open Book Signing Choosing Conferences to Attend Pitching Open Source at Work Misconceptions and Starting Small Avoiding Maintainer Burnout Handling AI Noise and Low Effort PRs DCO and Licensing Basics Better Communication and Reviews Node and Drupal Lessons Optimism for Open Source Future Resources Brian Muenzenmeyer https://brianmuenzenmeyer.com https://approachableopensource.com/ https://bsky.app/profile/brianmuenzenmeyer.com https://www.linkedin.com/in/brian-muenzenmeyer-91a77554/ https://www.renderatl.com/schedule upcoming https://nodeconf.eu/program upcoming spectrum of engagement https://approachableopensource.com/blog/2025-open-source-pace-layers/ change in contention https://brianmuenzenmeyer.com/posts/2018-i-maintainer/ burnout https://approachableopensource.com/read/the_spectrum_of_engagement/ https://approachableopensource.com/read/the_four_files_of_any_open_source_project/ LICENSE Hodag Cryptid https://en.wikipedia.org/wiki/Hodag https://www.rhinelanderchamber.com/about-the-hodag/ You should write a book All contributors spec Talk at all things apart DCO Developer Certificate of Origin Open source law policy and practice Sustain OSS Guests Brian Muenzenmeyer - brianmuenzenmeyer.com Hosts Nic Laflin - nLighteneddevelopment.com nicxvan John Picozzi - epam.com johnpicozzi Bernardo Martinez - bernardm28 JD Flynn - dorficus MOTW Correspondent Jacob Rockowitz - jrockowitz.com jrockowitz Brief description: AI Webform Generator enables site builders to create a Drupal Webform, or update an existing one, from plain-English instructions. It sends the request through the site's configured Drupal AI provider, validates the returned Webform definition, and saves the resulting form. Review the generated change before using the form. Module name/project name: AI Webform Generator (ai_webform_generator) Brief history Created on 2 July 2026 by chaitanyadessai (Chaitanya R Dessai). The current stable release is 1.0.2, released on 3 July 2026, and supports Drupal ^10 || ^11. Maintainership Appears actively maintained: Drupal.org lists an update on 24 July 2026. Maintainers: zeeshan_khan and chaitanyadessai. (Specbee) Security coverage: Yes. Stable releases are covered by Drupal's security advisory policy. Test coverage: Yes. Version 1.0.2 includes unit, kernel, and functional tests for prompt building, JSON validation, settings, route access, Webform building, and optional CAPTCHA elements. Documentation: Yes. The project page and module README cover requirements, configuration, usage, security considerations, and supported field types. Issues: 1 open issue, with 0 open bug reports (7 issues total). Usage stats: 1 site reports using this module. Module features and usage Creates complete Webforms and updates existing Webforms in place from natural-language prompts. Supports common Webform elements, including text, email, telephone, number, date, select, checkbox, radio, range, password, hidden, and managed-file elements. Validates the AI response before applying the Webform definition. Uses the existing Drupal AI provider configuration; API keys are not stored in this module's configuration. Provides configurable model, temperature, output-token, and per-user request limits to balance output quality and provider spend. Requires trusted users with both the generator permission and ordinary Webform edit access when changing an existing form. AI-Generate Notes, Review, and Recipe (used for testing) https://github.com/jrockowitz/drupal_playground/tree/main/recipes/drupal_playground_webform_ai AI-Generated Assessment Technical: The module separates AI generation, prompt building, JSON validation, and Webform construction into Drupal services. It uses the site's configured Drupal AI provider, validates a limited allowlist of Webform element types before saving, and exposes model, temperature, output-token, and per-user request-limit settings. Access and error handling: Generation requires its own permission, and updating an existing Webform also requires normal Webform update access. A per-user flood limit constrains provider spend; failures are logged, with detailed upstream errors shown only to generator administrators. Code quality: Version 1.0.2 uses strict types and separates form, service, validation, and persistence responsibilities. It includes unit, kernel, and functional coverage for core behavior. This assessment is a code review of the released module, not a security audit. Implementation: The module creates new Webforms and updates supported fields of existing Webforms in place, but saves the generated definition immediately without a preview, diff, or approval screen. Usefulness: The module is useful for quickly drafting straightforward Webforms and iterating on common field changes when a site builder reviews the result. Complex, highly customized, or regulated forms need especially careful manual review before publication. How to use it: Configure a chat-capable provider, select an existing Webform or choose to create one, describe the fields and validation in plain English, submit the request, and then review the saved Webform. For example, create a disposable contact Webform and ask the generator to add a required telephone field while preserving the existing fields. AI-generated source code: The module's runtime use of AI and its code style cannot establish whether its source was AI-generated or AI-assisted. Its public project metadata does not make an authorship claim, so this is unknown. Possible improvements: Add a preview/diff and explicit approval before saving; broaden support for advanced Webform structures and handlers; add optional, privacy-conscious prompt and response audit logs; and expand regression coverage for complex Webform updates. Next steps for adopters: Restrict generation to trusted roles, begin with a low request limit, test representative prompts outside production, and review every generated field, validation rule, confirmation message, and permission before publishing.
Watch the full episode on YouTube:We first covered Baseten last year when DeepSeek mania was at peak hype. Now they have raised a monster $13B round and become one of the new cohort of AI Infra decacorns that are (with Nvidia, Intel, and the semis complex) chief beneficiaries of the Inference Inflection. We return to Baseten at the peak of the 2026 edition of Open Weights debate. Ali has published a viral breakdown of Kimi K3:And since you last saw him, Philip has spoken at AI Engineer and written the definitive book on Inference Engineering spotted all over SF:Three years ago, inference engineering barely existed as a category.Today, it is one of the most critical disciplines in AI. Inference engineering inherently tackles a different question than standard model training: “How do you turn those weights from training into a product that is fast, reliable, and affordable at scale?” Focusing on these creates an entirely new optimization problem.In one recent GLM-5.2 experiment, quantizing more of the model actually preserved its benchmark quality while increasing throughput by 20%, because the errors introduced in different layers could cancel each other out.Inference is no longer just the final step after training. It is becoming its own engineering discipline, with its own research problems, infrastructure, and increasingly specialized roles.In this episode, Baseten's Philip Kiely and Ali Taha join swyx and Vibhu to explain what actually happens after a new open model is released and what it takes to turn “we generated a token” into a fast, reliable, production-ready API.We go deep on cache-aware routing, disaggregated prefill and decode, quantization, speculative decoding, KV-cache movement, model parallelism, GPU kernels, and the race to make frontier models up to 10× faster. Philip and Ali explain why inference optimizations can still produce gains of 20%, 100%, or even 200%; how quantization errors can cancel one another out; why identical weights can behave differently across clusters; and how Baseten grafted a Kimi vision encoder onto GLM-5.2 without changing the underlying language model.The conversation then expands beyond LLMs into NVIDIA Dynamo, mega kernels, Rubin, AI-specific chips, local inference, video generation, diffusion versus autoregressive models, and the enormous compute barrier to generating coherent long-form video. Finally, we explore the convergence of training and inference, continual learning through persistent KV cache, and the emerging loop where models help optimize the infrastructure that runs them.We discuss:* What happens when a 200,000-token request enters an inference system* Cache-aware routing and reusing previously computed KV cache* Why prefill and decode are increasingly handled by different GPUs* When dedicated deployments become cheaper and more reliable than shared APIs* How speculative decoding uses a smaller model to accelerate a larger one* Tool calling, structured outputs, and what LLMs actually do* What it takes to support a new open model on day zero* Grafting Kimi's vision encoder onto GLM-5.2* Retrofitting inefficient model layers with components from other architectures* Why models sometimes collapse into repeating the same token* How hardware, kernels, and race conditions create nondeterministic failures* Preserving model fidelity while making inference faster* How quantization errors can cancel each other out* Why inference optimizations still deliver gains of 20%, 100%, and 200%* How optimized serving can make a model up to 10× faster* NVIDIA Dynamo, KV-aware routing, and distributed model serving* Speculative decoding the speculative decoder* Why local AI is about making models less dumb while data-center AI is about making them less slow* Tensor, expert, and pipeline parallelism across GPUs* Hardware-aware model design, auto-tuning, and the case against mega kernels* Rubin and why inference is becoming a systems problem* Whether modern GPUs are evolving into programmable AI ASICs* Why enormous models like Kimi K3 require GB300-class hardware* Why open-source video generation still trails Veo, Kling, and other closed models* The quadratic attention bottleneck behind long-form AI video* Autoregressive video, real-time generation, and compounding quality drift* Why future video systems may combine autoregressive and diffusion architectures* Training for inference and inference for training* Continuous post-training, deployment, evaluation, and improvement loops* How GLM-5.2 helped optimize the kernels serving GLM-5.2 itself* Why faster networking could unlock dramatically faster decoding* Continual learning, KV-cache compaction, and persistent model memoryShow Notes* How to build a day-0 API for Kimi K3* 22580: From GPT2 to Kimi3, ExplainedPhilip Kiely* LinkedIn: https://www.linkedin.com/in/philipkiely* X: https://x.com/philipkiely* Inference Engineering: https://www.baseten.co/inference-engineering/Ali Taha* LinkedIn: https://www.linkedin.com/in/aliestaha/* X: https://x.com/waterloointernTimestamps00:00:00 Introduction and the 200K-Token Prompt00:03:18 Dedicated Deployments, Speculative Decoding, and Tool Calling00:11:26 Launching Production-Ready Open Models00:19:06 Model Retrofits, Failure Modes, and Nondeterminism00:28:22 Quantization and Canceling Errors00:32:15 The Race to 10× Faster Inference00:40:48 Dynamo, Speculation, and Local vs. Data-Center AI00:50:18 Model Parallelism, Auto-Tuning, and Mega Kernels01:00:55 Rubin, GPUs vs. ASICs, and Custom AI Chips01:10:03 Giant Models and the Limits of GPU Memory01:12:42 AI Video, Quadratic Attention, and Autoregressive Generation01:21:47 Audio, Images, and Diffusion Models01:27:32 Training, Self-Optimizing Models, and Continual Learning01:40:06 Closing ThoughtsTranscriptIntroduction: Baseten, Waterloo Intern, and Inference EngineeringSwyx [00:00:00]: Okay, we're here in the studio with Philip, old friend from Inference Engineering, the book, as well as Baseten and everything that you've done, you and I have done before, as well as Ali. Welcome.Ali [00:00:15]: Pleasure to meet you.Swyx [00:00:15]: Waterloo intern.Ali [00:00:16]: Waterloo intern, always.Swyx [00:00:17]: When did you get “Waterloo intern” as a handle?Ali [00:00:19]: As a handle? Oh.Ali [00:00:20]: I think the rebranding happened mid-March. When I saw it was open, I was like, “I have to take it. Up for grabs.”Philip [00:00:26]: The problem is that Ali is really good at his job and is not gonna be an intern much longer.Philip [00:00:30]: So we have to figure out who's gonna get the handle.Ali [00:00:33]: Well, I'll pass the torch over to the next intern.Swyx [00:00:34]: Oh, okay. It can be, like, you just pass it to another Waterloo grad.Ali [00:00:37]: To another Waterloo intern. No, bruh.Philip [00:00:39]: Yeah.Ali [00:00:39]: Intern.Swyx [00:00:40]: Intern, yeah.Ali [00:00:40]: And no.Philip [00:00:41]: You gotta get an intern from Waterloo.Ali [00:00:42]: Yeah, I've gotta get an intern from Waterloo.Swyx [00:00:44]: Right.Ali [00:00:44]: But they have to follow the path.Swyx [00:00:45]: Oh, it could, but it could come from Baseten, so it's like whoever Baseten gets from Waterloo.Ali [00:00:48]: Right.Swyx [00:00:49]: Has the title of Waterloo.Ali [00:00:50]: It stays in the ecosystem.Philip [00:00:51]: Exactly.Ali [00:00:52]: Halfway through the internship, you either get it or you're out.Philip [00:00:55]: You should also do, like, a big graduation ceremony where you change the handle.Ali [00:00:59]: Just say it.Philip [00:00:59]: For everybody.Swyx [00:01:00]: You guys are good at ceremonies, clearly. We had a nice launch of the book, very successful. But before we get into all that, I wanna start off with a fun question for you. Okay, you're an expert inference engineer. What happens when I send a long query, say two hundred thousand tokens into Baseten's inference? What's the process of query through GPU model routing, balancing, all that? What is all the stuff that we don't think about?Long Context Requests, KV Cache, and Cache-Aware RoutingPhilip [00:01:26]: With a long query specifically, the first thing that I'm gonna ask is, “Have you sent me this query before, or at least part of it?” and I really hope you have, because it's gonna be a lot easier for me and a lot cheaper for you. So the first thing that we're gonna look at is some cache-aware routing, where we're going to see, we probably have a number of instances, a number of replicas up serving whatever model you're hitting. We want to send this one to something with, number one, available prefill workers, and number two, ideally some cached input already there so that we can skip prefill on at least part of these two hundred thousand tokens. If you're doing two hundred thousand tokens, it's probably coding or a multi-turn agent or something where you would expect to have that cached. If you don't, we're gonna have to send it to a prefill worker. We've at least on certain models disaggregated prefill and decode, so you're going to have one set of GPUs that's solely going to process the input, create the KV cache, and get you your first token, and then that's going to be passed over to a separate set of GPUs, which is going to run decode. We're going to iteratively make those tokens. We're probably going to have some speculator model in front of that. I'm going to assume that you're doing coding, and because of that, our speculator model, which assumes you're doing coding, is gonna have a high draft token acceptance rate. If I'm wrong and you're asking me to summarize every Harry Potter book, it's gonna be slower. And then we stream that output to you and account for it, charge you, a couple of pennies and say, “Hey, would you like to send another one?”Swyx [00:03:04]: Except Baseten doesn't charge by pennies.Philip [00:03:07]: Well, yeah, we charge. I'm assuming that we're talking about the public model APIs. If you are setting up a dedicated deployment, then yeah, it's not pennies.Public APIs vs. Dedicated DeploymentsSwyx [00:03:18]: Yeah, one of the key differentiators when I was talking with Baseten initially was that people who want very high volume just need to rent by the box, ‘cause then it's up to you to figure out how to saturate the box.Ali [00:03:31]: And more often than not, it's, like, way cheaper if you're pushing, like, millions of tokens per hour, if you just pay per hour instead of pay per token.Philip [00:03:37]: Yeah, they do. I think that we've increasingly seen a lot of demand for the pay per token APIs, just because everyone wants to try open models, and then once they find a use case that's really sticky, then they move over to dedicated.Swyx [00:03:51]: Is there a best practice on when it's time to swap over?Philip [00:03:54]: Couple reasons. Yeah, reliability, that's a big one, right?Ali [00:03:57]: Like, if they have a very specific use case, they want you to train something specifically for them, like they want their own spec dec, for instance, for their own traffic.Swyx [00:04:04]: Spec dec is speculative decoding.Speculative Decoding and Custom SpeculatorsAli [00:04:05]: Speculative decoding, yeah.Swyx [00:04:07]: You have to explain.Ali [00:04:07]: Sorry. Like, speculative decoding is like, if you have a huge model, right? And so the model is going to be generating one token at a time every single turn, every single forward pass. So we attach, like, this little, like, parasite, like this layer that goes on top of the model, and this model just has to predict. It does three very fast autoregressive forward passes, and it will predict, like, three certain tokens, and then you do one forward stage over the entire original model in order to see if those predictions were correct or not, and then you accept them or you reject them. Now, this draft model is traffic specific, so if you, like, Philip said, if you're summarizing Harry Potter books, I can train exclusively that draft model on Harry Potter books, and I can guarantee you that I'm gonna accept the three tokens every single time. And so with that case, I increase your decode speed. I wouldn't be able to provide this to you if you're a shared endpointSwyx [00:04:53]: YeahAli [00:04:53]: ‘cause I have no idea if you're doing Harry Potter, if you're doing coding, if you're doing English. We don't know. Also, there was a thing in the book that mentioned that if they really cared about a specific threshold, chapter four, I think. Do you remember that?Philip [00:05:06]: Yeah. The things that you can do is you can set a specific, like, batch sizing, a specific, like, parallelism strategy if you're trying to optimize for, like, throughput versus latency. You can. Maybe a NVFP4 quant doesn't pass your benchmarks and you wanna run a model at higher precision, you could do that. There's just a bunch of reasons why you might wanna have your own endpoint and the biggest one, of course, just being, like, you don't have to deal with someone else doing a hundred million tokens of benchmarking traffic at the endpoint when you happen to be trying to serve your users.Swyx [00:05:40]: Yeah. I think one thing that is. That is a classic journey. Like, it's people is asking the, what happens when you type Google into the browser. Tool calling, is that just, you're generating JSON or is there more complication beyond that?Tool Calling, JSON, and Structured OutputsAli [00:05:58]: Certain customers that we have, they have their own post-trained models, and so they demand a tool calling that's not just, like parse a file or go find the weather. It's something that's very specific and you have to do post-training on this. And if the post-training on the model is not good or if the quantization after the post-training to get the inference to be fast, the model will struggle reading the JSON file and reading the tool calling. But it doesn't require its own like sandbox. It's not like it's going to use that tool calling to like escape a sandbox or like it doesn't have to be contained. It can just be a normal dedicated deployment. The challenge with tool calling more and more seems to be that the companies want certain tool calling which is a very sensitive thing to train. And because you're dealing with all of the JSON outputs, if it doesn't like close the end of the request in a very certain manner, you end up with a model that did the tool calling and like the thinking and so as a result of that, it didn't see the result and just hallucinated the result as it decoded. That seems to be the most challenging thing with tool calling, not really the sandboxes model.Philip [00:06:56]: Yeah, that's a challenge on the training side and then on the inference side, there's work that you can do to scope the possible output. So we published this at this point close to two years ago, the solution to this problem which is you make a state machine and you use that to constrain the output to a specific format. So this is the structured output problem. If you remember backSwyx [00:07:27]: Yeah, the specific grammar is,Philip [00:07:29]: Yeah, exactlySwyx [00:07:30]: GML had this thing.Philip [00:07:31]: Yeah. So it's like the old-school “make sure this is only JSON”, return only JSON orSwyx [00:07:38]: YeahPhilip [00:07:38]: Grandma's gonna die type of prompts.Swyx [00:07:39]: Is it BNF grammar? At some point OpenAI had released a thing that was like, yeah, if you want to constrain your output, write BNF grammar, back as NOR.Philip [00:07:47]: In our inference system, it's just a specified output format. And you get the guarantee that your output's gonna be structured along that format. And so applying that to tool calls can like help cut down on. You can still call the wrong tool or call no tool. It doesn't solve the certainty problem but it at least solves the output structuring problemSwyx [00:08:10]: YeahPhilip [00:08:10]: Within tool calls.Swyx [00:08:12]: And MCP is just another form of tool, right.Philip [00:08:14]: Yeah, exactly.Swyx [00:08:15]: As far as there's no special thing there.Philip [00:08:16]: The thing I'm always like explaining to people is the LLM is not capable of doing anything. It's only capable of making suggestions of what to do and then if those suggestions are formatted in a certain way and applied to a system that knows what to do with them, then an action occurs.Swyx [00:08:32]: Yeah. Part of the fun stuff is, this is solved outside of tool calling too. Like in an agent loop if the output is not correct or you're right, like reasoning, tool calling was done in the reasoning trace, just be like, “Oh, I don't know what to do. Let me just try again.” And it might get there after a few tries. And on your point of training, sometimes this is harder in smaller models, so you don't have the same exact quality outputAli [00:08:56]: Right.Swyx [00:08:57]: When you just swap from a big model, right?Ali [00:08:59]: Yeah. I will say that, before, I think we need to go back to inference engineering proper.Ali [00:09:04]: But, I had expected that something would replace JSON because it's hard to stream JSON ‘cause JSON must be complete and you must have open and close brackets and everything. So it's hard to parse something or validate something while it's being streamed. So people invented all sorts of things that are like, I forget the name of some of these alternatives, but it's something like TOML, something like YAML. But JSON seems to be dominant still.Philip [00:09:30]: The JSON outputs aren't that long, right? Like you could have a long-- ‘cause tool calls also contain the arguments in them and perhaps for a certain tool you might pass like a very long argument. But my impression of the median tool call is that it's a relatively small number of tokens, right? So I would expect that speculators are generally fairly good at something as formatted as JSON. And so you would have like a pretty fast decode step there and that the streaming wouldn't be as valuable, but maybe I'm wrong about that.Ali [00:10:02]: I think you're also bounded by the software or that the model is gonna integrate with if the software is built with JSON for the tool calls or if the company that you'- if your customer says that this is how our software works and our tools are interfaced with JSON, you can ask them to like, change their software and say like, “Yeah, this is gonna be better for the model.” but like with the right training shouldn't be that much of a difference. Also more profitable if it outputs more tokens probably.Swyx [00:10:25]: Depends on your business model.Swyx [00:10:27]: It really depends. But I will say that, as a writer with like experience a lot with generated output, I do try to move from text to JSON text which is very long JSON, right? Like there's paragraphs in every field because I'm trying to structure it, right?Philip [00:10:44]: Right.Swyx [00:10:44]: I want you to first make factual statements, then make opinions then make bullet point summaries, have dates, have entity references have your sources for references, all these things. Anyway, so these are things that like I think people who really experiment with structural output have to really care about. But, let's, let's recurse up the stack a little bit. Before we started recording, you mentioned something really cool, which is that there's a lot of engineering that-- inference engineering that goes on when a new model provider releases a new model, right? So let's call it GLM-5.2, Kimi K3. I had previously assumed, especially if it's like, well, GLM 5 to 5.1 to GLM-5.2, like that you've supported them before. Is it that much work?What It Takes to Support a New Open ModelAli [00:11:26]: It's a lot of work.Swyx [00:11:28]: Yeah. Okay. So like, a lot of people, all you guys, right whenever a new model launch like, people rush to say like, “Oh, Hugging Face supports this, Fireworks supports this, Spacetime supports this,” and I'm like, “Yeah, of course we support it.” But what goes into that? What goes intoPhilip [00:11:40]: I think it's more than just support it too, right? It benefits the consumer a lot. Like I think it was with Kimi K2.5 or GLM-5.2 the latest, there was an inference war, right? X provider is at 90 tokens a second. The next day we're at 150. The nextSwyx [00:11:55]: I kinda kicked that off with the GLM-5.2.Swyx [00:11:58]: I wrote a Twitter article about. It got like half a million views,Ali [00:12:02]: Based on being numberSwyx [00:12:03]: YeahAli [00:12:04]: Or it's for something else.Swyx [00:12:05]: Yeah. Which,Ali [00:12:06]: Oh my GodSwyx [00:12:07]: Which then got everyone really excited about, hey, how can we, bend tracks a little bit further and,Philip [00:12:14]: There's a difference between support the model, as in I can make a token out of this model, and support a model, as in I have a production-ready API from this model.Philip [00:12:26]: Getting to the point of I can make a token out of this model is not that hard because generally the, open source inference engines, vLLM, SGLang of the world oftentimes even receive weights ahead of time, maintainers do, or the people making the model merge PRs to ensure support. So you generally can, just get it working on the standard open source stack without too much pain in most cases. The challenge is, every inference company is gonna have own proprietary stack. Some open source components, some in-house stuff. And for any arbitrary model, there's going to be some new stuff. Sometimes you get lucky, like K, two five to two six was, like, pretty similar.Quantization, Speculators, and Production ReadinessAli [00:13:16]: Yeah. It was pure continued post-trainingPhilip [00:13:18]: YeahAli [00:13:18]: If I remember correctly.Philip [00:13:19]: Even in those cases, there's still stuff you have to do. You have to redo the quantization work. You're taking the model from. Generally, these models are not released in NVFP4, and we want them to be in NVFP4 for maximum Blackwell compatibility. So we have to perform that quantization, and, calibrate the quantization to make sure that we're not causing any regression in the model's intelligence. And then we also have to train the speculator, as we've talked about. Generally, we have. We have ZDR, zero data retention on our model APIs, so we don't know exactly the traffic that people are sending us, but we know what's popular. We know that coding use cases are popular. We know that agents, agentic use cases are popular. So we can get public data sets that are representative of that traffic and train general speculators. Now, with speculators today, you need to train the speculator using the base model itself because you're getting hidden states out of the model from running inference on these specific prompts, and that is the training data you use to create the speculator. So there's that process which you need the real model weights for. And then there's of course just the process of, standing up all the infrastructure behind it, loading all this stuff, testing it. And then when there's a new model with a newer architecture, I think that, like, the DeepSeek models tend to be the most challenging as they have, like, the most novel architectural stuff going on, model after model. But every new model has something. Kimi K2 had. Oh, sorry, GLM-5.2 hadAli [00:14:53]: Sparse attention.Philip [00:14:54]: Yeah,Ali [00:14:54]: YeahPhilip [00:14:54]: the DSA.Ali [00:14:55]: Right. Which is brought from DeepSeek.Philip [00:14:57]: Yeah. AndAli [00:14:59]: So you can copy-paste then?Philip [00:15:01]: It kindAli [00:15:01]: I don't know how this works.Philip [00:15:02]: So, like we had to, like, build support for that into our runtime. And you're right, like it is really interesting the way that all of these open source labs borrow from each other. For example, like GLM-5.2 doesn't have vision. So something that, Haley, a guy on our team, if we could take a look at this, he, like, grafted the Kimi vision encoder onto GLM-5.2.Retrofitting Vision into GLM-5.2Ali [00:15:27]: We'll be training the projector.Philip [00:15:28]: Exactly. So if you think about, like, the encoder, there's the encoder, which is the part that looks at the image and turns it into latent information, and then there's the projector which likeAli [00:15:38]: You can say latent space. It's okay.Philip [00:15:41]: And then there's the projector that maps it onto, the model itself, and then there's the model weights. You don't wanna mess with the model weights because you run a chance of making the model dumber at something else for the purpose of giving it vision. So instead, Haley started with just a projector, which is only a handful of millions of parameters.Ali [00:16:02]: That would be, yeah.Philip [00:16:02]: Yeah.Ali [00:16:03]: Can you show the training one?Ali [00:16:04]: Like the way it groksPhilip [00:16:05]: YeahAli [00:16:06]: Very interesting.Philip [00:16:06]: And maybeAli [00:16:07]: That right therePhilip [00:16:07]: Maybe Ali, you should take it from here. You've got a betterAli [00:16:10]: Ooh, double the sandPhilip [00:16:11]: Understanding of this than I do.Ali [00:16:11]: Yeah. You can see, like, he. The way he trained this is really cool. At the beginning, he was training it using just like, “Here's a picture of a mountain. Can you describe what's in this mountain?” And that caused it just like the first, learning walls. Like here you can see this all we're trying to teach it is to translate the encoded. Like it's already taken the encoder from Kimi K. It's taken the image. It'Philip [00:16:31]: Yeah. FrozenAli [00:16:31]: FrozenPhilip [00:16:32]: With adapter.Ali [00:16:32]: Exactly.Philip [00:16:33]: Yeah.Ali [00:16:33]: So the brain is frozen and the eyes are frozen. It's just we're tryingPhilip [00:16:37]: AlignAli [00:16:38]: Interconnect between the eye and the brain, right? So the projector. And so you take the tokens and then he's like, “Oh, can you describe what's in this image?” And he's like, “Oh, it's a mountain,” or it's a person or it's a human, whatever the case is. But that didn't cause complete understanding. So he changed it such that every image was associated with a data set of questions. Like, does this image have a white male? Does this image have birds in the top corner? Does this image have a scientist in it? All of that stuff. And it would have to answer questions correctly. And using not just training on describing an image, but being able to answer question, another question, answer over time. Like you can see the grokking, which is like genuinely insane, that retrofitting vision into a large LLM can learn to that extent. And even for images that it doesn't perform well on, for instance, if you ask it a picture of like Stephen Hawking, “Who is this?” Maybe it doesn't get it, but it will say something like, “This is Albert Einstein.” Like it still understandsPhilip [00:17:25]: Close enoughAli [00:17:26]: That this is a scientist who is a man who has, some significant achievements, all that stuff. So that's like really cool.Philip [00:17:32]: Yeah. So, we've covered Hao Tian before, who the author of the LLaVA paper that did this, a while ago. And I think that's very foundational work for anyone who hasn't done vision work before.Ali [00:17:41]: Same with the CLIP and MetaCLIP, where you go from just captioning to building out questionsPhilip [00:17:47]: RightAli [00:17:47]: Off the image and how much better you can get performance.Philip [00:17:50]: Right. Right. Right. Yeah. But what's, what's so exciting about this is if you look at a model like this. Now, this is a little bit more of a research project. It's not. It got to 56% on MMLU Pro, I think. So not quite frontier. But if you're running this model, you haven't suffered any loss on your GLM-5.2 quality. If you don't have an image, it'll just behave exactly the way it used to. And ultimatelyAli [00:18:14]: Which in the inference code you literally do not include the other part, right?Philip [00:18:18]: Yeah. You would just skip the encoder if you don't have an image input.Ali [00:18:22]: Okay.Philip [00:18:22]: Just confirming.Philip [00:18:23]: YeahAli [00:18:23]: Does it affect a lot on the overall inference side? Like you're not adding much, you're adding a very small vision encoder. These are typically likePhilip [00:18:30]: They're super fineAli [00:18:31]: Less than a billion parameters, right?Philip [00:18:32]: Yeah. It's, - There's a little bit less standardization among vision encodersSwyx [00:18:37]: YeahPhilip [00:18:37]: So the support matrix can be a little bit, sparser. But overall, yeah, it's a pretty, it's a pretty minor component of the overall system. And ultimately what you get out of the system is all of a sudden you have Kimi Vision, GLM weights, and DeepSeek attention all in one model.Open Source Model Grafting and Franken-MergesPhilip [00:18:56]: And that's, I think, a lot of the power and beauty of open source, is that you can take all of these different components and combine them together into a system that's better than anyoneSwyx [00:19:05]: YeahPhilip [00:19:05]: Can be individually.Swyx [00:19:06]: People used to say that you would also do Franken-merges where you would take likePhilip [00:19:10]: YeahSwyx [00:19:10]: Layers from each model.Swyx [00:19:11]: Does anyone do that anymore?Ali [00:19:13]: Well, to your point previously when you were mentioning like, the work that goes into supporting a model when it first comes out, like GLM-5.2 or MiniMax M3 or whatever the case is. Sometimes you do have to like, you do have to switch out some things. Like, for instance, the MiniMax M3 head uses full attention, and with full attention you end up with this like insane bottleneck in spec dec ‘cause you're doing auto-regressive token generation for three tokens, and you're doing this like N squared over all of the tokens that are in your sequence. Your KV cache is like very large because it's not sparse, it's not top K. So we find it better to like, okay, we're gonna replace this, we're gonna replace this layer with a layer from another model that's using like GQA, for instance. And then just with the right training, you can get it to have the same acceptance rate. So it is very possible to retrofit layers from other models and very much needed. If a layer is like inefficient, the training just becomes the challenge, like how do you ensure that you train it properly? Which again to your earlier point is like the mesh between training and inference. As in like you need very good training in order to do fast inference. That's like, I feel like more and more becoming true.Swyx [00:20:21]: Yeah. Anything else on the support side when you say like get it to fully production ready?Loop Detection, Race Conditions, and Non-DeterminismPhilip [00:20:26]: Yeah. I think that there's also a question of just, we can test a model to a pretty extensive degree, but we're trying to get it out quickly and then you see a bunch of other people test it and you get interesting results. There was an issue with, GLM briefly where we had some like mode collapses where it would just output the same token over and over again for certain prompts on certain temperatures. Like once you expose an endpoint to the real world, there's going to be, so many more varieties of things given to it that you're able to, discover and patch things. So it's not just a, day zero process, it's then like for the first week, for the first month, if a model remains popular, like how do you both fix bugs and then continue to push the envelope on performance?Ali [00:21:21]: What do you mean you don't want your model outputting S?Swyx [00:21:24]: Is there loop detection on that stuff, by the way? It still happens like quite a lot, which is surprising.Ali [00:21:30]: We have like we, in our endpoint, like if a model was to output the same token like four plus times, we just cut the generation. We say like, “Oh, sorry, this-- Like try again,” or like we will reprocess the request. ‘Cause we know then, like if it, like if, yeah, it's four times the same token, it's probably collapsed.Swyx [00:21:45]: Yeah. Is there a way to opt out in case I really want that?Ali [00:21:48]: You want that?Ali [00:21:50]: I think there's a way that we have to handle it. I'm not exactly certain, but I feel like in certain models, like when they output something like you can imagine, like a table for instance, and so they want, they wanna draw like 12 dashes and 12 dashes. Yeah, I think there's a way for that to happen. I think we only do it on certain tokens. Like we exclude certain special characters.Swyx [00:22:07]: Yeah.Ali [00:22:07]: So we only do it on like certain like S is the most common almost. GLM-5.2Swyx [00:22:11]: OhAli [00:22:11]: And I think it was DSV 4 as well. Like you'd just have like looping issues where like you literallySwyx [00:22:17]: ItAli [00:22:17]: Just have like S.Swyx [00:22:18]: Yeah. Is there a special, something special about S? No, just randomlyAli [00:22:21]: It just seems to be the one token involved.Swyx [00:22:23]: Yeah. And it'Philip [00:22:24]: Is thereSwyx [00:22:24]: And it's only temperature 0Ali [00:22:27]: NoSwyx [00:22:27]: Even at other temperaturesAli [00:22:27]: Even at like 0.9 or whatever, it will still, it will still collapse.Swyx [00:22:30]: That's weird, right?Ali [00:22:30]: It's, it is an inference problem to be honest, like a software problem. Like oftentimes, the image you run will-- like NVIDIA will release an image for instance, and if we will upstream the changes from their latest TensorRT-LLM image into our stack, we'll find that it fixes it. Or oftentimes this will only happen in an inference engine that you're using like SGLang. But if you were to switch to vLLM, that isn't the case. So it seems to be like an extremely like deterministic software issue and not really a model issue. It's not like a weights problem. Like I'- we'll say like, “Oh, it's a problem with the quant. We did PTQ wrong,” right? But that isn't, that doesn't make sense because the same weights used with a different inference engine does not repeat the problem. And sometimes it's, the kernels that are being used in the backend have like these very subtle sometimes race conditions, where if you were to use this model hosted on one cluster, you will never get this problem.Swyx [00:23:19]: Oh my God.Ali [00:23:19]: But if you host it on a different cluster, you will. And the reason is the KV cache transfer from a node to node in that one cluster is using a slower interconnect than the node to node in another cluster. So that exposes the race, whereas in another cluster it doesn't. So then you end up just like, okay, this model is not gonna be hosted on this cluster. We're gonna host it on, another cluster because that cluster exposed that problem. But then it ends up with like, okay, is it the software? Is it the model weights or is it the hardware?Swyx [00:23:42]: There is a thing about this with temperature 0 still not being deterministic, right?Ali [00:23:46]: Right.Swyx [00:23:46]: Mostly because of hardware. Even at temperature 0 same model, you won't always get the same output.Swyx [00:23:52]: Even-- But I'm surprised by the race condition one because, I thought PyTorch was a graph that like guarantees that you at least, execute things in the right order.Ali [00:24:02]: Well, yeah, true. Like I'm not, I'm not saying that there is. Like well, you have things like PTL optimizations where like you can start a kernel before the end of the previous kernel, and that's like ‘cause you want to do that because there'sSwyx [00:24:12]: It's like pipeliningAli [00:24:12]: Expense. Exactly.Swyx [00:24:13]: Yeah.Ali [00:24:13]: But it'- But you don't do it cleanly. Like you overlap a little bit of the execution. No, it is very possible that the kernel itself, like that one block that is supposed to be running in this instance of time, that kernel itself has a race condition. For instance, like a missing barrier. Like often if you're designing a kernel and you want it to make it to be very fast, if you don't test it extensively, you'll, you'll have certain threads access data points from registers before they've been written to by other threadsSwyx [00:24:36]: YeahAli [00:24:36]: For example, because like your barrier is wrong or your synchronization was wrong. But yeah, like the testing itself is very difficult in those like, andSwyx [00:24:42]: And there's no like borrow checkerAli [00:24:45]: What does that mean?Swyx [00:24:46]: Like Rust. Like the. If you're trying to have like memory safety It sounds like a comparable problem.Ali [00:24:52]: Well, yes, but you're working in CUDA, right, NVIDIA GPUs. Like- You just need a higher level language like modular Maybe that's what modular is supposed to do. I don't know.Quantization Quality and Vendor FidelityVibhu [00:25:00]: How do you see keeping quality of the model? So you talked about all these steps of, okay, you gotta do quantization, train your own speculative decoderAli [00:25:07]: RightVibhu [00:25:07]: Run on different hardware. Looking at other model providers, okay, you kicked off a inference speed race on the consumer end. What goes into keeping quality the same across them, right? Sure, you can run benchmarksAli [00:25:22]: YeahVibhu [00:25:22]: But, like, how do you determine how much quantization are there standards? What goes intoPhilip [00:25:27]: There's a few things on quality. Most inference optimizations are lossless. KV caching, for example. You are just recomputing or preventing recomputing the same values. Speculation, of course, if a draft token is wrong, it gets rejected. The main lossy optimization is quantization. And that really comes down to, number one, data format, number two, which parts of the model you choose to quantize, which layers, and number three, like doing a lot of calibration on the quantized weights, to ensure that you're preserving all the outliers. There's other tricks that you can do, though. A big one is long context, ‘cause one thing you asked at, right at the beginning is, “Oh, what's gonna happen if I send a 200,000 token request in?” So with a long input sequence, you need to, store a lot more information. You need to process a lot more tokens. And so even if a model has a context of a certain length, you might, as an inference provider, choose to build an API with a shorter context length, and of course a full length one as well. Because if someone doesn't need the full million token context, for example, you can get them better performance. I don't know if that's exactly like quality of the model. The way that I think about quality is to what degree are we faithfully serving the original model? If you think of a golden implementation of a model that performs exactly the way the model is designed to perform, I think of quality as how close are we getting to that, 100% fidelity of the model.Philip [00:27:13]: You can also, of course, think about quality from the training side and how do you push yourself past 100%. But when I think about purely inference optimizations, it's getting faster while staying as close to that 100% fidelity mark as possible. And certainly our standard internally is that, like you should not be able to tell the difference between our API and a, official API. I think Kimi in particular does a good job of vendor benchmarking hereAli [00:27:41]: YesPhilip [00:27:41]: Where they haveAli [00:27:42]: They released an actual vendor benchmark.Philip [00:27:43]: Exactly, yeah.Ali [00:27:44]: ‘Cause they accused, some people, Amazon? There was some provider that was not doing very well on Kimi's benchmark.Philip [00:27:50]: Yeah.Philip [00:27:51]: So, with Reflect we probablyVibhu [00:27:52]: This was a long time ago, right?Philip [00:27:54]: No.Ali [00:27:54]: Yeah, like threeVibhu [00:27:55]: They alsoAli [00:27:55]: Four, five months agoVibhu [00:27:57]: This also happened with, I don't remember which model, but they pulled out quite a few, and then they started a whole chart about this. It might have beenPhilip [00:28:03]: Kimi Vendor Verifier.Ali [00:28:04]: Yeah.Philip [00:28:05]: Yeah.Ali [00:28:05]: Yeah, ‘cause you, ‘cause you'd be pissed, right? Like if you'Philip [00:28:07]: Yeah.Ali [00:28:07]: If like if I'm a consumer and I'm using like Amazon's endpoint for instance, and I've used Kimi and I'm like, “Oh my God, like this is bad,” I'm not gonna say, “Oh, Amazon quantized the model in a bad way.” I'm gonna say, “Oh, Kimi sucks.” Right?Philip [00:28:17]: Yeah.Ali [00:28:17]: So it seems like that makes sense.Philip [00:28:19]: Yeah, they care. They care.Vibhu [00:28:21]: Justifiably.Ali [00:28:21]: Yeah, justifiably.Vibhu [00:28:22]: This is probably a stupid question, but just checking, has anything improved from main quantization?Philip [00:28:28]: Yeah.Vibhu [00:28:28]: Like, is quantization always strictly worse?Ali [00:28:30]: Well technicallyVibhu [00:28:32]: NoAli [00:28:32]: It's a lossy. QuantizationPhilip [00:28:33]: YeahAli [00:28:33]: Is a lossy, it's a lossy implementation.Philip [00:28:36]: Speed improvesVibhu [00:28:36]: Speed improves.Ali [00:28:37]: It the number, likeVibhu [00:28:38]: No, I' always look for inverse scaling laws.Philip [00:28:40]: Yeah.Ali [00:28:40]: Yeah.Vibhu [00:28:40]: This is something I learned from Noam Brown, where like things that normally act in one direction sometimes do.Philip [00:28:45]: Well, technically when you run a benchmark, because these models are deterministic, sometimes your,Ali [00:28:52]: YeahPhilip [00:28:52]: NVFP4 quant is like, two basis points higher than yourAli [00:28:56]: No, it's noise. It's noise.Philip [00:28:57]: Yeah, exactly. I'm like, yeah, it's, it's within. That's why I always say within margin of error.Philip [00:29:01]: And I stopped saying that because everyone assumes that what is, well, within some margin of error, we're barely inside of that to the worst, so we're saying. But yeah, sometimes it's just like, gives you a higher output score. But like Ali said, that's noise. To my knowledge, you're not necessarily making the results better. You're just trying to, again, like keep your fidelity as close to 100% to the original model.Layer Selection, KL Divergence, and Better QuantizationAli [00:29:27]: There is, to your point, research that we did on MP. I don't know if you are able to pullPhilip [00:29:31]: YeahAli [00:29:32]: A tweet we did. One of our research interns, Joshua, I think it's a tweet on how we have 20% better quantized GLM-5.2 than NVIDIA. Essentially what we found throughout like this month research is, okay, quantization is a lossy. It's. You're compressing the data from, occupying 16 bits to occupying, four bits, for instance. And so you're losing some information, and you're trying to minimize that. And so when I say that I'm gonna quantize the model, my job becomes how do I find the layers that I can quantize, and how to find the layers to not. For instance, with image models, I don't quantize modulation layers, and I don't quantize out projections because those two are. Like out projection is what you see as the user. Modulation is what the model sees or understands. Right, exactly. And so to his paper, do you have the. It doesn't have the. Yeah. It's a long paper. I don't know if I can findVibhu [00:30:25]: If there's a part to search or it's probably in the thread.Ali [00:30:28]: It's probably in the thread.Vibhu [00:30:29]: Yeah.Ali [00:30:29]: But the long and the short is it is very possible that quantizing more of the model makes the results. Like if I have a model that I quantize layers one, five, and 10, and another model where I only quantize layers one and It is possible that the model in which I quantized more information is going to perform better because the quantization errors have canceled out. And so what Joshua showed in his mathematical proof where he had like a verifier in, is that you can predict which layers are going to have quantization errors that will cancel out with each other, and you choose to quantize those layers. And so the result of doing this mathematical quantization is you end up with a model that's 20% more quantized than another provider, so you get 20% more throughput of it because there's more layers than running an NVFP4, and your quality is better than that other quant because the layers that you chose to quantize have their errors cancel out, like one layer skewed to the right one layer skewed to the left, one layer skewed to the right. Your final logits distribution is more similar to the original distribution of the model, so you have better fidelity. And so the way we proved this was with KL divergence. So instead of just scoring on the benchmarks, we scored the KL divergence between the logit distribution of the quantized model and the logit distribution of the original full precision model, and we showed that with this technique we get. If your probability distribution on the logits which token it wants to select is more of the same as the original model, you're probably gonna end up staying true to the original model. So yeah, so it seems like previously before this, it seemed like the industry was, well, the more you quantize, the worse it's gonna be, ‘cause the more loss you introduce. That's not exactly, not necessarily true. So yeah, doesn't improve it, but can cancel out.Philip [00:31:57]: I think it might be this, but reminds me a good bit about pruning where you can prune off certain layers.Philip [00:32:03]: But very interesting. Didn't know this was a whole paper you guys put out.Ali [00:32:06]: It's. Fun fact, it was originally 72 pages, this paper, and then we decidedPhilip [00:32:11]: WowAli [00:32:11]: We can't tell. We couldn't release it. So it's now 45.Swyx [00:32:15]: Still 39 pages, so very substantive. We talked about evals and all these things and, like what's possible in terms of speedup? Like it's like probably like the numberInference Speedups and BenchmarkingSwyx [00:32:25]: Thing that people do wanna care about, and it's something that you wrote about in your post. Like official API is 70 tokens per second, and you push it up to 90. Is that like a normal thing?Philip [00:32:36]: So what's cool about working in inference, the reason that I think inference is going to be a useful place to do engineering for a long time, is that if you look at highly optimized domains like, say, finance, if you're in finance, you measure how much better you got in basis points. It's like, “Oh, I got five basis points better, like twentieth of 1% better,” that's huge news because everything is so optimized. When we publish optimizations, it's 20%, it's 100% it's 200%. So there's still probably like a lot further to go, honestly. Like you'll, you'll know that inference is pretty much solved when researchers start publishing about how they got 1% faster at something.Swyx [00:33:19]: Which by the way, because I am from the finance background, in the ‘70s, that was the margin at the time. When you did quantitative finance research, you would findAli [00:33:27]: And like 20%, tens of percent.Swyx [00:33:29]: That's. Yes.Philip [00:33:29]: Yeah.Swyx [00:33:30]: And now it'Philip [00:33:31]: Tiny fractionsSwyx [00:33:32]: For those people interested, look up Andrew Lo's paper. He had a really interesting illustration of quant, stat arb, distribution, narrowing down from like those kinds of 20% differences in the ‘70s, down to nothing today, which is very cool.Philip [00:33:48]: Exactly, and we're at the beginning of the same type of thing. Now benchmarking is hard. I think anyone will tell you that, and benchmarking provider speeds is hard because there's so many variables that go into it. What hardware are you using? How much load do you have on the system? What's the exact nature of the prompts and input and output sequence lengths? All that stuff. But overall, when you start stacking these improvements, you're looking at multiples. You can look at it. The most common form, of course, is TPS, tokens per second, which is bad naming by us in the industry, ‘cause there's two tokens per second. There's tokens per second, the throughput number, and the latency number.Ali [00:34:31]: TTMT, yeah.Philip [00:34:32]: Like total tokens per second out of the, out of the GPU as a throughput number. Most people only care about tokens per second as the latency number, which we should call ITL, intertoken latency, but we don't.Philip [00:34:44]: Anyway, so you can imagine a standard API without many optimizations for a 1 trillion parameter model operating somewhere in the 30 to 50 tokens per second range for reasonable traffic profile. And we generally see the goal of, pushing to 10X that. But, not necessarily day zero, but by stacking enough optimizations, if you have, say like four optimizations, each of which doubles performance. Or sorry, three optimizations, each of which doubles performance, then you stack that up, that's an 8X gain. That's the order of magnitude that we're working with in this space. We're trying to make things substantially faster, not just go from like 70 to 90.Swyx [00:35:38]: Are you saying you've. You have done that?Philip [00:35:40]: So let's say you have as a reasonable baseline, 30 or 40 tokens per second. You can achieve 10X that. So like on GLM-5.2, if you run it unquantized, perhaps on H100s even, and you're just using an off-the-shelf inference engine with no particular optimizations, no speculator, nothing extra around like KV routing, no disaggregation, you're, you're probably, yeah, looking at that like 30 to 40. You think that's like a reasonable baseline?Swyx [00:36:12]: Right. Right.Philip [00:36:12]: To get to something like 10X, there's a lot of trade-offs that you're making. If we're running at more like a 300, 400 tokens per second range, you are using the best hardware possible. You have a optimized speculator. You have done all of your quantization work. You are Seeing a pretty high cache hit rate. You are running with a reasonably small batch size and a parallelism configuration that is tuned for latency versus throughput, but it is possible. So the spreads that you see if you, like, go on artificial analysis or you go on OpenRouter and you look at, the worst provider to the best provider, oftentimes can hit that range. 10X is of course very aggressive. It's oftentimes maybe more of a four to six times improvement. But that's the performance that makes us really excited, is when we can get these huge gains, not just go from 70 to 90 tokens.Stacking Optimizations: NVFP4, Speculation, and DisaggregationAli [00:37:19]: It's also, like, hardware dependent. Like, ifPhilip [00:37:20]: YeahAli [00:37:20]: If you have a thing where you're serving it on just, like, a node of H100s and then you throw, like, you shard the model across, like, four nodes of B200s. Like, you can definitely increase the speed with just throwing more hardware at it. Like, normalizing for the same exact hardware and the same number of GPUs.Philip [00:37:35]: Yeah. Then you're looking at, like, a two to 4X improvementAli [00:37:38]: Right. RightPhilip [00:37:38]: Depending on the inference optimizations. So yeah, it's. Some of it's, what's the call, and some of it's who's the driver.Vibhu [00:37:46]: If you break down the two to 4X, say the example is run GLM-5.2Ali [00:37:51]: YeahVibhu [00:37:51]: On B200sAli [00:37:53]: YeahVibhu [00:37:53]: Single node, right? What's, like, the cost trade-off for effort to get, like, the last bit of juice out versus what should people just think of, right?Ali [00:38:01]: Spectre quantization. Yeah.Vibhu [00:38:03]: Spectre quantization.Ali [00:38:04]: That's, that's, that's like 95%. LikeVibhu [00:38:06]: And how far does that get you? And how easy is that for the average person to do? So say right I wanna throw the weights of GLM-5.2 on a node of B200s, how easy is it to find speculative decoder- decoder model or already quantized model? How much work goes into it?Philip [00:38:23]: If you're doing it up front, it's quite a lot of work. If you're doing it today, there's going to be people who have published things that you can just, you can just grab some NVFP4 weights. You can grab a speculator. Yeah, if we're thinking about, like, what are the 2Xs we're stacking, going from, BF16 to NVFP4 is, it's not quite a 2X, right? It's like. I think it's about, like, 30 to 40%, from 16 to 8, and then another 30 to 40% multiplied from, 8 to 4. So that doesn't quite get you a 2X, but, like, roughly a 2X. Speculator, roughly a 2X. Disagg on top of that if you're able to get enough hardware and put enough traffic through it, another roughly a 2X. And then you add in some, double-digit percent increase from having just a better runtime with, the latest kernels and stuff behind it. And that's how it stacks up.Ali [00:39:21]: YeahPhilip [00:39:21]: So building each of those, like, building the, quantized weights is, for someone who really knows what they're doing, hours to days of work. Building the speculator, again, like, hours to days of work. And the, disagg setup, hours to days. Well okay, but like once you haveAli [00:39:39]: Once set up. Once set up. YeahPhilip [00:39:40]: Yeah, getting disagg working for the first time, I'm saying, of course, is very difficult.Philip [00:39:44]: The marginal implementationAli [00:39:48]: Like, if you're just grabbing, like if you are a person, like just a normal consumer who has access to, like, a node of B200s and you're wondering, “How can I just host it myself?” You don't need to quantize the model yourself. There's always gonna be, like, an open source quantized checkpoint. NVIDIA's gonna push one out if no one else does. You. Usually, the providers will have their own spec dec that they've trained as well. You don't need to train your own spec dec. You can just use that as well.Philip [00:40:09]: Yeah. Like, GLM-5.2 has its own MTP.Ali [00:40:13]: Right. Right.Vibhu [00:40:14]: What's multi token prediction?Philip [00:40:15]: Yes.Ali [00:40:16]: I'm justVibhu [00:40:16]: Can you explain that?Ali [00:40:16]: I'm just an expert.Ali [00:40:18]: I can do it for you in case I get it wrong?Vibhu [00:40:20]: No.Vibhu [00:40:21]: Yeah, you should correct if we're wrong, but their multi-token prediction can be used for self-speculative decoding.Ali [00:40:27]: I'm not sure. I'm not gonna correct that.Vibhu [00:40:28]: Okay. I'm semi-confident in thatAli [00:40:30]: Okay. YeahVibhu [00:40:30]: But someone can check. But it's useful to paint the story of, okay, not just the average person, but say a company wants to switch from serverless inference I wanna throw this up on. I wanna rent some GPUs, throw it up. These are the steps you take to do significantly faster than just put it behind vLLM.Ali [00:40:48]: Right.Vibhu [00:40:49]: I was waiting for a mention of Dynamo.Vibhu [00:40:51]: I feel like, that's supposed to be the baseline that you measure against.Dynamo, KV Routing, and Disaggregation ToolkitsPhilip [00:40:55]: I would think of Dynamo as less of a box system and more of a toolkit for building with. So when we talk about doing aware routing, when we talk about doing KV offloading, when we talk about doing, PD disaggregation, Dynamo fundamentally is. By the way, Dynamo is an open source library from NVIDIA.Ali [00:41:17]: We've done a pod with KylePhilip [00:41:18]: OkayAli [00:41:19]: Kyle Cranin.Philip [00:41:19]: Cool. So then your listeners know then that it supports all the different inference frameworks. And it is multi hardware, which is interesting.Ali [00:41:28]: But it's just a router, it's not like an optimizer layer.Philip [00:41:30]: Yeah. All it does, like, what Dynamo is good at, it is a library for moving information around your cluster, around your hardware. So if you have, KV cache on one place and you need it to be somewhere else, Dynamo coordinates NIXL for you to move that around.Philip [00:41:49]: That doesn't mean that, like, out of the box, you just say, “Pip install Dynamo,” and then you get, like, a massive performance speed up. It's more of a developer toolkit.Ali [00:42:01]: Yeah. I would have said it would. It comes with a set of defaults that you can then swap out.Philip [00:42:06]: It does. If the industry at large, I think, was, like, rolling out all of these deployments, standard, then I think it would be, like, a credible baseline. But, we've got to, we've got to benchmark against, like, what we're seeing in the wild.Speculative Decoding Methods: Medusa, EAGLE, n-Gram, and Spec-SpecVibhu [00:42:23]: I did wanna talk a little bit more about PD disagg, because that is probably, like, number three after quantized and speculative decoding. In your book though, I was just gonna pull out the book.Philip [00:42:31]: Yeah.Vibhu [00:42:32]: Like section 522 on Medusa, 523 on EAGLEPhilip [00:42:35]: YeahVibhu [00:42:36]: 524 on gram.Philip [00:42:37]: It's 55, would be disaggregationAli [00:42:42]: Yeah. Well, no, I just wanted to dwell a little bitPhilip [00:42:44]: YeahAli [00:42:44]: The other. Like, so what do you choose to include? What do you choose to not to include? Because there was all these other techniques.Philip [00:42:51]: Yeah.Ali [00:42:51]: Are these still relevant? Because I think they came out, like, a year and a half ago maybe.Vibhu [00:42:55]: Medusa is quite old.Philip [00:42:56]: Yeah, Medusa's old.Ali [00:42:58]: It was old.Vibhu [00:42:58]: But is it in the book as a good, here'sPhilip [00:43:01]: BaselineVibhu [00:43:01]: Baseline vanilla understand it?Philip [00:43:02]: Like you should know this.Vibhu [00:43:03]: Like I read the paper, I'm like, “ it makes so much sense.”Philip [00:43:05]: Yeah.Philip [00:43:05]: So with the book, I had a couple goals. One was to give people just a working vocabulary for the space as a whole, and the other was to give them some intuition about how each of these techniques works. As I mentioned in my AI Engineer talk, which is the first public addendum to this, the speculation space has moved much faster than everything else. So yeah, even at the time that I wrote the book Medusa, I very much included as a way for people to understand how the space evolved rather than what the most modern technique is. And now of course, there's DFlash, dSpark. There's, there's newer techniques even than EAGLE, although EAGLE is still very commonly used.Ali [00:43:51]: SpecSpecta.Philip [00:43:52]: Yes. Speculative decoding.Vibhu [00:43:54]: What canAli [00:43:56]: Oh, it's a paper by Tri Dao and it's like, it's doing speculative decodingVibhu [00:44:00]: HuhAli [00:44:01]: For the speculative decoder.Philip [00:44:02]: Oh, in spec- oh my God.Ali [00:44:02]: It's literally just an another. It's like, yeah, that's the most simple way to explain it, and it seems like he got trivial speed ups there. But it seems that the complexity with training, it's almost like in our mind at least, it's almost as complex as training GANs. Like it's like a very delicate balance and oftentimes you, it's just but yeah, it's literally speculative decoding on speculative decoding.Vibhu [00:44:21]: Speculative.Ali [00:44:22]: Yeah. We saw this paper.Vibhu [00:44:24]: It's interesting, right?Ali [00:44:24]: Yeah.Vibhu [00:44:24]: I wouldn't even expect it to be very particular to train, I wouldAli [00:44:29]: Right.Vibhu [00:44:29]: The naive part of me is like, okay, train speculative decoder.Ali [00:44:32]: But like, and it makes sense, like the whole idea of speculative decoding is you. It's like, it's like almost like the iPhone auto predict version but for a normal model, right? Like you're just, you're just, generating three tokens and you're like, okay, I'll do prefill on them. And so you save those three turns for your original model. Now your speculative decoder is doing three turns of auto regression, so why not just have an even smaller model?Ali [00:44:53]: The other question there is what are the size of speculators? So say forPhilip [00:44:58]: Right. It's like a billion parameters.Ali [00:45:01]: Like for MiniMax, it's. Yeah. It's like one layer. It's like one 60th of the original model usually.Philip [00:45:06]: Yeah. I think we should do a paper when we get back to the office.Philip [00:45:10]: SpeculativeAli [00:45:11]: SpeculativePhilip [00:45:11]: Decoding.Ali [00:45:13]: No, it's, it does seem like how, when do you stop? But then it also seems like if you're able to train spec-spec decode for instance, right? Like if you're able to have a small model that is accurately predicts what the intermediate speculator is gonna predict, that is able to predict what the original target model's gonna predict, then why not just use that smallest model directly, right?Vibhu [00:45:34]: Yeah. This isAli [00:45:35]: Like it seems likeVibhu [00:45:35]: Adjacent to the routing problem.Ali [00:45:36]: Right.Vibhu [00:45:36]: Yeah.Ali [00:45:36]: Right.Philip [00:45:37]: The thing with speculators is one of the practical constraints on using them is that you do have to run a small model on the same hardware that you're running the big model on. There is a orchestration and resource competition problem inherent in that, and that is one of the constraints on speculation in general, is that draft tokens cost resources to create and cost software complexity to manage. And so if you have like infinitely recursive speculators, you add in quite a bit of that complexity on the actual implementation within the inference engine as well, not just in the training process.Vibhu [00:46:17]: I was gonna say, I would wonder if you could do similar, like distillation and pruning of, it's the same thing, it's just a model. Can we not just distill a lot of the weights, quantize the speculator, out of my domain? The question that also comes up is, this is all for big server workloads, right? How much of this applies to, say I have this MacBook, I wanna run Gemma really efficiently. Similar problems, not the same?Local AI vs. Data Center InferencePhilip [00:46:45]: Pretty different. I talked to Selo, about this on his podcast a couple weeks ago. The difference between inference engineering for the data center and for production workloads versus inference engineering for local AI, is that we start with fundamentally like different constraints and different goals. With local AI, it's how do I fit this model onto my hardware and then make it less dumb? And with data center influence, it's how do I load this model and then make it less slow? And we care about less dumb, and they care about less slow. But the local AI inference engineering ecosystem, I think has a lot for us to learn from in the data center space. They are experts in various forms of quantization, including dynamic quantization that we just don't touch, in the pruning, in the distillation, in the, layer removal. There'Ali [00:47:42]: Layer removal matters less.Philip [00:47:43]: Yeah. There'Ali [00:47:44]: No one loves pruning really.Philip [00:47:45]: Yeah. Well, but the, but they doVibhu [00:47:46]: Which is surprising, right? But that's, that's a whole different thingPhilip [00:47:48]: Just to fit something on the laptop.Ali [00:47:50]: Right.Philip [00:47:50]: So yeah, it's a, it's an interesting, it's an interesting space. Not necessarily that like their techniques make sense for us to do in the data center, because we have different resources and different goals, but more that the process as well as the openness of that field is something to, admire.Ali [00:48:12]: Yeah. Like to your point, like, certain optimizations that would. Like for instance, Turbo Quantum Sharper, like it made such huge hype on that and we did like a whole deep dive on Twitter and like said, what is it? How does it work? Why is it good or not? And it took off and it was implemented on local devices because your memory bandwidth is so slow on like a MacBook, for instance. But try putting the same thing on like an NVIDIA GPU on a B200 Turbo quant would not be. Like, it would not be used. Like, NVIDIA - Like, NVIDIA made it clear that this is not a good optimization, and we've seen it firsthand where the overhead of doing dequantization, quantization of, in the kernel itself with turbo quant kernel, each end is much slower than the time that you save from doing the bandwidth. ‘Cause on the B200s, you have like 3.5 terabytes per second. You don't need decrease the storage that much. You don't need to do, FP4 KV cache. You don't need to use a requant. There's, there's, there's better optimizations to be made. But on Edge devices, it's extremely important, it's extremely useful. So, seems to be, like, different optimizations there, but then they're all uniquely combined with like all you wanna quantize the model, you wanna do speculative decoding, like certain common prefixes with bothPhilip [00:49:18]: Principles.Ali [00:49:19]: Yeah, exactly. Exactly. Exactly.Philip [00:49:20]: They also do a lot of work on, model parallelism, especially over, heterogeneous topology, where you have, some sparks and they are wired together with, Ethernet, DGX sparks.Ali [00:49:35]: Yeah, this is the Exo Labs guys.Philip [00:49:36]: Yeah. You have, a nu
Every once in awhile we release a new video free for all to see, and today is that day! Please enjoy, and if you find this interesting you may want to consider becoming a member: https://www.pointfree.co/pricing. --- We explore how encoding data to JSON does not hinder our ability to query it from SQLite's powerful tools. We can filter, sort, and even _section_ results by location data stored as JSON in a table column, and all with type safe and schema safe guarantees at compile time.
This week, the guys are talking about a duress-wipe phone case that's landed someone federal charges, GCC's new policy on AI-generated code, and yet another AUR malware wave that's got Arch disabling package adoptions entirely. There's a from-scratch Rust rewrite of the X server called YServer, ShadowFetch Linux brings local AI to the desktop, Nouveau is enabling atomic mode setting by default, and GOG is finally building an official Linux client. For tips, we have DNS Globe for watching DNS propagation, the bash builtin complete for custom tab-autocompletion, uv for fast Python environment and package management, and Miller for slicing and converting CSV, TSV, and JSON data. You can view the show notes at http://bit.ly/4vZejHf, and have a great week! Host: Jonathan Bennett Co-Hosts: Jeff Massie, Rob Campbell, and Ken McDonald Download or subscribe to Untitled Linux Show at https://twit.tv/shows/untitled-linux-show Join Club TWiT for Ad-Free Podcasts! Support what you love and get ad-free audio and video feeds, a members-only Discord, and exclusive content. Join today: https://twit.tv/clubtwit Club TWiT members can discuss this episode and leave feedback in the Club TWiT Discord.
This week, the guys are talking about a duress-wipe phone case that's landed someone federal charges, GCC's new policy on AI-generated code, and yet another AUR malware wave that's got Arch disabling package adoptions entirely. There's a from-scratch Rust rewrite of the X server called YServer, ShadowFetch Linux brings local AI to the desktop, Nouveau is enabling atomic mode setting by default, and GOG is finally building an official Linux client. For tips, we have DNS Globe for watching DNS propagation, the bash builtin complete for custom tab-autocompletion, uv for fast Python environment and package management, and Miller for slicing and converting CSV, TSV, and JSON data. You can view the show notes at http://bit.ly/4vZejHf, and have a great week! Host: Jonathan Bennett Co-Hosts: Jeff Massie, Rob Campbell, and Ken McDonald Download or subscribe to Untitled Linux Show at https://twit.tv/shows/untitled-linux-show Join Club TWiT for Ad-Free Podcasts! Support what you love and get ad-free audio and video feeds, a members-only Discord, and exclusive content. Join today: https://twit.tv/clubtwit Club TWiT members can discuss this episode and leave feedback in the Club TWiT Discord. Sponsor: bitwarden.com/twit
Addy and Joey recap SIGGRAPH — covering EVRN's image-to-3D tool, ComfyCode.ai's metadata-driven JSON prompting, and Beeble's SwitchX 2.0 push toward pro-grade outputs. Neill Blomkamp releases a fully AI-generated short film, Flux 3 enters the video mod...
David Soria Parra is an Engineering Lead at Anthropic and one of the core maintainers of the Model Context Protocol (MCP). We explore the biggest evolution of the protocol since its launch, and why MCP is becoming the foundation for the next generation of AI agents.We discuss why MCP is moving toward stateless communication, what developers misunderstand about state, sessions, and transport layers, and how lessons from real-world deployments at massive scale have shaped the protocol's future. We also dive into MCP v2, SDK migrations, protocol design, extension architecture, governance, developer experience, and how Anthropic thinks about balancing simplicity with long-term flexibility.Along the way, we explore progressive disclosure, tool search, programmatic tool calling, context bloat, forward compatibility, long-running AI tasks, protocol evolution, open-source governance, observability, and why the future of AI infrastructure will depend on designing protocols that can evolve without breaking the ecosystem.Timestamps:[00:00] Introduction[01:59] Why MCP Had to Become Stateless[04:28] The Tradeoffs of Stateless Design[06:13] What We Learned About Agent State[08:04] Sessions, Models & Implicit State[09:33] Migrating to MCP v2[12:19] Lessons from HTTP & Open Source Standards[18:16] Shipping Fast Without Breaking Everything[20:35] The Future Complexity of MCP[22:44] Core Features vs Extensions[26:47] Progressive Disclosure Explained[28:16] Solving Context Bloat[30:50] Why Tool Search Beats Progressive Disclosure[32:10] The Biggest MCP Anti-Pattern[34:25] Designing for Forward Compatibility[38:41] Why "Tasks" Matter[40:53] JSON, Tokens & Better Tool Calling[44:44] Observability & Tracing AI Agents[47:34] Will MCP Ever Be Finished?[50:22] What's Next for MCP
In MobileViews 620, Jon Westfall recorded a special solo "vidcast,"(since I was not available for recording a podcast this week) recording a walk-and-talk along an historic railroad trail in Cleveland, Mississippi. Filming entirely on his Insta360 Luna Ultra with a neck mount and the creator pack microphone, Jon used the scenic Sunday walk to share his recent deep dive into data sovereignty and the process of building his own local alternatives to popular subscription apps. The core of Jon's summer project was migrating away from the Day One journaling app to avoid its $25 yearly fee and proprietary cloud storage. Using ChatGPT and Codex, he generated scripts to convert his Day One JSON export into future-proof Markdown files managed within an Obsidian vault. He then successfully replicated Day One's best features, using Apple Shortcuts and Python to ingest text snippets, process daily photos, perform offline audio transcriptions, and even selectively transcode video files larger than 25MB down to mobile-friendly sizes. Expanding his DIY software suite, Jon also automated his personal relationship management and location tracking. He built a script that "interviews" him weekly to automatically update his self-hosted Monica CRM and Obsidian vault with details about his interactions with family and friends. Furthermore, to reclaim his travel history after Google restricted the web version of Google Maps Timeline, Jon coded a tool to parse his device's local JSON location data into detailed, daily Markdown travel logs. With his self-hosted documentation ecosystem fully functional, Jon is taking it on the road for late-summer travel and will return to the podcast in mid-August.
A JSON bug is about to rock the Java world, scam compounds continue in Myanmar despite the junta crackdown, and Google has a new APT naming scheme. Show notes Risky Bulletin: A JSON RCE bug is about to rock the Java world
In this episode, we break down two governance proposals currently up for vote (ADR 29 and ADR 31), dive into revenue sharing and dynamic fees, and cover the latest updates across several upcoming chains.Swap now https://swap.thorchain.org/ THORChain is a decentralized crypto exchange. THORChain is the first and biggest DEX for Bitcoin. You can use any self custody wallet to swap and there's no KYC required.Timestamps:00:00:00 Intro00:03:00 Miradex shoutout00:05:00 Marketing update00:09:00 XMR publication strategy00:10:00 Chad met someone with close ties to Trezor00:11:00 Robinhood made its platform available for AI agents00:12:00 AI compatibility with STO00:13:00 Chad B talks about the conference00:14:00 Chad B has someone who wants to meet Kenton for a documentary00:17:00 THORDex.io may be a scam — be careful00:19:00 THORWallet won a patent case over the THORChain brand00:21:00 ADR 29 discussion (revenue sharing)00:22:00 Some integrators prefer revenue sharing over affiliate fees00:25:00 SwapKit revenue sharing looks promising and could drive more volume00:27:00 More integration tools will make Randy's BD role easier00:28:00 44 new partners want widget/API access00:31:00 Analytics tab is ready for integrators to track generated revenue00:33:00 Why API access is whitelisted00:34:00 Integrators can run their own infrastructure instead of using the API00:35:00 ADR 31 discussion00:36:00 Formalizing the relationship helps integrators build a roadmap00:37:00 Kenton explains why he changed his opinion on the app layer since ADR 2000:39:00 Permissioned access is more secure00:40:00 Merger discussion00:43:00 Why merging would be complicated00:45:00 The power of collaboration00:54:00 When will the protocol return to normal?00:55:00 XMR update — lots of edge cases and fixes00:56:00 When is Zcash coming?00:57:00 Once churn is working, Zcash can be added00:58:00 Bittensor (TAO) update01:00:00 Dynamic fee update — Symbiosis was disabled due to a bug01:02:00 Screen share of Rayyk's affiliate page01:04:00 90% of generated fees came from revenue sharing — huge volume increase01:06:00 THORChain is performing much better with dynamic fees01:07:00 Promising initial data01:09:00 THORChain vs. Chainflip pool price comparison01:11:00 App Layer design to improve price execution01:14:00 Great job, Chad!01:18:00 Signal vs. noise considerations with revenue sharing01:20:00 SwapKit partnership could be much more effective01:23:00 It will also drive more volume for SwapKit01:27:00 Start simple and keep the variables limited01:29:00 Deepen the Tron.USDT pool? (POL?)01:31:00 How have Rapid Swaps been performing? Are we ready for version 3?01:37:00 Devel's proposal and its synergy with dynamic fees01:39:00 MEV concerns01:43:00 A tip mechanism may address MEV concerns01:49:00 Let's finish our current ADRs01:52:00 Maximum effective bond — increase to 2 million?01:54:00 Multi-node operators vs. single-node operators with a single vault01:57:00 When can integrators expect revenue sharing?02:00:00 Revenue sharing presents a lot of opportunity02:04:00 AI agents and revenue sharing02:06:00 Separate STO API interface for AI agents using JSON
¿Todavía usando awk para extraer información de ps aux o df? En 2026 hay herramientas mucho mejores. En este episodio te presento tres herramientas que forman un tridente imbatible para trabajar con información del sistema en formato JSON: jc, jq y gron.jc es un conversor de comandos Linux a JSON. Se instala con pip y un simple pipe convierte la salida de ps, df, free, ss, systemctl, lsblk y hasta 60 comandos más en JSON estructurado. Olvídate de awk y de los scripts frágiles que se rompen cuando cambia el orden de las columnas. Con jc, el JSON no depende del formato de salida. Tiene parsers específicos para cada comando, incluyendo crontab, last, lsof, pip list, lsmod, date y más. Si no encuentra el parser que necesitas, puedes crear el tuyo. Está escrito en Python y tiene licencia MIT.jq es la navaja suiza de los JSON. Te permite filtrar, ordenar, agrupar, seleccionar y transformar cualquier JSON con una sintaxis potente. Está escrito en Go y es maduro, estable y rapidísimo. Combinado con jc, puedes listar los procesos que más RAM consumen, los discos por encima del 80% de ocupación o los servicios que han fallado, todo en una sola línea. Si la sintaxis te parece liosa, puedes pedirle a cualquier modelo de lenguaje que te genere la expresión que necesitas.gron es el menos conocido pero igual de útil. Aplana un JSON convirtiendo cada valor en una línea independiente con su ruta completa. ¿Para qué sirve? Para poder usar grep directamente sobre un JSON. Si alguna vez has hecho un curl a una API y has intentado hacer grep sobre el resultado, sabes que no funciona porque todo está en una línea. Con gron, cada valor tiene su propia línea y puedes buscar con grep. Además permite la operación inversa con --ungron: modificas el JSON aplanado con sed y lo reconstruyes.En el episodio presento sysreport.py, un script en Python que junta toda la información del sistema en un solo JSON usando jc y luego te permite hacer preguntas en lenguaje natural usando Llama 3.2 con Ollama. Le preguntas qué procesos consumen más RAM, qué servicios están caídos o si hay algún disco lleno, y él te responde en lenguaje natural. Todo corriendo en local, sin gastar un euro en APIs.El script se puede usar como API local, se combina con watch para monitorización en tiempo real y con notify-send para notificaciones en el escritorio. Además se integra directamente con el nightly-runner del episodio 815 para incluir el estado del sistema en el resumen matutino.Capítulos:0:00 - Introducción: el problema de la salida en texto plano2:30 - jc: convierte comandos Linux a JSON5:00 - jq: la navaja suiza de los JSON8:00 - gron: haz greppable cualquier JSON10:30 - Combinando jc y jq para consultas del sistema13:00 - sysreport.py: el script que lo junta todo16:00 - Preguntando al sistema en lenguaje natural19:00 - Monitorización con watch y notificaciones21:00 - Ventajas: local, sin coste y sin dependencias23:00 - Cierre: el tridente JSONMás información y enlaces en las notas del episodio
Topics covered in this episode: django-orjson Best Django Redis configuration for speed and size Linus Torvalds puts the foot down against Anti-AI Kernel Maintainers Django Steering Council backs the Triptych Project Extras Joke Watch on YouTube About the show Sponsored by us! Support our work through: Our courses at Talk Python Consulting from Six Feet Up Connect with the hosts Michael: Mastodon / BlueSky / X / LinkedIn Calvin: Mastodon / BlueSky / X / LinkedIn Show: Mastodon / BlueSky / X Join us on YouTube at pythonbytes.fm/live to be part of the audience. Usually Tuesday at 7am PT. Older video versions available there too. Michael #1: django-orjson Adam Johnson dropped django-orjson - drop-in replacements for the Django and DRF pieces that touch JSON, swapping stdlib json for orjson, the Rust-based library. Headline numbers: 10x faster serialization, 2x faster deserialization. The interesting question is why this needs to be a package at all. pip install orjson is the easy part. Adam's actual pitch: adopting it "isn't easy, especially when your framework uses json in many different parts." Django scatters JSON across JsonResponse, the test client and test case classes, the json_script template tag, and more. There's no single hook to grab, so you get a library that catches them all. Adam is refreshingly honest about the scale of the win. His words: "While database queries tend to dominate the typical Django application's runtime, the time spent in serialization and deserialization can still be significant." He calls it "a nearly free performance win" - not "this will 10x your app." That's a claim about cost, not magnitude, and it's worth keeping those straight. Worth flagging what the post doesn't cover: caveats. There are none in the article, but orjson has real ones. Django and Flask both render datetimes as RFC 822 HTTP-date (Wed, 15 Jul 2026 12:00:00 GMT); orjson does ISO 8601. It can't do ensure_ascii, it rejects NaN and Infinity (which stdlib happily emits), and it raises on Decimal. If you've got a JS client parsing dates, that's a wire-format change. Who should actually take this? If you're a DRF shop shoveling JSON all day, yes - it's cheap and it's real. If your app mostly renders HTML templates, you're optimizing a slice of runtime that's already near zero. The problem Adam's package solves doesn't exist in Flask or Quart. They already centralize every JSON operation - jsonify, request.get_json(), the test client, the |tojson filter - behind one provider object at app.json. So there's no library to install. It's about ten lines: import orjson from quart.json.provider import JSONProvider # or flask.json.provider class OrjsonProvider(JSONProvider): def dumps(self, obj, **kwargs) -> str: return orjson.dumps(obj).decode() # provider must return str def loads(self, s, **kwargs): return orjson.loads(s) app.json = OrjsonProvider(app) The numbers on talkpython.fm Evaluated it, measured it, and skipped it. The biggest JSON payload we serve is our MCP server returning a cached episode transcript, about 139 KB. Swapping the provider saves 0.119 milliseconds per request. That total response takes 1.1 ms We got 4.1x, not 10x - and the reason is the good lesson. Payload shape decides your speedup. The 10x is for structure-heavy data, lots of small keys where stdlib burns time in Python-level dispatch per item. Our hot payload is one giant transcript string, so the work is escaping and memcpy Calvin #2: Best Django Redis configuration for speed and size Peter Bengtsson revisits a classic: his 2017 "Fastest Redis configuration for Django" benchmark now has a 2026 update posted this week. The 2017 post pitted django-redis serializers (json, ujson, msgpack, pickle) and compressors (zlib, lzma) against each other; conclusion was msgpack + zlib as the sweet spot - avoid the json serializer, it's fat and slow. The 2026 update narrows focus to just compressors: default (no compression), zlib, lzma, and newcomer zstd. New results: lzma compresses best but is slowest; zstd is the fastest compressor on Ubuntu; differences between them are very small. Big takeaway across both: compression buys you a lot of space (2–3.5x smaller) for very little speed cost - worth it for Redis where memory is the constraint. Caveat from the author: results depend heavily on your data - his test stores short strings of numbers, so benchmark your own workload. Michael #3: Linus Torvalds puts the foot down against Anti-AI Kernel Maintainers Write up on Ars. Really good coverage by Maximillian: Time to wake up (for some) Torvalds said that “Linux is not one of those anti-AI projects, and if somebody has issues with that, they can do the open-source thing and fork it. Or just walk away.” I agree with Max, putting your head in the sand and waiting for AI to go away will likely mean you won't be working professionally in software development in the coming years. The statement came amid a lengthy thread arguing about the use of Sashiko, an “agentic Linux kernel code review system” that its creators claim can, in tests, independently find 53.6 percent of the bugs that would end up being fixed by human coders in later commits. “We're not forcing anybody to use [LLM tools], but I will very loudly ignore people who try to argue against other people from using it,” Torvalds said. “Anybody who points to the problems at AI had better be looking in the mirror and pointing at themselves at the same time,” Torvalds wrote. Calvin #4: Django Steering Council backs the Triptych Project Django Steering Council issued a Letter of Collaboration backing Carson Gross & Alex Petros's funding bid for the Triptych Project - three proposals to make HTML more expressive natively, in every browser. The three additions: PUT/PATCH/DELETE methods for forms, button actions (buttons that fire HTTP requests without a wrapping form), and partial page replacement. Distills the core ideas from HTMX/Unpoly/Turbo into the HTML standard itself - no JS, no library, nothing to ship or maintain. Current focus is button actions (WHATWG #12330): Logout instead of wrapping a button in a form. Relevant to Django directly - think the admin submit row and disguised delete links; Django 6.0's template partials were already inspired by these patterns. How to help: companies can send non-binding letters of support on letterhead; individuals can read the proposals and weigh in on the WHATWG issues. Extras Calvin: DOOMQL - A playable first-person shooter whose framebuffer is a SQL query. Michael: Granian 2.7.9 fixes WSGI threadpool scheduler starvation/underscaling Welcome Calvin post Joke: Solving all bugs
An airhacks.fm conversation with Stanislav Bashkyrtsev about: discussion about testing terminology and the difference between unit tests, component tests, System Tests, and integration tests, defining component tests as in-process invocations without HTTP, using RestAssured with MockMvc-style direct endpoint calls, avoiding mocks in favor of real system tests, why code coverage is a misused management metric, the anti-pattern of using reflection to inflate coverage, distinguishing line and branch coverage from actual verification, using coverage from system tests to detect dead code for pruning, mutation testing with PIT to measure assertion quality, testing Quarkus applications, the default Guice and Guava dependencies in Quarkus RESTEasy, starting a new microservice with a separate system-test module, calling endpoints over HTTP with the MicroProfile REST Client or the Java HTTP client, deploying Quarkus on AWS Lambda as a production-like environment, backward compatibility testing with multiple production versions, turning system tests into stress and load tests, testing connection pools and metrics under load, introducing a test-only private API to verify state changes in serverless systems, contract-driven work in large consulting projects, generating JSON and JSONB directly in PostgreSQL and returning it over JDBC, mapping database rows to Java records instead of DTOs, running GraalVM inside the Oracle Database for stored procedures and table triggers, the pendulum between database-centric and application-centric logic, the convergence of SQL and NoSQL databases, CI/CD pipelines with Jenkins and manual production deployment steps, avoiding Jenkins access to production via CGI shell scripts behind nginx, AWS CodePipeline and CodeBuild with CDK-defined infrastructure, event-driven pipelines triggered by S3 put-object events, multi-account roles with short-lived STS credentials, the size of the AWS SDK and reducing it by excluding unused HTTP clients, health checks and Kubernetes liveness and readiness probes, why health checks make little sense for short-lived Lambdas, a version endpoint for deployment smoke tests Stanislav Bashkyrtsev on twitter: @sbashkirtsev
Topics covered in this episode: The trusted-publishing debate: how to do it right vs. why you shouldn't trust it JupyterLab 4.6 and Notebook 7.6 are out! Tau – new small, readable terminal coding agent Django Tasks and Django 6.1 Extras Joke Watch on YouTube About the show Sponsored by us! Support our work through: Our courses at Talk Python Consulting from Six Feet Up Connect with the hosts Michael: Mastodon / BlueSky / X / LinkedIn Calvin: Mastodon / BlueSky / X / LinkedIn Show: Mastodon / BlueSky / X Join us on YouTube at pythonbytes.fm/live to be part of the audience. Usually Tuesday at 7am PT. Older video versions available there too. Finally, if you want an artisanal, hand-crafted digest of every week of the show notes in email form? Add your name and email to our friends of the show list, we'll never share it. Calvin #1: The trusted-publishing debate: how to do it right vs. why you shouldn't trust it https://snarky.ca/how-to-publish-to-pypi-using-github-actions-securely/ (Brett Cannon) and https://blog.yossarian.net/2026/07/07/You-shouldnt-trust-trusted-publishing (William Woodruff) Trusted Publishing (PyPI's OIDC-based auth scheme, also now used by npm, RubyGems, crates.io, NuGet) replaces long-lived API tokens with short-lived, auto-scoped credentials tied to CI/CD machine identity. Yossarian's post: it's purely an authentication mechanism between a machine identity and a package — it says nothing about package safety or quality. PyPI deliberately avoids any "verified/trusted" badge for it, unlike its verified-URL checkmarks. Same logic applies to PyPI attestations: anyone can sign with any machine identity they control, so an attestation's presence isn't itself a trust signal. Bottom line from that post: don't confuse "trusted" (machine-to-machine) with "trustworthy" (human judgment about the package). Snarky.ca's companion piece is more practical: given GitHub Actions compromises in the news, the real fix is 3 concrete steps — run zizmor to lock down workflow permissions/checkout credentials and pin actions to commit hashes, adopt Trusted Publishing to eliminate stored PyPI tokens, and require manual approval via a GitHub environment before any publish job runs. Takeaway for listeners: Trusted Publishing is good hygiene for how you authenticate to PyPI, but it's not a substitute for securing your CI pipeline itself — or for actually vetting the packages you install. Michael #2: JupyterLab 4.6 and Notebook 7.6 are out! Michał Krassowski's rundown - a chunky minor release: 68 features, 97 bug fixes, 95 contributors, one of the biggest ever. Scratchpad console (Notebook 7.6 headliner) - a console next to your notebook sharing its kernel, for throwaway experiments. Ctrl+B. Jump to last-edited cell - new commands hop through recently edited cells. File browser glow-up - Date Created column, editable breadcrumbs with Tab-completion, and Open in Terminal. Debugger - sources open in the main area, floating step/continue overlay, live kernel-sources filter. Custom layouts (Lab) - activity bar top/bottom, draggable panels, four-way tab splits, per-panel Ctrl+scroll zoom. ~5x faster extension builds - webpack → Rspack, and jupyter-builder means no full Lab install needed to build extensions. Keyboard/a11y - add shortcuts from the UI (no JSON), Find & Replace in Edit menu (Ctrl+H). Calvin #3: Tau – new small, readable terminal coding agent Tau – new small, readable terminal coding agent (Python 3.12+), built as both a working tool and a teaching project for how coding agents work under the hood Install via uv tool install tau-ai, pipx, or pip; ships a tau CLI Three-layer architecture: tau_ai (provider-neutral model layer) → tau_agent (reusable "brain": messages, tools, events, loop) → tau_coding (CLI/TUI, file & shell tools, sessions) Supports OpenAI, Anthropic, OpenAI Codex, OpenRouter, Hugging Face, and custom/local OpenAI-compatible endpoints Built-in tools (read/write/edit/bash), durable JSONL sessions with resume/branching, project instructions via AGENTS.md, and context compaction Core harness is UI-agnostic — same brain can power the TUI, print mode, or a custom frontend — usable as a standalone library too Michael #4: Django Tasks and Django 6.1 Django 6.0 finally ships first-party background tasks (django.tasks) - out of Jake Howard's DEP 14, accepted May 2024, after two decades of everyone bolting on Celery/RQ/Huey. It's an API, not a worker. Django handles task definition, validation, queuing, and result storage - it does not execute them. You bring the backend. The default backend traps people. ImmediateBackend runs tasks inline on the request thread and blocks until done - so out of the box .enqueue() backgrounds nothing (a 5-second task means a 5-second response). The other built-in, DummyBackend, runs nothing at all. Both are dev/test only. Nice API otherwise: slap @task on a function, call .enqueue(), get back a TaskResult you look up later by id - with async twins like aenqueue(). Gotcha: args and return values must survive a JSON round-trip, so a tuple sneakily comes back as a list. The community local backend to know: django-tasks-local by Chris Beaven (SmileyChris). A ThreadPoolExecutor backend that gives real background threads with zero infrastructure - no Redis, no Celery, no database - plus a ProcessPoolBackend for CPU-bound work → github.com/lincolnloop/django-tasks-local Its catch: results live in memory, so pending tasks vanish on restart or deploy. Great for dev and low-traffic production; for persistence, drop to Jake Howard's django-tasks (DatabaseBackend + worker command). Extras Calvin: Fixing the dictionary with Python 3.14 — Hugo van Kemenade stumbled on - and got fixed - a markup bug in the OED's own citation of a 1706 use of the pi symbol. Michael: Bunny DNS is now free Jokes: What's the object-oriented way to become wealthy? Inheritance To understand what recursion is... You must first understand what recursion is 3 SQL statements walk into a NoSQL bar. Soon, they walk out They couldn't find a table.
Founder/ Data S ience, AI & Economic Consultant at Analytics TX LLC Consulting practice focused on analytics, economic analysis, and executive advisory. • Design consulting frameworks to audit enterprise datasets, resolve data silos, build Python analytics workflows, predictive analysis, and statistical models, translating complex data into decision systems for founders and leadership teams. • Advise organizations on analytics infrastructure, AI tool selection, and pipeline architecture; deliver executive training in statistical reasoning, business analytics, and economic indicators. • Provide statistical and economic expert analysis used in U.S. litigation, including modeling, statistical evaluation of claims, and expert reports. Statistical modeling and damages analyses have contributed to multi-million-dollar settlements and financial exposure reductions in complex litigation matters. (Clients confidential) • Built Post it Save it App, a production SaaS platform that ingests LinkedIn post and profile data via OAuth API and processes it into structured performance analytics dashboards and reports (JSON, Excel, HTML). Deliberately built without AI, delivering accurate, deterministic content analytics at a fraction of the cost of AI-based alternatives or reliance on LLM uncertainties. • Built an end-to-end course creation system for the Professional Certificate in Business Analytics: Data-Informed Decision Making, parsing source materials (PDFs, Word, images) and transforming them into full course content including slides, AI-generated visuals, and voiceover-driven video modules using image generation and ffmpeg pipelines. Follow her on the author page on Amazon where she has published her book:https://www.amazon.com/Invisible-Hand-Visible-Profit-Decisions/dp/B0GY7V14VL Linkedin: https://www.linkedin.com/in/kruti-lehenbauer/ ***********Susanne Mueller / www.susannemueller.biz TEDX Talk, May 2022: Running and Life: 5KM Formula for YOUR Successhttps://www.youtube.com/watch?v=oT_5Er1cLvY Join Substack: https://substack.com/@susannemuellernyc?Enjoy one coaching session for free if you are a yearly subscriber. 800+ weekly blogs / 500+ podcasts / 1 Ironman Triathlon / 5 half ironman races / 26 marathon races / 4 books / 1 Mt. Kilimanjaro / 1 TEDx Talk
In Episode 73, Salma Aboukarr joins the show. A creative director and founder, she explains how she moved from painstaking CGI workflows in Blender and 3Ds Max to AI-native campaigns for brands including Coca-Cola, Panasonic, and Google Labs. She breaks down her viral IKEA exploding-room video that helped brands see the commercial potential of generative AI video, the detailed JSON prompting method behind it, and the modern AI creative stack she uses across Claude, Midjourney, Nano Banana, Seedance, FAL.ai, Z-Image Turbo, Qwen, style LoRAs, and custom AI agents.The conversation goes deep on AI advertising, product fidelity, photorealistic skin, color correction, video upscaling, automated client workflows, and why high-end AI work still depends on original concepts, trained taste, and obsessive finishing. They debate whether AI can truly be original, why technical teams struggle to manufacture taste, how creators survive a feed flooded with AI content, and why the next creative moat may come from the experiences, references, and strange little details nobody else can copy.---⏱️ Fast Hour00:00 Meet Salma Aboukarr02:17 From CGI agency to AI-first studio04:44 Product fidelity before AI got good07:34 The duct-taped road to photorealism11:10 Art direction beyond basic prompting14:28 Salma's current AI creative stack16:08 How she stress-tests every new model20:20 Why color correction still matters22:25 The IKEA video that changed everything27:13 Going viral and handling AI backlash30:38 Originality as the next creative moat33:05 Can AI actually be original?38:52 Inside an AI-native creative agency39:55 Z-Image Turbo and aesthetic base models43:06 From client brief to automated pipeline47:54 Style LoRAs for brand consistency49:11 Claude, MCP, FAL, and leaving ComfyUI51:07 The model that cut a day to 15 minutes55:13 Why AI content stopped feeling special01:00:57 The value trapped in AI archives01:02:06 Create for yourself or the audience?01:04:31 Can engineers manufacture taste?01:08:38 Finding inspiration outside the feed01:11:57 Jackie Chan and thumbnail fuel01:15:29 Final lessons from a creative trailblazer#AICreative #AIAdvertising #GenerativeAI #AIVideo #CreativeDirection #Midjourney #ClaudeAI #NanoBanana #SeedanceAI #AIWorkflow #AIAgency #AIContentCreation #BrandMarketing #CreativeTechnology #ProductPhotography #AIBranding #FutureOfAdvertising #FastHours
Jake shares lessons from rebuilding a permissions system around flexible roles, temporary permission leases, and delegated user management. He also discusses adding Apple Pay and Google Pay, plus designing automated text-message payment flows that use numbered choices instead of complicated keywords.Michael talks about building lender integrations, the importance of adding real context to API documentation, and using Claude to turn JSON specifications into PHP DTOs and enums. He also reflects on his team's new helpdesk rotation, where direct exposure to users has uncovered long-standing bugs, inefficient manual processes, and opportunities for developers to better understand the people using their software.They finish with the challenges of parenting teenagers, preparations for Laracon US in Boston, including a search for coffee, donuts, bagels, and cannoli. (00:00) - Club World Cup surprises and sports talk (03:12) - UFC, LeBron and basketball moves (06:33) - Rebuilding roles and permissions (08:39) - Permission leases and delegated management (10:02) - Apple Pay, Google Pay and text-based payments (11:50) - Lender integrations and better API documentation (14:08) - A new developer and the helpdesk rotation (17:13) - Automated tickets and failed queue jobs (18:54) - The address autocomplete bug (19:58) - When tiny code changes create hidden failures (23:02) - Why support requests need a real ticket (25:17) - Developers learning directly from customers (27:53) - Cross-department training and business context (30:51) - Building trust beyond Slack (32:02) - Weekly stress, workload and personal check-ins (35:18) - Parenting teenagers and setting curfews (37:10) - Planning for Laracon US in Boston (40:00) - Australian and American school calendars (42:51) - Vacations, PTO and wrapping up
Malcolm Matalka joins William and Eyvonne to challenge the narrative that Infrastructure as Code (IaC) is dead. Malcolm argues that the real value of IaC was never the syntax, but state and governance. Together they examine whether the state was a file problem at all, or a distributed systems problem in a JSON costume. Episode... Read more »
Malcolm Matalka joins William and Eyvonne to challenge the narrative that Infrastructure as Code (IaC) is dead. Malcolm argues that the real value of IaC was never the syntax, but state and governance. Together they examine whether the state was a file problem at all, or a distributed systems problem in a JSON costume. Episode... Read more »
Topics covered in this episode: Backup Docker volumes locally or to any S3 Pyodide 314.0 Release nb-cli: A Command-Line Interface for AI Agents and Notebook Automation Hindsight Agent Memory That Learns Extras Joke Watch on YouTube About the show Sponsored by us! Support our work through: Our courses at Talk Python AWS Community Day Midwest tomorrow Wednesday the 24th in downtown Indianapolis, Six Feet Up is sponsoring and there are 2 Sixies presenting Connect with the hosts Michael: Mastodon / BlueSky / X / LinkedIn Calvin: Mastodon / BlueSky / X / LinkedIn Show: Mastodon / BlueSky / X Join us on YouTube at pythonbytes.fm/live to be part of the audience. Usually Tuesday at 7am PT. Older video versions available there too. Finally, if you want an bonus digest of every week of the show notes in email form? Add your name and email to our friends of the show list, we'll never share it. Michael #1: Backup Docker volumes locally or to any S3 Via Bryan Weber (thanks Bryan!), who spotted it over on Virtualization HowTo. Find Bryan at bryanwweber.com. offen/docker-volume-backup is a lightweight companion container that backs up the volumes your apps actually depend on, then ships them somewhere safe. It's tiny: written in Go and about 25MB compressed, roughly 1/20th the size of the shell-based image (jareware/docker-volume-backup) that inspired it. Drop it into your docker compose file as a backup service, mount the volumes you care about as read-only, and you're off. Push backups to a pile of destinations: a local directory, plus any S3, WebDAV, Azure Blob Storage, Dropbox, Google Drive, or SSH-compatible target. Mix and match as many as you want in one run. Recurring cron-style backups in a Compose setup, or one-off backups straight from the Docker CLI. Production-friendly touches worth calling out: Rotates away old backups so you don't quietly fill the disk. GPG encryption for your archives. Notifications on finished and failed runs (so you find out about failures before you need the backup). Stop a container during backup for a consistent snapshot using a simple docker-volume-backup.stop-during-backup=true label, then auto-restart it. Run custom commands during the backup lifecycle (great for a database dump before the file copy). Docker Swarm support, plus arm64 and arm/v7 builds. Hello, Raspberry Pi homelab. Fun aside from Bryan: he searched our back catalog for this tool and the search came back so fast he thought it hadn't run. Love to hear it. Calvin #2: Pyodide 314.0 Release PEP 783 is the real news — Pyodide maintainers used to hand-build 300+ packages. Now anyone can publish Pyodide wheels to PyPI with cibuildwheel. The version jump from 0.29 to 314.0 is intentional — it now tracks the Python version, so 314.x = Python 3.14. Binary compatibility is locked per Python cycle, meaning packages you build today won't break on the next Pyodide release. sqlite3, ssl, and lzma are back in the default stdlib — no more await pyodide.loadPackage("sqlite3"). Bigger download, but a much smoother experience for newcomers. bigint precision bug is fixed — values above 2^53 were silently losing precision when crossing the Python/JS boundary. The new JsBigInt type makes the roundtrip correct. Worth flagging if anyone is doing numeric work in a browser app. Experimental TCP sockets in Node.js — you can now connect Pyodide to a real database (MySQL, PostgreSQL, Redis tested) when running server-side. Blurs the line between "Python in the browser" and "Python runtime anywhere Wasm runs." Michael #3: nb-cli: A Command-Line Interface for AI Agents and Notebook Automation From Piyush Jain (Jupyter and LangChain maintainer) on the Jupyter blog: nb-cli: A Command-Line Interface for AI Agents and Notebook Automation. nb-cli is an experimental, Rust-based CLI to read, write, execute, and search Jupyter notebooks. The premise: agents are great at CLIs but terrible at hand-editing the nested JSON in an .ipynb, so let them operate on the notebook from the outside instead of running inside it. Works with or without a Jupyter server. No server? It reads/writes .ipynb files directly and talks to kernels over ZeroMQ. Connected to a live JupyterLab, your edits show up instantly via Y.js (the same CRDT Jupyter uses). Smart output format: instead of token-heavy JSON or ambiguous plain markdown, it uses @@cell / @@output sentinels with inline metadata. Less wasted context, unambiguous structure, and it degrades gracefully on truncation. The payoff is composability. "Add a summary section and run it" becomes one shell pipeline instead of six agent tool calls. And nb search notebook.ipynb --with-errors returns only the failing cells, so the agent skips the cells that worked. Claude Code tie-in: it ships as an agent skill. npx skills install jupyter-ai-contrib/nb-cli and your agent can drive notebooks via nb. Out of jupyter-ai-contrib, which aims to become an official Jupyter AI subproject. Still early (crates.io is at v0.0.5), so kick the tires before anything load-bearing. See also marimo-pair. Calvin #4: Hindsight Agent Memory That Learns AI agents forget everything between sessions — Hindsight gives them persistent memory that learns over time Simple three-method API: retain(), recall(), reflect() — store, retrieve, and reason over memories TEMPR retrieval runs semantic, keyword, graph, and temporal search in parallel for accurate results Automatically consolidates related facts into durable observations instead of piling up duplicates pip install hindsight-all runs the entire server in-process; integrates with LangChain, LlamaIndex, Pydantic AI, CrewAI, and more Extras Calvin: Clanker: A Word For The Machine **Ponytail — You know him. Long ponytail. Oval glasses. Has been at the company longer than the version control** **Klangk: Multi-User AI Sandboxing, Collaboration and Coding Platform** Cursor announces Origin performative-ui to quick start your new idea Michael: Astral Joins OpenAI: The Interview SpaceX to acquire Cursor And OpenAI renews Open Source support Portuguese subtitles are now available for Talk Python courses DSF is hiring including Six Feet Up support Joke: Oh Babe…