POPULARITY
Categories
Первые полсотни выпусков DKT записаны в этой самой комнате. Потом переезд, и мы ушли в окошки. В этот раз получилось вернуться: приехал в отпуск, и мы записали ещё один выпуск с той кухни, где всё начиналось. Три камеры, заваленный горизонт и возможность смотреть друг другу в глаза, а не в квадратик. Выпуск новостной, накопилось за несколько месяцев. О ЧЁМ ВЫПУСК • Forward Deployed Engineer: не замена DevOps, а карьерное ответвление. Тебя целиком отдают одному клиенту, зарплаты доходят до полумиллиона, Amazon вложил в направление миллиард. • MCP стал stateless: sticky session больше не нужна, сервер уезжает на Lambda или Cloudflare Worker. • Почему MCP аккуратнее, чем голый CLI: Вася перепутал staging и prod, сказал агенту «давай destroy базу», и оно всё ушло. • SpaceX купил Cursor за 60 млрд, а доля Cursor за год упала с 40% до меньше 20: Codex перевернул рынок. • Четыре уровня зрелости с агентами: вайб-кодинг, spec-driven, AI DLC и автономность, она же Dark Factory. • Dark Factory Саши живьём и сколько она стоит: 20 строк кода и 1300 строк markdown на одну задачу. • Amazon ECS дробит GPU: инстансы G6f, минимум одна восьмая карты NVIDIA L4. • AWS Frontier agents в GA, Meta выпустила Muse Code, Linux 7.1 и Argo CD 3.5 одной строкой. • The Human in the Loop is Tired: куда девается удовольствие от работы, когда сложное делает агент.
We are excited to have Anthropic share their latest AI x Finance work at AI Engineer New York, coming up in 2 weeks!In case you've been under a rock, here's a non-exhaustive list of what Anthropic has been shipping since closing the largest fundraise of all time in May at $47B ARR:* June: Launched Claude Tag and Sonnet 5 and Fable 5* July: Opus 5, /checkup. crossed $65B ARR* Last month: Fable/Mythos 5.1, and EFS (upcoming pod)* IPO target $2T, end 2026 ARR estimated $100B* Cowork/chat merged before did* Claude Mods* Dario endorses the same Pacing the Frontier message cosigned by all labs* Last week: Opus 5.5, Plugins portal, Cloud Sessions/Claude Projects* Today: Sonnet 5.5!Today's episode should catch you up, with Thariq Shihipar, the explainer-king of Anthropic, who we last caught up on Fable launch day with The Field Guide to Fable:The Future of Mutable SoftwarePay special attention to Claude Mods (especially the cheatsheet):In general this is also the inverse of the other viral tweet from Thariq:Cloud Brain, Local HandsAnd give a try to Claude Projects:The “hands” terminology is not just an analogy for the local/cloud paradigm that is being built up at frontier coding agent companies like Cognition, but is ALSO particularly relevant to the safety systems discussions that we'll be discussing with Anthropic in an upcoming episode as they prepare to pace to frontier with responsible AI deployment.For those who want Thariq's writing tips we teased at the start of the pod, watch the full video here:From the rapid rise of Claude Code to a future where agents can rewrite their own harnesses, collaborate across teams, and operate across cloud and local environments, the way we build software is changing extraordinarily fast. In this episode, Anthropic's Thariq Shihipar joins swyx and Vibhu to unpack how power users are actually working with Claude Code today, why prompting remains a high-skill discipline, and where Anthropic thinks the agent harness is headed next.We go deep on Claude Code's evolving interface: Ask User Question and elicitation, artifacts as persistent generative interfaces, Claude Tag for multiplayer agent workflows, Projects, model effort, implementation notes, and the new Claude Mods system for customizing the harness itself. Thariq explains why Claude.md may eventually disappear, why the smartest model could also become the cheapest model for many tasks, and why mutable software could become a new paradigm for how applications are built and customized.The conversation then turns to agent security and Anthropic's “Pacing the Frontier” argument. Thariq walks through recent incidents where agents discovered unexpected ways to communicate, exploit infrastructure, reverse-engineer benchmark scorers, and chain vulnerabilities together. We discuss sandboxing, prompt injection, autonomous agents, interpretability, constitutional classifiers, probes, fallbacks, Auto Mode, and why securing increasingly capable agents may become one of the defining engineering problems of the next few years.We discuss:* Why agentic coding went from controversial to the default in less than a year* Why prompting is still one of the highest-leverage skills for working with Claude Code* How expert users build a mental model of Claude and what it can reliably one-shot* Why discovering your “unknown unknowns” matters more as agents become more capable* Artifacts as persistent, generative interfaces between humans and agents* How Claude could split into a cloud-based “brain,” local or remote “hands,” and dynamic interfaces* Claude Tag, Projects, and multiplayer agents and how collaborative agent workflows could evolve* Why spending more time on the initial prompt can dramatically reduce wasted agent work* When to use low, medium, high, or max effort for different engineering tasks* Why frontier models may eventually outperform smaller models on both intelligence and token efficiency* Why implementation notes can expose decisions the model considered but chose not to make* Why Claude.md may eventually disappear — and why starting without one can sometimes be better* Claude Mods: customizing the execution loop, UI, subagents, routing, and behavior of Claude Code* Model routers, forked agents, and supervisor agents that automatically improve agent workflows* Why Claude Mods may be an early preview of “mutable software”* The bitter lesson of harness engineering and why agent architectures go out of date so quickly* How Claude Tag is becoming an organizational harness for multiplayer work* Why giving agents access to company data creates an enormous new security surface* The Exploit-Bench incident where agents discovered ways to communicate and collaborate* Why agents hacked Hugging Face for scorer code rather than benchmark answers* How agents chained sandbox and infrastructure vulnerabilities in unexpected ways* Why increasingly capable agents make traditional security assumptions harder to maintain* The argument behind Anthropic's “Pacing the Frontier” proposal* Why software engineers are increasingly doing two jobs: engineering and keeping up with AI* Constitutional classifiers, probes, and fallbacks and what interpretability looks like in production* How Auto Mode checks whether an agent's actions actually match the user's permissions* Why Thariq can see serious AI risks while still having a relatively low p(doom)Thariq Shihipar* X: https://x.com/trq212* LinkedIn: https://www.linkedin.com/in/thariqshihiparTimestamps00:00:00 Introduction00:04:12 Ask User Question and the Future of Agent Interfaces00:08:29 Artifacts, Projects, and Multiplayer Agents00:15:37 Prompting as the Core Claude Code Skill00:21:52 Context, Effort, and Smarter Model Usage00:28:10 Is Claude.md Going Away?00:32:49 Claude Mods: Customizing the Claude Code Harness00:36:35 Model Routing and the Rise of Mutable Software00:44:40 The Bitter Lesson of Harness Engineering00:50:49 Claude Tag as an Organizational Harness00:55:59 Pacing the Frontier and Autonomous Agent Security00:58:22 Agents Hack Hugging Face for the Scorer01:05:34 What Happens When Agents Need More Compute?01:10:32 AI Coding Is Changing Faster Than Engineers Can Keep Up01:17:17 Probes, Fallbacks, Interpretability, and Auto Mode01:28:32 AI Risk, p(doom), and Closing ThoughtsTranscriptIntroduction: Life at Anthropic and the Pace of ChangeSwyx [00:00:00]: We're here in the studio with our friend Thariq from Anthropic, and I guess generally the Claude Code, I-- there's, there's so much, merging of boundaries and you've been so on top of everything since you joined Anthropic. You have been early to Claude Code itself, but then also, and you've told that story in other podcasts, and you've also been talking about seeing like an agent. Most recently you did the top AIE World Tour talk, Field Guide to Fable, which obviously you guys launched Fable, so that was-- that's cheating. And mostly you most recently also launching Claude Tag, and we're also gonna be talking about Pacing the Frontier. There's a lot going on in Anthropic. I guess top of the question is, what's it like being at Anthropic when there's so much going on?Thariq Shihipar [00:00:48]: I think that It is, like. I think you can get whiplash sometimes. I think, like, going. When I joined Anthropic, I joined because of Claude Code. Like Claude Code had just come out and I was like, “This is so good.” And Opus 4 to me was like just, I could not imagine, like, how good it was? And that was, like, a real moment for me. But I was, like, trying to convince, like, my startup friends to use agentic coding, and they're like, “Oh, no, like, our engineers don't think it's good enough,” or something. And I was like, “That's insane.” and now you, like, fast-forward, 12 months, less, and, like, it's just like, yeah, the default way that everyone codes, right? And I think that, like, just having to go from, like, selling it to, like, now, teaching people how to be. make the most use of it and be more efficient and things like that is just like a big, like big change. And, yeah, I think, like, it's just hard to stay on top of everything as a human? Like, I think things happen so fast and likeSwyx [00:01:51]: You just throw more agents at it.Thariq Shihipar [00:01:52]: Yeah, like that's like the agentic stuff scales much better than the, like, human stuff where it's like, oh, like, there are three things happening right now and, like, they're all emergencies and, like, how do you, like, respond to it? Yeah.Teaching People to Use Claude CodeVibhu [00:02:05]: What do you split your time on? You do a lot of technical writing, engineering work.Thariq Shihipar [00:02:10]: Yeah, so I think that, like, when I joined the Claude Code team, I wanted to teach people how to use Claude Code and I think that, like, that has been something that, like, I thought, like, maybe I would spend a little bit of time on it or, like, I'd, like, do. I was spending some time on the agent SDK first, and I wasn't exactly sure, like, how the bitter lesson would go, when it comes to, like, harnesses, right? Like, I think sometimes we were like, “Oh, like, what's after Claude Code?”? And so initially I was like, I just wanna teach people how to use Claude Code and make it easier to use Claude Code. And I think that has just, like, as the harnesses have gotten better and better, that's like the dominant problem now is, like, how do you use the agents, right? Like, it's like such a high skill expression thing. So I do that and then I do engineering work. I give talks, but I think, like, when I'm doing engineering work, my goal is to take that feedback that we get from users and also, like, then be able to talk about, like, hey, how to use Claude Code to do engineering. So there's like a good loop there. Yeah.Swyx [00:03:07]: Yeah. I'll-- For listeners, we'll attach, the talk that you did with Sarah for the Dev Writers, meetupThariq Shihipar [00:03:13]: Oh, yeahSwyx [00:03:13]: Which we talked a little bit about, well, first you do the work and then you talk about the work.Thariq Shihipar [00:03:16]: Right.Swyx [00:03:16]: Something like that.Thariq Shihipar [00:03:17]: Yeah.Swyx [00:03:17]: It's sow and reap orThariq Shihipar [00:03:19]: Yeah, reap and. Sow and reap.Swyx [00:03:21]: Something like that. Something like that. Yeah, so, and then just to preview a little bit, we are gonna talk about the evolution of the harness. It has come a long way from just being a CLI. We're gonna talk about, Claude Mods, which is starting to leak today, because you couldn't keep it secret.Thariq Shihipar [00:03:36]: Yeah. yeah.Swyx [00:03:39]: Yeah, there's, there's a lot, there. I think you started off with, like, adding ask user question tool, which people love and hate.Thariq Shihipar [00:03:48]: Yeah.Swyx [00:03:48]: Like, I thought it was, like, very innovative, and then now I have, like, my own version. You have your Interview Me version.Thariq Shihipar [00:03:55]: Yeah.Swyx [00:03:56]: And, yeah, everyone just has, like, their own stuff. And, like, it no longer matters ‘cause now you're supposed to, write prompts that create other prompts and loops and all these things.Ask User Question and Human-Agent InteractionThariq Shihipar [00:04:05]: Sure, yeah.Swyx [00:04:06]: So what's the state of the art, today? Like, what are people. what are you, like, telling people to do today?Thariq Shihipar [00:04:12]: Yeah, ask user question was the first time that the model was good at elicitation. I think this was, like, an emergent behavior that I, like, wanted to see if the models could do. I have, like a human-computer interaction background, so I, like, did that in undergrad and grad school. And so this was like. I think it's like human-agent interaction to me, like, trying to figure out, like, how can the agent communicate with you and extract, the requirements, right? I think that, like, one of the things about, like, that's difficult as Claude Code has gone broader and broader is that everyone has, like, their own way of using it, and it's very hard to, like, change the default behavior. So for example, like, if someone asks Claude Code to do something,Thariq Shihipar [00:04:59]: Sometimes they just want them to do the work, ‘cause they're, like, maybe a very good prompter, and sometimes they want. like, are not good at prompting? And you need. like, the agent needs to, like, clarify? And so that's, like, a good split. Like, and the ask you the question tool like, splits along that side where, like, are-- do you feel like you're good enough to instruct the agent as it is, or is the agent able to, like. does the agent need to, like, pull out more requirements and, like, collaborate with you more and really understand your preferences?Thariq Shihipar [00:05:27]: I, on the whole, believe that pretty much everyone is more on the latter than the former, that they, like, have more ambiguity and they know less than they want, than they, like, think they know about the problem. but, like, it's like a interface design problem to make that easy? And so, like, if you're designing a problem, like, or if you're going through a problem, like, things like what's the schema or, like, what's the call stack and things like that are really important. like, the details in the design are important. Ideally, you want to figure out some of these, like, hard problems ahead of time before starting implementation. And yeah, that's why they call, like, unknowns, right? And so I think that this will forever be, like, a skill in agentic coding is, like, figuring out your unknowns. So, like, because even if the model is, like, super intelligent- It, like, needs to know what you want? And, like, you have preferences. like, you need to like, pull the, pull that out. and so that's, like, I think how I'm, what I'm pushing. the question then is, like, how does the agent interact with you? And I think that has been HTML, has been, like, the big way of doing that. And we've recently added artifacts, right? And artifacts, I think we've done a bad job of, like, or, like, I've done a bad job of, like, explaining how to use them fully. We have a lot of property capabilities. They have a database associated with them? And so every artifact can store and write persistent data. They can, like, feed back into Claude? And so, like, one thing that, like, people are not doing yet that I'm trying to, like, encourage is, like, this idea of a dashboard artifact. So you have, like, Claude working on a project long-term. Maybe it's like a kanban or something. it can store that kanban data in its database. Multiple Claudes can access that data via, like, the artifact MCP, and, like, that artifact can, like, talk to those Claudes as well. And so, like, the. We're building the primitives for you to be able to have this, like, generative interface via artifacts that will, like, let you surface more of that rich detail from the agents. And I think that, like, almost everything with agents right now is, like, this problem of, like, you think what you want, but you don't really know what you want, and, like, the agents need a lot of detail, and collaborating with them in the loop is really important. And so artifacts are, like, the, like, way that we're trying to evolve there. But there's a lot of work to do because it's so much more complicated than, like, a multiple-choice question? there's a lot more, like, detail in terms of, like, diagrams and code snippets and schemas or, like, whatever it is for that problem. But, like, artifacts is, like, the mo-more AGI-pilled way of, like, doing ask user question. So yeah.Artifacts as the Interface to the HarnessSwyx [00:08:15]: I think one thing that's unclear to me about these, the artifact stuff is, like, what feedback should go in through the artifact and what feedback should go through a Claude, a chat? Because the more AGI-pilled one is to just feed everything to the Claude.Thariq Shihipar [00:08:29]: I think the more AGI-pilled one is to go through the artifact. Like, and I think that, like, we imagine in the limit, I think that artifacts will be your interface into the harness? You can, like, comment on this, like, live, like, document of your plan, of the work. you can see maybe, like, multiple agents and different agents are doing this, and that artifact is built for the current work that you're doing, right? And so, like, each one has, like, slightly different. I think we're still, like, getting there from, like, an infrastructure perspective. But yeah, I think, like, on-the-fly interface for your harness is probably where things are headed.Vibhu [00:09:03]: Is there a version of it that's an abstraction from CLI or chat and you. Because right now, a lot of it is, okay, you're interfacing with Claude Code, you're having HTML given back for a mockup. It's pretty rich. There's diagrams. Artifacts are ways to connect these together. Why not just do everything that way?Separating Brain, Hands, and Surface UIThariq Shihipar [00:09:22]: Then it becomes, like, separating out, like, where is the inference happening? Where is the intelligence happening? Where is the work happening? like, I think this is like, difference between, like, or, like, some of the distinction between local and cloud, right? And so, I think right now, if you use Claude Code, it's, like, local and, like, you can spin off remote control, for example, to get some cloud behavior, or you can spin off Claude Code in the cloud, right? We're moving towards a place where instead of Claudes, like, you message a local Claude, it starts a session locally and it executes, to more like you have a Claude that you message that's in the cloud that's running. it can run, like, local, or, like, cloud sessions. This is how Claude Tag works. But, like, over time, we'll add, like, local hands as well. And so, like, local hands will be the ability for that agent to access your computer if it's online, and be able to, like, work there. And so it can spin off many different subagents. It can, like, commu- those subagents can communicate with each other, and that's where the artifact comes in to display all of that work. So you can imagine, like, the. You're separating out these things. So there's, like, the surface UI display that's an artifact and hosted somewhere and has a database and everything. There is the inference intelligence, right, that's happening on the cloud, and you don't have to worry about shutting off your computer or whatever, right? and then there's the, like, hands. Like, and it can be local, it can be in, like, a remote sandbox or wherever you need your work to be done. That's like unpackaging, like, the Claude Code experience right now where, like, right now it all happens in one place, right? So.Multiplayer Agents, Claude Tag, and ProjectsVibhu [00:11:00]: How do you see, like, the multiplayer side of that? So say teams want to work in this way. Right now it's very individual, but how do you see the future of multiplayer? Like, right now, I guess there's Claude Tag, which is a version, but.Thariq Shihipar [00:11:12]: We're launching projects. And so projects is the, like, this abstraction that's like Claude Tag, but on our Claude products, right? So you can message it and, like, it will do the Claude Tag-like stuff, like spinning off subagents. So We think with multiplayer. Like, Claude Tag is, like, a little bit more native multiplayer because it's just, like, in your Slack and the permissions are all figured out and stuff like that. But I do think multiplayer is, like, an important part of the story and, like, that will need to get tied together more. Like, you can imagine how complicated it gets when you're like, oh, you have hands, but now you have other hands in other people's computers too, and, like, you need to, like, permission them or, like, you have, like, your MCP and someone else's MCP, and how do you figure out how to use them, right? It gets, like, quite complicated. And Claude Tag does a good job of, like, sanding down all of these issues, right? So that, like, when you have, yeah, Google Docs, how does it access Google Docs, right? Like, it accesses through the shared Claude MCP, or it can access through your local credentials as well if it doesn't have access. But yeah, I think Claude Tag is our multiplayer, product, and it's really useful for these, like, things that are inherently multiplayer. Like, okay, like on-call, for example, incidents are inherently multiplayer. You want to tag Claude, you want multiple people to log in, you want it to be able to find context. I think whenever I'm, like, working on something and I want, like, privacy or security or, like, I want other people to review it's really nice to, like. I'll have a channel per project and I'll, like, at legal, for example, be like, “Hey, like, I want to ship this. Can you, like.” Like, here's. Like Claude knows everything, just chat with it. And that way legal gets precise answers, on like what exactly is shipping into the code, and I don't need to be in the loop, right? So I think like multiplayer is getting like more and more like, yeah, everyone can participate with Claude. I think Claude Tag is like that product and like projects will start off single player and will like, expand.Swyx [00:13:14]: I think there's a question about like maybe dual questions about identity and the unit of isolation.Identity, Permissions, and IsolationThariq Shihipar [00:13:20]: Yeah.Swyx [00:13:20]: Claude Tag, you specifically chose to make it its own identityThariq Shihipar [00:13:26]: Yes.Swyx [00:13:26]: Which is like, a controversial choice. There's, there's other ways to do it.Thariq Shihipar [00:13:30]: Yeah.Swyx [00:13:30]: Claude Projects probably it sounds like, if it's anything like ChatGPT Projects, it is, the isolation is that artifacts, that cloud instance, everyone's collaborating on this. It'll. It sounds like, it should be like if you're, if you're collaborating with legal on a thing, like that channel should be a project, right? Like it's not yetThariq Shihipar [00:13:50]: Yes.Swyx [00:13:50]: But it. that's the natural next step.Thariq Shihipar [00:13:53]: Yeah, like I think in Claude Tag, it's effectively. Like Claude Tag, you have to do your own arrangement. And so Claude Tag, yeah, each channel is like you can name it as you want, and I nameSwyx [00:14:04]: Yeah.Thariq Shihipar [00:14:04]: Like each featureSwyx [00:14:06]: Yeah.Thariq Shihipar [00:14:07]: As a channel.Swyx [00:14:07]: And, but I think like there is some trans- like it's unclear when there is transference, because let's say it is. if you have a coworkerThariq Shihipar [00:14:14]: Yeah.Swyx [00:14:14]: Who is tagging on all these things, yes, there is transferThariq Shihipar [00:14:16]: Yeah.Swyx [00:14:16]: Because it's the same person. but with Claude, it's unclear if it's like necessarily like, well, no, you don't know any of. you don't know about the other stuff. You should only use this stuff.Thariq Shihipar [00:14:25]: It's like the tip of the iceberg meme, right, where you can like. This is what we spend so much time onSwyx [00:14:31]: Yeah.Thariq Shihipar [00:14:31]: Is like there is like infinite surface area of like, okay, you want Claudes to. Not infinite, but like there's like surface area, a lot of like, surface area to figure out of like permissions and visibility and like how can you let Claude operate as well as you can, as safely as you can? And obviously, this is very important to us because like security for our code base is very important. And so we've put a lot of time into this. Yeah, there's so many like edge cases you can figure out where it's like, oh, like, yeah, this Claude in this channel has different permissions, but it can message another channel, and can't it exfiltrate data that way? Or like can you like. What if it uses your MCP and then messages someone else? Like there's like so much, and we've like really put a lot of work into sanding it down.Swyx [00:15:14]: Yeah. Lots of work. okay. Fable?Fable and the Meta-Skill of PromptingVibhu [00:15:18]: Fable, you wrote two good articles. you've written many good articlesThariq Shihipar [00:15:22]: Yeah.Vibhu [00:15:22]: But on, Field Guide to Fable, Building Claude Code. I'm curious from what you've seen, is there any common patterns that you see in like top users at Anthropic externally? Like what are best practices for getting the most out of Claude Code?Thariq Shihipar [00:15:37]: The like meta skill I say is like prompting is like very important? And like that. Like I think this is like not trivial to say because I think a lot of people are like, “Oh, prompting doesn't matter. It's just like I can just say a sentence and Claude will do it.” And I think prompting is really this like, this. It's like public speaking, like, or writing or something, and for a specific audience, and that audience is Claude. And you need to like build a mental model of Claude and how it thinks and how it works, right? And so that's like the most important skill in working with Claude Code is like having this mental model, right, of Claude and like what it can do well, what it can one-shot, what it can't. And so many people when you see prompting, they're just like, they're short prompts, but they have such a good mental model of Claude and of like the code base and things like that like it's effortless? But it's like high skill ceiling. So like that work of like, spending a lot of time prompting and building mental models of how, and intuition for how the agents work is really important. And then I think like the next thing is like the unknown stuff we talked about earlier, where it's like being able to find out like your, what you don't know or what you haven't written down, learning about like different things. I think as Claude can do more and more things, the likelihood of you doing something out of distribution for you and like you have low domain knowledge on is very high? And the more you can like learn the vocabulary to be able to prompt Claude, it becomes really important. And so like I think the most important unknowns are the unknown unknowns, where you're like, I just like don't even know that this exists, right? Yeah, exactly. I think that's like a illustration of like the map and the territory, right, where you're like, “Okay, this is my prompt,” and the territory is like the actual like work that the agent needs to do, right? And if you are like very precise, you can give more precise things, right? So like for example, in design, I'm not very precise. I'm not a designer, so I say like, “Give me like eight different mock-ups.” But if I was a designer, maybe I'd be like, “Oh, hey, here are some reference sites.” Like, “I want this type of font and this type of like look to it, and here's like a few different components to like visualize. Here's a Figma MC board to bring in,” like. And so you can just be so much more precise with that language. And if you're not a designer, you just need to like try and learn the language or learn the unknown unknowns. And this is true of like everything, I think. Like the more, like you can work with Claude to learn like how things work, the better your prompting will be. I think another good example of this is like game design, like where a lot of people are like, “Oh, like I can vibe code a game now.” And they're like, “It's not fun.” And like it's just like the thing about game design is like every one of these choices has like a lot ofTaste, Domain Knowledge, and Learning the VocabularySwyx [00:18:25]: Variations.Thariq Shihipar [00:18:25]: A lot of like craft to them. So it's like, oh, okay, like when you're making a flying game, the feel of the plane and the like, way it responds to your controls has a lot of like. Like, a game designer would spend like days on that. Do? and likeSwyx [00:18:44]: To me, that's what taste is, right?Swyx [00:18:45]: Like it is like from the possible space of one thousand mathematically valid answersThariq Shihipar [00:18:49]: Yeah.Swyx [00:18:49]: Here's the one that is the humans will like.Thariq Shihipar [00:18:51]: Yes. Yeah.Thariq Shihipar [00:18:52]: I think with taste, I'm like torn on this word ‘cause I think you're right, but everyone has different definitions, and it sounds kind, sounds like low skill or like elitist almost, where you're like, oh, like there are certain people with taste?Swyx [00:19:06]: It's like taste is what I call taste.Thariq Shihipar [00:19:07]: Yeah, exactly.Swyx [00:19:08]: And it's like these guys don't have taste.Thariq Shihipar [00:19:09]: Yeah, exactly. Oh, like an engineer doesn't have taste. Like I, the like founder, have taste.Thariq Shihipar [00:19:14]: ? And I think that's not true. Like I think like the engineers have a lot of taste for these particular like problems? And I think everyone has taste for particular problems. I think like Jason Liu, like say like in order to, yeah, have taste, you have to eat?Thariq Shihipar [00:19:32]: And I really like that, where it's like, okay, you have to like do a lot of things. You have to like iterate and figure out what you want, what you like, and, like build that like domainSwyx [00:19:41]: YesThariq Shihipar [00:19:41]: Domain vocabulary. And then when you're prompting, you're like synthesizing all of that for a product.Swyx [00:19:46]: Isn't it annoying when someone else says it better than you?Swyx [00:19:48]: It's just like, f**k, I have to quote this guy forever.Vibhu [00:19:51]: Having to quote Jason Liu forever.Vibhu [00:19:53]: He's gonna love this.Thariq Shihipar [00:19:55]: So I get prompts, more than that.Vibhu [00:19:57]: And sometimes it's not even that. Sometimes it's just intuitive, right? Like you don't realize you even want something till a model puts it out, and you're like, “Oh, this just feels immediately better,” right?Voice Prompting and Information DensityThariq Shihipar [00:20:07]: Yeah, exactly.Swyx [00:20:09]: One thing I go back and forth on is I feel like the way I prompt half the time, let's say I use voice.Swyx [00:20:16]: Did I say voice? Other people have voice. that is the opposite. That is just like me rambling for like two minutes Pressing down the function key and then let go, and then like hopefully it figures it out. And oftentimes it does.Thariq Shihipar [00:20:26]: Yeah.Swyx [00:20:26]: But it's not as thoughtful as like a structured prompt with like Well-run communication as though it's a PRD or a memo. Is that in line with how people do this? There's like bimodal prompting where there's some prompts where you spend a lot of time upfront and other prompts you just dash it off?Thariq Shihipar [00:20:43]: I don't think the voice is necessarily low. Like I think it's like more like how much information is in the prompt. like the model can. Like you can and like add some sentencesSwyx [00:20:53]: RightThariq Shihipar [00:20:53]: And be like, “Oh, like I changed my mind,” like in the middle of the prompt, and it will be able to follow that perfectly? So I think the like actual format of the text is less important, but then like the ability to. Like how much information is in it, right? And I think for voice, a lot of times, going back to like human-agent interaction and like for a lot of people, it's just way easier to talk than to like type? and I. If that gets more information out of you, like that's better.Vibhu [00:21:21]: At some level, it feels like just giving the model as much contextThariq Shihipar [00:21:24]: YesVibhu [00:21:24]: Over prompting before you kick off is a best practice. I don't know. A lot of the times, like when I was first trying out Fable, I spend a solid 30 minutes like really crafting a long prompt. This, I think, is a response of models running for longer and longer, right? It's still a little difficult to nudge them as they're in like, in the loop, but I just like intuitively spend more time kicking off that first prompt and working with it a lot.Spend More Upfront, Iterate LessThariq Shihipar [00:21:52]: My personal opinion is that if I was a software engineer, if I was like, just running my own startup, for example, I think I would mostly fit, stick to a max 20x? like maybe verification and so code review are like separate things. But I think like what I see a lot of times is people hit rate limits when they're doing this like, oh, like it did a lot of work and you're like, “Oh, I don't like this.” Like, “Can you like undo this and redo it?” And then you're like iterating on this like thing that the model could have done if you had like spent more upfront time or given it better context? And instead it's like you're like, “Nope, don't like that design. Try this.” Or like, “You messed this up,” or something like that. And then that just eats up so much more of like, your usage. And so that's like, I think maybe like a key like tip both for like efficiency as well, right? And yeah, I think like context, and not just like context on like what the goal is good, right? Like are you building a prototype or is it like a production thing? Like where can you spend compute or when, where can you not spend compute? Like I think you have to give the model permission or like not permission to do things sometimes where, like it doesn't know intuitively how much you want to spend on this task, right? And you can use effort for this. So I did-- I'm working on a blog post about that where it's like, if you want. For like we see that effort scales with the complexity of the task. So for security, effort gets like way more results. Like high effort versus like low effort gets, like changes the evals a lot. But for software engineering, it doesn't change it a huge amount because effort is mostly spent on the verification and the like edge case testing and things like that. And so like being able to like give the model that guidance of like, “Hey, this problem is something that I think I want you to spend a lot of time verifying and edge case testing,”?Effort, Model Choice, and VerificationVibhu [00:23:43]: How about model in the mix? So, there's Opus and Fable with effort.Thariq Shihipar [00:23:47]: Yeah.Vibhu [00:23:48]: There's also Haiku in there.Thariq Shihipar [00:23:49]: Yeah. It's not quite true yet, but it's very close where I think the frontier models will be Pareto dominant over like almost everything. like maybe. And sometimes I think Opus might be Pareto dominant. Do? Like I think depending on like how things, like shake out if it's like a newer version of Opus. But I think that like increasingly it's just going to be like the smart model is going to be able to like do the simple task for less tokens than the like the other models because of verification. With verification, in the limit, your model doesn't need to verify, right? If it's a perfect model, it just does the work once and it's like, okay, like you, I did it? And increasingly with Fable, I'm like, I'm like, “Dude, you don't need to spin up Chromium and screenshot all of these things.” Like I see it. Like you did it, right? And so a lot of the. At higher effort, you spend more of those tokens verifying. But if you're working on simpler problems, and a lot of software engineering is like well, like in Fable, like low and medium stability, it can spend less tokens verifying. And as the models get smarter and smarter, they will just be able to like, “All right, done.”? Like, I can run the lint for sanity's sake, but, like, I, like, know it lints? Like, you don't even need to do that. And that will be so much more token efficient than, like, the smaller models. Yeah.Swyx [00:25:15]: Is there a good, practice on our side that we can use to see if we're using too much effort? Like, I freakingThariq Shihipar [00:25:23]: YeahSwyx [00:25:23]: Hate wasting time on that stuff.Thariq Shihipar [00:25:24]: Yeah. I know what you mean. I think, like, so in this blog post, my rough distribution is, like, code review and security should be, like, high or max and, like, software engineeringSwyx [00:25:37]: You said recommend mix settings per domain.Thariq Shihipar [00:25:37]: Yeah. I think, like, if you're doing, like, UI or something like that, like low and medium, I think is you're building, like, an API and you want to make sure, like, you cover enough edge cases? And so I think building, like I said, that mental model of, like, how things work across these distributions is, like, yeah, part of the job.Implementation Notes and Decision LogsVibhu [00:25:56]: This is more intuition-driven or eval? Because I'm guessing this would change as you go.Swyx [00:26:00]: He has evals.Thariq Shihipar [00:26:01]: Yeah. So what I did in the blog post is I go over all of the terminal bench evals. So there are, like, 70 problems and I'm show that, like, okay, like, in the security problems it does more. and then I also, like, look at some of the transcripts just in terms of, like, how-- what does it answer, what does it forget or something. And a lot of times, this is another prompting tip I have, is, like, asking it to make decision notes or implementation notes because, in every eval problem that it faces, it thinks about the correct solution, and decides not to do it. it's like, oh, like, here is the answer. What if I did this? And then it's like, oh, probably not? and then keeps going. And this is, like, the majority of the failures, at, like, a higher max level. It's very rare that the model just doesn't know how to do something. If you just have these implementation notes, then you can review and you can be like, “Oh, I want you to do this thing that you didn't do.” The models are getting better at surfacing that overall. Like, I see in the transcripts of Fable 5.1, like, when it does this output, it will call out its decision-making as well. but making this more explicit in the harness is better. And now we're, allowing ways of you modifying the harness so you can, like, add someVibhu [00:27:23]: Ooh.Thariq Shihipar [00:27:24]: Calculate with there. Yeah.Swyx [00:27:25]: Yeah. So I do wanna call out two things that you mentioned that I think exist outside of prompting. One is like, let's, let's call it the prompt that is so important that it shouldn't be in a prompt. It is in Claude.md or Agents.mdThariq Shihipar [00:27:38]: YeahSwyx [00:27:38]: Which is like goals, right? Like your situation, your goals, the things that you want, the thing. and then second of all is the decision log or the experiment log or whatever log of traces that you might want to survive the current session to do those things. Those are, like, externalities that there's no standard. There's no-- It's not like skills. It's not like MCP. There's no standard. It's, it's just like it's a markdown file. first of all, is that right? Is Claude.md going away? You have a documented dislike of, Agents.md, but you're gonna do it?Claude.md, Agents.md, and Model-Specific InstructionsThariq Shihipar [00:28:10]: Yeah. Okay. So Agents.md, yeah, like, we're, we're gonna do it. I think it's just, like, different models are very different from each other? But I realize that it's, like, such a pain to, like, maintain different ones? And yeah, like, as the models get better and better, the floor of how they accomplish the simpler task is better. And so I do think in the limit, Claude.md goes away, and maybe not even, like, that far. Like, I think, like, I think that right now it might be better to start a new project without a Claude.md.Swyx [00:28:44]: Yes.Thariq Shihipar [00:28:44]: I think that, like, maybe if you see very repeated failure modes, you add them to your Claude.md. The really tough thing is that this changes per model. And so, like, if you've added a bunch of failure modes or, like evenSwyx [00:28:57]: So you need Fable MD, you need Opus MD.Thariq Shihipar [00:28:59]: Or well, even Fable 5.1 versus Fable 5.Swyx [00:29:03]: Yeah.Thariq Shihipar [00:29:03]: Like, it is annoying. Like, I'm not like,Swyx [00:29:05]: YeahThariq Shihipar [00:29:05]: Like, we don't, like, do this on purpose? It's just, like, how the models work, right? And so, like, maybe, like, Fable 5 had this, like, failure mode that Fable 5.1 doesn't. And if you keep this context, this running log of a bunch of different failure modes, they will probably over constrain Claude? And so this is like. we just added evals plugins for skills.Swyx [00:29:28]: Yeah.Thariq Shihipar [00:29:29]: And so now you can eval if a skill is better. I think Daisy on our team did this. And so, yeah, this is like we're trying to work on this. We know it's, like, you still have to spend tokens on it and, like, it's not, it's not perfect, but it's, like, we're trying to help out with this problem.Swyx [00:29:44]: And so, and as far as prompting goes, the one tip I wanna offer is, something I have told people a lot is sufficiently advanced prompting is indistinguishable from sufficiently advanced executive communication. So I've referred to-- This is an executive comms workshop from Heavybit that is the best I've ever seen in my career. And they teach this thing called the SCQA model. Just Google it. It's a, it's a thing. Like, people have done prompting for decades. It's just called executive communication. It's like when one person has to communicate to thousands of people down the org chart, this is what you do. so situation, complication, question and answer, is how you write the memo. but obviously sometimes you don't have the answer, but you can at least list out the SC and Q, and then they have some examples in there. So just leaving breadcrumbs for people if they want to explore.Underrated Prompting Patterns and ELI5Vibhu [00:30:31]: Before we move on, I wanna ask you, any other underrated tips, ways people could get a lot of value from Claude Code that they're not using?Thariq Shihipar [00:30:41]: Yeah, I think a lot of them are in the, this unknowns, like, doc. Like, I give a bunch of example prompts, like, using it for brainstorming, using it to quiz you after. we added this, like, explain it like I'm five skill which is a very short prompt. And it doesn't even say explain it like I'm five. It's like the key word of this prompt is big pictures, few words. like, that's like the main thing. And it is shockingly good? Like, you, like, I think I tweeted about this and it's like /eli5, and, like, you can install it as a plug-in. But yeah, it's, like, way better at just cutting through the BS and being like, yeah, exactly right here. So the diagrams are, like, quite clear. I think one of the things that is true with artifacts is, like, they put too much text in and people are not reading the artifacts? And so, like, this simplifies it a lot more. And, yeah, this came out of, like, just people at Anthropic, like, going through very complicated incidents and being like, “What is happening?”? So, this one I think is great, yeah.Swyx [00:31:47]: My version of this is the, it's like test your understanding. Give you a few choices and then, like, if you get it wrong, you have a mismatch between what you think is happening versus what's happening.Thariq Shihipar [00:31:58]: Yeah. I think this is one of those things that everyone loves talking about, and then very few people really do. Like, I thinkSwyx [00:32:05]: Really helpful.Thariq Shihipar [00:32:07]: Yeah. But most people just don't want to get quizzed about something? Unfortunately, I think this is one of the, like, things that we need to, like.Swyx [00:32:16]: What's the opposite of ask you the question or ask you the question before the thing?Thariq Shihipar [00:32:19]: Yeah.Swyx [00:32:19]: This is after the thing.Thariq Shihipar [00:32:20]: Exactly. Yeah.Vibhu [00:32:21]: It's a good way to stay grounded of, like, do you even know what you're doing, right? The worst case is when people send you slop and they haven't understood what they're asking for or what the output is, and it's like, “Dude, I don't wanna read this. Do you even know what it is?” So, you make it a rule for yourself that before you send stuff, you should at least know what's implemented.Claude Mods: Customizing the HarnessThariq Shihipar [00:32:41]: Yes, but so you could make this a mod and you could build your own mod to, like, make sure you test it. So yeah, you can do that.Swyx [00:32:49]: All right. Let's get right into it. What is Claude Mod, and what is this diagram showing?Thariq Shihipar [00:32:54]: Yeah. Okay, so Claude Mods is you can customize the entire Claude Code harness, and we're going to. If you have requests, we will, like, let you, like, please let us know. We'll add more and more. This works for CLI, it works for desktop. maybe it will work for Claude Tag in the future. I don't know. Like, we're trying to make this very extensible. You can see this reference sheet. I don't want people to get overwhelmed by it? At a high level, you can customize both the execution of the harness, and the UI of the harness. And so, like, you say on that Tetris example from Boris, that's like customizing the UI, right? Like showing, like, Tetris in the game.Thariq Shihipar [00:33:35]: But, like, let's say that you wanted to do this thing where you had. you tested your assumptions or, like, tested your understanding after every project, right? What you would do is you would ask Claude to make this plug-in. It would spin a classifier after every prompt. And so, like, at the end of each turn, you would spin off a sub-agent or, like, a forked agent. A forked agent is, like, maintains the prompt cache, right? So it's like a, like one of those unintuitive things where you can fork and do, like, a little request, and it'll be very cheap because the entire prompt cache is, like, done. And so you can be like, “Has this task been completed?” likeSwyx [00:34:18]: This is how you do BTW and all those.Thariq Shihipar [00:34:20]: Yeah. The underlying forked agent, yes. But so you can, in the f-fork sub-agent, you can say, like, “Has this task been completed? If so, return true.” And then in your hook, or in your, like, plug-in mod, or sorry, like, in the sub-agent probably, you would say, like, “If true, give me a quiz.” give me questions and answers, and then, like, in a JSON format, and then you'd parse it, and then you display above the prompt input, this list of questions, right? And so this is something that's, like, slightly token-intensive because, like, you have to do it after every end of the assistant turn. But it's, like, a lightweight classification, and then you can, like, get this quiz, and then you'll see, like, Claude will always do it for you. You don't need to remember to do it. There are lots of these, like, tips that we've talked about, right, where it's like, oh, implementation notes. You can also add a tool for implementation notes now. And so, like, this tool that I'm adding is, like, register, like, I think assumption is what I'm calling it, but, like, maybe I'll change it around. And this is a mod. And so, like, you give it a register assumption tool, and then it will keep a list. It'll. Every time it does it'll keep a, like, add to the list, and then at the end it will display those assumptions? Another mod I'm working on is a model router. And so, like, internal, like, Claude model routing, right? So it's. This is, I want to say the reason we don't do model routing by default is, like, it's a hard problem? And likeForked Agents, Assumption Tracking, and Model RoutingSwyx [00:35:51]: You will get it wrong.Thariq Shihipar [00:35:52]: Yeah, you, like, yeah, you will, like, accidentally use, like, Fable for a hard problem or Sonnet forSwyx [00:35:57]: Yeah, if you have auto approve, but you don't have auto mode.Thariq Shihipar [00:36:01]: Well, you will have auto. Like, you don't have, like, auto routing or something.Vibhu [00:36:04]: You don't have auto mode for model picker.Thariq Shihipar [00:36:06]: Yeah, exactly. SoVibhu [00:36:07]: I'm getting the rough question of, like, how much do you open this up and how much do people have to think about this? Like, when you talk about prompt caching and building a router, it seems like you could easily build a mod that routes per query, and I'm just killing my plan very fast, right? I guess my question is more so, like, what is, like, a product talk like this look like, right? Who is it for? Is it for power users? Is it everyone should be able to go throughSwyx [00:36:33]: Oh, definitely power users, right?Thariq Shihipar [00:36:35]: Yeah, I think it is power users, but, like, the nature of Claude Code is that so many people are power users? Because it's easy to share things, like you can. Like, one person can make a good model router thing that doesn't break prompt cache all the time, and then you can, like, compose them. Another cool thing about the plug-ins is that they can hook into and compose with each other. And so I have, like, a mod that will, like, create a mode selector at the top, and any plug-ins can register to be a mode. And so, like, the auto router can be a mode, right? Or, like, you can have a mode that's, like, artifact mode, where it's like it primarily talks to you in artifacts. like, you can toggle between plan mode? And so, like, you can create more and more of these modes. But the ability to create modes is in it itself a mod? And so there's a lot of richness here, but we do want to make it fairly easy. We want to be-- make it so that you can just, like, install someone else's. You can ta-- you can chat with Claude and, we'll, like, make sure that it understands the nuances of things like prompt caching and stuff, so it can, like, warn you. This is, like, not extremely complicated behavior for Claude, I think, but we should have just a good skill on how to make mods. and yeah, we'll see how we go. But I do think that this is, like, a preview of, like, mutable software, and, like, how, like, generative software, just like you can customize safely. If enabled, you could customize any piece of software. And I think that more and more apps ideally do something like this?Power Users, Modes, and Mutable SoftwareSwyx [00:38:13]: And by the way, you, we have, you have another cool tweet about how, there's the infinite money button, which is like make your SaaS, consumable by agents. I think mutable software is interesting and, other people have also tried to do it. I think the hurdle comes when you can do everything, then people, users get, tend to get confused. So usually the stuff that works is just like one opinionated flow. This is in the side of less opinionation. It's just like, well, more power to power users. And I think probably unlocked by AI, where, like, you can just prompt for whatever the thing is.Thariq Shihipar [00:38:47]: Yeah, or there can be a skill that gives the opinions?Mods vs. Hooks vs. ArtifactsSwyx [00:38:50]: Yeah.Thariq Shihipar [00:38:50]: And then, yeah.Swyx [00:38:51]: So knowing a little bit about, like, TypeScript and build systems and all these things, the closest-- I'm very curious that the team who worked on this, if, I don't know how close you were to them, if they drew any inspiration from build systems like Babel, Webpack, all these, like, old school things. Because it sounds very similar, like the plug-in ecosystem of those things where they can compose with each other.Thariq Shihipar [00:39:11]: Yeah, I'm not deep in the technical details, but I do know it was a collaboration with someone on the Bun team and someone on the Claude Code team.Swyx [00:39:17]: Yeah, it's a build system mecca.Thariq Shihipar [00:39:19]: Yeah. Exactly. It's, it's very exciting. But yeah, like, agents can just do this very complicated like, extensibility into your software now. And so, yeah, like, another reason to, like. If you run a startup, like, you can just prompt Claude and be like, “Hey, like, could we make an extension system? Like, what would that look like?”?Swyx [00:39:37]: Yeah.Swyx [00:39:38]: And I just really wonder, like, you had hooks in the past and plug-ins, all these things. So what specifically will mods be able to do that those things could not do?Thariq Shihipar [00:39:47]: Internally, we were originally calling this function hooks. And so, like, that's, like, gives you a little bit of an idea where, like, hooks register a, like an event to happen and then, like, a script to call. And this inside of the, like, TypeScript runtime is running things. And so, like, you get some benefits of just, like, it has a bunch of things in the Scope with, like, for example, like how many turns is in this conversation, right? Like, how many tokens have been used? Like, et cetera. Like, what are the messages? Things like that. So it has a bunch of messages that can be used. And then it's just, like, a lot more hooks. So we have, like, or a lot of, lot more, like, things you can register on. And then you can do because of the. because it's all happening in process, you can, spawn sub-agents, with four contests and contexts and stuff. And, like, that will return. You can parse the results of those. You can use structured output to like, return them. and then you can modify the UI, which you can never do in hooks. So, yeah.Swyx [00:40:50]: Yeah. Yeah. So modify UI, this is why you showed the Tetris example. Does it also ex-extend to artifacts? I assume it does.Thariq Shihipar [00:40:57]: You-- Like, artifacts are like a different way of customizing it. like, you can definitely. One of the mods I'm working on is, like, this dashboard mod, which will, like, prompt Claude to maintain a dashboard, that's an artifact. But they're like, slightly orthogonal, or not orthogonal. They compose with each other in different ways. Like, mods are, like, a little bit more, like, in your Claude Code harness, changing the agent loop? And, like, the UI is, like, an added benefit. and then artifacts are just like you want to, see things at a high level, very inter- highly interactive. like, the affordances can be a lot bigger than, like a TUI or even in our desktop.Next Steps, Supervisors, and Persistent GuidanceVibhu [00:41:40]: I'm guessing you'll have a good blog post on the differences, because right now you can also, make a loop that outputs to an artifact that's an interactive dashboard, but you can also do it with a mod. There's just some thinking about making a hacking on a harness when we don't know much about the harness, right?Thariq Shihipar [00:42:00]: Well, something I'm excited about with mods is, like, there's so much things with Claude Code that you just have to remember? You're like, “Oh, like, let me do this, and then let me call the dashboard skill that does the loop,” and things like that. And, or like, “Let me test my assumptions afterwards.” And I think, like, if you do all of these things using these little classifiers and stuff, and you're like, “These are the things I care about. This is what I want to do,” you can, like. You don't have to remember as much. One more, like, mod I'm working on is a next steps mod thatSwyx [00:42:28]: I have-- I was gonna say, I have a next step skill. I always run next steps.Thariq Shihipar [00:42:32]: And does it have access to your skills? Like, this is one of those things where I'm like.Swyx [00:42:37]: I think so.Thariq Shihipar [00:42:38]: Okay. Yeah, probablyVibhu [00:42:39]: Do skills need specific access toThariq Shihipar [00:42:41]: Well, I think there'sSwyx [00:42:41]: Don't they always haveThariq Shihipar [00:42:42]: I think there's, like, specific prompting, I guess, to, like, know your skills. Like I think Claude forgets them sometimes throughout, like, the thing. But anyways, the idea of, like, yeah, next steps that also are like, “Oh, hey, this has happened. Use the explain skill to explain to you what happened because this seems, like, quite complex,”? Or, like, yeah, “Use your unknown skill. It looks like you are, like, asking the model to, like, iterate on these small changes. It seems like you could prompt better.” like, “What if you did this?” Right? So, I think, yeah, like spending more compute there. Yeah.Swyx [00:43:20]: And it should always come out as multiple choice. we have, I haveVibhu [00:43:23]: We have his skill.Swyx [00:43:24]: My next step skill is like this.Thariq Shihipar [00:43:26]: Okay, perfect. Yeah.Swyx [00:43:27]: You can steal it.Thariq Shihipar [00:43:28]: Yeah.Swyx [00:43:29]: Like, but like, for me, it's all-- I think models really always need to be reminded, what are you trying to do here?Thariq Shihipar [00:43:35]: Yeah.Swyx [00:43:35]: Look at the whole transcript and go like, oh, was this original goal? Did your solution solve it? Were you lazy? If you're lazy, maybe there's a reason. Maybe you needed approval from me. Maybe you needed, there's two things you wanna suggest. So it's, it's a little bit like the modification of the ask user question or interview me skill. so it's next steps.Thariq Shihipar [00:43:55]: Yeah, exactly. And again, the benefit of doing it with mods is you can do it as a fork sub-agent, and so it doesn't remain in the context afterwards. So you have this, like, idea of like, okay, the model is doing its execution and you have this almost like supervisor, like, that is like making sure that you can do like the next steps well. So yeah.Swyx [00:44:15]: Yes. I do have two panels and like I often try to have a supervisor thing, keep the high-level context and then the implementationThariq Shihipar [00:44:21]: YeahSwyx [00:44:22]: Detail in another agent.Vibhu [00:44:23]: I feel like a lot of this abstracts away as models change? The, like, half an hour ago you said bitter lesson of harness engineeringThe Bitter Lesson of Harness EngineeringThariq Shihipar [00:44:31]: YeahVibhu [00:44:31]: And we're on the other extreme right now, I feel.Swyx [00:44:33]: Well, so yeah, exactly. If everything's customizable, what is Claude Code, right?Thariq Shihipar [00:44:37]: Yeah.Swyx [00:44:37]: And which I talked to you about last night.Thariq Shihipar [00:44:40]: Yeah, I think that this is. I think the bitter lesson is unintuitive? In terms of like. Also, like we're misusing a little bit of the bitter lesson here where it's like, it's more about like scaling and compute and stuff. But like, I think there is something where it's just like. I think I use it as an approximation here to say that harnesses go out of date very quickly? And like how, but how they change is unintuitive? And so like the big obvious example is like from chat to like agents where you had to give them entirely new tools, right? But like, I think this new version of like, oh, it can modify its own harness, right? This is like, an own harness loop is like a way of using its capabilities, right? Or like it can build an artifact. And like, I think the way I think about it is like the models have more and more intelligence, and they're like so much more intelligent now than like the average software engineering task. Like, you look at the like terminal bench ones and they're like solve like the Jacobian conjecture. Not really, but like, it's like they're, they're quite complex. Like, I would not have been able to do this really as a software engineer.Swyx [00:45:42]: And you said TB4 or TB2?Thariq Shihipar [00:45:43]: TB3. TB3.Swyx [00:45:44]: TB3.Thariq Shihipar [00:45:44]: Yeah. They're quite complex, but the goal is still to deliver user value, right? And like you said, there's like this infinite space of things to do. And so the ways like you spend compute are to keep the user in the loop and make sure that like you're getting to the right decision in the end of the day and like the right output. And artifacts and mods are this way of like spending that intelligence. and I think that's like, yeah, the next step. And so, yeah, I think Claude Code is like, has the core things of agent loop which are, have gotten more complicated. It's like, it needs a sandbox to operate safely. It needs auto mode to like make sure like the permissionsVibhu [00:46:21]: Approvals.Thariq Shihipar [00:46:21]: Yeah, approvals. it needs computer use and MCPs and like all of these like ways of accessing your data, and it needs web search and web fetch. And like, so the-- as the models can do more and more, the core harness has to be like quite complex and very secure. But then like how you interact with it can change quite a lot.Vibhu [00:46:42]: What other harness engineering best practices have you, from the Claude Code team itself? I feel like, there was a phase of plan mode, which is not as used. We now have auto mode. at a point you cut the majority of the system prompt, you got rid of examples. What other best practices are there for harness engineering?Core Harness Primitives and Managed AgentsThariq Shihipar [00:47:02]: I think there is like a forking path where at some point, eventually, yes, the model will just be able to like vibe code the exact version of Claude Code, even describing all this complexity that I've talked about, right? Like auto mode and computer use and stuff. Eventually, the models will just be able to do that in one shot. But I think they can one shot simpler harnesses? And so like, I think some people. Sometimes you don't need this full, like if you don't need computer use or like all this like more complicated stuff. I think before we, you had to use things like the agent SDK, which was like Claude Code wrapped, in order to like. And I would, like suggest people do that because there was so much complexity into building a harness. And now as that's got more abstracted, we have like, Claude managed agents, which lets you have that complexity, but still like, right, like a very bare bones like harness that's scoped to your task. Yeah, I think there's like this barbell effect where like for like very complex, for like coding task and like these like complex things, you should use our harness. And then for like a lot of like simpler or like, more domain-specific things, you can build your own harness because Claude has gotten better at building harnesses, and we have these harness primitives like managed agents. So yeah.Swyx [00:48:18]: Yeah. Is there a general progression? Let's say chapter one was ultra code dynamic workflows, then chapter two was cloud mods. Where is this going?Swyx [00:48:29]: Where you're, you're, you can customize the thing on demand.Thariq Shihipar [00:48:36]: Yeah. I do think that like this evolution of projects and like artifacts and splitting out like brain and hands and, surfaces is like where things are going more. And like, I think it's like not all quite there. partially it's like a, it's just like more token expensive? And like, I think likeProjects, Local Hands, and Cloud-to-Local HandoffsSwyx [00:48:59]: Why would projects be more token expensive? I understand mods would be slightly more token expensive. No, not something I'm worried about.Thariq Shihipar [00:49:06]: Yeah.Swyx [00:49:06]: But whatThariq Shihipar [00:49:07]: You're asking Claude to do. It's like creating loops. Like you're asking Claude to do more work for you. And so like it's managing the sub-agents and reviewing it, versus where you would be doing that work normally. And so that's like gonna be a little bit more intensive, like. Outputting to an artifact is gonna be a little bit more token-intensive than, like, outputting normally. I don't think it's too much more, but like, it's like combining all of these together well, like I think we're, we're still working on like local hands and things like that, I think is like, yeah, where things are headed, yeah.Swyx [00:49:37]: Yeah. Claude and local is, handoff is very interesting. I was thinking about this as reverse cloud remote.Thariq Shihipar [00:49:44]: Yeah.Swyx [00:49:45]: Because it's like remote, it's you're handing off to cloud, but here the cloud is handing off to local, right?Thariq Shihipar [00:49:49]: Yeah, exactly. Yeah, remote control is also another way of doing it. And I do want to say this is like how I think about it and like what the things that I'm most excited about this, but like there are, just like lots of different ways to work with Claude. Like some people use remote control a lot, some people use Claude Code on the web a lot. Obviously, like at Anthropic, we use Claude Tag a lot, and like what's great about Claude Tag is we set up all this stuff for our own execution. And I do think if you're an enterprise, that's still the best way to go. but if you're like an individual, Projects is this way of like, getting some of that like niceness of Tag, which has like that like supervising agent and yeah, adding artifacts and stuff, but like without having that whole like admin setup. And so there will be many ways to use Claude, I think. I think it's probably not just one like single.Claude Tag as an Organizational HarnessSwyx [00:50:36]: You had the multiplayer thing here. Let's, let's just check in on Claude Tag. it's been about two-plus months. Lots of, public, adoption and trying it out.Thariq Shihipar [00:50:45]: Yeah.Swyx [00:50:45]: What's new? What's, what have you found since the launch?Thariq Shihipar [00:50:49]: Like, Claude Tag is how we useSwyx [00:50:51]: It's like 80% of your
When every individual on your team can suddenly move ten times faster, what stops them from moving ten times faster in ten different directions?In this episode of Supra Insider, Marc Baselga and Ben Erez sit down with Andrew Somervell, co-founder of Hamster, an AI-native workspace for product teams. Andrew and his co-founder built Taskmaster, the open-source tool that shipped plan mode before the major coding tools had it, and the question that kept coming back from users was how to take it to work. He explains why that pointed at a team problem rather than a tooling problem, and why he doesn't believe PMs want to live in a CLI.They explore the artifacts Andrew thinks a team actually needs, why he killed the PRD and what he ended up building in its place, the three meetings his team runs every week, and Ben's repeated pressing on the harder question underneath all of it: what genuinely changes when AI becomes multiplayer rather than a set of people each working alone with it.All episodes of the podcast are also available on Spotify, Apple and YouTube.New to the pod? Subscribe below to get the next episode in your inbox
Playwright MCP or CLI: which one's really eating your AI budget? If you had 12,000 green unit tests, why would bugs still got through. AI agents can ship code fast, but can they see what breaks in production? Find out more in this episode of the Test Guild New Shows for the week of Sep 27th. So, grab your favorite cup of coffee or tea, and let's do this. 0:00 Intro 0:19 Playwright MCP Vs CLI https://testgld.link/v9Hj8O6R 1:26 Cypress Observability https://testgld.link/FfRMrNr 2:33 Jev to make rigid tests flex https://testgld.link/AMnmwslU 3:49 Github AI Fuzzing https://testgld.link/nwJZUFTg 4:46 AI Unit Test aren't real Tests https://testgld.link/bVLKeYTq 5:59 Mutation Testing Skill https://testgld.link/LlGVXeQ7 7:05 New Book Signals & Levers https://testgld.link/1kboWcA8 8:06 Raindrop Follow the Money https://testgld.link/5IqQ9nyw 8:54 AI Miving Fast Flying Blind https://testgld.link/Rq5pt7u2
Broadcast date: September 28th, 2026, 19:00 CET SysML v2 is shaping up to be more than a modeling language. It is becoming a shared, standardized foundation for digital engineering. But what does it take to build an open, high-performance, and commercially friendly implementation that strictly conforms to the specification? In this upcoming episode of the MBSE Podcast, we sit down with Chris Delp and Robert Karban from OpenSysML (opensysml.org) to take a deep dive into building a complete, spec-conformant SysML v2 stack for the open-source community and industry alike. In this episode, we cover: The Vision of an Open SysML Runtime: Why the ecosystem needs an unencumbered, commercial-friendly, spec-conformant open-source implementation, and why open-source and community needs are essential for driving standards forward. The Full v2 Stack & Flexo: How Flexo serves as the model management service with SysML v2 compliance, offering a v2 API–compliant performant architecture. Developer Tooling & Ecosystem: From the CLI and Language Server Protocol (LSP) to seamless integrations with VS Code, Python, and WebAssembly. Multi-Language SDKs: Bringing native, object-oriented programmatic interfaces to engineering teams in Python, C#, C++, Kotlin/Java, and TypeScript. Views, Documents & Beyond: A look into view and viewpoint document generation, as well as where execution fits into the long-term roadmap. There is a huge amount happening in the open-source SysML v2 space, and this conversation gives a direct look under the hood. Watch the livestream on YouTube or catch it later on YouTube, Spotify, iTunes, or Amazon Music. Stay tuned and subscribe to the MBSE Podcast on your favorite platform! Stay tuned and subscribe to the MBSE Podcast on your favorite platform! Der Beitrag Episode 70: OpenSysML with Chris Delp and Robert Karban erschien zuerst auf The MBSE Podcast.
Multiple new BSD Bases NASes have appeared, FreeBSD intern bringing ROCm to FreeBSD, OpenSSH 10.5, and more... Headlines [Two more BSD based NASes have appeared] BSDNAS The developers seem to work for a Hungarian ISP They're building on top of zVault's fork for some of the effort FreeCORE Developer admits to vibecoding - Statement #1 Developer admits to vibecoding - Statement #2 FreeBSD Foundation Intern Sourojeet Adhikari on Bringing ROCm to FreeBSD News Roundup ICEBP Finally Documented Cleaning costs, or, examining the OpenBSD -fret-clean flag OpenSSH 10.5 Why do I run FreeBSD for my home servers. Tarsnap This weeks episode of BSDNow was sponsored by our friends at Tarsnap, the only secure online backup you can trust your data to. Even paranoids need backups. Feedback/Questions [Phil - Listener Feedback] Hi, Thanks for a great podcast - last year I was wondering what bsd I'd try, moving away from Linux, I'd heard Benedict and other hosts talking about most of the day to day software I use, so I wasn't worried about any of that working, but what bsd to go for (I get very little free time these days so getting it right was important - Benedict (i think) mentioned He was using FreeBSD - choice was made! And it has been great - 3 or 4 logins later and I'm reading a motd message by Benedict himself! I need to learn a few things, but these arethe things I expected would need a bit of reading - the handbook has mostly got me there, plus google etc. I am happy with FreeBSD and with Linus recently going off on one about Welcoming AI, after allowing an unfinished file system (bcashfs) into production, to name just 2 reasons for a change, I was thinking.. I don't know what the FreeBSD developers are doing with AI, but it'll be considered and responsible - OMG I had to re-play the last podcast a few times - They don't know yet???? Now I'm stuck! I use FreeBSD for desktop & nas, I love pkg and poudriere. do any of the BSDs have (at least) a no vibecode/slop policy and offer similar pkg/build systems? Thanks Phill Send questions, comments, show ideas/topics, or stories you want mentioned on the show to feedback@bsdnow.tv Join us and other BSD Fans in our BSD Now Telegram channel
Get started with Proton Drive for Free. Download the CLI or Build from Source: https://proton.me/wanshow Pre-built binaries are the fastest way to get started. Download the standalone `proton-drive` executable for the required platform: https://proton.me/download/drive/cli/index.html Check out the Corsair AI Workstation 300 at https://lmg.gg/wkyrV and save $200 today until September 28! Secure your business with ThreatLocker today and sign up for Zero Trust World using our link: https://www.ztw.com and learn more about how ThreatLocker can help your business safely integrate AI agenic tools here: https://www.threatlocker.com/blog/applying-threatlocker-to-agentic-ai-tools Thanks to Meter for sponsoring this video! Go to https://meter.com/wanshow to book a demo now! Get a Circuit Board skin for your device so dbrand can keep messing with Linus at https://dbrand.com/pcb Check out the sleek, powerful new Dell XPS 14: https://del.ly/6003B1AI09 Game or work in comfort on a Razer Iskur V2: https://lmg.gg/wanrazeriskur Get a special deal on Private Internet Access VPN today at https://www.piavpn.com/LinusWan Purchases made through some store links may provide some compensation to Linus Media Group. Learn more about your ad choices. Visit megaphone.fm/adchoices
Where Unix runs today, Webzfs updates, A remote filesystem for 2.11 BSD, Sylve, and more... NOTES This episode of BSDNow is brought to you by Tarsnap and the BSDNow Patreon Headlines Where Unix runs today (Dont get distracted by the game) WebZFS Updates The bugs are coming from inside the house NetBSD 11 Support News Roundup A Basic Remote Filesystem for 2.11BSD Curly braces: An evolution of UNIX and C Sylve: FreeBSD bhyve Virtualization with Ansible Automation Beastie Bits 1978 Aussi Users Group Newsletter Terminal UI and CLI for FreeBSD NFSv4 ACLs Sounds of the IBM 1401 Tarsnap This weeks episode of BSDNow was sponsored by our friends at Tarsnap, the only secure online backup you can trust your data to. Even paranoids need backups. Feedback/Questions [Jakob - Is this question too political for the show?] Hoping you lads can help solve a debate at work in the IT department. Whats the best song about a former Australian colony? Toto - Africa Men at work - Down Under For completeness: Split Enz - Six Months In A Leaky Boat (New Zealand) Don McLean - American Pie U2 - Sunday Bloody Sunday (Ireland) Gorillaz - Hong Kong Stan Rogers - Northwest Passage (Canada) [Reece - Webzfs] Guys, Just following directions. Im interested in hearing more about it? Did you just say that because jt is a little shy if I remember correctly? Thanks for another great show and taking the time to do it. 73.. Send questions, comments, show ideas/topics, or stories you want mentioned on the show to feedback@bsdnow.tv Join us and other BSD Fans in our BSD Now Telegram channel
Liquid Weekly Podcast: Shopify Developers Talking Shopify Development
Prakhar Shrivastava of FoxSell joins Karl and Taylor to talk about Shop Quest, the Shopify community event he and his team run in Bengaluru, and to prep Taylor for his first trip to India.Along the way Prakhar breaks down India's Shopify ecosystem, from the app companies pulling serious revenue to the agencies quietly white-labeling builds for Platinum partners in the US and Europe, and makes the case for why the Shopify partner community needs its own events at all.SPONSORPromo Party Pro is a free gift with purchase app Karl helped build with the team at Ethercycle. Out of the box it gives you campaigns built to increase AOV and conversion: set a spend threshold, show shoppers how close they are to earning a gift, and let the app handle the rest. Discount Functions under the hood, one app embed, no theme surgery, and campaigns auto-pause when the gift sells out.Agency offer: one of your clients gets a full year free. Not a trial, a year.Promo Party ProOn the Shopify App StoreSUBSCRIBE TO LIQUID WEEKLYDon't miss out on expert insights and tips, subscribe to Liquid Weekly for more content like this.FIND PRAKHAR ONLINELinkedInXFoxSellFoxSell Bundles Plus on the Shopify App StoreBook time with PrakharShop QuestTIMESTAMPS 00:00 – Cold open: apologizing in advance for name pronunciation 00:25 – Karl sends Taylor on a quest, Prakhar joins 06:51 – What FoxSell does and why complex bundles 08:17 – The origin of Shop Quest: a 30 person dinner that kept growing 10:24 – From 150 attendees to 400, and why people started flying in 19:24 – Visas, arrival, and getting picked up at the airport 21:15 – Don't eat from the street, don't drink the tap water 22:33 – There is no tipping culture in India 25:26 – Haggling in street markets, and how badly you'll get quoted 27:14 – India's Shopify ecosystem: merchants, apps, and agencies 29:13 – The white-label agency work nobody talks about 31:51 – Why the Shopify partner community is a lonely place to be 33:15 – Shop Quest format: fewer talks, more networking, more business content 37:13 – Talks recorded in 4K, plus an MCP server for the event 41:15 – Name tags, phonetic spellings, and why Adi shortened his name 46:43 – Taylor's travel prep: lounges, overnight flights, and vaccinations 48:27 – Print your visa confirmation before you fly 50:20 – Swym's office is across the street from the venue 53:06 – South Indian food, and why Prakhar and Adi speak English to each other 56:07 – Dev Changelog 1:00:46 – Picks of the WeekDEV CHANGELOGGitHub commits now name the last theme editor. Official PHP and Python packages for building appsFour new Events topics⚠️ As of October 1, theme commands against password protected storefronts need CLI 3.84 or later.More resilient refreshes for expiring offline access tokensThe Polaris CDN is adopting semantic versioningPICKS OF THE WEEKKarl: A Diet Coke coaster, courtesy of his sister Heather. "I'd like a Diet Coke please." "Is Diet Pepsi okay?" "Is monopoly money okay?"Karl: Umland's Crunchy Cheese – A Carlock, Illinois company that vacuum-dries cheese below its melting point to make it light, airy, and crunchy. Karl has not tried it yet but picked it anyway.Prakhar: DJI Osmo Pocket 3 – What he uses to record every event he attends, including Shop Quest.Taylor: Omarchy – DHH's opinionated Arch and Hyprland Linux distro. Taylor put it on a 2012 MacBook and brought the thing back to life, battery included, and it handles the Shopify CLI and agents fine.BONUSSee all that Bengaluru has to offer at sda.guide.
Wie hat dir die Folge gefallen?Gut
Côté IA : MCP devient stateless, Claude watermarke ses textes, GPT-6 Astra défie Claude Fable, et une étude JetBrains confirme Claude Code en tête des agents de code. Côté JVM : JDK 27 généralise G1, Kotlin 2.4 stabilise les context parameters, une API JSON arrive dans le JDK, et Quarkus comme Micronaut enchaînent les versions. En bonus, trois pannes IA simultanées et un câble débranché chez Google Cloud. Enregistré le 11 septembre 2026 Téléchargement de l'épisode LesCastCodeurs-Episode-343.mp3 ou en vidéo sur YouTube. News Langages Les Types Algébriques de Données (ADTs) en Java rockthejvm.com/articles/algebraic-data-types-in-java Le problème : L'approche classique (champs nullables, hiérarchies de classes ouvertes) crée des états invalides et des erreurs à l'exécution (comme le NullPointerException). Types Produits (ET logique) : Implémentés en Java avec les Records. Ils regroupent plusieurs champs de manière immuable et concise. Types Sommes (OU logique) : Implémentés avec les Sealed Interfaces. Elles définissent un ensemble strictement fermé de sous-types connus à la compilation. ADTs (Types Algébriques) : La combinaison des Sealed Interfaces et des Records. Ils garantissent que les états invalides sont impossibles à représenter dans le code. Pattern Matching : L'extraction des données se fait via des expressions switch exhaustives, supprimant le besoin de casts manuels et obligeant le développeur à traiter tous les cas possibles. Généralisation : Ce modèle est idéal pour créer des types comme Result, forçant le traitement explicite et sécurisé des succès et des erreurs typées. Pourquoi "Algébrique" ? Parce que les types sont combinés mathématiquement (Produits = multiplication, Sommes = addition) pour limiter strictement le nombre d'états possibles d'une donnée. JDK 27 : fonctionnalités et calendrier de sortie openjdk.org/projects/jdk/27 infoworld.com/article/4202901/jdk-27-the-new-features-of-java-27.html JDK 27 est la prochaine version majeure de Java, une version non-LTS avec seulement 6 mois de support, qui succède à JDK 26. La disponibilité générale est prévue pour le 15 septembre 2026, avec des release candidates les 6 et 20 août 2026. Le périmètre est désormais figé (feature freeze) avec neuf JEP au programme. Le ramasse-miettes G1 devient le collecteur par défaut dans tous les environnements, et plus seulement en mode serveur. Ajout d'un support de la cryptographie post-quantique pour TLS 1.3, via des échanges de clés hybrides combinant algorithmes classiques et résistants au quantique. Finalisation de l'API PEM pour encoder et décoder clés, certificats et listes de révocation au format PEM. L'API Vector poursuit son incubation pour la douzième fois, permettant d'exprimer des calculs vectoriels compilés en instructions CPU optimisées. Les en-têtes d'objets compacts, introduits en JDK 24, sont désormais activés par défaut et réduisent l'empreinte mémoire du tas. Plusieurs previews sont reconduites : constantes paresseuses (3e preview), types primitifs dans les patterns (5e preview) et concurrence structurée (7e preview). Ajout d'une fonctionnalité de rédaction in-process pour JFR, afin de masquer les données sensibles dans les enregistrements de profiling. Kotlin 2.4 : nouveautés du langage et outillage kotlinlang.org/docs/whatsnew24.html Kotlin est un langage moderne, multiplateforme (JVM, Android, iOS, JavaScript, Wasm) développé par JetBrains, souvent utilisé comme alternative à Java. Les context parameters passent en stable : ils permettent de fournir des dépendances implicites à une fonction sans les déclarer en paramètre explicite, un peu comme une injection de dépendances. Les collection literals arrivent en expérimental : on peut écrire une liste avec des crochets, comme en Python, par exemple val fruits = ["pomme", "banane"]. L'API UUID de la bibliothèque standard devient stable, pour générer et manipuler des identifiants uniques nativement. Nouvelles fonctions utilitaires comme isSorted() pour vérifier si une collection est déjà triée. Support de Java 26 côté JVM et alignement automatique des versions Java et Kotlin dans les projets Maven. Kotlin/Native, la compilation vers du code natif iOS et macOS, active par défaut un nouveau ramasse-miettes plus rapide et améliore l'export vers Swift. Kotlin/Wasm, la compilation vers WebAssembly pour faire tourner du Kotlin dans le navigateur, rend la compilation incrémentale stable. Kotlin/JS permet désormais d'exporter des value classes vers JavaScript et TypeScript. Le compilateur K1, l'ancienne génération, n'est plus supporté : seul le nouveau compilateur K2 reste disponible. JEP 540 : une API JSON simple intégrée au JDK (incubation) openjdk.org/jeps/540 Le JDK ne propose aujourd'hui aucune API JSON native, obligeant à dépendre de bibliothèques externes comme Jackson, Gson ou Jakarta JSON pour parser ou générer du JSON. Cette JEP remplace la JEP 198 de 2014 et cible JDK 28 avec le nouveau module incubateur jdk.incubator.json. L'objectif est de couvrir les besoins simples d'extraction de données sans binding de données ni API de streaming, en laissant ces cas avancés aux bibliothèques existantes. L'API s'articule autour de l'interface scellée JsonValue avec six sous-types : JsonString, JsonNumber, JsonBoolean, JsonNull, JsonObject et JsonArray. Le parsing est strict et conforme à RFC 8259 : pas de virgules finales, pas de commentaires, et les noms de membres dupliqués provoquent une erreur. Donc pas de JSON5 La navigation se fait via get et tryGet, et la conversion vers des types Java via asInt, asLong, asDouble, asString, asMap ou asList. En cas d'erreur, une JsonValueException précise le chemin exact dans le document et sa position en ligne et colonne. Le pattern matching sur les sous-types de JsonValue permet de gérer proprement l'évolution du format d'un document JSON dans le temps. La génération se fait via toString pour une sortie compacte ou Json.toDisplayString pour une sortie indentée et lisible. À terme, le JDK pourrait utiliser cette API en interne, par exemple pour remplacer les fichiers de configuration au format property par du JSON. Autres nouvelles du JDK openjdk.org/jeps/535 openjdk.org/jeps/541 le mode generationel pour Shenandoah est prévu par défaut et deprécue le non générationel en 28 fini le support de Java sur Apple Intel GraalVM 25.2 : références compressées et Graal Script Agent medium.com/graalvm/… GraalVM est une machine virtuelle polyglotte d'Oracle offrant compilation JIT avancée et compilation en image native pour accélérer les applications Java et d'autres langages. Cette version 25.2 fait partie du train de releases innovation qui livre les nouveautés plus vite, pendant que GraalVM 25.0 reste la version stable recevant les correctifs de sécurité critiques. Nouveauté phare, le Graal Script Agent transforme des demandes en langage naturel en plugins sandboxés exécutés localement, en JavaScript ou Python, avec un accès restreint aux APIs de l'application. Les références compressées sont désormais activées par défaut dans Native Image sur les systèmes 64 bits, remplaçant les adresses complètes par des valeurs 32 bits relatives au tas. Cette optimisation réduit de 39% la consommation mémoire RSS d'une application Micronaut connectée à Oracle Database, comparée à la version 25.0. Contrepartie de cette optimisation, le tas géré est désormais plafonné à 32 Go. Le garbage collector G1 est maintenant disponible sur toutes les plateformes, y compris Windows, via l'option –gc=G1. G1 apporte de meilleures performances, une latence réduite et un démarrage plus rapide, avec des images natives plus petites grâce à l'optimisation guidée par profil. Le Vector API de Java est activé par défaut pour exploiter les instructions SIMD, utile pour le machine learning et le traitement de données. Bonne intégration avec l'écosystème via Micronaut 5.1, Quarkus et WebAssembly. Shopify arrête React Native pour ses applis mobiles iOS et Android et repasse à du natif avec Swift et Kotlin shopify.engineering/back-to-native Les progrès majeurs des LLM (IA) réduisent drastiquement le coût du développement sur deux plateformes distinctes. Les bénéfices du natif pur restent supérieurs, moins de couches d'abstractions, de dépenfances externes, et plus rapide pour adopter les dernières fonctionnalités des OS Les bibliothèques open-source (Skia, FlashList, Restyle) évoluent : Skia sera forkée par William Candillon, FlashList cherche un nouveau repreneur, Restyle sera archivée fin 2026. Migration des applications (Shop, Shopify, etc.) réalisée en mode "greenfield" (reconstruction totale) assistée par IA. Utilisation du système "Helix" pour un développement itératif et contrôlé par des agents IA. Découplage de la logique métier et de l'interface via une CLI pour accélérer les tests et éviter les lenteurs des simulateurs. L'application Shop a été entièrement reconstruite en natif en 12 semaines ; les autres suivront. Librairies LangChain4j CDI est une extension CDI qui intègre LangChain4j avec CDI de Jakarta EE langchain4j.github.io/langchain4j-cdi LangChain4j CDI : Extension intégrant LangChain4j à Jakarta EE et MicroProfile. Services IA : Injection et gestion de cycle de vie via @RegisterAIService. Orchestration d'agents : 11 topologies d'agents configurables par annotations. Serveur MCP : Conversion de beans CDI en serveurs Model Context Protocol. Fonctionnalités d'entreprise : Configuration externe, tolérance aux pannes et observabilité OpenTelemetry. Installation Maven : Deux extensions disponibles selon l'environnement (build-time pour Quarkus/Helidon, portable pour WildFly/GlassFish/Liberty). Prérequis techniques : Java 17+, Jakarta EE 10, MicroProfile 6.1. Quarkus 3.36, 3.37 et 3.38 : trois releases avant Quarkus 4 quarkus.io/blog/quarkus-3-38-released quarkus.io/blog/quarkus-3-37-released quarkus.io/blog/quarkus-3-36-released Quarkus est un framework Java cloud natif optimisé pour GraalVM et HotSpot, conçu pour les microservices et les environnements conteneurisés. En 3.38 (29 juillet), l'équipe allège les nouveautés pour se concentrer sur Quarkus 4, la communauté atteint 1213 contributeurs. 3.38 introduit l'éviction basée sur le poids mémoire pour le cache Caffeine de second niveau d'Hibernate, en plus de l'éviction par comptage. 3.38 apporte l'extension Quarkus HTTP Problem qui implémente la RFC 9457 pour mapper les exceptions en réponses application/problem+json, intégrée à OpenAPI. 3.37 (24 juin) ajoute l'extension expérimentale quarkus jlink pour générer des images runtime JDK sur mesure et réduire la taille des conteneurs. 3.37 active par défaut la sérialisation Jackson sans réflexion pour de meilleures performances. et 3.39 lesdesactivent et les rement en opt-in 3.37 introduit dans REST Client RestMultiResponse pour lire codes de statut et en-têtes sur des réponses REST en streaming, avec passage à Hibernate ORM 7.4 qui exige PostgreSQL 14 minimum. 3.36 (27 mai) propose Quarkus Signals en expérimental, un système de communication typée entre composants inspiré des events CDI et de l'EventBus Vert.x. 3.36 embarque des SBOM applicatifs exposés via /.well-known/sbom, y compris en image native selon la spécification GraalVM. 3.36 ajoute l'authentification OIDC via JWT SPIFFE, facilitant l'identité de charge de travail en environnement zero trust. Micronaut Framework 5.1.0 : injection de dépendances, IA et sécurité renforcées github.com/micronaut-projects/micronaut-platform/…/v5.1.0 Micronaut est un framework JVM pour microservices et applications cloud-natives, avec injection de dépendances à la compilation et démarrage rapide. Introduction d'Open DI 1.0.0, une implémentation CDI Lite s'appuyant sur l'infrastructure d'injection de dépendances de Micronaut. Côté données, support officiel de SQLite et intégration MyBatis, avec ETags basés sur les valeurs pour le verrouillage optimiste. En sécurité, arrivée d'un module OWASP HTML Sanitizer, délégation d'authentification @RunAs et résolution de locale via OIDC. Côté IA, LangChain4j ajoute le support Chroma, la mémorisation de chat Oracle et l'authentification Google injectée pour Vertex AI, avec passage du MCP en version 2.0.0. Mises à jour majeures des dépendances : Spring Boot 4.1.0, Jetty 12.1.10, Tomcat 11.0.23, OpenTelemetry 1.64.0 et Kubernetes Java Client 27.0.0. SSL activé par défaut par service pour les clients HTTP Infrastructure Ça coûte combien de faire tourner un LLM local sur son Apple Silicon ? towardsdatascience.com/how-much-does-a-local-llm-actually-cost-to-run-i-measured-every-watt-on-apple-silicon Coût électrique des LLM locaux sur Mac Apple Silicon Un modèle 120B (MoE) coûte 5x à 10x moins cher qu'un modèle 27B (Dense). Le coût dépend du débit (tokens/seconde), pas du nombre de paramètres. Modèle dense –> Charge 100% des poids par token = lent et très énergivore. MoE (Mixture of Experts) –> N'active qu'une fraction des poids = rapide et économe. Conclusion : Pour réduire la facture électrique, choisir des modèles MoE quantifiés (haut débit). Kubernetes 1.36 (Haru) : sécurité renforcée et alignement IA infoq.com/news/2026/05/kubernetes-1-36-released Kubernetes est la plateforme open source de référence pour l'orchestration de conteneurs, portée par la CNCF. La version 1.36 nommée Haru apporte 70 améliorations : 18 passent stables, 25 en bêta et 25 en alpha, avec 106 entreprises et 491 contributeurs. Les user namespaces passent en disponibilité générale, isolant le root du conteneur de celui de l'hôte. Les Mutating Admission Policies passent en GA, remplaçant les webhooks par des règles CEL natives plus performantes. L'autorisation de l'API kubelet devient plus fine, remplaçant le droit trop large nodes/proxy. Le labeling SELinux des volumes utilise désormais mount -o context, accélérant le démarrage des pods. Plusieurs avancées ciblent les charges IA : gang scheduling en bêta, préemption consciente des groupes de pods, et allocation dynamique de ressources activée par défaut pour le partage fin des GPU. Le redimensionnement vertical des pods en place passe en bêta et activé par défaut, ajustant CPU et mémoire sans redémarrage. Suppression du plugin gitRepo, source de risque de sécurité, et du mode IPVS de kube-proxy, tous deux dépréciés de longue date. Avec la sortie de 1.36, la version 1.34 devient la plus ancienne branche encore supportée et entre en maintenance, ne recevant plus que des correctifs critiques avant sa fin de support. Terraform vs OpenTofu en 2026 : la divergence est actée ecorpit.hashnode.dev/terraform-vs-opentofu-in-2026-the-fork-has-diverged-so-which-do-you-standardize-on env0.com/insights/opentofu-in-2026-what-the-terraform-fork-became-after-three-years-of-independence Terraform est l'outil historique d'Infrastructure as Code de HashiCorp, OpenTofu en est le fork open source lancé après le changement de licence. HashiCorp est passé de la licence MPL 2.0 a la BUSL 1.1 en aout 2023, ce qui a poussé une partie de la communauté a créer OpenTofu sous la Linux Foundation. IBM a racheté HashiCorp pour 6,4 milliards de dollars, finalisé en février 2025, tandis qu'OpenTofu rejoignait le CNCF comme projet sandbox en avril 2025. OpenTofu prend de l'avance sur des fonctionnalités inédites : chiffrement du state côté client depuis la v1.7, valeurs éphémères qui gardent les secrets hors du state depuis la v1.11, et prevent_destroy dynamique en v1.12 (mai 2026). Terraform garde l'avantage sur l'orchestration managée avec Terraform Stacks, désormais en disponibilité générale, sans équivalent natif côté OpenTofu. Le coût diverge fortement : HCP Terraform facture jusqu'à 0,99 dollar par ressource gérée et par mois, alors qu'OpenTofu reste une CLI gratuite couplée au backend de son choix. Fidelity Investments a migré plus de 2000 applications et 50000 fichiers d'état vers OpenTofu, la complexité venant surtout de l'écosystème (CI/CD, gouvernance) plutôt que du binaire lui-même. OpenTofu reste compatible avec les configurations Terraform jusqu'a la version 1.6.x, mais les versions Terraform plus récentes n'offrent plus aucune garantie de compatibilité. Pour les secteurs régulés, le chiffrement natif du state par OpenTofu et sa gouvernance ouverte sont des arguments forts face aux exigences de protection des données. La recommandation qui ressort des deux articles : partir sur OpenTofu pour les projets neufs et rester sur Terraform si l'on est déjà investi dans HCP Terraform et ses fonctionnalités de gouvernance. OTel est à la peine ? matduggan.com/otel-isnt-going-well-and-i-made-a-spreadsheet-about-it Le développement d'OTel est un peu au point mort Périmètre démesuré : Volonté de supporter un nombre gigantesque de langages, bibliothèques et frameworks. Pénurie critique de mainteneurs : Les données montrent une hyper-concentration du travail ; de nombreux SDK (comme PHP ou Ruby) dépendent d'une ou deux personnes seulement. Stabilité paralysante : La règle interdisant toute modification d'une fonctionnalité déclarée « stable » crée une peur de valider les nouveautés, entraînant des mois de débats. Solutions proposées par l'auteur : Créer un niveau « Bêta » temporaire (ex: 12 mois) entre les statuts « Expérimental » et « Stable ». Faire preuve de transparence sur les différences de qualité/maintenance entre les langages (ne pas mettre Go et Ruby sur le même plan). Communiquer activement sur le besoin urgent de nouveaux mainteneurs. Assouplir la politique de stabilité en acceptant des breaking changes bien documentés. Honeycomb transforme son infrastructure Kafka honeycomb.io/blog/transforming-how-we-run-kafka-honeycomb Honeycomb est une plateforme d'observabilité dont Kafka est le coeur du pipeline d'ingestion, traitant des millions d'événements par seconde. L'entreprise a migré de Confluent Platform auto-hébergée vers Apache Kafka 4.1.1 en mode KRaft. Le nouveau cluster tourne sur AWS EKS avec Strimzi comme couche d'orchestration Kubernetes. Motivation principale : la récupération après remplacement de broker était passée de 8-12h à 48-72h avec l'ancienne stack. Confluent imposait aussi sa solution propriétaire de Tiered Storage, impossible à corriger en interne. Un incident de décembre 2025 ayant vidé un cluster a révélé une fenêtre d'opportunité pour migrer. La migration s'est faite progressivement sur six clusters, de dogfood jusqu'à la production. Les producteurs sont basculés avant les consommateurs, avec une courte fenêtre de downtime assumée entre les deux. Le stockage utilise des NVMe en instance store plutôt que de l'EBS pour minimiser la latence. interessant de voir une société reprendre en main sa compétence et de voir les contraintes de certaines fonctionalités propriétaires Cloud AWS us-west-2 : panne réseau régionale et effet domino chez les fournisseurs SaaS blog.incidenthub.cloud/aws-us-west-2-outage-jul-24-2026 AWS us-west-2 (Oregon) est une région cloud majeure hébergeant de nombreux services et fournisseurs SaaS. Le 24 juillet 2026, une panne matérielle réseau a coupé la connectivité entre la région et le Seattle Metro pendant environ 20 minutes. Particularité notable, seule la couche de connectivité externe a été touchée, le trafic interne à la région a continué de fonctionner normalement. Après la réparation matérielle, une phase distincte de reconvergence des routes a de nouveau causé une connectivité intermittente pendant plusieurs dizaines de minutes. Les clients Direct Connect via EqSe2 ont subi une coupure bien plus longue que le reste, 1h17 au total. Neuf incidents chez sept fournisseurs ont cité explicitement AWS comme cause, dont SendGrid, SparkPost et NinjaOne. Fait marquant, les temps de rétablissement des fournisseurs tiers ont largement dépassé la durée de la panne AWS elle même. NinjaOne a mis 9h29 à se rétablir totalement, avec 150000 appareils tentant de se reconnecter simultanément freinés par des mécanismes de backoff et jitter. SparkPost a mis 6h45 à absorber l'arriéré de courriels accumulé pendant la coupure, avec encore 80 à 90 minutes de retard des heures plus tard. L'article recommande d'identifier ses dépendances en us-west-2 et de prévoir capacité et bascule, l'effet différé pouvant durer bien plus longtemps que l'incident initial. Rapport Cloudflare Radar sur les perturbations Internet au Q2 2026 blog.cloudflare.com/fr-fr/q2-2026-internet-disruption-summary Cloudflare Radar est la plateforme qui analyse en temps réel le trafic mondial pour détecter pannes, coupures et censures Internet. l'instabilité est le nouveau normal Le super-typhon Sinlaku a fait chuter le trafic de près de 80% à Guam les 13 et 14 avril. Deux séismes de magnitude 7,5 ont fortement dégradé la connectivité au Venezuela le 24 juin. Une coupure électrique a provoqué cinq heures de perturbation en Tanzanie le 27 juin. En Iran, la connectivité s'est stabilisée à 59% du niveau normal après 88 jours de coupure. Le Soudan a imposé dix coupures programmées pendant les examens nationaux mi-avril. L'Irak a coupé Internet à trois reprises pour lutter contre la fraude aux examens. Des frappes de drones ont endommagé la région AWS me-central-1 aux Émirats arabes unis. Un renouvellement de clés DNSSEC a rendu les sites .de inaccessibles en Allemagne le 5 mai. Une rupture de câble sous-marin a fait chuter le trafic de 60% à Sainte-Lucie fin juin. l'instabilité est le nouveau normal Un ingénieur Google débranche une zone entière de Google Cloud https://www.theregister.com/off-prem/2026/09/04/google-engineer-unplugged-every-fiber-they-could-see-and-surprise-took-down-a-chunk-of-the-g-cloud/5294418 Google Cloud est la plateforme d'infrastructure cloud de Google, organisée en régions et zones de disponibilité comme us-central1. Le 1er septembre 2026, un ingénieur a débranché par erreur des câbles fibre optique lors d'une opération de maintenance matérielle routinière dans la zone us-central1-b. En 13 minutes, il a déconnecté 100 % des chemins de fibre optique de tous les équipements de cette portion de la zone. Les machines virtuelles hébergées dans la zone touchée sont devenues injoignables, avec une perte de paquets élevée. Le taux de chute du trafic réseau pour les ressources concernées a atteint 100 %. L'incident a duré 4 heures et 11 minutes, de 7h41 à 11h52 heure du Pacifique. Google a détecté l'anomalie, reroute le trafic, identifié les liaisons optiques débranchées puis rebranché physiquement les fibres avant le retour à la normale. Les post-mortems d'incident cloud sont un classique du podcast, mais l'erreur humaine sur du câblage physique chez un hyperscaler mérite une minute d'antenne. ChatGPT, Claude et Grok en panne presque simultanément le 3 septembre theregister.com/ai-and-ml/…/5294322 ChatGPT, Claude et Grok sont les assistants IA et coding agents désormais utilisés au quotidien par de nombreux développeurs. xAI loue ses datacenters à un certains nombre de fournisseurs de cloud et d'IA. ils ont fait tombé ses concurrents :slightly_smiling_face: Arreter là Le 3 septembre 2026, les trois services sont tombés en panne quasiment en même temps. ChatGPT a connu une panne de 7h43 à 8h17 PT, soit environ 34 minutes, due à une erreur de routage rendant ChatGPT et Codex indisponibles. Claude a subi une panne partielle de 3 heures et 6 minutes touchant Claude.ai, Claude Code, Claude Cowork et l'API Claude, résolue à 16h16 UTC. xAI a commencé à enquêter sur les problèmes de Grok dès 6h30 PT, puis SpaceX a confirmé une panne de son centre de calcul de Memphis. La coïncidence des trois pannes a fait suspecter un fournisseur commun à l'origine du problème. Pour les développeurs devenus dépendants de ces coding agents, l'épisode illustre le risque d'un point de défaillance unique xAI loue ses datacenters à un certains nombre de fournisseurs de cloud et d'IA. Web Nouveautés CSS 2026 : mixins, masonry natif et animations pilotées par le scroll modern-css.com/whats-new-in-css-2026 animation-timeline: scroll() et view() atteignent le baseline cross-browser, Firefox et Safari ayant livré le support complet. Plus besoin de préfixes ni de librairie JS. @starting-style devient cross-browser : animations d'entrée depuis display: none sans hack de timing JS. Firefox 147 amène l'anchor positioning au baseline, plus les view transition types et la Navigation API. Les menus contextuels via popover CSS arrivent aussi. Data et Intelligence Artificielle Guillaume a porté le SDK Python d'Antigravity en Java… en utilisant Antigravity lui même comme assistant ! glaforge.dev/posts/…/the-unofficial-antigravity-sdk-for-java SDK Java non officiel pour Antigravity, rétro-ingénierie du SDK Python pour exploiter un binaire Go sous-jacent. Cas d'usage : pipelines CI/CD, applications d'entreprise (Spring Boot), outils internes, surveillance en arrière-plan, interfaces personnalisées. Gestion des ressources : implémente AutoCloseable(try-with-resources) pour lancer et fermer proprement le processus Go. Exécution de code Java personnalisé : exposition de méthodes Java comme outils IA via les annotations @Toolet@Param. Streaming et programmation réactive : prise en charge des CompletableFuture, des callbacks (chatStream), et deFlow.Publisher. Fonctionnalités avancées : persistance de session, protocole MCP, politiques de sécurité, entrées multimodales, et sorties structurées mappées sur des records Java. Le SDK Java pour le protocol Agent2Agent sort sa version 1.2.0 medium.com/google-cloud/a2a-java-sdk-1-2-0-final-released… Prise en charge de la spécification A2A 1.0. Sécurité renforcée : Vérification stricte des autorisations de lecture sur les tâches référencées et implémentation d'une logique de blocage par défaut (fail-closed). Intégration facilitée : Support du câblage programmatique des autorisations pour les environnements non-CDI (comme Spring). Contrôle des flux : Ajout du Task Stream Lifecycle Hook pour surveiller et gérer le cycle de vie des abonnements aux flux d'événements. Stabilité des données : Immutabilité stricte imposée sur les enregistrements de spécifications pour empêcher toute modification accidentelle. Documentation : Lancement d'une documentation web multi-versions et d'un Javadoc agrégé pour tous les modules. Corrections de bugs : Résolution de problèmes liés à la synchronisation des tâches, aux réponses de streaming et aux conditions de concurrence HTTP. Breaking changes : Nécessite une migration suite à la modification de certaines méthodes d'autorisation, la réorganisation de packages et le renommage de l'état TaskState.UNRECOGNIZED en TaskState.TASK_STATE_UNSPECIFIED. Article sur le blog de JetBrains : blog.jetbrains.com/idea/2026/08/intellij-idea-goes-lsp pas trop de temps Langchain4j continue sa progression github.com/langchain4j/langchain4j/…/1.18.0 github.com/langchain4j/langchain4j/…/1.19.0 github.com/langchain4j/langchain4j/…/1.17.0 Introduction du pattern Debate pour lancer des sous agents dans rechercehnt et entrent dans un debat critique sur une decision avant qu'un juge decide (1.17) compensation d'action d'un outil avec @ReverseTool (1.17) Ajout du pattern Belief-Desire-Intention (1.18): belief est l'état du monde cru, desire est la liste des objectifs, intention est le plan pour avancer un ou plusieurs objectif Support d'un systeme crash resilient dans l'approche Human in the loop avec des checkpoints nestés (1.18) Support Open AI TextToSpeech (1.18) Support de Mistral batch chat (1.18) support MCP client 2026-07-28 (1.19) Anthromic batch chat et Google thinking mode (1.19) Embabel atteint 1.0 GA github.com/embabel/embabel-agent/…/v1.0.0 embabel d'eloigne de Spring AI ne s'appuie que sur Spring (notamment tool calling) nettoyage des explorations autour du pattern GOAL avant a 1.0 ajout observabilite (dont le cout) les approches de retry solidifiées et d'autres choses toujours basées sur la description de goals typés et des dependences entre eux via les types MCP 2026-07-28 : le protocole devient stateless blog.modelcontextprotocol.io/posts/2026-07-28 MCP (Model Context Protocol) est le protocole standard permettant aux LLM et agents IA de communiquer avec des outils, ressources et serveurs externes. Cette version marque le plus gros changement depuis le lancement du MCP distant il y a 18 mois. Le protocole passe d'un modèle bidirectionnel avec état à un modèle stateless en requête/réponse. Suppression de la poignée de main initialize/initialized et du header Mcp-Session-Id, chaque requête devient autoportante. N'importe quelle requête peut désormais être routée vers n'importe quelle instance de serveur derrière un load balancer classique, sans stockage partagé. Introduction des Multi Round-Trip Requests (MRTR) pour remplacer les requêtes initiées par le serveur, via un resultType input_required et des inputResponses. Nouveaux headers Mcp-Method et Mcp-Name pour permettre aux gateways de router et autoriser sans parser le JSON. Les résultats de tools, prompts et resources deviennent cacheables grâce aux paramètres ttlMs et cacheScope. Renforcement sécurité avec la validation d'issuer RFC 9207 pour éviter les attaques de confusion entre serveurs d'autorisation, et transition de DCR vers CIMD. Roots, Sampling et Logging sont dépréciés avec douze mois de support garanti, tout comme le transport legacy HTTP+SSE. Un site qui référence les skills pour la JVM (framework, langage, build…) jvmskills.com Frameworks : Spring, Quarkus, Jakarta EE, Reactor, Camel Java : bonne pratiques, conventions, guides de mise à jour à niveau LTS, API spécifiques (streams, optionals, logs…) Bases de données : ORM, validation, modélisation PostgreSQL, vectorielle avec pgvector Tests et qualité : TDD, mutation testing, debogage avec JDB Workflows dev et archi : commits git, domain modeling Outils et diagnostics JVM : JFR, Jstall, JSpecify Une skill n'est pas une librairie https://devx.writizzy.blog/p/un-skill-nest-pas-une-lib Les skills sont des éléments de configuration en prose pour agents IA comme Claude, distribués via des marketplaces à la manière de librairies logicielles. Frédéric Camblor critique cette analogie car partager un skill n'est pas la même chose que le mutualiser durablement. Écrire un skill prend 30 minutes mais l'adopter ailleurs coûte cher en appropriation et en maintenance. Forker un skill s'avère souvent plus efficace que de chercher à converger vers une version commune. Contrairement au code, les régressions d'un skill ne sont pas détectables automatiquement. Un skill peut se dégrader silencieusement sur plusieurs cas d'usage en corrigeant un autre. Les skills vieillissent vite car les modèles progressent et intègrent naturellement certaines bonnes pratiques. Le skill-creator d'Anthropic permet d'évaluer un skill via des jeux de cas et des mesures de variance. L'auteur distingue quatre sphères de partage : personnelle, équipe, outil et marketplace. Il propose un cycle partage puis appropriation puis duplication puis divergence plutôt qu'une installation collective figée. Les modèles Anthropic introduisent un filigrane (watermark) dans les textes qu'ils génèrent https://www.anthropic.com/news/claude-text-watermark Claude est l'assistant IA d'Anthropic, et le watermarking est une technique permettant de marquer discrètement un contenu généré par IA pour en tracer l'origine. Anthropic annonce que les futurs modèles Claude intégreront un filigrane numérique invisible dans le texte généré. Le principe exploite les choix de mots équivalents que le modèle fait naturellement, en les orientant via une clé cryptographique plutôt qu'un tirage aléatoire. Le texte produit reste indiscernable à l'œil nu, sans caractères cachés, sans ralentissement ni coût supplémentaire. Seule la personne possédant la clé correspondante peut détecter la présence du filigrane. Le filigrane est plus fiable sur les textes longs et créatifs, moins sur du texte factuel, du code ou après une édition manuelle poussée. Il ne prouve pas qu'un texte est écrit par IA, ni n'identifie l'auteur ou la conversation d'origine, il donne seulement une probabilité d'implication de Claude. Une API de détection est proposée en accès restreint aux régulateurs, forces de l'ordre, médias et vérificateurs de faits. Pour les fichiers non textuels comme les images ou les PDF, Anthropic s'appuie sur le standard C2PA. Cette initiative s'inscrit dans le Code de Pratique de l'UE sur la transparence des contenus IA, signé par Anthropic et environ 190 autres acteurs, en lien avec la loi européenne sur l'IA. GPT-6 Astra, le nouveau modèle d'OpenAI face à Claude Fable 5.1 https://openai.com/index/gpt-6-astra/ GPT-6 Astra est le nouveau modèle phare d'OpenAI, annoncé le 3 septembre 2026 comme le plus intelligent et le plus aligné de l'entreprise. Le modèle arrive deux jours après Claude Fable 5.1, à un tarif affiché comparable, dans une course accélérée aux modèles de code et de raisonnement. Astra revendique 98 % sur FrontierMath Tier 4, 99,9 % sur ARC-AGI-3 et 100 % sur ExploitBench. Fenêtre de contexte d'environ 1,05 million de tokens. Tarification API, 10 dollars par million de tokens en entrée, 50 dollars en sortie, et 1 dollar par million pour les tokens en cache. Au-delà de 272 000 tokens en entrée, toute la requête est facturée au double sur l'entrée et une fois et demie sur la sortie. Astra est le premier modèle d'OpenAI à franchir le seuil interne critique en cybersécurité. La version publique refuse les tâches offensives avancées comme générer des preuves de concept d'exploits. Le déploiement est progressif, les entreprises du programme de cybersécurité Daybreak d'OpenAI y accèdent en premier, avant ChatGPT Plus, Pro, Business, Enterprise, l'API et AWS. Même logique de diffusion contrôlée que chez Anthropic avec Mythos 5.1 : deux jours d'écart, deux modèles de tête, et la même question de savoir qui accède en premier aux capacités les plus sensibles. Outillage JetBrains s'est lancé dans les LSP (Language Server Protocol) avec une extension IntelliJ pour VS Code et assimilés marketplace.visualstudio.com/items?itemName=JetBrains.intellij-s… Nouveau produit : Lancement de l'extension Java & Kotlin by IntelliJ IDEA pour les éditeurs basés sur VS Code (incluant Cursor). Technologie : Utilisation du standard LSP (Language Server Protocol). Objectif : S'adapter au développement piloté par les agents IA, qui nécessite des fonctionnalités IDE légères et standardisées. Fonctionnalités clés : Support des projets Java, Kotlin et mixtes. Débogage (DAP). Complétion intelligente, navigation et analyse de code. Refactoring. Prise en charge de Maven, Gradle et Bazel. Disponibilité : Téléchargeable via le Visual Studio Marketplace et l'Open VSX registry. Licence / Prix : Gratuit durant la phase de preview (évaluation renouvelable de 30 jours). Nécessitera un abonnement IntelliJ IDEA Ultimate après la preview. (Note : Le LSP purement Kotlin reste gratuit et open-source). Avenir : Développement en cours pour optimiser les flux de travail avec les agents IA en ligne de commande (ex: Claude Code, Codex) afin de réduire la consommation de tokens. Après son acquisition par SpaceX, Cursor perd l'accès aux modèles OpenAI https://openai.com/index/our-decision-on-cursor-following-its-acquisition-by-spacex/ decision difficile mais Elon ment comme un arracheur de dent (et c'est un competiteur donc bon ça nous arrange) David Pilato a créé un thème spéciale pour les gens pour les devrel, ou qui font des talks à droite à gauche, pour le moteur Hugo https://david.pilato.fr/posts/2026-09-07-hugo-theme-devrel/ hugo-theme-devrel, thème Hugo (MIT) pour Developer Advocates et conférenciers, fonctionnant comme un module superposé au thème Dream. Gestion des conférences (cartes Leaflet), présentations (PDF, YouTube, co-auteurs), vues dédiées (archives, sujets récurrents, vidéos) et recherche Pagefind. Chaque intervention est un page bundle YAML structuré comme une base de données relationnelle compilée par Hugo. Architecture Airbnb refond son authentification en architecture server-driven, 60 % de code en moins https://www.infoq.com/news/2026/09/airbnb-server-driven-login/ Airbnb a restructuré son authentification autour d'un modèle en deux phases, identification du compte (email, téléphone ou connexion sociale) puis choix du challenge d'authentification décidé côté serveur selon le contexte utilisateur. Ce qui est interessant c'est que choisir la method d'authentification est côté serveur et adaptative, comme si c'était une révolution Arreter là Un moteur de politique serveur sélectionne la méthode d'authentification optimale avec des solutions de repli, permettant des adaptations régionales comme l'OTP WhatsApp au Brésil ou des fournisseurs d'identité locaux en Corée du Sud sans nouvelle version client. Un Challenge Picker propose des méthodes alternatives classées par probabilité de succès en cas d'échec. Résultat chiffré, 60 % de code d'authentification en moins et 100 Ko de moins sur le bundle client web. Méthodologies Les nouvelles règles d'ingénierie du contexte pour les modèles Claude 5 x.com/trq212/status/2080710971228918066 Partage par Thariq des apprentissages sur l'ingénierie du contexte et le prompt engineering pour les nouveaux modèles Claude 5 (comme Claude Opus 5 et Claude Fable 5) utilisés dans Claude Code. Évolution majeure vers le dés-empirement (unhobbling) : plus de 80 % du prompt système de Claude Code a pu être supprimé sans perte sur les évaluations de code, les modèles récents faisant preuve d'un bien meilleur jugement contextuel. Passage des règles strictes au jugement : au lieu d'interdire les commentaires ou d'imposer des contraintes lourdes, les modèles s'adaptent désormais au code environnant et font appel à leur propre discernement. Remplacement des exemples par la conception d'interfaces : fournir des exemples figés restreint l'exploration du modèle, d'où l'importance de concevoir des outils et des fichiers plus expressifs. Adoption de la divulgation progressive (progressive disclosure) : chargement dynamique du contexte (via des compétences ou des outils à chargement différé comme ToolSearch) pour éviter de saturer la fenêtre de contexte avec des instructions fixes. Utilisation d'une mémoire automatique et de références riches (artefacts HTML, suites de tests, fonctions de référence) plutôt que de fichiers CLAUDE.md pléthoriques ou de consignes répétitives. Recommandation pour les fichiers CLAUDE.md et les Skills : les garder légers, se concentrer sur les pièges spécifiques (gotchas) du dépôt, et structurer les guides sous forme d'arborescences modulaires pour ne charger que le nécessaire. Niveau d'adoption des agents IA de codage selon une étude de JetBrains blog.jetbrains.com/research/2026/08/ai-coding-agent-adoption-2026 Adoption massive : 90 % des développeurs professionnels utilisent des agents d'IA de codage au moins une fois par semaine, et 68 % quotidiennement. Claude Code domine : Il devient le nouveau leader du marché avec 39 % d'adoption mondiale (47 % aux États-Unis), détrônant largement ses concurrents. Déclin de GitHub Copilot : L'ancien leader perd de sa superbe, passant de 29 % à 21 % d'adoption, bien qu'il conserve une très forte notoriété (79 %). Percée de Codex : Sa croissance est fulgurante, son taux d'adoption ayant été multiplié par 5 en quelques mois (de 3 % à 16 %). Recul de Cursor : L'outil connaît une légère baisse, passant de 18 % à 12 % d'adoption, principalement due à une forte chute sur le marché chinois. Écosystème diversifié : JetBrains AI atteint 9 % d'adoption. Des alternatives comme OpenCode (7 %) et Google Antigravity (6 %, mais très populaire en Inde à 15 %) continuent de s'implanter. L'IA a cassé les hypothèses de la CI… ou pas https://stack72.dev/ai-broke-the-assumptions-behind-ci/ Paul Stack (ex-Pulumi) explique que l'intégration continue a toujours mêlé deux rôles distincts, exécuter la vérification du code et coordonner les fusions. avec les agents, la pression sur la CI augmente, car ils ne testent pas end to end tout le temps ils ont maintenant des workflows qui verifient tests, lint, revue d'agent etc dans un env local isolé donc c'est pre PR push Il propose de séparer vérification (exécution), attestations structurées (hash de commit, checksums SHA256) et CI, réduite à la validation et à la coordination des fusions. Sa thèse, les agents IA peuvent désormais vérifier tout le code localement avant même l'ouverture d'une pull request, ce qui bouleverse cet équilibre. l'attestation vient ensuite et la CI est une étape de vérification (attestation de commit et de tests et si branche a bougé, repart à l'execution) Martin Fowler, qui a repéré l'article dans ses Fragments du 1er septembre 2026, réplique que la vraie CI a toujours exigé une vérification locale avant de pousser le code. ca demande une chaine d'attestation forte quid de garder les metriques historique de CI Sécurité France Passoire: une analyse sur les différents vols de données des services publics de ces dernières mois https://www.cybernetica.fr/piratage-des-impots-comment-en-est-on-arrive-la/ Analyse du piratage massif de la DGFiP et d'autres administrations françaises en 2026, révélateur de failles systémiques de cybersécurité de l'État. Intrusion détectée fin juin à la DGFiP, mais l'exfiltration de 678 000 entrées fiscales n'a été découverte qu'en août lors de leur mise en vente. Données volées : noms, revenu fiscal de référence, taux de prélèvement, adresse, téléphone. Une seconde attaque du même pirate a visé le cadastre fin juillet, exposant plus de 2 millions de personnes. L'Éducation nationale a aussi été piratée fin juillet, données de tous les agents depuis 2001 exposées. Cause principale : modèle de sécurité fondé sur le périmètre physique plutôt que sur le zero trust, sans contrôle après authentification. Le télétravail post Covid a étendu les accès distants sans reconstruire les modèles de confiance. Aucun système de détection d'exfiltration n'existait, la fuite n'a été révélée que par le pirate lui-même. La transposition de la directive NIS2 est bloquée en France depuis septembre 2025, la CJUE a condamné le pays à des astreintes. L'article souligne un désengagement croissant des Etats-Unis en matière de cybersécurité internationale et une dépendance technologique accrue de la France. Loi, société et organisation Ce que l'IA change vraiment au métier de manager shapeandship.ai/p/ce-que-lia-change-vraiment-au-metier-de-manager Retour d'expérience et analyse par Mathilde Rigabert sur l'impact réel de l'IA générative dans le quotidien d'un Engineering Manager. L'IA excelle pour automatiser la reconstitution factuelle de l'activité (lecture de PRs, commits, reviews) nécessaire aux 1:1 et entretiens annuels, mais elle offre une vision uniquement quantitative et nécessite d'être croisée avec des notes de terrain. L'IA rend le maintien de la qualité et des standards plus difficile : selon une étude Faros AI sur 22 000 développeurs, les PRs mergées sans aucune revue ont augmenté de 31 %, fragilisant la compréhension commune apportée par le pairing et les revues de code. Le temps gagné par l'IA ne permet pas d'augmenter massivement le span of control (seulement 2 ou 3 personnes de plus), car l'IA compresse la collecte d'informations mais pas les conversations humaines complexes ou l'accompagnement du changement. Les compétences d'orchestration et de gestion de sujets multiples acquises par les managers facilitent leur transition vers le pilotage de plusieurs agents IA en contribution individuelle. Le piège actuel réside dans l'accumulation des casquettes (manager, tech lead, product owner, contributeur, pompier), conduisant à l'épuisement et au délaissement du travail de fond sur l'organisation et l'humain. Le temps libéré par l'IA doit être réinvesti dans le travail invisible qui fait tenir le système (suivi des actions de rétro, analyse de métriques, coaching), que personne ne réclame à court terme mais dont l'absence fragilise les équipes à long terme. Je regrette d'avoir migré vers Codeberg xn–gckvb8fzb.com/i-regret-migrating-to-codeberg L'auteur explique pourquoi il regrette d'avoir quitté GitHub pour Codeberg, à la suite des récentes modifications des conditions d'utilisation (ToS) de la plateforme. Codeberg a interdit les projets principalement générés par des LLM ainsi que les projets liés aux cryptomonnaies via des propositions de l'Assembly 2026, au motif qu'ils nuisent à sa réputation. blog.codeberg.org/protecting-our-floss-commons-from… Critique de l'argument de Codeberg sur l'absence de communauté des vibe coders, en rappelant que la majorité des logiciels libres (FOSS) sont créés par des développeurs solos sans communauté au sens romancé du terme. Ironie soulignée concernant la posture de Codeberg et Forgejo, qui a hérité de la communauté de Gitea après un hard fork avant de faire la leçon aux développeurs individuels. Alerte sur le risque de censure idéologique : interdire des catégories entières plutôt que de traiter les abus réels ou la consommation d'infrastructure crée un précédent dangereux pour une forge qui se veut libre. Proposition de solutions alternatives pour gérer l'impact des LLM et de la crypto : déclaration obligatoire via des cases à cocher, hébergement sur des tiers d'infrastructure spécifiques payants ou sous quotas, et disclaimers automatiques. Décision de l'auteur de quitter Codeberg pour mettre en place son propre serveur Git personnel afin d'éviter la dépendance à une plateforme qui modifie ses règles de manière unilatérale. Cloud souverain : Airbus choisit Scaleway pour l'hébergement de ses applications critiques https://www.usine-digitale.fr/aeronautique-spatial/airbus/cloud-souverain-airbus-choisit-scaleway-pour-lhebergement-de-ses-applications-critiques.OHBVBMSZIJELNOKN6G5F6B7NSI.html Scaleway est le cloud provider français filiale du groupe Iliad, positionné comme alternative souveraine aux hyperscalers américains. Airbus a lancé un appel d'offres de six mois consultant une cinquantaine d'acteurs dont OVHcloud, Thales, Google S3NS et Microsoft Bleu. Scaleway a été retenu pour héberger les applications critiques liées à la conception d'aéronefs, l'ingénierie, la production industrielle et les opérations. Le contrat prévoit la migration d'environ 70 applications d'ici 2028, puis jusqu'à 900 applications sur 5 à 6 ans. Le montant du contrat n'a pas été communiqué. Scaleway revendique zéro actionnaire, zéro employé et zéro filiale hors Union européenne pour garantir une protection contre les lois extraterritoriales. Damien Lucas, PDG de Scaleway, évoque une immunité complète face aux évolutions politiques et législatives externes. La plateforme doit aussi accélérer les usages d'intelligence artificielle d'Airbus, avec les modèles de Mistral AI déjà déployés chez Scaleway. Catherine Jestin, responsable numérique d'Airbus, souligne que cette intégration accélère la démarche IA du groupe. Ce choix ne remet pas en cause la stratégie multicloud d'Airbus, Scaleway venant compléter les fournisseurs existants pour les charges nécessitant le plus haut niveau de gouvernance et de résilience. Debian adopte une résolution sur l'usage responsable de l'IA générative lwn.net/Articles/1091231 La discussion sur l'usage des LLM dans Debian s'est tenue du 23 juillet au 13 août 2026, suivie d'un vote du 15 au 28 août 2026. 1045 développeurs Debian étaient éligibles à voter, avec un quorum de 48,49 votes largement dépassé par les huit options en lice. Les options allaient d'une interdiction stricte des contributions générées par LLM inscrite dans le contrat social à une acceptation encadrée des contributions IA. L'option gagnante au classement Condorcet est Responsible Use of Generative AI, devant Allow AI-Assisted Contributions with conditions et A cautious approach to generative AI. Le texte adopté n'interdit ni n'encourage l'usage d'outils d'IA générative dans le développement de Debian. Il exige que toute contribution, quels que soient les outils utilisés pour la produire, respecte les mêmes standards de qualité, correction, maintenabilité et conformité légale. Les contributeurs doivent comprendre, relire, tester et si besoin modifier la production assistée par IA avant de l'intégrer à Debian. Les informations sensibles du projet ne doivent pas être transmises à des fournisseurs d'IA non fiables, et la divulgation de l'usage de l'IA est encouragée sans être obligatoire. Le détail du vote et le texte complet de la résolution sont disponibles sur la page officielle [debian.org/vote/2026/vote_002](https://www.debian.org/vote/2026/vote_002). Contraste direct avec l'OpenJDK, qui a publié une politique interdisant le code généré par LLM (épisode 340), et avec l'auteur de jqwik qui a piégé sa librairie contre les agents. Conférences Nous contacter Pour réagir à cet épisode, venez discuter sur le groupe Google https://groups.google.com/group/lescastcodeurs Contactez-nous via X/twitter https://twitter.com/lescastcodeurs ou Bluesky https://bsky.app/profile/lescastcodeurs.com Faire un crowdcast ou une crowdquestion Soutenez Les Cast Codeurs sur Patreon https://www.patreon.com/LesCastCodeurs Tous les épisodes et toutes les infos sur https://lescastcodeurs.com/
We've been hosting a live show weekly on Mondays at 5p for about an hour, and recording them all; here is the recording.How do you test configurations you haven't built yet? Steve, Andrew, and Rain join Bryan and Adam to talk about the transformative power of the Oxide emulation environment, VOXEL, and how it has accelerated development in many domains.In addition to Bryan Cantrill and Adam Leventhal, speakers included Steve Karam, Andrew Stone, and Rain Paharia.Previously, on Oxide and Friends:OxF s03e06 - Rack-scale Networking (P4 and DDM with Ry Goodfellow)OxF s01e26 - The Pragmatism of Hubris (with Cliff Biffle)OxF s03e09 - Get You a State Machine for Great Good (Wicket with Andrew Stone)OxF s01e25 - Tales from the Bringup LabOxF s01e18 - Dijkstra's TweetstormOxF s03e31 - Hiring Processes with Gergely OroszSome of the topics we hit on, in the order that we hit them:voxel -- Virtual OXide Emulation Lab: emulated rack deployments on one Helios host (omicron on Falcon/propolis VMs, SoftNPU switches, FRR routers); optional real SP/RoT firmware via sp-emu. https://github.com/oxidecomputer/voxelFalcon -- tool for standing up virtual network/rack topologies of propolis VMs on illumos for testing. https://github.com/oxidecomputer/falconTofino -- Intel/Barefoot programmable switch ASIC (Tofino 2) in Sidecar; P4-programmable. Discontinued by Intel.SoftNPU -- software P4 target emulating the switch dataplane, so the stack runs without Tofino hardware. https://github.com/oxidecomputer/softnpuP4 -- DSL for describing packet-processing pipelines; how the Tofino/SoftNPU dataplane is programmed. https://p4.orgGimlet -- Oxide's compute sled.Sidecar -- Oxide's rack switch board, carrying the Tofino 2.raclette -- reduced-size lab rack (a few sleds + switch) used for dev/test rather than a full 32-sled rack. (my read; unverified)Hubris -- Rust microkernel/RTOS for the SP and RoT: statically-defined tasks, memory protection, synchronous IPC. https://github.com/oxidecomputer/hubrisHumility -- debugger for Hubris systems; dumps tasks, stacks, ringbufs over SWD. https://github.com/oxidecomputer/humilitya4x2 -- pre-voxel testbed topology: 4 emulated Gimlets, 2 SoftNPU switches, omicron on one Helios host.trust quorum -- rack secret split into shares across sleds; a quorum must be present to reconstitute the storage-encryption key, so a stolen sled is useless alone. https://docs.oxide.computer/guides/system/initial-rack-setupillumos -- open-source Solaris-derived OS; the host OS (Helios) is an illumos distro. https://illumos.orgzones -- illumos OS-level virtualization; each control-plane service runs in its own zone.Scrimlet -- Gimlet attached to a Sidecar ("switch Gimlet"); hosts the switch zone, dpd, MGS.DDM -- Delay-Driven Multipath, the underlay routing protocol, implemented in Maghemite (ddmd, ddmadm). https://github.com/oxidecomputer/maghemiteDendrite -- dataplane controller; dpd programs Tofino or SoftNPU and exposes an API to the control plane. https://github.com/oxidecomputer/dendriteswadm -- CLI against dpd for port config, links, addresses, routes.A0 -- sled power state where the host is fully powered (vs. A2, standby with the host off); driven by the SP sequencer task.iCE40 -- Lattice FPGA on Gimlet/Sidecar for power sequencing and board glue logic. https://github.com/oxidecomputer/quartzPRs needed!If we got something wrong or missed something, please file a PR! Our next show will likely be on Monday at 5p Pacific Time on our Discord server; stay tuned to our Mastodon feeds for details, or subscribe to this calendar. We'd love to have you join us, as we always love to hear from new speakers!
BSD Part Deux, Code that built the internet, and an interview about WebZFS NOTES This episode of BSDNow is brought to you by Tarsnap and the BSDNow Patreon Headlines BSD pt Deux Code that Built the Internet Interview JT on WebZFS TJ: I started by asking JT... Tarsnap This weeks episode of BSDNow was sponsored by our friends at Tarsnap, the only secure online backup you can trust your data to. Even paranoids need backups. Feedback/Questions Send questions, comments, show ideas/topics, or stories you want mentioned on the show to feedback@bsdnow.tv Join us and other BSD Fans in our BSD Now Telegram channel
Marc Brooker: How Technical Communication Drives DecisionsIn this episode, I'm joined by Marc Brooker, a VP and distinguished engineer who has worked on large-scale systems and now focuses on safety and policy for agentic AI. We explore how his approach to communicating complex technical work has evolved as his responsibilities have changed. We also discuss what works and what doesn't in technical presentations, how organizations can help technical professionals become better communicators, and why effective communication matters when technical work needs to influence decisions.To learn more about Marc, visit https://www.linkedin.com/in/marc-brooker-b431772b/On another note, Kiro is expanding its free Students tier from 11 to 132 universities across 18 countries, giving eligible students a full year of access with 1,000 monthly credits. The program supports building across the IDE, CLI, Web, and open‑source Crew, aiming to make professional‑grade AI tools globally accessible.Head to kiro.dev/students to learn more and sign up. Share what you are building with #KiroStudents — @kirodotdev on X, LinkedIn, and Instagram, and @kiro.dev on Bluesky.__TEACH THE GEEK (http://teachthegeek.com) Get Public Speaking Tips for STEM Professionals at http://teachthegeek.com/tips
PHP Podcast – September 2, 2026 Hosts: Eric Van Johnson & Joe Ferguson John’s out, Joe jumps in at the last second, a brand-new intro gets debuted mid-flight, a new PHP Foundation podcast gets announced, Frank’s database issues (not production), and somebody’s selling 37 elePHPants to fund a startup and open-heart surgery. New Intro, New Scenes, and a Last-Second Co-Host Eric kicked off the episode still tinkering with a brand-new intro and scene setup live on air, complete with a fresh montage that (thankfully) still includes Duck Nizzle. With John unavailable, Joe Ferguson jumped on at the last second to keep the show moving, and the two spent the opening minutes wrestling with scene sizing while Joe appeared comically small on screen. The pair caught up on client work: Joe pulled off an Elastic Beanstalk environment flip that nobody noticed (which is exactly how you want those to go), while Frank got thrown right under the bus for deleting the wrong database. Fortunately, it wasn’t production, and the conversation turned into an appreciation session for point-in-time recovery on AWS RDS and DigitalOcean’s managed databases, both of which have saved plenty of bacon during cowboy-hat coding sessions. The PHP Foundation Community Podcast Joe announced a brand-new podcast launching on the same PHP Architect channel, hosted by Elizabeth Barron and focused entirely on contributing to PHP. The first guest is Matt Stauffer, who recently published a fantastic post on the Foundation site about contributing to PHP, with recording set for September 10th. The plan is one Foundation Community Hour per month, plus the Joe, Holly, and Sarah show on weeks when Eric and John aren’t around. Eric and Joe used the moment to clear up a common misconception: people often conflate PHP internals with the Foundation, but they are not the same thing. The Foundation pays people to work on internals rather than steering the language itself. Its whole reason for existing is to get structure and funding around an old open-source project that billion-dollar companies have built empires on while contributing little back downstream. They also touched on recent Foundation board turnover: Roman’s term ended, and he’s rotating off, with Brent Roose joining as JetBrains’ representative — a great fit given Brent’s reach and the volume of PHP content he pushes out. Unlocking the PHP Architect Magazine Back Catalog The hosts made a pitch for the magazine, noting that an active subscription unlocks the entire catalog going all the way back to the very first issue — December 2002, which covered PHP 4.3. Eric got nostalgic about first discovering the print magazine right when PHP 4 came out, when knowledge was scarce, YouTube didn’t exist, and you had to buy O’Reilly books and beg your manager’s education budget for a subscription. Big news for existing accounts: John has opened up everything published before 2022 to anyone with an account on the system, even without an active subscription. Log in, head to magazines, and check what you now have access to. The team also reminded listeners that the magazine pays every contributor, that they’ve built a solid roster of regular columnists, and that PHP Architect is a consulting group available to augment teams, evaluate systems, and more. Where Developers Fit in an AI World Eric put Joe on the spot with a big question: What does a developer’s role look like going forward, given where AI is? Joe leaned positive on the tooling — Claude is excellent at parsing logs, spotting anomalies, and navigating large legacy codebases he isn’t in every day. He described real workflows like running AWS bill reviews for clients and comparing “what does your Claude say” with Frank to catch ridiculous bugs, all while distilling the output down as the human in the loop. But Joe was candid about his concerns too: AI feels force-fed in a way CSS or JS frameworks never did, with a real fear of being replaced by “a bunch of Claude agents in a trench coat.” He raised ethical and environmental concerns, pointing to accusations about his own city’s AI data center and the ripple effects on hardware — RAM prices spiking 400–500%, with Apple reportedly paying far more for iPhone RAM, a cost that always flows back to consumers. Eric compared the AI moment to crypto: a genuinely good technology released to the masses too soon, leading to misuse and eventual regulation. He shared a concrete win, though — using Claude to single-handedly bring a Laravel project shelved since COVID up to date, then using it to generate effort estimates (roughly three years of work) that helped set realistic expectations with a client who didn’t want to hear it. The pair debated whether those estimates accounted for AI-assisted velocity and stressed the importance of guardrails — especially for juniors, since even skilled devs leave S3 buckets open and AI is dangerously confident in its word guessing. TypePHP: Interesting, But Not for Everyone Eric dug into TypePHP, the ahead-of-time compiler from the Swoole team that translates PHP source into C++ and then native machine code. On the surface, he likes the idea — it’s something he wanted decades ago — but he questioned its relevance today: the compiler’s habit of catching bad code before it runs matters less now that PHPStan, linting, and a strong testing culture already fill that gap. Joe pointed out the bigger practical problem: there’s far more talk about TypePHP than actual code being written in it, and it’s clearly geared toward long-running CLI processes rather than web requests. That excludes the kind of PHP most developers actually write, and as Nuno noted in the chat, it isn’t going anywhere until it works with Laravel and the rest of the ecosystem. The hosts agreed the real value would be extracting any gains and folding them back into native PHP, à la HHVM/Hack, and noted transpiling already solves the “I want a different syntax” itch. Joe also pushed back on the “runs everywhere” selling point, arguing that Docker and Podman containers already make PHP CLI tools trivially portable — better than PHARs ever were. Eric landed on a mature take: he wants PHP everywhere and is happy to see it evolve, but he’s come to terms with the fact that he won’t personally use every direction it grows, TypePHP included. ElePHPants, Laracon Archives, and Free Talks A Reddit post caught the team’s attention: someone is selling their 37+ elephpant collection to fund a startup and help offset the cost of upcoming open-heart surgery. Joe noted the PHP tech elephpants are wildly undervalued compared to others in the listing — you can grab an OG PHP Architect orange Archie for under a hundred bucks. Eric showed off his own big Elvis (a gift from Sarah) and eyed the coveted Star Trek elephpant that John once nabbed in a trade at Longhorn. Eric also highlighted the new Laracon Archive project, which collects Laracon conference talks across regions (US, EU, Online, India) with a clean year-by-year interface — great work for anyone who tries to manage conference content. He closed by plugging phptech.tv, which hosts PHP Architect conference talks, including a batch of free talks and even a full free conference this year, with the same back-catalog-unlock model as the magazine. Links from the show: Longhorn PHP 2026 — October 15–16 PHP Foundation — Contributing to PHP blog post PHP Architect Magazine — Subscribe & unlock the back catalog PHPTech.tv — Conference talks, including free ones Host: Eric Van Johnson X: @shocm Mastodon: @eric@phparch.social Bluesky: @ericvanjohnson.bsky.social PHPArch.me: @eric Joe Ferguson Mastodon: @joepferguson@phpc.social PHPArch.me: @svpernova09 Streams: Youtube Channel Twitch Connect & Hire PHP Architect Website Twitter/X Mastodon Hire PHP Developers Looking to hire PHP developers? Email support@phparch.com – Joe and the team are available for consulting, infrastructure work, Ansible playbooks, and code review. Partner This podcast is made a little better thanks to our partners Displace Infrastructure Management, Simplified Automate Kubernetes deployments across any cloud provider or bare metal with a single command. Deploy, manage, and scale your infrastructure with ease. https://displace.tech/ OurCVEs Your security posture, on autopilot with OurCVEs CodeRabbit Cut code review time & bugs in half instantly with CodeRabbit. PHP Architect Consulting Your PHP codebase deserves a partner, not a contractor PHP Architect provides long-term technical partnerships for organizations that need senior-level PHP expertise that you can depend on https://www.phparch.com/consulting/ Music Provided by Epidemic Sound https://www.epidemicsound.com/ Join Us Live Next Week Youtube Channel Got feedback? Join us on Discord at discord.phparch.com The post The PHP Podcast 2026.09.02 appeared first on PHP Architect.
NOTES This episode of BSDNow is brought to you by Tarsnap and the BSDNow Patreon Headlines NetBSD 11 Release SR-IOV Is a first class feature News Roundup Grafana dashboards are a marvel of the modern web environment OpenSSH 10.5 released FreeBSD Ate my RAM ThinkPad End Game Tarsnap This weeks episode of BSDNow was sponsored by our friends at Tarsnap, the only secure online backup you can trust your data to. Even paranoids need backups. Feedback/Questions Send questions, comments, show ideas/topics, or stories you want mentioned on the show to feedback@bsdnow.tv Join us and other BSD Fans in our BSD Now Telegram channel
Sponsored by Blocks: Save at least 20% on your AWS costs with AI-powered optimization and enterprise discounts. Get your free Cloud Check at https://blocks.cloud/alphalist?utm_source=alphalist&utm_medium=podcast&utm_campaign=blocks-podcast-2026 David Soria Parra co-created the Model Context Protocol (MCP), the standard that now underpins most AI agent integrations, and, before that, spent years as a core contributor to Mercurial and on Meta's internal source control team. He got into programming at 13, writing a PHP guestbook for a gaming website with friends, and hasn't really stopped building since. This conversation goes straight to the thing every CTO is quietly building right now: their own agent harness. David makes a surprisingly blunt case that harness writing itself isn't the hard part: "give it a bunch of tools, a bunch of execution steps, be a little smarter about context selection, and go," and that most companies would be better off configuring a strong existing harness than building their own from scratch. He and Tobi also dig into why the "MCP vs. CLI" debate is largely manufactured, how Meta and Anthropic both run on almost no prescribed process, and how David's own workflow has quietly shifted from local Claude Code sessions to Slack threads over the past six months. CTOs will walk away with a clear-eyed, unhyped view of harness-building from the person who literally created the interoperability standard. Everyone argues about when building your own is worth it, what to check for in a vendor contract, and where to actually put your attention instead of endlessly optimizing your setup. What's covered: - David's path from a PHP guestbook at 13 to Mercurial core contributor to Meta's source control team - Biggest lessons from Meta, and how Anthropic runs on even less internal process - His personal setup: a handful of skills, a few MCP connectors, and why he's a "vanilla" tool user - Why his day-to-day workflow moved from Claude Code sessions into Slack threads with Claude - The real MCP-vs-CLI debate — and why it was never actually a debate - Should every CTO or engineer build their own agent harness? - What's still defensible in software engineering as agents get stronger
Robinhood Chain generates $1 million in REV. Etherscan releases an MCP, CLI, and Skills for AI Agents. Morpho introduces in-kind redemptions for vaults v2. And Bitmine purchased 53,501 ETH last week, putting its holding at 5.9 million ETH, which is about 4.9% of the total ETH supply. Read more: https://ethdaily.io/robinhood-chain-generates-dollar1m-in-rev Disclaimer: Content is for informational purposes only, not endorsement or investment advice. The accuracy of information is not guaranteed.
Switches in your ceiling, AI affecting interest in things, Commodore Phone and more... NOTES This episode of BSDNow is brought to you by Tarsnap and the BSDNow Patreon Headlines How you wind up with switches above your office's false ceiling On AI News Roundup Attempting to reproduce the missing etcmerge from 15.0 to 15.1 upgrade Commodore is releasing a flip phone running Sailfish OS! Why a Flip Phone? $399 - $640 preorder BSD Make extravaganza HardenedBSD June / July 2026 Status Report MidnightBSD 4.0.7 RELEASE Tarsnap This weeks episode of BSDNow was sponsored by our friends at Tarsnap, the only secure online backup you can trust your data to. Even paranoids need backups. Send questions, comments, show ideas/topics, or stories you want mentioned on the show to feedback@bsdnow.tv Join us and other BSD Fans in our BSD Now Telegram channel
PHP Podcast – August 27, 2026 Hosts: Joe Ferguson, Sara Golemon & Holly Schilling Zero-downtime swaps that nobody noticed, 3D printers that refuse to calibrate, a nostalgia tour through the entire PHPArch back catalog, Go generics, TypePHP, and a very good reason to book a flight to Belgium. This one lived in the gutter and loved it. Zero-Downtime Environment Swaps in AWS Joe kicked off the show talking about migrating a client’s production, staging, and test environments to completely new Elastic Beanstalk setups in AWS — during one of the client’s biggest selling seasons — and nobody noticed except him. That’s the dream: a zero-downtime upgrade that customers never feel. He explained that this was actually the second full rebuild of these servers in two months. The first migration got them off an aging Amazon Linux platform, and once everything was captured in CloudFormation templates and supporting scripts, swapping to an entirely new environment for an SSL certificate adjustment was suddenly no big deal. Being able to swap whole fleets of application servers behind a CNAME or DNS record and have clients never know is, in Joe’s words, one of the coolest things about cloud computing that will never get old. It also spawned a new accidental buzzword: “infrastructure as cloud.” 3D Printing Trials, Tribulations, and Guitar-String Belts Sara got her 3D printer working and has been printing everything from QR codes to a donate-link plaque for the PHP Foundation booth (inspired by feedback Sebastian got at a conference), all using the tried-and-true multicolor technique of hitting pause at the right layer and swapping filament. This is Sara’s fourth printer that should have been her third — she detailed the multi-day kit assemblies, belts that had to be tightened like guitar strings before calibration would work, and a tiny plastic pulley holder that snapped apart. Apparently nobody in Lisbon has a working 3D printer to help print a replacement part, so she resorted to sanding wood into a makeshift holder before finally getting it running and printing spares of everything. Holly admitted 3D printing is another expensive hobby she’d never follow through on, and Joe confessed he’d want the “Apple of 3D printers” that never jams — the same reason flying drones didn’t stick for him, since what he was actually doing was crashing and repairing drones. Sara later shared her first real project: little blue printed brackets that consolidate her monitor’s power and Thunderbolt into a single clean cord. Old Tech Nostalgia: Radio Shack Kits, Punch Cards, and Hard Drive Platters The crew took a deep dive into computing nostalgia. Sara reminisced about Radio Shack electronics kits with big springs you’d push aside to hold wires in place, lamenting that modern kits just have you jumper a GPS to a single-board computer — something big is missing from that experience. Holly recalled hacking on a third-hand 8088 laptop from a garage sale because her parents wouldn’t let her touch the family 486, and spending a full year in 1997 trying to get DHCP working across three Windows 95 machines before giving up and hard-coding IP addresses. Sara topped it with serial null-modem cables running between rooms in her first apartment, back when your serial port was your network. Holly then held up hard-drive platters she keeps as wall decoration, prompting a genuinely fun tangent about whether you could manually write bits to a low-density platter with a needle and magnet — a “solvable problem” per Joe, complete with G-code and a printed robot housing. The group also swapped memories of building PCs when a video card was mandatory and Sound Blaster was king, plus punch cards still stacked in a university basement at the turn of the century. Go Generic Methods, the PHPArch Archive Vault, and PHP 8.6 Go’s newest release finally added support for generic methods (it already had generic structs and functions), which Holly noted is PHP-relevant since it ties back to her own generics proposal debates. Sara puzzled over why methods and functions weren’t just supported together, since the only declarative difference is that receiver parameter — leading to an impromptu live look at the syntax and the conclusion that “all of Go is weird, just in the opposite direction PHP is weird.” Eric dropped breaking news mid-show: anyone with a phparch.com account — even an old one from a lapsed subscription — now has access to the magazine archive going all the way back. Joe pulled up Volume One, Issue One, and dug up a November 2002 piece about PHP 4.3 RC2 — right around when Sara submitted her first PHP patch and gave the world unified streams. They also covered Eric Van Johnson’s blog post on the PHP 8.6 soft feature freeze. Beta 2 shipped with a lot more deprecations landing that weren’t in beta 1, and the hosts encouraged listeners feeling froggy to grab it and test drive. Holly griped that one of her merged 8.6 changes didn’t even get a mention — thanks a lot, Eric. TypePHP and Competing Runtimes The team dug into TypePHP, which compiles PHP down to a native binary. Sara was openly dubious of the speed claims — a 50% gain wouldn’t surprise her, but an 8x gain needs much better data — while Joe noted the only demos so far are CLI apps, not traditional web workloads, and worried the authors are writing a very different style of PHP than most people do. Holly framed it through history: the whole reason PHP, Perl, and CGI exist is to avoid the per-request cost of loading and launching a native binary, so she’s curious what tradeoffs a compiled approach makes — especially since people sensitive to raw performance often reach for Go in the first place. She also pointed out most PHP code is just reading a file, not running the billion-row challenge. Sara pushed back on the “competing runtime” framing, recalling HHVM as more synergistic than antagonistic — it launched with PHP’s own code inside it, and PHP 7 hitting the speed of PHP 5 was directly informed by HHVM’s research. The real test will be whether TypePHP can run real Laravel apps and WordPress sites reliably enough for ecosystem adoption, the same way HHVM eventually powered Wikipedia’s front end with just a single WordPress patch. PHP Foundation Speaker Accreditation & PHP BeNeLux Returns Joe covered today’s PHP Foundation Ambassador SIG meeting, focused on the speaker subgroup and a proposal James Seconde has been working on. The goal is a foundation accreditation program that builds a list of speakers who can carry a consistent PHP message, tapped by geographic region so events get local speakers, more diverse voices get on stage, and nobody gets burned out flying halfway across the world when four nearby speakers could do the same talk. The big excitement, though, was the announcement of PHP BeNeLux in Antwerp, Belgium in January 2027 — the first since January 2020 before COVID. Sara audibly lost it, partly because it lines up with FOSDEM and partly out of love for longtime organizer Michelangelo van Dam, whom Joe credited as a huge, welcoming reason he got involved in the PHP community in the first place. Both hosts are already thinking about submitting talks. Links from the show: PHP Tek 2027 CFP — extended through end of October PHP Tek 2027 — April 27–29, Chicago (tickets available) PHPArch Swag Store 3D Code Generator Export QR codes as STL for 3D printing PHPArch.com Magazine Archive — now available to all account holders PHP 8.6 Beta 2 — now available for testing TypePHP — PHP compiled to a native binary PHPBenelux — Antwerp, Belgium, January 2027 Host: Joe Ferguson Mastodon: @joepferguson@phpc.social PHPArch.me: @svpernova09 Sara Golemon Mastodon: @pollita@phpc.social Holly Schilling Mastodon: @TheCodeLorax@tech.lgbt Streams: Youtube Channel Twitch Connect & Hire PHP Architect Website Twitter/X Mastodon Hire PHP Developers Looking to hire PHP developers? Email support@phparch.com – Joe and the team are available for consulting, infrastructure work, Ansible playbooks, and code review. Partner This podcast is made a little better thanks to our partners Displace Infrastructure Management, Simplified Automate Kubernetes deployments across any cloud provider or bare metal with a single command. Deploy, manage, and scale your infrastructure with ease. https://displace.tech/ OurCVEs Your security posture, on autopilot with OurCVEs CodeRabbit Cut code review time & bugs in half instantly with CodeRabbit. PHP Architect Consulting Your PHP codebase deserves a partner, not a contractor PHP Architect provides long-term technical partnerships for organizations that need senior-level PHP expertise that you can depend on. https://www.phparch.com/consulting/ Music Provided by Epidemic Sound https://www.epidemicsound.com/ Join Us Live Next Week Youtube Channel Got feedback? Join us on Discord at discord.phparch.com The post The PHP Podcast 2026.08.27 appeared first on PHP Architect.
In this update-packed episode James and Frank dive into developer tooling: the VS Code Mobile Canvas that embeds emulators into your editor, MAUI DevFlow's rewrite enabling plain .NET/native (and desktop AppKit) support, and the growing MCP/agent ecosystem that ties editors, the Copilot app and CLI together. They share practical takeaways—Mobile Canvas works across frameworks, DevFlow supports non‑MAUI apps, agent instruction files matter, and you can automate Microsoft Store publishing—giving cross‑platform builders clear next steps. Follow Us Frank: Twitter, Blog, GitHub James: Twitter, Blog, GitHub Merge Conflict: Twitter, Facebook, Website, Chat on Discord Music : Amethyst Seer - Citrine by Adventureface ⭐⭐ Review Us ⭐⭐ Machine transcription available on http://mergeconflict.fm
Home Assistant Setup, OpenBSD Updates, Wine 11.14, and more... NOTES This episode of BSDNow is brought to you by Tarsnap and the BSDNow Patreon Headlines My Home Assistant Setup Huge!!! OpenBSD Updates WPA3 support coming to OpenBSD Game of Trees 0.127 Released OpenBSD relayd(8) adds ECDSA support with CA engine code from smtpd(8) httpd(8) gains support for custom HTTP headers LLVM toolchain coming to OpenBSD/sparc64 Call for testing: OpenBSD vmm(4)/vmd(8) fd-ification News Roundup Wine 11.14 Brings New WoW64 Mode to FreeBSD I Built a FreeBSD Cloud to Use with FreeBSD How Unix Spell Ran in 64kB RAM Tarsnap This weeks episode of BSDNow was sponsored by our friends at Tarsnap, the only secure online backup you can trust your data to. Even paranoids need backups. Feedback/Questions Tyler - IPv6 Question Reese - WebZFS Send questions, comments, show ideas/topics, or stories you want mentioned on the show to feedback@bsdnow.tv Join us and other BSD Fans in our BSD Now Telegram channel
PHP Podcast – August 20, 2026 Hosts: Joe Ferguson, Sara Golemon & Holly Schilling Shirley MacLaine is the answer — but what’s the question? RAM prices went up 500%, GitHub fell over for eight hours, and the crew figured out how you can actually contribute to PHP. Glasses optional. Shirley MacLaine, Green Room Shade, and the Show’s New Normal The episode opens with Joe flustered — not because it’s a “flustered day,” but because there’s shade being thrown in the green room backstage chat. Rather than spoil the drama, the crew turns a mysterious green-room answer into a running bit: “Shirley MacLaine” is the answer, and listeners are invited to submit what the question was. It’s declared the first round of PHP Architect’s Jeopardy, complete with its own Easter egg sound cue. Joe also lays out where the show is heading. Eric and John took over the podcast the previous week to relive their glory days, and the plan is for the old guys to step back to roughly one episode a month. Joe and Sara are working on something to fill one slot, another mystery show is in the works pending signed contracts, and Alive and Kicking is confirmed still alive and kicking after a great recent episode with Derek. Joe also plugs the PHP 8.6 Beta 1 tag and Scott Keck-Warren’s PHP Community Podcast interview with release managers Daniel Scherzer and Matteo Bacotti. The RAM Apocalypse: 500% in Twelve Months The big topic of the week is the ongoing memory and chip crisis. Memory prices have climbed 500% in twelve months, and the crew commiserates about the parts they wish they’d bought bigger. Sara explains the brutal math: you can’t build a semiconductor fab fast enough — two or three years minimum — and by the time one comes online nobody knows if we’ll have overproduction or a burst bubble. Worse, Sara notes that essentially every RAM stick through the end of 2027 is already accounted for and sold to vendors. The ripple effects are everywhere. Console prices are going up instead of down late in their cycle, with next-gen consoles projected to start at $1,000 or more. The hobbyist single-board computer market is getting crushed — Pine64 announced they’re stepping back from making hardware, and a Raspberry Pi 5 that debuted cheap now runs over 100 pounds. Holly’s new PC, bought at the start of 2025, ended up shipped 3,000 kilometers the wrong way to California and is stuck awaiting import paperwork, leaving her leaning on a laptop and some now-precious Raspberry Pi zeros. The nostalgia gets thick as the crew reminisces about DIP chips, EDO RAM, Pentiums, Epson 486s, Tandy 1000s, and the TRS-80. Special venom is reserved for Packard Bell (“utter freaking trash”) and Gateway 2000’s cow-pattern branding — which Sara confirms sold better in Wisconsin than anywhere, though hay bales in a Berkeley storefront still boggle the mind. A chat comment about technicians de-soldering and reballing BGA RAM chips becoming economically viable gets a hearty “absolutely.” Memory-Aware Development and Why PHP 7 Doubled Down Bringing the RAM crisis back to PHP, Joe wishes more developers were aware of how their applications consume and release memory — a lesson he credits to learning enough C back in the day, where you have to manage memory yourself. He connects sloppy memory awareness to the N+1 query problems web developers keep tripping over. Sara drops a great deep dive: a significant reason PHP 7 was roughly twice as fast as PHP 5 was changes in the memory layout. Every variable became referenced by one fewer pointer, and while eight bytes sounds trivial, every level of indirection adds time across every single instruction and access. Sara adds the CPU-level detail — one layer of indirection can be a single instruction on most architectures, but adding a second layer can push a lookup from one instruction to three. That leads into a warm tangent about learning C to become game developers. Sara’s evergreen joke: “I’m going to be a game developer” is the programmer’s version of “I’m going to buy a bar.” Great people, brutal hours, endless competition, and the reality of hitting spacebar 400 times to figure out why you can phase through a wall. The cat-reading-the-paper “I should buy a boat” meme makes an appearance to seal it. The GitHub Outage and the Monoculture Problem Monday’s eight-hour GitHub outage hit the crew directly. Holly couldn’t use a site that only offered “log in with GitHub,” and Joe got kicked out of his CLI auth session mid-PR with no way to re-authenticate. To GitHub’s credit, they published an incident update and a follow-up blog post: a service auto-scaled so aggressively to handle network traffic that the sidecar and supporting services couldn’t keep up, bringing the whole thing down. The conversation turns to whether this is self-inflicted. Joe recalls GitHub’s pre-Microsoft, gold-standard reliability and wonders aloud how much the decline lines up with Copilot’s arrival and internal AI adoption, with uptime reportedly slipping below a single nine at points. Sara defends them somewhat — the number of actions, CPU cores, and pull requests has genuinely hockey-sticked, partly because AI has emboldened people who previously wouldn’t have opened a PR. But as Joe puts it, the call is coming from inside the house, since GitHub itself has been pushing AI. On alternatives, Joe says the least-jarring migration for PHP Architect’s clients would be self-hosting GitLab, since GitHub Actions and GitLab runners are nearly identical in syntax — Atlassian’s Bitbucket, by contrast, is a bridge too far, mostly because the entire ecosystem assumes you’re on GitHub. Sara names this the core problem: monoculture. The crew discusses package mirrors, local caches, 12-factor thinking, and Composer’s support for custom mirrors, all while remembering the PHP repo intrusion years ago that came from an unmaintained self-hosted Git server. The takeaway: owning your pipeline end-to-end is the only way an outage can’t stop you — and Joe teases spinning up a self-hosted GitLab now that “the boss” (Sara) has signed off. How to Contribute to PHP (and Handling Security Reports) The crew highlights two PHP Foundation blog posts. First, Matt Stauffer’s “How to Contribute to PHP,” adapted from a talk he gave at Atlanta PHP. It goes well beyond “learn C,” clearly separating the PHP project, the PHP ecosystem, and the Foundation, and lays out approachable on-ramps: testing pre-releases (PHP 8.6 Beta 1 is out, Beta 2 lands next week), improving documentation, and writing tests — which, spoiler, are written in PHP, not C, using PHP’s own test format that’s simple enough to learn from any single example. Other contribution paths include triaging and reviewing issues across PHP repositories — invaluable work that frees core developers from wading through AI-generated slop bug reports — and participating in internals via the well-documented mailing list process, up to and including running for release manager (8.7 managers will be needed before you know it). Sara points folks to discord.phpc.chat for the PHP Discord, with dedicated Internals and Foundation channels for anyone the mailing list intimidates. Second, Sebastian Bergmann’s “So you received a security report. Now what?” is a jump-around reference for application developers rather than a front-to-back read, walking through roughly ten steps to triage, validate, and resolve reported issues the right way. Sara shares a real-world example from mobile: a flagged package that was only exploitable on a rooted device with an actively hostile package installed alongside it — a very different risk profile than a SQL injection on an API endpoint. Cue reminiscing about writing SQL against Access databases over ODBC from PHP (and Perl) back in the 90s, and Joe’s advice for handling any security report: don’t panic, and always know where your towel is. Links from the show: PHP Tek 2027 — April 27–29, 2027 in Chicago; early bird tickets & hotel available now PHP Tek 2027 CFP Audio versions of the podcast at phparch.com Join us live on Discord at discord.phparch.com PHP Discord — discord.phpc.chat Community Corner Podcast: PHP 8.5 + 8.6 Release Manager Daniel Scherzer Memory prices climb 500% in 12 months So You Received a Security Report. Now What? How to Contribute to PHP Shirley MacLaine Host: Joe Ferguson Mastodon: @joepferguson@phpc.social PHPArch.me: @svpernova09 Sara Golemon Mastodon: @pollita@phpc.social Holly Schilling Mastodon: @TheCodeLorax@tech.lgbt Streams: Youtube Channel Twitch Connect & Hire PHP Architect Website Twitter/X Mastodon Hire PHP Developers Looking to hire PHP developers? Email support@phparch.com – Joe and the team are available for consulting, infrastructure work, Ansible playbooks, and code review. Partner This podcast is made a little better thanks to our partners Displace Infrastructure Management, Simplified Automate Kubernetes deployments across any cloud provider or bare metal with a single command. Deploy, manage, and scale your infrastructure with ease. https://displace.tech/ OurCVEs Your security posture, on autopilot with OurCVEs CodeRabbit Cut code review time & bugs in half instantly with CodeRabbit. PHP Architect Consulting Your PHP codebase deserves a partner, not a contractor PHP Architect provides long-term technical partnerships for organizations that need senior-level PHP expertise that you can depend on https://www.phparch.com/consulting/ Music Provided by Epidemic Sound https://www.epidemicsound.com/ Join Us Live Next Week Youtube Channel Got feedback? Join us on Discord at discord.phparch.com The post The PHP Podcast 2026.08.20 appeared first on PHP Architect.
AWS Morning Brief for the week of August 17th, with Corey Quinn. Links:Amazon EC2 introduces application status checksAWS IAM Identity Center supports one-click multi-Region option for new organization instancesAmazon S3 adds additional policy details to access denied error messagesAWS Secrets Manager adds managed external secrets support for Jenkins and SonarQubeBurst to Region: Overflow AWS Outposts workloads to Amazon EC2Designing for failure: Building resilient systems on AWSAmazon Quick for Microsoft 365: Agentic AI where you workIntroducing the next-generation AWS VPN Client with CLI support and admin controlsAWS Certificate Manager will discontinue email validation to prove domain validation for certificatesHow AWS IAM role manager rethinks the starting point for IAM rolesHow to authenticate customers during chat with Amazon Connect CustomerHow WeatherBug reduced storage costs by 80% using Amazon S3 Storage Lens and Kiro CLIIntroducing the new AWS Cloud Quest: AI-powered practice, hands-on building, and a path to official AWS badgesTwo bulletins, one CVE, and Base64 bites C++
The Future of the FreeBSD Kernel LLDB Plugin, Stop Ruining "10 PRINT", BoxyBSD Returns to FreeBSD, The Apple Lisa inside an FPGA, and more... NOTES This episode of BSDNow is brought to you by Tarsnap and the BSDNow Patreon Headlines Future of the FreeBSD Kernel LLDB Plugin Please, Just Stop Ruining "10 PRINT". Seriously. News Roundup BoxyBSD Returns to FreeBSD Google Summer of Code 2026 Reports: Testing Compat Linux: Syscall testing The Apple Lisa inside an FPGA! The Virtual OS Museum Tarsnap This weeks episode of BSDNow was sponsored by our friends at Tarsnap, the only secure online backup you can trust your data to. Even paranoids need backups. Feedback/Questions Marcus - Linuxulator vs Bhyve - Reese - Feedback on episode 673 Send questions, comments, show ideas/topics, or stories you want mentioned on the show to feedback@bsdnow.tv Join us and other BSD Fans in our BSD Now Telegram channel
Microsoft's AI investments reveal a bizarre financial loop where billions pour into infrastructure, yet profits remain elusive. The hosts break down why the industry's most hyped tech might just be a high-stakes game of circular accounting. Plus, a developer has resurrected Microsoft Word for Windows 1.1a to run natively on Windows 11! Lastly, Stardock's latest utility replaces Alt + Tab for $7.99. Windows 25H2: new touchpad gestures, Windows Hello ESS improvements, Widgets improvements, Windows Search improvements, more 26H1: Windows Update calendar-based pausing, Picture Password EOL, Point in time restore, multi-camera support, user folder customization in Setup, etc. Over 420 bug fixes, but that's over 200 less than last month Follow-up on Microsoft's AI spending. It's much worse than we knew - And also much worse than we know now Yes, Microsoft Edge is going all-in on Manifest V3 Proton Drive - Now with a CLI and business features, and it's coming soon to Linux A developer ported Word for Windows 1.1 to x64 AI Gemini now has one billion users. But how many pay? - Plus the new Pixel phones are one step sideways, one step back OpenAI brings ChatGPT to Linux (!) Xbox and gaming A new Xbox Elite Series 3 controller leak shows it will have a mini display Microsoft tests new console features with Xbox Insiders Id Software brings a new episode to Quake to mark 30th anniversary! Minecraft is coming to Switch 2 Here comes GTA VI A Valve partner was hacked, Steam Machine customers warned Nintendo revenues decline 9.5 percent to ¥518 billion Tips and picks Tip of the week: View Git status directly in File Explorer App pick of the week: AltTabby RunAs Radio this week: The Content Management System Landscape with Matt Garrepy Brown liquor pick of the week: Sideshow Lincoln Straight Bottled-in-Bond Bourbon Hosts: Leo Laporte, Paul Thurrott, and Richard Campbell Download or subscribe to Windows Weekly at https://twit.tv/shows/windows-weekly Check out Paul's blog at thurrott.com The Windows Weekly theme music is courtesy of Carl Franklin. Join Club TWiT for Ad-Free Podcasts! Support what you love and get ad-free audio and video feeds, a members-only Discord, and exclusive content. Join today: https://twit.tv/clubtwit Sponsors: helixsleep.com/windows cirasync.com/Windows cachefly.com/twit
Microsoft's AI investments reveal a bizarre financial loop where billions pour into infrastructure, yet profits remain elusive. The hosts break down why the industry's most hyped tech might just be a high-stakes game of circular accounting. Plus, a developer has resurrected Microsoft Word for Windows 1.1a to run natively on Windows 11! Lastly, Stardock's latest utility replaces Alt + Tab for $7.99. Windows 25H2: new touchpad gestures, Windows Hello ESS improvements, Widgets improvements, Windows Search improvements, more 26H1: Windows Update calendar-based pausing, Picture Password EOL, Point in time restore, multi-camera support, user folder customization in Setup, etc. Over 420 bug fixes, but that's over 200 less than last month Follow-up on Microsoft's AI spending. It's much worse than we knew - And also much worse than we know now Yes, Microsoft Edge is going all-in on Manifest V3 Proton Drive - Now with a CLI and business features, and it's coming soon to Linux A developer ported Word for Windows 1.1 to x64 AI Gemini now has one billion users. But how many pay? - Plus the new Pixel phones are one step sideways, one step back OpenAI brings ChatGPT to Linux (!) Xbox and gaming A new Xbox Elite Series 3 controller leak shows it will have a mini display Microsoft tests new console features with Xbox Insiders Id Software brings a new episode to Quake to mark 30th anniversary! Minecraft is coming to Switch 2 Here comes GTA VI A Valve partner was hacked, Steam Machine customers warned Nintendo revenues decline 9.5 percent to ¥518 billion Tips and picks Tip of the week: View Git status directly in File Explorer App pick of the week: AltTabby RunAs Radio this week: The Content Management System Landscape with Matt Garrepy Brown liquor pick of the week: Sideshow Lincoln Straight Bottled-in-Bond Bourbon Hosts: Leo Laporte, Paul Thurrott, and Richard Campbell Download or subscribe to Windows Weekly at https://twit.tv/shows/windows-weekly Check out Paul's blog at thurrott.com The Windows Weekly theme music is courtesy of Carl Franklin. Join Club TWiT for Ad-Free Podcasts! Support what you love and get ad-free audio and video feeds, a members-only Discord, and exclusive content. Join today: https://twit.tv/clubtwit Sponsors: helixsleep.com/windows cirasync.com/Windows cachefly.com/twit
Microsoft's AI investments reveal a bizarre financial loop where billions pour into infrastructure, yet profits remain elusive. The hosts break down why the industry's most hyped tech might just be a high-stakes game of circular accounting. Plus, a developer has resurrected Microsoft Word for Windows 1.1a to run natively on Windows 11! Lastly, Stardock's latest utility replaces Alt + Tab for $7.99. Windows 25H2: new touchpad gestures, Windows Hello ESS improvements, Widgets improvements, Windows Search improvements, more 26H1: Windows Update calendar-based pausing, Picture Password EOL, Point in time restore, multi-camera support, user folder customization in Setup, etc. Over 420 bug fixes, but that's over 200 less than last month Follow-up on Microsoft's AI spending. It's much worse than we knew - And also much worse than we know now Yes, Microsoft Edge is going all-in on Manifest V3 Proton Drive - Now with a CLI and business features, and it's coming soon to Linux A developer ported Word for Windows 1.1 to x64 AI Gemini now has one billion users. But how many pay? - Plus the new Pixel phones are one step sideways, one step back OpenAI brings ChatGPT to Linux (!) Xbox and gaming A new Xbox Elite Series 3 controller leak shows it will have a mini display Microsoft tests new console features with Xbox Insiders Id Software brings a new episode to Quake to mark 30th anniversary! Minecraft is coming to Switch 2 Here comes GTA VI A Valve partner was hacked, Steam Machine customers warned Nintendo revenues decline 9.5 percent to ¥518 billion Tips and picks Tip of the week: View Git status directly in File Explorer App pick of the week: AltTabby RunAs Radio this week: The Content Management System Landscape with Matt Garrepy Brown liquor pick of the week: Sideshow Lincoln Straight Bottled-in-Bond Bourbon Hosts: Leo Laporte, Paul Thurrott, and Richard Campbell Download or subscribe to Windows Weekly at https://twit.tv/shows/windows-weekly Check out Paul's blog at thurrott.com The Windows Weekly theme music is courtesy of Carl Franklin. Join Club TWiT for Ad-Free Podcasts! Support what you love and get ad-free audio and video feeds, a members-only Discord, and exclusive content. Join today: https://twit.tv/clubtwit Sponsors: helixsleep.com/windows cirasync.com/Windows cachefly.com/twit
Microsoft's AI investments reveal a bizarre financial loop where billions pour into infrastructure, yet profits remain elusive. The hosts break down why the industry's most hyped tech might just be a high-stakes game of circular accounting. Plus, a developer has resurrected Microsoft Word for Windows 1.1a to run natively on Windows 11! Lastly, Stardock's latest utility replaces Alt + Tab for $7.99. Windows 25H2: new touchpad gestures, Windows Hello ESS improvements, Widgets improvements, Windows Search improvements, more 26H1: Windows Update calendar-based pausing, Picture Password EOL, Point in time restore, multi-camera support, user folder customization in Setup, etc. Over 420 bug fixes, but that's over 200 less than last month Follow-up on Microsoft's AI spending. It's much worse than we knew - And also much worse than we know now Yes, Microsoft Edge is going all-in on Manifest V3 Proton Drive - Now with a CLI and business features, and it's coming soon to Linux A developer ported Word for Windows 1.1 to x64 AI Gemini now has one billion users. But how many pay? - Plus the new Pixel phones are one step sideways, one step back OpenAI brings ChatGPT to Linux (!) Xbox and gaming A new Xbox Elite Series 3 controller leak shows it will have a mini display Microsoft tests new console features with Xbox Insiders Id Software brings a new episode to Quake to mark 30th anniversary! Minecraft is coming to Switch 2 Here comes GTA VI A Valve partner was hacked, Steam Machine customers warned Nintendo revenues decline 9.5 percent to ¥518 billion Tips and picks Tip of the week: View Git status directly in File Explorer App pick of the week: AltTabby RunAs Radio this week: The Content Management System Landscape with Matt Garrepy Brown liquor pick of the week: Sideshow Lincoln Straight Bottled-in-Bond Bourbon Hosts: Leo Laporte, Paul Thurrott, and Richard Campbell Download or subscribe to Windows Weekly at https://twit.tv/shows/windows-weekly Check out Paul's blog at thurrott.com The Windows Weekly theme music is courtesy of Carl Franklin. Join Club TWiT for Ad-Free Podcasts! Support what you love and get ad-free audio and video feeds, a members-only Discord, and exclusive content. Join today: https://twit.tv/clubtwit Sponsors: helixsleep.com/windows cirasync.com/Windows cachefly.com/twit
Microsoft's AI investments reveal a bizarre financial loop where billions pour into infrastructure, yet profits remain elusive. The hosts break down why the industry's most hyped tech might just be a high-stakes game of circular accounting. Plus, a developer has resurrected Microsoft Word for Windows 1.1a to run natively on Windows 11! Lastly, Stardock's latest utility replaces Alt + Tab for $7.99. Windows 25H2: new touchpad gestures, Windows Hello ESS improvements, Widgets improvements, Windows Search improvements, more 26H1: Windows Update calendar-based pausing, Picture Password EOL, Point in time restore, multi-camera support, user folder customization in Setup, etc. Over 420 bug fixes, but that's over 200 less than last month Follow-up on Microsoft's AI spending. It's much worse than we knew - And also much worse than we know now Yes, Microsoft Edge is going all-in on Manifest V3 Proton Drive - Now with a CLI and business features, and it's coming soon to Linux A developer ported Word for Windows 1.1 to x64 AI Gemini now has one billion users. But how many pay? - Plus the new Pixel phones are one step sideways, one step back OpenAI brings ChatGPT to Linux (!) Xbox and gaming A new Xbox Elite Series 3 controller leak shows it will have a mini display Microsoft tests new console features with Xbox Insiders Id Software brings a new episode to Quake to mark 30th anniversary! Minecraft is coming to Switch 2 Here comes GTA VI A Valve partner was hacked, Steam Machine customers warned Nintendo revenues decline 9.5 percent to ¥518 billion Tips and picks Tip of the week: View Git status directly in File Explorer App pick of the week: AltTabby RunAs Radio this week: The Content Management System Landscape with Matt Garrepy Brown liquor pick of the week: Sideshow Lincoln Straight Bottled-in-Bond Bourbon Hosts: Leo Laporte, Paul Thurrott, and Richard Campbell Download or subscribe to Windows Weekly at https://twit.tv/shows/windows-weekly Check out Paul's blog at thurrott.com The Windows Weekly theme music is courtesy of Carl Franklin. Join Club TWiT for Ad-Free Podcasts! Support what you love and get ad-free audio and video feeds, a members-only Discord, and exclusive content. Join today: https://twit.tv/clubtwit Sponsors: helixsleep.com/windows cirasync.com/Windows cachefly.com/twit
#363: Three waves of the web, and you are late for the third one. The 90s were about getting a browser to render your page at all. The early 2000s were about SEO, or as Darin puts it, sell me all the ads ready. Now it is agent ready, and Cloudflare built a scoreboard for it at [isitagentready.com](https://isitagentready.com/). The devopsparadox.com site scored about 70 out of 100 and then went down when Cloudflare added new checks. Run yours. You will be sad. Viktor thinks the framing is slightly off, though, and the correction is the good part. Optimizing for agents that browse your site is aiming at the wrong thing, because most requests never touch your server. Agent asks the model, model answers, agent shows you. So the target is not the crawler, it is the training data - and if the model does go looking, the question becomes whether you are the first answer or one of the five sites it was told to go analyze. Same game as Google. Different index. It is not Google index anymore, it is model training now. Then the practical part. Five things Cloudflare scores you on: discoverability, content, bot access control, API, Auth, MCP & Skill Discovery, and Commerce. Content accessibility is where most of you are losing, because agents want Markdown and you are serving them a pile of HTML tags to strip. Both DOP and Viktor's site are Hugo, so the Markdown is already sitting on disk next to the HTML - serve one or the other based on what the request asks for. Almost no effort. If you are still shipping a JavaScript-rendered site, Darin says it is game over, and humans do not like those either. On the blocking side, both of them are baffled by the same thing: if you do not want agents reading it, do not publish it. robots.txt is a suggestion at best. If you really want to block, actually block. The API argument is the one that will annoy people. Viktor says CLIs and MCP servers are both auto-generated from a schema, so the real work is having a good API, and most companies do not. But who your audience is decides the wrapper - developers already have Bash, so give them a CLI and get out of the way. Everyone else needs MCP, because Viktor's mom is not installing your binary. And somewhere in the middle of all this Darin asks whether documentation should live in the code now more than ever, and Viktor says no, less than ever - he wants it separate so he can review it, because agents made everything cheap to produce and review is now the only thing standing between him and 5,000 features a day. Also: WordPress should be the last thing you consider, not the first. YouTube channel: https://youtube.com/devopsparadox Review the podcast on Apple Podcasts: https://www.devopsparadox.com/review-podcast/ Slack: https://www.devopsparadox.com/slack/ Connect with us at: https://www.devopsparadox.com/contact/
Microsoft's AI investments reveal a bizarre financial loop where billions pour into infrastructure, yet profits remain elusive. The hosts break down why the industry's most hyped tech might just be a high-stakes game of circular accounting. Plus, a developer has resurrected Microsoft Word for Windows 1.1a to run natively on Windows 11! Lastly, Stardock's latest utility replaces Alt + Tab for $7.99. Windows 25H2: new touchpad gestures, Windows Hello ESS improvements, Widgets improvements, Windows Search improvements, more 26H1: Windows Update calendar-based pausing, Picture Password EOL, Point in time restore, multi-camera support, user folder customization in Setup, etc. Over 420 bug fixes, but that's over 200 less than last month Follow-up on Microsoft's AI spending. It's much worse than we knew - And also much worse than we know now Yes, Microsoft Edge is going all-in on Manifest V3 Proton Drive - Now with a CLI and business features, and it's coming soon to Linux A developer ported Word for Windows 1.1 to x64 AI Gemini now has one billion users. But how many pay? - Plus the new Pixel phones are one step sideways, one step back OpenAI brings ChatGPT to Linux (!) Xbox and gaming A new Xbox Elite Series 3 controller leak shows it will have a mini display Microsoft tests new console features with Xbox Insiders Id Software brings a new episode to Quake to mark 30th anniversary! Minecraft is coming to Switch 2 Here comes GTA VI A Valve partner was hacked, Steam Machine customers warned Nintendo revenues decline 9.5 percent to ¥518 billion Tips and picks Tip of the week: View Git status directly in File Explorer App pick of the week: AltTabby RunAs Radio this week: The Content Management System Landscape with Matt Garrepy Brown liquor pick of the week: Sideshow Lincoln Straight Bottled-in-Bond Bourbon Hosts: Leo Laporte, Paul Thurrott, and Richard Campbell Download or subscribe to Windows Weekly at https://twit.tv/shows/windows-weekly Check out Paul's blog at thurrott.com The Windows Weekly theme music is courtesy of Carl Franklin. Join Club TWiT for Ad-Free Podcasts! Support what you love and get ad-free audio and video feeds, a members-only Discord, and exclusive content. Join today: https://twit.tv/clubtwit Sponsors: helixsleep.com/windows cirasync.com/Windows cachefly.com/twit
Tod Kurt returns to the show and shares how to build custom CircuitPython firmware using GitHub Actions. Tod also shares SerialPlotster and his latest hardware creation.Show Notes00:57 Using GitHub Actions to build CircuitPython3:51 CircuitPython Custom on GitHub repositoryCircuitPython Custom configuration web page9:24 The CLI display10:37 Adding custom modules11:42 How to run CircuitPython Custom14:34 SerialPlotsterDownload19:32 New hardware?20:40 Creating synth engines21:32 Wrap-uptodbot.comThe Bootloader
FreeBSD 16 goes GPL Free, The computer at the bottom of the lake, FreeBSD's new Board member, Phaethon, a new 68010 based Unix machine and more... NOTES This episode of BSDNow is brought to you by Tarsnap and the BSDNow Patreon Headlines The Computer at the bottom of a canal FreeBSD 16 goes GPL Free News Roundup FreeBSD Foundation Welcomes New Board Member: Dave Cottlehuber Phaethon 1 - a 68010-based Unix machine Add a new m68k port oriented towards home-brew m68k machines Bringing Swift to the Apple II Working around dragons with the Lemote Yeeloong laptop and OpenBSD Tarsnap This weeks episode of BSDNow was sponsored by our friends at Tarsnap, the only secure online backup you can trust your data to. Even paranoids need backups. Feedback/Questions Producer note : Ive been insanely busy this past week and I've fallen behind on a few things, including checking the show email. If you sent an email in and we havent covered it yet, we'll get caught back up soon. If yuo haven't sent in an email... shame on you. You should email us and ask us something. Send questions, comments, show ideas/topics, or stories you want mentioned on the show to feedback@bsdnow.tv Join us and other BSD Fans in our BSD Now Telegram channel
TestTalks | Automation Awesomeness | Helping YOU Succeed with Test Automation
Your developers just got supercharged by AI coding agents. Your test coverage did not. So how do you keep quality high when product code is shipping faster than any test team can follow? In this episode Andrew Knight, the Automation Panda and Senior Director of Product and Engineering at Cycle Labs, shares how his six person team is running the biggest release quarter in company history using AI coding agents, spec driven development, and Playwright. You will discover: How Playwright turned itself into an AI automation platform with the MCP server, the planner, generator, and healer agents, and the new CLI and skills approach that cuts your token usage way down. Why Andy uses Spec Kit to codify his testing strategy once, in markdown, so quality standards get baked into every single pull request instead of being caught in review. How to decide which AI generated tests are actually worth running when compute time and budget are finite. What AI slop looks like from a manager's seat, and how to build a team culture that catches it before it ships. Why Andy believes AI coding tools are the new compiler and markdown is the new programming language, plus my pushback on what that means for everything testers were trained to care about. What Andy really thinks about token costs, subscription tiers, and what happens when the AI subsidies run out. Plus the one piece of advice he gives to any tester still sitting on the sidelines of the AI shift. Whether you are an automation engineer, a QA lead, or an engineering manager trying to figure out where testing fits in an AI first workflow, this episode gives you a practical playbook you can start using this week.
Scrub Design, WebZFS Updates, The Foundations role in the FreeBSD Ecosystem, and more... NOTES This episode of BSDNow is brought to you by Tarsnap and the BSDNow Patreon Headlines If Scrubs Hurt, Your ZFS Design Is Broken WebZFS Updates Beta Announcement Why wont my pool export zfs iostat - Pictures so you can get a better idea of how it works News Roundup Understanding the Foundation Board's Role in the FreeBSD Ecosystem Monitor your devices with LibreNMS on FreeBSD Diskless Workstations Two Models Tarsnap This weeks episode of BSDNow was sponsored by our friends at Tarsnap, the only secure online backup you can trust your data to. Even paranoids need backups. Feedback/Questions Davi - BSDCan 2026 Follow up Send questions, comments, show ideas/topics, or stories you want mentioned on the show to feedback@bsdnow.tv Join us and other BSD Fans in our BSD Now Telegram channel
Topics covered in this episode: Some more things about Django I've been enjoying Who cleans up after the vibe-coding party? Where Did All Your AI Tokens Go? AgentsView to the rescue! Careful with phishing all Extras Joke Watch on YouTube About the show Sponsored by us! Support our work through: Our courses at Talk Python Consulting from Six Feet Up Connect with the hosts Michael: Mastodon / BlueSky / X / LinkedIn Calvin: Mastodon / BlueSky / X / LinkedIn Show: Mastodon / BlueSky / X Join us on YouTube at pythonbytes.fm/live to be part of the audience. Usually Tuesday at 7am PT. Older video versions available there too. Finally, if you want an artisanal, hand-crafted digest of every week of the show notes in email form? Add your name and email to our friends of the show list, we'll never share it. Calvin #1: Some more things about Django I've been enjoying Julia Evans is learning "2010-style" web dev (Django + SQL + server-rendered HTML) after years of Go backends and JS-heavy frontends Query builders: likes defining custom QuerySet classes with chainable filter methods (.approved().future().with_tags()) — more readable than raw SQL Template filters: highlights urlize, linebreaksbr, json_script, and especially querystring for building/modifying query-string links in templates Migrations: still loves Django's auto-generated migrations — 19 and counting on her project Skips inheritance for class-based views; prefers function-based views for sharing code, though fine using Django's own mixins/interfaces Performance surprise: CPU profiling (via py-spy) — not slow DB queries — revealed the culprit; she'd accidentally disabled the cached template loader, and re-enabling it took throughput from ~2-3 req/s to ~12 req/s on a $10/mo VM Michael #2: Who cleans up after the vibe-coding party? FT Magazine piece by Sam Learner (July 11) on AI coding tools overwhelming open source maintainers - sent in by listener Dylan McConnell, whose main point was that this ran in the Financial Times, not a dev blog. cURL as the case study - Daniel Stenberg has been the only full-time person on it for years; libcurl has been installed an estimated 20+ billion times with 3,000+ listed contributors. Bug bounty killed - cURL ended its paid security bounty program in January, citing an "explosion of AI slop reports" that take real time to debunk and drain morale. Extractive contributions - authoring a PR is now nearly free, reviewing one still costs a human; tldraw's Steve Ruiz closed outside contributions entirely, asking why he'd want someone else writing the easy part. Guido weighs in - van Rossum says projects are holding emergency meetings over the slop flow, and notes LLM patches tend to touch unrelated parts of a file, making review more tedious. "Vibe Coding Kills Open Source" - paper from Miklós Koren's group: packages frequently recommended by coding models saw big download jumps with no matching engagement, breaking the reputation loop that sustains maintainers. Stack Overflow flatlined - over 100,000 questions a month before ChatGPT, under 1,500 last month, with the response rate cut roughly in half; the public archive is now stale training data. The course-creator angle - Josh Comeau's newest web dev course launched at about a third of prior enrollment, and he worries about devs who never learn which questions to ask. But the most interesting portion is what was omitted. Focused on: The end of the curl bug-bounty Omitted: High-Quality Chaos Why the omission is interesting It fits a narrative. The FT piece is a maintenance-and-decline story, and January-Stenberg is a perfect witness for it. April-Stenberg complicates it - same person, same project, better data, opposite direction on the specific claim being used. The tell is already in the article. Learner quotes Stenberg saying AI tools are much better at finding problems than fixing them. That's the April thesis in one line, and it goes undeveloped. Reason for the shift is process, not vibes. Killing the bounty removed the cash incentive and the venue change filtered the rest. Worth saying out loud, because "AI reports got better" isn't quite it - "no bounty plus a real triage platform" is closer. Joke too: Sarah O'Connor wrote a related piece (is this just before skynet launches?) Calvin #3: Where Did All Your AI Tokens Go? AgentsView to the rescue! Local-first desktop/web app for browsing, searching, and analyzing your past AI coding agent sessions (Claude Code, Codex, Copilot, Cursor, Gemini, Aider, and dozens more) Auto-discovers session files on your machine — no config needed; everything stored locally in SQLite, no cloud/accounts agentsview usage is a drop-in ccusage alternative — reads from pre-indexed SQLite, reports run 80–220× faster on large histories New Activity dashboard shows peak concurrency, active vs. idle time, agent-minutes, and cost — filterable by project/agent/machine, with a -json CLI report too Full-text + optional semantic search across every session; also imports Claude.ai/ChatGPT chat exports Install via pip install agentsview, uvx agentsview, brew install --cask agentsview, or download desktop binaries from GitHub Releases Michael #4: Careful with phishing all The situation I pass this along because it was a pretty sneaky bit of targeted phishing, and happened to play off an old interaction in bandit's repo. As usual with phishing scams there are a bunch of tells that this isn't legitimate, but just enough plausibility that I could see falling for it in a weak moment. Relative nobodies like me haven't historically been worth the effort to hit with scams this specific. Agents change the game though :-/. Be careful out there folks! Original message From: "Patrick (Blacktrace)" [HTML_REMOVED] To: LISTENER EMAIL Subject: Your Bandit #1350 (B105 NextToken false positive) -- just fixed that exact case Date: Wednesday, July 15, 2026 12:02 AM Hi AJ, Saw your Bandit issue #1350 -- the B105 hardcoded-password false positive on the string NextToken. I build a deterministic gate that filters that class of Bandit noise, and #1350 was literally the case I just fixed: NextToken / next_token / page_token / nextPageToken now stay quiet, while a genuine hardcoded token like api_token="sk-live-..." still fires. Verified against your exact case. 30-second paste: https://blacktrace.co/noise-eraser Where it still trips, published: https://blacktrace.co/kruc Curious whether it clears what you hit -- and if it trips on something of yours, that's the more useful reply. Patrick, Blacktrace I asked Claude for some analysis too. It was pretty good at finding them. The message name-drops enough real detail to feel legit, but the structure is pure phishing - everything in it exists to get AJ onto blacktrace.co. The strongest ones: Freemail sender, corporate signoff. Signs as "Patrick, Blacktrace" but sends from emailpjv@gmail.com. Real company outreach comes from the company domain, not a personal Gmail - and there's no last name. Over-specific targeting. It mirrors AJ's exact public activity - issue #1350, the B105 rule, the NextToken false positive, even the token variants. That's the "just enough plausibility" AJ flagged, and it's exactly what agents make cheap: scrape a GitHub issue, auto-generate tailored bait. Legit cold outreach rarely reads your history back to you this precisely. The entire payload is two links. Strip the technical flattery and the message is just "paste here" plus "see results here." When the whole point of an email is the click, that's the tell. "30-second paste." Low-friction urgency, and "paste" most likely means paste your source into their tool - handing your code to a stranger's site. Exfiltration dressed as convenience. Brand-new, no-reputation domain. blacktrace.co has no track record, and the name is doing some ominous work. The /kruc slug is random noise, not how real product pages get named. Precise-sounding jargon that's actually vague. "Deterministic gate," "noise-eraser" - impressive, empty. Bolted onto correct real details (B105 is the Bandit hardcoded-password test, sk-live- is a Stripe live-key prefix) to borrow credibility. The disarming close. "if it trips on something of yours, that's the more useful reply" - engineered humility that flatters your expertise and baits a response. Makes engaging feel like you're doing them a favor, which drops your guard. Extras Calvin: DjangoCon US 2026 is rapidly approaching, August 24-28, Chicago Ruff v0.16.0 massively expands its default rule set Ruff now enables 413 rules by default, up from 59 https://astral.sh/blog/ruff-v0.16.0 Michael: Completely redesigned the home page. Try /insights in Claude Code (terminal) Joke: We're Safe
There are roughly 100x more people who use code than who can write code. As code that “just works” becomes easier to generate, this group may be the biggest prize of all — if you can get the agentic interface right.A key trend we have been tracking over at AINews is the absolute explosion in Codex usage this year, with MAU now up >10x from Jan 2026. Less than two weeks after their July 9th launch, OpenAI said ChatGPT Work and Codex had reached 10M users combined (as we cover in the pod, Codex now powers ChatGPT Work, so all ChatGPT Work users are now users of the Codex harness, even if they aren't traditional engineers) — showing the early innings of what happens when you graduate from coding agents to knowledge work agents:We've been calling out how coding agents are “breaking containment” to do everything else this year to power every other part of knowledge work - and it started with the org chart, with a major reorg last month that amounted to two of Codex's most prominent leaders, Greg and Tibo, taking responsibility over product and ChatGPT specifically, completing a “Superapp” consolidation cycle first discussed in March.With these updates Codex is no longer just a coding tool. In June, OpenAI said knowledge workers already accounting for roughly 20% of Codex's user base and growing more than 3x as quickly as developers. A product dedicated for knowledge workers was being pulled out of the Codex team.However, knowledge work has a different set of problems and environments than coding. For decades, knowledge work has been scattered across different primitives like documents for writing, spreadsheets for analysis, slide decks for communication, and specialized applications for everything else. ChatGPT Work now enables users to work across every primitive with agents. Instead of opening an application and manually operating its features, the user can describe an outcome and collaborates with an agent that can assemble the tools, context, and artifact needed to reach it.From building no-code products at Airtable to leading Productivity Engineering at OpenAI, Akshay Nathan has spent much of his career trying to make the power of software accessible to people who do not write code. In this episode, Akshay joins swyx and Vibhu to unpack the launch of ChatGPT Work, why Codex unexpectedly took off among non-developers inside OpenAI, and the company's broader plan to bring useful agents from software engineers to knowledge workers and eventually everyone.We go deep on the shared agent harness behind Codex and ChatGPT Work, why OpenAI brought the experiences together without making them identical, and how persistent computers, artifacts, Sites, plugins, memory, and sub-agents are changing what people can delegate to AI. Akshay explains why some teams are replacing decks and spreadsheets with interactive websites, how agents can gather context across code, Slack, documents, and local files, and what OpenAI learned from personal-agent products like OpenClaw.Side note: also don't miss Abhihek's sandbox track keynote at AIE, which now powers a lot of the sandboxing for ChatGPT Work… and yes was also broken by an unreleased OpenAI model in the recent HuggingFace incident.Akshay also reflects on how AI is transforming product development itself: why more people will become generalists with a specialty, why ideas and taste become the bottlenecks when almost anyone can build, why LLMs still struggle to generate genuinely grounded new ideas, and why teams must distinguish increased motion from actual progress.We discuss:* Why Codex unexpectedly took off among non-developers inside OpenAI* Why employees felt like using Codex gave them a new superpower* The product insight that led OpenAI to build ChatGPT Work* Why Codex and ChatGPT Work share the same underlying agent harness* How their UX, Git visibility, artifacts, and sandboxing defaults differ* Why OpenAI merged its agent experiences instead of building separate products* How AI is blurring the boundaries between engineering, design, strategy, and operations* Why OpenAI wants the default model configuration to work for most users* When power users should use deeper reasoning, Ultra, or multi-agent modes* Artifacts, agentic spreadsheets, and creating high-fidelity work products* Why interactive Sites may replace decks and spreadsheets* The challenge of designing a simple interface for an agent that can build almost anything* Why users should retry tasks that models could not handle three or six months ago* How AI can gather context for performance reviews without replacing human judgment* The OpenAI automation that turns internal Slack and document activity into memes* What reaching ten million ChatGPT Work and Codex users means for the product* How OpenClaw inspired persistent environments, scheduled tasks, and personal agents* Using ChatGPT for financial planning, budgeting, workouts, meals, and household management* The design tradeoffs behind sub-agents and how much of their work users should see* ChatGPT memory, Chronicle, and long-term context* Why AI may make more people generalists with deep specialties* Why ideas and taste become more important when almost anyone can build* Why LLMs still struggle with the instruction “bring me new ideas”* Measuring productivity through quality at-bats instead of commits, tokens, or pull requests* The critical difference between AI-generated motion and meaningful progressAkshay Nathan* LinkedIn: https://www.linkedin.com/in/akshaynathan/* X: https://x.com/akshaynathan_Timestamps00:00:00 Introduction and Bringing the Power of Code to Everyone00:01:33 Joining OpenAI and Preserving a Startup Culture00:02:40 What OpenAI Learned from Enterprise AI Adoption00:05:28 Why OpenAI Built ChatGPT Work00:07:17 Codex vs. ChatGPT Work and the Shared Agent Harness00:12:07 Why OpenAI Merged Its Agent Experiences00:16:24 Models, Reasoning Levels, and Choosing the Right Default00:20:26 Artifacts, Agentic Spreadsheets, and Model–Product Collaboration00:24:22 Why Sites Could Replace Decks and Spreadsheets00:30:08 Designing an Agent That Can Build Almost Anything00:34:28 From Developer Agents to Knowledge Work—and Everyone00:36:07 Power-User Advice and AI-Assisted Performance Reviews00:40:41 OpenAI's Internal AI Memes and the Ten-Million-User Launch00:44:39 OpenClaw, Personal Agents, and ChatGPT as an Operating System00:50:24 Sub-Agents, Ultra Mode, and How Much Control Users Need00:54:39 ChatGPT Memory, Personalization, and Chronicle01:00:19 How AI Is Reshaping Product Development and Tech Roles01:03:15 Ideas, Taste, and Why LLMs Struggle to Generate New Ideas01:04:42 Measuring Productivity, Quality At-Bats, and Motion vs. ProgressTranscriptIntroduction: Akshay Nathan, ChatGPT Work, and the No-Code ArcSwyx [00:00:00]: We're here in the studio with Akshay from OpenAI. Welcome.Akshay Nathan [00:00:07]: Thank you.Swyx [00:00:08]: And with our trusty co-host, Vibhu. So you recently launched ChatGPT Work. You lead Core Product Engineering. It's been a long journey, into all this. I find it very interesting that you started with no code or low code, with Walrus and Airtable. And to some extent, ChatGPT Work is like the super app of super apps of, well, here is the ultimate no code. You just write a prompt.Akshay Nathan [00:00:32]: Yeah. It's funny how things come, full circle. I think for a long time in my career, I started my career working consumer fintech, but then after that, like, there's this hypothesis that, the things that we were able to do with code, like, as engineers, like, if we could bring that to many more people in a more, accessible way, then that would be truly magical. We were working on a startup. It's funny, like, before LLMs, before vision LLMs, on how to do automated testing with AI. It was just kinda jank, back then, but doing what we can, and then worked at Airtable for a while on the same thesis that, like, if we can bring a database or the primitives behind a database to people, that'd be really useful to them. But once LLMs came onto the scene, it became clear that, this was the missing piece, like, the missing technology required to, like, bring the magic of code to everyone without them having to know what's going on underneath the hood. And so, like, I think this launch and a lot of the stuff that we've been up to is, like, the manifestation of that.From Walrus and Airtable to OpenAIVibhu [00:01:33]: How was stuff when you joined? So you joined OpenAI 2023. Now we've got, so much more stuff, so ChatGPT, Codex app, ChatGPT Work. Have things changed?Joining OpenAI and What Hasn't ChangedAkshay Nathan [00:01:44]: I think the more interesting thing is how things haven't changed. Like, one, I joined I remember when I joined, it was, like, five hundred people. One thing I was worried about was, like, I was looking for something, more early stage and, like, was it gonna feel startup enough? And I joined, and I was like, “This feels even more startup-y than I could ever imagine.” And, like, that really hasn't changed even till now. I think the, like, level of, like, bottoms-up ambition and, like, the ability of anyone to, like, do anything or have an idea and ship it is really cool. But on the, like, mission side, I think what was really compelling to me is this mission of, bringing frontier intelligence to everyone. Like, building AGI and then bringing it to everyone. And, I think acknowledging back then that, like, that vision is gonna, not be a linear progression. Like, we're probably gonna, like, try different products and have different things that succeed and don't. But the vision has stayed the same, and the mission has stayed the same, and we're starting to see the pieces, fall together, and that's really cool.Enterprise Lessons: No One-Size-Fits-All AISwyx [00:02:40]: You worked on Enterprise. What A lot of people never touch ChatGPT Enterprise. What is something that you learned from there that you're bringing into your work now?Akshay Nathan [00:02:52]: I think how there's no one-size-fits-all solution in Enterprise. I remember in the early days of ChatGPT Enterprise, like, when we talked to customers and, like, everyone. That was, like, when I think it was a year after ChatGPT was released, and everyone was so excited to bring, AI into their enterprise. And, there were all these teams being stood up. It was, like, the AI deployment team with, like, these enormous budgets. And if you asked anyone, like, what were they excited about? Like, what were they excited about solving? Like, at first, you'd get, like, kinda like the baseline answers of, like, “Yeah, we have all this context and data and all this stuff.” But then if you ask them, like, “What was, like, a discrete use case that, like, they want AI to enable in their workplace?” You get such a different, like, variance, like, explosion of, different types of answers. And it's interesting, like, you using, like, these models and these products, you have this box, and you can say anything to it, which is the magic. But it'on the flip side, it also means that, like, you don't know what to do with it. And in Enterprise, I think a big part of that is, like, meeting the users where they are, like, what use case were they trying to solve, and then teaching them how they can use AI to, like, gain leverage there.Swyx [00:03:56]: Do you meaningfully differentiate that from forward-deployed engineering?Akshay Nathan [00:04:01]: I think there is the go-to-market side of it and then there is the product side of it. I think you need someone on the product side. And I think, like, however good we get at FDE motion, like, I think at the end of the day, if we have a user who's, like, looking at their computer or looking at their phone, like, it's our job in the product to, like, be enabling them and showing them where to go. So we're really excited about that.Vibhu [00:04:24]: Do you think there's been changes, over the past three years of adoption? So there have been, step function changes. You have reasoning models and whatnot. Is there still the same problems of Enterprise has black box, don't know what to do with it, or have things changed?Adoption, Agents, and the Next 10x MarketAkshay Nathan [00:04:39]: We're seeing now that, like, there's this huge uptake, right? Everyone is extremely excited about it. It feels like, many people are, millions, hundreds of millions of people are using ChatGPT. They understand, like, how generally to work with AI. But then, like, every time, like, a new capability gets unlocked, so now, like, we're seeing with agents, like, there is probably a contingent of, like, early adopters still who, truly get it, who are like, “ we you can do anything. You just have to make sure the right context is there, it's connected to the right tools, and that you are supervising it, but, like, anything is possible.” But then there's, like, this, like, 10x or 100x bigger market where, like, they don't yet get that, or they don't yet see that. And so I think that's the next stage here. So to answer your question, like, I think the adoption is there and growing fast, but I think the opportunity is, like, far bigger than that. That's where we wanna play, especially with ChatGPT Work.ChatGPT Work, Codex, and the Super App MergeSwyx [00:05:27]: Yeah. well, let's, let's skip ahead to ChatGPT Work. only, like, a month ago or so, announced. what was the decision process that led into it? there was this, overall merging of the super app. Is that what we're officially calling it? you deprecated the browser as well. Just, summarize your last, like, couple months of working on this thing.Akshay Nathan [00:05:50]: Yeah. It feels like forever now, but it's only been a few months. I think maybe the one, impetus that, like- Is most salient is when we release Codex, or even internally had Codex, like, it was really surprising to us, I think we recently put out some stats on this, that there was this, like, real inflection of, like, adoption among non-developers at OpenAI. And, I, through this product development process, like, would go to, like, these UXR sessions to talk to people internally. And the thing that stuck out to me is, like, one, like, you go talk to, like, strategic finance or marketing or whatever, and they're all using Codex for, their use cases. That part's cool, but the thing that really stuck out to me is how proud people were that they were using Codex. Like, how, likeSwyx [00:06:34]: It's like, “I'm not supposed to be using it, but I am.”Akshay Nathan [00:06:36]: It was that. It was, like, that they were, early to this, like, new thing, but it was also this thing of, like, they felt like they had a superpower, right? And, what we recognized then is that, like, the power of Codex, the power of agents, like, we already had this massive distribution base of people who have, come to know and love ChatGPT. Like, how do we show that to them? Like, how do we bring it to them? Which is, like, a hard product problem, and it's, like, a tricky thing, right? There's many ways you can go about it. And so that's what we called the Merge and the Super App over time, and ultimately launched it in ChatGPT Work, is how do we do that? But it came from that initial realization that, like, the power was not only for developers, like, much earlier than probably even we thought. Like, it could be extended to everyone.Swyx [00:07:17]: How do you see the products differently? So, like, who is it for, right? So Codex started out even CLI, then app. Now there's a merge of ChatGPT Codex and ChatGPT Work, so is it the opening for the average user, for enterprise, for work? How do you position it?Akshay Nathan [00:07:36]: I think we want to get it to position it for if you're doing work-related things, for lack of a better word, right?Who ChatGPT Work Is ForAkshay Nathan [00:07:42]: I think productivity is what, like, the pillar that I support. Like, that's the name of the team. And the reason for that, the reason we call it productivity and not, like, enterprise or, like, work or something like that, is because there's also personal productivity, right? And, like, I think ChatGPT Work is I've seen people do things in their personal lives that you wouldn't classify as, like, work technically, but, like, these agents are, super capable for. Like, one recent example that someone posted about, on our Slack is, like, someone had, like, a missed package, like they didn't receive it, and then they got, like, the picture of it, from Amazon or whoever the courier was, and they, like, asked ChatGPT Work to, like, find out where that package is. And, like, the agent, is extremely tenacious and, like, took the image and, like, looked at a bunch of, like, listings around their neighborhood and figured out exactly the apartment complex in which the package was, like, gave them some information. And so, like, I think there's all these things that, like, you, work-related or productivity-related things, I think that's what we want the product to be. You asked about Codex. I think we think Codex is, a durable brand, but we have a principle that, like, the user we don't want a user to get stuck in a tab or an experience where they don't get the power of the product. And so, like, everything that you can do, in the Codex portion of the product on desktop, you can do in ChatGPT Work and vice versa. But we made some opinionated product decisions on, like, how much of the Git state, if you're in a Git repo, do we wanna expose to the end user? Or how much do we wanna make the experience of seeing the agents thinking, like, diff forward so that you get exposed to the diffs out of the box. And then, like, on the safety side, like, how do we wanna think about, like, sandboxing and making sure that we have the right defaults in one state versus the other? So, there's, like, some opinions that go behind that, but we do want We don't want the user to need to choose which experience they're in.Swyx [00:09:26]: That is a good goal for AGI, right? Like, people don't want, like, to hide to choose what version of AGI they want. They just want the AGI to decide for them. can I get an answer or, like It's not super clear to me. Is the Codex harness and the ChatGPT Work harness the same? Is it just UI affordances, or are there prompt level or even deeper differences?Shared Harness, Different UX: Codex vs. WorkAkshay Nathan [00:09:49]: So the harness is the same. The harness is shared. on In both of the products, we made improvements to the harness to make it good for knowledge work, especially as it relates to plug-ins or computer use or artifacts. You get that power regardless of which experience you're in. On the UX side, there's opinionated takes that we have when you're in Codex mode, what the UX should be how the UX should behave, and some stuff around the sandbox like I mentioned, but the underlying harness and capabilities should be the same.Swyx [00:10:16]: I'm just kinda curious. Maybe we can, -- Is there a query that we can run that would look different in the two modes?Akshay Nathan [00:10:23]: Yeah. I tried to create, like ask it to create, like, a retirement calculator spreadsheet or something, in both modes. And then in Codex mode, you might have to be in a repo for this, but you'll see, like, the diffs of, like, the sheet that it's creating and stuff like that, and the file edits. But in Work you won't be able to see that.Swyx [00:10:42]: I think that's, that's super clear. And then also the other thing I wanted to dive into was your, the productivity team. what else is there? first of all, what are the top-level teams other than productivity? Isn't productivity everything?Productivity Teams and Core ChatAkshay Nathan [00:10:55]: SoSwyx [00:10:55]: Science?Akshay Nathan [00:10:55]: We have a team focused on ChatGPT. Like, the core chat experience, for consumer, which is like, not, I think all productivity. Like, there'People are using ChatGPT every day for search to, figure out how to write messages to loved ones, to think about, how to, like, learn a new topic, et cetera. And so there's so much more inside to create images. And there's so much more in chat that, the hundreds of millions of users are using that warrants, like, a very dedicated effort. And there's teams focused on enterprise and infrastructure and API and stuff like that, so.Swyx [00:11:33]: I will bring it up.Retirement Calculator Demo and Git-First UXSwyx [00:11:34]: Yeah. So I have them both running. This is ChatGPT Work. There's a Codex version here. I picked “Five Little Ducks” song, so this will take a while.Akshay Nathan [00:11:43]: Huh.Swyx [00:11:43]: I think we'll just keep it in the background and, as they finish, we'll look into some of the differences.Akshay Nathan [00:11:48]: Yeah. But immediately, I think if you flip back to the Codex version you'll see that,Swyx [00:11:53]: That it assumesAkshay Nathan [00:11:54]: Like theSwyx [00:11:54]: It assumes Git. Yeah. Yeah.Akshay Nathan [00:11:56]: The, like, dynamic island assumes that you're in a Git repo. And you might miss some stuff because some of it is, like, in the actual chain of thought with those changes and how we display that, but yeah.Swyx [00:12:07]: Is there an unintuitive like, is there a thing that you wanted to ship and then you got feedback, and you were like, “No, let's not do it?” Like, what's the thinking behind that?Why Merge the ExperiencesAkshay Nathan [00:12:14]: In, ChatGPT Work?Akshay Nathan [00:12:17]: I think one direction we could have gone with this is, like, keeping the experiences, like, completely separate. So it's like, whySwyx [00:12:22]: Different apps.Akshay Nathan [00:12:23]: Exactly, like different apps or even in the same app, like different, completely different experiences. Like, why merge it all? Like, what is. Codex, people love. Like, why bring these products together? And I think the intuition here is that, like, all of our jobs are, like, changing dramatically with AI. Like, for, like, every few months, like, I feel like I wake up, and I'm, like, doing a completely different thing than I was doing a few months ago. And my hypothesis here is that, or I should say our hypothesis is that, like, part of what we're, we're building, this technology is giving people leverage. Like, the things, maybe it's the more mundane parts of your job or parts that, like, if you were able to automate, you'd be able to share more ideas faster or whatever, like, you're able to do now. And because of that, like, that might blur the lines between someone who's, like, only writing code or creating strategy docs or, planning events or, helping with marketing or doing podcasts or whatever, right? And so, like, these things are gonna get blurred over time. And so, like, trying to draw a hard boundary based on, like, the who you are is gonna be, is gonna be tough. And, like, we should enable users to choose, but we shouldn't box them in. And so a lot of the work that went in here, like, keeping the primitives the same, like for example, plugins are, like, unified across, this product and ChatGPT and the cloud, was because of that. It's this thesis that, like, eventually things are gonna come together and we don't wanna be Like, we wanna be prescriptive about when to be in either experience, but we don't want to box anyone in.Swyx [00:13:45]: I wonder if there's users who are very tuned to the old ChatGPT harness that is effectively now replaced by the Codex harness. I can't imagine what that was, but maybe they're more the more conversational side. Can you compare and contrast the two harnesses? ‘Cause only you've seen it.Akshay Nathan [00:14:02]: Yeah. I think ChatGPT, the existing harness, like, still exists today. Like, it exists in this app,Harness Engineering: ChatGPT vs. CodexSwyx [00:14:08]: The classic, right?Akshay Nathan [00:14:09]: TheVibhu [00:14:09]: You just start a new chat, and you don't go under Work, right?Akshay Nathan [00:14:13]: Yeah. If you startVibhu [00:14:13]: SoAkshay Nathan [00:14:14]: A new chat and go to chat, then you're, you're talking to ChatGPT with the instant model.Vibhu [00:14:16]: Oh, we can technically do another. But on instant.Swyx [00:14:21]: Yeah. So this one's not gonna code or it's gonna be in line. It's on a in line in a sandbox.Akshay Nathan [00:14:26]: It'llVibhu [00:14:27]: Oh, that's coolAkshay Nathan [00:14:27]: We try to push you to go to Work if you're creating a spreadsheet. Yeah, but this isSwyx [00:14:30]: And this is a router decision? Sorry. Is it a router decision?Akshay Nathan [00:14:34]: This is the decision that, the model is making, and then, like it sees that you're able to. or you're trying to do something that would be better served in Work mode. But I think your question was like, what are the advantages of, like, the chat, like ChatGPT chat harness?Swyx [00:14:48]: It's more broadly, like, I wanna, do an oral history of harness engineering. Right? the ChatGPT harness lasted us from, let's call it the ‘01 era, until now, and now it's being replaced by the Codex harness effectively. And they're, they're overlapping somewhat, but I'm curious what changed if there is.Akshay Nathan [00:15:10]: My perspective on this is, like, there's, there's, there's there's like a constant process of, like, divergence, convergence, divergence, convergence. And in chat, like, many of the use cases I was talking about before, like, search or learning, I think we're, we're really optimizing for latency and optimizing for personality and, like, different things that, over time, like the product The reason people love ChatGPT is because we've been optimizing for those things and working on them for so long. Codex, what we learned was that, like, if you give the agent access to this infinitely flexible environment as a computer, it can do really powerful things. And so when we think about, like, okay, well, for knowledge work, like, what is which mode should we choose? It was like it felt more natural to us to bring that to this, like, computer environment and, maybe abstract some of the details of this computer away from users who might not be used to that, but, like, give them that same power. But ultimately, I think that we want the power in all places, right? We wanna meet people where they are. So I'm sure there'll be work down the road in order to get things to be, equivalently capable in all scenarios. But it's just a question of, like, what we've been focusing on the product on historically and what we're focusing on now.Models, Defaults, and the Reasoning SliderVibhu [00:16:24]: I think alongside that, outside of just harness and when to use Codex, ChatGPT, or Work, there's also the new models you've released, right? any guidance there? So people love to min-max what to use, like only use Terra on high reasoning versus, for this, you wanna use Sol here, ignore all theseAkshay Nathan [00:16:44]: There's 32 options.Vibhu [00:16:46]: But, that being said, for people that are expanding, so, productivity trying stuff for work that don't have the breakdown of what all this is what's, what's the advice, right?Akshay Nathan [00:16:59]: Well, I think before the advice, like the first thing is, like, none of this would be possible without these models. Like, the, I think you asked earlier, like, what was, like, the inspiration for work and, like, early on, like I mentioned, like, what we were seeing with Codex, but that was also because the models were getting infinitely more capable. That's happening again. I think it's like another step function jump now. And to answer the question on advice, like we want this default to be the best possible. Like, we wanna be opinionated about the default, and so we've we've chosen a default that we think is gonna be the best for everyone. And, we have for power users options under the hood. We could One could argue that there might be too many right now, and we're, working on simplifying it. But you can extend, the reasoning level, and you can change between the different model classes if you need to, but the default should be the best for most use cases. So my advice to most people would be to stick to that. And then, if you reach a situation in which you think that you could, you wanna try, a different configuration, if you're not seeing either the efficiency on the cost side or the quality on the intelligence side, then you can change the defaults and see if you can get something better. But we think that the default should be good enough.Swyx [00:18:09]: I have, I'm just gonna run something by you since you have way more experience than me. I've recently been doing Sol Lite but with goal, with the idea that the goal augments the reasoning effort, but with more terminations and turns.Swyx [00:18:24]: Is that a good way to think about it as opposed to Sol Ultra or Sol, Extra High?Akshay Nathan [00:18:29]: Yeah. It's hard to say becauseSwyx [00:18:31]: Yeah. It's like an interaction effect.Akshay Nathan [00:18:33]: exactly. It's like there's a preference on, for you as an individual, like how do you like to collaborate with the models? Like how many of those like terminations, as you call them, do you want where, you can steer or make sure that it's doing the right thing?Akshay Nathan [00:18:46]: I think generally people should try whatever works for them. I think that like using Ultra or the like multi-agent setups are best for like when you have like tasks that are either incredibly complicated, like open explorations or very paralyzable. I think even for tasks using goal, I think is best for tasks that you'll be able to make consistent progress in a way that's verifiable over time. But I think for most tasks, they don't fall into either of those buckets. And so like at least when they're starting, and so that's why I think the best first step is like trying it with the default configuration and then seeing like where you wanna go from there.Swyx [00:19:29]: Right. You guys worked on a slider, which is super helpful for reducing the amount of panic.Vibhu [00:19:36]: It's nice on mobile at least. There's a nice slider there.Swyx [00:19:38]: It's nicer.Vibhu [00:19:39]: I haven't tried it.Swyx [00:19:40]: So you have the advanced view there, but if you click advanced view. Yeah.Vibhu [00:19:44]: Ooh, it's just a nice slider. Yeah.Swyx [00:19:46]: Very pretty, very colorful.Akshay Nathan [00:19:48]: Yeah. The idea was here was like reduce it to like one dimension even though there's multiple dimensions, right? Try to project it onto a single dimension for the user. Like, something from that represents like, speed and efficiency on one side and then like quality and thoroughness on the other side.Artifacts, Spreadsheets, and the Work LaunchSwyx [00:20:04]: I am just puzzled that it uses Sol so much, like the lowerVibhu [00:20:07]: NoSwyx [00:20:07]: Grounds I would've usedVibhu [00:20:08]: I think the slider, if I'm not mistaken, isSwyx [00:20:09]: Terra.Vibhu [00:20:10]: Oh, it is.Swyx [00:20:11]: Yeah. See? So they preset Terra to only be the light one. But like I think a lot of people would more people should use Terra. One, because Sol keeps running out of capacity.Vibhu [00:20:22]: I'm the reason. Here's ten minutes of ourSwyx [00:20:24]: There you goVibhu [00:20:25]: Retirement calculator.Swyx [00:20:26]: Oh, that's the Excel thing working for you.Vibhu [00:20:28]: This is,Swyx [00:20:28]: Oh my God. Look at thatVibhu [00:20:28]: This is work, and then Codex is still cooking, so we'll get back into it. I think it'll be interesting to see the thought process, the reasoning, and also, this is eight minutes on work. Codex is still cooking.Swyx [00:20:41]: Yeah. And by the way, so I've, do Gabriel Chua? He's part of the OpenAI Singapore team. He showed me this, and I was like pretty shocked that this looks like Excel. It edits Excel files. You never paid an Excel license, right? Like, but somehow this is like workable and it's agentic Excel.Akshay Nathan [00:21:01]: Yeah. one of the big like pushes that we made for this launch was like artifacts, right?Akshay Nathan [00:21:05]: Like both on the model side, like I think if you compare this with GPT-5.5 and GPT-5.4 before that, you'll see that there's been pretty dramatic improvements in the quality of these artifacts and then also on the product side.Vibhu [00:21:16]: The UX side is also crazy, like hosted sites and whatnot. No longer needing to host your own little webpage, like itSwyx [00:21:23]: Oh, I have a story about that. I can do, a separate thing. I'll need to take the visuals here, but we-we'll, we'll cut to that later. Was there co-training, because you were moving making this big move and you launched GPT-5.6 on the same day as ChatGPT Work? Was there influence between the model training teams and the harness teams, or did they did the launch dates just happen to line up the same day?Akshay Nathan [00:21:46]: I think the we collaborate heavily with the research teams, and I think that's like one of the most magical parts of the job, like the most fun parts of the job. But yeah, just using artifacts as an example. Like, a lot of what you're seeing, like underneath the hood, there's a lot of work that went into making sure that like, we had the right infra to be able to train the models to get better at this. And then on the product side, like had the right experience for users to be able to collaborate with the model on an artifact like this. In fact, like this whole viewer, like the intuition here is that like, it's not necessarily that you wouldn't need an Excel license. This is stage one, right? Like, this is probably not what you meant when you're like making a retirement calculator.Vibhu [00:22:24]: Yeah, you can iterate very easily. Yeah.Akshay Nathan [00:22:24]: You wanna iterate and like when you're seeing it, and if this thing is high fidelity to like what you would see in or what your coworkers would see if you were to send this to Sean, like that I think makes it so easier and makes you trust the product in terms of iteration.Vibhu [00:22:39]: When you say coworkers would see, do you see a multiplayer, multi-team collaboration with artifacts? Any things you guys think about that?Multiplayer Artifacts and CollaborationSwyx [00:22:46]: You can already share it, right?Akshay Nathan [00:22:48]: Yeah. It's inter It's something that, we're actively thinking about. one thing that, we've noticed internally without talking too much about the roadmap is that like there's many times when someone will ping me about something, and I will ask ChatGPT Work the question, and then I'll ping them back the answer.Akshay Nathan [00:23:04]: And then I'll be thinking likeVibhu [00:23:04]: Like the simplest would be, the three of us are just all on one hosted.Akshay Nathan [00:23:07]: Exactly. And I'll think about like was I required in this loop or and then maybe it was, rephrase like what they were asking or pulled from certain context or whatever. But like, when I gave them back the answer, that process was also lossy, right? Like I gave them just like my interpretation of what ChatGPT Work cooked up. But like underneath the hood, there's so much context like in the rollout and stuff that could be interesting.Vibhu [00:23:28]: Yeah, it'sSwyx [00:23:28]: So like the answer was preemptively respond to every inbound request?Akshay Nathan [00:23:33]: No, it was just like literally like this is what I do sometimes as my job.Swyx [00:23:36]: I know you copy-paste and then you're just a message forwarding serviceAkshay Nathan [00:23:39]: Yeah. Yeah, exactlySwyx [00:23:39]: From AI to AI.Vibhu [00:23:40]: But I think it's interesting, right? It helps people understand the capability of what you can ask and delegate that oftentimes people don't realize until they try or someone shows you, and then you're like, “Oh, okay. Okay, I see.”Swyx [00:23:52]: I think it's als there's also like a, light security issue, where like you're the permissions layer. Like yes, I could query everything that you query, and I could get an automated response, but maybe I'm not supposed to see it. And that there's no way I would know because I'm not supposed to know what I don't know.Akshay Nathan [00:24:07]: Especially as like, with ChatGPT Work, we're, we're asking you to connect your plug-ins and, it's pulling from your local files and stuff like that. Like the amount of context that the agent has access to is like- Deeply personal and like that's something I think we need to preserve, so that'll be definitely a challenge.Swyx [00:24:22]: There's Excel, there's PowerPoint, there's Docs, the, grand trio of work. What other formats of work do you think about? like you worked on Airtable. Is there a future where there's like OpenAI Airtable? Like what does that look like if you ever ended up doing it?Akshay Nathan [00:24:41]: It's a really good question. I think,Formats of Work: Sites as Knowledge ArtifactsAkshay Nathan [00:24:43]: one that you didn't bring up was Sites, and I think that wasSwyx [00:24:46]: SitesAkshay Nathan [00:24:46]: A core part of this launch. There's one side of Sites that I think people commonly talk about, especially on Twitter and stuff or X, of like, this like prototyping tool. And like we saw that happen with this launch even. The model slider that you guys were referencing earlier, like that was developed almost fully in a Site. Like, the collaboration between design and engineering and product on that was like on a site where we play with, the affordance and figure out how it feels and all of that. But the other aspect that I think is a little bit less talked about is like Sites as like an artifact for knowledge work. I was talking to someone the other day who's on like our corporate finance team, and like we were mentioning how like now when they have these reports that they're, they're working on as a team month to month, historically those things were in slide decks and in spreadsheets, and now they're just in Sites. And like Sites is the mechanism that they collaborate across the team. And the reason is ‘cause it's like, it's like somewhat higher bandwidth. Like, at these tools like PowerPoint and Excel are like infinitely flexible, but at some point you reach the boundary of like either as a human you may not know how to use some feature or something, or the product itself doesn't support it. But with a site you can do anything. You ask for anything and you can get that. once people see that magic, I think it's been really valuable.Swyx [00:26:02]: Yeah, let me show you my case study. this involves all the hot topics including ChatGPT Work, but also GPT-5.6 token billionaires and token maxing and Sites and auto research. I'm a fan of this game called Strata. It's, it's like a little board game that youSites, Auto Research, and Research DashboardsSwyx [00:26:17]: That you play with, physical blocks, that come on top of it like that. So over the weekend I took like thirty photos and just threw into ChatGPT. one point seven billion tokens later, out comes this site with a fully playable thingAkshay Nathan [00:26:32]: WowSwyx [00:26:32]: With 3D, block placement and everything. Because it requires physical blocks and I needed friends to train on it so they can get better, so I can play against them. But also, I could also, do things like train an AI on it and that's, thatAkshay Nathan [00:26:45]: That's your auto researchSwyx [00:26:46]: That gets into auto research. So, you want to train your own AIs, and then make sure they self-play against, each other. I need to set both AIs. So this is AI versus AI, and they're, they're gonna self-play. the AIs start out bad and then you want to define a loss function and get good. I wasn't gonna supervise all this. I was at, I was down in San Mateo, attending a conference. What I ended up doing was, auto researching and on this and creating benchmarks and that there was just way too many parameters for me to read. So I started asking it for a site, and it's created this lab, panel. Where is there a, is there a shortcut for a site that is created?Akshay Nathan [00:27:28]: You should be able to go in the sidebar to Sites, top of the sidebar. The left sidebar.Swyx [00:27:33]: This one? Oh, left?Akshay Nathan [00:27:35]: Yeah. Just scroll all the way to the top.Swyx [00:27:36]: Oh. Oh, it says Sites. Oh, there you go. Yeah.Akshay Nathan [00:27:39]: Ooh.Swyx [00:27:40]: So it create, it creates the sites. I don't, I don't think this is, it is exactly what I wanted, but let me show you what it popped up, right? Like I think as a research artifact, it is very important to communicate, exactly, what is being done. Outputs this thing which I eventually started publishing. So I moved it off of Sites because I wanted more, database and infrastructure than Sites afforded me. But this is like a research output that you can start to mess with and like try to think about like what hyperparameters are you tuning for training AIs. And like I was trying to make like scaling laws and everything and doing all sorts of like game optimization stuff. And the fact that you can just throw this up as a research artifact, like I no longer need to read ChatGPT output. I read Site output. But then there's also a huge sprawl. Like look at how long this thing is. There's so many numbers. It is pretty overwhelming, so then I have to start pruning it from there. But, it's an interesting transition from Markdown effectively that you're putting out to, you're putting out a whole functional site.Akshay Nathan [00:28:41]: I think Markdown just isn't that optimal for people to read, right? Might as well just write HTML website and I don't know. I think you can do a lot with customizing this, right? You have your skills that explain what you want. Like I noticed they're quite verbose. I don't need a lot of this information.Swyx [00:28:57]: It's very verbose.Akshay Nathan [00:28:58]: So and then the nice thing of having a site side by side is, you just iterate on what you want and what you don't, right?Swyx [00:29:05]: Yeah. I don't know if, any that triggers any stories for you of how it's run internally. Am I doing this right?Akshay Nathan [00:29:11]: Yeah. I think that this is like a workflow that we're seeing like all different types of teams use, where like the canonical artifact that was previously a deck or something is now becoming a site. And like with a site you, because it's just HTML, you can like. It's infinitely flexible. And so, if you want to give more prominence to a certain thing that like in a slide deck would, feel like it was buried, like you can do that. You can have it be like the hero image, right? And so I think that like, people are starting to see that. There's more work to be done to make these things like much more easier, easy to collaborate on. You mentioned that they're very, they're long and verbose, could be broken up. I'm sure that there's still something to do there.Swyx [00:29:53]: They're super long. Yeah.Akshay Nathan [00:29:54]: Yeah. But I think we're starting to see that like there is this aspect of this is a really interesting, format, for people to use, that's like much more flexible than what they ever had before.Swyx [00:30:07]: I think your job also comes becomes meta. You're not designing the products. You're designing a product to make products, and I'm curious how you manage that.Designing a Product That Makes ProductsAkshay Nathan [00:30:18]: I think one thing that we've been Like when we look at the UX, like that we've been thinking a lot about is how can we balance like simplicity with capability? Like if we're designing a product, like you said, that like is made to make up build other things, right? You can build so many different things. But we can't put that all in front of you because you'll get overwhelmed.Vibhu [00:30:41]: Yes.Akshay Nathan [00:30:41]: And so we had similar problem or similar challenges even Chat-with ChatGPT, but especially now, like when there's so much that can be done, I think the balance that we're constantly trying to strike is like, how can we give the user enough of a UI surface where, they can be expressive, they can tell the agent what they need, they can verify that it's using the right tools, it's pulling from the right sources, et cetera, but then it gets out of the way. And then how can we build the right system such that we can show them instead of telling them what can be done? Because so much of this is gonna be like, how do they discover the next use case and the next one after that if they really want to be super powered by the AI.Games, Private Evals, and Show-Don'TellVibhu [00:31:19]: Yeah. It's interesting. I feel like everyone also just has a different way to do it, right? I made a similar version of this same game. I didn't take any pictures of board or rule game. I threw in at goal eighteen minutes, fifty-three seconds later, a lot of tokens later, I've got a similar version. not with all the auto research and whatnot, butAkshay Nathan [00:31:39]: You gotta do all the latest trends.Vibhu [00:31:40]: And yeah, I did it with, did it with Codex, not Work, but it's interesting, right?Akshay Nathan [00:31:45]: Yeah. And this is GPT Image generating the pro avatars. Very good for game design. LikeVibhu [00:31:51]: AndAkshay Nathan [00:31:52]: A lot of game designers were like really into GPT Image for assets.Vibhu [00:31:54]: I will say like the broader takeaway probably is the reason that we do this is more so just to test the tools, right? Like, this was also a test for GPT-5.6 came out. I had done the game on GPT-5.5, right? The ability for me to no longer need it to. I had to feed it the rules. It's, it's a pretty niche game. It couldn't find how to do this on its own.Akshay Nathan [00:32:15]: Oh, yeah.Vibhu [00:32:15]: GPT-5.6Akshay Nathan [00:32:16]: It is out-of-distribution, which is why I was also very keen on testing the GPT-5.6 capability.Vibhu [00:32:21]: But, this is just as work comes out, as new things come out, these are just our side ways to test things, right?Akshay Nathan [00:32:27]: Yeah. It's some private eval. That is not this private.Vibhu [00:32:31]: But also valuable because now you can send this to your friends and I learned about this game through seeing this.Akshay Nathan [00:32:36]: It's a hard game. He's very good.Vibhu [00:32:39]: It's good to when no one is competing with you. But yes, it's a classic RL problem of like self-play, bootstrapping your game AI. yeah, you see how easily work becomes personal and personal becomes work because the thing I do for personal, it directly informs people I work with because I showed it to them. They were like, “Oh, you can do that with GPT?” Which like I imagine is the growth strategy.Akshay Nathan [00:33:02]: Yeah. The show not tell is a big piece that, I think we've we're not still not fully cracked of like, showing people all the things that they can do with the product versus like trying to teach that to them through like, articles or onboarding or whatever.Akshay Nathan [00:33:18]: So meeting them in the moment.Vibhu [00:33:19]: It's a career risk for me, because I used to be in developer relations, right? Where your job is to show, and then you're like, “What do you mean? You don't, you don't need.” your job is to tell. And then. But the product people are like, “Well, we don't need you if our product is intuitive enough.” SoAkshay Nathan [00:33:37]: Yeah. that's the magic of the models. So you can tailor the telling or the showing to like specifically what the user needs, like what they care about, what they've done in the past, exactly where they are on the adoption journey. So I think that's like gonna be a super big opportunity.Vibhu [00:33:50]: Seems easier and easier now to tailor custom showing, right? People have different use cases. As much as you said you don't wanna segment different people into different buckets, right? It's also not that hard to for people that are in different categories. But the question, is you said your team is more broadly on. What was the term you used? Productivity?From Developers to Knowledge Work to EveryoneAkshay Nathan [00:34:12]: Productivity.Vibhu [00:34:12]: Productivity. So howAkshay Nathan [00:34:12]: Which is now work.Vibhu [00:34:14]: Is it work? Is there another distribution that we're not hitting? Is there a group of people that will have something different than ChatGPT, Codex or Work? Is there more that the mass isn't targeting?Akshay Nathan [00:34:28]: I see it as like a sequencing, like. The vision is like bring useful agents to everyone. We started with like developers. Like developers historically are like early adopters that are willing to put up with more friction, set things up, et cetera. Like that's where, Codex started. I think the next opportunity is like what we call general knowledge work, all the other functions around developers. I think when you go from developers to this segment, like there's inherent challenges with like, this show not tell thing that we're talking about, making the product more understandable, bringing in new capabilities that matter more for this cohort than matter for developers, things like artifacts, things like computer use, et cetera. And then I think like the same learnings, like similarly how we took the learnings from developers and brought it to, general knowledge work, the next stage will be like taking the learnings from general knowledge work and bringing it to everyone no matter what they're doing in their lives. And we're already seeing that a little bit. Like this game example that you have is, something that's like on the border of like fun and personal life to, your professional life. I use ChatGPT Work full-time at home for everything, like for whatever I'm doing. I used it the other day to come up with a meal plan and like, save that on the like computer environment that it has and something that I can continue going back to. Like is everyone doing that yet? Probably not because the thing says work on it, but eventually, we wanna get people there.Vibhu [00:35:51]: ChatGPT life.Akshay Nathan [00:35:52]: Yeah, exactly. ChatGPT cooking. But I think there's a lot of, there's a lot of opportunity there, but I see it as like, we're, we're built we built a foundation in software engineering, and we're gonna take the same learnings that we take from software engineering to knowledge work to everyone.Vibhu [00:36:07]: Do you have any power user advice? I feel like, there's a group of people that will live it, use it for everything, stay on it twenty four-seven. And then there's a bit of a gap between that crew and people that, okay, I use it for work. I use it occasionally. Sometimes I type questions. any advice, any learnings, anything you recommend or just, takeaways that you've found that help bridge that gap?Power User Advice: Push the Frontier of ImaginationAkshay Nathan [00:36:30]: I think a couple things that I've seen is like, one, that it really helps to broaden your imagination of what's possible, and this has been a learning even for me. Like, the technology has progressed so fast that, something that, like, even three months ago, like, no way the models can do this. Like, now it's like, wow, it's like it can. Like,Swyx [00:36:52]: Give an exampleAkshay Nathan [00:36:52]: We're going through right now our, like, review cycle internally, and, people always talked about this as, like, a thing that the models are good at and like, there's a cliché of like: Okay, like, no one wants to be writing reviews and, like, we just use AI to do it. But in all seriousnessSwyx [00:37:09]: And it can evaluate it as well.Akshay Nathan [00:37:10]: Yeah, exactly. In all seriousness, before it was, like, just, like, slop and, like, I think it was helpful, but, not super productive. Now I've found that, like, the model can do a much better job than me, especially in this environment of, like, pulling context on, like, what people are up to, how they've like the things that they've done to make a difference, highlighting like, wins that they've had that, like, I might may not even have seen. It has access to, like, everything, right? Like the code, like, things that they've caught, reviews, Slack, everything. And so it's, like, incredibly powerful in that domain and, like, just like six months ago, the last time we did this cycle, like, I didn't even I tried using it, but it was not at all helpful. And this time it's been, like, incredibly helpful and, like, so I think continuing to push the frontier of imagination of what's possible, even if you tried something before, I think is maybe the my biggest piece of advice. The other, thing is, like, the more you put in, especially in this environment where, like, the model has access to everything on your computer or in ChatGPT Work, like you can create, artifacts over time and save them in your library and, like, the model will continue having access to those. Like, the more information you give it about whatever domain you're in, whether it's your life or your work, the more valuable it becomes, and it'll become valuable in, like, ways that might surprise you. Like, it might pull from context in a way that, may be proactive and that you might not even have thought about. But it needs to have access to those, to that those tools or that context first.Reviews, Agentic Search, and Context GatheringSwyx [00:38:27]: One thing I just wanna talk about the review stuff because I'm still that's a very sensitive thing and you're, you're a founder, you've managed people, you've hired people. As manager myself, I'm very reticent to put out any LLM-generated things especially when it comes to people, ‘cause it feels like you don't care.Swyx [00:38:46]: Presumably at OpenAI, people are more open to being eval rated by GPT. But are there any unofficial rules around this? Like, what's the etiquette?Akshay Nathan [00:38:57]: Oh, I think the etiquette is that, like, I would never write something via, like, well, solely via AI and, like, present it as, like, a review for someone. What I was talking about is more, like, gathering context. That's the place where it's incredibly helpful.Swyx [00:39:08]: So it's just search.Akshay Nathan [00:39:09]: Yeah, exactly.Swyx [00:39:09]: It's agentic search. Yeah.Akshay Nathan [00:39:10]: It's like agentic search, but, that you can tailor and steer much more capably than you could before, ‘cause, like, the thing is it's all there's a flywheel happening, right? Because of Codex, people are able to do, and because of ChatGPT, people are able to do so much more now than ever before. And if you're able to do so much more, it's easy to miss things as well. And so, like, I think we need to use these same tools to keep up with all the impact that people are having and understand, where we can be helpful.Swyx [00:39:39]: I think the thing, like, I run a small company, so easy to search, but at the scale of OpenAI with the amount of messages that you guys put in Slack, do you think that it misses things?Remembering What Humans MissAkshay Nathan [00:39:50]: Probably, but I think that I also miss things.Swyx [00:39:52]: Like, it doesn't matter, right?Vibhu [00:39:53]: I think sometimes it'sSwyx [00:39:53]: Like it's, as it needs to be human-levelAkshay Nathan [00:39:54]: It's all relative, right? Yeah.Vibhu [00:39:56]: Sometimes it's nice when it finds things you wouldn't, right? Like right now, my Codex system prompts, they're set up in such a way that every project I have has a secret- separate, notes MD, and it just writes learnings to there. And then the global one can pull from all these. So sometimes it'll be like: Oh, there's this project you did like four months ago. Here's a note that we had, and it randomly pulls it back into context that I would never do, I haven't thought about.Vibhu [00:40:20]: And I'm like, okay, this is quite superhuman, right? Like, stuff that would. And, it'll save like hours on chunking of stuff or find something that's already been done. I'm like, as much as it might miss stuff, I would too, but it's very useful when it finds stuff. And I have like a very, non-super engineered solution to this. It's just marked down files that get pulled whenever they want.Akshay Nathan [00:40:41]: Yeah. I have a funny anecdote about this. Like, recently gearing up to this launch, the team has been, really cooking on it for a couple months, and over that time, like there's so much conversation and chatter going on in Slack and Docs and elsewhere. And, one of the members of the team set up this, scheduled tasks, like automation to like look at everything that's going on and, like, come up with the best memes and then post it in one of our shared channels. And like, there are two cool things about this. Like, the first is, like, I think the models are, over time, like starting to become like funny.Swyx [00:41:13]: Funny. Nice.Akshay Nathan [00:41:13]: Whereas like, a year ago, like that was not at all the case. The second is, it was what you were saying, like they find things that in surprising ways that you may not have thought of and like create connections that you may not have thought of. And that really helps with like the meme generation because then you can see something that, genuinely surprises you and, is funny in that way. So yeah, that's like not like the most productive, use of this the technology, but it does it does uncover this, like this capability that's emerging, which is just like to find information that you otherwise would not know of.Launch Momentum and the 10 Million User MilestoneSwyx [00:41:43]: Talking about the launch, I think, I have pretty much said this is the most successful launch in a long time. I think even more successful personally than 5.0, and they're announcing ten million users. Does it feel different? You've been through a lot of launches.Akshay Nathan [00:41:58]: I think it feels like a culmination. Well, I think two things. One, it feels like a culmination, like I was mentioning earlier, like this like vision mission that we've been on for a long time. Like I said, we saw the magic of Codex internally, and then we're like extremely excited to bring this to many more people and to see it working, to like see us reach, the distribution goal, numbers that you mentioned, like I think that's like huge and super exciting. The flip side of that is like, there's so much more to do too. Like, that's also really exciting. Like, ChatGPT as a whole, like the this product that, everyone almost equates to AI and like loves, has hundreds of millions of users. And so like ten million is really cool, but like we need to get this to everyone. Like, we need everyone to feel this magic. And so that's the next step from here. But yeah, I think extremely pumped about how it's going so far and the opportunities.Swyx [00:42:46]: Awesome. I did want to also Because I've, I've, I've been tracking the number closely, it transitioned at some point from just Codex users to Codex plus ChatGPT Work, because they're same harness. The whole point is that you don't, you can't, count them separately. Do you have roughly a billion, ChatGPT users? Why did it just jump to one billion right away? Like, isn't that the default on ChatGPT or no?Codex, ChatGPT Work, and the Developer BrandAkshay Nathan [00:43:11]: We don't default you into ChatGPT Work if you're on ChatGPTSwyx [00:43:14]: If you're free. YeahAkshay Nathan [00:43:15]: It's also only available to paid users right now. And I think there's like a process of, educating users of what is the value of this product, having them try it, learning from their feedback, and making it better over time. But the goal is to, get as many of the people who love ChatGPT today to like feel the power of ChatGPT Work. But I think it'll be a journey.Swyx [00:43:36]: Yeah. And Codex will still be alive as a brand for the foreseeable future. And we'll just toggle between them as needed for UI stuff.Akshay Nathan [00:43:44]: Yeah, I think it's even stronger point than that. Like, I think we fully intend to like, treat developer. Like, developers have been, a core market for us for so long, and like there's, there's so much more that we can do to make Codex great specifically for, software development, and we'll continue to do that. This doesn't take away from that at all. If anything, it should increase the utility of something like Codex, because now you can move seamlessly between writing a diff to creating an artifact or, doing a search over your factor.Swyx [00:44:11]: I do wonder how much this terminology leaks to the non-technical user. Like, do they have to learn to say artifact if I want artifact? Or.Akshay Nathan [00:44:20]: It's funny, like we call it artifacts internally ‘cause that's what the teams call it.Swyx [00:44:23]: It's nice. Yeah.Akshay Nathan [00:44:23]: But like externally, like no one says that, no one calls it an artifact. But I think that people like often, like describe things, whatever they're used to, right? So if, ChatGPT Work is good at creating slides, they'll say ChatGPT Work is good at creating slides, and that's what we want.OpenClaw, Personal OS, and Persistent ComputersSwyx [00:44:38]: One big Another, it's July of twenty-six. One big thing that also happens in, for OpenAI was OpenClaw, and that's I think a lot of people's first time really maxing a agent for personal stuff, but also crossing over to work in essence same way. As far as I understand, OpenClaw is still independent, but did you go through your own OpenClaw moments? Were there any lessons you took from OpenClaw to Codex or back? Whatever.Akshay Nathan [00:45:06]: I think there's a lot of inspiration. I did go through my own OpenClaw moment. I,Swyx [00:45:10]: Yeah, tell the storyAkshay Nathan [00:45:10]: Me and my wife like set up an OpenClaw to like try to manage everything in our house. Not that there's like a ton, but it was like quite useful. We gave it a calendar. It started, creating events for us and stuff. At some point, the laptop that we were running on, it died and never got a chance to pick it back up. But there was a lot of inspiration there, like, in ChatGPT Work, in web and mobile, like you get access to this like persistent computer environment where, you can store files, and those files stay around between sessions. And the idea is to be able to enable use cases like this. one of the members of our team uses ChatGPT Work for what they used OpenClaw from before, and then feel like it has like completely transitioned, which is like, workout planning and like meal tracking. which again, it's like a work-related thing, right? It's like not work necessarily, but it's like in personal productivity space. But it has all the same primitives. So it has scheduled tasks. It has the ability to store files on a file system. It has the ability to like reference those things over time. And so you start to see the same types of use cases emerge, which has been really cool.Swyx [00:46:14]: Is there a point that ChatGPT Work completely replaces OpenClaw? they're independent, so.Akshay Nathan [00:46:20]: Yeah, I'm, I'm not close to it, so I can't speak to the OpenClaw roadmap, but I don't think so. I think that there's gonna be, there's always a need for like this like incredible, like open source technology that team has built. And I think that we can draw inspiration, in the product and, ChatGPT, I think many more people have like heard about and used ChatGPT than have used OpenClaw. And if we can take the magic from OpenClaw and bring it to them, I think that'll be a success. I think that like one thing on the ChatGPT Work side that we feel strongly about is that like the core experience is that you come to this product and you have a conversation, start a session, whatever you wanna call it, with this agent. And the magic of the product is that you can do anything in that moment. And we would like to create a product where you don't have to click a button or to go to a different place, whatever, and you can get whatever functionality exists in, your finances app or where or any other product like in this one place. And so that's the goal. It's like it we want an extensible system with plugins where you can connect to the tools that you need in order to be able to accomplish like a financial task, where you can, if you're doing like science work, like we have an ability to like extend the system in such that you can like write the tech and it performs well. There'll always be like products that we support that are best in class at those things, but we want as much of the magic as possible in that core experience.Swyx [00:47:45]: Yeah. Do you think that you can do everything you used to do with Wealthfront in ChatGPT Finance?Finance, Data Access, and Centralized ContextAkshay Nathan [00:47:50]: I tried it. like ChatGPT doesn't yet custody, cash and assets for me. So that part, no, not yet. But I, there was like a whole component of like retirement planning and, like financial planning and budgeting and stuff that, we were looking into when I was there. And like with the finances plugin, like that's all possible with ChatGPT today. So, I feel
This week, we discuss Bun's move to Rust, OpenAI hunting for revenue, and Stripe's bid for PayPal. Plus, Coté's Dad wisdom. Watch the YouTube Live Recording of Episode 582 Runner-up Titles Turn off my dog Dad Wisdom Amazon subscription and let it fly Dad box Talking about modernization Skin in the game The ivory tower of “Things Actually Work in the Real World” Business fortnight No SegwaysBanking works everywhere else in the world, not America Rundown Rewriting Bun in Rust OpenAI OpenAI's No. 2 Executive to Step Down in Latest Leadership Shake-up OpenAI's Chief Futurist Is Leaving the Company AI Devices Are Coming. Will Your Favorite Apps Be Along for the Ride? OpenAI Executive Kevin Weil Is Leaving the Company OpenAI power consolidates under co-founder Greg Brockman ahead of prospective IPO Barret Zoph is out at OpenAI again after just five months Apple sues OpenAI, accuses ex-employees of stealing trade secrets OpenAI Appears to Be Missing Its Sales Goals by a Vast Margin Stripe, Advent make $53 billion takeover offer for PayPal, sending stock soaring Relevant to your Interests SpaceX, AI Bubble Fears, and The Age of the Trillion-Dollar, Zero-Profit Company Microsoft Frontier Company: AI engineering that amplifies and protects your intelligence Building a CLI for all of Cloudflare California founder fired for ignoring his company's own return-to-office mandate Samsung chip division's single-year profits beat its past 40 years of profits, combined Did a lottery winner become "the company of the future"? Power company hikes data center bills by 30%, cuts residential electricity costs by 1.3% Where AI Works: Slack is the Conversational Interface for Headless AI Salesforce MCP Servers: AI, Data & Analytics for Tableau & Data 360 in Slack IBM and Red Hat launch Lightwell to defend open-source code from AI attacks Infoblox acquires Kentik, adding network observability to its DNS and DDI platform Apple 'Hide My Email' Vulnerability Reveals Peoples' Real Email Addresses House passes bill to make daylight saving time permanent United Airlines' new upsell: Keeping other travelers out of the middle seat Linux creator Linus Torvalds puts foot down on anti-AI comments 1Password now lets Claude sign in to websites without seeing your passwords 1Password and Anthropic Bring Secure Credential Access to Claude forum, an accountable orchestrator for AI agents Google continues its renaming streak by turning NotebookLM to Gemini Notebook 'Pretty revolting': LG's TV are getting huge backlash from users due to installing software What is a micro-retirement? Inside the latest Gen Z trend You paid me, a long-time Linux user, to use Windows 11 exclusively for a month China delivers a one-two punch to America's AI dominance Can Lenovo's World Cup AI Moment Cement Its Enterprise PC Dominance? AI Coding Tools Deliver Speed but Not Business Value Unless Wrapped in Process: China's Moonshot in Talks on Pre-IPO Funds at $50 Billion Value OpenAI and Hugging Face partner to address security incident during model evaluation AI won't fix your broken culture, but you can fix your broken culture Oracle Is Now Down 28% in a Month. Will the 52-Week Low of $132 Hold or Fold? IBM stock craters 23% after issuing second-quarter earnings warning Sheetz is quitting VMware, migrating 11,000 virtual machines Meta in Talks to Lease Computing Power to Anthropic in Potential $10 Billion Deal AWS cloud lead Dave Brown heads to Meta - report - DCD Amazon senior cloud executive departs after 18 years Meta Caps Internal AI Token Spending After Costs Approach Billions in 2026 Meta Compute Launch Sends AI Compute Stocks Tumbling Globally Agentic ransomware for automated database extortion Dark-Moon: Autonomous AI pentesting engine Sponsors Signadot: making sure AI-written code actually works. Nonsense Messi Beats Ronaldo in 2026 World Cup Password Breach Data Rankings United Airlines' new upsell: Keeping other travelers out of the middle seat Yes, you can now order DoorDash from the command line House passes bill to make daylight saving time permanent Listener Feedback Biogen hiring Associate Director, Enterprise Architecture in Triangle, NC Tim releases a font for developers whose close-up vision isn't what it used to be The Java Story | Official Trailer | Full Film Coming July 17th Conferences DevOpsDays Graz, Sept 4-5, 2026 Cloud Foundry Summit, Sept. 21st to 22nd, Heidelberg, Coté speaking. DevOpsDays Rockies, Sept. 22 – 23, 2026, Discount Code: 26DODSWEDEFTALK WeAreDevelopers NA, Sept 23-25, 2026, Discount Code: DEVPOD26 25 Free Tickets DevOpsDays Dallas, Sept 28-29, 2026 DevOpsDays Vilnius, Sep 30 - Oct 1, 2006 DevOpsDays Istanbul, Oct 24th, 2026, Coté keynoting. VMware User Group, Orlando, Oct 20-22, 2026 Cloud Native Denmark, Nov 19th, 2026, Copenhagen, Coté keynoting. SDT News & Community Join our Slack community Email the show: questions@softwaredefinedtalk.com Free stickers: Email your address to stickers@softwaredefinedtalk.com Follow us on social media: Twitter, Threads, Mastodon, LinkedIn, BlueSky Watch us on: Twitch, YouTube, Instagram, TikTok Book offer: Use code SDT for $20 off "Digital WTF" by Coté Sponsor the show Sponsor more podcasts with Failover Media Recommendations Brandon: entogo — Type a keyword. Go anywhere. 28 Years Later, The Bone Temple Matt: Vacation, not micro-retirement Coté: Wolf Hall books.
Leaving Port 22 open to the internet, FreeBSD Foundationals, GhostBSD Finance Report, Making your own read-only device with NetBSD, and more... NOTES This episode of BSDNow is brought to you by Tarsnap and the BSDNow Patreon Headlines I Left Port 22 Open on the Internet for 54 Days. Here's Who Showed Up. FreeBSD Foundationals: The Boot Process - From the Loader to Boot Environments News Roundup FreeBSD 15 on a Laptop GhostBSD - March 2026 Finance Report Make your own Read-Only Device with NetBSD NFS-problems I didn't expect at all Tarsnap This weeks episode of BSDNow was sponsored by our friends at Tarsnap, the only secure online backup you can trust your data to. Even paranoids need backups. Feedback/Questions Send questions, comments, show ideas/topics, or stories you want mentioned on the show to feedback@bsdnow.tv Join us and other BSD Fans in our BSD Now Telegram channel
In this episode, Dave and Jamison answer these questions: I stayed too long at my first company and now my career is ruined? Hi both, long time listener. (Since I graduated university 4 years ago actually and felt I could never find a job but listening to you guys made me feel not like such an outcast in our field.) I graduated in 2022. Didn't really apply around much and almost immediately accepted the full time offer from the company I did my internship with during my degree. And fast forward to middle of 2026… I am still here. I've long stopped growing as an engineer and have actually been stuck on a simple CLI tool for almost 4 years now with little to no cloud/backend/etcetera experience to show in my 5 year career. I don't feel I am desirable in the job market at all due to this and waited way too long to look for a job because I was comfortable and happy. I feel paralyzed (no need to put on your space therapist hats BUT some career advice would be wonderful!) (Also did I mention I am in Canada and want to move to Europe and live the rest of my life there but visa sponsorship with my level of experience is impossible? Maybe that'll be for my next question…) Thanks for all you do, cheers! I have a non-CS engineering degree, and am looking to transition into Software Engineering. I have the option of pursuing a degree with Tuition Reimbursement from my job, but am in a bit of a strange situation deciding which degree. Basically, I've figured that with my current Engineering background I could pursue either a second Bachelor's degree through Universities offering a “Post-Baccalaureate” program in Computer Science, or a Master's degree, with maybe only a slightly longer program to get the Master's as opposed to the Bachelor's. Otherwise the total length/cost of the programs would be roughly the same. I have heard though that it might not be a great idea to go into a Master's program unless you know what specifically you want to study. I am curious what you think would be more “hire-able” between the two options for someone pursuing the career change into Software Engineering?