POPULARITY
Categories
When The Jetsons imagined life in 2062, Rosie the robot maid seemed like pure science fiction. Today, chatbots can answer questions, analyze data and even crack jokes. But they still can't load the dishwasher. In this episode of Building 32, MIT CSAIL associate professor Vincent Sitzmann explains why building AI that understands and interacts with the physical world is a fundamentally different challenge. Sitzmann explains why today's AI still struggles with seemingly simple physical tasks and how researchers are teaching machines to see, predict, and interact with the world around them. In this episode: 00:00 - Intro 03:50 - From electrical engineering to computer science 05:05 - What is computer vision? 08:32 - Limitations of computer vision 09:41 - How do we fix the limitations of computer vision? 12:41 - Why can't LLMs solve computer vision? 18:31 - How robots can learn like babies 20:57 - The future of robots in homes 23:59 - Continual learning and exploration 27:25 - Outro Episode Resources:
Series: PraiseService: You Are From GodType: You Are From GodSpeaker: Scott Taylor & Tyler HallJoin us as we discover the hospitality of God and why He deserves your offering of praise. What implications does our submission to God have for how we interact with one another in the body of Christ and outside of our spiritual family? You Are From God is a podcast dedicated to the teaching of the Bible and the Christian faith. This YouTube series follows a Bible reading plan for 2026 curated by the West Mason Church of Christ. Our mission is to help Christians or anyone who wants to learn about about the love of God and root their identity…
Insider Program: https://www.jochumstrength.comJochum Strength Gear: https://www.bonfire.com/store/jochum-strength/1:1 Consults: https://calendly.com/jochumstrength/consultInstagram: https://www.instagram.com/austinjochum/Rival Nutrition: Use code JST20 for 20% off at https://rstr.co/rivalnutrition/jstIvan Escott joins the podcast to talk Olympic weightlifting, Garage Strength, skill acquisition, athleticism, and the long-term game of actually loving training.Ivan breaks down his journey from Division I swimming to building a garage gym during COVID, learning Olympic lifts from Garage Strength YouTube videos, moving close to the gym, investing in coaching, and eventually becoming part of the Garage Strength team.We get into the quest to clean and jerk 405, why Olympic lifts can transfer so well to sprint swimming, how to use Olympic lifts for athletes versus Olympic lifters, VBT for pulls and strength work, and why variation keeps training fun and productive.Ivan also talks about late-start Olympic lifters, building realistic goals, finding limiting factors, choosing the right variations, and what it is actually like training inside the Garage Strength culture.00:00 Welcome to the podcast02:00 Quest to clean and jerk 40503:15 From Division I swimming to Olympic lifting09:00 Going all in at Garage Strength13:30 Olympic lifting for sprint swimming17:45 Olympic lifts for athletes vs Olympic lifters22:30 VBT, pulls, and minimum velocity28:00 Building technique through variation34:00 Athlete days, jumps, and athleticism40:30 Coaching late-start Olympic lifters55:30 Garage Strength cultureExtras:Keep the training fun, chase new variations, and stay a student of the game.
Watch the full episode on YouTube:We first covered Baseten last year when DeepSeek mania was at peak hype. Now they have raised a monster $13B round and become one of the new cohort of AI Infra decacorns that are (with Nvidia, Intel, and the semis complex) chief beneficiaries of the Inference Inflection. We return to Baseten at the peak of the 2026 edition of Open Weights debate. Ali has published a viral breakdown of Kimi K3:And since you last saw him, Philip has spoken at AI Engineer and written the definitive book on Inference Engineering spotted all over SF:Three years ago, inference engineering barely existed as a category.Today, it is one of the most critical disciplines in AI. Inference engineering inherently tackles a different question than standard model training: “How do you turn those weights from training into a product that is fast, reliable, and affordable at scale?” Focusing on these creates an entirely new optimization problem.In one recent GLM-5.2 experiment, quantizing more of the model actually preserved its benchmark quality while increasing throughput by 20%, because the errors introduced in different layers could cancel each other out.Inference is no longer just the final step after training. It is becoming its own engineering discipline, with its own research problems, infrastructure, and increasingly specialized roles.In this episode, Baseten's Philip Kiely and Ali Taha join swyx and Vibhu to explain what actually happens after a new open model is released and what it takes to turn “we generated a token” into a fast, reliable, production-ready API.We go deep on cache-aware routing, disaggregated prefill and decode, quantization, speculative decoding, KV-cache movement, model parallelism, GPU kernels, and the race to make frontier models up to 10× faster. Philip and Ali explain why inference optimizations can still produce gains of 20%, 100%, or even 200%; how quantization errors can cancel one another out; why identical weights can behave differently across clusters; and how Baseten grafted a Kimi vision encoder onto GLM-5.2 without changing the underlying language model.The conversation then expands beyond LLMs into NVIDIA Dynamo, mega kernels, Rubin, AI-specific chips, local inference, video generation, diffusion versus autoregressive models, and the enormous compute barrier to generating coherent long-form video. Finally, we explore the convergence of training and inference, continual learning through persistent KV cache, and the emerging loop where models help optimize the infrastructure that runs them.We discuss:* What happens when a 200,000-token request enters an inference system* Cache-aware routing and reusing previously computed KV cache* Why prefill and decode are increasingly handled by different GPUs* When dedicated deployments become cheaper and more reliable than shared APIs* How speculative decoding uses a smaller model to accelerate a larger one* Tool calling, structured outputs, and what LLMs actually do* What it takes to support a new open model on day zero* Grafting Kimi's vision encoder onto GLM-5.2* Retrofitting inefficient model layers with components from other architectures* Why models sometimes collapse into repeating the same token* How hardware, kernels, and race conditions create nondeterministic failures* Preserving model fidelity while making inference faster* How quantization errors can cancel each other out* Why inference optimizations still deliver gains of 20%, 100%, and 200%* How optimized serving can make a model up to 10× faster* NVIDIA Dynamo, KV-aware routing, and distributed model serving* Speculative decoding the speculative decoder* Why local AI is about making models less dumb while data-center AI is about making them less slow* Tensor, expert, and pipeline parallelism across GPUs* Hardware-aware model design, auto-tuning, and the case against mega kernels* Rubin and why inference is becoming a systems problem* Whether modern GPUs are evolving into programmable AI ASICs* Why enormous models like Kimi K3 require GB300-class hardware* Why open-source video generation still trails Veo, Kling, and other closed models* The quadratic attention bottleneck behind long-form AI video* Autoregressive video, real-time generation, and compounding quality drift* Why future video systems may combine autoregressive and diffusion architectures* Training for inference and inference for training* Continuous post-training, deployment, evaluation, and improvement loops* How GLM-5.2 helped optimize the kernels serving GLM-5.2 itself* Why faster networking could unlock dramatically faster decoding* Continual learning, KV-cache compaction, and persistent model memoryShow Notes* How to build a day-0 API for Kimi K3* 22580: From GPT2 to Kimi3, ExplainedPhilip Kiely* LinkedIn: https://www.linkedin.com/in/philipkiely* X: https://x.com/philipkiely* Inference Engineering: https://www.baseten.co/inference-engineering/Ali Taha* LinkedIn: https://www.linkedin.com/in/aliestaha/* X: https://x.com/waterloointernTimestamps00:00:00 Introduction and the 200K-Token Prompt00:03:18 Dedicated Deployments, Speculative Decoding, and Tool Calling00:11:26 Launching Production-Ready Open Models00:19:06 Model Retrofits, Failure Modes, and Nondeterminism00:28:22 Quantization and Canceling Errors00:32:15 The Race to 10× Faster Inference00:40:48 Dynamo, Speculation, and Local vs. Data-Center AI00:50:18 Model Parallelism, Auto-Tuning, and Mega Kernels01:00:55 Rubin, GPUs vs. ASICs, and Custom AI Chips01:10:03 Giant Models and the Limits of GPU Memory01:12:42 AI Video, Quadratic Attention, and Autoregressive Generation01:21:47 Audio, Images, and Diffusion Models01:27:32 Training, Self-Optimizing Models, and Continual Learning01:40:06 Closing ThoughtsTranscriptIntroduction: Baseten, Waterloo Intern, and Inference EngineeringSwyx [00:00:00]: Okay, we're here in the studio with Philip, old friend from Inference Engineering, the book, as well as Baseten and everything that you've done, you and I have done before, as well as Ali. Welcome.Ali [00:00:15]: Pleasure to meet you.Swyx [00:00:15]: Waterloo intern.Ali [00:00:16]: Waterloo intern, always.Swyx [00:00:17]: When did you get “Waterloo intern” as a handle?Ali [00:00:19]: As a handle? Oh.Ali [00:00:20]: I think the rebranding happened mid-March. When I saw it was open, I was like, “I have to take it. Up for grabs.”Philip [00:00:26]: The problem is that Ali is really good at his job and is not gonna be an intern much longer.Philip [00:00:30]: So we have to figure out who's gonna get the handle.Ali [00:00:33]: Well, I'll pass the torch over to the next intern.Swyx [00:00:34]: Oh, okay. It can be, like, you just pass it to another Waterloo grad.Ali [00:00:37]: To another Waterloo intern. No, bruh.Philip [00:00:39]: Yeah.Ali [00:00:39]: Intern.Swyx [00:00:40]: Intern, yeah.Ali [00:00:40]: And no.Philip [00:00:41]: You gotta get an intern from Waterloo.Ali [00:00:42]: Yeah, I've gotta get an intern from Waterloo.Swyx [00:00:44]: Right.Ali [00:00:44]: But they have to follow the path.Swyx [00:00:45]: Oh, it could, but it could come from Baseten, so it's like whoever Baseten gets from Waterloo.Ali [00:00:48]: Right.Swyx [00:00:49]: Has the title of Waterloo.Ali [00:00:50]: It stays in the ecosystem.Philip [00:00:51]: Exactly.Ali [00:00:52]: Halfway through the internship, you either get it or you're out.Philip [00:00:55]: You should also do, like, a big graduation ceremony where you change the handle.Ali [00:00:59]: Just say it.Philip [00:00:59]: For everybody.Swyx [00:01:00]: You guys are good at ceremonies, clearly. We had a nice launch of the book, very successful. But before we get into all that, I wanna start off with a fun question for you. Okay, you're an expert inference engineer. What happens when I send a long query, say two hundred thousand tokens into Baseten's inference? What's the process of query through GPU model routing, balancing, all that? What is all the stuff that we don't think about?Long Context Requests, KV Cache, and Cache-Aware RoutingPhilip [00:01:26]: With a long query specifically, the first thing that I'm gonna ask is, “Have you sent me this query before, or at least part of it?” and I really hope you have, because it's gonna be a lot easier for me and a lot cheaper for you. So the first thing that we're gonna look at is some cache-aware routing, where we're going to see, we probably have a number of instances, a number of replicas up serving whatever model you're hitting. We want to send this one to something with, number one, available prefill workers, and number two, ideally some cached input already there so that we can skip prefill on at least part of these two hundred thousand tokens. If you're doing two hundred thousand tokens, it's probably coding or a multi-turn agent or something where you would expect to have that cached. If you don't, we're gonna have to send it to a prefill worker. We've at least on certain models disaggregated prefill and decode, so you're going to have one set of GPUs that's solely going to process the input, create the KV cache, and get you your first token, and then that's going to be passed over to a separate set of GPUs, which is going to run decode. We're going to iteratively make those tokens. We're probably going to have some speculator model in front of that. I'm going to assume that you're doing coding, and because of that, our speculator model, which assumes you're doing coding, is gonna have a high draft token acceptance rate. If I'm wrong and you're asking me to summarize every Harry Potter book, it's gonna be slower. And then we stream that output to you and account for it, charge you, a couple of pennies and say, “Hey, would you like to send another one?”Swyx [00:03:04]: Except Baseten doesn't charge by pennies.Philip [00:03:07]: Well, yeah, we charge. I'm assuming that we're talking about the public model APIs. If you are setting up a dedicated deployment, then yeah, it's not pennies.Public APIs vs. Dedicated DeploymentsSwyx [00:03:18]: Yeah, one of the key differentiators when I was talking with Baseten initially was that people who want very high volume just need to rent by the box, ‘cause then it's up to you to figure out how to saturate the box.Ali [00:03:31]: And more often than not, it's, like, way cheaper if you're pushing, like, millions of tokens per hour, if you just pay per hour instead of pay per token.Philip [00:03:37]: Yeah, they do. I think that we've increasingly seen a lot of demand for the pay per token APIs, just because everyone wants to try open models, and then once they find a use case that's really sticky, then they move over to dedicated.Swyx [00:03:51]: Is there a best practice on when it's time to swap over?Philip [00:03:54]: Couple reasons. Yeah, reliability, that's a big one, right?Ali [00:03:57]: Like, if they have a very specific use case, they want you to train something specifically for them, like they want their own spec dec, for instance, for their own traffic.Swyx [00:04:04]: Spec dec is speculative decoding.Speculative Decoding and Custom SpeculatorsAli [00:04:05]: Speculative decoding, yeah.Swyx [00:04:07]: You have to explain.Ali [00:04:07]: Sorry. Like, speculative decoding is like, if you have a huge model, right? And so the model is going to be generating one token at a time every single turn, every single forward pass. So we attach, like, this little, like, parasite, like this layer that goes on top of the model, and this model just has to predict. It does three very fast autoregressive forward passes, and it will predict, like, three certain tokens, and then you do one forward stage over the entire original model in order to see if those predictions were correct or not, and then you accept them or you reject them. Now, this draft model is traffic specific, so if you, like, Philip said, if you're summarizing Harry Potter books, I can train exclusively that draft model on Harry Potter books, and I can guarantee you that I'm gonna accept the three tokens every single time. And so with that case, I increase your decode speed. I wouldn't be able to provide this to you if you're a shared endpointSwyx [00:04:53]: YeahAli [00:04:53]: ‘cause I have no idea if you're doing Harry Potter, if you're doing coding, if you're doing English. We don't know. Also, there was a thing in the book that mentioned that if they really cared about a specific threshold, chapter four, I think. Do you remember that?Philip [00:05:06]: Yeah. The things that you can do is you can set a specific, like, batch sizing, a specific, like, parallelism strategy if you're trying to optimize for, like, throughput versus latency. You can. Maybe a NVFP4 quant doesn't pass your benchmarks and you wanna run a model at higher precision, you could do that. There's just a bunch of reasons why you might wanna have your own endpoint and the biggest one, of course, just being, like, you don't have to deal with someone else doing a hundred million tokens of benchmarking traffic at the endpoint when you happen to be trying to serve your users.Swyx [00:05:40]: Yeah. I think one thing that is. That is a classic journey. Like, it's people is asking the, what happens when you type Google into the browser. Tool calling, is that just, you're generating JSON or is there more complication beyond that?Tool Calling, JSON, and Structured OutputsAli [00:05:58]: Certain customers that we have, they have their own post-trained models, and so they demand a tool calling that's not just, like parse a file or go find the weather. It's something that's very specific and you have to do post-training on this. And if the post-training on the model is not good or if the quantization after the post-training to get the inference to be fast, the model will struggle reading the JSON file and reading the tool calling. But it doesn't require its own like sandbox. It's not like it's going to use that tool calling to like escape a sandbox or like it doesn't have to be contained. It can just be a normal dedicated deployment. The challenge with tool calling more and more seems to be that the companies want certain tool calling which is a very sensitive thing to train. And because you're dealing with all of the JSON outputs, if it doesn't like close the end of the request in a very certain manner, you end up with a model that did the tool calling and like the thinking and so as a result of that, it didn't see the result and just hallucinated the result as it decoded. That seems to be the most challenging thing with tool calling, not really the sandboxes model.Philip [00:06:56]: Yeah, that's a challenge on the training side and then on the inference side, there's work that you can do to scope the possible output. So we published this at this point close to two years ago, the solution to this problem which is you make a state machine and you use that to constrain the output to a specific format. So this is the structured output problem. If you remember backSwyx [00:07:27]: Yeah, the specific grammar is,Philip [00:07:29]: Yeah, exactlySwyx [00:07:30]: GML had this thing.Philip [00:07:31]: Yeah. So it's like the old-school “make sure this is only JSON”, return only JSON orSwyx [00:07:38]: YeahPhilip [00:07:38]: Grandma's gonna die type of prompts.Swyx [00:07:39]: Is it BNF grammar? At some point OpenAI had released a thing that was like, yeah, if you want to constrain your output, write BNF grammar, back as NOR.Philip [00:07:47]: In our inference system, it's just a specified output format. And you get the guarantee that your output's gonna be structured along that format. And so applying that to tool calls can like help cut down on. You can still call the wrong tool or call no tool. It doesn't solve the certainty problem but it at least solves the output structuring problemSwyx [00:08:10]: YeahPhilip [00:08:10]: Within tool calls.Swyx [00:08:12]: And MCP is just another form of tool, right.Philip [00:08:14]: Yeah, exactly.Swyx [00:08:15]: As far as there's no special thing there.Philip [00:08:16]: The thing I'm always like explaining to people is the LLM is not capable of doing anything. It's only capable of making suggestions of what to do and then if those suggestions are formatted in a certain way and applied to a system that knows what to do with them, then an action occurs.Swyx [00:08:32]: Yeah. Part of the fun stuff is, this is solved outside of tool calling too. Like in an agent loop if the output is not correct or you're right, like reasoning, tool calling was done in the reasoning trace, just be like, “Oh, I don't know what to do. Let me just try again.” And it might get there after a few tries. And on your point of training, sometimes this is harder in smaller models, so you don't have the same exact quality outputAli [00:08:56]: Right.Swyx [00:08:57]: When you just swap from a big model, right?Ali [00:08:59]: Yeah. I will say that, before, I think we need to go back to inference engineering proper.Ali [00:09:04]: But, I had expected that something would replace JSON because it's hard to stream JSON ‘cause JSON must be complete and you must have open and close brackets and everything. So it's hard to parse something or validate something while it's being streamed. So people invented all sorts of things that are like, I forget the name of some of these alternatives, but it's something like TOML, something like YAML. But JSON seems to be dominant still.Philip [00:09:30]: The JSON outputs aren't that long, right? Like you could have a long-- ‘cause tool calls also contain the arguments in them and perhaps for a certain tool you might pass like a very long argument. But my impression of the median tool call is that it's a relatively small number of tokens, right? So I would expect that speculators are generally fairly good at something as formatted as JSON. And so you would have like a pretty fast decode step there and that the streaming wouldn't be as valuable, but maybe I'm wrong about that.Ali [00:10:02]: I think you're also bounded by the software or that the model is gonna integrate with if the software is built with JSON for the tool calls or if the company that you'- if your customer says that this is how our software works and our tools are interfaced with JSON, you can ask them to like, change their software and say like, “Yeah, this is gonna be better for the model.” but like with the right training shouldn't be that much of a difference. Also more profitable if it outputs more tokens probably.Swyx [00:10:25]: Depends on your business model.Swyx [00:10:27]: It really depends. But I will say that, as a writer with like experience a lot with generated output, I do try to move from text to JSON text which is very long JSON, right? Like there's paragraphs in every field because I'm trying to structure it, right?Philip [00:10:44]: Right.Swyx [00:10:44]: I want you to first make factual statements, then make opinions then make bullet point summaries, have dates, have entity references have your sources for references, all these things. Anyway, so these are things that like I think people who really experiment with structural output have to really care about. But, let's, let's recurse up the stack a little bit. Before we started recording, you mentioned something really cool, which is that there's a lot of engineering that-- inference engineering that goes on when a new model provider releases a new model, right? So let's call it GLM-5.2, Kimi K3. I had previously assumed, especially if it's like, well, GLM 5 to 5.1 to GLM-5.2, like that you've supported them before. Is it that much work?What It Takes to Support a New Open ModelAli [00:11:26]: It's a lot of work.Swyx [00:11:28]: Yeah. Okay. So like, a lot of people, all you guys, right whenever a new model launch like, people rush to say like, “Oh, Hugging Face supports this, Fireworks supports this, Spacetime supports this,” and I'm like, “Yeah, of course we support it.” But what goes into that? What goes intoPhilip [00:11:40]: I think it's more than just support it too, right? It benefits the consumer a lot. Like I think it was with Kimi K2.5 or GLM-5.2 the latest, there was an inference war, right? X provider is at 90 tokens a second. The next day we're at 150. The nextSwyx [00:11:55]: I kinda kicked that off with the GLM-5.2.Swyx [00:11:58]: I wrote a Twitter article about. It got like half a million views,Ali [00:12:02]: Based on being numberSwyx [00:12:03]: YeahAli [00:12:04]: Or it's for something else.Swyx [00:12:05]: Yeah. Which,Ali [00:12:06]: Oh my GodSwyx [00:12:07]: Which then got everyone really excited about, hey, how can we, bend tracks a little bit further and,Philip [00:12:14]: There's a difference between support the model, as in I can make a token out of this model, and support a model, as in I have a production-ready API from this model.Philip [00:12:26]: Getting to the point of I can make a token out of this model is not that hard because generally the, open source inference engines, vLLM, SGLang of the world oftentimes even receive weights ahead of time, maintainers do, or the people making the model merge PRs to ensure support. So you generally can, just get it working on the standard open source stack without too much pain in most cases. The challenge is, every inference company is gonna have own proprietary stack. Some open source components, some in-house stuff. And for any arbitrary model, there's going to be some new stuff. Sometimes you get lucky, like K, two five to two six was, like, pretty similar.Quantization, Speculators, and Production ReadinessAli [00:13:16]: Yeah. It was pure continued post-trainingPhilip [00:13:18]: YeahAli [00:13:18]: If I remember correctly.Philip [00:13:19]: Even in those cases, there's still stuff you have to do. You have to redo the quantization work. You're taking the model from. Generally, these models are not released in NVFP4, and we want them to be in NVFP4 for maximum Blackwell compatibility. So we have to perform that quantization, and, calibrate the quantization to make sure that we're not causing any regression in the model's intelligence. And then we also have to train the speculator, as we've talked about. Generally, we have. We have ZDR, zero data retention on our model APIs, so we don't know exactly the traffic that people are sending us, but we know what's popular. We know that coding use cases are popular. We know that agents, agentic use cases are popular. So we can get public data sets that are representative of that traffic and train general speculators. Now, with speculators today, you need to train the speculator using the base model itself because you're getting hidden states out of the model from running inference on these specific prompts, and that is the training data you use to create the speculator. So there's that process which you need the real model weights for. And then there's of course just the process of, standing up all the infrastructure behind it, loading all this stuff, testing it. And then when there's a new model with a newer architecture, I think that, like, the DeepSeek models tend to be the most challenging as they have, like, the most novel architectural stuff going on, model after model. But every new model has something. Kimi K2 had. Oh, sorry, GLM-5.2 hadAli [00:14:53]: Sparse attention.Philip [00:14:54]: Yeah,Ali [00:14:54]: YeahPhilip [00:14:54]: the DSA.Ali [00:14:55]: Right. Which is brought from DeepSeek.Philip [00:14:57]: Yeah. AndAli [00:14:59]: So you can copy-paste then?Philip [00:15:01]: It kindAli [00:15:01]: I don't know how this works.Philip [00:15:02]: So, like we had to, like, build support for that into our runtime. And you're right, like it is really interesting the way that all of these open source labs borrow from each other. For example, like GLM-5.2 doesn't have vision. So something that, Haley, a guy on our team, if we could take a look at this, he, like, grafted the Kimi vision encoder onto GLM-5.2.Retrofitting Vision into GLM-5.2Ali [00:15:27]: We'll be training the projector.Philip [00:15:28]: Exactly. So if you think about, like, the encoder, there's the encoder, which is the part that looks at the image and turns it into latent information, and then there's the projector which likeAli [00:15:38]: You can say latent space. It's okay.Philip [00:15:41]: And then there's the projector that maps it onto, the model itself, and then there's the model weights. You don't wanna mess with the model weights because you run a chance of making the model dumber at something else for the purpose of giving it vision. So instead, Haley started with just a projector, which is only a handful of millions of parameters.Ali [00:16:02]: That would be, yeah.Philip [00:16:02]: Yeah.Ali [00:16:03]: Can you show the training one?Ali [00:16:04]: Like the way it groksPhilip [00:16:05]: YeahAli [00:16:06]: Very interesting.Philip [00:16:06]: And maybeAli [00:16:07]: That right therePhilip [00:16:07]: Maybe Ali, you should take it from here. You've got a betterAli [00:16:10]: Ooh, double the sandPhilip [00:16:11]: Understanding of this than I do.Ali [00:16:11]: Yeah. You can see, like, he. The way he trained this is really cool. At the beginning, he was training it using just like, “Here's a picture of a mountain. Can you describe what's in this mountain?” And that caused it just like the first, learning walls. Like here you can see this all we're trying to teach it is to translate the encoded. Like it's already taken the encoder from Kimi K. It's taken the image. It'Philip [00:16:31]: Yeah. FrozenAli [00:16:31]: FrozenPhilip [00:16:32]: With adapter.Ali [00:16:32]: Exactly.Philip [00:16:33]: Yeah.Ali [00:16:33]: So the brain is frozen and the eyes are frozen. It's just we're tryingPhilip [00:16:37]: AlignAli [00:16:38]: Interconnect between the eye and the brain, right? So the projector. And so you take the tokens and then he's like, “Oh, can you describe what's in this image?” And he's like, “Oh, it's a mountain,” or it's a person or it's a human, whatever the case is. But that didn't cause complete understanding. So he changed it such that every image was associated with a data set of questions. Like, does this image have a white male? Does this image have birds in the top corner? Does this image have a scientist in it? All of that stuff. And it would have to answer questions correctly. And using not just training on describing an image, but being able to answer question, another question, answer over time. Like you can see the grokking, which is like genuinely insane, that retrofitting vision into a large LLM can learn to that extent. And even for images that it doesn't perform well on, for instance, if you ask it a picture of like Stephen Hawking, “Who is this?” Maybe it doesn't get it, but it will say something like, “This is Albert Einstein.” Like it still understandsPhilip [00:17:25]: Close enoughAli [00:17:26]: That this is a scientist who is a man who has, some significant achievements, all that stuff. So that's like really cool.Philip [00:17:32]: Yeah. So, we've covered Hao Tian before, who the author of the LLaVA paper that did this, a while ago. And I think that's very foundational work for anyone who hasn't done vision work before.Ali [00:17:41]: Same with the CLIP and MetaCLIP, where you go from just captioning to building out questionsPhilip [00:17:47]: RightAli [00:17:47]: Off the image and how much better you can get performance.Philip [00:17:50]: Right. Right. Right. Yeah. But what's, what's so exciting about this is if you look at a model like this. Now, this is a little bit more of a research project. It's not. It got to 56% on MMLU Pro, I think. So not quite frontier. But if you're running this model, you haven't suffered any loss on your GLM-5.2 quality. If you don't have an image, it'll just behave exactly the way it used to. And ultimatelyAli [00:18:14]: Which in the inference code you literally do not include the other part, right?Philip [00:18:18]: Yeah. You would just skip the encoder if you don't have an image input.Ali [00:18:22]: Okay.Philip [00:18:22]: Just confirming.Philip [00:18:23]: YeahAli [00:18:23]: Does it affect a lot on the overall inference side? Like you're not adding much, you're adding a very small vision encoder. These are typically likePhilip [00:18:30]: They're super fineAli [00:18:31]: Less than a billion parameters, right?Philip [00:18:32]: Yeah. It's, - There's a little bit less standardization among vision encodersSwyx [00:18:37]: YeahPhilip [00:18:37]: So the support matrix can be a little bit, sparser. But overall, yeah, it's a pretty, it's a pretty minor component of the overall system. And ultimately what you get out of the system is all of a sudden you have Kimi Vision, GLM weights, and DeepSeek attention all in one model.Open Source Model Grafting and Franken-MergesPhilip [00:18:56]: And that's, I think, a lot of the power and beauty of open source, is that you can take all of these different components and combine them together into a system that's better than anyoneSwyx [00:19:05]: YeahPhilip [00:19:05]: Can be individually.Swyx [00:19:06]: People used to say that you would also do Franken-merges where you would take likePhilip [00:19:10]: YeahSwyx [00:19:10]: Layers from each model.Swyx [00:19:11]: Does anyone do that anymore?Ali [00:19:13]: Well, to your point previously when you were mentioning like, the work that goes into supporting a model when it first comes out, like GLM-5.2 or MiniMax M3 or whatever the case is. Sometimes you do have to like, you do have to switch out some things. Like, for instance, the MiniMax M3 head uses full attention, and with full attention you end up with this like insane bottleneck in spec dec ‘cause you're doing auto-regressive token generation for three tokens, and you're doing this like N squared over all of the tokens that are in your sequence. Your KV cache is like very large because it's not sparse, it's not top K. So we find it better to like, okay, we're gonna replace this, we're gonna replace this layer with a layer from another model that's using like GQA, for instance. And then just with the right training, you can get it to have the same acceptance rate. So it is very possible to retrofit layers from other models and very much needed. If a layer is like inefficient, the training just becomes the challenge, like how do you ensure that you train it properly? Which again to your earlier point is like the mesh between training and inference. As in like you need very good training in order to do fast inference. That's like, I feel like more and more becoming true.Swyx [00:20:21]: Yeah. Anything else on the support side when you say like get it to fully production ready?Loop Detection, Race Conditions, and Non-DeterminismPhilip [00:20:26]: Yeah. I think that there's also a question of just, we can test a model to a pretty extensive degree, but we're trying to get it out quickly and then you see a bunch of other people test it and you get interesting results. There was an issue with, GLM briefly where we had some like mode collapses where it would just output the same token over and over again for certain prompts on certain temperatures. Like once you expose an endpoint to the real world, there's going to be, so many more varieties of things given to it that you're able to, discover and patch things. So it's not just a, day zero process, it's then like for the first week, for the first month, if a model remains popular, like how do you both fix bugs and then continue to push the envelope on performance?Ali [00:21:21]: What do you mean you don't want your model outputting S?Swyx [00:21:24]: Is there loop detection on that stuff, by the way? It still happens like quite a lot, which is surprising.Ali [00:21:30]: We have like we, in our endpoint, like if a model was to output the same token like four plus times, we just cut the generation. We say like, “Oh, sorry, this-- Like try again,” or like we will reprocess the request. ‘Cause we know then, like if it, like if, yeah, it's four times the same token, it's probably collapsed.Swyx [00:21:45]: Yeah. Is there a way to opt out in case I really want that?Ali [00:21:48]: You want that?Ali [00:21:50]: I think there's a way that we have to handle it. I'm not exactly certain, but I feel like in certain models, like when they output something like you can imagine, like a table for instance, and so they want, they wanna draw like 12 dashes and 12 dashes. Yeah, I think there's a way for that to happen. I think we only do it on certain tokens. Like we exclude certain special characters.Swyx [00:22:07]: Yeah.Ali [00:22:07]: So we only do it on like certain like S is the most common almost. GLM-5.2Swyx [00:22:11]: OhAli [00:22:11]: And I think it was DSV 4 as well. Like you'd just have like looping issues where like you literallySwyx [00:22:17]: ItAli [00:22:17]: Just have like S.Swyx [00:22:18]: Yeah. Is there a special, something special about S? No, just randomlyAli [00:22:21]: It just seems to be the one token involved.Swyx [00:22:23]: Yeah. And it'Philip [00:22:24]: Is thereSwyx [00:22:24]: And it's only temperature 0Ali [00:22:27]: NoSwyx [00:22:27]: Even at other temperaturesAli [00:22:27]: Even at like 0.9 or whatever, it will still, it will still collapse.Swyx [00:22:30]: That's weird, right?Ali [00:22:30]: It's, it is an inference problem to be honest, like a software problem. Like oftentimes, the image you run will-- like NVIDIA will release an image for instance, and if we will upstream the changes from their latest TensorRT-LLM image into our stack, we'll find that it fixes it. Or oftentimes this will only happen in an inference engine that you're using like SGLang. But if you were to switch to vLLM, that isn't the case. So it seems to be like an extremely like deterministic software issue and not really a model issue. It's not like a weights problem. Like I'- we'll say like, “Oh, it's a problem with the quant. We did PTQ wrong,” right? But that isn't, that doesn't make sense because the same weights used with a different inference engine does not repeat the problem. And sometimes it's, the kernels that are being used in the backend have like these very subtle sometimes race conditions, where if you were to use this model hosted on one cluster, you will never get this problem.Swyx [00:23:19]: Oh my God.Ali [00:23:19]: But if you host it on a different cluster, you will. And the reason is the KV cache transfer from a node to node in that one cluster is using a slower interconnect than the node to node in another cluster. So that exposes the race, whereas in another cluster it doesn't. So then you end up just like, okay, this model is not gonna be hosted on this cluster. We're gonna host it on, another cluster because that cluster exposed that problem. But then it ends up with like, okay, is it the software? Is it the model weights or is it the hardware?Swyx [00:23:42]: There is a thing about this with temperature 0 still not being deterministic, right?Ali [00:23:46]: Right.Swyx [00:23:46]: Mostly because of hardware. Even at temperature 0 same model, you won't always get the same output.Swyx [00:23:52]: Even-- But I'm surprised by the race condition one because, I thought PyTorch was a graph that like guarantees that you at least, execute things in the right order.Ali [00:24:02]: Well, yeah, true. Like I'm not, I'm not saying that there is. Like well, you have things like PTL optimizations where like you can start a kernel before the end of the previous kernel, and that's like ‘cause you want to do that because there'sSwyx [00:24:12]: It's like pipeliningAli [00:24:12]: Expense. Exactly.Swyx [00:24:13]: Yeah.Ali [00:24:13]: But it'- But you don't do it cleanly. Like you overlap a little bit of the execution. No, it is very possible that the kernel itself, like that one block that is supposed to be running in this instance of time, that kernel itself has a race condition. For instance, like a missing barrier. Like often if you're designing a kernel and you want it to make it to be very fast, if you don't test it extensively, you'll, you'll have certain threads access data points from registers before they've been written to by other threadsSwyx [00:24:36]: YeahAli [00:24:36]: For example, because like your barrier is wrong or your synchronization was wrong. But yeah, like the testing itself is very difficult in those like, andSwyx [00:24:42]: And there's no like borrow checkerAli [00:24:45]: What does that mean?Swyx [00:24:46]: Like Rust. Like the. If you're trying to have like memory safety It sounds like a comparable problem.Ali [00:24:52]: Well, yes, but you're working in CUDA, right, NVIDIA GPUs. Like- You just need a higher level language like modular Maybe that's what modular is supposed to do. I don't know.Quantization Quality and Vendor FidelityVibhu [00:25:00]: How do you see keeping quality of the model? So you talked about all these steps of, okay, you gotta do quantization, train your own speculative decoderAli [00:25:07]: RightVibhu [00:25:07]: Run on different hardware. Looking at other model providers, okay, you kicked off a inference speed race on the consumer end. What goes into keeping quality the same across them, right? Sure, you can run benchmarksAli [00:25:22]: YeahVibhu [00:25:22]: But, like, how do you determine how much quantization are there standards? What goes intoPhilip [00:25:27]: There's a few things on quality. Most inference optimizations are lossless. KV caching, for example. You are just recomputing or preventing recomputing the same values. Speculation, of course, if a draft token is wrong, it gets rejected. The main lossy optimization is quantization. And that really comes down to, number one, data format, number two, which parts of the model you choose to quantize, which layers, and number three, like doing a lot of calibration on the quantized weights, to ensure that you're preserving all the outliers. There's other tricks that you can do, though. A big one is long context, ‘cause one thing you asked at, right at the beginning is, “Oh, what's gonna happen if I send a 200,000 token request in?” So with a long input sequence, you need to, store a lot more information. You need to process a lot more tokens. And so even if a model has a context of a certain length, you might, as an inference provider, choose to build an API with a shorter context length, and of course a full length one as well. Because if someone doesn't need the full million token context, for example, you can get them better performance. I don't know if that's exactly like quality of the model. The way that I think about quality is to what degree are we faithfully serving the original model? If you think of a golden implementation of a model that performs exactly the way the model is designed to perform, I think of quality as how close are we getting to that, 100% fidelity of the model.Philip [00:27:13]: You can also, of course, think about quality from the training side and how do you push yourself past 100%. But when I think about purely inference optimizations, it's getting faster while staying as close to that 100% fidelity mark as possible. And certainly our standard internally is that, like you should not be able to tell the difference between our API and a, official API. I think Kimi in particular does a good job of vendor benchmarking hereAli [00:27:41]: YesPhilip [00:27:41]: Where they haveAli [00:27:42]: They released an actual vendor benchmark.Philip [00:27:43]: Exactly, yeah.Ali [00:27:44]: ‘Cause they accused, some people, Amazon? There was some provider that was not doing very well on Kimi's benchmark.Philip [00:27:50]: Yeah.Philip [00:27:51]: So, with Reflect we probablyVibhu [00:27:52]: This was a long time ago, right?Philip [00:27:54]: No.Ali [00:27:54]: Yeah, like threeVibhu [00:27:55]: They alsoAli [00:27:55]: Four, five months agoVibhu [00:27:57]: This also happened with, I don't remember which model, but they pulled out quite a few, and then they started a whole chart about this. It might have beenPhilip [00:28:03]: Kimi Vendor Verifier.Ali [00:28:04]: Yeah.Philip [00:28:05]: Yeah.Ali [00:28:05]: Yeah, ‘cause you, ‘cause you'd be pissed, right? Like if you'Philip [00:28:07]: Yeah.Ali [00:28:07]: If like if I'm a consumer and I'm using like Amazon's endpoint for instance, and I've used Kimi and I'm like, “Oh my God, like this is bad,” I'm not gonna say, “Oh, Amazon quantized the model in a bad way.” I'm gonna say, “Oh, Kimi sucks.” Right?Philip [00:28:17]: Yeah.Ali [00:28:17]: So it seems like that makes sense.Philip [00:28:19]: Yeah, they care. They care.Vibhu [00:28:21]: Justifiably.Ali [00:28:21]: Yeah, justifiably.Vibhu [00:28:22]: This is probably a stupid question, but just checking, has anything improved from main quantization?Philip [00:28:28]: Yeah.Vibhu [00:28:28]: Like, is quantization always strictly worse?Ali [00:28:30]: Well technicallyVibhu [00:28:32]: NoAli [00:28:32]: It's a lossy. QuantizationPhilip [00:28:33]: YeahAli [00:28:33]: Is a lossy, it's a lossy implementation.Philip [00:28:36]: Speed improvesVibhu [00:28:36]: Speed improves.Ali [00:28:37]: It the number, likeVibhu [00:28:38]: No, I' always look for inverse scaling laws.Philip [00:28:40]: Yeah.Ali [00:28:40]: Yeah.Vibhu [00:28:40]: This is something I learned from Noam Brown, where like things that normally act in one direction sometimes do.Philip [00:28:45]: Well, technically when you run a benchmark, because these models are deterministic, sometimes your,Ali [00:28:52]: YeahPhilip [00:28:52]: NVFP4 quant is like, two basis points higher than yourAli [00:28:56]: No, it's noise. It's noise.Philip [00:28:57]: Yeah, exactly. I'm like, yeah, it's, it's within. That's why I always say within margin of error.Philip [00:29:01]: And I stopped saying that because everyone assumes that what is, well, within some margin of error, we're barely inside of that to the worst, so we're saying. But yeah, sometimes it's just like, gives you a higher output score. But like Ali said, that's noise. To my knowledge, you're not necessarily making the results better. You're just trying to, again, like keep your fidelity as close to 100% to the original model.Layer Selection, KL Divergence, and Better QuantizationAli [00:29:27]: There is, to your point, research that we did on MP. I don't know if you are able to pullPhilip [00:29:31]: YeahAli [00:29:32]: A tweet we did. One of our research interns, Joshua, I think it's a tweet on how we have 20% better quantized GLM-5.2 than NVIDIA. Essentially what we found throughout like this month research is, okay, quantization is a lossy. It's. You're compressing the data from, occupying 16 bits to occupying, four bits, for instance. And so you're losing some information, and you're trying to minimize that. And so when I say that I'm gonna quantize the model, my job becomes how do I find the layers that I can quantize, and how to find the layers to not. For instance, with image models, I don't quantize modulation layers, and I don't quantize out projections because those two are. Like out projection is what you see as the user. Modulation is what the model sees or understands. Right, exactly. And so to his paper, do you have the. It doesn't have the. Yeah. It's a long paper. I don't know if I can findVibhu [00:30:25]: If there's a part to search or it's probably in the thread.Ali [00:30:28]: It's probably in the thread.Vibhu [00:30:29]: Yeah.Ali [00:30:29]: But the long and the short is it is very possible that quantizing more of the model makes the results. Like if I have a model that I quantize layers one, five, and 10, and another model where I only quantize layers one and It is possible that the model in which I quantized more information is going to perform better because the quantization errors have canceled out. And so what Joshua showed in his mathematical proof where he had like a verifier in, is that you can predict which layers are going to have quantization errors that will cancel out with each other, and you choose to quantize those layers. And so the result of doing this mathematical quantization is you end up with a model that's 20% more quantized than another provider, so you get 20% more throughput of it because there's more layers than running an NVFP4, and your quality is better than that other quant because the layers that you chose to quantize have their errors cancel out, like one layer skewed to the right one layer skewed to the left, one layer skewed to the right. Your final logits distribution is more similar to the original distribution of the model, so you have better fidelity. And so the way we proved this was with KL divergence. So instead of just scoring on the benchmarks, we scored the KL divergence between the logit distribution of the quantized model and the logit distribution of the original full precision model, and we showed that with this technique we get. If your probability distribution on the logits which token it wants to select is more of the same as the original model, you're probably gonna end up staying true to the original model. So yeah, so it seems like previously before this, it seemed like the industry was, well, the more you quantize, the worse it's gonna be, ‘cause the more loss you introduce. That's not exactly, not necessarily true. So yeah, doesn't improve it, but can cancel out.Philip [00:31:57]: I think it might be this, but reminds me a good bit about pruning where you can prune off certain layers.Philip [00:32:03]: But very interesting. Didn't know this was a whole paper you guys put out.Ali [00:32:06]: It's. Fun fact, it was originally 72 pages, this paper, and then we decidedPhilip [00:32:11]: WowAli [00:32:11]: We can't tell. We couldn't release it. So it's now 45.Swyx [00:32:15]: Still 39 pages, so very substantive. We talked about evals and all these things and, like what's possible in terms of speedup? Like it's like probably like the numberInference Speedups and BenchmarkingSwyx [00:32:25]: Thing that people do wanna care about, and it's something that you wrote about in your post. Like official API is 70 tokens per second, and you push it up to 90. Is that like a normal thing?Philip [00:32:36]: So what's cool about working in inference, the reason that I think inference is going to be a useful place to do engineering for a long time, is that if you look at highly optimized domains like, say, finance, if you're in finance, you measure how much better you got in basis points. It's like, “Oh, I got five basis points better, like twentieth of 1% better,” that's huge news because everything is so optimized. When we publish optimizations, it's 20%, it's 100% it's 200%. So there's still probably like a lot further to go, honestly. Like you'll, you'll know that inference is pretty much solved when researchers start publishing about how they got 1% faster at something.Swyx [00:33:19]: Which by the way, because I am from the finance background, in the ‘70s, that was the margin at the time. When you did quantitative finance research, you would findAli [00:33:27]: And like 20%, tens of percent.Swyx [00:33:29]: That's. Yes.Philip [00:33:29]: Yeah.Swyx [00:33:30]: And now it'Philip [00:33:31]: Tiny fractionsSwyx [00:33:32]: For those people interested, look up Andrew Lo's paper. He had a really interesting illustration of quant, stat arb, distribution, narrowing down from like those kinds of 20% differences in the ‘70s, down to nothing today, which is very cool.Philip [00:33:48]: Exactly, and we're at the beginning of the same type of thing. Now benchmarking is hard. I think anyone will tell you that, and benchmarking provider speeds is hard because there's so many variables that go into it. What hardware are you using? How much load do you have on the system? What's the exact nature of the prompts and input and output sequence lengths? All that stuff. But overall, when you start stacking these improvements, you're looking at multiples. You can look at it. The most common form, of course, is TPS, tokens per second, which is bad naming by us in the industry, ‘cause there's two tokens per second. There's tokens per second, the throughput number, and the latency number.Ali [00:34:31]: TTMT, yeah.Philip [00:34:32]: Like total tokens per second out of the, out of the GPU as a throughput number. Most people only care about tokens per second as the latency number, which we should call ITL, intertoken latency, but we don't.Philip [00:34:44]: Anyway, so you can imagine a standard API without many optimizations for a 1 trillion parameter model operating somewhere in the 30 to 50 tokens per second range for reasonable traffic profile. And we generally see the goal of, pushing to 10X that. But, not necessarily day zero, but by stacking enough optimizations, if you have, say like four optimizations, each of which doubles performance. Or sorry, three optimizations, each of which doubles performance, then you stack that up, that's an 8X gain. That's the order of magnitude that we're working with in this space. We're trying to make things substantially faster, not just go from like 70 to 90.Swyx [00:35:38]: Are you saying you've. You have done that?Philip [00:35:40]: So let's say you have as a reasonable baseline, 30 or 40 tokens per second. You can achieve 10X that. So like on GLM-5.2, if you run it unquantized, perhaps on H100s even, and you're just using an off-the-shelf inference engine with no particular optimizations, no speculator, nothing extra around like KV routing, no disaggregation, you're, you're probably, yeah, looking at that like 30 to 40. You think that's like a reasonable baseline?Swyx [00:36:12]: Right. Right.Philip [00:36:12]: To get to something like 10X, there's a lot of trade-offs that you're making. If we're running at more like a 300, 400 tokens per second range, you are using the best hardware possible. You have a optimized speculator. You have done all of your quantization work. You are Seeing a pretty high cache hit rate. You are running with a reasonably small batch size and a parallelism configuration that is tuned for latency versus throughput, but it is possible. So the spreads that you see if you, like, go on artificial analysis or you go on OpenRouter and you look at, the worst provider to the best provider, oftentimes can hit that range. 10X is of course very aggressive. It's oftentimes maybe more of a four to six times improvement. But that's the performance that makes us really excited, is when we can get these huge gains, not just go from 70 to 90 tokens.Stacking Optimizations: NVFP4, Speculation, and DisaggregationAli [00:37:19]: It's also, like, hardware dependent. Like, ifPhilip [00:37:20]: YeahAli [00:37:20]: If you have a thing where you're serving it on just, like, a node of H100s and then you throw, like, you shard the model across, like, four nodes of B200s. Like, you can definitely increase the speed with just throwing more hardware at it. Like, normalizing for the same exact hardware and the same number of GPUs.Philip [00:37:35]: Yeah. Then you're looking at, like, a two to 4X improvementAli [00:37:38]: Right. RightPhilip [00:37:38]: Depending on the inference optimizations. So yeah, it's. Some of it's, what's the call, and some of it's who's the driver.Vibhu [00:37:46]: If you break down the two to 4X, say the example is run GLM-5.2Ali [00:37:51]: YeahVibhu [00:37:51]: On B200sAli [00:37:53]: YeahVibhu [00:37:53]: Single node, right? What's, like, the cost trade-off for effort to get, like, the last bit of juice out versus what should people just think of, right?Ali [00:38:01]: Spectre quantization. Yeah.Vibhu [00:38:03]: Spectre quantization.Ali [00:38:04]: That's, that's, that's like 95%. LikeVibhu [00:38:06]: And how far does that get you? And how easy is that for the average person to do? So say right I wanna throw the weights of GLM-5.2 on a node of B200s, how easy is it to find speculative decoder- decoder model or already quantized model? How much work goes into it?Philip [00:38:23]: If you're doing it up front, it's quite a lot of work. If you're doing it today, there's going to be people who have published things that you can just, you can just grab some NVFP4 weights. You can grab a speculator. Yeah, if we're thinking about, like, what are the 2Xs we're stacking, going from, BF16 to NVFP4 is, it's not quite a 2X, right? It's like. I think it's about, like, 30 to 40%, from 16 to 8, and then another 30 to 40% multiplied from, 8 to 4. So that doesn't quite get you a 2X, but, like, roughly a 2X. Speculator, roughly a 2X. Disagg on top of that if you're able to get enough hardware and put enough traffic through it, another roughly a 2X. And then you add in some, double-digit percent increase from having just a better runtime with, the latest kernels and stuff behind it. And that's how it stacks up.Ali [00:39:21]: YeahPhilip [00:39:21]: So building each of those, like, building the, quantized weights is, for someone who really knows what they're doing, hours to days of work. Building the speculator, again, like, hours to days of work. And the, disagg setup, hours to days. Well okay, but like once you haveAli [00:39:39]: Once set up. Once set up. YeahPhilip [00:39:40]: Yeah, getting disagg working for the first time, I'm saying, of course, is very difficult.Philip [00:39:44]: The marginal implementationAli [00:39:48]: Like, if you're just grabbing, like if you are a person, like just a normal consumer who has access to, like, a node of B200s and you're wondering, “How can I just host it myself?” You don't need to quantize the model yourself. There's always gonna be, like, an open source quantized checkpoint. NVIDIA's gonna push one out if no one else does. You. Usually, the providers will have their own spec dec that they've trained as well. You don't need to train your own spec dec. You can just use that as well.Philip [00:40:09]: Yeah. Like, GLM-5.2 has its own MTP.Ali [00:40:13]: Right. Right.Vibhu [00:40:14]: What's multi token prediction?Philip [00:40:15]: Yes.Ali [00:40:16]: I'm justVibhu [00:40:16]: Can you explain that?Ali [00:40:16]: I'm just an expert.Ali [00:40:18]: I can do it for you in case I get it wrong?Vibhu [00:40:20]: No.Vibhu [00:40:21]: Yeah, you should correct if we're wrong, but their multi-token prediction can be used for self-speculative decoding.Ali [00:40:27]: I'm not sure. I'm not gonna correct that.Vibhu [00:40:28]: Okay. I'm semi-confident in thatAli [00:40:30]: Okay. YeahVibhu [00:40:30]: But someone can check. But it's useful to paint the story of, okay, not just the average person, but say a company wants to switch from serverless inference I wanna throw this up on. I wanna rent some GPUs, throw it up. These are the steps you take to do significantly faster than just put it behind vLLM.Ali [00:40:48]: Right.Vibhu [00:40:49]: I was waiting for a mention of Dynamo.Vibhu [00:40:51]: I feel like, that's supposed to be the baseline that you measure against.Dynamo, KV Routing, and Disaggregation ToolkitsPhilip [00:40:55]: I would think of Dynamo as less of a box system and more of a toolkit for building with. So when we talk about doing aware routing, when we talk about doing KV offloading, when we talk about doing, PD disaggregation, Dynamo fundamentally is. By the way, Dynamo is an open source library from NVIDIA.Ali [00:41:17]: We've done a pod with KylePhilip [00:41:18]: OkayAli [00:41:19]: Kyle Cranin.Philip [00:41:19]: Cool. So then your listeners know then that it supports all the different inference frameworks. And it is multi hardware, which is interesting.Ali [00:41:28]: But it's just a router, it's not like an optimizer layer.Philip [00:41:30]: Yeah. All it does, like, what Dynamo is good at, it is a library for moving information around your cluster, around your hardware. So if you have, KV cache on one place and you need it to be somewhere else, Dynamo coordinates NIXL for you to move that around.Philip [00:41:49]: That doesn't mean that, like, out of the box, you just say, “Pip install Dynamo,” and then you get, like, a massive performance speed up. It's more of a developer toolkit.Ali [00:42:01]: Yeah. I would have said it would. It comes with a set of defaults that you can then swap out.Philip [00:42:06]: It does. If the industry at large, I think, was, like, rolling out all of these deployments, standard, then I think it would be, like, a credible baseline. But, we've got to, we've got to benchmark against, like, what we're seeing in the wild.Speculative Decoding Methods: Medusa, EAGLE, n-Gram, and Spec-SpecVibhu [00:42:23]: I did wanna talk a little bit more about PD disagg, because that is probably, like, number three after quantized and speculative decoding. In your book though, I was just gonna pull out the book.Philip [00:42:31]: Yeah.Vibhu [00:42:32]: Like section 522 on Medusa, 523 on EAGLEPhilip [00:42:35]: YeahVibhu [00:42:36]: 524 on gram.Philip [00:42:37]: It's 55, would be disaggregationAli [00:42:42]: Yeah. Well, no, I just wanted to dwell a little bitPhilip [00:42:44]: YeahAli [00:42:44]: The other. Like, so what do you choose to include? What do you choose to not to include? Because there was all these other techniques.Philip [00:42:51]: Yeah.Ali [00:42:51]: Are these still relevant? Because I think they came out, like, a year and a half ago maybe.Vibhu [00:42:55]: Medusa is quite old.Philip [00:42:56]: Yeah, Medusa's old.Ali [00:42:58]: It was old.Vibhu [00:42:58]: But is it in the book as a good, here'sPhilip [00:43:01]: BaselineVibhu [00:43:01]: Baseline vanilla understand it?Philip [00:43:02]: Like you should know this.Vibhu [00:43:03]: Like I read the paper, I'm like, “ it makes so much sense.”Philip [00:43:05]: Yeah.Philip [00:43:05]: So with the book, I had a couple goals. One was to give people just a working vocabulary for the space as a whole, and the other was to give them some intuition about how each of these techniques works. As I mentioned in my AI Engineer talk, which is the first public addendum to this, the speculation space has moved much faster than everything else. So yeah, even at the time that I wrote the book Medusa, I very much included as a way for people to understand how the space evolved rather than what the most modern technique is. And now of course, there's DFlash, dSpark. There's, there's newer techniques even than EAGLE, although EAGLE is still very commonly used.Ali [00:43:51]: SpecSpecta.Philip [00:43:52]: Yes. Speculative decoding.Vibhu [00:43:54]: What canAli [00:43:56]: Oh, it's a paper by Tri Dao and it's like, it's doing speculative decodingVibhu [00:44:00]: HuhAli [00:44:01]: For the speculative decoder.Philip [00:44:02]: Oh, in spec- oh my God.Ali [00:44:02]: It's literally just an another. It's like, yeah, that's the most simple way to explain it, and it seems like he got trivial speed ups there. But it seems that the complexity with training, it's almost like in our mind at least, it's almost as complex as training GANs. Like it's like a very delicate balance and oftentimes you, it's just but yeah, it's literally speculative decoding on speculative decoding.Vibhu [00:44:21]: Speculative.Ali [00:44:22]: Yeah. We saw this paper.Vibhu [00:44:24]: It's interesting, right?Ali [00:44:24]: Yeah.Vibhu [00:44:24]: I wouldn't even expect it to be very particular to train, I wouldAli [00:44:29]: Right.Vibhu [00:44:29]: The naive part of me is like, okay, train speculative decoder.Ali [00:44:32]: But like, and it makes sense, like the whole idea of speculative decoding is you. It's like, it's like almost like the iPhone auto predict version but for a normal model, right? Like you're just, you're just, generating three tokens and you're like, okay, I'll do prefill on them. And so you save those three turns for your original model. Now your speculative decoder is doing three turns of auto regression, so why not just have an even smaller model?Ali [00:44:53]: The other question there is what are the size of speculators? So say forPhilip [00:44:58]: Right. It's like a billion parameters.Ali [00:45:01]: Like for MiniMax, it's. Yeah. It's like one layer. It's like one 60th of the original model usually.Philip [00:45:06]: Yeah. I think we should do a paper when we get back to the office.Philip [00:45:10]: SpeculativeAli [00:45:11]: SpeculativePhilip [00:45:11]: Decoding.Ali [00:45:13]: No, it's, it does seem like how, when do you stop? But then it also seems like if you're able to train spec-spec decode for instance, right? Like if you're able to have a small model that is accurately predicts what the intermediate speculator is gonna predict, that is able to predict what the original target model's gonna predict, then why not just use that smallest model directly, right?Vibhu [00:45:34]: Yeah. This isAli [00:45:35]: Like it seems likeVibhu [00:45:35]: Adjacent to the routing problem.Ali [00:45:36]: Right.Vibhu [00:45:36]: Yeah.Ali [00:45:36]: Right.Philip [00:45:37]: The thing with speculators is one of the practical constraints on using them is that you do have to run a small model on the same hardware that you're running the big model on. There is a orchestration and resource competition problem inherent in that, and that is one of the constraints on speculation in general, is that draft tokens cost resources to create and cost software complexity to manage. And so if you have like infinitely recursive speculators, you add in quite a bit of that complexity on the actual implementation within the inference engine as well, not just in the training process.Vibhu [00:46:17]: I was gonna say, I would wonder if you could do similar, like distillation and pruning of, it's the same thing, it's just a model. Can we not just distill a lot of the weights, quantize the speculator, out of my domain? The question that also comes up is, this is all for big server workloads, right? How much of this applies to, say I have this MacBook, I wanna run Gemma really efficiently. Similar problems, not the same?Local AI vs. Data Center InferencePhilip [00:46:45]: Pretty different. I talked to Selo, about this on his podcast a couple weeks ago. The difference between inference engineering for the data center and for production workloads versus inference engineering for local AI, is that we start with fundamentally like different constraints and different goals. With local AI, it's how do I fit this model onto my hardware and then make it less dumb? And with data center influence, it's how do I load this model and then make it less slow? And we care about less dumb, and they care about less slow. But the local AI inference engineering ecosystem, I think has a lot for us to learn from in the data center space. They are experts in various forms of quantization, including dynamic quantization that we just don't touch, in the pruning, in the distillation, in the, layer removal. There'Ali [00:47:42]: Layer removal matters less.Philip [00:47:43]: Yeah. There'Ali [00:47:44]: No one loves pruning really.Philip [00:47:45]: Yeah. Well, but the, but they doVibhu [00:47:46]: Which is surprising, right? But that's, that's a whole different thingPhilip [00:47:48]: Just to fit something on the laptop.Ali [00:47:50]: Right.Philip [00:47:50]: So yeah, it's a, it's an interesting, it's an interesting space. Not necessarily that like their techniques make sense for us to do in the data center, because we have different resources and different goals, but more that the process as well as the openness of that field is something to, admire.Ali [00:48:12]: Yeah. Like to your point, like, certain optimizations that would. Like for instance, Turbo Quantum Sharper, like it made such huge hype on that and we did like a whole deep dive on Twitter and like said, what is it? How does it work? Why is it good or not? And it took off and it was implemented on local devices because your memory bandwidth is so slow on like a MacBook, for instance. But try putting the same thing on like an NVIDIA GPU on a B200 Turbo quant would not be. Like, it would not be used. Like, NVIDIA - Like, NVIDIA made it clear that this is not a good optimization, and we've seen it firsthand where the overhead of doing dequantization, quantization of, in the kernel itself with turbo quant kernel, each end is much slower than the time that you save from doing the bandwidth. ‘Cause on the B200s, you have like 3.5 terabytes per second. You don't need decrease the storage that much. You don't need to do, FP4 KV cache. You don't need to use a requant. There's, there's, there's better optimizations to be made. But on Edge devices, it's extremely important, it's extremely useful. So, seems to be, like, different optimizations there, but then they're all uniquely combined with like all you wanna quantize the model, you wanna do speculative decoding, like certain common prefixes with bothPhilip [00:49:18]: Principles.Ali [00:49:19]: Yeah, exactly. Exactly. Exactly.Philip [00:49:20]: They also do a lot of work on, model parallelism, especially over, heterogeneous topology, where you have, some sparks and they are wired together with, Ethernet, DGX sparks.Ali [00:49:35]: Yeah, this is the Exo Labs guys.Philip [00:49:36]: Yeah. You have, a nu
What did you think of today's message? Support the showWith Northgate Online, you can join us every Sunday live at 9:00a and 11:00a, and our gatherings are available on-demand starting at 7p! Join us at https://thisis.churchSubscribe to our channel to see more messages from Northgate: https://www.youtube.com/@Northgate2201 —If you would like to give, visit https://thisis.church/give/—Check out our Care Ministries for prayer, food pantry, memorial services and more at https://thisis.church/care—You are welcome at Northgate just like you are. Life may be going great for you or you may have hurts, hang-ups, and habits. No matter where you are on your spiritual journey, you are welcome at Northgate. We value the process of journey. We believe in the transformative power of Christ. Northgate has a clear vision of transforming our homes, communities, and world by Pursuing God, Building Community, and Unleashing Compassion.—Follow Northgate on Instagram: https://instgram.com/ngatecfFollow Northgate on Facebook: https://www.facebook.com/ThisIsNorthgate/Follow Larry Davis: https://www.instagram.com/sirlawrencedavisSubscribe to Northgate's Podcast (Apple): https://podcasts.apple.com/us/podcast/northgate/id1583512612Subscribe to Northgate's Podcast (Google): https://podcasts.google.com/feed/aHR0cHM6Ly9mZWVkcy5idXp6c3Byb3V0LmNvbS81ODE2ODAucnNzShare your experience with Northgate by leaving a review: https://g.page/r/CRHE7UBydhxzEBM/review...
The Name Above Every Name Conference Three The Broken Heart and the Descent into the Heart Learning to Stand Before God “The sacrifice acceptable to God is a broken spirit; a broken and contrite heart, O God, Thou wilt not despise.” — Psalm 51:17 There comes a moment in every authentic spiritual life when we discover that we cannot pray as we imagined we would. We begin with generous resolutions. We set aside time. We struggle to be faithful. We seek recollection. Yet gradually another discovery is made. The deepest obstacle to prayer is not distraction. It is ourselves. Not because we are evil. But because our hearts have become divided. Part of us longs for God. Another part clings to the world. Part of us desires silence. Another part fears it. Part of us seeks Christ. Another part still wishes to preserve itself. Prayer slowly uncovers this division. Not to discourage us. But to heal us. ⸻⸻⸻ Archimandrite Zacharias often says that the purpose of the spiritual life is not simply to become moral. It is to acquire a heart capable of standing before the living God. That is a very different thing. Many people imagine holiness as flawless behavior. The saints speak instead of relationship. A man may appear outwardly blameless while never allowing God near his heart. Another may come before God in tears, unable even to lift up his eyes, and return home justified. The difference is not perfection. It is truth. God can only fill the heart that has ceased pretending. ⸻⸻⸻ This is why repentance occupies such a central place in the Jesus Prayer. “Have mercy on me.” These words are not born of self-hatred. Nor are they the cry of someone obsessed with guilt. They are the language of love. 2 Only someone who has glimpsed the beauty of Christ can grieve over the ways he has turned from Him. Repentance is therefore not primarily looking at ourselves. It is looking at Christ. His light reveals everything. Not harshly. Tenderly. Like the rising sun that slowly illumines a valley still covered in morning mist. Nothing is forced. Everything is revealed in love. ⸻⸻⸻ The Fathers distinguish between remorse and repentance. Remorse circles endlessly around itself. Repentance turns toward God. Remorse says, “I have failed.” Repentance says, “Lord, have mercy.” Remorse ends in discouragement. Repentance opens the door to hope. This distinction is essential. Many sincere Christians carry burdens Christ never asked them to carry. 3 They believe continual self-accusation is humility. It is not. True humility forgets itself because it has become absorbed in Christ. The humble man knows his poverty. But he knows God's mercy even more. ⸻⸻⸻ Archimandrite Zacharias writes often about standing before God. This expression appears again and again. It is wonderfully simple. Prayer is not first speaking. Nor thinking. Nor even feeling. Prayer is standing. Standing before the Face of Christ. Standing in truth. Standing in poverty. Standing in hope. Everything else gradually unfolds from this. Sometimes words disappear. Sometimes thoughts become quiet. Sometimes only tears remain. 4 The Fathers are not disturbed by this. The heart has begun to speak. ⸻⸻⸻ One of the greatest gifts God grants us is the knowledge of our own weakness. Strangely, we usually resist this gift. We prefer strength. Competence. Success. Spiritual progress. God often gives something different. He allows us to encounter ourselves. Not to humiliate us. But to make us truthful. Without this knowledge, compassion remains shallow. We judge others because we have never seen ourselves. Once we have stood honestly before God, judgment becomes almost impossible. We recognize the same poverty within ourselves. And mercy begins to flow naturally. ⸻⸻⸻ The publican in the Temple becomes the icon of this prayer. He possesses nothing. No arguments. 5 No achievements. No excuses. He simply stands. “God, be merciful to me.” Nothing more. The Fathers never tire of returning to this image because it reveals the true posture of prayer. The Kingdom belongs to those who have nothing left except God. ⸻⸻⸻ There is another mystery hidden here. As repentance deepens, the heart begins to enlarge. This seems impossible. One might imagine continual repentance making a person smaller. The opposite occurs. Pride narrows the heart. Humility expands it. The more we stand honestly before Christ, the more room He creates within us. We begin to carry others. Their grief no longer feels foreign. Their wounds become our concern. Prayer quietly widens until it embraces the whole Adam. This is why the saints wept. 6 Not because they despaired. But because their hearts had become vast. ⸻⸻⸻ Archimandrite Zacharias frequently speaks of bearing shame. This language can trouble modern ears. Yet he uses it with profound care. The shame he describes is not humiliation inflicted by others. Nor is it psychological self-contempt. It is the willingness to stand before God without hiding. Adam covered himself. The saints uncover themselves before Christ. Nothing remains concealed. Nothing defended. Nothing explained away. And because everything is brought into the light, everything becomes capable of healing. The heart finally stops fleeing. It remains. ⸻⸻⸻ This is perhaps the deepest work of the Jesus Prayer. Not continual repetition. Continual return. 7 Every invocation is another act of trust. Another refusal to hide. Another opening of the heart. The prayer slowly becomes the place where Christ and the soul meet. No visions are required. No extraordinary experiences. Only faithfulness. Day after day. Year after year. Until standing before Christ becomes more natural than standing before ourselves. ⸻⸻⸻ There comes a point when prayer grows wonderfully simple. The words become fewer. The heart quieter. Silence itself begins to speak. This is not emptiness. It is fullness. The Presence for which words were preparing us has quietly arrived. The Fathers never encourage us to seek such moments. They teach us simply to remain faithful. Grace knows its own time. The flower does not open because we command it. 8 It opens because spring has come. So too the heart. ⸻⸻⸻ Perhaps this is the greatest lesson Archimandrite Zacharias offers us. The Jesus Prayer is not given so that we may become experts in prayer. It is given so that we may become truthful before God. The heart that continually cries, “Lord Jesus Christ, Son of God, have mercy on me,” gradually loses the need to defend itself. It no longer fears weakness. It no longer hides failure. It no longer seeks admiration. It desires only to remain before Christ. And this is already the beginning of heaven. For heaven is nothing other than the soul standing unveiled before the Face of the One who has loved it from all eternity. As we leave this conference, let us ask for only one grace. Not extraordinary prayer. Not lofty experiences. Simply the courage to remain before Christ without masks. To allow His light to reveal us. To allow His mercy to heal us. 9 To allow His love to enlarge our hearts until they become capable of bearing both our own poverty and the suffering of the whole world. For this is the hidden work of the holy Name. It leads us, not away from ourselves, but through the truth of ourselves into the inexhaustible mercy of God. 10
The Briefing's Nick Pitts talks about the rise of teen takeovers taking place and the connection between our current attention span with our upbringing. Luke Moon of Generation Zion and The Philos Project shares what to know about the rising following of Nick Fuentes and why the slow walk toward anti-Jewish rhetoric, can soon take over. The Reconnect with Carmen and all Faith Radio are made possible by your support. Give now: Click here
https://garykaltbaum.com/ The opinions you hear on BizTalkRadio, BizTV, or BizTalkPodcasts are those of the hosts, callers, and guests and do not necessarily reflect those of BizTalkRadio, BizTV, or BizTalkPodcasts, its management or advertisers. The information on BizTalkRadio does not constitute a recommendation, offer, or solicitation to buy or sell any product or securities. Please consult a professional before investing.
This sermon uses the agricultural metaphors in Isaiah 28 to illustrate God's ongoing work in the lives of believers after salvation. The preacher explains that just as a farmer employs different methods to thresh various grains, God applies specific pressures and trials to remove the pride and self-reliance that hinder spiritual growth. These difficulties are not intended to destroy but to enhance the believer's character, ensuring they become useful for God's glory. The message encourages the audience to view adversity as a necessary part of sanctification rather than a sign of abandonment or punishment. Ultimately, believers are called to surrender their outer shells so that God can reveal and use the valuable fruit within.
The Value of Continual PursuitSermon Notes
Christians have been saved by grace alone through faith alone, but to what end? This passage challenges us to not just enjoy God's forgiveness but to seek the expansion of the kingdom of God among the nations. Notes: With the temple built... ...the kingdom expands (vv1-6, 17-18) ...the nations ascend (vv7-11) ...the worship perpetuates (vv12-15)
Living Way Community Church
Living Way Community Church
Start Your Transformation Now In this episode of The Jim Fortin Podcast, Jim explores a provocative idea: the listener has no birth date. While the body has an incarnation date, the soul — the true source of power — never dies and therefore was never born. Jim explains that most people live exclusively from their 3D, mechanistic identity, completely disconnected from the continuous soul-consciousness that is the actual source of all power, peace, and creation. Jim unpacks why chasing external power — money, status, recognition — is an outdated model of consciousness that is rapidly collapsing as humanity evolves into higher levels of awareness. He reveals the single word that secretly sabotages most people's ability to create what they want, and offers a radically different way to relate to desire itself. This episode challenges listeners to stop chasing and start remembering who they truly are. What You'll Discover in This Episode: (00:00) You have no birth date — Jim introduces the central idea that while the body has an incarnation date, the soul is continuous and was never born, and explains why this is the seat of true power. (03:24) Continual vs. continuous — Jim distinguishes between the body's continual, repeating incarnations and the soul's continuous, unbroken existence, and why identifying with soul changes everything. (10:13) There is no point A and point B — Jim dismantles the manifestation myth that separates a person from their desire, using the metaphor of a small wave and a big wave both being made of the same substance. (15:56) Old power structures are crumbling — Jim explains why external forms of power (money, politics, status) are an outdated 3D model collapsing as consciousness evolves into a 5D paradigm. (23:15) The hidden danger of "I need" — Jim reveals how the word "need" instantly creates separation from the thinking substance that creates all things, and why this keeps people stuck. (27:10) To need nothing is to have everything — Jim shares the epiphany that anything available to others is available to the listener, and closes with a final truth meant to be felt rather than analyzed. Listen, apply, and enjoy! Transformational Takeaway The body has a birth date, but the soul does not — and remembering this is where real power begins. Most people sabotage their own desires by living from "I need," a word that instantly creates separation from the one substance that all things, including the listener, are made of. There is no point A and point B, no chasing, no struggling to close a gap that was never truly there. The invitation in this episode is simple but profound: stop relating to life from need, and start relating to it from wholeness. To need nothing is to have everything. Mentioned Resources: The Science of Getting Rich by Wallace D. Wattles Seat of the Soul by Gary Zukav Disclosure: Some of the links above are affiliate links, meaning, at no additional cost to you, I will earn a small commission if you make a purchase. Let's Connect: Instagram | Facebook | YouTube | LinkedIn LIKED THE EPISODE? If you're the kind of person who likes to help others, then share this with your friends and family. If you have found value, they will too. Please leave a review on Apple Podcasts so we can reach more people. Listening on Spotify? Please leave a comment below. We would love to hear from you! With gratitude, Jim
What if better management starts with seeing connections that were there all along? In this conversation, Balaji Reddie and Andrew Stotz unpack one of the most powerful and overlooked ideas in management: that everything is connected. Drawing on Dr. W. Edwards Deming's systems thinking, he explains why the problems we face today were set in motion long before we noticed them, and why the solutions are rarely where we think to look. Along the way, a cup of coffee becomes a window into five years of invisible effort, hydrogen and oxygen defy everything we'd expect, and a classroom game reveals that cooperation isn't a soft ideal — it might just be human nature. Whether you run a team of two or a global organization, this episode will quietly shift the way you read a situation, ask a question, or respond to a problem. TRANSCRIPT 0:00:02.2 Andrew Stotz: My name is Andrew Stotz, and I'll be your host as we dive deeper into the teachings of Dr. W. Edwards Deming. Today, I'm continuing my discussions with Balaji Reddie, an educator and trainer in the teachings of Dr. Deming and quality management generally. The topic for today is connectedness, which will help you see things you normally would not see. Balaji, how are you? 0:00:30.7 Balaji Reddie: I'm good, I'm fine. 0:00:32.5 Andrew Stotz: Nice to connect with you. 0:00:34.1 Balaji Reddie: Yes, same here. 0:00:35.3 Andrew Stotz: Connectedness day. 0:00:36.5 Balaji Reddie: Yeah, connectedness, absolutely. Yeah, we chose this term very carefully, right? 0:00:42.6 Andrew Stotz: I suspect you have, given your precision. [laughter] 0:00:47.4 Balaji Reddie: Yeah, to see things that you normally would not see, let me attribute this to Henry Neave, who said that he heard Deming say this during his last three or four seminars in England, right? In Britain rather. Yeah, I think it was mostly in England. And he said, when he heard him say that, he said, "I am not here to teach you anything new. I'm here to make you see things that you normally would not see." And I would like to just expand on that and say when you see things differently, obviously, you observe, you record different things. And when you record different things, you ask different questions. When you ask different questions, you get different answers. And when you get different answers, you draw different conclusions. When you draw different conclusions, you take different decisions. And when you take different decisions, you get different results. It's insanity to expect different results by asking the same questions every single time. So management is not so much about giving the right answers as much as it is about asking the correct questions. And I think Dr. Deming helps you do that. He helps you ask the correct questions. 0:02:02.2 Balaji Reddie: Now coming to the word connectedness. We spoke about the System of Profound Knowledge, a unified theory of leadership and management, where Dr. Deming brings together four sciences. And the first one he says about systems thinking, or what he called as appreciation for a system, appreciation for a fact that everything is systemic. There's nothing that happens in isolation, right? And the second one, of course, he said was understanding variation, then understanding theory of knowledge and understanding psychology. But he said, all four are equally important, et cetera. And then comes the, I won't say contradictory, but a chapter before Profound Knowledge on systems especially. So that's where I decided to use the word connectedness, that all these are connected. What did he mean by that? So I believe he wanted people to learn about connectedness and systems first. 0:03:01.6 Balaji Reddie: So let's dive right there. What did he always say about systems, or connectedness as we say here? And let's get into his definition for system. And he says that "a system is a network," right, "of interdependent components that work together to achieve the aim of the system." And then he says "every system must have an aim. Without an aim, there is no system." And very well articulated throughout this book which he wrote, he always began with aim of the chapter. Even in Out of the Crisis, he always started with aim of a chapter. So network of interdependent components. So let's just break it down, because the three important words are network, interdependent, and aim. And he says aim comes first because without the aim, there is no system. 0:04:00.6 Balaji Reddie: Right. And he goes on to tell us what the aim of a system should be. Okay. He says here that the recommended aim should be for everybody to gain. Now, in one of the paragraphs there, management's job, where he says should be to make it clear, the first step is clarification. Everyone in the organization must understand the aim and how to direct their efforts towards the aim, right? And everyone must understand, okay, this is important. The danger and loss to the whole organization from a team that seeks to become a selfish, independent profit center. Right. So he says here we need to give both, why we need to do what we need to do. And he says the aim precedes the system. Now, I would like to leave this discussion right here. We will cover this in our 14 Points, our session on the 14 Points, because I believe the word aim and purpose are, they go hand in hand. Aim gives us a direction and purpose gives us the reason why we exist, right? And so the direction of all the efforts, you can see the development. He says here, he gives some examples that the system must create something of value. In other words, results. The intended results along with consideration of recipients and of cost mold the aim of the system. So he said it should be something that makes life better for everyone. 0:05:39.8 Balaji Reddie: Now, he was talking about a man-made system here. Remember, he wrote this for organizations, but I think I mentioned this before that this can equally be applied to natural systems, if you want to really study them in detail. I remember talking to Barbara Lawton about this, and I said, "Natural systems, they do have an aim. It's just that we've not understood it." The purpose and aim of, in fact, that was the basis for my paper, which I delivered at the Deming Research Seminar in 2002. And the title of my paper was Deming and Ecology. And I was trying to put together the thoughts of the physicist Fritjof Capra, who wrote some wonderful books like The Tao of Physics, The Turning Point, The Web of Life, and The Hidden Connections. And The Systems View of Life is also one of the sub-books or booklets that he came out with. But these were his wonderful books. At the time, he had not yet written The Hidden Connections, right? So based on his three books that I had read up until then, of course, one was called Uncommon Wisdom, where he reproduced some of the interviews he had with some systems thinkers in the world, including people not from conventional systems thinking. For example, he also reproduced his interaction with the then Prime Minister or the ex-Prime Minister of India, Mrs. Indira Gandhi, and how he spent time with her, because he was over here in India for quite some time studying the language Sanskrit and studying the texts over here, because he learned about systems thinking even deeper. 0:07:25.1 Balaji Reddie: So over there, he wanted to, he raised a question, what is the purpose? So I wanted to use Dr. Deming's theories to understand the purpose of a natural system and what is the purpose of man [laughter] in this natural system. So that was my theory. And that's when this aim thing became more and more clear. And Barbara said, "No, you're right. There is an aim. It's just that we're trying to understand it. Without that, you won't have the system working so beautifully. There is something that is at work." Right. And he says here, management of a system. And then he says that a system includes the future. And now here's where, when you talk about system includes the future, he mentions Peter Senge in his book, if you know about The Fifth Discipline: The Art and Practice of the Learning Organization. And that's where this fantastic statement comes in because he says delayed effects. So although, Deming never used these words directly, but here I am putting among the first principles of systems thinking is that cause and effect are not closely related in time or space. Right. And that gives a lot of frightening conclusions, or inferences that you can draw from this. If I say that, by the way, that cause and effect are not closely related in time or space, and if I ask you what you understand by this statement, your first knee-jerk reaction would be what I do today, the consequences would be felt in another time at another place. Absolutely right. That's exactly what this means. But the converse is also true. What we get to see in front of our eyes, the origins are at another time and another place. 0:09:15.9 Balaji Reddie: So we only see the event, but systems thinking teaches us there are no events in this world. There are only eventualities. Now, this comes into play when you start doing a study, typically when people learn about quality, and then the first thing they learn is to, okay, identify a problem and try to solve it, right? And then they start asking the question, "Why did this happen?" And then you draw a fishbone diagram, which is, you know, one of the most popular tools. But the problem there is you need to arrange the causes logically and not categorically. And you also need to take into account that these causes are also related to each other. You can't isolate them. And the fact that, there could be an element of time over here, right, that this is a delayed effect that you're seeing. So the effect comes much later. And sometimes it could be weeks and months. So you have to take that into account, right? So the cause and effect are not true, which means then the next corollary from this statement is that seemingly disconnected events turn out to be connected. And so when you are studying something that's happening in front of your eyes, very quickly you ask the question, "What else happened a month ago, a week ago, that we seem to have missed? Are we missing something here?" So that's one of those things that we need to get into, right? We need to go backwards in time and not look at the localized stuff here. We need to look in space all around and try to figure out where and how we need to change things. 0:11:06.1 Andrew Stotz: It's also interesting to think about all the things that have to have happened for an outcome to occur today the way you want it, that is the right outcome. 0:11:19.8 Balaji Reddie: Yes. 0:11:21.4 Andrew Stotz: And a good example I use for that, and I was just talking two different times this week with my team, is coffee. Coffee is a really interesting one, because when you go to a coffee shop, they make you a cup of coffee, they give you a cup of coffee, but what you don't realize is that there was a farmer, took him five years to grow that tree. And then they had to pick it at the right time. You pick it at the wrong time, it goes wrong. 0:11:47.4 Balaji Reddie: Oh, yeah. 0:11:48.3 Andrew Stotz: They put it into some water. They leave it in too long, it's fermented. They leave it in too short, it doesn't, the shell doesn't come off. And then if they, once they've worked on it, if they have their grinders that are getting off the mucilage off of it, if they're too tight, well, then you're going to get a little slice off the top of the bean. And then when that goes to the roasting machine, that's going to burn right there because there's no protection in that area. But they got to get all of those things right all the way to that barista. And if that barista just basically doesn't check the water temperature, it's done. 0:12:29.5 Balaji Reddie: It's done. [laughter] 0:12:29.9 Andrew Stotz: You throw it away. And when you think about it, I have another lesson that I give to young people about capitalism, that how all of these decisions are happening all along the way. And when that happens, there's waste in the system. And that waste goes back all the way back to the effort of the farmer five years ago when they planted that tree. 0:12:51.8 Balaji Reddie: That's right. 0:12:52.3 Andrew Stotz: And so reducing waste is not only good for you and good for costs in general, but it has a respect for the process, a respect for the system. 0:13:04.0 Balaji Reddie: Yes. Connectedness. Reconnected. [laughter] 0:13:07.6 Andrew Stotz: Well, and the other thing I tell them is I say, and here you have a coffee bean coming from this country that's then being shipped to this country. They never met. 0:13:16.8 Balaji Reddie: Yeah. 0:13:17.1 Andrew Stotz: And that's capitalism. You trade. In fact, they may not speak the same language. 0:13:23.1 Balaji Reddie: Absolutely. 0:13:23.5 Andrew Stotz: Or they may even be enemies. They may hate each other. 0:13:26.9 Balaji Reddie: Yeah. 0:13:27.4 Andrew Stotz: And yet they trade. 0:13:29.5 Balaji Reddie: Yeah. 0:13:31.1 Andrew Stotz: And so it's a voluntary connectedness that happens through capitalism. Anyways. 0:13:37.2 Balaji Reddie: Yeah. The right quality and uniformity are foundations for commerce, prosperity, and peace. [laughter] That's Deming for you. 0:13:46.1 Andrew Stotz: Yeah. 0:13:46.3 Balaji Reddie: I think he chose that sentence so beautifully. It goes into the Deming Medal. He knew what he was talking about. So he was talking about systems, and I think profound knowledge came towards the end of his life, and rightly so. It came as a revelation, as a catharsis. I think he went through all of that, especially, I think it was between '88 and '89, if I'm not mistaken. Jaki Graham talks about this very, very funnily. She says that they were giving him pure oxygen and maybe that made his brain cells go alive. [laughter] And he could think so clearly. He was getting pure oxygen. And so his thought process... 0:14:31.0 Andrew Stotz: By the way, for the listeners out there, you may feel like I feel when I think about how Dr. Deming didn't really come out with the System of Profound Knowledge until he was, early '90s, let's say, '90s. You think to yourself, okay, I still have time to come up with my incredible insight on... 0:14:52.4 Balaji Reddie: Correct. 0:14:52.9 Andrew Stotz: Bringing together all my experience. 0:14:55.1 Balaji Reddie: Yeah. So because I think I shared this with you before that when you read Mary Walton's book and he says, "No, I don't think I'm ready to establish an institute," and he only did it one month before he passed. So I think by that time he was quite sure what he created was quite amazing, and it'll stand the test of time, right? And although, the title of the book is The New Economics for Industry, Government, Education, I believe it is industry, government, education, and healthcare. Because when he was in Japan, he said Japan must see itself as a system. And when he got the Deming Medal, that's what he said. Oh, sorry, when he got the Emperor's Medal, the Second Order of the Sacred Treasure. And he was talking to the Japanese Prime Minister. And he said, "You all must see yourselves as a system." And the four sectors of industry, government, education, and healthcare must work together. So he meant it as in the broadest sense of the term, right? And that's why this word system. So the first thing, of course, cause and effect are not closely related. Seemingly disconnected events turn out to be connected. The third important thing here in systems thinking is the reductionist thinking that we have. 0:16:12.0 Balaji Reddie: We believe the word analysis means dissection. And if you look deeply and read the book Mind and the World Order, what is the purpose of analysis? The purpose of analysis is not dissection. That's just the act of breaking a system down. But the purpose of analysis is interrelationships, establishing the interrelationships or understanding the connectedness. And I find it quite funny because people say analysis is dissection, and so they came up with a new term, really not necessary, synthesis, bringing it all together. There was no need for that term. You're not engaging in what, tautology? [laughter] You didn't need that. Analysis itself means dissecting and understanding relationships, right? 0:17:02.2 Andrew Stotz: Hmm. 0:17:03.3 Balaji Reddie: And so he said here, we have that reductionist thinking, but we forget that in a system, you have what are called synergistic relationships. You cannot understand a system by breaking it down into its components and studying the components separately and then believing that the components and their attributes and what they do gives the output of what we see. No, you could be completely wrong here. It's not what they do separately. Okay, it's a good thing to understand the components of a system and study them and maybe even establish connectedness, but sometimes what the results you see are quite contrary to what they could be giving separately. I'll give you an example here. Take the case of water. Water is composed of hydrogen and oxygen. You cannot study and draw conclusions about water by studying the attributes of hydrogen and oxygen separately. You'd be making a mistake, a grave error. Okay, let's do that. Hydrogen, a highly inflammable gas, catches fire like this. Oxygen feeds fire. But you put them together and they give you something that quenches fire. Hydrogen is not wet. Oxygen is not wet. But water is wet. [laughter] How do you explain this? 0:18:25.9 Balaji Reddie: All right. If you take the case of sugar, it's a hydrocarbon. Try tasting hydrogen. All right, try tasting carbon. Now put them together and see what you get. Now, can it not be then that you are studying a system and when you break things down and you're studying the components separately, you could be making the same error? It's not what they do separately, it's what they do together. It's the same with sections in a company, departments in a company, companies themselves. Today, we use this sophisticated term, supply chain, call it what you will. I love to use the word provider network. I don't like using the word supplier any longer. I like to use the word provider, because today we're living in a different world. We're getting activities done by someone else, and you can't say he's supplying me this activity. He's providing me this service, right? So provider is a better term to use than supplier. And it's no longer a chain. It's a network, right? It's more lateral. It's all over the place. It's not linear. You're doing a huge disservice by using the term supply chain, because it's not bringing out what exactly is happening right there. 0:19:38.3 Balaji Reddie: And today's supply chains or provider networks are globally dispersed. So you need to have something that connects them. You just said that, the farmer could be in a different country, speaking a different language, maybe he didn't even know where it is being used, and they could hate each other, but they have to work, [laughter] you need to work together. So there's some kind of connectedness that comes in once again. But here's where people make mistakes. So you're talking about synergistic relationships. So sometimes in a team, now here's where all that incentive pay and choosing employee of the month can destroy a system. 0:20:18.8 Balaji Reddie: It could have been the entire team that did well, and they did well because they worked as a team. And it was synergistic, where some components had to, I can't use the right word, but maybe had to underperform so that the system came out ahead. It's like the organs of the human body, right? That's an interconnected, interdependent system. You have a primary function of each of the organs, but there are many other functions we are still discovering, right? For example, if you take the case of ears, besides hearing, they're responsible for balance in our body, right? Eyes, besides sight, they are also responsible for balance in our body. If you try this yoga position, you know, there's a yoga, so to say, yeah, you stand on one leg. And then you lift one leg off the ground and put the ball of your heel on the knee, the other leg's knee. And then you hold your hands above your head to form like a namaste or a triangle. As long as your eyes are open, you stand upright. Shut your eyes, you fall. It's a system and we don't even realize it, right? So this is something we're still... 0:21:38.3 Andrew Stotz: You also fall in that posture when your teacher says, as mine did once, "Now you're in the position. Now raise your other leg." [laughter] 0:21:54.5 Andrew Stotz: My yoga teacher didn't come yesterday, but I did my sessions every day. But I haven't been able to do that one yet. 0:22:03.6 Balaji Reddie: We had that joke in school. Why does the crane stand on one leg? [laughter] 'Cause if it raised the other leg, it'd fall. [laughter] So it continues standing on one leg. So yes. So you have the synergistic relationships which exist, all right? And it's so difficult then to establish these interrelationships. And sometimes they're visible, sometimes they're not, all right? And sometimes you get to see the results right up front. Sometimes you get to see them in time delay. And that is why Deming said continual improvement. It never stops. Continual learning. You can never say you have complete knowledge of a process or a system. You can never say that, right? And now comes the best part. When you say here that all these are interconnected and there's synergy and then there's cause and effect not closely related in time and space, the output of a system, that is why, is subject to variation. And that's why the study of variation becomes important, right? 0:23:16.2 Balaji Reddie: And when you draw a diagram saying, suppose you say this is a system and all the components are there and you draw the interconnectedness and you draw an arrow and you say, "Okay, you're getting the system is subject to variation," right? Can you with absolute certainty say that because here, because the funny thing is each of the component's performance is also subject to variation, right? So can you say that, this combination gave me this output? You can never say that. You can only say it's probable, might be somewhere here. And so I think Walter Shewhart's greatness was in recognizing that what is natural to a system and what is not natural. When do you start asking the question, "What do I know about this?" When you cannot predict. When you can predict with a certain amount of certainty that this is going to happen. 0:24:14.6 Balaji Reddie: It is between these two limits. Well, then that's when you say, "Okay, I think I know quite a bit about this, but I don't know much. I want to reduce those limits, because I want to be as certain as possible when I make a prediction." And that was what was continual learning. It was not so much about cost. It was as much as it had to do about understanding and predicting. I know what's going to happen next. To a certain extent, this will happen. This will happen, right? And that's where the understanding of variation comes in. Okay, and then he brought in the third, that in variation, you saw that there were what he called as random variation and assignable variation, or what Deming had common and special cause variation. Common causes belong to the system and special causes are alien to the system under study, right? And so say it's easier to identify the special causes because they tell you something is not right. And so you can say, "I know it happened." You can isolate and say, "This is the cause," right? It makes it easier for you then to isolate that, study it, but... 0:25:30.1 Andrew Stotz: I just had that happen this morning. 0:25:32.4 Balaji Reddie: Okay. 0:25:33.7 Andrew Stotz: I'll give you an example. 0:25:35.2 Balaji Reddie: Yeah, sure. 0:25:36.0 Andrew Stotz: I've been working with my accounting team and what we're trying to do is make sure we have on-time and accurate monthly financial statements. And I've chosen those words very, very carefully. And what I've done with the team is ask them, there's no pressure in this. All I'm asking you to do is pick a day that you think we will be able to have on-time and accurate monthly financial statements. And then they pick their days for each month to the end of the year, which pretty much certainty accounting is a repeating of a similar function. And so then as I was talking with them throughout that day, I said, "Why are these boxes piled up around here?" And they said, "Well, the room that we store them in is getting full." The paper, 'cause you have a regulation in Thailand, you have to have your accounting documents for up to six or eight years on premises in case the revenue department comes. So I said, "Well, show me that room." And we went to the room and it was an absolute mess. It's like people have been dumping boxes and papers in there for five years. 0:26:51.6 Balaji Reddie: Wow. 0:26:52.6 Andrew Stotz: And I said, "Okay. You got to clean up this room this month." They're like, "Ah, it's going to take us two months." I said, "Look, this is critical." And they agreed. And so they have now completed that task, but it consumed so much of their time to do that. And it looks beautiful now. But they're two days late on the on-time and accurate monthly financials. And if we were to look at that on a control chart, we would see that it's a big deviation from what they've been able to do. But that special cause is almost meaningless, just as meaningless as the random variation in the sense that this was just a simple situation where there was a trade-off, they made the trade-off, and it's an acceptable trade-off. 0:27:43.1 Balaji Reddie: Yeah, I'd like to put it this way. When you know why it has happened, it's not that harmful, right? As long as you can predict when we do something. It's the unpredictability that makes a special cause harmful. And that's why when I used the term that Deming's thinking was about reducing or eliminating harmful variation, that is things which get out of your hand and you usually don't know why it's happening. So once you get to know, I believe that this could be connected, yeah, that's a good place to start, right? 0:28:15.7 Andrew Stotz: Yeah. Now, someone that doesn't understand the situation and look back over time may not understand what that cause was, which I'm thinking that I want to develop kind of a note-taking system where we actually write down what was that so that we have a record and an understanding of it. But... 0:28:37.1 Balaji Reddie: Yeah. And so when you say here that there's so many, that complete knowledge is not possible, and that's where you start thinking that was Deming being a pessimist to say complete knowledge is not possible, unknown and unknowable. But no, he was just being realistic. But he did not leave us right there. He said the way you can make the unknown as known as possible is through the theory of knowledge. And that's how the whole thing is. It's beautifully done, what the theory of knowledge is. But he says here, very simply, one of those exercises which you see in the new edition of The New Economics on page 58. But if you go to the old edition, right? That's, I believe, on... Where would that be? Page 84. The example he gives there, effect, the plans and their effect, area A, B, and C, the matrix diagram, if you see that? 0:29:50.0 Andrew Stotz: Yeah. 0:29:50.5 Balaji Reddie: That he says was devised by Henry. And then, of course, he picked on that and put it into his book. I got a chance to actually work on this in the form of a game. And the game, I would break up the audience into four or five groups of five people each and then say, "You're department A, department B, department C," and then have a host of projects lined up for them. And I say, "Now you're going to be ranked and rated based on how you do as a department." And so they would choose those projects where only they were benefiting, right? And obviously, that's what you see in the first table. A chooses three projects, and B chooses two projects, and C chooses three projects. And you see the plus signs there indicating that it's for them, for them alone. And then they are exposed to the fact that, wait a minute, now we're going to give you some extra information. You chose these projects because they were good for you, but now we're going to tell you that these were bad for the others, right? And only one of them was good. 0:30:59.9 Balaji Reddie: And I say, "Okay, now here's this extra information you've been given. Now what are you going to do?" All right. And I said, "I'm giving you five minutes to... Because the company is going to do badly, right? And even though you get a good rating, but the company does badly and eventually the company might have to shut down. So what would you want to do now? I'm giving you five minutes to rectify this." You know, Andrew, not even five minutes was necessary. They would come back within two or three minutes and say, "Let's just choose those projects where everyone gains." And I was surprised the first time that happened. And that's why Deming said cooperation is the key and it is natural. It's just that we have brought in this competition thing. So competition is not natural. He says cooperation is natural. [laughter] 0:31:53.9 Andrew Stotz: And that, in Asia, and I can speak to Thailand, what I like to tell people here in Thailand for managers and people running business is that the core strength of Thai people is that they want to work together. So when you break them up and you put them into groups and then you incentivize them either by group or by individual, you destroy one of their core strengths. 0:32:28.7 Balaji Reddie: Right. 0:32:29.1 Andrew Stotz: That's not a core strength of American. A lot of Americans don't want to work together. They just want to do their own thing. Right? They've been trained to be independent and they don't want that hassle. But in Asia, particularly in Thailand, where we're struggling to bring more and more value, what's happened is the KPI mentality of divide and conquer really is damaging. And it brings me to another point or question that I wanted to mention, and that is about the natural... My mom has a saying, and it's "inch by inch, it's a cinch, yard by yard, it's hard." And what she was telling me was to break things down into manageable parts and then work on the next right part, and then you will start to make progress. And we have studied under Adam Smith and the concept of how division of labor brings value, can increase output. And so there's a whole side of things. And then we're incentivized as individuals, particularly in the West, but let's say generally that, if it's to be, it's going to be up to you and you got to make yourself performance and your department's performance. I don't care about the others. I care about what you're doing with yours. 0:32:54.6 Andrew Stotz: So there's an accountability aspect and it all comes together to give people the perception that, hey, I need to break this down into smaller parts. And so there's just this natural inclination to do that. And if you were to say to people, you can't break it down into individual parts, that also doesn't work. You have to be able to understand the different parts and their interconnectedness, going back to what you're saying. 0:34:26.8 Balaji Reddie: Yes. 0:34:27.1 Andrew Stotz: But there's a lot of different things that I'm saying here. But I think the thing that I'm saying is that the pressure on most people is to divide and conquer, to say, "I'm going to go and look at this department and you and your performance." And what Deming was saying was, you've got to look wider than that if you're really going to be successful. Tell me where I'm right or wrong or give me some guidance. 0:34:53.5 Balaji Reddie: Yeah. I mean, that's exactly how it is. And I just told you that when I saw this happen, it was like an eye-opener. And then what happened after that was, okay, there were some other projects which where we were not going to do well, but the company's going to come out ahead. So why don't we choose those? And that's exactly what you see in the table, that they choose certain projects where they themselves are... Now that's real optimization, right? That's getting the best for everyone. 0:35:23.4 Andrew Stotz: Yeah. 0:35:24.8 Balaji Reddie: And today it's become an overused and abused term. For anything they use the word optimize. I think Deming was way very clear about this 30 years ago when he came up with this idea. And he said it's for everybody to gain. It's not compromise, it's optimization. Right? So I, of course, have a different take on optimize and we'll see that next week when we come into four parts of profound knowledge. So I'd like to end here with talking about cooperation is the key and I think we realized why. 0:35:59.1 Andrew Stotz: Okay. 0:35:59.8 Balaji Reddie: Deming says it's not about each other, it's win-win. The system should... And he gives that lovely example in the same chapter, the last example, where he talks about the truck company, the two truckers who used to keep their establishment open. Alternately overnight so that both could gain and they would use each other's resources. He said, "I was surprised when I called the pickup truck and I saw that the truck was from his competitor. He sent his competitor's truck to pick up my car." That was a classic case of cooperation. So since he ends with that, he ends that chapter, I'd like to also speak about that. Cooperation is the key because we are dealing with something very complex, right? And none of us is as strong as all of us. It's not just about, like you said, because cause and effect are not closely related, it is imperative now that we need to work together. So that was with systems thinking. Cause and effect are not closely related. There exist synergistic relationships. Seemingly disconnected events turn out to be connected. And outputs are the result of a myriad of inputs. You can never say you know everything. So that was systems thinking for you as they apply to the elements of profound knowledge as well. 0:37:33.3 Andrew Stotz: Great. 0:37:33.7 Balaji Reddie: Next week, we'll get into profound knowledge. 0:37:37.1 Andrew Stotz: You know, I think my ending thought on this goes back to the concept of meditation. 0:37:45.4 Balaji Reddie: Okay. 0:37:45.9 Andrew Stotz: Concept of insight meditation or Vipassana or... 0:37:52.6 Balaji Reddie: Vipassana. 0:37:53.7 Andrew Stotz: Yeah. So part of what I understand about meditation is the idea of clearing our mind in this moment. It may also be just observing our mind and just observing all that's happening and going on, which is fine. And let's say that a meditation master may be able to clear their mind. But let's just say that what we're trying to do is come down to this instant, this moment. And I remember during COVID when a lot of pressures were on in my business and stuff, I used to meditate and focus on one saying, which is, "I'm okay right now at this moment." 0:38:36.7 Balaji Reddie: Right. 0:38:37.1 Andrew Stotz: And why this is interesting is because let's say this moment is time zero. What you've just explained is that there's about 50 million things before time zero that have been connected in one way or another that got you to time zero. 0:38:55.1 Balaji Reddie: Yes. 0:38:55.3 Andrew Stotz: And you also said that a system, as Deming said, is a system includes the future. 0:39:00.2 Balaji Reddie: Future. 0:39:01.2 Andrew Stotz: And that's going to be a function of all the different things that happen after time zero, the interconnected things. And it just made me realize, ah, so the value of meditation, I can see it even more now when I think about all the connections on both the future side and the past side. And so it just helped me think about it from that perspective, a little bit different, but anyways. 0:39:23.4 Balaji Reddie: Yep. 0:39:24.9 Andrew Stotz: So I just want to thank you for taking the time. And in fact, what I was thinking is that I need to do more interviews with our advisors and other board members of the Institute, including yourself, because they've all got a lot to say. And I've got some discussions going on with that right now. So I'm looking forward to that. And so I appreciate the discussion and the journey that you're taking us on. And just for the listeners out there, remember to go to deming.org and jump into DemingNEXT to continue your journey. 0:40:03.8 Balaji Reddie: Yes, absolutely. 0:40:05.3 Andrew Stotz: There's so much, so much to learn and explore. So I'm going to wrap it up there. This is your host, Andrew Stotz, and I'll leave you with one of my favorite quotes from Dr. Deming, and that is, "People are entitled," that includes you... 0:40:21.7 Balaji Reddie: Yes. 0:40:22.4 Andrew Stotz: Are entitled to joy in work.
Thomas Ahle wants Normal Computing to be the Lovable for chip design: type your intent, and a swarm of agents carries it from design through optimisation, formalisation and verification to tape-out. To get there, his team at wrote their own open-source Verilog simulator, 580,000 lines in 43 days, because commercial EDA verifiers run about $10,000 per core and there are no decent open-source compilers to build on.That sets up the question Tim keeps pressing: if an agent can produce a chip design, a proof, or a working program, how do you actually know it is correct? Passing 70% of tests is not the same as being right, and a single fabricated bug can cost a company a fortune. They dig into ProgramBench (rebuild a program from its tests, roughly 0% success), the difference between structure and competence, and the "understanding debt" you take on when nobody reads the code.From there: auto-formalisation in Lean and the AlphaProof trick of training on prove-or-disprove; why there is no single true representation of a spec (Petri nets, TLA+, Erik Curiel's "math does not represent"); and thermodynamic computing, where Normal Computing's CN101 chip is built so that its physical noise *is* the computation, settling a stochastic differential equation in hardware to invert a matrix. Plus Bayesian uncertainty, specialisation, the Chomsky hierarchy, AI slop, and whether performance is all that matters.Recorded in Zurich.Disclosure: Normal Computing paid our production and travel costs for this show. We retained full editorial control. They did not see the video before publication, and we did not show it to them or discuss it with them beforehand.---TIMESTAMPS:00:00:00 Meet Thomas Ahle: the Lovable for chip design00:03:41 Why hardware needs formal verification00:06:36 Ten thousand dollars per core and a six-month agent run00:07:40 Rebuilding programs from tests: ProgramBench and zero percent00:12:15 Structure vs competence: can you learn a program from behavior?00:15:27 Continual learning, abstraction, and Claude as an ecosystem00:23:17 Autoformalization and the AlphaProof trick00:29:31 No single true representation: specs, Petri nets and TLA+00:34:43 Thermodynamic computing: when noise is the computation00:37:32 Bayesian uncertainty in the age of token streams00:41:12 Hybrid compute: vibe-coding loops, binaries and Stockfish00:44:44 Co-design, central-AI apps and API pricing00:49:45 Chain of thoughtlessness and the Chomsky hierarchy00:53:40 AI psychosis, slop and the broken social contract00:57:34 Typing it yourself, teamwork and performance vs competence---REFERENCES:person:[00:00:10] Thomas Ahlehttps://thomasahle.comorganization:[00:00:27] Normal Computinghttps://normalcomputing.com/paper:[00:11:21] ProgramBench: Can Language Models Rebuild Programs From Scratch?https://arxiv.org/abs/2605.03546[00:31:55] Autoformalizing Memory Device Specifications with Agentshttps://arxiv.org/abs/2605.00058[00:35:20] Thermo AI and the Fluctuation Frontierhttps://arxiv.org/abs/2302.06584[00:36:40] Thermo Comp System for AI Applicationshttps://arxiv.org/abs/2312.04836[00:37:05] Thermodynamic Linear Algebrahttps://arxiv.org/abs/2308.05660[00:44:50] An efficient probabilistic hardware architecture for diffusion-like modelshttps://arxiv.org/abs/2510.23972tool:other:[00:01:00] Building an Open-Source Verilog Simulator with AI: 580K Lines in 43 Dayshttps://normalcomputing.com/blog/building-an-open-source-verilog-simulator-with-ai-580k-lines-in-43-days[00:02:55] Normal Computing Announces Tape-Out of the World's First Thermodynamic Computing Chip (CN101)https://www.normalcomputing.com/blog/normal-computing-announces-tape-out-of-worlds-first-thermodynamic-computing-chip[00:32:02] DRAMBench: Autoformalizing DRAM Specifications with Timed Petri Netshttps://www.iese.fraunhofer.de/blog/drambench-autoformalizing-dram-specifications/---ReScript: https://app.rescript.info/share/ff9684a112ab37744096adaeb097a263
Pastor Jonathan Phillips teaches that worship is more than a song. Drawing from Psalm 66 and Romans 12:1, he unpacks what it looks like to live a worship life beyond Sunday morning. Whether it is in the extraordinary, the mundane, or the painful moments of life, worship is a continual response to who God is. If you have only experienced worship as a moment, this message will change how you see the rest of your week.Watch on YouTube: https://youtu.be/fHi4MIRLgeE
Pastor Jonathan Phillips teaches that worship is more than a song. Drawing from Psalm 66 and Romans 12:1, he unpacks what it looks like to live a worship life beyond Sunday morning. Whether it is in the extraordinary, the mundane, or the painful moments of life, worship is a continual response to who God is. If you have only experienced worship as a moment, this message will change how you see the rest of your week.Watch on YouTube: https://youtu.be/fHi4MIRLgeE
Pastor Wayne Van Gelderen shares biblical truth that will bring hope and comfort in these uncertain days. May we draw closer to God through this time and impact those around us for eternity. https://fallsbaptist.org https://baptistcollege.org https://www.theegeneration.org https://ontovictorypress.com If you'd like to support this ministry - https://fallsbaptist.org/give/
Introduction - What is meant by inner, Christian experience? I. The Experience of the Anguished II. The Experience of the Merry in Christ III. The Continual Feast Application - Christ makes your heart merry and feasting, there is always more to discover
Text: 1 Peter 2:4-12; 4:10-11 Preacher: Derek Baker
What does it take to build economies that are truly inclusive—and built for long-term impact? In this episode of Develop This!, Dennis Fraise speaks with Philip Gaskin to explore his diverse career spanning the private sector, entrepreneurship, and community development—and how those experiences shape his approach to economic transformation today. Philip shares how lessons from the private sector can directly inform economic development strategy, especially when it comes to innovation, ecosystem building, and driving measurable community impact. The conversation also highlights the importance of addressing systemic barriers that limit access to capital and opportunity in underserved communities. A key focus is the role of economic developers as connectors—bridging policy, private sector insight, and community needs to build stronger, more resilient local economies. The discussion also touches on Kansas City's evolving economic landscape and how regional ecosystems can serve as powerful models for inclusive growth and entrepreneurial support. Key Takeaways Private sector experience can strengthen economic development strategy Inclusive growth requires addressing access to capital and opportunity gaps Economic developers play a key role in policy and collaboration Strong ecosystems drive innovation and community transformation Kansas City offers a model for regional economic growth Continuous transformation is essential for long-term impact Key Topics Covered Private sector lessons for economic development Community and economic transformation Inclusive entrepreneurship and access to capital Role of economic developers in policy and collaboration Kansas City's ecosystem and growth model Sound Bites "Crazy times call for crazy organizations" "Communities need to embrace their potential" "Continual work on transforming communities"
A common trap in the Christian life is the “graduation mindset”: I got baptized, received First Communion, got confirmed… I'm good. Joe Rockey and Father Boniface Hicks argue that this is not only false—it quietly starves your soul. This episode is a practical invitation and blueprint for continual conversion: ongoing reaffirmation with Jesus that turns faith from a box you checked into a life you live.Father lays out a simple foundation that makes growth sustainable: Sunday Mass, monthly confession, daily prayer (15 minutes to an hour), spiritual reading, and a dose of silence. Once those basics are in place, faith begins to “take on a life of its own.” You start pulling on a thread—an event, a parish opportunity, a lead—and it opens doors you didn't plan: Bible study, new friendships, new discoveries, deeper prayer, real formation. And God isn't passive in any of it—He attracts, invites, and prepares opportunities without manipulating your freedom.Joe adds what this looks like in real practice: don't stay a passive listener to Scripture. Put yourself in the scene. Notice the emotions that aren't written down. Ask what the apostles needed their readers to understand and why. That habit of deeper attention builds a stronger interior life—and even changes how you hear the homily at Mass. The call is simple: keep going deeper, because depth is what breaks the “I did this once, I'm done” illusion.Key IdeasThe “I'm done” mindset (post-sacraments) is spiritually costly; the antidote is ongoing conversion.A durable foundation: Sunday Mass + monthly confession + daily prayer + spiritual reading + silence.Growth often starts with a small “thread” (event/opportunity) that becomes a habit and opens unexpected doors.God draws without coercion: invitation, attraction, prepared opportunities—no manipulation.Go deeper in Scripture by entering the scene: emotions, relationships, motives—not just facts.Links & References (official/source only)Hallow (official):https://hallow.com/Bible in a Year (Ascension, official):https://ascensionpress.com/pages/bibleinayearCatechism in a Year (Ascension, official):https://ascensionpress.com/pages/catechisminayearJeff Cavins (official):https://www.jeffcavins.com/CTA: If this helped, please leave a review or share this episode with a friend.Questions or thoughts? Email FatherAndJoe@gmail.com .Tags (comma-separated)Father and Joe, Joe Rockey, Father Boniface Hicks, continual conversion, ongoing conversion, sacraments, baptism, first communion, confirmation, Sunday Mass, confession, monthly confession, daily prayer, spiritual reading, silence, Scripture, Bible study, catechism, formation, discipleship, Catholic life, parish life, retreat, pilgrimage, parish mission, Eucharistic adoration, holy hour, daily Mass, Hallow app, Bible in a Year, Catechism in a Year, Jeff Cavins, homily, spiritual growth, curiosity, habits, events to habits, freedom, God's invitation
The new AIEWF website is live! CFPs close in 2 days and we will run our first New Engineer Orientation this weekend, get your tickets booked ASAP as they -will- sell out. Take the AI Engineering Survey and get >$2k in credits and free AIE WF tickets!One of the central tensions in the agents industry is that even while there are major decacorn agent labs like Sierra, Decagon, Notion and Cursor being built up, it is also true that it has never been easier to DIY agents, with a plethora of agent frameworks like LangGraph and Pydantic and Flue, and managed agents from Anthropic and Gemini and Amazon. There has been a wave of companies building their own background agents from Shopify to Stripe to Paradigm to Razorpay, and even Cognition's friends Ramp have built their own coding agent with other friend Modal.You'd think Cognition might feel a bit threatened, but they're not - even after all this, they were way oversubscribed for the $1B Series D they just announced:Walden Yan, coiner of context engineering and Chief Product Officer/Cofounder of Cognition, invited OpenInspect's Cole Murray to talk about why the Devin is in the Details.Full conversation live on the pod today: In retrospect, async agents were the most AGI pilled bet you could make in 2024 - the models weren't good enough yet to vibecode, and people didn't trust AI enough to let it rip, nobody (including early Cognition) was sure about the form factors. Now it is obvious:* The first wave of AI coding tools made the developer faster but remain heavily in the loop. Copilor and Cursor's tab autocomplete are prime examples However, the workflow was still heavily centered around and bottlenecked by the developer's local workflow: a developer in an IDE, watching the model, accepting or rejecting changes, and pushing code one interaction at a time.* The second wave was local agents: Claude Code, Windsurf, Cursor's agents pane: first one and increasingly many terminals all running concurrently.* The current Age of Async Agents points to a different future focused more on agent orchestration which drives end-to-end development.According to previous guest Steve Yegge, there are finer-grained 8 levels to agent adoption, but we have collapsed it into three.As Cursor's Michael Truell put it in The third era of AI software development:Cursor is no longer primarily about writing code. It is about helping developers build the factory that creates their software. This factory is made up of fleets of agents that they interact with as teammates: providing initial direction, equipping them with the tools to work independently, and reviewing their work.The agent should not sit solely inside the developer's flow. It should be setup to work in the background so that you can give it a task, a repo, a machine, a shell, a browser, tests, memory, and review loops to go do the work somewhere else.In less than a year, the sentiment has shifted from avoiding multi-agent systems:to suggesting approaches that actually work:From coining “context engineering” to building the infrastructure behind Devin's 7x PR growth and jump from 16% to 80% of commits across Cognition repos, Walden Yan has had a front-row seat to the background-agent shift. In this episode, Cognition co-founder and CPO Walden Yan joins swyx alongside Cole Murray, creator of OpenInspect, to unpack why everyone is building their own Devin, what changed after the December 2025 model inflection, and why “spec to pull request” is now becoming a real production workflow.We go deep on the architecture of background agents: harness-in-the-box vs out-of-the-box, why Devin separates the “brain” from the machine, why repo setup is still one of the hardest problems, why Docker is not always enough, and how full VMs, snapshots, scoped secrets, GitHub bots, Slack integrations, and video-based testing all fit together. Walden and Cole also dig into memory, MCP limitations, multi-agent orchestration, AI code review, SRE auto-triage, PMs shipping code from Slack, Windsurf 2.0, hybrid frontier/sub-frontier systems, and the real failure mode of uncontrolled vibe coding: your codebase regressing to your worst engineer.And as agents eat software… and software eats the world… you can draw the conclusion on what is next:We discuss:* Why the engineering world is waking up to background agents and cloud agents* The December 2025 model inflection that made spec-to-PR workflows practical* Devin's 7x merged PR growth and rise from 16% to 80% of commits* Why Cole built OpenInspect as an open-source background-agent system* The economics of $20/seat agent products and why monetization is tricky* What Cognition actually sells beyond Devin: infra, onboarding, integrations, and adoption* Harness in the box vs out of the box, and why architecture matters* Why Devin separates the brain from the machine for security and permissions* Repo setup, scoped secrets, Docker Compose, and agent-ready dev environments* Why full VMs matter when agents need to run real applications and test them* Android, macOS, Windows, nested virtualization, and machine-specific agent work* Why testing is much harder than “computer use”* Screenshots, video verification, and the “I know it works” merge moment* GitHub UX, Devin Review, AI reviewers, and agents responding to PR comments* Why MCP alone is not enough for first-class Slack and enterprise integrations* Memory, Knowledge, skills, Claude.md, and why retrieval is still unsolved* Devin's auto-generated memories and the challenge of memory pruning* Always-on agents as permanent PMs for issues, tickets, and product areas* Sub-agents, meta-Devin management, and what multi-agent systems actually add* Why pure auto-merge vibe coding breaks down after about two weeks* AI code smells, lint rules, reward hacking, and Semgrep for agent-written code* GitAI, inline context, and preserving the “why” behind code changes* Local testing, mock servers, older codebases, and preparing companies for agents* Windsurf 2.0 and the handoff between local foreground agents and cloud background agents* SRE auto-triage, support workflows, and agents as first responders* PMs, marketing, and non-engineers creating pull requests from Slack* AI agent budgets, $1k-$5k per engineer spend, and hybrid frontier/sub-frontier systems* The rise of autonomous coding factories and who Cognition is hiringWalden Yan* X: https://x.com/walden_yan* LinkedIn: https://www.linkedin.com/in/waldenyan/Cole Murray* X: https://x.com/_colemurray* LinkedIn: https://www.linkedin.com/in/colemurray/* OpenInspect / Background Agents: https://github.com/ColeMurray/background-agentsTimestamps00:00:00 Introduction00:00:43 Why Everyone Is Building Their Own Devin00:01:57 Devin's 2025 Ramp: 7x PR Growth and 80% of Commits00:03:49 OpenInspect and the Rise of Open-Source Background Agents00:07:59 What Cognition Actually Sells Beyond Devin00:09:56 Background Agent Architecture: Harness In vs Out of the Box00:12:08 Separating the Brain from the Machine00:14:07 Repo Setup, Secrets, Docker, and Full VMs00:19:13 Why Testing Is Harder Than Computer Use00:22:40 Video Verification and the “I Know It Works” Merge Moment00:23:19 GitHub UX, Devin Review, and AI Code Review00:25:42 MCP, Slack, and Enterprise Agent Integrations00:28:59 Memory, Knowledge, and Always-On Agents00:36:16 Sub-Agents, Multi-Agent Orchestration, and Meta-Devin00:43:55 Vibe Coding, Auto-Merge, and Codebase Decay00:48:38 Agent Infra, VPCs, Cloud Providers, and Fast VM Restore00:52:25 AI Code Smells, Reward Hacking, and Code Review Systems00:56:10 Making Codebases Agent-Ready00:58:30 Windsurf 2.0 and the Local-to-Cloud Agent Handoff01:01:15 SRE Auto-Triage, PMs Shipping Code, and Agent Use Cases01:04:32 Agent Budgets, Hybrid Models, and Autonomous Coding Factories01:06:51 Hiring at Cognition and OpenInspect Consulting01:07:45 OutroTranscriptIntroduction: Walden Yan, Cole Murray, and Context EngineeringSwyx [00:00:00]: All right, we're in the studio with Walden Yan, co-founder of Cognition, CPO.Walden [00:00:08]: Happy to be here.Swyx [00:00:09]: Which is a cool title. And coiner of context engineering.Walden [00:00:15]: Although I think there are many people who'd used the terms in various ways beforehand, but I did find that people, both internally and externally, enjoyed the upgrade from prompt engineering or model wrapping into maybe a more thoughtful way to build agents.Swyx [00:00:33]: For those who haven't caught up on that, I have on screen the Don't Build Multi-Agents post, which you should go read on and we might refer to, and Cole Murray, who created OpenInspect.Cole [00:00:43]: Great to be here.Swyx [00:00:43]: So let's talk about it. Everyone is building their own Devins. What's going on?The December Shift: From Handholding Models to Autonomous PRsCole [00:00:51]: So I think the engineering world is waking up to this idea of background agents, cloud agents, whatever you'd like to call it. And I think we saw a shift around the December timeframe of 2025, where the models Opus 4.5 and GPT 5.2, they reached a capability where we moved away from handholding the model and being able to actually more or less autonomously drive the model. And what I mean by that is that we could pretty much go from a specification to a completed pull request, assuming the spec was good enough, with very little friction. And that paradigm alone, I think, changed a lot of how we interact with agents, and opened this world where background agents became more practical.Swyx [00:01:41]: I think for Cole, everyone experienced this in December, but I feel like there was just this increasing ramp, right? There was this moment which was, I think, Sonnet 3.7, where, You guys rewrote Devin in one night or something. So describe 2025 or how it felt from your side.Walden [00:02:01]: In retrospect, we always thought it was ramping up, but then even now, over the last three, four months from today, it's been ramping up even faster. So it's almost funny to be talking about how, big of a leap Sonnet 3.7 was, and honestly, a lot of it was stripping out parts of Devin that were no longer needed with that jump in of intelligence. But I also just think that a lot of the recent leaps, especially, you look at, models like Opus and the latest GPT models, they are reaching levels of autonomy where people are actually finding that they actually can just be hands-off. And people who were once debating, “Oh, do I need to be in the weeds with my model in the IDE? Can I just completely move it off into the cloud?” That's a more serious conversation, and we've seen that in all of our growth charts. Internally there's this funny graph where our usage has, of PRs, our merged PRs, has grown 7X since I forget what it was called.Swyx [00:02:57]: I think Dev, maybe tweeted that. Yes.Walden [00:03:01]: it grew like 7X over, the last, I think it was, two months, three months, something like that. And then you see our engineering headcount growth. It's, gone up by, 10% or something.Swyx [00:03:11]: We were, we were afraid To release this. So this is Devin commit percentages on all Devin repos, was 16% in January and now 80% in March.Walden [00:03:25]: It's a big shift right now. And so it makes sense that a lot of people are now thinking about, buying Devin, but also maybe, trying to build their own and there's Lots of I have a lot of fun building Devin, so I can see why other people would want to build their own cloud agents as well. Matt, well, maybe it's good to hear, what initially inspired you to try to build OpenInspect?OpenInspect: Ramp, Cloud Agents, and Open SourceCole [00:03:49]: OpenInspect came about, through primarily my clients observing how they were using tools like Claude, OpenAI's Codex at the time, and seeing some of the friction that they were having with it. Primarily the Claude was being used through Slack, and a big issue they ran into was that the sessions that were launched were specific to whoever called it via Slack. And so if a PM was the one who invoked the session and they would then go to pass context to engineering can't see the session. And that in itself was a deal breaker because the PM, “Hey, engineering, can you jump in?” But there's nothing to jump in on unless they're copy-pasting out or the single response that came back. And so seeing some of these problems, I had built a similar architecture internally, just to experiment with, test out different ideas as this trend of moving off of localhost was starting to become, And as Ramp released their blog post, I had a lot of the pieces for this already in place, and just thought it would be funny to, see what Claude could do just purely from the blog post. And on my X account, there's actually a thread of where I live tweeted, going through thisCole [00:05:14]: comparing GPT and Claude as both of them are going through it.Swyx [00:05:17]: On the announcement thing or something else?Cole [00:05:19]: right after it got released. We can put it in the show notes. Yeah, it was helpful that I had already knew how to verify the system. I knew what I was looking for. I think Ramp did a great job of really illustrating, the technical aspects of how to build something. It was much more than just like, “Hey, we built a great system.” It was, “And here's how you can build it too.” And so, I resonated a lot with that, just with the problems that I was already seeing, and I thought that, looking around, I didn't really see anything in the open source community that, met this type of system. I think there's a lot that run, in localhost like Superset, Conductor, and many others.But nothing that was actually running in the cloud. And so, I built it, and I thought it was interesting to just open source it and allow anyone to then have a foundation that they can mix and match on top of.The Business of Background Agents: Open Source vs. DevinSwyx [00:06:16]: So literally after Devin was launched was, there was OpenDevin Which became All Hands. I don't know if you tried that orWalden [00:06:22]: I was going to say, one of the things that interested me a lot with OpenInspect was, you didn't try to go make it then something you monetize. There are a lot of, I think, these open source projects would then go and really try to, raise VSwyx [00:06:36]: That's why no OpenDevin. Yeah.Walden [00:06:38]: yeah, and how did you think about that? I thought that was very interesting.Cole [00:06:44]: I thought, and just what I had seen across my clients, was that having a background agent system is going to become a critical infrastructure within their company. And so because of that, I think that I wanted to open source it so that they could fork it and put in whatever customization they wanted. To that question though, I get asked all, “Oh, are you going to raise? Are you going to turn this into a service?”Walden [00:07:08]: I'm sure you've gotten offers.Cole [00:07:09]: but primarily I don't want to do that for a few reasons. One, I think that I don't want to compete for, $20 a seat. I think that is just a really difficult business. I think it's very easy to copy the main pieces of it. Again, I built this fairly quickly. And I think because you are not owning, I guess, the entire stack, it's hard to monetize. You have money being made at the sandbox layer with Daytona, E2b, many other players. You have money being made at the model layer. And you sit in this weird in-between gray area where what are you actually selling? You're selling, I guess, the infrastructure. You're selling, the integrations maybe.Swyx [00:07:55]: let's ask the guy. What are you What are you selling?Walden [00:07:59]: Well, yeah, there's multiple layers to this in practice, and actually it's funny you mentioned the infrastructure, ‘cause when we got started building Devin as well, we had to go figure out how to make the infrastructure as well because,Swyx [00:08:10]: You had to build this two years before everyone else,?Swyx [00:08:15]: Including, the model sideWalden [00:08:17]: It was not, it was not very polished at the start, when we just built it off of raw VMs from cloud providers like EC2, the boot up time was so slow, I think, And especially then, turning off the machines, saving them, and then to be able to bring them back up again when the, when you want Devin to wake up again later. It would just be out cold for like 10 minutes because that's just how long these systems took. They were not built for this repeated down and up usage. And so we actually had to go do all of that. And as a result now, one thing we offer when we go and sell Devin to people is, you don't have to worry about all the compute side of things. We'll make it work. We'll make it work in your cloud if you want it to. But aside from the product, and I want to go into the agents and the tuning of the intelligence part later, but I think a big part of what we do at Cognition as well is to just make sure that your company learns and uses and adopts these coding agents. ‘Cause I think for especially the largest enterprises in the world, you find that there is a lot of people who want to move over to using AI for their day-to-day workloads. But because of the way projects are planned, because, not everyone is literate in using AI in these ways, having a team of engineers who can actually go in and onboard you, set up all the integrations you need, the automations you need to really get to that level of, leverage with AI, is super helpful. And so We do that. We show thought partners to the customers that we work with as well.Swyx [00:09:56]: So let's talk about, architectural stuff. I think that's always, that is something that was the topic of conversation between the two of you. Is this, the mental model that you want to start with or something else? I'll just leave the floor open to you guys.Agent Architecture: Harness in the Box vs. Out of the BoxCole [00:10:11]: I think, maybe we can start here as just a general what are the pieces of a background agent system. And then maybe we can go into some of the nuances of, Decisions that you can make.Swyx [00:10:22]: But I guess I also Like, what, maybe what Walden is saying is the agent is like in this open code box, I guess. Right? This is infra, and then there's, that's the agent. And you had this discussion about whether you put the agent in here or in Out externally. Can you tease that out?Cole [00:10:39]: In a background agent systems, you have a decision to make of where the agent is actually going to run. This is typically described as the harness in the box or out of the box. With running the agent in the box, you're making some trade-offs by doing that. The negative trade-off you're making is primarily security. Because the agent is running in that box, unless you otherwise design it, all of your secrets need to go into that box as well. And given the nature of AI, it can be unpredictable, and you could very easily end up accidentally exfilling your secrets, or other unintended behavior. Now, the out of the box is the idea that we are going to have the actual agent running not directly in the sandbox, and we will have, quote-unquote, the brain of the agent running in some type of worker, control plane. That sandbox then is going to serve as the hands where the brain is basically operating and making tool calls into that environment to manipulate it. I guess other trade-off that you're making between the two systems is that, in my opinion, running it out of the box is much more complex because, you have state that has to be managed, whereas if you're running it in the box, all of the state of that agent is actually in the box, and yes, it's you could persist it elsewhere, but it's all localized and you have less concerns to worry about.Walden [00:12:08]: I think a lot of that, what you mentioned, is why we actually from the start built Devin to what we called separate the brain from the machine. The other thing that this allows you to do is reuse any existing infrastructure you have for dev boxes Perhaps. And so you don't have to worry as much about making a new type of dev box that has all the dependencies the brain needs, as you mentioned, the secrets the brain needs as well. One thing that we've seen some customers run into is, you have a GitHub app and you want Devin, your agent, whatever, be able to interact with GitHub through this application, but then you have different users with different actual permissions. If they are all interacting through the same GitHub app and there's no actual, separation between the system that decides, what it does and the actual secrets on the machine, then you run into an issue where, okay, it's hard to do the separation. But in practice, with Devin, it's much easier because we just say whatever you put on the machine, that is, the scope of basically what the user is free to do, what the agent is free to do. So only put the most scoped secrets on that machine, and then the brain is fully not accessible from the machine. So you don't have to worry about messing with the, any of the most secure parts of the brain if the user is free to do whatever they want with the machine.Swyx [00:13:31]: I was going to just bring, I have this, chart from OpenAI, where I don't know if this is, in the box, out of the box. That is something that they do use to describe it. And then also recently Anthropic did, managed agentsSwyx [00:13:44]: Which is, this is their thing. I don't know. It's all, it's all variations of the same pattern, right?Cole [00:13:49]: So this would be out of the box.Swyx [00:13:51]: Which, is preferable for them because it's less work?Cole [00:13:56]: I would say it's more work.Swyx [00:13:58]: It's more work?Cole [00:13:58]: But it, in my opinion, it is the better architecture of the two. It's just, you're taking on a bit of complexity by doing that.Repo Setup, Docker, and VM-Based Development EnvironmentsWalden [00:14:07]: One thing I've not seen a lot of other players do well is how do you manage what's actually on the box? And this can be complex for many reasons. Let's say you have a big repository that's changing and updating a lot with changing dependencies. How do you make sure that the working environment of the agent actually stays up to date, has all the credentials it needs to, let's say, run the app and test it, and all the things you want your autonomousSwyx [00:14:34]: So a repo setup.Walden [00:14:35]: Exactly. So in, internally At Cognition, we call this repo setup.Cole [00:14:39]: The hardest part ofWalden [00:14:40]: It's been a perennial problem since the start of the company, of how do we help people get this set up? Because not everyone just has, working cloud environments working out of the box. And do you find this to be a common problem withSwyx [00:14:53]: How do you solve it?Walden [00:14:53]: Your clients?Cole [00:14:54]: This is a very common problem, and through my consulting, this is a lot of what I help teams do. A lot of teams don't really have great developer environment setups, if any. A lot of the times it's, “Go talk to Bob and get the secrets,” and that obviously doesn't work when the agent needs to actually set this up. And so a lot of that, most teams are using Docker Compose or some type of microservices. And so for theSwyx [00:15:19]: Even in prod?Cole [00:15:20]: Not in prod. With the OpenInspect, you are using this primarily to interact, and make code changes. There is other use cases, but you can hook, whether through CLI, MCPs, other tools, you can then hook that into your production systems primarily for, SRE type use cases. But you are not, necessarily, trying to test your prod internal microservice through the system.Walden [00:15:48]: And you mentioned Docker Compose. I think one direction we saw some of our friends take early on was, using Docker containers as the level of abstraction for their models. There's lots of reasons, I think, why Docker containers are not great. One thing is, Docker container's not really a true security boundary, for one. But the other is, if you are running real applications, a lot of times those applications use Docker, and then you have to think about Docker in Docker, which is, really weird. And so I think part of, the really hard challenge of getting VMs to work, why did we do that? Well, it was because we realized that you actually needed, full VMs to be able to do these types of things. And especially nowadays where there's actually value in running the application and clicking around and sending you screen recordings of these things. The value just, keeps adding on top of that. But it is a decision I see people run into when they try to build their own systems, is, “Oh, do we, in addition to this, do we put the agent in the machine or out of the machine? Do we use Docker? Do we use something else?” What do you recommend people nowadays?Cole [00:16:57]: I think Docker is a good solution for maybe not running the agent, but running your infrastructure, because that is more or less the same setup your engineers are probably already using. If they're not, then I don't know what they're using. But they're probably already using Docker Compose.Swyx [00:17:14]: I've always had a small candle for web containers. I don't know if you guys have tried them before.Swyx [00:17:19]: To me, they were, supposed to be like Docker Light.Cole [00:17:22]: Is it?Swyx [00:17:22]: I don't know.Cole [00:17:22]: No, I haven't tried it. But yeah, I think any environment that you've set up that is a good experience for your developer naturally lends itself to being easy to set up for the agent. And once you figure out that local developer story, you've more or less solved the agent in a sandbox, environment setup. OpenInspect does have hooks as well, where you can, run a setup SH script that will pre-install everything. You can then pre-snapshot that build so it starts instantly, and then there is a second hook to actually then, restore the state of the sandbox when it comes back. And so you can already have all of those microservices running and basically get the same experience that you would on your machine within the sandbox.Testing Agents: Computer Use, Screenshots, and Real App WorkflowsWalden [00:18:08]: Another thing that we've been thinking a lot about is like Different VM service offerings. Have you had customers where they needed like macOS specific VMs or like Windows specificWalden [00:18:20]: VMs?Walden [00:18:22]: There are like many technologies in the world that only work on specific types of machines, right? If you're building a.NET application that has to run on Windows or like, maybe more commonly if you want to build iOS or macOS Does that workSwyx [00:18:32]: Does Commission supportSwyx [00:18:33]: Choices like that?Walden [00:18:35]: The fundamental architecture we do, because we do the separation, it does support, but the actual work in progress is happening right now on these. Another thing that we've actually recently added support now for, it's in beta, is doing Android development. To do that, we needed to support, I think, nested virtualization within our machines because the VM itself is like a, is a virtualized Firecracker instance, and then you had to then run another Android emulator inside. And there's like weird performance issues that like, it, which is why it's like still in beta. We have to think through these problems, but it unlocks a lot for anyone who wants to do Android development.Swyx [00:19:13]: I was trying to find like a reference video for the testing thing. I couldn't find it, but I think you worked on the testing, capability. Why call it testing and not like computer use or I don't know, it's, what's the general Category of problem?Walden [00:19:26]: I think that when people think about the ability of an AI to run your app and test it, I think they actually over-index on the computer use part of it because computer use in my mind is the literal, okay, you want what button you want to click. Can you emit the right coordinates to go click that button? I think testing is actually a really interesting likeWalden [00:19:48]: Problem-solving, challenge for these AIs because if you wanted to do arbitrary testing, imagine you make a change that spans the frontend and the backend, maybe, even some other like even more deeply nested service. To actually test that change, we have to reason through what-- how do you first run these applications to orchestrate with each other with the right version of the code? Then, okay, how do I trigger the feature or how do I make the thing actually happen? And this can get arbitrarily hard, maybe you have to be an admin. Maybe a certain thing has to be feature flagged on. Maybe, you have to like run two sessions and then send us a very specific word into one of them to trigger a specific behavior. And figuring out how do you do that requires a lot of code base context, requires, a lot of orchestration that we've specifically done. And in some cases, we found that you actually, no one frontier model can actually do this full end-to-end task itself.Walden [00:20:42]: We've seen cases where we actually had to orchestrate different frontier models together to solve this problem together. That is where we spend most of our time when we think about this testing problem, not so much the computer use part. Computer use for what it's worth has gotten a lot better with recent models and it's made that part of the job certainly easier.Swyx [00:20:58]: Especially with like even 4.7, that they released yesterday, apparently like way better in terms of the vision stuff, which is going to be encompassing computer use.Walden [00:21:08]: Having evals for all these as well is something that like takes a while to build up. And having the evals be right is tricky as well. Do you ever see like, clients who are building their own agents have to start standing up evals to make sure things don't regress?Swyx [00:21:25]: Not so much evals in the traditional sense, but specific to the testing part that has just gone in. I just added support for screenshots And in theory you can also do video. I need to put in a plugin to do that. But they do show up natively, and it was a very heavily requested feature, especially after Cursor's recording came out. I think that was very enlightening for everyone of like, “Oh, this is a very good feature to actually have.”, I think with Devin you guys have had this for a while.Swyx [00:21:57]: Oh, yeah. See how screenshots work. Yeah, I don't know if there's anything, super and not obvious. It's like once what feature to build, you can just prompt it and it Will mostly work.Walden [00:22:09]: I think to Walden's point, though, the computer use is a subset of the larger testing problem, and I think that's very specific to the code base that you're working and it's not something that, out of the box that you could just solve it. The-- you do need the code base context to actually know how to test it. And I think in the case of a background agent system, you fortunately do have that code base locally that what is changing and could then inspect it and use that to drive the model.Swyx [00:22:40]: For those who haven't seen it before, this is an example of how it works. You, after the PR is done, you click testing approved, and then it sends you back a video. What I really like is that it labels, It's very small here, but it actually labels what it's testing. And then it-- and then you actually see the cursor and everything. So I don't know, yeah, the engineering in this, just Whatever you want to show. ‘cause this is like, this is one of those like, oh, few of the AGI moments, right? ‘cause Once I look at this, I actually don't I wish I can just merge inside Of Slack instead of going to GitHub ‘cause I don't need to see the code. I know it works.Walden [00:23:19]: Maybe a new feature in Cursor. Yeah, the annotations at the bottom was also a big difference for me when I, when I added those.Swyx [00:23:27]: It's just like, what am I looking at? What are you trying to demonstrate?Walden [00:23:30]: Exactly. There's a surprisingly long tail of small details that ends up making a big difference for this end metric of like how fast do you actually merge the code in. One experience that we spent a lot of time tuning early on was what is the right experience on GitHub for these tools. Because I think, most tools out there when you build the agent, you'll think about, oh, it'll create the PR for you. We try to take that a step further and say, “Oh, what if we actually made sure you could interact Devin, with direct Devin directly on GitHub?” And so we made sure that you can comment on GitHub, and Devin would actually receive those comments and address them back. But there's actually quite a bit of tuning you have to do here because you can imagine that actually like-We recently have Devin Review, for example. Devin Review will post comments on his own PR And then Devin has to then goGitHub Workflows: Devin Review, Comments, and PR AutomationSwyx [00:24:23]: He answers his own comments, which is Really loopy. So like, yeah, I like that it just updates here that it's, that I have commented But usually it's just me saying like, “Hey, merged, fix any merge conflicts.”Walden [00:24:37]: The, so when Devin fixes his own comments, you might be scared that, oh, maybe I'll infinite loop. But we've put a lot of work into making sure it doesn't, both by making sure that the comments are high signal, but also that the agent is thoughtful about what comments it immediately goes and tries to fix, and what comments it's like, “Wait a second, I think you're wrong.” Actually, that's one of my favorite moments is when Devin tells me that I'm wrong, when I try to get it to do something different. But tuning that behavior, actually makes a big difference in terms of how useful the actual GitHub experience is.Cole [00:25:06]: I think to touch on that as well, I think having the AI reviewer integrated into the system is a critical part of this background system. OpenInspect does have that. It has a GitHub code reviewer that you can control the prompt. It does do comments as well. It doesn't do them automatically yet. The capability is there, but it's not fully used.Swyx [00:25:27]: So you have to ask for it?Cole [00:25:28]: you do, yeah. You can tag it on GitHub, and then whatever you named your, GitHub bot, it will then follow up on it. It will then, if you have merge conflicts or whatever you have asked it to resolve, it will then resolve it, but it doesn't do it automatically yet.Integrations: Slack, MCP, and First-Party Agent InterfacesWalden [00:25:42]: Well, I'm curious, what is, the most common thing that people end up requesting, that they still need on top of OpenInspect when you help them go implement it?Cole [00:25:52]: I think a lot of it comes down to actually integrating it into the company. It's one thing to have the background agent system set up, but if it isn't actually integrated into your larger ecosystem, it isn't that useful. It is useful to be able to kick off sessions, but what we really want to be able to do is hook it into all of our other systems, whether that is the production database with read-only credentials, the logs, a Confluence or internal knowledge-based system. I think that is where I see the huge leap for companies, and that can be a challenge for companies as well who are maybe not familiar with exactly how to approach it, especially if they're in environments that have more compliance type things where, access control can be pretty big and how do you deliberately think about these problems, I find to be, one of the problems that comes with a system like this.Walden [00:26:46]: The thing we found is So, MCPs, obviously it has been like this, really big explosion of, oh, you can go, integrate it with all these different things. But to actually get the integration right and the and get the right experience, oftentimes we found that we had to go build our own ad hoc things. I think Slack is a great example of this. You could give your agent a Slack MCP and okay, it can post messages back to you on Slack. But we actually use Devin like a coworker in Slack, and that's how it's been built from the ground up. But to do that, you actually need to, support webhooks that come back, right? And then Devin has to respond in a natural way and then hopefully don't spam your threads too much and annoy the people in your company. So you got to tune that experience just right. Especially when there's a lot of back and forths, we find that we actually have to go beyond the simple MCP integrations in these places.Swyx [00:27:39]: I just pulled up the MCP marketplace. I know this is a Fair amount of work. Is the answer to eventually take first party control of all the top MCPs? Is that theWalden [00:27:48]: I would love a world where you could have something that's more expressive than MCP. That, goes both ways, not just a set of tools, but a proper system that interacts back and lets it Have the right experience with all these interfaces.Swyx [00:28:03]: So there actually is sampling in the MCP spec, but nobody Uses it, right?Walden [00:28:07]: And so I think that's the other part is, actually we found that when the MCP spec starts to get too complicated, it starts to lose its original promise of Being like a simple one-step connect. Now then we have to go figure out how to support all these different variations of things and It starts to look a lot like just building the first party integrations in a lot of these cases now.Cole [00:28:29]: I think it matters, too, how critical it is to your company, right? If this is something that nearly every session is going through, it probably makes sense to own it so that you can make optimizations on top of it Versus just whatever is off the shelf.Swyx [00:28:43]: Awesome. Other than MCPs, what else, sorry, well, I don't know if that's Narrowing in too much on, integrations. But what else? What other elements of building OpenInspect or Devin that you guys really sink on?Memory and Knowledge: What Agents Should RememberCole [00:28:59]: I think, a problem that comes up very frequently is this idea of memories or knowledge base.Swyx [00:29:05]: Oh, boy. How do you solve it?Cole [00:29:08]: so not solved yet, is the short answer.Cole [00:29:11]: it's something, there's a open issue for it, someone asking about it.Swyx [00:29:16]: There's, I, D Wiki hasn't indexed anything about memory yet.Cole [00:29:20]: how I'm seeing it solved across my clients is primarily through skills. I find that skills can be a good gap within that or updating Claude MD, but I think memory as a whole is a pretty unsolved problem, and it is why I've been hesitant to add it. I think there is parts of memory and that can be addressed, but I think as a whole it's a very difficult retrieval problem.Swyx [00:29:44]: Oh my God. RAMP didn't write anything about memory? I see zero search results.Walden [00:29:50]: No. Memory can be quite tricky to get right because it's the retrieval, but also the generation of the memories that can be really tricky. You don't want it to just like Remember very specific details.Swyx [00:29:59]: Walk us through the Devin memory journey because I know there's been a journey.Walden [00:30:03]: the first version of memory that like stuck around for a while was A system we have called Knowledge. And the idea was we wanted it to pick up things over time and not need the user to be proactive about teaching Devin things. So, okay, any time you remind Devin, “Wait, no, that's not quite the way you're supposed to use Git”Like, we actually want Devin to say, “Hey, do you want me to actually just remember this for the future?” And for you to just basically quickly approve or reject and for it to build up over time. ‘Cause I find that, 95%, I think, or some crazy stat like that of the memories that Devin has are all through these auto-generated things. Very few people actually just want to sit down and write big docs on Here's how you're supposed to work with the technology, et cetera. The generation and the retrieval has been something that we've been trying to tune a lot over the years. Generation, you don't want it to remember something like, if you asked one time to like, “Oh, please open as a draft PR,” you don't want to be like, “Oh, everyone forever now should get their PRs as draft PRs.” But you do want some, conveyor. Maybe you want to say like, “Oh, Cole generally likes, things to be created as draft PRs.” Same with retrieval, if you have thousands of these memories, how do you actually make sure they're retrieved at the right time? And that can be quite tricky to do right without exploding the context with a bunch of useful yeah, useless information. Surprising amount of just, eval work to just make sure that, memory is, remains a reliable system as new models come and go.Cole [00:31:31]: Do you have anything that you could share on, memory pruning? And like the temporal aspect of memory?Swyx [00:31:36]: Deleting and forgetting?Walden [00:31:39]: The, today, the, So the things they could do is it could edit memories. And so if your memory used to say like, “Oh, Cole likes to open everything as like a draft PR,” then you can imagine, “No, don't do that.” And then it'll say, “Oh, do you want me to update the memory to be Cole now want everything as, open PRs?” I think that at the same time we don't know if this is going to be the final version of the system. Whatever we have here will probably, translate into the new system that we'll be coming up with. But I think one big difference between two years ago and today is these agents are really good at using anything that resembles a file system natively. And so part of us are, is thinking, “Oh, should we rebuild memories to feel more like a file system that we let the agent navigate on its own?” That's been an interesting exploration. Also similar ideas in the scale space.Swyx [00:32:35]: I am pulling up OpenClaude's memory thing right now. So memory, OpenClaude has like this like daily memory journal thing, right? And you can I mean, that is a file system you can grep through and is a source of truth. I don't know if it's the best. It's probably super noisy, but at least, if you lose something you can discover it or you can apply some, forgetting algorithm to, more ancient memories that don't get recalled again or something. I don't know.Walden [00:33:01]: One thing we've been trying to do to push the boundaries of how you use agents at your company is letting an agent basically have a very similar file, a memory.md or something, and just like be your permanent PM for a specific set of issues maybe. So we have like some Slack channels internally, maybe a Slack channel dedicated to, a specific product like DeepWiki maybe. And you can imagine that, or you want a Devin that never stops, it's just always awake, but it has this like memory dock that it can just maintain for itself about, okay, what are like the number one priorities of what we have to fix and prioritize? Who is responsible for some upcoming work? Maybe they'll even Devin will even tag you on some recurring basis. And so it's been an interesting move to see, okay, how can we actually use Devin for more than just engineering? Can we actually upstream above the engineering process and maybe it's just Devin creating tickets, which then maybe some humans do, but then maybe other Devins do.Swyx [00:34:00]: One of my more fun automations is go research competitors and just suggest stuff to me on a weekly basis. That's the automation. I can't find it right now, but basically it just like, “Look at competitors and suggest things.” “And here are three things that you've suggested that I don't want any more of,” and you just stick that in the prompts. But like I wish actually So for like when I, for example, when I reject a PR, I wish that it updated memory so that I can then just not have to go up, go back and update the scheduled, sync, but anyway, feature request.Walden [00:34:31]: what? We might change it soon. I guess OpenInspect, in the time you've been around, has there been anything you tried to implement but then you had to like undo and like do a different way?OpenInspect Architecture: Webhooks, Control Planes, and Agent StateCole [00:34:41]: Nothing yet, but something that is on my mind. The initial way that I built it was that each of the integrations lives as its own package. And so you have The Slack bot, which is what's handling the webhooks, and then is basically interacting with the control plane. As I'm seeing the system starting to be more integrated, specifically with the GitHub bot integration, I'm considering bringing that all into the central control plane because especially now I want to start, And a request that I'm getting is the ability to monitor, the actual, pull requests being merged, as well as just tracking ofSwyx [00:35:19]: What do I have open?Cole [00:35:21]: What do I have open? How many of these are getting merged? How many comments are showing up? To just understand the health of the system. And so in the case of a GitHub app, you only have one webhook. And so then it's a question of do I put that webhook in that GitHub bot package? That's weird. It doesn't really make sense to live there because that package is more for like the code reviewer. Or do I like centralize it? So that's something that's on my mind of, making that decision. I think the other one we touched on earlier is the harness in the box versus out of the box. I think long term the architecture will eventually come back out of the box. Some of the newer tools that I've added are calling back into the control plane so that you don't have the secrets in the sandbox. And so I think long term I probably will pull the actual, agent out of the box, but I think for now it's fine.Subagents and Multi-Agent Systems: When Parallelism Helps or HurtsSwyx [00:36:16]: Just, a quick question on pulling the agent out of the box. I'm One thing I'm very bullish on this year is agents calling other agents or spawning sub-agents or Whatever you want to call it. Does that make it harder or easier? I can't tell. Because if the harness is in the box, you can just spin up more boxes. If the harness is outside the box, then you're, it's less easy because you are, you have a unicorn pet of a, of a harness that's, living outside the box.Cole [00:36:45]: In theory it would be the same way, right? Whether, one agent has launched many, sub-sessions within it, OpenInspect, for example, can launch sub-sessions and actually create other environments and then monitor them. In the case where it is out of the box, that would basically just be an additional session that's running. And so that session is also running outside of the box. It's running in your worker plane, wherever you're running this. And then you really just have to think about how does your top level agent then interact with it. I do think it can be more complex, just ‘cause again, you have now a more difficult architecture. But I think if you figured it out once, it's probably fine.Swyx [00:37:26]: Well, then I'm just, throwing it open to you in terms of, I call this like meta Devin management. Which is like the, Devin's calling Devins or Devin scheduling Devins or querying trajectories or anything like that. What have you built or unshipped, anything?Cole [00:37:46]: I think one of the surprising things we've seen is that a lot of the ways that, these, separate agents work with each other, and you want them to, parallelize their work, has still mostly followed the same manager sub-agents regime. And a lot of people I think are excited about this world where you have swarms of agents that, talk with each other all over the place. We've actually given Devin an MCP so they can just go arbitrarily message other Devins And create new Devins, et cetera. But I guess, it somehow creates, a really chaotic world in that sense. And so we've still found that most practical use on a day-to-day basis has been one single Devin.Cole [00:38:33]: Figuring out how to segregate the work and get, have other Devins work on it in, a relatively isolated sense, each with their own boxes Not sharing machines, so there's, a very little room for conflict is the regime that you have to create today.Swyx [00:38:50]: I'll call out, the experiments from Cursor, right? This is Wilson Lin's work on Single agent to multi-agent, and you're obviously famously on the side of don't build multi-agent. But they went through the whole thing, only to arrive at, this Which is exactly what Devin has, I think.Cole [00:39:08]: I think there will be a revision to that post at some point AboutSwyx [00:39:12]: Tell us about itCole [00:39:12]: I think multi-agents were very much not at all possible a year ago. You do see more multi-agent experiments today, but you can argue, are they really multi-agents, or are they just just, tool calls,? There are people who, will create sub-agents to go look for XYZ file, XYZ implementation. Has really nice context management benefits because all of the tool calls and tokens that it spends then get collapsed back to just the answer for the main agent. There's a lot of benefits to doing this. We basically have Devin do this with Deep Bookie, make a call out to Deep Bookie, give you back the results, but that feels like a tool call,? It's not like these, two collaborators actually talking back with each, back and forth with each other. But I think the thing that gives me the most bullishness that multi-agents might actually be possible is actually what I said earlier about Devin will actually sometimes tell me I'm wrong and push back, and I think that demonstrates a level of maturity and communication today that makes a multi-agent world possible. One, can two agents who have seen different information come back to each other and actually figure out who is right, what is the correct implementation? They're not just, yes men. Claude, I guess is like, used to just say, what is it? “You're right,” or,Swyx [00:40:25]: “You're absolutely right.”Cole [00:40:26]: “You're absolutely right.” Yeah.Swyx [00:40:28]: The Have you seen, did you seeCole [00:40:29]: The age is overSwyx [00:40:30]: The Codex app troll in Topic? This is the Codex app. Inside of Settings, there's a little, there's a little Easter egg, right? So if you go to, the Themes or Appearance, right? There's all these, color codes, and the top is absolutely, and it's the Topic's colors. Which is such a troll. Anyway.Model Behavior: Pushback, Adversarial Prompts, and Agent SkepticismCole [00:40:53]: I love that Easter egg. Did you discover that yourself?Swyx [00:40:54]: No, it was, someone was, tweeting about it And I was like, I was like, “Is this true?” Because, sometimes people just tweet stuff to, get a rise out of you. But yeah, there you go, in Topic colors.Cole [00:41:06]: Yeah. So yeah, we're out of this regime where, it just says you're absolutely right, and they can have real conversations and real back and forths.Swyx [00:41:13]: You can prompt it as well to be more adversarial or whatever. Yeah. Okay. Yeah, that, I mean, to me, that is more intelligence, right? That is not just something that's, a dumb tool, it's actually pushing back on you I think. Yeah.Cole [00:41:24]: when you mentioned, of course, the blog posts. There was one blog they had where they fed a swarm of agents together and built a browser.Swyx [00:41:34]: That was I think that was the one.Cole [00:41:36]: You can have, likeSwyx [00:41:37]: I think it's the same oneCole [00:41:37]: Creation of it. We found a surprising success of, don't do a swarm or anything, just have one Devin, it does its own context management. Just let it keep running for a while and give it some crazy tasks. I think we asked it to, rebuild, a Windows OS system. And it managed to do it just like, going on for long enough. It'sSwyx [00:41:55]: Was this Andrew's thing?Cole [00:41:58]: there were lots of demos that we ended up not posting, ‘cause at some point we'd just be posting way too much a bunch of, Demos. But I love that because it shows that I think the multi-agent thing still has, a bit of exciting sexiness to it, which is maybe still beyond still, the actual delta it adds to the capabilities of these systems. But it's absolutely the future. I think we're heading in that direction and we can see the progress being made there already.Swyx [00:42:25]: If I were to, make one super minor pushback because I don't feel that confident about it yetCole [00:42:33]: Go for itSwyx [00:42:33]: But I've had Ryan Lopopolo from OpenAI on the pod And he's a super slop cannon, right? Oh my God, that's my coding agent being done. I downloaded this, Peon Ping. I don't know if you guys have heard this. It takes like-, sound packs from popular games like, Command and Conquer and Warcraft, and then it plays it whenever it's done. And so it's like, “Work,” or whatever, “At your command,” or something. Anyway, what I got from the Cursor code base and from Ryan's thing was that there's a slop cannon approach where you try to loosen the single agent's, bottleneck, and I feel like that is, probably an, a very important thing to try to figure out. I don't think anyone's, really solved it. Because then you just have more reviewer slop on top of the agent slop To try to wrangle it all. Ryan will probably very strongly object that I say that he hasn't solved it, but he thinks he's He thinks he's completely solved it. But I think it's still I think it's, very important, ‘cause, that is a bottleneck, right? I feel Devin is slow sometimes Because I'm like, well, yeah, this is very readable and very sensible, but also it is slower than it could be if I just, I want a button to just say, “Just ramp this up 1,000 next parallel, in parallel and just, see what happens,”? And I don't know if that's, feasible at some point in the future.Code Review, Entropy, and AI SlopWalden [00:43:55]: I And we've also run experiments internally where we've basically tried to build entire products, true products that we knew we would eventually ship, but for now, let's try to see if we can do it just by purely, vibe coding on top of each other, auto merge, no code review at all. And then there's this benchmark of how many weeks can you go onto this for Before you say, “We have the trashiest code base.”Walden [00:44:18]: “Let's actually rewrite it from scratch.”Swyx [00:44:19]: Start a new factory, yeah. What'd you find?Walden [00:44:21]: I think we found that the state-of-the-art in December was you can probably, run this for about two weeks. By the end of those two weeks, you'd find that, hey, you want to, change the color of a button. Well, it turns out this button is implemented in, 10 different places, and they, have All these different variations, and oh, you forgot one of them, and actually it's a slightly different color in one spot. And you're like, “Okay, this is too much to work with. Let's actually try to do code review at the same time.” And make sure that we're on top of our software, actually cleaning it up a bit And making sure it's done in a scalable way.Cole [00:44:54]: I think building on that, the idea of, you don't have to look at code, I think is generally a bad idea. And the meme that I have for thatWalden [00:45:03]: What timeline, all right, is Do you think that statement will be true on?Cole [00:45:06]: I think probably for a while it'll be true that you should continue to look at your code. A problem that I see a lot of teams run into that I work with who are embracing AI native, AI first coding, is The meme that I have is that your code base regresses to your worst engineer, because that engineer who is, very gung-ho about AI and is not auditing their code, their pattern starts cementing into the code, and now the AI is referencing their patterns. And so now their if/else block that, is 20 if/elses back and forth, the AI is seeing that as the pattern of how things are done and starts to then exponentially grow this slop. And I find to your point, a pretty good approach to that is having scheduled cleanup, whether by humans or through systems, that are looking for duplication. They then address that. You'll end up with like 12 helpers for how to format a date. And you need to address that, because otherwise it will continue to sprawl.Swyx [00:46:09]: Within balance, I think it's fine to have some duplication, and then sometimes To have garbage collection, right? Yeah. The What I've been, talking about with a lot of engineering leaders is that you want to be very strict about the boundaries between modules, and it's your job as an architect, as a CTO, whatever, to say like, “Okay, here's the hard contract between you guys and you guys. Whatever you do inside this black box is your business. You do whatever. But between these guys, let's be, really damn clear, and any movement must be signed off by a human or me,” or. Then, and like that's that. I don't know if you have any other modifications or advice.Walden [00:46:44]: Well, I guess generally on the topic of, where humans can be useful, I found that ‘cause, some of these, really deep infra problems, sometimes just having a human that just has, really deep expertise can make a big difference. I've actually seen this come into play when actually building agents. So we've had a few friends now, try building their own coding agents, and I think one same problem that I recurringly heard a lot of them run into was this problem of like, “Oh, Grep is really slow on our agents' machines.” And so a lot of them, I assume because they're using AI and they themselves don't have, super deep infra background knowledge, say, “Okay, we're going to go build our own custom Grep index. It's going to be really fast,” and use that as a way around this problem. When we ran into this problem About like, maybe like a year and a half ago when we were, in the early days of building Devin, we obviously didn't have AI then. We just asked our, how to, how to do this. You can just swap out a new Grep index, so.Infrastructure Details: Grep, File Systems, and SandboxesSwyx [00:47:45]: What do you mean you hand-coded Devin? What?Walden [00:47:48]: It's like, can you believe we hand-wrote this code? And we had, our infra people who are really amazing, they were looking into it and they're like, “Oh, what? We realized that actually the root cause of this problem is actually super simple, but like fine-grain detail,” which is that a lot of these virtual machines actually underlying them don't use real file systems. They use these, network file systems where things are actually cached over the network actually in S3. So when you're Grepping, you're actually making network calls Every time you're doing these things, and that's why Grep is extremely slow on these machines. And so again, goes back to, what is all of the crazy infra work that we had to do to actually get these machines working. If you try to do this yourself, there are tons of small details like this, and so we had to eventually go swap out that network file system. ButSwyx [00:48:35]: I think there's a write-up about it, right? Silas did one about the virtual file system.Walden [00:48:38]: Oh, that was a whole other thing. TheSwyx [00:48:39]: Oh, that's a different thingWalden [00:48:40]: The BlockDev file storage formatSwyx [00:48:42]: I'll bring it upWalden [00:48:42]: Which is, a file system format that we built so that the VMs could be spun up and down very quickly. Basically, the intuition behind this is-Imagine you have, a terabyte of disk, and your agent only, wrote, a hundred lines of code on top of that disk. How long does it, say, take to, save and re-bring up that disk? And most systems, because you're not optimizing for this case, it's just, on the order of a terabyte of work because you have to Save all of that and bring it back up. In our system, we try to build a file system that incrementally builds on top of each other. So every time you save and bring the machine back up, you're only doing work that is proportional to effectively the diff in the file system. And so this, shaves off a lot of time in the boot-up process of Devin. I think we This is actually now outdated. We have a newer system inside of Devin. But yeah, there's a lot of tiny details you have to get right here to actually get the day-to-day experience of Devin to be good.Swyx [00:49:39]: It's, not technically agents, but it is agent infra, and when you sell an agent as a company, you sell agent plus agent infra.Walden [00:49:46]: At least the way we do it be And the other The nice thing about having the agent infra being done together is, you We get to deploy Devin in whatever environment we want now. We don't need to wait for some underlying infra provider to also go and support VPC or on-prem or FedGovCloud, for instance. So we can actually go and figure out, okay, since we own the infrastructure, how can we get that set up for you?Cloud Providers: Modal, Daytona, and Enterprise SandboxesSwyx [00:50:12]: Whereas you're Cloudflare dependent.Cole [00:50:15]: so Cloudflare runs the control plane. The sandboxes, Modal is supported. A contributor just added Daytona. E2B is on the roadmap, and I think there's an abstraction in place that if any contributor wants to add a new provider, they can add that in.Walden [00:50:32]: Well, what are, How are the customers you work with Do they generally try to then go set up a contract with another one of these third-party providers? Do they try to do the VMs in-house?Cole [00:50:44]: most of them I see using Modal. I think Modal has a greatWalden [00:50:48]: Shout out Modal.Swyx [00:50:48]: Shout out Modal.Cole [00:50:50]: I think Modal has a great offering. It captures all of the sandbox pieces you need, snapshots being a pretty big piece of that, and given that they also offer GPUs, I think it's a pretty nice offering as a whole.Swyx [00:51:04]: no debate there.Walden [00:51:07]: Modal is great, especially, I think their container offering is, the most natural, and so especially if you are willing to, forego, the full VM requirements Modal is, a really vast place you can spin something up on.Swyx [00:51:20]: Is there a point So Modal's very Python, and I feel like most workload, has really shifted to JavaScript. I don't know if you guys Get the same feeling. So, okay, when I started Landspace and IE and all these things, I was like 50/50 Python and JS, right? That's roughly. I think that's wrong now. I think JS has won. I don't know if you guys Like, I Maybe I'm overstating it, and maybe for cognition, there's, C# and Java and what have you. But for, new greenfield apps, do you feel that Do you get that sense? Does it matter?Cole [00:51:52]: I think that most of the libraries that I see in this space are Python native first, especially in theCole [00:51:58]: Observability space. That said, I think that there is a pretty big appeal of having your entire system in one language. Especially when you have both your frontend and backend communicating, you can have one central type Which is very nice.Swyx [00:52:11]: That's my case against Modal, which is Then you have to run JS. You can run JS inside Modal. It's just, one extra step That, isn't native to the runtime. I don't know ifWalden [00:52:22]: I don't knowSwyx [00:52:23]: Reviews. Do you have numbers? I don't know.Walden [00:52:25]: the one thing I don't like about Python is whenever AI, whenever it writes Python, it always does, the weirdest patterns, andSwyx [00:52:32]: Oh, because it's, mixing two and three or what?Walden [00:52:34]: I think it's something mixing two and three, yeah. The I don't know if you see this. It always tries to do, has attribute on objects as likeCole [00:52:41]: Oh, my God.Walden [00:52:41]: But it's like But that you shouldn't be doing that. It should error if there wasSwyx [00:52:45]: Because it's training on library code?Cole [00:52:47]: I think it's more of, likeCole [00:52:48]: From what I've seen, it's more of, a reward hacking mechanism where it doesn't want to basicallyWalden [00:52:54]: It'll never error.Cole [00:52:54]: It doesn't want the code to fail. And so it Even when it knows it has the attribute, it'll call getattr on a, and for a lot of my clients who have moved towards more autonomous coding, we've put that in as a lint rule That if you do getattr, your pull request is going to fail.Slop Signatures: Comments, Backwards Compatibility, and TypesSwyx [00:53:12]: Ooh, this is a fun topic. Can you tell me more about this? What else is a sign of AI coding that you have to put guards in?Walden [00:53:21]: So we were talking just before this about Opus 4.7. One of the things this new model likes to do is it writes lots of comments. Not like, it'll, comment every line, but it'll write, paragraph, PRDs, on top of every function. But I will say, to its credit, these aren't slop, descriptions like they were before. “Oh, here's what this function does.” It's like, “Oh, here's actually the r
A verse by verse study through the book of 1 John with Pastor Kevin Edwards of Calvary Chapel Clayton, NC. https://www.calvaryclayton.com
SHABBAT DAY LESSON — LEVITICUS 24Teachers: Kerry & Karen BattleWHAT WE COVERLeviticus 24 reveals the difference between maintained holiness and gradual covenant decay.This chapter is not merely about lamps, bread, punishment, or judicial law.The chapter exposes:continual covenant consciousnesscontinual maintenancecontinual remembrancethe weight of speechpublic corruptioninward decaydesensitization toward holinessequal justice before YahuahThe Continual Lamp and Maintained IlluminationLeviticus 24:1–4The lamp was commanded to burn continually before Yahuah.The oil had to remain pure.The priests were responsible for maintaining the light continually.This section reveals that holiness requires continual maintenance.The chapter exposes a terrifying reality:People rarely drift into darkness suddenly.Usually:maintenance weakens firstreverence weakens nextcompromise spreads quietlythen corruption manifests publiclyThe Holy Bread and Continual RemembranceLeviticus 24:5–9The bread remained continually before Yahuah as a memorial.This section teaches:continual remembrancecontinual dependencecovenant awarenessdisciplined orderThe chapter reveals that people often collapse because continual remembrance weakens over time.Familiarity slowly destroys reverence.What people stop honoring continually, they eventually begin treating casually.The Blasphemer and the Exposure of Inward CorruptionLeviticus 24:10–16The blasphemer did not begin with outward speech first.The mouth exposed corruption already forming inwardly.This section reveals:hardened dishonorpublic corruptioninward decayconflict exposing formationdesensitization toward holinessPressure exposed what had already been maintained secretly in the heart.The chapter teaches that public corruption is often the final visible stage of inward compromise long tolerated privately.Equal Justice and Covenant OrderLeviticus 24:17–23The same law applied equally.The same accountability applied equally.This section reveals:judicial consistencyrestrained justicecovenant orderequal standards before YahuahHoliness without justice becomes hypocrisy.Justice without holiness becomes brutality.WHY THIS MESSAGE MATTERSLeviticus 24 exposes how covenant decay develops gradually inside individuals and communities.The chapter teaches:what is continually maintained reveals what is truly honoredtolerated compromise reshapes consciencefamiliarity weakens reverenceneglected maintenance spreads darkness quietlyspeech eventually exposes inward formationcommunities decay collectively when corruption becomes normalizedThis chapter destroys emotional religion.Holiness is not occasional emotion.Holiness requires continual maintenance before Yahuah.SCRIPTURE REFERENCESLeviticus 24Leviticus 24:1–4Leviticus 24:5–9Leviticus 24:10–16Leviticus 24:17–23Exodus 27:20–21Psalm 119:105Proverbs 6:23Deuteronomy 8:11–14Matthew 12:34–37Matthew 12:31–32James 3Isaiah 5:201 Corinthians 5Galatians 6:7ABOUT AHAVA ~ LOVE ASSEMBLYWe teach the pure Word of Yahuah. No religion. No traditions. No compromise.Teaching is established by Scripture only: line upon line, precept upon precept, with covenant understanding rooted in the Hebrew thought-world of the text.SUPPORT THE WORK — GIVE VIA ZELLEZelle QR available at: ahavaloveministry.comZelle only.FINAL WORDHoliness rarely collapses suddenly.It usually decays gradually through:neglected maintenanceweakened remembrancetolerated compromisefamiliaritydesensitization toward holinessWhat is continually maintained reveals what is truly honored before Yahuah.FINAL HEART CHECKWhat are you continually maintaining?Has reverence become continual, or merely emotional?Have you slowly become desensitized toward holiness?What emerges from your mouth during pressure?What does your speech reveal has been growing inwardly for years?Are you preserving holiness collectively, or silently tolerating corruption?
There is a major sin issue Paul addresses in 1 Corinthians 5 and he tells the church, who is boasting about this particular sin, that this man should be banned from their assembly, he should be turned over to Satan, and they should not even eat a meal with this guy! Sin is serious and is not a toy to be played with! We are to repent of our sin, which means turn away from it, and we must die to sin (Romans 6). When we give our lives to Christ, the old man is buried and we come out of that water a new creation, filled with he Holy Spirit. He begins to transform our lives into he image of Christ...if we let him. Do not choose sin over righteousness!
This episode of Own Your Power explores the life-changing impact of continual personal growth, challenging listeners to break free from stagnation, fear, and comfort to unlock their true potential. It emphasizes that real fulfillment and happiness come from progress, not perfection, and that growth requires intentional action, pushing through discomfort, and redefining limiting beliefs. By focusing on self-improvement, embracing challenges, and committing to becoming a better version of yourself every day, you can create the freedom, success, and extraordinary life you truly desire. Own your power with this Success Tip. For more about Rod and his real estate investing journey go to www.rodkhleif.com
Partakers of the Divine Nature2 Peter 1:1-11 Partakers of the Divine Nature: 2 Peter 2:1-11 Chris Moore Message Slides For the bulletin in PDF form, click here. Our Passive Role in Spiritual Growth (vv. 1-4) - Our Growth Begins with God's Initiative (vv. 1-3) - Our Growth Continues as we Hear & Believe (v. 4)Our Active Role in Spiritual Growth (vv. 5-11) - True Spiritual Growth is Comprehensive (vv. 5-7) - True Spiritual Growth is Continual (v. 8) - True Spiritual Growth is Essential (vv. 9-11)Discussion Questions1. Think about a time in your life when you experienced significant spiritual growth. What were some of the key elements that led to your growth?Our Passive Role in Spiritual Growth (2 Peter 1:1-4)2. What does it mean that God “has granted to us all things that pertain to life and godliness” (1:3)? Does this mean we don't need any outside resources like books or other people? How would you explain this?3. One way we grow and become “partakers of the divine nature” is by hearing and believing His “precious and very great promises” (1:4). What promises has God made that encourage you most? 4. What are some practical ways we can make sure we continue believing and trusting so we keep growing in Christ? Our Active Role in Spiritual Growth (2 Peter 1:5-11)5. Peter mentions eight characteristics in vv. 5-7. We are supposed to “make every effort” to grow in these areas. In which of these areas do you need to grow?6. In v.8 we learn we should be “increasing” in these eight areas and in our spiritual growth. Are you in a season of spiritual growth right now? If not, what next steps can you take?7. Do you tend to emphasize the active role or the passive role in your spiritual growth? Which aspect needs some extra attention right now? Pray for the Unreached: The Kapu in India The Kapu are a large Hindu community in southern India, historically known as farmers and protectors, with many still engaged in agriculture today. While some have moved into business, education, and leadership, many remain in rural areas with limited access to resources. Though the Bible and gospel tools are available, only a small percentage follow Christ. Pray for a growing movement of disciples among the Kapu, for hearts to be open to the gospel, and for both their spiritual and physical needs to be met.FinancesWeekly Budget 34,615Giving For 04/19 40,027Giving For 04/26 20,288YTD Budget 1,488,462Giving 1,796,478 OVER/(UNDER) 308,016 Fellowship Baby DedicationFellowship is grateful to partner with parents as they dedicate their children to the Lord. On Sunday, May 10, during both services, we will set aside a special time for families to dedicate their children before the Fellowship body. If you would like to participate, please email Lisa at lgerdes@fellowshipconway.org and include your preferred service time. The dedication will take place at the beginning of each service.New to Fellowship?We are so glad that you chose to worship with our Fellowship Family this morning. If you are joining us for the first time or have been checking us out for a few weeks, we are excited you are here and would love to meet you. Please fill out the “Connect Card” and bring it to the Connection Center in the Atrium, we would love to say “hi” and give you a gift. Fellowship Kids VBS - There's No Place Like Rome…That's why we want your kids to join us for an exciting Bible-times adventure with the Underground Church in ancient Rome! They will explore authentic Marketplace shops, visit the Apostle Paul (who's under house arrest), sneak to the cave where the Underground Church meets, take part in games, dance to lively Bible songs, and sample tasty tidbits as they discover more about the early church. Join us June 22-26, 9:00 am- 12:00 pm. This is for kids currently in Kindergarten through 4th grade. Register by May 24 at ffellowshipconway.org/register. FSM 2026 Fellowship GraduatesWe're excited to celebrate our 2026 high school seniors! If Fellowship Bible Church is your home, we'd love to honor you during our Sunday morning services on May 17 at fellowshipconway.org/register. Send five photos for the senior slideshow to Casey Goode at cgoode@fellowshipconway.org by May 1. Fellowship Women's Bible Study - Knowing GodJoin us for “Knowing God,” a 4 week study of The Trinity by Rebecca Carter & Heather Harrison. We'll meet Tuesday nights at 6:30pm, beginning June 2nd at Fellowship. Register at fellowshipconway.org/women. Text Shanna at 336-0332 to reserve free childcare by May 25th.Fellowship Kids Summer Volunteers We are wrapping up the school year and preparing to head into summer. We are continuing our three-year journey through the Bible, and we need your help to do so. Our children are ready to learn about and worship Jesus. We have a place for everyone...Nursery, classroom leaders, hall monitors, storytellers, and special needs buddies. The summer sessions are May 31 through August 9. Contact Heather today at hfulmer@fellowshipconway.org or Ashley at aoverstreet@fellowshipconway.orgMission fundraiser at StobysTHANK YOU FELLOWSHIP!!! With your support (and many, many pancakes made and eaten) we were able to raise around $2,500 towards our youth/college trip to the Czech Republic. We appreciate you and your generosity!Prayer During ServiceWe love to pray for one another. Our prayer team will have people at the front of the Auditorium under the signs Hope and Love to pray for you after the message. Please feel free to walk up to them for prayer or encouragement during the first worship song after the message.
2 Peter 2:9 TPTIf the Lord YAHWEH rescued Lot, he knows how to continually rescue the godly from their trials and to reserve the ungodly for punishment on the day of judgment.
God's people can close their day with thankful praise because the Lord has all along again provided for them in every way. Transition from Morning to Evening with Continual Praise for God's Provision Throughout Your Day.
Shintaro and David discuss Ken Gunji, a high-level Japanese judoka training at the dojo, and the massive impact he's had on the room. They dive into technical exchanges, differences in gripping systems, and how having an elite training partner elevates both coaching and learning.00:00 Intro to Ken Gunji00:10 Size, athleticism, and presence on the mat00:47 Role in Japan training trip and dojo connections01:11 Keio University and academic background02:20 Language differences: Japanese vs English communication02:50 Asahi Kasei career and All Japan Pro Tournament success04:05 Middle school All Japan wrestling champion04:36 How Gunji elevates the training environment05:13 High-level technical discussions and idea exchange06:31 Importance of having a true “judo sounding board”07:49 Translating high-level concepts into teaching08:42 Consistency of having a world-class athlete in the dojo09:27 Value of hands-on learning vs instructionals09:54 Differences in gripping systems (tight vs loose sleeve control)10:23 Timing-based entries vs control-based approaches10:50 Why certain techniques are hard to teach across styles11:40 Training drills: Ouchi gari finishing mechanics12:38 Gunji's transitions: cutback to Uchimata13:09 Importance of precise, actionable feedback13:54 Problems with vague coaching (“no kuzushi”)14:24 Importance of high-level training partners14:37 Uchimata timing and hop sequencing strategy15:05 Matching opponent movement to avoid counters15:34 Continual learning—even at elite levels15:48 Mutual learning between high-level athletes16:08 Training with Gunji at the dojo16:17 Outro
AI has been trained like software. But what if it should be grown like life? In this episode of Eye on AI, Craig Smith sits down with Sebastian Risi, professor and leading researcher in neuroevolution and artificial life, to explore a fundamentally different approach to building intelligence, one inspired by how nature evolves, grows, and adapts. Sebastian explains why traditional AI systems are limited by fixed architectures and one-time training, and how evolutionary algorithms can create systems that continuously learn, self-organize, and even grow their own neural structures over time. They dive into concepts like plastic neural networks that keep updating during their lifetime, AI systems that can recover from damage, and models that develop from a single "cell" into complex structures, similar to biological organisms. The conversation also explores how combining large language models with evolutionary search could unlock more creative and open-ended problem solving, from merging specialized models to building AI systems capable of generating and testing scientific ideas. If you want to understand where AI is headed beyond today's transformer models, and why the future may look more like living systems than software, this episode offers a clear and thought-provoking perspective. Subscribe for more conversations with the people building the future of AI and emerging technology. Stay Updated: Craig Smith on X: https://x.com/craigss Eye on A.I. on X: https://x.com/EyeOn_AI (00:00) Why copy nature's evolution for AI (01:20) What neuroevolution actually means (05:52) How evolutionary search replaces gradients (08:03) Plastic neural networks and continuous learning (11:53) Growing neural networks like living systems (18:08) Scaling challenges and limits of growth (23:16) Can evolving systems replace LLM training (27:28) Continual learning and model merging (30:27) Artificial life, self-repair, and resilience (35:10) AI scientists and evolution with LLMs
This week, we have tech editors Ronan Mc Laughlin and Josh Weinberg coming in from opposite ends of the world, both in dimly lit hotel rooms, and without their usual podcast recording equipment. It's only up from here! In the first section, Dave Rome chats to Ronan about his previous day touring Pirelli. Then Dave reminds us of the real-world compromises that 1x shifting continues to battle. Also, you'll hear a rant, and Ronan has a Good Thing that many probably already own. Next stop is Dave catching up with our US tech editor Josh Weinberg, who finds himself in Taiwan for the Taipei Cycle Show. And finally, members of Escape Collective (who get access to everything we do at Escape Collective), can tune in for our popular Ask a Wrench segment, where Dave and pro mechanic Zach Edwards answer technical questions from members. Finally, the day has come for Geek Warning to be a motion picture. You'll now find episodes on YouTube, too. Happy Geeking! Time stamps: 4:40 - Touring Pirelli HQ 12:00 - Pirelli's new flagship race tyre 18:00 - Wolf Tooth's new flagship Mark Zero range 21:40 - The real world compromises of 1x shifting 34:00 - Content creation that has Dave ranting 40:20 - Ronan's Good Thing that you may already own - 44:45 - Taipei Cycle Show with Josh 52:00 - 32er manufacturing is brewing, 3D printing, and other trends 1:07:00 - Ask a Wrench (members only) 1:08:00 - Shimano brakes gone bad if not used 1:17:30 - Stuck tyre bead on rim. Solutions? 1:26:00 - Re-using old brake hoses during re-routing
Got a question? Let us know!Step Ten: Ongoing InventoryThis week on Made for Mondays, Heather is joined by Jamey, Tyler, and RaChelle to talk about Step 10 — continuing to take personal inventory and promptly admitting when we're wrong. Before diving into the conversation, the group catches up about the weekend and reflects on the Bible Reading Challenge, continuing through Deuteronomy and Mark.Then the conversation turns to Sunday's message.SUNDAY DISHThe reaction to “taking inventory” When people hear the phrase take inventory, reactions vary. For some it sounds freeing and clarifying. For others it feels exhausting or intimidating. The group reflects on why honest self-examination can feel uncomfortable—even though it's meant to lead to freedom.The illusion of “arriving” Jamey pointed out that following Jesus doesn't make us sinless—it makes us forgiven. Yet many Christians quietly assume maturity means we should eventually stop struggling. That expectation can create pressure to hide our struggles instead of bringing them honestly before God and trusted community.“When,” not “if” Step 10 uses the phrase when we were wrong, not if. That small word reminds us that spiritual growth doesn't eliminate mistakes—it teaches us how to respond when they happen. Honest acknowledgment of failure doesn't lower the bar for holiness; it keeps us grounded in humility and grace.Living one day at a time Jamey shared the illustration of eating a lifetime's worth of food one day at a time. In the same way, spiritual growth becomes overwhelming when we try to think about the entire journey at once. Focusing on today helps us stay connected to Jesus in the present instead of discouraged by the past or anxious about the future.Honest community The group reflects on a story shared Sunday about an older man who openly admitted his ongoing struggles. Moments like that show the power of honesty in community. When people feel safe enough to tell the truth about their lives, it creates space for real growth without pretending we've already arrived.Practicing Step 10 Jamey described three ways to practice this step:Spot-check inventory — pausing in the moment when something feels offDaily inventory — reflecting on the day with GodPeriodic inventory — stepping back occasionally for deeper reflectionFor someone feeling overwhelmed, the best place to start may simply be a daily moment of reflection with God—asking where things went well, where we missed the mark, and where grace is needed.Final ReflectionRegularly admitting when we're wrong doesn't push us farther from Jesus—it keeps us close to Him. Honest reflection reminds us that growth isn't about perfection, but about continually returning to grace.Join Us SundayThat's all we have time for today, friends! Join us THIS Sunday at 9 and 10:45 AM as we continue taking the next step toward healing and freedom together. If you can't make it in person, watch on YouTube at 1 PM.These Steps may be challenging, but they're shaping something good.Remember: it works if you work it. Go be love, everybody—we'll see you next week!Stay Connected Website: https://believerschurch.org/ Bible Reading Plan: https://believerschurch.org/bible-reading-plan/ Believers Facebook: https://www.facebook.com/believerschurch.va/ Believers Instagram: https://www.instagram.com/believers_church/ Subscribe to The Outlet: https://believerschurch.us13.list-manage.com/subscribe?u=66f00f86238de86688d2480e6&id=729c3f381f
The Health Detective Podcast is entering a new chapter. In this transition episode, Michele Scarlet steps in as the new host and shares what listeners can expect moving forward. This podcast is dedicated to real health transformations; the kind that happen when we start asking deeper questions and looking beyond surface-level answers. Together, we'll explore the patterns behind chronic symptoms, investigate root causes, and continually evolve our understanding of functional health. Moving forward, you can expect conversations around: • Functional health transformations and case-based insights • The art of asking better questions in root cause medicine • Continual education and evolving clinical thinking • Interviews with practitioners pushing the boundaries of functional health • Occasional insights into the business side of health coaching and building a practice Because healing doesn't happen when we stop learning; it happens when we stay curious. If you're a practitioner, health coach, or someone passionate about understanding the body and uncovering real solutions, you're in the right place. Subscribe so you don't miss what's ahead. Until the next investigation… keep asking deeper questions.
Bex Moorhouse is Global Head of Strategy, Ops Excellence & Performance - Procurement & REWS at WPP where she is passionate about helping to prioritize employee experience and create environments where people thrive. Mike Petrusky asks Bex about her new role and her passion for cultivating such engaging environments where all employees feel seen, valued, and empowered. She believes that empathy and curiosity are crucial "power skills" for workplace professionals, enabling them to adapt, learn, and lead effectively in changing environments, so she shares that storytelling and clear communication are essential. Bex and Mike discuss how real estate and facility management teams often need to better articulate their value in business terms and focus less on justification and more on impact for end users. Continual adaptation, embracing new technologies, and focusing on both operational excellence and human experience will keep the FM profession relevant, so Mike and Bex offer the inspiration you will need to be a Workplace Innovator in your organization! Connect with Bex on LinkedIn: https://www.linkedin.com/in/bexmoorhouse/ Learn more about Bex's work: https://www.bexmoorhouse.com/ Find out more about WPP: https://www.wpp.com/en-us Watch the podcast on YouTube: https://www.youtube.com/playlist?list=PLSkmmkVFvM4H3pwnlU2AuqynuRDpvnh4J Discover free resources and explore past interviews at: https://eptura.com/discover-more/podcasts/workplace-innovator/ Learn more about Eptura™: https://eptura.com/ Connect with Mike on LinkedIn: https://www.linkedin.com/in/mikepetrusky/
The readings for this homily: https://bible.usccb.org/bible/readings/030926.cfmFather Matthew Tomeny, MIC, opens with a memorable story from Venerable Archbishop Fulton Sheen, who once welcomed a drunk woman into Saint Patrick's Cathedral in New York City. Rather than turning her away, he offered her tea and promised not to ask her to go to confession — until she returned sober and ready to encounter God's mercy.Father Matthew connects this to the Scripture reading of Naaman the leper, who expected an extraordinary cure but was healed by the simple act of dipping seven times in the Jordan River. Salvation does not require grand quests or heroic feats. Instead, the Sacraments of the Church provide the ordinary means by which God cleanses our souls and restores our union with Him.Through Baptism, Jesus washes away our sins. Through the Sacrament of Reconciliation, He continues to cleanse us when we fall. And through the Eucharist, we express that communion in the most intimate way possible. Father Matthew emphasizes that holiness is intended for all people, regardless of their past. Just as Archbishop Sheen did not write off the drunk woman, neither should we write off anyone who struggles.Continual repentance—the virtue of penance—keeps our hearts aligned with God's will. When we are in order with God, trials lose their power to derail us. Take advantage of these simple ways to holiness and share that satisfaction with others. ★ Support this podcast ★
The Philadelphia Flyers are having a tough season to say the least. As the team is still in their rebuild mode, we've seen some good hockey. We've also seen some not to good hockey. What might this team and front office do before the trade deadline? This week, Daniel Esche from BrotherlyPuck.com and the Brotherly Pod podcast joined us for a great discussion about this team as well as some Lehigh Valley Phantoms storylines as well.But first, the guys dove into how the Sixers are having major issues with rebounding - a theme that's not getting enough attention this season. (Approx. 5:00)From there, they got into how the Eagles should be handling certain aspects of the offseason. Should they truly try to acquire Maxx Crosby and what should the front office do about the tight end position?(Approx. 16:40)The guys then talked about the positive vibes stemming from Phillies Spring Training as Andrew Painter and Bryce Harper have been showing us some good stuff! (Approx. 28:10)What they threw down on the Table this week was a great and in-depth conversation with Daniel Esche from Brotherly Puck about this Flyers team. The team has been in a holding pattern. What could Danny Briere do at the trade deadline to truly help this team during the rebuild? Is Rick Tocchet the guy to coach this team to the next level? All of this and much more this week on the Table! (Approx. 36:35)SUBSCRIBE on YouTube: youtube.com/@thephiladelphiasportstableHead over to our website for all of our podcasts and more: philadelphiasportstable.comFollow us on Threads: @philadelphiasportstableFollow us on Twitter/X: @PhiladelphiaPSTFollow us on Instagram: @philadelphiasportstable.Follow us on Facebook: facebook.com/PhiladelphiaSportsTable
Alisa is on fire with a message that's resonating with so many: understanding and finding peace with both food noise—the mental chatter around eating—and body noise—the persistent thoughts about how we look and show up in the world. In this episode, she breaks down why these inner messages exist, how they connect to deeper feelings of safety, and what it takes to respond with awareness, truth, and grace. What You'll Learn: What food noise really is: Persistent mental chatter about food that can feel overwhelming. Why it exists: Not all food noise is bad! Hunger is actually a gift—a signal from your body to nourish yourself. The 7 forms of hunger: From true, homeostatic hunger to emotional, social, environmental, habitual, survival, and reward hunger. Understanding which is which can help you respond with awareness rather than guilt. The connection to safety: Increased food noise often points to a deeper sense of insecurity or loss of safety. Body noise: Continual thoughts about our appearance and how we look, and why awareness is the first step in responding differently. Soul work: How sin and brokenness in our world have disrupted our connection to God, ourselves, and others, and why body and food struggles often show up the moment we feel "something is wrong" with our bodies. Hope in community: RW+ offers a space to explore, process, and grow together. Scripture Highlight: Genesis 3:9 reminds us that God seeks us even in our brokenness—and invites us into healing, not shame.
#310 In this episode, Thom Plummer shares insights on how gym owners can build lasting businesses by focusing on purpose, continuous learning, and strategic planning. Discover how to avoid common pitfalls, develop emotional connections, and reinvent your approach every few years for lasting success and fulfillment. Main Topics Covered: The importance of starting with a clear life and business purpose Mastering core skills: sales, marketing, and finance Developing emotional ties through storytelling and communication The significance of strategic exit planning and long-term vision Reinventing your business approach every 3-5 years Trends shaping the future of gyms, including aging populations and AI impacts Personal character traits that foster ongoing growth and leadership Timestamps: 00:00 - Introduction and guest credibility 02:29 - Beginnings in the gym industry in 1977 04:13 - Lessons learned from past entrepreneurial ventures 07:15 - Mastering business fundamentals for gym owners 09:14 - Horizontal vs. vertical management in gyms 11:19 - The importance of daily sales and marketing discipline 13:21 - Planning your business with an end goal in mind 17:30 - The significance of understanding your exit strategy 20:39 - Trends: aging populations and classic gym models return 23:10 - The dangers of chasing other people's dreams 25:22 - The power of storytelling and emotional connection 27:05 - Developing communication skills for business growth 29:24 - The impact of reading and continual learning 32:37 - AI, data, and future trends in the fitness industry 35:52 - Reinventing the business every 3-5 years 39:00 - Personal stories of transformation and leadership 43:41 - Building genuine client relationships and community 47:36 - The value of personal integrity and serving others 51:22 - Continual self-education and pulling the thread of knowledge 54:08 - Resources and strategies for improving writing skills 58:23 - The importance of storytelling and human touch in marketing 61:00 - Final reflections on purposeful living and legacy Resources & Links: The Experience Economy by Pine & Gilmore Stage Not Age by Christina Roulstone Anne Handley's Writing Courses Gotham Writers Workshop SNHU Writing Programs Connect with Thom Plummer: Perform Better Speaker Schools Additional Highlights: Emphasizing that true progress comes from intentional story-driven communication. The necessity of planning the business's end game to align daily actions. Reinvention as key to staying relevant amidst changing industry trends. The importance of character traits like coachability, humility, and service. This episode underscores that long-term success in the gym industry hinges on purpose, strategic management, continual learning, and authentic human connection. Implement these principles to elevate your business and life.
In this introspective episode of the Uncommon Wealth Podcast, host Phillip Ramsey addresses a compelling issue that many entrepreneurs face: the perils of anchoring one's identity solely in their business endeavors. Through reflective storytelling and practical insights, Ramsey unpacks the emotional and psychological challenges that can arise when deeply invested business leaders suddenly find themselves in a post-sale identity vacuum.Dive into a nuanced conversation about the intertwined nature of identity and entrepreneurship. Phillip Ramsey shares the compelling narrative of a former business owner who experienced a complete mental breakdown after selling his business. This episode delves into how this profound experience reveals the often-overlooked psychological stress related to business ownership and identity. Ramsey emphasizes the importance of acknowledging and addressing these issues early on, suggesting mindful consideration of one's identity beyond their business achievements. With expert analysis and real-life examples, Phillip provides a roadmap for aligning personal identity with broader life roles, such as being a parent, partner, or spiritual being, thus paving the way for a healthier, more balanced approach to business and life.Key Takeaways: Entrepreneurs often face identity crises post-business sale, highlighting the danger of tying identity solely to business roles. Business ownership, much like parenting, involves deep emotional investment, which can lead to identity issues when those roles change or end. Maintaining a multi-faceted identity is crucial for personal well-being; this includes identifying roles beyond being a business owner. Continual self-reflection and asking "Who am I?" can help build a comprehensive identity that is not reliant on external success or business status. Emphasizing personal roles, such as being a friend, family member, or spiritual individual, can offer stability and fulfillment beyond business achievements.Notable Quotes: "A lot of times business owners and people who are starting a business struggle with putting all of their identity in their business." "You pour yourself into a business as much as you do a child; when it's gone, you can feel lost, no matter your bank account balance." "When you think about your identity, try to think about something other than what your business is." "I want to have my identity rooted in something that can never be taken away."
I talk about how being emotionally accessible all the time is often mistaken for maturity or leadership, but taken too far it becomes a liability. When I'm always available to absorb other people's emotions, my own clarity and authority start to fade. Constant access doesn't build real connection, it trains people to depend on me while draining my energy and decision making. In this episode, I explain why boundaries around emotional access protect your presence and help you stay strong and focused. Show Notes: [02:25]#1 Constant access trains people to offload their emotional regulation onto you. [08:40]#2 Accessibility dilutes signal and presence. [12:44]#3 Continual access creates emotional debt. [15:36]Recap Next Steps: --- Power Presence is not taught. It is enforced. If you are operating in environments where hesitation costs money, authority, or leverage, the Power Presence Mastermind exists as a controlled setting for discipline, execution, and consequence-based decision-making. Details live here: http://PowerPresenceProtocol.com/Mastermind This Masterclass is the public record of standards. Private enforcement happens elsewhere. All episodes and the complete archive: → WorkOnYourGamePodcast.com
Nathan Lambert and Sebastian Raschka are machine learning researchers, engineers, and educators. Nathan is the post-training lead at the Allen Institute for AI (Ai2) and the author of The RLHF Book. Sebastian Raschka is the author of Build a Large Language Model (From Scratch) and Build a Reasoning Model (From Scratch). Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep490-sc See below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc. Transcript: https://lexfridman.com/ai-sota-2026-transcript CONTACT LEX: Feedback – give feedback to Lex: https://lexfridman.com/survey AMA – submit questions, videos or call-in: https://lexfridman.com/ama Hiring – join our team: https://lexfridman.com/hiring Other – other ways to get in touch: https://lexfridman.com/contact SPONSORS: To support this podcast, check out our sponsors & get discounts: Box: Intelligent content management platform. Go to https://box.com/ai Quo: Phone system (calls, texts, contacts) for businesses. Go to https://quo.com/lex UPLIFT Desk: Standing desks and office ergonomics. Go to https://upliftdesk.com/lex Fin: AI agent for customer service. Go to https://fin.ai/lex Shopify: Sell stuff online. Go to https://shopify.com/lex CodeRabbit: AI-powered code reviews. Go to https://coderabbit.ai/lex LMNT: Zero-sugar electrolyte drink mix. Go to https://drinkLMNT.com/lex Perplexity: AI-powered answer engine. Go to https://perplexity.ai/ OUTLINE: (00:00) – Introduction (01:39) – Sponsors, Comments, and Reflections (16:29) – China vs US: Who wins the AI race? (25:11) – ChatGPT vs Claude vs Gemini vs Grok: Who is winning? (36:11) – Best AI for coding (43:02) – Open Source vs Closed Source LLMs (54:41) – Transformers: Evolution of LLMs since 2019 (1:02:38) – AI Scaling Laws: Are they dead or still holding? (1:18:45) – How AI is trained: Pre-training, Mid-training, and Post-training (1:51:51) – Post-training explained: Exciting new research directions in LLMs (2:12:43) – Advice for beginners on how to get into AI development & research (2:35:36) – Work culture in AI (72+ hour weeks) (2:39:22) – Silicon Valley bubble (2:43:19) – Text diffusion models and other new research directions (2:49:01) – Tool use (2:53:17) – Continual learning (2:58:39) – Long context (3:04:54) – Robotics (3:14:04) – Timeline to AGI (3:21:20) – Will AI replace programmers? (3:39:51) – Is the dream of AGI dying? (3:46:40) – How AI will make money? (3:51:02) – Big acquisitions in 2026 (3:55:34) – Future of OpenAI, Anthropic, Google DeepMind, xAI, Meta (4:08:08) – Manhattan Project for AI (4:14:42) – Future of NVIDIA, GPUs, and AI compute clusters (4:22:48) – Future of human civilization
WANTED: Developers and STEM experts! Get paid to create benchmarks and improve AI models. Sign up for Alignerr using our link: https://alignerr.com/?referral-source=briankeating One of the most powerful AI systems we've ever built is succeeding for reasons we still don't understand. And worse, they may succeed for reasons that might lock us into the wrong future for humanity. Today's guest is Anil Ananthaswamy, an award-winning science writer and one of the clearest thinkers on the mathematical foundations of machine learning. In this conversation, we're not just talking about new demos, incremental improvements, or updates on new models being released. We're asking even harder questions: Why does the mathematics of machine learning work at all? How do these models succeed when they suffer from problems like overparameterization and lack of training data? And are large language models revealing deep structure, or are they just producing very convincing illusions and causing us to face an increasingly AI-slop-driven future? KEY TAKEAWAYS 00:00 — Book explores why ML works through math 02:47 — Perceptron proof shows simple math guarantees learning 05:11 — Early AI failed due to single-layer limits 07:12 — Nonlinear limits caused the first AI winter 09:04 — Backpropagation revived neural networks 10:59 — GPUs + big data enabled deep learning 15:25 — AI success risks technological lock-in 17:30 — LLMs lack human-like learning and embodiment 22:57 — High-dimensional spaces power ML behavior 27:36 — Data saturation may slow future gains 31:11 — Continual learning is still missing in AI 33:46 — Neuromorphic chips promise energy efficiency 41:49 — Overparameterized models still generalize well 45:05 — SGD succeeds via randomness in complex landscapes 48:27 — Perceptrons remain the core of modern neural net - Additional resources: Anil's NEW Book "Why Machines Learn: The Elegant Math Behind Modern AI": https://www.amazon.com/Why-Machines-Learn-Elegant-Behind/dp/0593185749 Get My NEW Book: Focus Like a Nobel Prize Winner: https://www.amazon.com/dp/B0FN8DH6SX?ref_=pe_93986420_775043100 Please join my mailing list here
→ Watch on YouTube → Detailed Show Notes → Timestamps: (00:00) An overview of the The Old Testament.(04:27) Bryce breaks down the Old Testament into nine time periods.(09:41) Canonization and the creation of the Greek Septuagint. The authors of the New Testament quoted the Greek version of the Old Testament. The version that Jesus used is unknown.(13:01) The Old Testament is not one book written by a single author. It is an anthology of books written over centuries by individuals. Later authors sometimes rejected and edited earlier authors. Remnants remain that multiple Gods participated in the creation. The first word in Genesis invites us to consider the Grand Pre-mortal Council.(20:03) The creation accounts address why the world was created, not how. The purpose of the earth is to create eternal families.(33:17) Ancient cultures shared the creation story at their temples during the New Year. The time-honored principle of marriage and family are connected to the purposes of creation.(36:21) Chaos played a role in the formation of the earth. God transformed unorganized matter into something beautiful. The cosmology of the Bible invites modern readers to think about scripture differently.(41:18) The temple takes us to the creation and the creation takes us to the temple because each us back to God.(48:25) Moses 2 and 3 can be read as a single account. This is contrasted with the documentary hypothesis, where scholars believe Genesis 1 and 2 come from two different sources because of differences in the text.(54:21) Seven eternal lessons in the creation account invite us to find success during our time on earth. Finding balance between work and rest.(1:00:29) Seek spiritual things first, then temporal.(1:01:50) Cherubim and a flaming sword foiled Satan's plan and preserved the space between the trees. God protected our probationary estate. God knew from the very beginning that we would sin and need time to repent.(1:13:11) Satan's Plan B is to get us to take away our own or someone else's probationary state. Toxic perfectionism is addressed. Continual progression is what matters to God.(1:16:55) The river flowing out of Eden as a symbol for the division found in mortality. We must find ways to be unified.(1:23:08) Instead of focusing on what we are missing, our focus should be on all our blessings.(1:28:23) Pardes is an acronym to describe the ways of reading scripture: Peshat, remez, derash, and sod. The plain reading (peshat), allegorical or hidden reading (remez), the moral or imperative sense or application (derash), and the mystical, esoteric, or temple reading (sod).(1:30:13) The rib in Genesis 2.22 symbolizes partnership in marriage.(1:38:44) “Helper” or ʿēzer as found in Genesis 2.18 has often been misread and used to subjugate women. The term translated as “help meet” actually denotes the kind of powerful help that God gives. Eve's position next to Adam places both in a setting as having dominion over the whole earth. Eve is called “Zoe” in the LXX, the mother of all the living ones. There is no kingdom without Eve.(1:47:02) Fig leaves can represent covering our sins with bigger and bigger lies.(1:52:47) Garments are a piece of the temple that we wear to remind us of our connection to the Savior's atonement. → For more of Bryce Dunford’s podcast classes, click here. → Enroll in Institute → YouTube → Apple Podcasts → Spotify → Amazon Music → Facebook The post Ep 354 | Genesis 1-2; Moses 2-3; Abraham 4-5, Come Follow Me 2026 (January 12-18) appeared first on LDS Scripture Teachings.
Drawing from Maslow's hierarchy of needs, this episode breaks down what it truly means to live with integrity, creativity, compassion and authenticity while freeing yourself from the opinions of others. It dives into acceptance, spontaneity, gratitude and purpose as essential practices for growth, emphasizing that self actualization is not a destination but a daily commitment to personal and spiritual evolution. By focusing inward, embracing empathy and maximizing your own potential, this conversation challenges you to live more fully, lead with compassion and show up each day as a better version of yourself. Own your power with this Success Tip. For more about Rod and his real estate investing journey go to www.rodkhleif.com