POPULARITY
On Point news analyst Jack Beatty has an epiphany: Retrofitting should be the defining goal of our civilization. *** Thank you for listening. Help power On Point by making a donation here: wbur.org/giveonpoint
Watch the full episode on YouTube:We first covered Baseten last year when DeepSeek mania was at peak hype. Now they have raised a monster $13B round and become one of the new cohort of AI Infra decacorns that are (with Nvidia, Intel, and the semis complex) chief beneficiaries of the Inference Inflection. We return to Baseten at the peak of the 2026 edition of Open Weights debate. Ali has published a viral breakdown of Kimi K3:And since you last saw him, Philip has spoken at AI Engineer and written the definitive book on Inference Engineering spotted all over SF:Three years ago, inference engineering barely existed as a category.Today, it is one of the most critical disciplines in AI. Inference engineering inherently tackles a different question than standard model training: “How do you turn those weights from training into a product that is fast, reliable, and affordable at scale?” Focusing on these creates an entirely new optimization problem.In one recent GLM-5.2 experiment, quantizing more of the model actually preserved its benchmark quality while increasing throughput by 20%, because the errors introduced in different layers could cancel each other out.Inference is no longer just the final step after training. It is becoming its own engineering discipline, with its own research problems, infrastructure, and increasingly specialized roles.In this episode, Baseten's Philip Kiely and Ali Taha join swyx and Vibhu to explain what actually happens after a new open model is released and what it takes to turn “we generated a token” into a fast, reliable, production-ready API.We go deep on cache-aware routing, disaggregated prefill and decode, quantization, speculative decoding, KV-cache movement, model parallelism, GPU kernels, and the race to make frontier models up to 10× faster. Philip and Ali explain why inference optimizations can still produce gains of 20%, 100%, or even 200%; how quantization errors can cancel one another out; why identical weights can behave differently across clusters; and how Baseten grafted a Kimi vision encoder onto GLM-5.2 without changing the underlying language model.The conversation then expands beyond LLMs into NVIDIA Dynamo, mega kernels, Rubin, AI-specific chips, local inference, video generation, diffusion versus autoregressive models, and the enormous compute barrier to generating coherent long-form video. Finally, we explore the convergence of training and inference, continual learning through persistent KV cache, and the emerging loop where models help optimize the infrastructure that runs them.We discuss:* What happens when a 200,000-token request enters an inference system* Cache-aware routing and reusing previously computed KV cache* Why prefill and decode are increasingly handled by different GPUs* When dedicated deployments become cheaper and more reliable than shared APIs* How speculative decoding uses a smaller model to accelerate a larger one* Tool calling, structured outputs, and what LLMs actually do* What it takes to support a new open model on day zero* Grafting Kimi's vision encoder onto GLM-5.2* Retrofitting inefficient model layers with components from other architectures* Why models sometimes collapse into repeating the same token* How hardware, kernels, and race conditions create nondeterministic failures* Preserving model fidelity while making inference faster* How quantization errors can cancel each other out* Why inference optimizations still deliver gains of 20%, 100%, and 200%* How optimized serving can make a model up to 10× faster* NVIDIA Dynamo, KV-aware routing, and distributed model serving* Speculative decoding the speculative decoder* Why local AI is about making models less dumb while data-center AI is about making them less slow* Tensor, expert, and pipeline parallelism across GPUs* Hardware-aware model design, auto-tuning, and the case against mega kernels* Rubin and why inference is becoming a systems problem* Whether modern GPUs are evolving into programmable AI ASICs* Why enormous models like Kimi K3 require GB300-class hardware* Why open-source video generation still trails Veo, Kling, and other closed models* The quadratic attention bottleneck behind long-form AI video* Autoregressive video, real-time generation, and compounding quality drift* Why future video systems may combine autoregressive and diffusion architectures* Training for inference and inference for training* Continuous post-training, deployment, evaluation, and improvement loops* How GLM-5.2 helped optimize the kernels serving GLM-5.2 itself* Why faster networking could unlock dramatically faster decoding* Continual learning, KV-cache compaction, and persistent model memoryShow Notes* How to build a day-0 API for Kimi K3* 22580: From GPT2 to Kimi3, ExplainedPhilip Kiely* LinkedIn: https://www.linkedin.com/in/philipkiely* X: https://x.com/philipkiely* Inference Engineering: https://www.baseten.co/inference-engineering/Ali Taha* LinkedIn: https://www.linkedin.com/in/aliestaha/* X: https://x.com/waterloointernTimestamps00:00:00 Introduction and the 200K-Token Prompt00:03:18 Dedicated Deployments, Speculative Decoding, and Tool Calling00:11:26 Launching Production-Ready Open Models00:19:06 Model Retrofits, Failure Modes, and Nondeterminism00:28:22 Quantization and Canceling Errors00:32:15 The Race to 10× Faster Inference00:40:48 Dynamo, Speculation, and Local vs. Data-Center AI00:50:18 Model Parallelism, Auto-Tuning, and Mega Kernels01:00:55 Rubin, GPUs vs. ASICs, and Custom AI Chips01:10:03 Giant Models and the Limits of GPU Memory01:12:42 AI Video, Quadratic Attention, and Autoregressive Generation01:21:47 Audio, Images, and Diffusion Models01:27:32 Training, Self-Optimizing Models, and Continual Learning01:40:06 Closing ThoughtsTranscriptIntroduction: Baseten, Waterloo Intern, and Inference EngineeringSwyx [00:00:00]: Okay, we're here in the studio with Philip, old friend from Inference Engineering, the book, as well as Baseten and everything that you've done, you and I have done before, as well as Ali. Welcome.Ali [00:00:15]: Pleasure to meet you.Swyx [00:00:15]: Waterloo intern.Ali [00:00:16]: Waterloo intern, always.Swyx [00:00:17]: When did you get “Waterloo intern” as a handle?Ali [00:00:19]: As a handle? Oh.Ali [00:00:20]: I think the rebranding happened mid-March. When I saw it was open, I was like, “I have to take it. Up for grabs.”Philip [00:00:26]: The problem is that Ali is really good at his job and is not gonna be an intern much longer.Philip [00:00:30]: So we have to figure out who's gonna get the handle.Ali [00:00:33]: Well, I'll pass the torch over to the next intern.Swyx [00:00:34]: Oh, okay. It can be, like, you just pass it to another Waterloo grad.Ali [00:00:37]: To another Waterloo intern. No, bruh.Philip [00:00:39]: Yeah.Ali [00:00:39]: Intern.Swyx [00:00:40]: Intern, yeah.Ali [00:00:40]: And no.Philip [00:00:41]: You gotta get an intern from Waterloo.Ali [00:00:42]: Yeah, I've gotta get an intern from Waterloo.Swyx [00:00:44]: Right.Ali [00:00:44]: But they have to follow the path.Swyx [00:00:45]: Oh, it could, but it could come from Baseten, so it's like whoever Baseten gets from Waterloo.Ali [00:00:48]: Right.Swyx [00:00:49]: Has the title of Waterloo.Ali [00:00:50]: It stays in the ecosystem.Philip [00:00:51]: Exactly.Ali [00:00:52]: Halfway through the internship, you either get it or you're out.Philip [00:00:55]: You should also do, like, a big graduation ceremony where you change the handle.Ali [00:00:59]: Just say it.Philip [00:00:59]: For everybody.Swyx [00:01:00]: You guys are good at ceremonies, clearly. We had a nice launch of the book, very successful. But before we get into all that, I wanna start off with a fun question for you. Okay, you're an expert inference engineer. What happens when I send a long query, say two hundred thousand tokens into Baseten's inference? What's the process of query through GPU model routing, balancing, all that? What is all the stuff that we don't think about?Long Context Requests, KV Cache, and Cache-Aware RoutingPhilip [00:01:26]: With a long query specifically, the first thing that I'm gonna ask is, “Have you sent me this query before, or at least part of it?” and I really hope you have, because it's gonna be a lot easier for me and a lot cheaper for you. So the first thing that we're gonna look at is some cache-aware routing, where we're going to see, we probably have a number of instances, a number of replicas up serving whatever model you're hitting. We want to send this one to something with, number one, available prefill workers, and number two, ideally some cached input already there so that we can skip prefill on at least part of these two hundred thousand tokens. If you're doing two hundred thousand tokens, it's probably coding or a multi-turn agent or something where you would expect to have that cached. If you don't, we're gonna have to send it to a prefill worker. We've at least on certain models disaggregated prefill and decode, so you're going to have one set of GPUs that's solely going to process the input, create the KV cache, and get you your first token, and then that's going to be passed over to a separate set of GPUs, which is going to run decode. We're going to iteratively make those tokens. We're probably going to have some speculator model in front of that. I'm going to assume that you're doing coding, and because of that, our speculator model, which assumes you're doing coding, is gonna have a high draft token acceptance rate. If I'm wrong and you're asking me to summarize every Harry Potter book, it's gonna be slower. And then we stream that output to you and account for it, charge you, a couple of pennies and say, “Hey, would you like to send another one?”Swyx [00:03:04]: Except Baseten doesn't charge by pennies.Philip [00:03:07]: Well, yeah, we charge. I'm assuming that we're talking about the public model APIs. If you are setting up a dedicated deployment, then yeah, it's not pennies.Public APIs vs. Dedicated DeploymentsSwyx [00:03:18]: Yeah, one of the key differentiators when I was talking with Baseten initially was that people who want very high volume just need to rent by the box, ‘cause then it's up to you to figure out how to saturate the box.Ali [00:03:31]: And more often than not, it's, like, way cheaper if you're pushing, like, millions of tokens per hour, if you just pay per hour instead of pay per token.Philip [00:03:37]: Yeah, they do. I think that we've increasingly seen a lot of demand for the pay per token APIs, just because everyone wants to try open models, and then once they find a use case that's really sticky, then they move over to dedicated.Swyx [00:03:51]: Is there a best practice on when it's time to swap over?Philip [00:03:54]: Couple reasons. Yeah, reliability, that's a big one, right?Ali [00:03:57]: Like, if they have a very specific use case, they want you to train something specifically for them, like they want their own spec dec, for instance, for their own traffic.Swyx [00:04:04]: Spec dec is speculative decoding.Speculative Decoding and Custom SpeculatorsAli [00:04:05]: Speculative decoding, yeah.Swyx [00:04:07]: You have to explain.Ali [00:04:07]: Sorry. Like, speculative decoding is like, if you have a huge model, right? And so the model is going to be generating one token at a time every single turn, every single forward pass. So we attach, like, this little, like, parasite, like this layer that goes on top of the model, and this model just has to predict. It does three very fast autoregressive forward passes, and it will predict, like, three certain tokens, and then you do one forward stage over the entire original model in order to see if those predictions were correct or not, and then you accept them or you reject them. Now, this draft model is traffic specific, so if you, like, Philip said, if you're summarizing Harry Potter books, I can train exclusively that draft model on Harry Potter books, and I can guarantee you that I'm gonna accept the three tokens every single time. And so with that case, I increase your decode speed. I wouldn't be able to provide this to you if you're a shared endpointSwyx [00:04:53]: YeahAli [00:04:53]: ‘cause I have no idea if you're doing Harry Potter, if you're doing coding, if you're doing English. We don't know. Also, there was a thing in the book that mentioned that if they really cared about a specific threshold, chapter four, I think. Do you remember that?Philip [00:05:06]: Yeah. The things that you can do is you can set a specific, like, batch sizing, a specific, like, parallelism strategy if you're trying to optimize for, like, throughput versus latency. You can. Maybe a NVFP4 quant doesn't pass your benchmarks and you wanna run a model at higher precision, you could do that. There's just a bunch of reasons why you might wanna have your own endpoint and the biggest one, of course, just being, like, you don't have to deal with someone else doing a hundred million tokens of benchmarking traffic at the endpoint when you happen to be trying to serve your users.Swyx [00:05:40]: Yeah. I think one thing that is. That is a classic journey. Like, it's people is asking the, what happens when you type Google into the browser. Tool calling, is that just, you're generating JSON or is there more complication beyond that?Tool Calling, JSON, and Structured OutputsAli [00:05:58]: Certain customers that we have, they have their own post-trained models, and so they demand a tool calling that's not just, like parse a file or go find the weather. It's something that's very specific and you have to do post-training on this. And if the post-training on the model is not good or if the quantization after the post-training to get the inference to be fast, the model will struggle reading the JSON file and reading the tool calling. But it doesn't require its own like sandbox. It's not like it's going to use that tool calling to like escape a sandbox or like it doesn't have to be contained. It can just be a normal dedicated deployment. The challenge with tool calling more and more seems to be that the companies want certain tool calling which is a very sensitive thing to train. And because you're dealing with all of the JSON outputs, if it doesn't like close the end of the request in a very certain manner, you end up with a model that did the tool calling and like the thinking and so as a result of that, it didn't see the result and just hallucinated the result as it decoded. That seems to be the most challenging thing with tool calling, not really the sandboxes model.Philip [00:06:56]: Yeah, that's a challenge on the training side and then on the inference side, there's work that you can do to scope the possible output. So we published this at this point close to two years ago, the solution to this problem which is you make a state machine and you use that to constrain the output to a specific format. So this is the structured output problem. If you remember backSwyx [00:07:27]: Yeah, the specific grammar is,Philip [00:07:29]: Yeah, exactlySwyx [00:07:30]: GML had this thing.Philip [00:07:31]: Yeah. So it's like the old-school “make sure this is only JSON”, return only JSON orSwyx [00:07:38]: YeahPhilip [00:07:38]: Grandma's gonna die type of prompts.Swyx [00:07:39]: Is it BNF grammar? At some point OpenAI had released a thing that was like, yeah, if you want to constrain your output, write BNF grammar, back as NOR.Philip [00:07:47]: In our inference system, it's just a specified output format. And you get the guarantee that your output's gonna be structured along that format. And so applying that to tool calls can like help cut down on. You can still call the wrong tool or call no tool. It doesn't solve the certainty problem but it at least solves the output structuring problemSwyx [00:08:10]: YeahPhilip [00:08:10]: Within tool calls.Swyx [00:08:12]: And MCP is just another form of tool, right.Philip [00:08:14]: Yeah, exactly.Swyx [00:08:15]: As far as there's no special thing there.Philip [00:08:16]: The thing I'm always like explaining to people is the LLM is not capable of doing anything. It's only capable of making suggestions of what to do and then if those suggestions are formatted in a certain way and applied to a system that knows what to do with them, then an action occurs.Swyx [00:08:32]: Yeah. Part of the fun stuff is, this is solved outside of tool calling too. Like in an agent loop if the output is not correct or you're right, like reasoning, tool calling was done in the reasoning trace, just be like, “Oh, I don't know what to do. Let me just try again.” And it might get there after a few tries. And on your point of training, sometimes this is harder in smaller models, so you don't have the same exact quality outputAli [00:08:56]: Right.Swyx [00:08:57]: When you just swap from a big model, right?Ali [00:08:59]: Yeah. I will say that, before, I think we need to go back to inference engineering proper.Ali [00:09:04]: But, I had expected that something would replace JSON because it's hard to stream JSON ‘cause JSON must be complete and you must have open and close brackets and everything. So it's hard to parse something or validate something while it's being streamed. So people invented all sorts of things that are like, I forget the name of some of these alternatives, but it's something like TOML, something like YAML. But JSON seems to be dominant still.Philip [00:09:30]: The JSON outputs aren't that long, right? Like you could have a long-- ‘cause tool calls also contain the arguments in them and perhaps for a certain tool you might pass like a very long argument. But my impression of the median tool call is that it's a relatively small number of tokens, right? So I would expect that speculators are generally fairly good at something as formatted as JSON. And so you would have like a pretty fast decode step there and that the streaming wouldn't be as valuable, but maybe I'm wrong about that.Ali [00:10:02]: I think you're also bounded by the software or that the model is gonna integrate with if the software is built with JSON for the tool calls or if the company that you'- if your customer says that this is how our software works and our tools are interfaced with JSON, you can ask them to like, change their software and say like, “Yeah, this is gonna be better for the model.” but like with the right training shouldn't be that much of a difference. Also more profitable if it outputs more tokens probably.Swyx [00:10:25]: Depends on your business model.Swyx [00:10:27]: It really depends. But I will say that, as a writer with like experience a lot with generated output, I do try to move from text to JSON text which is very long JSON, right? Like there's paragraphs in every field because I'm trying to structure it, right?Philip [00:10:44]: Right.Swyx [00:10:44]: I want you to first make factual statements, then make opinions then make bullet point summaries, have dates, have entity references have your sources for references, all these things. Anyway, so these are things that like I think people who really experiment with structural output have to really care about. But, let's, let's recurse up the stack a little bit. Before we started recording, you mentioned something really cool, which is that there's a lot of engineering that-- inference engineering that goes on when a new model provider releases a new model, right? So let's call it GLM-5.2, Kimi K3. I had previously assumed, especially if it's like, well, GLM 5 to 5.1 to GLM-5.2, like that you've supported them before. Is it that much work?What It Takes to Support a New Open ModelAli [00:11:26]: It's a lot of work.Swyx [00:11:28]: Yeah. Okay. So like, a lot of people, all you guys, right whenever a new model launch like, people rush to say like, “Oh, Hugging Face supports this, Fireworks supports this, Spacetime supports this,” and I'm like, “Yeah, of course we support it.” But what goes into that? What goes intoPhilip [00:11:40]: I think it's more than just support it too, right? It benefits the consumer a lot. Like I think it was with Kimi K2.5 or GLM-5.2 the latest, there was an inference war, right? X provider is at 90 tokens a second. The next day we're at 150. The nextSwyx [00:11:55]: I kinda kicked that off with the GLM-5.2.Swyx [00:11:58]: I wrote a Twitter article about. It got like half a million views,Ali [00:12:02]: Based on being numberSwyx [00:12:03]: YeahAli [00:12:04]: Or it's for something else.Swyx [00:12:05]: Yeah. Which,Ali [00:12:06]: Oh my GodSwyx [00:12:07]: Which then got everyone really excited about, hey, how can we, bend tracks a little bit further and,Philip [00:12:14]: There's a difference between support the model, as in I can make a token out of this model, and support a model, as in I have a production-ready API from this model.Philip [00:12:26]: Getting to the point of I can make a token out of this model is not that hard because generally the, open source inference engines, vLLM, SGLang of the world oftentimes even receive weights ahead of time, maintainers do, or the people making the model merge PRs to ensure support. So you generally can, just get it working on the standard open source stack without too much pain in most cases. The challenge is, every inference company is gonna have own proprietary stack. Some open source components, some in-house stuff. And for any arbitrary model, there's going to be some new stuff. Sometimes you get lucky, like K, two five to two six was, like, pretty similar.Quantization, Speculators, and Production ReadinessAli [00:13:16]: Yeah. It was pure continued post-trainingPhilip [00:13:18]: YeahAli [00:13:18]: If I remember correctly.Philip [00:13:19]: Even in those cases, there's still stuff you have to do. You have to redo the quantization work. You're taking the model from. Generally, these models are not released in NVFP4, and we want them to be in NVFP4 for maximum Blackwell compatibility. So we have to perform that quantization, and, calibrate the quantization to make sure that we're not causing any regression in the model's intelligence. And then we also have to train the speculator, as we've talked about. Generally, we have. We have ZDR, zero data retention on our model APIs, so we don't know exactly the traffic that people are sending us, but we know what's popular. We know that coding use cases are popular. We know that agents, agentic use cases are popular. So we can get public data sets that are representative of that traffic and train general speculators. Now, with speculators today, you need to train the speculator using the base model itself because you're getting hidden states out of the model from running inference on these specific prompts, and that is the training data you use to create the speculator. So there's that process which you need the real model weights for. And then there's of course just the process of, standing up all the infrastructure behind it, loading all this stuff, testing it. And then when there's a new model with a newer architecture, I think that, like, the DeepSeek models tend to be the most challenging as they have, like, the most novel architectural stuff going on, model after model. But every new model has something. Kimi K2 had. Oh, sorry, GLM-5.2 hadAli [00:14:53]: Sparse attention.Philip [00:14:54]: Yeah,Ali [00:14:54]: YeahPhilip [00:14:54]: the DSA.Ali [00:14:55]: Right. Which is brought from DeepSeek.Philip [00:14:57]: Yeah. AndAli [00:14:59]: So you can copy-paste then?Philip [00:15:01]: It kindAli [00:15:01]: I don't know how this works.Philip [00:15:02]: So, like we had to, like, build support for that into our runtime. And you're right, like it is really interesting the way that all of these open source labs borrow from each other. For example, like GLM-5.2 doesn't have vision. So something that, Haley, a guy on our team, if we could take a look at this, he, like, grafted the Kimi vision encoder onto GLM-5.2.Retrofitting Vision into GLM-5.2Ali [00:15:27]: We'll be training the projector.Philip [00:15:28]: Exactly. So if you think about, like, the encoder, there's the encoder, which is the part that looks at the image and turns it into latent information, and then there's the projector which likeAli [00:15:38]: You can say latent space. It's okay.Philip [00:15:41]: And then there's the projector that maps it onto, the model itself, and then there's the model weights. You don't wanna mess with the model weights because you run a chance of making the model dumber at something else for the purpose of giving it vision. So instead, Haley started with just a projector, which is only a handful of millions of parameters.Ali [00:16:02]: That would be, yeah.Philip [00:16:02]: Yeah.Ali [00:16:03]: Can you show the training one?Ali [00:16:04]: Like the way it groksPhilip [00:16:05]: YeahAli [00:16:06]: Very interesting.Philip [00:16:06]: And maybeAli [00:16:07]: That right therePhilip [00:16:07]: Maybe Ali, you should take it from here. You've got a betterAli [00:16:10]: Ooh, double the sandPhilip [00:16:11]: Understanding of this than I do.Ali [00:16:11]: Yeah. You can see, like, he. The way he trained this is really cool. At the beginning, he was training it using just like, “Here's a picture of a mountain. Can you describe what's in this mountain?” And that caused it just like the first, learning walls. Like here you can see this all we're trying to teach it is to translate the encoded. Like it's already taken the encoder from Kimi K. It's taken the image. It'Philip [00:16:31]: Yeah. FrozenAli [00:16:31]: FrozenPhilip [00:16:32]: With adapter.Ali [00:16:32]: Exactly.Philip [00:16:33]: Yeah.Ali [00:16:33]: So the brain is frozen and the eyes are frozen. It's just we're tryingPhilip [00:16:37]: AlignAli [00:16:38]: Interconnect between the eye and the brain, right? So the projector. And so you take the tokens and then he's like, “Oh, can you describe what's in this image?” And he's like, “Oh, it's a mountain,” or it's a person or it's a human, whatever the case is. But that didn't cause complete understanding. So he changed it such that every image was associated with a data set of questions. Like, does this image have a white male? Does this image have birds in the top corner? Does this image have a scientist in it? All of that stuff. And it would have to answer questions correctly. And using not just training on describing an image, but being able to answer question, another question, answer over time. Like you can see the grokking, which is like genuinely insane, that retrofitting vision into a large LLM can learn to that extent. And even for images that it doesn't perform well on, for instance, if you ask it a picture of like Stephen Hawking, “Who is this?” Maybe it doesn't get it, but it will say something like, “This is Albert Einstein.” Like it still understandsPhilip [00:17:25]: Close enoughAli [00:17:26]: That this is a scientist who is a man who has, some significant achievements, all that stuff. So that's like really cool.Philip [00:17:32]: Yeah. So, we've covered Hao Tian before, who the author of the LLaVA paper that did this, a while ago. And I think that's very foundational work for anyone who hasn't done vision work before.Ali [00:17:41]: Same with the CLIP and MetaCLIP, where you go from just captioning to building out questionsPhilip [00:17:47]: RightAli [00:17:47]: Off the image and how much better you can get performance.Philip [00:17:50]: Right. Right. Right. Yeah. But what's, what's so exciting about this is if you look at a model like this. Now, this is a little bit more of a research project. It's not. It got to 56% on MMLU Pro, I think. So not quite frontier. But if you're running this model, you haven't suffered any loss on your GLM-5.2 quality. If you don't have an image, it'll just behave exactly the way it used to. And ultimatelyAli [00:18:14]: Which in the inference code you literally do not include the other part, right?Philip [00:18:18]: Yeah. You would just skip the encoder if you don't have an image input.Ali [00:18:22]: Okay.Philip [00:18:22]: Just confirming.Philip [00:18:23]: YeahAli [00:18:23]: Does it affect a lot on the overall inference side? Like you're not adding much, you're adding a very small vision encoder. These are typically likePhilip [00:18:30]: They're super fineAli [00:18:31]: Less than a billion parameters, right?Philip [00:18:32]: Yeah. It's, - There's a little bit less standardization among vision encodersSwyx [00:18:37]: YeahPhilip [00:18:37]: So the support matrix can be a little bit, sparser. But overall, yeah, it's a pretty, it's a pretty minor component of the overall system. And ultimately what you get out of the system is all of a sudden you have Kimi Vision, GLM weights, and DeepSeek attention all in one model.Open Source Model Grafting and Franken-MergesPhilip [00:18:56]: And that's, I think, a lot of the power and beauty of open source, is that you can take all of these different components and combine them together into a system that's better than anyoneSwyx [00:19:05]: YeahPhilip [00:19:05]: Can be individually.Swyx [00:19:06]: People used to say that you would also do Franken-merges where you would take likePhilip [00:19:10]: YeahSwyx [00:19:10]: Layers from each model.Swyx [00:19:11]: Does anyone do that anymore?Ali [00:19:13]: Well, to your point previously when you were mentioning like, the work that goes into supporting a model when it first comes out, like GLM-5.2 or MiniMax M3 or whatever the case is. Sometimes you do have to like, you do have to switch out some things. Like, for instance, the MiniMax M3 head uses full attention, and with full attention you end up with this like insane bottleneck in spec dec ‘cause you're doing auto-regressive token generation for three tokens, and you're doing this like N squared over all of the tokens that are in your sequence. Your KV cache is like very large because it's not sparse, it's not top K. So we find it better to like, okay, we're gonna replace this, we're gonna replace this layer with a layer from another model that's using like GQA, for instance. And then just with the right training, you can get it to have the same acceptance rate. So it is very possible to retrofit layers from other models and very much needed. If a layer is like inefficient, the training just becomes the challenge, like how do you ensure that you train it properly? Which again to your earlier point is like the mesh between training and inference. As in like you need very good training in order to do fast inference. That's like, I feel like more and more becoming true.Swyx [00:20:21]: Yeah. Anything else on the support side when you say like get it to fully production ready?Loop Detection, Race Conditions, and Non-DeterminismPhilip [00:20:26]: Yeah. I think that there's also a question of just, we can test a model to a pretty extensive degree, but we're trying to get it out quickly and then you see a bunch of other people test it and you get interesting results. There was an issue with, GLM briefly where we had some like mode collapses where it would just output the same token over and over again for certain prompts on certain temperatures. Like once you expose an endpoint to the real world, there's going to be, so many more varieties of things given to it that you're able to, discover and patch things. So it's not just a, day zero process, it's then like for the first week, for the first month, if a model remains popular, like how do you both fix bugs and then continue to push the envelope on performance?Ali [00:21:21]: What do you mean you don't want your model outputting S?Swyx [00:21:24]: Is there loop detection on that stuff, by the way? It still happens like quite a lot, which is surprising.Ali [00:21:30]: We have like we, in our endpoint, like if a model was to output the same token like four plus times, we just cut the generation. We say like, “Oh, sorry, this-- Like try again,” or like we will reprocess the request. ‘Cause we know then, like if it, like if, yeah, it's four times the same token, it's probably collapsed.Swyx [00:21:45]: Yeah. Is there a way to opt out in case I really want that?Ali [00:21:48]: You want that?Ali [00:21:50]: I think there's a way that we have to handle it. I'm not exactly certain, but I feel like in certain models, like when they output something like you can imagine, like a table for instance, and so they want, they wanna draw like 12 dashes and 12 dashes. Yeah, I think there's a way for that to happen. I think we only do it on certain tokens. Like we exclude certain special characters.Swyx [00:22:07]: Yeah.Ali [00:22:07]: So we only do it on like certain like S is the most common almost. GLM-5.2Swyx [00:22:11]: OhAli [00:22:11]: And I think it was DSV 4 as well. Like you'd just have like looping issues where like you literallySwyx [00:22:17]: ItAli [00:22:17]: Just have like S.Swyx [00:22:18]: Yeah. Is there a special, something special about S? No, just randomlyAli [00:22:21]: It just seems to be the one token involved.Swyx [00:22:23]: Yeah. And it'Philip [00:22:24]: Is thereSwyx [00:22:24]: And it's only temperature 0Ali [00:22:27]: NoSwyx [00:22:27]: Even at other temperaturesAli [00:22:27]: Even at like 0.9 or whatever, it will still, it will still collapse.Swyx [00:22:30]: That's weird, right?Ali [00:22:30]: It's, it is an inference problem to be honest, like a software problem. Like oftentimes, the image you run will-- like NVIDIA will release an image for instance, and if we will upstream the changes from their latest TensorRT-LLM image into our stack, we'll find that it fixes it. Or oftentimes this will only happen in an inference engine that you're using like SGLang. But if you were to switch to vLLM, that isn't the case. So it seems to be like an extremely like deterministic software issue and not really a model issue. It's not like a weights problem. Like I'- we'll say like, “Oh, it's a problem with the quant. We did PTQ wrong,” right? But that isn't, that doesn't make sense because the same weights used with a different inference engine does not repeat the problem. And sometimes it's, the kernels that are being used in the backend have like these very subtle sometimes race conditions, where if you were to use this model hosted on one cluster, you will never get this problem.Swyx [00:23:19]: Oh my God.Ali [00:23:19]: But if you host it on a different cluster, you will. And the reason is the KV cache transfer from a node to node in that one cluster is using a slower interconnect than the node to node in another cluster. So that exposes the race, whereas in another cluster it doesn't. So then you end up just like, okay, this model is not gonna be hosted on this cluster. We're gonna host it on, another cluster because that cluster exposed that problem. But then it ends up with like, okay, is it the software? Is it the model weights or is it the hardware?Swyx [00:23:42]: There is a thing about this with temperature 0 still not being deterministic, right?Ali [00:23:46]: Right.Swyx [00:23:46]: Mostly because of hardware. Even at temperature 0 same model, you won't always get the same output.Swyx [00:23:52]: Even-- But I'm surprised by the race condition one because, I thought PyTorch was a graph that like guarantees that you at least, execute things in the right order.Ali [00:24:02]: Well, yeah, true. Like I'm not, I'm not saying that there is. Like well, you have things like PTL optimizations where like you can start a kernel before the end of the previous kernel, and that's like ‘cause you want to do that because there'sSwyx [00:24:12]: It's like pipeliningAli [00:24:12]: Expense. Exactly.Swyx [00:24:13]: Yeah.Ali [00:24:13]: But it'- But you don't do it cleanly. Like you overlap a little bit of the execution. No, it is very possible that the kernel itself, like that one block that is supposed to be running in this instance of time, that kernel itself has a race condition. For instance, like a missing barrier. Like often if you're designing a kernel and you want it to make it to be very fast, if you don't test it extensively, you'll, you'll have certain threads access data points from registers before they've been written to by other threadsSwyx [00:24:36]: YeahAli [00:24:36]: For example, because like your barrier is wrong or your synchronization was wrong. But yeah, like the testing itself is very difficult in those like, andSwyx [00:24:42]: And there's no like borrow checkerAli [00:24:45]: What does that mean?Swyx [00:24:46]: Like Rust. Like the. If you're trying to have like memory safety It sounds like a comparable problem.Ali [00:24:52]: Well, yes, but you're working in CUDA, right, NVIDIA GPUs. Like- You just need a higher level language like modular Maybe that's what modular is supposed to do. I don't know.Quantization Quality and Vendor FidelityVibhu [00:25:00]: How do you see keeping quality of the model? So you talked about all these steps of, okay, you gotta do quantization, train your own speculative decoderAli [00:25:07]: RightVibhu [00:25:07]: Run on different hardware. Looking at other model providers, okay, you kicked off a inference speed race on the consumer end. What goes into keeping quality the same across them, right? Sure, you can run benchmarksAli [00:25:22]: YeahVibhu [00:25:22]: But, like, how do you determine how much quantization are there standards? What goes intoPhilip [00:25:27]: There's a few things on quality. Most inference optimizations are lossless. KV caching, for example. You are just recomputing or preventing recomputing the same values. Speculation, of course, if a draft token is wrong, it gets rejected. The main lossy optimization is quantization. And that really comes down to, number one, data format, number two, which parts of the model you choose to quantize, which layers, and number three, like doing a lot of calibration on the quantized weights, to ensure that you're preserving all the outliers. There's other tricks that you can do, though. A big one is long context, ‘cause one thing you asked at, right at the beginning is, “Oh, what's gonna happen if I send a 200,000 token request in?” So with a long input sequence, you need to, store a lot more information. You need to process a lot more tokens. And so even if a model has a context of a certain length, you might, as an inference provider, choose to build an API with a shorter context length, and of course a full length one as well. Because if someone doesn't need the full million token context, for example, you can get them better performance. I don't know if that's exactly like quality of the model. The way that I think about quality is to what degree are we faithfully serving the original model? If you think of a golden implementation of a model that performs exactly the way the model is designed to perform, I think of quality as how close are we getting to that, 100% fidelity of the model.Philip [00:27:13]: You can also, of course, think about quality from the training side and how do you push yourself past 100%. But when I think about purely inference optimizations, it's getting faster while staying as close to that 100% fidelity mark as possible. And certainly our standard internally is that, like you should not be able to tell the difference between our API and a, official API. I think Kimi in particular does a good job of vendor benchmarking hereAli [00:27:41]: YesPhilip [00:27:41]: Where they haveAli [00:27:42]: They released an actual vendor benchmark.Philip [00:27:43]: Exactly, yeah.Ali [00:27:44]: ‘Cause they accused, some people, Amazon? There was some provider that was not doing very well on Kimi's benchmark.Philip [00:27:50]: Yeah.Philip [00:27:51]: So, with Reflect we probablyVibhu [00:27:52]: This was a long time ago, right?Philip [00:27:54]: No.Ali [00:27:54]: Yeah, like threeVibhu [00:27:55]: They alsoAli [00:27:55]: Four, five months agoVibhu [00:27:57]: This also happened with, I don't remember which model, but they pulled out quite a few, and then they started a whole chart about this. It might have beenPhilip [00:28:03]: Kimi Vendor Verifier.Ali [00:28:04]: Yeah.Philip [00:28:05]: Yeah.Ali [00:28:05]: Yeah, ‘cause you, ‘cause you'd be pissed, right? Like if you'Philip [00:28:07]: Yeah.Ali [00:28:07]: If like if I'm a consumer and I'm using like Amazon's endpoint for instance, and I've used Kimi and I'm like, “Oh my God, like this is bad,” I'm not gonna say, “Oh, Amazon quantized the model in a bad way.” I'm gonna say, “Oh, Kimi sucks.” Right?Philip [00:28:17]: Yeah.Ali [00:28:17]: So it seems like that makes sense.Philip [00:28:19]: Yeah, they care. They care.Vibhu [00:28:21]: Justifiably.Ali [00:28:21]: Yeah, justifiably.Vibhu [00:28:22]: This is probably a stupid question, but just checking, has anything improved from main quantization?Philip [00:28:28]: Yeah.Vibhu [00:28:28]: Like, is quantization always strictly worse?Ali [00:28:30]: Well technicallyVibhu [00:28:32]: NoAli [00:28:32]: It's a lossy. QuantizationPhilip [00:28:33]: YeahAli [00:28:33]: Is a lossy, it's a lossy implementation.Philip [00:28:36]: Speed improvesVibhu [00:28:36]: Speed improves.Ali [00:28:37]: It the number, likeVibhu [00:28:38]: No, I' always look for inverse scaling laws.Philip [00:28:40]: Yeah.Ali [00:28:40]: Yeah.Vibhu [00:28:40]: This is something I learned from Noam Brown, where like things that normally act in one direction sometimes do.Philip [00:28:45]: Well, technically when you run a benchmark, because these models are deterministic, sometimes your,Ali [00:28:52]: YeahPhilip [00:28:52]: NVFP4 quant is like, two basis points higher than yourAli [00:28:56]: No, it's noise. It's noise.Philip [00:28:57]: Yeah, exactly. I'm like, yeah, it's, it's within. That's why I always say within margin of error.Philip [00:29:01]: And I stopped saying that because everyone assumes that what is, well, within some margin of error, we're barely inside of that to the worst, so we're saying. But yeah, sometimes it's just like, gives you a higher output score. But like Ali said, that's noise. To my knowledge, you're not necessarily making the results better. You're just trying to, again, like keep your fidelity as close to 100% to the original model.Layer Selection, KL Divergence, and Better QuantizationAli [00:29:27]: There is, to your point, research that we did on MP. I don't know if you are able to pullPhilip [00:29:31]: YeahAli [00:29:32]: A tweet we did. One of our research interns, Joshua, I think it's a tweet on how we have 20% better quantized GLM-5.2 than NVIDIA. Essentially what we found throughout like this month research is, okay, quantization is a lossy. It's. You're compressing the data from, occupying 16 bits to occupying, four bits, for instance. And so you're losing some information, and you're trying to minimize that. And so when I say that I'm gonna quantize the model, my job becomes how do I find the layers that I can quantize, and how to find the layers to not. For instance, with image models, I don't quantize modulation layers, and I don't quantize out projections because those two are. Like out projection is what you see as the user. Modulation is what the model sees or understands. Right, exactly. And so to his paper, do you have the. It doesn't have the. Yeah. It's a long paper. I don't know if I can findVibhu [00:30:25]: If there's a part to search or it's probably in the thread.Ali [00:30:28]: It's probably in the thread.Vibhu [00:30:29]: Yeah.Ali [00:30:29]: But the long and the short is it is very possible that quantizing more of the model makes the results. Like if I have a model that I quantize layers one, five, and 10, and another model where I only quantize layers one and It is possible that the model in which I quantized more information is going to perform better because the quantization errors have canceled out. And so what Joshua showed in his mathematical proof where he had like a verifier in, is that you can predict which layers are going to have quantization errors that will cancel out with each other, and you choose to quantize those layers. And so the result of doing this mathematical quantization is you end up with a model that's 20% more quantized than another provider, so you get 20% more throughput of it because there's more layers than running an NVFP4, and your quality is better than that other quant because the layers that you chose to quantize have their errors cancel out, like one layer skewed to the right one layer skewed to the left, one layer skewed to the right. Your final logits distribution is more similar to the original distribution of the model, so you have better fidelity. And so the way we proved this was with KL divergence. So instead of just scoring on the benchmarks, we scored the KL divergence between the logit distribution of the quantized model and the logit distribution of the original full precision model, and we showed that with this technique we get. If your probability distribution on the logits which token it wants to select is more of the same as the original model, you're probably gonna end up staying true to the original model. So yeah, so it seems like previously before this, it seemed like the industry was, well, the more you quantize, the worse it's gonna be, ‘cause the more loss you introduce. That's not exactly, not necessarily true. So yeah, doesn't improve it, but can cancel out.Philip [00:31:57]: I think it might be this, but reminds me a good bit about pruning where you can prune off certain layers.Philip [00:32:03]: But very interesting. Didn't know this was a whole paper you guys put out.Ali [00:32:06]: It's. Fun fact, it was originally 72 pages, this paper, and then we decidedPhilip [00:32:11]: WowAli [00:32:11]: We can't tell. We couldn't release it. So it's now 45.Swyx [00:32:15]: Still 39 pages, so very substantive. We talked about evals and all these things and, like what's possible in terms of speedup? Like it's like probably like the numberInference Speedups and BenchmarkingSwyx [00:32:25]: Thing that people do wanna care about, and it's something that you wrote about in your post. Like official API is 70 tokens per second, and you push it up to 90. Is that like a normal thing?Philip [00:32:36]: So what's cool about working in inference, the reason that I think inference is going to be a useful place to do engineering for a long time, is that if you look at highly optimized domains like, say, finance, if you're in finance, you measure how much better you got in basis points. It's like, “Oh, I got five basis points better, like twentieth of 1% better,” that's huge news because everything is so optimized. When we publish optimizations, it's 20%, it's 100% it's 200%. So there's still probably like a lot further to go, honestly. Like you'll, you'll know that inference is pretty much solved when researchers start publishing about how they got 1% faster at something.Swyx [00:33:19]: Which by the way, because I am from the finance background, in the ‘70s, that was the margin at the time. When you did quantitative finance research, you would findAli [00:33:27]: And like 20%, tens of percent.Swyx [00:33:29]: That's. Yes.Philip [00:33:29]: Yeah.Swyx [00:33:30]: And now it'Philip [00:33:31]: Tiny fractionsSwyx [00:33:32]: For those people interested, look up Andrew Lo's paper. He had a really interesting illustration of quant, stat arb, distribution, narrowing down from like those kinds of 20% differences in the ‘70s, down to nothing today, which is very cool.Philip [00:33:48]: Exactly, and we're at the beginning of the same type of thing. Now benchmarking is hard. I think anyone will tell you that, and benchmarking provider speeds is hard because there's so many variables that go into it. What hardware are you using? How much load do you have on the system? What's the exact nature of the prompts and input and output sequence lengths? All that stuff. But overall, when you start stacking these improvements, you're looking at multiples. You can look at it. The most common form, of course, is TPS, tokens per second, which is bad naming by us in the industry, ‘cause there's two tokens per second. There's tokens per second, the throughput number, and the latency number.Ali [00:34:31]: TTMT, yeah.Philip [00:34:32]: Like total tokens per second out of the, out of the GPU as a throughput number. Most people only care about tokens per second as the latency number, which we should call ITL, intertoken latency, but we don't.Philip [00:34:44]: Anyway, so you can imagine a standard API without many optimizations for a 1 trillion parameter model operating somewhere in the 30 to 50 tokens per second range for reasonable traffic profile. And we generally see the goal of, pushing to 10X that. But, not necessarily day zero, but by stacking enough optimizations, if you have, say like four optimizations, each of which doubles performance. Or sorry, three optimizations, each of which doubles performance, then you stack that up, that's an 8X gain. That's the order of magnitude that we're working with in this space. We're trying to make things substantially faster, not just go from like 70 to 90.Swyx [00:35:38]: Are you saying you've. You have done that?Philip [00:35:40]: So let's say you have as a reasonable baseline, 30 or 40 tokens per second. You can achieve 10X that. So like on GLM-5.2, if you run it unquantized, perhaps on H100s even, and you're just using an off-the-shelf inference engine with no particular optimizations, no speculator, nothing extra around like KV routing, no disaggregation, you're, you're probably, yeah, looking at that like 30 to 40. You think that's like a reasonable baseline?Swyx [00:36:12]: Right. Right.Philip [00:36:12]: To get to something like 10X, there's a lot of trade-offs that you're making. If we're running at more like a 300, 400 tokens per second range, you are using the best hardware possible. You have a optimized speculator. You have done all of your quantization work. You are Seeing a pretty high cache hit rate. You are running with a reasonably small batch size and a parallelism configuration that is tuned for latency versus throughput, but it is possible. So the spreads that you see if you, like, go on artificial analysis or you go on OpenRouter and you look at, the worst provider to the best provider, oftentimes can hit that range. 10X is of course very aggressive. It's oftentimes maybe more of a four to six times improvement. But that's the performance that makes us really excited, is when we can get these huge gains, not just go from 70 to 90 tokens.Stacking Optimizations: NVFP4, Speculation, and DisaggregationAli [00:37:19]: It's also, like, hardware dependent. Like, ifPhilip [00:37:20]: YeahAli [00:37:20]: If you have a thing where you're serving it on just, like, a node of H100s and then you throw, like, you shard the model across, like, four nodes of B200s. Like, you can definitely increase the speed with just throwing more hardware at it. Like, normalizing for the same exact hardware and the same number of GPUs.Philip [00:37:35]: Yeah. Then you're looking at, like, a two to 4X improvementAli [00:37:38]: Right. RightPhilip [00:37:38]: Depending on the inference optimizations. So yeah, it's. Some of it's, what's the call, and some of it's who's the driver.Vibhu [00:37:46]: If you break down the two to 4X, say the example is run GLM-5.2Ali [00:37:51]: YeahVibhu [00:37:51]: On B200sAli [00:37:53]: YeahVibhu [00:37:53]: Single node, right? What's, like, the cost trade-off for effort to get, like, the last bit of juice out versus what should people just think of, right?Ali [00:38:01]: Spectre quantization. Yeah.Vibhu [00:38:03]: Spectre quantization.Ali [00:38:04]: That's, that's, that's like 95%. LikeVibhu [00:38:06]: And how far does that get you? And how easy is that for the average person to do? So say right I wanna throw the weights of GLM-5.2 on a node of B200s, how easy is it to find speculative decoder- decoder model or already quantized model? How much work goes into it?Philip [00:38:23]: If you're doing it up front, it's quite a lot of work. If you're doing it today, there's going to be people who have published things that you can just, you can just grab some NVFP4 weights. You can grab a speculator. Yeah, if we're thinking about, like, what are the 2Xs we're stacking, going from, BF16 to NVFP4 is, it's not quite a 2X, right? It's like. I think it's about, like, 30 to 40%, from 16 to 8, and then another 30 to 40% multiplied from, 8 to 4. So that doesn't quite get you a 2X, but, like, roughly a 2X. Speculator, roughly a 2X. Disagg on top of that if you're able to get enough hardware and put enough traffic through it, another roughly a 2X. And then you add in some, double-digit percent increase from having just a better runtime with, the latest kernels and stuff behind it. And that's how it stacks up.Ali [00:39:21]: YeahPhilip [00:39:21]: So building each of those, like, building the, quantized weights is, for someone who really knows what they're doing, hours to days of work. Building the speculator, again, like, hours to days of work. And the, disagg setup, hours to days. Well okay, but like once you haveAli [00:39:39]: Once set up. Once set up. YeahPhilip [00:39:40]: Yeah, getting disagg working for the first time, I'm saying, of course, is very difficult.Philip [00:39:44]: The marginal implementationAli [00:39:48]: Like, if you're just grabbing, like if you are a person, like just a normal consumer who has access to, like, a node of B200s and you're wondering, “How can I just host it myself?” You don't need to quantize the model yourself. There's always gonna be, like, an open source quantized checkpoint. NVIDIA's gonna push one out if no one else does. You. Usually, the providers will have their own spec dec that they've trained as well. You don't need to train your own spec dec. You can just use that as well.Philip [00:40:09]: Yeah. Like, GLM-5.2 has its own MTP.Ali [00:40:13]: Right. Right.Vibhu [00:40:14]: What's multi token prediction?Philip [00:40:15]: Yes.Ali [00:40:16]: I'm justVibhu [00:40:16]: Can you explain that?Ali [00:40:16]: I'm just an expert.Ali [00:40:18]: I can do it for you in case I get it wrong?Vibhu [00:40:20]: No.Vibhu [00:40:21]: Yeah, you should correct if we're wrong, but their multi-token prediction can be used for self-speculative decoding.Ali [00:40:27]: I'm not sure. I'm not gonna correct that.Vibhu [00:40:28]: Okay. I'm semi-confident in thatAli [00:40:30]: Okay. YeahVibhu [00:40:30]: But someone can check. But it's useful to paint the story of, okay, not just the average person, but say a company wants to switch from serverless inference I wanna throw this up on. I wanna rent some GPUs, throw it up. These are the steps you take to do significantly faster than just put it behind vLLM.Ali [00:40:48]: Right.Vibhu [00:40:49]: I was waiting for a mention of Dynamo.Vibhu [00:40:51]: I feel like, that's supposed to be the baseline that you measure against.Dynamo, KV Routing, and Disaggregation ToolkitsPhilip [00:40:55]: I would think of Dynamo as less of a box system and more of a toolkit for building with. So when we talk about doing aware routing, when we talk about doing KV offloading, when we talk about doing, PD disaggregation, Dynamo fundamentally is. By the way, Dynamo is an open source library from NVIDIA.Ali [00:41:17]: We've done a pod with KylePhilip [00:41:18]: OkayAli [00:41:19]: Kyle Cranin.Philip [00:41:19]: Cool. So then your listeners know then that it supports all the different inference frameworks. And it is multi hardware, which is interesting.Ali [00:41:28]: But it's just a router, it's not like an optimizer layer.Philip [00:41:30]: Yeah. All it does, like, what Dynamo is good at, it is a library for moving information around your cluster, around your hardware. So if you have, KV cache on one place and you need it to be somewhere else, Dynamo coordinates NIXL for you to move that around.Philip [00:41:49]: That doesn't mean that, like, out of the box, you just say, “Pip install Dynamo,” and then you get, like, a massive performance speed up. It's more of a developer toolkit.Ali [00:42:01]: Yeah. I would have said it would. It comes with a set of defaults that you can then swap out.Philip [00:42:06]: It does. If the industry at large, I think, was, like, rolling out all of these deployments, standard, then I think it would be, like, a credible baseline. But, we've got to, we've got to benchmark against, like, what we're seeing in the wild.Speculative Decoding Methods: Medusa, EAGLE, n-Gram, and Spec-SpecVibhu [00:42:23]: I did wanna talk a little bit more about PD disagg, because that is probably, like, number three after quantized and speculative decoding. In your book though, I was just gonna pull out the book.Philip [00:42:31]: Yeah.Vibhu [00:42:32]: Like section 522 on Medusa, 523 on EAGLEPhilip [00:42:35]: YeahVibhu [00:42:36]: 524 on gram.Philip [00:42:37]: It's 55, would be disaggregationAli [00:42:42]: Yeah. Well, no, I just wanted to dwell a little bitPhilip [00:42:44]: YeahAli [00:42:44]: The other. Like, so what do you choose to include? What do you choose to not to include? Because there was all these other techniques.Philip [00:42:51]: Yeah.Ali [00:42:51]: Are these still relevant? Because I think they came out, like, a year and a half ago maybe.Vibhu [00:42:55]: Medusa is quite old.Philip [00:42:56]: Yeah, Medusa's old.Ali [00:42:58]: It was old.Vibhu [00:42:58]: But is it in the book as a good, here'sPhilip [00:43:01]: BaselineVibhu [00:43:01]: Baseline vanilla understand it?Philip [00:43:02]: Like you should know this.Vibhu [00:43:03]: Like I read the paper, I'm like, “ it makes so much sense.”Philip [00:43:05]: Yeah.Philip [00:43:05]: So with the book, I had a couple goals. One was to give people just a working vocabulary for the space as a whole, and the other was to give them some intuition about how each of these techniques works. As I mentioned in my AI Engineer talk, which is the first public addendum to this, the speculation space has moved much faster than everything else. So yeah, even at the time that I wrote the book Medusa, I very much included as a way for people to understand how the space evolved rather than what the most modern technique is. And now of course, there's DFlash, dSpark. There's, there's newer techniques even than EAGLE, although EAGLE is still very commonly used.Ali [00:43:51]: SpecSpecta.Philip [00:43:52]: Yes. Speculative decoding.Vibhu [00:43:54]: What canAli [00:43:56]: Oh, it's a paper by Tri Dao and it's like, it's doing speculative decodingVibhu [00:44:00]: HuhAli [00:44:01]: For the speculative decoder.Philip [00:44:02]: Oh, in spec- oh my God.Ali [00:44:02]: It's literally just an another. It's like, yeah, that's the most simple way to explain it, and it seems like he got trivial speed ups there. But it seems that the complexity with training, it's almost like in our mind at least, it's almost as complex as training GANs. Like it's like a very delicate balance and oftentimes you, it's just but yeah, it's literally speculative decoding on speculative decoding.Vibhu [00:44:21]: Speculative.Ali [00:44:22]: Yeah. We saw this paper.Vibhu [00:44:24]: It's interesting, right?Ali [00:44:24]: Yeah.Vibhu [00:44:24]: I wouldn't even expect it to be very particular to train, I wouldAli [00:44:29]: Right.Vibhu [00:44:29]: The naive part of me is like, okay, train speculative decoder.Ali [00:44:32]: But like, and it makes sense, like the whole idea of speculative decoding is you. It's like, it's like almost like the iPhone auto predict version but for a normal model, right? Like you're just, you're just, generating three tokens and you're like, okay, I'll do prefill on them. And so you save those three turns for your original model. Now your speculative decoder is doing three turns of auto regression, so why not just have an even smaller model?Ali [00:44:53]: The other question there is what are the size of speculators? So say forPhilip [00:44:58]: Right. It's like a billion parameters.Ali [00:45:01]: Like for MiniMax, it's. Yeah. It's like one layer. It's like one 60th of the original model usually.Philip [00:45:06]: Yeah. I think we should do a paper when we get back to the office.Philip [00:45:10]: SpeculativeAli [00:45:11]: SpeculativePhilip [00:45:11]: Decoding.Ali [00:45:13]: No, it's, it does seem like how, when do you stop? But then it also seems like if you're able to train spec-spec decode for instance, right? Like if you're able to have a small model that is accurately predicts what the intermediate speculator is gonna predict, that is able to predict what the original target model's gonna predict, then why not just use that smallest model directly, right?Vibhu [00:45:34]: Yeah. This isAli [00:45:35]: Like it seems likeVibhu [00:45:35]: Adjacent to the routing problem.Ali [00:45:36]: Right.Vibhu [00:45:36]: Yeah.Ali [00:45:36]: Right.Philip [00:45:37]: The thing with speculators is one of the practical constraints on using them is that you do have to run a small model on the same hardware that you're running the big model on. There is a orchestration and resource competition problem inherent in that, and that is one of the constraints on speculation in general, is that draft tokens cost resources to create and cost software complexity to manage. And so if you have like infinitely recursive speculators, you add in quite a bit of that complexity on the actual implementation within the inference engine as well, not just in the training process.Vibhu [00:46:17]: I was gonna say, I would wonder if you could do similar, like distillation and pruning of, it's the same thing, it's just a model. Can we not just distill a lot of the weights, quantize the speculator, out of my domain? The question that also comes up is, this is all for big server workloads, right? How much of this applies to, say I have this MacBook, I wanna run Gemma really efficiently. Similar problems, not the same?Local AI vs. Data Center InferencePhilip [00:46:45]: Pretty different. I talked to Selo, about this on his podcast a couple weeks ago. The difference between inference engineering for the data center and for production workloads versus inference engineering for local AI, is that we start with fundamentally like different constraints and different goals. With local AI, it's how do I fit this model onto my hardware and then make it less dumb? And with data center influence, it's how do I load this model and then make it less slow? And we care about less dumb, and they care about less slow. But the local AI inference engineering ecosystem, I think has a lot for us to learn from in the data center space. They are experts in various forms of quantization, including dynamic quantization that we just don't touch, in the pruning, in the distillation, in the, layer removal. There'Ali [00:47:42]: Layer removal matters less.Philip [00:47:43]: Yeah. There'Ali [00:47:44]: No one loves pruning really.Philip [00:47:45]: Yeah. Well, but the, but they doVibhu [00:47:46]: Which is surprising, right? But that's, that's a whole different thingPhilip [00:47:48]: Just to fit something on the laptop.Ali [00:47:50]: Right.Philip [00:47:50]: So yeah, it's a, it's an interesting, it's an interesting space. Not necessarily that like their techniques make sense for us to do in the data center, because we have different resources and different goals, but more that the process as well as the openness of that field is something to, admire.Ali [00:48:12]: Yeah. Like to your point, like, certain optimizations that would. Like for instance, Turbo Quantum Sharper, like it made such huge hype on that and we did like a whole deep dive on Twitter and like said, what is it? How does it work? Why is it good or not? And it took off and it was implemented on local devices because your memory bandwidth is so slow on like a MacBook, for instance. But try putting the same thing on like an NVIDIA GPU on a B200 Turbo quant would not be. Like, it would not be used. Like, NVIDIA - Like, NVIDIA made it clear that this is not a good optimization, and we've seen it firsthand where the overhead of doing dequantization, quantization of, in the kernel itself with turbo quant kernel, each end is much slower than the time that you save from doing the bandwidth. ‘Cause on the B200s, you have like 3.5 terabytes per second. You don't need decrease the storage that much. You don't need to do, FP4 KV cache. You don't need to use a requant. There's, there's, there's better optimizations to be made. But on Edge devices, it's extremely important, it's extremely useful. So, seems to be, like, different optimizations there, but then they're all uniquely combined with like all you wanna quantize the model, you wanna do speculative decoding, like certain common prefixes with bothPhilip [00:49:18]: Principles.Ali [00:49:19]: Yeah, exactly. Exactly. Exactly.Philip [00:49:20]: They also do a lot of work on, model parallelism, especially over, heterogeneous topology, where you have, some sparks and they are wired together with, Ethernet, DGX sparks.Ali [00:49:35]: Yeah, this is the Exo Labs guys.Philip [00:49:36]: Yeah. You have, a nu
In this episode of Roofing Road Trips®, Heidi J. Ellsworth sits down with Dale Nelson from Roof Hugger to explore the key differences between roof coatings and metal roof retrofit solutions. From understanding what a retrofit system is to discussing purlin strengthening, waterproofing, code compliance and long-term performance, they'll break down the factors contractors and building owners should consider when evaluating their options. Tune in to learn how the right solution can extend roof life, improve building performance and create new opportunities for roofing contractors. Learn more at RoofersCoffeeShop.com! https://www.rooferscoffeeshop.com/ Are you a contractor looking for resources? Become an R-Club Member today! https://www.rooferscoffeeshop.com/rcs-club-sign-up Sign up for the Week in Roofing! https://www.rooferscoffeeshop.com/sign-up Learn more about Roof Hugger here! https://www.rooferscoffeeshop.com/directory/roofhugger Follow Us! https://www.facebook.com/rooferscoffeeshop/ https://www.linkedin.com/company/rooferscoffeeshop-com https://x.com/RoofCoffeeShop https://www.instagram.com/rooferscoffeeshop/ https://www.youtube.com/channel/UCAQTC5U3FL9M-_wcRiEEyvw https://www.pinterest.com/rcscom/ https://www.tiktok.com/@rooferscoffeeshop https://www.rooferscoffeeshop.com/rss #RoofersCoffeeShop #MetalCoffeeShop #AskARoofer #CoatingsCoffeeShop #RoofingProfessionals #RoofingContractors #RoofingIndustry #RoofHugger
Assets under management in the impact and ESG space are close to an all-time high, somewhere between $1.5 and $1.6 trillion. So why is capital actually leaving the sector? In this episode of the State of Sustainability, host Saif Hameed takes a detour from resilience and volatility to dig into what's going wrong, and where the genuine opportunities still exist.The short answer: the industry has an identity problem. ESG labels get applied to mainstream tech stocks on the basis that they carry low environmental risk. Returns across impact funds are wildly inconsistent and the gap between what the sector promises and what it can actually deliver has become impossible to ignore.Saif traces how impact investing developed along two tracks. Private markets had early pioneers like the Acumen Fund building something genuinely mission-led. Public markets then borrowed the language through ESG frameworks, with rather looser results.Three lessons from that history:ESG was never designed to do this job. It emerged in the 1990s as a risk-assessment tool, not as a valid basis for investment inclusion. Retrofitting it into a mainstream strategy was always going to cause problems.The trade-off question needs an honest answer. You cannot simultaneously maximise financial returns and social impact without giving something up. The industry has spent years avoiding that conversation.Measurement doesn't scale. Comparing affordable housing projects with renewable energy infrastructure under a single performance framework produces numbers that mean very little.Where does that leave things? Saif argues the sector needs a significant rebrand and a serious recalibration of financial expectations. But there are three areas where impact investing still has real potential: generating brand equity and strategic value for corporate venture capital; venture philanthropy, where charitable capital gets recycled for compounding impact rather than disappearing into operational costs; and catalytic first-loss capital inside blended finance structures run by multilateral development institutions.The through-line: accepting below-market returns might be the only way to preserve what impact investing was actually supposed to be.What are your thoughts on this? I'd love to hear from you. Email Saif@altruistiq.comReady to transform your sustainability reporting? Start your journey at Altruistiq.comThis podcast is produced by The Podcast Coach.
Comments/ideas: ACFpod@outlook.comCooling is responsible for 15 per cent of global emissions and uses nearly two thirds of the electricity in commercial buildings. In this episode, Sam Ringwaldt from Conry Tech explains how modular micro units can cut cooling energy by 70 per cent and increase asset valuations by 18 per cent. We explore the rise of Comfort as a Service, the next generation of deep‑tech retrofits, and what this means for commercial buildings and AI data centres across the Asia Pacific region. It is a clear and practical look at why energy efficiency is becoming a financial strategy for the climate sector rather than simply an engineering decision.REF: Conry Tech, ABOUT SAM: Sam Ringwaldt is a Founder and the CEO of Conry Tech. Sam is an experienced industry leader, with 20 years of experience in building up HVAC companies, growing teams, and promoting new HVAC technologies worldwide. Sam was responsible for introducing Turbocor Technology into the North American and Australasian markets, driving its growth till it became today's dominant HVAC technology, and was able to lead both governments and the private sector to embrace the new technology, adjusting building standards, and driving new frontiers of sustainability and energy efficiency.HOST, PRODUCTION, ARTWORK: Joseph Jacobelli | MUSIC: Ep76 onward excerpts from Vivaldi's La Follia, played by Luca Jacobelli.
Data centers have moved from largely invisible digital infrastructure to a highly visible source of public debate as artificial intelligence accelerates demand for power, fiber, and compute capacity. The modern data center is now being built closer to population centers to support low-latency services, bringing critical infrastructure into direct contact with residential communities for the first time. This shift has elevated concerns around electricity pricing, land use, water consumption, and environmental impact—while policy frameworks and energy markets struggle to adapt at the same pace.The core issue driving today's tension is not simply whether data centers should exist, but how the costs and benefits of the modern data center are allocated. Do data centers represent a net burden on local communities, or can they function as a mechanism for modernizing the electric grid, stabilizing local tax bases, and expanding pathways into skilled technical work—if governed with the right market structures and incentives?That's the tension at the heart of this episode of Straight Outta Crumpton, hosted by Greg Crumpton, with guest Julia Chuang, Associate Professor of Sociology at the University of Maryland. Together, they unpack how media narratives shape public perception, why energy-market structure changes the “who pays” debate, and what it will take to train—and retain—the specialized workforce needed to build, retrofit, and operate the digital backbone of the AI era.What you'll learn…Energy prices aren't a universal data-center story—they're a market-structure story. Chuang explains how regulated, vertically integrated utility markets (like Virginia) create a perception of “free riding,” while more deregulated states can allow data centers to bring power on-site, build microgrids, and even sell power back—changing the public cost equation.The jobs debate is real, but incomplete. Data centers may not employ huge headcounts once operational, but the construction cycle can stretch 5–8 years for large campuses—and the bigger labor crunch is the shortage of specialized electricians, HVAC, and critical infrastructure talent trained for modern, high-density compute.Retrofitting legacy facilities is the next wave hiding in plain sight. The core constraint of upgrading older colocation sites is power provisioning. Many legacy designs were built around roughly 100 watts per square foot and cannot be scaled up overnight, because local transformers, feeders, and transmission capacity are often insufficient. As a result, operators are forced into creative hybrid approaches—combining limited high-density zones with lower-density legacy space—and, in some cases, consolidating power by acquiring neighboring leases.Julia Chuang is an Associate Professor of Sociology at the University of Maryland whose work focuses on institutions, groups, and how large systems shape behavior and outcomes. Her earlier research examined land use and industrial development in China, including factories, construction, and real estate—ground-level industries that, like today's data centers, reshape communities through capital, policy, and infrastructure. She now applies that lens to the U.S. data center boom, attending industry conferences and conducting interviews across the ecosystem to understand how data centers affect energy markets, local communities, and the politics of infrastructure.
In this episode, we speak with Bruce Fleming, CEO of Montana Renewables, the leading producer of sustainable aviation fuel in the United States. Rather than building a new SAF facility from the ground up, Montana Renewables converted a portion of an existing crude oil refinery in Great Falls, Montana to process feedstocks including used cooking oil, agricultural waste, and emerging crops into SAF, renewable diesel, and renewable hydrogen.Fleming discusses:The retrofit model: Why converting an existing crude oil refinery is a fundamentally different capital proposition from building a greenfield SAF plant.Feedstock agnosticism: Why the company is indifferent to which renewable feedstocks run through its system, including its claim to be the first producer to have made SAF from camelina, a cover crop that does not compete with food production and carries a very low carbon intensity score.The investment drought: Why private capital has effectively exited the SAF space, and the direct link between regulatory unpredictability and the absence of long-term investment. The book-and-claim efficiency case: Why separating the physical movement of SAF from its associated emissions certificate could save a dollar per gallon in logistics costs, and how a global book-and-claim system could accelerate SAF adoption.Fuelling small, not large: Why the immediate SAF opportunity lies in general aviation and regional operations rather than long-haul commercial carriers.If you LOVED this episode, you'll also love the conversation we had with Adam Klauber, Chief Sustainability Officer at World Energy, who shares how the book-and-claim model has evolved from a concept to a practical mechanism for scaling corporate demand for SAF. Check it out here.Learn more about the innovators who are navigating the industry's challenges to make sustainable aviation a reality, in our new book ‘Sustainability in the Air: Volume 2'. Click here to learn more.Feel free to reach out via email to podcast@simpliflying.com. For more content on sustainable aviation, visit our website green.simpliflying.com and join the movement. It's about time.Links & More: Montana Renewables Sustainable Aviation Fuel - Montana Renewables Montana Renewables and World Energy join forces to drive efficiency and scale in sustainable aviation fuel (SAF) deliveries - World Energy Montana Renewables launches MaxSAF™ Blended, accelerating SAF supply - ChemAnalyst
Season 6, Episode 9: Welcome back to a new episode of Keeping it Real with Dr. Kuehl. This week, Dr. Chris Kuehl talks to members about what else is going on in the world besides oil.ASA Chief Economist Dr. Chris Kuehl is back with his weekly economic update podcast. In Season 6, Episode 9 (10:58 in length), ASA Chief Economist Dr. Chris Kuehl talks to members about what besides oil is happening and what the ASA member should look out for.Construction - should ASA members be concerned?Is the data clear / positive within the construction sector?What is the average salary to live comfortably per state?Is remote work still motivation as where people are moving?What is affecting / driving residential growth?Where is the shift with Gen Z & millennials?Is there a surge with non-residential?Data centers - what is developing with these?Water - what is the demand? Is this a challenge to keep up?What is the controversy with data centers?? What makes people in the area testy?Retrofitting that is happening on the construction side - is this the death of offices?Office buildings are back - but made differently... why?Health care, why is this a big driver for the boomer generation?Ask Dr. Kuehl a Question!Have a question or topic for Chris Kuehl that you would like answered on this podcast? Email it to Brianna Dovichi at bdovichi@asa.net.
Ready to dive into the nitty-gritty of making your home earthquake-ready? Eric G and John Dudley are here to chat about the ins and outs of retrofitting your castle against those pesky seismic shakes. But wait, there's more! We're not just stopping at earthquake prep; we're also tackling the often-overlooked world of sprinklers—yes, those things that keep your lawn from turning into a desert! Whether you're fixing up a broken sprinkler head or dreaming of a whole new irrigation system, we've got tips and tricks that won't break the bank. Plus, we'll sprinkle in some witty banter about the importance of using the right materials to keep your home safe and sound. So grab your hard hat and your gardening gloves because we're about to get to work!Takeaways:Earthquake retrofitting isn't just for the 'big one'—it can save your home from all sorts of disasters, so don't ignore your foundation!Sprinkler systems can save you a ton on water bills, and with smart technology, you can program the whole thing from your phone—how cool is that?When it comes to retrofitting, hiring an engineer is a smart move; trust me, their plans will keep your house standing when the earth starts to shake!If your sprinkler heads are as old as your grandma's recipes, it might be time for an upgrade to avoid wasting water and money—nobody wants a lake in their backyard!Using the right fasteners in construction is crucial; those flimsy drywall screws won't save your house during an earthquake, so invest in some serious hardware!Don't skimp on your home's safety; retrofitting can even help lower your insurance rates, so it's not just an expense, it's an investment in peace of mind!Links referenced in this episode:regreen.comotolawn.comaroundthehouseonline.comaroundthehousehq.comCompanies mentioned in this episode:RegreenRain BirdOtolawn.comThanks for listening to Around the house if you want to hear more please subscribe so you get notified of the latest episode as it posts at https://around-the-house-with-e.captivate.fm/listenIf you want to join the Around the House Insider for access to the back catalog, Exclusive Content and a direct email to Eric G and access to the show early https://around-the-house-with-e.captivate.fm/support We love comments and we would love reviews on how this information has helped you on your house! Thanks for listening! For more information about the show head to https://aroundthehouseonline.com/Information given on the Around the House Show should not be considered construction or design advice for your specific project, nor is it intended to replace consulting at your home or jobsite by a building professional. The views and opinions expressed by those interviewed on the podcast are those of the guests and do not necessarily reflect the views and opinions of the Around the House Show.Mentioned in this episode:Subscribe to the podcast Make sure and Subscribe on your favorite podcast player or the link below! Podcast Subscribe 2026Check out our New YouTube channel @AroundtheHouse HQ Make sure you subscribe and RING THE BELL for our brand new channel with 4k content! Click the link to take you there! YouTube Around the House HQInstaBid: Stop losing jobs to slow estimates Turn 3 hours of manual estimating into 5 minutes. Real material prices. Real labor rates. Professional PDF quotes delivered instantly. Try it free at instabid.pro. Use code ATH50 for 50% off your first month. That's instabid.pro — code ATH50InstabidSubscribe to the podcast Make sure and Subscribe on your favorite podcast player or the link below! Podcast Subscribe 2026Take a second and leave us a review on your favorite podcast player! Quick favor—if you're enjoying the show, the absolute best way you can support us is by leaving a quick review on your favorite podcast player. InstaBid: Stop losing jobs to slow estimates Turn 3 hours of manual estimating into 5 minutes. Real material prices. Real labor rates. Professional PDF quotes delivered instantly. Try it free at instabid.pro. Use code ATH50 for 50% off your first month. That's instabid.pro — code ATH50InstabidTake a second and leave us a review on your favorite podcast player! Quick favor—if you're enjoying the show, the absolute best way you can support us is by leaving a quick review on your favorite podcast player.
Earlier this month a report from the ESRI found that we are “lagging considerably behind” our targets for decarbonising residential heat. Since then, we have also heard about the cumbersome nature of applying for grants to retrofit homes. So, is the current system fit for purpose? Or even worse, is it failing?Eamon Ryan, former Leader of the Green Party joins Ciara to discuss.
Brian Turner is the CEO of OTI, a master systems integrator that connects HVAC, lighting, and building systems into unified platforms to drive operational and energy efficiency, and ESG performance across the built environment. With over 30 years of experience in building automation, Brian has worked closely with major property owners to advance system integration and building data strategies. Prior to OTI, he led Controlco and later BuildingsIOT, where he helped shape the evolution of smart building infrastructure and IoT-driven operations. At OTI, Brian is focused on transforming how buildings operate by bringing together fragmented systems into a single, intelligent layer—making buildings smarter, more efficient, and easier to manage.(01:40) - What's OTI (02:32) - 30 Years of Building Technology(04:13) - The Retrofit Opportunity(05:55) - Why LEED Performance Slips(07:03) - Investing in Energy Technology(11:17) - A Unified Brain for Buildings(14:48) - How OTI Engages Clients(16:08) - Product Roadmap(18:55) - Regulation vs Real Market Drivers(19:59) - Comfort Across Asset Classes(25:27) - Feature: Blueprint: The Future of Real Estate 2026 in Vegas on Sep. 22-24(26:19) - Local Law 97(27:49) - When Owners Take Action on Energy(34:33) - Split Incentives in Leases(38:12) - Passive House Needs Operational Technology(40:20) - Wireless & Lower Retrofit Costs(44:49) - Where AI Adds Value(48:20) - Collaboration Superpower: Albert Einstein and Claude Shannon
To watch a video version of this podcast, click here: https://youtu.be/U0ALmS9vUC0In this episode, Reuben Saltzman and Tessa Murry talk with Sophie Ashley of Energy Vanguard about her journey from hands‑on carpentry and post‑Katrina rebuilding work to becoming an HVAC designer for high‑performance homes. Sophie shares how her field experience shaped her understanding of building science and why proper load calculations, ventilation strategies, and dehumidification planning are essential for modern airtight homes.The conversation also explores the challenges of open‑cell spray foam, moisture buildup in encapsulated attics, and what builders and inspectors often overlook in new construction. Sophie breaks down heat‑pump retrofits, electrification trends, and the importance of balancing comfort, durability, and system design—offering practical, science‑based insights for anyone working with or living in high‑performance homes.Here's the link to Inspector Empire Builder: https://www.iebcoaching.com/eventsYou can check out Energy Vanguard website here: https://www.energyvanguard.com/TakeawaysTight, high‑performance homes often require dedicated dehumidification, even in northern climates.Open‑cell spray foam allows moisture movement, which can raise attic humidity and impact roof decks.Proper HVAC design requires accurate load calculations, not rule‑of‑thumb sizing.Balanced ventilation (HRVs/ERVs) is essential in tight homes; Minnesota enforces some of the strictest standards.Retrofitting heat pumps into existing homes requires duct evaluation—it's not a simple swap.Many builder issues stem from overlooked details: attic access leaks, duct issues, missing covers, and ceiling‑plane air leaks.Electrification is growing, but homeowners must understand system impacts and design considerations.Chapters00:00 — Introduction02:00 — Sophie's Background & Career Path05:00 — High‑Performance Building & HVAC Design11:00 — Ventilation, ERVs & Climate Differences15:00 — Dehumidification in Airtight Homes17:00 — Moisture Problems with Open‑Cell Foam22:00 — Solutions: Conditioning Attics & Diffusion Ports26:00 — Heat Pumps, Dual‑Fuel & Proper Sizing31:00 — Electrification Trends38:00 — Common New‑Construction Issues47:00 — Field Lessons & Moisture Failures52:00 — How to Reach Sophie53:00 — Closing Remarks
Dr. Ciaran Byrne, Director of Retrofitting, SEAI, joins the panel of Paul McAuliffe, Fianna Fáil TD for Dublin North-West, Jen Cummins, Social Democrats TD for Dublin South-Central and Roderic O'Gorman, Green Party Leader & TD for Dublin West.
Ireland is nowhere near meeting its retrofit targets, that's according to a new ESRI report out today.So why are people not retrofitting?For more on this, Shane is joined by Claire McManus, Director of JFOC Architects and Housing Spokesperson for the Royal Institute of Architects of Ireland.
According to new Amárach research carried out on behalf of the Department of Enterprise, Tourism and Employment more than four in five businesses (85%) say sustainability is important to the day-to-day running of their business and have considered retrofitting . The findings of the second phase of SME Sustainability Research – Wave 2 were announced by the Minister for Enterprise, Tourism and Employment Peter Burke T.D. and are in line with the previous year's findings. The value of retrofitting The survey of 344 SMEs shows that two in five had taken steps such as insulating their buildings or changing their windows in the past two years to improve their energy efficiency. Speaking at the launch, Minister Burke said by doing so these businesses would also be cutting their energy costs and would become more competitive: "It's really encouraging to see businesses reducing their costs by tackling the energy usage in their buildings. There is however another sizeable cohort of businesses (44%) who cite upfront investment costs as a barrier to becoming more sustainable. "That's why I'd ask SMEs to avail of the Local Enterprise Offices' Energy Efficiency Grant (EEG) and the SEAI's Building Energy Upgrade Scheme (BEUS) to buy energy efficient equipment and to retrofit their buildings. I changed the terms and conditions of the energy efficient grant last year so that a 75% grant is now available, up to a maximum of €10,000, which can make a huge difference to energy bills. In 2025, 681 small business were approved for the EEG at estimated value of €5.7 million, while 186 BEUS grants with an estimated value of €3.36 million were approved." Minister Burke announced the research at Wholesome Kitchen, Dominick St, Mullingar which had recently used the Climate Toolkit 4 Business to understand their environmental impact. Businesses can now also use the Toolkit to measure their Scope 1, 2 and 3 emissions. Minister Burke said by estimating their environmental impact, SMEs can start to tackle it: "Through this research we can see that businesses are also concerned that their staff may not implement sustainability measures. The Toolkit is free so anyone can use it to understand their business's carbon footprint, and it will provide information on where to access the Government's sustainability and energy supports." This year's survey included questions on the potential of the circular economy to Irish businesses. Minister of State for Employment, Small Business & Retail and Circular Economy Alan Dillon T.D. said it's clear that businesses are seeing the enormous value of re-using, recycling and minimising waste: "Not only did more than one in three (35%) respondents say that they already participate in the circular economy, of those that don't, a quarter are interested in doing so. By supporting businesses to reuse resources, reduce waste and keep materials in circulation for longer, they will not only become more sustainable they will cut costs and become more competitive." Key Findings 85% of businesses say sustainability is important to their business on a day-to-day basis, maintaining the high levels recorded in the 2024 research. Businesses said that making a positive difference (35%) and saving money (34%) were the top motivations in becoming sustainable. Just over a quarter of business say that climate change is currently affecting their operations, rising significantly among larger firms and those operating for more than 20 years. Among affected businesses, adverse weather is now the dominant impact, reflecting the growing reality of extreme weather events. Most sustainability action is concentrated in practical, cost-effective areas: waste reduction (49%), energy efficiency (44%), and renewable energy adoption (33%) remain the most common measures adopted by businesses. The main barrier for organisations to act more sustainably remained upfront investment costs (22%), although at a lower rate compared to 2024. This research was under...
Susan's career journey in sustainable construction @ 0:00 Susan Heinking has a background in architecture and has been working in the construction industry for the past 10 years, with a focus on sustainable building practices. She discusses how her career has evolved from architecture to construction, with a consistent emphasis on designing and building environmentally-friendly, energy-efficient structures. The shift in attitudes towards climate change and sustainability @ 3:20 Susan describes how attitudes towards climate change and the importance of sustainability have shifted over the course of her career. In the early years, there was more skepticism, but now there is a much greater awareness and demand for sustainable building practices, as the impacts of climate change have become more evident. Challenges of retrofitting vs. building new @ 3:59 Susan discusses the tradeoffs between retrofitting existing buildings versus building new, more energy-efficient structures. Existing buildings can often be made more sustainable, but there is also a cultural preference for new, "shiny" buildings. She highlights the need to balance these considerations and find the most responsible approach for each project. The role of government regulations and incentives @ 12:00 Susan explains how government regulations and incentives have impacted the sustainability efforts in the construction industry, sometimes helping and sometimes hindering progress. She discusses how she has adapted her approach to focus more on the business case for sustainability, rather than relying solely on government mandates. Emerging trends and the role of technology @ 18:39 Looking to the future, Susan discusses the increasing collaboration and standardization happening within the construction industry to drive sustainability efforts. She sees potential for AI and other technologies to help streamline processes and improve efficiency, while still allowing for customization to meet the needs of individual clients and projects. Recap and next steps @ 24:52 Michael and Susan wrap up the conversation, with Susan providing information on how listeners can connect with her and learn more about her work in sustainable construction. https://PepperConstruction.com
Retrofitting is an instrumental step in reducing the carbon footprint of a city's building stock. It also extends the life of a building and has a lower environmental impact than demolishing inefficient properties and building anew. Even a new development, such as the East Village in Stratford London, although just 12 years old, is still largely heated by fossil fuel. Adaptable designs are critical to bring future improvements to existing structures. Marion Baeli is a pioneer of sustainable architecture, her practice identified easy-to-deliver improvements to energy use on one of the buildings in the development, at the same time as adding capacity that could finance the project. Guest Marion Baeli, Principal, Sustainability Transformation at 10 Design Partner Egis is a leading global architectural, consulting, construction engineering, operations and mobility services firm. Egis creates and operates intelligent infrastructure and buildings that both respond to the climate emergency and contribute to balanced, sustainable and resilient development.Its 22,000 employees operate across over 100 countries, deploying their expertise to develop and deliver cutting-edge innovations and solutions for clients. Through the wide range of its activities, Egis plays a central role in the collective organisation of society and the living environment of citizens all over the world.The post #359f Sustainability and Adaptation in East London first appeared on Engineering Matters.
The Elephant In The Room Property Podcast | Inside Australian Real Estate
What if our homes did more than just provide shelter? What if they could actually contribute to the health of the planet and the people living within them? In this episode, we sit down with Caroline Pidcock, a visionary architect and champion of regenerative design, to explore why Australia's current approach to housing is falling short—and how we can change it.Caroline shares her deep expertise on the "Circular Economy" and why we must transition from merely being "less bad" to being "positively good" for our environment. We dive into the hidden health risks of poorly designed homes, the reality of building for extreme weather, and why the "bigger is better" mindset in Australian property is a trap.What we explore in this conversation:Regenerative vs. Sustainable: Why doing "zero harm" isn't enough anymore.The Circular Economy: How to treat buildings as material banks for the future.Health and Architecture: The impact of light, air quality, and materials on your daily well-being.Building Standards: A look at why Australian regulations are trailing behind global leaders.Retrofitting for Resilience: Practical ways to improve existing homes for a changing climate.Whether you are a homeowner, an investor, or simply curious about the future of our cities, this conversation will challenge you to think differently about the spaces we inhabit. Hit play to learn how we can build a future that thrives!Episode Highlights00:00 — Welcome: Rethinking How We Build01:13 — Caroline Pidcock: Beyond Sustainability04:18 — Fixing the Flaws in Modern Design07:06 — Regenerative Design in Action17:17 — Policy Shifts for a Livable Future20:47 — Growth vs. the Environment23:23 — Hard Lessons from Failed Developments26:08 — How Our Cities are Evolving27:47 — The Reality of Melbourne's Planning31:43 — Regional Living & Staying Connected33:08 — Leading the Charge for Urban Change35:49 — Simple Tools for Sustainable Living37:20 — The Hidden Hurdles of Rezoning40:54 — How Density Affects Our Communities48:23 — Final Thoughts: A Legacy for the FutureAbout the GuestCaroline Pidcock is a renowned Australian architect and advocate who has dedicated her career to sustainable and regenerative design. With decades of experience across residential and commercial projects, she is a past President of the Australian Institute of Architects (NSW Chapter) and the Australian Sustainable Built Environment Council (ASBEC).Caroline is a leading voice in the "Living Building Challenge" and is deeply committed to the principles of the circular economy. Her work focuses on creating spaces that are not only carbon-neutral but also enhance the biological and social systems they inhabit. Recognized for her leadership in climate action within the property industry, she continues to influence policy and practice to ensure a resilient and healthy built environment for future generations.Connect with CarolineCaroline Pidcock's LinkedIn
Humans are amazing pattern-matching machines. However, sometimes we extrapolate from a single anecdote and build upon a shoddy foundation. More in this week's episode.3+2 going out of business saleLock special pricing for Unplugged with Aviv: https://avivby.gumroad.com/l/unpluggedGrab a copy of my books, Capitalizing Your Technology and The Tech Executive Operating System.Subscribe to the best newsletter for tech executives.For any questions or comments, reach out to me directly: aviv@avivbenyosef.com
Ever opened a pump lid and watched the pool start emptying onto the pad? We've been there, and today we map out the simple field habits that stop the flood, speed up service, and keep clients happy. From spotting below-waterline equipment to shutting down both sides of the system, we share practical, low-cost tricks that save a service day—think tennis balls in skimmers, expanding chamois in return stubs, and a checklist that prevents air leaks and lost prime.We also dig into cleaner selection with real-world guidance that cuts through confusion. On plaster and pebble, geared suction units like the Hayward PoolCleaner or Polaris Atlas/Max deliver reliable coverage, with wide-body options gliding over tall anti-vortex main drains. On vinyl and fiberglass, bouncing diaphragm cleaners shine, climbing walls and handling slopes where geared units often stall. If pressure is your plan, know the plumbing: most Polaris pressure models require a dedicated booster pump; the Polaris 360 is the rare return-side exception that runs without one when returns are set up correctly.To round it out, we clarify the heat pump vs gas heater puzzle. A heat pump needs a dedicated 220–230V electrical circuit and real amperage headroom; a gas heater needs a properly sized gas line and, often, an upgraded meter from the utility. Retrofitting either after a build adds cost and complexity, so we lay out what to check before promising a swap. The goal: fewer surprises at the pad, better system performance, and faster visits that impress clients.• Identifying equipment set below the waterline• Shutting both suction and return before opening lids• Using tennis balls and chamois rags to stop flow• Managing dual skimmers for vacuuming and cleaners• Choosing cleaners for plaster, pebble, vinyl and fiberglass• Navigating anti-vortex main drains with wide-body units• Understanding pressure cleaners and booster pumps• Differentiating heat pumps and gas heaters requirements• Estimating real costs for electrical and gas line runs• VerifyinSend us a textSupport the Pool Guy Podcast Show Sponsors! HASA https://bit.ly/HASAThe Bottom Feeder. Save $100 with Code: DVB100https://store.thebottomfeeder.com/Try Skimmer FREE for 30 days:https://getskimmer.com/poolguy Get UPA Liability Insurance $64 a month! https://forms.gle/F9YoTWNQ8WnvT4QBAPool Guy Coaching: https://bit.ly/40wFE6y
It has been decades since the last significant earthquake in the United States. Yet, there are earthquake risks across the USA. In this podcast we learn about how homeowners and businesses can take proactive steps to improve the seismic survivability of their properties. The podcast guest is Kyle Tourjé, a second-generation contractor specializing in structural retrofitting, repair, and geohazard mitigation. As Executive Vice President of Alpha Structural, Inc., he oversees all engineering and construction operations. Having grown up in the trade and with over 15 years of experience, including personally repairing and inspecting over 6,000 structures, Kyle combines hands-on construction and field engineering expertise with leadership in real estate and disaster response. His work bridges the gap between engineering solutions and the realities of property ownership and management, code compliance, and disaster response. He has a background spanning construction, insurance claims, and litigation support, he applies practical solutions to California's evolving structural, legal, and environmental challenges. Kyle's focus is on advancing straightforward, lasting solutions that improve safety and resilience for communities across the region. For more about Alpha Structural, visit http://www.alphastructural.comPlease visit our sponsors!L3Harris Technologies' BeOn PPT App. Learn more about this amazing product here: www.l3harris.com Visit The Readiness Lab and learn about our Next Level Emergency Management training! https://www.thereadinesslab.com/Impulse: Bleeding Control Kits by professionals for professionals: www.dobermanemg.com/impulseDoberman Emergency Management Group provides subject matter experts in planning and training: www.dobermanemg.comCheck out how you can use digital twins in your training, exercising, and planning using RSET https://rset.com/ For sponsorship requests, check out our Sponsorship Portfolio here or email us at contact@thereadinesslab.com
(00:00:00) The Importance of Infrastructure in AI Computing (00:04:53) Challenges of Power Consumption in AI (00:11:11) Retrofitting vs. New Data Centers for AI (00:20:28) Optimizing Power Distribution for High-Density Racks (00:25:10) Emerging Cooling Technologies for AI Workloads (00:29:22) Structured Cabling Solutions for AI (00:35:59) Future-Proofing Data Centers for AI Adoption (00:38:21) Motivation and Passion in AI Infrastructure In this conversation, Todd Reed speaks with Bob Wagner, Senior Development Manager at Panduit, about the critical infrastructure supporting AI computing. They explore the challenges of power consumption and heat generation in data centers, the importance of optimizing power distribution, and the emerging cooling technologies necessary for managing AI workloads.The discussion also covers the differences between retrofitting existing data centers and building new ones, the role of structured cabling in simplifying installations, and strategies for future-proofing data centers to meet the demands of AI. Bob shares his passion for innovation and problem-solving in the rapidly evolving landscape of AI infrastructure.Thank you for listening and please take a moment to subscribe, rate, and review our show on your favorite app.To get a hold of us here at Keepin' The Lights On, please email: podcast@graybar.comThank you to our sponsor, Panduit: https://www.graybar.com/manufacturers/panduit/c/sup-panduit?utm_source=Podcast&utm_medium=Ep+61+AI+Cooling&utm_campaign=podcast-main-page&utm_id=PodcastPanduitTo reach Bob Wagner on LinkedIn: https://www.linkedin.com/in/bob-wagner-57a5b46/Learn more about Panduit: https://www.graybar.com/manufacturers/panduit/c/sup-panduit?utm_source=Podcast&utm_medium=Ep+61+AI+Cooling&utm_campaign=podcast-main-page&utm_id=PodcastPanduitMeson Sabika (Spanish Tapas): www.Mesonsabika.comHesed House, a shelter for the unhoused in Aurora, IL: www.hesedhouse.orgWatch on YouTube: https://youtu.be/wKU1tu7yJIATakeaways Distribution is crucial for the infrastructure of AI computing.AI computing is leading to unprecedented power consumption challenges.Retrofitting existing data centers for AI is complex and requires careful planning.High-density racks require optimized power distribution solutions.Emerging cooling technologies are essential for managing heat in AI workloads.Structured cabling solutions can simplify installation and maintenance in data centers.Future-proofing data centers involves planning for power and cooling needs.The demand for AI is driving innovation in data center infrastructure.Collaboration and planning are key to addressing the challenges of AI computing.Passion for problem-solving drives innovation in AI infrastructure.
In this episode we spoke with Mike Rohrmoser, VP of Product Management for OEM Solutions at Digi, a global provider of mission-critical IoT connectivity products and services. We explored how manufacturers are addressing labor shortages with IoT and automation, the trade-offs between retrofitting existing factories and building new ones, the evolving sensor and connectivity landscape, and practical steps to scale IoT pilots into production. Key insights: • Retrofitting existing plants is often the smarter move. Brownfield upgrades can cost 40–60% less than new builds and achieve faster returns when paired with business-focused use cases and retrofit connectivity. • Sensors and networks must be judged as a whole system. Industrial buyers weigh accuracy, deployment simplicity, and lifetime cost over unit price, with wireless IO-Link and LTE Cat 1 gaining traction and 5G RedCap on the horizon. • Edge AI is real, but focused. Today it is most effective in computer vision for quality inspection and counting, while new designs anticipate broader workloads as adoption matures. • GenAI augments people, not machines. Its strengths are in analysis, documentation, and device management, while safety-critical real-time control remains firmly in the domain of conventional automation. • Scaling pilots requires proving value early. Many initiatives stall when they start with technology instead of problems; success depends on production-ready components, operator trust, and leadership alignment. IoT ONE database: https://www.iotone.com/case-studies The Industrial IoT Spotlight podcast is produced by Asia Growth Partners (AGP): https://asiagrowthpartners.com/
Anaheim hotel workers could get affordable housing help after city leaders green lit a proposal Tuesday evening. The state is offering to help retrofit some houses for earthquakes. The L.A. Sparks will have a new training facility, which it said is the largest investment to date for a women's sports team. Plus, more.Support The L.A. Report by donating at LAist.com/join and by visiting https://laist.comVisit www.preppi.com/LAist to receive a FREE Preppi Emergency Kit (with any purchase over $100) and be prepared for the next wildfire, earthquake or emergency! Support the show: https://laist.com
Ciaran Byrne, Director of National Retrofit, SEAI and Brian McIntyre, Programme Manager, SEAI
Trasformare in dual fuel i veicoli per gli autotrasporti pesanti, in modo che possano funzionare anche con un mix di idrogeno e gasolio, come già si fa con alcuni camion che funzionano con un mix di gasolio e metano (quando quest'ultimo è disponibile). Si tratta di una soluzione che numerose compagnie di trasporti in Europa, USA e Australia hanno sperimentato nel corso del 2025: non ottimale, ma semplice e che permette di decarbonizzare, in toto o in parte, i veicoli pesanti già esistenti, gradualmente e senza “strappi tecnologici”. Anche se il costo dell’idrogeno Green rimane una barriera non indifferente. Ne parliamo con Fernando Ortenzi, ricercatore ENEA e responsabile del progetto IPCEI H2 Technology per i veicoli pesanti.
Trasformare in dual fuel i veicoli per gli autotrasporti pesanti, in modo che possano funzionare anche con un mix di idrogeno e gasolio, come già si fa con alcuni camion che funzionano con un mix di gasolio e metano (quando quest'ultimo è disponibile). Si tratta di una soluzione che numerose compagnie di trasporti in Europa, USA e Australia hanno sperimentato nel corso del 2025: non ottimale, ma semplice e che permette di decarbonizzare, in toto o in parte, i veicoli pesanti già esistenti, gradualmente e senza “strappi tecnologici”. Anche se il costo dell'idrogeno Green rimane una barriera non indifferente. Ne parliamo con Fernando Ortenzi, ricercatore ENEA e responsabile del progetto IPCEI H2 Technology per i veicoli pesanti.
Marie Donnelly, Chairperson of the Climate Change Advisory Council, calls on the government to improve its home improvement grants for retrofitting.
The Government must improve its supports for retrofitting, heat pumps and solar PV panels. That's the call from the Climate Change Advisory Council, whose Chair Marie Donnelly who explained it all to Shane.
Part 2: Beyond Lesson Plans: Balancing Delays & Highlighting Progress Teaching today means far more than covering math problems and reading lists. It's managing the ripple effects of the pandemic, lingering academic delays, and the daily pressures kids bring with them. Alongside the frustrations are signs of progress, as schools adapt with new resources and evolving approaches. In part two of this story, we cover how teachers are finding the silver lining in these challenges and what are some key focuses heading into this school year. The Hidden Housing Crisis For America's Seniors For millions of older Americans, the dream of aging in place is colliding with the reality of inaccessible and unaffordable housing. Retrofitting homes is often out of reach financially, downsizing isn't the easy fix it appears to be, and without these changes, independence and safety become harder to hold onto in later life. We cover this quiet crisis and what resources are available to take proactive steps early on. Viewpoints Explained: Why Are So Few Women In This Industry? Just 12 percent of police officers are women and only 3 percent are in leadership positions in America. We cover one initiative that's focused on driving more women into this public-facing sector. Culture Crash: The Magic Of Film: Why 70mm Screenings Outshine Digital While digital dominates the box office and on streaming platforms, the texture and scale of 70mm film screenings continue to drive movie lovers to the theater. We cover this art form and why we're a fan. Learn more about your ad choices. Visit megaphone.fm/adchoices
For millions of older Americans, the dream of aging in place is colliding with the reality of inaccessible and unaffordable housing. Retrofitting homes is often out of reach financially, downsizing isn't the easy fix it appears to be, and without these changes, independence and safety become harder to hold onto in later life. We cover this quiet crisis and what resources are available to take proactive steps early on. Learn More: https://viewpointsradio.org/the-hidden-housing-crisis-for-americas-seniors Learn more about your ad choices. Visit megaphone.fm/adchoices
This week on Everybody in the Pool, we're celebrating our 100th episode with a look at what matters most: your actions.Since this show began a little over two years ago, the goal has been simple — to spotlight innovation, ingenuity, and capital coming together to tackle the climate crisis. Hope is stronger than fear, but hope alone isn't a plan. This milestone episode is about agency — the choices we make in our own lives, and how together, those choices add up to systemic change.Listeners wrote in and sent voice memos sharing the climate actions they've taken:Investing through platforms like Climatize to fund renewable energy projectsMoving retirement savings and banking into fossil fuel–free funds and community credit unionsCutting back on red meat, shifting diets, and sourcing local foodTackling food waste with apps like FlashFood and composting with Mill (our presenting sponsor for this week's episode)Retrofitting homes with solar, heat pumps, and energy efficiency upgradesRethinking careers, transportation, and even family planning with the climate in mindAlong the way, we revisit powerful clips from past episodes and highlight the ripple effects of these solutions — from decarbonizing finance to building circular food systems.Thank you to everyone who has listened, shared, and taken action. This episode is a reminder that we are not helpless — our feedback, votes, purchases, and investments all send signals that drive change. Drops become a flood.Thanks to Mill for sponsoring this week's episode! Get $75 off yours with my custom link! https://www.mill.com/lp/mollywood?utm_source=newsletter-sponsorship&utm_medium=partnership&utm_campaign=everbodyinthepool &utm_content=mollywoodAll episodes: https://www.everybodyinthepool.com/Subscribe to the Everybody in the Pool newsletter: https://www.mollywood.co/Become a member and get an ad-free version of the podcast: https://everybodyinthepool.supercast.com/Please subscribe and tell your friends about Everybody in the Pool! Send feedback or become a sponsor at in@everybodyinthepool.com! Hosted on Acast. See acast.com/privacy for more information.
We talk a lot about new technology on the podcast and you may have caught yourself thinking "that doesn't apply to me" a time or two. Brenton Peters is here today to let you in on the things he's learned along the way while retrofitting his aged equipment with the latest and greatest technology. From his G5 to AutoPath, Turn Automation to SF-RTK, and everything between, he is a walking testament to the fact that far-out technology isn't all that far-fetched. This is one you can't miss!
In this educational session, Adam from National Comfort Institute (NCI) delivers a comprehensive deep dive into Fan Law 2 and its practical applications for residential HVAC systems at the 6th Annual HVACR Training Symposium. Adam begins by establishing the fundamental concepts of CFM (cubic feet per minute) and static pressure, explaining how these measurements relate to system performance. He shares a humbling personal story about learning to measure gas pressure from a homeowner, emphasizing that even experienced technicians can benefit from understanding basic measurement principles. The presentation focuses heavily on Fan Law 2, which allows technicians to predict how changes in airflow will affect static pressure in a non-proportional relationship - a critical concept for equipment sizing and replacement decisions. The core of the presentation revolves around practical applications of Fan Law 2 in real-world scenarios. Adam demonstrates how to calculate pressure drops across filters, evaporator coils, and entire duct systems when airflow changes occur. He emphasizes that static pressure increases exponentially when airflow increases, which explains why oversized systems often perform poorly. Through detailed examples using actual field measurements, he shows how a 16% increase in airflow can result in a 33% increase in static pressure, highlighting the importance of proper system sizing. Perhaps most importantly, Adam presents a systematic approach to equipment selection that goes beyond simply matching tonnage. He demonstrates how contractors can "back into" total external static pressure calculations by carefully selecting low-pressure-drop components like evaporator coils and filters. This methodology allows technicians to predict system performance before installation, preventing the common scenario where new equipment sounds "like a rocket ship" due to excessive static pressure. The presentation concludes with a compelling comparison showing how proper component selection can reduce system static pressure from over 1.0 inches to 0.64 inches while maintaining the same capacity and airflow. Topics Covered Static Pressure Fundamentals Definition and measurement using manometers Inches of water column explained Relationship between static pressure and system performance Fan Law 2 Mathematics Breaking down the intimidating formula into simple terms Step-by-step calculation examples Common mistakes when squaring numbers in calculations Practical Applications Filter pressure drop calculations at different airflows Evaporator coil pressure drop analysis Total External Static Pressure (TESP) predictions Duct system pressure calculations Equipment Selection Strategy How to select evaporator coils based on pressure drop ratings Filter sizing for optimal pressure drop Using manufacturer data sheets effectively AHRI matchup considerations beyond just capacity Real-World Problem Solving Preventing "rocket ship" installations Retrofitting existing systems with proper calculations Downsizing benefits for static pressure reduction System commissioning and performance verification Professional Development Moving beyond equipment replacement guesswork Using measurement tools like True Flow Grid Understanding manufacturer specifications Elevating installation quality through proper system design Have a question that you want us to answer on the podcast? Submit your questions at https://www.speakpipe.com/hvacschool. Purchase your tickets or learn more about the 7th Annual HVACR Training Symposium at https://hvacrschool.com/symposium. Subscribe to our podcast on your iPhone or Android. Subscribe to our YouTube channel. Check out our handy calculators here or on the HVAC School Mobile App for Apple and Android
Paul Cunningham, Political Correspondent, discusses news that a €2 billion commitment to retrofit residential homes by the year 2030, is now under review by Environment and Energy minister, Darragh O'Brien
Could something as seemingly simple as air quality management cut your PRRS outbreak risk in half? The latest research suggests exactly that – and it's changing how producers think about biosecurity investments.A groundbreaking study from the University of Minnesota has revealed that properly implemented air filtration systems reduce PRRS outbreak risks by 51-58% compared to non-filtered farms. This comprehensive research analyzed data from the Morrison Swine Health Monitoring Project, representing about 60% of US breeding herds over a 15+ year period. What makes this study particularly valuable is its consideration of both positive and negative pressure filtration systems, along with sophisticated controls for regional pig density and spatial correlation factors.For producers weighing the investment, the findings provide clear ROI calculation guidance. With implementation costs ranging from $250-500 per sow and filter lifespans typically reaching 4-6 years, the protection against costly PRRS outbreaks makes a compelling business case – particularly in pig-dense regions like Southeast Iowa and Minnesota. Retrofitting existing facilities often requires upgrading fan capacity and improving building seals, but these investments extend facility lifespans by 10-20 years while dramatically reducing disease risks.Robert Langenhorst, technical service manager with American Air Filter, and Dr. Xiaomei Yue of the University of Minnesota, emphasize that filtration must be viewed as one layer in a comprehensive biosecurity approach. Regular maintenance, inspection for damage, and proper sealing are essential for system effectiveness. As the industry increasingly looks to protect nurseries and growing facilities in addition to sow farms, this research provides timely guidance for strategic disease prevention through improved air quality management.
Welcome back to Architecture 5 10 20! I'm your host, Guy Geier, Managing Partner of FXCollaborative Architects in New York. My guests for this podcast are pioneers and visionaries shaping the future of the built environment across various disciplines. Join me in exploring their remarkable journeys, discovering how they reach their current heights, and envisioning what lies ahead in the next 5, 10, and 20 years. I am thrilled to welcome Adam Fisher to the podcast for this episode! Adam is the Managing Director of Sustainable Brokerage at JLL, and he joins me in this episode to discuss his background and how he became focused on the intersection of sustainability and real estate. After first studying mechanical engineering with a focus on energy efficiency, he made a transition into sustainability consulting and eventually joined JLL to help build out their sustainability strategies. Listen in as Adam explains the concept of "sustainable transaction strategies", which goes beyond just green leasing to look at the entire lifecycle of a real estate transaction, including understanding the client's sustainability targets, identifying sustainable spaces and landlord partners, using sustainability in lease negotiations, and aligning on aspects such as but not limited to waste management. A key challenge that Adam highlights in the episode is the gap between intentions and what is actually enforced when it comes to sustainability, and he strongly advocates for more specific, discrete clauses that outline clear responsibilities and verification for both landlords and tenants. Looking ahead to the future, Adam is optimistic, seeing sustainability becoming more effectively integrated into core business processes rather than being a siloed function. Facility managers and building operators will need to be brought into the conversation and empowered to make sustainable decisions. Overall, the real estate industry has a significant opportunity to drive meaningful change through informed, forward-looking decisions around sustainability, and Adam really helps drive this home. His insights into sustainable real estate transaction strategies emphasize the importance of a holistic process that embeds sustainability throughout the transaction cycle. As sustainability becomes further integrated into the real estate process and regulatory measures continue to be implemented, the industry has a unique opportunity and responsibility to drive meaningful change through informed, forward-looking decisions. Enjoy my conversation with Adam Fisher! Time stamps: [02:04] - Adam Fisher reveals how having watched the documentary An Inconvenient Truth sparked his journey into sustainable real estate solutions. [04:37] - Hear how, at JLL, Adam bridged sustainability and brokerage by launching a unified advisory business model. [06:54] - Sustainable transactions require aligning corporate goals, building selection, and lease terms from the very beginning. [10:15] - Adam argues that most lease clauses lack teeth, so he advocates for enforceable commitments between parties. [12:11] - Real estate sustainability work requires expertise in buildings, people, and persuasive communication and not just green knowledge. [14:44] - Adam points out that strong early coordination among all stakeholders prevents surprises during lease negotiations and construction planning. [16:46] - Companies set sustainability goals but unfortunately rarely integrate them into real estate. [18:33] - Despite political backlash, many firms are quietly maintaining long-term sustainability and decarbonization commitments. [20:39] - Adam points out that sustainability efforts have shifted from short-term savings to long-term asset value and risk avoidance. [22:34] - Retrofitting buildings now prioritizes long-term value, tenant expectations, and risk avoidance. [24:11] - High tenant demand for sustainable real estate far exceeds current supply. [26:37] - Adam argues that sustainability needs to be integrated into core business operations and not treated as a separate initiative. [29:15] - Education is crucial so building operators understand and properly use high-performance systems as are intended. [30:38] - Engaging operators directly is ke or else smart building strategies will fail despite great planning. Links / Resources:Guy Geier Instagram | Twitter Adam Fisher / JLL Adam's LinkedIn | JLL Website | JLL LinkedIn
Antoine Larvol, CTO of Windar Photonics, discusses how their continuous wave LiDAR technology enhances wind turbine performance through optimization and monitoring, increasing AEP and reducing loads, particularly for legacy turbines. Sign up now for Uptime Tech News, our weekly email update on all things wind technology. This episode is sponsored by Weather Guard Lightning Tech. Learn more about Weather Guard's StrikeTape Wind Turbine LPS retrofit. Follow the show on Facebook, YouTube, Twitter, Linkedin and visit Weather Guard on the web. And subscribe to Rosemary Barnes' YouTube channel here. Have a question we can answer on the show? Email us! Welcome to Uptime Spotlight, shining light on wind. Energy's brightest innovators. This is the Progress Powering Tomorrow. Alright, we're here in Phoenix, a CP, clean power, uh, 2025. So I'm, uh. Sitting with Antoine Larvol from, he's a CTO from Windar. Yep. Welcome to the show. Thank you. Uh, we've been, uh, happy enough to get actually to sit inside your booth where it's nice and qui. Quiet and isn't it nice? Yeah. We got glass behind the camera here and people are walking by, walking by, walking by. Um, so this morning, uh, we, we talked yesterday a little bit about what wind photonics does. Yep. Of course, from our, uh, some of our other friends around the world. We've heard about some, some campaigns you've done in the United States, which have been. Really successful. So yeah, congrat good. Congratulations there. Yeah, thank you. Um, and, and as, as a lot of things in the wind industry, Windar, photonics based in Denmark. Antoine Larvol: Yeah. Joel Saxum: So you guys, uh, bring it, bring in that Danish [00:01:00]technology. We're here, of course, bringing it to the US market at a CP, the American Clean Power Show. So welcome to the States. Thank you. Um, it's a short one, but a Antoine Larvol: good one. Yeah, yeah, yeah, Joel Saxum: exactly. So, so I want to talk a little bit about what Windar photonics and, and it is a LIDAR based sensor, correct? Antoine Larvol: Yes. Right. So. We do continuous wave base, uh, lidar. Yep. Uh, main product is a two beam version mm-hmm. Where you shoot, uh, at 80 meters in front of the turbine. Mm-hmm. And you basically alternate from one beam to the other. And measure wind speed and direction upfront, the, the turbine among others. Joel Saxum: Right. So we're talking about, uh, if you, if you're in the wind industry, you've ever seen these lidar units that are put actually, you're the cell mounted, correct? Yes. Okay. Yeah. So, and, and, uh, we're looking more on the optimization, retrofit monitoring side of things. Yeah, Antoine Larvol: exactly. So we've never been a resource assessment company. Yeah. Or we don't look at power curve verification and stuff like that. We really [00:02:00] focus on. Retrofitting those, existing turbines. And then add value to In terms of information to, the customer, Yeah. With the mon monitoring side of things. Yeah. And, from day one, that's been the goal of Windar Making something cheap, robust. That can just stay there and measure with good availability, wind speed, and direction coming to your turbine. Joel Saxum: I love it. so we wanna squeeze as much as we can outta these turbines. And you guys are increasing AEP that's, the name of the game. Yeah. Right. Increasing AEP below rated. and then above rated you decrease loads. Increase uptime. and we basically do that by going on the line of the wind direction. that you then feed to the turbine controller and then we can actually adjust the, yaw position of the turbine according to our information. So I want to talk a little bit, we, we chatted a little bit offline about the, technology behind it, right? Yep. And people in the wind industry, if you're around the wind industry around resourcing or you're around optimization, you've heard [00:03:00] lidar. Yep. You know what I mean? And,
In this episode of the Industrial Advisors podcast, the hosts discuss the concept of transforming obsolete industrial spaces into desirable tenant properties through retrofitting. They examine examples of successful projects in Tacoma, Auburn, and South Seattle, such as IRG's redevelopment of the Super Value site and the LIFT project at the old Sears building. The conversation also covers the significance of such projects in land-scarce markets, the trends of downsizing large industrial parks like 212 Business Park, and the ongoing viability of these transformations in the current market. The episode concludes with a review of a current project involving the former Ardagh glass plant and the potential for future redevelopment opportunities. 00:00 Introduction to Obsolete Spaces 00:27 Examples of Successful Retrofitting Projects 01:14 Market Dynamics and Challenges 01:31 Case Study: 212 Business Park 02:56 Viability in Today's Market 04:25 Conclusion and Final Thoughts You can find every episode of this show on Apple Podcasts, Spotify or YouTube, For more, visit industrialadvisors.com
Samantha Libreri, Eastern Correspondent, reports on calls for local authority tenancy to be brought under the remit of the Residential Tenancies Board.
Toby Cambray talks about the experience of working on his own retrofit project, and the lessons learnt in the process. Check out the show notes for more information.
Former Chief Assistant U.S. Attorney and National Review Contributing Editor Andy McCarthy joins Greg for Tuesday's 3 Martini Lunch. Today, they tackle Harvard's complaining about the Trump administration's demands, the legal fight over an illegal immigrant deported to El Salvador, and fresh evidence that the left's climate promises don't add up.First, they highlight Harvard's response to having more than $2 billion in federal grants frozen for refusing to comply with Trump administration orders targeting antisemitism and more. Andy questions why a wealthy institution like Harvard relies on taxpayer dollars at all, but also warns against government overreach—regardless of party. Meanwhile, Greg notes the irony of Harvard objecting to federal pressure when the left regularly uses it to punish conservative schools and organizations.Next, they examine the legal fight surrounding Kilmar Abrego Garcia, an illegal immigrant deported to El Salvador and now imprisoned there. A federal judge has ruled that Garcia must be returned to the U.S. due to a violation of his due process rights. Andy explains why the court is legally correct and critiques the White House's arguments on the issue. Greg questions why the U.S. government makes it so easy for people to enter illegally but so difficult to remove them. Andy points to the man he says is responsible for this mess.Finally, they dissect a Washington Post analysis revealing that many “green” home retrofits don't produce energy savings for years—or even decades. Andy argues this underscores how the environmental left's agenda is filled with economic and environmental contradictions. He also points out that some of these so-called green initiatives are actually terrible for the environment in many ways.Please visit our great sponsors:Oracle will cut your cloud bill in HALF —new US customers only, offer ends May 31st! Check eligibility: https://oracle.com/MARTINIThis spring, get up to 50% off select plants at Fast Growing Trees with code MARTINI, plus an extra 15% off at checkout on your first purchase! Visit https://fastgrowingtrees.com/MartiniThis podcast is sponsored by BetterHelp. Your well-being is worth it. Visit https://BetterHelp.com/3ML to get 10% off your first month
Send me a messageIn this episode of Climate Confident, I sit down with Puja Balachander, CEO and co-founder of UpGreen, to explore how commercial landlords and asset managers can accelerate energy efficiency retrofits while keeping costs down.Buildings account for nearly 40% of global carbon emissions, yet many remain inefficient due to financial and logistical barriers. UpGreen tackles this by reducing upfront retrofit costs and enabling landlords to recapture savings from tenants, turning sustainability upgrades into a viable business strategy.We discuss:Why 87% of UK commercial buildings must undergo energy upgrades within the next five years to meet regulations.How UpGreen's model cuts retrofit costs by up to 80% while recovering 60% of expenses through tenant savings.The hidden inefficiencies preventing widespread adoption of energy retrofits, despite their cost-effectiveness.The challenges of scaling retrofits across different markets, from the UK's public energy performance data to Germany's fragmented regulations.The future of retrofits beyond energy efficiency, including climate adaptation measures for flood and heat resilience.This episode offers practical insights for commercial landlords, sustainability professionals, and policymakers looking to unlock the full potential of building decarbonisation.
Industrial Talk is onsite at PowerGen and talking to Bert Warner, Director of Commercial BD with Propane Education & Research Council about "abundant fuel for an expanding energy market". Scott MacKenzie interviews Bert Warner, Director of Commercial Business Development at the Propane Education and Research Council (PERC), at the Power Gen event in Dallas, Texas. Bert discusses the abundance of propane in the U.S., with 40 billion gallons extracted annually, of which only 10 billion are used domestically. He emphasizes propane's green energy credentials, resiliency, and cost-effectiveness compared to natural gas and electricity. Bert highlights the importance of responsible energy diversification and the need for greater public education on energy sources. He encourages listeners to visit propane.com for more information. Action Items [ ] Provide more information about propane and its benefits to the public [ ] Connect with Bert Warner on LinkedIn to discuss propane further Outline Introduction and Welcome Scott MacKenzie introduces the Industrial Talk Podcast, emphasizing its focus on industry innovations and trends. Scott highlights the importance of the PowerGen event in Dallas, Texas, and its significance for the power and fuel industries. Scott introduces Bert Warner, Director of Commercial Business Development for the Propane Education and Research Council (PERC). Bert provides a brief background on his role and the council's mission to promote propane use in commercial sectors, particularly in power generation. Current Energy Crisis and Propane's Role Bert discusses the ongoing energy crisis, emphasizing the growing gap between energy demand and supply. Scott and Bert agree on the need for a diverse energy mix, including natural gas, hydrogen, and propane. Bert stresses the importance of keeping propane in the energy conversation due to its potential to address the current crisis. Scott and Bert discuss the societal tendency to react to crises rather than being proactive in energy solutions. Propane's Abundance and Green Energy Aspects Bert shares statistics on propane extraction and usage in the US, highlighting its abundance. Bert explains that the US uses only a small fraction of the propane extracted, with a significant amount exported. Bert emphasizes propane's green energy credentials compared to the national average electric grid. Bert discusses the resiliency of propane, particularly in providing energy to remote or growing areas where traditional infrastructure is costly to extend. Propane's Flexibility and Market Competitiveness Bert explains the flexibility of propane in terms of delivery and usage, making it an attractive option for various sectors. Bert compares propane to diesel, highlighting its environmental friendliness and lower maintenance requirements. Scott and Bert discuss the growing conversation around decarbonization and the role of propane in this context. Bert emphasizes the need for responsible energy diversification to meet high energy demands without compromising on environmental impact. Retrofitting and Fuel Switching Scott inquires about the ease of retrofitting systems to use propane instead of diesel or natural gas. Bert explains that retrofitting involves bringing in new tanks and piping, but the process is not overly complex. Bert clarifies that it is not necessarily fuel...
Dean says to free your mind when remodeling your home to help gain more creativity. Dean talks retrofitting an attic with fire& ember protective vents, how to remove gray spots/mold on a vinyl tub wall, advice on a chipping living room ceiling, Dean talks about replacing a mirrored-closet door, installing a sliding door in a mobile home, and marrying two different materials for countertops. Dean believes that there are no bad ideas, just decisions... the thinking and planning needs to be expanded to be able to get inspired
In this episode of the IoT For All Podcast, Fabrizio Del Maffeo, co-founder and CEO of Axelera AI, joins Ryan Chacon to discuss edge AI. The conversation covers the importance and benefits of edge AI, such as reduced latency, real-time decision-making, and enhanced privacy, optimizing algorithms and hardware design for edge devices, the potential of AI in various industries, the role of cloud computing, retrofitting existing solutions with AI, and the impact of generative AI.Fabrizio Del Maffeo is co-founder and CEO of Axelera AI, a Netherlands-based startup building scalable hardware for AI at the edge. Fabrizio leads a world-class executive team, board of directors, and advisors from top AI Fortune 500 companies. Previously, Fabrizio was Vice President and Managing Director of AAEON Technology Europe, the AI and IoT computing company within the ASUS Group. Fabrizio graduated with a Master's degree in telecommunication engineering from Milan Politecnico University.Axelera AI is on a mission to provide rapid access to advanced Edge AI-native hardware and software solutions for companies of all sizes across a range of market verticals and place AI in the hands of those who could not otherwise afford it. They do this by delivering faster, more efficient, and easy-to-use inference acceleration while minimizing power and cost. To do this, their platform is purpose-built to support AI strategies across a wide-range of industries while seamlessly integrating with existing technologies.Discover more about IoT athttps://www.iotforall.comFind IoT solutions:https://marketplace.iotforall.comMore about Axelera AI:https://www.axelera.aiConnect with Fabrizio:https://www.linkedin.com/in/delmaffeo/(00:00) Intro(00:10) Fabrizio Del Maffeo and Axelera AI(01:20) What is edge AI?(02:30) Benefits and challenges of edge computing(05:17) Privacy and compliance in edge AI(06:26) Future of edge computing and AI(08:18) Retrofitting existing edge devices with AI(11:02) Role of cloud computing(12:26) Impact of generative AI(15:24) Industry insights from recent events(17:11) Learn more and follow upSubscribe to the Channel:https://bit.ly/2NlcEwmJoin Our Newsletter:https://newsletter.iotforall.comFollow Us on Social:https://linktr.ee/iot4all
Click this link to learn more about the Business Mastery Class for Solo Inspectors:https://events.iebcoaching.com/BusinessMasteryforSoloInspectors25In this episode, Reuben Saltzman and Tessa Murry welcome Philippe Heller, a seasoned San Diego home inspector. Philippe shares his journey from corporate life to running a successful home inspection business, emphasizing fire safety in California. They discuss new regulations on defensible space, fire-hardening features, retrofitting older homes, and the role of specialized fire protection companies. The conversation covers air quality concerns, evolving building codes, and fire-resistant materials. Philippe also highlights advanced fire protection systems, personal fire defense strategies, and opportunities for home inspectors to adapt and innovate. Here's the link to check Inspector Empire Builder: https://www.iebcoaching.com.You can find Philippe at https://sdinspect.com.TakeawaysPhilippe Heller transitioned from a corporate job to home inspections.The importance of fire safety regulations in California.Defensible space is crucial for homes in fire-prone areas.Home inspectors can provide valuable insights into fire safety.Philippe's company became the largest home inspection firm in San Diego.Insurance companies are starting to consider fire safety policies.New building codes require fire-hardening features in homes.Home inspectors need to adapt to changing regulations.Philippe's journey reflects the entrepreneurial spirit.The podcast emphasizes the importance of community and support in business.Home fire hardening features are essential for safety.Retrofitting older homes can significantly reduce fire risk.Specialized companies offer valuable services for home protection.Air quality is a major concern, especially during wildfire seasons.Building codes have evolved in response to past fire disasters.Fire-resistant materials are crucial for modern home construction.Advanced fire protection systems can enhance home safety.Personal fire defense strategies can be lifesaving during emergencies.Home inspection services vary greatly by region and need.There are numerous opportunities for home inspectors to innovate and expandtheir services.Chapters02:05 Special Guest Introduction: Philippe Heller04:40 Philippe's Journey into Home Inspections12:50 Tanya's Role and Company Growth14:40 Defensible Home Services and Fire Safety19:10 California's Fire Safety Regulations22:59 Fire Hardening Features in High-Risk Areas25:06 Home Fire Hardening Features26:12 Retrofitting Older Homes for Fire Safety27:43 Specialized Companies for Home Protection28:50 Air Quality and Ventilation Concerns30:30 California's Strict Air Quality Regulations31:52 Building Code Changes Post-Fires32:59 Fire-Resistant Building Materials34:36 Advanced Fire Protection Systems36:55 Personal Fire Defense Strategies39:25 Home Inspection Services and Pricing41:54 Regional Differences in Home Inspections43:49 Opportunities for Home Inspectors
With the Japanese taking control around the Pacific in early 1941, it became apparent that more resources and ships would be needed if there was any hope to defend against and defeat those forces. It was determined that several previously manufactured vessels could be converted to better suit the needs for this type of warfare. This is why a Cleveland class light cruiser was turned into an aircraft carrier, becoming the USS Princeton (nicknamed “Sweet P”). From humble beginnings it had incredible exploits in the Pacific Theater of World War II. In this episode we explore what life was like aboard this vessel from the people who were aboard, ” detailing various battles in the campaign against the Japanese, every day decisions, and technical aspects of such a ship. We're joined by David Leick, author of “USS Princeton: The Life and Loss of ‘Sweet P,'” to see an account of one of the first light aircraft carriers through to its eventual sinking.See omnystudio.com/listener for privacy information.
This week, a show taped live at Syracuse University on September 30 with Associate Professor Dimitar Gueorguiev, author of the excellent Retrofitting Leninism: Participation Without Democracy in China. We discuss his book, his recent paper exploring hawkishness in Chinese public opinion, and his thoughts about the upcoming U.S. presidential election.1:59 Syracuse University's MAX 132 class ("the globalization class")4:10 Dimitar's background and how he became interested in China 7:44 How the genre of authoritarian resilience took off 14:26 China's understanding of democracy (whole-process democracy)17:40 Features of Leninism that have allowed the Chinese Communist Party to survive21:21 Why China in the 1980s and '90s admired Singaporea's authoritarian PAP 23:37 The idea of the mass line27:16 China's sentiment analysis through technology, and using bottom-up information as performance evaluation 34:03 The COVID-19 pandemic and the confirmation bias of the regime-type explanation37:37 The National People's Congress and the Chinese People's Political Consultative Conference (CPPCC)40:14 Dimitar's research on hawkishness in China: how he got the data, what drives Chinese hawkishness, and the national security vs. economic lens 51:08 Why those who are dissatisfied with the government lean more hawkish and those who are satisfied with the government lean more dovish 56:30 The upcoming U.S. election: how things may play out under the two different administrations, and understanding Chinese preferences Recommendations:Dimitar: The TV series The Expanse (2015-2022)Kaiser: Anthea Roberts' Six Faces of Globalization: Who Wins, Who Loses, and Why It Matters; and the documentary Wise Guy: David Chase and The Sopranos (2024)See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.