POPULARITY
Categories
The disappearance of Kiely Rodni shocked the country, but the discovery of her vehicle by Adventures With Purpose raised difficult questions about how she was found after an extensive law-enforcement search. In this Police Off the Cuff livestream, we examine the search, the recovery of Kiely and her vehicle from Prosser Creek Reservoir, and what investigators may have missed along the way. From a retired law-enforcement perspective, we break down the investigative decisions, the evidence, and the lessons this tragic case offers for future missing-person investigations. Date Found: August 21, 2022 by the volunteer dive group Adventures with Purpose. Location: Submerged inside her vehicle in the Prosser Creek Reservoir near Truckee, California. Last Seen: August 6, 2022 at a large gathering at the Prosser Family Campground. Ruling: Her death was officially ruled an accidental drowning by the Nevada County Sheriff's Office Coroner. [1, 2, 3] Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Alison Curtis is joined by Cork singer-songwriter Darren Kiely, catching him early as his career gathers serious pace. Darren chats about growing up “stereotypically Irish”, learning fiddle and guitar, and the musical roots in his family—including a brilliant Dubliners connection!He also talks about spending time gigging in New York and Nashville, and how he follows what feels right in his music.Plus, Darren shares huis influences from Bell X1 to The Lumineers, and performs his song “Wait” live in our EP snug studio.Darren Kiely tour info: https://www.darrenkiely.com/tour/
We're excited to introduce you to LumiVita, where vitality, longevity, and wellness are reimagined. With beautiful new locations in Buckhead and Midtown, LumiVita is creating a modern approach to feeling and living your best.In this episode, CEO and Co-Founder Deirdra Kiely joins us to share more about the vision behind LumiVita, the innovative services they offer, and how their approach can help you invest in your health, wellness, and longevity. Tune in and discover what makes LumiVita such an exciting addition to Atlanta's wellness landscape. Enjoy!
Watch the full episode on YouTube:We first covered Baseten last year when DeepSeek mania was at peak hype. Now they have raised a monster $13B round and become one of the new cohort of AI Infra decacorns that are (with Nvidia, Intel, and the semis complex) chief beneficiaries of the Inference Inflection. We return to Baseten at the peak of the 2026 edition of Open Weights debate. Ali has published a viral breakdown of Kimi K3:And since you last saw him, Philip has spoken at AI Engineer and written the definitive book on Inference Engineering spotted all over SF:Three years ago, inference engineering barely existed as a category.Today, it is one of the most critical disciplines in AI. Inference engineering inherently tackles a different question than standard model training: “How do you turn those weights from training into a product that is fast, reliable, and affordable at scale?” Focusing on these creates an entirely new optimization problem.In one recent GLM-5.2 experiment, quantizing more of the model actually preserved its benchmark quality while increasing throughput by 20%, because the errors introduced in different layers could cancel each other out.Inference is no longer just the final step after training. It is becoming its own engineering discipline, with its own research problems, infrastructure, and increasingly specialized roles.In this episode, Baseten's Philip Kiely and Ali Taha join swyx and Vibhu to explain what actually happens after a new open model is released and what it takes to turn “we generated a token” into a fast, reliable, production-ready API.We go deep on cache-aware routing, disaggregated prefill and decode, quantization, speculative decoding, KV-cache movement, model parallelism, GPU kernels, and the race to make frontier models up to 10× faster. Philip and Ali explain why inference optimizations can still produce gains of 20%, 100%, or even 200%; how quantization errors can cancel one another out; why identical weights can behave differently across clusters; and how Baseten grafted a Kimi vision encoder onto GLM-5.2 without changing the underlying language model.The conversation then expands beyond LLMs into NVIDIA Dynamo, mega kernels, Rubin, AI-specific chips, local inference, video generation, diffusion versus autoregressive models, and the enormous compute barrier to generating coherent long-form video. Finally, we explore the convergence of training and inference, continual learning through persistent KV cache, and the emerging loop where models help optimize the infrastructure that runs them.We discuss:* What happens when a 200,000-token request enters an inference system* Cache-aware routing and reusing previously computed KV cache* Why prefill and decode are increasingly handled by different GPUs* When dedicated deployments become cheaper and more reliable than shared APIs* How speculative decoding uses a smaller model to accelerate a larger one* Tool calling, structured outputs, and what LLMs actually do* What it takes to support a new open model on day zero* Grafting Kimi's vision encoder onto GLM-5.2* Retrofitting inefficient model layers with components from other architectures* Why models sometimes collapse into repeating the same token* How hardware, kernels, and race conditions create nondeterministic failures* Preserving model fidelity while making inference faster* How quantization errors can cancel each other out* Why inference optimizations still deliver gains of 20%, 100%, and 200%* How optimized serving can make a model up to 10× faster* NVIDIA Dynamo, KV-aware routing, and distributed model serving* Speculative decoding the speculative decoder* Why local AI is about making models less dumb while data-center AI is about making them less slow* Tensor, expert, and pipeline parallelism across GPUs* Hardware-aware model design, auto-tuning, and the case against mega kernels* Rubin and why inference is becoming a systems problem* Whether modern GPUs are evolving into programmable AI ASICs* Why enormous models like Kimi K3 require GB300-class hardware* Why open-source video generation still trails Veo, Kling, and other closed models* The quadratic attention bottleneck behind long-form AI video* Autoregressive video, real-time generation, and compounding quality drift* Why future video systems may combine autoregressive and diffusion architectures* Training for inference and inference for training* Continuous post-training, deployment, evaluation, and improvement loops* How GLM-5.2 helped optimize the kernels serving GLM-5.2 itself* Why faster networking could unlock dramatically faster decoding* Continual learning, KV-cache compaction, and persistent model memoryShow Notes* How to build a day-0 API for Kimi K3* 22580: From GPT2 to Kimi3, ExplainedPhilip Kiely* LinkedIn: https://www.linkedin.com/in/philipkiely* X: https://x.com/philipkiely* Inference Engineering: https://www.baseten.co/inference-engineering/Ali Taha* LinkedIn: https://www.linkedin.com/in/aliestaha/* X: https://x.com/waterloointernTimestamps00:00:00 Introduction and the 200K-Token Prompt00:03:18 Dedicated Deployments, Speculative Decoding, and Tool Calling00:11:26 Launching Production-Ready Open Models00:19:06 Model Retrofits, Failure Modes, and Nondeterminism00:28:22 Quantization and Canceling Errors00:32:15 The Race to 10× Faster Inference00:40:48 Dynamo, Speculation, and Local vs. Data-Center AI00:50:18 Model Parallelism, Auto-Tuning, and Mega Kernels01:00:55 Rubin, GPUs vs. ASICs, and Custom AI Chips01:10:03 Giant Models and the Limits of GPU Memory01:12:42 AI Video, Quadratic Attention, and Autoregressive Generation01:21:47 Audio, Images, and Diffusion Models01:27:32 Training, Self-Optimizing Models, and Continual Learning01:40:06 Closing ThoughtsTranscriptIntroduction: Baseten, Waterloo Intern, and Inference EngineeringSwyx [00:00:00]: Okay, we're here in the studio with Philip, old friend from Inference Engineering, the book, as well as Baseten and everything that you've done, you and I have done before, as well as Ali. Welcome.Ali [00:00:15]: Pleasure to meet you.Swyx [00:00:15]: Waterloo intern.Ali [00:00:16]: Waterloo intern, always.Swyx [00:00:17]: When did you get “Waterloo intern” as a handle?Ali [00:00:19]: As a handle? Oh.Ali [00:00:20]: I think the rebranding happened mid-March. When I saw it was open, I was like, “I have to take it. Up for grabs.”Philip [00:00:26]: The problem is that Ali is really good at his job and is not gonna be an intern much longer.Philip [00:00:30]: So we have to figure out who's gonna get the handle.Ali [00:00:33]: Well, I'll pass the torch over to the next intern.Swyx [00:00:34]: Oh, okay. It can be, like, you just pass it to another Waterloo grad.Ali [00:00:37]: To another Waterloo intern. No, bruh.Philip [00:00:39]: Yeah.Ali [00:00:39]: Intern.Swyx [00:00:40]: Intern, yeah.Ali [00:00:40]: And no.Philip [00:00:41]: You gotta get an intern from Waterloo.Ali [00:00:42]: Yeah, I've gotta get an intern from Waterloo.Swyx [00:00:44]: Right.Ali [00:00:44]: But they have to follow the path.Swyx [00:00:45]: Oh, it could, but it could come from Baseten, so it's like whoever Baseten gets from Waterloo.Ali [00:00:48]: Right.Swyx [00:00:49]: Has the title of Waterloo.Ali [00:00:50]: It stays in the ecosystem.Philip [00:00:51]: Exactly.Ali [00:00:52]: Halfway through the internship, you either get it or you're out.Philip [00:00:55]: You should also do, like, a big graduation ceremony where you change the handle.Ali [00:00:59]: Just say it.Philip [00:00:59]: For everybody.Swyx [00:01:00]: You guys are good at ceremonies, clearly. We had a nice launch of the book, very successful. But before we get into all that, I wanna start off with a fun question for you. Okay, you're an expert inference engineer. What happens when I send a long query, say two hundred thousand tokens into Baseten's inference? What's the process of query through GPU model routing, balancing, all that? What is all the stuff that we don't think about?Long Context Requests, KV Cache, and Cache-Aware RoutingPhilip [00:01:26]: With a long query specifically, the first thing that I'm gonna ask is, “Have you sent me this query before, or at least part of it?” and I really hope you have, because it's gonna be a lot easier for me and a lot cheaper for you. So the first thing that we're gonna look at is some cache-aware routing, where we're going to see, we probably have a number of instances, a number of replicas up serving whatever model you're hitting. We want to send this one to something with, number one, available prefill workers, and number two, ideally some cached input already there so that we can skip prefill on at least part of these two hundred thousand tokens. If you're doing two hundred thousand tokens, it's probably coding or a multi-turn agent or something where you would expect to have that cached. If you don't, we're gonna have to send it to a prefill worker. We've at least on certain models disaggregated prefill and decode, so you're going to have one set of GPUs that's solely going to process the input, create the KV cache, and get you your first token, and then that's going to be passed over to a separate set of GPUs, which is going to run decode. We're going to iteratively make those tokens. We're probably going to have some speculator model in front of that. I'm going to assume that you're doing coding, and because of that, our speculator model, which assumes you're doing coding, is gonna have a high draft token acceptance rate. If I'm wrong and you're asking me to summarize every Harry Potter book, it's gonna be slower. And then we stream that output to you and account for it, charge you, a couple of pennies and say, “Hey, would you like to send another one?”Swyx [00:03:04]: Except Baseten doesn't charge by pennies.Philip [00:03:07]: Well, yeah, we charge. I'm assuming that we're talking about the public model APIs. If you are setting up a dedicated deployment, then yeah, it's not pennies.Public APIs vs. Dedicated DeploymentsSwyx [00:03:18]: Yeah, one of the key differentiators when I was talking with Baseten initially was that people who want very high volume just need to rent by the box, ‘cause then it's up to you to figure out how to saturate the box.Ali [00:03:31]: And more often than not, it's, like, way cheaper if you're pushing, like, millions of tokens per hour, if you just pay per hour instead of pay per token.Philip [00:03:37]: Yeah, they do. I think that we've increasingly seen a lot of demand for the pay per token APIs, just because everyone wants to try open models, and then once they find a use case that's really sticky, then they move over to dedicated.Swyx [00:03:51]: Is there a best practice on when it's time to swap over?Philip [00:03:54]: Couple reasons. Yeah, reliability, that's a big one, right?Ali [00:03:57]: Like, if they have a very specific use case, they want you to train something specifically for them, like they want their own spec dec, for instance, for their own traffic.Swyx [00:04:04]: Spec dec is speculative decoding.Speculative Decoding and Custom SpeculatorsAli [00:04:05]: Speculative decoding, yeah.Swyx [00:04:07]: You have to explain.Ali [00:04:07]: Sorry. Like, speculative decoding is like, if you have a huge model, right? And so the model is going to be generating one token at a time every single turn, every single forward pass. So we attach, like, this little, like, parasite, like this layer that goes on top of the model, and this model just has to predict. It does three very fast autoregressive forward passes, and it will predict, like, three certain tokens, and then you do one forward stage over the entire original model in order to see if those predictions were correct or not, and then you accept them or you reject them. Now, this draft model is traffic specific, so if you, like, Philip said, if you're summarizing Harry Potter books, I can train exclusively that draft model on Harry Potter books, and I can guarantee you that I'm gonna accept the three tokens every single time. And so with that case, I increase your decode speed. I wouldn't be able to provide this to you if you're a shared endpointSwyx [00:04:53]: YeahAli [00:04:53]: ‘cause I have no idea if you're doing Harry Potter, if you're doing coding, if you're doing English. We don't know. Also, there was a thing in the book that mentioned that if they really cared about a specific threshold, chapter four, I think. Do you remember that?Philip [00:05:06]: Yeah. The things that you can do is you can set a specific, like, batch sizing, a specific, like, parallelism strategy if you're trying to optimize for, like, throughput versus latency. You can. Maybe a NVFP4 quant doesn't pass your benchmarks and you wanna run a model at higher precision, you could do that. There's just a bunch of reasons why you might wanna have your own endpoint and the biggest one, of course, just being, like, you don't have to deal with someone else doing a hundred million tokens of benchmarking traffic at the endpoint when you happen to be trying to serve your users.Swyx [00:05:40]: Yeah. I think one thing that is. That is a classic journey. Like, it's people is asking the, what happens when you type Google into the browser. Tool calling, is that just, you're generating JSON or is there more complication beyond that?Tool Calling, JSON, and Structured OutputsAli [00:05:58]: Certain customers that we have, they have their own post-trained models, and so they demand a tool calling that's not just, like parse a file or go find the weather. It's something that's very specific and you have to do post-training on this. And if the post-training on the model is not good or if the quantization after the post-training to get the inference to be fast, the model will struggle reading the JSON file and reading the tool calling. But it doesn't require its own like sandbox. It's not like it's going to use that tool calling to like escape a sandbox or like it doesn't have to be contained. It can just be a normal dedicated deployment. The challenge with tool calling more and more seems to be that the companies want certain tool calling which is a very sensitive thing to train. And because you're dealing with all of the JSON outputs, if it doesn't like close the end of the request in a very certain manner, you end up with a model that did the tool calling and like the thinking and so as a result of that, it didn't see the result and just hallucinated the result as it decoded. That seems to be the most challenging thing with tool calling, not really the sandboxes model.Philip [00:06:56]: Yeah, that's a challenge on the training side and then on the inference side, there's work that you can do to scope the possible output. So we published this at this point close to two years ago, the solution to this problem which is you make a state machine and you use that to constrain the output to a specific format. So this is the structured output problem. If you remember backSwyx [00:07:27]: Yeah, the specific grammar is,Philip [00:07:29]: Yeah, exactlySwyx [00:07:30]: GML had this thing.Philip [00:07:31]: Yeah. So it's like the old-school “make sure this is only JSON”, return only JSON orSwyx [00:07:38]: YeahPhilip [00:07:38]: Grandma's gonna die type of prompts.Swyx [00:07:39]: Is it BNF grammar? At some point OpenAI had released a thing that was like, yeah, if you want to constrain your output, write BNF grammar, back as NOR.Philip [00:07:47]: In our inference system, it's just a specified output format. And you get the guarantee that your output's gonna be structured along that format. And so applying that to tool calls can like help cut down on. You can still call the wrong tool or call no tool. It doesn't solve the certainty problem but it at least solves the output structuring problemSwyx [00:08:10]: YeahPhilip [00:08:10]: Within tool calls.Swyx [00:08:12]: And MCP is just another form of tool, right.Philip [00:08:14]: Yeah, exactly.Swyx [00:08:15]: As far as there's no special thing there.Philip [00:08:16]: The thing I'm always like explaining to people is the LLM is not capable of doing anything. It's only capable of making suggestions of what to do and then if those suggestions are formatted in a certain way and applied to a system that knows what to do with them, then an action occurs.Swyx [00:08:32]: Yeah. Part of the fun stuff is, this is solved outside of tool calling too. Like in an agent loop if the output is not correct or you're right, like reasoning, tool calling was done in the reasoning trace, just be like, “Oh, I don't know what to do. Let me just try again.” And it might get there after a few tries. And on your point of training, sometimes this is harder in smaller models, so you don't have the same exact quality outputAli [00:08:56]: Right.Swyx [00:08:57]: When you just swap from a big model, right?Ali [00:08:59]: Yeah. I will say that, before, I think we need to go back to inference engineering proper.Ali [00:09:04]: But, I had expected that something would replace JSON because it's hard to stream JSON ‘cause JSON must be complete and you must have open and close brackets and everything. So it's hard to parse something or validate something while it's being streamed. So people invented all sorts of things that are like, I forget the name of some of these alternatives, but it's something like TOML, something like YAML. But JSON seems to be dominant still.Philip [00:09:30]: The JSON outputs aren't that long, right? Like you could have a long-- ‘cause tool calls also contain the arguments in them and perhaps for a certain tool you might pass like a very long argument. But my impression of the median tool call is that it's a relatively small number of tokens, right? So I would expect that speculators are generally fairly good at something as formatted as JSON. And so you would have like a pretty fast decode step there and that the streaming wouldn't be as valuable, but maybe I'm wrong about that.Ali [00:10:02]: I think you're also bounded by the software or that the model is gonna integrate with if the software is built with JSON for the tool calls or if the company that you'- if your customer says that this is how our software works and our tools are interfaced with JSON, you can ask them to like, change their software and say like, “Yeah, this is gonna be better for the model.” but like with the right training shouldn't be that much of a difference. Also more profitable if it outputs more tokens probably.Swyx [00:10:25]: Depends on your business model.Swyx [00:10:27]: It really depends. But I will say that, as a writer with like experience a lot with generated output, I do try to move from text to JSON text which is very long JSON, right? Like there's paragraphs in every field because I'm trying to structure it, right?Philip [00:10:44]: Right.Swyx [00:10:44]: I want you to first make factual statements, then make opinions then make bullet point summaries, have dates, have entity references have your sources for references, all these things. Anyway, so these are things that like I think people who really experiment with structural output have to really care about. But, let's, let's recurse up the stack a little bit. Before we started recording, you mentioned something really cool, which is that there's a lot of engineering that-- inference engineering that goes on when a new model provider releases a new model, right? So let's call it GLM-5.2, Kimi K3. I had previously assumed, especially if it's like, well, GLM 5 to 5.1 to GLM-5.2, like that you've supported them before. Is it that much work?What It Takes to Support a New Open ModelAli [00:11:26]: It's a lot of work.Swyx [00:11:28]: Yeah. Okay. So like, a lot of people, all you guys, right whenever a new model launch like, people rush to say like, “Oh, Hugging Face supports this, Fireworks supports this, Spacetime supports this,” and I'm like, “Yeah, of course we support it.” But what goes into that? What goes intoPhilip [00:11:40]: I think it's more than just support it too, right? It benefits the consumer a lot. Like I think it was with Kimi K2.5 or GLM-5.2 the latest, there was an inference war, right? X provider is at 90 tokens a second. The next day we're at 150. The nextSwyx [00:11:55]: I kinda kicked that off with the GLM-5.2.Swyx [00:11:58]: I wrote a Twitter article about. It got like half a million views,Ali [00:12:02]: Based on being numberSwyx [00:12:03]: YeahAli [00:12:04]: Or it's for something else.Swyx [00:12:05]: Yeah. Which,Ali [00:12:06]: Oh my GodSwyx [00:12:07]: Which then got everyone really excited about, hey, how can we, bend tracks a little bit further and,Philip [00:12:14]: There's a difference between support the model, as in I can make a token out of this model, and support a model, as in I have a production-ready API from this model.Philip [00:12:26]: Getting to the point of I can make a token out of this model is not that hard because generally the, open source inference engines, vLLM, SGLang of the world oftentimes even receive weights ahead of time, maintainers do, or the people making the model merge PRs to ensure support. So you generally can, just get it working on the standard open source stack without too much pain in most cases. The challenge is, every inference company is gonna have own proprietary stack. Some open source components, some in-house stuff. And for any arbitrary model, there's going to be some new stuff. Sometimes you get lucky, like K, two five to two six was, like, pretty similar.Quantization, Speculators, and Production ReadinessAli [00:13:16]: Yeah. It was pure continued post-trainingPhilip [00:13:18]: YeahAli [00:13:18]: If I remember correctly.Philip [00:13:19]: Even in those cases, there's still stuff you have to do. You have to redo the quantization work. You're taking the model from. Generally, these models are not released in NVFP4, and we want them to be in NVFP4 for maximum Blackwell compatibility. So we have to perform that quantization, and, calibrate the quantization to make sure that we're not causing any regression in the model's intelligence. And then we also have to train the speculator, as we've talked about. Generally, we have. We have ZDR, zero data retention on our model APIs, so we don't know exactly the traffic that people are sending us, but we know what's popular. We know that coding use cases are popular. We know that agents, agentic use cases are popular. So we can get public data sets that are representative of that traffic and train general speculators. Now, with speculators today, you need to train the speculator using the base model itself because you're getting hidden states out of the model from running inference on these specific prompts, and that is the training data you use to create the speculator. So there's that process which you need the real model weights for. And then there's of course just the process of, standing up all the infrastructure behind it, loading all this stuff, testing it. And then when there's a new model with a newer architecture, I think that, like, the DeepSeek models tend to be the most challenging as they have, like, the most novel architectural stuff going on, model after model. But every new model has something. Kimi K2 had. Oh, sorry, GLM-5.2 hadAli [00:14:53]: Sparse attention.Philip [00:14:54]: Yeah,Ali [00:14:54]: YeahPhilip [00:14:54]: the DSA.Ali [00:14:55]: Right. Which is brought from DeepSeek.Philip [00:14:57]: Yeah. AndAli [00:14:59]: So you can copy-paste then?Philip [00:15:01]: It kindAli [00:15:01]: I don't know how this works.Philip [00:15:02]: So, like we had to, like, build support for that into our runtime. And you're right, like it is really interesting the way that all of these open source labs borrow from each other. For example, like GLM-5.2 doesn't have vision. So something that, Haley, a guy on our team, if we could take a look at this, he, like, grafted the Kimi vision encoder onto GLM-5.2.Retrofitting Vision into GLM-5.2Ali [00:15:27]: We'll be training the projector.Philip [00:15:28]: Exactly. So if you think about, like, the encoder, there's the encoder, which is the part that looks at the image and turns it into latent information, and then there's the projector which likeAli [00:15:38]: You can say latent space. It's okay.Philip [00:15:41]: And then there's the projector that maps it onto, the model itself, and then there's the model weights. You don't wanna mess with the model weights because you run a chance of making the model dumber at something else for the purpose of giving it vision. So instead, Haley started with just a projector, which is only a handful of millions of parameters.Ali [00:16:02]: That would be, yeah.Philip [00:16:02]: Yeah.Ali [00:16:03]: Can you show the training one?Ali [00:16:04]: Like the way it groksPhilip [00:16:05]: YeahAli [00:16:06]: Very interesting.Philip [00:16:06]: And maybeAli [00:16:07]: That right therePhilip [00:16:07]: Maybe Ali, you should take it from here. You've got a betterAli [00:16:10]: Ooh, double the sandPhilip [00:16:11]: Understanding of this than I do.Ali [00:16:11]: Yeah. You can see, like, he. The way he trained this is really cool. At the beginning, he was training it using just like, “Here's a picture of a mountain. Can you describe what's in this mountain?” And that caused it just like the first, learning walls. Like here you can see this all we're trying to teach it is to translate the encoded. Like it's already taken the encoder from Kimi K. It's taken the image. It'Philip [00:16:31]: Yeah. FrozenAli [00:16:31]: FrozenPhilip [00:16:32]: With adapter.Ali [00:16:32]: Exactly.Philip [00:16:33]: Yeah.Ali [00:16:33]: So the brain is frozen and the eyes are frozen. It's just we're tryingPhilip [00:16:37]: AlignAli [00:16:38]: Interconnect between the eye and the brain, right? So the projector. And so you take the tokens and then he's like, “Oh, can you describe what's in this image?” And he's like, “Oh, it's a mountain,” or it's a person or it's a human, whatever the case is. But that didn't cause complete understanding. So he changed it such that every image was associated with a data set of questions. Like, does this image have a white male? Does this image have birds in the top corner? Does this image have a scientist in it? All of that stuff. And it would have to answer questions correctly. And using not just training on describing an image, but being able to answer question, another question, answer over time. Like you can see the grokking, which is like genuinely insane, that retrofitting vision into a large LLM can learn to that extent. And even for images that it doesn't perform well on, for instance, if you ask it a picture of like Stephen Hawking, “Who is this?” Maybe it doesn't get it, but it will say something like, “This is Albert Einstein.” Like it still understandsPhilip [00:17:25]: Close enoughAli [00:17:26]: That this is a scientist who is a man who has, some significant achievements, all that stuff. So that's like really cool.Philip [00:17:32]: Yeah. So, we've covered Hao Tian before, who the author of the LLaVA paper that did this, a while ago. And I think that's very foundational work for anyone who hasn't done vision work before.Ali [00:17:41]: Same with the CLIP and MetaCLIP, where you go from just captioning to building out questionsPhilip [00:17:47]: RightAli [00:17:47]: Off the image and how much better you can get performance.Philip [00:17:50]: Right. Right. Right. Yeah. But what's, what's so exciting about this is if you look at a model like this. Now, this is a little bit more of a research project. It's not. It got to 56% on MMLU Pro, I think. So not quite frontier. But if you're running this model, you haven't suffered any loss on your GLM-5.2 quality. If you don't have an image, it'll just behave exactly the way it used to. And ultimatelyAli [00:18:14]: Which in the inference code you literally do not include the other part, right?Philip [00:18:18]: Yeah. You would just skip the encoder if you don't have an image input.Ali [00:18:22]: Okay.Philip [00:18:22]: Just confirming.Philip [00:18:23]: YeahAli [00:18:23]: Does it affect a lot on the overall inference side? Like you're not adding much, you're adding a very small vision encoder. These are typically likePhilip [00:18:30]: They're super fineAli [00:18:31]: Less than a billion parameters, right?Philip [00:18:32]: Yeah. It's, - There's a little bit less standardization among vision encodersSwyx [00:18:37]: YeahPhilip [00:18:37]: So the support matrix can be a little bit, sparser. But overall, yeah, it's a pretty, it's a pretty minor component of the overall system. And ultimately what you get out of the system is all of a sudden you have Kimi Vision, GLM weights, and DeepSeek attention all in one model.Open Source Model Grafting and Franken-MergesPhilip [00:18:56]: And that's, I think, a lot of the power and beauty of open source, is that you can take all of these different components and combine them together into a system that's better than anyoneSwyx [00:19:05]: YeahPhilip [00:19:05]: Can be individually.Swyx [00:19:06]: People used to say that you would also do Franken-merges where you would take likePhilip [00:19:10]: YeahSwyx [00:19:10]: Layers from each model.Swyx [00:19:11]: Does anyone do that anymore?Ali [00:19:13]: Well, to your point previously when you were mentioning like, the work that goes into supporting a model when it first comes out, like GLM-5.2 or MiniMax M3 or whatever the case is. Sometimes you do have to like, you do have to switch out some things. Like, for instance, the MiniMax M3 head uses full attention, and with full attention you end up with this like insane bottleneck in spec dec ‘cause you're doing auto-regressive token generation for three tokens, and you're doing this like N squared over all of the tokens that are in your sequence. Your KV cache is like very large because it's not sparse, it's not top K. So we find it better to like, okay, we're gonna replace this, we're gonna replace this layer with a layer from another model that's using like GQA, for instance. And then just with the right training, you can get it to have the same acceptance rate. So it is very possible to retrofit layers from other models and very much needed. If a layer is like inefficient, the training just becomes the challenge, like how do you ensure that you train it properly? Which again to your earlier point is like the mesh between training and inference. As in like you need very good training in order to do fast inference. That's like, I feel like more and more becoming true.Swyx [00:20:21]: Yeah. Anything else on the support side when you say like get it to fully production ready?Loop Detection, Race Conditions, and Non-DeterminismPhilip [00:20:26]: Yeah. I think that there's also a question of just, we can test a model to a pretty extensive degree, but we're trying to get it out quickly and then you see a bunch of other people test it and you get interesting results. There was an issue with, GLM briefly where we had some like mode collapses where it would just output the same token over and over again for certain prompts on certain temperatures. Like once you expose an endpoint to the real world, there's going to be, so many more varieties of things given to it that you're able to, discover and patch things. So it's not just a, day zero process, it's then like for the first week, for the first month, if a model remains popular, like how do you both fix bugs and then continue to push the envelope on performance?Ali [00:21:21]: What do you mean you don't want your model outputting S?Swyx [00:21:24]: Is there loop detection on that stuff, by the way? It still happens like quite a lot, which is surprising.Ali [00:21:30]: We have like we, in our endpoint, like if a model was to output the same token like four plus times, we just cut the generation. We say like, “Oh, sorry, this-- Like try again,” or like we will reprocess the request. ‘Cause we know then, like if it, like if, yeah, it's four times the same token, it's probably collapsed.Swyx [00:21:45]: Yeah. Is there a way to opt out in case I really want that?Ali [00:21:48]: You want that?Ali [00:21:50]: I think there's a way that we have to handle it. I'm not exactly certain, but I feel like in certain models, like when they output something like you can imagine, like a table for instance, and so they want, they wanna draw like 12 dashes and 12 dashes. Yeah, I think there's a way for that to happen. I think we only do it on certain tokens. Like we exclude certain special characters.Swyx [00:22:07]: Yeah.Ali [00:22:07]: So we only do it on like certain like S is the most common almost. GLM-5.2Swyx [00:22:11]: OhAli [00:22:11]: And I think it was DSV 4 as well. Like you'd just have like looping issues where like you literallySwyx [00:22:17]: ItAli [00:22:17]: Just have like S.Swyx [00:22:18]: Yeah. Is there a special, something special about S? No, just randomlyAli [00:22:21]: It just seems to be the one token involved.Swyx [00:22:23]: Yeah. And it'Philip [00:22:24]: Is thereSwyx [00:22:24]: And it's only temperature 0Ali [00:22:27]: NoSwyx [00:22:27]: Even at other temperaturesAli [00:22:27]: Even at like 0.9 or whatever, it will still, it will still collapse.Swyx [00:22:30]: That's weird, right?Ali [00:22:30]: It's, it is an inference problem to be honest, like a software problem. Like oftentimes, the image you run will-- like NVIDIA will release an image for instance, and if we will upstream the changes from their latest TensorRT-LLM image into our stack, we'll find that it fixes it. Or oftentimes this will only happen in an inference engine that you're using like SGLang. But if you were to switch to vLLM, that isn't the case. So it seems to be like an extremely like deterministic software issue and not really a model issue. It's not like a weights problem. Like I'- we'll say like, “Oh, it's a problem with the quant. We did PTQ wrong,” right? But that isn't, that doesn't make sense because the same weights used with a different inference engine does not repeat the problem. And sometimes it's, the kernels that are being used in the backend have like these very subtle sometimes race conditions, where if you were to use this model hosted on one cluster, you will never get this problem.Swyx [00:23:19]: Oh my God.Ali [00:23:19]: But if you host it on a different cluster, you will. And the reason is the KV cache transfer from a node to node in that one cluster is using a slower interconnect than the node to node in another cluster. So that exposes the race, whereas in another cluster it doesn't. So then you end up just like, okay, this model is not gonna be hosted on this cluster. We're gonna host it on, another cluster because that cluster exposed that problem. But then it ends up with like, okay, is it the software? Is it the model weights or is it the hardware?Swyx [00:23:42]: There is a thing about this with temperature 0 still not being deterministic, right?Ali [00:23:46]: Right.Swyx [00:23:46]: Mostly because of hardware. Even at temperature 0 same model, you won't always get the same output.Swyx [00:23:52]: Even-- But I'm surprised by the race condition one because, I thought PyTorch was a graph that like guarantees that you at least, execute things in the right order.Ali [00:24:02]: Well, yeah, true. Like I'm not, I'm not saying that there is. Like well, you have things like PTL optimizations where like you can start a kernel before the end of the previous kernel, and that's like ‘cause you want to do that because there'sSwyx [00:24:12]: It's like pipeliningAli [00:24:12]: Expense. Exactly.Swyx [00:24:13]: Yeah.Ali [00:24:13]: But it'- But you don't do it cleanly. Like you overlap a little bit of the execution. No, it is very possible that the kernel itself, like that one block that is supposed to be running in this instance of time, that kernel itself has a race condition. For instance, like a missing barrier. Like often if you're designing a kernel and you want it to make it to be very fast, if you don't test it extensively, you'll, you'll have certain threads access data points from registers before they've been written to by other threadsSwyx [00:24:36]: YeahAli [00:24:36]: For example, because like your barrier is wrong or your synchronization was wrong. But yeah, like the testing itself is very difficult in those like, andSwyx [00:24:42]: And there's no like borrow checkerAli [00:24:45]: What does that mean?Swyx [00:24:46]: Like Rust. Like the. If you're trying to have like memory safety It sounds like a comparable problem.Ali [00:24:52]: Well, yes, but you're working in CUDA, right, NVIDIA GPUs. Like- You just need a higher level language like modular Maybe that's what modular is supposed to do. I don't know.Quantization Quality and Vendor FidelityVibhu [00:25:00]: How do you see keeping quality of the model? So you talked about all these steps of, okay, you gotta do quantization, train your own speculative decoderAli [00:25:07]: RightVibhu [00:25:07]: Run on different hardware. Looking at other model providers, okay, you kicked off a inference speed race on the consumer end. What goes into keeping quality the same across them, right? Sure, you can run benchmarksAli [00:25:22]: YeahVibhu [00:25:22]: But, like, how do you determine how much quantization are there standards? What goes intoPhilip [00:25:27]: There's a few things on quality. Most inference optimizations are lossless. KV caching, for example. You are just recomputing or preventing recomputing the same values. Speculation, of course, if a draft token is wrong, it gets rejected. The main lossy optimization is quantization. And that really comes down to, number one, data format, number two, which parts of the model you choose to quantize, which layers, and number three, like doing a lot of calibration on the quantized weights, to ensure that you're preserving all the outliers. There's other tricks that you can do, though. A big one is long context, ‘cause one thing you asked at, right at the beginning is, “Oh, what's gonna happen if I send a 200,000 token request in?” So with a long input sequence, you need to, store a lot more information. You need to process a lot more tokens. And so even if a model has a context of a certain length, you might, as an inference provider, choose to build an API with a shorter context length, and of course a full length one as well. Because if someone doesn't need the full million token context, for example, you can get them better performance. I don't know if that's exactly like quality of the model. The way that I think about quality is to what degree are we faithfully serving the original model? If you think of a golden implementation of a model that performs exactly the way the model is designed to perform, I think of quality as how close are we getting to that, 100% fidelity of the model.Philip [00:27:13]: You can also, of course, think about quality from the training side and how do you push yourself past 100%. But when I think about purely inference optimizations, it's getting faster while staying as close to that 100% fidelity mark as possible. And certainly our standard internally is that, like you should not be able to tell the difference between our API and a, official API. I think Kimi in particular does a good job of vendor benchmarking hereAli [00:27:41]: YesPhilip [00:27:41]: Where they haveAli [00:27:42]: They released an actual vendor benchmark.Philip [00:27:43]: Exactly, yeah.Ali [00:27:44]: ‘Cause they accused, some people, Amazon? There was some provider that was not doing very well on Kimi's benchmark.Philip [00:27:50]: Yeah.Philip [00:27:51]: So, with Reflect we probablyVibhu [00:27:52]: This was a long time ago, right?Philip [00:27:54]: No.Ali [00:27:54]: Yeah, like threeVibhu [00:27:55]: They alsoAli [00:27:55]: Four, five months agoVibhu [00:27:57]: This also happened with, I don't remember which model, but they pulled out quite a few, and then they started a whole chart about this. It might have beenPhilip [00:28:03]: Kimi Vendor Verifier.Ali [00:28:04]: Yeah.Philip [00:28:05]: Yeah.Ali [00:28:05]: Yeah, ‘cause you, ‘cause you'd be pissed, right? Like if you'Philip [00:28:07]: Yeah.Ali [00:28:07]: If like if I'm a consumer and I'm using like Amazon's endpoint for instance, and I've used Kimi and I'm like, “Oh my God, like this is bad,” I'm not gonna say, “Oh, Amazon quantized the model in a bad way.” I'm gonna say, “Oh, Kimi sucks.” Right?Philip [00:28:17]: Yeah.Ali [00:28:17]: So it seems like that makes sense.Philip [00:28:19]: Yeah, they care. They care.Vibhu [00:28:21]: Justifiably.Ali [00:28:21]: Yeah, justifiably.Vibhu [00:28:22]: This is probably a stupid question, but just checking, has anything improved from main quantization?Philip [00:28:28]: Yeah.Vibhu [00:28:28]: Like, is quantization always strictly worse?Ali [00:28:30]: Well technicallyVibhu [00:28:32]: NoAli [00:28:32]: It's a lossy. QuantizationPhilip [00:28:33]: YeahAli [00:28:33]: Is a lossy, it's a lossy implementation.Philip [00:28:36]: Speed improvesVibhu [00:28:36]: Speed improves.Ali [00:28:37]: It the number, likeVibhu [00:28:38]: No, I' always look for inverse scaling laws.Philip [00:28:40]: Yeah.Ali [00:28:40]: Yeah.Vibhu [00:28:40]: This is something I learned from Noam Brown, where like things that normally act in one direction sometimes do.Philip [00:28:45]: Well, technically when you run a benchmark, because these models are deterministic, sometimes your,Ali [00:28:52]: YeahPhilip [00:28:52]: NVFP4 quant is like, two basis points higher than yourAli [00:28:56]: No, it's noise. It's noise.Philip [00:28:57]: Yeah, exactly. I'm like, yeah, it's, it's within. That's why I always say within margin of error.Philip [00:29:01]: And I stopped saying that because everyone assumes that what is, well, within some margin of error, we're barely inside of that to the worst, so we're saying. But yeah, sometimes it's just like, gives you a higher output score. But like Ali said, that's noise. To my knowledge, you're not necessarily making the results better. You're just trying to, again, like keep your fidelity as close to 100% to the original model.Layer Selection, KL Divergence, and Better QuantizationAli [00:29:27]: There is, to your point, research that we did on MP. I don't know if you are able to pullPhilip [00:29:31]: YeahAli [00:29:32]: A tweet we did. One of our research interns, Joshua, I think it's a tweet on how we have 20% better quantized GLM-5.2 than NVIDIA. Essentially what we found throughout like this month research is, okay, quantization is a lossy. It's. You're compressing the data from, occupying 16 bits to occupying, four bits, for instance. And so you're losing some information, and you're trying to minimize that. And so when I say that I'm gonna quantize the model, my job becomes how do I find the layers that I can quantize, and how to find the layers to not. For instance, with image models, I don't quantize modulation layers, and I don't quantize out projections because those two are. Like out projection is what you see as the user. Modulation is what the model sees or understands. Right, exactly. And so to his paper, do you have the. It doesn't have the. Yeah. It's a long paper. I don't know if I can findVibhu [00:30:25]: If there's a part to search or it's probably in the thread.Ali [00:30:28]: It's probably in the thread.Vibhu [00:30:29]: Yeah.Ali [00:30:29]: But the long and the short is it is very possible that quantizing more of the model makes the results. Like if I have a model that I quantize layers one, five, and 10, and another model where I only quantize layers one and It is possible that the model in which I quantized more information is going to perform better because the quantization errors have canceled out. And so what Joshua showed in his mathematical proof where he had like a verifier in, is that you can predict which layers are going to have quantization errors that will cancel out with each other, and you choose to quantize those layers. And so the result of doing this mathematical quantization is you end up with a model that's 20% more quantized than another provider, so you get 20% more throughput of it because there's more layers than running an NVFP4, and your quality is better than that other quant because the layers that you chose to quantize have their errors cancel out, like one layer skewed to the right one layer skewed to the left, one layer skewed to the right. Your final logits distribution is more similar to the original distribution of the model, so you have better fidelity. And so the way we proved this was with KL divergence. So instead of just scoring on the benchmarks, we scored the KL divergence between the logit distribution of the quantized model and the logit distribution of the original full precision model, and we showed that with this technique we get. If your probability distribution on the logits which token it wants to select is more of the same as the original model, you're probably gonna end up staying true to the original model. So yeah, so it seems like previously before this, it seemed like the industry was, well, the more you quantize, the worse it's gonna be, ‘cause the more loss you introduce. That's not exactly, not necessarily true. So yeah, doesn't improve it, but can cancel out.Philip [00:31:57]: I think it might be this, but reminds me a good bit about pruning where you can prune off certain layers.Philip [00:32:03]: But very interesting. Didn't know this was a whole paper you guys put out.Ali [00:32:06]: It's. Fun fact, it was originally 72 pages, this paper, and then we decidedPhilip [00:32:11]: WowAli [00:32:11]: We can't tell. We couldn't release it. So it's now 45.Swyx [00:32:15]: Still 39 pages, so very substantive. We talked about evals and all these things and, like what's possible in terms of speedup? Like it's like probably like the numberInference Speedups and BenchmarkingSwyx [00:32:25]: Thing that people do wanna care about, and it's something that you wrote about in your post. Like official API is 70 tokens per second, and you push it up to 90. Is that like a normal thing?Philip [00:32:36]: So what's cool about working in inference, the reason that I think inference is going to be a useful place to do engineering for a long time, is that if you look at highly optimized domains like, say, finance, if you're in finance, you measure how much better you got in basis points. It's like, “Oh, I got five basis points better, like twentieth of 1% better,” that's huge news because everything is so optimized. When we publish optimizations, it's 20%, it's 100% it's 200%. So there's still probably like a lot further to go, honestly. Like you'll, you'll know that inference is pretty much solved when researchers start publishing about how they got 1% faster at something.Swyx [00:33:19]: Which by the way, because I am from the finance background, in the ‘70s, that was the margin at the time. When you did quantitative finance research, you would findAli [00:33:27]: And like 20%, tens of percent.Swyx [00:33:29]: That's. Yes.Philip [00:33:29]: Yeah.Swyx [00:33:30]: And now it'Philip [00:33:31]: Tiny fractionsSwyx [00:33:32]: For those people interested, look up Andrew Lo's paper. He had a really interesting illustration of quant, stat arb, distribution, narrowing down from like those kinds of 20% differences in the ‘70s, down to nothing today, which is very cool.Philip [00:33:48]: Exactly, and we're at the beginning of the same type of thing. Now benchmarking is hard. I think anyone will tell you that, and benchmarking provider speeds is hard because there's so many variables that go into it. What hardware are you using? How much load do you have on the system? What's the exact nature of the prompts and input and output sequence lengths? All that stuff. But overall, when you start stacking these improvements, you're looking at multiples. You can look at it. The most common form, of course, is TPS, tokens per second, which is bad naming by us in the industry, ‘cause there's two tokens per second. There's tokens per second, the throughput number, and the latency number.Ali [00:34:31]: TTMT, yeah.Philip [00:34:32]: Like total tokens per second out of the, out of the GPU as a throughput number. Most people only care about tokens per second as the latency number, which we should call ITL, intertoken latency, but we don't.Philip [00:34:44]: Anyway, so you can imagine a standard API without many optimizations for a 1 trillion parameter model operating somewhere in the 30 to 50 tokens per second range for reasonable traffic profile. And we generally see the goal of, pushing to 10X that. But, not necessarily day zero, but by stacking enough optimizations, if you have, say like four optimizations, each of which doubles performance. Or sorry, three optimizations, each of which doubles performance, then you stack that up, that's an 8X gain. That's the order of magnitude that we're working with in this space. We're trying to make things substantially faster, not just go from like 70 to 90.Swyx [00:35:38]: Are you saying you've. You have done that?Philip [00:35:40]: So let's say you have as a reasonable baseline, 30 or 40 tokens per second. You can achieve 10X that. So like on GLM-5.2, if you run it unquantized, perhaps on H100s even, and you're just using an off-the-shelf inference engine with no particular optimizations, no speculator, nothing extra around like KV routing, no disaggregation, you're, you're probably, yeah, looking at that like 30 to 40. You think that's like a reasonable baseline?Swyx [00:36:12]: Right. Right.Philip [00:36:12]: To get to something like 10X, there's a lot of trade-offs that you're making. If we're running at more like a 300, 400 tokens per second range, you are using the best hardware possible. You have a optimized speculator. You have done all of your quantization work. You are Seeing a pretty high cache hit rate. You are running with a reasonably small batch size and a parallelism configuration that is tuned for latency versus throughput, but it is possible. So the spreads that you see if you, like, go on artificial analysis or you go on OpenRouter and you look at, the worst provider to the best provider, oftentimes can hit that range. 10X is of course very aggressive. It's oftentimes maybe more of a four to six times improvement. But that's the performance that makes us really excited, is when we can get these huge gains, not just go from 70 to 90 tokens.Stacking Optimizations: NVFP4, Speculation, and DisaggregationAli [00:37:19]: It's also, like, hardware dependent. Like, ifPhilip [00:37:20]: YeahAli [00:37:20]: If you have a thing where you're serving it on just, like, a node of H100s and then you throw, like, you shard the model across, like, four nodes of B200s. Like, you can definitely increase the speed with just throwing more hardware at it. Like, normalizing for the same exact hardware and the same number of GPUs.Philip [00:37:35]: Yeah. Then you're looking at, like, a two to 4X improvementAli [00:37:38]: Right. RightPhilip [00:37:38]: Depending on the inference optimizations. So yeah, it's. Some of it's, what's the call, and some of it's who's the driver.Vibhu [00:37:46]: If you break down the two to 4X, say the example is run GLM-5.2Ali [00:37:51]: YeahVibhu [00:37:51]: On B200sAli [00:37:53]: YeahVibhu [00:37:53]: Single node, right? What's, like, the cost trade-off for effort to get, like, the last bit of juice out versus what should people just think of, right?Ali [00:38:01]: Spectre quantization. Yeah.Vibhu [00:38:03]: Spectre quantization.Ali [00:38:04]: That's, that's, that's like 95%. LikeVibhu [00:38:06]: And how far does that get you? And how easy is that for the average person to do? So say right I wanna throw the weights of GLM-5.2 on a node of B200s, how easy is it to find speculative decoder- decoder model or already quantized model? How much work goes into it?Philip [00:38:23]: If you're doing it up front, it's quite a lot of work. If you're doing it today, there's going to be people who have published things that you can just, you can just grab some NVFP4 weights. You can grab a speculator. Yeah, if we're thinking about, like, what are the 2Xs we're stacking, going from, BF16 to NVFP4 is, it's not quite a 2X, right? It's like. I think it's about, like, 30 to 40%, from 16 to 8, and then another 30 to 40% multiplied from, 8 to 4. So that doesn't quite get you a 2X, but, like, roughly a 2X. Speculator, roughly a 2X. Disagg on top of that if you're able to get enough hardware and put enough traffic through it, another roughly a 2X. And then you add in some, double-digit percent increase from having just a better runtime with, the latest kernels and stuff behind it. And that's how it stacks up.Ali [00:39:21]: YeahPhilip [00:39:21]: So building each of those, like, building the, quantized weights is, for someone who really knows what they're doing, hours to days of work. Building the speculator, again, like, hours to days of work. And the, disagg setup, hours to days. Well okay, but like once you haveAli [00:39:39]: Once set up. Once set up. YeahPhilip [00:39:40]: Yeah, getting disagg working for the first time, I'm saying, of course, is very difficult.Philip [00:39:44]: The marginal implementationAli [00:39:48]: Like, if you're just grabbing, like if you are a person, like just a normal consumer who has access to, like, a node of B200s and you're wondering, “How can I just host it myself?” You don't need to quantize the model yourself. There's always gonna be, like, an open source quantized checkpoint. NVIDIA's gonna push one out if no one else does. You. Usually, the providers will have their own spec dec that they've trained as well. You don't need to train your own spec dec. You can just use that as well.Philip [00:40:09]: Yeah. Like, GLM-5.2 has its own MTP.Ali [00:40:13]: Right. Right.Vibhu [00:40:14]: What's multi token prediction?Philip [00:40:15]: Yes.Ali [00:40:16]: I'm justVibhu [00:40:16]: Can you explain that?Ali [00:40:16]: I'm just an expert.Ali [00:40:18]: I can do it for you in case I get it wrong?Vibhu [00:40:20]: No.Vibhu [00:40:21]: Yeah, you should correct if we're wrong, but their multi-token prediction can be used for self-speculative decoding.Ali [00:40:27]: I'm not sure. I'm not gonna correct that.Vibhu [00:40:28]: Okay. I'm semi-confident in thatAli [00:40:30]: Okay. YeahVibhu [00:40:30]: But someone can check. But it's useful to paint the story of, okay, not just the average person, but say a company wants to switch from serverless inference I wanna throw this up on. I wanna rent some GPUs, throw it up. These are the steps you take to do significantly faster than just put it behind vLLM.Ali [00:40:48]: Right.Vibhu [00:40:49]: I was waiting for a mention of Dynamo.Vibhu [00:40:51]: I feel like, that's supposed to be the baseline that you measure against.Dynamo, KV Routing, and Disaggregation ToolkitsPhilip [00:40:55]: I would think of Dynamo as less of a box system and more of a toolkit for building with. So when we talk about doing aware routing, when we talk about doing KV offloading, when we talk about doing, PD disaggregation, Dynamo fundamentally is. By the way, Dynamo is an open source library from NVIDIA.Ali [00:41:17]: We've done a pod with KylePhilip [00:41:18]: OkayAli [00:41:19]: Kyle Cranin.Philip [00:41:19]: Cool. So then your listeners know then that it supports all the different inference frameworks. And it is multi hardware, which is interesting.Ali [00:41:28]: But it's just a router, it's not like an optimizer layer.Philip [00:41:30]: Yeah. All it does, like, what Dynamo is good at, it is a library for moving information around your cluster, around your hardware. So if you have, KV cache on one place and you need it to be somewhere else, Dynamo coordinates NIXL for you to move that around.Philip [00:41:49]: That doesn't mean that, like, out of the box, you just say, “Pip install Dynamo,” and then you get, like, a massive performance speed up. It's more of a developer toolkit.Ali [00:42:01]: Yeah. I would have said it would. It comes with a set of defaults that you can then swap out.Philip [00:42:06]: It does. If the industry at large, I think, was, like, rolling out all of these deployments, standard, then I think it would be, like, a credible baseline. But, we've got to, we've got to benchmark against, like, what we're seeing in the wild.Speculative Decoding Methods: Medusa, EAGLE, n-Gram, and Spec-SpecVibhu [00:42:23]: I did wanna talk a little bit more about PD disagg, because that is probably, like, number three after quantized and speculative decoding. In your book though, I was just gonna pull out the book.Philip [00:42:31]: Yeah.Vibhu [00:42:32]: Like section 522 on Medusa, 523 on EAGLEPhilip [00:42:35]: YeahVibhu [00:42:36]: 524 on gram.Philip [00:42:37]: It's 55, would be disaggregationAli [00:42:42]: Yeah. Well, no, I just wanted to dwell a little bitPhilip [00:42:44]: YeahAli [00:42:44]: The other. Like, so what do you choose to include? What do you choose to not to include? Because there was all these other techniques.Philip [00:42:51]: Yeah.Ali [00:42:51]: Are these still relevant? Because I think they came out, like, a year and a half ago maybe.Vibhu [00:42:55]: Medusa is quite old.Philip [00:42:56]: Yeah, Medusa's old.Ali [00:42:58]: It was old.Vibhu [00:42:58]: But is it in the book as a good, here'sPhilip [00:43:01]: BaselineVibhu [00:43:01]: Baseline vanilla understand it?Philip [00:43:02]: Like you should know this.Vibhu [00:43:03]: Like I read the paper, I'm like, “ it makes so much sense.”Philip [00:43:05]: Yeah.Philip [00:43:05]: So with the book, I had a couple goals. One was to give people just a working vocabulary for the space as a whole, and the other was to give them some intuition about how each of these techniques works. As I mentioned in my AI Engineer talk, which is the first public addendum to this, the speculation space has moved much faster than everything else. So yeah, even at the time that I wrote the book Medusa, I very much included as a way for people to understand how the space evolved rather than what the most modern technique is. And now of course, there's DFlash, dSpark. There's, there's newer techniques even than EAGLE, although EAGLE is still very commonly used.Ali [00:43:51]: SpecSpecta.Philip [00:43:52]: Yes. Speculative decoding.Vibhu [00:43:54]: What canAli [00:43:56]: Oh, it's a paper by Tri Dao and it's like, it's doing speculative decodingVibhu [00:44:00]: HuhAli [00:44:01]: For the speculative decoder.Philip [00:44:02]: Oh, in spec- oh my God.Ali [00:44:02]: It's literally just an another. It's like, yeah, that's the most simple way to explain it, and it seems like he got trivial speed ups there. But it seems that the complexity with training, it's almost like in our mind at least, it's almost as complex as training GANs. Like it's like a very delicate balance and oftentimes you, it's just but yeah, it's literally speculative decoding on speculative decoding.Vibhu [00:44:21]: Speculative.Ali [00:44:22]: Yeah. We saw this paper.Vibhu [00:44:24]: It's interesting, right?Ali [00:44:24]: Yeah.Vibhu [00:44:24]: I wouldn't even expect it to be very particular to train, I wouldAli [00:44:29]: Right.Vibhu [00:44:29]: The naive part of me is like, okay, train speculative decoder.Ali [00:44:32]: But like, and it makes sense, like the whole idea of speculative decoding is you. It's like, it's like almost like the iPhone auto predict version but for a normal model, right? Like you're just, you're just, generating three tokens and you're like, okay, I'll do prefill on them. And so you save those three turns for your original model. Now your speculative decoder is doing three turns of auto regression, so why not just have an even smaller model?Ali [00:44:53]: The other question there is what are the size of speculators? So say forPhilip [00:44:58]: Right. It's like a billion parameters.Ali [00:45:01]: Like for MiniMax, it's. Yeah. It's like one layer. It's like one 60th of the original model usually.Philip [00:45:06]: Yeah. I think we should do a paper when we get back to the office.Philip [00:45:10]: SpeculativeAli [00:45:11]: SpeculativePhilip [00:45:11]: Decoding.Ali [00:45:13]: No, it's, it does seem like how, when do you stop? But then it also seems like if you're able to train spec-spec decode for instance, right? Like if you're able to have a small model that is accurately predicts what the intermediate speculator is gonna predict, that is able to predict what the original target model's gonna predict, then why not just use that smallest model directly, right?Vibhu [00:45:34]: Yeah. This isAli [00:45:35]: Like it seems likeVibhu [00:45:35]: Adjacent to the routing problem.Ali [00:45:36]: Right.Vibhu [00:45:36]: Yeah.Ali [00:45:36]: Right.Philip [00:45:37]: The thing with speculators is one of the practical constraints on using them is that you do have to run a small model on the same hardware that you're running the big model on. There is a orchestration and resource competition problem inherent in that, and that is one of the constraints on speculation in general, is that draft tokens cost resources to create and cost software complexity to manage. And so if you have like infinitely recursive speculators, you add in quite a bit of that complexity on the actual implementation within the inference engine as well, not just in the training process.Vibhu [00:46:17]: I was gonna say, I would wonder if you could do similar, like distillation and pruning of, it's the same thing, it's just a model. Can we not just distill a lot of the weights, quantize the speculator, out of my domain? The question that also comes up is, this is all for big server workloads, right? How much of this applies to, say I have this MacBook, I wanna run Gemma really efficiently. Similar problems, not the same?Local AI vs. Data Center InferencePhilip [00:46:45]: Pretty different. I talked to Selo, about this on his podcast a couple weeks ago. The difference between inference engineering for the data center and for production workloads versus inference engineering for local AI, is that we start with fundamentally like different constraints and different goals. With local AI, it's how do I fit this model onto my hardware and then make it less dumb? And with data center influence, it's how do I load this model and then make it less slow? And we care about less dumb, and they care about less slow. But the local AI inference engineering ecosystem, I think has a lot for us to learn from in the data center space. They are experts in various forms of quantization, including dynamic quantization that we just don't touch, in the pruning, in the distillation, in the, layer removal. There'Ali [00:47:42]: Layer removal matters less.Philip [00:47:43]: Yeah. There'Ali [00:47:44]: No one loves pruning really.Philip [00:47:45]: Yeah. Well, but the, but they doVibhu [00:47:46]: Which is surprising, right? But that's, that's a whole different thingPhilip [00:47:48]: Just to fit something on the laptop.Ali [00:47:50]: Right.Philip [00:47:50]: So yeah, it's a, it's an interesting, it's an interesting space. Not necessarily that like their techniques make sense for us to do in the data center, because we have different resources and different goals, but more that the process as well as the openness of that field is something to, admire.Ali [00:48:12]: Yeah. Like to your point, like, certain optimizations that would. Like for instance, Turbo Quantum Sharper, like it made such huge hype on that and we did like a whole deep dive on Twitter and like said, what is it? How does it work? Why is it good or not? And it took off and it was implemented on local devices because your memory bandwidth is so slow on like a MacBook, for instance. But try putting the same thing on like an NVIDIA GPU on a B200 Turbo quant would not be. Like, it would not be used. Like, NVIDIA - Like, NVIDIA made it clear that this is not a good optimization, and we've seen it firsthand where the overhead of doing dequantization, quantization of, in the kernel itself with turbo quant kernel, each end is much slower than the time that you save from doing the bandwidth. ‘Cause on the B200s, you have like 3.5 terabytes per second. You don't need decrease the storage that much. You don't need to do, FP4 KV cache. You don't need to use a requant. There's, there's, there's better optimizations to be made. But on Edge devices, it's extremely important, it's extremely useful. So, seems to be, like, different optimizations there, but then they're all uniquely combined with like all you wanna quantize the model, you wanna do speculative decoding, like certain common prefixes with bothPhilip [00:49:18]: Principles.Ali [00:49:19]: Yeah, exactly. Exactly. Exactly.Philip [00:49:20]: They also do a lot of work on, model parallelism, especially over, heterogeneous topology, where you have, some sparks and they are wired together with, Ethernet, DGX sparks.Ali [00:49:35]: Yeah, this is the Exo Labs guys.Philip [00:49:36]: Yeah. You have, a nu
When violent insurrectionists stormed the United States Capitol on 6th January 2021, they ruined a birthday forever. Kiely and Jessie are best friends who share the same birthday and the same big dreams of theatrical stardom. On the fateful day of the insurrection, their birthday goes from bad to worse when allegations emerge of their involvement in the attempted coup. Faced with their special day going down in infamy and spending the rest of their lives in prison, can Kiely and Jessie sing and dance their way to innocence?
Tommy Rooney and Jonathan Higgins were on hand to get word from both camps as Limerick and Galway prepare to face off on Sunday at Croke Park for the Liam McCarthy Cup. Hear from Limerick player Will O'Donoghue and manager John Kiely, then hear from Galway manager Micheal Donoghue and later players Cathal and Padraic MannionHurling on Off The Ball with Applegreen
Tommy Rooney is joined by Arthur O'Dea and Stephen Doyle for Monday's Newsround following a massive weekend of sporting action, which kicked into action as Pico Lopes and Cape Verde bowed out of the World Cup following an instant classic Round of 32 match against World Champions Argentina. The weekend then finally ended following 2 brilliant All-Ireland Hurling Semi-Finals with England hanging on at the Azteca against Mexico and advancing to the last 8. The Newsround on Off The Ball - Weekdays - 7PM Viagra Connect 50mg film-coated tablets. Contains sildenafil. For adult men with erectile dysfunction. Subject to suitability. Maximum dosage one 50mg tablet per day. Always read the label.
Since the late nineteenth century, amusement parks have been providing countless hours of enjoyment for people all around the world. Often driven by the latest technology and advances in mechanical engineering, the thrill rides at parks like Disney Land, Great America, and other independent parks offer a controlled environment to experience terror and excitement. While these rides, and the parks in general, are very safe and held to strict safety standards, there are times when the unthinkable happens—a cable snaps, a safety harness breaks—and the once safe ride becomes a nightmare for passengers. Far more often than not, tragic amusement park accidents are the result of human foolishness or, far less often, operator error. But other times, they are a bizarre fluke; a one in a million mechanical problem no one saw coming. Either way, the results can be shocking, horrifying, and even deadly. MENTIONED IN THIS EPISODE Get Tickets for Alaina's Book Tour for THE BUTCHER LEGACY! Get Tickets to our MORBID LIVE show at Radio City Music Hall with Special Guest Jonathan Van Ness! References Akst, Daniel. 1982. "Short circuit found in fatal amusement ride." The Record (Hackensack, NJ), August 5: 3. Anaheim Bulletin. 1973. "D'land visitor drowning victim." Anaheim Bulletin, June 23: 1. Associated Press. 1980. "Roller coaster death probed." Free Lance (Hollister, CA), April 3: 10. —. 1998. "Disney visitor had no chance, surgeon says." Sacramento Bee, December 28: 4. Brown, Lee. 1964. "2 youths tell story of fatal 'bobsled' ride." The Independent (Long Beach, CA), May 22: 17. Daily News. 1983. "A ride to the courthouse." Daily News (New York, NY), July 3: 32. Daily Record. 1982. "Electrical shock killed man on Action Park ride." Daily Record (Morristown, NJ), August 1: 2. Fisher, Joseph. 1980. "Man who fell from alpine slide dies after several days in coma." Daily Record (Morristown, NJ), Juky 17: 1. Futia, Michael, and John Mintz. 1982. "Death doesn't cut lines for thrill rides." The Record (Hackensack, NJ), August 2: 13. Gaura, Maria. 1998. "Coaster victim's death witnessed by family." San Francisco Chronicle, September 11: 13. Gaura, Maria, and Manny Fernandez. 1998. "Victim's kin mull suit against Great America." San Francisco Chronicle, Seoptember 9: 1. Haefele, Marc. 1980. "Dangers cited by slide employees." Daily Record (Morristown, NJ), August 14: 19. Hatfield, Larry. 1980. "Roller coaster crash caused by 'phantom'." San Francisco Examiner, May 1980: 3. Hoover, Ken, and Sabin Russell. 1999. "Fall from ride kills boy at Great America." San Francisco Chronicle, August 23: 1. Kiely, Eugene. 1987. "Prosecutor: Action Park drowning accidental." The Record (Hackensack NJ), July 21: 28. Los Angeles Times. 1964. "Boy criticallt hurt on ride at Disneyland." Los Angeles Times, May 17: 3. —. 1966. "He tried to join his friends." Los Angeles Times, June 19: 3. —. 1964. "Inquest ruled out in fatal Disneyland fall." Los Angeles Times, May 27: 35. Lyman, Julie, Kevin Fagan, and Bill Workman. 1999. "Questions linger in amusement park death." San Francisco Chronicle , November 6: 1. Mulvihill, Andy. 2020. "Remembering Action Park, New Jersey's Deranged Theme Park, "Where You're the Center of the Accident"." Esquire, July 2. Press-Telegram. 1964. "Boy badly hurt in tumble from Disney bobsled." Press-Telegram (Long Beach, CA), May 16: 13. —. 1966. "Monorail victim crashing party?" Press-Telegram (Long Beach, CA), June 19: 4. —. 1964. "Bobsled rider's death probed." Press-Telegram, May 20: 39. Reckard, Scott, and Tracy Weber. 1998. "Autopsy sheds light on Disneyland fatality." Los Angeles Times, December 31: 31. Soiffer, Bill. 1980. "Brakes suspected in coaster tragedy." San Francisco Chronicle, March 31: 3. Stolztfus, Duane. 1984. "Water slide blamed for son's death." Daily Record (Morristown, NJ), August 28: 11. Webber, Tracy. 1999. "Fatal accident at Disneyland in '98 still haunts family." Los Angeles Times, December 13: 110. Yi, Daniel, and Robert Ourlian. 1998. "Man dies 2 days after being injured at Disneyland." Los Angeles Times, December 27: 76. Cowritten by Alaina Urquhart, Ash Kelley & Dave White (Since 10/2022)Produced & Edited by Mikie Sirois (Since 2023)Research by Dave White (Since 10/2022), Alaina Urquhart & Ash KelleyListener Correspondence & Collaboration by Debra LallyListener Tale Video Edited by Aidan McElman (Since 6/2025) Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Senkt vegane Ernährung den Testosteronspiegel – oder ist das nur ein Fitness-Mythos? In dieser Folge sprechen wir über Testosteron, Soja, Cholesterin, Muskelaufbau, Krafttraining, Energiedefizite, Blutwerte und den wachsenden Trend rund um TRT. Wissenschaftlich eingeordnet, praxisnah erklärt und mit Blick darauf, was für vegane Sportlerinnen und Sportler wirklich relevant ist. ------------------------------------------------------------------------ Dominiks Buch zur pflanzenbasierten Sporternährung im UTB-Verlag: https://www.utb.de/doi/book/10.36198/9783838560328 Dominiks Gesundheitscommunity: www.gsundes-hannover.de Dominiks Online-Knie-Kurs: https://gsundes-hannover.de/knieschmerzen/ Dominiks Online-Rücken-Kurs: https://copecart.com/products/34bd5abb/checkout Marcs veganes Online-Fitness-Coaching: https://vegainer-academy.com/ Marcs Online-Kurs: https://www.copecart.com/products/a50f88f2/checkout ------------------------------------------------------------------------ Dieser Podcast wird unterstützt von der Firma Watson Nutrition. Die Firma bietet als einzige umfassend laborgeprüfte Nahrungsergänzungsmittel für eine optimierte Nährstoffversorgung. Zum Angebot zählen Multi-Supplemente, Mono-Supplemente, Sportsupplemente wie Kreatin oder auch Proteinriegel, Shakes und essenzielle Aminosäuren Mit dem Code veganperformance erhältst du 5 % Rabatt auf deine Bestellung. Zur Firmenwebseite: Watson Nutrition ------------------------------------------------------------------------ Quellen: Wissenschaftliche Studien, Reviews und Leitlinien Allen, N. E., Appleby, P. N., Davey, G. K., & Key, T. J. (2000). Hormones and diet: Low insulin-like growth factor-I but normal bioavailable androgens in vegan men. British Journal of Cancer, 83(1), 95–97. Baillargeon, J., Kuo, Y. F., Westra, J. R., Urban, R. J., & Goodwin, J. S. (2018). Testosterone prescribing in the United States, 2002–2016. JAMA, 320(2), 200–202. Bhasin, S., Storer, T. W., Berman, N., Callegari, C., Clevenger, B., Phillips, J., Bunnell, T. J., Tricker, R., Shirazi, A., & Casaburi, R. (1996). The effects of supraphysiologic doses of testosterone on muscle size and strength in normal men. The New England Journal of Medicine, 335(1), 1–7. Bhasin, S., Brito, J. P., Cunningham, G. R., Hayes, F. J., Hodis, H. N., Matsumoto, A. M., Snyder, P. J., Swerdloff, R. S., Wu, F. C., & Yialamas, M. A. (2018). Testosterone therapy in men with hypogonadism: An Endocrine Society clinical practice guideline. The Journal of Clinical Endocrinology & Metabolism, 103(5), 1715–1744. Christou, M. A., Christou, P. A., Markozannes, G., Tsatsoulis, A., Mastorakos, G., & Tigas, S. (2017). Effects of anabolic androgenic steroids on the reproductive system of athletes and recreational users: A systematic review and meta-analysis. Sports Medicine, 47(9), 1869–1883. Cinar, V., Polat, Y., Baltaci, A. K., & Mogulkoc, R. (2011). Effects of magnesium supplementation on testosterone levels of athletes and sedentary subjects at rest and after exhaustion. Biological Trace Element Research, 140(1), 18–23. Corona, G., Rastrelli, G., Monami, M., Saad, F., Luconi, M., Lucchese, M., Facchiano, E., Sforza, A., Forti, G., Mannucci, E., & Maggi, M. (2013). Body weight loss reverts obesity-associated hypogonadotropic hypogonadism: A systematic review and meta-analysis. European Journal of Endocrinology, 168(6), 829–843. Demay, M. B., Pittas, A. G., Bikle, D. D., Diab, D. L., Kiely, M. E., Lazaretti-Castro, M., Lips, P., Mitchell, D. M., Murad, M. H., Powers, S., Rao, S. D., Scragg, R., Tayek, J. A., Valent, A. M., Walsh, J. M. E., & McCartney, C. R. (2024). Vitamin D for the prevention of disease: An Endocrine Society clinical practice guideline. The Journal of Clinical Endocrinology & Metabolism, 109(8), 1907–1947. Dubin, J. M., Jesse, E., Fantus, R. J., Bennett, N. E., Brannigan, R. E., Thirumavalavan, N., & Halpern, J. A. (2022). Guideline-discordant care among direct-to-consumer testosterone therapy platforms. JAMA Internal Medicine, 182(12), 1321–1323. European Association of Urology. (2026). Male hypogonadism. In EAU guidelines on sexual and reproductive health. Guisado-Cuadrado, I., Recacha-Ponce, P., Peinado, A. B., & Romero-Parra, N. (2026). Biochemical responses to experimentally induced short-term low energy availability in athletes: A systematic review. Scandinavian Journal of Medicine & Science in Sports, 36(3), Article e70249. Key, T. J. A., Roe, L., Thorogood, M., Moore, J. W., Clark, G. M. G., & Wang, D. Y. (1990). Testosterone, sex hormone-binding globulin, calculated free testosterone, and oestradiol in male vegans and omnivores. British Journal of Nutrition, 64(1), 111–119. Leproult, R., & Van Cauter, E. (2011). Effect of 1 week of sleep restriction on testosterone levels in young healthy men. JAMA, 305(21), 2173–2174. Lincoff, A. M., Bhasin, S., Flevaris, P., Mitchell, L. M., Basaria, S., Boden, W. E., Cunningham, G. R., Granger, C. B., Khera, M., Thompson, I. M., Wang, Q., Wolski, K., Davey, D., Kalahasti, V., Khan, N., Miller, M. G., Snabes, M. C., Chan, A., Dubcenco, E., Li, X., et al. (2023). Cardiovascular safety of testosterone-replacement therapy. The New England Journal of Medicine, 389(2), 107–117. Messina, M. (2010). Soybean isoflavone exposure does not have feminizing effects on men: A critical examination of the clinical evidence. Fertility and Sterility, 93(7), 2095–2104. Morden, N. E., Woloshin, S., Brooks, C. G., & Schwartz, L. M. (2019). Trends in testosterone prescribing for age-related hypogonadism in men with and without heart disease. JAMA Internal Medicine, 179(3), 446–448. Morton, R. W., Sato, K., Gallaugher, M. P. B., Oikawa, S. Y., McNicholas, P. D., Fujita, S., & Phillips, S. M. (2018). Muscle androgen receptor content but not systemic hormones is associated with resistance training-induced skeletal muscle hypertrophy in healthy, young men. Frontiers in Physiology, 9, Article 1373. Mountjoy, M., Ackerman, K. E., Bailey, D. M., Burke, L. M., Constantini, N., Hackney, A. C., Heikura, I. A., Melin, A., Pensgaard, A. M., Stellingwerff, T., Sundgot-Borgen, J. K., Torstveit, M. K., Jacobsen, A. U., Verhagen, E., Budgett, R., Engebretsen, L., & Erdener, U. (2023). 2023 International Olympic Committee's consensus statement on Relative Energy Deficiency in Sport. British Journal of Sports Medicine, 57(17), 1073–1097. Mulhall, J. P., Trost, L. W., Brannigan, R. E., Kurtz, E. G., Redmon, J. B., Chiles, K. A., Lightner, D. J., Miner, M. M., Murad, M. H., Nelson, C. J., Platz, E. A., Ramanathan, L. V., & Lewis, R. W. (2018). Evaluation and management of testosterone deficiency: AUA guideline. The Journal of Urology, 200(2), 423–432. Prasad, A. S., Mantzoros, C. S., Beck, F. W. J., Hess, J. W., & Brewer, G. J. (1996). Zinc status and serum testosterone levels of healthy adults. Nutrition, 12(5), 344–348. Rao, P. K., Boulet, S. L., Mehta, A., Hotaling, J., Eisenberg, M. L., Honig, S. C., Warner, L., Kissin, D. M., Nangia, A. K., & Ross, L. S. (2017). Trends in testosterone replacement therapy use from 2003 to 2013 among reproductive-age men in the United States. The Journal of Urology, 197(4), 1121–1126. Reed, K. E., Camargo, J., Hamilton-Reeves, J., Kurzer, M., & Messina, M. (2021). Neither soy nor isoflavone intake affects male reproductive hormones: An expanded and updated meta-analysis of clinical studies. Reproductive Toxicology, 100, 60–67. Sagoe, D., Molde, H., Andreassen, C. S., Torsheim, T., & Pallesen, S. (2014). The global epidemiology of anabolic-androgenic steroid use: A meta-analysis and meta-regression analysis. Annals of Epidemiology, 24(5), 383–398. Travison, T. G., Araujo, A. B., O'Donnell, A. B., Kupelian, V., & McKinlay, J. B. (2007). A population-level decline in serum testosterone levels in American men. The Journal of Clinical Endocrinology & Metabolism, 92(1), 196–202. Travison, T. G., Vesper, H. W., Orwoll, E., Wu, F., Kaufman, J. M., Wang, Y., Lapauw, B., Fiers, T., Matsumoto, A. M., & Bhasin, S. (2017). Harmonized reference ranges for circulating testosterone levels in men of four cohort studies in the United States and Europe. The Journal of Clinical Endocrinology & Metabolism, 102(4), 1161–1173. Wankhede, S., Langade, D., Joshi, K., Sinha, S. R., & Bhattacharyya, S. (2015). Examining the effect of Withania somnifera supplementation on muscle strength and recovery: A randomized controlled trial. Journal of the International Society of Sports Nutrition, 12, Article 43. West, D. W. D., & Phillips, S. M. (2012). Associations of exercise-induced hormone profiles and gains in strength and hypertrophy in a large cohort after weight training. European Journal of Applied Physiology, 112(7), 2693–2702. Whittaker, J., & Wu, K. (2021). Low-fat diets and testosterone in men: Systematic review and meta-analysis of intervention studies. The Journal of Steroid Biochemistry and Molecular Biology, 210, Article 105878. Positionspapiere, Behörden und Informationsquellen Deutsche Gesellschaft für Ernährung. (2024). DGE veröffentlicht neues Positionspapier zu veganer Ernährung. Deutsche Gesellschaft für Ernährung. National Institutes of Health, Office of Dietary Supplements. (n.d.). Vitamin B12: Fact sheet for health professionals. Abgerufen am 21. Mai 2026. National Institutes of Health, Office of Dietary Supplements. (n.d.). Vitamin D: Fact sheet for health professionals. Abgerufen am 21. Mai 2026. U.S. Food and Drug Administration. (2025, 28. Februar). FDA issues class-wide labeling changes for testosterone products. U.S. Food and Drug Administration. World Anti-Doping Agency. (2026). The 2026 prohibited list. World Anti-Doping Agency. ‘They've invented a spurious pseudo-disease': Why are so many men being told they have low testosterone? (2026, 10. Mai). The Guardian.
We pick up where we left off with Kiely, Jenna, and their sperm donor Dominick. This time it's Jenna's turn to get pregnant. And, unlike last time, things do not go according to plan. ⭐️ This episode originally ran on November 30, 2016 and is a favorite from the archives. We hope you enjoy, and we'll be back next week with a brand new episode. … • Join LST+ for community and access to You Know What, another show in the Longest Shortest universe! • Follow us on Instagram • Sign up for our newsletter, where we recommend other parenting + reproductive health media • Buy books by LST guests (your purchase supports the show!) • Website: longestshortesttime.com Learn more about your ad choices. Visit megaphone.fm/adchoices
If you aren't the one educating your users on the fundamentals of AI, your competitors will happily do it for you. This week on Dev Interrupted, Andrew sits down with Philip Kiely, Head of AI Education at Baseten and author of Inference Engineering, to discuss why the secret to winning the AI market is owning the educational narrative through active market development. They explore the rise of the "Double-T" shaped engineer, the hidden complexities of scaling the inference stack, and why the most successful AI companies treat developer education as a mission-critical go-to-market motion.Read the guide: The APEX FrameworkFollow the show:Subscribe to our Substack Follow us on LinkedInSubscribe to our YouTube ChannelLeave us a ReviewFollow the hosts:Follow AndrewFollow BenFollow DanFollow today's stories:Baseten: Explore the inference platform where Philip serves as Head of AI Education.Inference Engineering: Download the free PDF or order a paper copy of Philip's comprehensive guide to the AI infrastructure stack.LinkedIn: Philip Kiely X/Twitter: @philipkiely Website: philipkiely.com OFFERSStart Free Trial: Get started with LinearB's AI productivity platform for free.Book a Demo: Learn how you can ship faster, improve DevEx, and lead with confidence in the AI era.LEARN ABOUT LINEARBAI Code Reviews: Automate reviews to catch bugs, security risks, and performance issues before they hit production.AI & Productivity Insights: Go beyond DORA with AI-powered recommendations and dashboards to measure and improve performance.AI-Powered Workflow Automations: Use AI-generated PR descriptions, smart routing, and other automations to reduce developer toil.MCP Server: Interact with your engineering data using natural language to build custom reports and get answers on the fly.
This Week in Machine Learning & Artificial Intelligence (AI) Podcast
In this episode, Philip Kiely, head of AI education at Baseten, joins us to unpack the fast-evolving discipline of inference engineering. We explore why inference has become the stickiest and most critical workload in AI, how it blends GPU programming, applied research, and large-scale distributed systems, and where the line sits between inference and model serving. Philip shares how research-to-production can move in hours, not months, and why understanding “the knobs” of inference—batching, quantization, speculation, and KV cache reuse—lets teams design better products and SLAs. We trace the inference maturity journey from closed APIs to dedicated deployments and in-house platforms, discuss GPU lifecycles, and survey today's runtime landscape, including vLLM, SGLang, and TensorRT LLM. Finally, we look ahead to agents and multimodality, making the case for specialized, workload-specific runtimes when performance and efficiency matter most. The complete show notes for this episode can be found at https://twimlai.com/go/766.
Henry Shefflin talks to us about his first win as under 20 manager while Taggy and Ronnie analyse their win in Kildare.Plus, John Kiely and Ben O'Connor react to the League Final.The KCLR Hurling Podcast brought to you by Morrissey Motors Peugeot Kilkenny.
Interview by Kris PetersPatient Sixty-Seven have become one of the most compelling and community‑driven voices in modern metalcore - a band built not on hype or industry shortcuts, but on heart, resilience, and the belief that heavy music can be a lifeline. Emerging from the isolated but fiercely creative city of Perth, Australia, P67 have spent the past decade turning personal struggle into connection, and connection into a movement that now has reached metalcore fans across the globe.From their earliest releases, Patient Sixty-Seven stood out for their emotional honesty - songs that didn't shy away from fear, grief, or self‑doubt, but instead embraced them as part of the human experience. That vulnerability resonated deeply, helping the band build a loyal, engaged, and heartfelt community long before the industry took notice. Fans didn't just listen; they shared stories, found comfort in the lyrics, and formed bonds with each other that extended far beyond the music.Now, Patient Sixty-Seven are stepping into their most significant chapter yet. In May this year, the band will join Of Mice & Men and Crystal Lake on a major Australian tour - a career‑defining moment that places them alongside some of the most influential names in modern heavy music. It's a testament to how far they've come, and a signal of where they're headed next.HEAVY caught up with vocalist Tom Kiely to find out more. One of the topics of discussion is how the band approaches major International supports. Do they go out there to warm the crowd up and play a role, or do they attack it with more vigour and go out there with a view to blowing everyone else off stage?"I think for us, we just want to be ourselves," Tom measured. "I think obviously we want to make sure that we bring a high level of energy and intensity, because we know that ultimately our role on the tour is to get the crowd warmed up; to get the crowd moving; to get the crowd excited for the bands that are coming after us. By doing that it leans nicely into what we like to do anyway, which is play with a lot of energy and get the crowd involved. We try to be interactive and try and bring that spark to the stage and after our set's finished, hopefully people are even more excited for the next few bands.Opening is always tricky because you know the crowd's definitely still getting warmed up and maybe not moving as much, so it's our job to shake off any cobwebs people have if they haven't been to a show in a while. We do what we can to get people banging their heads and maybe getting a mosh pit going. We find that a lot of the times once you start talking to the crowd and interacting with them there's a lot of people who are ready to get moving. If we can get a few mosh pits going, that'll be a highlight for sure (laughs)."In the full interview, Tom talked more about the run of shows with Of Mice & Men and Crystal Lake, where they fit in with the line-up, what to expect from their live show and what three songs concert goers can listen to in order to get to know the band before the shows.He also spoke about curating a set list to appeal to fans of the headliners while also playing their strongest material, how far advanced work is on their new album, what direction it is going to take musically and more.Of Mice & Men 2026 Australian Tour Dates With Crystal LakeTuesday 5th May – PERTH, Magnet HouseThursday 7th May – ADELAIDE, Lion Arts FactoryFriday 8th May – MELBOURNE, 170 RussellSaturday 9th May – SYDNEY, Manning BarSunday 10th May – BRISBANE, The TriffidTickets https://thephoenix.au/of-mice-and-men/Become a supporter of this podcast: https://www.spreaker.com/podcast/heavy-music-interviews--2687660/support.
Anya chats with Lead Culinary Instructor Ben Kiely of Pacific Institute of Culinary Arts (PICA) about the next generation of kitchen professionals, the current realities of the industry, how teaching culinary arts has evolved over the years, kitchen culture and the ongoing changes, PICA's programs and what current students are looking to do once they graduate, and so much more.
Patrick McKenzie (patio11) and Philip Kiely, early employee at Baseten, discuss the inference stack: the critical layer of software and hardware that sits between a model's weights and a user's prompt. They cover inference engineering, how intermediate layers are evolving over a technical stack that is changing every six months, and how sophisticated organizations are actually consuming LLMs beyond just writing their questions into chatbot apps.–Full transcript available here: www.complexsystemspodcast.com/inference-engineering-with-philip-kiely/–Presenting Sponsors: Mercury, Meter, & GranolaComplex Systems is presented by Mercury—radically better banking for founders. Mercury offers the best wire experience anywhere: fast, reliable, and free for domestic U.S. wires, so you can stay focused on growing your business. Apply online in minutes at mercury.com.Networking infrastructure has a way of accumulating technical debt faster than almost anything else in IT. Meter handles the full stack (wired, wireless, and cellular) as a single integrated solution: designed, deployed, and managed end-to-end so there's only one vendor to call when something goes wrong. Visit meter.com/complexsystems to book a demo. If meetings consistently leave you with hazy action items and lost context, Granola handles the transcription so you can actually participate and gives you searchable notes afterward. Try it free at granola.ai/complexsystems with code COMPLEXSYSTEMS–Links:Download Inference Engineering: https://www.baseten.com/inference-engineering/ Philip's website: https://philipkiely.com/ Stripe's Emily Sands on Complex Systems: https://www.complexsystemspodcast.com/episodes/the-past-present-and-future-of-ai-with-stripe/ Des Traynor on Complex Systems: https://www.complexsystemspodcast.com/episodes/des-traynor/ –Timestamps:(00:00) Intro(00:30) The AI deployment pipeline(03:04) Evolution of abstraction layers in engineering(05:14) Defining inference and model weights(08:45) Architecture of language and diffusion models(10:11) AI adoption in the broader economy(11:30) The shift toward agentic workflows and RL(14:55) Function calling and real-world actions(20:10) Sponsors: Mercury | Meter(22:59) Technologies for agentic tools: MCP and skills(25:32) The craft of writing a harness(29:56) Using AI for automated proofreading and tool creation(34:12) Balancing LLMs with deterministic code(37:31) Observability and chain of thought reasoning(39:31) Sponsor: Granola(41:21) Observability and chain of thought reasoning(50:45) Speculative decoding and hidden states(55:37) The value of smaller, task-specific models(59:55) Internal competencies versus buying solutions(01:09:27) Self-publishing a technical book in record time(01:23:20) Wrap
This week on the show, Scott talks to Philip Kiley about his new book, Inference Engineering. Inference Engineering is your guide to becoming an expert in inference. It contains everything that Philip has learned in four years of working at Baseten. This book is based on the hundreds of thousands of words of documentation, blogs, and talks he's written on inference; interviews with dozens of experts from our engineering team; and countless conversations with customers and builders around the world. https://www.baseten.co/inference-engineering/
What does cybersecurity really mean for today's CPA firms? In this episode, we sit down with Luke Kiely, Chief Information Security Officer at SmartVault and Chief Security Officer at ComplyWise, to explore why cybersecurity is no longer just an IT issue, but a firm-wide responsibility.Luke breaks down how most breaches still begin with a simple email and a distracted click, why busy season increases vulnerability, and the practical safeguards firms can put in place without a massive IT budget.This episode offers clear, actionable insight into protecting client data and securing the future of your firm.Resources:Luke Kiely LinkedIn ProfileSmartVaultComplyWiseFTC Safeguards Rule OverviewIRS Publication 4557 – Safeguarding Taxpayer Data
Software Engineering Radio - The Podcast for Professional Software Developers
Philip Kiely, software developer relations lead at Baseten, speaks with host Jeff Doolittle about multi-agent AI, emphasizing how to build AI-native software beyond simple ChatGPT wrappers. Kiely advocates for composing multiple models and agents that take action to achieve complex user goals, rather than just producing information. He explains the transition from off-the-shelf models to custom solutions, driven by needs for domain-specific quality, latency improvements, and economic sustainability, which introduces the engineering challenge of inference engineering. Kiely stresses that AI engineering is primarily software engineering with new challenges, requiring robust observability and careful consideration of trust and safety through evals and alignment. He recommends an approach of iterative experimentation to get started with multi-agent AI systems. Brought to you by IEEE Computer Society and IEEE Software magazine.
Gemma, Natee, and Marc are back for another attempt at entertaining prehistorically inclined people with carefully edited commentary and interviews with people who actually know what they're talking about. In this episode, David Armsby's back and better than ever, Darren Naish's hair is slicker than ever, and Gemma interviews palaeobotanist and artist Julianne Kiely, who's here to save us all from painfully generic and/or inaccurate flora in palaeoart. Will Gemma dare challenge Julianne's assertion that plants are cooler than animals? How handsome is a cockroach? Can we still be confident that an illustration of a dinosaur was AI-generated? Find out...by listening. Show Notes At Chasmosaurs.com!
On this week's episode we have long time chef and educator Ben Kiely. Head instructor at the Pacific Institute of Culinary Arts here in Vancouver. His many years of cooking in England and across Europe led him to Greece where he found himself falling for a Canadian woman and he followed her back to Vancouver and they started a family. A story as old as time. Especially for a cook. I hope you enjoy Ben's insights on this business and industry at large. I really enjoyed chatting with him. Send us your feedback
In this episode of The Winston Marshall Show, I sit down with Father Benedict Kiely, founder of Nasarean.org and one of the world's leading voices on the persecution of Christians.We discuss the genocide unfolding in Nigeria, where thousands of Christians are being murdered each year by Islamist militias while Western governments and media look away. Father Kiely exposes how the massacre of Christians is dismissed as “climate change” violence, a lie repeated by politicians and journalists unwilling to name Islamist extremism.From ISIS's resurgence across Africa and the Middle East to the silent persecution of Christians in Iraq under Iranian-backed militias, Kiely lays bare a pattern of denial stretching from Abuja to Washington. He explains how Christianity faces extinction in its ancient homelands, the failure of the West's moral leadership, and why Europe's collapse of faith has left it powerless to confront evil.We also explore stories of hope, the revival of Christianity in post-communist Albania, the endurance of believers speaking Aramaic in Iraq, and why, despite centuries of persecution, faith refuses to die.All this: Nigeria's genocide, Islamist expansion, the silence of the West, and the rebirth of Christianity where it was once crushed.-----------------------------------------------------------------------------------------------------------------------To see more exclusive content and interviews consider subscribing to my substack here: https://www.winstonmarshall.co.uk/-----------------------------------------------------------------------------------------------------------------------FOLLOW ME ON SOCIAL MEDIA:Substack: https://www.winstonmarshall.co.uk/X: https://twitter.com/mrwinmarshallInsta: https://www.instagram.com/winstonmarshallLinktree: https://linktr.ee/winstonmarshall----------------------------------------------------------------------------------------------------------------------Chapters 00:00 Introduction 01:18 Father Benedict Keeley's Role and Media Coverage02:54 Genocide in Nigeria and Media Bias08:41 Islamist Extremism Across Africa26:42 Situation in Syria and the Role of Al-Qaeda33:04 Christian Persecution in Sudan 40:08 Christian Persecution in Europe and the Role of the Media1:04:06 Challenges Faced by Christians in the UK1:11:59 Final Thoughts Hosted on Acast. See acast.com/privacy for more information.
#surrogacy #ivf #surrogate Grace's Instagram: https://www.instagram.com/graces_surro_journey?igsh=MzRlODBiNWFlZA== Kiely's Instagram: https://www.instagram.com/thatsurrogatelife?igsh=NTc4MTIwNjQ2YQ==What happens when the transfers are done, the last delivery is behind you, and the identity you wore with pride suddenly shifts? We sit down with a seasoned pair of voices—Grace, a two‑time carrier, and Kiely, a four‑time carrier—to talk candidly about “retirement” from surrogacy, the choice to step back versus being told no by clinics, and the surprising ways purpose expands after the final journey.The conversation moves from age cutoffs and ACOG guidance on C‑sections to the emotional calculus of ending on your own terms. Grace shares how preparing for her last pregnancy shaped a peaceful exit, while Kiely explains why she wanted the decision to be hers, then channeled that energy into writing, mentoring, and creating Send a Friend, a grassroots program that brings an experienced surrogate to support first‑time postpartum surrogates. Along the way, we reflect on what surrogacy teaches our families: that love builds families in many forms; that children can learn to answer strangers with clarity and pride; and that empathy deepens when you stand next to someone who longs for a child money can't buy.We also get practical. Expect frank talk about changing insurance rules, clinic discretion, compensation norms, and how to research without falling into bias. You'll hear why Facebook groups can be both a lifeline and a minefield, and how to gather perspective that sets realistic expectations for matching, protocols, and postpartum recovery. Most of all, we reframe the label “retired.” You're not done—you're a surrogate emerita, carrying wisdom forward through advocacy, education, and community.If this conversation resonates, follow the show, share it with someone curious about surrogacy, and leave a review with the insight you wish every new surrogate knew. Your voice helps more families—and more surrogates—find their path.My Mom Is Brave- https://a.co/d/bENg23CSend a friend- https://thatsurrogatelife.com/send-a-friend?fbclid=IwRlRTSAN6_4JleHRuA2FlbQIxMABzcnRjBmFwcF9pZAo2NjI4NTY4Mzc5AAEeWkNyXW2MxUwkTjDVHsBzCcmsoautbyibsqDBUbhCpFOcLozud1tg5LweaWw_aem_XCw185aF7SEW5njM85Y4MASend us a texthttps://stopsitsurrogate.com
In late 2018, the Connecticut Environmental Conservation Police uncovered a chilling case involving a group of young trophy hunters. Over just a few months, they had illegally taken at least 19 deer - often during nighttime hunts near residential neighborhoods. What started as a routine investigation quickly unraveled into something far more disturbing: secret planning sessions, a manifesto detailing their exploits, and a twisted tribute to the grandfather who taught them to night hunt. Join Investigator Patrick Kiely as he recounts the unbelievable story of the “Killing Krew Klan.” Our Sponsors: Thin Green Line Podcast Don Noyes Chevrolet North American Game Warden Museum Hunt Regs WiseEye SecureIt Gun Storage XS Sights “A Cowboy in the Woods” Book Maine's Operation Game Thief International Wildlife Crimestoppers Here's what we discuss: · An area known for night hunting · Spotting night hunters requires patience and timing · The state's healthy deer population is tempting for poachers · A patrol officer spots suspicious signs · The initial arrest leads to more questions · Cell phones: everyone documents everything · “I wouldn't even call them hunters; they were trophy poachers.” · It definitely wasn't squirrels · The group is released but phones are seized · A stunning discovery · “It was an every-night occurrence.” · The group frequently hunted near houses · None of 19 deer were registered · The puzzle pieces: pictures, locations and times · The serial poaching had gone on for years, and had grown · Group relied on thinly stretched law enforcement · A specific 16-point buck and an unlikely story · US Fish and Wildlife joins the investigation · Cell phone metadata pinpoints locations and times · “Not a care in the world.” · A handwritten manifesto is found · The ‘zombie' deer · Timing was perfect – and lucky · Even illegal roadkill wasn't off limits to the ‘Klan' · $100 does for sale, and banquet hall venison · Multiple deer were taken nightly · Managing investigations and public perception · Hunters had noticed a decline · “It was a joke to them.” · Many state charges were misdemeanors · Local hunters weigh in · Technology has changed investigation strategies · Limitation statutes prevented even more charges · Getting buy-in from other agencies · Balancing criminal and wildlife investigations can be a challenge · The cell phones were crucial · Rising bear population has led to conflicts · Educating the public · Staffing numbers are on the rise · “It was a learning experience for all of us.” Credits Hosts: Wayne Saunders and John Nores Producer: Jay Ammann Warden's Watch logo & Design: Ashley Hannett Research / Content Coordinator: Stacey DesRoches Subscribe: Apple Podcasts Spotify Amazon Google Waypoint Stitcher TuneIn Megaphone Find More Here: Website Warden's Watch / TGL Store Facebook Facebook Fan Page Instagram Threads YouTube RSS Learn more about your ad choices. Visit megaphone.fm/adchoices
Hiring, leading, and keeping your team focused is one of the toughest parts of growing a landscaping company. Add in the unique challenges of ADHD and neurodivergence, and it can feel overwhelming. But what if those differences could actually be a strength in your company?In this episode of The Landscaper's Guide, Jack Jostes interviews LeanScaper Chief Customer Experience Officer Kristen Kiely. Kristen has nearly a decade of industry experience with LMN, The Grounds Guys, and now LeanScaper, and her husband runs a $5M design-build and snow company in Toronto. She shares how neurodivergence, structure, and operational systems can actually fuel growth for landscaping companies.You'll discover:Why ADHD and neurodivergence can be a superpower in the landscape industryHow tools like calendaring, nutrition, and routines improve focus and leadershipWhy efficient operations and financial discipline drive profitability and scalabilityWhether you're managing a $2M company or pushing past $10M, this episode will help you think differently about your team and your growth.Show Notes:Watch the full episode + see the transcript: https://landscapersguide.com/podcast/ Get your free beef jerky sample: https://landscapersguide.com/toolbox See upcoming live and virtual events: https://landscapersguide.com/eventsConnect with Kristen: Instagram: https://instagram.com/kristencxo LinkedIn: https://www.linkedin.com/in/kristenkielycfe/ LeanScaper: https://leanscaper.com
Tricia Kiely is a feminist comedian from Boston living in Southern NH. They call her the pun princess for her word-nerd stylings and quick wit.In addition to being silly royalty she performs and teaches improv with Nashua's Strugglebus Improv, hosts joke writing workshops and produces a monthly variety show in Nashua. She believes everyone has a funny bone and would love to help you find yours (respectfully of course)Catch her quick wit and silly quips all over the Merrimack Valley and beyond.
Noel catches up with Mark Kiely. The actor is probably best known for his role as Gil Meyers on Beverly Hills, 90210. Mark also had a recurring role on 24. His other TV roles include NYPD Blue, Lois and Clark, CSI, The Shield and more. His movies roles include Bruce Almighty, The Edge and The Judge. Mark took a break from acting to become a competitive swim coach in Rhode Island.
The Catholic cardinal Jorge Mario Bergolio ascended to the papacy in 2013. In honor of Saint Francis of Assisi, he chose as his papal name Francis. For a dozen years he was the head of the Catholic Church and a major figure in the moral and cultural life of the West. After a prolonged illness, Pope Francis died on April 21 of this year. There are over 1.4 billion Catholics in the world, and they play a significant role in the production of Western culture and Western opinion. The foundational structures of Europe are derivative of, or inseparably woven into, the history of the Catholic Church. And whether the pope strengthens or undermines the moral confidence of Western nations matters: it mattered during the papacy of John Paul II during the cold war; it mattered in the confrontation with jihadist terror during the papacy of Benedict XVI; and it cannot but be a factor in the horizons of Western civilization. This podcast focuses on a particular dimension of the late Pope Francis's legacy, namely, how he engaged the Jewish people, Israel, and the Middle East. To discuss the legacy of Pope Francis, the Church's engagement in the Middle East, and who might be the next Catholic pope, Mosaic's editor Jonathan Silver sat down with Father Benedict Kiely. Kiely was born in London, ordained a Catholic priest in Canterbury, and has spent most of his ministry in the United States. In 2014, he founded Nasarean.org, a charity that supports persecuted Christians around the world, and especially in Iraq, Syria, and Lebanon. One of his aims is to see the church grow closer to its Middle Eastern roots, and that means, in some grand spiritual way, closer too to its Jewish roots. For Catholics, the question of the Church's attitude toward Zionism and Israel is not perhaps among the most pressing of ecclesiastical priorities. One would not expect it to weigh heavily on the Vatican's conclave in the election of the next pope. This conversation thus takes the perspective of an outsider. Moreover, there are very deep theological matters that will always divide the Catholic Church from the Jewish people. And some of those very deep theological matters also shape the way that Catholics tend to think about Zionism and the modern state of Israel. The Jewish people are animated by a belief in covenantal chosenness, and a sense of sacred obligation to uphold God's ways in their actions, in their families, and in their nation. That obligation is structured by tradition and law, and it is expressed nationally in the people of Israel, which, after a long hiatus in exile, again has a sovereign state in the land of its fathers. For Catholics, of course, the Church is the new Israel, and despite very welcome and laudable developments since the promulgation of Nostra Aetate in 1965, that is an unbridgeable theological chasm. Nonetheless, friendship between Christians and Jews is essential to revitalizing our shared civilization and passing it on to future generations. Musical selections in this podcast are drawn from the Quintet for Clarinet and Strings, op. 31a, composed by Paul Ben-Haim and performed by the ARC Ensemble.
Reach Out Via Text!In this episode of the Growing Green Podcast, host Jeremiah Jennings sits down with Kristen Kiely, Chief Customer Experience Officer at Leanscaper, to dive deep into building scalable systems for your landscaping business. Kristen shares lessons from her time at LMN and The Grounds Guys, and how Leanscaper is empowering business owners—no matter the size—to implement proven frameworks that drive serious results. From step-by-step SOPs to accessing expert advisors, this conversation unpacks what it takes to grow with clarity and confidence. Whether you're just getting started or you're a $10M+ company ready to fine-tune your operations, Kristen offers a fresh perspective on why intentional systems and a growth-minded community are game-changers. Get ready to learn, laugh, and leave inspired to take that next leap forward with Growing Green Landscapes leading the way.Support the show 10% off LMN Software- https://lmncompany.partnerlinks.io/growinggreenpodcast Signup for our Newsletter- https://mailchi.mp/942ae158aff5/newsletter-signup Book A Consult Call-https://stan.store/GrowingGreenPodcast Lawntrepreneur Academy-https://www.lawntrepreneuracademy.com/ The Landscaping Bookkeeper-https://thelandscapingbookkeeper.com/ Instagram- https://www.instagram.com/growinggreenlandscapes/ Email-ggreenlandscapes@gmail.com Growing Green Website- https://www.growinggreenlandscapes.com/
Have you tried turning the EPA off and on again? Register to vote (or check to make sure you're registered): https://www.climatechangemakers.org/quick-register-to-vote Check out the Evergreen Action Plan 2.0: https://evergreenaction.com/initiatives/a-bold-climate-plan-evergreen-action-plan-2And if you want to get more involved, Climate Changemakers also lined up this Vote Forward page, where you can write letters and send them to potentially critical prospective voters in swing states. You can write about whatever you want, but if you're reading this, you might want to focus on the exciting and actually very good climate action that could come from the next climate-focused US President: https://votefwd.org/climatechangemakers BONUS EPISODES available on Patreon (https://www.patreon.com/deniersplaybook) SOCIALS & MORE (https://linktr.ee/deniersplaybook) [For sponsorship inquiries, please contact climatetown@no-logo.co]DISCLAIMER: Some media clips have been edited for length and clarity.CREDITS Created by: Rollie Williams, Nicole Conlan & Ben BoultHosts: Rollie Williams & Nicole ConlanExecutive Producer: Ben Boult Post-production: Jubilaria Media Producers: Irene Plagianos, Miranda Manganaro, Daniella Philipson Researchers: Carly Rizzuto, Canute Haroldson & James Crugnale Art: Jordan Doll Music: Tony Domenick Special thanks: The Civil Liberties Defense CenterIn partnership with: Evergreen ActionSOURCESProject 2025 | Presidential Transition Project. (2024). The Heritage Foundation.Dans, P., & Groves, S. (Eds.). (2023). Mandate for Leadership: The Conservative Promise. The Heritage Foundation.The Heritage Foundation. (2023, April 28). Project 2025: Staffing the Next Conservative Administration | #Heritage50 [Video]. YouTube.ProPublica. (2024, August 10). Project 2025 Private Training Video: Left-Wing Code Words and Language [Video]. YouTube.MSNBC. (2024, June 22). Top Project 2025 architect talks conservative blueprint for Trump second term [Video]. YouTube.CNN. (2024, July 4). Pro-Trump think tank leader makes ominous threat about ‘second American Revolution' [Video]. YouTube.The Kevin Roberts Show. (2024, September 24). New York Times Climate Forward | Dr. Kevin Roberts [Video]. YouTube.Heritage Response Room. (2018, January 25). Tommy Binion discusses how Trump has embraced 64% of Heritage policy recommendations on Fox Business [Video]. YouTube.Washington Post. (2017, October 18). Watch Trump's full speech to the Heritage Foundation [Video]. YouTube.CNN. (2024, July 11). Evidence shatters Trump's claims about his ties to Project 2025 [Video]. YouTube.Fox News. (2024, July 25). Trump dispels myths on Project 2025: 'I have nothing to do' with it [Video]. YouTube.Trump, D. J. [@realDonaldTrump]. (2018, February 28). The Heritage Foundation has just stated that 64% of the Trump Agenda is already done, faster than even Ronald Reagan [Tweet]. X.Trump, D. J. [@realDonaldTrump]. (2024, July 5). I know nothing about Project 2025. I have no idea who is behind it. Truth Social.Trump, D. J. [@realDonaldTrump]. (2024, July 11). I know nothing about Project 2025. I have not seen it, have no idea who is in charge of it. Truth Social.Wiles, S., & LaCivita, C. (2024, July 30). Trump Campaign Statement on Project 2025's Demise. Donald J Trump for President.The Heritage Foundation. (2024). About Heritage. The Heritage Foundation.Blasko, A. (2004, June 7). REAGAN AND HERITAGE: A Unique Partnership. The Heritage Foundation.Malcolm, J., Slattery, E., & Bates, T. (2017, February 1). A Closer Look at Neil Gorsuch, an Excellent Choice for the Supreme Court. The Heritage Foundation.Heritage Expert Helps Shape Supreme Court Nominee List. (2016, September 14). The Heritage Foundation.Heritage Analysis of Trump Administration's First Year Draws High-Profile Attention. (2018, February 28). The Heritage Foundation.Kavanaugh Was Included on the List The Heritage Foundation Helped Compile. (2018, August 31). The Heritage Foundation.Supreme Court Nominee Brett Kavanaugh Was Included on the List The Heritage Foundation Helped Compile. (2018, August 31). The Heritage Foundation.Heritage Pulls Out All Stops for Amy Coney Barrett's Confirmation. (2020, October 27). The Heritage Foundation.McMurry, Evan. (2015, July 19). Fox Panel Dines Out on Trump's Comments: ‘Despicable,' ‘Clown.' Mediaite.NowThis Impact. (2024, July 1). BET Awards Host Taraji P. Henson: 'Project 2025 Is Not a Game' [Video]. YouTube.Crowley, Kinsey. (2024, July 11). How did 'Project 2025' talk erupt? BET Awards host Taraji P. Henson's comments offer clues. USA Today.Project 2025 - Explore. (2024). Google Trends.Noor, Dharna. (2023, July 31). Inside the Republican Plot to Dismantle US Environmental Policy. Mother Jones. Phillips-Fein, Kim. (2024, June 4). The Mandate for Leadership, Then and Now. The Nation.Gertz, Matt. (2024, July 8). Donald Trump on Heritage's Kevin Roberts, who oversees Project 2025: “He's going to be so incredible.” Media Matters.Westervelt, Amy. (2024, July 22). Newsletter: Everything You Need to Know About Project 2025's Plan for the EPA. Drilled.MacGillis, Alec. (2024, August 1). The Man Behind Project 2025's Most Radical Plans. ProPublica.Kroll, A., & Surgery, N. (2024, August 10). Inside Project 2025's Secret Training Videos. ProPublica.Costello, T., & Lawrence, C. (2024, August 15). Undercover in Project 2025. Centre for Climate Reporting.Olmsted, Edith. (2024, September 25). Ex-Project 2025 Leader Brags Trump's Policy Mirrors Theirs. The New Republic. Anthony, Jason. (2024, July 18).Project 2025 in the Real World. Field Guide to the Anthropocene.Guides: Public Policy Research Think Tanks 2019: Top Think Tanks - Worldwide (US and non-US). (2019). Penn Libraries; University of Pennsylvania.Ball, Molly. (2013, September 25). The Fall of the Heritage Foundation and the Death of Republican Ideas. The Atlantic.Mahler, Jonathan. (2018, June 20). How One Conservative Think Tank Is Stocking Trump's Government. The New York Times. Waldman, Scott. (2023, July 28). Conservatives have already written a climate plan for Trump's second term. Politico.Waldman, Scott. (2023, September 26). Conservatives have already written a climate plan for Trump's second term. E&E News.Associated Press. (2024, July 7). Leader of the pro-Trump Project 2025 suggests there will be a new American Revolution. Politico.Contorno, Steve. (2024, July 11). Trump claims not to know who is behind Project 2025. A CNN review found at least 140 people who worked for him are involved. CNN.Tufts, Sierra. (2024, July 11). ‘I know nothing about Project 2025'; Trump posts on his social media site. Wane.Beckwith, Ryan Teague. (2024, July 12). Project 2025's plan to criminalize porn has a sinister subplot. MSNBC.Hale, Z., & Tiernan, T. (2024, July 29). US ELECTIONS: Project 2025 blueprint envisions major rollbacks on US energy, climate policy. S&P Global.Ordoñez, Franco. (2024, July 30). Project 2025's director steps down, but the think tank says work will go on. NPR.Smith, M., & Swenson, A. (2024, July 30). Vance praises a key leader behind Project 2025, a conservative effort Trump has disavowed. AP.Asiedu, Kwasi Gyamfi. (2024, August 14). J.D. Vance wrote the foreword for Project 2025's Kevin Roberts' upcoming book. PolitiFact.Devine, C., Tolan, C., Ash, A., & Lah, K. (2024, August 15). Hidden-camera video shows Project 2025 co-author discussing his secret work preparing for a second Trump term. CNN.Durkee, Alison. (2024, August 15). What We Know About Trump's Link To Project 2025—As Author Claims Ex-President ‘Blessed It' In Secret Recording. Forbes.Kelly, John. (2024, August 22). Hundreds of proposals in Project 2025 match Trump's policies. CBS News.Kiely, E., Gore, D., & Farley, R. (2024, September 10). A Guide to Project 2025. FactCheck.org.The War on Cars. (2024, September 17). Project 2025 and the Stakes for Transportation [Video]. YouTube.NowThis Impact. (2024, July 8). Project 2025 Would Be Terrible for the Climate [Video]. YouTube.The Wall Street Journal. (2024, October 3). Why the Presidential Race Is Fixating on Project 2025 Now | WSJ [Video]. YouTube.CNN. (2024, Aug 21). Kenan Thompson tells friends about Project 2025 in DNC skit [Video]. YouTube.EPA Press Office. (2020, March 17). Mandy Gunasekara Sworn in as EPA Chief of Staff. US EPA.C-SPAN. (2015, February 26). Sen. James Inhofe (R-OK) Snowball in the Senate (C-SPAN) [Video]. YouTube.DeSmog. (2024). Energy 45 Fund. Energy 45. (2020, February 19). Wayback Machine - Internet Archive.Wikipedia Contributors. (2019). Mexico City policy. Wikipedia; Wikimedia Foundation.Project 2025. (2024). Kamala Harris for President: Official Campaign Website.The Second Half Of The Decisive Decade: Potential U.S. Pathways On Climate, Jobs, And Health. (2024, August 12). Energy Innovation: Policy and Technology.The Next President Needs a Bold Climate Roadmap: Meet the Evergreen Action Plan 2.0. (2024). Evergreen Action.We Wrote the Climate Playbook in 2020. Biden Has Made Significant Progress–And There's More Opportunity Ahead. (2024, July 11). Evergreen Action.See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
In an eye-opening episode, Michael Knowles sits down with Fr. Kiely to shed light on a pressing issue often overlooked by mainstream media: 'The Hidden War On Christians Around the World.' This powerful interview delves into the harrowing stories of persecution that millions of Christians face globally, exploring the complexities and the resilience of faith under fire. Fr. Kiely, a dedicated advocate for persecuted Christians, brings to the forefront the struggles and injustices faced by believers in various corners of the world. From the Middle East to Africa, from Asia to Latin America, this conversation uncovers the trials and tribulations of those who endure oppression for their faith.
Join Michael Knowles in a profoundly moving and eye-opening episode titled 'Persecuted Christians and the Church,' featuring his special guest, Fr. Kiely. In this important discussion, they delve into the rarely discussed but critical issue of Christian persecution around the world, shedding light on the challenges and threats faced by believers and the global church. Fr. Kiely brings to the table his deep insights and firsthand experiences, providing an in-depth look into the lives of Christians who live under constant threat for their faith. Together, they explore the historical context, current situations, and what the future holds for religious freedom globally.