POPULARITY
Watch the full episode on YouTube:We first covered Baseten last year when DeepSeek mania was at peak hype. Now they have raised a monster $13B round and become one of the new cohort of AI Infra decacorns that are (with Nvidia, Intel, and the semis complex) chief beneficiaries of the Inference Inflection. We return to Baseten at the peak of the 2026 edition of Open Weights debate. Ali has published a viral breakdown of Kimi K3:And since you last saw him, Philip has spoken at AI Engineer and written the definitive book on Inference Engineering spotted all over SF:Three years ago, inference engineering barely existed as a category.Today, it is one of the most critical disciplines in AI. Inference engineering inherently tackles a different question than standard model training: “How do you turn those weights from training into a product that is fast, reliable, and affordable at scale?” Focusing on these creates an entirely new optimization problem.In one recent GLM-5.2 experiment, quantizing more of the model actually preserved its benchmark quality while increasing throughput by 20%, because the errors introduced in different layers could cancel each other out.Inference is no longer just the final step after training. It is becoming its own engineering discipline, with its own research problems, infrastructure, and increasingly specialized roles.In this episode, Baseten's Philip Kiely and Ali Taha join swyx and Vibhu to explain what actually happens after a new open model is released and what it takes to turn “we generated a token” into a fast, reliable, production-ready API.We go deep on cache-aware routing, disaggregated prefill and decode, quantization, speculative decoding, KV-cache movement, model parallelism, GPU kernels, and the race to make frontier models up to 10× faster. Philip and Ali explain why inference optimizations can still produce gains of 20%, 100%, or even 200%; how quantization errors can cancel one another out; why identical weights can behave differently across clusters; and how Baseten grafted a Kimi vision encoder onto GLM-5.2 without changing the underlying language model.The conversation then expands beyond LLMs into NVIDIA Dynamo, mega kernels, Rubin, AI-specific chips, local inference, video generation, diffusion versus autoregressive models, and the enormous compute barrier to generating coherent long-form video. Finally, we explore the convergence of training and inference, continual learning through persistent KV cache, and the emerging loop where models help optimize the infrastructure that runs them.We discuss:* What happens when a 200,000-token request enters an inference system* Cache-aware routing and reusing previously computed KV cache* Why prefill and decode are increasingly handled by different GPUs* When dedicated deployments become cheaper and more reliable than shared APIs* How speculative decoding uses a smaller model to accelerate a larger one* Tool calling, structured outputs, and what LLMs actually do* What it takes to support a new open model on day zero* Grafting Kimi's vision encoder onto GLM-5.2* Retrofitting inefficient model layers with components from other architectures* Why models sometimes collapse into repeating the same token* How hardware, kernels, and race conditions create nondeterministic failures* Preserving model fidelity while making inference faster* How quantization errors can cancel each other out* Why inference optimizations still deliver gains of 20%, 100%, and 200%* How optimized serving can make a model up to 10× faster* NVIDIA Dynamo, KV-aware routing, and distributed model serving* Speculative decoding the speculative decoder* Why local AI is about making models less dumb while data-center AI is about making them less slow* Tensor, expert, and pipeline parallelism across GPUs* Hardware-aware model design, auto-tuning, and the case against mega kernels* Rubin and why inference is becoming a systems problem* Whether modern GPUs are evolving into programmable AI ASICs* Why enormous models like Kimi K3 require GB300-class hardware* Why open-source video generation still trails Veo, Kling, and other closed models* The quadratic attention bottleneck behind long-form AI video* Autoregressive video, real-time generation, and compounding quality drift* Why future video systems may combine autoregressive and diffusion architectures* Training for inference and inference for training* Continuous post-training, deployment, evaluation, and improvement loops* How GLM-5.2 helped optimize the kernels serving GLM-5.2 itself* Why faster networking could unlock dramatically faster decoding* Continual learning, KV-cache compaction, and persistent model memoryShow Notes* How to build a day-0 API for Kimi K3* 22580: From GPT2 to Kimi3, ExplainedPhilip Kiely* LinkedIn: https://www.linkedin.com/in/philipkiely* X: https://x.com/philipkiely* Inference Engineering: https://www.baseten.co/inference-engineering/Ali Taha* LinkedIn: https://www.linkedin.com/in/aliestaha/* X: https://x.com/waterloointernTimestamps00:00:00 Introduction and the 200K-Token Prompt00:03:18 Dedicated Deployments, Speculative Decoding, and Tool Calling00:11:26 Launching Production-Ready Open Models00:19:06 Model Retrofits, Failure Modes, and Nondeterminism00:28:22 Quantization and Canceling Errors00:32:15 The Race to 10× Faster Inference00:40:48 Dynamo, Speculation, and Local vs. Data-Center AI00:50:18 Model Parallelism, Auto-Tuning, and Mega Kernels01:00:55 Rubin, GPUs vs. ASICs, and Custom AI Chips01:10:03 Giant Models and the Limits of GPU Memory01:12:42 AI Video, Quadratic Attention, and Autoregressive Generation01:21:47 Audio, Images, and Diffusion Models01:27:32 Training, Self-Optimizing Models, and Continual Learning01:40:06 Closing ThoughtsTranscriptIntroduction: Baseten, Waterloo Intern, and Inference EngineeringSwyx [00:00:00]: Okay, we're here in the studio with Philip, old friend from Inference Engineering, the book, as well as Baseten and everything that you've done, you and I have done before, as well as Ali. Welcome.Ali [00:00:15]: Pleasure to meet you.Swyx [00:00:15]: Waterloo intern.Ali [00:00:16]: Waterloo intern, always.Swyx [00:00:17]: When did you get “Waterloo intern” as a handle?Ali [00:00:19]: As a handle? Oh.Ali [00:00:20]: I think the rebranding happened mid-March. When I saw it was open, I was like, “I have to take it. Up for grabs.”Philip [00:00:26]: The problem is that Ali is really good at his job and is not gonna be an intern much longer.Philip [00:00:30]: So we have to figure out who's gonna get the handle.Ali [00:00:33]: Well, I'll pass the torch over to the next intern.Swyx [00:00:34]: Oh, okay. It can be, like, you just pass it to another Waterloo grad.Ali [00:00:37]: To another Waterloo intern. No, bruh.Philip [00:00:39]: Yeah.Ali [00:00:39]: Intern.Swyx [00:00:40]: Intern, yeah.Ali [00:00:40]: And no.Philip [00:00:41]: You gotta get an intern from Waterloo.Ali [00:00:42]: Yeah, I've gotta get an intern from Waterloo.Swyx [00:00:44]: Right.Ali [00:00:44]: But they have to follow the path.Swyx [00:00:45]: Oh, it could, but it could come from Baseten, so it's like whoever Baseten gets from Waterloo.Ali [00:00:48]: Right.Swyx [00:00:49]: Has the title of Waterloo.Ali [00:00:50]: It stays in the ecosystem.Philip [00:00:51]: Exactly.Ali [00:00:52]: Halfway through the internship, you either get it or you're out.Philip [00:00:55]: You should also do, like, a big graduation ceremony where you change the handle.Ali [00:00:59]: Just say it.Philip [00:00:59]: For everybody.Swyx [00:01:00]: You guys are good at ceremonies, clearly. We had a nice launch of the book, very successful. But before we get into all that, I wanna start off with a fun question for you. Okay, you're an expert inference engineer. What happens when I send a long query, say two hundred thousand tokens into Baseten's inference? What's the process of query through GPU model routing, balancing, all that? What is all the stuff that we don't think about?Long Context Requests, KV Cache, and Cache-Aware RoutingPhilip [00:01:26]: With a long query specifically, the first thing that I'm gonna ask is, “Have you sent me this query before, or at least part of it?” and I really hope you have, because it's gonna be a lot easier for me and a lot cheaper for you. So the first thing that we're gonna look at is some cache-aware routing, where we're going to see, we probably have a number of instances, a number of replicas up serving whatever model you're hitting. We want to send this one to something with, number one, available prefill workers, and number two, ideally some cached input already there so that we can skip prefill on at least part of these two hundred thousand tokens. If you're doing two hundred thousand tokens, it's probably coding or a multi-turn agent or something where you would expect to have that cached. If you don't, we're gonna have to send it to a prefill worker. We've at least on certain models disaggregated prefill and decode, so you're going to have one set of GPUs that's solely going to process the input, create the KV cache, and get you your first token, and then that's going to be passed over to a separate set of GPUs, which is going to run decode. We're going to iteratively make those tokens. We're probably going to have some speculator model in front of that. I'm going to assume that you're doing coding, and because of that, our speculator model, which assumes you're doing coding, is gonna have a high draft token acceptance rate. If I'm wrong and you're asking me to summarize every Harry Potter book, it's gonna be slower. And then we stream that output to you and account for it, charge you, a couple of pennies and say, “Hey, would you like to send another one?”Swyx [00:03:04]: Except Baseten doesn't charge by pennies.Philip [00:03:07]: Well, yeah, we charge. I'm assuming that we're talking about the public model APIs. If you are setting up a dedicated deployment, then yeah, it's not pennies.Public APIs vs. Dedicated DeploymentsSwyx [00:03:18]: Yeah, one of the key differentiators when I was talking with Baseten initially was that people who want very high volume just need to rent by the box, ‘cause then it's up to you to figure out how to saturate the box.Ali [00:03:31]: And more often than not, it's, like, way cheaper if you're pushing, like, millions of tokens per hour, if you just pay per hour instead of pay per token.Philip [00:03:37]: Yeah, they do. I think that we've increasingly seen a lot of demand for the pay per token APIs, just because everyone wants to try open models, and then once they find a use case that's really sticky, then they move over to dedicated.Swyx [00:03:51]: Is there a best practice on when it's time to swap over?Philip [00:03:54]: Couple reasons. Yeah, reliability, that's a big one, right?Ali [00:03:57]: Like, if they have a very specific use case, they want you to train something specifically for them, like they want their own spec dec, for instance, for their own traffic.Swyx [00:04:04]: Spec dec is speculative decoding.Speculative Decoding and Custom SpeculatorsAli [00:04:05]: Speculative decoding, yeah.Swyx [00:04:07]: You have to explain.Ali [00:04:07]: Sorry. Like, speculative decoding is like, if you have a huge model, right? And so the model is going to be generating one token at a time every single turn, every single forward pass. So we attach, like, this little, like, parasite, like this layer that goes on top of the model, and this model just has to predict. It does three very fast autoregressive forward passes, and it will predict, like, three certain tokens, and then you do one forward stage over the entire original model in order to see if those predictions were correct or not, and then you accept them or you reject them. Now, this draft model is traffic specific, so if you, like, Philip said, if you're summarizing Harry Potter books, I can train exclusively that draft model on Harry Potter books, and I can guarantee you that I'm gonna accept the three tokens every single time. And so with that case, I increase your decode speed. I wouldn't be able to provide this to you if you're a shared endpointSwyx [00:04:53]: YeahAli [00:04:53]: ‘cause I have no idea if you're doing Harry Potter, if you're doing coding, if you're doing English. We don't know. Also, there was a thing in the book that mentioned that if they really cared about a specific threshold, chapter four, I think. Do you remember that?Philip [00:05:06]: Yeah. The things that you can do is you can set a specific, like, batch sizing, a specific, like, parallelism strategy if you're trying to optimize for, like, throughput versus latency. You can. Maybe a NVFP4 quant doesn't pass your benchmarks and you wanna run a model at higher precision, you could do that. There's just a bunch of reasons why you might wanna have your own endpoint and the biggest one, of course, just being, like, you don't have to deal with someone else doing a hundred million tokens of benchmarking traffic at the endpoint when you happen to be trying to serve your users.Swyx [00:05:40]: Yeah. I think one thing that is. That is a classic journey. Like, it's people is asking the, what happens when you type Google into the browser. Tool calling, is that just, you're generating JSON or is there more complication beyond that?Tool Calling, JSON, and Structured OutputsAli [00:05:58]: Certain customers that we have, they have their own post-trained models, and so they demand a tool calling that's not just, like parse a file or go find the weather. It's something that's very specific and you have to do post-training on this. And if the post-training on the model is not good or if the quantization after the post-training to get the inference to be fast, the model will struggle reading the JSON file and reading the tool calling. But it doesn't require its own like sandbox. It's not like it's going to use that tool calling to like escape a sandbox or like it doesn't have to be contained. It can just be a normal dedicated deployment. The challenge with tool calling more and more seems to be that the companies want certain tool calling which is a very sensitive thing to train. And because you're dealing with all of the JSON outputs, if it doesn't like close the end of the request in a very certain manner, you end up with a model that did the tool calling and like the thinking and so as a result of that, it didn't see the result and just hallucinated the result as it decoded. That seems to be the most challenging thing with tool calling, not really the sandboxes model.Philip [00:06:56]: Yeah, that's a challenge on the training side and then on the inference side, there's work that you can do to scope the possible output. So we published this at this point close to two years ago, the solution to this problem which is you make a state machine and you use that to constrain the output to a specific format. So this is the structured output problem. If you remember backSwyx [00:07:27]: Yeah, the specific grammar is,Philip [00:07:29]: Yeah, exactlySwyx [00:07:30]: GML had this thing.Philip [00:07:31]: Yeah. So it's like the old-school “make sure this is only JSON”, return only JSON orSwyx [00:07:38]: YeahPhilip [00:07:38]: Grandma's gonna die type of prompts.Swyx [00:07:39]: Is it BNF grammar? At some point OpenAI had released a thing that was like, yeah, if you want to constrain your output, write BNF grammar, back as NOR.Philip [00:07:47]: In our inference system, it's just a specified output format. And you get the guarantee that your output's gonna be structured along that format. And so applying that to tool calls can like help cut down on. You can still call the wrong tool or call no tool. It doesn't solve the certainty problem but it at least solves the output structuring problemSwyx [00:08:10]: YeahPhilip [00:08:10]: Within tool calls.Swyx [00:08:12]: And MCP is just another form of tool, right.Philip [00:08:14]: Yeah, exactly.Swyx [00:08:15]: As far as there's no special thing there.Philip [00:08:16]: The thing I'm always like explaining to people is the LLM is not capable of doing anything. It's only capable of making suggestions of what to do and then if those suggestions are formatted in a certain way and applied to a system that knows what to do with them, then an action occurs.Swyx [00:08:32]: Yeah. Part of the fun stuff is, this is solved outside of tool calling too. Like in an agent loop if the output is not correct or you're right, like reasoning, tool calling was done in the reasoning trace, just be like, “Oh, I don't know what to do. Let me just try again.” And it might get there after a few tries. And on your point of training, sometimes this is harder in smaller models, so you don't have the same exact quality outputAli [00:08:56]: Right.Swyx [00:08:57]: When you just swap from a big model, right?Ali [00:08:59]: Yeah. I will say that, before, I think we need to go back to inference engineering proper.Ali [00:09:04]: But, I had expected that something would replace JSON because it's hard to stream JSON ‘cause JSON must be complete and you must have open and close brackets and everything. So it's hard to parse something or validate something while it's being streamed. So people invented all sorts of things that are like, I forget the name of some of these alternatives, but it's something like TOML, something like YAML. But JSON seems to be dominant still.Philip [00:09:30]: The JSON outputs aren't that long, right? Like you could have a long-- ‘cause tool calls also contain the arguments in them and perhaps for a certain tool you might pass like a very long argument. But my impression of the median tool call is that it's a relatively small number of tokens, right? So I would expect that speculators are generally fairly good at something as formatted as JSON. And so you would have like a pretty fast decode step there and that the streaming wouldn't be as valuable, but maybe I'm wrong about that.Ali [00:10:02]: I think you're also bounded by the software or that the model is gonna integrate with if the software is built with JSON for the tool calls or if the company that you'- if your customer says that this is how our software works and our tools are interfaced with JSON, you can ask them to like, change their software and say like, “Yeah, this is gonna be better for the model.” but like with the right training shouldn't be that much of a difference. Also more profitable if it outputs more tokens probably.Swyx [00:10:25]: Depends on your business model.Swyx [00:10:27]: It really depends. But I will say that, as a writer with like experience a lot with generated output, I do try to move from text to JSON text which is very long JSON, right? Like there's paragraphs in every field because I'm trying to structure it, right?Philip [00:10:44]: Right.Swyx [00:10:44]: I want you to first make factual statements, then make opinions then make bullet point summaries, have dates, have entity references have your sources for references, all these things. Anyway, so these are things that like I think people who really experiment with structural output have to really care about. But, let's, let's recurse up the stack a little bit. Before we started recording, you mentioned something really cool, which is that there's a lot of engineering that-- inference engineering that goes on when a new model provider releases a new model, right? So let's call it GLM-5.2, Kimi K3. I had previously assumed, especially if it's like, well, GLM 5 to 5.1 to GLM-5.2, like that you've supported them before. Is it that much work?What It Takes to Support a New Open ModelAli [00:11:26]: It's a lot of work.Swyx [00:11:28]: Yeah. Okay. So like, a lot of people, all you guys, right whenever a new model launch like, people rush to say like, “Oh, Hugging Face supports this, Fireworks supports this, Spacetime supports this,” and I'm like, “Yeah, of course we support it.” But what goes into that? What goes intoPhilip [00:11:40]: I think it's more than just support it too, right? It benefits the consumer a lot. Like I think it was with Kimi K2.5 or GLM-5.2 the latest, there was an inference war, right? X provider is at 90 tokens a second. The next day we're at 150. The nextSwyx [00:11:55]: I kinda kicked that off with the GLM-5.2.Swyx [00:11:58]: I wrote a Twitter article about. It got like half a million views,Ali [00:12:02]: Based on being numberSwyx [00:12:03]: YeahAli [00:12:04]: Or it's for something else.Swyx [00:12:05]: Yeah. Which,Ali [00:12:06]: Oh my GodSwyx [00:12:07]: Which then got everyone really excited about, hey, how can we, bend tracks a little bit further and,Philip [00:12:14]: There's a difference between support the model, as in I can make a token out of this model, and support a model, as in I have a production-ready API from this model.Philip [00:12:26]: Getting to the point of I can make a token out of this model is not that hard because generally the, open source inference engines, vLLM, SGLang of the world oftentimes even receive weights ahead of time, maintainers do, or the people making the model merge PRs to ensure support. So you generally can, just get it working on the standard open source stack without too much pain in most cases. The challenge is, every inference company is gonna have own proprietary stack. Some open source components, some in-house stuff. And for any arbitrary model, there's going to be some new stuff. Sometimes you get lucky, like K, two five to two six was, like, pretty similar.Quantization, Speculators, and Production ReadinessAli [00:13:16]: Yeah. It was pure continued post-trainingPhilip [00:13:18]: YeahAli [00:13:18]: If I remember correctly.Philip [00:13:19]: Even in those cases, there's still stuff you have to do. You have to redo the quantization work. You're taking the model from. Generally, these models are not released in NVFP4, and we want them to be in NVFP4 for maximum Blackwell compatibility. So we have to perform that quantization, and, calibrate the quantization to make sure that we're not causing any regression in the model's intelligence. And then we also have to train the speculator, as we've talked about. Generally, we have. We have ZDR, zero data retention on our model APIs, so we don't know exactly the traffic that people are sending us, but we know what's popular. We know that coding use cases are popular. We know that agents, agentic use cases are popular. So we can get public data sets that are representative of that traffic and train general speculators. Now, with speculators today, you need to train the speculator using the base model itself because you're getting hidden states out of the model from running inference on these specific prompts, and that is the training data you use to create the speculator. So there's that process which you need the real model weights for. And then there's of course just the process of, standing up all the infrastructure behind it, loading all this stuff, testing it. And then when there's a new model with a newer architecture, I think that, like, the DeepSeek models tend to be the most challenging as they have, like, the most novel architectural stuff going on, model after model. But every new model has something. Kimi K2 had. Oh, sorry, GLM-5.2 hadAli [00:14:53]: Sparse attention.Philip [00:14:54]: Yeah,Ali [00:14:54]: YeahPhilip [00:14:54]: the DSA.Ali [00:14:55]: Right. Which is brought from DeepSeek.Philip [00:14:57]: Yeah. AndAli [00:14:59]: So you can copy-paste then?Philip [00:15:01]: It kindAli [00:15:01]: I don't know how this works.Philip [00:15:02]: So, like we had to, like, build support for that into our runtime. And you're right, like it is really interesting the way that all of these open source labs borrow from each other. For example, like GLM-5.2 doesn't have vision. So something that, Haley, a guy on our team, if we could take a look at this, he, like, grafted the Kimi vision encoder onto GLM-5.2.Retrofitting Vision into GLM-5.2Ali [00:15:27]: We'll be training the projector.Philip [00:15:28]: Exactly. So if you think about, like, the encoder, there's the encoder, which is the part that looks at the image and turns it into latent information, and then there's the projector which likeAli [00:15:38]: You can say latent space. It's okay.Philip [00:15:41]: And then there's the projector that maps it onto, the model itself, and then there's the model weights. You don't wanna mess with the model weights because you run a chance of making the model dumber at something else for the purpose of giving it vision. So instead, Haley started with just a projector, which is only a handful of millions of parameters.Ali [00:16:02]: That would be, yeah.Philip [00:16:02]: Yeah.Ali [00:16:03]: Can you show the training one?Ali [00:16:04]: Like the way it groksPhilip [00:16:05]: YeahAli [00:16:06]: Very interesting.Philip [00:16:06]: And maybeAli [00:16:07]: That right therePhilip [00:16:07]: Maybe Ali, you should take it from here. You've got a betterAli [00:16:10]: Ooh, double the sandPhilip [00:16:11]: Understanding of this than I do.Ali [00:16:11]: Yeah. You can see, like, he. The way he trained this is really cool. At the beginning, he was training it using just like, “Here's a picture of a mountain. Can you describe what's in this mountain?” And that caused it just like the first, learning walls. Like here you can see this all we're trying to teach it is to translate the encoded. Like it's already taken the encoder from Kimi K. It's taken the image. It'Philip [00:16:31]: Yeah. FrozenAli [00:16:31]: FrozenPhilip [00:16:32]: With adapter.Ali [00:16:32]: Exactly.Philip [00:16:33]: Yeah.Ali [00:16:33]: So the brain is frozen and the eyes are frozen. It's just we're tryingPhilip [00:16:37]: AlignAli [00:16:38]: Interconnect between the eye and the brain, right? So the projector. And so you take the tokens and then he's like, “Oh, can you describe what's in this image?” And he's like, “Oh, it's a mountain,” or it's a person or it's a human, whatever the case is. But that didn't cause complete understanding. So he changed it such that every image was associated with a data set of questions. Like, does this image have a white male? Does this image have birds in the top corner? Does this image have a scientist in it? All of that stuff. And it would have to answer questions correctly. And using not just training on describing an image, but being able to answer question, another question, answer over time. Like you can see the grokking, which is like genuinely insane, that retrofitting vision into a large LLM can learn to that extent. And even for images that it doesn't perform well on, for instance, if you ask it a picture of like Stephen Hawking, “Who is this?” Maybe it doesn't get it, but it will say something like, “This is Albert Einstein.” Like it still understandsPhilip [00:17:25]: Close enoughAli [00:17:26]: That this is a scientist who is a man who has, some significant achievements, all that stuff. So that's like really cool.Philip [00:17:32]: Yeah. So, we've covered Hao Tian before, who the author of the LLaVA paper that did this, a while ago. And I think that's very foundational work for anyone who hasn't done vision work before.Ali [00:17:41]: Same with the CLIP and MetaCLIP, where you go from just captioning to building out questionsPhilip [00:17:47]: RightAli [00:17:47]: Off the image and how much better you can get performance.Philip [00:17:50]: Right. Right. Right. Yeah. But what's, what's so exciting about this is if you look at a model like this. Now, this is a little bit more of a research project. It's not. It got to 56% on MMLU Pro, I think. So not quite frontier. But if you're running this model, you haven't suffered any loss on your GLM-5.2 quality. If you don't have an image, it'll just behave exactly the way it used to. And ultimatelyAli [00:18:14]: Which in the inference code you literally do not include the other part, right?Philip [00:18:18]: Yeah. You would just skip the encoder if you don't have an image input.Ali [00:18:22]: Okay.Philip [00:18:22]: Just confirming.Philip [00:18:23]: YeahAli [00:18:23]: Does it affect a lot on the overall inference side? Like you're not adding much, you're adding a very small vision encoder. These are typically likePhilip [00:18:30]: They're super fineAli [00:18:31]: Less than a billion parameters, right?Philip [00:18:32]: Yeah. It's, - There's a little bit less standardization among vision encodersSwyx [00:18:37]: YeahPhilip [00:18:37]: So the support matrix can be a little bit, sparser. But overall, yeah, it's a pretty, it's a pretty minor component of the overall system. And ultimately what you get out of the system is all of a sudden you have Kimi Vision, GLM weights, and DeepSeek attention all in one model.Open Source Model Grafting and Franken-MergesPhilip [00:18:56]: And that's, I think, a lot of the power and beauty of open source, is that you can take all of these different components and combine them together into a system that's better than anyoneSwyx [00:19:05]: YeahPhilip [00:19:05]: Can be individually.Swyx [00:19:06]: People used to say that you would also do Franken-merges where you would take likePhilip [00:19:10]: YeahSwyx [00:19:10]: Layers from each model.Swyx [00:19:11]: Does anyone do that anymore?Ali [00:19:13]: Well, to your point previously when you were mentioning like, the work that goes into supporting a model when it first comes out, like GLM-5.2 or MiniMax M3 or whatever the case is. Sometimes you do have to like, you do have to switch out some things. Like, for instance, the MiniMax M3 head uses full attention, and with full attention you end up with this like insane bottleneck in spec dec ‘cause you're doing auto-regressive token generation for three tokens, and you're doing this like N squared over all of the tokens that are in your sequence. Your KV cache is like very large because it's not sparse, it's not top K. So we find it better to like, okay, we're gonna replace this, we're gonna replace this layer with a layer from another model that's using like GQA, for instance. And then just with the right training, you can get it to have the same acceptance rate. So it is very possible to retrofit layers from other models and very much needed. If a layer is like inefficient, the training just becomes the challenge, like how do you ensure that you train it properly? Which again to your earlier point is like the mesh between training and inference. As in like you need very good training in order to do fast inference. That's like, I feel like more and more becoming true.Swyx [00:20:21]: Yeah. Anything else on the support side when you say like get it to fully production ready?Loop Detection, Race Conditions, and Non-DeterminismPhilip [00:20:26]: Yeah. I think that there's also a question of just, we can test a model to a pretty extensive degree, but we're trying to get it out quickly and then you see a bunch of other people test it and you get interesting results. There was an issue with, GLM briefly where we had some like mode collapses where it would just output the same token over and over again for certain prompts on certain temperatures. Like once you expose an endpoint to the real world, there's going to be, so many more varieties of things given to it that you're able to, discover and patch things. So it's not just a, day zero process, it's then like for the first week, for the first month, if a model remains popular, like how do you both fix bugs and then continue to push the envelope on performance?Ali [00:21:21]: What do you mean you don't want your model outputting S?Swyx [00:21:24]: Is there loop detection on that stuff, by the way? It still happens like quite a lot, which is surprising.Ali [00:21:30]: We have like we, in our endpoint, like if a model was to output the same token like four plus times, we just cut the generation. We say like, “Oh, sorry, this-- Like try again,” or like we will reprocess the request. ‘Cause we know then, like if it, like if, yeah, it's four times the same token, it's probably collapsed.Swyx [00:21:45]: Yeah. Is there a way to opt out in case I really want that?Ali [00:21:48]: You want that?Ali [00:21:50]: I think there's a way that we have to handle it. I'm not exactly certain, but I feel like in certain models, like when they output something like you can imagine, like a table for instance, and so they want, they wanna draw like 12 dashes and 12 dashes. Yeah, I think there's a way for that to happen. I think we only do it on certain tokens. Like we exclude certain special characters.Swyx [00:22:07]: Yeah.Ali [00:22:07]: So we only do it on like certain like S is the most common almost. GLM-5.2Swyx [00:22:11]: OhAli [00:22:11]: And I think it was DSV 4 as well. Like you'd just have like looping issues where like you literallySwyx [00:22:17]: ItAli [00:22:17]: Just have like S.Swyx [00:22:18]: Yeah. Is there a special, something special about S? No, just randomlyAli [00:22:21]: It just seems to be the one token involved.Swyx [00:22:23]: Yeah. And it'Philip [00:22:24]: Is thereSwyx [00:22:24]: And it's only temperature 0Ali [00:22:27]: NoSwyx [00:22:27]: Even at other temperaturesAli [00:22:27]: Even at like 0.9 or whatever, it will still, it will still collapse.Swyx [00:22:30]: That's weird, right?Ali [00:22:30]: It's, it is an inference problem to be honest, like a software problem. Like oftentimes, the image you run will-- like NVIDIA will release an image for instance, and if we will upstream the changes from their latest TensorRT-LLM image into our stack, we'll find that it fixes it. Or oftentimes this will only happen in an inference engine that you're using like SGLang. But if you were to switch to vLLM, that isn't the case. So it seems to be like an extremely like deterministic software issue and not really a model issue. It's not like a weights problem. Like I'- we'll say like, “Oh, it's a problem with the quant. We did PTQ wrong,” right? But that isn't, that doesn't make sense because the same weights used with a different inference engine does not repeat the problem. And sometimes it's, the kernels that are being used in the backend have like these very subtle sometimes race conditions, where if you were to use this model hosted on one cluster, you will never get this problem.Swyx [00:23:19]: Oh my God.Ali [00:23:19]: But if you host it on a different cluster, you will. And the reason is the KV cache transfer from a node to node in that one cluster is using a slower interconnect than the node to node in another cluster. So that exposes the race, whereas in another cluster it doesn't. So then you end up just like, okay, this model is not gonna be hosted on this cluster. We're gonna host it on, another cluster because that cluster exposed that problem. But then it ends up with like, okay, is it the software? Is it the model weights or is it the hardware?Swyx [00:23:42]: There is a thing about this with temperature 0 still not being deterministic, right?Ali [00:23:46]: Right.Swyx [00:23:46]: Mostly because of hardware. Even at temperature 0 same model, you won't always get the same output.Swyx [00:23:52]: Even-- But I'm surprised by the race condition one because, I thought PyTorch was a graph that like guarantees that you at least, execute things in the right order.Ali [00:24:02]: Well, yeah, true. Like I'm not, I'm not saying that there is. Like well, you have things like PTL optimizations where like you can start a kernel before the end of the previous kernel, and that's like ‘cause you want to do that because there'sSwyx [00:24:12]: It's like pipeliningAli [00:24:12]: Expense. Exactly.Swyx [00:24:13]: Yeah.Ali [00:24:13]: But it'- But you don't do it cleanly. Like you overlap a little bit of the execution. No, it is very possible that the kernel itself, like that one block that is supposed to be running in this instance of time, that kernel itself has a race condition. For instance, like a missing barrier. Like often if you're designing a kernel and you want it to make it to be very fast, if you don't test it extensively, you'll, you'll have certain threads access data points from registers before they've been written to by other threadsSwyx [00:24:36]: YeahAli [00:24:36]: For example, because like your barrier is wrong or your synchronization was wrong. But yeah, like the testing itself is very difficult in those like, andSwyx [00:24:42]: And there's no like borrow checkerAli [00:24:45]: What does that mean?Swyx [00:24:46]: Like Rust. Like the. If you're trying to have like memory safety It sounds like a comparable problem.Ali [00:24:52]: Well, yes, but you're working in CUDA, right, NVIDIA GPUs. Like- You just need a higher level language like modular Maybe that's what modular is supposed to do. I don't know.Quantization Quality and Vendor FidelityVibhu [00:25:00]: How do you see keeping quality of the model? So you talked about all these steps of, okay, you gotta do quantization, train your own speculative decoderAli [00:25:07]: RightVibhu [00:25:07]: Run on different hardware. Looking at other model providers, okay, you kicked off a inference speed race on the consumer end. What goes into keeping quality the same across them, right? Sure, you can run benchmarksAli [00:25:22]: YeahVibhu [00:25:22]: But, like, how do you determine how much quantization are there standards? What goes intoPhilip [00:25:27]: There's a few things on quality. Most inference optimizations are lossless. KV caching, for example. You are just recomputing or preventing recomputing the same values. Speculation, of course, if a draft token is wrong, it gets rejected. The main lossy optimization is quantization. And that really comes down to, number one, data format, number two, which parts of the model you choose to quantize, which layers, and number three, like doing a lot of calibration on the quantized weights, to ensure that you're preserving all the outliers. There's other tricks that you can do, though. A big one is long context, ‘cause one thing you asked at, right at the beginning is, “Oh, what's gonna happen if I send a 200,000 token request in?” So with a long input sequence, you need to, store a lot more information. You need to process a lot more tokens. And so even if a model has a context of a certain length, you might, as an inference provider, choose to build an API with a shorter context length, and of course a full length one as well. Because if someone doesn't need the full million token context, for example, you can get them better performance. I don't know if that's exactly like quality of the model. The way that I think about quality is to what degree are we faithfully serving the original model? If you think of a golden implementation of a model that performs exactly the way the model is designed to perform, I think of quality as how close are we getting to that, 100% fidelity of the model.Philip [00:27:13]: You can also, of course, think about quality from the training side and how do you push yourself past 100%. But when I think about purely inference optimizations, it's getting faster while staying as close to that 100% fidelity mark as possible. And certainly our standard internally is that, like you should not be able to tell the difference between our API and a, official API. I think Kimi in particular does a good job of vendor benchmarking hereAli [00:27:41]: YesPhilip [00:27:41]: Where they haveAli [00:27:42]: They released an actual vendor benchmark.Philip [00:27:43]: Exactly, yeah.Ali [00:27:44]: ‘Cause they accused, some people, Amazon? There was some provider that was not doing very well on Kimi's benchmark.Philip [00:27:50]: Yeah.Philip [00:27:51]: So, with Reflect we probablyVibhu [00:27:52]: This was a long time ago, right?Philip [00:27:54]: No.Ali [00:27:54]: Yeah, like threeVibhu [00:27:55]: They alsoAli [00:27:55]: Four, five months agoVibhu [00:27:57]: This also happened with, I don't remember which model, but they pulled out quite a few, and then they started a whole chart about this. It might have beenPhilip [00:28:03]: Kimi Vendor Verifier.Ali [00:28:04]: Yeah.Philip [00:28:05]: Yeah.Ali [00:28:05]: Yeah, ‘cause you, ‘cause you'd be pissed, right? Like if you'Philip [00:28:07]: Yeah.Ali [00:28:07]: If like if I'm a consumer and I'm using like Amazon's endpoint for instance, and I've used Kimi and I'm like, “Oh my God, like this is bad,” I'm not gonna say, “Oh, Amazon quantized the model in a bad way.” I'm gonna say, “Oh, Kimi sucks.” Right?Philip [00:28:17]: Yeah.Ali [00:28:17]: So it seems like that makes sense.Philip [00:28:19]: Yeah, they care. They care.Vibhu [00:28:21]: Justifiably.Ali [00:28:21]: Yeah, justifiably.Vibhu [00:28:22]: This is probably a stupid question, but just checking, has anything improved from main quantization?Philip [00:28:28]: Yeah.Vibhu [00:28:28]: Like, is quantization always strictly worse?Ali [00:28:30]: Well technicallyVibhu [00:28:32]: NoAli [00:28:32]: It's a lossy. QuantizationPhilip [00:28:33]: YeahAli [00:28:33]: Is a lossy, it's a lossy implementation.Philip [00:28:36]: Speed improvesVibhu [00:28:36]: Speed improves.Ali [00:28:37]: It the number, likeVibhu [00:28:38]: No, I' always look for inverse scaling laws.Philip [00:28:40]: Yeah.Ali [00:28:40]: Yeah.Vibhu [00:28:40]: This is something I learned from Noam Brown, where like things that normally act in one direction sometimes do.Philip [00:28:45]: Well, technically when you run a benchmark, because these models are deterministic, sometimes your,Ali [00:28:52]: YeahPhilip [00:28:52]: NVFP4 quant is like, two basis points higher than yourAli [00:28:56]: No, it's noise. It's noise.Philip [00:28:57]: Yeah, exactly. I'm like, yeah, it's, it's within. That's why I always say within margin of error.Philip [00:29:01]: And I stopped saying that because everyone assumes that what is, well, within some margin of error, we're barely inside of that to the worst, so we're saying. But yeah, sometimes it's just like, gives you a higher output score. But like Ali said, that's noise. To my knowledge, you're not necessarily making the results better. You're just trying to, again, like keep your fidelity as close to 100% to the original model.Layer Selection, KL Divergence, and Better QuantizationAli [00:29:27]: There is, to your point, research that we did on MP. I don't know if you are able to pullPhilip [00:29:31]: YeahAli [00:29:32]: A tweet we did. One of our research interns, Joshua, I think it's a tweet on how we have 20% better quantized GLM-5.2 than NVIDIA. Essentially what we found throughout like this month research is, okay, quantization is a lossy. It's. You're compressing the data from, occupying 16 bits to occupying, four bits, for instance. And so you're losing some information, and you're trying to minimize that. And so when I say that I'm gonna quantize the model, my job becomes how do I find the layers that I can quantize, and how to find the layers to not. For instance, with image models, I don't quantize modulation layers, and I don't quantize out projections because those two are. Like out projection is what you see as the user. Modulation is what the model sees or understands. Right, exactly. And so to his paper, do you have the. It doesn't have the. Yeah. It's a long paper. I don't know if I can findVibhu [00:30:25]: If there's a part to search or it's probably in the thread.Ali [00:30:28]: It's probably in the thread.Vibhu [00:30:29]: Yeah.Ali [00:30:29]: But the long and the short is it is very possible that quantizing more of the model makes the results. Like if I have a model that I quantize layers one, five, and 10, and another model where I only quantize layers one and It is possible that the model in which I quantized more information is going to perform better because the quantization errors have canceled out. And so what Joshua showed in his mathematical proof where he had like a verifier in, is that you can predict which layers are going to have quantization errors that will cancel out with each other, and you choose to quantize those layers. And so the result of doing this mathematical quantization is you end up with a model that's 20% more quantized than another provider, so you get 20% more throughput of it because there's more layers than running an NVFP4, and your quality is better than that other quant because the layers that you chose to quantize have their errors cancel out, like one layer skewed to the right one layer skewed to the left, one layer skewed to the right. Your final logits distribution is more similar to the original distribution of the model, so you have better fidelity. And so the way we proved this was with KL divergence. So instead of just scoring on the benchmarks, we scored the KL divergence between the logit distribution of the quantized model and the logit distribution of the original full precision model, and we showed that with this technique we get. If your probability distribution on the logits which token it wants to select is more of the same as the original model, you're probably gonna end up staying true to the original model. So yeah, so it seems like previously before this, it seemed like the industry was, well, the more you quantize, the worse it's gonna be, ‘cause the more loss you introduce. That's not exactly, not necessarily true. So yeah, doesn't improve it, but can cancel out.Philip [00:31:57]: I think it might be this, but reminds me a good bit about pruning where you can prune off certain layers.Philip [00:32:03]: But very interesting. Didn't know this was a whole paper you guys put out.Ali [00:32:06]: It's. Fun fact, it was originally 72 pages, this paper, and then we decidedPhilip [00:32:11]: WowAli [00:32:11]: We can't tell. We couldn't release it. So it's now 45.Swyx [00:32:15]: Still 39 pages, so very substantive. We talked about evals and all these things and, like what's possible in terms of speedup? Like it's like probably like the numberInference Speedups and BenchmarkingSwyx [00:32:25]: Thing that people do wanna care about, and it's something that you wrote about in your post. Like official API is 70 tokens per second, and you push it up to 90. Is that like a normal thing?Philip [00:32:36]: So what's cool about working in inference, the reason that I think inference is going to be a useful place to do engineering for a long time, is that if you look at highly optimized domains like, say, finance, if you're in finance, you measure how much better you got in basis points. It's like, “Oh, I got five basis points better, like twentieth of 1% better,” that's huge news because everything is so optimized. When we publish optimizations, it's 20%, it's 100% it's 200%. So there's still probably like a lot further to go, honestly. Like you'll, you'll know that inference is pretty much solved when researchers start publishing about how they got 1% faster at something.Swyx [00:33:19]: Which by the way, because I am from the finance background, in the ‘70s, that was the margin at the time. When you did quantitative finance research, you would findAli [00:33:27]: And like 20%, tens of percent.Swyx [00:33:29]: That's. Yes.Philip [00:33:29]: Yeah.Swyx [00:33:30]: And now it'Philip [00:33:31]: Tiny fractionsSwyx [00:33:32]: For those people interested, look up Andrew Lo's paper. He had a really interesting illustration of quant, stat arb, distribution, narrowing down from like those kinds of 20% differences in the ‘70s, down to nothing today, which is very cool.Philip [00:33:48]: Exactly, and we're at the beginning of the same type of thing. Now benchmarking is hard. I think anyone will tell you that, and benchmarking provider speeds is hard because there's so many variables that go into it. What hardware are you using? How much load do you have on the system? What's the exact nature of the prompts and input and output sequence lengths? All that stuff. But overall, when you start stacking these improvements, you're looking at multiples. You can look at it. The most common form, of course, is TPS, tokens per second, which is bad naming by us in the industry, ‘cause there's two tokens per second. There's tokens per second, the throughput number, and the latency number.Ali [00:34:31]: TTMT, yeah.Philip [00:34:32]: Like total tokens per second out of the, out of the GPU as a throughput number. Most people only care about tokens per second as the latency number, which we should call ITL, intertoken latency, but we don't.Philip [00:34:44]: Anyway, so you can imagine a standard API without many optimizations for a 1 trillion parameter model operating somewhere in the 30 to 50 tokens per second range for reasonable traffic profile. And we generally see the goal of, pushing to 10X that. But, not necessarily day zero, but by stacking enough optimizations, if you have, say like four optimizations, each of which doubles performance. Or sorry, three optimizations, each of which doubles performance, then you stack that up, that's an 8X gain. That's the order of magnitude that we're working with in this space. We're trying to make things substantially faster, not just go from like 70 to 90.Swyx [00:35:38]: Are you saying you've. You have done that?Philip [00:35:40]: So let's say you have as a reasonable baseline, 30 or 40 tokens per second. You can achieve 10X that. So like on GLM-5.2, if you run it unquantized, perhaps on H100s even, and you're just using an off-the-shelf inference engine with no particular optimizations, no speculator, nothing extra around like KV routing, no disaggregation, you're, you're probably, yeah, looking at that like 30 to 40. You think that's like a reasonable baseline?Swyx [00:36:12]: Right. Right.Philip [00:36:12]: To get to something like 10X, there's a lot of trade-offs that you're making. If we're running at more like a 300, 400 tokens per second range, you are using the best hardware possible. You have a optimized speculator. You have done all of your quantization work. You are Seeing a pretty high cache hit rate. You are running with a reasonably small batch size and a parallelism configuration that is tuned for latency versus throughput, but it is possible. So the spreads that you see if you, like, go on artificial analysis or you go on OpenRouter and you look at, the worst provider to the best provider, oftentimes can hit that range. 10X is of course very aggressive. It's oftentimes maybe more of a four to six times improvement. But that's the performance that makes us really excited, is when we can get these huge gains, not just go from 70 to 90 tokens.Stacking Optimizations: NVFP4, Speculation, and DisaggregationAli [00:37:19]: It's also, like, hardware dependent. Like, ifPhilip [00:37:20]: YeahAli [00:37:20]: If you have a thing where you're serving it on just, like, a node of H100s and then you throw, like, you shard the model across, like, four nodes of B200s. Like, you can definitely increase the speed with just throwing more hardware at it. Like, normalizing for the same exact hardware and the same number of GPUs.Philip [00:37:35]: Yeah. Then you're looking at, like, a two to 4X improvementAli [00:37:38]: Right. RightPhilip [00:37:38]: Depending on the inference optimizations. So yeah, it's. Some of it's, what's the call, and some of it's who's the driver.Vibhu [00:37:46]: If you break down the two to 4X, say the example is run GLM-5.2Ali [00:37:51]: YeahVibhu [00:37:51]: On B200sAli [00:37:53]: YeahVibhu [00:37:53]: Single node, right? What's, like, the cost trade-off for effort to get, like, the last bit of juice out versus what should people just think of, right?Ali [00:38:01]: Spectre quantization. Yeah.Vibhu [00:38:03]: Spectre quantization.Ali [00:38:04]: That's, that's, that's like 95%. LikeVibhu [00:38:06]: And how far does that get you? And how easy is that for the average person to do? So say right I wanna throw the weights of GLM-5.2 on a node of B200s, how easy is it to find speculative decoder- decoder model or already quantized model? How much work goes into it?Philip [00:38:23]: If you're doing it up front, it's quite a lot of work. If you're doing it today, there's going to be people who have published things that you can just, you can just grab some NVFP4 weights. You can grab a speculator. Yeah, if we're thinking about, like, what are the 2Xs we're stacking, going from, BF16 to NVFP4 is, it's not quite a 2X, right? It's like. I think it's about, like, 30 to 40%, from 16 to 8, and then another 30 to 40% multiplied from, 8 to 4. So that doesn't quite get you a 2X, but, like, roughly a 2X. Speculator, roughly a 2X. Disagg on top of that if you're able to get enough hardware and put enough traffic through it, another roughly a 2X. And then you add in some, double-digit percent increase from having just a better runtime with, the latest kernels and stuff behind it. And that's how it stacks up.Ali [00:39:21]: YeahPhilip [00:39:21]: So building each of those, like, building the, quantized weights is, for someone who really knows what they're doing, hours to days of work. Building the speculator, again, like, hours to days of work. And the, disagg setup, hours to days. Well okay, but like once you haveAli [00:39:39]: Once set up. Once set up. YeahPhilip [00:39:40]: Yeah, getting disagg working for the first time, I'm saying, of course, is very difficult.Philip [00:39:44]: The marginal implementationAli [00:39:48]: Like, if you're just grabbing, like if you are a person, like just a normal consumer who has access to, like, a node of B200s and you're wondering, “How can I just host it myself?” You don't need to quantize the model yourself. There's always gonna be, like, an open source quantized checkpoint. NVIDIA's gonna push one out if no one else does. You. Usually, the providers will have their own spec dec that they've trained as well. You don't need to train your own spec dec. You can just use that as well.Philip [00:40:09]: Yeah. Like, GLM-5.2 has its own MTP.Ali [00:40:13]: Right. Right.Vibhu [00:40:14]: What's multi token prediction?Philip [00:40:15]: Yes.Ali [00:40:16]: I'm justVibhu [00:40:16]: Can you explain that?Ali [00:40:16]: I'm just an expert.Ali [00:40:18]: I can do it for you in case I get it wrong?Vibhu [00:40:20]: No.Vibhu [00:40:21]: Yeah, you should correct if we're wrong, but their multi-token prediction can be used for self-speculative decoding.Ali [00:40:27]: I'm not sure. I'm not gonna correct that.Vibhu [00:40:28]: Okay. I'm semi-confident in thatAli [00:40:30]: Okay. YeahVibhu [00:40:30]: But someone can check. But it's useful to paint the story of, okay, not just the average person, but say a company wants to switch from serverless inference I wanna throw this up on. I wanna rent some GPUs, throw it up. These are the steps you take to do significantly faster than just put it behind vLLM.Ali [00:40:48]: Right.Vibhu [00:40:49]: I was waiting for a mention of Dynamo.Vibhu [00:40:51]: I feel like, that's supposed to be the baseline that you measure against.Dynamo, KV Routing, and Disaggregation ToolkitsPhilip [00:40:55]: I would think of Dynamo as less of a box system and more of a toolkit for building with. So when we talk about doing aware routing, when we talk about doing KV offloading, when we talk about doing, PD disaggregation, Dynamo fundamentally is. By the way, Dynamo is an open source library from NVIDIA.Ali [00:41:17]: We've done a pod with KylePhilip [00:41:18]: OkayAli [00:41:19]: Kyle Cranin.Philip [00:41:19]: Cool. So then your listeners know then that it supports all the different inference frameworks. And it is multi hardware, which is interesting.Ali [00:41:28]: But it's just a router, it's not like an optimizer layer.Philip [00:41:30]: Yeah. All it does, like, what Dynamo is good at, it is a library for moving information around your cluster, around your hardware. So if you have, KV cache on one place and you need it to be somewhere else, Dynamo coordinates NIXL for you to move that around.Philip [00:41:49]: That doesn't mean that, like, out of the box, you just say, “Pip install Dynamo,” and then you get, like, a massive performance speed up. It's more of a developer toolkit.Ali [00:42:01]: Yeah. I would have said it would. It comes with a set of defaults that you can then swap out.Philip [00:42:06]: It does. If the industry at large, I think, was, like, rolling out all of these deployments, standard, then I think it would be, like, a credible baseline. But, we've got to, we've got to benchmark against, like, what we're seeing in the wild.Speculative Decoding Methods: Medusa, EAGLE, n-Gram, and Spec-SpecVibhu [00:42:23]: I did wanna talk a little bit more about PD disagg, because that is probably, like, number three after quantized and speculative decoding. In your book though, I was just gonna pull out the book.Philip [00:42:31]: Yeah.Vibhu [00:42:32]: Like section 522 on Medusa, 523 on EAGLEPhilip [00:42:35]: YeahVibhu [00:42:36]: 524 on gram.Philip [00:42:37]: It's 55, would be disaggregationAli [00:42:42]: Yeah. Well, no, I just wanted to dwell a little bitPhilip [00:42:44]: YeahAli [00:42:44]: The other. Like, so what do you choose to include? What do you choose to not to include? Because there was all these other techniques.Philip [00:42:51]: Yeah.Ali [00:42:51]: Are these still relevant? Because I think they came out, like, a year and a half ago maybe.Vibhu [00:42:55]: Medusa is quite old.Philip [00:42:56]: Yeah, Medusa's old.Ali [00:42:58]: It was old.Vibhu [00:42:58]: But is it in the book as a good, here'sPhilip [00:43:01]: BaselineVibhu [00:43:01]: Baseline vanilla understand it?Philip [00:43:02]: Like you should know this.Vibhu [00:43:03]: Like I read the paper, I'm like, “ it makes so much sense.”Philip [00:43:05]: Yeah.Philip [00:43:05]: So with the book, I had a couple goals. One was to give people just a working vocabulary for the space as a whole, and the other was to give them some intuition about how each of these techniques works. As I mentioned in my AI Engineer talk, which is the first public addendum to this, the speculation space has moved much faster than everything else. So yeah, even at the time that I wrote the book Medusa, I very much included as a way for people to understand how the space evolved rather than what the most modern technique is. And now of course, there's DFlash, dSpark. There's, there's newer techniques even than EAGLE, although EAGLE is still very commonly used.Ali [00:43:51]: SpecSpecta.Philip [00:43:52]: Yes. Speculative decoding.Vibhu [00:43:54]: What canAli [00:43:56]: Oh, it's a paper by Tri Dao and it's like, it's doing speculative decodingVibhu [00:44:00]: HuhAli [00:44:01]: For the speculative decoder.Philip [00:44:02]: Oh, in spec- oh my God.Ali [00:44:02]: It's literally just an another. It's like, yeah, that's the most simple way to explain it, and it seems like he got trivial speed ups there. But it seems that the complexity with training, it's almost like in our mind at least, it's almost as complex as training GANs. Like it's like a very delicate balance and oftentimes you, it's just but yeah, it's literally speculative decoding on speculative decoding.Vibhu [00:44:21]: Speculative.Ali [00:44:22]: Yeah. We saw this paper.Vibhu [00:44:24]: It's interesting, right?Ali [00:44:24]: Yeah.Vibhu [00:44:24]: I wouldn't even expect it to be very particular to train, I wouldAli [00:44:29]: Right.Vibhu [00:44:29]: The naive part of me is like, okay, train speculative decoder.Ali [00:44:32]: But like, and it makes sense, like the whole idea of speculative decoding is you. It's like, it's like almost like the iPhone auto predict version but for a normal model, right? Like you're just, you're just, generating three tokens and you're like, okay, I'll do prefill on them. And so you save those three turns for your original model. Now your speculative decoder is doing three turns of auto regression, so why not just have an even smaller model?Ali [00:44:53]: The other question there is what are the size of speculators? So say forPhilip [00:44:58]: Right. It's like a billion parameters.Ali [00:45:01]: Like for MiniMax, it's. Yeah. It's like one layer. It's like one 60th of the original model usually.Philip [00:45:06]: Yeah. I think we should do a paper when we get back to the office.Philip [00:45:10]: SpeculativeAli [00:45:11]: SpeculativePhilip [00:45:11]: Decoding.Ali [00:45:13]: No, it's, it does seem like how, when do you stop? But then it also seems like if you're able to train spec-spec decode for instance, right? Like if you're able to have a small model that is accurately predicts what the intermediate speculator is gonna predict, that is able to predict what the original target model's gonna predict, then why not just use that smallest model directly, right?Vibhu [00:45:34]: Yeah. This isAli [00:45:35]: Like it seems likeVibhu [00:45:35]: Adjacent to the routing problem.Ali [00:45:36]: Right.Vibhu [00:45:36]: Yeah.Ali [00:45:36]: Right.Philip [00:45:37]: The thing with speculators is one of the practical constraints on using them is that you do have to run a small model on the same hardware that you're running the big model on. There is a orchestration and resource competition problem inherent in that, and that is one of the constraints on speculation in general, is that draft tokens cost resources to create and cost software complexity to manage. And so if you have like infinitely recursive speculators, you add in quite a bit of that complexity on the actual implementation within the inference engine as well, not just in the training process.Vibhu [00:46:17]: I was gonna say, I would wonder if you could do similar, like distillation and pruning of, it's the same thing, it's just a model. Can we not just distill a lot of the weights, quantize the speculator, out of my domain? The question that also comes up is, this is all for big server workloads, right? How much of this applies to, say I have this MacBook, I wanna run Gemma really efficiently. Similar problems, not the same?Local AI vs. Data Center InferencePhilip [00:46:45]: Pretty different. I talked to Selo, about this on his podcast a couple weeks ago. The difference between inference engineering for the data center and for production workloads versus inference engineering for local AI, is that we start with fundamentally like different constraints and different goals. With local AI, it's how do I fit this model onto my hardware and then make it less dumb? And with data center influence, it's how do I load this model and then make it less slow? And we care about less dumb, and they care about less slow. But the local AI inference engineering ecosystem, I think has a lot for us to learn from in the data center space. They are experts in various forms of quantization, including dynamic quantization that we just don't touch, in the pruning, in the distillation, in the, layer removal. There'Ali [00:47:42]: Layer removal matters less.Philip [00:47:43]: Yeah. There'Ali [00:47:44]: No one loves pruning really.Philip [00:47:45]: Yeah. Well, but the, but they doVibhu [00:47:46]: Which is surprising, right? But that's, that's a whole different thingPhilip [00:47:48]: Just to fit something on the laptop.Ali [00:47:50]: Right.Philip [00:47:50]: So yeah, it's a, it's an interesting, it's an interesting space. Not necessarily that like their techniques make sense for us to do in the data center, because we have different resources and different goals, but more that the process as well as the openness of that field is something to, admire.Ali [00:48:12]: Yeah. Like to your point, like, certain optimizations that would. Like for instance, Turbo Quantum Sharper, like it made such huge hype on that and we did like a whole deep dive on Twitter and like said, what is it? How does it work? Why is it good or not? And it took off and it was implemented on local devices because your memory bandwidth is so slow on like a MacBook, for instance. But try putting the same thing on like an NVIDIA GPU on a B200 Turbo quant would not be. Like, it would not be used. Like, NVIDIA - Like, NVIDIA made it clear that this is not a good optimization, and we've seen it firsthand where the overhead of doing dequantization, quantization of, in the kernel itself with turbo quant kernel, each end is much slower than the time that you save from doing the bandwidth. ‘Cause on the B200s, you have like 3.5 terabytes per second. You don't need decrease the storage that much. You don't need to do, FP4 KV cache. You don't need to use a requant. There's, there's, there's better optimizations to be made. But on Edge devices, it's extremely important, it's extremely useful. So, seems to be, like, different optimizations there, but then they're all uniquely combined with like all you wanna quantize the model, you wanna do speculative decoding, like certain common prefixes with bothPhilip [00:49:18]: Principles.Ali [00:49:19]: Yeah, exactly. Exactly. Exactly.Philip [00:49:20]: They also do a lot of work on, model parallelism, especially over, heterogeneous topology, where you have, some sparks and they are wired together with, Ethernet, DGX sparks.Ali [00:49:35]: Yeah, this is the Exo Labs guys.Philip [00:49:36]: Yeah. You have, a nu
New Hampshire Unscripted talks with the performance arts movers and shakers
(Celebrating 80 grand years of community radio!) Rubin Lichtenstein comes back home to the WKXL NH Unscripted studio with some great updates about his business “Rubin's Hot Sauce”. Rubin, CTV (Concord Community TV) and the restaurant 90 Low, located in Eagle SQ., got together and created a cooking show. It's been getting great reviews and personally I'm loving it. Rubin has news about working with Hannaford Supermarkets, gives us updates on some of the fairs and festivals he'll be doing this year and we get to discuss his 5 and 10yr plans for the biz.
calvaryhouston.com
Olaf und Tom waren diesmal beim Vollplaybacktheater (VPT) zu Gast. Anlass war die aktuelle Show „Der Fluch des Rubins“, mit der das VPT derzeit wieder auf Tour ist. Vor Ort ergab sich die Gelegenheit zu einem ausführlichen Gespräch mit Christoph und David vom VPT. Dabei ging es um die Entstehung der Show, die besondere Arbeitsweise des Ensembles und natürlich um die Frage, wie aus den bekannten Hörspielvorlagen ein Abend auf der Bühne wird, der irgendwo zwischen Theater, Hörspiel und herrlichem Chaos liegt.
En début d'émission, hommage à Dieudonné Larose, voix emblématique de la musique haïtienne. L'artiste est décédé le 9 janvier 2026 à l'âge de 80 ans. La chorale des Chérubins Gospel a été élue meilleure chorale de France, le 26 décembre 2025, dans une émission diffusée sur France 3. Ils répondent aux questions de Claudy Siar, Warra Charlotte Gomis et Queen Stelyna. Playlist du 12 janvier : Hommage à Dieudonné Larose, figure incontournable de la musique haïtienne. Décédé le 9 janvier au Canada, Dieudonné Larose occupe une place importante dans le konpa qui était pour lui une musique de revendications. Dieudonné Larose - Haïti (2003) Dieudonné Larose et Missile 727 - Mandela (1992) Dieudonné Larose - Hommage à Nemours (1999). Depuis 40 ans, la Chorale des Chérubins Gospel, composée de 100 artistes, transmet un message universel d'amour au rythme de sonorités venues d'Afrique, des Caraïbes et États-Unis. Ils se sont produits aux côtés de Céline Dion, Catherine Lara, Maurane, Jeane Manson, Francis Cabrel ou encore Alain Souchon. Les Chérubins Gospel - Medley explosif Les Chérubins Gospel - Happy day Les Chérubins Gospel - When I think Les Chérubins Gospel - Marche près de moi. Plus d'informations sur Les Chérubins Gospel Retrouvez la playlist officielle de RFI Musique.
En début d'émission, hommage à Dieudonné Larose, voix emblématique de la musique haïtienne. L'artiste est décédé le 9 janvier 2026 à l'âge de 80 ans. La chorale des Chérubins Gospel a été élue meilleure chorale de France, le 26 décembre 2025, dans une émission diffusée sur France 3. Ils répondent aux questions de Claudy Siar, Warra Charlotte Gomis et Queen Stelyna. Playlist du 12 janvier : Hommage à Dieudonné Larose, figure incontournable de la musique haïtienne. Décédé le 9 janvier au Canada, Dieudonné Larose occupe une place importante dans le konpa qui était pour lui une musique de revendications. Dieudonné Larose - Haïti (2003) Dieudonné Larose et Missile 727 - Mandela (1992) Dieudonné Larose - Hommage à Nemours (1999). Depuis 40 ans, la Chorale des Chérubins Gospel, composée de 100 artistes, transmet un message universel d'amour au rythme de sonorités venues d'Afrique, des Caraïbes et États-Unis. Ils se sont produits aux côtés de Céline Dion, Catherine Lara, Maurane, Jeane Manson, Francis Cabrel ou encore Alain Souchon. Les Chérubins Gospel - Medley explosif Les Chérubins Gospel - Happy day Les Chérubins Gospel - When I think Les Chérubins Gospel - Marche près de moi. Plus d'informations sur Les Chérubins Gospel Retrouvez la playlist officielle de RFI Musique.
(00:00) Reaction to Pacers going up 2-1 (00:13:11) Are the Knicks handling head coach search wrong? (00:25:36) Expect Rodgers and Steelers to have high-powered offense? (00:38:31) How embarrassing would losing series be for Thunder? (00:52:56) 49ers as good as anyone with a healthy CMC? (01:03:03) Fanatics CEO Michael Rubin joins the show (01:22:13) Which team needs Durant the most? (01:26:33) Anyone switching to pick the Pacers? Learn more about your ad choices. Visit podcastchoices.com/adchoices
If you've ever felt torn between wanting to improve yourself and just wanting to be… this episode is for you. I sat down with Gretchen Rubin—best-selling author, researcher, and self-knowledge expert—to unpack some of her most powerful insights from her book, “Secrets of Adulthood.” We're talking about real-life advice you can actually use… the kind that makes you go, “Oh, that's me.”
WATCH FULL VIDEO INTERVIEW HEREIn this episode of Lab Rats to Unicorns, John Flavin is joined by Dr. Kate Rubins, NASA astronaut, microbiologist, Army Major, and the first person to sequence DNA in space. With a PhD in cancer biology from Stanford and groundbreaking research on viruses like Ebola and smallpox, Dr. Rubins has spent her career at the intersection of biology, innovation, and exploration. From her early work in infectious disease labs to running her own research facility in the Democratic Republic of Congo—and ultimately conducting over 200 experiments aboard the International Space Station—Dr. Rubins exemplifies what happens when science meets mission. In this episode, Kate reflects on her path to becoming an astronaut (spoiler: she applied as a joke), what it's like to conduct molecular biology experiments in microgravity, and how the constraints of space research are driving innovation in diagnostics, biotech, and health access on Earth.
In early 2021, as Dr. Kate Rubins was floating above Earth in the International Space Station, she decided she wanted to give back to the country that had given so much to her. She immediately commissioned for the Army Reserves, and today is both prepping for NASA's upcoming moon missions while also doing microbiological research and training for the Army. Hosts LTG (Ret.) Leslie C. Smith and SMA (Ret.) Dan Dailey sit down with MAJ Rubins to discuss her career as a microbiologist, what lessons she's learned in the Reserves that she applies to her NASA work and what it's like to tie your shoelaces in zero gravity. Guest: MAJ Kate Rubins, PhD, U.S. Army Reserve and NASA Astronaut Has a member of the Army positively changed your life? Now is your chance to thank them publicly with a shoutout via our Hooah Hotline and have it possibly appear on an upcoming episode of AUSA's Army Matters podcast! AUSA's Army Matters podcast can also be heard on Wreaths Across America Radio on Monday at 8 pm Eastern. You can find Wreaths Across America Radio on the iHeart Radio app, the Audacy app, and the TuneIn app. Search the word Wreath. Donate: If you are interested in supporting AUSA's educational programs, such as this podcast, please visit www.ausa.org/donate. Feedback: How are we doing? Email us at podcast@ausa.org. Disclaimer: The appearance of U.S. Department of Defense (DoD) visual information does not imply or constitute DoD endorsement. AUSA's Army Matters podcast primary purpose is to entertain. The podcast does not constitute advice or services. While guests are invited to listen, listeners please note that you are not being provided professional advice from the podcast or the guests. The views and opinions of our guests do not necessarily reflect the views of AUSA.
Où se trouvent les Chérubins ? Se font-ils face ? Si l'un des Chérubins représente le peuple juif, et l'autre Hachem, que signifie alors le fait qu'ils se regardent ? Quelle est la nature du lien le plus élevé entre D.ieu et le 'Am Israël ?
Le 21 janvier est la journée Internationale des câlins. Ce sera l'occasion de rappeler ce qu'ils apportent aux enfants et aux jeunes. Avec Sylvie Anzalone, porte-parole de l'ONE. Merci pour votre écoute Tendances Première, c'est également en direct tous les jours de la semaine de 10h à 11h30 sur www.rtbf.be/lapremiere Retrouvez tous les épisodes de Tendances Première sur notre plateforme Auvio.be : https://auvio.rtbf.be/emission/11090 Et si vous avez apprécié ce podcast, n'hésitez pas à nous donner des étoiles ou des commentaires, cela nous aide à le faire connaître plus largement.
Ce soir dans L'Église d'Aujourd'hui, nous découvrons Chérubins, une nouvelle plateforme dédiée aux contes pour enfants. Son co-fondateur, Thibaut Despierres, nous présente cette application accessible sur smartphone, tablette et ordinateur, qui propose des histoires audio de qualité, fondées sur des valeurs chrétiennes. Les premières séries se concentrent sur trois thématiques : les Vies Saintes, les Épopées de l'Ancien Testament, et les Témoins de l'Histoire. Chérubins sera officiellement lancée le 1er décembre, en début de l'Avent. L'Eglise d'aujourd'hui est une émission qui invite à découvrir les mille visages des chrétiens de nos jours. L'Eglise d'aujourd'hui est présentée par Matteo Ghisalberti et proposée par le diocèse de Monaco. Elle est diffusée sur RMC le samedi à minuit après l'After Foot (20h-minuit).
Gravity dominates every moment of our experience here on Earth. We may take it for granted, but NASA astronaut Kate Rubins assuredly does not. She knows firsthand the fun and challenges of living in microgravity. During her time in space, Rubins conducted important experiments so that someday humans can handle even longer missions — like heading to Mars.
Astronaut and molecular biologist Dr. Kate Rubins shares her groundbreaking work on the International Space Station, from being the first to conduct DNA sequencing in space to advancing biotechnology in a unique and challenging environment. Dr. Rubins explains how space affects biological processes, the tools being developed to study these effects, and how these advancements could revolutionize industries on Earth. Her insights into the future of space travel and exploration, including the potential for sustainable life on Mars, offer a glimpse into the exciting intersection of biology and space science. Grow Everything brings the bioeconomy to life. Hosts Karl Schmieder and Erum Azeez Khan share stories and interview the leaders and influencers changing the world by growing everything. Biology is the oldest technology. And it can be engineered. What are we growing? Learn more at www.messaginglab.com/groweverything Chapters: 00:00:00 – The Next Giant Leap: Returning to the Moon After 50 Years 00:00:18 – Adventures & Anecdotes: Nova Scotia to Space Conversations 00:01:57 – Aliens & Engineering: A Dive Into the Sci-Fi Universe and Genetics 00:05:25 – Space Dreams: Why Exploration Fuels Human Curiosity 00:08:27 – Meet Dr. Kate Rubins: The Astronaut Changing the Game in Space Biology 00:11:43 – The Real Lab in Space: Overcoming Challenges and Pushing Innovations 00:18:45 – First PCR in Space: How Dr. Kate Rubins Made History with DNA Sequencing in Microgravity 00:23:33 – Building for Mars: Synthetic Biology's Role in the Red Planet Mission 00:25:03 – Closed Loop Systems: The Future of Space Sustainability 00:27:41 – Is There Life on Mars? Exploring the Possibilities 00:30:01 – Human Engineering: Could We Modify Ourselves for Space Survival? 00:31:13 – Earth and Space: Integrating Biotech for Space and Beyond 00:34:26 – The Moon Beckons: Our Next Step in Human Exploration 00:35:37 – Unlocking Water on the Moon: What It Means for Future Missions 00:38:56 – Looking Forward: Space, Technology, and the Future of Humanity Topics Covered: biotech, bioengineering, precision fermentation, epigenetics, optogenetics, light, biosolutions, cellular control, photomolecular biology Episode Links: Kate Rubins NASA The Next 500 Years by Chris Mason Have a question or comment? Message us here: Text or Call (804) 505-5553 Instagram / TikTok / Twitter / LinkedIn / Youtube / GrowEverything website Email: groweverything@messaginglab.com Music by: Nihilore Production by: Amplafy Media
Navigating Entrepreneurship and Facing Reality with Adam Rubins In this episode of No Bullsh*t Talks, join Adam Rubins, former head of EMEA marketing at Disney and founder of Now Next Why, for an enlightening conversation. Adam shares his intricate journey of selling his business, addressing the challenges and emotional toll it takes on an entrepreneur. The discussion emphasises the importance of understanding the 'why' in business, misconceptions around agency growth, and the necessity of mental health awareness for leaders. Furthermore, the episode prepares listeners to face current challenging times by being realistic, stressing that acknowledging the truth empowers us to tackle upcoming challenges. Key insights include building a values-led organization, the significance of mentoring, and the importance of asking for help. Timestamps 00:00 Introduction to No Bullsh*t Talks 01:12 Meet Adam Rubins 02:19 The Disney Experience: Lessons and Challenges 05:15 Transition to Entrepreneurship 09:20 The Reality of Running a Business 14:24 The Obsession with Growth and Scaling 16:14 Challenges in the Agency World 28:00 The Importance of Collaboration and Ecosystems 36:18 Navigating Business Acquisitions 38:47 The Emotional Rollercoaster of Selling a Business 41:25 The Reality of Post-Sale Life 44:18 The Importance of Mental Health for Business Leaders 50:31 Generational Shifts in Leadership 01:00:42 Advice for Current and Aspiring Agency Owners 01:11:16 Final Thoughts and Optimism for the Future
Och vi bara fortsätter … Johan ska på konferens med Kurt. Har vi fått nog av journalister som chefer? Lena Andersson får kalla handen och Handels rektor Lars Strannegård varma famnen. Vissa är aktuella för alla jobb. Thomas Gür sätter sitt vassa finger på att det faktiskt måste finnas något slags gemenskap för att det ska bli ett samhälle. Lite reklam för att ni ska prenumerera på oss, så att ni slipper reklam. Hamas tycker att det är toppen med alla döda, eftersom det försätter Israel i en taskig sits. Wall Street Journal har Hamasledningens mejl. Vi börjar tala lite om vår vistelse på hotell Alhambra Palace, men fransar snart ut i hotell i allmänhet och boken som Johan hela tiden tjatar om att han ska skriva. Vi minns våra vistelser på Soho House, särskilt på South beach i Miami, där en bartender knackade på varje eftermiddag och blandade en drink. Som hos mormor? Kanske om din mormor hette Vanderbildt, eller Onassis. Susanna läser Rick Rubin och gillar det, trots alla klyschor. Grejen med klyschor är ju att de ofta är sanna, om än slitna. Rubin ser för övrigt ut precis som Fröding på sjukbädden. Hur påverkas vi egentligen av att leva i fantasivärldar hela tiden? Jamen, är det så nytt, undrar Susanna, och börjar tala om grottmänniskor. Sådana som hennes morbror, Lage Lindell, ville måla för. Vi är rörande överens om att folk inte har en aning om vad de vill ha. det borde även tidningar lära sig. Och så det eländiga EU-valet och särskilt dess dåliga förlorare: Jimmie, förstås, men också den auktoritära liberalen Macron. Det börjar bli dags att skriva hans historia. Become a member at https://plus.acast.com/s/hakeliuspopova. Hosted on Acast. See acast.com/privacy for more information.
Houston, we definitely do NOT have a problem…with interviewing Dr. Kate Rubins, NASA astronaut. Dr. Rubins is a virologist who has spent over 300 days in space, performing experiments aboard the International Space Station, where she was the first person to sequence DNA in space. We caught up with Dr. Rubins at the Neutral Buoyancy Lab in Houston, where she discusses what it felt like the first time she saw the earth from space, some of the difficulties in performing research without gravity, how to study the microbiome of the ISS, how the international inhabitants of the ISS communicate with each other, and the spur-of-the-moment event that led to her becoming an astronaut. This episode was supported by Cestodium, a new weight-loss program.* Participants: Karl Klose, Ph.D. (UTSA) Kate Rubins, Ph.D. (NASA) Janakiram Seshu, Ph.D. (UTSA) Jesus Romo, Ph.D. (UTSA) *The recorded ads heard on microTalk are for parody purposes only, there are no actual products for sale.
Na, ist euch bei dem Intro auch ein Schauer über den Rücken gelaufen? Perfekt, für die aktuelle Folge durften wir nämlich Tim Grobe interviewen. Er ist Schauspieler und Synchronsprecher. Bei den Drei ??? spricht er Dr. Shaitan aus „Der dunkle Taipan“. Doch, wie ist er überhaupt zu der Rolle gekommen? Echte Fans werden hier Parallelen zu Fall 25. „und die singende Schlange“ erkennen. Dort wurde die Rolle gesprochen von Lutz Mackensy. Aber auch in bspw. „Geisterbucht“ (Mr. Smith), „Schattenwelt“ (Lemuel Garvine) und in „Die Yacht des Verrats“ (Thabani) hat Tim Auftritte. Wie passend also, dass wir dieses Interview bei der Hörspielkönigin im Studio aufnehmen durften. Apropos „Der dunke Taipan“… Hier erzählt uns Tim noch einige Anekdoten rund um die Tour! So durfte er u.a. mit Matthias Keller und Kathrin Fröhlich zusammenarbeiten. Tim ist auch ein wahres Multitalent auf der Bühne! Warum? Na, er durfte noch ein Paar mehr Rollen während den Aufführungen sprechen. Mehr dazu im Interview. Darüber hinaus hat Tim auch das Hörbuch zum 5. Fall „und der Fluch des Rubins“ eingelesen. Wie gestalteten sich hierzu wohl die Aufnahmen? Auch seine weiteren Sprecherrollen werden hier beleuchtet (z.B. in Naruto, oder als Stimme von Mike Tyson). Besucht uns auch gerne auf www.instagram.com/diggytalk/ für weiteren Content „rund um die digitale Welt“. Copyright 2023 Diggytalk – Diggytalk ist eine eingetragene Marke von Dominik Grote - Hosted on Acast. See acast.com/privacy for more information.
Jump in with Carlos Juico and Gavin Ruta on episode 139 of Jumpers Jump. This episode we discuss: New guest Django, Gavin's hallucinations while being sick, Having split personality disorder, The greatest performer of our generation, Women claims person isn't real on flight, Michael Rubin hosts all white party for elites, Influencer kills 12 people while driving at high speeds, Getting more then one tattoo, Do you step into changing people's timeline, Growing up in different environments, having direction for your goals, abnormal lessons, Damar Hamlin incident and much more! Follow the podcast: @JumpersPodcast Follow Carlos: @CarlosJuico Follow Gavin: @GavinRutaa Check out the podcast on YouTube: https://bit.ly/JumpersJumpYT Thanks to our Sponsors: Get 20% Off and Free Shipping with the code JUMPERS at Manscaped.com Take advantage of this special financing offer at https://NetSuite.com/JUMPERS Learn more about your ad choices. Visit podcastchoices.com/adchoices
The boys are joined by Mitch Tischler of Monumental Sports Network to hear his thoughts on Roulliers retirement, Josh harris at Rubins, and Dotson poised for a breakout year? Then they're joined by Rick Doc Wa;ler before answering Fan questions!!Support the show
Unser Sales Manager
Are you looking to sell your creative agency? Or are you looking to acquire one? You need to strap yourself in because it's gonna get bumpy. Unless… You take on board the wisdom contained in the next 60 minutes of this interview with Adam Rubins, who's experienced the highs and lows, the mental stresses and strains of the agency sales process and lived to tell the tale. Not only that but he's built a new business around helping agency owners avoid the common pitfalls of selling or acquiring a creative agency. Plus Hollywood The future of TV streaming Mental welfare at work Wearing masks The value of new business to your agency's price tag How to make your agency more attractive to a buyer Agency employment laws US and UK agencies - divided by a common employment language And sharing an office with Sean Connery! Learn more about your ad choices. Visit megaphone.fm/adchoices
During his time in college, Dan Rubins wanted to get involved in a music program, but he found that many of them were generating competition among students, and he was seeking connection and community. And so, Hear Your Song was born: a non-profit serving kids ages 6 to 18 with serious illnesses and complex health needs to make their voices heard through collaborative songwriting. Based in NYC but available virtually nationwide, Hear Your Song helps children with a wide range of physical and mental-health-related diagnoses share their medical journey or emotions through song. It is a child-driven process where children share what they want and choose how to define themselves. From silly songs to emotional songs, from explaining the journey of a particular illness to songs about inequality, the children Dan and his team have worked with never fail to impress with their creativity, maturity, and determination. This project helps children feel empowered and in control of what is usually an uncontrollable and disheartening situation. Today, Dan has joined me on this episode to share more about his work, how you can use music as a tool, and how you can get involved in this beautiful project. I highly recommend checking out their YouTube channel to listen to many beautiful and moving songs they've created alongside these brave children. Key Takeaways with Dan Rubins The power of music for personal development. How music is used as a learning and advocating tool. How children with a diagnosis are empowered to define themselves outside of their disease. The beautiful ways that Hear Your Song has provided an emotional outlet for children in difficult medical journeys. The connection and relationships that are being built with others who are on similar journeys. Why creating a safe and empowering environment for children to find their voice is so important. Show Notes: Get Full Access to the Show Notes by visiting: MatteasJoy.org/57. Rate & Review If you enjoyed today's episode of The Joy In The Journey, hit the subscribe button on Apple Podcasts, Spotify, Stitcher, or wherever you listen, so future episodes are automatically downloaded directly to your device. You can also help by providing an honest rating & review over on Apple Podcasts. Reviews go a long way in helping us build awareness so that we can impact even more people. THANK YOU!
In this Industry Talent special, we hear from Darren and Liz. Both shared a background in marketing and the agency world, before joining forces in 2019 to launch Conker, a company that finds talent for senior roles. They discuss qualifications, how consumer habits have changed, the demand for a new type of leader, and how their secret to success is something called a Partnership Charter. Hosted on Acast. See acast.com/privacy for more information.
Kaffeeklatsch vom Gebrauchtwarencenter - Drei Fragezeichen Hörspiel-Podcast
Endlich mal wieder Mathildakommando. Holt euch eure Kaffeetassen und macht euch bereit für einen Überraschungsbesuch in der Zentrale. Wie gut bekommt man einen Stockdegen und was hat es eigentlich mit dem August auf sich? Wie groß war denn nun der Scheck? Wer hat jetzt auch Bock auf Buchstabensuppe und wann geht eigentlich der nächste Flug nach Naknek. Viel Spaß mit der neuen Folge --- Send in a voice message: https://podcasters.spotify.com/pod/show/kaffee-klatsch/message
“At the start of the pandemic, my co-founder and I realized that it was the opportunity to really start expanding this mission to help kids with serious illnesses have the power and choice through a creative process that they so often don't get to have in their daily lives,” shares Dan Rubins. Dan is the co-founder and executive director of Hear Your Song, a non-profit organization empowering kids with serious health conditions to make their voices heard through collaborative songwriting. Today, Dan joins host Tiffany Zehara to talk about how the organization got started and how it works to empower kids globally. Kids with severe illnesses and diseases lose a lot of agency over their own lives. They may spend a lot of time inpatient in hospitals and don't have as much opportunity to express themselves creatively as other kids. Hear Your Song gives kids an opportunity to be involved in every step of the song creation process from lyrics and vocals to beats and instruments. It also provides an online community so that kids and their families can feel more connected. Music can be incredibly healing, but the creation of music itself also provides an opportunity for empowerment. Tune into today's episode of Humanitarian Entrepreneur Podcast for a talk with special guest Dan Rubins to learn more about the work Hear Your Song is doing to improve the lives of children with serious and chronic illnesses. Quotes “At the start of the pandemic, my co-founder and I realized that it was the opportunity to really start expanding this mission to help kids with serious illnesses have the power and choice through a creative process that they so often don't get to have in their daily lives.” (3:28-3:46 | Dan) “It's really all about giving kids as many choices and as much control of every step of the process as we possibly can.” (7:10-7:18 | Dan) “We're hoping to this year take on some more multilingual partnerships as well, because that's something we're really excited about giving kids the opportunity to write songs in whatever language they feel most comfortable in.” (8:14-8:26 | Dan) “One of the things that's been wonderful on the community side has been really getting to know kids and families and bringing them together in ways that we would probably never have thought of if we were working purely in person.” (12:22-12:36 | Dan) “We also have what we call Cheer Your Song showcases where we have kids sharing their song virtually live with some of the volunteers who worked on their song. And we invite big audiences to watch that as well and put comments in the chat and sort of respond to kids songs and ask questions in real time. And those kinds of things have really been amazing for just making kids and families feel that they're part of a wider community. And a lot of kids listen to each other's songs, which is really cool. All the songs are up on our Youtube channel.” (13:05-13:40 | Dan) Connect with Dan Rubins: Website: hearyoursong.org Facebook: https://www.facebook.com/hearyoursonghys Instagram/Twitter/TikTok: @HearYourSongHYS Youtube: www.youtube.com/hearyoursong Spotify: https://open.spotify.com/artist/61cbxgJyTl1HTAkewAvTla?si=EUPInWLsTJ68E9h6-9kgpg To connect with Tiffany to solve problems or affect the kind of change you want: https://calendly.com/humanitarianentrepreneur/discovery-call Website: https://humanitarian-entrepreneur.com Podcast production and show notes provided by HiveCast.fm
En Tunisie, la communauté subsaharienne est souvent confrontée à des difficultés économiques. La garderie Les Chérubins à Tunis offre à la communauté la garde des enfants subsahariens à un prix symbolique. Une initiative solidaire mise en place par une femme migrante ivoirienne. De notre correspondante à Tunis, Derrière une porte sans pancartes, entre des maisons, les voix des enfants de 3 mois à 7 ans résonnent dans la garderie Les Chérubins, située au cœur du quartier populaire de Bhar Lazreg à Tunis. Isabelle Bessan, Ivoirienne installée en Tunisie, a lancé cette crèche en 2018, au départ dans un garage, puis dans ce local, pour accueillir les enfants de migrants subsahariens du quartier. « À la crèche, nous accueillons les enfants à partir de deux mois. On apprend à l'enfant à s'asseoir, faire les quatre pattes et les premiers pas pour pouvoir marcher. Et ensuite, à l'âge de deux ans et demi, nous donnons des bases à l'enfant, c'est-à-dire que nous lui apprenons à former les chiffres et les lettres » Cet enfant récite un sketch sur la thématique de l'intégration, chère à Isabelle pour sensibiliser les enfants face aux risques d'attaques racistes ou discriminatoires. « C'est vrai que ce sont des enfants, mais je me dis que parfois, ils tombent de haut, en entendant le mot "migrant", explique-t-elle. C'est en ce sens-là que j'ai composé ce petit sketch pour leur expliquer ce qu'était un migrant, ce qu'était l'intégration et le pays où ils vivent. » Une réalité migratoire, qu'Isabelle cache aussi aux enfants lorsqu'elle devient trop dure. Elle confie qu'une dizaine de ceux qui venaient à la crèche ont traversé clandestinement la mer Méditerranée avec leurs parents cette année. Deux d'entre eux sont morts dans un naufrage. Pour les familles qui restent en Tunisie, le contexte économique et social est de plus en plus difficile à vivre. La communauté compte entre 30 000 et 50 000 personnes. La majorité, sans papiers ni cartes de séjour, travaille comme main-d'œuvre non déclarée et faiblement rémunérée avec un minimum de droits en Tunisie. « Il y a beaucoup de Subsahariens maintenant en Tunisie, et c'est difficile de trouver du travail », confirme Jessica Cohebi, ivoirienne et mère d'un petit garçon qui fréquente la crèche. Le lieu n'a pas encore de statut juridique clair, donc Isabelle Bessan demande aux parents une somme symbolique pour payer le loyer, les charges. « Sur le plan financier, c'est un grand soulagement parce qu'on ne paye pas grand-chose en fait, on donne juste ce que l'on peut donner, explique Jessica. Et sur le plan social, affectif, il y a un soulagement aussi de savoir qu'ici, il est un peu en famille. » La garderie repose beaucoup sur la solidarité des trois femmes qui y travaillent, mais les besoins restent importants. Isabelle lance fréquemment des appels aux dons sur les réseaux sociaux pour des denrées alimentaires et des jouets.
After a long and grueling day at the studio, Rick Rubin finally comes home to find his young son playing with his toy truck in the driveway. Trying to act normal, Rick tells him to come in and have dinner with him, but his son has other plans.
The San Antonio Philharmonic launched in August with a mission to open young minds and spirits for life. Jeremy Brimhall, Director of Education and Community Engagement at the San Antonio Philharmonic, and Peter Rubins, Vice President of the board of the San Antonio Philharmonic, join us to share about opportunities in the 2022–23 season for students and young children—and their families—to enjoy beautiful classical music at Young People's Concerts, meet musicians at branches of the San Antonio Public Library, and more.
THE EMBC NETWORK featuring: ihealthradio and worldwide podcasts
Dan Rubins is the Co-Founder and Executive Director of Hear Your Song, a 501(c)(3) organization that empowers children and teens with serious illnesses and complex health needs to make their voices heard through collaborative songwriting. After launching Hear Your Song as an undergraduate organization while a sophomore at Yale University, Dan led the organization's national and virtual expansion in 2020 in response to the COVID-19 pandemic. Since then, Hear Your Song has helped over 250 kids ages 6-18 in 27 states and 6 countries write their own songs with the support of hundreds of volunteer musicians around the world. A musical theater and opera composer himself, Dan also holds an MA in Elementary Inclusive Education from Teachers College, Columbia University and an MA in Shakespeare Studies from King's College London/Shakespeare's Globe. About Hear Your Song, Inc.: Hear Your Song, Inc. is a 501(c)(3) nonprofit organization based in New York City that empowers children and teens with serious illnesses and complex health needs to make their voices heard through collaborative songwriting. Hear Your Song provides power and choice — and a microphone — to young people with a wide range of diagnoses, both mental and physical health conditions, who are so often deprived of both power and choice in living their day-to-day life and in managing their health care journeys. Hear Your Song gives kids who are often cast to the margins the chance to define themselves beyond their diagnoses, embraced and validated by a community that responds with caring, collaborative creativity to help each young person tell their stories through song. Hear Your Song first launched as an undergraduate organization at Yale University in May 2014, piloting partnerships at Yale-New Haven Children's Hospital and Elizabeth Seton Children's in Yonkers, NY. Six years later, at the start of the pandemic, we realized that it was more important than ever to allow children with serious illnesses to share their stories and find a community of support while most isolated and at risk. So, in March 2020, Hear Your Song began a national expansion, offering virtual sessions to kids wherever they are, whether they're staying at a hospital or receiving treatment or recovering at home. Hear Your Song has now supported over 200 children ages 6-18 in writing their own songs. At the same time, we've grown six undergraduate-led, campus-based chapters; engaged hundreds of volunteer musicians around the globe; and built over a dozen partnerships with children's hospitals, specialized schools/camps, and other nonprofits, including the Montefiore Medical Center, Newton-Wellesley Hospital, Double H Ranch Camp, The ELM Project, and the Crohn's and Colitis Foundation. We focus our outreach to new communities on populations often overlooked in pediatric arts programming, especially kids with diagnoses that disproportionately impact communities of color.
THE EMBC NETWORK featuring: ihealthradio and worldwide podcasts
Dan Rubins is the Co-Founder and Executive Director of Hear Your Song, a 501(c)(3) organization that empowers children and teens with serious illnesses and complex health needs to make their voices heard through collaborative songwriting. After launching Hear Your Song as an undergraduate organization while a sophomore at Yale University, Dan led the organization's national and virtual expansion in 2020 in response to the COVID-19 pandemic. Since then, Hear Your Song has helped over 250 kids ages 6-18 in 27 states and 6 countries write their own songs with the support of hundreds of volunteer musicians around the world. A musical theater and opera composer himself, Dan also holds an MA in Elementary Inclusive Education from Teachers College, Columbia University and an MA in Shakespeare Studies from King's College London/Shakespeare's Globe. About Hear Your Song, Inc.: Hear Your Song, Inc. is a 501(c)(3) nonprofit organization based in New York City that empowers children and teens with serious illnesses and complex health needs to make their voices heard through collaborative songwriting. Hear Your Song provides power and choice — and a microphone — to young people with a wide range of diagnoses, both mental and physical health conditions, who are so often deprived of both power and choice in living their day-to-day life and in managing their health care journeys. Hear Your Song gives kids who are often cast to the margins the chance to define themselves beyond their diagnoses, embraced and validated by a community that responds with caring, collaborative creativity to help each young person tell their stories through song. Hear Your Song first launched as an undergraduate organization at Yale University in May 2014, piloting partnerships at Yale-New Haven Children's Hospital and Elizabeth Seton Children's in Yonkers, NY. Six years later, at the start of the pandemic, we realized that it was more important than ever to allow children with serious illnesses to share their stories and find a community of support while most isolated and at risk. So, in March 2020, Hear Your Song began a national expansion, offering virtual sessions to kids wherever they are, whether they're staying at a hospital or receiving treatment or recovering at home. Hear Your Song has now supported over 200 children ages 6-18 in writing their own songs. At the same time, we've grown six undergraduate-led, campus-based chapters; engaged hundreds of volunteer musicians around the globe; and built over a dozen partnerships with children's hospitals, specialized schools/camps, and other nonprofits, including the Montefiore Medical Center, Newton-Wellesley Hospital, Double H Ranch Camp, The ELM Project, and the Crohn's and Colitis Foundation. We focus our outreach to new communities on populations often overlooked in pediatric arts programming, especially kids with diagnoses that disproportionately impact communities of color.
We wrote a song! You'll meet Catherine, the new co-host of the podcast! Dan Rubins and Jake Gluckman from Hear Your Song, Inc. join us and help write a song about Richard's recent hospital stay that features Ivy the elephant and her love of coffee. Hear Your Song, Inc. is a 501(c)(3) nonprofit organization based in New York City that empowers children and teens with serious illnesses and complex health needs to make their voices heard through collaborative songwriting. Hear Your Song provides power and choice — and a microphone — to young people with a wide range of diagnoses, both mental and physical health conditions, who are so often deprived of both power and choice in living their day-to-day life and in managing their health care journeys. Hear Your Song gives kids who are often cast to the margins the chance to define themselves beyond their diagnoses, embraced and validated by a community that responds with caring, collaborative creativity to help each young person tell their stories through song. Hear Your Song first launched as an undergraduate organization at Yale University in May 2014, piloting partnerships at Yale-New Haven Children's Hospital and Elizabeth Seton Children's in Yonkers, NY. Six years later, at the start of the pandemic, we realized that it was more important than ever to allow children with serious illnesses to share their stories and find a community of support while most isolated and at risk. So, in March 2020, Hear Your Song began a national expansion, offering virtual sessions to kids wherever they are, whether they're staying at a hospital or receiving treatment or recovering at home. Hear Your Song has now supported over 200 children ages 6-18 in writing their own songs. At the same time, we've grown six undergraduate-led, campus-based chapters; engaged hundreds of volunteer musicians around the globe; and built over a dozen partnerships with children's hospitals, specialized schools/camps, and other nonprofits, including the Montefiore Medical Center, Newton-Wellesley Hospital, Double H Ranch Camp, The ELM Project, and the Crohn's and Colitis Foundation. We focus our outreach to new communities on populations often overlooked in pediatric arts programming, especially kids with diagnoses that disproportionately impact communities of color. Hear Your Song website Watch the podcast on YouTube --- Support this podcast: https://podcasters.spotify.com/pod/show/artsforthehealthofit/support
Hear Your Song is a non-profit organisation that empowers children and teens with serious illnesses and complex health needs to make their voices heard through collaborative songwriting. Kids living with significant health challenges need the chance to show the world — and sometimes to hear for themselves, too — that they are more than their diagnoses.In live, collaborative songwriting sessions, Hear Your Song volunteers work with children and teens to guide them through the process of writing their own song lyrics. Using the kid songwriter's ideas for musical style, melody, instrumentation, and tempo, volunteer composers and musicians then set those words to music and record the song to be heard, celebrated, and shared.Hear Your Song partners with pediatric hospitals, camps, schools, and other nonprofit programs that serve kids experiencing serious illnesses and complex health needs. Hear Your Song's volunteers collaborate with kids through campus-based chapters and at the organization's national level. All of our programs are available to our families and partner organizations free of charge. See acast.com/privacy for privacy and opt-out information. Become at member at: https://plus.acast.com/s/tobyonathursday.
Der Bobcast und der Fluch des RubinsEs herrscht Verwirrung bei der Besprechung der Die drei ??? Kult-Folge: was um alles in der Welt ist „sülzieren“? Warum machen die „Fünf Freunde“ nicht mit? Und wieso hilft der Zwerg nicht…?! Bei den Aufnahmen für Die drei ??? und der Fluch des Rubins“ ist Jens Wawrczeck in eine Sprecherin verliebt, in weiteren Rollen brillieren Joachim Wolff, Reinhilt Schneider und Gottfried Kramer, der als Mister Rhandur einen Mord begeht (oder doch nicht?). Buchautor André Marx berichtet von der Entwicklung seiner Rubin-Fortsetzung „Feuriges Auge“ und verrät dabei auch, warum ihn die Schwarzbärte nerven. Heikedine Körting hat ihren Sprecher:innen vor 40 Jahren 150 Tannenbäume geschenkt - im Bobcast klären wir, wo noch alle Nadeln dran sind. Und Andreas Fröhlich outet sich im Gespräch mit Kai Schwind: er latscht zu gern über fränkische Felder! Aber warum? Gäste in dieser Podcast-Folge: Heikedine Körting und André Marx „Haschimitenfürst – Der Bobcast“ ist ein Podcast von EUROPA, a division of Sony Music Entertainment Germany GmbHIdee: Andreas Fröhlich/ Regie & Konzeption: Ralf Podszus/ Moderation: Kai Schwind und Andreas Fröhlich/ Titelmusik: Jan-Friedrich Conrad/ Redaktion: Jens Nimmerrichter/Produktion: Carina Schwarz/ Management & Koordination: Nina Schulze PellengahrRedaktion Sony: Maike Müller/ Covermotiv: Aiga Rasch (Illustrationen), Tom Presting (Gestaltung), Christian Hartman, Haakon Dueland (Fotos)/ Eine Produktion von Podever See acast.com/privacy for privacy and opt-out information.
Join us for the latest episode of The Hamilton Review Podcast! In this conversation, Dr. Bob has a heartwarming discussion with Dan Rubins, Co-Founder and Executive Director of Hear Your Song, a 501(c)(3) organization that empowers children and teens with serious illnesses and complex health needs to make their voices heard through collaborative songwriting. Dan talks with Dr. Bob about how he developed Hear Your Song as an undergraduate at Yale University and the wonderful work that the organization is doing to benefit children. Don't miss this special episode friends and we thank you for listening! Dan Rubins is the Co-Founder and Executive Director of Hear Your Song, a 501(c)(3) organization that empowers children and teens with serious illnesses and complex health needs to make their voices heard through collaborative songwriting. After launching Hear Your Song as an undergraduate organization while a sophomore at Yale University, Dan led the organization's national and virtual expansion in 2020 in response to the COVID-19 pandemic. Since then, Hear Your Song has helped over 200 kids ages 6-18 in 27 states write their own songs with the support of hundreds of volunteer musicians around the world. A musical theater and opera composer himself, Dan also holds an MA in Elementary Inclusive Education from Teachers College, Columbia University and an MA in Shakespeare Studies from King's College London/Shakespeare's Globe. Dan lives in New York City, where he was formerly a 4th and 5th grade teacher. How to contact Dan Rubins: SOCIAL MEDIA LINKS www.youtube.com/hearyoursong Facebook/Instagram/TikTok/Twitter handles: @HearYourSongHYS Website: www.hearyoursong.org How to contact Dr. Bob: Dr. Bob on YouTube: https://www.youtube.com/channel/UChztMVtPCLJkiXvv7H5tpDQ Dr. Bob on Instagram: https://www.instagram.com/drroberthamilton/ Dr. Bob on Facebook: https://www.facebook.com/bob.hamilton.1656 Dr. Bob's Seven Secrets Of The Newborn website: https://7secretsofthenewborn.com/ Dr. Bob's website: https://roberthamiltonmd.com/ Pacific Ocean Pediatrics: http://www.pacificoceanpediatrics.com/
August braucht die Hilfe der drei ???. Er hat von seinem verstorbenen Großonkel den wertvollen Rubin "Feuriges Auge" geerbt. Doch auf dem Stein liegt ein Fluch: Jeder seiner Besitzer ist dem Tode geweiht. Um den Rubin zu finden und den Fluch zu brechen, müssen Justus, Peter und Bob eine ganze Kette rätselhafter Zusammenhänge lösen. Und sie sind nicht die Einzigen, die hinter dem Edelstein her sind. Sollen sie die Drohungen des geheimnisvollen Mr Rhandur ernst nehmen? Eine spannende Schatzsuche beginnt ... Julia Schütze #Whisper2Me www.juliaschuetze.at/whisper2me www.kosmos.de
August braucht die Hilfe der drei ???. Er hat von seinem verstorbenen Großonkel den wertvollen Rubin "Feuriges Auge" geerbt. Doch auf dem Stein liegt ein Fluch: Jeder seiner Besitzer ist dem Tode geweiht. Um den Rubin zu finden und den Fluch zu brechen, müssen Justus, Peter und Bob eine ganze Kette rätselhafter Zusammenhänge lösen. Und sie sind nicht die Einzigen, die hinter dem Edelstein her sind. Sollen sie die Drohungen des geheimnisvollen Mr Rhandur ernst nehmen? Eine spannende Schatzsuche beginnt ... Julia Schütze #Whisper2Me www.juliaschuetze.at/whisper2me www.kosmos.de
Terouma: Quelle photo mettre dans sa chambre à coucher ? le puissant message des chérubins bébés placés sur l'arche sainte ! Inscrivez vous à notre Whatsapp pour les cours et replay: https://chabad77.org/etorah-cours-rav-yossi-amar Pour les autres cours ou replay téléchargez l'application ETORAH. Nouveau! Podcast ETORAH sur SPOTIFY - APPLE - GOOGLE ! Ecouter & partagez ! MERCI
In this episode of The Hockey Writers Maple Leafs Lounge, our Toronto Maple Leafs writing crew members Peter Baracchini and Alex Hobson get together to discuss the 'gong show' against the Winnipeg Jets where Jason Spezza was handed a six-game suspension, the injury to Rasmus Sandin, the Leafs' quick starts and horrible finishes recently, Kristians Rubins impact on the defence, overall lack of officiating against the Jets, Nick Ritchie and Ondrej Kase's performance in the top six and more. Our Maple Leafs Lounge crew are great writers: Kevin Armstrong - Peter Baracchini - Alex Hobson And, make sure to check out all of our great Maple Leafs content Follow The Hockey Writers: Twitter - Instagram - Facebook Sign up for the "Morning Skate" newsletter Join us in the Hockey Lounge on Discord to talk Maple Leafs and all things hockey Graphics by Vince Richard
Josh & Jeanne Rubin, who you might know as @realfoodganstas on Instagram, are some of the true OG's in the metabolic health world. These two have paved the way in not only thyroid & metabolic health, but in how to create a sustainable health journey that you truly enjoy. The Rubins have 20 years of clinical work under their belt and have pursued education through a wide range of programs, including the CHECK institute (Paul Chek) & the RCP institute (Morley Robbins). While the Rubins encourage a food-first approach to healing, they emphasize that it can't stop there. We HAVE to address the mind. We HAVE to address the nervous system. "The biology of the individual cannot be separated from the social and psychological aspects of their life." Our mindset shapes our biology. We can be checking boxes all day long, but we can get lost in the process and create more chaos in the stress of "doing all the right things." Has your pursuit of health become unhealthy? Find out in our season 2 finale! Join us as we sit down & discuss the following: Josh & Jeanne's health philosophy: it always comes back to the cell Introduction to the nervous system How our nervous system impacts how we see the world How our nervous system gets shaped as a child How developmental trauma occurs How trauma impacts our healing The importance of "play" in reconnecting to our bodies Getting rid of the chaos in our lives Preserving the innocence & play of our children The emotional component of dis-ease The importance of simplicity on a health journey Where to start on your journey to avoid getting overwhelmed *Not medical advice. This podcast & episode are for inspirational + educational purposes only* Where to find the Rubins: Josh & Jeanne's Instagram Josh & Jeanne's Website Where to find us: Kori's Instagram Fallon's Instagram Restore your metabolism: Freely Rooted Fallon's Table Our FREE downloads: Restore Your Metabolism: Free 5 Step Guide Metabolic Foods Guide
We're all trying our best to figure out how social media and politics work in 2021 and yet, it seems impossible when people, apps, and products continually get canceled. But don't worry because Mock and Daisy help clear things up this week as they discuss all things involving big tech and politics with special guest Dave Rubin from The Rubin Report. During their interview, they talk about everything from the flaws in big tech, who could be running in 2024, Rubins history, and oh Clyde (his rescue dog). Please visit our great sponsors:My Pillowhttps://www.mypillow.com/chicksNow get BOGO Giza Dream Sheets with promo code CHICKS. Genucelhttps://lovegenucel.com/chicksNow until Christmas, get 60% off Most Popular Genucel Package: lovegenucel.com/CHICKSThe Association of Mature American Citizenshttps://amac.us/chicksThe benefits of membership are great, but the cause is even greater.Acre Goldhttps://getacregold.com/chicksVisit GetAcreGold.com/CHICKS and start investing in physical Gold today!Calibrate Healthhttps://joincalibrate.comGet $50 off the 1-year metabolic reset with code CHICKS.Ruff Greenshttps://ruffgreens.com/chicksRuff Greens, they make any pet food….better. Get yours today!Omaha Steakshttps://omahasteaks.comUse keyword CHICKS for the Perfect Gift Package and 8 FREE burgers. Cozy Earthhttps://CozyEarth.comEnter promo code CHICKS and save 35%.
The Athletic Maple Leafs Reporter Joshua Kloke joins Game Play to tee up tonight's game between the Toronto Maple Leafs and the Columbus Blue Jackets. Kloke discusses how Toronto will handle the slew of injuries they are dealing with, what fans should expect from recent call ups Kristians Rubins, Alex Steeves, and Alex Biega, when Josh Ho-Sang might get a shot with the big club, and more.
This episode we take a little look into the world of recruitment: What are businesses looking for when hiring C-Suite executives and what are the C-Suite executives looking for when seeking new opportunities, has this changed since the global pandemic has affected the traditional working models? We asked Liz Jones and Daren Rubins, Founders of Conker, a specialist executive search company that specialises in the media and marcom industries to help us out with the answers! ABOUT CONKER Conker is an Executive Search business that finds transformational talent for senior roles in agencies, media owners, digital platforms and client brands. Our remit stretches across commercial, strategic, technical, operational and specialist areas - and often in combination. Learn more about Conker https://conkerwithus.com ABOUT LIZ JONES Liz is an experienced commercial leader who understands how talent drives better business outcomes. She is a strong networker and collaborator; passionate about contributing to positive change and business results through diversity, inclusion and well-being. Alongside her most recent role at Dentsu Aegis Network (CEO of B2B), she was also the Executive Sponsor for Diversity and Inclusion and together with HR built the first set of complete diversity metrics for business to inform 2020 goals focused on gender, agile working and ethnicity. Liz led the successful roll-out of unconscious bias training to the leadership team and she also launched the first LGBT+ resource group and gender equality network at Dentsu Aegis Network. She has always enjoyed building diverse, high performing teams. At PSI over eight years she hired over 20 individuals from different backgrounds and markets and in doing so doubled the revenue and profit of the business. A proud WACL member, she drove the recent initiative with LinkedIn using data for the first time to look at gender parity in the Media, Marketing and Communications industry. She has always mentored formally and informally via NABS, WACL, IPA and the wider industry. She is currently on the WACL Exec (Communications and Voice) and was a main committee member of the IAA for many years. Liz is also a NABs Stronger than Summer committee member and mentors for most of the initiatives in the media industry. ABOUT DAREN RUBINS Daren's entire career has been built around helping companies and people to achieve their potential. Across 30 incredibly successful years in advertising, Daren has been responsible for hiring and/or leading some of today's most successful names in the marketing industry. His belief is that motivated talent, in the right environment, can achieve anything. In his 10 years as MD and then CEO of PHD in the UK, the agency experienced unprecedented success. Highlights included; Pitching and winning Cadbury, Kraft, Expedia, Dyson, Twitter, Confused.com, Purple Bricks, Viacom, British Heart Foundation, P&O Cruises, VW Group, Magners and many more. Winning eight individual Media Agency of the Year accolades, seven Grand Prix awards and two Cannes Lions Gold awards. Achieving four consecutive top 100 Best Companies Awards, including 25th, 18th and 16th positions. Daren also co-created the industry-accreditation Advanced Media Certificate, for which he received an IPA Fellowship and has spoken on numerous platforms about talent and diversity, including the 3% Conference, WACL, Token Man, Omniwomen, Bloom and HeForShe. In 2017, Daren became CEO at The Lighthouse Company, where he placed 18 business leaders in as many months. In his spare time, Daren mentors individuals at all levels and also the runs the Marketing Advisory Group for the JW3 Community Centre.
Becca Teich (@w0rmwrm) returns to the pod to join Gemma and Phoebe as they deep dive into queer theorist Gayle Rubin's opus to think though sex work, kink, and heterosexuality as a cultural normalizing device in Vanderpump Rules. Bonus Feature: A brief conversation about Camille Donatacci's Playboy's Playmates Exposed photoshoot feat. a Miata. Readings: Gayle Rubin, Deviations: https://read.dukeupress.edu/books/book/1560/DeviationsA-Gayle-Rubin-Reader "Contagion, Brenda Iijima Interviews Rebecca Teich" Follow Money Can't Buy You Class on Instagram @moneycantbuyyouclass_pod & @sad_porous_grad on twitter
Tidligere i år og 2020 udkom der fire nye album, Johnny Cash, Forever Words Expanded med ny musik til gamle digte og tekster af Johnny Cash. POVs musikredaktør Jan Eriksen blev så begejstret for disse album, at han kontaktede en af Danmarks største Johnny Cash-entusiaster. Dennis Greis Lydom, der holder foredrag og optræder med Cash-musik. Resultatet af denne podcast, hvor de tager på vandring ned gennem de seks American Recordings album, der udkom i den sene del af Cash's karriere. De to sidste udkom efter Johnny Cash's død i 2003. Efter musikpodcasten Mediano Music er blevet en del af POV hedder det fremover Mediano Music POVcast. I 1969 udgav Johnny Cash albummet At San Quentin, der blev det mest solgte album det år - i hele verden. 20 år senere var Johnny Cash på mange måder rangeret ud på et sidespor i sin karriere og til dels også menneskeligt. Godt nok havde han fundet sin gud og var kommet clean ud af sin misbrugsbehandling i 1983, men det kneb indimellem med at holde sig på stien. Efter sin anden fyring, denne gang fra selskabet Mercury, sang Cash Bob Dylans "It Ain't Me Babe" under en hyldestkoncert til den gamle ven, som Cash bl.a. havde indspillet sammen med på Dylans Nashville Skyline. Da pladeselskabsmanden Rick Rubin overværede koncerten, opstod der en drøm i ham. En drøm om at give Johnny Cash den respekt og de sange, han som en af de tidlige rockpionerer havde fortjent. Den var Cash helt med på, faktisk havde han en liste med 100 sange, som han godt kunne tænke sig at indspille. Så de mødtes og begyndte at indspille i Rubins stue. Resten er musikhistorie. Nu også fortalt i denne POVcast: På grund af en fejl i en ellers helt nyindkøbt mikser er lyden lidt ulden, men kan sagtens høres. Det beklager vi. Og så skal det tilføjes at The Class of '55, 'de fire vise mænd', som Jan Eriksen nævner i podcasten, er Jerry Lee Lewis, Carl Perkins, Roy Orbison og Johnny Cash. Nogle gange kan de ske, at det løber af med nørder, når de får talt sig varme. American Recordings American II: Unchained American III: Solitary Man American IV: The Man Comes Around American V: A Hundred Highways American VI: Ain't No Grave
This is the first of a 3 part mini-series on mental health, with my guest Adam Rubins. At the age of 12yrs Adam Rubins was set upon and beaten up in his local synagogue. How is it that a chance event in someone's childhood can set up a chain reaction that ultimately leads to a lifetime of battling depression? Despite this childhood trauma, Adam Rubins has operated at a high level of performance and success in the high octane, high-stress world of entertainment and communications. Adam's career has spanned working as the marketing director of the Walt Disney Company and CEO of a successful marketing agency Way to Blue. Now committed to the cause of promoting mental health and wellbeing in the workplace, Adam's story is a brutally honest and revealing insight into the challenges faced by many people today in their working lives. Adam talks about the particular pressures of expectation that contributed to his depression and burnout. His advice to people who are struggling with mental health issues And his call on businesses to challenge the notion that shareholder value and employee wellbeing are at odds. Having spent much of my career in advertising and suffered from depression, this is a subject dear to my heart. This episode is supported by the Alliance of Independent Agencies, which are partnering with Turning The Tables on the mini-series. The Alliance is an organization that takes the mental health and wellbeing of people in communications agencies seriously. Their Wellbeing Action Group promotes the importance of creating safe environments and building people's resilience. They even have an annual Festival of Happiness as well as promoting training for Mental First Aid Champions. How good is that! This is the kind of support and agenda that every industry and community needs. I would particularly like to thank Graham Kemp, Clive Mishon, and everyone at the Alliance for supporting this mini-series. https://allindependentagencies.org/ You can contact Adam on LinkedIn References Mickel Therapy Share your experiences about this episode on the Turning the Tables podcast community page on Facebook. And on Instagram TurningtheTablespodcast Episode Credits Editor and sound engineer: Tim White email: showupnow@gmail.com Host: Simon Ratcliffe Music: Broken Elegance -Unconditionally River Meditation - Audioautix When I'm gone LiQWYD Scott Buckley - Life is AB5 Piano strings - Tim White
Folge 5 ist ein Jubiläum - wir sprechen nämlich dieses Mal über die Folge Nr. 5 (und der Fluch des Rubins) sowie dessen Fortsetzung in Folge 200 (Feuriges Auge). Letztere ist die Längste der drei Fragezeichen Folgen überhaupt und beschert Euch 15 Sonderminuten in dieser Folge. Und fünf Folgen Podcast sind doch ebenfalls ein kleines Jubiläum. Wir sprechen von zerkratzten Fensterscheiben, mit Tesa reparierten Kassetten und Michael "Morton" Knight. Viel Spaß beim reinhören. ??? ??? ??? ??? ??? ??? ??? ??? ??? Besprochene Folgen: #5 und der Fluch des Rubins (https://dreifragezeichen.de/produktwelt/details/und-der-fluch-des-rubins) #200 Feuriges Auge (https://dreifragezeichen.de/produktwelt/details/feuriges-auge) ??? ??? ??? ??? ??? ??? ??? ??? ??? Hinterlasst gern eine Bewertung, abonniert uns oder schreibt uns eine E-Mail, wenn Ihr Fragen, Wünsche oder Anregungen habt. Neue Folgen gibts immer am 3. des Monats. Kontakt: mail@machdenverstaerkeran.de
In dieser Folge irren Edgar Wallace seine Nachbarn durch den Londoner Nebel - ob vor Blindheit oder Alkohol ist nicht ganz klar. Aber zum Glück hat uns ein netter Herr in einem Kastenwagen aufgegabelt und vor einem Blindenheim abgesetzt. Dort hat uns ein augenscheinlich blinder Pfarrer freundlicherweise zum gemeinsamen Musikhören eingeladen: Dadada daaaaaaam! Viel Spaß mit unserer Besprechung zu "Die Toten Augen von London" - The Dark Eyes of London / Die toten Augen von London Buch - Simpsons: Moe Szyslak - Hui Buh - ??? Vampir im Internet - Point Whitmark - Ich habe Sie nicht erwartet, Mr. Bond - Nosferatu (Schatten) - Film Es: „Hello George“ - James Bond Intro - Joachim Wolff bei den drei ???: und der Fluch des Rubins / und der sprechende Totenkopf - Der Beisser James Bond - Beethovens 5. aus der Sicht eines Sportreporters Wir freuen uns über Eure Fragen, Anregungen und Kommentare!
Topics!1. Dog clothes2. Dishwashing gloves3. Shopping with headphones4. Chick Corea 1941-20215. The Weeknd at the Super Bowl6. Noah S/S 20217. Wet socks8. Sorel slippers9. Graco carseats10. Chuck Johnson's The Cinder Grove11. The difference between masters and publishing rights12. Public domain13. Trader Joe's tortilla chips14. French fries at Mexican restaurants15. Underrated/overrated/properly rated: Munchkins flavors16. Underrated/overrated/properly rated: Seinfeld characters17. Smoked beers (Pipeworks, Alarmist)18. Mary Wilson 1944-202119. Rubins and Salmieri's Dragons Love Tacos20. Valentine's Day