Podcasts about Taha

  • 846PODCASTS
  • 2,286EPISODES
  • 30mAVG DURATION
  • 5WEEKLY NEW EPISODES
  • Aug 14, 2026LATEST

POPULARITY

20192020202120222023202420252026

Categories



Best podcasts about Taha

Show all podcasts related to taha

Latest podcast episodes about Taha

Qur'an Conversations
S4 E23: The Thief of Joy (TaHa 131-133) | Quran Conversations Season 4

Qur'an Conversations

Play Episode Listen Later Aug 14, 2026 69:09


What happens when we spend so much time looking at what others have that we lose sight of what Allah has already given us?In this episode of Quran Conversations, Dalia Mogahed is joined by her teacher, Imam Mohamed Magid, to reflect on verses 131–133 of Surah TaHa.The conversation begins with Allah's warning not to gaze longingly at the beauty and luxuries others have been given in this world. What looks like success may simply be a test, and what Allah provides is ultimately better and more lasting. Imam Magid reflects on the danger of becoming captivated by appearances, wealth, status, and material success, especially when what we admire may have been built through injustice or harm.In this episode, you will learn:

Yeni Şafak Podcast
Taha Kılınç - Hamas'ın çizgisi

Yeni Şafak Podcast

Play Episode Listen Later Aug 12, 2026 5:00


Merhum İsmail Heniyye, Aksâ Tufanı'nın başlangıcından kısa bir süre sonra, Dünya Müslüman Âlimler Birliği'nin 9-10 Ocak 2024'te Katar'ın başkenti Doha'da düzenlenen toplantısında yaptığı konuşmada şunları söylemişti:

Büchermarkt - Deutschlandfunk
Karosh Taha: "Gulistan"

Büchermarkt - Deutschlandfunk

Play Episode Listen Later Aug 11, 2026 6:50


Schoeß, Marie www.deutschlandfunk.de, Büchermarkt

Büchermarkt - Deutschlandfunk
Büchermarkt 11.08.2026: Longlist 2026, Karosh Taha und Sonya Walger

Büchermarkt - Deutschlandfunk

Play Episode Listen Later Aug 11, 2026 19:44


Fuhrig, Dirk www.deutschlandfunk.de, Büchermarkt

Vikerhommiku intervjuud
Tarmo Jüristo: opositsioonierakonnad praegu valitsusse ei taha

Vikerhommiku intervjuud

Play Episode Listen Later Aug 11, 2026 9:42


Lesestoff | rbbKultur
Karosh Taha: "Gulistan"

Lesestoff | rbbKultur

Play Episode Listen Later Aug 11, 2026 5:42


"Gulistan" ist kein fiktives Land. Es ist der Name der Hauptfigur in Karosh Tahas gleichnamigen Roman. Diese Gulistan lebt in einer Diktatur, die der im Irak Saddam Husseins ähnelt. Als Kurdin wird sie politisch verfolgt. Und macht sich trotzdem auf die Suche nach ihrem Mann, als er spurlos verschwindet. Wie Karosh Taha das Leben im Extrem und seine Konsequenzen für die Figuren erzählt, berichtet Marlen Hobrack in der Buchkritik auf radio3.

Qur'an Conversations
S4 E22: The Danger You Don't Notice (TaHa 128-130) | Quran Conversations Season 4

Qur'an Conversations

Play Episode Listen Later Aug 7, 2026 61:20


What if the greatest warning signs in your life are the ones you've learned to ignore?In this episode of Quran Conversations, Dalia Mogahed is joined once again by Imam Kebba Sallah to reflect on verses 128-130 of Surah TaHa.Allah asks a powerful question: Has it not become clear to them how many generations before them We destroyed, through whose homes they now walk? The verse reminds us that history is not merely something to study—it is something meant to transform us. Every ruined civilization, every forgotten nation, and every warning from the past is an invitation to examine our own hearts before it's too late.In this episode, you will learn:

Yeni Şafak Podcast
Taha Kılınç-Arapgir'den dünyaya

Yeni Şafak Podcast

Play Episode Listen Later Aug 7, 2026 5:12


Fetih ve diğer Filistin fraksiyonları, 1960'ların ikinci yarısından itibaren Ürdün'de yoğunlaşmış, oradan İsrail'e “fedâî” eylemleri düzenliyordu. Hamas'ın 7 Ekim 2023'te başlattığı “Aksâ Tufanı Operasyonu” gibi düşünün.

Yeni Şafak Podcast
Taha Kılınç - Sebte kumpası

Yeni Şafak Podcast

Play Episode Listen Later Aug 5, 2026 4:57


İspanya'nın başkenti Madrid'deki tek sinagogda, 31 Mart 1992 günü yaklaşık iki saat süren bir merasim vardı. İspanya Kralı Juan Carlos, eşi Kraliçe Sofia, İsrail Cumhurbaşkanı Hayim Herzog, İspanya Yahudi Cemaati Başkanı Samuel Toledano ve kalabalık bir davetli topluluğunun katıldığı merasim, Yahudilerin İspanya'dan sürgün edildiği meşhur “Elhamra Bildirgesi”nin resmen iptali münasebetiyle düzenleniyordu. 31 Mart 1492'de Kral Ferdinand ve Kraliçe İsabella'nın Müslümanların elinden aldıktan hemen sonra Gırnâta'daki (bugün İspanya'nın Granada şehri) Elhamra Sarayı'nın Elçiler Salonu'nda deklare ettikleri fermanla İspanya Yahudileri (Sefaradlar) dünyanın dört bir tarafına sürgün edilmişti. Tam 500 yıl sonra, bir başka İspanya kral ve kraliçesi, ülkelerinin başkentinde Yahudileri ağırlıyordu.

Latent Space: The AI Engineer Podcast — CodeGen, Agents, Computer Vision, Data Science, AI UX and all things Software 3.0

Watch the full episode on YouTube:We first covered Baseten last year when DeepSeek mania was at peak hype. Now they have raised a monster $13B round and become one of the new cohort of AI Infra decacorns that are (with Nvidia, Intel, and the semis complex) chief beneficiaries of the Inference Inflection. We return to Baseten at the peak of the 2026 edition of Open Weights debate. Ali has published a viral breakdown of Kimi K3:And since you last saw him, Philip has spoken at AI Engineer and written the definitive book on Inference Engineering spotted all over SF:Three years ago, inference engineering barely existed as a category.Today, it is one of the most critical disciplines in AI. Inference engineering inherently tackles a different question than standard model training: “How do you turn those weights from training into a product that is fast, reliable, and affordable at scale?” Focusing on these creates an entirely new optimization problem.In one recent GLM-5.2 experiment, quantizing more of the model actually preserved its benchmark quality while increasing throughput by 20%, because the errors introduced in different layers could cancel each other out.Inference is no longer just the final step after training. It is becoming its own engineering discipline, with its own research problems, infrastructure, and increasingly specialized roles.In this episode, Baseten's Philip Kiely and Ali Taha join swyx and Vibhu to explain what actually happens after a new open model is released and what it takes to turn “we generated a token” into a fast, reliable, production-ready API.We go deep on cache-aware routing, disaggregated prefill and decode, quantization, speculative decoding, KV-cache movement, model parallelism, GPU kernels, and the race to make frontier models up to 10× faster. Philip and Ali explain why inference optimizations can still produce gains of 20%, 100%, or even 200%; how quantization errors can cancel one another out; why identical weights can behave differently across clusters; and how Baseten grafted a Kimi vision encoder onto GLM-5.2 without changing the underlying language model.The conversation then expands beyond LLMs into NVIDIA Dynamo, mega kernels, Rubin, AI-specific chips, local inference, video generation, diffusion versus autoregressive models, and the enormous compute barrier to generating coherent long-form video. Finally, we explore the convergence of training and inference, continual learning through persistent KV cache, and the emerging loop where models help optimize the infrastructure that runs them.We discuss:* What happens when a 200,000-token request enters an inference system* Cache-aware routing and reusing previously computed KV cache* Why prefill and decode are increasingly handled by different GPUs* When dedicated deployments become cheaper and more reliable than shared APIs* How speculative decoding uses a smaller model to accelerate a larger one* Tool calling, structured outputs, and what LLMs actually do* What it takes to support a new open model on day zero* Grafting Kimi's vision encoder onto GLM-5.2* Retrofitting inefficient model layers with components from other architectures* Why models sometimes collapse into repeating the same token* How hardware, kernels, and race conditions create nondeterministic failures* Preserving model fidelity while making inference faster* How quantization errors can cancel each other out* Why inference optimizations still deliver gains of 20%, 100%, and 200%* How optimized serving can make a model up to 10× faster* NVIDIA Dynamo, KV-aware routing, and distributed model serving* Speculative decoding the speculative decoder* Why local AI is about making models less dumb while data-center AI is about making them less slow* Tensor, expert, and pipeline parallelism across GPUs* Hardware-aware model design, auto-tuning, and the case against mega kernels* Rubin and why inference is becoming a systems problem* Whether modern GPUs are evolving into programmable AI ASICs* Why enormous models like Kimi K3 require GB300-class hardware* Why open-source video generation still trails Veo, Kling, and other closed models* The quadratic attention bottleneck behind long-form AI video* Autoregressive video, real-time generation, and compounding quality drift* Why future video systems may combine autoregressive and diffusion architectures* Training for inference and inference for training* Continuous post-training, deployment, evaluation, and improvement loops* How GLM-5.2 helped optimize the kernels serving GLM-5.2 itself* Why faster networking could unlock dramatically faster decoding* Continual learning, KV-cache compaction, and persistent model memoryShow Notes* How to build a day-0 API for Kimi K3* 22580: From GPT2 to Kimi3, ExplainedPhilip Kiely* LinkedIn: https://www.linkedin.com/in/philipkiely* X: https://x.com/philipkiely* Inference Engineering: https://www.baseten.co/inference-engineering/Ali Taha* LinkedIn: https://www.linkedin.com/in/aliestaha/* X: https://x.com/waterloointernTimestamps00:00:00 Introduction and the 200K-Token Prompt00:03:18 Dedicated Deployments, Speculative Decoding, and Tool Calling00:11:26 Launching Production-Ready Open Models00:19:06 Model Retrofits, Failure Modes, and Nondeterminism00:28:22 Quantization and Canceling Errors00:32:15 The Race to 10× Faster Inference00:40:48 Dynamo, Speculation, and Local vs. Data-Center AI00:50:18 Model Parallelism, Auto-Tuning, and Mega Kernels01:00:55 Rubin, GPUs vs. ASICs, and Custom AI Chips01:10:03 Giant Models and the Limits of GPU Memory01:12:42 AI Video, Quadratic Attention, and Autoregressive Generation01:21:47 Audio, Images, and Diffusion Models01:27:32 Training, Self-Optimizing Models, and Continual Learning01:40:06 Closing ThoughtsTranscriptIntroduction: Baseten, Waterloo Intern, and Inference EngineeringSwyx [00:00:00]: Okay, we're here in the studio with Philip, old friend from Inference Engineering, the book, as well as Baseten and everything that you've done, you and I have done before, as well as Ali. Welcome.Ali [00:00:15]: Pleasure to meet you.Swyx [00:00:15]: Waterloo intern.Ali [00:00:16]: Waterloo intern, always.Swyx [00:00:17]: When did you get “Waterloo intern” as a handle?Ali [00:00:19]: As a handle? Oh.Ali [00:00:20]: I think the rebranding happened mid-March. When I saw it was open, I was like, “I have to take it. Up for grabs.”Philip [00:00:26]: The problem is that Ali is really good at his job and is not gonna be an intern much longer.Philip [00:00:30]: So we have to figure out who's gonna get the handle.Ali [00:00:33]: Well, I'll pass the torch over to the next intern.Swyx [00:00:34]: Oh, okay. It can be, like, you just pass it to another Waterloo grad.Ali [00:00:37]: To another Waterloo intern. No, bruh.Philip [00:00:39]: Yeah.Ali [00:00:39]: Intern.Swyx [00:00:40]: Intern, yeah.Ali [00:00:40]: And no.Philip [00:00:41]: You gotta get an intern from Waterloo.Ali [00:00:42]: Yeah, I've gotta get an intern from Waterloo.Swyx [00:00:44]: Right.Ali [00:00:44]: But they have to follow the path.Swyx [00:00:45]: Oh, it could, but it could come from Baseten, so it's like whoever Baseten gets from Waterloo.Ali [00:00:48]: Right.Swyx [00:00:49]: Has the title of Waterloo.Ali [00:00:50]: It stays in the ecosystem.Philip [00:00:51]: Exactly.Ali [00:00:52]: Halfway through the internship, you either get it or you're out.Philip [00:00:55]: You should also do, like, a big graduation ceremony where you change the handle.Ali [00:00:59]: Just say it.Philip [00:00:59]: For everybody.Swyx [00:01:00]: You guys are good at ceremonies, clearly. We had a nice launch of the book, very successful. But before we get into all that, I wanna start off with a fun question for you. Okay, you're an expert inference engineer. What happens when I send a long query, say two hundred thousand tokens into Baseten's inference? What's the process of query through GPU model routing, balancing, all that? What is all the stuff that we don't think about?Long Context Requests, KV Cache, and Cache-Aware RoutingPhilip [00:01:26]: With a long query specifically, the first thing that I'm gonna ask is, “Have you sent me this query before, or at least part of it?” and I really hope you have, because it's gonna be a lot easier for me and a lot cheaper for you. So the first thing that we're gonna look at is some cache-aware routing, where we're going to see, we probably have a number of instances, a number of replicas up serving whatever model you're hitting. We want to send this one to something with, number one, available prefill workers, and number two, ideally some cached input already there so that we can skip prefill on at least part of these two hundred thousand tokens. If you're doing two hundred thousand tokens, it's probably coding or a multi-turn agent or something where you would expect to have that cached. If you don't, we're gonna have to send it to a prefill worker. We've at least on certain models disaggregated prefill and decode, so you're going to have one set of GPUs that's solely going to process the input, create the KV cache, and get you your first token, and then that's going to be passed over to a separate set of GPUs, which is going to run decode. We're going to iteratively make those tokens. We're probably going to have some speculator model in front of that. I'm going to assume that you're doing coding, and because of that, our speculator model, which assumes you're doing coding, is gonna have a high draft token acceptance rate. If I'm wrong and you're asking me to summarize every Harry Potter book, it's gonna be slower. And then we stream that output to you and account for it, charge you, a couple of pennies and say, “Hey, would you like to send another one?”Swyx [00:03:04]: Except Baseten doesn't charge by pennies.Philip [00:03:07]: Well, yeah, we charge. I'm assuming that we're talking about the public model APIs. If you are setting up a dedicated deployment, then yeah, it's not pennies.Public APIs vs. Dedicated DeploymentsSwyx [00:03:18]: Yeah, one of the key differentiators when I was talking with Baseten initially was that people who want very high volume just need to rent by the box, ‘cause then it's up to you to figure out how to saturate the box.Ali [00:03:31]: And more often than not, it's, like, way cheaper if you're pushing, like, millions of tokens per hour, if you just pay per hour instead of pay per token.Philip [00:03:37]: Yeah, they do. I think that we've increasingly seen a lot of demand for the pay per token APIs, just because everyone wants to try open models, and then once they find a use case that's really sticky, then they move over to dedicated.Swyx [00:03:51]: Is there a best practice on when it's time to swap over?Philip [00:03:54]: Couple reasons. Yeah, reliability, that's a big one, right?Ali [00:03:57]: Like, if they have a very specific use case, they want you to train something specifically for them, like they want their own spec dec, for instance, for their own traffic.Swyx [00:04:04]: Spec dec is speculative decoding.Speculative Decoding and Custom SpeculatorsAli [00:04:05]: Speculative decoding, yeah.Swyx [00:04:07]: You have to explain.Ali [00:04:07]: Sorry. Like, speculative decoding is like, if you have a huge model, right? And so the model is going to be generating one token at a time every single turn, every single forward pass. So we attach, like, this little, like, parasite, like this layer that goes on top of the model, and this model just has to predict. It does three very fast autoregressive forward passes, and it will predict, like, three certain tokens, and then you do one forward stage over the entire original model in order to see if those predictions were correct or not, and then you accept them or you reject them. Now, this draft model is traffic specific, so if you, like, Philip said, if you're summarizing Harry Potter books, I can train exclusively that draft model on Harry Potter books, and I can guarantee you that I'm gonna accept the three tokens every single time. And so with that case, I increase your decode speed. I wouldn't be able to provide this to you if you're a shared endpointSwyx [00:04:53]: YeahAli [00:04:53]: ‘cause I have no idea if you're doing Harry Potter, if you're doing coding, if you're doing English. We don't know. Also, there was a thing in the book that mentioned that if they really cared about a specific threshold, chapter four, I think. Do you remember that?Philip [00:05:06]: Yeah. The things that you can do is you can set a specific, like, batch sizing, a specific, like, parallelism strategy if you're trying to optimize for, like, throughput versus latency. You can. Maybe a NVFP4 quant doesn't pass your benchmarks and you wanna run a model at higher precision, you could do that. There's just a bunch of reasons why you might wanna have your own endpoint and the biggest one, of course, just being, like, you don't have to deal with someone else doing a hundred million tokens of benchmarking traffic at the endpoint when you happen to be trying to serve your users.Swyx [00:05:40]: Yeah. I think one thing that is. That is a classic journey. Like, it's people is asking the, what happens when you type Google into the browser. Tool calling, is that just, you're generating JSON or is there more complication beyond that?Tool Calling, JSON, and Structured OutputsAli [00:05:58]: Certain customers that we have, they have their own post-trained models, and so they demand a tool calling that's not just, like parse a file or go find the weather. It's something that's very specific and you have to do post-training on this. And if the post-training on the model is not good or if the quantization after the post-training to get the inference to be fast, the model will struggle reading the JSON file and reading the tool calling. But it doesn't require its own like sandbox. It's not like it's going to use that tool calling to like escape a sandbox or like it doesn't have to be contained. It can just be a normal dedicated deployment. The challenge with tool calling more and more seems to be that the companies want certain tool calling which is a very sensitive thing to train. And because you're dealing with all of the JSON outputs, if it doesn't like close the end of the request in a very certain manner, you end up with a model that did the tool calling and like the thinking and so as a result of that, it didn't see the result and just hallucinated the result as it decoded. That seems to be the most challenging thing with tool calling, not really the sandboxes model.Philip [00:06:56]: Yeah, that's a challenge on the training side and then on the inference side, there's work that you can do to scope the possible output. So we published this at this point close to two years ago, the solution to this problem which is you make a state machine and you use that to constrain the output to a specific format. So this is the structured output problem. If you remember backSwyx [00:07:27]: Yeah, the specific grammar is,Philip [00:07:29]: Yeah, exactlySwyx [00:07:30]: GML had this thing.Philip [00:07:31]: Yeah. So it's like the old-school “make sure this is only JSON”, return only JSON orSwyx [00:07:38]: YeahPhilip [00:07:38]: Grandma's gonna die type of prompts.Swyx [00:07:39]: Is it BNF grammar? At some point OpenAI had released a thing that was like, yeah, if you want to constrain your output, write BNF grammar, back as NOR.Philip [00:07:47]: In our inference system, it's just a specified output format. And you get the guarantee that your output's gonna be structured along that format. And so applying that to tool calls can like help cut down on. You can still call the wrong tool or call no tool. It doesn't solve the certainty problem but it at least solves the output structuring problemSwyx [00:08:10]: YeahPhilip [00:08:10]: Within tool calls.Swyx [00:08:12]: And MCP is just another form of tool, right.Philip [00:08:14]: Yeah, exactly.Swyx [00:08:15]: As far as there's no special thing there.Philip [00:08:16]: The thing I'm always like explaining to people is the LLM is not capable of doing anything. It's only capable of making suggestions of what to do and then if those suggestions are formatted in a certain way and applied to a system that knows what to do with them, then an action occurs.Swyx [00:08:32]: Yeah. Part of the fun stuff is, this is solved outside of tool calling too. Like in an agent loop if the output is not correct or you're right, like reasoning, tool calling was done in the reasoning trace, just be like, “Oh, I don't know what to do. Let me just try again.” And it might get there after a few tries. And on your point of training, sometimes this is harder in smaller models, so you don't have the same exact quality outputAli [00:08:56]: Right.Swyx [00:08:57]: When you just swap from a big model, right?Ali [00:08:59]: Yeah. I will say that, before, I think we need to go back to inference engineering proper.Ali [00:09:04]: But, I had expected that something would replace JSON because it's hard to stream JSON ‘cause JSON must be complete and you must have open and close brackets and everything. So it's hard to parse something or validate something while it's being streamed. So people invented all sorts of things that are like, I forget the name of some of these alternatives, but it's something like TOML, something like YAML. But JSON seems to be dominant still.Philip [00:09:30]: The JSON outputs aren't that long, right? Like you could have a long-- ‘cause tool calls also contain the arguments in them and perhaps for a certain tool you might pass like a very long argument. But my impression of the median tool call is that it's a relatively small number of tokens, right? So I would expect that speculators are generally fairly good at something as formatted as JSON. And so you would have like a pretty fast decode step there and that the streaming wouldn't be as valuable, but maybe I'm wrong about that.Ali [00:10:02]: I think you're also bounded by the software or that the model is gonna integrate with if the software is built with JSON for the tool calls or if the company that you'- if your customer says that this is how our software works and our tools are interfaced with JSON, you can ask them to like, change their software and say like, “Yeah, this is gonna be better for the model.” but like with the right training shouldn't be that much of a difference. Also more profitable if it outputs more tokens probably.Swyx [00:10:25]: Depends on your business model.Swyx [00:10:27]: It really depends. But I will say that, as a writer with like experience a lot with generated output, I do try to move from text to JSON text which is very long JSON, right? Like there's paragraphs in every field because I'm trying to structure it, right?Philip [00:10:44]: Right.Swyx [00:10:44]: I want you to first make factual statements, then make opinions then make bullet point summaries, have dates, have entity references have your sources for references, all these things. Anyway, so these are things that like I think people who really experiment with structural output have to really care about. But, let's, let's recurse up the stack a little bit. Before we started recording, you mentioned something really cool, which is that there's a lot of engineering that-- inference engineering that goes on when a new model provider releases a new model, right? So let's call it GLM-5.2, Kimi K3. I had previously assumed, especially if it's like, well, GLM 5 to 5.1 to GLM-5.2, like that you've supported them before. Is it that much work?What It Takes to Support a New Open ModelAli [00:11:26]: It's a lot of work.Swyx [00:11:28]: Yeah. Okay. So like, a lot of people, all you guys, right whenever a new model launch like, people rush to say like, “Oh, Hugging Face supports this, Fireworks supports this, Spacetime supports this,” and I'm like, “Yeah, of course we support it.” But what goes into that? What goes intoPhilip [00:11:40]: I think it's more than just support it too, right? It benefits the consumer a lot. Like I think it was with Kimi K2.5 or GLM-5.2 the latest, there was an inference war, right? X provider is at 90 tokens a second. The next day we're at 150. The nextSwyx [00:11:55]: I kinda kicked that off with the GLM-5.2.Swyx [00:11:58]: I wrote a Twitter article about. It got like half a million views,Ali [00:12:02]: Based on being numberSwyx [00:12:03]: YeahAli [00:12:04]: Or it's for something else.Swyx [00:12:05]: Yeah. Which,Ali [00:12:06]: Oh my GodSwyx [00:12:07]: Which then got everyone really excited about, hey, how can we, bend tracks a little bit further and,Philip [00:12:14]: There's a difference between support the model, as in I can make a token out of this model, and support a model, as in I have a production-ready API from this model.Philip [00:12:26]: Getting to the point of I can make a token out of this model is not that hard because generally the, open source inference engines, vLLM, SGLang of the world oftentimes even receive weights ahead of time, maintainers do, or the people making the model merge PRs to ensure support. So you generally can, just get it working on the standard open source stack without too much pain in most cases. The challenge is, every inference company is gonna have own proprietary stack. Some open source components, some in-house stuff. And for any arbitrary model, there's going to be some new stuff. Sometimes you get lucky, like K, two five to two six was, like, pretty similar.Quantization, Speculators, and Production ReadinessAli [00:13:16]: Yeah. It was pure continued post-trainingPhilip [00:13:18]: YeahAli [00:13:18]: If I remember correctly.Philip [00:13:19]: Even in those cases, there's still stuff you have to do. You have to redo the quantization work. You're taking the model from. Generally, these models are not released in NVFP4, and we want them to be in NVFP4 for maximum Blackwell compatibility. So we have to perform that quantization, and, calibrate the quantization to make sure that we're not causing any regression in the model's intelligence. And then we also have to train the speculator, as we've talked about. Generally, we have. We have ZDR, zero data retention on our model APIs, so we don't know exactly the traffic that people are sending us, but we know what's popular. We know that coding use cases are popular. We know that agents, agentic use cases are popular. So we can get public data sets that are representative of that traffic and train general speculators. Now, with speculators today, you need to train the speculator using the base model itself because you're getting hidden states out of the model from running inference on these specific prompts, and that is the training data you use to create the speculator. So there's that process which you need the real model weights for. And then there's of course just the process of, standing up all the infrastructure behind it, loading all this stuff, testing it. And then when there's a new model with a newer architecture, I think that, like, the DeepSeek models tend to be the most challenging as they have, like, the most novel architectural stuff going on, model after model. But every new model has something. Kimi K2 had. Oh, sorry, GLM-5.2 hadAli [00:14:53]: Sparse attention.Philip [00:14:54]: Yeah,Ali [00:14:54]: YeahPhilip [00:14:54]: the DSA.Ali [00:14:55]: Right. Which is brought from DeepSeek.Philip [00:14:57]: Yeah. AndAli [00:14:59]: So you can copy-paste then?Philip [00:15:01]: It kindAli [00:15:01]: I don't know how this works.Philip [00:15:02]: So, like we had to, like, build support for that into our runtime. And you're right, like it is really interesting the way that all of these open source labs borrow from each other. For example, like GLM-5.2 doesn't have vision. So something that, Haley, a guy on our team, if we could take a look at this, he, like, grafted the Kimi vision encoder onto GLM-5.2.Retrofitting Vision into GLM-5.2Ali [00:15:27]: We'll be training the projector.Philip [00:15:28]: Exactly. So if you think about, like, the encoder, there's the encoder, which is the part that looks at the image and turns it into latent information, and then there's the projector which likeAli [00:15:38]: You can say latent space. It's okay.Philip [00:15:41]: And then there's the projector that maps it onto, the model itself, and then there's the model weights. You don't wanna mess with the model weights because you run a chance of making the model dumber at something else for the purpose of giving it vision. So instead, Haley started with just a projector, which is only a handful of millions of parameters.Ali [00:16:02]: That would be, yeah.Philip [00:16:02]: Yeah.Ali [00:16:03]: Can you show the training one?Ali [00:16:04]: Like the way it groksPhilip [00:16:05]: YeahAli [00:16:06]: Very interesting.Philip [00:16:06]: And maybeAli [00:16:07]: That right therePhilip [00:16:07]: Maybe Ali, you should take it from here. You've got a betterAli [00:16:10]: Ooh, double the sandPhilip [00:16:11]: Understanding of this than I do.Ali [00:16:11]: Yeah. You can see, like, he. The way he trained this is really cool. At the beginning, he was training it using just like, “Here's a picture of a mountain. Can you describe what's in this mountain?” And that caused it just like the first, learning walls. Like here you can see this all we're trying to teach it is to translate the encoded. Like it's already taken the encoder from Kimi K. It's taken the image. It'Philip [00:16:31]: Yeah. FrozenAli [00:16:31]: FrozenPhilip [00:16:32]: With adapter.Ali [00:16:32]: Exactly.Philip [00:16:33]: Yeah.Ali [00:16:33]: So the brain is frozen and the eyes are frozen. It's just we're tryingPhilip [00:16:37]: AlignAli [00:16:38]: Interconnect between the eye and the brain, right? So the projector. And so you take the tokens and then he's like, “Oh, can you describe what's in this image?” And he's like, “Oh, it's a mountain,” or it's a person or it's a human, whatever the case is. But that didn't cause complete understanding. So he changed it such that every image was associated with a data set of questions. Like, does this image have a white male? Does this image have birds in the top corner? Does this image have a scientist in it? All of that stuff. And it would have to answer questions correctly. And using not just training on describing an image, but being able to answer question, another question, answer over time. Like you can see the grokking, which is like genuinely insane, that retrofitting vision into a large LLM can learn to that extent. And even for images that it doesn't perform well on, for instance, if you ask it a picture of like Stephen Hawking, “Who is this?” Maybe it doesn't get it, but it will say something like, “This is Albert Einstein.” Like it still understandsPhilip [00:17:25]: Close enoughAli [00:17:26]: That this is a scientist who is a man who has, some significant achievements, all that stuff. So that's like really cool.Philip [00:17:32]: Yeah. So, we've covered Hao Tian before, who the author of the LLaVA paper that did this, a while ago. And I think that's very foundational work for anyone who hasn't done vision work before.Ali [00:17:41]: Same with the CLIP and MetaCLIP, where you go from just captioning to building out questionsPhilip [00:17:47]: RightAli [00:17:47]: Off the image and how much better you can get performance.Philip [00:17:50]: Right. Right. Right. Yeah. But what's, what's so exciting about this is if you look at a model like this. Now, this is a little bit more of a research project. It's not. It got to 56% on MMLU Pro, I think. So not quite frontier. But if you're running this model, you haven't suffered any loss on your GLM-5.2 quality. If you don't have an image, it'll just behave exactly the way it used to. And ultimatelyAli [00:18:14]: Which in the inference code you literally do not include the other part, right?Philip [00:18:18]: Yeah. You would just skip the encoder if you don't have an image input.Ali [00:18:22]: Okay.Philip [00:18:22]: Just confirming.Philip [00:18:23]: YeahAli [00:18:23]: Does it affect a lot on the overall inference side? Like you're not adding much, you're adding a very small vision encoder. These are typically likePhilip [00:18:30]: They're super fineAli [00:18:31]: Less than a billion parameters, right?Philip [00:18:32]: Yeah. It's, - There's a little bit less standardization among vision encodersSwyx [00:18:37]: YeahPhilip [00:18:37]: So the support matrix can be a little bit, sparser. But overall, yeah, it's a pretty, it's a pretty minor component of the overall system. And ultimately what you get out of the system is all of a sudden you have Kimi Vision, GLM weights, and DeepSeek attention all in one model.Open Source Model Grafting and Franken-MergesPhilip [00:18:56]: And that's, I think, a lot of the power and beauty of open source, is that you can take all of these different components and combine them together into a system that's better than anyoneSwyx [00:19:05]: YeahPhilip [00:19:05]: Can be individually.Swyx [00:19:06]: People used to say that you would also do Franken-merges where you would take likePhilip [00:19:10]: YeahSwyx [00:19:10]: Layers from each model.Swyx [00:19:11]: Does anyone do that anymore?Ali [00:19:13]: Well, to your point previously when you were mentioning like, the work that goes into supporting a model when it first comes out, like GLM-5.2 or MiniMax M3 or whatever the case is. Sometimes you do have to like, you do have to switch out some things. Like, for instance, the MiniMax M3 head uses full attention, and with full attention you end up with this like insane bottleneck in spec dec ‘cause you're doing auto-regressive token generation for three tokens, and you're doing this like N squared over all of the tokens that are in your sequence. Your KV cache is like very large because it's not sparse, it's not top K. So we find it better to like, okay, we're gonna replace this, we're gonna replace this layer with a layer from another model that's using like GQA, for instance. And then just with the right training, you can get it to have the same acceptance rate. So it is very possible to retrofit layers from other models and very much needed. If a layer is like inefficient, the training just becomes the challenge, like how do you ensure that you train it properly? Which again to your earlier point is like the mesh between training and inference. As in like you need very good training in order to do fast inference. That's like, I feel like more and more becoming true.Swyx [00:20:21]: Yeah. Anything else on the support side when you say like get it to fully production ready?Loop Detection, Race Conditions, and Non-DeterminismPhilip [00:20:26]: Yeah. I think that there's also a question of just, we can test a model to a pretty extensive degree, but we're trying to get it out quickly and then you see a bunch of other people test it and you get interesting results. There was an issue with, GLM briefly where we had some like mode collapses where it would just output the same token over and over again for certain prompts on certain temperatures. Like once you expose an endpoint to the real world, there's going to be, so many more varieties of things given to it that you're able to, discover and patch things. So it's not just a, day zero process, it's then like for the first week, for the first month, if a model remains popular, like how do you both fix bugs and then continue to push the envelope on performance?Ali [00:21:21]: What do you mean you don't want your model outputting S?Swyx [00:21:24]: Is there loop detection on that stuff, by the way? It still happens like quite a lot, which is surprising.Ali [00:21:30]: We have like we, in our endpoint, like if a model was to output the same token like four plus times, we just cut the generation. We say like, “Oh, sorry, this-- Like try again,” or like we will reprocess the request. ‘Cause we know then, like if it, like if, yeah, it's four times the same token, it's probably collapsed.Swyx [00:21:45]: Yeah. Is there a way to opt out in case I really want that?Ali [00:21:48]: You want that?Ali [00:21:50]: I think there's a way that we have to handle it. I'm not exactly certain, but I feel like in certain models, like when they output something like you can imagine, like a table for instance, and so they want, they wanna draw like 12 dashes and 12 dashes. Yeah, I think there's a way for that to happen. I think we only do it on certain tokens. Like we exclude certain special characters.Swyx [00:22:07]: Yeah.Ali [00:22:07]: So we only do it on like certain like S is the most common almost. GLM-5.2Swyx [00:22:11]: OhAli [00:22:11]: And I think it was DSV 4 as well. Like you'd just have like looping issues where like you literallySwyx [00:22:17]: ItAli [00:22:17]: Just have like S.Swyx [00:22:18]: Yeah. Is there a special, something special about S? No, just randomlyAli [00:22:21]: It just seems to be the one token involved.Swyx [00:22:23]: Yeah. And it'Philip [00:22:24]: Is thereSwyx [00:22:24]: And it's only temperature 0Ali [00:22:27]: NoSwyx [00:22:27]: Even at other temperaturesAli [00:22:27]: Even at like 0.9 or whatever, it will still, it will still collapse.Swyx [00:22:30]: That's weird, right?Ali [00:22:30]: It's, it is an inference problem to be honest, like a software problem. Like oftentimes, the image you run will-- like NVIDIA will release an image for instance, and if we will upstream the changes from their latest TensorRT-LLM image into our stack, we'll find that it fixes it. Or oftentimes this will only happen in an inference engine that you're using like SGLang. But if you were to switch to vLLM, that isn't the case. So it seems to be like an extremely like deterministic software issue and not really a model issue. It's not like a weights problem. Like I'- we'll say like, “Oh, it's a problem with the quant. We did PTQ wrong,” right? But that isn't, that doesn't make sense because the same weights used with a different inference engine does not repeat the problem. And sometimes it's, the kernels that are being used in the backend have like these very subtle sometimes race conditions, where if you were to use this model hosted on one cluster, you will never get this problem.Swyx [00:23:19]: Oh my God.Ali [00:23:19]: But if you host it on a different cluster, you will. And the reason is the KV cache transfer from a node to node in that one cluster is using a slower interconnect than the node to node in another cluster. So that exposes the race, whereas in another cluster it doesn't. So then you end up just like, okay, this model is not gonna be hosted on this cluster. We're gonna host it on, another cluster because that cluster exposed that problem. But then it ends up with like, okay, is it the software? Is it the model weights or is it the hardware?Swyx [00:23:42]: There is a thing about this with temperature 0 still not being deterministic, right?Ali [00:23:46]: Right.Swyx [00:23:46]: Mostly because of hardware. Even at temperature 0 same model, you won't always get the same output.Swyx [00:23:52]: Even-- But I'm surprised by the race condition one because, I thought PyTorch was a graph that like guarantees that you at least, execute things in the right order.Ali [00:24:02]: Well, yeah, true. Like I'm not, I'm not saying that there is. Like well, you have things like PTL optimizations where like you can start a kernel before the end of the previous kernel, and that's like ‘cause you want to do that because there'sSwyx [00:24:12]: It's like pipeliningAli [00:24:12]: Expense. Exactly.Swyx [00:24:13]: Yeah.Ali [00:24:13]: But it'- But you don't do it cleanly. Like you overlap a little bit of the execution. No, it is very possible that the kernel itself, like that one block that is supposed to be running in this instance of time, that kernel itself has a race condition. For instance, like a missing barrier. Like often if you're designing a kernel and you want it to make it to be very fast, if you don't test it extensively, you'll, you'll have certain threads access data points from registers before they've been written to by other threadsSwyx [00:24:36]: YeahAli [00:24:36]: For example, because like your barrier is wrong or your synchronization was wrong. But yeah, like the testing itself is very difficult in those like, andSwyx [00:24:42]: And there's no like borrow checkerAli [00:24:45]: What does that mean?Swyx [00:24:46]: Like Rust. Like the. If you're trying to have like memory safety It sounds like a comparable problem.Ali [00:24:52]: Well, yes, but you're working in CUDA, right, NVIDIA GPUs. Like- You just need a higher level language like modular Maybe that's what modular is supposed to do. I don't know.Quantization Quality and Vendor FidelityVibhu [00:25:00]: How do you see keeping quality of the model? So you talked about all these steps of, okay, you gotta do quantization, train your own speculative decoderAli [00:25:07]: RightVibhu [00:25:07]: Run on different hardware. Looking at other model providers, okay, you kicked off a inference speed race on the consumer end. What goes into keeping quality the same across them, right? Sure, you can run benchmarksAli [00:25:22]: YeahVibhu [00:25:22]: But, like, how do you determine how much quantization are there standards? What goes intoPhilip [00:25:27]: There's a few things on quality. Most inference optimizations are lossless. KV caching, for example. You are just recomputing or preventing recomputing the same values. Speculation, of course, if a draft token is wrong, it gets rejected. The main lossy optimization is quantization. And that really comes down to, number one, data format, number two, which parts of the model you choose to quantize, which layers, and number three, like doing a lot of calibration on the quantized weights, to ensure that you're preserving all the outliers. There's other tricks that you can do, though. A big one is long context, ‘cause one thing you asked at, right at the beginning is, “Oh, what's gonna happen if I send a 200,000 token request in?” So with a long input sequence, you need to, store a lot more information. You need to process a lot more tokens. And so even if a model has a context of a certain length, you might, as an inference provider, choose to build an API with a shorter context length, and of course a full length one as well. Because if someone doesn't need the full million token context, for example, you can get them better performance. I don't know if that's exactly like quality of the model. The way that I think about quality is to what degree are we faithfully serving the original model? If you think of a golden implementation of a model that performs exactly the way the model is designed to perform, I think of quality as how close are we getting to that, 100% fidelity of the model.Philip [00:27:13]: You can also, of course, think about quality from the training side and how do you push yourself past 100%. But when I think about purely inference optimizations, it's getting faster while staying as close to that 100% fidelity mark as possible. And certainly our standard internally is that, like you should not be able to tell the difference between our API and a, official API. I think Kimi in particular does a good job of vendor benchmarking hereAli [00:27:41]: YesPhilip [00:27:41]: Where they haveAli [00:27:42]: They released an actual vendor benchmark.Philip [00:27:43]: Exactly, yeah.Ali [00:27:44]: ‘Cause they accused, some people, Amazon? There was some provider that was not doing very well on Kimi's benchmark.Philip [00:27:50]: Yeah.Philip [00:27:51]: So, with Reflect we probablyVibhu [00:27:52]: This was a long time ago, right?Philip [00:27:54]: No.Ali [00:27:54]: Yeah, like threeVibhu [00:27:55]: They alsoAli [00:27:55]: Four, five months agoVibhu [00:27:57]: This also happened with, I don't remember which model, but they pulled out quite a few, and then they started a whole chart about this. It might have beenPhilip [00:28:03]: Kimi Vendor Verifier.Ali [00:28:04]: Yeah.Philip [00:28:05]: Yeah.Ali [00:28:05]: Yeah, ‘cause you, ‘cause you'd be pissed, right? Like if you'Philip [00:28:07]: Yeah.Ali [00:28:07]: If like if I'm a consumer and I'm using like Amazon's endpoint for instance, and I've used Kimi and I'm like, “Oh my God, like this is bad,” I'm not gonna say, “Oh, Amazon quantized the model in a bad way.” I'm gonna say, “Oh, Kimi sucks.” Right?Philip [00:28:17]: Yeah.Ali [00:28:17]: So it seems like that makes sense.Philip [00:28:19]: Yeah, they care. They care.Vibhu [00:28:21]: Justifiably.Ali [00:28:21]: Yeah, justifiably.Vibhu [00:28:22]: This is probably a stupid question, but just checking, has anything improved from main quantization?Philip [00:28:28]: Yeah.Vibhu [00:28:28]: Like, is quantization always strictly worse?Ali [00:28:30]: Well technicallyVibhu [00:28:32]: NoAli [00:28:32]: It's a lossy. QuantizationPhilip [00:28:33]: YeahAli [00:28:33]: Is a lossy, it's a lossy implementation.Philip [00:28:36]: Speed improvesVibhu [00:28:36]: Speed improves.Ali [00:28:37]: It the number, likeVibhu [00:28:38]: No, I' always look for inverse scaling laws.Philip [00:28:40]: Yeah.Ali [00:28:40]: Yeah.Vibhu [00:28:40]: This is something I learned from Noam Brown, where like things that normally act in one direction sometimes do.Philip [00:28:45]: Well, technically when you run a benchmark, because these models are deterministic, sometimes your,Ali [00:28:52]: YeahPhilip [00:28:52]: NVFP4 quant is like, two basis points higher than yourAli [00:28:56]: No, it's noise. It's noise.Philip [00:28:57]: Yeah, exactly. I'm like, yeah, it's, it's within. That's why I always say within margin of error.Philip [00:29:01]: And I stopped saying that because everyone assumes that what is, well, within some margin of error, we're barely inside of that to the worst, so we're saying. But yeah, sometimes it's just like, gives you a higher output score. But like Ali said, that's noise. To my knowledge, you're not necessarily making the results better. You're just trying to, again, like keep your fidelity as close to 100% to the original model.Layer Selection, KL Divergence, and Better QuantizationAli [00:29:27]: There is, to your point, research that we did on MP. I don't know if you are able to pullPhilip [00:29:31]: YeahAli [00:29:32]: A tweet we did. One of our research interns, Joshua, I think it's a tweet on how we have 20% better quantized GLM-5.2 than NVIDIA. Essentially what we found throughout like this month research is, okay, quantization is a lossy. It's. You're compressing the data from, occupying 16 bits to occupying, four bits, for instance. And so you're losing some information, and you're trying to minimize that. And so when I say that I'm gonna quantize the model, my job becomes how do I find the layers that I can quantize, and how to find the layers to not. For instance, with image models, I don't quantize modulation layers, and I don't quantize out projections because those two are. Like out projection is what you see as the user. Modulation is what the model sees or understands. Right, exactly. And so to his paper, do you have the. It doesn't have the. Yeah. It's a long paper. I don't know if I can findVibhu [00:30:25]: If there's a part to search or it's probably in the thread.Ali [00:30:28]: It's probably in the thread.Vibhu [00:30:29]: Yeah.Ali [00:30:29]: But the long and the short is it is very possible that quantizing more of the model makes the results. Like if I have a model that I quantize layers one, five, and 10, and another model where I only quantize layers one and It is possible that the model in which I quantized more information is going to perform better because the quantization errors have canceled out. And so what Joshua showed in his mathematical proof where he had like a verifier in, is that you can predict which layers are going to have quantization errors that will cancel out with each other, and you choose to quantize those layers. And so the result of doing this mathematical quantization is you end up with a model that's 20% more quantized than another provider, so you get 20% more throughput of it because there's more layers than running an NVFP4, and your quality is better than that other quant because the layers that you chose to quantize have their errors cancel out, like one layer skewed to the right one layer skewed to the left, one layer skewed to the right. Your final logits distribution is more similar to the original distribution of the model, so you have better fidelity. And so the way we proved this was with KL divergence. So instead of just scoring on the benchmarks, we scored the KL divergence between the logit distribution of the quantized model and the logit distribution of the original full precision model, and we showed that with this technique we get. If your probability distribution on the logits which token it wants to select is more of the same as the original model, you're probably gonna end up staying true to the original model. So yeah, so it seems like previously before this, it seemed like the industry was, well, the more you quantize, the worse it's gonna be, ‘cause the more loss you introduce. That's not exactly, not necessarily true. So yeah, doesn't improve it, but can cancel out.Philip [00:31:57]: I think it might be this, but reminds me a good bit about pruning where you can prune off certain layers.Philip [00:32:03]: But very interesting. Didn't know this was a whole paper you guys put out.Ali [00:32:06]: It's. Fun fact, it was originally 72 pages, this paper, and then we decidedPhilip [00:32:11]: WowAli [00:32:11]: We can't tell. We couldn't release it. So it's now 45.Swyx [00:32:15]: Still 39 pages, so very substantive. We talked about evals and all these things and, like what's possible in terms of speedup? Like it's like probably like the numberInference Speedups and BenchmarkingSwyx [00:32:25]: Thing that people do wanna care about, and it's something that you wrote about in your post. Like official API is 70 tokens per second, and you push it up to 90. Is that like a normal thing?Philip [00:32:36]: So what's cool about working in inference, the reason that I think inference is going to be a useful place to do engineering for a long time, is that if you look at highly optimized domains like, say, finance, if you're in finance, you measure how much better you got in basis points. It's like, “Oh, I got five basis points better, like twentieth of 1% better,” that's huge news because everything is so optimized. When we publish optimizations, it's 20%, it's 100% it's 200%. So there's still probably like a lot further to go, honestly. Like you'll, you'll know that inference is pretty much solved when researchers start publishing about how they got 1% faster at something.Swyx [00:33:19]: Which by the way, because I am from the finance background, in the ‘70s, that was the margin at the time. When you did quantitative finance research, you would findAli [00:33:27]: And like 20%, tens of percent.Swyx [00:33:29]: That's. Yes.Philip [00:33:29]: Yeah.Swyx [00:33:30]: And now it'Philip [00:33:31]: Tiny fractionsSwyx [00:33:32]: For those people interested, look up Andrew Lo's paper. He had a really interesting illustration of quant, stat arb, distribution, narrowing down from like those kinds of 20% differences in the ‘70s, down to nothing today, which is very cool.Philip [00:33:48]: Exactly, and we're at the beginning of the same type of thing. Now benchmarking is hard. I think anyone will tell you that, and benchmarking provider speeds is hard because there's so many variables that go into it. What hardware are you using? How much load do you have on the system? What's the exact nature of the prompts and input and output sequence lengths? All that stuff. But overall, when you start stacking these improvements, you're looking at multiples. You can look at it. The most common form, of course, is TPS, tokens per second, which is bad naming by us in the industry, ‘cause there's two tokens per second. There's tokens per second, the throughput number, and the latency number.Ali [00:34:31]: TTMT, yeah.Philip [00:34:32]: Like total tokens per second out of the, out of the GPU as a throughput number. Most people only care about tokens per second as the latency number, which we should call ITL, intertoken latency, but we don't.Philip [00:34:44]: Anyway, so you can imagine a standard API without many optimizations for a 1 trillion parameter model operating somewhere in the 30 to 50 tokens per second range for reasonable traffic profile. And we generally see the goal of, pushing to 10X that. But, not necessarily day zero, but by stacking enough optimizations, if you have, say like four optimizations, each of which doubles performance. Or sorry, three optimizations, each of which doubles performance, then you stack that up, that's an 8X gain. That's the order of magnitude that we're working with in this space. We're trying to make things substantially faster, not just go from like 70 to 90.Swyx [00:35:38]: Are you saying you've. You have done that?Philip [00:35:40]: So let's say you have as a reasonable baseline, 30 or 40 tokens per second. You can achieve 10X that. So like on GLM-5.2, if you run it unquantized, perhaps on H100s even, and you're just using an off-the-shelf inference engine with no particular optimizations, no speculator, nothing extra around like KV routing, no disaggregation, you're, you're probably, yeah, looking at that like 30 to 40. You think that's like a reasonable baseline?Swyx [00:36:12]: Right. Right.Philip [00:36:12]: To get to something like 10X, there's a lot of trade-offs that you're making. If we're running at more like a 300, 400 tokens per second range, you are using the best hardware possible. You have a optimized speculator. You have done all of your quantization work. You are Seeing a pretty high cache hit rate. You are running with a reasonably small batch size and a parallelism configuration that is tuned for latency versus throughput, but it is possible. So the spreads that you see if you, like, go on artificial analysis or you go on OpenRouter and you look at, the worst provider to the best provider, oftentimes can hit that range. 10X is of course very aggressive. It's oftentimes maybe more of a four to six times improvement. But that's the performance that makes us really excited, is when we can get these huge gains, not just go from 70 to 90 tokens.Stacking Optimizations: NVFP4, Speculation, and DisaggregationAli [00:37:19]: It's also, like, hardware dependent. Like, ifPhilip [00:37:20]: YeahAli [00:37:20]: If you have a thing where you're serving it on just, like, a node of H100s and then you throw, like, you shard the model across, like, four nodes of B200s. Like, you can definitely increase the speed with just throwing more hardware at it. Like, normalizing for the same exact hardware and the same number of GPUs.Philip [00:37:35]: Yeah. Then you're looking at, like, a two to 4X improvementAli [00:37:38]: Right. RightPhilip [00:37:38]: Depending on the inference optimizations. So yeah, it's. Some of it's, what's the call, and some of it's who's the driver.Vibhu [00:37:46]: If you break down the two to 4X, say the example is run GLM-5.2Ali [00:37:51]: YeahVibhu [00:37:51]: On B200sAli [00:37:53]: YeahVibhu [00:37:53]: Single node, right? What's, like, the cost trade-off for effort to get, like, the last bit of juice out versus what should people just think of, right?Ali [00:38:01]: Spectre quantization. Yeah.Vibhu [00:38:03]: Spectre quantization.Ali [00:38:04]: That's, that's, that's like 95%. LikeVibhu [00:38:06]: And how far does that get you? And how easy is that for the average person to do? So say right I wanna throw the weights of GLM-5.2 on a node of B200s, how easy is it to find speculative decoder- decoder model or already quantized model? How much work goes into it?Philip [00:38:23]: If you're doing it up front, it's quite a lot of work. If you're doing it today, there's going to be people who have published things that you can just, you can just grab some NVFP4 weights. You can grab a speculator. Yeah, if we're thinking about, like, what are the 2Xs we're stacking, going from, BF16 to NVFP4 is, it's not quite a 2X, right? It's like. I think it's about, like, 30 to 40%, from 16 to 8, and then another 30 to 40% multiplied from, 8 to 4. So that doesn't quite get you a 2X, but, like, roughly a 2X. Speculator, roughly a 2X. Disagg on top of that if you're able to get enough hardware and put enough traffic through it, another roughly a 2X. And then you add in some, double-digit percent increase from having just a better runtime with, the latest kernels and stuff behind it. And that's how it stacks up.Ali [00:39:21]: YeahPhilip [00:39:21]: So building each of those, like, building the, quantized weights is, for someone who really knows what they're doing, hours to days of work. Building the speculator, again, like, hours to days of work. And the, disagg setup, hours to days. Well okay, but like once you haveAli [00:39:39]: Once set up. Once set up. YeahPhilip [00:39:40]: Yeah, getting disagg working for the first time, I'm saying, of course, is very difficult.Philip [00:39:44]: The marginal implementationAli [00:39:48]: Like, if you're just grabbing, like if you are a person, like just a normal consumer who has access to, like, a node of B200s and you're wondering, “How can I just host it myself?” You don't need to quantize the model yourself. There's always gonna be, like, an open source quantized checkpoint. NVIDIA's gonna push one out if no one else does. You. Usually, the providers will have their own spec dec that they've trained as well. You don't need to train your own spec dec. You can just use that as well.Philip [00:40:09]: Yeah. Like, GLM-5.2 has its own MTP.Ali [00:40:13]: Right. Right.Vibhu [00:40:14]: What's multi token prediction?Philip [00:40:15]: Yes.Ali [00:40:16]: I'm justVibhu [00:40:16]: Can you explain that?Ali [00:40:16]: I'm just an expert.Ali [00:40:18]: I can do it for you in case I get it wrong?Vibhu [00:40:20]: No.Vibhu [00:40:21]: Yeah, you should correct if we're wrong, but their multi-token prediction can be used for self-speculative decoding.Ali [00:40:27]: I'm not sure. I'm not gonna correct that.Vibhu [00:40:28]: Okay. I'm semi-confident in thatAli [00:40:30]: Okay. YeahVibhu [00:40:30]: But someone can check. But it's useful to paint the story of, okay, not just the average person, but say a company wants to switch from serverless inference I wanna throw this up on. I wanna rent some GPUs, throw it up. These are the steps you take to do significantly faster than just put it behind vLLM.Ali [00:40:48]: Right.Vibhu [00:40:49]: I was waiting for a mention of Dynamo.Vibhu [00:40:51]: I feel like, that's supposed to be the baseline that you measure against.Dynamo, KV Routing, and Disaggregation ToolkitsPhilip [00:40:55]: I would think of Dynamo as less of a box system and more of a toolkit for building with. So when we talk about doing aware routing, when we talk about doing KV offloading, when we talk about doing, PD disaggregation, Dynamo fundamentally is. By the way, Dynamo is an open source library from NVIDIA.Ali [00:41:17]: We've done a pod with KylePhilip [00:41:18]: OkayAli [00:41:19]: Kyle Cranin.Philip [00:41:19]: Cool. So then your listeners know then that it supports all the different inference frameworks. And it is multi hardware, which is interesting.Ali [00:41:28]: But it's just a router, it's not like an optimizer layer.Philip [00:41:30]: Yeah. All it does, like, what Dynamo is good at, it is a library for moving information around your cluster, around your hardware. So if you have, KV cache on one place and you need it to be somewhere else, Dynamo coordinates NIXL for you to move that around.Philip [00:41:49]: That doesn't mean that, like, out of the box, you just say, “Pip install Dynamo,” and then you get, like, a massive performance speed up. It's more of a developer toolkit.Ali [00:42:01]: Yeah. I would have said it would. It comes with a set of defaults that you can then swap out.Philip [00:42:06]: It does. If the industry at large, I think, was, like, rolling out all of these deployments, standard, then I think it would be, like, a credible baseline. But, we've got to, we've got to benchmark against, like, what we're seeing in the wild.Speculative Decoding Methods: Medusa, EAGLE, n-Gram, and Spec-SpecVibhu [00:42:23]: I did wanna talk a little bit more about PD disagg, because that is probably, like, number three after quantized and speculative decoding. In your book though, I was just gonna pull out the book.Philip [00:42:31]: Yeah.Vibhu [00:42:32]: Like section 522 on Medusa, 523 on EAGLEPhilip [00:42:35]: YeahVibhu [00:42:36]: 524 on gram.Philip [00:42:37]: It's 55, would be disaggregationAli [00:42:42]: Yeah. Well, no, I just wanted to dwell a little bitPhilip [00:42:44]: YeahAli [00:42:44]: The other. Like, so what do you choose to include? What do you choose to not to include? Because there was all these other techniques.Philip [00:42:51]: Yeah.Ali [00:42:51]: Are these still relevant? Because I think they came out, like, a year and a half ago maybe.Vibhu [00:42:55]: Medusa is quite old.Philip [00:42:56]: Yeah, Medusa's old.Ali [00:42:58]: It was old.Vibhu [00:42:58]: But is it in the book as a good, here'sPhilip [00:43:01]: BaselineVibhu [00:43:01]: Baseline vanilla understand it?Philip [00:43:02]: Like you should know this.Vibhu [00:43:03]: Like I read the paper, I'm like, “ it makes so much sense.”Philip [00:43:05]: Yeah.Philip [00:43:05]: So with the book, I had a couple goals. One was to give people just a working vocabulary for the space as a whole, and the other was to give them some intuition about how each of these techniques works. As I mentioned in my AI Engineer talk, which is the first public addendum to this, the speculation space has moved much faster than everything else. So yeah, even at the time that I wrote the book Medusa, I very much included as a way for people to understand how the space evolved rather than what the most modern technique is. And now of course, there's DFlash, dSpark. There's, there's newer techniques even than EAGLE, although EAGLE is still very commonly used.Ali [00:43:51]: SpecSpecta.Philip [00:43:52]: Yes. Speculative decoding.Vibhu [00:43:54]: What canAli [00:43:56]: Oh, it's a paper by Tri Dao and it's like, it's doing speculative decodingVibhu [00:44:00]: HuhAli [00:44:01]: For the speculative decoder.Philip [00:44:02]: Oh, in spec- oh my God.Ali [00:44:02]: It's literally just an another. It's like, yeah, that's the most simple way to explain it, and it seems like he got trivial speed ups there. But it seems that the complexity with training, it's almost like in our mind at least, it's almost as complex as training GANs. Like it's like a very delicate balance and oftentimes you, it's just but yeah, it's literally speculative decoding on speculative decoding.Vibhu [00:44:21]: Speculative.Ali [00:44:22]: Yeah. We saw this paper.Vibhu [00:44:24]: It's interesting, right?Ali [00:44:24]: Yeah.Vibhu [00:44:24]: I wouldn't even expect it to be very particular to train, I wouldAli [00:44:29]: Right.Vibhu [00:44:29]: The naive part of me is like, okay, train speculative decoder.Ali [00:44:32]: But like, and it makes sense, like the whole idea of speculative decoding is you. It's like, it's like almost like the iPhone auto predict version but for a normal model, right? Like you're just, you're just, generating three tokens and you're like, okay, I'll do prefill on them. And so you save those three turns for your original model. Now your speculative decoder is doing three turns of auto regression, so why not just have an even smaller model?Ali [00:44:53]: The other question there is what are the size of speculators? So say forPhilip [00:44:58]: Right. It's like a billion parameters.Ali [00:45:01]: Like for MiniMax, it's. Yeah. It's like one layer. It's like one 60th of the original model usually.Philip [00:45:06]: Yeah. I think we should do a paper when we get back to the office.Philip [00:45:10]: SpeculativeAli [00:45:11]: SpeculativePhilip [00:45:11]: Decoding.Ali [00:45:13]: No, it's, it does seem like how, when do you stop? But then it also seems like if you're able to train spec-spec decode for instance, right? Like if you're able to have a small model that is accurately predicts what the intermediate speculator is gonna predict, that is able to predict what the original target model's gonna predict, then why not just use that smallest model directly, right?Vibhu [00:45:34]: Yeah. This isAli [00:45:35]: Like it seems likeVibhu [00:45:35]: Adjacent to the routing problem.Ali [00:45:36]: Right.Vibhu [00:45:36]: Yeah.Ali [00:45:36]: Right.Philip [00:45:37]: The thing with speculators is one of the practical constraints on using them is that you do have to run a small model on the same hardware that you're running the big model on. There is a orchestration and resource competition problem inherent in that, and that is one of the constraints on speculation in general, is that draft tokens cost resources to create and cost software complexity to manage. And so if you have like infinitely recursive speculators, you add in quite a bit of that complexity on the actual implementation within the inference engine as well, not just in the training process.Vibhu [00:46:17]: I was gonna say, I would wonder if you could do similar, like distillation and pruning of, it's the same thing, it's just a model. Can we not just distill a lot of the weights, quantize the speculator, out of my domain? The question that also comes up is, this is all for big server workloads, right? How much of this applies to, say I have this MacBook, I wanna run Gemma really efficiently. Similar problems, not the same?Local AI vs. Data Center InferencePhilip [00:46:45]: Pretty different. I talked to Selo, about this on his podcast a couple weeks ago. The difference between inference engineering for the data center and for production workloads versus inference engineering for local AI, is that we start with fundamentally like different constraints and different goals. With local AI, it's how do I fit this model onto my hardware and then make it less dumb? And with data center influence, it's how do I load this model and then make it less slow? And we care about less dumb, and they care about less slow. But the local AI inference engineering ecosystem, I think has a lot for us to learn from in the data center space. They are experts in various forms of quantization, including dynamic quantization that we just don't touch, in the pruning, in the distillation, in the, layer removal. There'Ali [00:47:42]: Layer removal matters less.Philip [00:47:43]: Yeah. There'Ali [00:47:44]: No one loves pruning really.Philip [00:47:45]: Yeah. Well, but the, but they doVibhu [00:47:46]: Which is surprising, right? But that's, that's a whole different thingPhilip [00:47:48]: Just to fit something on the laptop.Ali [00:47:50]: Right.Philip [00:47:50]: So yeah, it's a, it's an interesting, it's an interesting space. Not necessarily that like their techniques make sense for us to do in the data center, because we have different resources and different goals, but more that the process as well as the openness of that field is something to, admire.Ali [00:48:12]: Yeah. Like to your point, like, certain optimizations that would. Like for instance, Turbo Quantum Sharper, like it made such huge hype on that and we did like a whole deep dive on Twitter and like said, what is it? How does it work? Why is it good or not? And it took off and it was implemented on local devices because your memory bandwidth is so slow on like a MacBook, for instance. But try putting the same thing on like an NVIDIA GPU on a B200 Turbo quant would not be. Like, it would not be used. Like, NVIDIA - Like, NVIDIA made it clear that this is not a good optimization, and we've seen it firsthand where the overhead of doing dequantization, quantization of, in the kernel itself with turbo quant kernel, each end is much slower than the time that you save from doing the bandwidth. ‘Cause on the B200s, you have like 3.5 terabytes per second. You don't need decrease the storage that much. You don't need to do, FP4 KV cache. You don't need to use a requant. There's, there's, there's better optimizations to be made. But on Edge devices, it's extremely important, it's extremely useful. So, seems to be, like, different optimizations there, but then they're all uniquely combined with like all you wanna quantize the model, you wanna do speculative decoding, like certain common prefixes with bothPhilip [00:49:18]: Principles.Ali [00:49:19]: Yeah, exactly. Exactly. Exactly.Philip [00:49:20]: They also do a lot of work on, model parallelism, especially over, heterogeneous topology, where you have, some sparks and they are wired together with, Ethernet, DGX sparks.Ali [00:49:35]: Yeah, this is the Exo Labs guys.Philip [00:49:36]: Yeah. You have, a nu

Yeni Şafak Podcast
Taha Kılınç-Japon Ahmed'in öyküsü

Yeni Şafak Podcast

Play Episode Listen Later Jul 31, 2026 5:00


Geçtiğimiz cuma günü -24 Temmuz 2026- Lübnan'ın başkenti Beyrut'ta, sıra dışı bir cenaze töreni vardı: Uzun yıllardır Beyrut'un kenar mahallelerinde yaşamını sürdüren ve bir gün önce 78 yaşında vefat eden Kozo Okamoto adlı bir Japon'un Filistin bayrağına sarılı naaşı Filistinli gençlerin omuzlarında taşındı ve Filistinli şehitlerin kabirlerine ev sahipliği yapan Makbaratu'ş-Şuhedâ'da (Şehitler Kabristanı) toprağa verildi. Kozo Okamoto'nun öyküsü, olağanüstü ayrıntılarla ve sürprizlerle dolu bir Ortadoğu hikâyesiydi:

Qur'an Conversations
S4 E21: The More You Return, the More It Gives (TaHa 127) | Quran Conversations Season 4

Qur'an Conversations

Play Episode Listen Later Jul 30, 2026 64:02


What does it actually mean to build a relationship with the Qur'an, beyond simply reading its words?In this episode of Quran Conversations, Dalia Mogahed is joined by Dr. Hatim Yousef to reflect on how we can become true companions of the Qur'an, before turning to verse 127 of Surah TaHa and its warning about repeatedly turning away from Allah's signs.Dr. Hatim Yousef is an Islamic scholar, educator, and linguist who studied Islamic jurisprudence, Quranic sciences, tajweed, and hifdh under traditional scholars in the UAE and Jordan. He has memorized the Quran and holds several ijazas. Since moving to the United States in 2008, he has served as an Imam and instructor at the ADAMS Center in Virginia, teaching tafsir, hadith, seerah, Arabic grammar, and tajweed. Fluent in Arabic and English, Dr. Hatim also holds a degree in English and Translation, an MA in English and Literature, and a PhD in Linguistics.Dr. Hatim shares practical advice for anyone wanting to understand and reflect on the Qur'an more deeply. It begins with something simple: spending time with it. Reciting it every day, learning its language, studying with knowledgeable teachers, and allowing that relationship to grow consistently over time. As he explains, the more time you spend with the Qur'an, the more the Qur'an gives you.In this episode, you will learn:

Yeni Şafak Podcast
Taha Kılınç - Utanmadınız mı?

Yeni Şafak Podcast

Play Episode Listen Later Jul 29, 2026 4:58


Çin yönetimi, Türkiye'den bir grup “gazeteci”yi yine Doğu Türkistan'a götürmüş; yedirmiş, içirmiş, gezdirmiş. Son zamanlarda gittikçe yoğunlaşan bu türden programların amacı elbette “bölgede herhangi bir problemin olmadığını” ispat gayreti. (İsrail'in sizi alıp Kudüs'e, Gazze'ye, Râmallah'a götürmesi, gezdirmesi, yedirip içirmesi; bu sırada da “terörle ve radikallikle mücadele” konusunda attığı adımları anlatması gibi düşünün.)

Qur'an Conversations
S4 E20: The Blindness We Don't See (TaHa 124–126) | Quran Conversations Season 4

Qur'an Conversations

Play Episode Listen Later Jul 24, 2026 55:26


What if the greatest form of blindness isn't losing your sight—but losing your connection with Allah?In this episode of Quran Conversations, Dalia Mogahed is joined by Imam Mohamed Magid to reflect on verses 124–126 of Surah TaHa.Allah warns that whoever turns away from His remembrance will experience a life of constriction and hardship—not necessarily through a lack of wealth or comfort, but through an absence of inner peace. These verses invite us to reconsider where true happiness comes from, how remembrance shapes our identity, and why spiritual blindness is far more dangerous than physical blindness.Together, Dalia and Imam Mohamed explore the meaning of dhikr, the emptiness that nothing in this world can fill except Allah, and how making Allah the center of our lives transforms the way we navigate hardship, success, and everything in between.In this episode, you will learn:

Yeni Şafak Podcast
Taha Kılınç - Bir bereket ocağı

Yeni Şafak Podcast

Play Episode Listen Later Jul 24, 2026 5:33


Emevî Camii'nin batı tarafında, Memlûk Sultanı Zâhir Baybars'ın kabrine ev sahipliği yapan Zâhiriyye Medresesi'nin tam karşısındaki çarşıya giriyoruz. Birkaç dakika boyunca elektronik eşya dükkânlarının, oyuncakçıların, baharatçıların, ayakkabıcıların, ayaküstü atıştırmalık satan mekânların, kitapçıların önünden geçtikten sonra Dârulhadis en-Nûriyye'nin kapısını görüyoruz. 1170 tarihli bu yapı, İslâm tarihinde hadis ilmine has okul anlamındaki “dârulhadis” müesseselerinin ilk örneği. İsmini, kurucusu Nûreddîn Mahmûd Zengî'den alıyor. Çarşıda yürümeye devam ediyoruz. Çünkü hedefimiz, başka ve çok daha ünlü bir dârulhadis: Eşrefiyye.

Les Nuits de France Culture
Taha Hussein, doyen des lettres arabes et ami de la France

Les Nuits de France Culture

Play Episode Listen Later Jul 23, 2026 64:55


durée : 01:04:55 - Les Nuits de France Culture - par : Albane Penaranda - En mars 1974 sur France Culture, la journaliste Hélène Tournaire rend hommage au grand universitaire et romancier égyptien Taha Hussein, quelques mois après sa disparition. Avec la participation de Maurice Couve de Murville, Jacques Berque et Moenis Taha Hussein, fils de l'écrivain. - équipe : Mathias Le Gargasson, Antoine Dhulster, Rafik Zénine, Vincent Abouchar, Emily Vallat, Hassane M'Béchour, INA Vous aimez ce podcast ? Pour écouter tous les épisodes sans limite, rendez-vous sur Radio France

Musik ist Trumpf
Zu Gast: Taha von Kenitra!

Musik ist Trumpf

Play Episode Listen Later Jul 22, 2026 80:20


In der letzten Folge vor der Sommerpause ist Taha von Kenitra ist zu Gast bei Musik ist Trumpf. Kenitra, benannt nach Tahas Heimatstadt in Marokko, ist eine Band aus Münster deren neue Single „Go home“ es Till & Henning sehr angetan hat. Mit Taha reden sie über Musik, Heimat, Integration und Assimilation…und die Musik, die Taha geprägt hat. Die Songs zur Folge: 1) Amsterdam / Jaques Brel2) I think of you / Rodriguez3) Hold on / Angus & Julia Stone4) I / Kendrick Lamar5) Ya Rayah / Rachid Taha Sponsor:Der Guitar Summit 2026 findet auch dieses Jahr in Mannheim statt vom 25.9.-27.9!Workshops, Masterclasses, Konzerte und eine groß Auswahl von Ausstellern warten auf Euch! Auf der Rockantenne Stage gibt es viele Interviews mit Stars & Menschen aus der Gitarrenwelt. Natürlich ist Till auch wieder am Start mit Till & Tone und Musik ist Trumpf! Sichert euch jetzt das 3 Tagesticket mit 10% Rabatt: geht einfach zu www.guitarsummit.de (http://www.guitarsummit.de/) und bucht dort das 3 Tages-Ticket. Dann den Code TRUMPF10 eingeben und 10% mit Musik ist Trumpf sparen! Hosted on Acast. See acast.com/privacy for more information.

Yeni Şafak Podcast
Taha Kılınç - Tatlı bir Şam akşamı

Yeni Şafak Podcast

Play Episode Listen Later Jul 22, 2026 5:17


Uzun ve yorucu hastalıklar geçiren bir yakınınızın adım adım iyileştiğini gördüğünüzde ne hissederseniz, ben de Şâm-ı Şerîf'e her gidişimde aynı hislerle doluyorum. 2024'ün son günlerinden itibaren bir ayağım sürekli orada ve her seferinde istikbale dair umutlarım artmış olarak İstanbul'a dönüyorum. Geçtiğimiz cumayla pazartesi arasında (17-20 Temmuz) dört dostumuzla birlikte gerçekleştirdiğim seyahat de bu şekildeydi, hamd olsun. Anlatacak çok şey var, ama bu yazıyı güzel bir tevafukun üzerine bina edeceğim.

Yeni Şafak Podcast
Taha Kılınç - Ne yüzle?

Yeni Şafak Podcast

Play Episode Listen Later Jul 18, 2026 5:33


Bazı konular var, köşe yazısına dönüştürmek üzere bir kenara ayırıyorum; sonra gündem değişiyor veya daha öncelikli bazı şeyleri yazmak icap ediyor, ama konu kendi iç dünyamda güncelliğini bir türlü yitirmiyor. Hatta bazen öfkemi bile kabartıyor, yazılmadan orada öylece durdukça. Mesela, Fransa Cumhurbaşkanı Emmanuel Macron'un, NATO zirvesi için Türkiye'ye gelmeden önce 6-7 Eylül'de gerçekleştirdiği Suriye ziyareti sırasında sarf ettiği şu cümleler tam da bu kabilden: “Ülkelerimiz arasında uzun yıllara dayanan çok boyutlu ilişkiler var. Buna binaen Suriye'de, önceki rejim döneminde kapatılan Fransız-Hristiyan okullarını yeniden açmak istiyoruz. Bu okullar, Suriye çapında genç kız ve erkek öğrencilere eğitim verecek.”

Qur'an Conversations
S4:E19 - When Your Mistake Becomes Your Turning Point (TaHa 121-124) | Quran Conversations Season 4

Qur'an Conversations

Play Episode Listen Later Jul 17, 2026 68:54


What if your greatest mistake became the very thing that brought you closer to Allah? In this episode of Quran Conversations, Dalia Mogahed is joined by Imam Kebba Sallah to reflect on verses 121–123 of Surah TaHa. Imam Kebba Sallah was born and raised in the Washington D.C. area. He is the community Imam for the ADAMS Sully branch in Chantilly, Virginia. Imam Kebba has completed memorization of the Holy Quran and has ijaazah in Qiraat Hafs An Assim. He has studied Islamic sciences in the traditional method with local and international teachers. He is a student of Imam Mohamed Magid and other U.S. based scholars. He studied classical Arabic language at Qasid Institute in Amman, Jordan and continues to pursue advanced Arabic mastery. As the story of Adam (AS) continues, these verses move beyond the mistake itself to reveal something even more profound: Allah's mercy after the fall. Adam's story reminds us that failure is never the end. What defines us is not whether we make mistakes, but how we respond to them. Together, Dalia and Imam Kebba explore the symbolism of clothing in the Qur'an, the connection between taqwa and spiritual protection, the role of regret in sincere repentance, and why our mistakes can become the very means through which Allah elevates us. They also reflect on the lifelong battle against Shaytan and the peace that comes from following Allah's guidance. In this episode, you will learn: 

Yeni Şafak Podcast
Mahmut Ay - İnanca, düşünceye ve eyleme ahlâkî bir nefes üflemeye çalışan bilge: Taha Abdurrahman

Yeni Şafak Podcast

Play Episode Listen Later Jul 17, 2026 12:40


Modern dünyanın dayattığı şartların etkisi altındaki İslam dünyası, günümüzde tarihin en derin krizlerinden birini yaşamaktadır. Bu kriz; ekonomik krizlerin, siyasî çalkantıların ya da askerî yenilgilerin çok ötesinde, her şeyi kuşatan varoluşsal bir “ahlak krizi”dir. Çağımızda dinin şekilsel ritüellerinin, zâhirî formlarının, mabetlerinin ve dinî konuşmaların “görünürleşmesine” karşılık sokağın, mabedin, tekkenin, ticaretin, siyasetin ve insan ilişkilerinin içindeki ahlakî özün gittikçe “buharlaşmasına” şahitlik ediyoruz.

TuttoSvenskan
#662 Taha Ali direkt från Grekland | Bjerkebo om ständiga flyttrykten!

TuttoSvenskan

Play Episode Listen Later Jul 17, 2026 78:11


Vi ringer upp allas vår Taha Ali som får berätta om MFF-avskedet och känslorna kring det hela. Men också om första intrycket av PAOK. Isak Bjerkebo om drömmålet mot BP! Bajens Boudri-bomb! Djurgårn på mittfältsjakt, och AIK:s kuppförsök av Sagoe Jr!Medverkande:August Spångberg, Freddie Arnesson, Josip Ladan, Marcus Thapper, Taha Ali & Isak Bjerkebo90MinSvenskan görs i samarbete med:ATG:Läs om våra senaste tankar gällande spel på: https://www.atg.se/k2618 år gäller för spel och stödlinjen.se finns om du upplever minsta problematik med spelande.Tillsammanslag Big 9: https://www.atg.se/tillsammans/lagsid...TV4 Play:TV4 Play Sport Fotboll via 90MinSvenskan! Streama Allsvenskan, Superettan, La Liga och Serie A med vår dunderdeal hos TV4 där ni får 349kr/mån i 6 månader. Signa upp här: https://www.tv4play.se/kampanj/svenskanSociala Medier:Instagram - https://www.instagram.com/90MinSvenskan/X/Twitter - https://x.com/90MinSvenskanTikTok - https://www.tiktok.com/@90MinSvenskanTidskoder:00:00 Intro10:08 Samtal med Taha Ali29:27 Samtal med Isak Bjerkebo41:23 Bajens bud på Amin Boudri53:17 Malmös nya värvning 56:22 Silly runt Djurgården01:00:03 Sagoe till AIK?01:04:20 Ibrahim till Häcken 01:07:00 Helgens matcher 01:15:27 Avrundning Hosted on Acast. See acast.com/privacy for more information.

Qur'an Conversations
S4 E18: The Garden You're Looking For (TaHa 117–120) | Quran Conversations Season 4

Qur'an Conversations

Play Episode Listen Later Jul 16, 2026 59:19


What if the life you've been chasing was never meant to satisfy you?In this episode of Quran Conversations, Dalia Mogahed is joined once again by Ustadha Samia Mubarak to reflect on verses 117–120 of Surah TaHa.As Allah warns Adam about Shaytan's deception, these verses reveal a timeless truth about the human condition. Shaytan's strategy has never changed: convince us that fulfillment lies just beyond what Allah has already given us. The promise of "more" becomes the very thing that robs us of gratitude, contentment, and closeness to Allah.Together, Dalia and Samia explore the illusion of the perfect life, why the Qur'an is the closest thing to Jannah we can experience in this world, and how Shaytan exploits our deepest fears and desires to pull us away from Allah.In this episode, you will learn:

TD Ameritrade Network
SPCX Falls Below IPO Price: What's Behind Price Action & Outlook Impact

TD Ameritrade Network

Play Episode Listen Later Jul 16, 2026 8:56


Taha Ahmed talks about Wall Street's largest IPO, SpaceX (SPCX), after shares of the Elon Musk-led firm came back to Earth, now trading below the stock's initial IPO price of $135. He points to the company's future aspirations as the crux of bullish momentum, meaning any hit to its outlook will create ripple effects in the stock. Taha believes SpaceX needs time for valuations to come into check but sees the AI trade offering the next step toward long-term profitability. ======== Schwab Network ========Empowering every investor and trader, every market day.Subscribe to the Market Minute newsletter - https://schwabnetwork.com/subscribeDownload the iOS app - https://apps.apple.com/us/app/schwab-network/id1460719185Download the Amazon Fire Tv App - https://www.amazon.com/TD-Ameritrade-Network/dp/B07KRD76C7Watch on Sling - https://watch.sling.com/1/asset/191928615bd8d47686f94682aefaa007/watchWatch on Vizio - https://www.vizio.com/en/watchfreeplus-exploreWatch on DistroTV - https://www.distro.tv/live/schwab-network/Follow us on X – https://twitter.com/schwabnetworkFollow us on Facebook – https://www.facebook.com/schwabnetworkFollow us on LinkedIn - https://www.linkedin.com/company/schwab-network/About Schwab Network - https://schwabnetwork.com/about

SBS Urdu - ایس بی ایس اردو
'I only did my duty': Bondi Junction attack hero Muhammad Taha awarded Australia's third-highest civilian honour, the Bravery Medal - "میں نے صرف اپنا فرض نبھایا" — بونڈائی جنکشن حملے کے ہیرو محمد طہ

SBS Urdu - ایس بی ایس اردو

Play Episode Listen Later Jul 14, 2026 11:58


Muhammad Taha, who confronted the attacker during the deadly 2024 Bondi Junction stabbing, was recently presented with the Bravery Medal by the Governor-General. The award is Australia's third-highest civilian honour for bravery. In an exclusive interview with SBS Urdu, Taha recalled the horrific events of that day, spoke about helping his injured colleague and reflected on the support he received from the community. - بونڈائی جنکشن میں 2024 کے مہلک چاقو حملے میں مسلح شخص کا مقابلہ کرنے والے محمد طہٰ کو حال ہی میں گورنر جنرل کے ہاتھوں بریوری میڈل سے نوازا گیا ہے، جو آسٹریلیا کا تیسرا بڑا سولین اعزاز برائے بہادری ہے۔ SBS اردو سے خصوصی گفتگو میں طہٰ نے اس ہولناک دن کی یادیں، اپنے زخمی ساتھی کی مدد اور کمیونٹی کی جانب سے ملنے والی حمایت کا ذکر کیا۔______________ہر بدھ اور جمعہ کا پورا پروگرام اس لنک پرسنئے , اردو پرگرام سننے کے دیگر طریقے SBS Audio کے نام سے موجود ہماری موبائیل ایپ ایپل (آئی فون) یا اینڈرائیڈ , ڈیوائیسزپرانسٹال کیجئے۔ پوڈکاسٹ سننے کے لئے نیچے دئے اپنے پسندیدہ پوڈ کاسٹ پلیٹ فارم کا انتخاب کیجئے:  Spotify Podcast , Apple Podcasts - ______________ہر بدھ اور جمعہ کا پورا پروگرام اس لنک پرسنئے , اردو پرگرام سننے کے دیگر طریقے SBS Audio کے نام سے موجود ہماری موبائیل ایپ ایپل (آئی فون) یا اینڈرائیڈ , ڈیوائیسزپرانسٹال کیجئے۔ پوڈکاسٹ سننے کے لئے نیچے دئے اپنے پسندیدہ پوڈ کاسٹ پلیٹ فارم کا انتخاب کیجئے:  Spotify Podcast , Apple Podcasts

Yeni Şafak Podcast
Taha Kılınç - El-Emîru'l-Vâlid

Yeni Şafak Podcast

Play Episode Listen Later Jul 14, 2026 5:08


Katar eski Emiri Şeyh Hamed bin Halîfe Âl-i Sâni, 74 yaşında hayata gözlerini yumdu. Son dönemde sağlık sorunları yaşadığı bilinen Şeyh Hamed, tahtını 2013 yılında oğlu Şeyh Temîm'e devretmişti. O tarihten günümüze, unvanı “El-Emîru'l-Vâlid” idi, yani Baba Emir. Şeyh Hamed, Katar halkının her anlamda babasıydı. Örneğin, kendisine karşı darbe girişiminde bulunan bir grup askerin ailelerine maaş bağlatıp “Eşlerinin ve çocuklarının ne suçu var? Onlar bize emanet” diyecek kadar âlicenaptı.

Masty o Rasty | پادکست فارسی مستی و راستی

https://www.instagram.com/taha.mazaheri/Taha Mazaheri is a guitarist who currently plays in King Raam's band. This is part 2 of his story. contact me at https://t.me/queenraaminfo@kingraam.comTo learn more about psychedelic therapy go to my brother Mehran's page at: https://www.mindbodyintegration.ca/ or to https://www.somaretreats.org for his next retreat.https://www.instagram.com/ravannamacommunity/https://www.somaretreats.org***Masty o Rasty is not responsible for, or condone, the views and opinions expressed by our guests ******مستی و راستی هیچگونه مسولیتی در برابر نظرها و عقاید مهمان‌های برنامه ندارد.***Support the showhttps://paypal.me/raamemamiVenmo + Revolut: @KingRaamhttps://www.instagram.com/kingraam Hosted on Acast. See acast.com/privacy for more information.

Medyascope.tv Podcast
Türkiye'de servet neden sermayeye dönüşemiyor? | Taha Akyol anlatıyor | Sağduyu

Medyascope.tv Podcast

Play Episode Listen Later Jul 8, 2026 48:58


Servet ile sermaye arasındaki fark nedir? Türkiye neden güçlü sermaye üretemiyor? Sağduyu programında Tarık Çelenk, konuğu gazeteci ve yazar Taha Akyol ile Osmanlı'dan Cumhuriyet'e uzanan süreçte Türkiye'nin ekonomik dönüşümünü, girişimcilik kültürünü, burjuvaziyi, Anadolu sermayesini ve kalkınmanın önündeki tarihsel engelleri konuşuyor. Programda; servet-sermaye ayrımı, Sabri Ülgener'in görüşleri, Osmanlı ekonomisi, İttihat ve Terakki'nin milli iktisat politikası, Cumhuriyet'in ekonomik temelleri, MÜSİAD-TÜSİAD tartışmaları, muhafazakâr sermaye, girişimcilik kültürü ve Türkiye'nin neden nitelikli sermaye üretemediği kapsamlı şekilde değerlendiriliyor. Learn more about your ad choices. Visit megaphone.fm/adchoices

Qur'an Conversations
S4 E17: The Difference Between Adam and Iblis (TaHa 115-116) | Quran Conversations Season 4

Qur'an Conversations

Play Episode Listen Later Jul 3, 2026 69:11


Why do some mistakes bring us closer to Allah, while others pull us further away?In this episode of Quran Conversations, Dalia Mogahed is joined by Dr. Hatim Yousef—Islamic scholar, linguist, Qur'an teacher, and translator—for a reflection on verses 115-116 of Surah TaHa.Drawing on his background in linguistics, Qur'anic sciences, and tafsir, Dr. Hatim explores the story of Adam and Iblis through both the language of the Qur'an and its spiritual lessons. Together, he and Dalia reflect on the beauty of the Qur'an's sounds, the honour Allah bestowed upon humanity, the danger of arrogance, and the transformative power of sincere repentance.This conversation reminds us that the Qur'an is not simply a book to be understood intellectually—it is a living conversation with Allah, meant to heal, awaken, and transform the heart.In this episode, you will learn:

Masty o Rasty | پادکست فارسی مستی و راستی

https://www.instagram.com/taha.mazaheri/Taha Mazaheri is a guitarist who currently plays in King Raam's band. In this episode he tells the story of how he became a musician in Iran and how he developed his sound and technique. contact me at https://t.me/queenraaminfo@kingraam.comTo learn more about psychedelic therapy go to my brother Mehran's page at: https://www.mindbodyintegration.ca/ or to https://www.somaretreats.org for his next retreat.https://www.instagram.com/ravannamacommunity/https://www.somaretreats.org***Masty o Rasty is not responsible for, or condone, the views and opinions expressed by our guests ******مستی و راستی هیچگونه مسولیتی در برابر نظرها و عقاید مهمان‌های برنامه ندارد.***Support the showhttps://paypal.me/raamemamiVenmo + Revolut: @KingRaam https://www.instagram.com/kingraam Hosted on Acast. See acast.com/privacy for more information.

Qur'an Conversations
S4 E16: The Small Deeds That Save You (TaHa 113-114) | Quran Conversations Season 4

Qur'an Conversations

Play Episode Listen Later Jun 26, 2026 51:21


How much is enough to earn Allah's pleasure? And could the smallest act of sincerity outweigh a lifetime of mistakes?In this episode of Quran Conversations, Dalia Mogahed and Sheikh Mohammed Magid reflect on verses 113-114 of Surah TaHa, where Allah reassures the believers that no righteous deed will ever be lost and reveals His extraordinary generosity in rewarding even the smallest acts of goodness.Their discussion explores the vast difference between Allah's justice and His mercy. While every bad deed is counted as one, every sincere good deed is multiplied many times over. They reflect on why our worries, grief, and emotional struggles are never wasted, how ordinary moments of daily life become acts of worship through intention, and why sincerity matters more than quantity.The conversation then turns to the Qur'an itself—why it was revealed in Arabic, why its language is part of its miracle, and how engaging with the Qur'an transforms not only what we know, but who we become.In this episode, you will learn:

#Sillypodden
Jävlades Alexander Isak med Taha Ali?

#Sillypodden

Play Episode Listen Later Jun 22, 2026 19:29


Det var deppig stämning på Sveriges träning efter förlusten mot Nederländerna. Den blev gladare efter Taha Alis nya presskonferens-tabbe. Mer från avsnittet: *Är stjärnans hemliga team avslöjat? *Graham Potter har lindat spelarna runt sitt finger *Förbundsbossen om buropen mot Infantino Medverkande: Per Bohman, Malin Wahlberg och Linn Nordström Kontakt: podcast@aftonbladet.se Ansvarig utgivare: Lotta Folcker

Qur'an Conversations
S4:E15 - The Good Deed You Thought Didn't Matter (TaHa 112–113) | Quran Conversations Season 4

Qur'an Conversations

Play Episode Listen Later Jun 19, 2026 51:21


Can one sincere act change your entire standing before Allah?In this episode of Quran Conversations, Dalia Mogahed is joined by Imam Mohamed Magid to reflect on verses 112–113 of Surah TaHa.After discussing accountability and the consequences of wrongdoing, these verses shift our attention to hope. Allah reassures believers that no righteous deed will ever be lost, overlooked, or diminished. Even the smallest act of goodness—done with sincerity and faith—can carry immense weight with Allah.The conversation explores Allah's generosity in multiplying good deeds, the hidden rewards behind everyday acts of service, and the countless opportunities Allah places in our lives to draw closer to Him. Dalia and Imam Magid also reflect on the Qur'an as a living miracle, the significance of its Arabic language, and why the Qur'an uses diverse methods of guidance, warning, stories, and reminders to awaken the human heart.In this episode, you will learn:

Qur'an Conversations
S4 E14: The Weight of What We Carry (TaHa 110–111) | Quran Conversations Season 4

Qur'an Conversations

Play Episode Listen Later Jun 12, 2026 51:19


What if the thing you're trying hardest to hide is already completely known to Allah?In this episode of Quran Conversations, Dalia Mogahed is joined by Imam Muhammad Magid to reflect on verses 110–111 of Surah TaHa.These verses shift our attention to two profound realities: Allah's infinite knowledge and humanity's complete humility before Him on the Day of Judgment. Allah knows what came before us, what lies ahead of us, and everything in between—while our own knowledge remains limited and incomplete.Together, Dalia and Imam Magid explore what it means to be truly known by Allah, the power of naming and understanding our experiences, and why the Day of Judgment is ultimately a day of complete exposure, accountability, and justice.In this episode, you will learn:

Knowledge Cast by Enterprise Knowledge
Wael Taha - Vice President of Enterprise Architecture at Brown Brothers Harriman

Knowledge Cast by Enterprise Knowledge

Play Episode Listen Later Jun 11, 2026 32:29


Enterprise Knowledge's Lulit Tesfaye, VP of Knowledge & Data Services, speaks with Wael Taha, VP of Enterprise Architecture at Brown Brothers Harriman. Over the past 12 years, Wael has dedicated his career to architecting D&A solutions and directing D&A professionals in building and evolving modern D&A platforms, delivering significant business benefits and solving problems that were invincible with traditional D&A platforms and techniques.In their conversation, Lulit and Wael discuss Wael's journey from business analytics and client consulting to "embracing data as an enterprise asset," promoting confidence in BBH's data, and how to translate the semantic layer into business value for different audiences, with and without AI as a consideration. They also chat about the key challenges facing data teams these days, focusing on compliance, security, and governance.Note: The views expressed by guests are their own, and do not necessarily reflect the views of their organization.---Comment below to let us know how your organization is embracing data as a knowledge asset, or ask any questions you have for our Knowledge Cast hosts! They may be answered in a future episode.---To learn more about Enterprise Knowledge, visit us at: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠enterprise-knowledge.com⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠.EK's Knowledge Base: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://enterprise-knowledge.com/knowledge-base/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Contact Us: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://enterprise-knowledge.com/contact-us/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠LinkedIn: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.linkedin.com/company/enterprise-knowledge-llc/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Twitter/X: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://twitter.com/ekconsulting⁠⁠⁠⁠Ad Music License: Upbeat Corporate by JP Bianchini https://jpbianchini.comCreative Commons — Attribution 3.0 Unported — CC BY 3.0Free Download / Stream: https://www.audiolibrary.com.co/jp-bianchini/upbeat-corporateMusic promoted by Audio Library https://youtu.be/Hiq53IyRnGI

Bangla Waz
কিয়ামত খুব নিকটে আবু ত্বহা মুহাম্মদ আদনান abu taha muhammad adnan bangla waz 2026 বাংলা ওয়াজ

Bangla Waz

Play Episode Listen Later Jun 11, 2026 35:55


Qur'an Conversations
S4 E13: The Day No One Can Ignore the Call (TaHa 108–109) | | Quran Conversations

Qur'an Conversations

Play Episode Listen Later Jun 5, 2026 59:49


What happens when every illusion of power disappears?In this episode of Quran Conversations, Dalia Mogahed is joined by her teacher, Imam Muhammad Magid, for a reflection on verses 108–109 of Surah TaHa.These verses continue the Qur'an's powerful depiction of the Day of Judgment—a day when every human being who has ever lived will stand before Allah. No one will be able to ignore the summons. No one will be able to resist, delay, negotiate, or escape.As the Qur'an describes humanity following the caller without deviation, Dalia and Imam Magid explore what it means to live in a world where we often ignore Allah's call, yet are heading toward a day when responding will no longer be optional. They reflect on the complete silence of that gathering, the collapse of worldly status, and the profound significance of Allah describing Himself as Ar-Rahman in a scene of overwhelming judgment and power.In this episode, you will learn:

#Sillypodden
Bank om Taha Alis succé: ”Ett stort fuck off”

#Sillypodden

Play Episode Listen Later Jun 4, 2026 22:19


Inside Blågult tar ner det skakiga krysset mot Grekland i VM-genrepet. Det positiva: Taha Alis magiska inhopp. Det negativa: typ allt annat. Medverkande: Per Bohman, Simon Bank och Malin Wahlberg. Kontakt: podcast@aftonbladet.se Ansvarig utgivare: Lotta Folcker Innehåller klipp från Viaplay

Kerem Önder
Kendini ve aileni gelen felaketten koru? - Tahrim 6-7 tefsiri / Kerem Önder

Kerem Önder

Play Episode Listen Later May 29, 2026 60:33


“Ey iman edenler! Kendinizi ve ailenizi, yakıtı insanlar ve taşlar olan ateşten koruyun.O ateşin başında gayet katı, çetin, Allah'ın kendilerine verdiği emirlere karşı gelmeyen ve kendilerineemredilen şeyi yapan melekler vardır.” Tahrim 6“Ey inkâr edenler! Bu gün özür dilemeyin! Siz ancak yapmakta olduklarınızın karşılığını görüyorsunuz.” 7Dünyadaki ateşler taşı yakamaz. Cehennem ateşi nasıl bir ateş ki yakıtı insan ve taş?“Keşşâfda, bu ifadeye, "Günahları terkedip, taatları yapmak ve ailenizi, kendinizi sorumlu tuttuğunuzşeylerle sorumlu tutmanız suretiyle, koruyun" manası verilmiştir. Yine bu ifadeye, "Kendinizi, nefsinizidavet ettiği şeylerden koruyun. Çünkü nefis, size kötü şeyleri emreder" manası verilmiştir."Bu ateşin başında iri gövdeli, sert tabiatlı melekler vardır." Cenâb-ı Hakk bu ifadeyle, on dokuzzebânî (cehennem bekçisi melek) ile onların yardımcılarını kastetmiştir. Bunlar, alabildiğine büyük, haşin vesert meleklerdir. Onların bu şekilde yaratılmış olmaları yadırganacak bir şey değildir. Yahut da onlar, Allah'ındüşmanlarına karşı, alabildiğine şiddetli, Allah dostlarına karşı alabildiğine merhametli oldukları için,hilkatlerinde değil de işlerinde böyle serttirler. Nitekim Hak teâlâ (mü'minleri vasfederken), "(Onlar),kâfirlere karşı alabildiğine sert; birbirlerine karşı ise son derece merhametlidirler" (Fetih 29) buyurmuştur.Ayetteki, "Ne emrolundular ise onu yaparlar" ifadesi. işin gerektirdiği ortamdan ötürü onların, çok çetin vesert olduklarına delâlet eder. Çünkü onlar Allah'ın emirlerini yerine getirme ve düşmanlarından intikamalma hususunda asla şefkatli davranmazlar. Burada, meleklerin, âhirette, Allahü teâlâ'nın emir ve yasaklarıile mükellef olmaya devam ettiklerine bir işaret vardır. Çünkü onlardan sâdır olacak isyan. Allah'ın emir veyasağına muhalefet olur."Ey kâfirler, bugün özür dilemeyin" buyurmuştur ki bu, "Onlara, "Bugün mazeret beyan etmeyin" denilecek"takdirindedir. Çünkü mazeret beyan atmek, tevbe etmek demektir. Tevbe ise, cehenneme girdikten sonra,artık makbul değildir. Binâenaleyh mazeret beyan etmek, onlara bir fayda vermez. yani, "Sizin yapmışolduğunuz o kötü amelleriniz, hikmet-i ilahiyye gereği size bu azabı gerekli kılmıştır."Allahü teâlâ, "Eğer yapamazsanız, ki kesinlikle yapamayacaksınız, o halde yakıtı insanlar ve taşlar olan oateşten korununuz. Zira o, kâfirler için hazırlanmıştır" (Bakara, 24) buyurarak, cehennemin kâfirler içinyaratıldığını bildirmiştir. Öyleyse burada, mü'minlere böyle hitap etmesinin hikmeti nedir? Deriz ki:Fasıkların derekeleri, kâfirlerin derekelerinin üstündedir. Çünkü fasıklar da kâfirlerle birlikte, aynı yerde,yani cehennemdedirler. Bundan dolayı mü'minlere, "Bu ateşin kendileri için hazırlandığı kimselerlebirlikte olmamanız için, fısk-ı fücurdan (günahtan) alabildiğine kaçınmak suretiyle kendinizi bu ateştenkoruyun" denmiştir. Bu hitapla, Cenâb-ı Hakk'ın onlara, mürtedlikten korunmalarını emretmiş olması dauzak bir ihtimal değildir.” RaziAli (radıyallahü anh), Katade ve Mücahid şöyle demişlerdir: Yaptığınız işlerle kendinizi koruyunuz, onlarayapacağınız tavsiyelerle de aile halkınızı koruyunuz.el-Kuşeyrî'nin zikrettiğine göre bu âyet-i kerîme nazil olunca, Ömer (radıyallahü anh) şöyle demiş:Ey Allah'ın Rasûlü! Haydi kendimizi koruduk diyelim. Peki aile halkımıza ne yapabiliriz?' Peygamber şöylebuyurdu: "Allah'ın size yasakladığı şeylerden onları alıkoyarsınız, Allah'ın emrettiklerini onlara daemredersiniz.""Sen aile halkına namazı emret, kendin de sabırla ona devam et" (Taha, 20/132)"Önce yakın akrabanı uyar." (eş-Şuara, 26/214)"Çocuklarınıza yedi yaşında namaz kılmalarını emrediniz. (Kılmazlarsa) on yaşında onları (hafifçe)dövünüz ve yataklarını birbirinden ayırınız." Ebû Dâvûd, I, 133; Hâkim“Hepiniz çobansınız ve hepiniz sürüsünden sorumludur. İnsanların başındaki İmâm (İslâm devletininyöneticisi) bir çobandır ve o, onlardan sorumludur. Adam aile halkı üzerinde bir çobandır ve onlardansorumludur."

Qur'an Conversations
S4 E12: When Truth Demands Surrender (TaHa 105) | | Quran Conversations

Qur'an Conversations

Play Episode Listen Later May 22, 2026 75:12


Why do some hearts surrender to the truth while others resist it, even when they recognise it?In this episode of Quran Conversations, we are joined by Ustadh Fahd Yasin. Ustadh Fahd Yasin has been studying Quranic/Classical Arabic for the past decade. He has received ijazas (certifications) in Tajwid and grammar, and he is certified in Quranic Arabic linguistics. His current interests are in Quranic Analysis, Arabic Grammar, Rhetoric, Tazkiyah, and Tafsir. He has been a Quranic Arabic instructor at Fawakih Institute for the past 5 years. Ustadh Fahd Yasin is passionate about spreading Quranic linguistics to all of his students and everyone he meets! He wishes for everyone to experience the Light of the Quran and taste of the Quran's Secrets, Nuances, and Linguistic Subtleties.Dalia Mogahed and Ustadh Fahd Yasin reflect on Surah TaHa, ayah 105, exploring the destruction of the mountains on the Day of Judgment and the deeper meanings hidden within the Qur'an's precise language.What begins as a linguistic discussion unfolds into something far more personal: a reflection on certainty, ego, accountability, and the condition of the human heart. Why did the Qur'an choose mountains as its symbol? Why did the Quraysh feel so threatened by the Qur'an? And what separates the people who surrender to truth from those who fight against it?This episode explores how the Qur'an challenges the very things we rely on for stability and security, reminding us that even the mountains will one day disappear like dust.In this episode, you will learn:

Lundh
273 Re:Lundh – Taha Ali

Lundh

Play Episode Listen Later May 21, 2026 67:36


Häromveckan blev Taha Ali uttagen i den svenska VM-truppen. Malmö-dribblerns väg till elitfotbollen och landslaget är otrolig. När jag poddintervjuade Taha Ali i februari 2025 berättade han om den krokiga vägen fram till att kunna leva på fotbollen, om bakslagen som ung när han blev ratad, om att bröderna fick honom att fortsätta när han ville sluta, om motgången i allsvenska Örebro, om när han själv erbjöd sig till Landskrona men fick nej, om räddningen i Västerås SK under Kalle Karlsson och om skiftet till Helsingborg.Malmö FF-spelaren talade även om han inte förstod varför Henrik Rydström inte spelade honom 90 minuter, om besvikelsen över att han gjorde färre mål, om den virala succén efter han ordnade segern mot Blåvitt som gav guld, om känslan av att vinna titlar, om hur han blivit så vass på att dribbla, om intresset från utlandet, om målet att gå till en topp fem liga, om att känna brådska sett till åldern, om målet han haft att få en chans för Sverige, om stoltheten över det somaliska ursprunget och om tankarna på dess landslag. Hosted on Acast. See acast.com/privacy for more information.

Sweat Elite
Running the Palestine Marathon: Ahmad Taha (2:48 marathoner) on Racing, Training in Ramallah, and the Right to Movement

Sweat Elite

Play Episode Listen Later May 18, 2026 53:05


Running Through Checkpoints - The Reality Of The Palestine Marathon Recording from Ramallah after the Palestine Marathon, Matt sits down with Palestinian runner Ahmad Taher alongside Reem to unpack a race experience that goes far beyond performance. With multiple postponements and uncertainty around travel, the trip was only locked in seven days out, setting the tone for a week defined by unpredictability both on and off the course. Ahmad Taha Instagram: https://www.instagram.com/taha_runs/ Reem Ali Instagram: https://www.instagram.com/reemali378/ Be coached by Matt: https://www.sweatelitecoaching.com/coaching-2026 Join the Shareholders Club / Private Podcast Feed: https://www.sweatelite.co/shareholders Matt Instagram: https://www.instagram.com/mattinglisfox/ Matt Training Log - Strava: https://www.strava.com/athletes/6248359/ Contact Matt: matt@sweatelite.co Ahmad placed third on a brutally hilly course with around 550m of elevation gain, but the result came with mixed emotions after battling a pre-race cold and overthinking during the race. He shares how he nearly stepped off the course around 22km before regrouping to finish, reflecting on the mental side of racing in tough conditions. The conversation goes deeper into Ahmad's journey, starting with the first Palestine Marathon in 2013 where he ran with minimal preparation, through to racing internationally in Spain, Barcelona, and Paris. He explains how the marathon itself was founded through the Right to Movement initiative, with a course that passes the Church of the Nativity, the separation wall, refugee camps, and military checkpoints - highlighting the realities of restricted movement in daily life. Ahmad and Reem speak openly about what it's like training in Palestine, including encounters with soldiers, raids, tear gas, and the constant need to check news before heading out to run. Despite this, Ahmad remains committed to improving within Palestine, aiming toward the national marathon record of around 2:29 and targeting future races in Amman and San Sebastian. They also reflect on the emotional moment of fellow runner Mohammed Alasi returning to competition to place second after 32 months in administrative detention, adding further weight to what this race represents beyond sport. Topics: 00:00 - Welcome to Ramallah 01:03 - Last Minute Race Trip 03:15 - Ahmad Race Recap 03:39 - Hills and Head Games 08:16 - First Palestine Marathon 11:04 - Spain and Marathon Mania 12:53 - Training in Palestine 16:04 - Race Origins and Meaning 20:49 - Checkpoints and Daily Life 24:56 - Areas A B C Explained 26:55 - Running Under Occupation 28:37 - Guns on the Run 29:41 - Raids and Tear Gas 31:53 - Checking News for Safety 33:40 - Why Train in Palestine 37:19 - Friend Returns From Prison 41:56 - Life in a Refugee Camp 43:18 - Foreign Passports and Checkpoints 47:24 - Marathon Course Sabotage 48:56 - Coaching and Next Goals 51:13 - Records and Sweet Farewell

Qur'an Conversations
S4 E11: The Moment You Realise You Were Wrong (TaHa 102–104) | Quran Conversations

Qur'an Conversations

Play Episode Listen Later May 15, 2026 57:16


What happens when everything you ignored becomes impossible to deny?In this episode of Quran Conversations, Dalia Mogahed is joined by Talha Ghannam. Talha is a Mathematics and Economics graduate, Islamic scholar, entrepreneur and community activist. He studied under leading scholars in the UK, Syria and Egypt, completed a seven-year Alimiyah course, and now focuses on purification of the heart and the Quran. Known for his Quranic reflection and tafsir content, his videos have reached millions, helping people connect deeply with the Quran. He is the founder of Quran Club, which has surpassed 500,000 downloads, co-founder of ClassTutor, supporting over 2,000 students with 150+ teachers, and co-founder of the Centre for Islam and Medicine, exploring contemporary bioethics through Islamic tradition.In this episode, Dalia and Talha reflect on verses 102–104 of Surah TaHa. A vivid, unsettling glimpse into the Day of Judgment.These verses don't just describe an event. They immerse you in it. Through sound, imagery, and subtle language, the Qur'an pulls you into a moment where control disappears, illusions collapse, and reality is fully exposed.This is not a distant scene. It is a mirror of what we are becoming.In this episode, you will learn:

Qur'an Conversations
S4 E10: The Cost of Ignoring the Truth (TaHa 99–101) | Quran Conversations

Qur'an Conversations

Play Episode Listen Later May 8, 2026 58:04


What is the Qur'an really asking from you, and what happens when you turn away from it?In this episode of Quran Conversations, Dalia Mogahed and Mohammed Magid reflect on verses 99–101 of Surah TaHa. A moment where the story pauses, and the Qur'an speaks directly to you.After the powerful narrative of Musa (peace be upon him), these verses shift from storytelling to meaning. They reveal why these stories are told, what they demand from us, and the consequences of ignoring their message.This is where the Qur'an stops being history, and becomes a mirror.In this episode, you will learn:

Qur'an Conversations
S4 E9: The Idol You Didn't Realise You Were Worshipping (TaHa 97–98) | Quran Conversations

Qur'an Conversations

Play Episode Listen Later May 4, 2026 57:15


What does it take to truly break free from false idols—whether they are people, ideas, or desires?In this episode of Quran Conversations, Dalia Mogahed and Mohammed Magid reflect on verses 97–98 of Surah TaHa. A powerful turning point where Prophet Musa (peace be upon him) confronts As-Samiri and dismantles the illusion of the golden calf.This moment goes beyond punishment. It exposes the reality of false devotion, the psychology of influence, and the path to spiritual liberation. From social isolation as consequence, to the public destruction of the idol, the Qur'an offers a profound blueprint for breaking free from manipulation, materialism, and misplaced worship.In this episode, you will learn:

Qur'an Conversations
S4 E8: Guarding Unity in Times of Crisis (TaHa 94-96) | Quran Conversations

Qur'an Conversations

Play Episode Listen Later Apr 17, 2026 63:38


What happens when two righteous leaders respond differently to the same crisis?In this episode of Qur'an Conversations, Dalia Mogahed and Sheikh Mohammed Majed reflect on Ayah 94 of Surah TaHa—a powerful moment where Prophet Musa (peace be upon him) confronts his brother Harun after the people fall into worship of the golden calf.This interaction reveals something deeply human: frustration, grief, restraint, and the weight of leadership. It also offers timeless lessons about obedience, accountability, and navigating conflict without tearing communities apart.In this episode, you will learn:

Qur'an Conversations
S4 E7: Disagreeing with Compassion (TaHa 89-93) | Quran Conversations

Qur'an Conversations

Play Episode Listen Later Apr 10, 2026 62:56


What happens when devotion is misplaced—and correction is met with resistance?In this episode of Qur'an Conversations, we reflect on Surah TaHa (20:89–93) and the painful moment when Bani Israel refuse to abandon the golden calf—even after being reminded of Allah's mercy. Through the dialogue between Prophet Musa and Prophet Harun (peace be upon them both), the Qur'an gives us a powerful lens into misplaced devotion, spiritual blindness, leadership under pressure, and how to repair relationships even after rupture You will learn:

The East is a Podcast
"War, War until Victory!": Iran resists ZioAmerican aggression w/ Sara Larijani and Taha Zeinali

The East is a Podcast

Play Episode Listen Later Mar 23, 2026 78:13


Sara Larijani is a PostDoc fellow in Political Geography at the University of Tehran. Her research is on British colonialism in Iran's oil frontiers. Taha Zeinali (@tahazeinalih) is a doctoral researcher in development studies at the Autonomous University of Barcelona. His research focus is imperialist hybrid war, sanctions and sovereign development in Iran. They are both co-founders of the Center for Resistance, Sovereignty and Development Studies at the University of Tehran. Watch the video edition on The East Is a Podcast YouTube channel https://youtu.be/PKc1RPqnca8 Consider supporting the show www.patreon.com/east_podcast

YAP - Young and Profiting
Hala Taha: The Mindset That Turned Rejection into a Multi-Million Dollar Business | Human Behavior | YAPClassic

YAP - Young and Profiting

Play Episode Listen Later Feb 25, 2026 54:09


Hala Taha's mindset was pushed to its breaking point by relentless rejection, discrimination, and loss. After three years of unpaid sacrifice at Hot 97, being fired and blackballed, repeatedly passed over for promotion, and ultimately losing her father to COVID, she had every reason to quit. But instead of waiting for permission, she rebuilt her psychology from the ground up, stacked her unique strengths, and carved out her own path. In this MIT keynote speech, Hala shares her raw, unfiltered come-up story and the exact mindset shifts that fueled her self-improvement and helped her build a profitable life and business against all odds. In this episode, Hala will discuss: (00:00) Introduction (05:18) Her Father's Grit and Palestinian Roots (08:41) Growing Up Between Two Worlds (15:51) Hot 97: Working for Free and Getting Blackballed (21:16) The Sorority of Hip Hop and MTV Rejection (28:35) Losing Her Father to COVID-19 in 2020 (37:54) Her Secrets to Profiting in Life (43:47) Audience Q&A Hala Taha is the host of Young and Profiting, a top 10 business and entrepreneurship podcast on Apple and Spotify. She's the founder and CEO of YAP Media, an award-winning social media and podcast production agency, as well as the YAP Media Network, where she helps renowned podcasters like Russell Brunson, Jenna Kutcher, and Neil Patel grow and monetize their shows. Through her work, Hala has become one of the most influential creator-entrepreneurs in podcasting. Sponsored By: Indeed - Get a $75 sponsored job credit to boost your job's visibility at Indeed.com/profiting Shopify - Start your $1/month trial at Shopify.com/profiting. Spectrum Business - Keep your business connected seamlessly with fast, reliable Internet, Phone, TV, and Mobile services. Visit https://spectrum.com/Business to learn more. Northwest Registered Agent - Build your brand and get your complete business identity in just 10 clicks and 10 minutes at northwestregisteredagent.com/paidyap Framer - Publish beautiful and production-ready websites. Go to Framer.com/profiting and get 30% off their Framer Pro annual plan. Quo - Run your business communications the smart way. Try Quo for free, plus get 20% off your first 6 months when you go to quo.com/profiting Working Genius - Take the Working Genius assessment and discover your natural gifts and thrive at work. Go to workinggenius.com and get 20% off with code PROFITING Experian - Manage and cancel your unwanted subscriptions and reduce your bills. Get started now with the Experian App and let your Big Financial Friend do the work for you. See experian.com for details. Huel -  Get all the daily nutrients you need with Huel. Grab Huel today and get 15% OFF with my code PROFITING at huel.com/PROFITING.  Resources Mentioned: Hala's Podcast, Young and Profiting: bit.ly/_YAP-apple  Hala's Agency, YAP Media: yapmedia.com    Active Deals - youngandprofiting.com/deals  Key YAP Links Reviews - ratethispodcast.com/yap YouTube - youtube.com/c/YoungandProfiting Newsletter - youngandprofiting.co/newsletter  LinkedIn - linkedin.com/in/htaha/ Instagram - instagram.com/yapwithhala/ Social + Podcast Services: yapmedia.com Transcripts - youngandprofiting.com/episodes-new  Entrepreneurship, Entrepreneurship Podcast, Business, Business Podcast, Self Improvement, Self-Improvement, Personal Development, Starting a Business, Strategy, Investing, Sales, Selling, Psychology, Productivity, Entrepreneurs, AI, Artificial Intelligence, Technology, Marketing, Negotiation, Money, Finance, Side Hustle, Startup, Mental Health, Career, Leadership, Mindset, Health, Growth Mindset, Habits, Positivity, Human Nature, Human Psychology, Critical Thinking, Robert Greene, Chris Voss, Robert Cialdini 

Business Made Simple with Donald Miller
#57: Hala Taha—Catapult Your Personal Brand With This Simple Messaging Formula

Business Made Simple with Donald Miller

Play Episode Listen Later Feb 2, 2026 40:53


You might think posting daily, chasing trends, and being "everywhere" is the key to growing your personal brand. But what if the real problem isn't visibility, it's clarity? Most people trying to build a brand skip the hardest part: figuring out exactly what they stand for and how to communicate it consistently. Without that foundation, every post feels like a shot in the dark. And if your message is muddy, even the best content won't convert. So how do you build a brand that actually resonates, builds trust, and drives results?   In this episode, Donald Miller sits down with "Podcast Princess" Hala Taha, host of Young and Profiting (YAP) Podcast, to unpack the messaging strategies that built her powerhouse personal brand. They dive into her step-by-step framework for personal brand clarity, how to create sticky content themes, and why it pays to stand against something just as much as you stand for something. You'll also hear how Hala used these strategies to grow a thriving business, land sponsorships, and inspire the next generation of creators. Whether you want to go big or just be "five miles famous," this episode shows you how to get known and stay known.     Check out more from Hala on her website: https://youngandprofiting.com   --   Click HERE to get in-person help creating your marketing at the next available StoryBrand Your Business LIVE event!   Click HERE to find a StoryBrand certified marketing coach to help you grow your business!   Learn how to make your marketing and messaging work using a proven framework in the updated book, Building a StoryBrand 2.0. Order it now on Amazon  or wherever you buy books!