City-state in ancient Greece
POPULARITY
Categories
Andrew Bayliss, author of Sparta: The Rise and Fall of an Ancient Superpower, Part One: Bayliss details Sparta's emergence as a classical Greek superpower, emphasizing its unique geography in the Eurotas valley. Bayliss explains the origin of Sparta's dual-kingship compromise and its complete dependence on "helots" (enslaved cousins), who were subjugated through systemic state terror and regularized beatings. The discussion explores the rigorous Agoge system, which began at age seven and subjected young boys to physical training, starvation, and tactical stealing. Additionally, Spartan women were trained to be strong, athletic mothers. The segment concludes with the Persian Wars, highlighting the legendary stand of the 300 at Thermopylae and the naval battle of Salamis. (1)
Andrew Bayliss, Part Two: Bayliss covers the Peloponnesian War, sparked by Athenian leader Pericles' provocations. The Spartans respond with siege tactics, while Athens suffers from a devastating plague that claims Pericles' life. Prominent figures like Lysander emerge, securing Persian funds to build a competitive fleet and capture the Athenian navy. However, Sparta struggles to govern, plagued by declining citizen numbers. The empire collapses after its defeat by the Thebans at Leuctra. The segment concludes by examining the historical validity of the "Thucydides Trap" in 21st-century geopolitics. (2)
Despite the space sector's consistent growth and greater importance as a critical infrastructure sector, finding cybersecurity talent is still proving to be a challenge. Host Maria Varmazis and Nick Cohen, Principal Engineer at the Aerospace Corporation, sat down to discuss the current state of space's cybersecurity workforce. The two look at how the existing workforce is not large enough to effectively cover the industry and its infrastructure and how solutions like the Aerospace Village at DEF CON 34 have been aiming to address expertise shortcomings. Key sources: The SPARTA Framework SpaceCOP SpaceTrail StarPWN The Aerospace Village SPARTA 4.0 Debut Experience the convergence of space and cyber threats at DEF CON 34 SPARTA v4.0 — Breaking New Ground Like what you heard? Be sure to subscribe to our free Signals and Space Briefing, our Sunday newsletter covering the intersection of cybersecurity and space. Subscribe at: https://thecyberwire.com/newsletters/signals-and-space Is there a topic or person you'd like to hear on our show? You can send your questions and feedback to space@n2k.com. You can also fill our our audience survey: https://www.surveymonkey.com/r/NJYCN2P T-Minus: Space-Cyber Briefing is a production of N2K CyberWire. N2K is your nexus for discovery and connection for people, technology, and ideas shaping the future of secure innovation. Learn how at n2k.com.
It's all questions on totally random subject! This episode's topic: CONFIDENCE ROUND Sponsored by LearnClash, the quiz duel app that helps you remember what you learn: learnclash.com/partners/budds Fact of the Day: The infamous "This is Sparta!" line from 300 was originally said in a low, monotone growl. After doing several takes of that, Gerard Butler asked to do one more take and used the now-iconic delivery. Triple Connections: Liberty, Storm, Sun THE FIRST TRIVIA QUESTION STARTS AT 01:32 SUPPORT THE SHOW MONTHLY, LISTEN AD-FREE FOR JUST $3 A MONTH: www.Patreon.com/TriviaWithBudds INSTANT DOWNLOAD DIGITAL TRIVIA GAMES ON ETSY, GRAB ONE NOW! GET A CUSTOM EPISODE FOR YOUR LOVED ONES: Email ryanbudds@gmail.com Theme song by www.soundcloud.com/Frawsty Bed Music: "Laser Groove" Kevin MacLeod (incompetech.com) Licensed under Creative Commons: By Attribution 4.0 License http://creativecommons.org/licenses/by/4.0/ http://TriviaWithBudds.comhttp://Facebook.com/TriviaWithBudds http://Instagram.com/ryanbudds Book a party, corporate event, or fundraiser anytime by emailing ryanbudds@gmail.com or use the contact form here: https://www.triviawithbudds.com/contact SPECIAL THANKS TO ALL MY AMAZING PATREON SUBSCRIBERS, INCLUDING: Samantha Wheeler Boomer Cates Mark Kloppenburg Cadi Snow-Brine Amber Shiels Alan Kreisel Rich Sommer Joe Heiman Waqas Ali Logan Booker Bringeka Sam Nathan Stenstrom Brooks Martin Robyn Price Gee Brian Clough Charles Glanville IV Lauren Schuette Evan Lemons AnneMarie Mattacchione Yves Bouyssounouse Kenny Zail York yates Gay Geek Fabulous Mollie Dominic Nathalie Avelar Natasha raina leslie gerhardt Diane White Youngblood Trophy Husband Trivia Lynnette Keel Lillian Campbell Jerry Loven Jamie Greig Gail Lancman Jeremy Yoder Adam Jacoby rondell Adam Suzan Tiffany Poplin Bill Bavar Sarah Daniel Hoisington Keith Martin Sue First Steve Hoeker Jessica Allen Lauren Glassman Brian Williams Brett Livaudais Linda Elswick Carter A. Fourqurean Justly Maya Brandon Lavin Kathy McHale Chuck Nealen Courtney French Nikki Long Mark Zarate Laura Palmer JT Dean Bratton Kristy Erin Burgess Trenton Sullivan Jen and Nic Michael Redman Timothy Heavner Jeff Foust Richard Lefdal Myles Bagby Jenna Leatherman Vernon Heagy Albert Thomas Kimberly Brown Tracy Oldaker Sara Zimmerman Madeleine Garvey Jenni Yetter Patrick Leahy Dillon Enderby James Brown Christy Shipley Clayton Polizzi Alexander Calder Ricky Carney Paul McLaughlin Willy Powell Robert Casey Matthew Frost Brian Salyer Greg Bristow Megan Donnelly Jim Fields Mo Martinez Luke Mckay Simon Time Feana Nevel Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Trapshooting legends Robert “Bob” Munson and Lou Ann “Annie” Munson join Richard and Zach for a special episode of Trap Talk From the Back Fence, recorded live during the 127th Grand American at the Trap Talk Studio in Building 106.Bob and Annie have spent more than five decades traveling the country, winning major championships, representing Minnesota, and setting a standard of excellence few shooters have ever matched.Bob's remarkable career includes hundreds of 100 straights, numerous Minnesota championships, the 1982 Grand American Clay Target Championship, and the 1983 Grand American High All-Around Championship with a 397x400.Annie became one of the early women to reach the 27-yard line, earned repeated All-American honors, and captured major victories including the 1984 Grand American Ladies High-Over-All and the 1979 Women's Doubles Championship.Together, Bob and Annie have won more husband-and-wife championships than anyone can accurately count. Their accomplishments are extraordinary, but their stories, memories, and perspective on the sport are what make this conversation truly special.This episode is a chance to sit down with two people who did not simply compete in trapshooting. They helped define an era of it.
Send us Fan MailIn this episode,, our stack of books is tied together with the common theme of being a bundle of Memoirs. Memoirs are one of our favorite types of non fiction and we are bringing you some good and unexpected ones today! Join us to dig a little deeper into this genre.Featured Books:Hello, Molly! A Memoir by Molly Shannon (LH)Actress of a Certain Age: My Twenty-year Trail to Overnight Success by Jeff Hiller (LH)Oh Brother: A Graphic Memoir by Georgina Chadderton (LP)Confessions of a Prairie Bitch: How I Survived Nellie Oleson and Learned to Love Being Hated by Alison Arngrim (LP)Book in Hand:The Lost Daughter of Sparta by Felicia Day (LH)Books Mentioned in This Episode:You're Never Weird on the Internet by Felicia DayThe Guild Series by Felicia DayAdditional Books That Go Along with Our Stack:The Year of Magical Thinking by Joan DidionKook by Peter HellerThe Glass Castle by Jeannette WallsThe Distance Between Us by Reyna GrandeThe Guild series by Felicia DayYou're Never Weird on the Internet (Almost) by Felicia DayWays to contact us:Join us on Patreon for extra content: https://www.patreon.com/c/BookBumblePodcastFollow us on Instagram - @thebookbumbleFacebook: Book BumbleOur website: https://thebookbumble.buzzsprout.comEmail: bookbumblepodcast@gmail.comSupport the showIf you like our podcast please help us: Rate and review us, subscribe, follow us on Insta, and join our Team Patreon! It won't be the same without you!
Mold in a Porsche. High spots on a fresh ceramic. And the one mistake that keeps detailers from getting real protection on the car.Marshall and Nick pull back the curtain on the kind of problems shop owners and serious enthusiasts run into every day - from dealing with mold remediation in modern vehicles to the exact reason coatings fail when people try to stretch product too far. If you've ever wondered whether you're using enough on the pad, enough on the panel, or enough on the process, this episode is going to hit home fast.Marshall also breaks down the real-world value of giving customers a path, not just a product - from maintenance kits to YouTube guidance to a complete ecosystem that helps shops protect cars and keep customers coming back. Nick adds the practical side of why some jobs should be left to the dealership, why overthinking tools can slow you down, and why “light usage” often just means not enough product on the surface.Then the conversation shifts to the bigger picture: Juice, Sliq , Spray Coat, Stack, and Eco One. What's the difference, when should you use each one, and how do dilution ratios and application style change the result? If you've ever second-guessed whether you're spraying enough, wiping enough, or mixing enough, this one clears it up.Essential listening for detailers, shop owners, and car lovers who want better results, fewer mistakes, and a simpler system that actually works.Chapter timestamps00:00 - Why sitting in a moldy car makes you rethink your whole day01:13 - Enzyme cleaning mold and why masks matter02:11 - Pulling mats, frunk talk, and when to replace instead of save04:25 - Why removing seats on modern cars is an insurance problem05:16 - Old-school flood-and-extract methods versus modern remediation07:56 - Dodge truck jokes and why truck owners get so defensive10:36 - Uno questions and the rule of loading the applicator thick12:19 - Why underapplying coating causes flashing and missed coverage15:54 - Why protection only works if you actually use enough of it17:14 - How Stack, Uno, Trey, and Sparta get used panel by panel19:19 - What a high spot is and why it is mostly just ugly21:39 - Nathan's customer gets a full maintenance path, not just a coating23:34 - Turning YouTube and the Facebook group into a product encyclopedia26:15 - Why shops can cut their cost in half by building kits into every job28:08 - Coating a new Traverse RS and the right order for trim, glass, and paint30:50 - Which detailing lights are worth buying and which ones are overhyped33:40 - Whether glass needs polishing before coating35:24 - How to fix a high spot with VLO, Lux, or a towel method38:56 - Looking back and forward while you coat so you do not miss anything43:46 - Juice versus Slick versus Spray Coat versus Stack48:46 - Why product cost should not make you underapply50:23 - Why Nick does not like using rinseless or waterless on interiors55:10 - Eco One dilution ratios and how the 5-gallon cube works59:05 - Closing thoughts and the value of the HyperClean Specialists group
This week in nerd culture was absolutely packed. From huge casting news for Amazon Prime Video's God of War series to major GTA 6 gameplay leaks, surprise franchise revivals, MCU updates, DC success stories, and more, we're breaking down the biggest nerd news headlines of the week!First up, Dave Bautista is officially stepping in as Kratos, replacing Ryan Hurst in Amazon's upcoming live-action God of War series. Is Bautista the right choice to bring the Ghost of Sparta to life?We also dive into the latest Grand Theft Auto 6 leaks, including major new gameplay details that may reveal more about Rockstar Games' highly anticipated open-world sequel.Plus, Lanterns is proving to be a major success for HBO and DC Studios, giving James Gunn's new DC Universe a much needed win. And after years of rumors and speculation, National Treasure 3 has finally been confirmed by the director. Could Nicolas Cage really be returning as Benjamin Franklin Gates?Elsewhere in the nerd world:• Former Doctor Who writer Mark Gatiss believes the series could potentially disappear from television for an extended period — possibly even another 15–20 year hiatus.• Marvel Studios president Kevin Feige says the MCU could potentially make hundreds of X-Men movies, opening the door for decades of mutant stories.• Final Fantasy VII: Revelations may be getting a playable demo sooner than expected.• James Spader is returning as Ultron for Marvel Studios' upcoming VisionQuest series.• And despite struggling at the box office, Masters of the Universe has reportedly become a massive success on Amazon Prime Video.We break down all of these stories, share our reactions, and debate what they could mean for the future of gaming, Marvel, DC, television, and nerd culture as a whole.Which story was the biggest news of the week? Let us know in the comments!
The unforeseen consequences of this war led to the first scientific history book, the birth of philosophy. and the rise of Alexander the Great. Assassin's Creed Odyssey puts us at the heart of the expansive conflict that occurred when Athens went to war with Sparta: the Peloponnesian War. Guest host Tristan Hughes is joined by Dr Roel Konijnendijk to recount the causes, events and legacy of this pivotal war in the ancient Greek world.Echoes of History is a Ubisoft podcast, brought to you by History Hit. History Hit is part of Little Dot Studios, a Certified B Corp™.Hosted by: Tristan HughesEdited by: Michael McDaidProduced by: Robin McConnellSenior Producer: Anne-Marie LuffProduction Manager: Beth DonaldsonExecutive Producers: Etienne Bouvier, Julien Fabre, Steve Lanham, Jen BennettMusic:Delphi by The FlightPirates, Thugs and Bandits by The Flight, Michael GeorgiadesIf you liked this podcast please subscribe, share, rate & review. Take part in our listener survey here. You can watch this episode on YouTube.Tell us your favourite Assassin's Creed game or podcast episode at echoes-of-history@historyhit.com Hosted on Acast. See acast.com/privacy for more information.
Jeden se musel kát před tvrdým jádrem fanoušků, druhý na tribuně zuřivě slavil góly proti klubu, za který byl ještě nedávno připravený padnout. Matěj Jurásek a Jan Bořil. I jejich divoká vystoupení řeší fotbalový podcast MVP. Kde jsou rezervy mistrovské Slavie? Jak silný může být letos Liberec? Zůstane reprezentační kapitán Ladislav Krejčí ve druhé anglické lize? Pusťte si nový díl MVP! --- Fotbal ze všech možných i nemožných úhlů pohledu. MVP jsou bývalí fotbaloví profesionálové Karel Tvaroh, Antonín Rosa, Tomáš Kučera a zkušený novinář Jan Palička, šéf sportovní rubriky Seznam Zpráv. Společně s námi hledejte nejdůležitější hráče, trenéry, přestupy, akce, problémy. Do hloubky a s humorem. I vy můžete být MVP. Každé úterý na webu Seznam Zpráv. Odebírejte na Podcasty.cz, Apple Podcasts nebo Spotify. Sledujte nás na Stream.cz nebo YouTube. Bližší pohled do kabin MVP se vám nabízí na našem Instagramu. Máte návrh, jak podcast vylepšit? Nebo nás chcete pochválit? Pište na audio@sz.cz
Le nouveau podcast football du FC Copains
Keith welcomes Jim Ward back to the show and we discuss the new Sparta LP "Cut A Silhouette", the birth of the record beginning with Jim rediscovering music he grew up listening to, his own music and reconnecting with several of his former At The Drive In / Sparta bandmates, the creative process for "Cut A Silhouette", recording the record with J Robbins and the artists Jim collaborated with on the record including Frank Iero of My Chemical Romance. We also discuss Jim's history with At The Drive In, his thoughts and reflections on the band, the recent Sparta "Wiretap Scars" and "Porcelain" anniversary tours, the time Bono of U2 helped Jim in a time of need and more.
Grunden i det antika Greklands militära organisation var hopliten. Hopliten var alltid en man och medborgare i någon stadsstat. Beväpnad med ett längre spjut, en sköld och ett kortare svärd, formerades hopliterna i så kallade falanger.Bland de mest kända grekiska stadsstaterna är Aten och Sparta. Tack vare ett rikt och välutvecklat skriftspråk med många överlevda texter, i kombination med en rad nutida arkeologiska fynd så vet vi ganska mycket om hur krigen under dessa århundraden tedde sig.I detta avsnitt 46 av Militärhistoriepodden gräver Peter Bennesved och Martin Hårdstedt djupare i dessa antika fenomen. De går igenom såväl stadsstaternas struktur, ekonomiska förutsättningar, såväl som hopliternas utrustning och den krigsmentalitet som rådde. På vägen hanterar vi också kända händelser som slaget vid Marathon, Salamis, Mantinea, och den fatala atenska expeditionen till Sicilien (ett antikt ”Barbarossa”?) under Peloponnesiska krigets höjdpunkt.Vi har en klar bild över hur både samhället och krigföringen organiserades, är stadsstaterna i antikens Grekland. Det fanns ett stort antal av dem, utspridda över kust och landområdena runt Egeiska havet.Perioden mellan ca 500-300 år före vår tideräkning är också en mycket intressant period vad gäller krigskonstens utveckling. Det som har beskrivits som en av de viktigaste fördelarna med den västerländska krigföringen, disciplinen, har ett tydligt ursprung i de grekiska stadsstaternas sätt att kriga.Grunden i det antika Greklands militära organisation var hopliten. Hopliten var alltid en man och medborgare i någon stadsstat. Beväpnad med ett längre spjut, en sköld och ett kortare svärd, formerades hopliterna i så kallade falanger.Falangen var en kompakt fyrkantsformation i flera led med fördelen att de tillsammans med sina kamrater bildade en sköldmur med en svärm av utstickande spjutspetsar. Exakt hur falangstrider gick till vet vi inte, men klart är att grekerna uppvisar en enastående disciplin och taktiskt tänkande som fick stora konsekvenser för dem själva och omvärlden. Deras förmåga att hålla ihop sina trupper och slås till siste man ledde till en rad häpnadsväckande segrar mot bland annat det enorma Persiska imperiet, men det ledde också till svårigheter när greker möte greker i strid. Hur skulle man ta krigskonsten vidare när falang mötte falang, och hur skulle framgångarna till land omsättas till framgångar till havs?Bild: Chigi-vas med hopiliter som håller sköldar och svärd. Wikipedia, Public Domain. Hosted on Acast. See acast.com/privacy for more information.
In 480 BC, Athens and Sparta joined forces to repel one of the greatest invasion forces the world had seen. Yet just a few decades later, they faced off against each other – with dire results for the people of Greece. This Long Read written by Adrian Goldsworthy reveals how these great cities, once allies, became deadly foes. Today's feature originally appeared in the June 2026 issue of HistoryExtra magazine, and has been voiced in partnership with the RNIB. Learn more about your ad choices. Visit podcastchoices.com/adchoices
Episodio donde se reseña la película The Last House, Pari ya acabó House of the Dragon y no está conforme, opiniones sobre el metroidvania de God of War: Sons of Sparta y el juego de Indiana Jones con o sin la voz de Harrison Ford, el gameplay del nuevo juego de Halloween mete dudas a los co-capitanes, la carrera ficticia entre el campeón del mundo de comer hotdogs vs el hombre más veloz del mundo, reseña CON spoliers de Spider-Man: Brand New Day, y terminamos con los castings nuevos de X-Men y sus posibles historias. Escúchanos: Spotify / Apple Podcasts / YouTube Apóyanos: patreon.com/holamsupernova Síguenos: Instagram/ Twitter/ TikTok @holamsupernova Merch: holamsupernova.myshopify.com
Sparta Praha je po třech kolech v pořádném průšvihu.
Jim Ward returns to talk with Dave about Sparta's new album Cut a Silhouette, out now on Equal Vision Records. Ward shares how a manager's blunt critique sparked three of the record's best songs, working with Jawbox founder and producer J. Robbins as a shepherd and mentor, and bridging his earliest work with who he is now. The two discuss the ticket price fatigue reshaping the touring industry, co-founding the Girls Rock camp in El Paso, Hayley Williams, and the elusive "Jims" supergroup featuring Jim Adkins of Jimmy Eat World.https://www.sparta.band/directeditionpodcast.comJoin The Patreon - www.patreon.com/davengersdirectedition
In this message, Pastor Caleb discusses keeping the correct heart in one's relationship with Jesus Christ and in what one does for the kingdom of God. Christians should not be on autopilot when living their lives or doing things for God. May Christians maintain a correct heart in serving and living for God. Send us Fan MailSupport the showFor more information for our church visit AGCSparta.org.
In this message, Pastor Caleb discusses not allowing voices or opportunities to hinder or distract one from fulfilling the call of God on one's life. A good soldier of Jesus Christ does not allow distractions to pull him away from the mission that has been given. Send us Fan MailSupport the showFor more information for our church visit AGCSparta.org.
MOPs & MOEs is proudly sponsored by Teamworks — the performance operations platform trusted by elite military units and professional sports organizations worldwide. Teamworks brings your scheduling, communications, athlete monitoring, and readiness data into one unified system — so your leaders stay informed, your people stay connected, and your unit stays ready. No more scattered spreadsheets or missed messages. Just one platform built for organizations where performance is the mission. Learn more at teamworkstactical.comWe are also supported by TrainHeroic — the coaching and programming platform built for strength and conditioning coaches who train serious athletes. Whether you're programming for a military unit, a tactical team, or individual athletes, TrainHeroic gives you the tools to build and deliver professional training programs, track athlete progress, and communicate directly with your people — all through one app. Your athletes get world-class programming on their phone; you get the visibility to actually coach them. Start your free trial at trainheroic.comThe History of Exercise Is Weirder Than You Think — Bill Hayes, Author of SweatBill Hayes spent four years chasing the history of exercise across ancient Greece, Renaissance Italy, Kansas City, and the Vatican — and somehow ended up writing the most entertaining book about working out that's ever existed. Drew and Alex sit down with him to find out what he learned.What we get into:Gloios — ancient Greek athletes would scrape the sweat and oil off their bodies after competing, funnel it into a jar, and sell it at the gymnasium for medicinal purposes. Warts. Hemorrhoids. The price depended on how good the athlete was. Christiane Amanpour interviewed Bill on CNN and said nothing had ever seemed so gross to her. She was correct.Exercise began as war training — tenth century BC, not that different from what we do now. Running, endurance testing, strength work. Sparta was the only place that let women train. The Olympic Games formalized competition in the eighth century BC and athletes became cultural heroes — which is why their sweat was worth buying.Bill's dad was a West Point swimmer, Korean War vet, Purple Heart recipient who lost an eye in combat. Bill chose poetry over a military career — an act of rebellion he's pretty honest about. But his dad dragged him to the Spokane Athletic Club every Sunday after church, and that's where the love of exercise started.Girolamo Mercuriale — a Renaissance physician who wrote the first systematic text on exercise, De Arte Gymnastica, in 1569. Bill tracked down a transcription of his manuscripts in a library in Kansas City, Kansas, that hadn't been checked out in 35 years.The Vatican Library story — Bill had a Guggenheim Fellowship and letters of recommendation and still got interrogated in a small room because he didn't have a PhD. He got in. Everything he needed was already in Kansas City.Staying in it across a lifetime — injuries, aging, motivation, and the underrated importance of rest. Bill has herniated discs, a severed rotator cuff, and is still going. The answer isn't finding the perfect program. It's just not stopping.Mentioned in this episode:Sweat — Bill HayesGirolamo Mercuriale, De Arte Gymnastica, 1569Isthmia, Olympia, Delphi, Mycenae — Bill's Greek road trip following the original athletic circuitThe Metropolitan Museum of Art — has a strigil and gloios pot on display if you want to see it in personLong and Strong — the Mops and Moes training program on TrainHeroic Views expressed are those of the speakers and do not represent any official organization.
In this message, Pastor Caleb continues to discuss the types of sin. The next type of sin discussed is the sin of commission. Looking at what the Bible says about sins of commission can help Christians understand how sin operates to better overcome it. Send us Fan MailSupport the showFor more information for our church visit AGCSparta.org.
This captures the three biggest substantive threads — the Kohler/Izzo roster debate, on-court takeaways from Moneyball Pro-Am, and the Senate NIL legislationBecome a supporter of this podcast: https://www.spreaker.com/podcast/this-is-sparta-msu--5664600/support.
What an opening couple of games! We discuss the dismantling of Sparta, an excellent away day in Plzeň, have a fruity beer, delve into a little local West Bohemian history, and celebrate the return of the occasionally-missed Vrabec Review!0.00 - opening2.10 - Zbrojovka vs. Sparta Prague16.35 - Hot? Or Not?!23.25 - Viktoria Plzeň vs. Zbrojovka38.40 - Vrabec Review40.30 - What's the Deal With... Plzeň?47.25 - Beer of the Podcast51.40 - Slovan Liberec and Hradec Králové previews1.05.30 - outro
Paulo Fonseca, est-il le principal responsable ?Après la défaite frustrante face au Sparta Prague (1-2), les questions se multiplient autour de l'Olympique Lyonnais et de son entraîneur. Paulo Fonseca a-t-il sa part de responsabilité dans ce revers ? Ses choix tactiques étaient-ils les bons ? Les joueurs sont-ils davantage en cause ?Dans cette émission, l'équipe du Winamax FC revient en détail sur la prestation lyonnaise, analyse les décisions du coach portugais et tente de comprendre les raisons de cette contre-performance qui fait déjà beaucoup parler.Aussi, le WFC évoque l'actualité de ces derniers jours : Esteban Lepaul, qu'attendre du meilleur buteur de l'année dernière ? Peut-il confirmer son nouveau statut, voire rêver de sélection ? Enfin, retour sur l'arrivée de Mohamed Salah en Turquie, à Trabzonspor !Ce podcast est hébergé par Podcastics, la plateforme pour créer et diffuser votre podcast facilement.
Sparta porazila Zlín, ale klid na Letné rozhodně není!
Watch the full episode on YouTube:We first covered Baseten last year when DeepSeek mania was at peak hype. Now they have raised a monster $13B round and become one of the new cohort of AI Infra decacorns that are (with Nvidia, Intel, and the semis complex) chief beneficiaries of the Inference Inflection. We return to Baseten at the peak of the 2026 edition of Open Weights debate. Ali has published a viral breakdown of Kimi K3:And since you last saw him, Philip has spoken at AI Engineer and written the definitive book on Inference Engineering spotted all over SF:Three years ago, inference engineering barely existed as a category.Today, it is one of the most critical disciplines in AI. Inference engineering inherently tackles a different question than standard model training: “How do you turn those weights from training into a product that is fast, reliable, and affordable at scale?” Focusing on these creates an entirely new optimization problem.In one recent GLM-5.2 experiment, quantizing more of the model actually preserved its benchmark quality while increasing throughput by 20%, because the errors introduced in different layers could cancel each other out.Inference is no longer just the final step after training. It is becoming its own engineering discipline, with its own research problems, infrastructure, and increasingly specialized roles.In this episode, Baseten's Philip Kiely and Ali Taha join swyx and Vibhu to explain what actually happens after a new open model is released and what it takes to turn “we generated a token” into a fast, reliable, production-ready API.We go deep on cache-aware routing, disaggregated prefill and decode, quantization, speculative decoding, KV-cache movement, model parallelism, GPU kernels, and the race to make frontier models up to 10× faster. Philip and Ali explain why inference optimizations can still produce gains of 20%, 100%, or even 200%; how quantization errors can cancel one another out; why identical weights can behave differently across clusters; and how Baseten grafted a Kimi vision encoder onto GLM-5.2 without changing the underlying language model.The conversation then expands beyond LLMs into NVIDIA Dynamo, mega kernels, Rubin, AI-specific chips, local inference, video generation, diffusion versus autoregressive models, and the enormous compute barrier to generating coherent long-form video. Finally, we explore the convergence of training and inference, continual learning through persistent KV cache, and the emerging loop where models help optimize the infrastructure that runs them.We discuss:* What happens when a 200,000-token request enters an inference system* Cache-aware routing and reusing previously computed KV cache* Why prefill and decode are increasingly handled by different GPUs* When dedicated deployments become cheaper and more reliable than shared APIs* How speculative decoding uses a smaller model to accelerate a larger one* Tool calling, structured outputs, and what LLMs actually do* What it takes to support a new open model on day zero* Grafting Kimi's vision encoder onto GLM-5.2* Retrofitting inefficient model layers with components from other architectures* Why models sometimes collapse into repeating the same token* How hardware, kernels, and race conditions create nondeterministic failures* Preserving model fidelity while making inference faster* How quantization errors can cancel each other out* Why inference optimizations still deliver gains of 20%, 100%, and 200%* How optimized serving can make a model up to 10× faster* NVIDIA Dynamo, KV-aware routing, and distributed model serving* Speculative decoding the speculative decoder* Why local AI is about making models less dumb while data-center AI is about making them less slow* Tensor, expert, and pipeline parallelism across GPUs* Hardware-aware model design, auto-tuning, and the case against mega kernels* Rubin and why inference is becoming a systems problem* Whether modern GPUs are evolving into programmable AI ASICs* Why enormous models like Kimi K3 require GB300-class hardware* Why open-source video generation still trails Veo, Kling, and other closed models* The quadratic attention bottleneck behind long-form AI video* Autoregressive video, real-time generation, and compounding quality drift* Why future video systems may combine autoregressive and diffusion architectures* Training for inference and inference for training* Continuous post-training, deployment, evaluation, and improvement loops* How GLM-5.2 helped optimize the kernels serving GLM-5.2 itself* Why faster networking could unlock dramatically faster decoding* Continual learning, KV-cache compaction, and persistent model memoryShow Notes* How to build a day-0 API for Kimi K3* 22580: From GPT2 to Kimi3, ExplainedPhilip Kiely* LinkedIn: https://www.linkedin.com/in/philipkiely* X: https://x.com/philipkiely* Inference Engineering: https://www.baseten.co/inference-engineering/Ali Taha* LinkedIn: https://www.linkedin.com/in/aliestaha/* X: https://x.com/waterloointernTimestamps00:00:00 Introduction and the 200K-Token Prompt00:03:18 Dedicated Deployments, Speculative Decoding, and Tool Calling00:11:26 Launching Production-Ready Open Models00:19:06 Model Retrofits, Failure Modes, and Nondeterminism00:28:22 Quantization and Canceling Errors00:32:15 The Race to 10× Faster Inference00:40:48 Dynamo, Speculation, and Local vs. Data-Center AI00:50:18 Model Parallelism, Auto-Tuning, and Mega Kernels01:00:55 Rubin, GPUs vs. ASICs, and Custom AI Chips01:10:03 Giant Models and the Limits of GPU Memory01:12:42 AI Video, Quadratic Attention, and Autoregressive Generation01:21:47 Audio, Images, and Diffusion Models01:27:32 Training, Self-Optimizing Models, and Continual Learning01:40:06 Closing ThoughtsTranscriptIntroduction: Baseten, Waterloo Intern, and Inference EngineeringSwyx [00:00:00]: Okay, we're here in the studio with Philip, old friend from Inference Engineering, the book, as well as Baseten and everything that you've done, you and I have done before, as well as Ali. Welcome.Ali [00:00:15]: Pleasure to meet you.Swyx [00:00:15]: Waterloo intern.Ali [00:00:16]: Waterloo intern, always.Swyx [00:00:17]: When did you get “Waterloo intern” as a handle?Ali [00:00:19]: As a handle? Oh.Ali [00:00:20]: I think the rebranding happened mid-March. When I saw it was open, I was like, “I have to take it. Up for grabs.”Philip [00:00:26]: The problem is that Ali is really good at his job and is not gonna be an intern much longer.Philip [00:00:30]: So we have to figure out who's gonna get the handle.Ali [00:00:33]: Well, I'll pass the torch over to the next intern.Swyx [00:00:34]: Oh, okay. It can be, like, you just pass it to another Waterloo grad.Ali [00:00:37]: To another Waterloo intern. No, bruh.Philip [00:00:39]: Yeah.Ali [00:00:39]: Intern.Swyx [00:00:40]: Intern, yeah.Ali [00:00:40]: And no.Philip [00:00:41]: You gotta get an intern from Waterloo.Ali [00:00:42]: Yeah, I've gotta get an intern from Waterloo.Swyx [00:00:44]: Right.Ali [00:00:44]: But they have to follow the path.Swyx [00:00:45]: Oh, it could, but it could come from Baseten, so it's like whoever Baseten gets from Waterloo.Ali [00:00:48]: Right.Swyx [00:00:49]: Has the title of Waterloo.Ali [00:00:50]: It stays in the ecosystem.Philip [00:00:51]: Exactly.Ali [00:00:52]: Halfway through the internship, you either get it or you're out.Philip [00:00:55]: You should also do, like, a big graduation ceremony where you change the handle.Ali [00:00:59]: Just say it.Philip [00:00:59]: For everybody.Swyx [00:01:00]: You guys are good at ceremonies, clearly. We had a nice launch of the book, very successful. But before we get into all that, I wanna start off with a fun question for you. Okay, you're an expert inference engineer. What happens when I send a long query, say two hundred thousand tokens into Baseten's inference? What's the process of query through GPU model routing, balancing, all that? What is all the stuff that we don't think about?Long Context Requests, KV Cache, and Cache-Aware RoutingPhilip [00:01:26]: With a long query specifically, the first thing that I'm gonna ask is, “Have you sent me this query before, or at least part of it?” and I really hope you have, because it's gonna be a lot easier for me and a lot cheaper for you. So the first thing that we're gonna look at is some cache-aware routing, where we're going to see, we probably have a number of instances, a number of replicas up serving whatever model you're hitting. We want to send this one to something with, number one, available prefill workers, and number two, ideally some cached input already there so that we can skip prefill on at least part of these two hundred thousand tokens. If you're doing two hundred thousand tokens, it's probably coding or a multi-turn agent or something where you would expect to have that cached. If you don't, we're gonna have to send it to a prefill worker. We've at least on certain models disaggregated prefill and decode, so you're going to have one set of GPUs that's solely going to process the input, create the KV cache, and get you your first token, and then that's going to be passed over to a separate set of GPUs, which is going to run decode. We're going to iteratively make those tokens. We're probably going to have some speculator model in front of that. I'm going to assume that you're doing coding, and because of that, our speculator model, which assumes you're doing coding, is gonna have a high draft token acceptance rate. If I'm wrong and you're asking me to summarize every Harry Potter book, it's gonna be slower. And then we stream that output to you and account for it, charge you, a couple of pennies and say, “Hey, would you like to send another one?”Swyx [00:03:04]: Except Baseten doesn't charge by pennies.Philip [00:03:07]: Well, yeah, we charge. I'm assuming that we're talking about the public model APIs. If you are setting up a dedicated deployment, then yeah, it's not pennies.Public APIs vs. Dedicated DeploymentsSwyx [00:03:18]: Yeah, one of the key differentiators when I was talking with Baseten initially was that people who want very high volume just need to rent by the box, ‘cause then it's up to you to figure out how to saturate the box.Ali [00:03:31]: And more often than not, it's, like, way cheaper if you're pushing, like, millions of tokens per hour, if you just pay per hour instead of pay per token.Philip [00:03:37]: Yeah, they do. I think that we've increasingly seen a lot of demand for the pay per token APIs, just because everyone wants to try open models, and then once they find a use case that's really sticky, then they move over to dedicated.Swyx [00:03:51]: Is there a best practice on when it's time to swap over?Philip [00:03:54]: Couple reasons. Yeah, reliability, that's a big one, right?Ali [00:03:57]: Like, if they have a very specific use case, they want you to train something specifically for them, like they want their own spec dec, for instance, for their own traffic.Swyx [00:04:04]: Spec dec is speculative decoding.Speculative Decoding and Custom SpeculatorsAli [00:04:05]: Speculative decoding, yeah.Swyx [00:04:07]: You have to explain.Ali [00:04:07]: Sorry. Like, speculative decoding is like, if you have a huge model, right? And so the model is going to be generating one token at a time every single turn, every single forward pass. So we attach, like, this little, like, parasite, like this layer that goes on top of the model, and this model just has to predict. It does three very fast autoregressive forward passes, and it will predict, like, three certain tokens, and then you do one forward stage over the entire original model in order to see if those predictions were correct or not, and then you accept them or you reject them. Now, this draft model is traffic specific, so if you, like, Philip said, if you're summarizing Harry Potter books, I can train exclusively that draft model on Harry Potter books, and I can guarantee you that I'm gonna accept the three tokens every single time. And so with that case, I increase your decode speed. I wouldn't be able to provide this to you if you're a shared endpointSwyx [00:04:53]: YeahAli [00:04:53]: ‘cause I have no idea if you're doing Harry Potter, if you're doing coding, if you're doing English. We don't know. Also, there was a thing in the book that mentioned that if they really cared about a specific threshold, chapter four, I think. Do you remember that?Philip [00:05:06]: Yeah. The things that you can do is you can set a specific, like, batch sizing, a specific, like, parallelism strategy if you're trying to optimize for, like, throughput versus latency. You can. Maybe a NVFP4 quant doesn't pass your benchmarks and you wanna run a model at higher precision, you could do that. There's just a bunch of reasons why you might wanna have your own endpoint and the biggest one, of course, just being, like, you don't have to deal with someone else doing a hundred million tokens of benchmarking traffic at the endpoint when you happen to be trying to serve your users.Swyx [00:05:40]: Yeah. I think one thing that is. That is a classic journey. Like, it's people is asking the, what happens when you type Google into the browser. Tool calling, is that just, you're generating JSON or is there more complication beyond that?Tool Calling, JSON, and Structured OutputsAli [00:05:58]: Certain customers that we have, they have their own post-trained models, and so they demand a tool calling that's not just, like parse a file or go find the weather. It's something that's very specific and you have to do post-training on this. And if the post-training on the model is not good or if the quantization after the post-training to get the inference to be fast, the model will struggle reading the JSON file and reading the tool calling. But it doesn't require its own like sandbox. It's not like it's going to use that tool calling to like escape a sandbox or like it doesn't have to be contained. It can just be a normal dedicated deployment. The challenge with tool calling more and more seems to be that the companies want certain tool calling which is a very sensitive thing to train. And because you're dealing with all of the JSON outputs, if it doesn't like close the end of the request in a very certain manner, you end up with a model that did the tool calling and like the thinking and so as a result of that, it didn't see the result and just hallucinated the result as it decoded. That seems to be the most challenging thing with tool calling, not really the sandboxes model.Philip [00:06:56]: Yeah, that's a challenge on the training side and then on the inference side, there's work that you can do to scope the possible output. So we published this at this point close to two years ago, the solution to this problem which is you make a state machine and you use that to constrain the output to a specific format. So this is the structured output problem. If you remember backSwyx [00:07:27]: Yeah, the specific grammar is,Philip [00:07:29]: Yeah, exactlySwyx [00:07:30]: GML had this thing.Philip [00:07:31]: Yeah. So it's like the old-school “make sure this is only JSON”, return only JSON orSwyx [00:07:38]: YeahPhilip [00:07:38]: Grandma's gonna die type of prompts.Swyx [00:07:39]: Is it BNF grammar? At some point OpenAI had released a thing that was like, yeah, if you want to constrain your output, write BNF grammar, back as NOR.Philip [00:07:47]: In our inference system, it's just a specified output format. And you get the guarantee that your output's gonna be structured along that format. And so applying that to tool calls can like help cut down on. You can still call the wrong tool or call no tool. It doesn't solve the certainty problem but it at least solves the output structuring problemSwyx [00:08:10]: YeahPhilip [00:08:10]: Within tool calls.Swyx [00:08:12]: And MCP is just another form of tool, right.Philip [00:08:14]: Yeah, exactly.Swyx [00:08:15]: As far as there's no special thing there.Philip [00:08:16]: The thing I'm always like explaining to people is the LLM is not capable of doing anything. It's only capable of making suggestions of what to do and then if those suggestions are formatted in a certain way and applied to a system that knows what to do with them, then an action occurs.Swyx [00:08:32]: Yeah. Part of the fun stuff is, this is solved outside of tool calling too. Like in an agent loop if the output is not correct or you're right, like reasoning, tool calling was done in the reasoning trace, just be like, “Oh, I don't know what to do. Let me just try again.” And it might get there after a few tries. And on your point of training, sometimes this is harder in smaller models, so you don't have the same exact quality outputAli [00:08:56]: Right.Swyx [00:08:57]: When you just swap from a big model, right?Ali [00:08:59]: Yeah. I will say that, before, I think we need to go back to inference engineering proper.Ali [00:09:04]: But, I had expected that something would replace JSON because it's hard to stream JSON ‘cause JSON must be complete and you must have open and close brackets and everything. So it's hard to parse something or validate something while it's being streamed. So people invented all sorts of things that are like, I forget the name of some of these alternatives, but it's something like TOML, something like YAML. But JSON seems to be dominant still.Philip [00:09:30]: The JSON outputs aren't that long, right? Like you could have a long-- ‘cause tool calls also contain the arguments in them and perhaps for a certain tool you might pass like a very long argument. But my impression of the median tool call is that it's a relatively small number of tokens, right? So I would expect that speculators are generally fairly good at something as formatted as JSON. And so you would have like a pretty fast decode step there and that the streaming wouldn't be as valuable, but maybe I'm wrong about that.Ali [00:10:02]: I think you're also bounded by the software or that the model is gonna integrate with if the software is built with JSON for the tool calls or if the company that you'- if your customer says that this is how our software works and our tools are interfaced with JSON, you can ask them to like, change their software and say like, “Yeah, this is gonna be better for the model.” but like with the right training shouldn't be that much of a difference. Also more profitable if it outputs more tokens probably.Swyx [00:10:25]: Depends on your business model.Swyx [00:10:27]: It really depends. But I will say that, as a writer with like experience a lot with generated output, I do try to move from text to JSON text which is very long JSON, right? Like there's paragraphs in every field because I'm trying to structure it, right?Philip [00:10:44]: Right.Swyx [00:10:44]: I want you to first make factual statements, then make opinions then make bullet point summaries, have dates, have entity references have your sources for references, all these things. Anyway, so these are things that like I think people who really experiment with structural output have to really care about. But, let's, let's recurse up the stack a little bit. Before we started recording, you mentioned something really cool, which is that there's a lot of engineering that-- inference engineering that goes on when a new model provider releases a new model, right? So let's call it GLM-5.2, Kimi K3. I had previously assumed, especially if it's like, well, GLM 5 to 5.1 to GLM-5.2, like that you've supported them before. Is it that much work?What It Takes to Support a New Open ModelAli [00:11:26]: It's a lot of work.Swyx [00:11:28]: Yeah. Okay. So like, a lot of people, all you guys, right whenever a new model launch like, people rush to say like, “Oh, Hugging Face supports this, Fireworks supports this, Spacetime supports this,” and I'm like, “Yeah, of course we support it.” But what goes into that? What goes intoPhilip [00:11:40]: I think it's more than just support it too, right? It benefits the consumer a lot. Like I think it was with Kimi K2.5 or GLM-5.2 the latest, there was an inference war, right? X provider is at 90 tokens a second. The next day we're at 150. The nextSwyx [00:11:55]: I kinda kicked that off with the GLM-5.2.Swyx [00:11:58]: I wrote a Twitter article about. It got like half a million views,Ali [00:12:02]: Based on being numberSwyx [00:12:03]: YeahAli [00:12:04]: Or it's for something else.Swyx [00:12:05]: Yeah. Which,Ali [00:12:06]: Oh my GodSwyx [00:12:07]: Which then got everyone really excited about, hey, how can we, bend tracks a little bit further and,Philip [00:12:14]: There's a difference between support the model, as in I can make a token out of this model, and support a model, as in I have a production-ready API from this model.Philip [00:12:26]: Getting to the point of I can make a token out of this model is not that hard because generally the, open source inference engines, vLLM, SGLang of the world oftentimes even receive weights ahead of time, maintainers do, or the people making the model merge PRs to ensure support. So you generally can, just get it working on the standard open source stack without too much pain in most cases. The challenge is, every inference company is gonna have own proprietary stack. Some open source components, some in-house stuff. And for any arbitrary model, there's going to be some new stuff. Sometimes you get lucky, like K, two five to two six was, like, pretty similar.Quantization, Speculators, and Production ReadinessAli [00:13:16]: Yeah. It was pure continued post-trainingPhilip [00:13:18]: YeahAli [00:13:18]: If I remember correctly.Philip [00:13:19]: Even in those cases, there's still stuff you have to do. You have to redo the quantization work. You're taking the model from. Generally, these models are not released in NVFP4, and we want them to be in NVFP4 for maximum Blackwell compatibility. So we have to perform that quantization, and, calibrate the quantization to make sure that we're not causing any regression in the model's intelligence. And then we also have to train the speculator, as we've talked about. Generally, we have. We have ZDR, zero data retention on our model APIs, so we don't know exactly the traffic that people are sending us, but we know what's popular. We know that coding use cases are popular. We know that agents, agentic use cases are popular. So we can get public data sets that are representative of that traffic and train general speculators. Now, with speculators today, you need to train the speculator using the base model itself because you're getting hidden states out of the model from running inference on these specific prompts, and that is the training data you use to create the speculator. So there's that process which you need the real model weights for. And then there's of course just the process of, standing up all the infrastructure behind it, loading all this stuff, testing it. And then when there's a new model with a newer architecture, I think that, like, the DeepSeek models tend to be the most challenging as they have, like, the most novel architectural stuff going on, model after model. But every new model has something. Kimi K2 had. Oh, sorry, GLM-5.2 hadAli [00:14:53]: Sparse attention.Philip [00:14:54]: Yeah,Ali [00:14:54]: YeahPhilip [00:14:54]: the DSA.Ali [00:14:55]: Right. Which is brought from DeepSeek.Philip [00:14:57]: Yeah. AndAli [00:14:59]: So you can copy-paste then?Philip [00:15:01]: It kindAli [00:15:01]: I don't know how this works.Philip [00:15:02]: So, like we had to, like, build support for that into our runtime. And you're right, like it is really interesting the way that all of these open source labs borrow from each other. For example, like GLM-5.2 doesn't have vision. So something that, Haley, a guy on our team, if we could take a look at this, he, like, grafted the Kimi vision encoder onto GLM-5.2.Retrofitting Vision into GLM-5.2Ali [00:15:27]: We'll be training the projector.Philip [00:15:28]: Exactly. So if you think about, like, the encoder, there's the encoder, which is the part that looks at the image and turns it into latent information, and then there's the projector which likeAli [00:15:38]: You can say latent space. It's okay.Philip [00:15:41]: And then there's the projector that maps it onto, the model itself, and then there's the model weights. You don't wanna mess with the model weights because you run a chance of making the model dumber at something else for the purpose of giving it vision. So instead, Haley started with just a projector, which is only a handful of millions of parameters.Ali [00:16:02]: That would be, yeah.Philip [00:16:02]: Yeah.Ali [00:16:03]: Can you show the training one?Ali [00:16:04]: Like the way it groksPhilip [00:16:05]: YeahAli [00:16:06]: Very interesting.Philip [00:16:06]: And maybeAli [00:16:07]: That right therePhilip [00:16:07]: Maybe Ali, you should take it from here. You've got a betterAli [00:16:10]: Ooh, double the sandPhilip [00:16:11]: Understanding of this than I do.Ali [00:16:11]: Yeah. You can see, like, he. The way he trained this is really cool. At the beginning, he was training it using just like, “Here's a picture of a mountain. Can you describe what's in this mountain?” And that caused it just like the first, learning walls. Like here you can see this all we're trying to teach it is to translate the encoded. Like it's already taken the encoder from Kimi K. It's taken the image. It'Philip [00:16:31]: Yeah. FrozenAli [00:16:31]: FrozenPhilip [00:16:32]: With adapter.Ali [00:16:32]: Exactly.Philip [00:16:33]: Yeah.Ali [00:16:33]: So the brain is frozen and the eyes are frozen. It's just we're tryingPhilip [00:16:37]: AlignAli [00:16:38]: Interconnect between the eye and the brain, right? So the projector. And so you take the tokens and then he's like, “Oh, can you describe what's in this image?” And he's like, “Oh, it's a mountain,” or it's a person or it's a human, whatever the case is. But that didn't cause complete understanding. So he changed it such that every image was associated with a data set of questions. Like, does this image have a white male? Does this image have birds in the top corner? Does this image have a scientist in it? All of that stuff. And it would have to answer questions correctly. And using not just training on describing an image, but being able to answer question, another question, answer over time. Like you can see the grokking, which is like genuinely insane, that retrofitting vision into a large LLM can learn to that extent. And even for images that it doesn't perform well on, for instance, if you ask it a picture of like Stephen Hawking, “Who is this?” Maybe it doesn't get it, but it will say something like, “This is Albert Einstein.” Like it still understandsPhilip [00:17:25]: Close enoughAli [00:17:26]: That this is a scientist who is a man who has, some significant achievements, all that stuff. So that's like really cool.Philip [00:17:32]: Yeah. So, we've covered Hao Tian before, who the author of the LLaVA paper that did this, a while ago. And I think that's very foundational work for anyone who hasn't done vision work before.Ali [00:17:41]: Same with the CLIP and MetaCLIP, where you go from just captioning to building out questionsPhilip [00:17:47]: RightAli [00:17:47]: Off the image and how much better you can get performance.Philip [00:17:50]: Right. Right. Right. Yeah. But what's, what's so exciting about this is if you look at a model like this. Now, this is a little bit more of a research project. It's not. It got to 56% on MMLU Pro, I think. So not quite frontier. But if you're running this model, you haven't suffered any loss on your GLM-5.2 quality. If you don't have an image, it'll just behave exactly the way it used to. And ultimatelyAli [00:18:14]: Which in the inference code you literally do not include the other part, right?Philip [00:18:18]: Yeah. You would just skip the encoder if you don't have an image input.Ali [00:18:22]: Okay.Philip [00:18:22]: Just confirming.Philip [00:18:23]: YeahAli [00:18:23]: Does it affect a lot on the overall inference side? Like you're not adding much, you're adding a very small vision encoder. These are typically likePhilip [00:18:30]: They're super fineAli [00:18:31]: Less than a billion parameters, right?Philip [00:18:32]: Yeah. It's, - There's a little bit less standardization among vision encodersSwyx [00:18:37]: YeahPhilip [00:18:37]: So the support matrix can be a little bit, sparser. But overall, yeah, it's a pretty, it's a pretty minor component of the overall system. And ultimately what you get out of the system is all of a sudden you have Kimi Vision, GLM weights, and DeepSeek attention all in one model.Open Source Model Grafting and Franken-MergesPhilip [00:18:56]: And that's, I think, a lot of the power and beauty of open source, is that you can take all of these different components and combine them together into a system that's better than anyoneSwyx [00:19:05]: YeahPhilip [00:19:05]: Can be individually.Swyx [00:19:06]: People used to say that you would also do Franken-merges where you would take likePhilip [00:19:10]: YeahSwyx [00:19:10]: Layers from each model.Swyx [00:19:11]: Does anyone do that anymore?Ali [00:19:13]: Well, to your point previously when you were mentioning like, the work that goes into supporting a model when it first comes out, like GLM-5.2 or MiniMax M3 or whatever the case is. Sometimes you do have to like, you do have to switch out some things. Like, for instance, the MiniMax M3 head uses full attention, and with full attention you end up with this like insane bottleneck in spec dec ‘cause you're doing auto-regressive token generation for three tokens, and you're doing this like N squared over all of the tokens that are in your sequence. Your KV cache is like very large because it's not sparse, it's not top K. So we find it better to like, okay, we're gonna replace this, we're gonna replace this layer with a layer from another model that's using like GQA, for instance. And then just with the right training, you can get it to have the same acceptance rate. So it is very possible to retrofit layers from other models and very much needed. If a layer is like inefficient, the training just becomes the challenge, like how do you ensure that you train it properly? Which again to your earlier point is like the mesh between training and inference. As in like you need very good training in order to do fast inference. That's like, I feel like more and more becoming true.Swyx [00:20:21]: Yeah. Anything else on the support side when you say like get it to fully production ready?Loop Detection, Race Conditions, and Non-DeterminismPhilip [00:20:26]: Yeah. I think that there's also a question of just, we can test a model to a pretty extensive degree, but we're trying to get it out quickly and then you see a bunch of other people test it and you get interesting results. There was an issue with, GLM briefly where we had some like mode collapses where it would just output the same token over and over again for certain prompts on certain temperatures. Like once you expose an endpoint to the real world, there's going to be, so many more varieties of things given to it that you're able to, discover and patch things. So it's not just a, day zero process, it's then like for the first week, for the first month, if a model remains popular, like how do you both fix bugs and then continue to push the envelope on performance?Ali [00:21:21]: What do you mean you don't want your model outputting S?Swyx [00:21:24]: Is there loop detection on that stuff, by the way? It still happens like quite a lot, which is surprising.Ali [00:21:30]: We have like we, in our endpoint, like if a model was to output the same token like four plus times, we just cut the generation. We say like, “Oh, sorry, this-- Like try again,” or like we will reprocess the request. ‘Cause we know then, like if it, like if, yeah, it's four times the same token, it's probably collapsed.Swyx [00:21:45]: Yeah. Is there a way to opt out in case I really want that?Ali [00:21:48]: You want that?Ali [00:21:50]: I think there's a way that we have to handle it. I'm not exactly certain, but I feel like in certain models, like when they output something like you can imagine, like a table for instance, and so they want, they wanna draw like 12 dashes and 12 dashes. Yeah, I think there's a way for that to happen. I think we only do it on certain tokens. Like we exclude certain special characters.Swyx [00:22:07]: Yeah.Ali [00:22:07]: So we only do it on like certain like S is the most common almost. GLM-5.2Swyx [00:22:11]: OhAli [00:22:11]: And I think it was DSV 4 as well. Like you'd just have like looping issues where like you literallySwyx [00:22:17]: ItAli [00:22:17]: Just have like S.Swyx [00:22:18]: Yeah. Is there a special, something special about S? No, just randomlyAli [00:22:21]: It just seems to be the one token involved.Swyx [00:22:23]: Yeah. And it'Philip [00:22:24]: Is thereSwyx [00:22:24]: And it's only temperature 0Ali [00:22:27]: NoSwyx [00:22:27]: Even at other temperaturesAli [00:22:27]: Even at like 0.9 or whatever, it will still, it will still collapse.Swyx [00:22:30]: That's weird, right?Ali [00:22:30]: It's, it is an inference problem to be honest, like a software problem. Like oftentimes, the image you run will-- like NVIDIA will release an image for instance, and if we will upstream the changes from their latest TensorRT-LLM image into our stack, we'll find that it fixes it. Or oftentimes this will only happen in an inference engine that you're using like SGLang. But if you were to switch to vLLM, that isn't the case. So it seems to be like an extremely like deterministic software issue and not really a model issue. It's not like a weights problem. Like I'- we'll say like, “Oh, it's a problem with the quant. We did PTQ wrong,” right? But that isn't, that doesn't make sense because the same weights used with a different inference engine does not repeat the problem. And sometimes it's, the kernels that are being used in the backend have like these very subtle sometimes race conditions, where if you were to use this model hosted on one cluster, you will never get this problem.Swyx [00:23:19]: Oh my God.Ali [00:23:19]: But if you host it on a different cluster, you will. And the reason is the KV cache transfer from a node to node in that one cluster is using a slower interconnect than the node to node in another cluster. So that exposes the race, whereas in another cluster it doesn't. So then you end up just like, okay, this model is not gonna be hosted on this cluster. We're gonna host it on, another cluster because that cluster exposed that problem. But then it ends up with like, okay, is it the software? Is it the model weights or is it the hardware?Swyx [00:23:42]: There is a thing about this with temperature 0 still not being deterministic, right?Ali [00:23:46]: Right.Swyx [00:23:46]: Mostly because of hardware. Even at temperature 0 same model, you won't always get the same output.Swyx [00:23:52]: Even-- But I'm surprised by the race condition one because, I thought PyTorch was a graph that like guarantees that you at least, execute things in the right order.Ali [00:24:02]: Well, yeah, true. Like I'm not, I'm not saying that there is. Like well, you have things like PTL optimizations where like you can start a kernel before the end of the previous kernel, and that's like ‘cause you want to do that because there'sSwyx [00:24:12]: It's like pipeliningAli [00:24:12]: Expense. Exactly.Swyx [00:24:13]: Yeah.Ali [00:24:13]: But it'- But you don't do it cleanly. Like you overlap a little bit of the execution. No, it is very possible that the kernel itself, like that one block that is supposed to be running in this instance of time, that kernel itself has a race condition. For instance, like a missing barrier. Like often if you're designing a kernel and you want it to make it to be very fast, if you don't test it extensively, you'll, you'll have certain threads access data points from registers before they've been written to by other threadsSwyx [00:24:36]: YeahAli [00:24:36]: For example, because like your barrier is wrong or your synchronization was wrong. But yeah, like the testing itself is very difficult in those like, andSwyx [00:24:42]: And there's no like borrow checkerAli [00:24:45]: What does that mean?Swyx [00:24:46]: Like Rust. Like the. If you're trying to have like memory safety It sounds like a comparable problem.Ali [00:24:52]: Well, yes, but you're working in CUDA, right, NVIDIA GPUs. Like- You just need a higher level language like modular Maybe that's what modular is supposed to do. I don't know.Quantization Quality and Vendor FidelityVibhu [00:25:00]: How do you see keeping quality of the model? So you talked about all these steps of, okay, you gotta do quantization, train your own speculative decoderAli [00:25:07]: RightVibhu [00:25:07]: Run on different hardware. Looking at other model providers, okay, you kicked off a inference speed race on the consumer end. What goes into keeping quality the same across them, right? Sure, you can run benchmarksAli [00:25:22]: YeahVibhu [00:25:22]: But, like, how do you determine how much quantization are there standards? What goes intoPhilip [00:25:27]: There's a few things on quality. Most inference optimizations are lossless. KV caching, for example. You are just recomputing or preventing recomputing the same values. Speculation, of course, if a draft token is wrong, it gets rejected. The main lossy optimization is quantization. And that really comes down to, number one, data format, number two, which parts of the model you choose to quantize, which layers, and number three, like doing a lot of calibration on the quantized weights, to ensure that you're preserving all the outliers. There's other tricks that you can do, though. A big one is long context, ‘cause one thing you asked at, right at the beginning is, “Oh, what's gonna happen if I send a 200,000 token request in?” So with a long input sequence, you need to, store a lot more information. You need to process a lot more tokens. And so even if a model has a context of a certain length, you might, as an inference provider, choose to build an API with a shorter context length, and of course a full length one as well. Because if someone doesn't need the full million token context, for example, you can get them better performance. I don't know if that's exactly like quality of the model. The way that I think about quality is to what degree are we faithfully serving the original model? If you think of a golden implementation of a model that performs exactly the way the model is designed to perform, I think of quality as how close are we getting to that, 100% fidelity of the model.Philip [00:27:13]: You can also, of course, think about quality from the training side and how do you push yourself past 100%. But when I think about purely inference optimizations, it's getting faster while staying as close to that 100% fidelity mark as possible. And certainly our standard internally is that, like you should not be able to tell the difference between our API and a, official API. I think Kimi in particular does a good job of vendor benchmarking hereAli [00:27:41]: YesPhilip [00:27:41]: Where they haveAli [00:27:42]: They released an actual vendor benchmark.Philip [00:27:43]: Exactly, yeah.Ali [00:27:44]: ‘Cause they accused, some people, Amazon? There was some provider that was not doing very well on Kimi's benchmark.Philip [00:27:50]: Yeah.Philip [00:27:51]: So, with Reflect we probablyVibhu [00:27:52]: This was a long time ago, right?Philip [00:27:54]: No.Ali [00:27:54]: Yeah, like threeVibhu [00:27:55]: They alsoAli [00:27:55]: Four, five months agoVibhu [00:27:57]: This also happened with, I don't remember which model, but they pulled out quite a few, and then they started a whole chart about this. It might have beenPhilip [00:28:03]: Kimi Vendor Verifier.Ali [00:28:04]: Yeah.Philip [00:28:05]: Yeah.Ali [00:28:05]: Yeah, ‘cause you, ‘cause you'd be pissed, right? Like if you'Philip [00:28:07]: Yeah.Ali [00:28:07]: If like if I'm a consumer and I'm using like Amazon's endpoint for instance, and I've used Kimi and I'm like, “Oh my God, like this is bad,” I'm not gonna say, “Oh, Amazon quantized the model in a bad way.” I'm gonna say, “Oh, Kimi sucks.” Right?Philip [00:28:17]: Yeah.Ali [00:28:17]: So it seems like that makes sense.Philip [00:28:19]: Yeah, they care. They care.Vibhu [00:28:21]: Justifiably.Ali [00:28:21]: Yeah, justifiably.Vibhu [00:28:22]: This is probably a stupid question, but just checking, has anything improved from main quantization?Philip [00:28:28]: Yeah.Vibhu [00:28:28]: Like, is quantization always strictly worse?Ali [00:28:30]: Well technicallyVibhu [00:28:32]: NoAli [00:28:32]: It's a lossy. QuantizationPhilip [00:28:33]: YeahAli [00:28:33]: Is a lossy, it's a lossy implementation.Philip [00:28:36]: Speed improvesVibhu [00:28:36]: Speed improves.Ali [00:28:37]: It the number, likeVibhu [00:28:38]: No, I' always look for inverse scaling laws.Philip [00:28:40]: Yeah.Ali [00:28:40]: Yeah.Vibhu [00:28:40]: This is something I learned from Noam Brown, where like things that normally act in one direction sometimes do.Philip [00:28:45]: Well, technically when you run a benchmark, because these models are deterministic, sometimes your,Ali [00:28:52]: YeahPhilip [00:28:52]: NVFP4 quant is like, two basis points higher than yourAli [00:28:56]: No, it's noise. It's noise.Philip [00:28:57]: Yeah, exactly. I'm like, yeah, it's, it's within. That's why I always say within margin of error.Philip [00:29:01]: And I stopped saying that because everyone assumes that what is, well, within some margin of error, we're barely inside of that to the worst, so we're saying. But yeah, sometimes it's just like, gives you a higher output score. But like Ali said, that's noise. To my knowledge, you're not necessarily making the results better. You're just trying to, again, like keep your fidelity as close to 100% to the original model.Layer Selection, KL Divergence, and Better QuantizationAli [00:29:27]: There is, to your point, research that we did on MP. I don't know if you are able to pullPhilip [00:29:31]: YeahAli [00:29:32]: A tweet we did. One of our research interns, Joshua, I think it's a tweet on how we have 20% better quantized GLM-5.2 than NVIDIA. Essentially what we found throughout like this month research is, okay, quantization is a lossy. It's. You're compressing the data from, occupying 16 bits to occupying, four bits, for instance. And so you're losing some information, and you're trying to minimize that. And so when I say that I'm gonna quantize the model, my job becomes how do I find the layers that I can quantize, and how to find the layers to not. For instance, with image models, I don't quantize modulation layers, and I don't quantize out projections because those two are. Like out projection is what you see as the user. Modulation is what the model sees or understands. Right, exactly. And so to his paper, do you have the. It doesn't have the. Yeah. It's a long paper. I don't know if I can findVibhu [00:30:25]: If there's a part to search or it's probably in the thread.Ali [00:30:28]: It's probably in the thread.Vibhu [00:30:29]: Yeah.Ali [00:30:29]: But the long and the short is it is very possible that quantizing more of the model makes the results. Like if I have a model that I quantize layers one, five, and 10, and another model where I only quantize layers one and It is possible that the model in which I quantized more information is going to perform better because the quantization errors have canceled out. And so what Joshua showed in his mathematical proof where he had like a verifier in, is that you can predict which layers are going to have quantization errors that will cancel out with each other, and you choose to quantize those layers. And so the result of doing this mathematical quantization is you end up with a model that's 20% more quantized than another provider, so you get 20% more throughput of it because there's more layers than running an NVFP4, and your quality is better than that other quant because the layers that you chose to quantize have their errors cancel out, like one layer skewed to the right one layer skewed to the left, one layer skewed to the right. Your final logits distribution is more similar to the original distribution of the model, so you have better fidelity. And so the way we proved this was with KL divergence. So instead of just scoring on the benchmarks, we scored the KL divergence between the logit distribution of the quantized model and the logit distribution of the original full precision model, and we showed that with this technique we get. If your probability distribution on the logits which token it wants to select is more of the same as the original model, you're probably gonna end up staying true to the original model. So yeah, so it seems like previously before this, it seemed like the industry was, well, the more you quantize, the worse it's gonna be, ‘cause the more loss you introduce. That's not exactly, not necessarily true. So yeah, doesn't improve it, but can cancel out.Philip [00:31:57]: I think it might be this, but reminds me a good bit about pruning where you can prune off certain layers.Philip [00:32:03]: But very interesting. Didn't know this was a whole paper you guys put out.Ali [00:32:06]: It's. Fun fact, it was originally 72 pages, this paper, and then we decidedPhilip [00:32:11]: WowAli [00:32:11]: We can't tell. We couldn't release it. So it's now 45.Swyx [00:32:15]: Still 39 pages, so very substantive. We talked about evals and all these things and, like what's possible in terms of speedup? Like it's like probably like the numberInference Speedups and BenchmarkingSwyx [00:32:25]: Thing that people do wanna care about, and it's something that you wrote about in your post. Like official API is 70 tokens per second, and you push it up to 90. Is that like a normal thing?Philip [00:32:36]: So what's cool about working in inference, the reason that I think inference is going to be a useful place to do engineering for a long time, is that if you look at highly optimized domains like, say, finance, if you're in finance, you measure how much better you got in basis points. It's like, “Oh, I got five basis points better, like twentieth of 1% better,” that's huge news because everything is so optimized. When we publish optimizations, it's 20%, it's 100% it's 200%. So there's still probably like a lot further to go, honestly. Like you'll, you'll know that inference is pretty much solved when researchers start publishing about how they got 1% faster at something.Swyx [00:33:19]: Which by the way, because I am from the finance background, in the ‘70s, that was the margin at the time. When you did quantitative finance research, you would findAli [00:33:27]: And like 20%, tens of percent.Swyx [00:33:29]: That's. Yes.Philip [00:33:29]: Yeah.Swyx [00:33:30]: And now it'Philip [00:33:31]: Tiny fractionsSwyx [00:33:32]: For those people interested, look up Andrew Lo's paper. He had a really interesting illustration of quant, stat arb, distribution, narrowing down from like those kinds of 20% differences in the ‘70s, down to nothing today, which is very cool.Philip [00:33:48]: Exactly, and we're at the beginning of the same type of thing. Now benchmarking is hard. I think anyone will tell you that, and benchmarking provider speeds is hard because there's so many variables that go into it. What hardware are you using? How much load do you have on the system? What's the exact nature of the prompts and input and output sequence lengths? All that stuff. But overall, when you start stacking these improvements, you're looking at multiples. You can look at it. The most common form, of course, is TPS, tokens per second, which is bad naming by us in the industry, ‘cause there's two tokens per second. There's tokens per second, the throughput number, and the latency number.Ali [00:34:31]: TTMT, yeah.Philip [00:34:32]: Like total tokens per second out of the, out of the GPU as a throughput number. Most people only care about tokens per second as the latency number, which we should call ITL, intertoken latency, but we don't.Philip [00:34:44]: Anyway, so you can imagine a standard API without many optimizations for a 1 trillion parameter model operating somewhere in the 30 to 50 tokens per second range for reasonable traffic profile. And we generally see the goal of, pushing to 10X that. But, not necessarily day zero, but by stacking enough optimizations, if you have, say like four optimizations, each of which doubles performance. Or sorry, three optimizations, each of which doubles performance, then you stack that up, that's an 8X gain. That's the order of magnitude that we're working with in this space. We're trying to make things substantially faster, not just go from like 70 to 90.Swyx [00:35:38]: Are you saying you've. You have done that?Philip [00:35:40]: So let's say you have as a reasonable baseline, 30 or 40 tokens per second. You can achieve 10X that. So like on GLM-5.2, if you run it unquantized, perhaps on H100s even, and you're just using an off-the-shelf inference engine with no particular optimizations, no speculator, nothing extra around like KV routing, no disaggregation, you're, you're probably, yeah, looking at that like 30 to 40. You think that's like a reasonable baseline?Swyx [00:36:12]: Right. Right.Philip [00:36:12]: To get to something like 10X, there's a lot of trade-offs that you're making. If we're running at more like a 300, 400 tokens per second range, you are using the best hardware possible. You have a optimized speculator. You have done all of your quantization work. You are Seeing a pretty high cache hit rate. You are running with a reasonably small batch size and a parallelism configuration that is tuned for latency versus throughput, but it is possible. So the spreads that you see if you, like, go on artificial analysis or you go on OpenRouter and you look at, the worst provider to the best provider, oftentimes can hit that range. 10X is of course very aggressive. It's oftentimes maybe more of a four to six times improvement. But that's the performance that makes us really excited, is when we can get these huge gains, not just go from 70 to 90 tokens.Stacking Optimizations: NVFP4, Speculation, and DisaggregationAli [00:37:19]: It's also, like, hardware dependent. Like, ifPhilip [00:37:20]: YeahAli [00:37:20]: If you have a thing where you're serving it on just, like, a node of H100s and then you throw, like, you shard the model across, like, four nodes of B200s. Like, you can definitely increase the speed with just throwing more hardware at it. Like, normalizing for the same exact hardware and the same number of GPUs.Philip [00:37:35]: Yeah. Then you're looking at, like, a two to 4X improvementAli [00:37:38]: Right. RightPhilip [00:37:38]: Depending on the inference optimizations. So yeah, it's. Some of it's, what's the call, and some of it's who's the driver.Vibhu [00:37:46]: If you break down the two to 4X, say the example is run GLM-5.2Ali [00:37:51]: YeahVibhu [00:37:51]: On B200sAli [00:37:53]: YeahVibhu [00:37:53]: Single node, right? What's, like, the cost trade-off for effort to get, like, the last bit of juice out versus what should people just think of, right?Ali [00:38:01]: Spectre quantization. Yeah.Vibhu [00:38:03]: Spectre quantization.Ali [00:38:04]: That's, that's, that's like 95%. LikeVibhu [00:38:06]: And how far does that get you? And how easy is that for the average person to do? So say right I wanna throw the weights of GLM-5.2 on a node of B200s, how easy is it to find speculative decoder- decoder model or already quantized model? How much work goes into it?Philip [00:38:23]: If you're doing it up front, it's quite a lot of work. If you're doing it today, there's going to be people who have published things that you can just, you can just grab some NVFP4 weights. You can grab a speculator. Yeah, if we're thinking about, like, what are the 2Xs we're stacking, going from, BF16 to NVFP4 is, it's not quite a 2X, right? It's like. I think it's about, like, 30 to 40%, from 16 to 8, and then another 30 to 40% multiplied from, 8 to 4. So that doesn't quite get you a 2X, but, like, roughly a 2X. Speculator, roughly a 2X. Disagg on top of that if you're able to get enough hardware and put enough traffic through it, another roughly a 2X. And then you add in some, double-digit percent increase from having just a better runtime with, the latest kernels and stuff behind it. And that's how it stacks up.Ali [00:39:21]: YeahPhilip [00:39:21]: So building each of those, like, building the, quantized weights is, for someone who really knows what they're doing, hours to days of work. Building the speculator, again, like, hours to days of work. And the, disagg setup, hours to days. Well okay, but like once you haveAli [00:39:39]: Once set up. Once set up. YeahPhilip [00:39:40]: Yeah, getting disagg working for the first time, I'm saying, of course, is very difficult.Philip [00:39:44]: The marginal implementationAli [00:39:48]: Like, if you're just grabbing, like if you are a person, like just a normal consumer who has access to, like, a node of B200s and you're wondering, “How can I just host it myself?” You don't need to quantize the model yourself. There's always gonna be, like, an open source quantized checkpoint. NVIDIA's gonna push one out if no one else does. You. Usually, the providers will have their own spec dec that they've trained as well. You don't need to train your own spec dec. You can just use that as well.Philip [00:40:09]: Yeah. Like, GLM-5.2 has its own MTP.Ali [00:40:13]: Right. Right.Vibhu [00:40:14]: What's multi token prediction?Philip [00:40:15]: Yes.Ali [00:40:16]: I'm justVibhu [00:40:16]: Can you explain that?Ali [00:40:16]: I'm just an expert.Ali [00:40:18]: I can do it for you in case I get it wrong?Vibhu [00:40:20]: No.Vibhu [00:40:21]: Yeah, you should correct if we're wrong, but their multi-token prediction can be used for self-speculative decoding.Ali [00:40:27]: I'm not sure. I'm not gonna correct that.Vibhu [00:40:28]: Okay. I'm semi-confident in thatAli [00:40:30]: Okay. YeahVibhu [00:40:30]: But someone can check. But it's useful to paint the story of, okay, not just the average person, but say a company wants to switch from serverless inference I wanna throw this up on. I wanna rent some GPUs, throw it up. These are the steps you take to do significantly faster than just put it behind vLLM.Ali [00:40:48]: Right.Vibhu [00:40:49]: I was waiting for a mention of Dynamo.Vibhu [00:40:51]: I feel like, that's supposed to be the baseline that you measure against.Dynamo, KV Routing, and Disaggregation ToolkitsPhilip [00:40:55]: I would think of Dynamo as less of a box system and more of a toolkit for building with. So when we talk about doing aware routing, when we talk about doing KV offloading, when we talk about doing, PD disaggregation, Dynamo fundamentally is. By the way, Dynamo is an open source library from NVIDIA.Ali [00:41:17]: We've done a pod with KylePhilip [00:41:18]: OkayAli [00:41:19]: Kyle Cranin.Philip [00:41:19]: Cool. So then your listeners know then that it supports all the different inference frameworks. And it is multi hardware, which is interesting.Ali [00:41:28]: But it's just a router, it's not like an optimizer layer.Philip [00:41:30]: Yeah. All it does, like, what Dynamo is good at, it is a library for moving information around your cluster, around your hardware. So if you have, KV cache on one place and you need it to be somewhere else, Dynamo coordinates NIXL for you to move that around.Philip [00:41:49]: That doesn't mean that, like, out of the box, you just say, “Pip install Dynamo,” and then you get, like, a massive performance speed up. It's more of a developer toolkit.Ali [00:42:01]: Yeah. I would have said it would. It comes with a set of defaults that you can then swap out.Philip [00:42:06]: It does. If the industry at large, I think, was, like, rolling out all of these deployments, standard, then I think it would be, like, a credible baseline. But, we've got to, we've got to benchmark against, like, what we're seeing in the wild.Speculative Decoding Methods: Medusa, EAGLE, n-Gram, and Spec-SpecVibhu [00:42:23]: I did wanna talk a little bit more about PD disagg, because that is probably, like, number three after quantized and speculative decoding. In your book though, I was just gonna pull out the book.Philip [00:42:31]: Yeah.Vibhu [00:42:32]: Like section 522 on Medusa, 523 on EAGLEPhilip [00:42:35]: YeahVibhu [00:42:36]: 524 on gram.Philip [00:42:37]: It's 55, would be disaggregationAli [00:42:42]: Yeah. Well, no, I just wanted to dwell a little bitPhilip [00:42:44]: YeahAli [00:42:44]: The other. Like, so what do you choose to include? What do you choose to not to include? Because there was all these other techniques.Philip [00:42:51]: Yeah.Ali [00:42:51]: Are these still relevant? Because I think they came out, like, a year and a half ago maybe.Vibhu [00:42:55]: Medusa is quite old.Philip [00:42:56]: Yeah, Medusa's old.Ali [00:42:58]: It was old.Vibhu [00:42:58]: But is it in the book as a good, here'sPhilip [00:43:01]: BaselineVibhu [00:43:01]: Baseline vanilla understand it?Philip [00:43:02]: Like you should know this.Vibhu [00:43:03]: Like I read the paper, I'm like, “ it makes so much sense.”Philip [00:43:05]: Yeah.Philip [00:43:05]: So with the book, I had a couple goals. One was to give people just a working vocabulary for the space as a whole, and the other was to give them some intuition about how each of these techniques works. As I mentioned in my AI Engineer talk, which is the first public addendum to this, the speculation space has moved much faster than everything else. So yeah, even at the time that I wrote the book Medusa, I very much included as a way for people to understand how the space evolved rather than what the most modern technique is. And now of course, there's DFlash, dSpark. There's, there's newer techniques even than EAGLE, although EAGLE is still very commonly used.Ali [00:43:51]: SpecSpecta.Philip [00:43:52]: Yes. Speculative decoding.Vibhu [00:43:54]: What canAli [00:43:56]: Oh, it's a paper by Tri Dao and it's like, it's doing speculative decodingVibhu [00:44:00]: HuhAli [00:44:01]: For the speculative decoder.Philip [00:44:02]: Oh, in spec- oh my God.Ali [00:44:02]: It's literally just an another. It's like, yeah, that's the most simple way to explain it, and it seems like he got trivial speed ups there. But it seems that the complexity with training, it's almost like in our mind at least, it's almost as complex as training GANs. Like it's like a very delicate balance and oftentimes you, it's just but yeah, it's literally speculative decoding on speculative decoding.Vibhu [00:44:21]: Speculative.Ali [00:44:22]: Yeah. We saw this paper.Vibhu [00:44:24]: It's interesting, right?Ali [00:44:24]: Yeah.Vibhu [00:44:24]: I wouldn't even expect it to be very particular to train, I wouldAli [00:44:29]: Right.Vibhu [00:44:29]: The naive part of me is like, okay, train speculative decoder.Ali [00:44:32]: But like, and it makes sense, like the whole idea of speculative decoding is you. It's like, it's like almost like the iPhone auto predict version but for a normal model, right? Like you're just, you're just, generating three tokens and you're like, okay, I'll do prefill on them. And so you save those three turns for your original model. Now your speculative decoder is doing three turns of auto regression, so why not just have an even smaller model?Ali [00:44:53]: The other question there is what are the size of speculators? So say forPhilip [00:44:58]: Right. It's like a billion parameters.Ali [00:45:01]: Like for MiniMax, it's. Yeah. It's like one layer. It's like one 60th of the original model usually.Philip [00:45:06]: Yeah. I think we should do a paper when we get back to the office.Philip [00:45:10]: SpeculativeAli [00:45:11]: SpeculativePhilip [00:45:11]: Decoding.Ali [00:45:13]: No, it's, it does seem like how, when do you stop? But then it also seems like if you're able to train spec-spec decode for instance, right? Like if you're able to have a small model that is accurately predicts what the intermediate speculator is gonna predict, that is able to predict what the original target model's gonna predict, then why not just use that smallest model directly, right?Vibhu [00:45:34]: Yeah. This isAli [00:45:35]: Like it seems likeVibhu [00:45:35]: Adjacent to the routing problem.Ali [00:45:36]: Right.Vibhu [00:45:36]: Yeah.Ali [00:45:36]: Right.Philip [00:45:37]: The thing with speculators is one of the practical constraints on using them is that you do have to run a small model on the same hardware that you're running the big model on. There is a orchestration and resource competition problem inherent in that, and that is one of the constraints on speculation in general, is that draft tokens cost resources to create and cost software complexity to manage. And so if you have like infinitely recursive speculators, you add in quite a bit of that complexity on the actual implementation within the inference engine as well, not just in the training process.Vibhu [00:46:17]: I was gonna say, I would wonder if you could do similar, like distillation and pruning of, it's the same thing, it's just a model. Can we not just distill a lot of the weights, quantize the speculator, out of my domain? The question that also comes up is, this is all for big server workloads, right? How much of this applies to, say I have this MacBook, I wanna run Gemma really efficiently. Similar problems, not the same?Local AI vs. Data Center InferencePhilip [00:46:45]: Pretty different. I talked to Selo, about this on his podcast a couple weeks ago. The difference between inference engineering for the data center and for production workloads versus inference engineering for local AI, is that we start with fundamentally like different constraints and different goals. With local AI, it's how do I fit this model onto my hardware and then make it less dumb? And with data center influence, it's how do I load this model and then make it less slow? And we care about less dumb, and they care about less slow. But the local AI inference engineering ecosystem, I think has a lot for us to learn from in the data center space. They are experts in various forms of quantization, including dynamic quantization that we just don't touch, in the pruning, in the distillation, in the, layer removal. There'Ali [00:47:42]: Layer removal matters less.Philip [00:47:43]: Yeah. There'Ali [00:47:44]: No one loves pruning really.Philip [00:47:45]: Yeah. Well, but the, but they doVibhu [00:47:46]: Which is surprising, right? But that's, that's a whole different thingPhilip [00:47:48]: Just to fit something on the laptop.Ali [00:47:50]: Right.Philip [00:47:50]: So yeah, it's a, it's an interesting, it's an interesting space. Not necessarily that like their techniques make sense for us to do in the data center, because we have different resources and different goals, but more that the process as well as the openness of that field is something to, admire.Ali [00:48:12]: Yeah. Like to your point, like, certain optimizations that would. Like for instance, Turbo Quantum Sharper, like it made such huge hype on that and we did like a whole deep dive on Twitter and like said, what is it? How does it work? Why is it good or not? And it took off and it was implemented on local devices because your memory bandwidth is so slow on like a MacBook, for instance. But try putting the same thing on like an NVIDIA GPU on a B200 Turbo quant would not be. Like, it would not be used. Like, NVIDIA - Like, NVIDIA made it clear that this is not a good optimization, and we've seen it firsthand where the overhead of doing dequantization, quantization of, in the kernel itself with turbo quant kernel, each end is much slower than the time that you save from doing the bandwidth. ‘Cause on the B200s, you have like 3.5 terabytes per second. You don't need decrease the storage that much. You don't need to do, FP4 KV cache. You don't need to use a requant. There's, there's, there's better optimizations to be made. But on Edge devices, it's extremely important, it's extremely useful. So, seems to be, like, different optimizations there, but then they're all uniquely combined with like all you wanna quantize the model, you wanna do speculative decoding, like certain common prefixes with bothPhilip [00:49:18]: Principles.Ali [00:49:19]: Yeah, exactly. Exactly. Exactly.Philip [00:49:20]: They also do a lot of work on, model parallelism, especially over, heterogeneous topology, where you have, some sparks and they are wired together with, Ethernet, DGX sparks.Ali [00:49:35]: Yeah, this is the Exo Labs guys.Philip [00:49:36]: Yeah. You have, a nu
In this message, Pastor Caleb discusses the types of sin there are. Knowing what types of sin there are can help with a better understanding of God's definition of sin. May Christians learn more about sin to better overcome it as stronger believers. Send us Fan MailSupport the showFor more information for our church visit AGCSparta.org.
In this message, Pastor Caleb discusses the importance of taking time to be refreshed and recharged spiritually, especially after ministering or after attacks from the spiritual enemy. May Christians know or remember how to be refreshed and recharged so they do not experience burnout for God. Send us Fan MailSupport the showFor more information for our church visit AGCSparta.org.
In this message, Pastor Caleb discusses the people who are enemies of the cross of Christ who pervert the Gospel of Jesus. The Apostle Paul describes these individuals as dogs. May Christians know how to handle these types of people in their lives and remain stronger believers in God. Send us Fan MailSupport the showFor more information for our church visit AGCSparta.org.
Pericles was the foremost statesman of Athens' classical age, dominating the political sphere, and giving his name to the building programme that produced the architecture we see on the Acropolis today. But what was he really like beyond the orator and politician? Paul Cartledge, AG Leventis Professor of Greek Culture, honorary citizen of modern day Sparta and author of a new biography of Pericles. We discuss Pericles' mid-life crisis, his genius for giving speeches and we start with Paul's experience of the BBC radio programme In Our Time on which Paul often appears as a guest. Paul Cartledge Links Pericles: Statesman, Demagogue, Eccentric BBCR4 In Our Time: The Delian League History Book Club Book List Oliver Webb-Carter Links Substack Who Cares Who Wins? Paean to Patrick Leigh Fermor X Instagram Email me: owcpods@gmail.com Learn more about your ad choices. Visit podcastchoices.com/adchoices
Christopher Nolan and Matt Damon's epic triumph The Odyssey resurrects the ancient world that turbocharged Western civilization — democracy, theater, philosophy, architecture, art.If that era had a soundtrack, it would sound uncannily like Robin Batteau's magical, mystical Banned in Sparta: eleven songs mosaic‑ed from the shards of pottery and scraps of parchment left behind by Greece's great lyric poets. These were the original singer‑songwriters — Archilochus, Sappho, Alcaeus — performing with lyres the way Dylan or Taylor Swift perform with guitars today.Batteau built these songs like a paleontologist assembling a T. Rex, filling gaps in the ancient fragments with “musical bones” of his own devising. And to bring them fully to life, he chose voices from our own Golden Age of singer‑songwriting: Eric Andersen (“tremulous, touchingly chamber finale”), Carolyn Hester (who hosted Dylan's first recorded appearance), Livingston Taylor, Kate Taylor, Tom Paxton, rocker Robin Lane, plus Tony‑winning actor James Naughton and his children Greg and Keira, Batteau himself, and newcomer Matt Nakoa.The result is a contemporary folk mosaic that feels both ancient and modern — a soundtrack to the world Homer walked through.Become a supporter of this podcast: https://www.spreaker.com/podcast/arroe-collins-like-it-s-live--4113802/support.
Christopher Nolan's new film ‘The Odyssey' is a box-office hit, but classics scholar Emily Wilson, who translated Homer's poem from the original text, says the film's script is abysmally written, lacking in emotional depth and character development. She talks about her translation and divisive criticism of the film. Also, we hear from actor Jon Bernthal who plays Menelaus, King of Sparta in ‘The Odyssey' and the Punisher in the new Spider-Man movie. See pcm.adswizz.com for information about our collection and use of personal data for sponsorship and to manage your podcast sponsorship preferences.NPR Privacy Policy
Send a Message to the TeamFor the 300th Episode of the Podcast, the team looks at a different outcome of the Battle of Thermopylae.Panel:Evan, Kai, and ChrisYou can follow and interact with A Fork In Time on….Discord: https://discord.com/invite/xhZEmZMKFSFacebook: https://www.facebook.com/aforkintimeTwitter: @AFITPodcastOur YouTube ChannelIf you enjoy the podcast and want to support it financially, you can help by:Supporting us monthly via Patreon: https://www.patreon.com/aforkintime....or, make a one-time donation via Podfan to A Fork In TimeE-Mail: aforkintimepodcast@gmail.comSupport the show
Christopher Nolan and Matt Damon's epic triumph The Odyssey resurrects the ancient world that turbocharged Western civilization — democracy, theater, philosophy, architecture, art.If that era had a soundtrack, it would sound uncannily like Robin Batteau's magical, mystical Banned in Sparta: eleven songs mosaic‑ed from the shards of pottery and scraps of parchment left behind by Greece's great lyric poets. These were the original singer‑songwriters — Archilochus, Sappho, Alcaeus — performing with lyres the way Dylan or Taylor Swift perform with guitars today.Batteau built these songs like a paleontologist assembling a T. Rex, filling gaps in the ancient fragments with “musical bones” of his own devising. And to bring them fully to life, he chose voices from our own Golden Age of singer‑songwriting: Eric Andersen (“tremulous, touchingly chamber finale”), Carolyn Hester (who hosted Dylan's first recorded appearance), Livingston Taylor, Kate Taylor, Tom Paxton, rocker Robin Lane, plus Tony‑winning actor James Naughton and his children Greg and Keira, Batteau himself, and newcomer Matt Nakoa.The result is a contemporary folk mosaic that feels both ancient and modern — a soundtrack to the world Homer walked through.Become a supporter of this podcast: https://www.spreaker.com/podcast/arroe-collins-unplugged-totally-uncut--994165/support.
On this episode of The Hillsdale College Online Courses Podcast, Jeremiah and Juan discuss the ways that Homer depicts women in The Odyssey before introducing a lecture from Dr. Benedict Whalen. Penelope, the wife of Odysseus, uses her own cunning to deceive the suitors while she awaits her husband’s return. Telemachus, their son, journeys to Sparta seeking news of his father. The Odyssey tells the tale of the cunning and crafty hero of the Trojan War on his journey home. After ten years of war, Odysseus faces another ten years of fantastic temptations, deadly monsters, treacherous seas, and impious men. Meanwhile, Ithaca is threatened by suitors who plague his wife and plunder his wealth. When Odysseus reaches home, he must slay the suitors, test the loyalty of his friends and family, and restore order to his marriage and his kingdom.See omnystudio.com/listener for privacy information.
On this episode of The Hillsdale College Online Courses Podcast, Jeremiah and Juan discuss the ways that Homer depicts women in The Odyssey before introducing a lecture from Dr. Benedict Whalen. Penelope, the wife of Odysseus, uses her own cunning to deceive the suitors while she awaits her husband’s return. Telemachus, their son, journeys to Sparta seeking news of his father. The Odyssey tells the tale of the cunning and crafty hero of the Trojan War on his journey home. After ten years of war, Odysseus faces another ten years of fantastic temptations, deadly monsters, treacherous seas, and impious men. Meanwhile, Ithaca is threatened by suitors who plague his wife and plunder his wealth. When Odysseus reaches home, he must slay the suitors, test the loyalty of his friends and family, and restore order to his marriage and his kingdom.See omnystudio.com/listener for privacy information.
The crew spends significant time contrasting Michigan State's modest NIL/jersey patch revenue against Big Ten rivals like Ohio State and Illinois, framing it as a symptom of the program "playing in waters they shouldn't be playing in". They also debate two known Spartan candidates, Tom Dieters and Ashton Henderson, as preferred picks for the vacant Athletic Director role over outside hires with flashy but "non-functional" resumes. Lighter segments include Hall of Fame reactions, a tribute to late linebacker Gail Clark, and playful banter about basketball eligibility rumors involving Jaxon Kohler and Tre Holloman.Become a supporter of this podcast: https://www.spreaker.com/podcast/this-is-sparta-msu--5664600/support.
In this heartwarming episode, Taylor and Dylan Roberts, owners of Roberts Five Farm in Sparta, join the show to share the remarkable story behind one of the Upper Cumberland's most beloved seasonal destinations. What began as an unexpected idea sparked during a work trip overseas has grown into a thriving flower farm built on faith, family, and perseverance. The couple discusses the challenges of growing more than 100,000 tulips each year, balancing farm life with raising three young children, and how they hope every visitor leaves the farm with a sense of peace. They also preview exciting plans to expand with sunflowers and wildflowers while reflecting on the role faith continues to play in every season of their journey. Listen To The Local Matters Podcast Today! The UC Now · News Talk 94.1
Illinois Farm Bureau President Philip Nelson covers a number of topics and issues. Previewing the Grand American championships taking place at the World Shooting and Recreational Complex in Sparta with Randy Moeller, Executive Director of the Amateur Trapshooting Association.
Most of us have a bottle of olive oil in the pantry and almost no idea whether it's actually good. Today I'm sitting down with Dino Pierrakos of Laconiko, a fourth-generation olive oil producer from Sparta, Greece, to talk about what separates a real extra virgin olive oil from the stuff sitting on most grocery store shelves.We get into why harvest timing matters more than the expiration date on the label, what "monocultivar" actually means and why it's worth looking for, and the difference between polyphenols and oleocanthal, the compound getting real attention from researchers for its anti-inflammatory potential. Dino also walks through one of the most persistent myths in the wellness space: that cooking with olive oil destroys its benefits. The research tells a more nuanced story, and we break down what actually happens to olive oil at high heat.If you've ever stood in the olive oil aisle wondering what you're supposed to be looking for, this episode will change how you shop for good.Right now you can grab 10% off at Laconiko with the code "wellnstrong10"Suggested Resources:Laconiko olive oilIbuprofen-Like Activity in Extra-Virgin Olive OilOlecanthal in Inflammation & CancerRich oleocanthal and oleacein extra virgin olive oil and inflammatory and antioxidant status in people with obesity and prediabetesSend me a text!This episode is proudly sponsored by: SizzlefishLet's talk about fueling your body with the best nature has to offer. If you're looking for premium, sustainable seafood delivered straight to your door, you need to check out Sizzlefish! Head to sizzlefish.com and use my code “wellnstrong” at checkout for an exclusive discount on your first order. Trust me, you're going to taste the difference with Sizzlefish!Join the WellnStrong mailing list for exclusive content here!Want more of The How To Be WellnStrong Podcast? Subscribe to the YouTube channel.Follow Jacqueline:Instagram PinterestTikTokYoutubeTo access notes from the show & full transcripts, head over to WellnStrong's Podcast Page
CHANCE LIGA začala pořádným výbuchem!
This is a teaser of the bonus episode, "A League by Choice?" found over on Patreon.A voluntary alliance can feel like freedom right up until the day you try to leave. After the Greek victory over Persia in 479 BC, more than a hundred islands and coastal city-states align themselves with Athens in what becomes the Delian League. We walk through why that choice makes sense at the time and why it isn't as simple as “Athens forced an empire on unwilling allies.” We start with the road not taken: Sparta. Even with its wartime prestige, Sparta's priorities stay anchored in the Peloponnese, its naval capacity is limited, and the helot system makes long overseas commitments dangerous. Add the reputational blow of Pausanias in Byzantium, and the eastern Aegean has every reason to look elsewhere for steady leadership. Athens, newly empowered by a fleet funded through the Laurean silver mines and proven in the Persian Wars, can actually patrol the Aegean Sea, protect trade routes, deter Persian interference, and speak to a shared Ionian identity many communities recognize. Then we track how the league's mechanics reshape its politics. Contributions shift from ships and crews to tribute, and what looks like a convenient budget choice slowly centralizes force, money, and decision-making in Athenian hands. The revolts of Naxos and Thasos show the turning point: joining may be voluntary, but withdrawal becomes a trigger for siege, disarmament, and lost autonomy. Even so, many allies stay because the benefits remain real and rebellion is costly and uncertain. If you care about ancient Greece history, the Delian League, and how security pacts evolve into empires, you'll find plenty to argue with here. Subscribe, share the episode with a friend, and leave a review with your answer: when does an alliance stop being a choice?Support the show
It's Friday, July 24th, A.D. 2026. This is The Worldview in 5 Minutes heard on 140 radio stations and at www.TheWorldview.com. I'm Adam McManus. (Adam@TheWorldview.com) By Adam McManus British pro-lifer brutally murdered There's a brutal update about the murder of a retired British pro-life politician. Shockingly, the man accused of murdering former British politician Ann Widdecombe on July 8th struck her 21 times in the head with a hammer before taking her wallet and leaving her home, reports Human Events. After she failed to appear for a scheduled television interview that afternoon, her body was discovered the following day by her gardener. Joshua Kerry, age 28, appeared at Westminster Magistrates' Court on July 21st after being charged with murder in the death of the 78-year-old former Conservative Member of Parliament. Prosecutors said police are also investigating whether the killing had a political or terrorist motive. According to the Associated Press, Widdecombe was eating lunch at her home on the edge of Dartmoor National Park in southwest England, when Kerry arrived. Widdecombe served as a Conservative member of Parliament from 1987 until 2010. Listen to her articulate defense of the unborn baby in a 2020 debate hosted by the Queen's University Belfast Literary and Scientific Society. WIDDECOMBE: “When a child is born, you cannot take its life. Not one minute after birth can you take its life. It does not matter how disabled that child is, how profoundly sick it may be, how stressed the mother is, how poor the mother is, whether the mother has suddenly decided that she didn't, after all, want to have the child. You cannot take that child's life. At the moment of birth, it has full civil rights. It is protected as much as you and I are protected by the law. “Yet, one day earlier, that is not so. But hang on. It's the same child! Nothing has happened to change that child into something different, other than the fact that finally it can be seen, that it is quite visibly here among us. But it's been here among us in the days and the weeks and, indeed, the months before that birth. “So, I would say that the law, which has now been imposed, I can't put it any other way, on Northern Ireland is a law that actually is illogical. It is immoral and it is unkind. If a child can enjoy our protection one day, that same child should be able to enjoy our protection the day before!” Proverbs 31:8 says, “"Speak up for those who cannot speak for themselves.” The brutal death of pro-lifer Ann Widdecombe is a profound loss. Samaritan's Purse aids thousands after Venezuela earthquakes Evangelist Franklin Graham visited Venezuela this week to meet with earthquake survivors who were being aided by Samaritan's Purse after the June 24 back-to-back magnitude 7.5 and 7.2 earthquakes killed at least 5,000 people and injured even more, reports The Christian Post. Samaritan's Purse Emergency Field Hospital, which has 60 beds and two operating rooms, has helped 2,500 people since it opened on June 30th. Graham said, "These people are surrounded by so much death and destruction. We want to bring the light of Jesus Christ into these places of pain.” Thomas Ovington, the Team Leader in Venezuela, gave his overview of the earthquake trauma and the help that Samaritan's Purse is offering. OVINGTON: “We're seeing complete chaos around. There's fallen buildings. There's people sleeping in the streets. Football stadiums packed full of people. It's also the huge number of people that are missing. There's just an incredibly high level of desperation. “But it's also an amazing opportunity for us to serve. We have a lot of different capabilities. This is the emergency room that we're in right now. Two operating theaters in the back where we can do the surgeries. A lot of the needs we're starting to see are trauma from fallen buildings. There's crush injuries, broken bones. “Additionally to that, we're doing distributions of tarps, non-food items, and potable water. “At Samaritan's Purse, our number one priority is to share the Gospel of Jesus and the hope we have in Him. And so here, we're in the midst of an incredibly hopeless situation for tens of thousands of people. It's really an incredible opportunity for us to live out our mission, relieving suffering and doing it all in Jesus' name.” You can make a donation to help the efforts of Samaritan's Purse on the ground in Venezuela through a special link in our transcript today at www.TheWorldview.com. Trump blasts Democrat Communists as greatest threat to America Ahead of the November midterm elections, President Donald Trump launched a broadside against the Democratic Party, saying the up-and-coming Democratic Socialists are "communism by another name" and "the greatest threat I've ever seen to our country.,” reports Real Clear Politics. TRUMP: “While our administration is lifting up the next generation with the power of freedom and American capitalism, proven all over the world when it's done right, the radical Left is actively trying to destroy our children's future with the disaster known as communism. This isn't social democracy; these are communists by another name, right? “Communism is the single greatest threat to our country in its history, including even World War I and World War II, Pearl Harbor, 9/11. These people are stone cold crazy. “You know they want to give everything away. You want free rent? You want [a] free house? What they don't say is that a year from now, there'll be absolute squalor, turmoil, and death. And that's why we declare again today that the United States of America will never be a communist country!” NYC Mayor threatens to arrest Israeli Prime Minister as a “war criminal” New York City Mayor Zohran Mamdani, a socialist, has been threatening to arrest Israeli Prime Minister Benjamin Netanyahu when he addresses the United Nations this September in New York City. This was in response to the International Criminal Court's November 2024 issuance of a warrant for Netanyahu's arrest for what it alleges are war crimes and crimes against humanity. MAMDAMI: “Benjamin Netanyahu is a war criminal, the architect of a horrific genocide against the Palestinian people. As I've said, I agree with the [International Criminal Court] that Benjamin Netanyahu should be arrested and tried for his crimes. “My administration has reviewed every avenue available under applicable law to determine whether New York City could execute the International Criminal Court's arrest warrant if Benjamin Netanyahu came here. It is clear that we do not have the independent legal authority to enforce this warrant. The federal government, however, does, and I call on them to join the ICC and execute this warrant.” Appearing on American Family Radio, Mike Huckabee, the U.S. Ambassador to Israel, had some choice words for the New York City mayor. HUCKABEE: “He certainly doesn't have the legal right to arrest a diplomat because he has diplomatic immunity if he comes to the U.N. So, this is nonsense on the part of Mamdami. He ought to be focusing on fixing potholes, picking up the trash, keeping the parks clean, and the needles from drug users off the streets. But instead, he's injecting himself into something about which he is utterly ignorant. “When he says Benjamin Netanyahu is a war criminal, on what does he base that? Benjamin Netanyahu is the elected leader of a country that very much mirrors the United States in terms of complete, wall-to-wall freedom. And if he's talking about Gaza and says, ‘Oh, they're guilty of genocide. Israel didn't start the war in Gaza; Hamas did by murdering 1,200 innocent civilians and taking captive 250 people, including 16 Americans.” Ambassador Huckabee explained what a similar attack on America would look like when scaled to the size of our population. HUCKABEE: “If Americans had the similar kind of assault on us in one day, if you looked at the scale of population, it would be the equivalent of 40,000 Americans being murdered in a single day, and 10,000 people being taken hostage, raped, starved, beaten, and abused. And I wonder what would Americans demand? What would Americans do if we had 40,000 people slaughtered, including pregnant women butchered in front of their children? If we'd had that in one day, I wonder if America's response would be to just gently write, maybe a strongly worded letter and say, ‘Please give our 10,000 hostages back.' “The Israelis prosecuted the war against Hamas for the sole purpose of getting their hostages back, but Hamas wouldn't let them go. This war should never have taken that long.” Homosexual Fox News host buys surrogate child On July 15, Fox News political analyst Guy Benson, a self-proclaimed homosexual, announced on Instagram that he and his homosexual partner, Adam Wise, became so-called “fathers” again with a picture of himself holding a little girl named Madison Halsey, reports LifeSiteNews.com. This is the second child the two men have procured through surrogacy. Their son, Conrad James, was also birthed by an anonymous surrogate mother in November 2023. The culture war over surrogacy – the practice of renting the body of a woman to gestate a child created outside the womb through in vitro fertilization – has been growing louder over the past several years. The New York Post noted that Benson's announcement “drew an outpouring of well-wishes from prominent conservatives, including former U.N. Ambassador Nikki Haley, Fox News host Dana Perino, and political commentator Sarah Cupp.” On the other hand, Christian commentators such as Megan Basham emphasized the injustice. She wrote, “Benson is an incisive pundit. But there is nothing conservative about deliberately creating motherless children. I pray Benson, and those like him on the Right, would consider what they are telling the world about how unnecessary they believe mothers (or if we're talking about two women, fathers) are to the needs of children.” And Katy Faust, a children's rights activist and founder of Them Before Us, wrote, “What did you think was going to happen after we legalized gay marriage? … What Obergefell enabled was something else. It accelerated the legal replacement of biological parenthood with intent-based parenthood.” Romans 1:26-27 says, “God gave them over to shameful lusts. Even their women exchanged natural sexual relations for unnatural ones. In the same way, the men also abandoned natural relations with women and were inflamed with lust for one another. Men committed shameful acts with other men, and received in themselves the due penalty for their error.” 24 Worldview listeners gave $10,569.75 And finally, by Thursday night at 10:00pm Central, 24 Worldview listeners stepped up to the plate and invested their treasure to fund the six-member team behind The Worldview for another year. Our thanks to Elizabeth in Lubbock, Texas who gave $50, Craig in Rochester, New Hampshire who gave $86.75 as well as Robert in San Antonio, Texas and Jane in Shrewsbury, Pennsylvania – both of whom gave $100. We are grateful to God for Elaine in Hampton, South Carolina who gave $120, Jennifer in San Jose, California, Alexander in Greensburg, Pennsylvania, and Jason in Lakeland, Florida – each of whom pledged $10/month for 12 months for a gift of $120. We were touched by the generosity of Daniel in Santa Cruz, California who pledged $22/month for 12 months for a gift of $264, Sarah in Madera, California who gave $300, and George in Leesburg, Virginia who pledged $25/month for 12 months for a gift of $300. We appreciated the sacrifice of Fred in Cascade, Montana, Michelle in Sparta, Michigan, Todd in Davenport, Iowa, Mary in Highland, New York, Caroline in Live Oak, Florida, and Clarence in Stone Ridge, New York – each of whom gave $347 as well as Michele in Kindersley, Saskatchewan, Canada who pledged $28.92/month for 12 months for a gift of $347. And we were blown away by David in Tall Timbers, Maryland who pledged $30/month for 12 months for a gift of $360, Rosilin in Waynesville, Missouri who gave $500, Tammy in Cascade, Idaho who gave $1,200, Elizabeth in Carol Stream, Illinois who gave $1,200, Emmanuel in Madison, Alabama who gave $1,200, and Calvin in Waldorf, Maryland who gave $2,000. Those 24 gifts add up to $10,569.75. Ready for our new grand total? Drum roll please. (drum roll sound effect) $56,062.75 (sound effect of people cheering) Wow! That is the single best day so far in the month of July. More and more Worldview listeners are catching the vision, thanks in part to the McVeda Family in Great Falls, Montana. Can you feel the momentum? We need to raise $25,163.25 by midnight on Friday, July 24 But we still have a steep hill to climb. So, if you have not stepped up to the plate yet, we could really use your help. By tonight at 12 midnight, we need to raise a whopping $25,163.25 to hit our $81,226 goal! But listen to this. I got a call yesterday morning from a Worldview friend in Naples, Florida right before I swam my 72 laps at the outdoor pool at my gym. He said, “Adam, I will match, dollar for dollar, the next 12 people who give a one-time gift of $500.” He is willing to invest $6,000 to encourage 12 people to collectively give another $6,000! If those gifts come in today, then we have to raise the difference of about $13,000. Would you consider being one of 6 people to pledge $100/month for 12 months for a gift of $1,200, one of 5 people to pledge $50/month for 12 months for a gift of $600, or one of 10 people to pledge $25/month for 12 months for a gift of $300? Just go to TheWorldview.com, click on Give, select the dollar amount, and make sure to click on the “recurring” button if that's your wish. And remember, if you want to continue your monthly pledge to The Worldview that you started in a previous year, please let me know so we can count your generous ongoing gift toward our total. What does God want you to give? Go to TheWorldview.com and click on Give. Close And that's The Worldview on this Friday, July 24th, in the year of our Lord 2026. Subscribe for free by Spotify, Amazon Music, or by iTunes or email to our unique Christian newscast at www.TheWorldview.com. Plus, you can get the Generations app through Google Play or The App Store. I'm Adam McManus (Adam@TheWorldview.com). Seize the day for Jesus Christ.
A king is murdered, a throne becomes a prize, and the world's biggest empire has to prove it can still hold together. We pull away from the Peloponnesian War just long enough to track the Persian succession shocks that set the stage for Persia's return to Greek politics, starting with the assassination of Xerxes and the hard-won rise of Artaxerxes I. Along the way, we sort Babylonian cuneiform evidence from Greek accounts packed with intrigue, and we ask what those contradictions tell us about propaganda, perspective, and how ancient history gets written.From there, we follow the early tests of Artaxerxes' rule, especially the Egyptian revolt led by Inaros and the disastrous Athenian intervention in the Nile Delta. It's a sharp reminder that the Achaemenid Empire is not just a backdrop to Greek history: it has its own internal pressures, powerful nobles, and regional flashpoints that can either drain imperial attention or force rapid adaptation.The story then turns to the chaotic handover after Artaxerxes' death and the violent climb of Darius II. Once stability returns, Persia pivots west with a smarter approach than invasion: leverage. We dig into how Persian objectives in Asia Minor shape support for Sparta, why Tissaphernes and Pharnabazus pursue different strategies, and how later decisions, including the rise of Cyrus the Younger, help set up the flow of Persian silver that changes the naval war against Athens.If you enjoyed this deep dive into Persian successions, satrap politics, and the Persian Empire's role in the Peloponnesian War, subscribe, share the episode, and leave a review so more listeners can find the series. Support the show
Schae is a Vet Tech for My Pets Vet in Florence and lives is Sparta, KY where she owns a farm. Schae moved here from California in 2017 and absolutely loves the area! She's also a huge B-105 listener and really enjoys the Confession Calls and Beat the Bear and adds she knows way too much about Jesse Tack! LOL For her induction song she wanted to hear "Suds in the Bucket" by Sara Evans. Welcome to the B-105 Country Club, Schae!See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
In this episode with The Inspire Health Podcast, Dr. Jen White takes it down a notch to talk about grief, death, dying, and how to put a shattered heart back together. Prompted by an unexpected wave of recent losses, including an injured duck, a dear client, and ultimately, her beloved 15-year-old "priestess" cat, Sparta, Dr. Jen shares what it's like to be swallowed whole by the "grief portal". Balancing deep sorrow with mystical insights, a little bit of humor (including her mom's infamous spicy chicken sandwich), and profound spiritual lessons, she guides listeners through the tools that helped her transition from clinical depression to a place of joyful remembrance. Whether you are mourning the loss of a pet, a friend, or an immediate family member, this episode is a comforting companion for your healing journey. Themes: Give yourself permission to experience the heavy emotions. Don't try to outrun it or busy yourself; let the grieving process be "ugly". When you are ready, actively choose a ceremonial or therapeutic outlet to release the grief. This could be plant medicine, therapy, EMDR, boxing, or simply wailing on the earth in the woods. Shift the energy by celebrating your loved one. Build a beautiful altar, share funny stories with friends, or frame their photos to invite their joyful memory back into your daily life. Be open and patient. Loved ones and pets often send spectacular confirmations that they are still with us, whether it's an orange butterfly, a rainbow reflecting off a license plate, or a perfectly timed song. Dr. Jen highly recommends watching Anthony Chene Productions on YouTube. Hearing consistent accounts of the "other side" can provide profound comfort, shift your perspective on mortality, and ultimately change the way you live your life today. Connect with Jen:
They were told Sparta could never be defeated. Then 300 warriors, 150 pairs of same-sex lovers changed history forever.Forged in the shadow of tyranny, the Sacred Band of Thebes wasn't just an elite military unit, it was the first of its kind, made up of 150 pairs of devoted male lovers who fought side by side against the most feared army in the ancient world. In this episode, Jordi and Brad explore one of the greatest stories in queer history, from secret rebellions and impossible battlefield victories to the heartbreaking final stand that transformed these warriors into legends. It's a powerful story of courage, sacrifice, and love that proves history's most extraordinary heroes are often the ones left out of the history books. If you love LGBTQ+ true crime podcast storytelling, queer history, and true crime with a queer perspective that uncovers forgotten LGBTQ+ stories, this is an episode you won't want to miss.Hosted by Jordi and Brad, Beers With Queers brings chilling crimes, remarkable queer stories, and forgotten LGBTQ+ history back into the spotlight, all with a cold one in hand. Grab a drink, press play, and discover why the legend of the Sacred Band of Thebes still resonates more than 2,300 years later. Hosted on Acast. See acast.com/privacy for more information.
We are joined by Felicia Day (!!!!!!!!) to talk about her newest graphic novel, Lost Daughter of Sparta, as well as Greek mythology, mythological retellings, and how these stories help us look at how modern society treats women.Content Warning: This episode contains conversations about or mentions of misogyny, ageism, grief, infidelity, human sacrifice, death, sexual assault, kidnapping, queerphobia, GuestFelicia Day is an actor, producer, writer, and streamer. She is the author of the graphic novel, “Lost Daughter of Sparta”. You can support her upcoming Kickstarter for The Guild Reunion movie here. Housekeeping- Books: Check out our previous book recommendations, guests' books, and more at spiritspodcast.com/books- Call to Action: Send in those urban legend emails!- Submit Your Urban Legends Audio: Call us! 617-420-2344Find Us Online- Website & Transcripts: spiritspodcast.com- Patreon: patreon.com/spiritspodcast- Merch: spiritspodcast.com/merch- Instagram: instagram.com/spiritspodcast- Bluesky: bsky.app/profile/spiritspodcast.com- Twitter: twitter.com/spiritspodcast- Tumblr: spiritspodcast.tumblr.comCast & Crew- Co-Hosts: Julia Schifini and Amanda McLoughlin- Editor: Bren Frederick- Music: Brandon Grugle, based on "Danger Storm" by Kevin MacLeod- Artwork: Allyson Wakeman- Multitude: multitude.productionsAbout UsSpirits is a boozy podcast about mythology, legends, and folklore. Every episode, co-hosts Julia and Amanda mix a drink and discuss a new story or character from a wide range of places, eras, and cultures. Learn brand-new stories and enjoy retellings of your favorite myths, served over ice every week, on Spirits.See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
The Plot Begins: Rage and Divine Bargains. Guest: Professor Emily Wilson. The plot of the Iliad is ignited by a clash of egos between Agamemnon and Achilles. When Agamemnon is forced to return his own war prize to appease Apollo, he seizes Achilles' enslaved woman, Briseis, to recoup his lost face. This action causes Achilles to withdraw from the fighting, perversely restoring his honor by demonstrating how much the Greeks suffer without him. This human conflict is mirrored by divine bargaining; for instance, Hera is so intent on destroying Troy that she offers to let Zeus destroy three of her own beloved cities, including Sparta, in exchange for his cooperation. The Greek audience would have recognized the historical weight of these fallen cities. Wilson interprets Agamemnon not as a simple villain, but as a weak and struggling leader who often blames his poor decisions on divine delusion rather than taking personal responsibility. Despite his flaws, the poem illustrates the immense difficulty of maintaining power and making decisions under the influence of manipulative gods. 5