The universe as a complex and orderly system or entity
POPULARITY
Categories
Tom Cochrane is founder of Eleviam, a brand accelerator running more than forty CPG brands across Amazon, Walmart, and TikTok Shop. Before building the agency, he put nearly five million dollars of his own capital into Amazon inventory to prove the model himself.This episode tackles a problem every CPG brand now faces: platforms are getting more expensive, more competitive, and harder to attribute, even as consumer attention keeps fragmenting.Tom and Eitan dig into the halo effect between TikTok Shop and Amazon, why a third of TikTok-driven purchases convert somewhere else entirely, and how Amazon's new AI shopping layer is changing what actually belongs in a listing.Listeners walk away with a clearer view of how to sequence platform investment as a brand matures, and why profitability discipline matters more than chasing growth at any cost.Website: https://expanio.com/Podcast website: https://expanio.com/commerce-untold-podcast/Eitan Koter's LinkedIn: https://www.linkedin.com/in/eitankoter/YouTube: https://www.youtube.com/@CommerceUntoldGuest: Tom Cochrane, Founder, EleviamTom Cochrane's LinkedIn: https://www.linkedin.com/in/thomasgcochrane/Eleviam: https://eleviam.io/Key Takeaways: • Ninety percent of US consumers still validate a purchase on Amazon even when they discovered the product somewhere else. • Thirty-seven percent of TikTok Shop transactions convert off of TikTok, most often on Amazon. • Creator partnerships should be judged on proven GMV track record, not follower count. • Amazon's AI shopping layer now reads full PDP and A+ content, so listings need to answer real buyer objections instead of just stuffing keywords. • Brands should launch omnichannel from day one, but stay narrow with a hero SKU before expanding the catalog. • Agencies scale by building repeatable systems and pod structures before they scale revenue.Chapters:[00:19] Introduction to Tom Cochrane and background[02:35] Running Amazon and TikTok Shop together: the halo effect[05:36] Rising platform costs and the shift toward creator partnerships[08:08] Competition, agentic shopping, and Amazon's AI-driven listings (Cosmos, Rufus)[11:53] Matching platform strategy to a brand's lifecycle stage[19:41] CPG challenges today: LTV, margins, and building a moat[26:11] Inside Eleviam: ICP, delivery, and where e-commerce is headed
Come journey with me to the Cosmos and engage the heavily realms of Yahweh
Ryan and Alex explore the Western esoteric tradition and the ancient ideas that have shaped humanity's search for hidden knowledge. They dive into alchemical elements, Atlantean wisdom, Hermeticism, and the Tibetan Book of the Dead, examining what these traditions reveal about death, reincarnation, and the nature of consciousness.
Welcome back to Aquarium Drunkard Transmissions with Jason P. Woodbury. This week, we've got something extra special: a behind the scenes look at the forthcoming documentary, Sun Ra: Door of the Cosmos, a sonic cinema odyssey through the life, work and philosophy of the surrealist composer and musician. In addition to a conversation with director Drew DeNicola, who you may know from the 2012 doc Big Star: Nothing Can Hurt Me, this episode will present archival material gathered for the film, including interviews and statements from Ra himself, rare recordings of the Sun Ra Arkestra in action, and previous Transmissions guest Camae Ayewa, also known as Moor Mother, reading from Sun Ra's book of poetry, The Immeasurable Equation. The documentary has launched a Kickstarter campaign to raise finishing funds that will run for the month of September, offering unique array of premiums at all donation levels, including a selection of vinyl recordings from the Modern Harmonic label, and limited edition t-shirt and large format art prints from legendary jazz photographer, Veryl Oakland. Now, settle in as we explore the Sun Ra's esoteric ideas, hear his singular voice, and explore the Arkestra's vast lore and history. Crack open this door of the cosmos and enter in. Learn more about your ad choices. Visit megaphone.fm/adchoices
Come journey with me to the Cosmos and engage God
At nearly 100 years old, The Globus Family of Brands remains a global leader in escorted tours, independent travel packages, and river cruises. Tune in to learn about Globus, Cosmos, and Avalon Waterways. Plus, we'll discuss the hottest travel trends both on land and the rivers.
Dr Daniel Cunnama, an astronomer and founder of ZuluAlpha Remote Observatories, tells John Maytham what NASA's newly launched Nancy Grace Roman Space Telescope will be studying. Presenter John Maytham is an actor and author-turned-talk radio veteran and seasoned journalist. His show serves a round-up of local and international news coupled with the latest in business, sport, traffic and weather. The host’s eclectic interests mean the program often surprises the audience with intriguing book reviews and inspiring interviews profiling artists. A daily highlight is Rapid Fire, just after 5:30pm. CapeTalk fans call in, to stump the presenter with their general knowledge questions. Another firm favourite is the humorous Thursday crossing with award-winning journalist Rebecca Davis, called “Plan B”. Thank you for listening to a podcast from Afternoon Drive with John Maytham Listen live on Primedia+ weekdays from 15:00 and 18:00 (SA Time) to Afternoon Drive with John Maytham broadcast on CapeTalk https://buff.ly/NnFM3Nk For more from the show go to https://buff.ly/BSFy4Cn or find all the catch-up podcasts here https://buff.ly/n8nWt4x Subscribe to the CapeTalk Daily and Weekly Newsletters https://buff.ly/sbvVZD5 Follow us on social media: CapeTalk on Facebook: https://www.facebook.com/CapeTalk CapeTalk on TikTok: https://www.tiktok.com/@capetalk CapeTalk on Instagram: https://www.instagram.com/ CapeTalk on X: https://x.com/CapeTalk CapeTalk on YouTube: https://www.youtube.com/@CapeTalk567 See omnystudio.com/listener for privacy information.
The Summer of Dan Savings Event comes to an end with 1992's Godzilla vs. Mothra. A stray meteoroid reawakens the King of the Monsters and sets off a series of natural disasters across the Pacific Ocean. The Japanese government forces ex-archeologist Takuya Fujito to investigate one such disaster: a landslide in Indonesia that unearthed something strange. On the remote Infant Island, Takuya's expedition discovers a gigantic egg and tiny twin fairies called the Cosmos. Mothra, the Guardian of Earth we all know and love, is set to hatch from the egg in order to save humanity from their own hubris. However, the Cosmos warn that Mothra has a dark twin named Battra, whose patience for mankind has long since run out. The two Divine Moths are drawn to Japan by Godzilla's latest rampage, igniting a three-way battle to restore the balance of nature. The real estate market trembles, egg prices skyrocket, and a familiar duet starts playing on today's long-prophesied episode of Anime Was (Not) A Mistake!. Rate, Review, Subscribe, and Listen to Us on Podbean/iTunes/Stitcher/Spotify Follow us on Instagram:@animewasnotamistakepodcast Or on Facebook:@animewasnotamistakepod Music Provided: “Attack Godzilla!” (M16) - Godzilla (1954) Soundtrack OST Composed by Akira Ifukube “The Defense Corps Mobilize II” (M20) - Godzilla vs. Gigan Soundtrack OST Composed by Akira Ifukube & Kunio Miyauchi “Mechagodzilla Appears/Godzilla vs. Mechagodzilla” (M11 + M11EXT) Composed by Masaru Sato.
Scholar M. David Litwa joins the podcast to explore the intersections of Gnosticism, Hermeticism, and ancient astrology. David is a historian of early Christianity and Greco-Roman religions who recently published a new translation of the Corpus Hermeticum. Together, we trace how these esoteric traditions developed and interacted in the ancient world, particularly in the cultural melting pot of Alexandria. We discuss the emergence of Hermeticism as a synthesis of Egyptian religion and Greek philosophy, and how early Christian Gnosticism subsequently adapted many of these cosmological frameworks. A major focal point is how ancient astrology, including the ascent and descent of the soul through the planetary spheres, was integrated into these religious movements. We contrast the Gnostic view of the planets as restrictive archons keeping the soul imprisoned with the more sympathetic Hermetic view of the visible cosmos as a second God and part of divine providence. We also take a close look at the Peratics, a unique Gnostic group that integrated technical Hellenistic astrology directly into their religion in order to map the soul's escape from the material realm. Finally, we unpack the historical tensions between fate, providence, and free will, examining how early Christians navigated astrological determinism before figures like Augustine firmly pushed back against the practice. This conversation highlights the profound impact that Hellenistic astrology had on the broader philosophical and religious landscape of the ancient world. This is episode 546 of The Astrology Podcast. David's Website and YouTube Channel https://mdavidlitwa.com https://www.youtube.com/@m.davidlitwa Timestamps 00:00:00 Introduction & David's Background00:04:46 Alexandrian vs. Roman Christianity00:07:27 Astrology in the New Testament00:09:55 The Nag Hammadi Library00:13:35 Debating the Term "Gnosticism"00:20:57 Alexandria: Birthplace of Hermeticism00:22:52 Dating the Corpus Hermeticum00:27:29 Early Astrological Hermetica00:35:18 The Cosmos as the Second God00:46:15 Ascent of the Soul00:54:08 Planets as Malefic Archons00:59:42 The Gnostic Demiurge01:06:58 The Paratai & Astrological Salvation01:16:36 Astrology, Fate, & Free Will01:21:17 Baptism & Freedom from Fate01:30:21 Augustine's Anti-Astrology Bias01:38:09 "De-astrologizing" Hermetic Texts01:48:46 The Need for Astrological Literacy01:50:46 Conclusion & Resources Watch the Video Version of This Episode https://www.youtube.com/watch?v=7UZl_fMgGRs - Listen to the Audio Version of This Episode Listen to the audio version of this episode or download it as an MP3:
The cosmos contains everything we know, from planets and stars to galaxies stretching across unimaginable distances. This episode explores the structure of the universe, how our understanding of it has changed over time, and the enormous scales that shape our view of space. Along the way, you'll hear about galaxies, cosmic distances, the observable universe, and humanity's attempts to understand our place within it. It's steady and consistent, with no whispering and no sudden changes, just enough to give your mind something to follow as you wind down. Happy sleeping! Read with permission from Cosmos, Wikipedia (https://en.wikipedia.org/wiki/Cosmos), licensed under CC BY-SA 4.0. — Ad-free episodes: icantsleep.supportingcast.fmHave a topic in mind? Request a topic Learn more about your ad choices. Visit megaphone.fm/adchoices Learn more about your ad choices. Visit megaphone.fm/adchoices
An imaginative girl narrates in her captain's log as she, her father, and her dog pass through a car wash…and emerge on the other side of the galaxy. Genre: Science Fiction Excerpt:"Let me see where there's a good place to land in that jungle," I said, easing off the brake again. "Not a good idea." I smiled a little. "Oh no?" Meena explained. "Don't you know what they say about the jungles of Zarqlok? Pretty from above, deadly from below." "What does that mean?" "Predators. Venomous predators. Lots of them." The Wheel of Fiction Turns. What did it land on this time?Each Season 9 story follows a theme chosen by the Wheel of Fiction. Thirteen spokes. Eight are the themes from previous seasons. One is "Turn Again." One is a wild card. And three are covered in question marks and will be revealed when the wheel lands on them. See a story trailer and a (satisfying) video of the wheel turning here: Car Wash to the Cosmos This episode landed on a mystery spoke that was revealed to be TRAVELER. Ever pretend you're going at warp speed when you're sitting in your car during the drying phase of the car wash? MERCH!Interested in merch, like mugs and notebooks, featuring my artwork?Please visit my Store page for info on where you can buy: STORYFEATHER STORE NEWSLETTERSStoryfeather Gazette (if you'd like to keep up with the fiction I create) Fictioneer's Field Guide (if you'd like writing tips and guidance from me) Choose what you want. (Either way, you're choosing high jinks.) FICTION-WRITING ARTICLES & RESOURCES (New)Come and explore the Fictioneer's Field Station. I've also started posting some tips on Pinterest: pinterest.com/storyfeather/ MY FIRST BOOK (yay)Ever wonder how I've gotten all these hundreds of stories written? I have a method. You can learn it in my book called Fictioneer's Field Guide: A Game Plan for Writing Short Stories. It's now available from Amazon as an eBook, paperback, and hardcover. You can also get there from my Store page: STORYFEATHER STORE CREDITSStory: "Car Wash to the Cosmos" Copyright © 2022 by Nila L. PatelNarration, Episode Art, Editing, and Production: Nila L. Patel Music:"Space Discoveries" by ANDREW SITKOV (Intro)"Through the gravitational field" by TRG BANKS (Outro)"Abstract Vision #5" by ANDREW SITKOV (Outro) Music by TRG BANKS"Through the gravitational field""Insideoutworld""Instellar breakfast""After midnight""Above the Earth""Cigar""Keith in 1987" Music by LEE ROSEVERE"Cosmic Drifting" Music by ANDREW SITKOV (MuzStation Game Music)"Space Discoveries""A Long Way""First Contact" Tracks by Andrew Sitkov are part of a music and sound effects bundles I purchased from Humble Bundle and sourced from GameDev Market. Music by Andrew Sitkov is licensed from GameDev MarketMusic by TRG Banks is licensed under CC0 1.0 UniversalMusic by Lee Rosevere is licensed under CC BY 4.0Sound effects from AudioJungle, GameDevMarket, and Soundly (through Hindenburg)Vocal effects created with Audacity Changes made to the musical tracks? Just cropping of some to align with my narration. Find more music by Andrew Sitkov at gamedevmarket.net Find more music by TRG Banks at freemusicarchive.org/music/TRG_BanksFind more music by Lee Rosevere at freemusicarchive.org/music/lee-rosevere and leerosevere.bandcamp.comFind more stories by Nila at storyfeather.com Episode Art Description:Digital drawing. View inside the cabin of a car through the front passenger window. Foreground, a girl and a terrier sit in the front passenger seat, both facing forward and seen in three-quarters view. The girl, seen from shoulders up holds her left hand up to her lips. Her eyes are raised in a thoughtful expression. A braid falls over her right shoulder and seatbelt. She has a star sticker on her right cheek. She wears a headband with two antennae ending in balls with rings around them, like ringed planets. Her right hand, just visible is held against the chest of the terrier who has front paws propped against the window edge, and mouth open with tongue lolling out. A man, face in profile, sits in the driver's seat, both hands resting on the steering wheel. He looks at the girl and dog with a worried expression. Moon roof is open to a starry view of outer space. The driver's side window shows darkness with streaks of light. The middle panel between the seats shows glowing lights where the gear shift would be. Watermark of "Storyfeather" along side of the driver's seat.
We're digging out some of our favorite conversations from the archive. This one is from January 2025, with Julian Gough. Enjoy. Julian Gough sums up his career as follows: "I just sit in my room and write." Well, I think being an acclaimed children's author, novelist, stage playwright, poet and top-ten Irish musician is a little more impressive than he's letting on… Oh, and I didn't even mention that he wrote the ending to the computer game Minecraft! His current project, The Egg and The Rock, puts all of this to shame. This book, which Julian is writing in public on Substack, seeks to do no less than redescribe the universe, arguing that is not some random, dead, purposeless sack of chemicals, but instead a living, evolving organism. Julian joins me to discuss why the arc of human evolution bends towards man-made black holes, the hidden catastrophe at the heart of materialist science, the strange life of subterranean ice aliens, and MUCH more! This was such an interesting conversation - I can't wait for you to hear it. For the full transcript, episode takeaways, and bucketloads of other goodies designed to make you go, "Hmm, that's interesting!", check out our Substack. Important Links: Julian's Website The Egg and The Rock Julian's Twitter Show Notes: "I just sit in my room and write" Why write a book in public? Materialism & science's hidden catastrophe "The scientific method is in conflict with human nature" The faulty assumption at the heart of cosmology Big bangs, supermassive black holes & Darwinian evolution: A ~30 minute masterclass in cosmological natural selection "I'm predicting very, very large amounts of life in this universe" The strange life of subterranean ice aliens Could we spot man-made black holes? Bringing consciousness into physics Pulling back the curtain Julian as World Emperor MORE! Books & Articles Mentioned: The New Inquisition: Irrational Rationalism and the Citadel of Science; by Robert Anton Wilson Against Method: Outline of an Anarchistic Theory of Knowledge; by Paul Feyerabend What the Tortoise Said to Achilles; by Lewis Carroll The Life of the Cosmos; by Lee Smolin What Is Life? The Physical Aspect of the Living Cell; by Erwin Schrödinger Isis Unveiled: A Master-Key to the Mysteries of Ancient and Modern Science and Theology; by Helena Petrovna Blavatsky The Bhagavad Gita Did the Universe evolve?; by Lee Smolin The Great Filter - Are We Almost Past It?; by Robin Hanson
David Beth was born and raised in Africa, living all over the world in service of the spirits. He was initiated as a priest of Haitian Vodou (Houngan Asogwe) in Port au Prince, Haiti and is the current Heresiarch of the Kosmic Gnosis – an esoteric current focused on chthonic mysteries and magic. University educated in Germany and the USA, David is also the co-founder of esoteric press Theion Publishing. He has lectured and published his writings internationally.Book link: https://theionpublishing.com/shop/db-black-pilgrimage/Kosmic Gnosis site: https://kosmicgnosis.com/---Become part of the Hermitix community:The Hermitage: the-hermitage.mn.coHermitix Twitter - / hermitixpodcast Support Hermitix:Patreon - www.patreon.com/hermitix Donations: - https://www.paypal.me/hermitixpod
Beloved author Jerry Spinelli is back with a new book FIFTH GRADE TOP DOGS! He's the author of memorable books like Maniac Magee, Stargirl and Crash. Hear how Jerry structured his 44-year writing career and how Kelly read FOURTH GRADE RATS back when she was in fourth grade. Our Books for Children and Young Adults:Flying Lessons & Other Stories Edited by Ellen OhIsaiah Dunn Is My Hero by Kelly J. BaptistIsaiah Dunn Saves the Day by Kelly J. BaptistThe Electric Slide and Kai by Kelly J. Baptist; Illustrated by Darnell JohnsonThe Swag is in the Socks by Kelly J. BaptistEb & Flow by Kelly J. BaptistReady, Set, Dough! by Kelly J. BaptistThe Band in Our Basement by Kelly J. Baptist; illustrated by Jenin MohammedSee You in the Cosmos by Jack ChengThe Many Masks of Andy Zhou by Jack ChengJumped In by Patrick Flores-ScottAmerican Road Trip by Patrick Flores-ScottNo Going Back by Patrick Flores-ScottThe Griffins of Castle Cary by Heather ShumakerKelly J. Baptist: kellyiswrite.comJack Cheng: jackcheng.comPatrick Flores-Scott: patrickfloresscott.comHeather Shumaker: heathershumaker.comFind us online: Contact us: hello@booksmitten.usBluesky: @booksmitten.bsky.socialProduced by Patrick Flores-Scott and Heather Shumaker
Episode 150. A milestone worth celebrating - and worth spending on something that actually matters. Today LeeAnne teaches the science of aging well, straightfrom one of the most credible independent longevity researchers working right now.LeeAnne Hayden went to Nashville and sat in a room with5,000 people while Dr. Rhonda Patrick, PhD, founder of FoundMyFitness, broke down the four fundamental things most people are not getting enough of - and how those gaps are making aging harder than it needs to be. Not complicatedinterventions. Not expensive procedures. Four foundational levers that the research says matter more than almost anything else.In this episode LeeAnne teaches all four of them, shareswhere she stands on each one personally, and explains what happened when her husband John sat in the same room and decided at 67 that he was ready to start taking his health seriously.Her BioAge - the age her cells actually register ratherthan her chronological age - came back at 47 and a half. She is 55. That is what consistent daily inputs over years actually looks like.The four levers Dr. RhondaPatrick presented:- Vitamin D - not a vitamin but a steroid hormone that regulatesover 1,000 genes and 70% of the US population has insufficient levels- Omega-3 index - a measurable, testable number that correlateswith lower all-cause mortality, better brain health, and lower inflammation- Magnesium - involved in over 300 enzymatic reactions, stored in bone, widely deficient and dramatically underdiagnosed- NAD+ - required for cellular energy and DNA repair, declineswith age as inflammation burns through itAlso in this episode:- What BioAge actually is and how LeeAnne tested hers- Why vitamin D functions as a steroid hormone and what that means for women in menopause- The omega-3 comparison that stopped John cold- Why standard blood tests miss magnesium deficiency and what to do instead- The COSMOS study and why the 'multivitamins are expensive urine'narrative was wrong- What LeeAnne personally takes and why it landed differentlyafter hearing the science- A note for anyone who has had a cancer history and wants to know about NAD IV drips- Episode 150 - what 150 episodes actually means and where the show is goingSmall daily inputs compound over decades. You cannot fix in a year what has been declining since you were 20. But you can start now. The best time was 10 years ago. The second best time is right now.Connect with LeeAnne:- Instagram: @leeannehayden- Website: leeannehayden.com- Questions about what LeeAnne uses personally: reach out through the website or send a DM
Agradece a este podcast tantas horas de entretenimiento y disfruta de episodios exclusivos como éste. ¡Apóyale en iVoox! En este episodio de Universo de Misterios analizamos un reciente estudio que propone cómo detectar civilizaciones avanzadas que habrían agotado su combustible de fusión. Exploramos la importancia del deuterio y el helio en la atmósfera de exoplanetas como posibles firmas tecnológicas, y cómo la búsqueda de inteligencia extraterrestre (SETI) puede beneficiarse de estos hallazgos. Profundizamos en la escala Kardashev, que clasifica a las civilizaciones según su capacidad para aprovechar energía, y repasamos las implicaciones de estos conceptos para entender la evolución tecnológica en el cosmos. Gracias por escuchar Universo de Misterios. Recuerda que la ignorancia nunca es deseable. Nos escuchamos en el próximo viaje al universo que viene. Escucha el episodio completo en la app de iVoox, o descubre todo el catálogo de iVoox Originals
Zazar is an omnidimensional galactic being channeled by Dr. Marilyn Gewacke — a clinical psychologist and two-time award-winning author of "The ZaZar Transmissions." In this Portal To Ascension transmission, Zazar delivers what he calls "big news": the fifth-dimensional cities of Telos, beneath Mount Shasta, and Shambala, in the Himalayas, are merging with inner earth into one unified 5D civilization — pouring the higher blueprint of the new human into everyone on Earth. He says the ancient ones of Atlantis, Egypt, Lemuria, Stonehenge and Newgrange stored their lost wisdom as light codes for this exact moment, that the recent "Atlas" was a galactic craft carrying those codes, and that you hold dormant landing fields and star-seed DNA ready to activate. And the part nobody expects: as you ascend, your lowest parallel lives simply dissolve — you will pass people you once knew and have no memory of them at all.
When the James Webb Space Telescope started sending back images in 2022, something mysterious appeared: little red dots, speckled across the sky. They didn't look like any known celestial objects, and astronomers began racing to figure out what they were. There was a lot of hot debate, but researchers are making progress—the dots appear to be black holes, but black holes like we've never seen them before. In a recent study, astrophysicist Rohan Naidu and his team looked into one little red dot with the catchy name of MoM-BH*-1 and concluded it was a new cosmic category called a black hole star. Guest: Dr. Rohan Naidu is an assistant professor at the Institute for Astronomy at the University of Hawaiʻi. Transcripts for each segment will be available the week after the show airs on sciencefriday.com. Subscribe to this podcast. Follow our show on Instagram, TikTok, Facebook, and Bluesky @scifri and sign up for our newsletters. Got a science question that's keeping you up at night? Call us: 877-472-4374 Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
What do you say if two young Missionaries from the Church of Jesus Christ of Latter-day Saints suddenly show up at your doorstep? What do you say? How do you say it? Watchman Fellowship can help! If you know little or nothing about Latter-day Saints, the Book of Mormon, or their beliefs, doctrines, and practices, we have the perfect, handy little book, just for you! Watchman staff have combined their years of engagement and experience with Latter-day Saints into one compact, easy-to-read, easy-to-use, little book called Trading Questions - Twenty Essential Conversations to Have with Your Mormon Friend. We're featuring an exclusive new-release offer you can take advantage of right now, right here. In addition to an autographed signed copy of Trading Questions you'll also receive additional bonus material and resources on Latter-day Saints from Watchman Fellowship! Also be sure to check out our trip to Utah this September! Brady Blevins is the Senior Apologist for Watchman Fellowship. He has been married to his wife, Jennifer, for over 25 years, and they have three children: Drew, Madison, and Jett. He earned his doctorate from Southwestern Baptist Theological Seminary and has served as both a college professor and a pastor for more than 20 years.Daniel Ray is a Staff Apologist for Watchman Fellowship. He holds a master's degree from Houston Christian University and is the co-author and editor of The Story of the Cosmos: How the Heavens Declare the Glory of God. Daniel also serves as the primary host of Watchman Fellowship's podcast, Apologetics Profile.James K. Walker is the President of Watchman Fellowship and a former fourth-generation Mormon. He holds a master's degree from Criswell College and is the author of The Concise Guide to Today's Religions and Spirituality, What the Qur'an Really Teaches About Jesus, and The Truth Behind the Secret.Apologetics Profile podcast episodes about Latter-day Saint beliefs and practices. You'll find the episodes to our interview with Dr. Bowman in this link. Additional information featured in this broadcast. https://newsroom.churchofjesuschrist.org/article/history-of-missionary-work-in-the-churchhttps://churchhistorylibrary.churchofjesuschrist.org/guides/missions-and-missionaries?lang=eng#finding-local-unit-historical-recordshttps://provo.mtc.missionary.org/aboutAdditional Resources:FREE: We are also offering a subscription to our 4-page bimonthly Profiles here: www.watchman.org/FreePROFILE NOTEBOOK: Order the complete collection of Watchman Fellowship Profiles (two volumes totalling over 700 pages -- from Astrology to Zen Buddhism) in either printed or PDF formats here: www.watchman.org/NotebookSUPPORT: Help us create more content like this. Make a tax-deductible donation here: www.watchman.org/GiveApologetics Profile is a ministry of Watchman Fellowship For more information, visit www.watchman.org © 2026 Watchman Fellowship, Inc.
In this episode, we cover Lucid Group's second quarter 2026 earnings call and the company's new direction under CEO Silvio Napoli. Napoli lays out three priorities for Lucid—cash and cost, customer and quality, and culture and team—as the company works to reduce cash burn, improve liquidity, and better align production with demand. We look at the workforce reductions and changes at Lucid's Arizona factory, along with efforts to improve the ownership experience through expanded service support, mobile service capacity, and tighter software and quality controls. The call also provides updates on Gravity and several of Lucid's major projects, including the Uber and Nuro robotaxi program, the AMP-2 factory in Saudi Arabia, and the midsize platform and Cosmos prototype. Lucid isn't providing formal guidance yet, but the company expects second-half production to be below Q2 levels while deliveries should exceed production as it works through existing inventory. We also hear Napoli's broader assessment of where Lucid has fallen short and what needs to change if the company is going to turn its technology and products into more consistent execution. Support the Show Other Podcasts: Beyond the Post YouTube Beyond the Post Podcast Shuffle Playlist 918Digital Website Publications: Lucid Group Investor Relations News Links: Lucid Q2 2026 Earnings Call *Show Art Created By Dall-e Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
What are the limits of the human mind? Neil deGrasse Tyson, Chuck Nice, and Gary O'Reilly explore sensory perception, dreaming, synesthesia, time, and the flexibility of the human brain with neuroscientist David Eagleman. Why does time seem to slow down during an emergency? NOTE: StarTalk+ Patrons can listen to this entire episode commercial-free here: https://startalkmedia.com/show/your-inner-cosmos-with-david-eagleman/ Thanks to our Patrons Wynona Pyrtel, BarefootCajun, Jaap Ouwejan, Ananda Mills, Rob Schebel, Ray J, Larry Chan, Steve Pynn, Andy K, Chefscary, Rebecca Tibbs, DucksOnQuack, Deatric Wilson, Dominica Larissa, Amy Morgenthau, 芽美 湯浅, Lisa Jones, Mandie Mack, Noah Henscheid, Ethan Brady, Paul, Scott Brasfield, Alfonso Moreira, Denislav Tsankov, Chris Clawson, Miranda Kagy, Jack Frosty B.B., Rohan A, Kevin Turpin, michelle palmer, SMAYS, Francoise De Larkeen, David, Anita Taesali, Atum Ra, Ty, Joel, Emerson Castaneda, Julie Hendrickson, Mike Bayliff, John Dunn, Julian James, Davit Harutyunyan, Jakub Schmiedberger, Tristian Phillips, Tristian Phillips, Michael Pennywell, Doguhan Uluca, Aislinn P, Andrew Hatton, Misti, WB, Lindsey Walker, Clifton Beasley, Shawn Grider, Patrick Rizio, Neil LaGrace, Brooke, and keith becker for supporting us this week. Subscribe to SiriusXM Podcasts+ to listen to new episodes of StarTalk Radio ad-free and a whole week early.Start a free trial now on Apple Podcasts or by visiting siriusxm.com/podcastsplus. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Click the link in the description, or go to https://betterhelp.com/Jon and get 10% off your first month of therapy.
durée : 00:09:10 - Les Matins de France Culture - par : Julie Gacon - A l'occasion du Festival International des Jardins de Chaumont-sur-Loire, la réalisatrice Momoko Seto présente sa création : un jardin baptisé "Planètes". - équipe : Sam Baquiast, Anouk Milliot, Sarah Masson, Rodi Eken, Léa Capuano, Romain Rincon Hernandez, Salomé Erbs - invités : Momoko Seto Réalisatrice et plasticienne Vous aimez ce podcast ? Pour écouter tous les épisodes sans limite, rendez-vous sur Radio France
On todays podcast, we talk to Jayne Lambert from Cosmos about their Hero Tours, what destinations are selling well and not forgetting their Avalon rover cruises
What do you say if two young Missionaries from the Church of Jesus Christ of Latter-day Saints suddenly show up at your doorstep? Watchman Fellowship can help! If you know little or nothing about Latter-day Saints, the Book of Mormon, or their beliefs, doctrines, and practices, we have the perfect, handy little book, just for you! Watchman staff have combined their years of engagement and experience with Latter-day Saints into one compact, easy-to-read, easy-to-use, little book called Trading Questions - Twenty Essential Conversations to Have with Your Mormon Friend. We're featuring an exclusive new-release offer you can take advantage of right now, right here. In addition to an autographed signed copy of Trading Questions you'll also receive additional bonus material and resources on Latter-day Saints from Watchman Fellowship! Brady Blevins is the Senior Apologist for Watchman Fellowship. He has been married to his wife, Jennifer, for over 25 years, and they have three children: Drew, Madison, and Jett. He earned his doctorate from Southwestern Baptist Theological Seminary and has served as both a college professor and a pastor for more than 20 years.Daniel Ray is a Staff Apologist for Watchman Fellowship. He holds a master's degree from Houston Christian University and is the co-author and editor of The Story of the Cosmos: How the Heavens Declare the Glory of God. Daniel also serves as the primary host of Watchman Fellowship's podcast, Apologetics Profile.James K. Walker is the President of Watchman Fellowship and a former fourth-generation Mormon. He holds a master's degree from Criswell College and is the author of The Concise Guide to Today's Religions and Spirituality, What the Qur'an Really Teaches About Jesus, and The Truth Behind the Secret.Apologetics Profile podcast episodes about Latter-day Saint beliefs and practices. You'll find the episodes to our interview with Dr. Bowman in this link. Additional information featured in this broadcast. https://newsroom.churchofjesuschrist.org/article/history-of-missionary-work-in-the-churchhttps://churchhistorylibrary.churchofjesuschrist.org/guides/missions-and-missionaries?lang=eng#finding-local-unit-historical-recordshttps://provo.mtc.missionary.org/abouthttps://www.watchman.org/utahAdditional Resources:FREE: We are also offering a subscription to our 4-page bimonthly Profiles here: www.watchman.org/FreePROFILE NOTEBOOK: Order the complete collection of Watchman Fellowship Profiles (two volumes totalling over 700 pages -- from Astrology to Zen Buddhism) in either printed or PDF formats here: www.watchman.org/NotebookSUPPORT: Help us create more content like this. Make a tax-deductible donation here: www.watchman.org/GiveApologetics Profile is a ministry of Watchman Fellowship For more information, visit www.watchman.org © 2026 Watchman Fellowship, Inc.
Play NowThe Seibertron.com Twincast / Podcast starts out episode 407 by reflecting on new announcements from Takara Tomy for upcoming Transformers toys, and these begin with a look at the New Legends line which will provide an avenue for the release of the hard-to-find Earthspark Soundwave and the Earthspark Cosmos which Hasbro previously canceled. After that, Missing Link Soundwave and two new cassette molds to go with it are discussed, with the cast also providing their thoughts on the very high suggested retail price of the release as compared to previous entries in the Missing Link lineup. Listener questions come up next with a fun question about Optimus Prime's greatest shame, followed by one that prompts the team for their highlights of the Transformers: Energon toyline. At the end, the recurring "Bragging Rights" segment brings the episode to a close as the cast shares their most recent Transformers toy and product acquisitions.
Your very existence is a cosmic anomaly. This Cosmos summary reveals the surprising origins of all matter and why it matters.
Just a Little Background About dr.sky 2nd Guest Host in the 2nd Hour dr.karl Hricko and His Knowledge Around Astronomy. Giving a Heades up for Future Young Scientist That Joining Local Astronomy Clubs Is a Superior Way for Beginners to Learn About the Cosmos. To Then Talking About the Orgin About the Earth and the Differnt Theories Behind the Big Bang.
Michael runs a Dark Elf necromancer in Gate to Sovngarde, Ray goes deep on Nordic Souls and Baldur's Gate 3, and the crew debates AI mods on Nexus and Elder Scrolls 6.Dave's ES Stories• Dave's Elder Scrolls fan fiction — WordPress blog, Winterhold story, multiple new characters• AI-generated mods on Nexus — quality concerns, maintainability, vibe-coding debate• Gate to Sovngarde mod pack — Dark Elf necromancer build, Master difficulty (unintentional)• Chillwind Depths at level 2 — three in-game days, zombie/conjure wolf tactics• White River Watch — expanded dungeon, blind man encounter, level scaling• Legacy of the Dragonborn — keeping it vanilla-plus, museum as focus vs. mod pack integration• Nordic Souls mod pack — economy, At Your Own Pace guilds, A Horse's Life, Reading is Good, Security perk• Adamant hand-to-hand add-on — unarmed combat skill tree• Simon Magus mods — Shouts overhaul vs. Thunderchild, guild overhauls• Baldur's Gate 3 — Ray's first playthrough, turn-based learning curve, story hooks• Nova Field (Starfield mod) — alternate start, hardcore enhancement• Vic's gaming break — Gothic remake impressions, Doom: The Dark Ages anticipation• Valheim — Moder trapped in mistlands, seed problems, version 1.0 coming September• Mouse PI for Hire — 1930s noir detective mini-impressions• No Man's Sky — 10th anniversary, ‘Cosmos' update teaser, planetary rotation speculation• Fallout TV show season 2 — Fallout Feed reviews, Deathclaw episode highlight• Obsidian's new Fallout game — mixed reactions, studio acquisition concerns• Bethesda / Microsoft — studio layoffs, Elder Scrolls 6 wait, modding economics, PS6/new Xbox• Oblivion Remastered — fun but lacked longevity vs. a true new release• Memory, Sorrow and Thorn — Tad Williams trilogy recommendation from Vic• Reacher novels vs. Amazon series — Ray's mixed book experience• Sugar (Apple TV+) — Vic recommends, Colin Farrell, sci-fi twist• Rockford Files & Mannix — Ray's classic TV deep dive, LA nostalgia• Listener story — Dave's story about Alain DunstaneThanks for listening to this episode from ASAPodcasting. If you have a moment please consider leaving a review on iTunes. http://www.asapodcasting.com/#/skyrim/The Butcher, Baker and Candle Maker in Spaaace: A No Man's Sky podcast: https://anchor.fm/butcher-bakerThe Fallout Feed: https://feeds.buzzsprout.com/138919.rssVault Talk: https://podcasters.spotify.com/pod/show/asapodcasting CONTACT: askyrimaddictpodcast@gmail.com Join our Discord:https://discord.gg/cVSN65jFor Merchandise please visit the Tragically Optimistic store here:https://optimistic.threadless.com/collections/asapodcasting-showsPatreon Patreon.com/asapodcasting A big thanks to the UESP, Elder Scrolls Wiki and Imperial Library for their incredible resources.
Saturday Street Takeover in Dayton?; Downtown Hotel Shuts down; Idiot who gave 10-year old girl a lighter; Tupac's death car for sale; Disney Adults are a different breed; Nimrods; Intern Will's Punchline Report; Evening Edge Local Music Showcase featuring American Whale and their singe, "Idiot of the Cosmos.
In this episode, Ray Cochrane breaks down NVIDIA’s case for world action models, the shift that swaps a robot’s picture-describing backbone for one trained to predict what happens next. He also covers Perseverance closing in on the off-world driving record, a derelict SpaceX rocket stage hitting the Moon, and Anthropic’s rework of Claude Fable 5’s biology safeguards. Finally, he digs into Gemini Omni, Google’s undisclosed trip-planning rankings, the Danube’s record low, and iFixit’s call for Apple to unlock the iPad bootloader. – Want to start a podcast? It’s easy to get started! Sign up at Blubrry – Thinking of buying a Starlink? Use my link to support the show. Subscribe to the Newsletter. Email Ray if you want to get in touch! Like and Follow Geek News Central’s Facebook Page. Support my Show Sponsor: Best Godaddy Promo Codes Get 1Password Full Summary Cochrane opens with a personal update. Wildfires in Eastern Oregon made for a rough week of heavy smoke, and a local building burned down, which he calls a real tragedy. Meanwhile, his work at Blubrry has centered on PowerPress fixes, where reproducing customer-reported bugs remains the biggest headache. Support tickets rarely carry enough detail, and the errors themselves are often too vague to diagnose. Consequently, he is leaning toward a stronger logging and error layer, and he asks experienced developers to share what actually works for them. Beyond VLAs: NVIDIA’s Case for World Action Models The featured story comes from NVIDIA’s developer blog, and it answers a question sitting underneath this year’s robot news. Why do robot arms fall apart the moment anything changes? Move a cup six inches, swap its shape, or change the lighting, and a policy that worked perfectly in training fails. The answer, according to NVIDIA, is not the robot but the model underneath it. For the last few years, the dominant approach has been the vision-language-action model, or VLA, built on an AI that originally learned to describe pictures. Consequently, it recognizes a banana it has never seen, in a kitchen it has never seen, yet it has no idea what that banana will do next. As the article puts it, such a model “does not learn what happens to a mug when the gripper closes, how a towel folds, where an object lands when released.” Because the physics never arrives with the model, every scrap of it has to come out of hand-recorded demonstrations. The proposed fix swaps the foundation entirely. Instead of building on a model that learned to caption images, a world action model builds on one trained to predict how video continues, so the physics is already paid for. Notably, these models output an action and a prediction of what the robot’s cameras will see, in the same pass. Cochrane likens it to forethought, imagining your own motion as you make it. NVIDIA’s implementation is Cosmos 3, pretrained on roughly 767 million images and 348 million videos of real-world dynamics. It ships in 4, 16, and 64 billion parameter sizes named Edge, Nano, and Super, and it runs in real time on a Jetson Thor board bolted to the robot itself. Cochrane recalls his dad owning one of those Jetson boards, and he asks anyone working in robotics to explain how the throughput figures fit together. However, he closes on an open question: where did 348 million videos actually come from? For deeper detail, he points listeners to the source article and to NVIDIA researcher Jim Fan. Sponsor: GoDaddy Economy hosting $6.99/month, WordPress hosting $12.99/month, domains $11.99. Website builder trial available. Use codes at geeknewscentral.com/godaddy to support the show. Perseverance Closes In on the Off-World Driving Record Ars Technica reports that NASA’s Perseverance rover is about to take the record for most distance driven on another world. The mark sits at roughly 28 miles, set by NASA’s own Opportunity rover across more than fourteen years before it went quiet in 2018. As Cochrane works out on air, that averages about two miles a year. Perseverance will pass it in roughly five years instead. The difference is a navigation system called AutoNav. Since a radio signal takes several minutes to reach Mars, earlier rovers crept along pre-plotted routes and stopped every half meter to think. Perseverance carries a second computer dedicated to processing what its cameras see, so it plans while the wheels keep turning. Consequently, about ninety percent of its driving is autonomous, against roughly ten percent for Curiosity, and it averages around 110 meters an hour rather than 15 to 18. Cochrane notes researchers finding the rover at planned sites days ahead of schedule, and he wonders aloud whether world action models might drive the next one. A SpaceX Rocket Stage Slammed Into the Moon Next, Smithsonian Magazine covered the Falcon 9 upper stage that struck the Moon on August 5. That stage flew back in January 2025, carrying Firefly’s Blue Ghost and ispace’s Resilience landers, and it was never meant to end up there. SpaceX’s Julianna Scheiman says a mixture of solar activity and gravity nudged the derelict onto a lunar path after nineteen months adrift. Four tonnes of dead hardware arrived at about 5,400 miles per hour. Nobody watched it happen, and the reason is a nice bit of physics. It struck sunlit ground near a crater called Einstein, and no impact flash has ever been detected on the lit part of the Moon. However, the instruments caught the aftermath. South Korea’s Danuri orbiter imaged a dark new mark, while the European Southern Observatory’s Very Large Telescope picked up sodium and lithium in the plume, the lithium possibly shed by the rocket itself. Astrophysicist Jonathan McDowell quipped that he has “Sir Isaac Newton’s personal assurance that it did indeed hit the moon,” while planetary scientist Hannah Sargeant warns against making a habit of it. Cochrane points out the Apollo landing sites are still sitting up there. Anthropic Reworks Claude Fable 5’s Biology Safeguards Anthropic published a post on how Claude Fable 5 handles biology questions, and the bind is genuine. Biology is the textbook dual-use problem, since the knowledge behind reading your own lab results also helps someone build a weapon. Rather than refusing outright, a classifier watches for risky requests and quietly reroutes them to Claude Opus 5, a capable model without Fable 5’s biological depth. Anthropic calls that mechanism a fallback. The trouble was how often it fired on people doing nothing wrong. This update cut biology-related fallbacks by roughly 85 percent in Anthropic’s own testing, with expected overall drops of 67 percent on Claude.ai and 55 percent on Cowork. Genuinely dual-use territory still trips it, and Anthropic names virology, toxicology, and molecular design. Cochrane hit the old behavior himself and found it irritating, so he welcomes the refinement. Even so, he would rather see a false positive than a model helping someone produce a virus. Five Builders Put Gemini Omni Through Its Paces Google highlighted five builders working with Gemini Omni. Omni is a model rather than an app, and it generates video from text, images, other video, or audio, while also editing footage you already have. Google claims it “combines an intuitive understanding of physics with Gemini’s real-world knowledge,” citing gravity, kinetic energy, and fluid dynamics. As Cochrane observes, that is the same bet NVIDIA is making with robots, only pointed at video generation instead. He also flags a naming collision worth knowing about. NVIDIA calls its architecture an omni-model while Google’s product is simply Omni, two different things landing in the same week. Additionally, he encourages listeners to watch the demos, though he still senses a disconnect in AI-generated video and concedes that knowing its origin may color the impression. Gemini Wants to Plan Your Vacation Another Gemini piece, a how-to on trip planning, drew Cochrane’s sharpest take of the night. Gemini plugs straight into Google Maps, Flights, and Hotels, pulling live locations, reviews, and prices to build an itinerary. Switch on a feature called Personal Intelligence, and it reads across your Google apps, turning a messy trip-planning email chain into a clean master plan. Clever, but he calls it extremely concerning. Once these become services, he expects partnerships to quietly push particular hotels, restaurants, resorts, and destinations onto users. Notably, Google’s post never explains how any of it gets ranked, and the words sponsored, ad, affiliate, commission, and paid never appear once. There is no disclosure of a commercial arrangement, and no denial of one either. Meanwhile the post hands readers off to Viator to book tours without describing that relationship at all. Cochrane suspects the real effect shows up slowly, in the shape of small businesses continuing to disappear. The Senate Blocks a Rule on Who Controls Research Money Science reports that the Senate passed a temporary spending bill in the early hours of Saturday the 8th. The Senate’s version carries a one-paragraph rider the House version lacks, and that rider stops the White House Office of Management and Budget from finalizing a set of proposed rules. OMB builds the president’s budget, clears agency regulations, and controls how approved money actually reaches agencies. The bill itself is a stopgap, which prevents a shutdown without settling anything. The rules reach every organization that takes federal money, a pot of roughly $1.1 trillion across 41 agencies, about $150 billion of it research grants. They would let political appointees second-guess which grants get funded, allow awarded grants to be pulled when the work does not match presidential priorities, and put several countries off limits for research partnerships, China first among them. Senator Susan Collins pushed the block through after telling OMB director Russell Vought the proposal was deeply flawed, noting nearly 500,000 public comments, the vast majority opposed. However, the 90-6 vote is not law. Both chambers are on recess. Vought reportedly said the rule would not have been finalized before December anyway, and the block only lasts as long as the stopgap, which expires December 11. The Danube Falls to a Record Low ESA published a pair of Copernicus Sentinel-2 satellite images showing the same bend of the Danube, 45 kilometers upstream of Budapest, photographed a year apart. Cochrane calls the before-and-after shocking, going from green to brown completely. Wire reports put the Budapest gauge near 10 centimeters at the start of the month, about four inches of water, against a previous record of 33 centimeters set in 2018. The knock-on effects arrived fast. Budapest ran short on both power and drinking water, while the shrinking flow concentrated pollution in what remained. Romania hit record lows on its own stretch as well. Cochrane hopes the recovery is already underway. What a Heatwave Actually Does to the Power Grid That river story runs directly into a Carbon Brief factcheck. Nuclear plants cool themselves with river water, so when the Danube dropped, plants in Hungary and Romania throttled back and pulled roughly 2.5 gigawatts off the grid. Romania declared a state of alert in its energy sector, and its navy reportedly used explosives to steer more water toward a plant intake. Carbon Brief then walked through what heatwaves do to each way of making electricity. Nuclear loses efficiency when the cooling water is already warm, though its shutdowns are mostly regulatory rather than mechanical. Gas turbines pull in less air because hot air is thinner, costing capacity. Wind falls off hardest, since a heatwave is a big stalled dome of high pressure and nearly still air. Solar is the surprise: cells genuinely do get less efficient as they heat up, yet total output climbs anyway, with UK solar up 46 percent during a four-day June heatwave against the week before. Butterflies Are on the Move Everywhere A new study in Nature Ecology and Evolution covered 1,758 butterfly species, roughly one in ten of every species we have named. The team pulled 6,182 records from 105 countries, reading non-English research alongside 68 expert write-ups. Four out of five species pushed into new territory, and about 79 percent of the logged shifts traced back to climate change and extreme weather. Separately, 27 percent saw their range shrink somewhere, and 22 percent moved up or down a mountain slope chasing cooler air. That sounds like good news, and it really is not. Expansion means a boundary moved, not that butterflies are thriving, since a species can push its northern edge forward while its southern edge quietly collapses. Lead author Shawan Chowdhury says the shifts turn up on every continent where butterflies occur. Additionally, monitoring gaps leave Central Africa, Southeast Asia, New Guinea, and the Amazon Basin barely counted at all. Cochrane recalls hearing years ago that butterflies were disappearing in Hawaii, and he invites listeners spotting unfamiliar species locally to contribute what they see. Primates Make Friends Across Species Cochrane called this one a fun find. A study in the journal Primates, led by Cyril Grueter at Oxford, gathered 427 documented cases going back to the 1970s across 88 primate species and 127 partner species. Play and grooming dominated at 139 and 136 cases, alongside carrying, huddling, food sharing, and even adoption. Primates usually started the interactions themselves, with juveniles playing most, adult females handling grooming and caregiving, and adult males least likely to join in. The examples are remarkable. Japanese macaques on the island of Yakushima groom sika deer and climb on their backs, a silverback gorilla cradled a tiny wild bushbaby, and wild capuchins in Brazil adopted an infant marmoset in a bond that held for weeks. However, Grueter rejects the pet-keeping headline and prefers the hedged term proto-pet keeping. The actual claim is smaller and more interesting: curiosity, tolerance, caregiving, and play have roots running far deeper than humans do. iFixit Tells Apple to Unlock the iPad Finally, an opinion piece from Charlie Sorrel at iFixit struck a chord. This fall, iPadOS 27 drops support for a batch of older iPads, including the 8th-generation iPad, the third-generation Air, the fifth-generation mini, and the first-generation iPad Pros. Cochrane owns one of those Pros and reports it still works fine. Those devices will not break, but they stop getting OS and security updates until apps abandon them and the battery gives out. The obvious second life is Linux, except the bootloader stays locked. Apple’s iBoot will not load anything else, unlike a Mac, a PC, or most Android phones. Sorrel argues it “should be a user choice, not a vendor choice,” and Cochrane agrees flatly. You own the device, so why does Apple decide what runs on it? He compares the situation to jailbreaking, and he suspects most consumers have never pushed back simply because it never occurs to them. Nevertheless, he hopes an unlock eventually breathes new life into hardware that still works perfectly well. Cochrane wraps with housekeeping: become a GNC Insider at geeknewscentral.com/insider, email geeknews@gmail.com, subscribe to the newsletter, and grab a modern podcast app at podcastapps.com. He thanks GoDaddy for over twenty years of keeping the show on the air, and he signs off wishing listeners a wonderful evening. The post The Robot That Imagines First #1872 appeared first on Geek News Central.
Listen Ad Free https://www.solgoodmedia.com - Listen to hundreds of audiobooks, thousands of short stories, and ambient sounds all ad free!
Bill Nye & Brannon Braga on Science, Disasters, and the Stories That Save Lives In this special edition of Rewind, Tony explores the explosive, urgent, and deeply scientific world of The End Is Nye — a series that confronts real natural disasters and shows how science can help us survive, adapt, and rebuild. This episode features two conversations woven together: Bill Nye, host of The End Is Nye, on the science behind extreme weather, natural disasters, and the solutions that can protect us. Brannon Braga, executive producer, on crafting six disaster scenarios, balancing fear with hope, and how his past work — from Star Trek Voyager to Cosmos to The Orville — shaped the series' tone and mission. Together, these interviews reveal how science communication, cinematic storytelling, and real‑world urgency collide in one of the most ambitious disaster‑science series ever made. “Sci‑Fi Talk Plus — your universe, unlocked.”
Ryan and Alex explore the ninth text of the Corpus Hermeticum, “On Thought and Sense.” They discuss the principle of correspondence, the fractal nature of the universe's emanation, God-Gnosis, the path of devotion, the will of God, and the unfolding course of the Cosmos.
Listen Ad Free https://www.solgoodmedia.com - Listen to hundreds of audiobooks, thousands of short stories, and ambient sounds all ad free!
Listen Ad Free https://www.solgoodmedia.com - Listen to hundreds of audiobooks, thousands of short stories, and ambient sounds all ad free!
Dans notre système solaire comme au-delà, existe-t-il d'autres planètes ou exoplanètes susceptibles d'avoir abrité, ou d'abriter encore, une forme de vie ? Quelles pourraient être ces formes de vie et dans quelles conditions pourraient-elles apparaître ? Des astrophysiciennes croisent leurs regards sur la quête de mondes habitables et sur les recherches menées pour comprendre la place de la vie dans l'Univers. Ravis de vous retrouver autour de la question la plus vertigineuse qui soit : sommes-nous seuls dans l'univers ? Y a-t-il de la vie ailleurs dans le cosmos ? Et comment la rechercher ? Dans notre système solaire et au-delà, y a-t-il d'autres planètes, d'autres lunes, voire des exoplanètes où la vie a pu apparaitre ? Et quelles formes de vies ? Comment ces questions intersidérales et sidérantes, longtemps de l'ordre du fantasme ou de la science-fiction, sont-elles désormais au cœur des missions d'explorations spatiales ? Avec les astrophysiciennes Athéna Coustenis et Thérèse Encrenaz pour leur ouvrage L'énigme de la vie dans le cosmos paru aux Éditions Eyrolles. Musiques diffusées dans l'émission Florian Pellissier Quintet, Djeuhdjoah - Où c'est ? Qui sait ? Flore Benguigui, Sensible Notes - More Understanding Than a Man.
Dans notre système solaire comme au-delà, existe-t-il d'autres planètes ou exoplanètes susceptibles d'avoir abrité, ou d'abriter encore, une forme de vie ? Quelles pourraient être ces formes de vie et dans quelles conditions pourraient-elles apparaître ? Des astrophysiciennes croisent leurs regards sur la quête de mondes habitables et sur les recherches menées pour comprendre la place de la vie dans l'Univers. Ravis de vous retrouver autour de la question la plus vertigineuse qui soit : sommes-nous seuls dans l'univers ? Y a-t-il de la vie ailleurs dans le cosmos ? Et comment la rechercher ? Dans notre système solaire et au-delà, y a-t-il d'autres planètes, d'autres lunes, voire des exoplanètes où la vie a pu apparaitre ? Et quelles formes de vies ? Comment ces questions intersidérales et sidérantes, longtemps de l'ordre du fantasme ou de la science-fiction, sont-elles désormais au cœur des missions d'explorations spatiales ? Avec les astrophysiciennes Athéna Coustenis et Thérèse Encrenaz pour leur ouvrage L'énigme de la vie dans le cosmos paru aux Éditions Eyrolles. Musiques diffusées dans l'émission Florian Pellissier Quintet, Djeuhdjoah - Où c'est ? Qui sait ? Flore Benguigui, Sensible Notes - More Understanding Than a Man.
"Rachel Syme has interviewed all the grand dames of New York and Los Angeles: she's had Cosmos with Carol Burnett, she's accompanied Sarah Jessica Parker to the ballet, she's shopped with Parker Posey at Rachel Comey, she's Skyped with Patti LuPone and gabbed with Barbra Streisand on her landline. At a time when the glossy profile feels hopelessly compromised, Rachel uses her deep knowledge of musicals, fashion, and old Hollywood to put meat on the bones of fame and fortune—and Rachel loves a good yarn. Our conversation was full of behind-the-scenes stories, showing how an interview can transform a mere celebrity into an icon." — Merve EmreSee Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
In this Ask Me Anything episode, host Amelia Phillips answers a diverse range of listener questions on multivitamins, troubleshooting creatine side effects, and figuring out where the heck you are in perimenopause. First up, she dives into the eternal debate around multivitamins – are they essential for everyone or simply expensive wee? Experiencing bloating or digestive discomfort while using creatine? It's super common, but Amelia has some handy tips to sort you out. Finally, she helps listener Jodie figure out where exactly she is on her perimenopause journey 14 Day Vitality Reset: To find out more and join, visit: https://v360.health/14-day-vitality-reset/ To find out more about the Prevention Magazine World Menopause Day Lunch visit: https://www.worldmenopauseday.com.au/ For a deeper dive into multivitamins, listen to: FoundMyFitness Podcast – How to slow biological aging: https://open.spotify.com/episode/3jbmzKCoiG5JRZlekt6Ksy?si=db58b7abefe842a9 1:13:37 – Can a daily multivitamin slow aging? (discussion of the COSMOS trial)1:22:18 – Omega-3, vitamin D and exercise – what moves the needle? About the host: Amelia Phillips is an exercise scientist, nutritionist, and published researcher (BSc, MNut) with a career spanning 26 years in health. She is the co-founder of Vitality360, a functional health platform that helps people gain deep insights into their health and make targeted changes for lasting vitality.A respected media presenter, Amelia has been featured on Channel 9’s hit show Do You Want to Live Forever? and is dedicated to helping people build a life of energy, connection, and purpose at any age or stage of life.Instagram: @_amelia_phillipsHave a question? Email: ap@ameliaphillips.com.auFind out more at: www.ameliaphillips.com.auDiscover Vitality360: https://v360.health CREDITSHost: Amelia Phillips Audio Producer: Darren RothMusic: Matt Nicholich Production Partner: Nova Entertainment Pty Ltd Healthy Her acknowledges the Traditional Owners of the Land we have recorded this podcast on, the Gadigal people of the Eora Nation. We pay our respects to their Elders past and present and extend that respect to all Aboriginal and Torres Strait Islander cultures. See omnystudio.com/listener for privacy informationSee omnystudio.com/listener for privacy information.
The Richmond Kickers are rolling.
Listen Ad Free https://www.solgoodmedia.com - Listen to hundreds of audiobooks, thousands of short stories, and ambient sounds all ad free!
In today's Daily Fix:The cloud servers for Aliens: Fireteam Elite have been shut down, and Switch players who purchased the game no longer have access to it. Cloud versions need to be stored onto servers in order to run, and when gamers "buy" a copy, they're only buying a license to access the game, not to own the game in perpetuity. The shutdown and lack of any kind of refund or discount for the sequel have subsequently shined more light on Sony's decision to discontinue manufacturing physical game discs in 2028. In other news, Marvel's Spider-Man 2 for PS5 and Steam saw a surge in sales and player counts since the release of Spider-Man: Brand New Day in theaters over a week ago. And finally, No Man's Sky is celebrating 10 years, and a comeback story of sorts since the game's disasterous launch in 2016. A new free content update, Cosmos, was also announced.
Listen Ad Free https://www.solgoodmedia.com - Listen to hundreds of audiobooks, thousands of short stories, and ambient sounds all ad free!
Listen Ad Free https://www.solgoodmedia.com - Listen to hundreds of audiobooks, thousands of short stories, and ambient sounds all ad free!
Listen Ad Free https://www.solgoodmedia.com - Listen to hundreds of audiobooks, thousands of short stories, and ambient sounds all ad free!
In this episode, Harmony sits down with her dear friend Sarah Poe — a mother, spiritual guide, and self-described "midwife of the soul" — for a deep dive into the star wisdom systems reshaping how we understand ourselves. Sarah is the founder of Living Stars School and works at the intersection of AstroSophy, true sidereal astrology, Gene Keys, and Human Design, weaving them together into what she calls a "soul blueprint": a multidimensional map back to your innate gifts and divine nature. Sarah shares her own path from a Southern Baptist preacher's family in West Virginia to becoming a teacher of Christ-centered star wisdom — including the 3:33am wake-up calls that started it all. Harmony and Sarah unpack the real difference between tropical (Western) astrology, Vedic astrology, and true sidereal astrology, why the sun isn't where you think it is, and why Ophiuchus — the "missing" 13th sign — matters more than you'd guess. They also explore the Gene Keys, Human Design, the Magdalene mysteries, indigenous ritual (including the Lakota "spirit plate" practice), and what it really means to stop chasing someone else's path and come home to your own. In this episode: What AstroSophy is, and how it uses the life of Jesus/Yeshua to contemplate the stars Tropical vs. Vedic vs. true sidereal astrology — and the 24-degree gap between them The 72-year wobble of the Earth, and why Sarah calls it "the mother's heartbeat" Ophiuchus, the 13th sign: the healer, the serpent, the bridge between death and alignment How Gene Keys and Human Design shift when calculated through true sidereal placements Sarah's years on the Standing Rock reservation and the Lakota spirit plate ritual The Gene Keys collective cycle, Gene Key 22, and "the great remembering" Why you don't need to be everything — owning your one, unique genius Sarah's new Human Design + Gene Keys course launching this fall Links & Resources: Enroll in Sarah's Human Design & Gene Keys course, starting September 8 — https://www.poesforpeace.com/ Sarah's Living Stars Community: https://www.poesforpeace.com/living-stars-membership Harmony's intimate fall retreat (early November, west coast of Canada) — DM @harmonyslaterofficial for details Follow the show: @findingharmonypodcast · Follow Harmony: @harmonyslaterofficial Subscribe and turn on automatic downloads so you never miss an episode. The Inner Rejuvenation Codes: https://harmonyslater.kit.com/inner-rejuvenation-codes-mc Join the Lightworker Mastermind: https://harmonyslater.com/lightworker-mastermind FIND Harmony online: https://harmonyslater.com/ Harmony on IG: https://www.instagram.com/harmonyslaterofficial/ Finding Harmony Podcast on IG: https://www.instagram.com/findingharmonypodcast/ FREE Manifestation Activation: https://harmonyslater.kit.com/manifestation-activation
In this episode, Harmony sits down with her dear friend Sarah Poe — a mother, spiritual guide, and self-described "midwife of the soul" — for a deep dive into the star wisdom systems reshaping how we understand ourselves. Sarah is the founder of Living Stars School and works at the intersection of AstroSophy, true sidereal astrology, Gene Keys, and Human Design, weaving them together into what she calls a "soul blueprint": a multidimensional map back to your innate gifts and divine nature. Sarah shares her own path from a Southern Baptist preacher's family in West Virginia to becoming a teacher of Christ-centered star wisdom — including the 3:33am wake-up calls that started it all. Harmony and Sarah unpack the real difference between tropical (Western) astrology, Vedic astrology, and true sidereal astrology, why the sun isn't where you think it is, and why Ophiuchus — the "missing" 13th sign — matters more than you'd guess. They also explore the Gene Keys, Human Design, the Magdalene mysteries, indigenous ritual (including the Lakota "spirit plate" practice), and what it really means to stop chasing someone else's path and come home to your own. In this episode: What AstroSophy is, and how it uses the life of Jesus/Yeshua to contemplate the stars Tropical vs. Vedic vs. true sidereal astrology — and the 24-degree gap between them The 72-year wobble of the Earth, and why Sarah calls it "the mother's heartbeat" Ophiuchus, the 13th sign: the healer, the serpent, the bridge between death and alignment How Gene Keys and Human Design shift when calculated through true sidereal placements Sarah's years on the Standing Rock reservation and the Lakota spirit plate ritual The Gene Keys collective cycle, Gene Key 22, and "the great remembering" Why you don't need to be everything — owning your one, unique genius Sarah's new Human Design + Gene Keys course launching this fall Links & Resources: Enroll in Sarah's Human Design & Gene Keys course, starting September 8 — https://www.poesforpeace.com/ Sarah's Living Stars Community: https://www.poesforpeace.com/living-stars-membership Harmony's intimate fall retreat (early November, west coast of Canada) — DM @harmonyslaterofficial for details Follow the show: @findingharmonypodcast · Follow Harmony: @harmonyslaterofficial Subscribe and turn on automatic downloads so you never miss an episode. The Inner Rejuvenation Codes: https://harmonyslater.kit.com/inner-rejuvenation-codes-mc Join the Lightworker Mastermind: https://harmonyslater.com/lightworker-mastermind FIND Harmony online: https://harmonyslater.com/ Harmony on IG: https://www.instagram.com/harmonyslaterofficial/ Finding Harmony Podcast on IG: https://www.instagram.com/findingharmonypodcast/ FREE Manifestation Activation: https://harmonyslater.kit.com/manifestation-activation
Watch the full episode on YouTube:We first covered Baseten last year when DeepSeek mania was at peak hype. Now they have raised a monster $13B round and become one of the new cohort of AI Infra decacorns that are (with Nvidia, Intel, and the semis complex) chief beneficiaries of the Inference Inflection. We return to Baseten at the peak of the 2026 edition of Open Weights debate. Ali has published a viral breakdown of Kimi K3:And since you last saw him, Philip has spoken at AI Engineer and written the definitive book on Inference Engineering spotted all over SF:Three years ago, inference engineering barely existed as a category.Today, it is one of the most critical disciplines in AI. Inference engineering inherently tackles a different question than standard model training: “How do you turn those weights from training into a product that is fast, reliable, and affordable at scale?” Focusing on these creates an entirely new optimization problem.In one recent GLM-5.2 experiment, quantizing more of the model actually preserved its benchmark quality while increasing throughput by 20%, because the errors introduced in different layers could cancel each other out.Inference is no longer just the final step after training. It is becoming its own engineering discipline, with its own research problems, infrastructure, and increasingly specialized roles.In this episode, Baseten's Philip Kiely and Ali Taha join swyx and Vibhu to explain what actually happens after a new open model is released and what it takes to turn “we generated a token” into a fast, reliable, production-ready API.We go deep on cache-aware routing, disaggregated prefill and decode, quantization, speculative decoding, KV-cache movement, model parallelism, GPU kernels, and the race to make frontier models up to 10× faster. Philip and Ali explain why inference optimizations can still produce gains of 20%, 100%, or even 200%; how quantization errors can cancel one another out; why identical weights can behave differently across clusters; and how Baseten grafted a Kimi vision encoder onto GLM-5.2 without changing the underlying language model.The conversation then expands beyond LLMs into NVIDIA Dynamo, mega kernels, Rubin, AI-specific chips, local inference, video generation, diffusion versus autoregressive models, and the enormous compute barrier to generating coherent long-form video. Finally, we explore the convergence of training and inference, continual learning through persistent KV cache, and the emerging loop where models help optimize the infrastructure that runs them.We discuss:* What happens when a 200,000-token request enters an inference system* Cache-aware routing and reusing previously computed KV cache* Why prefill and decode are increasingly handled by different GPUs* When dedicated deployments become cheaper and more reliable than shared APIs* How speculative decoding uses a smaller model to accelerate a larger one* Tool calling, structured outputs, and what LLMs actually do* What it takes to support a new open model on day zero* Grafting Kimi's vision encoder onto GLM-5.2* Retrofitting inefficient model layers with components from other architectures* Why models sometimes collapse into repeating the same token* How hardware, kernels, and race conditions create nondeterministic failures* Preserving model fidelity while making inference faster* How quantization errors can cancel each other out* Why inference optimizations still deliver gains of 20%, 100%, and 200%* How optimized serving can make a model up to 10× faster* NVIDIA Dynamo, KV-aware routing, and distributed model serving* Speculative decoding the speculative decoder* Why local AI is about making models less dumb while data-center AI is about making them less slow* Tensor, expert, and pipeline parallelism across GPUs* Hardware-aware model design, auto-tuning, and the case against mega kernels* Rubin and why inference is becoming a systems problem* Whether modern GPUs are evolving into programmable AI ASICs* Why enormous models like Kimi K3 require GB300-class hardware* Why open-source video generation still trails Veo, Kling, and other closed models* The quadratic attention bottleneck behind long-form AI video* Autoregressive video, real-time generation, and compounding quality drift* Why future video systems may combine autoregressive and diffusion architectures* Training for inference and inference for training* Continuous post-training, deployment, evaluation, and improvement loops* How GLM-5.2 helped optimize the kernels serving GLM-5.2 itself* Why faster networking could unlock dramatically faster decoding* Continual learning, KV-cache compaction, and persistent model memoryShow Notes* How to build a day-0 API for Kimi K3* 22580: From GPT2 to Kimi3, ExplainedPhilip Kiely* LinkedIn: https://www.linkedin.com/in/philipkiely* X: https://x.com/philipkiely* Inference Engineering: https://www.baseten.co/inference-engineering/Ali Taha* LinkedIn: https://www.linkedin.com/in/aliestaha/* X: https://x.com/waterloointernTimestamps00:00:00 Introduction and the 200K-Token Prompt00:03:18 Dedicated Deployments, Speculative Decoding, and Tool Calling00:11:26 Launching Production-Ready Open Models00:19:06 Model Retrofits, Failure Modes, and Nondeterminism00:28:22 Quantization and Canceling Errors00:32:15 The Race to 10× Faster Inference00:40:48 Dynamo, Speculation, and Local vs. Data-Center AI00:50:18 Model Parallelism, Auto-Tuning, and Mega Kernels01:00:55 Rubin, GPUs vs. ASICs, and Custom AI Chips01:10:03 Giant Models and the Limits of GPU Memory01:12:42 AI Video, Quadratic Attention, and Autoregressive Generation01:21:47 Audio, Images, and Diffusion Models01:27:32 Training, Self-Optimizing Models, and Continual Learning01:40:06 Closing ThoughtsTranscriptIntroduction: Baseten, Waterloo Intern, and Inference EngineeringSwyx [00:00:00]: Okay, we're here in the studio with Philip, old friend from Inference Engineering, the book, as well as Baseten and everything that you've done, you and I have done before, as well as Ali. Welcome.Ali [00:00:15]: Pleasure to meet you.Swyx [00:00:15]: Waterloo intern.Ali [00:00:16]: Waterloo intern, always.Swyx [00:00:17]: When did you get “Waterloo intern” as a handle?Ali [00:00:19]: As a handle? Oh.Ali [00:00:20]: I think the rebranding happened mid-March. When I saw it was open, I was like, “I have to take it. Up for grabs.”Philip [00:00:26]: The problem is that Ali is really good at his job and is not gonna be an intern much longer.Philip [00:00:30]: So we have to figure out who's gonna get the handle.Ali [00:00:33]: Well, I'll pass the torch over to the next intern.Swyx [00:00:34]: Oh, okay. It can be, like, you just pass it to another Waterloo grad.Ali [00:00:37]: To another Waterloo intern. No, bruh.Philip [00:00:39]: Yeah.Ali [00:00:39]: Intern.Swyx [00:00:40]: Intern, yeah.Ali [00:00:40]: And no.Philip [00:00:41]: You gotta get an intern from Waterloo.Ali [00:00:42]: Yeah, I've gotta get an intern from Waterloo.Swyx [00:00:44]: Right.Ali [00:00:44]: But they have to follow the path.Swyx [00:00:45]: Oh, it could, but it could come from Baseten, so it's like whoever Baseten gets from Waterloo.Ali [00:00:48]: Right.Swyx [00:00:49]: Has the title of Waterloo.Ali [00:00:50]: It stays in the ecosystem.Philip [00:00:51]: Exactly.Ali [00:00:52]: Halfway through the internship, you either get it or you're out.Philip [00:00:55]: You should also do, like, a big graduation ceremony where you change the handle.Ali [00:00:59]: Just say it.Philip [00:00:59]: For everybody.Swyx [00:01:00]: You guys are good at ceremonies, clearly. We had a nice launch of the book, very successful. But before we get into all that, I wanna start off with a fun question for you. Okay, you're an expert inference engineer. What happens when I send a long query, say two hundred thousand tokens into Baseten's inference? What's the process of query through GPU model routing, balancing, all that? What is all the stuff that we don't think about?Long Context Requests, KV Cache, and Cache-Aware RoutingPhilip [00:01:26]: With a long query specifically, the first thing that I'm gonna ask is, “Have you sent me this query before, or at least part of it?” and I really hope you have, because it's gonna be a lot easier for me and a lot cheaper for you. So the first thing that we're gonna look at is some cache-aware routing, where we're going to see, we probably have a number of instances, a number of replicas up serving whatever model you're hitting. We want to send this one to something with, number one, available prefill workers, and number two, ideally some cached input already there so that we can skip prefill on at least part of these two hundred thousand tokens. If you're doing two hundred thousand tokens, it's probably coding or a multi-turn agent or something where you would expect to have that cached. If you don't, we're gonna have to send it to a prefill worker. We've at least on certain models disaggregated prefill and decode, so you're going to have one set of GPUs that's solely going to process the input, create the KV cache, and get you your first token, and then that's going to be passed over to a separate set of GPUs, which is going to run decode. We're going to iteratively make those tokens. We're probably going to have some speculator model in front of that. I'm going to assume that you're doing coding, and because of that, our speculator model, which assumes you're doing coding, is gonna have a high draft token acceptance rate. If I'm wrong and you're asking me to summarize every Harry Potter book, it's gonna be slower. And then we stream that output to you and account for it, charge you, a couple of pennies and say, “Hey, would you like to send another one?”Swyx [00:03:04]: Except Baseten doesn't charge by pennies.Philip [00:03:07]: Well, yeah, we charge. I'm assuming that we're talking about the public model APIs. If you are setting up a dedicated deployment, then yeah, it's not pennies.Public APIs vs. Dedicated DeploymentsSwyx [00:03:18]: Yeah, one of the key differentiators when I was talking with Baseten initially was that people who want very high volume just need to rent by the box, ‘cause then it's up to you to figure out how to saturate the box.Ali [00:03:31]: And more often than not, it's, like, way cheaper if you're pushing, like, millions of tokens per hour, if you just pay per hour instead of pay per token.Philip [00:03:37]: Yeah, they do. I think that we've increasingly seen a lot of demand for the pay per token APIs, just because everyone wants to try open models, and then once they find a use case that's really sticky, then they move over to dedicated.Swyx [00:03:51]: Is there a best practice on when it's time to swap over?Philip [00:03:54]: Couple reasons. Yeah, reliability, that's a big one, right?Ali [00:03:57]: Like, if they have a very specific use case, they want you to train something specifically for them, like they want their own spec dec, for instance, for their own traffic.Swyx [00:04:04]: Spec dec is speculative decoding.Speculative Decoding and Custom SpeculatorsAli [00:04:05]: Speculative decoding, yeah.Swyx [00:04:07]: You have to explain.Ali [00:04:07]: Sorry. Like, speculative decoding is like, if you have a huge model, right? And so the model is going to be generating one token at a time every single turn, every single forward pass. So we attach, like, this little, like, parasite, like this layer that goes on top of the model, and this model just has to predict. It does three very fast autoregressive forward passes, and it will predict, like, three certain tokens, and then you do one forward stage over the entire original model in order to see if those predictions were correct or not, and then you accept them or you reject them. Now, this draft model is traffic specific, so if you, like, Philip said, if you're summarizing Harry Potter books, I can train exclusively that draft model on Harry Potter books, and I can guarantee you that I'm gonna accept the three tokens every single time. And so with that case, I increase your decode speed. I wouldn't be able to provide this to you if you're a shared endpointSwyx [00:04:53]: YeahAli [00:04:53]: ‘cause I have no idea if you're doing Harry Potter, if you're doing coding, if you're doing English. We don't know. Also, there was a thing in the book that mentioned that if they really cared about a specific threshold, chapter four, I think. Do you remember that?Philip [00:05:06]: Yeah. The things that you can do is you can set a specific, like, batch sizing, a specific, like, parallelism strategy if you're trying to optimize for, like, throughput versus latency. You can. Maybe a NVFP4 quant doesn't pass your benchmarks and you wanna run a model at higher precision, you could do that. There's just a bunch of reasons why you might wanna have your own endpoint and the biggest one, of course, just being, like, you don't have to deal with someone else doing a hundred million tokens of benchmarking traffic at the endpoint when you happen to be trying to serve your users.Swyx [00:05:40]: Yeah. I think one thing that is. That is a classic journey. Like, it's people is asking the, what happens when you type Google into the browser. Tool calling, is that just, you're generating JSON or is there more complication beyond that?Tool Calling, JSON, and Structured OutputsAli [00:05:58]: Certain customers that we have, they have their own post-trained models, and so they demand a tool calling that's not just, like parse a file or go find the weather. It's something that's very specific and you have to do post-training on this. And if the post-training on the model is not good or if the quantization after the post-training to get the inference to be fast, the model will struggle reading the JSON file and reading the tool calling. But it doesn't require its own like sandbox. It's not like it's going to use that tool calling to like escape a sandbox or like it doesn't have to be contained. It can just be a normal dedicated deployment. The challenge with tool calling more and more seems to be that the companies want certain tool calling which is a very sensitive thing to train. And because you're dealing with all of the JSON outputs, if it doesn't like close the end of the request in a very certain manner, you end up with a model that did the tool calling and like the thinking and so as a result of that, it didn't see the result and just hallucinated the result as it decoded. That seems to be the most challenging thing with tool calling, not really the sandboxes model.Philip [00:06:56]: Yeah, that's a challenge on the training side and then on the inference side, there's work that you can do to scope the possible output. So we published this at this point close to two years ago, the solution to this problem which is you make a state machine and you use that to constrain the output to a specific format. So this is the structured output problem. If you remember backSwyx [00:07:27]: Yeah, the specific grammar is,Philip [00:07:29]: Yeah, exactlySwyx [00:07:30]: GML had this thing.Philip [00:07:31]: Yeah. So it's like the old-school “make sure this is only JSON”, return only JSON orSwyx [00:07:38]: YeahPhilip [00:07:38]: Grandma's gonna die type of prompts.Swyx [00:07:39]: Is it BNF grammar? At some point OpenAI had released a thing that was like, yeah, if you want to constrain your output, write BNF grammar, back as NOR.Philip [00:07:47]: In our inference system, it's just a specified output format. And you get the guarantee that your output's gonna be structured along that format. And so applying that to tool calls can like help cut down on. You can still call the wrong tool or call no tool. It doesn't solve the certainty problem but it at least solves the output structuring problemSwyx [00:08:10]: YeahPhilip [00:08:10]: Within tool calls.Swyx [00:08:12]: And MCP is just another form of tool, right.Philip [00:08:14]: Yeah, exactly.Swyx [00:08:15]: As far as there's no special thing there.Philip [00:08:16]: The thing I'm always like explaining to people is the LLM is not capable of doing anything. It's only capable of making suggestions of what to do and then if those suggestions are formatted in a certain way and applied to a system that knows what to do with them, then an action occurs.Swyx [00:08:32]: Yeah. Part of the fun stuff is, this is solved outside of tool calling too. Like in an agent loop if the output is not correct or you're right, like reasoning, tool calling was done in the reasoning trace, just be like, “Oh, I don't know what to do. Let me just try again.” And it might get there after a few tries. And on your point of training, sometimes this is harder in smaller models, so you don't have the same exact quality outputAli [00:08:56]: Right.Swyx [00:08:57]: When you just swap from a big model, right?Ali [00:08:59]: Yeah. I will say that, before, I think we need to go back to inference engineering proper.Ali [00:09:04]: But, I had expected that something would replace JSON because it's hard to stream JSON ‘cause JSON must be complete and you must have open and close brackets and everything. So it's hard to parse something or validate something while it's being streamed. So people invented all sorts of things that are like, I forget the name of some of these alternatives, but it's something like TOML, something like YAML. But JSON seems to be dominant still.Philip [00:09:30]: The JSON outputs aren't that long, right? Like you could have a long-- ‘cause tool calls also contain the arguments in them and perhaps for a certain tool you might pass like a very long argument. But my impression of the median tool call is that it's a relatively small number of tokens, right? So I would expect that speculators are generally fairly good at something as formatted as JSON. And so you would have like a pretty fast decode step there and that the streaming wouldn't be as valuable, but maybe I'm wrong about that.Ali [00:10:02]: I think you're also bounded by the software or that the model is gonna integrate with if the software is built with JSON for the tool calls or if the company that you'- if your customer says that this is how our software works and our tools are interfaced with JSON, you can ask them to like, change their software and say like, “Yeah, this is gonna be better for the model.” but like with the right training shouldn't be that much of a difference. Also more profitable if it outputs more tokens probably.Swyx [00:10:25]: Depends on your business model.Swyx [00:10:27]: It really depends. But I will say that, as a writer with like experience a lot with generated output, I do try to move from text to JSON text which is very long JSON, right? Like there's paragraphs in every field because I'm trying to structure it, right?Philip [00:10:44]: Right.Swyx [00:10:44]: I want you to first make factual statements, then make opinions then make bullet point summaries, have dates, have entity references have your sources for references, all these things. Anyway, so these are things that like I think people who really experiment with structural output have to really care about. But, let's, let's recurse up the stack a little bit. Before we started recording, you mentioned something really cool, which is that there's a lot of engineering that-- inference engineering that goes on when a new model provider releases a new model, right? So let's call it GLM-5.2, Kimi K3. I had previously assumed, especially if it's like, well, GLM 5 to 5.1 to GLM-5.2, like that you've supported them before. Is it that much work?What It Takes to Support a New Open ModelAli [00:11:26]: It's a lot of work.Swyx [00:11:28]: Yeah. Okay. So like, a lot of people, all you guys, right whenever a new model launch like, people rush to say like, “Oh, Hugging Face supports this, Fireworks supports this, Spacetime supports this,” and I'm like, “Yeah, of course we support it.” But what goes into that? What goes intoPhilip [00:11:40]: I think it's more than just support it too, right? It benefits the consumer a lot. Like I think it was with Kimi K2.5 or GLM-5.2 the latest, there was an inference war, right? X provider is at 90 tokens a second. The next day we're at 150. The nextSwyx [00:11:55]: I kinda kicked that off with the GLM-5.2.Swyx [00:11:58]: I wrote a Twitter article about. It got like half a million views,Ali [00:12:02]: Based on being numberSwyx [00:12:03]: YeahAli [00:12:04]: Or it's for something else.Swyx [00:12:05]: Yeah. Which,Ali [00:12:06]: Oh my GodSwyx [00:12:07]: Which then got everyone really excited about, hey, how can we, bend tracks a little bit further and,Philip [00:12:14]: There's a difference between support the model, as in I can make a token out of this model, and support a model, as in I have a production-ready API from this model.Philip [00:12:26]: Getting to the point of I can make a token out of this model is not that hard because generally the, open source inference engines, vLLM, SGLang of the world oftentimes even receive weights ahead of time, maintainers do, or the people making the model merge PRs to ensure support. So you generally can, just get it working on the standard open source stack without too much pain in most cases. The challenge is, every inference company is gonna have own proprietary stack. Some open source components, some in-house stuff. And for any arbitrary model, there's going to be some new stuff. Sometimes you get lucky, like K, two five to two six was, like, pretty similar.Quantization, Speculators, and Production ReadinessAli [00:13:16]: Yeah. It was pure continued post-trainingPhilip [00:13:18]: YeahAli [00:13:18]: If I remember correctly.Philip [00:13:19]: Even in those cases, there's still stuff you have to do. You have to redo the quantization work. You're taking the model from. Generally, these models are not released in NVFP4, and we want them to be in NVFP4 for maximum Blackwell compatibility. So we have to perform that quantization, and, calibrate the quantization to make sure that we're not causing any regression in the model's intelligence. And then we also have to train the speculator, as we've talked about. Generally, we have. We have ZDR, zero data retention on our model APIs, so we don't know exactly the traffic that people are sending us, but we know what's popular. We know that coding use cases are popular. We know that agents, agentic use cases are popular. So we can get public data sets that are representative of that traffic and train general speculators. Now, with speculators today, you need to train the speculator using the base model itself because you're getting hidden states out of the model from running inference on these specific prompts, and that is the training data you use to create the speculator. So there's that process which you need the real model weights for. And then there's of course just the process of, standing up all the infrastructure behind it, loading all this stuff, testing it. And then when there's a new model with a newer architecture, I think that, like, the DeepSeek models tend to be the most challenging as they have, like, the most novel architectural stuff going on, model after model. But every new model has something. Kimi K2 had. Oh, sorry, GLM-5.2 hadAli [00:14:53]: Sparse attention.Philip [00:14:54]: Yeah,Ali [00:14:54]: YeahPhilip [00:14:54]: the DSA.Ali [00:14:55]: Right. Which is brought from DeepSeek.Philip [00:14:57]: Yeah. AndAli [00:14:59]: So you can copy-paste then?Philip [00:15:01]: It kindAli [00:15:01]: I don't know how this works.Philip [00:15:02]: So, like we had to, like, build support for that into our runtime. And you're right, like it is really interesting the way that all of these open source labs borrow from each other. For example, like GLM-5.2 doesn't have vision. So something that, Haley, a guy on our team, if we could take a look at this, he, like, grafted the Kimi vision encoder onto GLM-5.2.Retrofitting Vision into GLM-5.2Ali [00:15:27]: We'll be training the projector.Philip [00:15:28]: Exactly. So if you think about, like, the encoder, there's the encoder, which is the part that looks at the image and turns it into latent information, and then there's the projector which likeAli [00:15:38]: You can say latent space. It's okay.Philip [00:15:41]: And then there's the projector that maps it onto, the model itself, and then there's the model weights. You don't wanna mess with the model weights because you run a chance of making the model dumber at something else for the purpose of giving it vision. So instead, Haley started with just a projector, which is only a handful of millions of parameters.Ali [00:16:02]: That would be, yeah.Philip [00:16:02]: Yeah.Ali [00:16:03]: Can you show the training one?Ali [00:16:04]: Like the way it groksPhilip [00:16:05]: YeahAli [00:16:06]: Very interesting.Philip [00:16:06]: And maybeAli [00:16:07]: That right therePhilip [00:16:07]: Maybe Ali, you should take it from here. You've got a betterAli [00:16:10]: Ooh, double the sandPhilip [00:16:11]: Understanding of this than I do.Ali [00:16:11]: Yeah. You can see, like, he. The way he trained this is really cool. At the beginning, he was training it using just like, “Here's a picture of a mountain. Can you describe what's in this mountain?” And that caused it just like the first, learning walls. Like here you can see this all we're trying to teach it is to translate the encoded. Like it's already taken the encoder from Kimi K. It's taken the image. It'Philip [00:16:31]: Yeah. FrozenAli [00:16:31]: FrozenPhilip [00:16:32]: With adapter.Ali [00:16:32]: Exactly.Philip [00:16:33]: Yeah.Ali [00:16:33]: So the brain is frozen and the eyes are frozen. It's just we're tryingPhilip [00:16:37]: AlignAli [00:16:38]: Interconnect between the eye and the brain, right? So the projector. And so you take the tokens and then he's like, “Oh, can you describe what's in this image?” And he's like, “Oh, it's a mountain,” or it's a person or it's a human, whatever the case is. But that didn't cause complete understanding. So he changed it such that every image was associated with a data set of questions. Like, does this image have a white male? Does this image have birds in the top corner? Does this image have a scientist in it? All of that stuff. And it would have to answer questions correctly. And using not just training on describing an image, but being able to answer question, another question, answer over time. Like you can see the grokking, which is like genuinely insane, that retrofitting vision into a large LLM can learn to that extent. And even for images that it doesn't perform well on, for instance, if you ask it a picture of like Stephen Hawking, “Who is this?” Maybe it doesn't get it, but it will say something like, “This is Albert Einstein.” Like it still understandsPhilip [00:17:25]: Close enoughAli [00:17:26]: That this is a scientist who is a man who has, some significant achievements, all that stuff. So that's like really cool.Philip [00:17:32]: Yeah. So, we've covered Hao Tian before, who the author of the LLaVA paper that did this, a while ago. And I think that's very foundational work for anyone who hasn't done vision work before.Ali [00:17:41]: Same with the CLIP and MetaCLIP, where you go from just captioning to building out questionsPhilip [00:17:47]: RightAli [00:17:47]: Off the image and how much better you can get performance.Philip [00:17:50]: Right. Right. Right. Yeah. But what's, what's so exciting about this is if you look at a model like this. Now, this is a little bit more of a research project. It's not. It got to 56% on MMLU Pro, I think. So not quite frontier. But if you're running this model, you haven't suffered any loss on your GLM-5.2 quality. If you don't have an image, it'll just behave exactly the way it used to. And ultimatelyAli [00:18:14]: Which in the inference code you literally do not include the other part, right?Philip [00:18:18]: Yeah. You would just skip the encoder if you don't have an image input.Ali [00:18:22]: Okay.Philip [00:18:22]: Just confirming.Philip [00:18:23]: YeahAli [00:18:23]: Does it affect a lot on the overall inference side? Like you're not adding much, you're adding a very small vision encoder. These are typically likePhilip [00:18:30]: They're super fineAli [00:18:31]: Less than a billion parameters, right?Philip [00:18:32]: Yeah. It's, - There's a little bit less standardization among vision encodersSwyx [00:18:37]: YeahPhilip [00:18:37]: So the support matrix can be a little bit, sparser. But overall, yeah, it's a pretty, it's a pretty minor component of the overall system. And ultimately what you get out of the system is all of a sudden you have Kimi Vision, GLM weights, and DeepSeek attention all in one model.Open Source Model Grafting and Franken-MergesPhilip [00:18:56]: And that's, I think, a lot of the power and beauty of open source, is that you can take all of these different components and combine them together into a system that's better than anyoneSwyx [00:19:05]: YeahPhilip [00:19:05]: Can be individually.Swyx [00:19:06]: People used to say that you would also do Franken-merges where you would take likePhilip [00:19:10]: YeahSwyx [00:19:10]: Layers from each model.Swyx [00:19:11]: Does anyone do that anymore?Ali [00:19:13]: Well, to your point previously when you were mentioning like, the work that goes into supporting a model when it first comes out, like GLM-5.2 or MiniMax M3 or whatever the case is. Sometimes you do have to like, you do have to switch out some things. Like, for instance, the MiniMax M3 head uses full attention, and with full attention you end up with this like insane bottleneck in spec dec ‘cause you're doing auto-regressive token generation for three tokens, and you're doing this like N squared over all of the tokens that are in your sequence. Your KV cache is like very large because it's not sparse, it's not top K. So we find it better to like, okay, we're gonna replace this, we're gonna replace this layer with a layer from another model that's using like GQA, for instance. And then just with the right training, you can get it to have the same acceptance rate. So it is very possible to retrofit layers from other models and very much needed. If a layer is like inefficient, the training just becomes the challenge, like how do you ensure that you train it properly? Which again to your earlier point is like the mesh between training and inference. As in like you need very good training in order to do fast inference. That's like, I feel like more and more becoming true.Swyx [00:20:21]: Yeah. Anything else on the support side when you say like get it to fully production ready?Loop Detection, Race Conditions, and Non-DeterminismPhilip [00:20:26]: Yeah. I think that there's also a question of just, we can test a model to a pretty extensive degree, but we're trying to get it out quickly and then you see a bunch of other people test it and you get interesting results. There was an issue with, GLM briefly where we had some like mode collapses where it would just output the same token over and over again for certain prompts on certain temperatures. Like once you expose an endpoint to the real world, there's going to be, so many more varieties of things given to it that you're able to, discover and patch things. So it's not just a, day zero process, it's then like for the first week, for the first month, if a model remains popular, like how do you both fix bugs and then continue to push the envelope on performance?Ali [00:21:21]: What do you mean you don't want your model outputting S?Swyx [00:21:24]: Is there loop detection on that stuff, by the way? It still happens like quite a lot, which is surprising.Ali [00:21:30]: We have like we, in our endpoint, like if a model was to output the same token like four plus times, we just cut the generation. We say like, “Oh, sorry, this-- Like try again,” or like we will reprocess the request. ‘Cause we know then, like if it, like if, yeah, it's four times the same token, it's probably collapsed.Swyx [00:21:45]: Yeah. Is there a way to opt out in case I really want that?Ali [00:21:48]: You want that?Ali [00:21:50]: I think there's a way that we have to handle it. I'm not exactly certain, but I feel like in certain models, like when they output something like you can imagine, like a table for instance, and so they want, they wanna draw like 12 dashes and 12 dashes. Yeah, I think there's a way for that to happen. I think we only do it on certain tokens. Like we exclude certain special characters.Swyx [00:22:07]: Yeah.Ali [00:22:07]: So we only do it on like certain like S is the most common almost. GLM-5.2Swyx [00:22:11]: OhAli [00:22:11]: And I think it was DSV 4 as well. Like you'd just have like looping issues where like you literallySwyx [00:22:17]: ItAli [00:22:17]: Just have like S.Swyx [00:22:18]: Yeah. Is there a special, something special about S? No, just randomlyAli [00:22:21]: It just seems to be the one token involved.Swyx [00:22:23]: Yeah. And it'Philip [00:22:24]: Is thereSwyx [00:22:24]: And it's only temperature 0Ali [00:22:27]: NoSwyx [00:22:27]: Even at other temperaturesAli [00:22:27]: Even at like 0.9 or whatever, it will still, it will still collapse.Swyx [00:22:30]: That's weird, right?Ali [00:22:30]: It's, it is an inference problem to be honest, like a software problem. Like oftentimes, the image you run will-- like NVIDIA will release an image for instance, and if we will upstream the changes from their latest TensorRT-LLM image into our stack, we'll find that it fixes it. Or oftentimes this will only happen in an inference engine that you're using like SGLang. But if you were to switch to vLLM, that isn't the case. So it seems to be like an extremely like deterministic software issue and not really a model issue. It's not like a weights problem. Like I'- we'll say like, “Oh, it's a problem with the quant. We did PTQ wrong,” right? But that isn't, that doesn't make sense because the same weights used with a different inference engine does not repeat the problem. And sometimes it's, the kernels that are being used in the backend have like these very subtle sometimes race conditions, where if you were to use this model hosted on one cluster, you will never get this problem.Swyx [00:23:19]: Oh my God.Ali [00:23:19]: But if you host it on a different cluster, you will. And the reason is the KV cache transfer from a node to node in that one cluster is using a slower interconnect than the node to node in another cluster. So that exposes the race, whereas in another cluster it doesn't. So then you end up just like, okay, this model is not gonna be hosted on this cluster. We're gonna host it on, another cluster because that cluster exposed that problem. But then it ends up with like, okay, is it the software? Is it the model weights or is it the hardware?Swyx [00:23:42]: There is a thing about this with temperature 0 still not being deterministic, right?Ali [00:23:46]: Right.Swyx [00:23:46]: Mostly because of hardware. Even at temperature 0 same model, you won't always get the same output.Swyx [00:23:52]: Even-- But I'm surprised by the race condition one because, I thought PyTorch was a graph that like guarantees that you at least, execute things in the right order.Ali [00:24:02]: Well, yeah, true. Like I'm not, I'm not saying that there is. Like well, you have things like PTL optimizations where like you can start a kernel before the end of the previous kernel, and that's like ‘cause you want to do that because there'sSwyx [00:24:12]: It's like pipeliningAli [00:24:12]: Expense. Exactly.Swyx [00:24:13]: Yeah.Ali [00:24:13]: But it'- But you don't do it cleanly. Like you overlap a little bit of the execution. No, it is very possible that the kernel itself, like that one block that is supposed to be running in this instance of time, that kernel itself has a race condition. For instance, like a missing barrier. Like often if you're designing a kernel and you want it to make it to be very fast, if you don't test it extensively, you'll, you'll have certain threads access data points from registers before they've been written to by other threadsSwyx [00:24:36]: YeahAli [00:24:36]: For example, because like your barrier is wrong or your synchronization was wrong. But yeah, like the testing itself is very difficult in those like, andSwyx [00:24:42]: And there's no like borrow checkerAli [00:24:45]: What does that mean?Swyx [00:24:46]: Like Rust. Like the. If you're trying to have like memory safety It sounds like a comparable problem.Ali [00:24:52]: Well, yes, but you're working in CUDA, right, NVIDIA GPUs. Like- You just need a higher level language like modular Maybe that's what modular is supposed to do. I don't know.Quantization Quality and Vendor FidelityVibhu [00:25:00]: How do you see keeping quality of the model? So you talked about all these steps of, okay, you gotta do quantization, train your own speculative decoderAli [00:25:07]: RightVibhu [00:25:07]: Run on different hardware. Looking at other model providers, okay, you kicked off a inference speed race on the consumer end. What goes into keeping quality the same across them, right? Sure, you can run benchmarksAli [00:25:22]: YeahVibhu [00:25:22]: But, like, how do you determine how much quantization are there standards? What goes intoPhilip [00:25:27]: There's a few things on quality. Most inference optimizations are lossless. KV caching, for example. You are just recomputing or preventing recomputing the same values. Speculation, of course, if a draft token is wrong, it gets rejected. The main lossy optimization is quantization. And that really comes down to, number one, data format, number two, which parts of the model you choose to quantize, which layers, and number three, like doing a lot of calibration on the quantized weights, to ensure that you're preserving all the outliers. There's other tricks that you can do, though. A big one is long context, ‘cause one thing you asked at, right at the beginning is, “Oh, what's gonna happen if I send a 200,000 token request in?” So with a long input sequence, you need to, store a lot more information. You need to process a lot more tokens. And so even if a model has a context of a certain length, you might, as an inference provider, choose to build an API with a shorter context length, and of course a full length one as well. Because if someone doesn't need the full million token context, for example, you can get them better performance. I don't know if that's exactly like quality of the model. The way that I think about quality is to what degree are we faithfully serving the original model? If you think of a golden implementation of a model that performs exactly the way the model is designed to perform, I think of quality as how close are we getting to that, 100% fidelity of the model.Philip [00:27:13]: You can also, of course, think about quality from the training side and how do you push yourself past 100%. But when I think about purely inference optimizations, it's getting faster while staying as close to that 100% fidelity mark as possible. And certainly our standard internally is that, like you should not be able to tell the difference between our API and a, official API. I think Kimi in particular does a good job of vendor benchmarking hereAli [00:27:41]: YesPhilip [00:27:41]: Where they haveAli [00:27:42]: They released an actual vendor benchmark.Philip [00:27:43]: Exactly, yeah.Ali [00:27:44]: ‘Cause they accused, some people, Amazon? There was some provider that was not doing very well on Kimi's benchmark.Philip [00:27:50]: Yeah.Philip [00:27:51]: So, with Reflect we probablyVibhu [00:27:52]: This was a long time ago, right?Philip [00:27:54]: No.Ali [00:27:54]: Yeah, like threeVibhu [00:27:55]: They alsoAli [00:27:55]: Four, five months agoVibhu [00:27:57]: This also happened with, I don't remember which model, but they pulled out quite a few, and then they started a whole chart about this. It might have beenPhilip [00:28:03]: Kimi Vendor Verifier.Ali [00:28:04]: Yeah.Philip [00:28:05]: Yeah.Ali [00:28:05]: Yeah, ‘cause you, ‘cause you'd be pissed, right? Like if you'Philip [00:28:07]: Yeah.Ali [00:28:07]: If like if I'm a consumer and I'm using like Amazon's endpoint for instance, and I've used Kimi and I'm like, “Oh my God, like this is bad,” I'm not gonna say, “Oh, Amazon quantized the model in a bad way.” I'm gonna say, “Oh, Kimi sucks.” Right?Philip [00:28:17]: Yeah.Ali [00:28:17]: So it seems like that makes sense.Philip [00:28:19]: Yeah, they care. They care.Vibhu [00:28:21]: Justifiably.Ali [00:28:21]: Yeah, justifiably.Vibhu [00:28:22]: This is probably a stupid question, but just checking, has anything improved from main quantization?Philip [00:28:28]: Yeah.Vibhu [00:28:28]: Like, is quantization always strictly worse?Ali [00:28:30]: Well technicallyVibhu [00:28:32]: NoAli [00:28:32]: It's a lossy. QuantizationPhilip [00:28:33]: YeahAli [00:28:33]: Is a lossy, it's a lossy implementation.Philip [00:28:36]: Speed improvesVibhu [00:28:36]: Speed improves.Ali [00:28:37]: It the number, likeVibhu [00:28:38]: No, I' always look for inverse scaling laws.Philip [00:28:40]: Yeah.Ali [00:28:40]: Yeah.Vibhu [00:28:40]: This is something I learned from Noam Brown, where like things that normally act in one direction sometimes do.Philip [00:28:45]: Well, technically when you run a benchmark, because these models are deterministic, sometimes your,Ali [00:28:52]: YeahPhilip [00:28:52]: NVFP4 quant is like, two basis points higher than yourAli [00:28:56]: No, it's noise. It's noise.Philip [00:28:57]: Yeah, exactly. I'm like, yeah, it's, it's within. That's why I always say within margin of error.Philip [00:29:01]: And I stopped saying that because everyone assumes that what is, well, within some margin of error, we're barely inside of that to the worst, so we're saying. But yeah, sometimes it's just like, gives you a higher output score. But like Ali said, that's noise. To my knowledge, you're not necessarily making the results better. You're just trying to, again, like keep your fidelity as close to 100% to the original model.Layer Selection, KL Divergence, and Better QuantizationAli [00:29:27]: There is, to your point, research that we did on MP. I don't know if you are able to pullPhilip [00:29:31]: YeahAli [00:29:32]: A tweet we did. One of our research interns, Joshua, I think it's a tweet on how we have 20% better quantized GLM-5.2 than NVIDIA. Essentially what we found throughout like this month research is, okay, quantization is a lossy. It's. You're compressing the data from, occupying 16 bits to occupying, four bits, for instance. And so you're losing some information, and you're trying to minimize that. And so when I say that I'm gonna quantize the model, my job becomes how do I find the layers that I can quantize, and how to find the layers to not. For instance, with image models, I don't quantize modulation layers, and I don't quantize out projections because those two are. Like out projection is what you see as the user. Modulation is what the model sees or understands. Right, exactly. And so to his paper, do you have the. It doesn't have the. Yeah. It's a long paper. I don't know if I can findVibhu [00:30:25]: If there's a part to search or it's probably in the thread.Ali [00:30:28]: It's probably in the thread.Vibhu [00:30:29]: Yeah.Ali [00:30:29]: But the long and the short is it is very possible that quantizing more of the model makes the results. Like if I have a model that I quantize layers one, five, and 10, and another model where I only quantize layers one and It is possible that the model in which I quantized more information is going to perform better because the quantization errors have canceled out. And so what Joshua showed in his mathematical proof where he had like a verifier in, is that you can predict which layers are going to have quantization errors that will cancel out with each other, and you choose to quantize those layers. And so the result of doing this mathematical quantization is you end up with a model that's 20% more quantized than another provider, so you get 20% more throughput of it because there's more layers than running an NVFP4, and your quality is better than that other quant because the layers that you chose to quantize have their errors cancel out, like one layer skewed to the right one layer skewed to the left, one layer skewed to the right. Your final logits distribution is more similar to the original distribution of the model, so you have better fidelity. And so the way we proved this was with KL divergence. So instead of just scoring on the benchmarks, we scored the KL divergence between the logit distribution of the quantized model and the logit distribution of the original full precision model, and we showed that with this technique we get. If your probability distribution on the logits which token it wants to select is more of the same as the original model, you're probably gonna end up staying true to the original model. So yeah, so it seems like previously before this, it seemed like the industry was, well, the more you quantize, the worse it's gonna be, ‘cause the more loss you introduce. That's not exactly, not necessarily true. So yeah, doesn't improve it, but can cancel out.Philip [00:31:57]: I think it might be this, but reminds me a good bit about pruning where you can prune off certain layers.Philip [00:32:03]: But very interesting. Didn't know this was a whole paper you guys put out.Ali [00:32:06]: It's. Fun fact, it was originally 72 pages, this paper, and then we decidedPhilip [00:32:11]: WowAli [00:32:11]: We can't tell. We couldn't release it. So it's now 45.Swyx [00:32:15]: Still 39 pages, so very substantive. We talked about evals and all these things and, like what's possible in terms of speedup? Like it's like probably like the numberInference Speedups and BenchmarkingSwyx [00:32:25]: Thing that people do wanna care about, and it's something that you wrote about in your post. Like official API is 70 tokens per second, and you push it up to 90. Is that like a normal thing?Philip [00:32:36]: So what's cool about working in inference, the reason that I think inference is going to be a useful place to do engineering for a long time, is that if you look at highly optimized domains like, say, finance, if you're in finance, you measure how much better you got in basis points. It's like, “Oh, I got five basis points better, like twentieth of 1% better,” that's huge news because everything is so optimized. When we publish optimizations, it's 20%, it's 100% it's 200%. So there's still probably like a lot further to go, honestly. Like you'll, you'll know that inference is pretty much solved when researchers start publishing about how they got 1% faster at something.Swyx [00:33:19]: Which by the way, because I am from the finance background, in the ‘70s, that was the margin at the time. When you did quantitative finance research, you would findAli [00:33:27]: And like 20%, tens of percent.Swyx [00:33:29]: That's. Yes.Philip [00:33:29]: Yeah.Swyx [00:33:30]: And now it'Philip [00:33:31]: Tiny fractionsSwyx [00:33:32]: For those people interested, look up Andrew Lo's paper. He had a really interesting illustration of quant, stat arb, distribution, narrowing down from like those kinds of 20% differences in the ‘70s, down to nothing today, which is very cool.Philip [00:33:48]: Exactly, and we're at the beginning of the same type of thing. Now benchmarking is hard. I think anyone will tell you that, and benchmarking provider speeds is hard because there's so many variables that go into it. What hardware are you using? How much load do you have on the system? What's the exact nature of the prompts and input and output sequence lengths? All that stuff. But overall, when you start stacking these improvements, you're looking at multiples. You can look at it. The most common form, of course, is TPS, tokens per second, which is bad naming by us in the industry, ‘cause there's two tokens per second. There's tokens per second, the throughput number, and the latency number.Ali [00:34:31]: TTMT, yeah.Philip [00:34:32]: Like total tokens per second out of the, out of the GPU as a throughput number. Most people only care about tokens per second as the latency number, which we should call ITL, intertoken latency, but we don't.Philip [00:34:44]: Anyway, so you can imagine a standard API without many optimizations for a 1 trillion parameter model operating somewhere in the 30 to 50 tokens per second range for reasonable traffic profile. And we generally see the goal of, pushing to 10X that. But, not necessarily day zero, but by stacking enough optimizations, if you have, say like four optimizations, each of which doubles performance. Or sorry, three optimizations, each of which doubles performance, then you stack that up, that's an 8X gain. That's the order of magnitude that we're working with in this space. We're trying to make things substantially faster, not just go from like 70 to 90.Swyx [00:35:38]: Are you saying you've. You have done that?Philip [00:35:40]: So let's say you have as a reasonable baseline, 30 or 40 tokens per second. You can achieve 10X that. So like on GLM-5.2, if you run it unquantized, perhaps on H100s even, and you're just using an off-the-shelf inference engine with no particular optimizations, no speculator, nothing extra around like KV routing, no disaggregation, you're, you're probably, yeah, looking at that like 30 to 40. You think that's like a reasonable baseline?Swyx [00:36:12]: Right. Right.Philip [00:36:12]: To get to something like 10X, there's a lot of trade-offs that you're making. If we're running at more like a 300, 400 tokens per second range, you are using the best hardware possible. You have a optimized speculator. You have done all of your quantization work. You are Seeing a pretty high cache hit rate. You are running with a reasonably small batch size and a parallelism configuration that is tuned for latency versus throughput, but it is possible. So the spreads that you see if you, like, go on artificial analysis or you go on OpenRouter and you look at, the worst provider to the best provider, oftentimes can hit that range. 10X is of course very aggressive. It's oftentimes maybe more of a four to six times improvement. But that's the performance that makes us really excited, is when we can get these huge gains, not just go from 70 to 90 tokens.Stacking Optimizations: NVFP4, Speculation, and DisaggregationAli [00:37:19]: It's also, like, hardware dependent. Like, ifPhilip [00:37:20]: YeahAli [00:37:20]: If you have a thing where you're serving it on just, like, a node of H100s and then you throw, like, you shard the model across, like, four nodes of B200s. Like, you can definitely increase the speed with just throwing more hardware at it. Like, normalizing for the same exact hardware and the same number of GPUs.Philip [00:37:35]: Yeah. Then you're looking at, like, a two to 4X improvementAli [00:37:38]: Right. RightPhilip [00:37:38]: Depending on the inference optimizations. So yeah, it's. Some of it's, what's the call, and some of it's who's the driver.Vibhu [00:37:46]: If you break down the two to 4X, say the example is run GLM-5.2Ali [00:37:51]: YeahVibhu [00:37:51]: On B200sAli [00:37:53]: YeahVibhu [00:37:53]: Single node, right? What's, like, the cost trade-off for effort to get, like, the last bit of juice out versus what should people just think of, right?Ali [00:38:01]: Spectre quantization. Yeah.Vibhu [00:38:03]: Spectre quantization.Ali [00:38:04]: That's, that's, that's like 95%. LikeVibhu [00:38:06]: And how far does that get you? And how easy is that for the average person to do? So say right I wanna throw the weights of GLM-5.2 on a node of B200s, how easy is it to find speculative decoder- decoder model or already quantized model? How much work goes into it?Philip [00:38:23]: If you're doing it up front, it's quite a lot of work. If you're doing it today, there's going to be people who have published things that you can just, you can just grab some NVFP4 weights. You can grab a speculator. Yeah, if we're thinking about, like, what are the 2Xs we're stacking, going from, BF16 to NVFP4 is, it's not quite a 2X, right? It's like. I think it's about, like, 30 to 40%, from 16 to 8, and then another 30 to 40% multiplied from, 8 to 4. So that doesn't quite get you a 2X, but, like, roughly a 2X. Speculator, roughly a 2X. Disagg on top of that if you're able to get enough hardware and put enough traffic through it, another roughly a 2X. And then you add in some, double-digit percent increase from having just a better runtime with, the latest kernels and stuff behind it. And that's how it stacks up.Ali [00:39:21]: YeahPhilip [00:39:21]: So building each of those, like, building the, quantized weights is, for someone who really knows what they're doing, hours to days of work. Building the speculator, again, like, hours to days of work. And the, disagg setup, hours to days. Well okay, but like once you haveAli [00:39:39]: Once set up. Once set up. YeahPhilip [00:39:40]: Yeah, getting disagg working for the first time, I'm saying, of course, is very difficult.Philip [00:39:44]: The marginal implementationAli [00:39:48]: Like, if you're just grabbing, like if you are a person, like just a normal consumer who has access to, like, a node of B200s and you're wondering, “How can I just host it myself?” You don't need to quantize the model yourself. There's always gonna be, like, an open source quantized checkpoint. NVIDIA's gonna push one out if no one else does. You. Usually, the providers will have their own spec dec that they've trained as well. You don't need to train your own spec dec. You can just use that as well.Philip [00:40:09]: Yeah. Like, GLM-5.2 has its own MTP.Ali [00:40:13]: Right. Right.Vibhu [00:40:14]: What's multi token prediction?Philip [00:40:15]: Yes.Ali [00:40:16]: I'm justVibhu [00:40:16]: Can you explain that?Ali [00:40:16]: I'm just an expert.Ali [00:40:18]: I can do it for you in case I get it wrong?Vibhu [00:40:20]: No.Vibhu [00:40:21]: Yeah, you should correct if we're wrong, but their multi-token prediction can be used for self-speculative decoding.Ali [00:40:27]: I'm not sure. I'm not gonna correct that.Vibhu [00:40:28]: Okay. I'm semi-confident in thatAli [00:40:30]: Okay. YeahVibhu [00:40:30]: But someone can check. But it's useful to paint the story of, okay, not just the average person, but say a company wants to switch from serverless inference I wanna throw this up on. I wanna rent some GPUs, throw it up. These are the steps you take to do significantly faster than just put it behind vLLM.Ali [00:40:48]: Right.Vibhu [00:40:49]: I was waiting for a mention of Dynamo.Vibhu [00:40:51]: I feel like, that's supposed to be the baseline that you measure against.Dynamo, KV Routing, and Disaggregation ToolkitsPhilip [00:40:55]: I would think of Dynamo as less of a box system and more of a toolkit for building with. So when we talk about doing aware routing, when we talk about doing KV offloading, when we talk about doing, PD disaggregation, Dynamo fundamentally is. By the way, Dynamo is an open source library from NVIDIA.Ali [00:41:17]: We've done a pod with KylePhilip [00:41:18]: OkayAli [00:41:19]: Kyle Cranin.Philip [00:41:19]: Cool. So then your listeners know then that it supports all the different inference frameworks. And it is multi hardware, which is interesting.Ali [00:41:28]: But it's just a router, it's not like an optimizer layer.Philip [00:41:30]: Yeah. All it does, like, what Dynamo is good at, it is a library for moving information around your cluster, around your hardware. So if you have, KV cache on one place and you need it to be somewhere else, Dynamo coordinates NIXL for you to move that around.Philip [00:41:49]: That doesn't mean that, like, out of the box, you just say, “Pip install Dynamo,” and then you get, like, a massive performance speed up. It's more of a developer toolkit.Ali [00:42:01]: Yeah. I would have said it would. It comes with a set of defaults that you can then swap out.Philip [00:42:06]: It does. If the industry at large, I think, was, like, rolling out all of these deployments, standard, then I think it would be, like, a credible baseline. But, we've got to, we've got to benchmark against, like, what we're seeing in the wild.Speculative Decoding Methods: Medusa, EAGLE, n-Gram, and Spec-SpecVibhu [00:42:23]: I did wanna talk a little bit more about PD disagg, because that is probably, like, number three after quantized and speculative decoding. In your book though, I was just gonna pull out the book.Philip [00:42:31]: Yeah.Vibhu [00:42:32]: Like section 522 on Medusa, 523 on EAGLEPhilip [00:42:35]: YeahVibhu [00:42:36]: 524 on gram.Philip [00:42:37]: It's 55, would be disaggregationAli [00:42:42]: Yeah. Well, no, I just wanted to dwell a little bitPhilip [00:42:44]: YeahAli [00:42:44]: The other. Like, so what do you choose to include? What do you choose to not to include? Because there was all these other techniques.Philip [00:42:51]: Yeah.Ali [00:42:51]: Are these still relevant? Because I think they came out, like, a year and a half ago maybe.Vibhu [00:42:55]: Medusa is quite old.Philip [00:42:56]: Yeah, Medusa's old.Ali [00:42:58]: It was old.Vibhu [00:42:58]: But is it in the book as a good, here'sPhilip [00:43:01]: BaselineVibhu [00:43:01]: Baseline vanilla understand it?Philip [00:43:02]: Like you should know this.Vibhu [00:43:03]: Like I read the paper, I'm like, “ it makes so much sense.”Philip [00:43:05]: Yeah.Philip [00:43:05]: So with the book, I had a couple goals. One was to give people just a working vocabulary for the space as a whole, and the other was to give them some intuition about how each of these techniques works. As I mentioned in my AI Engineer talk, which is the first public addendum to this, the speculation space has moved much faster than everything else. So yeah, even at the time that I wrote the book Medusa, I very much included as a way for people to understand how the space evolved rather than what the most modern technique is. And now of course, there's DFlash, dSpark. There's, there's newer techniques even than EAGLE, although EAGLE is still very commonly used.Ali [00:43:51]: SpecSpecta.Philip [00:43:52]: Yes. Speculative decoding.Vibhu [00:43:54]: What canAli [00:43:56]: Oh, it's a paper by Tri Dao and it's like, it's doing speculative decodingVibhu [00:44:00]: HuhAli [00:44:01]: For the speculative decoder.Philip [00:44:02]: Oh, in spec- oh my God.Ali [00:44:02]: It's literally just an another. It's like, yeah, that's the most simple way to explain it, and it seems like he got trivial speed ups there. But it seems that the complexity with training, it's almost like in our mind at least, it's almost as complex as training GANs. Like it's like a very delicate balance and oftentimes you, it's just but yeah, it's literally speculative decoding on speculative decoding.Vibhu [00:44:21]: Speculative.Ali [00:44:22]: Yeah. We saw this paper.Vibhu [00:44:24]: It's interesting, right?Ali [00:44:24]: Yeah.Vibhu [00:44:24]: I wouldn't even expect it to be very particular to train, I wouldAli [00:44:29]: Right.Vibhu [00:44:29]: The naive part of me is like, okay, train speculative decoder.Ali [00:44:32]: But like, and it makes sense, like the whole idea of speculative decoding is you. It's like, it's like almost like the iPhone auto predict version but for a normal model, right? Like you're just, you're just, generating three tokens and you're like, okay, I'll do prefill on them. And so you save those three turns for your original model. Now your speculative decoder is doing three turns of auto regression, so why not just have an even smaller model?Ali [00:44:53]: The other question there is what are the size of speculators? So say forPhilip [00:44:58]: Right. It's like a billion parameters.Ali [00:45:01]: Like for MiniMax, it's. Yeah. It's like one layer. It's like one 60th of the original model usually.Philip [00:45:06]: Yeah. I think we should do a paper when we get back to the office.Philip [00:45:10]: SpeculativeAli [00:45:11]: SpeculativePhilip [00:45:11]: Decoding.Ali [00:45:13]: No, it's, it does seem like how, when do you stop? But then it also seems like if you're able to train spec-spec decode for instance, right? Like if you're able to have a small model that is accurately predicts what the intermediate speculator is gonna predict, that is able to predict what the original target model's gonna predict, then why not just use that smallest model directly, right?Vibhu [00:45:34]: Yeah. This isAli [00:45:35]: Like it seems likeVibhu [00:45:35]: Adjacent to the routing problem.Ali [00:45:36]: Right.Vibhu [00:45:36]: Yeah.Ali [00:45:36]: Right.Philip [00:45:37]: The thing with speculators is one of the practical constraints on using them is that you do have to run a small model on the same hardware that you're running the big model on. There is a orchestration and resource competition problem inherent in that, and that is one of the constraints on speculation in general, is that draft tokens cost resources to create and cost software complexity to manage. And so if you have like infinitely recursive speculators, you add in quite a bit of that complexity on the actual implementation within the inference engine as well, not just in the training process.Vibhu [00:46:17]: I was gonna say, I would wonder if you could do similar, like distillation and pruning of, it's the same thing, it's just a model. Can we not just distill a lot of the weights, quantize the speculator, out of my domain? The question that also comes up is, this is all for big server workloads, right? How much of this applies to, say I have this MacBook, I wanna run Gemma really efficiently. Similar problems, not the same?Local AI vs. Data Center InferencePhilip [00:46:45]: Pretty different. I talked to Selo, about this on his podcast a couple weeks ago. The difference between inference engineering for the data center and for production workloads versus inference engineering for local AI, is that we start with fundamentally like different constraints and different goals. With local AI, it's how do I fit this model onto my hardware and then make it less dumb? And with data center influence, it's how do I load this model and then make it less slow? And we care about less dumb, and they care about less slow. But the local AI inference engineering ecosystem, I think has a lot for us to learn from in the data center space. They are experts in various forms of quantization, including dynamic quantization that we just don't touch, in the pruning, in the distillation, in the, layer removal. There'Ali [00:47:42]: Layer removal matters less.Philip [00:47:43]: Yeah. There'Ali [00:47:44]: No one loves pruning really.Philip [00:47:45]: Yeah. Well, but the, but they doVibhu [00:47:46]: Which is surprising, right? But that's, that's a whole different thingPhilip [00:47:48]: Just to fit something on the laptop.Ali [00:47:50]: Right.Philip [00:47:50]: So yeah, it's a, it's an interesting, it's an interesting space. Not necessarily that like their techniques make sense for us to do in the data center, because we have different resources and different goals, but more that the process as well as the openness of that field is something to, admire.Ali [00:48:12]: Yeah. Like to your point, like, certain optimizations that would. Like for instance, Turbo Quantum Sharper, like it made such huge hype on that and we did like a whole deep dive on Twitter and like said, what is it? How does it work? Why is it good or not? And it took off and it was implemented on local devices because your memory bandwidth is so slow on like a MacBook, for instance. But try putting the same thing on like an NVIDIA GPU on a B200 Turbo quant would not be. Like, it would not be used. Like, NVIDIA - Like, NVIDIA made it clear that this is not a good optimization, and we've seen it firsthand where the overhead of doing dequantization, quantization of, in the kernel itself with turbo quant kernel, each end is much slower than the time that you save from doing the bandwidth. ‘Cause on the B200s, you have like 3.5 terabytes per second. You don't need decrease the storage that much. You don't need to do, FP4 KV cache. You don't need to use a requant. There's, there's, there's better optimizations to be made. But on Edge devices, it's extremely important, it's extremely useful. So, seems to be, like, different optimizations there, but then they're all uniquely combined with like all you wanna quantize the model, you wanna do speculative decoding, like certain common prefixes with bothPhilip [00:49:18]: Principles.Ali [00:49:19]: Yeah, exactly. Exactly. Exactly.Philip [00:49:20]: They also do a lot of work on, model parallelism, especially over, heterogeneous topology, where you have, some sparks and they are wired together with, Ethernet, DGX sparks.Ali [00:49:35]: Yeah, this is the Exo Labs guys.Philip [00:49:36]: Yeah. You have, a nu
Listen Now to 020 WTFuture Watch 020 WTFuture Hey folks! We are coming to you from Boulder Creek today, joined by Bobby in sunny San Francisco, before we head off on our annual summer vacation to Montreal! In this week’s chat, we completely geeked out over the future of our “augmented brains” and the rise of local AI, like the mind-blowing new Bonsai tech that squeezes a massive 27-billion parameter model right onto an iPhone. It is a total game-changer because it acts as a private, “sovereign” AI companion that does not need the cloud, saves energy, and ensures nobody steals your ideas. We also had a great time joking about the “dead internet” theory now that bot traffic has surged 7,800%. But don’t worry, human creativity is still king—especially when we use AI as a collaborator to create hilarious, custom animations of floppy disks barfing green data goo, complete with gassy liquid sound fx! Looking to the stars, we were absolutely captivated by SpaceX’s Starship landing beautifully in the Indian Ocean and their wild plans to launch giant AI compute satellites to build an off-planet internet. We also discussed the hilarious twist of a 1972 Soviet Venus probe, Cosmos 482, finally crash-landing back on Earth after a 53-year detour in orbit. To top it all off, astronomers just found LHS1140B, a rocky exoplanet 48 light-years away that has a faint stream of helium, marking the first time we’ve detected an atmosphere on a habitable-zone rocky world! It’s a beautifully abundant, high-tech future ahead of us, and as we pack our bags for the Laurentian Mountains, we can’t wait to see what wildly optimistic paradigm shifts happen next. Enjoy!