Podcasts about intern

  • 5,338PODCASTS
  • 12,356EPISODES
  • 41mAVG DURATION
  • 2DAILY NEW EPISODES
  • Aug 27, 2026LATEST

POPULARITY

20192020202120222023202420252026

Categories




Best podcasts about intern

Show all podcasts related to intern

Latest podcast episodes about intern

Your Money. Your Mission.
Intern Takeover: Money and Career Advice for College Students

Your Money. Your Mission.

Play Episode Listen Later Aug 27, 2026 30:45


In this episode, CEO Jim Popp turns the tables and hands the mic to three of Johnson Financial Group's summer interns, Joey Flaherty, Will Cardwell and Alyssa Ray, who just wrapped eight weeks rotating across wealth, treasury and operations. Together they discuss the gap between what college teaches and what you actually need to know about money, why the best mentors are the ones you stumble into naturally and how human connection stays a competitive advantage in an AI-powered world. Jim also opens up about the career and financial advice he wishes he'd had in his twenties. In this episode:00:00 – 07:54: Intro and What They've Learned07:55 – 16:48: Finance and Career Advice Q&A with Jim16:49 – 23:06: The Intern's Perspectives on the Future of Work and AI23:07 – 25:29: How to Get Your First Job25:30 – 26:47: The Importance of Mentorship26:48 – 30:45: Final Takeaways and Advice for College Students Additional resources:Young Professionals Hub | Johnson Financial GroupYoung Professionals: Here's How to Manage Your Money | Johnson Financial GroupMoney Mindset: Building a Healthy Financial Foundation | Johnson Financial GroupThe Retirement Roadmap: A Guide for Each Decade | Johnson Financial Group

URMIA Matters
URMIA's Be the Change Scholarship for an Intern

URMIA Matters

Play Episode Listen Later Aug 27, 2026 32:57


In this episode of URMIA Matters, host Caitlin Cai sits down with Ashley Caldwell of Gannon University and former student intern Tomoki Kouchi to explore the impact of URMIA's Be the Change Student Internship program. Together, they share how the program created a meaningful, 13-week risk management internship with funding from URMIA's Be the Change Internship program. They discuss how that allowed Tomo to help develop a campus risk assessment tool and business continuity planning resources. Listeners will gain insight into the benefits of hosting a student intern, and why investing in the next generation of risk management professionals benefits both students and higher education institutions alike. We'll talk about how the internship application works, what the expectations are, what Tomoki and Ashley's experience was like, and what's happened since the internship ended.Show NotesURMIA ScholarshipsURMIA's Be the Change Scholarship – Student InternshipGuests Ashley Caldwell, Assistant to the Vice President for Finance and Campus Operations and Risk Management and Insurance Coordinator - Gannon UniversityTomoki Kouchi, Student/Former Intern – Gannon UniversityGuest HostCaitlin Cai, Risk and Insurance Program Manager - University of Tennessee System Connect with URMIA & URMIA with your network-Share /Tag in Social Media @urmianetwork-Not a member? Join ->www.urmia.org/join-Email | contactus@urmia.org Give URMIA Matters a boost:-Give the podcast a 5 star rating-Share the podcast - click that button!-Follow on your podcast platform - don't miss an episode!Thanks for listening to URMIA Matters!

Talklaunch with Ryan Estes
Haleigh & Maggie from Denver&CO - Is MENver Gone For Good? ++ Denver Food & Wine Week!

Talklaunch with Ryan Estes

Play Episode Listen Later Aug 26, 2026 46:05


Haleigh and Maggie join us from Denver&CO, a woman-owned and Colorado rooted events, marketing, & PR company that's behind some of the city's hottest ticket events! They're establishing their business as Denver's premiere community-first creative and marketing partner.   As always, we're also going over the best news and events on our radar this week as well.   We're looking for an Intern! Reach out to tell us how you can help!   Check out our new Community Events Page: https://realgooddenver.com/events   Follow RGD: YouTube: https://www.youtube.com/channel/UC8u8GmvBi6th6LOOMCuwJKw Instagram: https://www.instagram.com/real_good_denver/ TikTok: https://www.tiktok.com/@realgooddenver   Do you have a Denver event, cause, opening, or recommendation that you want to share with us? We want to hear from you! Tell us what's good at tom@kitcaster.com. We're opening up early access to a custom Denver job alert program through our newsletter thanks to https://www.jobstreamai.com/. Sign up at realgooddenver.com to be the first to know when it's ready!! Check out our new Community Events Page: https://realgooddenver.com/events   Guests Haleigh Watts & Maggie Berra from Denver&Co   Events Red Rocks Schedule Green Flags Only Denver Food and Wine Festival   Check out our new Best of Denver Series ​​Music produced by Troy Higgins Goodboytroy.com

Felger & Massarotti
Off-Air Show with Felger, Big Jim and Intern Aurora // Felger's old Boston Herald articles // August 25, 2026

Felger & Massarotti

Play Episode Listen Later Aug 25, 2026 42:45


On this week's Off-Air Show, we say goodbye to Intern Aurora, react to Felger's old Boston Herald stories and more!See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

Todd Durkin IMPACT Show
Meet My 1st Intern Class at IXP: Asking the Questions You've Always Wanted to Ask | Ep. 493

Todd Durkin IMPACT Show

Play Episode Listen Later Aug 24, 2026 62:56


This is a conversation with a special group of young professionals that I really enjoyed working with. This week on The IMPACT SHOW, I sat down with the very FIRST class of interns at IMPACT-X Performance—Briseyda, Ernesto, and Jacob—and let them fire away with whatever questions they had. And MAN… they brought the heat! From building a world-class culture and creating exceptional customer experiences, to finding your purpose, learning from mistakes, balancing family and business, pursuing a career in athletic performance, and even navigating AI and the future of coaching, this conversation went DEEP. These 3 interns have been part of history at IMPACT-X Performance as they've been with me almost since the beginning. I was fired-up to answer their questions and share some of the lessons and wisdom that came out of this experience. In This Episode, You'll Hear: Why culture is everything — and how you build a culture that people can FEEL the moment they walk through the doors. The truth about work-life balance — why I don't necessarily believe in "balance" and instead think about life as an orchestra that requires constant adjustment. The biggest mistakes I've made in nearly 30 years in the fitness industry — including trusting your gut, making decisions based solely on money, trying to do too much, and failing to surround yourself with the right people. Why your passion and purpose matter — and why financial success should be an outcome of doing meaningful work, not the only reason you do it. The power of mentorship — why you NEVER outgrow having a mentor and how the right people can help guide your path and accelerate your growth. Creating an exceptional customer experience — why it's never about you, your accolades, or your credentials… it's about THEM. What it takes to become a great athletic performance coach — from starting with youth athletes to mastering communication, building experience, and learning to speak the language of the sport. Why communication may be your most important coaching skill — and how learning to communicate with a 10-year-old can make you a better coach at every level. The difference between being a "life coach" and actually BEING a coach — and why experience, education, emotional intelligence, and a deep commitment to helping people matter. How physical training can evolve into life coaching — and why changing the body and changing the mind are so deeply connected. What I look for when building a team — credentials matter, but HEART, character, people skills, work ethic, and a genuine desire to change lives matter even more. Why you need to build a CAREER, not just get a job — and the vision behind creating opportunities that can take someone from intern to trainer, leader, business owner, and beyond. How I recharge and stay grounded — including "mellow yellow time," getting outside, breathwork, prayer, and remembering WHY I'm doing what I'm doing. The role of technology and AI in the fitness industry — and why technology should enhance the human experience, never replace the heart and soul behind it. The one word I want to be remembered for — IMPACT. What my younger self would tell me today — and why believing, serving, and continuing to be a difference-maker matters through every season of life. This episode is a reminder that success isn't built overnight. It's built through experience, mistakes, mentors, relationships, relentless learning, and a commitment to serving people at the highest level possible. Whether you're a young trainer just getting started, a seasoned entrepreneur, a coach, a parent, or simply someone trying to figure out your next chapter, there's something in this conversation for YOU. My challenge is simple: get your mind right, find your purpose, surround yourself with great people, keep learning, and never stop looking for ways to make an IMPACT. If this episode speaks to you, don't keep it to yourself! Share it with a coach, trainer, entrepreneur, friend, or someone who needs to hear these lessons. LIKE, SUBSCRIBE, and leave a REVIEW wherever you listen to podcasts, and let me know your biggest takeaway from the episode. And if you've got a question you want me to answer on a future Q&A episode, DM me and send it my way! Please tag me on IG at: @ToddDurkin Don't forget to subscribe to my TEXT COMMUNITY so that you can receive on-going motivational texts from Todd Durkin. You can sign-up for free by texting him at 619.304.2216. **** Speakers Course Would you like to become a more polished public speaker? Would you like to become a keynote speaker locally, nationally, or even internationally? Would you like to just become a more effective communicator to get your message across even more effectively? If the answer is YES to any of those questions, NOW is the time to take part in the Todd Durkin IMPACT Speaker Course…and it all starts on September 8th!!! To find out all about it or REGISTER NOW, visit… www.ToddDurkin.com

The Megyn Kelly Show
Chandra Levy - The Missing Intern and the Congressman: The FULL MK Confidential Series

The Megyn Kelly Show

Play Episode Listen Later Aug 23, 2026 133:24


Megyn Kelly brings you the full five episode Chandra Levy "MK Confidential" series, "The Missing Intern and the Congressman." What do you think happened? Tune in to a new edition of "MK Confidential" weeknights this summer on the Megyn Kelly podcast feeds and YouTube channel.     Follow The Megyn Kelly Show on all social platforms: YouTube: https://www.youtube.com/MegynKelly Twitter: http://Twitter.com/MegynKellyShow Instagram: http://Instagram.com/MegynKellyShow Facebook: http://Facebook.com/MegynKellyShow Find out more information at:https://www.devilmaycaremedia.com/megynkellyshow Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

Handelsblatt Crime - spannende Wirtschaftskriminalfälle unserer Zeit
Vorkasse, Prominente, Pleite – wie Teldafax jahrelang durchkam

Handelsblatt Crime - spannende Wirtschaftskriminalfälle unserer Zeit

Play Episode Listen Later Aug 23, 2026 71:11


Rudi Völler und Thomas Gottschalk warben für den Stromanbieter. Intern wusste der Vorstand längst, dass das Unternehmen zahlungsunfähig war. Wie ein System funktionierte, das auf Täuschung aufgebaut war.

Karson & Kennedy
Life Advice for Intern Evelyn

Karson & Kennedy

Play Episode Listen Later Aug 21, 2026 3:48


Life Advice for Intern Evelyn full 228 Fri, 21 Aug 2026 12:42:28 +0000 pLOMfoNYzD1mFmqjLl4kowLWDNGSaJjV latest,wbmx,society & culture Karson & Kennedy latest,wbmx,society & culture Life Advice for Intern Evelyn Karson & Kennedy are honest and open about the most intimate details of their personal lives. The show is fast paced and will have you laughing until it hurts one minute and then wiping tears away from your eyes the next. Some of K&K’s most popular features are Can’t Beat Kennedy, What Did Barrett Say, and The Dirty on the 30! 2024 © 2021 Audacy, Inc. Society & Culture https://player.amperwavepodcasting.com?feed-link=https%3A%2F%2Frs

Karson & Kennedy
Intern Evelyn Exit Interview

Karson & Kennedy

Play Episode Listen Later Aug 21, 2026 4:34


Intern Evelyn Exit Interview full 274 Fri, 21 Aug 2026 12:46:29 +0000 0KNnQ6IaVed3hpGfXCoM3et5WZte8Tk1 latest,wbmx,society & culture Karson & Kennedy latest,wbmx,society & culture Intern Evelyn Exit Interview Karson & Kennedy are honest and open about the most intimate details of their personal lives. The show is fast paced and will have you laughing until it hurts one minute and then wiping tears away from your eyes the next. Some of K&K’s most popular features are Can’t Beat Kennedy, What Did Barrett Say, and The Dirty on the 30! 2024 © 2021 Audacy, Inc. Society & Culture https://player.amperwavepodcasting.com?feed-link=https%3A%2F%2Frss

Karson & Kennedy
K&K Full Show - Intern Evelyn's Last Day and Kennedy's Impossible Parody! 08-21-26

Karson & Kennedy

Play Episode Listen Later Aug 21, 2026 64:11


K&K Full Show - Intern Evelyn's Last Day and Kennedy's Impossible Parody! 08-21-26 full 3851 Fri, 21 Aug 2026 14:10:33 +0000 sGthnCunySrddXSNpAlnpwMZQzLvUq2N society & culture Karson & Kennedy society & culture K&K Full Show - Intern Evelyn's Last Day and Kennedy's Impossible Parody! 08-21-26 Karson & Kennedy are honest and open about the most intimate details of their personal lives. The show is fast paced and will have you laughing until it hurts one minute and then wiping tears away from your eyes the next. Some of K&K’s most popular features are Can’t Beat Kennedy, What Did Barrett Say, and The Dirty on the 30! 2024 © 2021 Audacy, Inc. Society & Culture https://play

Couch and The Rube
Ep. 1165: Twitter Question Friday: Dialing in on MSU football and goodbye to Nolan the Intern

Couch and The Rube

Play Episode Listen Later Aug 21, 2026 130:26 Transcription Available


We answered your questions — on Michigan State football as the season nears, on new MSU AD Dan Bartholomae and how that'll change things in East Lansing, on Jaxon Kohler going to BYU, on the Tigers, the Lions, life and more, including a farewell to Nolan the Intern.

VPM Daily Newscast
8/21/26 - How are Virginia agencies using tech to hire workers?

VPM Daily Newscast

Play Episode Listen Later Aug 21, 2026 5:32


Read more from VPM News:  Democratic National Committee taps Virginia for early 2028 presidential primary⁠  How do Virginia agencies use tech to hire workers?  WATCH: Amending Virginia    Other links:  Records: public housing staffers accused of trading maintenance work for sexual favors (Richmond Times-Dispatch)*  Ellwood Thompson's parent company acquired by data center owner (The Richmonder)  VCU suspends fraternity after hazing, assault allegations (WRIC)  LPD pauses operation of Flock license plate reader cameras (The News & Advance)*  Intern at Virginia seminary finds original draft of letter by Martin Luther King Jr. (WTOP)  *This outlet uses a paywall.  Our award-winning work is made possible with your donations. Visit vpm.org/donate to support local journalism.

Speakernomics
Treat AI Like An Intern

Speakernomics

Play Episode Listen Later Aug 18, 2026 33:45


Unlock the secrets to advancing your speaking business with proven AI strategies and genuine relationship-building tactics. This episode delivers practical insights from the intersection of technology and human connections.* Why AI can't replace authentic sales relationships and where it truly adds value* How to use generative AI, Google, and LinkedIn for actionable research and outreach* Techniques for effective prompt engineering and maximizing AI-driven personalization* Free and essential research tools every speaker can use to stand out and win more gigs* The power of speaker referrals, building industry connections, and sustaining long-term successBecome an NSA Member! https://nsaspeaker.org/join/#membership Learn more about your ad choices. Visit megaphone.fm/adchoices

Talklaunch with Ryan Estes
Chicken Salad, Modern Day Camping, & The Future of Broncos Tailgating

Talklaunch with Ryan Estes

Play Episode Listen Later Aug 18, 2026 40:38


Tom is fresh off a mountain camping trip straight out of the future. We're taking a look at more renderings from the new Broncos Stadium project at Burnham Yards, and diving into the value of a good chicken salad.   As always, we're also going over the best news and events on our radar this week as well.   We're looking for an Intern! Reach out to tell us how you can help!   Check out our new Community Events Page: https://realgooddenver.com/events   Follow RGD: YouTube: https://www.youtube.com/channel/UC8u8GmvBi6th6LOOMCuwJKw Instagram: https://www.instagram.com/real_good_denver/ TikTok: https://www.tiktok.com/@realgooddenver   Do you have a Denver event, cause, opening, or recommendation that you want to share with us? We want to hear from you! Tell us what's good at tom@kitcaster.com. We're opening up early access to a custom Denver job alert program through our newsletter thanks to https://www.jobstreamai.com/. Sign up at realgooddenver.com to be the first to know when it's ready!! Check out our new Community Events Page: https://realgooddenver.com/events   News Broncos reveal new tailgating vibe Updated Burnham Yard Master Plan Cirrus Social Club   Events Red Rocks Schedule Harry Potter™: The Exhibition   Shoutouts Heidi's Brooklyn Deli Jackery 1500 Chicken Salad Chick Den Thai Riot BBQ ESP HiFi   Check out our new Best of Denver Series ​​Music produced by Troy Higgins Goodboytroy.com

Karson & Kennedy
Intern Evelyn Had A Seagull Land On Her Head

Karson & Kennedy

Play Episode Listen Later Aug 17, 2026 4:52


Intern Evelyn Had A Seagull Land On Her Head full 292 Mon, 17 Aug 2026 12:58:41 +0000 iGS51hcxcApujqc5eLFHfUwu4spiBzGH latest,wbmx,society & culture Karson & Kennedy latest,wbmx,society & culture Intern Evelyn Had A Seagull Land On Her Head Karson & Kennedy are honest and open about the most intimate details of their personal lives. The show is fast paced and will have you laughing until it hurts one minute and then wiping tears away from your eyes the next. Some of K&K’s most popular features are Can’t Beat Kennedy, What Did Barrett Say, and The Dirty on the 30! 2024 © 2021 Audacy, Inc. Society & Culture https://player.amperwavepodcasting.com?feed-link=h

The EdUp Experience
Why Every Student Must Intern - with Dr. Beth Ross, President, Emmanuel College

The EdUp Experience

Play Episode Listen Later Aug 16, 2026 51:36


It's YOUR time to #EdUp with Dr. Beth Ross, President, Emmanuel CollegeThis episode, President Series #499, is powered by ⁠⁠⁠Ellucian⁠⁠⁠ -Higher Ed's AI-Enriched Platform!This episode is sponsored by the InsightsEDU 2027 Conference - focused on attracting & retaining the modern learner - February 23-25 in Phoenix, AZ! Use code EDUP & save $100! Early-bird pricing ends December 15, 2026!This episode is brought to YOU by ⁠EdUp Leadership⁠⁠ - the only intelligence platform built exclusively from presidential conversations in higher ed!YOUR cohost is ⁠Jamie Ceman, Senior Executive VP of Reputation Services, EducationDynamicsYOUR host is ⁠⁠Dr. Joe Sallustio⁠⁠Could AI be the catalyst that finally proves the value of the liberal arts?How does a work study student who walked into the wrong office become a college president?Why does Emmanuel require every student, philosophy or biology, to complete an internship?Thank YOU so much for tuning in. Join us on the next episode for YOUR time to EdUp!Connect with YOUR EdUp Team - ⁠⁠⁠⁠⁠ ⁠⁠⁠⁠Elvin Freytes⁠⁠⁠⁠⁠⁠⁠⁠⁠ & ⁠⁠⁠⁠⁠⁠⁠⁠⁠ Dr. Joe Sallustio⁠⁠⁠⁠● Join YOUR EdUp community at The EdUp ExperienceWe make education YOUR business!P.S. Want access to the only intelligence platform built exclusively from presidential conversations in higher ed? Well, we have an app for that!Join EdUp Leadership!

The Best One Yet

Louis Vuitton & Porsche built a $2M car… Only two people will buy it, but *you* can own it.You've heard of FOMO, but what about FOMOOCH?... Fear of Missing Out On Chips.Interns graduated to dealmakers this summer… we got numbers to prove it.Plus, we're taking vacation for two weeks, but are teasing a new series right here next week.$POAHY $LVMUY $NVDAGrab your Tickets to the IPO Tour: Our In-Person OfferingSan Francisco 9/23: https://www.ticketmaster.com/event/1C0064AFB5F688BDBoston 10/14: https://tickets.citywinery.com/event/tboy-the-ipo-tour-in-person-offering-8cdhupSeattle 11/4 (21+): https://www.axs.com/events/1446394/the-best-one-yet-ticketsNEWSLETTER:https://tboypod.com/newsletter OUR 2ND SHOW:Want more business storytelling from us? Check our weekly deepdive show, The Best Idea Yet: The untold origin story of the products you're obsessed with. Listen for free to The Best Idea Yet: https://wondery.com/links/the-best-idea-yet/NEW LISTENERSFill out our 2 minute survey: https://qualtricsxm88y5r986q.qualtrics.com/jfe/form/SV_dp1FDYiJgt6lHy6GET ON THE POD: Submit a shoutout or fact: https://tboypod.com/shoutouts SOCIALS:Instagram: https://www.instagram.com/tboypod TikTok: https://www.tiktok.com/@tboypodYouTube: https://www.youtube.com/@tboypod Linkedin (Nick): https://www.linkedin.com/in/nicolas-martell/Linkedin (Jack): https://www.linkedin.com/in/jack-crivici-kramer/Anything else: https://tboypod.com/ About Us: The daily pop-biz news show making today's top stories your business. Formerly known as Robinhood Snacks, The Best One Yet is hosted by Jack Crivici-Kramer & Nick Martell. Hosted on Acast. See acast.com/privacy for more information.

The Evening Edge with Todd
The Evening Edge with Todd Hollst 8.14.2026

The Evening Edge with Todd

Play Episode Listen Later Aug 14, 2026 59:29


Saturday Street Takeover in Dayton?; Downtown Hotel Shuts down; Idiot who gave 10-year old girl a lighter; Tupac's death car for sale; Disney Adults are a different breed; Nimrods; Intern Will's Punchline Report; Evening Edge Local Music Showcase featuring American Whale and their singe, "Idiot of the Cosmos.

The Drive with Lon Tay & Derek Piper
08/13/26 Hour 2: Cardinals-Cubs Weekend Series, Illini in the NFL Preseason & A Toast to Quinntern

The Drive with Lon Tay & Derek Piper

Play Episode Listen Later Aug 13, 2026 53:28


Hour 2 of The Drive gets you ready for a big weekend of baseball as the St. Louis Cardinals head to Wrigley Field to take on the Chicago Cubs in a three-game series beginning Friday, August 14. The rivals meet with both teams looking to make a push in the NL Central, setting the stage for another heated weekend at Wrigley. The guys also check in on former Illinois football players making their mark in NFL preseason action. With preseason games underway, which Illini have the best opportunity to make an NFL roster, and who could be fighting for a bigger role? And, unfortunately, it's time for a Friday Toast with a twist: the crew raises a glass to Quinntern the Intern, who is wrapping up her final day with The Drive. Plenty of laughs, memories and well wishes as the team says goodbye. Follow The Drive on X, Instagram, and Facebook.

Karson & Kennedy
Intern Evelyn's Dad Will Die On This Hill

Karson & Kennedy

Play Episode Listen Later Aug 12, 2026 4:00


Intern Evelyn's Dad Will Die On This Hill full 240 Wed, 12 Aug 2026 13:09:10 +0000 xtjDMCuqw2koJNneP7eJ6A4TlD8d0ysy latest,wbmx,society & culture Karson & Kennedy latest,wbmx,society & culture Intern Evelyn's Dad Will Die On This Hill Karson & Kennedy are honest and open about the most intimate details of their personal lives. The show is fast paced and will have you laughing until it hurts one minute and then wiping tears away from your eyes the next. Some of K&K’s most popular features are Can’t Beat Kennedy, What Did Barrett Say, and The Dirty on the 30! 2024 © 2021 Audacy, Inc. Society & Culture https://player.amperwavepodcasting.com?feed-link=http

The Sound of Ideas
American Red Cross declares nationwide crisis, Northeast Ohio calls for donors

The Sound of Ideas

Play Episode Listen Later Aug 12, 2026 51:41


American Red Cross declares blood crisis The American Red Cross has declared a national blood crisis, only the second time in its history it has taken that step. The first was in January 2022 as a result of the COVID-19 pandemic. This time, blood donations have fallen to a four-year summer low, and demand remains elevated. The shortage is especially urgent for type O blood, as supplies of O positive fell below a one-day supply in late July. Wednesday on the "Sound of Ideas" hosted by Stephanie Haney, we'll look at what's driving this shortage, how it's affecting hospitals and patients, and how you can help rebuild the nation's blood supply. Guest:- Christina Peters, Northern Ohio Regional Communications Director, American Red Cross Global HIV/AIDS research continues despite cuts to U.S. funding Later in the hour, a conversation about HIV/AIDS funding and the latest in research and treatment options. In 2019, during his State of the Union address, President Donald Trump called for the elimination of HIV transmission in the U.S. by 2030. Two years later, the United Nations adopted a similar declaration globally. But in 2026, those goals are facing significant challenges. In March 2025, the Trump administration terminated a federal advisory committee on HIV prevention and treatment. By the end of the year, global government funding to low and middle-income countries fell 25% compared with 2024, according to KFF. This February, the administration rescinded $600 million in Centers for Disease Control grants supporting HIV prevention and surveillance programs. Then in April, Trump announced plans to remove all members of the HIV/AIDS Presidential Advisory Council. Those changes came shortly after the U.S. Department of Health and Human Services laid off more than 10,000 federal health employees, including staff working on infectious disease and HIV/AIDS. The U.S. set out to reduce new HIV infections by 75% by 2025, which would have brought the number of new cases down to about 9,300. But the latest CDC data from 2024 puts the number of cases closer to 39,000. So, is the 2030 goal line still within reach? The future is uncertain. Guest:- Asia Russell, Executive Director, Health GAP (Global Access Project) Ideastream's Summer Interns Last week, we said goodbye to our 2026 cohort of interns here at Ideastream Public Media. Throughout the summer, our nine Cleveland interns learned about their respective fields and honed their crafts throughout our multiple departments – including news, marketing and radio broadcast. To learn more about these interns– what they accomplished here, their career aspirations and what they took away from this experience – we brought them on the show. We heard from four interns last week, and on Wednesday we will hear from the remaining five. Guests:- Grace Claxon, News Intern, Ideastream Public Media- Benjamin Giesen, WCLV Intern, Ideastream Public Media- Kiera McGuire, Sound of Ideas Intern, Ideastream Public Media- Sonya Suri, Digital Content Intern, Ideastream Public Media- TJ Thomas, Intern from University School, Ideastream Public Media

The MM+M Podcast
MM+M summer intern Lola Offenback's (second) exit interview

The MM+M Podcast

Play Episode Listen Later Aug 12, 2026 40:21


Lola Offenback, MM+M's editorial fellow, is finishing her second summer working with us.  To quote Herman's Hermits: “Second verse, same as the first.” Lola has once again exceeded our expectations with her versatility and resourcefulness as a reporter. She returned to MM+M just before Cannes and helped us manage the wave of news coming out of the French Riviera. From there, she took on deeper reporting assignments exploring women's health trends, medical marketing internship programs and allergy-focused ad campaigns.  She also got to cover some in-person activations around New York and hosted a two-part episode of the podcast about health brands advertising during the FIFA World Cup ahead of the final match. We're thrilled to have one more sit-down conversation with Lola to discuss how the eight-week internship went, what she made of her second year covering the medical marketing industry and what comes next for her career. For our Trends segment, we're talking about the proposed changes to the vaccine schedule and the policy messaging from Trump administration health officials that is untethered from the medical establishment.  Check us out at: mmm-online.com Follow us: YouTube: @MMM-onlineTikTok: @MMMnewsInstagram: @MMMnewsonlineTwitter/X: @MMMnewsLinkedIn: MM+M To read more of the most timely, balanced and original reporting in medical marketing, subscribe here.Music: “Deep Reflection” by DP and Triple Scoop Music. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

The Cook & Joe Show
11AM - Nick Farabaugh says this is important week for Howard, Allar taking steps; Intern Dom joins the show to discuss Pitt-Penn State, McCarthy's Steelers

The Cook & Joe Show

Play Episode Listen Later Aug 11, 2026 43:21


Hour 2 with Joe Starkey: Max Iheanachor won't play Thursday. It seems the Steelers will figure out the QB groups in three ways, which means Aaron Rodgers might not play against Green Bay. Michael Pittman has a right leg injury and Nick thinks he could be back next week. DK Metcalf hit his shoulder hard at camp but Nick doesn't think it's serious. Nick thinks Drew Allar will play close to the full second half. Unless Will Howard has a big preseason, Mason Rudolph is likely going to be the backup. Intern Dom joined the show. We have a heated Drew Allar debate.

The Cook & Joe Show
Intern Dom joins the show to discuss Pitt-Penn State, McCarthy's Steelers

The Cook & Joe Show

Play Episode Listen Later Aug 11, 2026 18:20


Penn State head coach Matt Campbell went to Pitt for a year. Intern Dom joined the show. We have a heated Drew Allar debate. Dom wonders if Aaron Rodgers wants a trade in the middle of the season. 

Little Gym, Big Heart with Devin Gage
From Unpaid Intern to Director in 2 Years - Here's the Blueprint

Little Gym, Big Heart with Devin Gage

Play Episode Listen Later Aug 11, 2026 19:09


Discover the exact blueprint for climbing the ladder in the fitness industry, moving from an unpaid intern to a gym facility leader in exactly two years. In this episode, we sit down with Jose Rivera, the first employee at Engaged Personal Training to ascend through every single level of the organization. From an exercise science graduate unsure of his career path to taking the keys to our newest location, Jose shares the mindsets, strategies, and actions that made him entirely undeniable and highly promotable to ownership. Whether you are a gym owner looking to develop rockstar talent or a personal trainer eager to build a sustainable, high-growth career, this interview breaks down what it truly takes to succeed.

LSI Behind the Win
Intern Edition

LSI Behind the Win

Play Episode Listen Later Aug 11, 2026 24:54


Meet LSI Interns Carley, Chloe, and Batu as we discuss how they ended up at LSI and what they're getting from the experience. At LSI, they've found growth, interesting projects, hands-on experience, insight into government relations, and exposure to the interconnectedness of business and government. In this episode, learn more about the inner workings of LSI and our interns.

Holmberg's Morning Sickness
08-10-26 - Paying Tribute To Legendary PHX Radio DJ Bone Mama Who Died Over Wknd- Entertainment Drill - MON - More Stories Of Bone Mama And John's Blind Radio Intern Jason

Holmberg's Morning Sickness

Play Episode Listen Later Aug 10, 2026 22:14


Link Up w/The Morning Sickness Digitally All Over:Instagram: @hms_98_official, @bosskupd, @bretvesely, @dickToledoX/Twitter: @HMSon98, @DickToledo, @bretveselyFacebook: @HMSKUPDYouTube: @hmspodcast9320, @98kupdRequest/Call in/Wakeup Song line:(IN AZ) 602.585.9800More HMS: www.holmbergpodcast.com, www.98kupd.comEmail: dtoledo@98kupd.com, bvesely@98kupd.com, bbogen@98kupd.comSee Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

Fred + Angi On Demand
FULL 6 AM: Dentist Scamming & Fred Goes To His Former Intern's Wedding!

Fred + Angi On Demand

Play Episode Listen Later Aug 10, 2026 37:42 Transcription Available


Paulina has 12 cavities and Fred is convinced she might be getting scammed! And Fred went to his former intern's wedding and had a blast!See omnystudio.com/listener for privacy information.

Fred + Angi On Demand
Radio Blogs: Fred Goes To His Former Intern's Wedding!

Fred + Angi On Demand

Play Episode Listen Later Aug 10, 2026 6:31 Transcription Available


Fred went to his former intern's wedding and had a blast!See omnystudio.com/listener for privacy information.

Holmberg's Morning Sickness - Arizona
08-10-26 - Paying Tribute To Legendary PHX Radio DJ Bone Mama Who Died Over Wknd- Entertainment Drill - MON - More Stories Of Bone Mama And John's Blind Radio Intern Jason

Holmberg's Morning Sickness - Arizona

Play Episode Listen Later Aug 10, 2026 22:14


Link Up w/The Morning Sickness Digitally All Over:Instagram: @hms_98_official, @bosskupd, @bretvesely, @dickToledoX/Twitter: @HMSon98, @DickToledo, @bretveselyFacebook: @HMSKUPDYouTube: @hmspodcast9320, @98kupdRequest/Call in/Wakeup Song line:(IN AZ) 602.585.9800More HMS: www.holmbergpodcast.com, www.98kupd.comEmail: dtoledo@98kupd.com, bvesely@98kupd.com, bbogen@98kupd.comSee Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

Host Lucie Výborné
Mikrobiom si tvoříme tím, co jíme, říká gastroenterolog. Proč jsou onemocnění střev na vzestupu?

Host Lucie Výborné

Play Episode Listen Later Aug 8, 2026 28:41


„Když jíte zdravě a užíváte hodně vlákniny, můžete tímto způsobem života změnit svůj mikrobiom a na základě toho snížit riziko vzniku prekanceróz v oblasti zažívacího traktu, ale velmi pravděpodobně i jinde,“ říká přednosta Interní kliniky 2. lékařské fakulty Univerzity Karlovy a Fakultní nemocnice Motol Radan Keil. Jak stres ovlivňuje zažívání? U kterých pacientů se uplatňuje fekální bakterioterapie? A jak složitá je výroba probiotik? Poslechněte si rozhovor.Všechny díly podcastu Host Radiožurnálu můžete pohodlně poslouchat v mobilní aplikaci mujRozhlas pro Android a iOS nebo na webu mujRozhlas.cz.

Večerní Host Radiožurnálu
Mikrobiom si tvoříme tím, co jíme, říká gastroenterolog. Proč jsou onemocnění střev na vzestupu?

Večerní Host Radiožurnálu

Play Episode Listen Later Aug 8, 2026 28:41


„Když jíte zdravě a užíváte hodně vlákniny, můžete tímto způsobem života změnit svůj mikrobiom a na základě toho snížit riziko vzniku prekanceróz v oblasti zažívacího traktu, ale velmi pravděpodobně i jinde,“ říká přednosta Interní kliniky 2. lékařské fakulty Univerzity Karlovy a Fakultní nemocnice Motol Radan Keil. Jak stres ovlivňuje zažívání? U kterých pacientů se uplatňuje fekální bakterioterapie? A jak složitá je výroba probiotik? Poslechněte si rozhovor.Všechny díly podcastu Host Radiožurnálu můžete pohodlně poslouchat v mobilní aplikaci mujRozhlas pro Android a iOS nebo na webu mujRozhlas.cz.

The Megyn Kelly Show
The Missing Intern & The Congressman: The Man in the Park — Ep. 4 | MK Confidential

The Megyn Kelly Show

Play Episode Listen Later Aug 7, 2026 34:37


“MK Confidential" opens on the evening of May 14th, 2001. Thirty-year-old lawyer Halle Shilling ties back her hair, grabs her yellow Walkman, and drives down to Rock Creek Park for a run. In the parking lot near the old Peirce Mill, a young man sits on a curb, watching her start her jog. Minutes later he is behind her and closing, and then he has her off the trail and down in the trees. There is a knife. She screams, and twists, and works her fingernails into the soft flesh under his tongue until he lets go and runs. Halle Shilling saves her own life less than a mile from the spot where, two weeks earlier, Chandra Levy disappeared. Megyn Kelly follows the case as its focus shifts away from the congressman and onto the man in the park. A geographic profiler spreads a map of Rock Creek Park across a table and marks every attack on a lone woman in the spring of 2001. They cluster on one isolated stretch of trail. Then they stop — all at once, the week a nineteen-year-old day laborer named Ingmar Guandique is arrested on those same trails and taken to jail. He is already serving ten years for attacking Halle Shilling and another runner, Christy Wiegand. No one working the Levy case has ever seriously looked at Guandique. Detectives fly to a California prison and bluff that they have his DNA on Chandra's things. He does not say he never met her. He says: so what if I touched her. Prison guards find a photograph of Chandra Levy in Guandique's cell, clipped from a magazine, kept beside his bed. It is still not enough to charge a man with murder. What finally is enough is a cellmate named Armando Morales, who says Guandique confessed to murdering Chandra. Morales tells the jury he has never been an informant in his life; the jury says Guandique is “guilty.” Neither of those statements will stand the test of time. This is Episode 4 of the Chandra Levy story. Home Title Lock: Go to https://hometitlelock.com/megyn and use promo code MEGYN to get a FREE title history report and a FREE TRIAL of their Triple Lock Protection Supersure Insurance: Upgrade your business insurance to a year-round SuperAgency at https://Supersure.com/Megyn   Follow The Megyn Kelly Show on all social platforms:   YouTube: https://www.youtube.com/MegynKelly   Twitter: https://Twitter.com/MegynKellyShow   Instagram: https://Instagram.com/MegynKellyShow   Facebook: https://Facebook.com/MegynKellyShow   
Find out more information at:
 https://www.devilmaycaremedia.com/megynkellyshow Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

The Megyn Kelly Show
The Missing Intern & The Congressman: Life Sentence — Ep. 5 | MK Confidential

The Megyn Kelly Show

Play Episode Listen Later Aug 7, 2026 32:59


“MK Confidential" opens in a hotel lobby in Annapolis, Maryland, in July of 2016. A sliding glass door closes on a golden retriever named Buddy, and a stranger gathers the animal up and carries him inside. He tells the dog's owner his name is Phoenix. Over the next few days he tells her the rest — twenty years in prison, and still capable of doing horrific things when prompted. She gets frightened enough to start recording their conversations. By the time she stops, she has seven hours of a career criminal named Armando Morales, the witness who sent Ingmar Guandique to prison for sixty years for the murder of Chandra Levy. Megyn Kelly lays out how the murder conviction comes apart. On the stand, Morales swore he had never been an informant. Sitting in the government's own files was a letter proving otherwise — a document the prosecution never turned over, and the defense was entitled to have. A judge vacates the conviction and orders a new trial. Then, the tapes surface, and the government's star witness sounds nothing like the reformed man the jury voted to trust. On July 28th, 2016, prosecutors drop the murder charge. Guandique is deported to El Salvador, and disappears. A dismissal is not an acquittal, and it is not an exoneration. Megyn walks through what the evidence still shows, and what has never fit — the rough horse trail where Chandra's family says she would never have run, the workout shoes still sitting in her apartment, an unidentified man's DNA on her clothing, her missing keys and pinkie ring, never found. Chandra's aunt never believed Guandique killed her; neither did the family's private eye. In time, even Bob and Susan Levy stopped being sure. Twenty-five years after their daughter walked out of her apartment, the murder of Chandra Levy is, in the eyes of the law, unsolved. Her parents find consolation in ladybugs, the source of Chandra's childhood nickname, which find them wherever they go, exploring the world in honor of the daughter who left so much of it unseen. This is Episode 5, and the conclusion of the Chandra Levy story. Home Title Lock: Go to https://hometitlelock.com/megyn and use promo code MEGYN to get a FREE title history report and a FREE TRIAL of their Triple Lock Protection Birch Gold: Text MK to 989898 and get your free info kit on gold. Follow The Megyn Kelly Show on all social platforms: YouTube: https://www.youtube.com/MegynKelly Twitter: https://Twitter.com/MegynKellyShow Instagram: https://Instagram.com/MegynKellyShow Facebook: https://Facebook.com/MegynKellyShow 
Find out more information at:
 https://www.devilmaycaremedia.com/megynkellyshow Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

The Evening Edge with Todd
The Evening Edge with Todd Hollst 8.7.2026

The Evening Edge with Todd

Play Episode Listen Later Aug 7, 2026 56:50


Latest on Buc-ee's Vs. Beaver's Mini Mart; Cars into Buildings update; Sex at the bar?; Would you dig through a stuffed trash truck for $1 Million?; Beer News; Intern Will's Punchline Report; Evening Edge Local Music Showcase and Ramar Estates.

The Megyn Kelly Show
The Missing Intern & The Congressman: Buried in the Park — Ep. 3 | MK Confidential

The Megyn Kelly Show

Play Episode Listen Later Aug 6, 2026 27:27


"MK Confidential" opens on the morning of July 25th, 2001. Twenty-eight police recruits stand shoulder to shoulder along Glover Road in Rock Creek Park, beating the brush, calling Chandra Levy's name. The order from headquarters was to search a hundred yards from the roads and the trails...but most of the team hears roads. By one o'clock it is ninety-one degrees, the search is called off, and the recruits climb onto a bus. Ten months from now, Chandra will be found seventy-nine yards below a trail no one walked that day. Megyn Kelly lays out the summer the investigation came apart. Police fail to retrieve security footage from Chandra's apartment building, and it's overwritten. An untrained sergeant switches on Chandra's laptop and corrupts the drive. And on the trails of that same park, a man is dragging women into ravines. Two of them fight him off and live. When he is finally caught, a detective slides Chandra's missing flier across the table and asks whether he has ever seen her. Yes, he says — once, near the Peirce Mill parking lot. He thought she was pretty. The detective never writes it down. Then, the September 11th attacks take the cameras and the federal agents off of Chandra Levy. The tip line goes quiet. Gary Condit loses his primary election and leaves D.C. And on May 22nd, 2002, a man walking his dog steps off a trail, brushes away leaves, and finds a human skull. Two weeks later, the family's own investigators return to the released crime scene with garden rakes and find more. This is Episode 3 of the Chandra Levy story. Home Title Lock: Go to https://hometitlelock.com/megyn and use promo code MEGYN to get a FREE title history report and a FREE TRIAL of their Triple Lock Protection Birch Gold: Text MK to 989898 and get your free info kit on gold. Follow The Megyn Kelly Show on all social platforms: YouTube: https://www.youtube.com/MegynKelly Twitter: https://Twitter.com/MegynKellyShow Instagram: https://Instagram.com/MegynKellyShow Facebook: https://Facebook.com/MegynKellyShow 
Find out more information at:
 https://www.devilmaycaremedia.com/megynkellyshow Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

Industry Standard w/ Barry Katz
Zoe Friedman: Intern to Discovering New Talent & How Comedians Got on Letterman

Industry Standard w/ Barry Katz

Play Episode Listen Later Aug 6, 2026 69:41


Zoe Friedman has spent decades helping shape the comedy industry from behind the scenes. From discovering and developing stand up talent to booking for The Late Show with David Letterman, leading talent development at Comedy Central, and working with Jimmy Kimmel Live! Her influence has reached every level of comedy. Today, as the co-founder and CEO of Comedy Gives Back, she continues to advocate for comedians by providing resources, support, and a stronger foundation for the comedy community.

The Megyn Kelly Show
The Missing Intern & The Congressman: "Did You Kill Chandra Levy?" — Ep. 2 | MK Confidential

The Megyn Kelly Show

Play Episode Listen Later Aug 5, 2026 27:01


"MK Confidential" opens on a phone call. In the early summer of 2001, a reporter at NBC's Washington station picks up, and a police source gives her one sentence: a missing intern, a congressman, and an affair. Within days, nearly every camera in America has swung off the search for Chandra Levy and onto Gary Condit.   Megyn Kelly traces the paths of a 24-year-old from Modesto and a fifty-three-year-old congressman, from an unannounced office visit to a phone number, a private line with soft music on it, and a set of rules Chandra recites to exactly one person she trusts. Get off the elevator. Say you're visiting a sick friend. Never bring your ID.   Then, the story turns. A detective swabs a sitting congressman's cheek in a dark supermarket parking lot. A stranger watches Condit push something deep into a trash can in Alexandria and goes back to see what it was. Twenty-four million people watch him tell Connie Chung he did not kill her, and refuse to say what he did do. A man with something to hide is not the same as a man who knows what happened to Chandra Levy; but in the summer of 2001, the two look identical. This is Episode 2 of the Chandra Levy story.   Home Title Lock: Go to https://hometitlelock.com/megyn and use promo code MEGYN to get a FREE title history report and a FREE TRIAL of their Triple Lock Protection   Supersure Insurance: Upgrade your business insurance to a year-round SuperAgency at https://Supersure.com/Megyn   Follow The Megyn Kelly Show on all social platforms:   YouTube: https://www.youtube.com/MegynKelly   Twitter: https://Twitter.com/MegynKellyShow   Instagram: https://Instagram.com/MegynKellyShow   Facebook: https://Facebook.com/MegynKellyShow   
Find out more information at:
 https://www.devilmaycaremedia.com/megynkellyshow Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

On Our Mark: The Weatherby Podcast
On Our Mark: Episode 152 - Life as a Weatherby Intern: A Summer in Sheridan

On Our Mark: The Weatherby Podcast

Play Episode Listen Later Aug 5, 2026 32:58


Ever wondered what it's really like to spend a summer interning at Weatherby? On this episode of the On Our Mark Podcast, we're joined by Sales, Product Development, and Marketing Intern Mitch Gallion to hear about his experience behind the scenes at Weatherby. From working across multiple departments to the lessons he learned along the way, Mitch shares what it's like to be part of the team during one of the busiest times of the year. We also dive into the 4th Annual Weatherby Film Festival, surviving Wyoming's record-breaking summer heat, Mitch's passion for running, and his experience taking on the legendary Bighorn Trail Run. In this episode we discuss: - What it's like to intern at Weatherby - Lessons and takeaways from the internship - Behind the scenes of the 4th Annual Weatherby Film Festival - Wyoming's record-breaking summer heat - Running and endurance training - The Bighorn Trail Run

The Decibel
Canadian NATO intern accused of spying

The Decibel

Play Episode Listen Later Aug 5, 2026 20:31


A Canadian woman has been arrested and accused of spying for a third country during her internship at NATO's headquarters in Belgium. The woman had a work history with Canadian government agencies and was once found to have committed fraud when trying to apply for a job with the Canada Border Services Agency. This case is now raising questions about how the Canadian government vets its candidates, especially for roles that could have access to sensitive information. Greg Mercer is an investigative journalist with The Globe and one of a team of reporters working on this story. He joins the show to explain what we know so far. Questions? Comments? Ideas? E-mail us at thedecibel@globeandmail.com Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

Indie Mixtape
E173. Assigned Intern at Birth

Indie Mixtape

Play Episode Listen Later Aug 5, 2026 46:28


This time, Ty and HB are putting their noses to the grindstone at their part time jobs! ** Episode theme: Part time jobs Games played in the show: Urban Myth Dissolution Center Perfect Tides: Station to Station Thank Goodness You're Here! Discounty Citizen Sleeper 2: Starward Vector Ty HB Moonshot Network Edited by Wheels Check out our podcast host, Pinecast. Start your own podcast for free with no credit card required. If you decide to upgrade, use coupon code r-12f50f for 40% off for 4 months, and support Indie Mixtape.

The Megyn Kelly Show
The Missing Intern & The Congressman: The Vanishing — Ep. 1 | MK Confidential

The Megyn Kelly Show

Play Episode Listen Later Aug 4, 2026 22:13


"MK Confidential" opens on the morning the biggest missing-persons story in America is set to have its biggest day. It is September 11th, 2001. Bob and Susan Levy are flying to Chicago to tell their daughter's story on Oprah; the congressman's son is booked on The View. For four months the country has asked two questions about a 24-year-old intern named Chandra Levy: where is she, and was a sitting congressman involved? At 8:46 a.m., Flight 11 hits the north tower. The interviews never happen, and by nightfall the story that consumed America is gone. Megyn Kelly goes back four months, to a 24-year-old intern from Modesto who came to Washington wanting a badge and a career at the FBI. On May 1st, 2001, Chandra prices flights home, opens a map of Rock Creek Park's trails, and logs off at 12:24. No one ever hears from her again.  Police find her wallet, her ID, her phone, her laptop still open — only her keys and a ring missing.  Chandra's parents read her phone bills, and one number stops them cold: their own congressman, Gary Condit. Why is he calling their daughter, and where is she? This is Episode 1 of the Chandra Levy story. Home Title Lock: Go to https://hometitlelock.com/megyn and use promo code MEGYN to get a FREE title history report and a FREE TRIAL of their Triple Lock Protection Birch Gold: Text MK to 989898 and get your free info kit on gold. Follow The Megyn Kelly Show on all social platforms: YouTube: https://www.youtube.com/MegynKelly Twitter: https://Twitter.com/MegynKellyShow Instagram: https://Instagram.com/MegynKellyShow Facebook: https://Facebook.com/MegynKellyShow 
Find out more information at:
 https://www.devilmaycaremedia.com/megynkellyshow Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

Finding True Wealth Podcast with Nick Hopwood, CFP
What Our Peak Intern Learned That the Classroom Couldn't Teach

Finding True Wealth Podcast with Nick Hopwood, CFP

Play Episode Listen Later Aug 4, 2026 11:05


Join Preston Gee, CFP® of Peak Wealth Management as he sits down with Peak summer intern Joey Ingram to reflect on an incredible internship experience. As Joey prepares to return to Michigan State University, he shares what he learned during his time at Peak, the meaningful opportunities he was given, the skills he developed, and how this experience will shape his future career. If you're a student interested in finance, investing, internships, wealth management, or career development, this conversation offers valuable insight into what it's like to gain real-world experience in the financial planning industry. Catch up with Joey and the Peak team, and hear firsthand how hands-on learning can make a lasting impact. — ✅ Apply For A Free Retirement Planning Session ✅ peakwm.com/start-here ------------------------------- Peak Wealth Management is a financial planning and wealth management firm in Plymouth, MI. We believe by providing education and guidance, we inspire our clients to make great decisions so they can Retire With Peace of Mind Stay Connected With Us: Podbean YouTube Apple Facebook X Peak Wealth    

Latent Space: The AI Engineer Podcast — CodeGen, Agents, Computer Vision, Data Science, AI UX and all things Software 3.0

Watch the full episode on YouTube:We first covered Baseten last year when DeepSeek mania was at peak hype. Now they have raised a monster $13B round and become one of the new cohort of AI Infra decacorns that are (with Nvidia, Intel, and the semis complex) chief beneficiaries of the Inference Inflection. We return to Baseten at the peak of the 2026 edition of Open Weights debate. Ali has published a viral breakdown of Kimi K3:And since you last saw him, Philip has spoken at AI Engineer and written the definitive book on Inference Engineering spotted all over SF:Three years ago, inference engineering barely existed as a category.Today, it is one of the most critical disciplines in AI. Inference engineering inherently tackles a different question than standard model training: “How do you turn those weights from training into a product that is fast, reliable, and affordable at scale?” Focusing on these creates an entirely new optimization problem.In one recent GLM-5.2 experiment, quantizing more of the model actually preserved its benchmark quality while increasing throughput by 20%, because the errors introduced in different layers could cancel each other out.Inference is no longer just the final step after training. It is becoming its own engineering discipline, with its own research problems, infrastructure, and increasingly specialized roles.In this episode, Baseten's Philip Kiely and Ali Taha join swyx and Vibhu to explain what actually happens after a new open model is released and what it takes to turn “we generated a token” into a fast, reliable, production-ready API.We go deep on cache-aware routing, disaggregated prefill and decode, quantization, speculative decoding, KV-cache movement, model parallelism, GPU kernels, and the race to make frontier models up to 10× faster. Philip and Ali explain why inference optimizations can still produce gains of 20%, 100%, or even 200%; how quantization errors can cancel one another out; why identical weights can behave differently across clusters; and how Baseten grafted a Kimi vision encoder onto GLM-5.2 without changing the underlying language model.The conversation then expands beyond LLMs into NVIDIA Dynamo, mega kernels, Rubin, AI-specific chips, local inference, video generation, diffusion versus autoregressive models, and the enormous compute barrier to generating coherent long-form video. Finally, we explore the convergence of training and inference, continual learning through persistent KV cache, and the emerging loop where models help optimize the infrastructure that runs them.We discuss:* What happens when a 200,000-token request enters an inference system* Cache-aware routing and reusing previously computed KV cache* Why prefill and decode are increasingly handled by different GPUs* When dedicated deployments become cheaper and more reliable than shared APIs* How speculative decoding uses a smaller model to accelerate a larger one* Tool calling, structured outputs, and what LLMs actually do* What it takes to support a new open model on day zero* Grafting Kimi's vision encoder onto GLM-5.2* Retrofitting inefficient model layers with components from other architectures* Why models sometimes collapse into repeating the same token* How hardware, kernels, and race conditions create nondeterministic failures* Preserving model fidelity while making inference faster* How quantization errors can cancel each other out* Why inference optimizations still deliver gains of 20%, 100%, and 200%* How optimized serving can make a model up to 10× faster* NVIDIA Dynamo, KV-aware routing, and distributed model serving* Speculative decoding the speculative decoder* Why local AI is about making models less dumb while data-center AI is about making them less slow* Tensor, expert, and pipeline parallelism across GPUs* Hardware-aware model design, auto-tuning, and the case against mega kernels* Rubin and why inference is becoming a systems problem* Whether modern GPUs are evolving into programmable AI ASICs* Why enormous models like Kimi K3 require GB300-class hardware* Why open-source video generation still trails Veo, Kling, and other closed models* The quadratic attention bottleneck behind long-form AI video* Autoregressive video, real-time generation, and compounding quality drift* Why future video systems may combine autoregressive and diffusion architectures* Training for inference and inference for training* Continuous post-training, deployment, evaluation, and improvement loops* How GLM-5.2 helped optimize the kernels serving GLM-5.2 itself* Why faster networking could unlock dramatically faster decoding* Continual learning, KV-cache compaction, and persistent model memoryShow Notes* How to build a day-0 API for Kimi K3* 22580: From GPT2 to Kimi3, ExplainedPhilip Kiely* LinkedIn: https://www.linkedin.com/in/philipkiely* X: https://x.com/philipkiely* Inference Engineering: https://www.baseten.co/inference-engineering/Ali Taha* LinkedIn: https://www.linkedin.com/in/aliestaha/* X: https://x.com/waterloointernTimestamps00:00:00 Introduction and the 200K-Token Prompt00:03:18 Dedicated Deployments, Speculative Decoding, and Tool Calling00:11:26 Launching Production-Ready Open Models00:19:06 Model Retrofits, Failure Modes, and Nondeterminism00:28:22 Quantization and Canceling Errors00:32:15 The Race to 10× Faster Inference00:40:48 Dynamo, Speculation, and Local vs. Data-Center AI00:50:18 Model Parallelism, Auto-Tuning, and Mega Kernels01:00:55 Rubin, GPUs vs. ASICs, and Custom AI Chips01:10:03 Giant Models and the Limits of GPU Memory01:12:42 AI Video, Quadratic Attention, and Autoregressive Generation01:21:47 Audio, Images, and Diffusion Models01:27:32 Training, Self-Optimizing Models, and Continual Learning01:40:06 Closing ThoughtsTranscriptIntroduction: Baseten, Waterloo Intern, and Inference EngineeringSwyx [00:00:00]: Okay, we're here in the studio with Philip, old friend from Inference Engineering, the book, as well as Baseten and everything that you've done, you and I have done before, as well as Ali. Welcome.Ali [00:00:15]: Pleasure to meet you.Swyx [00:00:15]: Waterloo intern.Ali [00:00:16]: Waterloo intern, always.Swyx [00:00:17]: When did you get “Waterloo intern” as a handle?Ali [00:00:19]: As a handle? Oh.Ali [00:00:20]: I think the rebranding happened mid-March. When I saw it was open, I was like, “I have to take it. Up for grabs.”Philip [00:00:26]: The problem is that Ali is really good at his job and is not gonna be an intern much longer.Philip [00:00:30]: So we have to figure out who's gonna get the handle.Ali [00:00:33]: Well, I'll pass the torch over to the next intern.Swyx [00:00:34]: Oh, okay. It can be, like, you just pass it to another Waterloo grad.Ali [00:00:37]: To another Waterloo intern. No, bruh.Philip [00:00:39]: Yeah.Ali [00:00:39]: Intern.Swyx [00:00:40]: Intern, yeah.Ali [00:00:40]: And no.Philip [00:00:41]: You gotta get an intern from Waterloo.Ali [00:00:42]: Yeah, I've gotta get an intern from Waterloo.Swyx [00:00:44]: Right.Ali [00:00:44]: But they have to follow the path.Swyx [00:00:45]: Oh, it could, but it could come from Baseten, so it's like whoever Baseten gets from Waterloo.Ali [00:00:48]: Right.Swyx [00:00:49]: Has the title of Waterloo.Ali [00:00:50]: It stays in the ecosystem.Philip [00:00:51]: Exactly.Ali [00:00:52]: Halfway through the internship, you either get it or you're out.Philip [00:00:55]: You should also do, like, a big graduation ceremony where you change the handle.Ali [00:00:59]: Just say it.Philip [00:00:59]: For everybody.Swyx [00:01:00]: You guys are good at ceremonies, clearly. We had a nice launch of the book, very successful. But before we get into all that, I wanna start off with a fun question for you. Okay, you're an expert inference engineer. What happens when I send a long query, say two hundred thousand tokens into Baseten's inference? What's the process of query through GPU model routing, balancing, all that? What is all the stuff that we don't think about?Long Context Requests, KV Cache, and Cache-Aware RoutingPhilip [00:01:26]: With a long query specifically, the first thing that I'm gonna ask is, “Have you sent me this query before, or at least part of it?” and I really hope you have, because it's gonna be a lot easier for me and a lot cheaper for you. So the first thing that we're gonna look at is some cache-aware routing, where we're going to see, we probably have a number of instances, a number of replicas up serving whatever model you're hitting. We want to send this one to something with, number one, available prefill workers, and number two, ideally some cached input already there so that we can skip prefill on at least part of these two hundred thousand tokens. If you're doing two hundred thousand tokens, it's probably coding or a multi-turn agent or something where you would expect to have that cached. If you don't, we're gonna have to send it to a prefill worker. We've at least on certain models disaggregated prefill and decode, so you're going to have one set of GPUs that's solely going to process the input, create the KV cache, and get you your first token, and then that's going to be passed over to a separate set of GPUs, which is going to run decode. We're going to iteratively make those tokens. We're probably going to have some speculator model in front of that. I'm going to assume that you're doing coding, and because of that, our speculator model, which assumes you're doing coding, is gonna have a high draft token acceptance rate. If I'm wrong and you're asking me to summarize every Harry Potter book, it's gonna be slower. And then we stream that output to you and account for it, charge you, a couple of pennies and say, “Hey, would you like to send another one?”Swyx [00:03:04]: Except Baseten doesn't charge by pennies.Philip [00:03:07]: Well, yeah, we charge. I'm assuming that we're talking about the public model APIs. If you are setting up a dedicated deployment, then yeah, it's not pennies.Public APIs vs. Dedicated DeploymentsSwyx [00:03:18]: Yeah, one of the key differentiators when I was talking with Baseten initially was that people who want very high volume just need to rent by the box, ‘cause then it's up to you to figure out how to saturate the box.Ali [00:03:31]: And more often than not, it's, like, way cheaper if you're pushing, like, millions of tokens per hour, if you just pay per hour instead of pay per token.Philip [00:03:37]: Yeah, they do. I think that we've increasingly seen a lot of demand for the pay per token APIs, just because everyone wants to try open models, and then once they find a use case that's really sticky, then they move over to dedicated.Swyx [00:03:51]: Is there a best practice on when it's time to swap over?Philip [00:03:54]: Couple reasons. Yeah, reliability, that's a big one, right?Ali [00:03:57]: Like, if they have a very specific use case, they want you to train something specifically for them, like they want their own spec dec, for instance, for their own traffic.Swyx [00:04:04]: Spec dec is speculative decoding.Speculative Decoding and Custom SpeculatorsAli [00:04:05]: Speculative decoding, yeah.Swyx [00:04:07]: You have to explain.Ali [00:04:07]: Sorry. Like, speculative decoding is like, if you have a huge model, right? And so the model is going to be generating one token at a time every single turn, every single forward pass. So we attach, like, this little, like, parasite, like this layer that goes on top of the model, and this model just has to predict. It does three very fast autoregressive forward passes, and it will predict, like, three certain tokens, and then you do one forward stage over the entire original model in order to see if those predictions were correct or not, and then you accept them or you reject them. Now, this draft model is traffic specific, so if you, like, Philip said, if you're summarizing Harry Potter books, I can train exclusively that draft model on Harry Potter books, and I can guarantee you that I'm gonna accept the three tokens every single time. And so with that case, I increase your decode speed. I wouldn't be able to provide this to you if you're a shared endpointSwyx [00:04:53]: YeahAli [00:04:53]: ‘cause I have no idea if you're doing Harry Potter, if you're doing coding, if you're doing English. We don't know. Also, there was a thing in the book that mentioned that if they really cared about a specific threshold, chapter four, I think. Do you remember that?Philip [00:05:06]: Yeah. The things that you can do is you can set a specific, like, batch sizing, a specific, like, parallelism strategy if you're trying to optimize for, like, throughput versus latency. You can. Maybe a NVFP4 quant doesn't pass your benchmarks and you wanna run a model at higher precision, you could do that. There's just a bunch of reasons why you might wanna have your own endpoint and the biggest one, of course, just being, like, you don't have to deal with someone else doing a hundred million tokens of benchmarking traffic at the endpoint when you happen to be trying to serve your users.Swyx [00:05:40]: Yeah. I think one thing that is. That is a classic journey. Like, it's people is asking the, what happens when you type Google into the browser. Tool calling, is that just, you're generating JSON or is there more complication beyond that?Tool Calling, JSON, and Structured OutputsAli [00:05:58]: Certain customers that we have, they have their own post-trained models, and so they demand a tool calling that's not just, like parse a file or go find the weather. It's something that's very specific and you have to do post-training on this. And if the post-training on the model is not good or if the quantization after the post-training to get the inference to be fast, the model will struggle reading the JSON file and reading the tool calling. But it doesn't require its own like sandbox. It's not like it's going to use that tool calling to like escape a sandbox or like it doesn't have to be contained. It can just be a normal dedicated deployment. The challenge with tool calling more and more seems to be that the companies want certain tool calling which is a very sensitive thing to train. And because you're dealing with all of the JSON outputs, if it doesn't like close the end of the request in a very certain manner, you end up with a model that did the tool calling and like the thinking and so as a result of that, it didn't see the result and just hallucinated the result as it decoded. That seems to be the most challenging thing with tool calling, not really the sandboxes model.Philip [00:06:56]: Yeah, that's a challenge on the training side and then on the inference side, there's work that you can do to scope the possible output. So we published this at this point close to two years ago, the solution to this problem which is you make a state machine and you use that to constrain the output to a specific format. So this is the structured output problem. If you remember backSwyx [00:07:27]: Yeah, the specific grammar is,Philip [00:07:29]: Yeah, exactlySwyx [00:07:30]: GML had this thing.Philip [00:07:31]: Yeah. So it's like the old-school “make sure this is only JSON”, return only JSON orSwyx [00:07:38]: YeahPhilip [00:07:38]: Grandma's gonna die type of prompts.Swyx [00:07:39]: Is it BNF grammar? At some point OpenAI had released a thing that was like, yeah, if you want to constrain your output, write BNF grammar, back as NOR.Philip [00:07:47]: In our inference system, it's just a specified output format. And you get the guarantee that your output's gonna be structured along that format. And so applying that to tool calls can like help cut down on. You can still call the wrong tool or call no tool. It doesn't solve the certainty problem but it at least solves the output structuring problemSwyx [00:08:10]: YeahPhilip [00:08:10]: Within tool calls.Swyx [00:08:12]: And MCP is just another form of tool, right.Philip [00:08:14]: Yeah, exactly.Swyx [00:08:15]: As far as there's no special thing there.Philip [00:08:16]: The thing I'm always like explaining to people is the LLM is not capable of doing anything. It's only capable of making suggestions of what to do and then if those suggestions are formatted in a certain way and applied to a system that knows what to do with them, then an action occurs.Swyx [00:08:32]: Yeah. Part of the fun stuff is, this is solved outside of tool calling too. Like in an agent loop if the output is not correct or you're right, like reasoning, tool calling was done in the reasoning trace, just be like, “Oh, I don't know what to do. Let me just try again.” And it might get there after a few tries. And on your point of training, sometimes this is harder in smaller models, so you don't have the same exact quality outputAli [00:08:56]: Right.Swyx [00:08:57]: When you just swap from a big model, right?Ali [00:08:59]: Yeah. I will say that, before, I think we need to go back to inference engineering proper.Ali [00:09:04]: But, I had expected that something would replace JSON because it's hard to stream JSON ‘cause JSON must be complete and you must have open and close brackets and everything. So it's hard to parse something or validate something while it's being streamed. So people invented all sorts of things that are like, I forget the name of some of these alternatives, but it's something like TOML, something like YAML. But JSON seems to be dominant still.Philip [00:09:30]: The JSON outputs aren't that long, right? Like you could have a long-- ‘cause tool calls also contain the arguments in them and perhaps for a certain tool you might pass like a very long argument. But my impression of the median tool call is that it's a relatively small number of tokens, right? So I would expect that speculators are generally fairly good at something as formatted as JSON. And so you would have like a pretty fast decode step there and that the streaming wouldn't be as valuable, but maybe I'm wrong about that.Ali [00:10:02]: I think you're also bounded by the software or that the model is gonna integrate with if the software is built with JSON for the tool calls or if the company that you'- if your customer says that this is how our software works and our tools are interfaced with JSON, you can ask them to like, change their software and say like, “Yeah, this is gonna be better for the model.” but like with the right training shouldn't be that much of a difference. Also more profitable if it outputs more tokens probably.Swyx [00:10:25]: Depends on your business model.Swyx [00:10:27]: It really depends. But I will say that, as a writer with like experience a lot with generated output, I do try to move from text to JSON text which is very long JSON, right? Like there's paragraphs in every field because I'm trying to structure it, right?Philip [00:10:44]: Right.Swyx [00:10:44]: I want you to first make factual statements, then make opinions then make bullet point summaries, have dates, have entity references have your sources for references, all these things. Anyway, so these are things that like I think people who really experiment with structural output have to really care about. But, let's, let's recurse up the stack a little bit. Before we started recording, you mentioned something really cool, which is that there's a lot of engineering that-- inference engineering that goes on when a new model provider releases a new model, right? So let's call it GLM-5.2, Kimi K3. I had previously assumed, especially if it's like, well, GLM 5 to 5.1 to GLM-5.2, like that you've supported them before. Is it that much work?What It Takes to Support a New Open ModelAli [00:11:26]: It's a lot of work.Swyx [00:11:28]: Yeah. Okay. So like, a lot of people, all you guys, right whenever a new model launch like, people rush to say like, “Oh, Hugging Face supports this, Fireworks supports this, Spacetime supports this,” and I'm like, “Yeah, of course we support it.” But what goes into that? What goes intoPhilip [00:11:40]: I think it's more than just support it too, right? It benefits the consumer a lot. Like I think it was with Kimi K2.5 or GLM-5.2 the latest, there was an inference war, right? X provider is at 90 tokens a second. The next day we're at 150. The nextSwyx [00:11:55]: I kinda kicked that off with the GLM-5.2.Swyx [00:11:58]: I wrote a Twitter article about. It got like half a million views,Ali [00:12:02]: Based on being numberSwyx [00:12:03]: YeahAli [00:12:04]: Or it's for something else.Swyx [00:12:05]: Yeah. Which,Ali [00:12:06]: Oh my GodSwyx [00:12:07]: Which then got everyone really excited about, hey, how can we, bend tracks a little bit further and,Philip [00:12:14]: There's a difference between support the model, as in I can make a token out of this model, and support a model, as in I have a production-ready API from this model.Philip [00:12:26]: Getting to the point of I can make a token out of this model is not that hard because generally the, open source inference engines, vLLM, SGLang of the world oftentimes even receive weights ahead of time, maintainers do, or the people making the model merge PRs to ensure support. So you generally can, just get it working on the standard open source stack without too much pain in most cases. The challenge is, every inference company is gonna have own proprietary stack. Some open source components, some in-house stuff. And for any arbitrary model, there's going to be some new stuff. Sometimes you get lucky, like K, two five to two six was, like, pretty similar.Quantization, Speculators, and Production ReadinessAli [00:13:16]: Yeah. It was pure continued post-trainingPhilip [00:13:18]: YeahAli [00:13:18]: If I remember correctly.Philip [00:13:19]: Even in those cases, there's still stuff you have to do. You have to redo the quantization work. You're taking the model from. Generally, these models are not released in NVFP4, and we want them to be in NVFP4 for maximum Blackwell compatibility. So we have to perform that quantization, and, calibrate the quantization to make sure that we're not causing any regression in the model's intelligence. And then we also have to train the speculator, as we've talked about. Generally, we have. We have ZDR, zero data retention on our model APIs, so we don't know exactly the traffic that people are sending us, but we know what's popular. We know that coding use cases are popular. We know that agents, agentic use cases are popular. So we can get public data sets that are representative of that traffic and train general speculators. Now, with speculators today, you need to train the speculator using the base model itself because you're getting hidden states out of the model from running inference on these specific prompts, and that is the training data you use to create the speculator. So there's that process which you need the real model weights for. And then there's of course just the process of, standing up all the infrastructure behind it, loading all this stuff, testing it. And then when there's a new model with a newer architecture, I think that, like, the DeepSeek models tend to be the most challenging as they have, like, the most novel architectural stuff going on, model after model. But every new model has something. Kimi K2 had. Oh, sorry, GLM-5.2 hadAli [00:14:53]: Sparse attention.Philip [00:14:54]: Yeah,Ali [00:14:54]: YeahPhilip [00:14:54]: the DSA.Ali [00:14:55]: Right. Which is brought from DeepSeek.Philip [00:14:57]: Yeah. AndAli [00:14:59]: So you can copy-paste then?Philip [00:15:01]: It kindAli [00:15:01]: I don't know how this works.Philip [00:15:02]: So, like we had to, like, build support for that into our runtime. And you're right, like it is really interesting the way that all of these open source labs borrow from each other. For example, like GLM-5.2 doesn't have vision. So something that, Haley, a guy on our team, if we could take a look at this, he, like, grafted the Kimi vision encoder onto GLM-5.2.Retrofitting Vision into GLM-5.2Ali [00:15:27]: We'll be training the projector.Philip [00:15:28]: Exactly. So if you think about, like, the encoder, there's the encoder, which is the part that looks at the image and turns it into latent information, and then there's the projector which likeAli [00:15:38]: You can say latent space. It's okay.Philip [00:15:41]: And then there's the projector that maps it onto, the model itself, and then there's the model weights. You don't wanna mess with the model weights because you run a chance of making the model dumber at something else for the purpose of giving it vision. So instead, Haley started with just a projector, which is only a handful of millions of parameters.Ali [00:16:02]: That would be, yeah.Philip [00:16:02]: Yeah.Ali [00:16:03]: Can you show the training one?Ali [00:16:04]: Like the way it groksPhilip [00:16:05]: YeahAli [00:16:06]: Very interesting.Philip [00:16:06]: And maybeAli [00:16:07]: That right therePhilip [00:16:07]: Maybe Ali, you should take it from here. You've got a betterAli [00:16:10]: Ooh, double the sandPhilip [00:16:11]: Understanding of this than I do.Ali [00:16:11]: Yeah. You can see, like, he. The way he trained this is really cool. At the beginning, he was training it using just like, “Here's a picture of a mountain. Can you describe what's in this mountain?” And that caused it just like the first, learning walls. Like here you can see this all we're trying to teach it is to translate the encoded. Like it's already taken the encoder from Kimi K. It's taken the image. It'Philip [00:16:31]: Yeah. FrozenAli [00:16:31]: FrozenPhilip [00:16:32]: With adapter.Ali [00:16:32]: Exactly.Philip [00:16:33]: Yeah.Ali [00:16:33]: So the brain is frozen and the eyes are frozen. It's just we're tryingPhilip [00:16:37]: AlignAli [00:16:38]: Interconnect between the eye and the brain, right? So the projector. And so you take the tokens and then he's like, “Oh, can you describe what's in this image?” And he's like, “Oh, it's a mountain,” or it's a person or it's a human, whatever the case is. But that didn't cause complete understanding. So he changed it such that every image was associated with a data set of questions. Like, does this image have a white male? Does this image have birds in the top corner? Does this image have a scientist in it? All of that stuff. And it would have to answer questions correctly. And using not just training on describing an image, but being able to answer question, another question, answer over time. Like you can see the grokking, which is like genuinely insane, that retrofitting vision into a large LLM can learn to that extent. And even for images that it doesn't perform well on, for instance, if you ask it a picture of like Stephen Hawking, “Who is this?” Maybe it doesn't get it, but it will say something like, “This is Albert Einstein.” Like it still understandsPhilip [00:17:25]: Close enoughAli [00:17:26]: That this is a scientist who is a man who has, some significant achievements, all that stuff. So that's like really cool.Philip [00:17:32]: Yeah. So, we've covered Hao Tian before, who the author of the LLaVA paper that did this, a while ago. And I think that's very foundational work for anyone who hasn't done vision work before.Ali [00:17:41]: Same with the CLIP and MetaCLIP, where you go from just captioning to building out questionsPhilip [00:17:47]: RightAli [00:17:47]: Off the image and how much better you can get performance.Philip [00:17:50]: Right. Right. Right. Yeah. But what's, what's so exciting about this is if you look at a model like this. Now, this is a little bit more of a research project. It's not. It got to 56% on MMLU Pro, I think. So not quite frontier. But if you're running this model, you haven't suffered any loss on your GLM-5.2 quality. If you don't have an image, it'll just behave exactly the way it used to. And ultimatelyAli [00:18:14]: Which in the inference code you literally do not include the other part, right?Philip [00:18:18]: Yeah. You would just skip the encoder if you don't have an image input.Ali [00:18:22]: Okay.Philip [00:18:22]: Just confirming.Philip [00:18:23]: YeahAli [00:18:23]: Does it affect a lot on the overall inference side? Like you're not adding much, you're adding a very small vision encoder. These are typically likePhilip [00:18:30]: They're super fineAli [00:18:31]: Less than a billion parameters, right?Philip [00:18:32]: Yeah. It's, - There's a little bit less standardization among vision encodersSwyx [00:18:37]: YeahPhilip [00:18:37]: So the support matrix can be a little bit, sparser. But overall, yeah, it's a pretty, it's a pretty minor component of the overall system. And ultimately what you get out of the system is all of a sudden you have Kimi Vision, GLM weights, and DeepSeek attention all in one model.Open Source Model Grafting and Franken-MergesPhilip [00:18:56]: And that's, I think, a lot of the power and beauty of open source, is that you can take all of these different components and combine them together into a system that's better than anyoneSwyx [00:19:05]: YeahPhilip [00:19:05]: Can be individually.Swyx [00:19:06]: People used to say that you would also do Franken-merges where you would take likePhilip [00:19:10]: YeahSwyx [00:19:10]: Layers from each model.Swyx [00:19:11]: Does anyone do that anymore?Ali [00:19:13]: Well, to your point previously when you were mentioning like, the work that goes into supporting a model when it first comes out, like GLM-5.2 or MiniMax M3 or whatever the case is. Sometimes you do have to like, you do have to switch out some things. Like, for instance, the MiniMax M3 head uses full attention, and with full attention you end up with this like insane bottleneck in spec dec ‘cause you're doing auto-regressive token generation for three tokens, and you're doing this like N squared over all of the tokens that are in your sequence. Your KV cache is like very large because it's not sparse, it's not top K. So we find it better to like, okay, we're gonna replace this, we're gonna replace this layer with a layer from another model that's using like GQA, for instance. And then just with the right training, you can get it to have the same acceptance rate. So it is very possible to retrofit layers from other models and very much needed. If a layer is like inefficient, the training just becomes the challenge, like how do you ensure that you train it properly? Which again to your earlier point is like the mesh between training and inference. As in like you need very good training in order to do fast inference. That's like, I feel like more and more becoming true.Swyx [00:20:21]: Yeah. Anything else on the support side when you say like get it to fully production ready?Loop Detection, Race Conditions, and Non-DeterminismPhilip [00:20:26]: Yeah. I think that there's also a question of just, we can test a model to a pretty extensive degree, but we're trying to get it out quickly and then you see a bunch of other people test it and you get interesting results. There was an issue with, GLM briefly where we had some like mode collapses where it would just output the same token over and over again for certain prompts on certain temperatures. Like once you expose an endpoint to the real world, there's going to be, so many more varieties of things given to it that you're able to, discover and patch things. So it's not just a, day zero process, it's then like for the first week, for the first month, if a model remains popular, like how do you both fix bugs and then continue to push the envelope on performance?Ali [00:21:21]: What do you mean you don't want your model outputting S?Swyx [00:21:24]: Is there loop detection on that stuff, by the way? It still happens like quite a lot, which is surprising.Ali [00:21:30]: We have like we, in our endpoint, like if a model was to output the same token like four plus times, we just cut the generation. We say like, “Oh, sorry, this-- Like try again,” or like we will reprocess the request. ‘Cause we know then, like if it, like if, yeah, it's four times the same token, it's probably collapsed.Swyx [00:21:45]: Yeah. Is there a way to opt out in case I really want that?Ali [00:21:48]: You want that?Ali [00:21:50]: I think there's a way that we have to handle it. I'm not exactly certain, but I feel like in certain models, like when they output something like you can imagine, like a table for instance, and so they want, they wanna draw like 12 dashes and 12 dashes. Yeah, I think there's a way for that to happen. I think we only do it on certain tokens. Like we exclude certain special characters.Swyx [00:22:07]: Yeah.Ali [00:22:07]: So we only do it on like certain like S is the most common almost. GLM-5.2Swyx [00:22:11]: OhAli [00:22:11]: And I think it was DSV 4 as well. Like you'd just have like looping issues where like you literallySwyx [00:22:17]: ItAli [00:22:17]: Just have like S.Swyx [00:22:18]: Yeah. Is there a special, something special about S? No, just randomlyAli [00:22:21]: It just seems to be the one token involved.Swyx [00:22:23]: Yeah. And it'Philip [00:22:24]: Is thereSwyx [00:22:24]: And it's only temperature 0Ali [00:22:27]: NoSwyx [00:22:27]: Even at other temperaturesAli [00:22:27]: Even at like 0.9 or whatever, it will still, it will still collapse.Swyx [00:22:30]: That's weird, right?Ali [00:22:30]: It's, it is an inference problem to be honest, like a software problem. Like oftentimes, the image you run will-- like NVIDIA will release an image for instance, and if we will upstream the changes from their latest TensorRT-LLM image into our stack, we'll find that it fixes it. Or oftentimes this will only happen in an inference engine that you're using like SGLang. But if you were to switch to vLLM, that isn't the case. So it seems to be like an extremely like deterministic software issue and not really a model issue. It's not like a weights problem. Like I'- we'll say like, “Oh, it's a problem with the quant. We did PTQ wrong,” right? But that isn't, that doesn't make sense because the same weights used with a different inference engine does not repeat the problem. And sometimes it's, the kernels that are being used in the backend have like these very subtle sometimes race conditions, where if you were to use this model hosted on one cluster, you will never get this problem.Swyx [00:23:19]: Oh my God.Ali [00:23:19]: But if you host it on a different cluster, you will. And the reason is the KV cache transfer from a node to node in that one cluster is using a slower interconnect than the node to node in another cluster. So that exposes the race, whereas in another cluster it doesn't. So then you end up just like, okay, this model is not gonna be hosted on this cluster. We're gonna host it on, another cluster because that cluster exposed that problem. But then it ends up with like, okay, is it the software? Is it the model weights or is it the hardware?Swyx [00:23:42]: There is a thing about this with temperature 0 still not being deterministic, right?Ali [00:23:46]: Right.Swyx [00:23:46]: Mostly because of hardware. Even at temperature 0 same model, you won't always get the same output.Swyx [00:23:52]: Even-- But I'm surprised by the race condition one because, I thought PyTorch was a graph that like guarantees that you at least, execute things in the right order.Ali [00:24:02]: Well, yeah, true. Like I'm not, I'm not saying that there is. Like well, you have things like PTL optimizations where like you can start a kernel before the end of the previous kernel, and that's like ‘cause you want to do that because there'sSwyx [00:24:12]: It's like pipeliningAli [00:24:12]: Expense. Exactly.Swyx [00:24:13]: Yeah.Ali [00:24:13]: But it'- But you don't do it cleanly. Like you overlap a little bit of the execution. No, it is very possible that the kernel itself, like that one block that is supposed to be running in this instance of time, that kernel itself has a race condition. For instance, like a missing barrier. Like often if you're designing a kernel and you want it to make it to be very fast, if you don't test it extensively, you'll, you'll have certain threads access data points from registers before they've been written to by other threadsSwyx [00:24:36]: YeahAli [00:24:36]: For example, because like your barrier is wrong or your synchronization was wrong. But yeah, like the testing itself is very difficult in those like, andSwyx [00:24:42]: And there's no like borrow checkerAli [00:24:45]: What does that mean?Swyx [00:24:46]: Like Rust. Like the. If you're trying to have like memory safety It sounds like a comparable problem.Ali [00:24:52]: Well, yes, but you're working in CUDA, right, NVIDIA GPUs. Like- You just need a higher level language like modular Maybe that's what modular is supposed to do. I don't know.Quantization Quality and Vendor FidelityVibhu [00:25:00]: How do you see keeping quality of the model? So you talked about all these steps of, okay, you gotta do quantization, train your own speculative decoderAli [00:25:07]: RightVibhu [00:25:07]: Run on different hardware. Looking at other model providers, okay, you kicked off a inference speed race on the consumer end. What goes into keeping quality the same across them, right? Sure, you can run benchmarksAli [00:25:22]: YeahVibhu [00:25:22]: But, like, how do you determine how much quantization are there standards? What goes intoPhilip [00:25:27]: There's a few things on quality. Most inference optimizations are lossless. KV caching, for example. You are just recomputing or preventing recomputing the same values. Speculation, of course, if a draft token is wrong, it gets rejected. The main lossy optimization is quantization. And that really comes down to, number one, data format, number two, which parts of the model you choose to quantize, which layers, and number three, like doing a lot of calibration on the quantized weights, to ensure that you're preserving all the outliers. There's other tricks that you can do, though. A big one is long context, ‘cause one thing you asked at, right at the beginning is, “Oh, what's gonna happen if I send a 200,000 token request in?” So with a long input sequence, you need to, store a lot more information. You need to process a lot more tokens. And so even if a model has a context of a certain length, you might, as an inference provider, choose to build an API with a shorter context length, and of course a full length one as well. Because if someone doesn't need the full million token context, for example, you can get them better performance. I don't know if that's exactly like quality of the model. The way that I think about quality is to what degree are we faithfully serving the original model? If you think of a golden implementation of a model that performs exactly the way the model is designed to perform, I think of quality as how close are we getting to that, 100% fidelity of the model.Philip [00:27:13]: You can also, of course, think about quality from the training side and how do you push yourself past 100%. But when I think about purely inference optimizations, it's getting faster while staying as close to that 100% fidelity mark as possible. And certainly our standard internally is that, like you should not be able to tell the difference between our API and a, official API. I think Kimi in particular does a good job of vendor benchmarking hereAli [00:27:41]: YesPhilip [00:27:41]: Where they haveAli [00:27:42]: They released an actual vendor benchmark.Philip [00:27:43]: Exactly, yeah.Ali [00:27:44]: ‘Cause they accused, some people, Amazon? There was some provider that was not doing very well on Kimi's benchmark.Philip [00:27:50]: Yeah.Philip [00:27:51]: So, with Reflect we probablyVibhu [00:27:52]: This was a long time ago, right?Philip [00:27:54]: No.Ali [00:27:54]: Yeah, like threeVibhu [00:27:55]: They alsoAli [00:27:55]: Four, five months agoVibhu [00:27:57]: This also happened with, I don't remember which model, but they pulled out quite a few, and then they started a whole chart about this. It might have beenPhilip [00:28:03]: Kimi Vendor Verifier.Ali [00:28:04]: Yeah.Philip [00:28:05]: Yeah.Ali [00:28:05]: Yeah, ‘cause you, ‘cause you'd be pissed, right? Like if you'Philip [00:28:07]: Yeah.Ali [00:28:07]: If like if I'm a consumer and I'm using like Amazon's endpoint for instance, and I've used Kimi and I'm like, “Oh my God, like this is bad,” I'm not gonna say, “Oh, Amazon quantized the model in a bad way.” I'm gonna say, “Oh, Kimi sucks.” Right?Philip [00:28:17]: Yeah.Ali [00:28:17]: So it seems like that makes sense.Philip [00:28:19]: Yeah, they care. They care.Vibhu [00:28:21]: Justifiably.Ali [00:28:21]: Yeah, justifiably.Vibhu [00:28:22]: This is probably a stupid question, but just checking, has anything improved from main quantization?Philip [00:28:28]: Yeah.Vibhu [00:28:28]: Like, is quantization always strictly worse?Ali [00:28:30]: Well technicallyVibhu [00:28:32]: NoAli [00:28:32]: It's a lossy. QuantizationPhilip [00:28:33]: YeahAli [00:28:33]: Is a lossy, it's a lossy implementation.Philip [00:28:36]: Speed improvesVibhu [00:28:36]: Speed improves.Ali [00:28:37]: It the number, likeVibhu [00:28:38]: No, I' always look for inverse scaling laws.Philip [00:28:40]: Yeah.Ali [00:28:40]: Yeah.Vibhu [00:28:40]: This is something I learned from Noam Brown, where like things that normally act in one direction sometimes do.Philip [00:28:45]: Well, technically when you run a benchmark, because these models are deterministic, sometimes your,Ali [00:28:52]: YeahPhilip [00:28:52]: NVFP4 quant is like, two basis points higher than yourAli [00:28:56]: No, it's noise. It's noise.Philip [00:28:57]: Yeah, exactly. I'm like, yeah, it's, it's within. That's why I always say within margin of error.Philip [00:29:01]: And I stopped saying that because everyone assumes that what is, well, within some margin of error, we're barely inside of that to the worst, so we're saying. But yeah, sometimes it's just like, gives you a higher output score. But like Ali said, that's noise. To my knowledge, you're not necessarily making the results better. You're just trying to, again, like keep your fidelity as close to 100% to the original model.Layer Selection, KL Divergence, and Better QuantizationAli [00:29:27]: There is, to your point, research that we did on MP. I don't know if you are able to pullPhilip [00:29:31]: YeahAli [00:29:32]: A tweet we did. One of our research interns, Joshua, I think it's a tweet on how we have 20% better quantized GLM-5.2 than NVIDIA. Essentially what we found throughout like this month research is, okay, quantization is a lossy. It's. You're compressing the data from, occupying 16 bits to occupying, four bits, for instance. And so you're losing some information, and you're trying to minimize that. And so when I say that I'm gonna quantize the model, my job becomes how do I find the layers that I can quantize, and how to find the layers to not. For instance, with image models, I don't quantize modulation layers, and I don't quantize out projections because those two are. Like out projection is what you see as the user. Modulation is what the model sees or understands. Right, exactly. And so to his paper, do you have the. It doesn't have the. Yeah. It's a long paper. I don't know if I can findVibhu [00:30:25]: If there's a part to search or it's probably in the thread.Ali [00:30:28]: It's probably in the thread.Vibhu [00:30:29]: Yeah.Ali [00:30:29]: But the long and the short is it is very possible that quantizing more of the model makes the results. Like if I have a model that I quantize layers one, five, and 10, and another model where I only quantize layers one and It is possible that the model in which I quantized more information is going to perform better because the quantization errors have canceled out. And so what Joshua showed in his mathematical proof where he had like a verifier in, is that you can predict which layers are going to have quantization errors that will cancel out with each other, and you choose to quantize those layers. And so the result of doing this mathematical quantization is you end up with a model that's 20% more quantized than another provider, so you get 20% more throughput of it because there's more layers than running an NVFP4, and your quality is better than that other quant because the layers that you chose to quantize have their errors cancel out, like one layer skewed to the right one layer skewed to the left, one layer skewed to the right. Your final logits distribution is more similar to the original distribution of the model, so you have better fidelity. And so the way we proved this was with KL divergence. So instead of just scoring on the benchmarks, we scored the KL divergence between the logit distribution of the quantized model and the logit distribution of the original full precision model, and we showed that with this technique we get. If your probability distribution on the logits which token it wants to select is more of the same as the original model, you're probably gonna end up staying true to the original model. So yeah, so it seems like previously before this, it seemed like the industry was, well, the more you quantize, the worse it's gonna be, ‘cause the more loss you introduce. That's not exactly, not necessarily true. So yeah, doesn't improve it, but can cancel out.Philip [00:31:57]: I think it might be this, but reminds me a good bit about pruning where you can prune off certain layers.Philip [00:32:03]: But very interesting. Didn't know this was a whole paper you guys put out.Ali [00:32:06]: It's. Fun fact, it was originally 72 pages, this paper, and then we decidedPhilip [00:32:11]: WowAli [00:32:11]: We can't tell. We couldn't release it. So it's now 45.Swyx [00:32:15]: Still 39 pages, so very substantive. We talked about evals and all these things and, like what's possible in terms of speedup? Like it's like probably like the numberInference Speedups and BenchmarkingSwyx [00:32:25]: Thing that people do wanna care about, and it's something that you wrote about in your post. Like official API is 70 tokens per second, and you push it up to 90. Is that like a normal thing?Philip [00:32:36]: So what's cool about working in inference, the reason that I think inference is going to be a useful place to do engineering for a long time, is that if you look at highly optimized domains like, say, finance, if you're in finance, you measure how much better you got in basis points. It's like, “Oh, I got five basis points better, like twentieth of 1% better,” that's huge news because everything is so optimized. When we publish optimizations, it's 20%, it's 100% it's 200%. So there's still probably like a lot further to go, honestly. Like you'll, you'll know that inference is pretty much solved when researchers start publishing about how they got 1% faster at something.Swyx [00:33:19]: Which by the way, because I am from the finance background, in the ‘70s, that was the margin at the time. When you did quantitative finance research, you would findAli [00:33:27]: And like 20%, tens of percent.Swyx [00:33:29]: That's. Yes.Philip [00:33:29]: Yeah.Swyx [00:33:30]: And now it'Philip [00:33:31]: Tiny fractionsSwyx [00:33:32]: For those people interested, look up Andrew Lo's paper. He had a really interesting illustration of quant, stat arb, distribution, narrowing down from like those kinds of 20% differences in the ‘70s, down to nothing today, which is very cool.Philip [00:33:48]: Exactly, and we're at the beginning of the same type of thing. Now benchmarking is hard. I think anyone will tell you that, and benchmarking provider speeds is hard because there's so many variables that go into it. What hardware are you using? How much load do you have on the system? What's the exact nature of the prompts and input and output sequence lengths? All that stuff. But overall, when you start stacking these improvements, you're looking at multiples. You can look at it. The most common form, of course, is TPS, tokens per second, which is bad naming by us in the industry, ‘cause there's two tokens per second. There's tokens per second, the throughput number, and the latency number.Ali [00:34:31]: TTMT, yeah.Philip [00:34:32]: Like total tokens per second out of the, out of the GPU as a throughput number. Most people only care about tokens per second as the latency number, which we should call ITL, intertoken latency, but we don't.Philip [00:34:44]: Anyway, so you can imagine a standard API without many optimizations for a 1 trillion parameter model operating somewhere in the 30 to 50 tokens per second range for reasonable traffic profile. And we generally see the goal of, pushing to 10X that. But, not necessarily day zero, but by stacking enough optimizations, if you have, say like four optimizations, each of which doubles performance. Or sorry, three optimizations, each of which doubles performance, then you stack that up, that's an 8X gain. That's the order of magnitude that we're working with in this space. We're trying to make things substantially faster, not just go from like 70 to 90.Swyx [00:35:38]: Are you saying you've. You have done that?Philip [00:35:40]: So let's say you have as a reasonable baseline, 30 or 40 tokens per second. You can achieve 10X that. So like on GLM-5.2, if you run it unquantized, perhaps on H100s even, and you're just using an off-the-shelf inference engine with no particular optimizations, no speculator, nothing extra around like KV routing, no disaggregation, you're, you're probably, yeah, looking at that like 30 to 40. You think that's like a reasonable baseline?Swyx [00:36:12]: Right. Right.Philip [00:36:12]: To get to something like 10X, there's a lot of trade-offs that you're making. If we're running at more like a 300, 400 tokens per second range, you are using the best hardware possible. You have a optimized speculator. You have done all of your quantization work. You are Seeing a pretty high cache hit rate. You are running with a reasonably small batch size and a parallelism configuration that is tuned for latency versus throughput, but it is possible. So the spreads that you see if you, like, go on artificial analysis or you go on OpenRouter and you look at, the worst provider to the best provider, oftentimes can hit that range. 10X is of course very aggressive. It's oftentimes maybe more of a four to six times improvement. But that's the performance that makes us really excited, is when we can get these huge gains, not just go from 70 to 90 tokens.Stacking Optimizations: NVFP4, Speculation, and DisaggregationAli [00:37:19]: It's also, like, hardware dependent. Like, ifPhilip [00:37:20]: YeahAli [00:37:20]: If you have a thing where you're serving it on just, like, a node of H100s and then you throw, like, you shard the model across, like, four nodes of B200s. Like, you can definitely increase the speed with just throwing more hardware at it. Like, normalizing for the same exact hardware and the same number of GPUs.Philip [00:37:35]: Yeah. Then you're looking at, like, a two to 4X improvementAli [00:37:38]: Right. RightPhilip [00:37:38]: Depending on the inference optimizations. So yeah, it's. Some of it's, what's the call, and some of it's who's the driver.Vibhu [00:37:46]: If you break down the two to 4X, say the example is run GLM-5.2Ali [00:37:51]: YeahVibhu [00:37:51]: On B200sAli [00:37:53]: YeahVibhu [00:37:53]: Single node, right? What's, like, the cost trade-off for effort to get, like, the last bit of juice out versus what should people just think of, right?Ali [00:38:01]: Spectre quantization. Yeah.Vibhu [00:38:03]: Spectre quantization.Ali [00:38:04]: That's, that's, that's like 95%. LikeVibhu [00:38:06]: And how far does that get you? And how easy is that for the average person to do? So say right I wanna throw the weights of GLM-5.2 on a node of B200s, how easy is it to find speculative decoder- decoder model or already quantized model? How much work goes into it?Philip [00:38:23]: If you're doing it up front, it's quite a lot of work. If you're doing it today, there's going to be people who have published things that you can just, you can just grab some NVFP4 weights. You can grab a speculator. Yeah, if we're thinking about, like, what are the 2Xs we're stacking, going from, BF16 to NVFP4 is, it's not quite a 2X, right? It's like. I think it's about, like, 30 to 40%, from 16 to 8, and then another 30 to 40% multiplied from, 8 to 4. So that doesn't quite get you a 2X, but, like, roughly a 2X. Speculator, roughly a 2X. Disagg on top of that if you're able to get enough hardware and put enough traffic through it, another roughly a 2X. And then you add in some, double-digit percent increase from having just a better runtime with, the latest kernels and stuff behind it. And that's how it stacks up.Ali [00:39:21]: YeahPhilip [00:39:21]: So building each of those, like, building the, quantized weights is, for someone who really knows what they're doing, hours to days of work. Building the speculator, again, like, hours to days of work. And the, disagg setup, hours to days. Well okay, but like once you haveAli [00:39:39]: Once set up. Once set up. YeahPhilip [00:39:40]: Yeah, getting disagg working for the first time, I'm saying, of course, is very difficult.Philip [00:39:44]: The marginal implementationAli [00:39:48]: Like, if you're just grabbing, like if you are a person, like just a normal consumer who has access to, like, a node of B200s and you're wondering, “How can I just host it myself?” You don't need to quantize the model yourself. There's always gonna be, like, an open source quantized checkpoint. NVIDIA's gonna push one out if no one else does. You. Usually, the providers will have their own spec dec that they've trained as well. You don't need to train your own spec dec. You can just use that as well.Philip [00:40:09]: Yeah. Like, GLM-5.2 has its own MTP.Ali [00:40:13]: Right. Right.Vibhu [00:40:14]: What's multi token prediction?Philip [00:40:15]: Yes.Ali [00:40:16]: I'm justVibhu [00:40:16]: Can you explain that?Ali [00:40:16]: I'm just an expert.Ali [00:40:18]: I can do it for you in case I get it wrong?Vibhu [00:40:20]: No.Vibhu [00:40:21]: Yeah, you should correct if we're wrong, but their multi-token prediction can be used for self-speculative decoding.Ali [00:40:27]: I'm not sure. I'm not gonna correct that.Vibhu [00:40:28]: Okay. I'm semi-confident in thatAli [00:40:30]: Okay. YeahVibhu [00:40:30]: But someone can check. But it's useful to paint the story of, okay, not just the average person, but say a company wants to switch from serverless inference I wanna throw this up on. I wanna rent some GPUs, throw it up. These are the steps you take to do significantly faster than just put it behind vLLM.Ali [00:40:48]: Right.Vibhu [00:40:49]: I was waiting for a mention of Dynamo.Vibhu [00:40:51]: I feel like, that's supposed to be the baseline that you measure against.Dynamo, KV Routing, and Disaggregation ToolkitsPhilip [00:40:55]: I would think of Dynamo as less of a box system and more of a toolkit for building with. So when we talk about doing aware routing, when we talk about doing KV offloading, when we talk about doing, PD disaggregation, Dynamo fundamentally is. By the way, Dynamo is an open source library from NVIDIA.Ali [00:41:17]: We've done a pod with KylePhilip [00:41:18]: OkayAli [00:41:19]: Kyle Cranin.Philip [00:41:19]: Cool. So then your listeners know then that it supports all the different inference frameworks. And it is multi hardware, which is interesting.Ali [00:41:28]: But it's just a router, it's not like an optimizer layer.Philip [00:41:30]: Yeah. All it does, like, what Dynamo is good at, it is a library for moving information around your cluster, around your hardware. So if you have, KV cache on one place and you need it to be somewhere else, Dynamo coordinates NIXL for you to move that around.Philip [00:41:49]: That doesn't mean that, like, out of the box, you just say, “Pip install Dynamo,” and then you get, like, a massive performance speed up. It's more of a developer toolkit.Ali [00:42:01]: Yeah. I would have said it would. It comes with a set of defaults that you can then swap out.Philip [00:42:06]: It does. If the industry at large, I think, was, like, rolling out all of these deployments, standard, then I think it would be, like, a credible baseline. But, we've got to, we've got to benchmark against, like, what we're seeing in the wild.Speculative Decoding Methods: Medusa, EAGLE, n-Gram, and Spec-SpecVibhu [00:42:23]: I did wanna talk a little bit more about PD disagg, because that is probably, like, number three after quantized and speculative decoding. In your book though, I was just gonna pull out the book.Philip [00:42:31]: Yeah.Vibhu [00:42:32]: Like section 522 on Medusa, 523 on EAGLEPhilip [00:42:35]: YeahVibhu [00:42:36]: 524 on gram.Philip [00:42:37]: It's 55, would be disaggregationAli [00:42:42]: Yeah. Well, no, I just wanted to dwell a little bitPhilip [00:42:44]: YeahAli [00:42:44]: The other. Like, so what do you choose to include? What do you choose to not to include? Because there was all these other techniques.Philip [00:42:51]: Yeah.Ali [00:42:51]: Are these still relevant? Because I think they came out, like, a year and a half ago maybe.Vibhu [00:42:55]: Medusa is quite old.Philip [00:42:56]: Yeah, Medusa's old.Ali [00:42:58]: It was old.Vibhu [00:42:58]: But is it in the book as a good, here'sPhilip [00:43:01]: BaselineVibhu [00:43:01]: Baseline vanilla understand it?Philip [00:43:02]: Like you should know this.Vibhu [00:43:03]: Like I read the paper, I'm like, “ it makes so much sense.”Philip [00:43:05]: Yeah.Philip [00:43:05]: So with the book, I had a couple goals. One was to give people just a working vocabulary for the space as a whole, and the other was to give them some intuition about how each of these techniques works. As I mentioned in my AI Engineer talk, which is the first public addendum to this, the speculation space has moved much faster than everything else. So yeah, even at the time that I wrote the book Medusa, I very much included as a way for people to understand how the space evolved rather than what the most modern technique is. And now of course, there's DFlash, dSpark. There's, there's newer techniques even than EAGLE, although EAGLE is still very commonly used.Ali [00:43:51]: SpecSpecta.Philip [00:43:52]: Yes. Speculative decoding.Vibhu [00:43:54]: What canAli [00:43:56]: Oh, it's a paper by Tri Dao and it's like, it's doing speculative decodingVibhu [00:44:00]: HuhAli [00:44:01]: For the speculative decoder.Philip [00:44:02]: Oh, in spec- oh my God.Ali [00:44:02]: It's literally just an another. It's like, yeah, that's the most simple way to explain it, and it seems like he got trivial speed ups there. But it seems that the complexity with training, it's almost like in our mind at least, it's almost as complex as training GANs. Like it's like a very delicate balance and oftentimes you, it's just but yeah, it's literally speculative decoding on speculative decoding.Vibhu [00:44:21]: Speculative.Ali [00:44:22]: Yeah. We saw this paper.Vibhu [00:44:24]: It's interesting, right?Ali [00:44:24]: Yeah.Vibhu [00:44:24]: I wouldn't even expect it to be very particular to train, I wouldAli [00:44:29]: Right.Vibhu [00:44:29]: The naive part of me is like, okay, train speculative decoder.Ali [00:44:32]: But like, and it makes sense, like the whole idea of speculative decoding is you. It's like, it's like almost like the iPhone auto predict version but for a normal model, right? Like you're just, you're just, generating three tokens and you're like, okay, I'll do prefill on them. And so you save those three turns for your original model. Now your speculative decoder is doing three turns of auto regression, so why not just have an even smaller model?Ali [00:44:53]: The other question there is what are the size of speculators? So say forPhilip [00:44:58]: Right. It's like a billion parameters.Ali [00:45:01]: Like for MiniMax, it's. Yeah. It's like one layer. It's like one 60th of the original model usually.Philip [00:45:06]: Yeah. I think we should do a paper when we get back to the office.Philip [00:45:10]: SpeculativeAli [00:45:11]: SpeculativePhilip [00:45:11]: Decoding.Ali [00:45:13]: No, it's, it does seem like how, when do you stop? But then it also seems like if you're able to train spec-spec decode for instance, right? Like if you're able to have a small model that is accurately predicts what the intermediate speculator is gonna predict, that is able to predict what the original target model's gonna predict, then why not just use that smallest model directly, right?Vibhu [00:45:34]: Yeah. This isAli [00:45:35]: Like it seems likeVibhu [00:45:35]: Adjacent to the routing problem.Ali [00:45:36]: Right.Vibhu [00:45:36]: Yeah.Ali [00:45:36]: Right.Philip [00:45:37]: The thing with speculators is one of the practical constraints on using them is that you do have to run a small model on the same hardware that you're running the big model on. There is a orchestration and resource competition problem inherent in that, and that is one of the constraints on speculation in general, is that draft tokens cost resources to create and cost software complexity to manage. And so if you have like infinitely recursive speculators, you add in quite a bit of that complexity on the actual implementation within the inference engine as well, not just in the training process.Vibhu [00:46:17]: I was gonna say, I would wonder if you could do similar, like distillation and pruning of, it's the same thing, it's just a model. Can we not just distill a lot of the weights, quantize the speculator, out of my domain? The question that also comes up is, this is all for big server workloads, right? How much of this applies to, say I have this MacBook, I wanna run Gemma really efficiently. Similar problems, not the same?Local AI vs. Data Center InferencePhilip [00:46:45]: Pretty different. I talked to Selo, about this on his podcast a couple weeks ago. The difference between inference engineering for the data center and for production workloads versus inference engineering for local AI, is that we start with fundamentally like different constraints and different goals. With local AI, it's how do I fit this model onto my hardware and then make it less dumb? And with data center influence, it's how do I load this model and then make it less slow? And we care about less dumb, and they care about less slow. But the local AI inference engineering ecosystem, I think has a lot for us to learn from in the data center space. They are experts in various forms of quantization, including dynamic quantization that we just don't touch, in the pruning, in the distillation, in the, layer removal. There'Ali [00:47:42]: Layer removal matters less.Philip [00:47:43]: Yeah. There'Ali [00:47:44]: No one loves pruning really.Philip [00:47:45]: Yeah. Well, but the, but they doVibhu [00:47:46]: Which is surprising, right? But that's, that's a whole different thingPhilip [00:47:48]: Just to fit something on the laptop.Ali [00:47:50]: Right.Philip [00:47:50]: So yeah, it's a, it's an interesting, it's an interesting space. Not necessarily that like their techniques make sense for us to do in the data center, because we have different resources and different goals, but more that the process as well as the openness of that field is something to, admire.Ali [00:48:12]: Yeah. Like to your point, like, certain optimizations that would. Like for instance, Turbo Quantum Sharper, like it made such huge hype on that and we did like a whole deep dive on Twitter and like said, what is it? How does it work? Why is it good or not? And it took off and it was implemented on local devices because your memory bandwidth is so slow on like a MacBook, for instance. But try putting the same thing on like an NVIDIA GPU on a B200 Turbo quant would not be. Like, it would not be used. Like, NVIDIA - Like, NVIDIA made it clear that this is not a good optimization, and we've seen it firsthand where the overhead of doing dequantization, quantization of, in the kernel itself with turbo quant kernel, each end is much slower than the time that you save from doing the bandwidth. ‘Cause on the B200s, you have like 3.5 terabytes per second. You don't need decrease the storage that much. You don't need to do, FP4 KV cache. You don't need to use a requant. There's, there's, there's better optimizations to be made. But on Edge devices, it's extremely important, it's extremely useful. So, seems to be, like, different optimizations there, but then they're all uniquely combined with like all you wanna quantize the model, you wanna do speculative decoding, like certain common prefixes with bothPhilip [00:49:18]: Principles.Ali [00:49:19]: Yeah, exactly. Exactly. Exactly.Philip [00:49:20]: They also do a lot of work on, model parallelism, especially over, heterogeneous topology, where you have, some sparks and they are wired together with, Ethernet, DGX sparks.Ali [00:49:35]: Yeah, this is the Exo Labs guys.Philip [00:49:36]: Yeah. You have, a nu

Best of Roula & Ryan
7a Former 2006 Intern Brittany Behind The Scenes Of Her Job 07-30-26

Best of Roula & Ryan

Play Episode Listen Later Jul 30, 2026 12:57


Best of Roula & Ryan
6a Wild Wednesday Night, Scoop Surprisingly Rude Celebrities, Why Was Eric Late, and Catching Up With Intern Brittany 07-30-26

Best of Roula & Ryan

Play Episode Listen Later Jul 30, 2026 32:24


Al & Jerry's Postgame Podcast
Al & Jerry: People in their 30s living at home and would your trade your life with an intern's? -- plus, warmup

Al & Jerry's Postgame Podcast

Play Episode Listen Later Jul 29, 2026 64:23


Al & Jerry: People in their 30s living at home and would your trade your life with an intern's? -- plus, warmup

Al & Jerry's Postgame Podcast
Al & Jerry: People in their 30s living at home and would your trade your life with an intern's?

Al & Jerry's Postgame Podcast

Play Episode Listen Later Jul 29, 2026 22:24


Al & Jerry: People in their 30s living at home and would your trade your life with an intern's?

Boomer & Gio
People in Their 30s Living at Home and Would Your Trade Your Life With an Intern's? | 'Al & Jerry's Postgame Podcast'

Boomer & Gio

Play Episode Listen Later Jul 29, 2026 23:42


Al & Jerry: People in their 30s living at home and would your trade your life with an intern's?

The Federalist Radio Hour
'You're Wrong' With Mollie Hemingway and David Harsanyi, Ep. 207: Democrats' Communist Resurgence

The Federalist Radio Hour

Play Episode Listen Later Jul 22, 2026 59:00 Transcription Available


Join Federalist Editor-In-Chief Mollie Hemingway and Washington Examiner Senior Writer David Harsanyi as they recap their Independence Day celebrations, discuss Sen. Lindsey Graham's legacy, weigh in on the scandal that broke Graham Platner's Senate campaign, and rant about the communism resurgence fostered by Democrats. Mollie and David also review Little House on the Prairie, Midwinter Break, The Intern, Eternity, and Past Lives, and Mollie shares her favorite moments from her trip to Europe.Order and review Mollie's book Alito: The Justice Who Reshaped the Supreme Court and Restored the Constitution here.The Federalist Foundation is a nonprofit, and we depend entirely on our listeners and readers — not corporations. If you value fearless, independent journalism, please consider a tax-deductible gift today at TheFederalist.com/donate. Your support keeps us going.