Podcasts about Spec

  • 1,380PODCASTS
  • 4,059EPISODES
  • 44mAVG DURATION
  • 5WEEKLY NEW EPISODES
  • Aug 27, 2026LATEST

POPULARITY

20192020202120222023202420252026

Categories



Best podcasts about Spec

Show all podcasts related to spec

Latest podcast episodes about Spec

Investor Fuel Real Estate Investing Mastermind - Audio Version
How to Develop Lakefront Real Estate | Land Acquisition, Spec Homes & Investor Capital

Investor Fuel Real Estate Investing Mastermind - Audio Version

Play Episode Listen Later Aug 27, 2026 26:37


In this episode, real estate developer Guice Mercer shares insights on land development, luxury lakefront projects, managing market risks, and building strong investor relationships. Discover practical strategies for success in high-end real estate markets and the importance of personal relationships and faith in business. In this episode, we explore the entrepreneurial journey, challenges in capital raising, marketing strategies, project management, and the importance of faith and integrity in business. Our guest shares valuable insights from his diverse experiences in construction, racing, and real estate.   Learn More in the Blog Article → https://investorfuel.com/blog/waterfront-spec-home-development/   Professional Real Estate Investors - How we can help you: Investor Fuel Mastermind:  Learn more about the Investor Fuel Mastermind, including 100% deal financing, massive discounts from vendors and sponsors you're already using, our world class community of over 150 members, and SO much more here: http://www.investorfuel.com/apply   Investor Machine Marketing Partnership:  Are you looking for consistent, high quality lead generation? Investor Machine is America's #1 lead generation service professional investors. Investor Machine provides true 'white glove' support to help you build the perfect marketing plan, then we'll execute it for you…talking and working together on an ongoing basis to help you hit YOUR goals! Learn more here: http://www.investormachine.com   Coaching with Mike Hambright:  Interested in 1 on 1 coaching with Mike Hambright? Mike coaches entrepreneurs looking to level up, build coaching or service based businesses (Mike runs multiple 7 and 8 figure a year businesses), building a coaching program and more. Learn more here: https://investorfuel.com/coachingwithmike   Attend a Vacation/Mastermind Retreat with Mike Hambright: Interested in joining a "mini-mastermind" with Mike and his private clients on an upcoming "Retreat", either at locations like Cabo San Lucas, Napa, Park City ski trip, Yellowstone, or even at Mike's East Texas "Big H Ranch"? Learn more here: http://www.investorfuel.com/retreat   Property Insurance: Join the largest and most investor friendly property insurance provider in 2 minutes. Free to join, and insure all your flips and rentals within minutes! There is NO easier insurance provider on the planet (turn insurance on or off in 1 minute without talking to anyone!), and there's no 15-30% agent mark up through this platform!  Register here: https://myinvestorinsurance.com/   New Real Estate Investors - How we can work together: Investor Fuel Club (Coaching and Deal Partner Community): Looking to kickstart your real estate investing career? Join our one of a kind Coaching Community, Investor Fuel Club, where you'll get trained by some of the best real estate investors in America, and partner with them on deals! You don't need $ for deals…we'll partner with you and hold your hand along the way! Learn More here: http://www.investorfuel.com/club   —--------------------

Rocky Mountain ATV/MC Keefer Tested
Show #501 - 2027 Honda CRF450R Review

Rocky Mountain ATV/MC Keefer Tested

Play Episode Listen Later Aug 17, 2026 72:15


Keefer talks about riding the new Honda CRF450R and what the changes are like on the track compared to the older model Honda as well as other colored machines. Get the straight scoop on if this new Honda could be the machine you purchase in 2027. --- CHAPTERS --- 00:00 Show intro announcer 10:27 Millville Honda intro recap 17:22 Engine rideability & power character 24:53 ECU maps, clutch feel & vibration 33:28 Commercial break 34:23 Live ad reads - Blood Lubricants & Ride Engineering 38:40 Chassis & rigidity impressions 46:31 Cornering weight, stability & flickability 54:38 Suspension & fork/shock settings 59:05 Ergonomics - seat, bars & grips 1:03:11 Spec sheet & final verdict 1:10:42 Next show preview & sign-off

Microsoft Cloud IT Pro Podcast
Episode 434: Reading the Spec Sheet on Microsoft’s Agentic SOC

Microsoft Cloud IT Pro Podcast

Play Episode Listen Later Aug 13, 2026 47:34 Transcription Available


Welcome to Episode 434 of the Microsoft Cloud IT Pro Podcast. Microsoft announced Project Perception on 2026-07-27 with a lot of “workforce of AI agents” language. In this episode, Ben and Scott discuss what Perception really is, what it needs from your tenant, what it costs, and what happens the first time an agent asks you to approve something. Six agents. Seven playbooks. Every one manually triggered. Each agent gets its own Entra Agent ID and a permission grant. There’s a lot to unpack. Your support makes this show possible! Please consider becoming a premium member for access to live shows and more. Check out our membership options. Show Notes Save $200 at TechCon365 with code Ben200: https://www.techcon365.com/Seattle/ What is Project Perception? How Nationwide stays ahead of attackers with Project Perception Key concepts in Project Perception Agent categories in Project Perception Microsoft Security Copilot Security Compute Units and capacity Work with playbooks in Project Perception Prior podcast context Ep 431: Agent Governance Is the New App Governance Ep 419: Security and AI  Ep 416 and Ep 412  Ep 406: Agents of Insight  Ep 377: Microsoft Copilot for Security  Sponsors Nasuni is a leading unstructured data platform for enterprises where file data is mission-critical for both people and AI. Nasuni powers the operational file layer where work happens — helping organizations manage, protect, and activate data so teams can work smarter, reduce costs, and operate securely without limits. Intelligink — Would you like to become the irreplaceable Microsoft 365 resource for your organization? Let us know!

Nutty Bites
Scrubs

Nutty Bites

Play Episode Listen Later Aug 13, 2026 9:45


Scrubs is back, and while it took Tek and I a while to watch it, it seems like we finished the season just in time to get ready for season 2! Welcome to dog days of podcasting for 2026, where podcasters … Continue reading → The post Scrubs appeared first on NIMLAS Studios.

CHCH Podcasts
Newsmakers: The Hamilton Spectator marks 180th anniversary

CHCH Podcasts

Play Episode Listen Later Aug 12, 2026 37:06


The Hamilton Spectator is celebrating its 180th anniversary this year. Founded before Confederation and the incorporation of the City of Hamilton, journalists at the Spec have worked tirelessly to keep our community informed. Newsmakers Host Rick Zamperin speaks with the Spectator's editor-in-chief Cheryl Stepan and reporter Jon Wells.

Nutty Bites
Masters of the Universe

Nutty Bites

Play Episode Listen Later Aug 11, 2026 12:27


Masters of the Universe kinda snuck up on me, but I am glad we watched it. Great does of nostalgia. Welcome to dog days of podcasting for 2026, where podcasters from around the world do a podcast a day for the … Continue reading → The post Masters of the Universe appeared first on NIMLAS Studios.

Bardtenders
The Mixing Glass | Guest Shift - Dave Krysl | Spec

Bardtenders

Play Episode Listen Later Aug 10, 2026 84:30


You're listening to Bardtenders! In this episode of "The Mixing Glass", Dave Krysl shares his insight into the cocktail world and the importance of having a system for not just inventory, but also recipe management and education. ------------Dave is the co-founder and CEO of Spec, a platform helping bars and restaurants save time and money through Recipe Management and Inventory Control. He graduated from the Leeds School of Business and spent 9 years in New York City, where he fell in love with cocktail culture. Aside from his passion for elevating the food and beverage industry, he plays guitar in a metalcore band called Haste the Day and spends most of his free time playing golf.----------Don't miss out on any of the action!  Head to www.bardtender.com to stay up to date with all of the Bardtender content, find resources for mental and physical well-being, get access to education materials, and check out what all of our bards are up to!Support the show

Nutty Bites
Count Binface

Nutty Bites

Play Episode Listen Later Aug 9, 2026 6:28


I love a good joke, and how funny is it that Count Binface has a chance to get a real chunk of the vote on 13 August? Welcome to dog days of podcasting for 2026, where podcasters from around the world … Continue reading → The post Count Binface appeared first on NIMLAS Studios.

Nutty Bites
Super Mario Galaxy the Movie

Nutty Bites

Play Episode Listen Later Aug 9, 2026 9:26


Super Mario Galaxy is a love letter to Mario players and it brings some great princess empowerment to the story. Welcome to dog days of podcasting for 2026, where podcasters from around the world do a podcast a day for the … Continue reading → The post Super Mario Galaxy the Movie appeared first on NIMLAS Studios.

Nutty Bites
Night Waves

Nutty Bites

Play Episode Listen Later Aug 8, 2026 15:50


Nutty and Tek (Pictured above) went to Night Waves a 21 plus event at the water park with music and djs and fun. It was a blast. Welcome to dog days of podcasting for 2026, where podcasters from around the world … Continue reading → The post Night Waves appeared first on NIMLAS Studios.

Nutty Bites
Sterling Point

Nutty Bites

Play Episode Listen Later Aug 7, 2026 5:21


I cried, three tissue cries by the end of this series, but I dug it. Sometimes you need a good cry, and I am very thankful it didn’t end on a cliffhanger. Welcome to dog days of podcasting for 2026, where … Continue reading → The post Sterling Point appeared first on NIMLAS Studios.

The Drive
Spec Says – Cyrus Allen will have Under 375 Yards

The Drive

Play Episode Listen Later Aug 6, 2026 8:10


Steven Spector joined The Drive to explain why he is tampering down expectations for Chiefs 5th round rookie Cyrus Allen.

Entre Chaves
Snippet #27 Como decidir entre os principais frameworks spec-driven

Entre Chaves

Play Episode Listen Later Aug 6, 2026 10:48


Qual framework realmente faz sentido para o seu projeto? Neste Snippet, Luiz Fernando Reis, Gerente Técnico na dti digital, compara sete dos principais frameworks utilizados atualmente, destacando suas características, vantagens, limitações e os cenários em que cada um entrega melhores resultados. Se você está avaliando ferramentas para desenvolver com agentes de IA, este conteúdo ajuda a transformar essa escolha em uma decisão técnica, e não apenas em uma tendência. Dê o play e ouça agora!Assuntos abordados:Spec-Driven;Spec Kit;OpenSpec;Tessel;PlanDeck;Aider;Escolha de frameworks.Links importantes:Vagas disponíveisNewsletterDúvidas? Nos mande pelo LinkedinContato:  entrechaves@dtidigital.com.brO Entre Chaves é uma iniciativa da dti digital, uma empresa WPP #inteligenciaartificial

Turn Down for Watt
Rivian R2 What Every Buyer Needs to Know!

Turn Down for Watt

Play Episode Listen Later Aug 5, 2026 52:33


We kick off this episode with a look at the grand opening of Power Up America Site #1 in Kentucky, Then we're joined by Travis Ketchum (Stay Curious) for an in-depth conversation covering the Rivian R2. Travis shares his impressions of the new Rivian R2 and discusses what prospective buyers should realistically expect. We break down the recent 70 MPH highway range tests from Out of Spec and State of Charge, why the results differed from the EPA estimate, and what those numbers actually mean for everyday driving.We also discuss:• Is the LIDAR package worth it?• Buy or lease—which makes the most sense?• Travis's lease calculator and how it can help buyers compare optionsResources & Featured Links

StarTalk Radio
Cosmic Queries – Planet 9

StarTalk Radio

Play Episode Listen Later Aug 4, 2026 47:41


Could there be a hidden planet lurking beyond Pluto and Neptune? Neil deGrasse Tyson and comic co-host Chuck Nice tackle a wide-ranging grab bag of fan questions covering Planet 9, galaxy shapes, dark matter, the odds of getting hit by a meteorite, and more! NOTE: StarTalk+ Patrons can listen to this entire episode commercial-free. Thanks to our Patrons Glenn Fiskin, A. Certain Woman, Jeffrey Schwartz, Adam Watson, Micheal Lawson, Aaron Hardy, Terje Berg, Peter Wells, Tyler, Brody Johnson, Ali, NanZan, SAMuri, Marian Bieniek, Jeff Whitney, Richard Ruffner, Adrian J Batson, Holly Grant, Amy Braden, Andrew Puente, Srihari Ravi, Eliezar Vega, Southern Tauren, Adam Tate, John Johnson, Achim Ferrandina, Rebecca Crowe, Cody May, Gary, William Green, Max_Palmer09, Dylan, Jonas and Aisté, Frank Pucino, Isaac Kinsey, Sean Whitehall, Anthony Prisco, Nathan Thornton, Alex Grzesiak, Lora Vatalaro, Manny Neto, Eric M, Robert Davies, Billy Metcalfe, Gabriel Gish, Sherry Lee Rachar, Alexander Nemeth, Ra.Spec, and Atticus Thompson for supporting us this week. Subscribe to SiriusXM Podcasts+ to listen to new episodes of StarTalk Radio ad-free and a whole week early.Start a free trial now on Apple Podcasts or by visiting siriusxm.com/podcastsplus. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

planet ra pluto neptune neil degrasse tyson simplecast spec john johnson star talk kuiper belt william green eric m aist chuck nice startalk radio peter wells samuri adam watson jeffrey schwartz spiral galaxies cosmic queries kuiper belt objects
Nutty Bites
Hadestown

Nutty Bites

Play Episode Listen Later Aug 4, 2026 10:06


I dug Hadestown on the big screen and really want more recordings like this made available in theatres. Welcome to dog days of podcasting for 2026, where podcasters from around the world do a podcast a day for the month of … Continue reading → The post Hadestown appeared first on NIMLAS Studios.

CTO Morning Coffee
Model AI uciekł z sandboxa i zhakował Hugging Face | Brew #70s

CTO Morning Coffee

Play Episode Listen Later Aug 4, 2026 80:13


Mimo wakacji i lata w pełni CTO Morning Brew powraca w pełnym składzie! ☕ Wojtek, Tomek i Sebastian biorą na warsztat gorące jak druga fala upałów newsy z rynku tech, geopolityki i finansów.O czym rozmawiamy w 68. odcinku?

Latent Space: The AI Engineer Podcast — CodeGen, Agents, Computer Vision, Data Science, AI UX and all things Software 3.0

Watch the full episode on YouTube:We first covered Baseten last year when DeepSeek mania was at peak hype. Now they have raised a monster $13B round and become one of the new cohort of AI Infra decacorns that are (with Nvidia, Intel, and the semis complex) chief beneficiaries of the Inference Inflection. We return to Baseten at the peak of the 2026 edition of Open Weights debate. Ali has published a viral breakdown of Kimi K3:And since you last saw him, Philip has spoken at AI Engineer and written the definitive book on Inference Engineering spotted all over SF:Three years ago, inference engineering barely existed as a category.Today, it is one of the most critical disciplines in AI. Inference engineering inherently tackles a different question than standard model training: “How do you turn those weights from training into a product that is fast, reliable, and affordable at scale?” Focusing on these creates an entirely new optimization problem.In one recent GLM-5.2 experiment, quantizing more of the model actually preserved its benchmark quality while increasing throughput by 20%, because the errors introduced in different layers could cancel each other out.Inference is no longer just the final step after training. It is becoming its own engineering discipline, with its own research problems, infrastructure, and increasingly specialized roles.In this episode, Baseten's Philip Kiely and Ali Taha join swyx and Vibhu to explain what actually happens after a new open model is released and what it takes to turn “we generated a token” into a fast, reliable, production-ready API.We go deep on cache-aware routing, disaggregated prefill and decode, quantization, speculative decoding, KV-cache movement, model parallelism, GPU kernels, and the race to make frontier models up to 10× faster. Philip and Ali explain why inference optimizations can still produce gains of 20%, 100%, or even 200%; how quantization errors can cancel one another out; why identical weights can behave differently across clusters; and how Baseten grafted a Kimi vision encoder onto GLM-5.2 without changing the underlying language model.The conversation then expands beyond LLMs into NVIDIA Dynamo, mega kernels, Rubin, AI-specific chips, local inference, video generation, diffusion versus autoregressive models, and the enormous compute barrier to generating coherent long-form video. Finally, we explore the convergence of training and inference, continual learning through persistent KV cache, and the emerging loop where models help optimize the infrastructure that runs them.We discuss:* What happens when a 200,000-token request enters an inference system* Cache-aware routing and reusing previously computed KV cache* Why prefill and decode are increasingly handled by different GPUs* When dedicated deployments become cheaper and more reliable than shared APIs* How speculative decoding uses a smaller model to accelerate a larger one* Tool calling, structured outputs, and what LLMs actually do* What it takes to support a new open model on day zero* Grafting Kimi's vision encoder onto GLM-5.2* Retrofitting inefficient model layers with components from other architectures* Why models sometimes collapse into repeating the same token* How hardware, kernels, and race conditions create nondeterministic failures* Preserving model fidelity while making inference faster* How quantization errors can cancel each other out* Why inference optimizations still deliver gains of 20%, 100%, and 200%* How optimized serving can make a model up to 10× faster* NVIDIA Dynamo, KV-aware routing, and distributed model serving* Speculative decoding the speculative decoder* Why local AI is about making models less dumb while data-center AI is about making them less slow* Tensor, expert, and pipeline parallelism across GPUs* Hardware-aware model design, auto-tuning, and the case against mega kernels* Rubin and why inference is becoming a systems problem* Whether modern GPUs are evolving into programmable AI ASICs* Why enormous models like Kimi K3 require GB300-class hardware* Why open-source video generation still trails Veo, Kling, and other closed models* The quadratic attention bottleneck behind long-form AI video* Autoregressive video, real-time generation, and compounding quality drift* Why future video systems may combine autoregressive and diffusion architectures* Training for inference and inference for training* Continuous post-training, deployment, evaluation, and improvement loops* How GLM-5.2 helped optimize the kernels serving GLM-5.2 itself* Why faster networking could unlock dramatically faster decoding* Continual learning, KV-cache compaction, and persistent model memoryShow Notes* How to build a day-0 API for Kimi K3* 22580: From GPT2 to Kimi3, ExplainedPhilip Kiely* LinkedIn: https://www.linkedin.com/in/philipkiely* X: https://x.com/philipkiely* Inference Engineering: https://www.baseten.co/inference-engineering/Ali Taha* LinkedIn: https://www.linkedin.com/in/aliestaha/* X: https://x.com/waterloointernTimestamps00:00:00 Introduction and the 200K-Token Prompt00:03:18 Dedicated Deployments, Speculative Decoding, and Tool Calling00:11:26 Launching Production-Ready Open Models00:19:06 Model Retrofits, Failure Modes, and Nondeterminism00:28:22 Quantization and Canceling Errors00:32:15 The Race to 10× Faster Inference00:40:48 Dynamo, Speculation, and Local vs. Data-Center AI00:50:18 Model Parallelism, Auto-Tuning, and Mega Kernels01:00:55 Rubin, GPUs vs. ASICs, and Custom AI Chips01:10:03 Giant Models and the Limits of GPU Memory01:12:42 AI Video, Quadratic Attention, and Autoregressive Generation01:21:47 Audio, Images, and Diffusion Models01:27:32 Training, Self-Optimizing Models, and Continual Learning01:40:06 Closing ThoughtsTranscriptIntroduction: Baseten, Waterloo Intern, and Inference EngineeringSwyx [00:00:00]: Okay, we're here in the studio with Philip, old friend from Inference Engineering, the book, as well as Baseten and everything that you've done, you and I have done before, as well as Ali. Welcome.Ali [00:00:15]: Pleasure to meet you.Swyx [00:00:15]: Waterloo intern.Ali [00:00:16]: Waterloo intern, always.Swyx [00:00:17]: When did you get “Waterloo intern” as a handle?Ali [00:00:19]: As a handle? Oh.Ali [00:00:20]: I think the rebranding happened mid-March. When I saw it was open, I was like, “I have to take it. Up for grabs.”Philip [00:00:26]: The problem is that Ali is really good at his job and is not gonna be an intern much longer.Philip [00:00:30]: So we have to figure out who's gonna get the handle.Ali [00:00:33]: Well, I'll pass the torch over to the next intern.Swyx [00:00:34]: Oh, okay. It can be, like, you just pass it to another Waterloo grad.Ali [00:00:37]: To another Waterloo intern. No, bruh.Philip [00:00:39]: Yeah.Ali [00:00:39]: Intern.Swyx [00:00:40]: Intern, yeah.Ali [00:00:40]: And no.Philip [00:00:41]: You gotta get an intern from Waterloo.Ali [00:00:42]: Yeah, I've gotta get an intern from Waterloo.Swyx [00:00:44]: Right.Ali [00:00:44]: But they have to follow the path.Swyx [00:00:45]: Oh, it could, but it could come from Baseten, so it's like whoever Baseten gets from Waterloo.Ali [00:00:48]: Right.Swyx [00:00:49]: Has the title of Waterloo.Ali [00:00:50]: It stays in the ecosystem.Philip [00:00:51]: Exactly.Ali [00:00:52]: Halfway through the internship, you either get it or you're out.Philip [00:00:55]: You should also do, like, a big graduation ceremony where you change the handle.Ali [00:00:59]: Just say it.Philip [00:00:59]: For everybody.Swyx [00:01:00]: You guys are good at ceremonies, clearly. We had a nice launch of the book, very successful. But before we get into all that, I wanna start off with a fun question for you. Okay, you're an expert inference engineer. What happens when I send a long query, say two hundred thousand tokens into Baseten's inference? What's the process of query through GPU model routing, balancing, all that? What is all the stuff that we don't think about?Long Context Requests, KV Cache, and Cache-Aware RoutingPhilip [00:01:26]: With a long query specifically, the first thing that I'm gonna ask is, “Have you sent me this query before, or at least part of it?” and I really hope you have, because it's gonna be a lot easier for me and a lot cheaper for you. So the first thing that we're gonna look at is some cache-aware routing, where we're going to see, we probably have a number of instances, a number of replicas up serving whatever model you're hitting. We want to send this one to something with, number one, available prefill workers, and number two, ideally some cached input already there so that we can skip prefill on at least part of these two hundred thousand tokens. If you're doing two hundred thousand tokens, it's probably coding or a multi-turn agent or something where you would expect to have that cached. If you don't, we're gonna have to send it to a prefill worker. We've at least on certain models disaggregated prefill and decode, so you're going to have one set of GPUs that's solely going to process the input, create the KV cache, and get you your first token, and then that's going to be passed over to a separate set of GPUs, which is going to run decode. We're going to iteratively make those tokens. We're probably going to have some speculator model in front of that. I'm going to assume that you're doing coding, and because of that, our speculator model, which assumes you're doing coding, is gonna have a high draft token acceptance rate. If I'm wrong and you're asking me to summarize every Harry Potter book, it's gonna be slower. And then we stream that output to you and account for it, charge you, a couple of pennies and say, “Hey, would you like to send another one?”Swyx [00:03:04]: Except Baseten doesn't charge by pennies.Philip [00:03:07]: Well, yeah, we charge. I'm assuming that we're talking about the public model APIs. If you are setting up a dedicated deployment, then yeah, it's not pennies.Public APIs vs. Dedicated DeploymentsSwyx [00:03:18]: Yeah, one of the key differentiators when I was talking with Baseten initially was that people who want very high volume just need to rent by the box, ‘cause then it's up to you to figure out how to saturate the box.Ali [00:03:31]: And more often than not, it's, like, way cheaper if you're pushing, like, millions of tokens per hour, if you just pay per hour instead of pay per token.Philip [00:03:37]: Yeah, they do. I think that we've increasingly seen a lot of demand for the pay per token APIs, just because everyone wants to try open models, and then once they find a use case that's really sticky, then they move over to dedicated.Swyx [00:03:51]: Is there a best practice on when it's time to swap over?Philip [00:03:54]: Couple reasons. Yeah, reliability, that's a big one, right?Ali [00:03:57]: Like, if they have a very specific use case, they want you to train something specifically for them, like they want their own spec dec, for instance, for their own traffic.Swyx [00:04:04]: Spec dec is speculative decoding.Speculative Decoding and Custom SpeculatorsAli [00:04:05]: Speculative decoding, yeah.Swyx [00:04:07]: You have to explain.Ali [00:04:07]: Sorry. Like, speculative decoding is like, if you have a huge model, right? And so the model is going to be generating one token at a time every single turn, every single forward pass. So we attach, like, this little, like, parasite, like this layer that goes on top of the model, and this model just has to predict. It does three very fast autoregressive forward passes, and it will predict, like, three certain tokens, and then you do one forward stage over the entire original model in order to see if those predictions were correct or not, and then you accept them or you reject them. Now, this draft model is traffic specific, so if you, like, Philip said, if you're summarizing Harry Potter books, I can train exclusively that draft model on Harry Potter books, and I can guarantee you that I'm gonna accept the three tokens every single time. And so with that case, I increase your decode speed. I wouldn't be able to provide this to you if you're a shared endpointSwyx [00:04:53]: YeahAli [00:04:53]: ‘cause I have no idea if you're doing Harry Potter, if you're doing coding, if you're doing English. We don't know. Also, there was a thing in the book that mentioned that if they really cared about a specific threshold, chapter four, I think. Do you remember that?Philip [00:05:06]: Yeah. The things that you can do is you can set a specific, like, batch sizing, a specific, like, parallelism strategy if you're trying to optimize for, like, throughput versus latency. You can. Maybe a NVFP4 quant doesn't pass your benchmarks and you wanna run a model at higher precision, you could do that. There's just a bunch of reasons why you might wanna have your own endpoint and the biggest one, of course, just being, like, you don't have to deal with someone else doing a hundred million tokens of benchmarking traffic at the endpoint when you happen to be trying to serve your users.Swyx [00:05:40]: Yeah. I think one thing that is. That is a classic journey. Like, it's people is asking the, what happens when you type Google into the browser. Tool calling, is that just, you're generating JSON or is there more complication beyond that?Tool Calling, JSON, and Structured OutputsAli [00:05:58]: Certain customers that we have, they have their own post-trained models, and so they demand a tool calling that's not just, like parse a file or go find the weather. It's something that's very specific and you have to do post-training on this. And if the post-training on the model is not good or if the quantization after the post-training to get the inference to be fast, the model will struggle reading the JSON file and reading the tool calling. But it doesn't require its own like sandbox. It's not like it's going to use that tool calling to like escape a sandbox or like it doesn't have to be contained. It can just be a normal dedicated deployment. The challenge with tool calling more and more seems to be that the companies want certain tool calling which is a very sensitive thing to train. And because you're dealing with all of the JSON outputs, if it doesn't like close the end of the request in a very certain manner, you end up with a model that did the tool calling and like the thinking and so as a result of that, it didn't see the result and just hallucinated the result as it decoded. That seems to be the most challenging thing with tool calling, not really the sandboxes model.Philip [00:06:56]: Yeah, that's a challenge on the training side and then on the inference side, there's work that you can do to scope the possible output. So we published this at this point close to two years ago, the solution to this problem which is you make a state machine and you use that to constrain the output to a specific format. So this is the structured output problem. If you remember backSwyx [00:07:27]: Yeah, the specific grammar is,Philip [00:07:29]: Yeah, exactlySwyx [00:07:30]: GML had this thing.Philip [00:07:31]: Yeah. So it's like the old-school “make sure this is only JSON”, return only JSON orSwyx [00:07:38]: YeahPhilip [00:07:38]: Grandma's gonna die type of prompts.Swyx [00:07:39]: Is it BNF grammar? At some point OpenAI had released a thing that was like, yeah, if you want to constrain your output, write BNF grammar, back as NOR.Philip [00:07:47]: In our inference system, it's just a specified output format. And you get the guarantee that your output's gonna be structured along that format. And so applying that to tool calls can like help cut down on. You can still call the wrong tool or call no tool. It doesn't solve the certainty problem but it at least solves the output structuring problemSwyx [00:08:10]: YeahPhilip [00:08:10]: Within tool calls.Swyx [00:08:12]: And MCP is just another form of tool, right.Philip [00:08:14]: Yeah, exactly.Swyx [00:08:15]: As far as there's no special thing there.Philip [00:08:16]: The thing I'm always like explaining to people is the LLM is not capable of doing anything. It's only capable of making suggestions of what to do and then if those suggestions are formatted in a certain way and applied to a system that knows what to do with them, then an action occurs.Swyx [00:08:32]: Yeah. Part of the fun stuff is, this is solved outside of tool calling too. Like in an agent loop if the output is not correct or you're right, like reasoning, tool calling was done in the reasoning trace, just be like, “Oh, I don't know what to do. Let me just try again.” And it might get there after a few tries. And on your point of training, sometimes this is harder in smaller models, so you don't have the same exact quality outputAli [00:08:56]: Right.Swyx [00:08:57]: When you just swap from a big model, right?Ali [00:08:59]: Yeah. I will say that, before, I think we need to go back to inference engineering proper.Ali [00:09:04]: But, I had expected that something would replace JSON because it's hard to stream JSON ‘cause JSON must be complete and you must have open and close brackets and everything. So it's hard to parse something or validate something while it's being streamed. So people invented all sorts of things that are like, I forget the name of some of these alternatives, but it's something like TOML, something like YAML. But JSON seems to be dominant still.Philip [00:09:30]: The JSON outputs aren't that long, right? Like you could have a long-- ‘cause tool calls also contain the arguments in them and perhaps for a certain tool you might pass like a very long argument. But my impression of the median tool call is that it's a relatively small number of tokens, right? So I would expect that speculators are generally fairly good at something as formatted as JSON. And so you would have like a pretty fast decode step there and that the streaming wouldn't be as valuable, but maybe I'm wrong about that.Ali [00:10:02]: I think you're also bounded by the software or that the model is gonna integrate with if the software is built with JSON for the tool calls or if the company that you'- if your customer says that this is how our software works and our tools are interfaced with JSON, you can ask them to like, change their software and say like, “Yeah, this is gonna be better for the model.” but like with the right training shouldn't be that much of a difference. Also more profitable if it outputs more tokens probably.Swyx [00:10:25]: Depends on your business model.Swyx [00:10:27]: It really depends. But I will say that, as a writer with like experience a lot with generated output, I do try to move from text to JSON text which is very long JSON, right? Like there's paragraphs in every field because I'm trying to structure it, right?Philip [00:10:44]: Right.Swyx [00:10:44]: I want you to first make factual statements, then make opinions then make bullet point summaries, have dates, have entity references have your sources for references, all these things. Anyway, so these are things that like I think people who really experiment with structural output have to really care about. But, let's, let's recurse up the stack a little bit. Before we started recording, you mentioned something really cool, which is that there's a lot of engineering that-- inference engineering that goes on when a new model provider releases a new model, right? So let's call it GLM-5.2, Kimi K3. I had previously assumed, especially if it's like, well, GLM 5 to 5.1 to GLM-5.2, like that you've supported them before. Is it that much work?What It Takes to Support a New Open ModelAli [00:11:26]: It's a lot of work.Swyx [00:11:28]: Yeah. Okay. So like, a lot of people, all you guys, right whenever a new model launch like, people rush to say like, “Oh, Hugging Face supports this, Fireworks supports this, Spacetime supports this,” and I'm like, “Yeah, of course we support it.” But what goes into that? What goes intoPhilip [00:11:40]: I think it's more than just support it too, right? It benefits the consumer a lot. Like I think it was with Kimi K2.5 or GLM-5.2 the latest, there was an inference war, right? X provider is at 90 tokens a second. The next day we're at 150. The nextSwyx [00:11:55]: I kinda kicked that off with the GLM-5.2.Swyx [00:11:58]: I wrote a Twitter article about. It got like half a million views,Ali [00:12:02]: Based on being numberSwyx [00:12:03]: YeahAli [00:12:04]: Or it's for something else.Swyx [00:12:05]: Yeah. Which,Ali [00:12:06]: Oh my GodSwyx [00:12:07]: Which then got everyone really excited about, hey, how can we, bend tracks a little bit further and,Philip [00:12:14]: There's a difference between support the model, as in I can make a token out of this model, and support a model, as in I have a production-ready API from this model.Philip [00:12:26]: Getting to the point of I can make a token out of this model is not that hard because generally the, open source inference engines, vLLM, SGLang of the world oftentimes even receive weights ahead of time, maintainers do, or the people making the model merge PRs to ensure support. So you generally can, just get it working on the standard open source stack without too much pain in most cases. The challenge is, every inference company is gonna have own proprietary stack. Some open source components, some in-house stuff. And for any arbitrary model, there's going to be some new stuff. Sometimes you get lucky, like K, two five to two six was, like, pretty similar.Quantization, Speculators, and Production ReadinessAli [00:13:16]: Yeah. It was pure continued post-trainingPhilip [00:13:18]: YeahAli [00:13:18]: If I remember correctly.Philip [00:13:19]: Even in those cases, there's still stuff you have to do. You have to redo the quantization work. You're taking the model from. Generally, these models are not released in NVFP4, and we want them to be in NVFP4 for maximum Blackwell compatibility. So we have to perform that quantization, and, calibrate the quantization to make sure that we're not causing any regression in the model's intelligence. And then we also have to train the speculator, as we've talked about. Generally, we have. We have ZDR, zero data retention on our model APIs, so we don't know exactly the traffic that people are sending us, but we know what's popular. We know that coding use cases are popular. We know that agents, agentic use cases are popular. So we can get public data sets that are representative of that traffic and train general speculators. Now, with speculators today, you need to train the speculator using the base model itself because you're getting hidden states out of the model from running inference on these specific prompts, and that is the training data you use to create the speculator. So there's that process which you need the real model weights for. And then there's of course just the process of, standing up all the infrastructure behind it, loading all this stuff, testing it. And then when there's a new model with a newer architecture, I think that, like, the DeepSeek models tend to be the most challenging as they have, like, the most novel architectural stuff going on, model after model. But every new model has something. Kimi K2 had. Oh, sorry, GLM-5.2 hadAli [00:14:53]: Sparse attention.Philip [00:14:54]: Yeah,Ali [00:14:54]: YeahPhilip [00:14:54]: the DSA.Ali [00:14:55]: Right. Which is brought from DeepSeek.Philip [00:14:57]: Yeah. AndAli [00:14:59]: So you can copy-paste then?Philip [00:15:01]: It kindAli [00:15:01]: I don't know how this works.Philip [00:15:02]: So, like we had to, like, build support for that into our runtime. And you're right, like it is really interesting the way that all of these open source labs borrow from each other. For example, like GLM-5.2 doesn't have vision. So something that, Haley, a guy on our team, if we could take a look at this, he, like, grafted the Kimi vision encoder onto GLM-5.2.Retrofitting Vision into GLM-5.2Ali [00:15:27]: We'll be training the projector.Philip [00:15:28]: Exactly. So if you think about, like, the encoder, there's the encoder, which is the part that looks at the image and turns it into latent information, and then there's the projector which likeAli [00:15:38]: You can say latent space. It's okay.Philip [00:15:41]: And then there's the projector that maps it onto, the model itself, and then there's the model weights. You don't wanna mess with the model weights because you run a chance of making the model dumber at something else for the purpose of giving it vision. So instead, Haley started with just a projector, which is only a handful of millions of parameters.Ali [00:16:02]: That would be, yeah.Philip [00:16:02]: Yeah.Ali [00:16:03]: Can you show the training one?Ali [00:16:04]: Like the way it groksPhilip [00:16:05]: YeahAli [00:16:06]: Very interesting.Philip [00:16:06]: And maybeAli [00:16:07]: That right therePhilip [00:16:07]: Maybe Ali, you should take it from here. You've got a betterAli [00:16:10]: Ooh, double the sandPhilip [00:16:11]: Understanding of this than I do.Ali [00:16:11]: Yeah. You can see, like, he. The way he trained this is really cool. At the beginning, he was training it using just like, “Here's a picture of a mountain. Can you describe what's in this mountain?” And that caused it just like the first, learning walls. Like here you can see this all we're trying to teach it is to translate the encoded. Like it's already taken the encoder from Kimi K. It's taken the image. It'Philip [00:16:31]: Yeah. FrozenAli [00:16:31]: FrozenPhilip [00:16:32]: With adapter.Ali [00:16:32]: Exactly.Philip [00:16:33]: Yeah.Ali [00:16:33]: So the brain is frozen and the eyes are frozen. It's just we're tryingPhilip [00:16:37]: AlignAli [00:16:38]: Interconnect between the eye and the brain, right? So the projector. And so you take the tokens and then he's like, “Oh, can you describe what's in this image?” And he's like, “Oh, it's a mountain,” or it's a person or it's a human, whatever the case is. But that didn't cause complete understanding. So he changed it such that every image was associated with a data set of questions. Like, does this image have a white male? Does this image have birds in the top corner? Does this image have a scientist in it? All of that stuff. And it would have to answer questions correctly. And using not just training on describing an image, but being able to answer question, another question, answer over time. Like you can see the grokking, which is like genuinely insane, that retrofitting vision into a large LLM can learn to that extent. And even for images that it doesn't perform well on, for instance, if you ask it a picture of like Stephen Hawking, “Who is this?” Maybe it doesn't get it, but it will say something like, “This is Albert Einstein.” Like it still understandsPhilip [00:17:25]: Close enoughAli [00:17:26]: That this is a scientist who is a man who has, some significant achievements, all that stuff. So that's like really cool.Philip [00:17:32]: Yeah. So, we've covered Hao Tian before, who the author of the LLaVA paper that did this, a while ago. And I think that's very foundational work for anyone who hasn't done vision work before.Ali [00:17:41]: Same with the CLIP and MetaCLIP, where you go from just captioning to building out questionsPhilip [00:17:47]: RightAli [00:17:47]: Off the image and how much better you can get performance.Philip [00:17:50]: Right. Right. Right. Yeah. But what's, what's so exciting about this is if you look at a model like this. Now, this is a little bit more of a research project. It's not. It got to 56% on MMLU Pro, I think. So not quite frontier. But if you're running this model, you haven't suffered any loss on your GLM-5.2 quality. If you don't have an image, it'll just behave exactly the way it used to. And ultimatelyAli [00:18:14]: Which in the inference code you literally do not include the other part, right?Philip [00:18:18]: Yeah. You would just skip the encoder if you don't have an image input.Ali [00:18:22]: Okay.Philip [00:18:22]: Just confirming.Philip [00:18:23]: YeahAli [00:18:23]: Does it affect a lot on the overall inference side? Like you're not adding much, you're adding a very small vision encoder. These are typically likePhilip [00:18:30]: They're super fineAli [00:18:31]: Less than a billion parameters, right?Philip [00:18:32]: Yeah. It's, - There's a little bit less standardization among vision encodersSwyx [00:18:37]: YeahPhilip [00:18:37]: So the support matrix can be a little bit, sparser. But overall, yeah, it's a pretty, it's a pretty minor component of the overall system. And ultimately what you get out of the system is all of a sudden you have Kimi Vision, GLM weights, and DeepSeek attention all in one model.Open Source Model Grafting and Franken-MergesPhilip [00:18:56]: And that's, I think, a lot of the power and beauty of open source, is that you can take all of these different components and combine them together into a system that's better than anyoneSwyx [00:19:05]: YeahPhilip [00:19:05]: Can be individually.Swyx [00:19:06]: People used to say that you would also do Franken-merges where you would take likePhilip [00:19:10]: YeahSwyx [00:19:10]: Layers from each model.Swyx [00:19:11]: Does anyone do that anymore?Ali [00:19:13]: Well, to your point previously when you were mentioning like, the work that goes into supporting a model when it first comes out, like GLM-5.2 or MiniMax M3 or whatever the case is. Sometimes you do have to like, you do have to switch out some things. Like, for instance, the MiniMax M3 head uses full attention, and with full attention you end up with this like insane bottleneck in spec dec ‘cause you're doing auto-regressive token generation for three tokens, and you're doing this like N squared over all of the tokens that are in your sequence. Your KV cache is like very large because it's not sparse, it's not top K. So we find it better to like, okay, we're gonna replace this, we're gonna replace this layer with a layer from another model that's using like GQA, for instance. And then just with the right training, you can get it to have the same acceptance rate. So it is very possible to retrofit layers from other models and very much needed. If a layer is like inefficient, the training just becomes the challenge, like how do you ensure that you train it properly? Which again to your earlier point is like the mesh between training and inference. As in like you need very good training in order to do fast inference. That's like, I feel like more and more becoming true.Swyx [00:20:21]: Yeah. Anything else on the support side when you say like get it to fully production ready?Loop Detection, Race Conditions, and Non-DeterminismPhilip [00:20:26]: Yeah. I think that there's also a question of just, we can test a model to a pretty extensive degree, but we're trying to get it out quickly and then you see a bunch of other people test it and you get interesting results. There was an issue with, GLM briefly where we had some like mode collapses where it would just output the same token over and over again for certain prompts on certain temperatures. Like once you expose an endpoint to the real world, there's going to be, so many more varieties of things given to it that you're able to, discover and patch things. So it's not just a, day zero process, it's then like for the first week, for the first month, if a model remains popular, like how do you both fix bugs and then continue to push the envelope on performance?Ali [00:21:21]: What do you mean you don't want your model outputting S?Swyx [00:21:24]: Is there loop detection on that stuff, by the way? It still happens like quite a lot, which is surprising.Ali [00:21:30]: We have like we, in our endpoint, like if a model was to output the same token like four plus times, we just cut the generation. We say like, “Oh, sorry, this-- Like try again,” or like we will reprocess the request. ‘Cause we know then, like if it, like if, yeah, it's four times the same token, it's probably collapsed.Swyx [00:21:45]: Yeah. Is there a way to opt out in case I really want that?Ali [00:21:48]: You want that?Ali [00:21:50]: I think there's a way that we have to handle it. I'm not exactly certain, but I feel like in certain models, like when they output something like you can imagine, like a table for instance, and so they want, they wanna draw like 12 dashes and 12 dashes. Yeah, I think there's a way for that to happen. I think we only do it on certain tokens. Like we exclude certain special characters.Swyx [00:22:07]: Yeah.Ali [00:22:07]: So we only do it on like certain like S is the most common almost. GLM-5.2Swyx [00:22:11]: OhAli [00:22:11]: And I think it was DSV 4 as well. Like you'd just have like looping issues where like you literallySwyx [00:22:17]: ItAli [00:22:17]: Just have like S.Swyx [00:22:18]: Yeah. Is there a special, something special about S? No, just randomlyAli [00:22:21]: It just seems to be the one token involved.Swyx [00:22:23]: Yeah. And it'Philip [00:22:24]: Is thereSwyx [00:22:24]: And it's only temperature 0Ali [00:22:27]: NoSwyx [00:22:27]: Even at other temperaturesAli [00:22:27]: Even at like 0.9 or whatever, it will still, it will still collapse.Swyx [00:22:30]: That's weird, right?Ali [00:22:30]: It's, it is an inference problem to be honest, like a software problem. Like oftentimes, the image you run will-- like NVIDIA will release an image for instance, and if we will upstream the changes from their latest TensorRT-LLM image into our stack, we'll find that it fixes it. Or oftentimes this will only happen in an inference engine that you're using like SGLang. But if you were to switch to vLLM, that isn't the case. So it seems to be like an extremely like deterministic software issue and not really a model issue. It's not like a weights problem. Like I'- we'll say like, “Oh, it's a problem with the quant. We did PTQ wrong,” right? But that isn't, that doesn't make sense because the same weights used with a different inference engine does not repeat the problem. And sometimes it's, the kernels that are being used in the backend have like these very subtle sometimes race conditions, where if you were to use this model hosted on one cluster, you will never get this problem.Swyx [00:23:19]: Oh my God.Ali [00:23:19]: But if you host it on a different cluster, you will. And the reason is the KV cache transfer from a node to node in that one cluster is using a slower interconnect than the node to node in another cluster. So that exposes the race, whereas in another cluster it doesn't. So then you end up just like, okay, this model is not gonna be hosted on this cluster. We're gonna host it on, another cluster because that cluster exposed that problem. But then it ends up with like, okay, is it the software? Is it the model weights or is it the hardware?Swyx [00:23:42]: There is a thing about this with temperature 0 still not being deterministic, right?Ali [00:23:46]: Right.Swyx [00:23:46]: Mostly because of hardware. Even at temperature 0 same model, you won't always get the same output.Swyx [00:23:52]: Even-- But I'm surprised by the race condition one because, I thought PyTorch was a graph that like guarantees that you at least, execute things in the right order.Ali [00:24:02]: Well, yeah, true. Like I'm not, I'm not saying that there is. Like well, you have things like PTL optimizations where like you can start a kernel before the end of the previous kernel, and that's like ‘cause you want to do that because there'sSwyx [00:24:12]: It's like pipeliningAli [00:24:12]: Expense. Exactly.Swyx [00:24:13]: Yeah.Ali [00:24:13]: But it'- But you don't do it cleanly. Like you overlap a little bit of the execution. No, it is very possible that the kernel itself, like that one block that is supposed to be running in this instance of time, that kernel itself has a race condition. For instance, like a missing barrier. Like often if you're designing a kernel and you want it to make it to be very fast, if you don't test it extensively, you'll, you'll have certain threads access data points from registers before they've been written to by other threadsSwyx [00:24:36]: YeahAli [00:24:36]: For example, because like your barrier is wrong or your synchronization was wrong. But yeah, like the testing itself is very difficult in those like, andSwyx [00:24:42]: And there's no like borrow checkerAli [00:24:45]: What does that mean?Swyx [00:24:46]: Like Rust. Like the. If you're trying to have like memory safety It sounds like a comparable problem.Ali [00:24:52]: Well, yes, but you're working in CUDA, right, NVIDIA GPUs. Like- You just need a higher level language like modular Maybe that's what modular is supposed to do. I don't know.Quantization Quality and Vendor FidelityVibhu [00:25:00]: How do you see keeping quality of the model? So you talked about all these steps of, okay, you gotta do quantization, train your own speculative decoderAli [00:25:07]: RightVibhu [00:25:07]: Run on different hardware. Looking at other model providers, okay, you kicked off a inference speed race on the consumer end. What goes into keeping quality the same across them, right? Sure, you can run benchmarksAli [00:25:22]: YeahVibhu [00:25:22]: But, like, how do you determine how much quantization are there standards? What goes intoPhilip [00:25:27]: There's a few things on quality. Most inference optimizations are lossless. KV caching, for example. You are just recomputing or preventing recomputing the same values. Speculation, of course, if a draft token is wrong, it gets rejected. The main lossy optimization is quantization. And that really comes down to, number one, data format, number two, which parts of the model you choose to quantize, which layers, and number three, like doing a lot of calibration on the quantized weights, to ensure that you're preserving all the outliers. There's other tricks that you can do, though. A big one is long context, ‘cause one thing you asked at, right at the beginning is, “Oh, what's gonna happen if I send a 200,000 token request in?” So with a long input sequence, you need to, store a lot more information. You need to process a lot more tokens. And so even if a model has a context of a certain length, you might, as an inference provider, choose to build an API with a shorter context length, and of course a full length one as well. Because if someone doesn't need the full million token context, for example, you can get them better performance. I don't know if that's exactly like quality of the model. The way that I think about quality is to what degree are we faithfully serving the original model? If you think of a golden implementation of a model that performs exactly the way the model is designed to perform, I think of quality as how close are we getting to that, 100% fidelity of the model.Philip [00:27:13]: You can also, of course, think about quality from the training side and how do you push yourself past 100%. But when I think about purely inference optimizations, it's getting faster while staying as close to that 100% fidelity mark as possible. And certainly our standard internally is that, like you should not be able to tell the difference between our API and a, official API. I think Kimi in particular does a good job of vendor benchmarking hereAli [00:27:41]: YesPhilip [00:27:41]: Where they haveAli [00:27:42]: They released an actual vendor benchmark.Philip [00:27:43]: Exactly, yeah.Ali [00:27:44]: ‘Cause they accused, some people, Amazon? There was some provider that was not doing very well on Kimi's benchmark.Philip [00:27:50]: Yeah.Philip [00:27:51]: So, with Reflect we probablyVibhu [00:27:52]: This was a long time ago, right?Philip [00:27:54]: No.Ali [00:27:54]: Yeah, like threeVibhu [00:27:55]: They alsoAli [00:27:55]: Four, five months agoVibhu [00:27:57]: This also happened with, I don't remember which model, but they pulled out quite a few, and then they started a whole chart about this. It might have beenPhilip [00:28:03]: Kimi Vendor Verifier.Ali [00:28:04]: Yeah.Philip [00:28:05]: Yeah.Ali [00:28:05]: Yeah, ‘cause you, ‘cause you'd be pissed, right? Like if you'Philip [00:28:07]: Yeah.Ali [00:28:07]: If like if I'm a consumer and I'm using like Amazon's endpoint for instance, and I've used Kimi and I'm like, “Oh my God, like this is bad,” I'm not gonna say, “Oh, Amazon quantized the model in a bad way.” I'm gonna say, “Oh, Kimi sucks.” Right?Philip [00:28:17]: Yeah.Ali [00:28:17]: So it seems like that makes sense.Philip [00:28:19]: Yeah, they care. They care.Vibhu [00:28:21]: Justifiably.Ali [00:28:21]: Yeah, justifiably.Vibhu [00:28:22]: This is probably a stupid question, but just checking, has anything improved from main quantization?Philip [00:28:28]: Yeah.Vibhu [00:28:28]: Like, is quantization always strictly worse?Ali [00:28:30]: Well technicallyVibhu [00:28:32]: NoAli [00:28:32]: It's a lossy. QuantizationPhilip [00:28:33]: YeahAli [00:28:33]: Is a lossy, it's a lossy implementation.Philip [00:28:36]: Speed improvesVibhu [00:28:36]: Speed improves.Ali [00:28:37]: It the number, likeVibhu [00:28:38]: No, I' always look for inverse scaling laws.Philip [00:28:40]: Yeah.Ali [00:28:40]: Yeah.Vibhu [00:28:40]: This is something I learned from Noam Brown, where like things that normally act in one direction sometimes do.Philip [00:28:45]: Well, technically when you run a benchmark, because these models are deterministic, sometimes your,Ali [00:28:52]: YeahPhilip [00:28:52]: NVFP4 quant is like, two basis points higher than yourAli [00:28:56]: No, it's noise. It's noise.Philip [00:28:57]: Yeah, exactly. I'm like, yeah, it's, it's within. That's why I always say within margin of error.Philip [00:29:01]: And I stopped saying that because everyone assumes that what is, well, within some margin of error, we're barely inside of that to the worst, so we're saying. But yeah, sometimes it's just like, gives you a higher output score. But like Ali said, that's noise. To my knowledge, you're not necessarily making the results better. You're just trying to, again, like keep your fidelity as close to 100% to the original model.Layer Selection, KL Divergence, and Better QuantizationAli [00:29:27]: There is, to your point, research that we did on MP. I don't know if you are able to pullPhilip [00:29:31]: YeahAli [00:29:32]: A tweet we did. One of our research interns, Joshua, I think it's a tweet on how we have 20% better quantized GLM-5.2 than NVIDIA. Essentially what we found throughout like this month research is, okay, quantization is a lossy. It's. You're compressing the data from, occupying 16 bits to occupying, four bits, for instance. And so you're losing some information, and you're trying to minimize that. And so when I say that I'm gonna quantize the model, my job becomes how do I find the layers that I can quantize, and how to find the layers to not. For instance, with image models, I don't quantize modulation layers, and I don't quantize out projections because those two are. Like out projection is what you see as the user. Modulation is what the model sees or understands. Right, exactly. And so to his paper, do you have the. It doesn't have the. Yeah. It's a long paper. I don't know if I can findVibhu [00:30:25]: If there's a part to search or it's probably in the thread.Ali [00:30:28]: It's probably in the thread.Vibhu [00:30:29]: Yeah.Ali [00:30:29]: But the long and the short is it is very possible that quantizing more of the model makes the results. Like if I have a model that I quantize layers one, five, and 10, and another model where I only quantize layers one and It is possible that the model in which I quantized more information is going to perform better because the quantization errors have canceled out. And so what Joshua showed in his mathematical proof where he had like a verifier in, is that you can predict which layers are going to have quantization errors that will cancel out with each other, and you choose to quantize those layers. And so the result of doing this mathematical quantization is you end up with a model that's 20% more quantized than another provider, so you get 20% more throughput of it because there's more layers than running an NVFP4, and your quality is better than that other quant because the layers that you chose to quantize have their errors cancel out, like one layer skewed to the right one layer skewed to the left, one layer skewed to the right. Your final logits distribution is more similar to the original distribution of the model, so you have better fidelity. And so the way we proved this was with KL divergence. So instead of just scoring on the benchmarks, we scored the KL divergence between the logit distribution of the quantized model and the logit distribution of the original full precision model, and we showed that with this technique we get. If your probability distribution on the logits which token it wants to select is more of the same as the original model, you're probably gonna end up staying true to the original model. So yeah, so it seems like previously before this, it seemed like the industry was, well, the more you quantize, the worse it's gonna be, ‘cause the more loss you introduce. That's not exactly, not necessarily true. So yeah, doesn't improve it, but can cancel out.Philip [00:31:57]: I think it might be this, but reminds me a good bit about pruning where you can prune off certain layers.Philip [00:32:03]: But very interesting. Didn't know this was a whole paper you guys put out.Ali [00:32:06]: It's. Fun fact, it was originally 72 pages, this paper, and then we decidedPhilip [00:32:11]: WowAli [00:32:11]: We can't tell. We couldn't release it. So it's now 45.Swyx [00:32:15]: Still 39 pages, so very substantive. We talked about evals and all these things and, like what's possible in terms of speedup? Like it's like probably like the numberInference Speedups and BenchmarkingSwyx [00:32:25]: Thing that people do wanna care about, and it's something that you wrote about in your post. Like official API is 70 tokens per second, and you push it up to 90. Is that like a normal thing?Philip [00:32:36]: So what's cool about working in inference, the reason that I think inference is going to be a useful place to do engineering for a long time, is that if you look at highly optimized domains like, say, finance, if you're in finance, you measure how much better you got in basis points. It's like, “Oh, I got five basis points better, like twentieth of 1% better,” that's huge news because everything is so optimized. When we publish optimizations, it's 20%, it's 100% it's 200%. So there's still probably like a lot further to go, honestly. Like you'll, you'll know that inference is pretty much solved when researchers start publishing about how they got 1% faster at something.Swyx [00:33:19]: Which by the way, because I am from the finance background, in the ‘70s, that was the margin at the time. When you did quantitative finance research, you would findAli [00:33:27]: And like 20%, tens of percent.Swyx [00:33:29]: That's. Yes.Philip [00:33:29]: Yeah.Swyx [00:33:30]: And now it'Philip [00:33:31]: Tiny fractionsSwyx [00:33:32]: For those people interested, look up Andrew Lo's paper. He had a really interesting illustration of quant, stat arb, distribution, narrowing down from like those kinds of 20% differences in the ‘70s, down to nothing today, which is very cool.Philip [00:33:48]: Exactly, and we're at the beginning of the same type of thing. Now benchmarking is hard. I think anyone will tell you that, and benchmarking provider speeds is hard because there's so many variables that go into it. What hardware are you using? How much load do you have on the system? What's the exact nature of the prompts and input and output sequence lengths? All that stuff. But overall, when you start stacking these improvements, you're looking at multiples. You can look at it. The most common form, of course, is TPS, tokens per second, which is bad naming by us in the industry, ‘cause there's two tokens per second. There's tokens per second, the throughput number, and the latency number.Ali [00:34:31]: TTMT, yeah.Philip [00:34:32]: Like total tokens per second out of the, out of the GPU as a throughput number. Most people only care about tokens per second as the latency number, which we should call ITL, intertoken latency, but we don't.Philip [00:34:44]: Anyway, so you can imagine a standard API without many optimizations for a 1 trillion parameter model operating somewhere in the 30 to 50 tokens per second range for reasonable traffic profile. And we generally see the goal of, pushing to 10X that. But, not necessarily day zero, but by stacking enough optimizations, if you have, say like four optimizations, each of which doubles performance. Or sorry, three optimizations, each of which doubles performance, then you stack that up, that's an 8X gain. That's the order of magnitude that we're working with in this space. We're trying to make things substantially faster, not just go from like 70 to 90.Swyx [00:35:38]: Are you saying you've. You have done that?Philip [00:35:40]: So let's say you have as a reasonable baseline, 30 or 40 tokens per second. You can achieve 10X that. So like on GLM-5.2, if you run it unquantized, perhaps on H100s even, and you're just using an off-the-shelf inference engine with no particular optimizations, no speculator, nothing extra around like KV routing, no disaggregation, you're, you're probably, yeah, looking at that like 30 to 40. You think that's like a reasonable baseline?Swyx [00:36:12]: Right. Right.Philip [00:36:12]: To get to something like 10X, there's a lot of trade-offs that you're making. If we're running at more like a 300, 400 tokens per second range, you are using the best hardware possible. You have a optimized speculator. You have done all of your quantization work. You are Seeing a pretty high cache hit rate. You are running with a reasonably small batch size and a parallelism configuration that is tuned for latency versus throughput, but it is possible. So the spreads that you see if you, like, go on artificial analysis or you go on OpenRouter and you look at, the worst provider to the best provider, oftentimes can hit that range. 10X is of course very aggressive. It's oftentimes maybe more of a four to six times improvement. But that's the performance that makes us really excited, is when we can get these huge gains, not just go from 70 to 90 tokens.Stacking Optimizations: NVFP4, Speculation, and DisaggregationAli [00:37:19]: It's also, like, hardware dependent. Like, ifPhilip [00:37:20]: YeahAli [00:37:20]: If you have a thing where you're serving it on just, like, a node of H100s and then you throw, like, you shard the model across, like, four nodes of B200s. Like, you can definitely increase the speed with just throwing more hardware at it. Like, normalizing for the same exact hardware and the same number of GPUs.Philip [00:37:35]: Yeah. Then you're looking at, like, a two to 4X improvementAli [00:37:38]: Right. RightPhilip [00:37:38]: Depending on the inference optimizations. So yeah, it's. Some of it's, what's the call, and some of it's who's the driver.Vibhu [00:37:46]: If you break down the two to 4X, say the example is run GLM-5.2Ali [00:37:51]: YeahVibhu [00:37:51]: On B200sAli [00:37:53]: YeahVibhu [00:37:53]: Single node, right? What's, like, the cost trade-off for effort to get, like, the last bit of juice out versus what should people just think of, right?Ali [00:38:01]: Spectre quantization. Yeah.Vibhu [00:38:03]: Spectre quantization.Ali [00:38:04]: That's, that's, that's like 95%. LikeVibhu [00:38:06]: And how far does that get you? And how easy is that for the average person to do? So say right I wanna throw the weights of GLM-5.2 on a node of B200s, how easy is it to find speculative decoder- decoder model or already quantized model? How much work goes into it?Philip [00:38:23]: If you're doing it up front, it's quite a lot of work. If you're doing it today, there's going to be people who have published things that you can just, you can just grab some NVFP4 weights. You can grab a speculator. Yeah, if we're thinking about, like, what are the 2Xs we're stacking, going from, BF16 to NVFP4 is, it's not quite a 2X, right? It's like. I think it's about, like, 30 to 40%, from 16 to 8, and then another 30 to 40% multiplied from, 8 to 4. So that doesn't quite get you a 2X, but, like, roughly a 2X. Speculator, roughly a 2X. Disagg on top of that if you're able to get enough hardware and put enough traffic through it, another roughly a 2X. And then you add in some, double-digit percent increase from having just a better runtime with, the latest kernels and stuff behind it. And that's how it stacks up.Ali [00:39:21]: YeahPhilip [00:39:21]: So building each of those, like, building the, quantized weights is, for someone who really knows what they're doing, hours to days of work. Building the speculator, again, like, hours to days of work. And the, disagg setup, hours to days. Well okay, but like once you haveAli [00:39:39]: Once set up. Once set up. YeahPhilip [00:39:40]: Yeah, getting disagg working for the first time, I'm saying, of course, is very difficult.Philip [00:39:44]: The marginal implementationAli [00:39:48]: Like, if you're just grabbing, like if you are a person, like just a normal consumer who has access to, like, a node of B200s and you're wondering, “How can I just host it myself?” You don't need to quantize the model yourself. There's always gonna be, like, an open source quantized checkpoint. NVIDIA's gonna push one out if no one else does. You. Usually, the providers will have their own spec dec that they've trained as well. You don't need to train your own spec dec. You can just use that as well.Philip [00:40:09]: Yeah. Like, GLM-5.2 has its own MTP.Ali [00:40:13]: Right. Right.Vibhu [00:40:14]: What's multi token prediction?Philip [00:40:15]: Yes.Ali [00:40:16]: I'm justVibhu [00:40:16]: Can you explain that?Ali [00:40:16]: I'm just an expert.Ali [00:40:18]: I can do it for you in case I get it wrong?Vibhu [00:40:20]: No.Vibhu [00:40:21]: Yeah, you should correct if we're wrong, but their multi-token prediction can be used for self-speculative decoding.Ali [00:40:27]: I'm not sure. I'm not gonna correct that.Vibhu [00:40:28]: Okay. I'm semi-confident in thatAli [00:40:30]: Okay. YeahVibhu [00:40:30]: But someone can check. But it's useful to paint the story of, okay, not just the average person, but say a company wants to switch from serverless inference I wanna throw this up on. I wanna rent some GPUs, throw it up. These are the steps you take to do significantly faster than just put it behind vLLM.Ali [00:40:48]: Right.Vibhu [00:40:49]: I was waiting for a mention of Dynamo.Vibhu [00:40:51]: I feel like, that's supposed to be the baseline that you measure against.Dynamo, KV Routing, and Disaggregation ToolkitsPhilip [00:40:55]: I would think of Dynamo as less of a box system and more of a toolkit for building with. So when we talk about doing aware routing, when we talk about doing KV offloading, when we talk about doing, PD disaggregation, Dynamo fundamentally is. By the way, Dynamo is an open source library from NVIDIA.Ali [00:41:17]: We've done a pod with KylePhilip [00:41:18]: OkayAli [00:41:19]: Kyle Cranin.Philip [00:41:19]: Cool. So then your listeners know then that it supports all the different inference frameworks. And it is multi hardware, which is interesting.Ali [00:41:28]: But it's just a router, it's not like an optimizer layer.Philip [00:41:30]: Yeah. All it does, like, what Dynamo is good at, it is a library for moving information around your cluster, around your hardware. So if you have, KV cache on one place and you need it to be somewhere else, Dynamo coordinates NIXL for you to move that around.Philip [00:41:49]: That doesn't mean that, like, out of the box, you just say, “Pip install Dynamo,” and then you get, like, a massive performance speed up. It's more of a developer toolkit.Ali [00:42:01]: Yeah. I would have said it would. It comes with a set of defaults that you can then swap out.Philip [00:42:06]: It does. If the industry at large, I think, was, like, rolling out all of these deployments, standard, then I think it would be, like, a credible baseline. But, we've got to, we've got to benchmark against, like, what we're seeing in the wild.Speculative Decoding Methods: Medusa, EAGLE, n-Gram, and Spec-SpecVibhu [00:42:23]: I did wanna talk a little bit more about PD disagg, because that is probably, like, number three after quantized and speculative decoding. In your book though, I was just gonna pull out the book.Philip [00:42:31]: Yeah.Vibhu [00:42:32]: Like section 522 on Medusa, 523 on EAGLEPhilip [00:42:35]: YeahVibhu [00:42:36]: 524 on gram.Philip [00:42:37]: It's 55, would be disaggregationAli [00:42:42]: Yeah. Well, no, I just wanted to dwell a little bitPhilip [00:42:44]: YeahAli [00:42:44]: The other. Like, so what do you choose to include? What do you choose to not to include? Because there was all these other techniques.Philip [00:42:51]: Yeah.Ali [00:42:51]: Are these still relevant? Because I think they came out, like, a year and a half ago maybe.Vibhu [00:42:55]: Medusa is quite old.Philip [00:42:56]: Yeah, Medusa's old.Ali [00:42:58]: It was old.Vibhu [00:42:58]: But is it in the book as a good, here'sPhilip [00:43:01]: BaselineVibhu [00:43:01]: Baseline vanilla understand it?Philip [00:43:02]: Like you should know this.Vibhu [00:43:03]: Like I read the paper, I'm like, “ it makes so much sense.”Philip [00:43:05]: Yeah.Philip [00:43:05]: So with the book, I had a couple goals. One was to give people just a working vocabulary for the space as a whole, and the other was to give them some intuition about how each of these techniques works. As I mentioned in my AI Engineer talk, which is the first public addendum to this, the speculation space has moved much faster than everything else. So yeah, even at the time that I wrote the book Medusa, I very much included as a way for people to understand how the space evolved rather than what the most modern technique is. And now of course, there's DFlash, dSpark. There's, there's newer techniques even than EAGLE, although EAGLE is still very commonly used.Ali [00:43:51]: SpecSpecta.Philip [00:43:52]: Yes. Speculative decoding.Vibhu [00:43:54]: What canAli [00:43:56]: Oh, it's a paper by Tri Dao and it's like, it's doing speculative decodingVibhu [00:44:00]: HuhAli [00:44:01]: For the speculative decoder.Philip [00:44:02]: Oh, in spec- oh my God.Ali [00:44:02]: It's literally just an another. It's like, yeah, that's the most simple way to explain it, and it seems like he got trivial speed ups there. But it seems that the complexity with training, it's almost like in our mind at least, it's almost as complex as training GANs. Like it's like a very delicate balance and oftentimes you, it's just but yeah, it's literally speculative decoding on speculative decoding.Vibhu [00:44:21]: Speculative.Ali [00:44:22]: Yeah. We saw this paper.Vibhu [00:44:24]: It's interesting, right?Ali [00:44:24]: Yeah.Vibhu [00:44:24]: I wouldn't even expect it to be very particular to train, I wouldAli [00:44:29]: Right.Vibhu [00:44:29]: The naive part of me is like, okay, train speculative decoder.Ali [00:44:32]: But like, and it makes sense, like the whole idea of speculative decoding is you. It's like, it's like almost like the iPhone auto predict version but for a normal model, right? Like you're just, you're just, generating three tokens and you're like, okay, I'll do prefill on them. And so you save those three turns for your original model. Now your speculative decoder is doing three turns of auto regression, so why not just have an even smaller model?Ali [00:44:53]: The other question there is what are the size of speculators? So say forPhilip [00:44:58]: Right. It's like a billion parameters.Ali [00:45:01]: Like for MiniMax, it's. Yeah. It's like one layer. It's like one 60th of the original model usually.Philip [00:45:06]: Yeah. I think we should do a paper when we get back to the office.Philip [00:45:10]: SpeculativeAli [00:45:11]: SpeculativePhilip [00:45:11]: Decoding.Ali [00:45:13]: No, it's, it does seem like how, when do you stop? But then it also seems like if you're able to train spec-spec decode for instance, right? Like if you're able to have a small model that is accurately predicts what the intermediate speculator is gonna predict, that is able to predict what the original target model's gonna predict, then why not just use that smallest model directly, right?Vibhu [00:45:34]: Yeah. This isAli [00:45:35]: Like it seems likeVibhu [00:45:35]: Adjacent to the routing problem.Ali [00:45:36]: Right.Vibhu [00:45:36]: Yeah.Ali [00:45:36]: Right.Philip [00:45:37]: The thing with speculators is one of the practical constraints on using them is that you do have to run a small model on the same hardware that you're running the big model on. There is a orchestration and resource competition problem inherent in that, and that is one of the constraints on speculation in general, is that draft tokens cost resources to create and cost software complexity to manage. And so if you have like infinitely recursive speculators, you add in quite a bit of that complexity on the actual implementation within the inference engine as well, not just in the training process.Vibhu [00:46:17]: I was gonna say, I would wonder if you could do similar, like distillation and pruning of, it's the same thing, it's just a model. Can we not just distill a lot of the weights, quantize the speculator, out of my domain? The question that also comes up is, this is all for big server workloads, right? How much of this applies to, say I have this MacBook, I wanna run Gemma really efficiently. Similar problems, not the same?Local AI vs. Data Center InferencePhilip [00:46:45]: Pretty different. I talked to Selo, about this on his podcast a couple weeks ago. The difference between inference engineering for the data center and for production workloads versus inference engineering for local AI, is that we start with fundamentally like different constraints and different goals. With local AI, it's how do I fit this model onto my hardware and then make it less dumb? And with data center influence, it's how do I load this model and then make it less slow? And we care about less dumb, and they care about less slow. But the local AI inference engineering ecosystem, I think has a lot for us to learn from in the data center space. They are experts in various forms of quantization, including dynamic quantization that we just don't touch, in the pruning, in the distillation, in the, layer removal. There'Ali [00:47:42]: Layer removal matters less.Philip [00:47:43]: Yeah. There'Ali [00:47:44]: No one loves pruning really.Philip [00:47:45]: Yeah. Well, but the, but they doVibhu [00:47:46]: Which is surprising, right? But that's, that's a whole different thingPhilip [00:47:48]: Just to fit something on the laptop.Ali [00:47:50]: Right.Philip [00:47:50]: So yeah, it's a, it's an interesting, it's an interesting space. Not necessarily that like their techniques make sense for us to do in the data center, because we have different resources and different goals, but more that the process as well as the openness of that field is something to, admire.Ali [00:48:12]: Yeah. Like to your point, like, certain optimizations that would. Like for instance, Turbo Quantum Sharper, like it made such huge hype on that and we did like a whole deep dive on Twitter and like said, what is it? How does it work? Why is it good or not? And it took off and it was implemented on local devices because your memory bandwidth is so slow on like a MacBook, for instance. But try putting the same thing on like an NVIDIA GPU on a B200 Turbo quant would not be. Like, it would not be used. Like, NVIDIA - Like, NVIDIA made it clear that this is not a good optimization, and we've seen it firsthand where the overhead of doing dequantization, quantization of, in the kernel itself with turbo quant kernel, each end is much slower than the time that you save from doing the bandwidth. ‘Cause on the B200s, you have like 3.5 terabytes per second. You don't need decrease the storage that much. You don't need to do, FP4 KV cache. You don't need to use a requant. There's, there's, there's better optimizations to be made. But on Edge devices, it's extremely important, it's extremely useful. So, seems to be, like, different optimizations there, but then they're all uniquely combined with like all you wanna quantize the model, you wanna do speculative decoding, like certain common prefixes with bothPhilip [00:49:18]: Principles.Ali [00:49:19]: Yeah, exactly. Exactly. Exactly.Philip [00:49:20]: They also do a lot of work on, model parallelism, especially over, heterogeneous topology, where you have, some sparks and they are wired together with, Ethernet, DGX sparks.Ali [00:49:35]: Yeah, this is the Exo Labs guys.Philip [00:49:36]: Yeah. You have, a nu

Nutty Bites
Godzilla (1954)

Nutty Bites

Play Episode Listen Later Aug 3, 2026 7:17


In anticipation for Godzilla Minus Zero, Tek and I watched the original Godzilla and were in for a big surprise. Welcome to dog days of podcasting for 2026, where podcasters from around the world do a podcast a day for the … Continue reading → The post Godzilla (1954) appeared first on NIMLAS Studios.

The Aubrey Masango Show
Please respond with the quotation on the company letterhead bbbeee tax sbd form psira certificate spec is attached for

The Aubrey Masango Show

Play Episode Listen Later Aug 3, 2026 45:55 Transcription Available


Siyabonga Motha speaks to Dr William Mpofu, a political analyst about the local government elections and whether the opposition is doing enough to prove to South Africans that they are ready to govern. Tags: 702, Aubrey Masango show, Aubrey Masango, Bra Aubrey, Political Analysis, Siyabonga Motha, Dr William Mpofu, ANC, EFF, DA, MKP The Aubrey Masango Show is presented by late night radio broadcaster Aubrey Masango. Aubrey hosts in-depth interviews on controversial political issues and chats to experts offering life advice and guidance in areas of psychology, personal finance and more. All Aubrey’s interviews are podcasted for you to catch-up and listen. Thank you for listening to this podcast from The Aubrey Masango Show. Listen live on weekdays between 20:00 and 24:00 (SA Time) to The Aubrey Masango Show broadcast on 702 https://buff.ly/gk3y0Kj and on CapeTalk between 20:00 and 21:00 (SA Time) https://buff.ly/NnFM3Nk Find out more about the show here https://buff.ly/lzyKCv0 and get all the catch-up podcasts https://buff.ly/rT6znsn Subscribe to the 702 and CapeTalk Daily and Weekly Newsletters https://buff.ly/v5mfet Follow us on social media: 702 on Facebook: https://www.facebook.com/TalkRadio702 702 on TikTok: https://www.tiktok.com/@talkradio702 702 on Instagram: https://www.instagram.com/talkradio702/ 702 on X: https://x.com/Radio702 702 on YouTube: https://www.youtube.com/@radio702 CapeTalk on Facebook: https://www.facebook.com/CapeTalk CapeTalk on TikTok: https://www.tiktok.com/@capetalk CapeTalk on Instagram: https://www.instagram.com/ CapeTalk on X: https://x.com/CapeTalk CapeTalk on YouTube: https://www.youtube.com/@CapeTalk567See omnystudio.com/listener for privacy information.

Nutty Bites
Biscuit Bake Off

Nutty Bites

Play Episode Listen Later Aug 2, 2026 8:18


Who will win the Biscuit Bake Off? Red Lobster Cheddar Biscuits, Popeyes Louisiana style, or Mom’s classic? The answer may surprise you, but also, I learned a few extra tricks. Welcome to dog days of podcasting for 2026, where podcasters from … Continue reading → The post Biscuit Bake Off appeared first on NIMLAS Studios.

Nutty Bites
Highland Games – Welcome to DDoP 2026

Nutty Bites

Play Episode Listen Later Aug 2, 2026 13:05


Welcome to DDoP 2026! let’s talk about this year’s Highland Games. Welcome to dog days of podcasting for 2026, where podcasters from around the world do a podcast a day for the month of august. I really suggest you listen to … Continue reading → The post Highland Games – Welcome to DDoP 2026 appeared first on NIMLAS Studios.

Comics for Fun and Profit
Episode 1034: Episode 1034-Tougher Questions, DC Catalog Spec?, Bad Seeds All In + Sneak Peek at Next Week w/Kyle & Drew

Comics for Fun and Profit

Play Episode Listen Later Aug 1, 2026 66:02


Episode 1034-Tougher Questions, DC Catalog Spec?, Bad Seeds All In + Sneak Peek at Next Week w/Kyle & DrewDEATH VIGIL (2026) #1Stjepan Sejic (CA) Sozomaika&The Amazing Spider-Man #34 Ryan StegmanSongs by Drew - Bad Seed(s) Like & Subscribe on Youtube www.youtube.com/@comicsforfunandprofit5331Patreon https://www.patreon.com/comicsfunprofit  Merch https://comicsfunprofit.threadless.comYour Support Keeps Our Show Going On Our Way to Two Thousand EpisodesDonate Here https://bit.ly/36s7YeLAll the C4FaP links you could ever need  https://beacons.ai/comicsfunprofitListen To the Episode Here: https://comcsforfunandprofit.podomatic.com

RBN Energy Blogcast
A Matter of Trust – Variations in Crude Oil Quality Make On-Spec Delivery Critical for Global Refiners

RBN Energy Blogcast

Play Episode Listen Later Jul 30, 2026 11:12


Crude oil — often mistakenly described as a single, standardized commodity — is anything but uniform. Every stream has its own chemical fingerprint, including differences in density and sulfur content. Today, we'll dig into why crude quality matters and how it can make a real difference to the bottom line.

Mañanas BLU con Néstor Morales
“Teletrabajo por mantenimiento de planta Spec es una prueba piloto”: ministro de Trabajo

Mañanas BLU con Néstor Morales

Play Episode Listen Later Jul 30, 2026 10:41


See omnystudio.com/listener for privacy information.

The Design Pop
Pop In: Why Great Designers Keep Learning

The Design Pop

Play Episode Listen Later Jul 29, 2026 12:51


Summer has been moving at full speed. In this episode, Alexandra reflects on Chicago Design Week, a record-breaking year for POP into Excellence, and the conversations that stayed with her after the event ended. With more than 2,300 seats filled across 11 sessions, this year's event was a powerful reminder that growth does not have a finish line—and that many of the challenges designers face are shared across the industry. Curious about POP into Excellence? Take a look at the full 2026 lineup: https://www.thedesignpop.com/pop-into-excellence-2026 Want to explore past sessions? You'll find recordings from all three years of POP into Excellence, along with more information about the event here: https://www.thedesignpop.com/pop-into-excellence Looking for a little more time to learn? Throughout July, we're offering a 30-Day Learning Pass. For a one-time $50, you'll have 30 days to explore the complete POP into Excellence recording library plus more than 850 searchable training videos covering CET, Spec, Worksheet, Aline, and other dealer design technology. It's a simple way to experience The Design POP without an ongoing subscription. (And yes—even if you purchase it on July 31, you still get a full 30 days.) https://www.thedesignpop.com/offers/SoVFRjFt/checkout If you'd like to stick around after that... You can learn more about monthly and annual memberships here: https://www.thedesignpop.com/pricing Connect with Alexandra on LinkedIn Follow The Design POP on LinkedIn Access on-demand training at The Design POP. Questions? Email info@thedesignpop.com The Design Pop is an Imagine a Place Production (presented by OFS) Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

Mil etter mil - en podcast om bil
Base-spec Pajero og ødelagt SL

Mil etter mil - en podcast om bil

Play Episode Listen Later Jul 29, 2026 48:03


David sikler på en veldig vanlig og samtidig merkelig Mitsubishi Pajero på Finn, Gard har tatt sin Mercedes SL på verksted for å reparere én ting og samtidig få en annen ting ødelagt. Og vi diskuterer feriebiler. Hosted on Acast. See acast.com/privacy for more information.

Scrum Master Toolbox Podcast
The Too-Technical PO and the One Who Was Willing to Experiment | Alf Dobbert-Baums

Scrum Master Toolbox Podcast

Play Episode Listen Later Jul 24, 2026 14:34


Alf Dobbert-Baums: The Too-Technical PO and the One Who Was Willing to Experiment In this episode, we refer to Shift: From Product to People and the value of keeping the PO–developer conversation at the goal level. The Great Product Owner: The PO Who Was Willing to Drop the Spec and Trust the Developers Read the full Show Notes and search through the world's largest audio library on Agile and Scrum directly on the Scrum Master Toolbox Podcast website: http://bit.ly/SMTP_ShowNotes.   "Let's ditch the technical part, and let's see what happens." - Alf Dobbert-Baums   The Great PO, in Alf's experience, is willing to experiment with letting go. She stopped writing detailed technical specifications and started writing plain user stories. Then — and this is the harder move — she stopped throwing them over the fence and walked into the developers' space to have the conversation. Alf had coached her toward it, and he watched the trust compound. Because she trusted the developers on the how, she had less work to do (the developers were closer to the system anyway) and the developers had room to make real decisions. The conversation stayed where it belonged: on the goal. Alf names the pattern this great PO embodied — open to experiments, willing to say "let's see what happens," and quietly resisting the seductive pull of getting technical because it feels safer. The Scrum Master's job is partly to make that letting-go feel safe enough to try.   Self-reflection Question: Where is your PO still writing the "how" — and what would it take for them to trust the developers enough to write only the "why" and the "what"? The Bad Product Owner: The PO Who Thinks They Know More Than the Developers Read the full Show Notes and search through the world's largest audio library on Agile and Scrum directly on the Scrum Master Toolbox Podcast website: http://bit.ly/SMTP_ShowNotes.   "He said: 'developers can retrieve the information from the archive database.' The developer said: 'the information is not in the archive database.'" - Alf Dobbert-Baums   The anti-pattern is the PO who is too technical — the one who writes specs full of implementation detail and treats the developers as executors. Alf calls him John. John wrote a user story that told the developers exactly where to fetch the data from: the archive database. Then he tried to hand the story to Alf to present, the way a memo gets handed across a desk. Alf refused. He insisted John present it to the developers himself. John did. The first developer to speak said the information wasn't in the archive database. Alf hopes that moment humbled John a little. The point isn't whether John was right about the database — the point is the dynamic. When the PO writes the technical answer into the story, the developers stop being collaborators and become receivers. The translation work that creates real value — between user goal and technical option — never happens. The PO ends up with more work, less ownership from the team, and worse outcomes.   In this segment, we refer to Shift: From Product to People and the idea that great POs facilitate the conversation between user goal and team, rather than dictating it.   Self-reflection Question: When was the last time your PO wrote a technical detail into a story and the developers had to walk it back — and what would change if the PO trusted them to design the "how"?   [The Scrum Master Toolbox Podcast Recommends]

trust drop berlin experiments technical agile pos scrum alf spec usability scrum masters agile coach baums requirements engineering will angela scrum master toolbox podcast
Neil Rogers Show
Neil Rogers Show (November 15, 1999)

Neil Rogers Show

Play Episode Listen Later Jul 22, 2026 143:34


Jews named Chuck, marginal hockey, AOL, cable vs DSL, (no sound 2:47-3:30), Neil's foster kid. Adam from WIOD at Spec's with "I don't have time to care" lady 13:38

Beyond Coding
AWS Veteran: The New Software Development Life Cycle

Beyond Coding

Play Episode Listen Later Jul 22, 2026 113:17


"I need to stop using Opus. This doesn't work." That was Heitor Lessa's conclusion after a refactor cost him 200 million tokens, and it forced him to rebuild the entire agent workflow now available for 1400 engineers. Heitor spent 11 years at AWS, built Lambda Powertools to 230 billion API calls a week, and in this episode he walks through the full SDLC workflow on screen, from discovery to merge check.In this episode, we cover:The product loop: discovery, whiteboarding, and the /roadmap commandSpec-driven development with Open Spec and why vanilla setups failThree model tiers: SOTA for planning, mid-tier for implementation, cheap models for reviewsMerge checks with adversarial reviewers and attestations that catch agents fabricating test resultsThe /retro command: using the Socratic method to make your workflow more deterministicIf you're an engineer figuring out how to work with agents at team scale without losing trust in your codebase, this is the workflow to steal. This is also the first Beyond Coding episode with visuals on screen, so let me know what you think of the format.Timestamps:00:00:00 - The Math Doesn't Add Up00:00:43 - Amazon Hypergrowth: 11 Years, 8 Different Roles00:03:29 - Learning From the Trenches as a Technical Account Manager00:08:38 - Developer Identity and the Birth of Lambda Powertools00:10:20 - The Hard Parts of Working in Public00:13:12 - How Powertools Hit 230 Billion API Calls a Week00:16:42 - Career Advice: Learn Adjacent Roles, Not More Tech00:19:37 - When Leadership Decisions Don't Make Sense to You00:23:21 - The Product Loop Starts With Discovery00:25:22 - From Whiteboard to /roadmap00:27:37 - Why Humans Plan First and Agents Come Second00:30:33 - Commands vs Skills Across 32 Different Models00:33:38 - Adversarial Reviewers on Every Plan00:36:07 - The Socratic Method, Explained00:40:29 - Why He Only Takes Paper Notes00:44:43 - The Five-Line Paper Trick for High-Stakes Meetings00:48:18 - /new-work: Capturing Scope Creep Without Derailing00:54:03 - The Dev Loop Begins: Open Spec Explore00:56:34 - Three Model Tiers: SOTA, Mid, Cheap00:57:43 - The $5,000/Month Per Engineer Question00:58:57 - Guardrails vs Autonomy for 1,400 Engineers01:04:22 - Auto-Sizer: Does This Task Even Need a Spec?01:07:26 - Decision Fatigue and Why Frameworks Win01:09:10 - The Plan Phase: Specs, Design, Formal Verification01:13:07 - The Refactor That Cost 200 Million Tokens01:15:11 - When Agents Forge Evidence They Ran Your Tests01:17:27 - Local-First Architecture Explained01:23:04 - The Apply Phase: Fully Autonomous Loops01:24:30 - Coding Was Never the Bottleneck01:26:39 - Why This Workflow Is an Investment01:27:39 - Decision Logs and the /onboarding Command01:29:06 - Running Agents Locally With Enterprise Governance01:32:42 - Hooks: Making Quality Gates Deterministic01:36:02 - Merge Checks: 15 Adversarial Reviewers Per Change01:38:30 - /retro: Interviewing Yourself to Improve the Loop01:43:12 - Trust, Loss of Trust, and Recovery With Agents01:48:02 - Experience, Scars, and Critical Thinking01:49:32 - Why Right Now Is the Time to Experiment01:52:04 - Conviction Comes From Being in the Loop#softwareengineering #aiagents #aws

Podlodka Podcast
Podlodka #486 – Spec-Driven Development

Podlodka Podcast

Play Episode Listen Later Jul 20, 2026 137:45


В разработке с агентами нам больше всего не хватает контроля за результатом и уверенности в нем. Один из способов их получить – использовать подход Spec-Driven Development, в котором перед тем, как отправить агента писать код, мы готовим ему понятное описание задачи и с продуктовой, и с архитектурной стороны. Вместе с Алексеем Верховским, мейнтейнером фреймворка BMAD, мы разбираемся во всех прикладных вопросах вокруг разработки по SDD. А если вы хотите вкатиться не только в SDD, но и в другие практики AI разработки, вступайте в наше закрытое сообщество Podlodka AI Engineers Club – закрытый чат, еженедельные стримы, хакатоны. https://podlodka.ai Также ждем вас, ваши лайки, репосты и комменты в мессенджерах и соцсетях!
 Telegram-чат: https://t.me/podlodka Telegram-канал: https://t.me/podlodkanews Twitter-аккаунт: https://twitter.com/PodcastPodlodka Ведущие в выпуске: Егор Толстой, Стас Цыганов Полезные ссылки: BMAD Method https://github.com/bmad-code-org/BMAD-METHOD Quick Dev skill explainer – скилл, покрывающий весь цикл одной coding-сессии целиком https://docs.bmad-method.org/explanation/quick-dev/ OpenSpec https://github.com/Fission-AI/OpenSpec/ Lost in the Middle: How Language Models Use Long Contexts https://arxiv.org/abs/2307.03172 Classifier Context Rot: Monitor Performance Degrades with Context Length https://arxiv.org/abs/2605.12366 No Silver Bullet https://worrydream.com/refs/Brooks_1986_-_No_Silver_Bullet.pdf Tracer Bullets — The Pragmatic Programmer concept, explained https://www.barbarianmeetscoding.com/notes/books/pragmatic-programmer/tracer-bullets/ From Plan to Action: How Well Do Agents Follow the Plan? https://arxiv.org/abs/2604.12147 Cross-Context Review: Improving LLM Output Quality by Separating Production and Review Sessions https://arxiv.org/html/2603.12123 Understanding Spec-Driven-Development: Kiro, Spec-Kit, and Tessl https://martinfowler.com/articles/exploring-gen-ai/sdd-3-tools.html

Nutty Bites
The Great Outdoors

Nutty Bites

Play Episode Listen Later Jul 19, 2026 71:57


Tek is back and we talk a lot about camping and being outside. What it is like but also how geeks do it sometimes. Nutty Bites is a not-for-profit podcast. All donations go to directly to charitable organizations chosen throughout … Continue reading → The post The Great Outdoors appeared first on NIMLAS Studios.

Today in Lighting
Today in Lighting, 15 JUL 2026

Today in Lighting

Play Episode Listen Later Jul 15, 2026 1:52


We are sponsored by Luminii. Great lighting comes together through every layer, not just one fixture. Learn more at https://www.wacarchitectural.com/ Highlights include: The July Issue of The Spec will be Available Later this Morning Messe Frankfurt Acquires LiGHT London, Expands Lighting Empire Rexel USA Acquires DEE Electronics Electrical Trends: Q2 Lighting. There is a Pulse plus Acuity and Signify Share

The John Batchelor Show
S8 Ep1114: Preview for Later Today: Henry Sokolski analyzes the international rules regarding strikes on nuclear power plants. He explains that under certain conditions, targeting facilities like Iran's Bushehr is legally permissible if they support spec

The John Batchelor Show

Play Episode Listen Later Jul 10, 2026 2:36


Preview for Later Today: Henry Sokolski analyzes the international rules regarding strikes on nuclear power plants. He explains that under certain conditions, targeting facilities like Iran's Bushehr is legally permissible if they support specific military operations and missions.1783 COMET

Sherlock Holmes: Trifles

"It is not cold which makes me shiver" [SPEC]      We've all had moments when we've shivered. Perhaps it's from being exposed to the cold for a while, so the entire body spasms in an effort to create heat. Or maybe it's a tingle along our spine, brought on by something that scares us.   We trace the appearance of shivers and shivering in the Sherlock Holmes stories, to determine who was shivering and why. We even explore a little about origins of the word itself. It's just a Trifle.     If you have a question for us, please email us at trifles@ihearofsherlock.com. If you use your inquiry on the show, we'll send you a thank you gift.   Our Merch Store is open: Trifles mugs, notepads, and oval stickers can be yours (or someone else's, if you'd like to make it a gift). Start shopping today.    "Trifling Trifles" is back — short-form content that doesn't warrant a full episode. This time, "Lomax" emerges from the sub-library. Make sure you're signed up so you don't miss this is exclusive benefit for our paying subscribers.  Check it out (Patreon | Substack).     Leave Trifles a five-star rating on Apple Podcasts and Spotify; listen to this episode here or wherever you get podcasts   Links All of our social links: https://linktr.ee/ihearofsherlock Email us at trifles @ ihearofsherlock.com    Music credits Performers: Uncredited violinist, US Marine Chamber Orchestra Publisher Info.: Washington, DC: United States Marine Band. Copyright: Creative Commons Attribution 3.0      

The No Name RC Podcast
NNRC News #3 - BRCA, INS RD#3, J Robinson Top 1/8 Cali Driver & More

The No Name RC Podcast

Play Episode Listen Later Jul 3, 2026 106:40


T Time Stamps : 00:00 intro Max lefty catch up  25:26 - BRCA 1:8 Nats  30:07 TNR Arenacross  34:46 - Jermaine Robinson Fastest 1:8 Racer in Cali!  38:49 - Cali King of Club Racing  41:11- LCRC - Party Rock Race  42:40 -  JConcepts INS RD #3  44:10 - Schumacher dominates on carpet at INS 49:07 - Domnic Paccione Best Spec Driver in the World & Spec vs Mod tangents  1:03:20 - ⅛ needs a mechanical difference  1:08:20 - SWORKZ LCG Conversion  1:34:09 - More Vintage Releases 1:38:40 -Light Weight Gears

Infendo Radio | Nintendo Podcast
815 – Switch 2 Spec Rumors, Splatoon Raiders Direct, and Elephant Mario

Infendo Radio | Nintendo Podcast

Play Episode Listen Later Jul 3, 2026 58:33


Episode 815 is locked and loaded! This week, we break down everything from the fresh Splatoon Raiders Direct, look into the massive $500 Elephant Mario statue, and dive into some wild new Nintendo Switch 2 hardware rumors. ? Splatoon Raiders Direct brings new details for the upcoming Nintendo Switch 2 ? RUMOR: Could EU regulations force Nintendo to give the Switch 2 a removable battery and an updated LCD screen? ? First 4 Figures reveals a massive, 18-inch Elephant Mario statue that will cost a pretty penny ? In Change the System: – Brandon hits the skies in Star Fox, tries out Meccha Cameleon, and preps an Xbox One – Eugene dives into Mina The Hollower, revisits Star Fox, and talks Steam Machine dreams Let's jump in!

Ham Radio 2.0
E1763: I Went Wardriving for Meshtastic and Meshcore Nodes!

Ham Radio 2.0

Play Episode Listen Later Jul 2, 2026 12:30 Transcription Available


Join me as I hit the roads of Galveston Island for some real-world wardriving — scouting Meshtastic and Meshcore LoRa nodes to see which mesh network is more active. From checking a long-term Spec 5 relay to testing portable nodes and repeaters, this drive reveals the current state of off-grid comms on the coast. Perfect for Meshtastic users, Meshcore enthusiasts, and anyone into LoRa mesh networking, emergency communications, or ham radio tech.Become a supporter of this podcast: https://www.spreaker.com/podcast/ham-radio-2-0--2042782/support.

Dev Questions with Tim Corey
316. I Dislike Spec Driven Development - Here Is Why

Dev Questions with Tim Corey

Play Episode Listen Later Jul 2, 2026 11:39


What are your thoughts on spec driven development? Does writing better specs mean that I get better output from my AI system? Doesn't a good spec cut down on hallucinations? These are the questions we will answer in today's episode of DevQuestions.Website: https://www.iamtimcorey.com/  Ask Your Question: https://suggestions.iamtimcorey.com/ Sign Up to Get More Great Developer Content in Your Inbox: https://signup.iamtimcorey.com/

DevOps Paradox
DOP 357: What Is Spec-Driven Development?

DevOps Paradox

Play Episode Listen Later Jul 1, 2026 57:18


#357: Type a prompt, get code, fix the hallucinations, type another prompt. That is vibe coding, and it is a fine place to start. It is a terrible place to stay. So what comes next - and is spec-driven development actually it, or just waterfall wearing a new hat? Here is the reframe that runs the whole conversation: everybody already works from a spec. Even the person who swears they are winging it has a spec in their head - which language, where it runs, what it does. The real question was never specs or no specs. It is whether you write them like waterfall, one giant document before anyone touches code, or like agile, just enough to start and the rest discovered as you go. A design is only validated when you implement it - everything before that is an educated guess. So instead of spending a month on one detailed design, build five throwaway MVPs in a day. Fully operational. Frontend, backend, running in a cluster, connected to a database. Show them to customers. Pick the one that works. Then have the agent write the spec from the winning code, and throw the code away. The spec is the output, not the input. A PowerPoint took you a month and told the customer nothing. A working thing they can touch tells you everything. Viktor and Darin push on where this breaks. Over-specifying gives you a false sense of security - you are lying to yourself that you know everything up front, and you do not. Legacy systems? The code is the only complete spec - any document written thirty years ago is fiction. Performance? Measure it in production and be lightning-fast to react. Greenfield, CRUD, clear API contracts - those genuinely want a spec first. The part nobody on the org chart wants to hear: this does not delete the business analyst or the developer. It collapses the roles. The code monkey who pulls a Jira ticket, does the work, pushes it - that job is turning into tech lead, architect, product manager, all at once. Plan mode writes the spec with you, not for you. You write it to a file because you cannot review what you cannot see. And you review the tests harder than the code, because the tests are the spec made executable. Specs were always supposed to be living documents. Now there is finally no excuse.   YouTube channel: https://youtube.com/devopsparadox   Review the podcast on Apple Podcasts: https://www.devopsparadox.com/review-podcast/   Slack: https://www.devopsparadox.com/slack/   Connect with us at: https://www.devopsparadox.com/contact/

Turn Down for Watt
Truth About EV Ownership — This Trip Exposed the EV Divide!

Turn Down for Watt

Play Episode Listen Later Jun 24, 2026 64:04


After four years of EV ownership, I thought I had a pretty good understanding of the electric vehicle experience. Then I took a trip through New England, Nova Scotia, Prince Edward Island, and rural Maine—and what I found completely changed my perspective.In some places, EVs were everywhere. Charging was plentiful, electric vehicles felt mainstream, and road-tripping was easy. In other regions, EV adoption was much lower, fast charging options were limited, and the ownership experience looked very different.In this episode of the Turn Down For Watt Podcast, we discuss the real-world EV ownership experience, the challenges and opportunities facing rural America and remote communities, charging infrastructure, DC fast charging networks, cold-weather EV travel, road trips, and the future of electric vehicle adoption.We also discuss why companies like PowerUp America are investing in underserved markets and what this means for the next phase of EV growth.This isn't a pro-EV or anti-EV discussion—it's an honest conversation about what EV ownership actually looks like in different parts of North America today.⚡ Topics Covered:• EV ownership after 4 years• EV road trips and long-distance travel• Rural vs urban EV adoption• Charging deserts and charging hubs• DC fast charging infrastructure• ChargePoint and FLO charging networks• Cold-weather EV ownership• Nova Scotia, PEI, New England, and Kentucky observations• Ford F-150 Lightning ownership• EV charging expansion• PowerUp America Kentucky• The future of electric vehiclesIf you enjoy content from Out of Spec, Everyday Chris, Gjeebs, Trucked Up EVs, Electric Duo, State of Charge, Miss GoElectric, Transport Evolved, Battery Life, The EV Geek, or other EV creators, you'll enjoy this discussion about the realities of electric vehicle ownership beyond the headlines.

php[podcast] episodes from php[architect]
Community Corner: Spec-Driven Development with Holly Schilling

php[podcast] episodes from php[architect]

Play Episode Listen Later Jun 24, 2026 27:58


 In this episode, Scott talks Holly Schilling about her work on the php tek 2026 mobile app and spec-driven development. Links: Our Discord – https://discord.gg/aMTxunVx Buy our shirts – https://store.phparch.com/products/community-corner-podcast-t-shirt Holly’s Links: Discord: TheCodeLorax Mastodon: https://tech.lgbt/@TheCodeLorax Blog: https://EventuallyWrong.com Scott’s Links: Website – https://scott.keck-warren.com/ Bluesky – https://bsky.app/profile/scottkeckwarren.bsky.social LinkedIn – https://www.linkedin.com/in/scott-keck-warren-91689810/ Mastodon – https://phpc.social/@scottkeckwarren PHP Architect Social Media: X: https://x.com/phparch Mastodon: https://phparch.social/@phparch Bluesky: https://bsky.app/profile/phparch.com Discord: https://discord.phparch.com Subscribe to our magazine: https://www.phparch.com/subscribe/ Partners This podcast is made a little better thanks to our partners. Displace Infrastructure Management, Simplified Automate Kubernetes deployments across any cloud provider or bare metal with a single command. Deploy, manage, and scale your infrastructure with ease. https://displace.tech/ PHPScore Put Your Technical Debt on Autopay with PHPScore https://phpscore.com/ CodeRabit CodeRabbit – Cut code review time & bugs in half instantly with CodeRabbit. https://www.coderabbit.ai/ Music Provided by Epidemic Sound https://www.epidemicsound.com/ #phpc #php #communityCornerPodcast #podcast #phptek The post Community Corner: Spec-Driven Development with Holly Schilling appeared first on PHP Architect.

The Beautifully Broken Podcast
Stop Arguing Spec Sheets: The Truth About Red Light Therapy Devices, Dosing & Business Models with Scott Kennedy

The Beautifully Broken Podcast

Play Episode Listen Later Jun 15, 2026 69:28


Red light therapy is now at Target. It's on Amazon by the thousands. And with that explosion comes a mountain of marketing claims, spec sheet arguments, and consumer confusion that has Freddie and Scott Kennedy — founder of Light Path LED and one of the most science-grounded voices in photobiomodulation — more motivated than ever to cut through the noise. In this fifth conversation across seven years together, they break down exactly what happens at the cellular level when red and near-infrared light hits your body: how cytochrome C oxidase absorbs photons, why that turbine in your mitochondria spins faster, how nitric oxide release drives vasodilation and blood flow, and why the gap between 700 and 800 nanometers is essentially a dead zone. They also get real about power, irradiance, pulsing frequencies, and why the best panel in the world won't save a business model built on charging $100 a session for red light alone. The second half of this episode is a full product walkthrough of what Scott has been building — including the new Titan, a six-foot full-body panel with an electric adjustable stand designed to deliver what $60,000 red light beds do at a fraction of the cost, and the TLC dork cap — three years in development, covering the full head, jaw, cervical lymph nodes, and occipital region with red, near-infrared, and blue light at 10 and 40Hz pulsing. Freddie and Scott also explore the torch, intranasal red light delivery, dental and gum applications, stacking red light with peptides for injection site activation, and the honest conversation about what a sustainable red light business model actually looks like in 2026. Use code BEAUTIFULLYBROKEN at lightpathled.com for a discount.   Episode Highlights   [00:00] – Red light therapy goes mainstream and why consumer confusion is exploding [02:10] – Scott's dental laser background and how he discovered photobiomodulation [04:32] – Why red light supports the body instead of directly “treating” conditions [06:22] – How light increases ATP, blood flow, oxygenation, and downstream cellular benefits [10:00] – The difference between red light and near infrared light [12:11] – What happens at the cellular level with mitochondria, cytochrome C oxidase, and nitric oxide [18:45] – Why irradiance and power claims can be misleading for consumers [24:32] – Therapeutic dose, joules, timing, and why more is not always better [30:53] – The problem with too many wavelengths and marketing-based panel design [36:05] – Lightpath LED's wavelength choices, pulsing features, and near infrared focus [44:11] – The Titan panel and why Scott designed a simpler full-body red light option [45:51] – The Dork Cap, blue blockers, torch attachments, and practical red light tools for daily use   Upgrade Your Health   LightPathLED: https://lightpathled.pxf.io/c/3438432/2059835/25794 Code: beautifullybroken   MaxGen Labs: https://maxgenlabs.com/BEAUTIFULLYBROKEN Code: beautifullybroken   The Biological Blueprint Course: https://www.beautifullybroken.world/biological-blueprint Earn 200 in BitCoin + Change your health   BEAM Minerals: http://beamminerals.com/beautifullybroken Code: BEAUTIFULLYBROKEN   Silver Biotics Wound Healing Gel: https://bit.ly/3JnxyDD 30% off with Code: BEAUTIFULLYBROKEN   StemRegen: https://www.stemregen.co/products/stemregen?_ef_transaction_id=&oid=1&affid=52 Code: beautifullybroken CONNECT WITH FREDDIEWork with Me: https://www.beautifullybroken.world/biological-blueprintWebsite and Store: (http://www.beautifullybroken.world) Instagram: (https://www.instagram.com/freddie.kimmelYouTube: https://www.youtube.com/@beautifullybrokenworld Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

The Curious Builder
Losers are Winners | Hans Frees Lost the Spec Home, the Savings, and His Dignity. Then Rebuilt.

The Curious Builder

Play Episode Listen Later Jun 11, 2026 32:18


Hans Frees of Outdoor Escapes has been in business 25 years, filed Chapter 7 bankruptcy, lost money on a spec home during the 2008 crash, and came out the other side with better relationships, a better location, and a much healthier respect for staying in his lane. He and Mark swap war stories about building spec homes they probably shouldn't have, losing sleep over unpaid subcontractors, and why the small-town bank that bet on them at rock bottom is still their bank today. It's 32 minutes of hard-won wisdom from someone who learned most of it the expensive way. Support the show - https://www.curiousbuilderpodcast.com/shop See our upcoming live events - https://www.curiousbuilderpodcast.com/events The host of the Curious Builder Podcast is Mark D. Williams, the founder of Mark D. Williams Custom Homes Inc. They are an award-winning Twin Cities-based home builder, creating quality custom homes and remodels — one-of-a-kind dream homes of all styles and scopes. Whether you're looking to reimagine your current space or start fresh with a new construction, we build homes that reflect how you live your everyday life. Sponsors for the Episode:  Pella Website: https://www.pella.com/ppc/professionals/why-wood/  Where to find the Guest:  Website: https://www.outdoorexcapes.com/ Instagram: http://www.instagram.com/outdoorexcapes Facebook: https://www.facebook.com/outdoorexcapes/ Where to find the Host:  Website - https://www.mdwilliamshomes.com/  Podcast Website - https://www.curiousbuilderpodcast.com Instagram - https://www.instagram.com/markdwilliams_customhomes/  Facebook - https://www.facebook.com/MarkDWilliamsCustomHomesInc/  LinkedIn - https://www.linkedin.com/in/mark-williams-968a3420/  Houzz - https://www.houzz.com/pro/markdwilliamscustomhomes/mark-d-williams-custom-homes-inc

Off Market Operator
Too Much Doom & Gloom: Real Talk on the Market, Spec Builds & Why Going Too Big Too Fast Is Risky

Off Market Operator

Play Episode Listen Later Jun 9, 2026 28:30


Join the #1 real estate community for agents and investors: https://www.skool.com/offmarketmethod/about?ref=791b3644f63045c9a6d3d8634e57c1f1Want to SCALE your real estate business to $100k/month? Go here: https://easybuttonrealestate.com/Join Easy Button Real Estate LIVE in San Diego: https://easybuttonrealestate.com/liveSummary:We're back after a little hiatus, and we're not holding back. In this episode, Tucker and I catch up on what's been going on in our worlds: Tucker's pushing through his $5-6M spec build in Naples, I recently welcomed my first kid (and survived the first two weeks), and we both weigh in on where this market is actually headed right now: interest rates, the 10-year bond, geopolitical tension, and all the weird signals that make it nearly impossible to read the tea leaves right now.But the real conversation? We get into the Brandon Turner situation, one of the biggest real estate syndication stories circulating online right now. Whether you've been following it or this is your first time hearing about it, we break down the real lessons behind it: what happens when you skip too many rungs on the investment ladder, how social media celebrity can fast-track your way into deals you're not ready for, and why the people who truly build lasting wealth in this game rarely, if ever, make the news. If you're wholesaling, flipping, or thinking about raising capital, this one's for you.Connect with Cole Ruud-JohnsonInstagram: https://www.instagram.com/coleruudjohnsonTwitter: https://twitter.com/coleruudjohnsonLinkedIn: https://www.linkedin.com/in/coleruudjohnsonTikTok: https://www.tiktok.com/@coleruudjohnson

Podcasting 2.0
Episode 262: Podcleanse

Podcasting 2.0

Play Episode Listen Later Jun 5, 2026 111:43 Transcription Available


Podcasting 2.0 June 5th 2026 Episode 262 - "Podcleanse" Dave and Adam are joined by John Spurlock and throw a big idea into the boardroom: The Podcast Data Collective Shownotes ----------------------------------------------------------------------------------------------------------------------------------------- John Spurlock - Guest The man behind op3.dev and Livewire.io - From the Great State of New Jersey! ----------------------------------------------------------------------------------------------------------------------------------------- 01 - THE IMPRESSION HEIST — AMP TASK FORCE RATIFIES 4 EXPOSURE DEFINITIONS, NO DISSENTING VOTES Podnews press release Jun 4: AMP Task Force Introduces Cross-Platform Alternative to the Podcast "Download" — "unified impression guidance for audio and video, advancing impression-based measurement as the medium's primary transaction currency." Four exposure definitions ratified. JS Jun 4 quote: "the AMP Task Force ratified a new framework with four exposure definitions, with no dissenting votes." Podcast Play: 30 seconds of content played, audio or video, once per user per session. Podcast Audience: The number of unique users who had a Podcast Play. Ad Impression: A commercial begins playing for the user. Ad Audience: The number of users exposed to an Ad Impression. They wanted to 'hasten the demand' Backstory: AMP first emerged May 29 (Podnews) — same day PC20-261 aired — "to confront podcasting's measurement dilemma." @dave reaction Jun 4 16:12: "RE: [Podnews AMP story] More secretive, back room podcast 'industry' nonsense." PNWR Jun 5 confirms the cabal-composition critique — James and Sam open the show debating AMP. James: "they also want to define what an impression is" + "we don't have a definition of podcast." Sam: "I don't think podcasting is [defined], we can measure consumption." PNWR catches the gaps [0:09:00-0:09:30]: "Spotify yes, Acast no, Art19 missing… Apple is already doing that. Apple is already being cut [out]." Same observation @dave made — who's in the room and who isn't. @js replies @dave on AMP Jun 4: "@dave Dave there were no dissenting votes" — Mastodon-thread confirmation that JS + Dave are on the same page about the consensus-by-cabal red flag. Discussion: V4V counter-thesis — No Agenda is value-for-value (no impressions, no exposures). Open standards vs industry cabals. PNWR is independent-podcaster-aligned; AMP is platform-aligned. Podnews AMP Jun 4 press release Podnews AMP origin May 29 @dave Jun 4 reaction post JS Jun 4 quote post PNWR this week (Pod News Weekly Review) ----------------------------------------------------------------------------------------------------------------------------------------- 02 - THE OPEN COUNTERPART — PODCAST INDEX ISSUE #775 (PNWR + @DAVE BOTH ON IT) ----------------------------------------------------------------------------------------------------------------------------------------- 03 - THE WHY BEHIND IMPRESSIONS — "THE FIRST FOUR AND A HALF MINUTES" ----------------------------------------------------------------------------------------------------------------------------------------- 04 - THE PODCASTING 2.0 DATA COLLECTIVE — THE OPEN ANSWER TO AMP The Podcasting 2.0 Data Collective — the open, V4V-aligned answer to the AMP cabal. Not a consortium with ratified definitions and trade-press releases. A collective of open tools and honest sentinels: OP3 for analytics, Podverse + newpodcasts.net for corpus data, Podcast Index for the namespace, Issue #775 for client identification done right. Matthew 5:6 (KJV): "Blessed are they which do hunger and thirst after righteousness: for they shall be filled." The verse that frames the work. Open data, transparent measurement, value-for-value — righteousness in podcast governance. Those who hunger for it are the ones who'll be filled. The AMP cabal trades righteousness for an ad-tech seat at the table; the Data Collective just keeps the lights on. THE CHARTER — Adam's working document, June 5 2026 We hold more power than we give ourselves credit for. Definition of a Podcast: Syndicated delivery of media files with precise consumption data for all stakeholders. What we brought in (the Podcasting 2.0 namespace contributions): Transcripts Chapters Funding (V4V) Person Location …etc. Statistical relevance: Advertising is based on percentages. Collectively we have about 10% of all apps — statistically enough to be relevant. Godcaster app tracing proves we can measure important metrics. Data to aggregate and display: Follows Plays per episode Completion rate by time Strategy: Become the authoritative source by publishing open stats Monetize We will not be loved initially by the industry, because we will have the truth. Advertisers will love us though, as will Podcasters. Monetization: Data subscriptions Resellers (DJL) Ad Networks Podcasters themselves (consideration) Podcast Index has built the trust needed to house this data. We already have a data exchange relationship with the apps. op3.dev is critical in this equation to offset the old system for correlation. OP3 full podcast support landed this week [PNWR 1:53:00-1:54:30] — OP3.dev now has full episode-level + show-level analytics support for podcasts. Spec work also moving on private feeds (insecure feeds spec). Direct relevance to V4V infrastructure. @dave → @james Jun 5 11:50: "Do you have the daily lists that show up on newpodcasts.net available anywhere as a download? I'd love the full, historical list of feed urls that have appeared there if possible." Open-data request — corpus curation theme. @dave → @mitch May 30: "Would you be able to send me a flat list of all the feed urls in Podverse which have more than X number of subscribers/followers? Let's say more than 5?" Podverse data request — corpus quality. Anchor FM RSS restoration request — Fri 11:01 email to NA inbox (Lusso Lets). Listener can't retrieve feed data from Podcast Index. Adjacent infra beat — the unsung user-facing pain of corpus indexing. Discussion: corpus curation as a steady-state job (Dave's sentinel work) vs measurement standards (the AMP cabal) — which one keeps the ecosystem honest? The Data Collective doesn't ratify, it just shows up to maintain. Hunger and thirst. They shall be filled. OP3.dev — open podcast analytics ----------------------------------------------------------------------------------------------------------------------------------------- 05 - CAPTIVATE LAUNCHES DAX US — THE IMPRESSION ECONOMY IRL ----------------------------------------------------------------------------------------------------------------------------------------- 06 - BBC GOES ALL-IN ON CROSSED WIRES YEAR 3 — IPLAYER DEAL + "EDINBURGH OF PODCASTING" ----------------------------------------------------------------------------------------------------------------------------------------- 07 - STREAMING CONSOLIDATION — YOUTUBE MUSIC + TUBI + NETFLIX ALL WANT "PODCAST" ----------------------------------------------------------------------------------------------------------------------------------------- 08 - SUPPLY CHAIN SECURITY — VS CODE DELAYS, PHP FOUNDATION, SLSA LEVEL 3 IS NOT ENOUGH ----------------------------------------------------------------------------------------------------------------------------------------- 09 - AI BUBBLE PC20-FLAVOR — TOTO CHUCKS, MOTHER COMPUTERS, "NO 'I', ONLY MATH" ----------------------------------------------------------------------------------------------------------------------------------------- 10 - QUIPS / TRANSITIONS ----------------------------------------------------------------------------------------------------------------------------------------- Last Modified 06/05/2026 14:38:09 by Freedom Controller

The Drive
Spec Says – There have Only been a Handful of Unicorns in Sports

The Drive

Play Episode Listen Later May 21, 2026 10:41


On this week's edition of Spec says, the boss gave his list of unicorns we have seen inside of sports.

The Oakley Podcast
293: New vs Used Trucks: When to Trade & When to Hold

The Oakley Podcast

Play Episode Listen Later May 20, 2026 56:45


This week on the Oakley Podcast, Jeremy Kellett talks sits down with Todd Venable, General Manager at MHC Kenworth in Little Rock, to break down the current truck market, from how COVID, fuel prices, interest rates, and freight rates have impacted new and used truck values to why now may be a better time to trade than a year ago. They cover the Kenworth order board and pent-up demand, owner-operators tied to strong carriers versus independents struggling with freight and fuel, and how being leased to a company like Oakley helps with financing approvals. Todd explains the upcoming 2027 EPA emissions changes, new technology in trucks (safety systems, digital dashes, video mirrors), and the importance of dealer training and communication. They also dive into warranties and extended coverage, modern maintenance intervals, common mistakes when spec'ing a truck (especially PTO capability), the appeal of models like the W900 and T880, and how MHC, as a family-owned dealer group, is investing for the future while supporting Oakley's owner-operators. Key topics in today's conversation include: Welcome to Today's Episode with Todd Venable (0:43) State of the Truck Market, Order Board, and Demand (3:50) New Truck Orders, Fleets vs Owner-Operators, and Pent-Up Demand (5:28) Trading Out of High-Payment Trucks and Used Market Recovery (9:54) Independent Owner-Operators vs Leased-On with a Carrier (11:19) Financing Realities, Credit, Down Payments, and Lender Options (15:20) 2027 EPA Emissions Changes and What They Mean (18:26) Cummins Gas Medium-Duty Engines and Technology Shifts (21:19) Favorite and Most Hated New Safety Tech (Bendix Wingman, Alerts) (24:40) Warranties, Real-World Warranty Wins, and Why They Matter (29:08) Spec'ing a Truck: Common Mistakes and Factory Options (35:09) Making Spec'ing Less Intimidating with Line-By-Line Reviews (38:23) Best-Selling Kenworth Models: W900, T880, and Market Favorites (42:38) Future of MHC Kenworth, Alternative Powertrains, and Investment (46:07) How Oakley Owner-Operators Should Approach MHC and Contacts (51:43) Final Thoughts and Takeaways (52:12) Oakley Trucking is a family-owned and operated trucking company headquartered in North Little Rock, Arkansas. For more information, check out our show website: podcast.bruceoakley.com. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

The John Batchelor Show
S8 Ep845: PREVIEW for Later Today: The Strategy of Electoral Delays. Guest: Evan Ellis. Negotiations regarding election dates involve complex demands for voter role purges and machine replacements. Ellis examines how slow-walking these reforms serves spec

The John Batchelor Show

Play Episode Listen Later May 8, 2026 1:15


PREVIEW for Later Today: The Strategy of Electoral Delays. Guest: Evan Ellis. Negotiations regarding election dates involve complex demands for voter role purges and machine replacements. Ellis examines how slow-walking these reforms serves specific political interests while complicating international policy goals.1940 VENEZUELA

The Modern Maker Podcast
Working on Commission vs. Spec | Modern Maker Podcast ep. 248

The Modern Maker Podcast

Play Episode Listen Later May 1, 2026 15:07


Terrazzo Talk, Fun Money vs Easy Money, and Client Talk! Subscribe and don't for get to leave questions and topics in the comments below. Thanks for watching! #podcast #modernmaker #makerFull Episode Link: https://youtu.be/155eiGNa0us?si=Gkm5xKSdQXjefokz