Podcasts about Sparse

  • 244PODCASTS
  • 470EPISODES
  • 34mAVG DURATION
  • 1WEEKLY EPISODE
  • Aug 31, 2026LATEST

POPULARITY

20192020202120222023202420252026


Best podcasts about Sparse

Show all podcasts related to sparse

Latest podcast episodes about Sparse

The Ryan Kelley Morning After
Kinda Good, Kinda Not Good (Hour 1)

The Ryan Kelley Morning After

Play Episode Listen Later Aug 31, 2026 91:17


(00:00-38:33) Jaeger bombs with random guys at the bar. The return of QFTA. Things like this never happen to me, but there I was. Cardinals just gasping to the finish line. Leo Bernal, welcome to the club. Same ol LIberatore: kinda good, kinda not good. Welcome to Instagram and Facebook, Martin. Duck Lips. Dabo vs. Lane this weekend. How about that Beau Pribula. Audio of Dan Mullen talking to the UNLV fans about bringing the energy. Papers has takes on TCU Football. Always problems in Dublin. That Friday night in Lawrence is gonna set the tone for Colonel's season. Sparse crowds at Busch over the weekend. Jacob is on the line and he wants to talk to Jackson. Taking Jello shots with a spoon? Use your boy tongue.(38:41-1:05:04) Bring Brady Cook home to the Battlehawks. Does Ladue produce a lot of pool boys? Alright, Sparky, take off the Sambas. More people at STL City games than USC Football games. Gardetto's flavored confetti. The most well-healed fan base in St. Louis sports. Turning around the Tar Heels. Lesbians make everything better. Steve from Birmingham is on the line and has an honest, serious question for the dais. Sad, that's really sad. Secret takes.(1:05:14-1:31:08) Voice of the Blues Chris Kerber joins us. Well, we're getting close. You getting fired up, Kerbs? Cale Makar now making $20M/year. The current economics of the game. Doug's not happy with MIzzou's scheduling. A hard 4.See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

The Ryan Kelley Morning After
Secret Takes (Full Show)

The Ryan Kelley Morning After

Play Episode Listen Later Aug 31, 2026 194:15


Jaeger bombs with random guys at the bar. The return of QFTA. Things like this never happen to me, but there I was. Cardinals just gasping to the finish line. Leo Bernal, welcome to the club. Same ol LIberatore: kinda good, kinda not good. Welcome to Instagram and Facebook, Martin. Duck Lips. Dabo vs. Lane this weekend. How about that Beau Pribula. Audio of Dan Mullen talking to the UNLV fans about bringing the energy. Papers has takes on TCU Football. Always problems in Dublin. That Friday night in Lawrence is gonna set the tone for Colonel's season. Sparse crowds at Busch over the weekend. Jacob is on the line and he wants to talk to Jackson. Taking Jello shots with a spoon? Use your boy tongue.Bring Brady Cook home to the Battlehawks. Does Ladue produce a lot of pool boys? Alright, Sparky, take off the Sambas. More people at STL City games than USC Football games. Gardetto's flavored confetti. The most well-healed fan base in St. Louis sports. Turning around the Tar Heels. Lesbians make everything better. Steve from Birmingham is on the line and has an honest, serious question for the dais. Sad, that's really sad. Secret takes.Voice of the Blues Chris Kerber joins us. Well, we're getting close. You getting fired up, Kerbs? Cale Makar now making $20M/year. The current economics of the game. Doug's not happy with MIzzou's scheduling. A hard 4.Isn't it ironic? The Babe Ruth of Japan. Leo Bernal getting the call up to make his debut for the Cardinals. Rainiel Rodriguez just keeps raking. Meat sticks for lunch. Jackson did not build the TMA app. Still waiting on Steven Time's apology for the UNLV pick. Audio of Memphis coach Charles Huff talking about how his program doesn't do fun things. Keep the toy.The Colonel Gabe DeArmond makes his glorious return to the show. What stands out to Gabe coming out of fall camp? Lots of unknowns heading into the season. What do we know about Cayden Green's injury? Confidence level in Jamal Roberts if Hardy misses a significant amount of time. Drink was fired up at his weekend presser on the current state of college athletics. If one coach doesn't sign a guy, another coach will. College football fans have short memories. This is an "establish the floor" year for Drink. A crime against the media.What grade does Mizzou football get over the last decade? Loofah, wash cloth, or bare hands? Jackson's learned to control his little prostate. False tie accusations. Breathing on manual. Know before you go to Faurot.It's OK, Jackson. Bobby Boots can't hurt you anymore. Doug wants every game to be tiger striped. Timmy Trumpet. 5th place finish for Jackson's team at St. Mary's Sports Trivia. Mizzou needs to play Tina E on third downs. Watching video of old Tigers hoopers giving massages to each other.Big doings starting tomorrow for the September EMOTD competition. Design Aire EMOTD.A column on Yahoo Sports says TCU's Sonny Dyke's must be on the hot seat after the loss to UNC in Ireland. Someone in the media has to pay for this.Doug didn't go to prom. Maybe we'll throw him an adult prom. Guys in their homes slapping and tickling. No prom court at the U High. Well we all have stories. Getting a second job at Pietro's. Chairman's Soprano's watching continues. The 2011 bullpen phone fiasco.Audio of Alec Burleson talking about getting to 100 RBIs for the first time and being a mentor on a young club. First Cardinal with 100 since Albert. Looming work stoppage halting the momentum of the game. Will we see Wetherholt again in 2026? Doug doesn't go back and watch old games. Tony Bank.And the winner of the EMOTD is...See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

Politics Politics Politics
Florida Races Explained! Why Paxton vs. Talarico is Really Dallas vs. Austin (with Brian Brushwood)

Politics Politics Politics

Play Episode Listen Later Aug 18, 2026 88:38


Florida is a red state, but it's never without interest and intrigue. Byron Donalds is going to destroy everybody in the Republican primary for governor. He's been a Trump loyalist, has Trump's endorsement, and all the rumors about Casey DeSantis running amounted to a hill of beans. Alex Vindman will likely win the Democratic Senate primary before losing to Ashley Moody in November. But the more interesting races are farther down the ballot, where several people may be out of politics forever when this is over.Florida's newly redrawn 22nd District has seven Republicans competing for what'll likely be a safe Republican seat. It's also where I finally found my line on AI-generated political content. A website attacking Belinda Keiser includes an AI image depicting her having sex with Jeffrey Epstein. I've had a lot of tolerance for campaigns turning metaphors and fantasies into images, but I guess my line in the sand is a candidate getting piped by Jeffrey Epstein. Now we know. David Burke appeared to be surging at the end, but there isn't much reliable polling.Politics Politics Politics is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.The race I'm watching most closely is Florida's 7th District, where Corey Mills is seeking a third term against former television sports reporter Ryan Elijah. To say Mills is scandal-plagued is to say Pinocchio is only helping Monstro's dental health, babe. Mills has faced allegations involving two former girlfriends, while the House Ethics Committee is examining possible sexual misconduct, inaccurate financial disclosures, campaign-finance violations, improper gifts, special favors tied to his office, and misuse of congressional resources. He denies wrongdoing. Sparse public polling suggests he could lose, and Republicans in Washington are desperately hoping he does.In Florida's 20th District, redistricting pushed Debbie Wasserman Schultz into a traditionally Black seat. She's running against former representative Sheila Cherfilus-McCormick, Dale Holness, Elijah Manley, and Uncle Luke himself, Luther Campbell. With four Black candidates dividing the vote and no runoff in Florida primaries, Wasserman Schultz only needs a plurality. I think she'll win. Let the sun shine down and reveal the truth as Florida votes — and start a stopwatch when the polls close, because these races will probably be called within about 90 minutes. Shocking how fast that can go. Just saying.Congressional Democrats are preparing another war-powers challenge after Trump threatened to bomb Oman if it obstructs efforts to reopen the Strait of Hormuz. They've been looking for another non-affordability way to hit Trump over Iran, and this is a good strategy. It's a historically unpopular war, and they want to tie him to it every way they can. Meanwhile, the administration has paused construction of a border wall inside Big Bend National Park while CBP reviews the project. Not all border is created equal. Big Bend accounts for a small share of crossings, and there are easier, lighter ways to patrol an area where somebody might still have to walk for three days after crossing.ABC has also filed a First Amendment lawsuit accusing the FCC of retaliating against the network for its editorial decisions. FCC Chairman Brendan Carr has threatened to revoke ABC's license amid his fight with the network, particularly over Jimmy Kimmel's comments about Charlie Kirk's death. It's not surprising that ABC is bringing a legal challenge, even if only as a predicate to any FCC action. If you're ABC and the chairman of the FCC threatens your license, it would be malpractice not to take him seriously and prepare in case it actually happens.Chapters00:00:00 - Intro00:04:08 - Florida Races00:20:49 - Oman00:25:47 - Border Wall00:28:02 - ABC Sues FCC00:29:35 - Brian Brushwood on the Nuances of Texas' Electorate01:25:20 - Wrap-up This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.politicspoliticspolitics.com/subscribe

Latent Space: The AI Engineer Podcast — CodeGen, Agents, Computer Vision, Data Science, AI UX and all things Software 3.0

Watch the full episode on YouTube:We first covered Baseten last year when DeepSeek mania was at peak hype. Now they have raised a monster $13B round and become one of the new cohort of AI Infra decacorns that are (with Nvidia, Intel, and the semis complex) chief beneficiaries of the Inference Inflection. We return to Baseten at the peak of the 2026 edition of Open Weights debate. Ali has published a viral breakdown of Kimi K3:And since you last saw him, Philip has spoken at AI Engineer and written the definitive book on Inference Engineering spotted all over SF:Three years ago, inference engineering barely existed as a category.Today, it is one of the most critical disciplines in AI. Inference engineering inherently tackles a different question than standard model training: “How do you turn those weights from training into a product that is fast, reliable, and affordable at scale?” Focusing on these creates an entirely new optimization problem.In one recent GLM-5.2 experiment, quantizing more of the model actually preserved its benchmark quality while increasing throughput by 20%, because the errors introduced in different layers could cancel each other out.Inference is no longer just the final step after training. It is becoming its own engineering discipline, with its own research problems, infrastructure, and increasingly specialized roles.In this episode, Baseten's Philip Kiely and Ali Taha join swyx and Vibhu to explain what actually happens after a new open model is released and what it takes to turn “we generated a token” into a fast, reliable, production-ready API.We go deep on cache-aware routing, disaggregated prefill and decode, quantization, speculative decoding, KV-cache movement, model parallelism, GPU kernels, and the race to make frontier models up to 10× faster. Philip and Ali explain why inference optimizations can still produce gains of 20%, 100%, or even 200%; how quantization errors can cancel one another out; why identical weights can behave differently across clusters; and how Baseten grafted a Kimi vision encoder onto GLM-5.2 without changing the underlying language model.The conversation then expands beyond LLMs into NVIDIA Dynamo, mega kernels, Rubin, AI-specific chips, local inference, video generation, diffusion versus autoregressive models, and the enormous compute barrier to generating coherent long-form video. Finally, we explore the convergence of training and inference, continual learning through persistent KV cache, and the emerging loop where models help optimize the infrastructure that runs them.We discuss:* What happens when a 200,000-token request enters an inference system* Cache-aware routing and reusing previously computed KV cache* Why prefill and decode are increasingly handled by different GPUs* When dedicated deployments become cheaper and more reliable than shared APIs* How speculative decoding uses a smaller model to accelerate a larger one* Tool calling, structured outputs, and what LLMs actually do* What it takes to support a new open model on day zero* Grafting Kimi's vision encoder onto GLM-5.2* Retrofitting inefficient model layers with components from other architectures* Why models sometimes collapse into repeating the same token* How hardware, kernels, and race conditions create nondeterministic failures* Preserving model fidelity while making inference faster* How quantization errors can cancel each other out* Why inference optimizations still deliver gains of 20%, 100%, and 200%* How optimized serving can make a model up to 10× faster* NVIDIA Dynamo, KV-aware routing, and distributed model serving* Speculative decoding the speculative decoder* Why local AI is about making models less dumb while data-center AI is about making them less slow* Tensor, expert, and pipeline parallelism across GPUs* Hardware-aware model design, auto-tuning, and the case against mega kernels* Rubin and why inference is becoming a systems problem* Whether modern GPUs are evolving into programmable AI ASICs* Why enormous models like Kimi K3 require GB300-class hardware* Why open-source video generation still trails Veo, Kling, and other closed models* The quadratic attention bottleneck behind long-form AI video* Autoregressive video, real-time generation, and compounding quality drift* Why future video systems may combine autoregressive and diffusion architectures* Training for inference and inference for training* Continuous post-training, deployment, evaluation, and improvement loops* How GLM-5.2 helped optimize the kernels serving GLM-5.2 itself* Why faster networking could unlock dramatically faster decoding* Continual learning, KV-cache compaction, and persistent model memoryShow Notes* How to build a day-0 API for Kimi K3* 22580: From GPT2 to Kimi3, ExplainedPhilip Kiely* LinkedIn: https://www.linkedin.com/in/philipkiely* X: https://x.com/philipkiely* Inference Engineering: https://www.baseten.co/inference-engineering/Ali Taha* LinkedIn: https://www.linkedin.com/in/aliestaha/* X: https://x.com/waterloointernTimestamps00:00:00 Introduction and the 200K-Token Prompt00:03:18 Dedicated Deployments, Speculative Decoding, and Tool Calling00:11:26 Launching Production-Ready Open Models00:19:06 Model Retrofits, Failure Modes, and Nondeterminism00:28:22 Quantization and Canceling Errors00:32:15 The Race to 10× Faster Inference00:40:48 Dynamo, Speculation, and Local vs. Data-Center AI00:50:18 Model Parallelism, Auto-Tuning, and Mega Kernels01:00:55 Rubin, GPUs vs. ASICs, and Custom AI Chips01:10:03 Giant Models and the Limits of GPU Memory01:12:42 AI Video, Quadratic Attention, and Autoregressive Generation01:21:47 Audio, Images, and Diffusion Models01:27:32 Training, Self-Optimizing Models, and Continual Learning01:40:06 Closing ThoughtsTranscriptIntroduction: Baseten, Waterloo Intern, and Inference EngineeringSwyx [00:00:00]: Okay, we're here in the studio with Philip, old friend from Inference Engineering, the book, as well as Baseten and everything that you've done, you and I have done before, as well as Ali. Welcome.Ali [00:00:15]: Pleasure to meet you.Swyx [00:00:15]: Waterloo intern.Ali [00:00:16]: Waterloo intern, always.Swyx [00:00:17]: When did you get “Waterloo intern” as a handle?Ali [00:00:19]: As a handle? Oh.Ali [00:00:20]: I think the rebranding happened mid-March. When I saw it was open, I was like, “I have to take it. Up for grabs.”Philip [00:00:26]: The problem is that Ali is really good at his job and is not gonna be an intern much longer.Philip [00:00:30]: So we have to figure out who's gonna get the handle.Ali [00:00:33]: Well, I'll pass the torch over to the next intern.Swyx [00:00:34]: Oh, okay. It can be, like, you just pass it to another Waterloo grad.Ali [00:00:37]: To another Waterloo intern. No, bruh.Philip [00:00:39]: Yeah.Ali [00:00:39]: Intern.Swyx [00:00:40]: Intern, yeah.Ali [00:00:40]: And no.Philip [00:00:41]: You gotta get an intern from Waterloo.Ali [00:00:42]: Yeah, I've gotta get an intern from Waterloo.Swyx [00:00:44]: Right.Ali [00:00:44]: But they have to follow the path.Swyx [00:00:45]: Oh, it could, but it could come from Baseten, so it's like whoever Baseten gets from Waterloo.Ali [00:00:48]: Right.Swyx [00:00:49]: Has the title of Waterloo.Ali [00:00:50]: It stays in the ecosystem.Philip [00:00:51]: Exactly.Ali [00:00:52]: Halfway through the internship, you either get it or you're out.Philip [00:00:55]: You should also do, like, a big graduation ceremony where you change the handle.Ali [00:00:59]: Just say it.Philip [00:00:59]: For everybody.Swyx [00:01:00]: You guys are good at ceremonies, clearly. We had a nice launch of the book, very successful. But before we get into all that, I wanna start off with a fun question for you. Okay, you're an expert inference engineer. What happens when I send a long query, say two hundred thousand tokens into Baseten's inference? What's the process of query through GPU model routing, balancing, all that? What is all the stuff that we don't think about?Long Context Requests, KV Cache, and Cache-Aware RoutingPhilip [00:01:26]: With a long query specifically, the first thing that I'm gonna ask is, “Have you sent me this query before, or at least part of it?” and I really hope you have, because it's gonna be a lot easier for me and a lot cheaper for you. So the first thing that we're gonna look at is some cache-aware routing, where we're going to see, we probably have a number of instances, a number of replicas up serving whatever model you're hitting. We want to send this one to something with, number one, available prefill workers, and number two, ideally some cached input already there so that we can skip prefill on at least part of these two hundred thousand tokens. If you're doing two hundred thousand tokens, it's probably coding or a multi-turn agent or something where you would expect to have that cached. If you don't, we're gonna have to send it to a prefill worker. We've at least on certain models disaggregated prefill and decode, so you're going to have one set of GPUs that's solely going to process the input, create the KV cache, and get you your first token, and then that's going to be passed over to a separate set of GPUs, which is going to run decode. We're going to iteratively make those tokens. We're probably going to have some speculator model in front of that. I'm going to assume that you're doing coding, and because of that, our speculator model, which assumes you're doing coding, is gonna have a high draft token acceptance rate. If I'm wrong and you're asking me to summarize every Harry Potter book, it's gonna be slower. And then we stream that output to you and account for it, charge you, a couple of pennies and say, “Hey, would you like to send another one?”Swyx [00:03:04]: Except Baseten doesn't charge by pennies.Philip [00:03:07]: Well, yeah, we charge. I'm assuming that we're talking about the public model APIs. If you are setting up a dedicated deployment, then yeah, it's not pennies.Public APIs vs. Dedicated DeploymentsSwyx [00:03:18]: Yeah, one of the key differentiators when I was talking with Baseten initially was that people who want very high volume just need to rent by the box, ‘cause then it's up to you to figure out how to saturate the box.Ali [00:03:31]: And more often than not, it's, like, way cheaper if you're pushing, like, millions of tokens per hour, if you just pay per hour instead of pay per token.Philip [00:03:37]: Yeah, they do. I think that we've increasingly seen a lot of demand for the pay per token APIs, just because everyone wants to try open models, and then once they find a use case that's really sticky, then they move over to dedicated.Swyx [00:03:51]: Is there a best practice on when it's time to swap over?Philip [00:03:54]: Couple reasons. Yeah, reliability, that's a big one, right?Ali [00:03:57]: Like, if they have a very specific use case, they want you to train something specifically for them, like they want their own spec dec, for instance, for their own traffic.Swyx [00:04:04]: Spec dec is speculative decoding.Speculative Decoding and Custom SpeculatorsAli [00:04:05]: Speculative decoding, yeah.Swyx [00:04:07]: You have to explain.Ali [00:04:07]: Sorry. Like, speculative decoding is like, if you have a huge model, right? And so the model is going to be generating one token at a time every single turn, every single forward pass. So we attach, like, this little, like, parasite, like this layer that goes on top of the model, and this model just has to predict. It does three very fast autoregressive forward passes, and it will predict, like, three certain tokens, and then you do one forward stage over the entire original model in order to see if those predictions were correct or not, and then you accept them or you reject them. Now, this draft model is traffic specific, so if you, like, Philip said, if you're summarizing Harry Potter books, I can train exclusively that draft model on Harry Potter books, and I can guarantee you that I'm gonna accept the three tokens every single time. And so with that case, I increase your decode speed. I wouldn't be able to provide this to you if you're a shared endpointSwyx [00:04:53]: YeahAli [00:04:53]: ‘cause I have no idea if you're doing Harry Potter, if you're doing coding, if you're doing English. We don't know. Also, there was a thing in the book that mentioned that if they really cared about a specific threshold, chapter four, I think. Do you remember that?Philip [00:05:06]: Yeah. The things that you can do is you can set a specific, like, batch sizing, a specific, like, parallelism strategy if you're trying to optimize for, like, throughput versus latency. You can. Maybe a NVFP4 quant doesn't pass your benchmarks and you wanna run a model at higher precision, you could do that. There's just a bunch of reasons why you might wanna have your own endpoint and the biggest one, of course, just being, like, you don't have to deal with someone else doing a hundred million tokens of benchmarking traffic at the endpoint when you happen to be trying to serve your users.Swyx [00:05:40]: Yeah. I think one thing that is. That is a classic journey. Like, it's people is asking the, what happens when you type Google into the browser. Tool calling, is that just, you're generating JSON or is there more complication beyond that?Tool Calling, JSON, and Structured OutputsAli [00:05:58]: Certain customers that we have, they have their own post-trained models, and so they demand a tool calling that's not just, like parse a file or go find the weather. It's something that's very specific and you have to do post-training on this. And if the post-training on the model is not good or if the quantization after the post-training to get the inference to be fast, the model will struggle reading the JSON file and reading the tool calling. But it doesn't require its own like sandbox. It's not like it's going to use that tool calling to like escape a sandbox or like it doesn't have to be contained. It can just be a normal dedicated deployment. The challenge with tool calling more and more seems to be that the companies want certain tool calling which is a very sensitive thing to train. And because you're dealing with all of the JSON outputs, if it doesn't like close the end of the request in a very certain manner, you end up with a model that did the tool calling and like the thinking and so as a result of that, it didn't see the result and just hallucinated the result as it decoded. That seems to be the most challenging thing with tool calling, not really the sandboxes model.Philip [00:06:56]: Yeah, that's a challenge on the training side and then on the inference side, there's work that you can do to scope the possible output. So we published this at this point close to two years ago, the solution to this problem which is you make a state machine and you use that to constrain the output to a specific format. So this is the structured output problem. If you remember backSwyx [00:07:27]: Yeah, the specific grammar is,Philip [00:07:29]: Yeah, exactlySwyx [00:07:30]: GML had this thing.Philip [00:07:31]: Yeah. So it's like the old-school “make sure this is only JSON”, return only JSON orSwyx [00:07:38]: YeahPhilip [00:07:38]: Grandma's gonna die type of prompts.Swyx [00:07:39]: Is it BNF grammar? At some point OpenAI had released a thing that was like, yeah, if you want to constrain your output, write BNF grammar, back as NOR.Philip [00:07:47]: In our inference system, it's just a specified output format. And you get the guarantee that your output's gonna be structured along that format. And so applying that to tool calls can like help cut down on. You can still call the wrong tool or call no tool. It doesn't solve the certainty problem but it at least solves the output structuring problemSwyx [00:08:10]: YeahPhilip [00:08:10]: Within tool calls.Swyx [00:08:12]: And MCP is just another form of tool, right.Philip [00:08:14]: Yeah, exactly.Swyx [00:08:15]: As far as there's no special thing there.Philip [00:08:16]: The thing I'm always like explaining to people is the LLM is not capable of doing anything. It's only capable of making suggestions of what to do and then if those suggestions are formatted in a certain way and applied to a system that knows what to do with them, then an action occurs.Swyx [00:08:32]: Yeah. Part of the fun stuff is, this is solved outside of tool calling too. Like in an agent loop if the output is not correct or you're right, like reasoning, tool calling was done in the reasoning trace, just be like, “Oh, I don't know what to do. Let me just try again.” And it might get there after a few tries. And on your point of training, sometimes this is harder in smaller models, so you don't have the same exact quality outputAli [00:08:56]: Right.Swyx [00:08:57]: When you just swap from a big model, right?Ali [00:08:59]: Yeah. I will say that, before, I think we need to go back to inference engineering proper.Ali [00:09:04]: But, I had expected that something would replace JSON because it's hard to stream JSON ‘cause JSON must be complete and you must have open and close brackets and everything. So it's hard to parse something or validate something while it's being streamed. So people invented all sorts of things that are like, I forget the name of some of these alternatives, but it's something like TOML, something like YAML. But JSON seems to be dominant still.Philip [00:09:30]: The JSON outputs aren't that long, right? Like you could have a long-- ‘cause tool calls also contain the arguments in them and perhaps for a certain tool you might pass like a very long argument. But my impression of the median tool call is that it's a relatively small number of tokens, right? So I would expect that speculators are generally fairly good at something as formatted as JSON. And so you would have like a pretty fast decode step there and that the streaming wouldn't be as valuable, but maybe I'm wrong about that.Ali [00:10:02]: I think you're also bounded by the software or that the model is gonna integrate with if the software is built with JSON for the tool calls or if the company that you'- if your customer says that this is how our software works and our tools are interfaced with JSON, you can ask them to like, change their software and say like, “Yeah, this is gonna be better for the model.” but like with the right training shouldn't be that much of a difference. Also more profitable if it outputs more tokens probably.Swyx [00:10:25]: Depends on your business model.Swyx [00:10:27]: It really depends. But I will say that, as a writer with like experience a lot with generated output, I do try to move from text to JSON text which is very long JSON, right? Like there's paragraphs in every field because I'm trying to structure it, right?Philip [00:10:44]: Right.Swyx [00:10:44]: I want you to first make factual statements, then make opinions then make bullet point summaries, have dates, have entity references have your sources for references, all these things. Anyway, so these are things that like I think people who really experiment with structural output have to really care about. But, let's, let's recurse up the stack a little bit. Before we started recording, you mentioned something really cool, which is that there's a lot of engineering that-- inference engineering that goes on when a new model provider releases a new model, right? So let's call it GLM-5.2, Kimi K3. I had previously assumed, especially if it's like, well, GLM 5 to 5.1 to GLM-5.2, like that you've supported them before. Is it that much work?What It Takes to Support a New Open ModelAli [00:11:26]: It's a lot of work.Swyx [00:11:28]: Yeah. Okay. So like, a lot of people, all you guys, right whenever a new model launch like, people rush to say like, “Oh, Hugging Face supports this, Fireworks supports this, Spacetime supports this,” and I'm like, “Yeah, of course we support it.” But what goes into that? What goes intoPhilip [00:11:40]: I think it's more than just support it too, right? It benefits the consumer a lot. Like I think it was with Kimi K2.5 or GLM-5.2 the latest, there was an inference war, right? X provider is at 90 tokens a second. The next day we're at 150. The nextSwyx [00:11:55]: I kinda kicked that off with the GLM-5.2.Swyx [00:11:58]: I wrote a Twitter article about. It got like half a million views,Ali [00:12:02]: Based on being numberSwyx [00:12:03]: YeahAli [00:12:04]: Or it's for something else.Swyx [00:12:05]: Yeah. Which,Ali [00:12:06]: Oh my GodSwyx [00:12:07]: Which then got everyone really excited about, hey, how can we, bend tracks a little bit further and,Philip [00:12:14]: There's a difference between support the model, as in I can make a token out of this model, and support a model, as in I have a production-ready API from this model.Philip [00:12:26]: Getting to the point of I can make a token out of this model is not that hard because generally the, open source inference engines, vLLM, SGLang of the world oftentimes even receive weights ahead of time, maintainers do, or the people making the model merge PRs to ensure support. So you generally can, just get it working on the standard open source stack without too much pain in most cases. The challenge is, every inference company is gonna have own proprietary stack. Some open source components, some in-house stuff. And for any arbitrary model, there's going to be some new stuff. Sometimes you get lucky, like K, two five to two six was, like, pretty similar.Quantization, Speculators, and Production ReadinessAli [00:13:16]: Yeah. It was pure continued post-trainingPhilip [00:13:18]: YeahAli [00:13:18]: If I remember correctly.Philip [00:13:19]: Even in those cases, there's still stuff you have to do. You have to redo the quantization work. You're taking the model from. Generally, these models are not released in NVFP4, and we want them to be in NVFP4 for maximum Blackwell compatibility. So we have to perform that quantization, and, calibrate the quantization to make sure that we're not causing any regression in the model's intelligence. And then we also have to train the speculator, as we've talked about. Generally, we have. We have ZDR, zero data retention on our model APIs, so we don't know exactly the traffic that people are sending us, but we know what's popular. We know that coding use cases are popular. We know that agents, agentic use cases are popular. So we can get public data sets that are representative of that traffic and train general speculators. Now, with speculators today, you need to train the speculator using the base model itself because you're getting hidden states out of the model from running inference on these specific prompts, and that is the training data you use to create the speculator. So there's that process which you need the real model weights for. And then there's of course just the process of, standing up all the infrastructure behind it, loading all this stuff, testing it. And then when there's a new model with a newer architecture, I think that, like, the DeepSeek models tend to be the most challenging as they have, like, the most novel architectural stuff going on, model after model. But every new model has something. Kimi K2 had. Oh, sorry, GLM-5.2 hadAli [00:14:53]: Sparse attention.Philip [00:14:54]: Yeah,Ali [00:14:54]: YeahPhilip [00:14:54]: the DSA.Ali [00:14:55]: Right. Which is brought from DeepSeek.Philip [00:14:57]: Yeah. AndAli [00:14:59]: So you can copy-paste then?Philip [00:15:01]: It kindAli [00:15:01]: I don't know how this works.Philip [00:15:02]: So, like we had to, like, build support for that into our runtime. And you're right, like it is really interesting the way that all of these open source labs borrow from each other. For example, like GLM-5.2 doesn't have vision. So something that, Haley, a guy on our team, if we could take a look at this, he, like, grafted the Kimi vision encoder onto GLM-5.2.Retrofitting Vision into GLM-5.2Ali [00:15:27]: We'll be training the projector.Philip [00:15:28]: Exactly. So if you think about, like, the encoder, there's the encoder, which is the part that looks at the image and turns it into latent information, and then there's the projector which likeAli [00:15:38]: You can say latent space. It's okay.Philip [00:15:41]: And then there's the projector that maps it onto, the model itself, and then there's the model weights. You don't wanna mess with the model weights because you run a chance of making the model dumber at something else for the purpose of giving it vision. So instead, Haley started with just a projector, which is only a handful of millions of parameters.Ali [00:16:02]: That would be, yeah.Philip [00:16:02]: Yeah.Ali [00:16:03]: Can you show the training one?Ali [00:16:04]: Like the way it groksPhilip [00:16:05]: YeahAli [00:16:06]: Very interesting.Philip [00:16:06]: And maybeAli [00:16:07]: That right therePhilip [00:16:07]: Maybe Ali, you should take it from here. You've got a betterAli [00:16:10]: Ooh, double the sandPhilip [00:16:11]: Understanding of this than I do.Ali [00:16:11]: Yeah. You can see, like, he. The way he trained this is really cool. At the beginning, he was training it using just like, “Here's a picture of a mountain. Can you describe what's in this mountain?” And that caused it just like the first, learning walls. Like here you can see this all we're trying to teach it is to translate the encoded. Like it's already taken the encoder from Kimi K. It's taken the image. It'Philip [00:16:31]: Yeah. FrozenAli [00:16:31]: FrozenPhilip [00:16:32]: With adapter.Ali [00:16:32]: Exactly.Philip [00:16:33]: Yeah.Ali [00:16:33]: So the brain is frozen and the eyes are frozen. It's just we're tryingPhilip [00:16:37]: AlignAli [00:16:38]: Interconnect between the eye and the brain, right? So the projector. And so you take the tokens and then he's like, “Oh, can you describe what's in this image?” And he's like, “Oh, it's a mountain,” or it's a person or it's a human, whatever the case is. But that didn't cause complete understanding. So he changed it such that every image was associated with a data set of questions. Like, does this image have a white male? Does this image have birds in the top corner? Does this image have a scientist in it? All of that stuff. And it would have to answer questions correctly. And using not just training on describing an image, but being able to answer question, another question, answer over time. Like you can see the grokking, which is like genuinely insane, that retrofitting vision into a large LLM can learn to that extent. And even for images that it doesn't perform well on, for instance, if you ask it a picture of like Stephen Hawking, “Who is this?” Maybe it doesn't get it, but it will say something like, “This is Albert Einstein.” Like it still understandsPhilip [00:17:25]: Close enoughAli [00:17:26]: That this is a scientist who is a man who has, some significant achievements, all that stuff. So that's like really cool.Philip [00:17:32]: Yeah. So, we've covered Hao Tian before, who the author of the LLaVA paper that did this, a while ago. And I think that's very foundational work for anyone who hasn't done vision work before.Ali [00:17:41]: Same with the CLIP and MetaCLIP, where you go from just captioning to building out questionsPhilip [00:17:47]: RightAli [00:17:47]: Off the image and how much better you can get performance.Philip [00:17:50]: Right. Right. Right. Yeah. But what's, what's so exciting about this is if you look at a model like this. Now, this is a little bit more of a research project. It's not. It got to 56% on MMLU Pro, I think. So not quite frontier. But if you're running this model, you haven't suffered any loss on your GLM-5.2 quality. If you don't have an image, it'll just behave exactly the way it used to. And ultimatelyAli [00:18:14]: Which in the inference code you literally do not include the other part, right?Philip [00:18:18]: Yeah. You would just skip the encoder if you don't have an image input.Ali [00:18:22]: Okay.Philip [00:18:22]: Just confirming.Philip [00:18:23]: YeahAli [00:18:23]: Does it affect a lot on the overall inference side? Like you're not adding much, you're adding a very small vision encoder. These are typically likePhilip [00:18:30]: They're super fineAli [00:18:31]: Less than a billion parameters, right?Philip [00:18:32]: Yeah. It's, - There's a little bit less standardization among vision encodersSwyx [00:18:37]: YeahPhilip [00:18:37]: So the support matrix can be a little bit, sparser. But overall, yeah, it's a pretty, it's a pretty minor component of the overall system. And ultimately what you get out of the system is all of a sudden you have Kimi Vision, GLM weights, and DeepSeek attention all in one model.Open Source Model Grafting and Franken-MergesPhilip [00:18:56]: And that's, I think, a lot of the power and beauty of open source, is that you can take all of these different components and combine them together into a system that's better than anyoneSwyx [00:19:05]: YeahPhilip [00:19:05]: Can be individually.Swyx [00:19:06]: People used to say that you would also do Franken-merges where you would take likePhilip [00:19:10]: YeahSwyx [00:19:10]: Layers from each model.Swyx [00:19:11]: Does anyone do that anymore?Ali [00:19:13]: Well, to your point previously when you were mentioning like, the work that goes into supporting a model when it first comes out, like GLM-5.2 or MiniMax M3 or whatever the case is. Sometimes you do have to like, you do have to switch out some things. Like, for instance, the MiniMax M3 head uses full attention, and with full attention you end up with this like insane bottleneck in spec dec ‘cause you're doing auto-regressive token generation for three tokens, and you're doing this like N squared over all of the tokens that are in your sequence. Your KV cache is like very large because it's not sparse, it's not top K. So we find it better to like, okay, we're gonna replace this, we're gonna replace this layer with a layer from another model that's using like GQA, for instance. And then just with the right training, you can get it to have the same acceptance rate. So it is very possible to retrofit layers from other models and very much needed. If a layer is like inefficient, the training just becomes the challenge, like how do you ensure that you train it properly? Which again to your earlier point is like the mesh between training and inference. As in like you need very good training in order to do fast inference. That's like, I feel like more and more becoming true.Swyx [00:20:21]: Yeah. Anything else on the support side when you say like get it to fully production ready?Loop Detection, Race Conditions, and Non-DeterminismPhilip [00:20:26]: Yeah. I think that there's also a question of just, we can test a model to a pretty extensive degree, but we're trying to get it out quickly and then you see a bunch of other people test it and you get interesting results. There was an issue with, GLM briefly where we had some like mode collapses where it would just output the same token over and over again for certain prompts on certain temperatures. Like once you expose an endpoint to the real world, there's going to be, so many more varieties of things given to it that you're able to, discover and patch things. So it's not just a, day zero process, it's then like for the first week, for the first month, if a model remains popular, like how do you both fix bugs and then continue to push the envelope on performance?Ali [00:21:21]: What do you mean you don't want your model outputting S?Swyx [00:21:24]: Is there loop detection on that stuff, by the way? It still happens like quite a lot, which is surprising.Ali [00:21:30]: We have like we, in our endpoint, like if a model was to output the same token like four plus times, we just cut the generation. We say like, “Oh, sorry, this-- Like try again,” or like we will reprocess the request. ‘Cause we know then, like if it, like if, yeah, it's four times the same token, it's probably collapsed.Swyx [00:21:45]: Yeah. Is there a way to opt out in case I really want that?Ali [00:21:48]: You want that?Ali [00:21:50]: I think there's a way that we have to handle it. I'm not exactly certain, but I feel like in certain models, like when they output something like you can imagine, like a table for instance, and so they want, they wanna draw like 12 dashes and 12 dashes. Yeah, I think there's a way for that to happen. I think we only do it on certain tokens. Like we exclude certain special characters.Swyx [00:22:07]: Yeah.Ali [00:22:07]: So we only do it on like certain like S is the most common almost. GLM-5.2Swyx [00:22:11]: OhAli [00:22:11]: And I think it was DSV 4 as well. Like you'd just have like looping issues where like you literallySwyx [00:22:17]: ItAli [00:22:17]: Just have like S.Swyx [00:22:18]: Yeah. Is there a special, something special about S? No, just randomlyAli [00:22:21]: It just seems to be the one token involved.Swyx [00:22:23]: Yeah. And it'Philip [00:22:24]: Is thereSwyx [00:22:24]: And it's only temperature 0Ali [00:22:27]: NoSwyx [00:22:27]: Even at other temperaturesAli [00:22:27]: Even at like 0.9 or whatever, it will still, it will still collapse.Swyx [00:22:30]: That's weird, right?Ali [00:22:30]: It's, it is an inference problem to be honest, like a software problem. Like oftentimes, the image you run will-- like NVIDIA will release an image for instance, and if we will upstream the changes from their latest TensorRT-LLM image into our stack, we'll find that it fixes it. Or oftentimes this will only happen in an inference engine that you're using like SGLang. But if you were to switch to vLLM, that isn't the case. So it seems to be like an extremely like deterministic software issue and not really a model issue. It's not like a weights problem. Like I'- we'll say like, “Oh, it's a problem with the quant. We did PTQ wrong,” right? But that isn't, that doesn't make sense because the same weights used with a different inference engine does not repeat the problem. And sometimes it's, the kernels that are being used in the backend have like these very subtle sometimes race conditions, where if you were to use this model hosted on one cluster, you will never get this problem.Swyx [00:23:19]: Oh my God.Ali [00:23:19]: But if you host it on a different cluster, you will. And the reason is the KV cache transfer from a node to node in that one cluster is using a slower interconnect than the node to node in another cluster. So that exposes the race, whereas in another cluster it doesn't. So then you end up just like, okay, this model is not gonna be hosted on this cluster. We're gonna host it on, another cluster because that cluster exposed that problem. But then it ends up with like, okay, is it the software? Is it the model weights or is it the hardware?Swyx [00:23:42]: There is a thing about this with temperature 0 still not being deterministic, right?Ali [00:23:46]: Right.Swyx [00:23:46]: Mostly because of hardware. Even at temperature 0 same model, you won't always get the same output.Swyx [00:23:52]: Even-- But I'm surprised by the race condition one because, I thought PyTorch was a graph that like guarantees that you at least, execute things in the right order.Ali [00:24:02]: Well, yeah, true. Like I'm not, I'm not saying that there is. Like well, you have things like PTL optimizations where like you can start a kernel before the end of the previous kernel, and that's like ‘cause you want to do that because there'sSwyx [00:24:12]: It's like pipeliningAli [00:24:12]: Expense. Exactly.Swyx [00:24:13]: Yeah.Ali [00:24:13]: But it'- But you don't do it cleanly. Like you overlap a little bit of the execution. No, it is very possible that the kernel itself, like that one block that is supposed to be running in this instance of time, that kernel itself has a race condition. For instance, like a missing barrier. Like often if you're designing a kernel and you want it to make it to be very fast, if you don't test it extensively, you'll, you'll have certain threads access data points from registers before they've been written to by other threadsSwyx [00:24:36]: YeahAli [00:24:36]: For example, because like your barrier is wrong or your synchronization was wrong. But yeah, like the testing itself is very difficult in those like, andSwyx [00:24:42]: And there's no like borrow checkerAli [00:24:45]: What does that mean?Swyx [00:24:46]: Like Rust. Like the. If you're trying to have like memory safety It sounds like a comparable problem.Ali [00:24:52]: Well, yes, but you're working in CUDA, right, NVIDIA GPUs. Like- You just need a higher level language like modular Maybe that's what modular is supposed to do. I don't know.Quantization Quality and Vendor FidelityVibhu [00:25:00]: How do you see keeping quality of the model? So you talked about all these steps of, okay, you gotta do quantization, train your own speculative decoderAli [00:25:07]: RightVibhu [00:25:07]: Run on different hardware. Looking at other model providers, okay, you kicked off a inference speed race on the consumer end. What goes into keeping quality the same across them, right? Sure, you can run benchmarksAli [00:25:22]: YeahVibhu [00:25:22]: But, like, how do you determine how much quantization are there standards? What goes intoPhilip [00:25:27]: There's a few things on quality. Most inference optimizations are lossless. KV caching, for example. You are just recomputing or preventing recomputing the same values. Speculation, of course, if a draft token is wrong, it gets rejected. The main lossy optimization is quantization. And that really comes down to, number one, data format, number two, which parts of the model you choose to quantize, which layers, and number three, like doing a lot of calibration on the quantized weights, to ensure that you're preserving all the outliers. There's other tricks that you can do, though. A big one is long context, ‘cause one thing you asked at, right at the beginning is, “Oh, what's gonna happen if I send a 200,000 token request in?” So with a long input sequence, you need to, store a lot more information. You need to process a lot more tokens. And so even if a model has a context of a certain length, you might, as an inference provider, choose to build an API with a shorter context length, and of course a full length one as well. Because if someone doesn't need the full million token context, for example, you can get them better performance. I don't know if that's exactly like quality of the model. The way that I think about quality is to what degree are we faithfully serving the original model? If you think of a golden implementation of a model that performs exactly the way the model is designed to perform, I think of quality as how close are we getting to that, 100% fidelity of the model.Philip [00:27:13]: You can also, of course, think about quality from the training side and how do you push yourself past 100%. But when I think about purely inference optimizations, it's getting faster while staying as close to that 100% fidelity mark as possible. And certainly our standard internally is that, like you should not be able to tell the difference between our API and a, official API. I think Kimi in particular does a good job of vendor benchmarking hereAli [00:27:41]: YesPhilip [00:27:41]: Where they haveAli [00:27:42]: They released an actual vendor benchmark.Philip [00:27:43]: Exactly, yeah.Ali [00:27:44]: ‘Cause they accused, some people, Amazon? There was some provider that was not doing very well on Kimi's benchmark.Philip [00:27:50]: Yeah.Philip [00:27:51]: So, with Reflect we probablyVibhu [00:27:52]: This was a long time ago, right?Philip [00:27:54]: No.Ali [00:27:54]: Yeah, like threeVibhu [00:27:55]: They alsoAli [00:27:55]: Four, five months agoVibhu [00:27:57]: This also happened with, I don't remember which model, but they pulled out quite a few, and then they started a whole chart about this. It might have beenPhilip [00:28:03]: Kimi Vendor Verifier.Ali [00:28:04]: Yeah.Philip [00:28:05]: Yeah.Ali [00:28:05]: Yeah, ‘cause you, ‘cause you'd be pissed, right? Like if you'Philip [00:28:07]: Yeah.Ali [00:28:07]: If like if I'm a consumer and I'm using like Amazon's endpoint for instance, and I've used Kimi and I'm like, “Oh my God, like this is bad,” I'm not gonna say, “Oh, Amazon quantized the model in a bad way.” I'm gonna say, “Oh, Kimi sucks.” Right?Philip [00:28:17]: Yeah.Ali [00:28:17]: So it seems like that makes sense.Philip [00:28:19]: Yeah, they care. They care.Vibhu [00:28:21]: Justifiably.Ali [00:28:21]: Yeah, justifiably.Vibhu [00:28:22]: This is probably a stupid question, but just checking, has anything improved from main quantization?Philip [00:28:28]: Yeah.Vibhu [00:28:28]: Like, is quantization always strictly worse?Ali [00:28:30]: Well technicallyVibhu [00:28:32]: NoAli [00:28:32]: It's a lossy. QuantizationPhilip [00:28:33]: YeahAli [00:28:33]: Is a lossy, it's a lossy implementation.Philip [00:28:36]: Speed improvesVibhu [00:28:36]: Speed improves.Ali [00:28:37]: It the number, likeVibhu [00:28:38]: No, I' always look for inverse scaling laws.Philip [00:28:40]: Yeah.Ali [00:28:40]: Yeah.Vibhu [00:28:40]: This is something I learned from Noam Brown, where like things that normally act in one direction sometimes do.Philip [00:28:45]: Well, technically when you run a benchmark, because these models are deterministic, sometimes your,Ali [00:28:52]: YeahPhilip [00:28:52]: NVFP4 quant is like, two basis points higher than yourAli [00:28:56]: No, it's noise. It's noise.Philip [00:28:57]: Yeah, exactly. I'm like, yeah, it's, it's within. That's why I always say within margin of error.Philip [00:29:01]: And I stopped saying that because everyone assumes that what is, well, within some margin of error, we're barely inside of that to the worst, so we're saying. But yeah, sometimes it's just like, gives you a higher output score. But like Ali said, that's noise. To my knowledge, you're not necessarily making the results better. You're just trying to, again, like keep your fidelity as close to 100% to the original model.Layer Selection, KL Divergence, and Better QuantizationAli [00:29:27]: There is, to your point, research that we did on MP. I don't know if you are able to pullPhilip [00:29:31]: YeahAli [00:29:32]: A tweet we did. One of our research interns, Joshua, I think it's a tweet on how we have 20% better quantized GLM-5.2 than NVIDIA. Essentially what we found throughout like this month research is, okay, quantization is a lossy. It's. You're compressing the data from, occupying 16 bits to occupying, four bits, for instance. And so you're losing some information, and you're trying to minimize that. And so when I say that I'm gonna quantize the model, my job becomes how do I find the layers that I can quantize, and how to find the layers to not. For instance, with image models, I don't quantize modulation layers, and I don't quantize out projections because those two are. Like out projection is what you see as the user. Modulation is what the model sees or understands. Right, exactly. And so to his paper, do you have the. It doesn't have the. Yeah. It's a long paper. I don't know if I can findVibhu [00:30:25]: If there's a part to search or it's probably in the thread.Ali [00:30:28]: It's probably in the thread.Vibhu [00:30:29]: Yeah.Ali [00:30:29]: But the long and the short is it is very possible that quantizing more of the model makes the results. Like if I have a model that I quantize layers one, five, and 10, and another model where I only quantize layers one and It is possible that the model in which I quantized more information is going to perform better because the quantization errors have canceled out. And so what Joshua showed in his mathematical proof where he had like a verifier in, is that you can predict which layers are going to have quantization errors that will cancel out with each other, and you choose to quantize those layers. And so the result of doing this mathematical quantization is you end up with a model that's 20% more quantized than another provider, so you get 20% more throughput of it because there's more layers than running an NVFP4, and your quality is better than that other quant because the layers that you chose to quantize have their errors cancel out, like one layer skewed to the right one layer skewed to the left, one layer skewed to the right. Your final logits distribution is more similar to the original distribution of the model, so you have better fidelity. And so the way we proved this was with KL divergence. So instead of just scoring on the benchmarks, we scored the KL divergence between the logit distribution of the quantized model and the logit distribution of the original full precision model, and we showed that with this technique we get. If your probability distribution on the logits which token it wants to select is more of the same as the original model, you're probably gonna end up staying true to the original model. So yeah, so it seems like previously before this, it seemed like the industry was, well, the more you quantize, the worse it's gonna be, ‘cause the more loss you introduce. That's not exactly, not necessarily true. So yeah, doesn't improve it, but can cancel out.Philip [00:31:57]: I think it might be this, but reminds me a good bit about pruning where you can prune off certain layers.Philip [00:32:03]: But very interesting. Didn't know this was a whole paper you guys put out.Ali [00:32:06]: It's. Fun fact, it was originally 72 pages, this paper, and then we decidedPhilip [00:32:11]: WowAli [00:32:11]: We can't tell. We couldn't release it. So it's now 45.Swyx [00:32:15]: Still 39 pages, so very substantive. We talked about evals and all these things and, like what's possible in terms of speedup? Like it's like probably like the numberInference Speedups and BenchmarkingSwyx [00:32:25]: Thing that people do wanna care about, and it's something that you wrote about in your post. Like official API is 70 tokens per second, and you push it up to 90. Is that like a normal thing?Philip [00:32:36]: So what's cool about working in inference, the reason that I think inference is going to be a useful place to do engineering for a long time, is that if you look at highly optimized domains like, say, finance, if you're in finance, you measure how much better you got in basis points. It's like, “Oh, I got five basis points better, like twentieth of 1% better,” that's huge news because everything is so optimized. When we publish optimizations, it's 20%, it's 100% it's 200%. So there's still probably like a lot further to go, honestly. Like you'll, you'll know that inference is pretty much solved when researchers start publishing about how they got 1% faster at something.Swyx [00:33:19]: Which by the way, because I am from the finance background, in the ‘70s, that was the margin at the time. When you did quantitative finance research, you would findAli [00:33:27]: And like 20%, tens of percent.Swyx [00:33:29]: That's. Yes.Philip [00:33:29]: Yeah.Swyx [00:33:30]: And now it'Philip [00:33:31]: Tiny fractionsSwyx [00:33:32]: For those people interested, look up Andrew Lo's paper. He had a really interesting illustration of quant, stat arb, distribution, narrowing down from like those kinds of 20% differences in the ‘70s, down to nothing today, which is very cool.Philip [00:33:48]: Exactly, and we're at the beginning of the same type of thing. Now benchmarking is hard. I think anyone will tell you that, and benchmarking provider speeds is hard because there's so many variables that go into it. What hardware are you using? How much load do you have on the system? What's the exact nature of the prompts and input and output sequence lengths? All that stuff. But overall, when you start stacking these improvements, you're looking at multiples. You can look at it. The most common form, of course, is TPS, tokens per second, which is bad naming by us in the industry, ‘cause there's two tokens per second. There's tokens per second, the throughput number, and the latency number.Ali [00:34:31]: TTMT, yeah.Philip [00:34:32]: Like total tokens per second out of the, out of the GPU as a throughput number. Most people only care about tokens per second as the latency number, which we should call ITL, intertoken latency, but we don't.Philip [00:34:44]: Anyway, so you can imagine a standard API without many optimizations for a 1 trillion parameter model operating somewhere in the 30 to 50 tokens per second range for reasonable traffic profile. And we generally see the goal of, pushing to 10X that. But, not necessarily day zero, but by stacking enough optimizations, if you have, say like four optimizations, each of which doubles performance. Or sorry, three optimizations, each of which doubles performance, then you stack that up, that's an 8X gain. That's the order of magnitude that we're working with in this space. We're trying to make things substantially faster, not just go from like 70 to 90.Swyx [00:35:38]: Are you saying you've. You have done that?Philip [00:35:40]: So let's say you have as a reasonable baseline, 30 or 40 tokens per second. You can achieve 10X that. So like on GLM-5.2, if you run it unquantized, perhaps on H100s even, and you're just using an off-the-shelf inference engine with no particular optimizations, no speculator, nothing extra around like KV routing, no disaggregation, you're, you're probably, yeah, looking at that like 30 to 40. You think that's like a reasonable baseline?Swyx [00:36:12]: Right. Right.Philip [00:36:12]: To get to something like 10X, there's a lot of trade-offs that you're making. If we're running at more like a 300, 400 tokens per second range, you are using the best hardware possible. You have a optimized speculator. You have done all of your quantization work. You are Seeing a pretty high cache hit rate. You are running with a reasonably small batch size and a parallelism configuration that is tuned for latency versus throughput, but it is possible. So the spreads that you see if you, like, go on artificial analysis or you go on OpenRouter and you look at, the worst provider to the best provider, oftentimes can hit that range. 10X is of course very aggressive. It's oftentimes maybe more of a four to six times improvement. But that's the performance that makes us really excited, is when we can get these huge gains, not just go from 70 to 90 tokens.Stacking Optimizations: NVFP4, Speculation, and DisaggregationAli [00:37:19]: It's also, like, hardware dependent. Like, ifPhilip [00:37:20]: YeahAli [00:37:20]: If you have a thing where you're serving it on just, like, a node of H100s and then you throw, like, you shard the model across, like, four nodes of B200s. Like, you can definitely increase the speed with just throwing more hardware at it. Like, normalizing for the same exact hardware and the same number of GPUs.Philip [00:37:35]: Yeah. Then you're looking at, like, a two to 4X improvementAli [00:37:38]: Right. RightPhilip [00:37:38]: Depending on the inference optimizations. So yeah, it's. Some of it's, what's the call, and some of it's who's the driver.Vibhu [00:37:46]: If you break down the two to 4X, say the example is run GLM-5.2Ali [00:37:51]: YeahVibhu [00:37:51]: On B200sAli [00:37:53]: YeahVibhu [00:37:53]: Single node, right? What's, like, the cost trade-off for effort to get, like, the last bit of juice out versus what should people just think of, right?Ali [00:38:01]: Spectre quantization. Yeah.Vibhu [00:38:03]: Spectre quantization.Ali [00:38:04]: That's, that's, that's like 95%. LikeVibhu [00:38:06]: And how far does that get you? And how easy is that for the average person to do? So say right I wanna throw the weights of GLM-5.2 on a node of B200s, how easy is it to find speculative decoder- decoder model or already quantized model? How much work goes into it?Philip [00:38:23]: If you're doing it up front, it's quite a lot of work. If you're doing it today, there's going to be people who have published things that you can just, you can just grab some NVFP4 weights. You can grab a speculator. Yeah, if we're thinking about, like, what are the 2Xs we're stacking, going from, BF16 to NVFP4 is, it's not quite a 2X, right? It's like. I think it's about, like, 30 to 40%, from 16 to 8, and then another 30 to 40% multiplied from, 8 to 4. So that doesn't quite get you a 2X, but, like, roughly a 2X. Speculator, roughly a 2X. Disagg on top of that if you're able to get enough hardware and put enough traffic through it, another roughly a 2X. And then you add in some, double-digit percent increase from having just a better runtime with, the latest kernels and stuff behind it. And that's how it stacks up.Ali [00:39:21]: YeahPhilip [00:39:21]: So building each of those, like, building the, quantized weights is, for someone who really knows what they're doing, hours to days of work. Building the speculator, again, like, hours to days of work. And the, disagg setup, hours to days. Well okay, but like once you haveAli [00:39:39]: Once set up. Once set up. YeahPhilip [00:39:40]: Yeah, getting disagg working for the first time, I'm saying, of course, is very difficult.Philip [00:39:44]: The marginal implementationAli [00:39:48]: Like, if you're just grabbing, like if you are a person, like just a normal consumer who has access to, like, a node of B200s and you're wondering, “How can I just host it myself?” You don't need to quantize the model yourself. There's always gonna be, like, an open source quantized checkpoint. NVIDIA's gonna push one out if no one else does. You. Usually, the providers will have their own spec dec that they've trained as well. You don't need to train your own spec dec. You can just use that as well.Philip [00:40:09]: Yeah. Like, GLM-5.2 has its own MTP.Ali [00:40:13]: Right. Right.Vibhu [00:40:14]: What's multi token prediction?Philip [00:40:15]: Yes.Ali [00:40:16]: I'm justVibhu [00:40:16]: Can you explain that?Ali [00:40:16]: I'm just an expert.Ali [00:40:18]: I can do it for you in case I get it wrong?Vibhu [00:40:20]: No.Vibhu [00:40:21]: Yeah, you should correct if we're wrong, but their multi-token prediction can be used for self-speculative decoding.Ali [00:40:27]: I'm not sure. I'm not gonna correct that.Vibhu [00:40:28]: Okay. I'm semi-confident in thatAli [00:40:30]: Okay. YeahVibhu [00:40:30]: But someone can check. But it's useful to paint the story of, okay, not just the average person, but say a company wants to switch from serverless inference I wanna throw this up on. I wanna rent some GPUs, throw it up. These are the steps you take to do significantly faster than just put it behind vLLM.Ali [00:40:48]: Right.Vibhu [00:40:49]: I was waiting for a mention of Dynamo.Vibhu [00:40:51]: I feel like, that's supposed to be the baseline that you measure against.Dynamo, KV Routing, and Disaggregation ToolkitsPhilip [00:40:55]: I would think of Dynamo as less of a box system and more of a toolkit for building with. So when we talk about doing aware routing, when we talk about doing KV offloading, when we talk about doing, PD disaggregation, Dynamo fundamentally is. By the way, Dynamo is an open source library from NVIDIA.Ali [00:41:17]: We've done a pod with KylePhilip [00:41:18]: OkayAli [00:41:19]: Kyle Cranin.Philip [00:41:19]: Cool. So then your listeners know then that it supports all the different inference frameworks. And it is multi hardware, which is interesting.Ali [00:41:28]: But it's just a router, it's not like an optimizer layer.Philip [00:41:30]: Yeah. All it does, like, what Dynamo is good at, it is a library for moving information around your cluster, around your hardware. So if you have, KV cache on one place and you need it to be somewhere else, Dynamo coordinates NIXL for you to move that around.Philip [00:41:49]: That doesn't mean that, like, out of the box, you just say, “Pip install Dynamo,” and then you get, like, a massive performance speed up. It's more of a developer toolkit.Ali [00:42:01]: Yeah. I would have said it would. It comes with a set of defaults that you can then swap out.Philip [00:42:06]: It does. If the industry at large, I think, was, like, rolling out all of these deployments, standard, then I think it would be, like, a credible baseline. But, we've got to, we've got to benchmark against, like, what we're seeing in the wild.Speculative Decoding Methods: Medusa, EAGLE, n-Gram, and Spec-SpecVibhu [00:42:23]: I did wanna talk a little bit more about PD disagg, because that is probably, like, number three after quantized and speculative decoding. In your book though, I was just gonna pull out the book.Philip [00:42:31]: Yeah.Vibhu [00:42:32]: Like section 522 on Medusa, 523 on EAGLEPhilip [00:42:35]: YeahVibhu [00:42:36]: 524 on gram.Philip [00:42:37]: It's 55, would be disaggregationAli [00:42:42]: Yeah. Well, no, I just wanted to dwell a little bitPhilip [00:42:44]: YeahAli [00:42:44]: The other. Like, so what do you choose to include? What do you choose to not to include? Because there was all these other techniques.Philip [00:42:51]: Yeah.Ali [00:42:51]: Are these still relevant? Because I think they came out, like, a year and a half ago maybe.Vibhu [00:42:55]: Medusa is quite old.Philip [00:42:56]: Yeah, Medusa's old.Ali [00:42:58]: It was old.Vibhu [00:42:58]: But is it in the book as a good, here'sPhilip [00:43:01]: BaselineVibhu [00:43:01]: Baseline vanilla understand it?Philip [00:43:02]: Like you should know this.Vibhu [00:43:03]: Like I read the paper, I'm like, “ it makes so much sense.”Philip [00:43:05]: Yeah.Philip [00:43:05]: So with the book, I had a couple goals. One was to give people just a working vocabulary for the space as a whole, and the other was to give them some intuition about how each of these techniques works. As I mentioned in my AI Engineer talk, which is the first public addendum to this, the speculation space has moved much faster than everything else. So yeah, even at the time that I wrote the book Medusa, I very much included as a way for people to understand how the space evolved rather than what the most modern technique is. And now of course, there's DFlash, dSpark. There's, there's newer techniques even than EAGLE, although EAGLE is still very commonly used.Ali [00:43:51]: SpecSpecta.Philip [00:43:52]: Yes. Speculative decoding.Vibhu [00:43:54]: What canAli [00:43:56]: Oh, it's a paper by Tri Dao and it's like, it's doing speculative decodingVibhu [00:44:00]: HuhAli [00:44:01]: For the speculative decoder.Philip [00:44:02]: Oh, in spec- oh my God.Ali [00:44:02]: It's literally just an another. It's like, yeah, that's the most simple way to explain it, and it seems like he got trivial speed ups there. But it seems that the complexity with training, it's almost like in our mind at least, it's almost as complex as training GANs. Like it's like a very delicate balance and oftentimes you, it's just but yeah, it's literally speculative decoding on speculative decoding.Vibhu [00:44:21]: Speculative.Ali [00:44:22]: Yeah. We saw this paper.Vibhu [00:44:24]: It's interesting, right?Ali [00:44:24]: Yeah.Vibhu [00:44:24]: I wouldn't even expect it to be very particular to train, I wouldAli [00:44:29]: Right.Vibhu [00:44:29]: The naive part of me is like, okay, train speculative decoder.Ali [00:44:32]: But like, and it makes sense, like the whole idea of speculative decoding is you. It's like, it's like almost like the iPhone auto predict version but for a normal model, right? Like you're just, you're just, generating three tokens and you're like, okay, I'll do prefill on them. And so you save those three turns for your original model. Now your speculative decoder is doing three turns of auto regression, so why not just have an even smaller model?Ali [00:44:53]: The other question there is what are the size of speculators? So say forPhilip [00:44:58]: Right. It's like a billion parameters.Ali [00:45:01]: Like for MiniMax, it's. Yeah. It's like one layer. It's like one 60th of the original model usually.Philip [00:45:06]: Yeah. I think we should do a paper when we get back to the office.Philip [00:45:10]: SpeculativeAli [00:45:11]: SpeculativePhilip [00:45:11]: Decoding.Ali [00:45:13]: No, it's, it does seem like how, when do you stop? But then it also seems like if you're able to train spec-spec decode for instance, right? Like if you're able to have a small model that is accurately predicts what the intermediate speculator is gonna predict, that is able to predict what the original target model's gonna predict, then why not just use that smallest model directly, right?Vibhu [00:45:34]: Yeah. This isAli [00:45:35]: Like it seems likeVibhu [00:45:35]: Adjacent to the routing problem.Ali [00:45:36]: Right.Vibhu [00:45:36]: Yeah.Ali [00:45:36]: Right.Philip [00:45:37]: The thing with speculators is one of the practical constraints on using them is that you do have to run a small model on the same hardware that you're running the big model on. There is a orchestration and resource competition problem inherent in that, and that is one of the constraints on speculation in general, is that draft tokens cost resources to create and cost software complexity to manage. And so if you have like infinitely recursive speculators, you add in quite a bit of that complexity on the actual implementation within the inference engine as well, not just in the training process.Vibhu [00:46:17]: I was gonna say, I would wonder if you could do similar, like distillation and pruning of, it's the same thing, it's just a model. Can we not just distill a lot of the weights, quantize the speculator, out of my domain? The question that also comes up is, this is all for big server workloads, right? How much of this applies to, say I have this MacBook, I wanna run Gemma really efficiently. Similar problems, not the same?Local AI vs. Data Center InferencePhilip [00:46:45]: Pretty different. I talked to Selo, about this on his podcast a couple weeks ago. The difference between inference engineering for the data center and for production workloads versus inference engineering for local AI, is that we start with fundamentally like different constraints and different goals. With local AI, it's how do I fit this model onto my hardware and then make it less dumb? And with data center influence, it's how do I load this model and then make it less slow? And we care about less dumb, and they care about less slow. But the local AI inference engineering ecosystem, I think has a lot for us to learn from in the data center space. They are experts in various forms of quantization, including dynamic quantization that we just don't touch, in the pruning, in the distillation, in the, layer removal. There'Ali [00:47:42]: Layer removal matters less.Philip [00:47:43]: Yeah. There'Ali [00:47:44]: No one loves pruning really.Philip [00:47:45]: Yeah. Well, but the, but they doVibhu [00:47:46]: Which is surprising, right? But that's, that's a whole different thingPhilip [00:47:48]: Just to fit something on the laptop.Ali [00:47:50]: Right.Philip [00:47:50]: So yeah, it's a, it's an interesting, it's an interesting space. Not necessarily that like their techniques make sense for us to do in the data center, because we have different resources and different goals, but more that the process as well as the openness of that field is something to, admire.Ali [00:48:12]: Yeah. Like to your point, like, certain optimizations that would. Like for instance, Turbo Quantum Sharper, like it made such huge hype on that and we did like a whole deep dive on Twitter and like said, what is it? How does it work? Why is it good or not? And it took off and it was implemented on local devices because your memory bandwidth is so slow on like a MacBook, for instance. But try putting the same thing on like an NVIDIA GPU on a B200 Turbo quant would not be. Like, it would not be used. Like, NVIDIA - Like, NVIDIA made it clear that this is not a good optimization, and we've seen it firsthand where the overhead of doing dequantization, quantization of, in the kernel itself with turbo quant kernel, each end is much slower than the time that you save from doing the bandwidth. ‘Cause on the B200s, you have like 3.5 terabytes per second. You don't need decrease the storage that much. You don't need to do, FP4 KV cache. You don't need to use a requant. There's, there's, there's better optimizations to be made. But on Edge devices, it's extremely important, it's extremely useful. So, seems to be, like, different optimizations there, but then they're all uniquely combined with like all you wanna quantize the model, you wanna do speculative decoding, like certain common prefixes with bothPhilip [00:49:18]: Principles.Ali [00:49:19]: Yeah, exactly. Exactly. Exactly.Philip [00:49:20]: They also do a lot of work on, model parallelism, especially over, heterogeneous topology, where you have, some sparks and they are wired together with, Ethernet, DGX sparks.Ali [00:49:35]: Yeah, this is the Exo Labs guys.Philip [00:49:36]: Yeah. You have, a nu

Salon Rising
The hair catastrophe that launched a brand – with Simone from Apotecari

Salon Rising

Play Episode Listen Later Aug 2, 2026 60:32 Transcription Available


Shiny scalps. Sparse hairlines. Clients who used to have a mane and now don't. You're having these conversations every week and you need better answers than the ones on the shelf.Simone from Apotecari has them. She's a naturopath, twenty years qualified, with a career spent formulating supplements for brands already sitting in your bathroom cabinet. She also had a hair catastrophe of her own, went looking for something that would fix it from the inside, and found nothing worth taking. So she built it.These are clients you already have. The one whose hairline is quietly moving back. The one on a weight loss injection who hasn't thought to mention it. The one who's oily by day two and blames your shampoo. Simone explains what's happening in each of them, and why the answer is almost never another bottle on the shelf.What we coverThe bond builder era and what a lot of stylists were really doing with itWhy hair is the first thing your body stops feeding when nutrients are shortPerimenopause hair loss and why nobody connected the dots for so longShiny scalps, receding hairlines and what is actually happening in the follicleOily scalps, dandruff and why washing more often changes nothingWhy your scalp ages seven times faster than the skin on your faceThe difference between hair shedding and hair loss, and why the word you use changes everythingWhy the hair industry is not great at doing its own researchWhat a proper ninety day product trial looks like, with data instead of feelingsShe hid her hair for six months. Then she went and fixed it for everybody else.Frequently asked questionsWhy does hair fall out during perimenopause?Hormonal change alters how follicles behave, which is why thinning often shows down the centre part and around the hairline. It rarely happens alone, and it sits alongside other perimenopause symptoms that are just as real and just as connected.Does washing your hair more often make it less oily?No. Washing removes the sebum already on the scalp but does nothing to the rate your body produces it. Balancing oil production happens internally and usually takes two to three months.How long should you take a hair supplement before judging it?Give it ninety days minimum, because that tracks with the hair growth cycle. Take before photos and test porosity and elasticity first, so you are working from data rather than a feeling.Mentioned in this episode:https://www.salonrising.com

Low-Noise
Siouxsie and the Banshees – The Scream | Before Goth Had a Name?

Low-Noise

Play Episode Listen Later Jul 20, 2026 35:31


What happens when punk stops trying to shock and starts trying to haunt?In this episode, we revisit The Scream, the extraordinary debut album by Siouxsie & the Banshees. Nearly fifty years on, it still feels unlike anything else released during the first wave of punk. Sparse, unsettling and fiercely original, The Scream pointed towards post-punk before the genre even had a name.From John McKay's jagged guitar work and Kenny Morris's restless drumming to Siouxsie's commanding, enigmatic voice, we explore how the band created an atmosphere that was as important as the songs themselves.In this episode I am in discussion with Dr. Andrew Webber.https://www.facebook.com/share/1F15mx4ea3/https://buymeacoffee.com/lownoiseWhy buy me a coffee?Low Noise is proudly ad-free. If you would like to to say thank you for any of the content you have enjoyed (and help support the continuation of creating more), the above link provides a way to make a small donation of your choice (I also function on coffee!).

The David Pakman Show
The mistakes are adding up quickly

The David Pakman Show

Play Episode Listen Later Jul 1, 2026 79:28


-- On the Show: -- Marc Elias, American elections attorney and founder of Democracy Docket, joins us to discuss the recent Supreme Court decisions and Republican efforts to suppress the midterm vote -- Donald Trump claims Iran is requesting negotiations, but Iranian officials publicly deny any direct meeting -- Donald Trump argues Congress can end birthright citizenship through legislation even after the Supreme Court reaffirms protections -- Sparse crowds at Donald Trump's Great American State Fair repeatedly contradict officials who describe the event as packed and enthusiastic -- Donald Trump reports $1 billion in crypto-related income while many Americans continue struggling with grocery, housing, and utility costs -- Donald Trump defends making money while president, dismisses questions about his financial disclosure, and praises his Qatari-gifted jet -- Nearly two weeks after Mitch McConnell is hospitalized, his office still does not explain his condition or confirm whether he has been discharged -- Donald Trump announces the first Republican national midterm convention, making his presidency the central focus of the campaign -- Jeff Landry warns that transgender athletes threaten sports even though participation numbers remain extremely small across -- On the Bonus Show: Takeaways from the Colorado primaries, Tom Kean returns to Congress after going missing, frequent AI chatbot users are more likely to believe anti-vaccine myths, and much more...

The Hot Tub Podcast
340 - "Very sparse, coarse chest hair"

The Hot Tub Podcast

Play Episode Listen Later Jul 1, 2026 46:02 Transcription Available


Mauler doesn't want to touch your damp washcloth, Rush saves his fancy cane for only the nicest events, Jenni's robot toilet takes a tumble down the stairs, and Brady gets stuck on the stairwell behind his neighbour's pet pony. Love the podcast? Leave us a review!

The Leslie Marshall Show
Trump's Iran Deal Already Falling Apart?; Great American Fair's Sparse Crowds

The Leslie Marshall Show

Play Episode Listen Later Jun 30, 2026 41:13


The guest host for today's show is Brad Bannon. Brad runs Bannon Communications Research, a polling, message development and media firm which helps labor unions, progressive issue groups and Democratic candidates win public affairs and political campaigns. His show, 'Deadline D.C. with Brad Bannon,' airs every Monday from 3-4pm ET.  Brad is first joined by CNN Military Analyst, Col. Cedric Leighton (USAF-Ret.). The pair analyzes whether Trump's peace deal with Iran is already falling apart before the 60-day negotiation process has even been fulfilled. Brad also asks Cedric whether the Strait of Hormuz is actually still open for commercial shipping vessels. Additionally, they discuss the latest news on the war between Russia and Ukraine, and whether there is pressure building to remove Putin from power internally. Then, Tara Devlin, host of the 'TARABUSTER' podcast, joins Brad to discuss the bigger meaning of the sparse crowds at the Trump organized 'Great American Fair.' Col. Leighton's website is www.CedricLeighton.com and his handle on BlueSky is @cedricleighton.bsky.social. Tara Devlin is host of the progressive podcast 'TARABUSTER,' which can be found at www.tarabuster.com. Her handle on BlueSky is @tarabuster.bsky.social. Brad is on the National Journal's panel of political insiders, is an American political analyst for The Times of India TV, and is a national political analyst for WGN TV and Radio in Chicago and KNX Radio in Los Angeles. Brad also writes a political column every Sunday for 'The Hill.' You can read his new Substack called, 'The Bannon Ballot Blast,' at www.bradbannon.substack.com. His handle on BlueSky is @bradbannon.bsky.social.

Progressive Voices
Leslie Marshall Show - Trump's Iran Deal Already Falling Apart; Great American Fair's Sparse Crowds

Progressive Voices

Play Episode Listen Later Jun 30, 2026 41:13


The guest host for today's show is Brad Bannon. Brad runs Bannon Communications Research, a polling, message development and media firm which helps labor unions, progressive issue groups and Democratic candidates win public affairs and political campaigns. His show, 'Deadline D.C. with Brad Bannon,' airs every Monday from 3-4pm ET.  Brad is first joined by CNN Military Analyst, Col. Cedric Leighton (USAF-Ret.). The pair analyzes whether Trump's peace deal with Iran is already falling apart before the 60-day negotiation process has even been fulfilled. Brad also asks Cedric whether the Strait of Hormuz is actually still open for commercial shipping vessels. Additionally, they discuss the latest news on the war between Russia and Ukraine, and whether there is pressure building to remove Putin from power internally. Then, Tara Devlin, host of the 'TARABUSTER' podcast, joins Brad to discuss the bigger meaning of the sparse crowds at the Trump organized 'Great American Fair.' Col. Leighton's website is www.CedricLeighton.com and his handle on BlueSky is @cedricleighton.bsky.social. Tara Devlin is host of the progressive podcast 'TARABUSTER,' which can be found at www.tarabuster.com. Her handle on BlueSky is @tarabuster.bsky.social. Brad is on the National Journal's panel of political insiders, is an American political analyst for The Times of India TV, and is a national political analyst for WGN TV and Radio in Chicago and KNX Radio in Los Angeles. Brad also writes a political column every Sunday for 'The Hill.' You can read his new Substack called, 'The Bannon Ballot Blast,' at www.bradbannon.substack.com. His handle on BlueSky is @bradbannon.bsky.social.

The John Batchelor Show
S8 Ep971: Henry Sokolski analyzes China's nuclear capabilities, including missile silos and underground transport systems, while questioning their peer-to-peer ambitions. He also observes economic trends, noting that gas price fluctuations and sparse Cos

The John Batchelor Show

Play Episode Listen Later Jun 5, 2026 17:18


Henry Sokolski analyzes China's nuclear capabilities, including missile silos and underground transport systems, while questioning their peer-to-peer ambitions. He also observes economic trends, noting that gas price fluctuations and sparse Costco crowds suggest consumers are becoming increasingly budget-conscious and selective about their spending habits in the current economy.1958

Dr. Chapa’s Clinical Pearls.
Hantavirus & Preganncy FAQ

Dr. Chapa’s Clinical Pearls.

Play Episode Listen Later May 11, 2026 16:33


Hantavirus was first discovered in the early 1950s near the Hantaan River in South Korea. The US has seen this before: the 1993 Four Corners outbreak was the first recognition of the virus in the United States, causing a deadly respiratory syndrome. Now, Hantavirus is in the news again with 17 Americans currently (5.10.26) enroute back to the US for specialized observation. In this episode, we will briefly review what this virus does and cover the SPARSE data we have regarding hantavirus infection in pregnancy. 1. Gilson GJ, Maciulla JA, Nevils BG, et al. Hantavirus Pulmonary Syndrome Complicating Pregnancy. American Journal of Obstetrics and Gynecology. 1994. 2. 5.10.26: https://www.nbcnews.com/health/health-news/hantavirus-stricken-cruise-ship-arrives-tenerife-rcna3443183. Janwadkar RS, Ritchie HM, Johnson CA. Unexpected Challenges: A Case Report of Hantavirus Infection in a Pregnant Patient in a Rural Emergency Department. The Journal of Emergency Medicine. 2025.

Sausage of Science
SoS 277: Catalina Fernández discusses a new causal model of human growth using temporally sparse data

Sausage of Science

Play Episode Listen Later Apr 27, 2026 35:35


In this episode, Dr. Catalina Fernández explains a new theoretical model of human growth and its opportunities for cross-sectional and diverse samples. Dr. Catalina Fernández is an Assistant Professor in the Department of Anthropology at Florida Atlantic University (United States). Her research focuses on the role of food and diet in human adaptation and evolution among contemporary populations. Drawing on evolutionary and biocultural frameworks and employing mixed methods, her work investigates how subsistence strategies, nutritional histories, and the environment shape genetic, physiological, and cultural adaptations. She is particularly interested in questions related to the consequences of global market integration for human health and well-being among rural and small-scale societies. She has experience working with rural and Indigenous communities in Latin America, addressing issues related to environmental and dietary adaptations, nutrition transition, chronic disease risk, and population genetics. Her most recent research project investigates the causes of variation in child growth trajectories among non-Western populations, aiming to better inform public health interventions using culturally and environmentally appropriate strategies. Building on this work, she is developing a research program that examines the life-course health outcomes related to water and food security resulting from the climate change–driven expansion of the mining industry among indigenous communities in Chile. Contact Dr. Fernández at catafernandezh@gmail.com ------------------------------ Find the paper discussed in this episode: John A. Bunce, Catalina I. Fernández, Caissa Revilla-Minaya; A causal model of human growth and its estimation using temporally sparse data. R Soc Open Sci. 1 August 2025; 12 (8): 250084. https://doi.org/10.1098/rsos.250084 ------------------------------ Contact the Sausage of Science Podcast and the Human Biology Association: Facebook: facebook.com/groups/humanbiologyassociation/, Website: humbio.org Chris Lynn, Co-Host, Website: cdlynn.people.ua.edu/, E-mail: cdlynn@ua.edu, X:@Chris_Ly Mecca E. Howe, Co-Host, E-mail: howemecca@gmail.com, LinkedIn: https://www.linkedin.com/in/mecca-howe/

Dr Mary Travelbest Guide
Dr. Mary Travelbest - Thessaloniki Greece Part 1

Dr Mary Travelbest Guide

Play Episode Listen Later Mar 13, 2026 9:44


Where in the world am I? In San Diego, talking about Thessoloniki Greece, Part 1 Welcome to the  Dr. Mary Travelbest Guide podcast. I returned from a 90-day journey around the world, and I'm excited to connect with fellow travelers and share experiences for world peace. Here is an FAQ about plane or train travel, Thessoloniki Greece, Part 1, and also about a health issue you don't want when you travel. Give a listen. I guide you to solo travel experiences to bring out your best. The FAQ is: If you could take a plane or a train, which would it be and why? Answer:  If I have the choice between a plane and a train, Most of the time… I choose the train. Now let's be practical. If the distance is extreme — say, cross-country or intercontinental — the plane wins on efficiency. At this stage of life, I value my energy. Six hours in the air may beat twenty hours of transfers. But when are both realistic options? Train. Here's why. First, the train allows me to arrive gently. There's no stripping down at security, no liquid anxiety, no rushing to a distant gate. I walk onto the train. I keep my water. I keep my dignity. That matters. Second, the scenery. At 50+, we understand that the journey is not separate from the destination. On a train, I see villages, farmland, people waiting on platforms, laundry on balconies. I watch life unfold. A plane gives me clouds. Third, ease of movement. I can stand up. Walk. Stretch. Visit the café car. Talk to someone if I choose — or not. For solo women, that flexibility feels empowering. Fourth, arrival point. Trains typically drop you in the center of town. Planes drop you 40 minutes away, followed by taxis, shuttles, and more logistics. Simplicity wins. Now — here's where I get skeptical of my own bias. If I'm exhausted… If connections are complicated… If safety or night travel becomes a concern…Going from Oslo to Bergen this past summer, we had a 7-hour delay, stranded in Voss due to the heated tracks. That was not unusual, I later learned. Side note: I did enjoy my time in Voss and learned to slow down. If I anticipate a delay like this, I will absolutely take the plane. Comfort and safety override romance. So my answer? If time is short and distance is long,,,,, fly. If time is flexible and distance is reasonable, take the train and let the world move past your window. At this stage of life, we're not just getting somewhere. We're experiencing how we get there. And that is the difference.   60-second confidence challenge Your challenge today  Confidence Challenge in Greece and on trains. If you like today's Confidence Challenge, my book series delves deeper into train travel while walking through the 5 steps to solo travel, from easy to more challenging, with foreign-language communication tips. You can find the series at the link in the description.    See Book A for addressing this concern..  Find it on the website​​ at https://www.5stepstosolotravel.com/ or on Amazon. It's a several-part series. Today's destination is Thessaloniki, Greece Part 1 of 2   Greece: my bucket list trip: Arrival, Ancient Echoes, and Modern Reality Welcome to my planned Step 5 travel — the kind where you don't just visit a place… you live inside it. This week and next week, I'm taking you to Thessaloniki, Greece's second-largest city — layered with Roman ruins, Byzantine churches, Jewish history, and modern-day contradictions.    

Don't Fight Backprop: Goodfire's Vision for Intentional Design, w/ Dan Balsam & Tom McGrath

Play Episode Listen Later Mar 5, 2026 107:20


Dan Balsam and Tom McGrath from Goodfire return to explore the frontier of mechanistic interpretability and their new research pillar, Intentional Design. They explain the shift from sparse autoencoders to understanding geometric structure in latent spaces, and share a proof-of-concept method for reducing hallucinations using probes and RL. The conversation tackles concerns about reward hacking, principles for shaping the loss landscape instead of fighting backprop, and what this means for aligning powerful models. They also discuss recent Goodfire results on Alzheimer's prediction, disentangling memorization vs reasoning weights, and how they balance commercial growth with a public benefit mission. Nathan uses Granola to uncover blind spots in conversations and AI research. Try it at granola.ai/tcr with code TCR — and if you're already using it, test his blind spot recipe here: https://bit.ly/granolablindspot LINKS: Detecting PII for Rakuten Interpretability for Alzheimer's biomarker detection You and Your Research Agent Adversarial examples and superposition Discovering rare behaviors with model diff Priors in time for interpretability Belief dynamics in in-context learning Mixing mechanisms in language models Sparse autoencoder scaling with manifolds Sponsors: VCX: VCX, by Fundrise, is the public ticker for private tech, giving everyday investors access to high-growth private companies in AI, space, defense tech, and more. Learn how to invest at https://getvcx.com Claude: Claude is the AI collaborator that understands your entire workflow, from drafting and research to coding and complex problem-solving. Start tackling bigger problems with Claude and unlock Claude Pro's full capabilities at https://claude.ai/tcr Serval: Serval uses AI-powered automations to cut IT help desk tickets by more than 50%, freeing your team from repetitive tasks like password resets and onboarding. Book your free pilot and guarantee 50% help desk automation by week 4 at https://serval.com/cognitive Tasklet: Tasklet is an AI agent that automates your work 24/7; just describe what you want in plain English and it gets the job done. Try it for free and use code COGREV for 50% off your first month at https://tasklet.ai PRODUCED BY: https://aipodcast.ing

Hawaiian Concert Guide
Hawaiian Concert Guide Show 699 - 27 Pineapples

Hawaiian Concert Guide

Play Episode Listen Later Mar 1, 2026 106:31


Hawaiian Concert Guide – Show 699 Theme: He Mele Inoa Opening Set – Gregory Juan (Album: Kauluwehi) He Mele Inoa no Kauluwehi (1:49) Artist: Gregory Juan Album: Kauluwehi Language: Hawaiian We open Show 699 with a traditional mele inoa — a name chant honoring Kauluwehi. In Hawaiian culture, a mele inoa is more than a song; it is a formal proclamation of identity, lineage, and character. These chants carry mana (spiritual power) and often highlight the beauty, traits, and ancestral ties of the person being honored. Listen for: Traditional chant phrasing Sparse, respectful instrumentation Emphasis on pronunciation and cadence Honokahua Nani E (4:02) Artist: Gregory Juan Album: Kauluwehi Language: Hawaiian This song honors Honokahua, an area in West Maui known for its cultural and archaeological significance. The word nani means “beautiful,” and the song reflects deep admiration for the land. Themes: Love of place (mele ʻāina) Natural imagery Cultural remembrance Kamalei Kawaʻa – Album: Mānaiakalani Hālaulani (3:31) Artist: Kamalei Kawaʻa Album: Mānaiakalani Language: Hawaiian A graceful contemporary Hawaiian composition. The title suggests heavenly or chiefly associations (lani meaning heaven or royalty). Kamalei blends traditional phrasing with modern melodic structure. Clean acoustic arrangement Strong falsetto phrasing Contemporary Hawaiian production style Kālepa (3:22) Artist: Kamalei Kawaʻa Album: Mānaiakalani Language: Hawaiian “Kālepa” references a name — possibly a person or a poetic symbol. In many Hawaiian compositions, personal names stand in for cherished relationships or deeper metaphors. Storytelling lyric structure Light, flowing rhythm Clear enunciation of Hawaiian text Kawika Kahiapo – Album: Kuʻu Manaʻo Ka Makani Kaʻili Aloha (5:50) Artist: Kawika Kahiapo Album: Kuʻu Manaʻo Language: Hawaiian Translated as “The Wind That Snatches Away Love,” this song uses classic Hawaiian metaphor, where wind represents emotional change, separation, or longing. Rich acoustic guitar Emotional vocal phrasing Poetic metaphor rooted in natural forces Kaulana Makapuʻu (4:43) Artist: Kawika Kahiapo Album: Kuʻu Manaʻo Language: Hawaiian Makapuʻu on Oʻahu's eastern shoreline is known for its lighthouse and powerful ocean views. This mele celebrates place with vivid imagery — cliffs, winds, and sea spray. Pride of place Coastal imagery Deep knowledge of ʻāina Les Waikīkings – Album: Hapa Haole with a Twist Papio (2:13) Artist: Les Waikīkings Album: Hapa Haole with a Twist Genre: Exotica A playful instrumental shift. “Papio” refers to a young jackfish common in Hawaiian waters. This track blends vintage steel guitar textures and surf-era island rhythm. The Hukilau (1:57) Artist: Les Waikīkings Album: Hapa Haole with a Twist Genre: Exotica A classic hapa haole standard celebrating the communal fishing tradition of the hukilau. The hukilau emphasizes cooperation — everyone pulling the net together. Ho‘okena – Album: Ho‘okena 5 Hawaiian Soul (4:32) Artist: Ho‘okena Album: Ho‘okena 5 Language: Hawaiian Written by Jon Osorio, this powerful anthem honors George Helm, a key figure in the Hawaiian cultural renaissance and the movement to protect Kahoʻolawe. Sovereignty Cultural revival Protection of land Heha Waipiʻo (3:49) Artist: Ho‘okena Album: Ho‘okena 5 Language: Hawaiian A closing tribute to Waipiʻo Valley on Hawaiʻi Island — a place of dramatic cliffs, waterfalls, and deep historical significance. “Heha” conveys awe and admiration. Tight multi-part harmony Traditional lyrical cadence Deep connection to ʻāina Show 699 Flow Summary Traditional name chant and mele ʻāina Contemporary Hawaiian songwriting Emotional metaphor and wind imagery Retro hapa haole exotica interlude Cultural anthem and powerful harmonies A beautiful arc — from honoring a name, to honoring land, to honoring culture itself.

Adventure On Deck
Wide Open Fiction. Week 47: The American Short Story

Adventure On Deck

Play Episode Listen Later Feb 24, 2026 33:25


With only five weeks left in this year-long journey, I can feel the end approaching—less like a high-wire act and more like gathering momentum toward something unknown. Week 47 of Ted Gioia's Immersive Humanities course explores twentieth-century American fiction through short stories and novel excerpts, revealing a distinctly American voice: sharp dialogue, vivid settings, and an experimental edge.O. Henry, “The Gift of the Magi” (1906): A charming story of love and sacrifice.F. Scott Fitzgerald, “A Diamond as Big as the Ritz” (1922): Wealth, excess, and a surprising twist.Ernest Hemingway, “The Killers” (1927): Sparse, tension-filled dialogue.William Faulkner, The Sound and the Fury (1929, excerpt): Challenging, with shifting time and perspective.Ralph Ellison, Invisible Man (1947, excerpt): A powerful sense of invisibility and identity.Shirley Jackson, “The Lottery” (1948): Disturbing and unforgettable.Flannery O'Connor, “A Good Man is Hard to Find” (1955): A Southern Gothic tale with shocking turns.Together, these works feel spacious, restless, and distinctly American—and they remind me how much more willing I am now to embrace difficult, even strange, books.This is a year-long challenge! Join me next week for a little Magical Realism.LINKTed Gioia/The Honest Broker's 12-Month ImmersiveHumanities Course (paywalled!)My Amazon Book List (NOT an affiliate link)CONNECTThe complete list of Crack the Book Episodes: https://cheryldrury.substack.com/p/crack-the-book-start-here?r=u3t2rTo read more of my writing, visit my Substack - https://www.cheryldrury.substack.com.Follow me on Instagram - https://www.instagram.com/cldrury/LISTENSpotify - https://open.spotify.com/show/5GpySInw1e8IqNQvXow7Lv?si=9ebd5508daa245bdApple Podcasts - https://podcasts.apple.com/us/podcast/crack-the-book/id1749793321Captivate - https://crackthebook.captivate.fm

Latent Space: The AI Engineer Podcast — CodeGen, Agents, Computer Vision, Data Science, AI UX and all things Software 3.0

From rewriting Google's search stack in the early 2000s to reviving sparse trillion-parameter models and co-designing TPUs with frontier ML research, Jeff Dean has quietly shaped nearly every layer of the modern AI stack. As Chief AI Scientist at Google and a driving force behind Gemini, Jeff has lived through multiple scaling revolutions from CPUs and sharded indices to multimodal models that reason across text, video, and code.Jeff joins us to unpack what it really means to “own the Pareto frontier,” why distillation is the engine behind every Flash model breakthrough, how energy (in picojoules) not FLOPs is becoming the true bottleneck, what it was like leading the charge to unify all of Google's AI teams, and why the next leap won't come from bigger context windows alone, but from systems that give the illusion of attending to trillions of tokens.We discuss:* Jeff's early neural net thesis in 1990: parallel training before it was cool, why he believed scaling would win decades early, and the “bigger model, more data, better results” mantra that held for 15 years* The evolution of Google Search: sharding, moving the entire index into memory in 2001, softening query semantics pre-LLMs, and why retrieval pipelines already resemble modern LLM systems* Pareto frontier strategy: why you need both frontier “Pro” models and low-latency “Flash” models, and how distillation lets smaller models surpass prior generations* Distillation deep dive: ensembles → compression → logits as soft supervision, and why you need the biggest model to make the smallest one good* Latency as a first-class objective: why 10–50x lower latency changes UX entirely, and how future reasoning workloads will demand 10,000 tokens/sec* Energy-based thinking: picojoules per bit, why moving data costs 1000x more than a multiply, batching through the lens of energy, and speculative decoding as amortization* TPU co-design: predicting ML workloads 2–6 years out, speculative hardware features, precision reduction, sparsity, and the constant feedback loop between model architecture and silicon* Sparse models and “outrageously large” networks: trillions of parameters with 1–5% activation, and why sparsity was always the right abstraction* Unified vs. specialized models: abandoning symbolic systems, why general multimodal models tend to dominate vertical silos, and when vertical fine-tuning still makes sense* Long context and the illusion of scale: beyond needle-in-a-haystack benchmarks toward systems that narrow trillions of tokens to 117 relevant documents* Personalized AI: attending to your emails, photos, and documents (with permission), and why retrieval + reasoning will unlock deeply personal assistants* Coding agents: 50 AI interns, crisp specifications as a new core skill, and how ultra-low latency will reshape human–agent collaboration* Why ideas still matter: transformers, sparsity, RL, hardware, systems — scaling wasn't blind; the pieces had to multiply togetherShow Notes:* Gemma 3 Paper* Gemma 3* Gemini 2.5 Report* Jeff Dean's “Software Engineering Advice fromBuilding Large-Scale Distributed Systems” Presentation (with Back of the Envelope Calculations)* Latency Numbers Every Programmer Should Know by Jeff Dean* The Jeff Dean Facts* Jeff Dean Google Bio* Jeff Dean on “Important AI Trends” @Stanford AI Club* Jeff Dean & Noam Shazeer — 25 years at Google (Dwarkesh)—Jeff Dean* LinkedIn: https://www.linkedin.com/in/jeff-dean-8b212555* X: https://x.com/jeffdeanGoogle* https://google.com* https://deepmind.googleFull Video EpisodeTimestamps00:00:04 — Introduction: Alessio & Swyx welcome Jeff Dean, chief AI scientist at Google, to the Latent Space podcast00:00:30 — Owning the Pareto Frontier & balancing frontier vs low-latency models00:01:31 — Frontier models vs Flash models + role of distillation00:03:52 — History of distillation and its original motivation00:05:09 — Distillation's role in modern model scaling00:07:02 — Model hierarchy (Flash, Pro, Ultra) and distillation sources00:07:46 — Flash model economics & wide deployment00:08:10 — Latency importance for complex tasks00:09:19 — Saturation of some tasks and future frontier tasks00:11:26 — On benchmarks, public vs internal00:12:53 — Example long-context benchmarks & limitations00:15:01 — Long-context goals: attending to trillions of tokens00:16:26 — Realistic use cases beyond pure language00:18:04 — Multimodal reasoning and non-text modalities00:19:05 — Importance of vision & motion modalities00:20:11 — Video understanding example (extracting structured info)00:20:47 — Search ranking analogy for LLM retrieval00:23:08 — LLM representations vs keyword search00:24:06 — Early Google search evolution & in-memory index00:26:47 — Design principles for scalable systems00:28:55 — Real-time index updates & recrawl strategies00:30:06 — Classic “Latency numbers every programmer should know”00:32:09 — Cost of memory vs compute and energy emphasis00:34:33 — TPUs & hardware trade-offs for serving models00:35:57 — TPU design decisions & co-design with ML00:38:06 — Adapting model architecture to hardware00:39:50 — Alternatives: energy-based models, speculative decoding00:42:21 — Open research directions: complex workflows, RL00:44:56 — Non-verifiable RL domains & model evaluation00:46:13 — Transition away from symbolic systems toward unified LLMs00:47:59 — Unified models vs specialized ones00:50:38 — Knowledge vs reasoning & retrieval + reasoning00:52:24 — Vertical model specialization & modules00:55:21 — Token count considerations for vertical domains00:56:09 — Low resource languages & contextual learning00:59:22 — Origins: Dean's early neural network work01:10:07 — AI for coding & human–model interaction styles01:15:52 — Importance of crisp specification for coding agents01:19:23 — Prediction: personalized models & state retrieval01:22:36 — Token-per-second targets (10k+) and reasoning throughput01:23:20 — Episode conclusion and thanksTranscriptAlessio Fanelli [00:00:04]: Hey everyone, welcome to the Latent Space podcast. This is Alessio, founder of Kernel Labs, and I'm joined by Swyx, editor of Latent Space. Shawn Wang [00:00:11]: Hello, hello. We're here in the studio with Jeff Dean, chief AI scientist at Google. Welcome. Thanks for having me. It's a bit surreal to have you in the studio. I've watched so many of your talks, and obviously your career has been super legendary. So, I mean, congrats. I think the first thing must be said, congrats on owning the Pareto Frontier.Jeff Dean [00:00:30]: Thank you, thank you. Pareto Frontiers are good. It's good to be out there.Shawn Wang [00:00:34]: Yeah, I mean, I think it's a combination of both. You have to own the Pareto Frontier. You have to have like frontier capability, but also efficiency, and then offer that range of models that people like to use. And, you know, some part of this was started because of your hardware work. Some part of that is your model work, and I'm sure there's lots of secret sauce that you guys have worked on cumulatively. But, like, it's really impressive to see it all come together in, like, this slittily advanced.Jeff Dean [00:01:04]: Yeah, yeah. I mean, I think, as you say, it's not just one thing. It's like a whole bunch of things up and down the stack. And, you know, all of those really combine to help make UNOS able to make highly capable large models, as well as, you know, software techniques to get those large model capabilities into much smaller, lighter weight models that are, you know, much more cost effective and lower latency, but still, you know, quite capable for their size. Yeah.Alessio Fanelli [00:01:31]: How much pressure do you have on, like, having the lower bound of the Pareto Frontier, too? I think, like, the new labs are always trying to push the top performance frontier because they need to raise more money and all of that. And you guys have billions of users. And I think initially when you worked on the CPU, you were thinking about, you know, if everybody that used Google, we use the voice model for, like, three minutes a day, they were like, you need to double your CPU number. Like, what's that discussion today at Google? Like, how do you prioritize frontier versus, like, we have to do this? How do we actually need to deploy it if we build it?Jeff Dean [00:02:03]: Yeah, I mean, I think we always want to have models that are at the frontier or pushing the frontier because I think that's where you see what capabilities now exist that didn't exist at the sort of slightly less capable last year's version or last six months ago version. At the same time, you know, we know those are going to be really useful for a bunch of use cases, but they're going to be a bit slower and a bit more expensive than people might like for a bunch of other broader models. So I think what we want to do is always have kind of a highly capable sort of affordable model that enables a whole bunch of, you know, lower latency use cases. People can use them for agentic coding much more readily and then have the high-end, you know, frontier model that is really useful for, you know, deep reasoning, you know, solving really complicated math problems, those kinds of things. And it's not that. One or the other is useful. They're both useful. So I think we'd like to do both. And also, you know, through distillation, which is a key technique for making the smaller models more capable, you know, you have to have the frontier model in order to then distill it into your smaller model. So it's not like an either or choice. You sort of need that in order to actually get a highly capable, more modest size model. Yeah.Alessio Fanelli [00:03:24]: I mean, you and Jeffrey came up with the solution in 2014.Jeff Dean [00:03:28]: Don't forget, L'Oreal Vinyls as well. Yeah, yeah.Alessio Fanelli [00:03:30]: A long time ago. But like, I'm curious how you think about the cycle of these ideas, even like, you know, sparse models and, you know, how do you reevaluate them? How do you think about in the next generation of model, what is worth revisiting? Like, yeah, they're just kind of like, you know, you worked on so many ideas that end up being influential, but like in the moment, they might not feel that way necessarily. Yeah.Jeff Dean [00:03:52]: I mean, I think distillation was originally motivated because we were seeing that we had a very large image data set at the time, you know, 300 million images that we could train on. And we were seeing that if you create specialists for different subsets of those image categories, you know, this one's going to be really good at sort of mammals, and this one's going to be really good at sort of indoor room scenes or whatever, and you can cluster those categories and train on an enriched stream of data after you do pre-training on a much broader set of images. You get much better performance. If you then treat that whole set of maybe 50 models you've trained as a large ensemble, but that's not a very practical thing to serve, right? So distillation really came about from the idea of, okay, what if we want to actually serve that and train all these independent sort of expert models and then squish it into something that actually fits in a form factor that you can actually serve? And that's, you know, not that different from what we're doing today. You know, often today we're instead of having an ensemble of 50 models. We're having a much larger scale model that we then distill into a much smaller scale model.Shawn Wang [00:05:09]: Yeah. A part of me also wonders if distillation also has a story with the RL revolution. So let me maybe try to articulate what I mean by that, which is you can, RL basically spikes models in a certain part of the distribution. And then you have to sort of, well, you can spike models, but usually sometimes... It might be lossy in other areas and it's kind of like an uneven technique, but you can probably distill it back and you can, I think that the sort of general dream is to be able to advance capabilities without regressing on anything else. And I think like that, that whole capability merging without loss, I feel like it's like, you know, some part of that should be a distillation process, but I can't quite articulate it. I haven't seen much papers about it.Jeff Dean [00:06:01]: Yeah, I mean, I tend to think of one of the key advantages of distillation is that you can have a much smaller model and you can have a very large, you know, training data set and you can get utility out of making many passes over that data set because you're now getting the logits from the much larger model in order to sort of coax the right behavior out of the smaller model that you wouldn't otherwise get with just the hard labels. And so, you know, I think that's what we've observed. Is you can get, you know, very close to your largest model performance with distillation approaches. And that seems to be, you know, a nice sweet spot for a lot of people because it enables us to kind of, for multiple Gemini generations now, we've been able to make the sort of flash version of the next generation as good or even substantially better than the previous generations pro. And I think we're going to keep trying to do that because that seems like a good trend to follow.Shawn Wang [00:07:02]: So, Dara asked, so it was the original map was Flash Pro and Ultra. Are you just sitting on Ultra and distilling from that? Is that like the mother load?Jeff Dean [00:07:12]: I mean, we have a lot of different kinds of models. Some are internal ones that are not necessarily meant to be released or served. Some are, you know, our pro scale model and we can distill from that as well into our Flash scale model. So I think, you know, it's an important set of capabilities to have and also inference time scaling. It can also be a useful thing to improve the capabilities of the model.Shawn Wang [00:07:35]: And yeah, yeah, cool. Yeah. And obviously, I think the economy of Flash is what led to the total dominance. I think the latest number is like 50 trillion tokens. I don't know. I mean, obviously, it's changing every day.Jeff Dean [00:07:46]: Yeah, yeah. But, you know, by market share, hopefully up.Shawn Wang [00:07:50]: No, I mean, there's no I mean, there's just the economics wise, like because Flash is so economical, like you can use it for everything. Like it's in Gmail now. It's in YouTube. Like it's yeah. It's in everything.Jeff Dean [00:08:02]: We're using it more in our search products of various AI mode reviews.Shawn Wang [00:08:05]: Oh, my God. Flash past the AI mode. Oh, my God. Yeah, that's yeah, I didn't even think about that.Jeff Dean [00:08:10]: I mean, I think one of the things that is quite nice about the Flash model is not only is it more affordable, it's also a lower latency. And I think latency is actually a pretty important characteristic for these models because we're going to want models to do much more complicated things that are going to involve, you know, generating many more tokens from when you ask the model to do so. So, you know, if you're going to ask the model to do something until it actually finishes what you ask it to do, because you're going to ask now, not just write me a for loop, but like write me a whole software package to do X or Y or Z. And so having low latency systems that can do that seems really important. And Flash is one direction, one way of doing that. You know, obviously our hardware platforms enable a bunch of interesting aspects of our, you know, serving stack as well, like TPUs, the interconnect between. Chips on the TPUs is actually quite, quite high performance and quite amenable to, for example, long context kind of attention operations, you know, having sparse models with lots of experts. These kinds of things really, really matter a lot in terms of how do you make them servable at scale.Alessio Fanelli [00:09:19]: Yeah. Does it feel like there's some breaking point for like the proto Flash distillation, kind of like one generation delayed? I almost think about almost like the capability as a. In certain tasks, like the pro model today is a saturated, some sort of task. So next generation, that same task will be saturated at the Flash price point. And I think for most of the things that people use models for at some point, the Flash model in two generation will be able to do basically everything. And how do you make it economical to like keep pushing the pro frontier when a lot of the population will be okay with the Flash model? I'm curious how you think about that.Jeff Dean [00:09:59]: I mean, I think that's true. If your distribution of what people are asking people, the models to do is stationary, right? But I think what often happens is as the models become more capable, people ask them to do more, right? So, I mean, I think this happens in my own usage. Like I used to try our models a year ago for some sort of coding task, and it was okay at some simpler things, but wouldn't do work very well for more complicated things. And since then, we've improved dramatically on the more complicated coding tasks. And now I'll ask it to do much more complicated things. And I think that's true, not just of coding, but of, you know, now, you know, can you analyze all the, you know, renewable energy deployments in the world and give me a report on solar panel deployment or whatever. That's a very complicated, you know, more complicated task than people would have asked a year ago. And so you are going to want more capable models to push the frontier in the absence of what people ask the models to do. And that also then gives us. Insight into, okay, where does the, where do things break down? How can we improve the model in these, these particular areas, uh, in order to sort of, um, make the next generation even better.Alessio Fanelli [00:11:11]: Yeah. Are there any benchmarks or like test sets they use internally? Because it's almost like the same benchmarks get reported every time. And it's like, all right, it's like 99 instead of 97. Like, how do you have to keep pushing the team internally to it? Or like, this is what we're building towards. Yeah.Jeff Dean [00:11:26]: I mean, I think. Benchmarks, particularly external ones that are publicly available. Have their utility, but they often kind of have a lifespan of utility where they're introduced and maybe they're quite hard for current models. You know, I, I like to think of the best kinds of benchmarks are ones where the initial scores are like 10 to 20 or 30%, maybe, but not higher. And then you can sort of work on improving that capability for, uh, whatever it is, the benchmark is trying to assess and get it up to like 80, 90%, whatever. I, I think once it hits kind of 95% or something, you get very diminishing returns from really focusing on that benchmark, cuz it's sort of, it's either the case that you've now achieved that capability, or there's also the issue of leakage in public data or very related kind of data being, being in your training data. Um, so we have a bunch of held out internal benchmarks that we really look at where we know that wasn't represented in the training data at all. There are capabilities that we want the model to have. Um, yeah. Yeah. Um, that it doesn't have now, and then we can work on, you know, assessing, you know, how do we make the model better at these kinds of things? Is it, we need different kind of data to train on that's more specialized for this particular kind of task. Do we need, um, you know, a bunch of, uh, you know, architectural improvements or some sort of, uh, model capability improvements, you know, what would help make that better?Shawn Wang [00:12:53]: Is there, is there such an example that you, uh, a benchmark inspired in architectural improvement? Like, uh, I'm just kind of. Jumping on that because you just.Jeff Dean [00:13:02]: Uh, I mean, I think some of the long context capability of the, of the Gemini models that came, I guess, first in 1.5 really were about looking at, okay, we want to have, um, you know,Shawn Wang [00:13:15]: immediately everyone jumped to like completely green charts of like, everyone had, I was like, how did everyone crack this at the same time? Right. Yeah. Yeah.Jeff Dean [00:13:23]: I mean, I think, um, and once you're set, I mean, as you say that needed single needle and a half. Hey, stack benchmark is really saturated for at least context links up to 1, 2 and K or something. Don't actually have, you know, much larger than 1, 2 and 8 K these days or two or something. We're trying to push the frontier of 1 million or 2 million context, which is good because I think there are a lot of use cases where. Yeah. You know, putting a thousand pages of text or putting, you know, multiple hour long videos and the context and then actually being able to make use of that as useful. Try to, to explore the über graduation are fairly large. But the single needle in a haystack benchmark is sort of saturated. So you really want more complicated, sort of multi-needle or more realistic, take all this content and produce this kind of answer from a long context that sort of better assesses what it is people really want to do with long context. Which is not just, you know, can you tell me the product number for this particular thing?Shawn Wang [00:14:31]: Yeah, it's retrieval. It's retrieval within machine learning. It's interesting because I think the more meta level I'm trying to operate at here is you have a benchmark. You're like, okay, I see the architectural thing I need to do in order to go fix that. But should you do it? Because sometimes that's an inductive bias, basically. It's what Jason Wei, who used to work at Google, would say. Exactly the kind of thing. Yeah, you're going to win. Short term. Longer term, I don't know if that's going to scale. You might have to undo that.Jeff Dean [00:15:01]: I mean, I like to sort of not focus on exactly what solution we're going to derive, but what capability would you want? And I think we're very convinced that, you know, long context is useful, but it's way too short today. Right? Like, I think what you would really want is, can I attend to the internet while I answer my question? Right? But that's not going to happen. I think that's going to be solved by purely scaling the existing solutions, which are quadratic. So a million tokens kind of pushes what you can do. You're not going to do that to a trillion tokens, let alone, you know, a billion tokens, let alone a trillion. But I think if you could give the illusion that you can attend to trillions of tokens, that would be amazing. You'd find all kinds of uses for that. You would have attend to the internet. You could attend to the pixels of YouTube and the sort of deeper representations that we can find. You could attend to the form for a single video, but across many videos, you know, on a personal Gemini level, you could attend to all of your personal state with your permission. So like your emails, your photos, your docs, your plane tickets you have. I think that would be really, really useful. And the question is, how do you get algorithmic improvements and system level improvements that get you to something where you actually can attend to trillions of tokens? Right. In a meaningful way. Yeah.Shawn Wang [00:16:26]: But by the way, I think I did some math and it's like, if you spoke all day, every day for eight hours a day, you only generate a maximum of like a hundred K tokens, which like very comfortably fits.Jeff Dean [00:16:38]: Right. But if you then say, okay, I want to be able to understand everything people are putting on videos.Shawn Wang [00:16:46]: Well, also, I think that the classic example is you start going beyond language into like proteins and whatever else is extremely information dense. Yeah. Yeah.Jeff Dean [00:16:55]: I mean, I think one of the things about Gemini's multimodal aspects is we've always wanted it to be multimodal from the start. And so, you know, that sometimes to people means text and images and video sort of human-like and audio, audio, human-like modalities. But I think it's also really useful to have Gemini know about non-human modalities. Yeah. Like LIDAR sensor data from. Yes. Say, Waymo vehicles or. Like robots or, you know, various kinds of health modalities, x-rays and MRIs and imaging and genomics information. And I think there's probably hundreds of modalities of data where you'd like the model to be able to at least be exposed to the fact that this is an interesting modality and has certain meaning in the world. Where even if you haven't trained on all the LIDAR data or MRI data, you could have, because maybe that's not, you know, it doesn't make sense in terms of trade-offs of. You know, what you include in your main pre-training data mix, at least including a little bit of it is actually quite useful. Yeah. Because it sort of tempts the model that this is a thing.Shawn Wang [00:18:04]: Yeah. Do you believe, I mean, since we're on this topic and something I just get to ask you all the questions I always wanted to ask, which is fantastic. Like, are there some king modalities, like modalities that supersede all the other modalities? So a simple example was Vision can, on a pixel level, encode text. And DeepSeq had this DeepSeq CR paper that did that. Vision. And Vision has also been shown to maybe incorporate audio because you can do audio spectrograms and that's, that's also like a Vision capable thing. Like, so, so maybe Vision is just the king modality and like. Yeah.Jeff Dean [00:18:36]: I mean, Vision and Motion are quite important things, right? Motion. Well, like video as opposed to static images, because I mean, there's a reason evolution has evolved eyes like 23 independent ways, because it's such a useful capability for sensing the world around you, which is really what we want these models to be. So I think the only thing that we can be able to do is interpret the things we're seeing or the things we're paying attention to and then help us in using that information to do things. Yeah.Shawn Wang [00:19:05]: I think motion, you know, I still want to shout out, I think Gemini, still the only native video understanding model that's out there. So I use it for YouTube all the time. Nice.Jeff Dean [00:19:15]: Yeah. Yeah. I mean, it's actually, I think people kind of are not necessarily aware of what the Gemini models can actually do. Yeah. Like I have an example I've used in one of my talks. It had like, it was like a YouTube highlight video of 18 memorable sports moments across the last 20 years or something. So it has like Michael Jordan hitting some jump shot at the end of the finals and, you know, some soccer goals and things like that. And you can literally just give it the video and say, can you please make me a table of what all these different events are? What when the date is when they happened? And a short description. And so you get like now an 18 row table of that information extracted from the video, which is, you know, not something most people think of as like a turn video into sequel like table.Alessio Fanelli [00:20:11]: Has there been any discussion inside of Google of like, you mentioned tending to the whole internet, right? Google, it's almost built because a human cannot tend to the whole internet and you need some sort of ranking to find what you need. Yep. That ranking is like much different for an LLM because you can expect a person to look at maybe the first five, six links in a Google search versus for an LLM. Should you expect to have 20 links that are highly relevant? Like how do you internally figure out, you know, how do we build the AI mode that is like maybe like much broader search and span versus like the more human one? Yeah.Jeff Dean [00:20:47]: I mean, I think even pre-language model based work, you know, our ranking systems would be built to start. I mean, I think even pre-language model based work, you know, our ranking systems would be built to start. With a giant number of web pages in our index, many of them are not relevant. So you identify a subset of them that are relevant with very lightweight kinds of methods. You know, you're down to like 30,000 documents or something. And then you gradually refine that to apply more and more sophisticated algorithms and more and more sophisticated sort of signals of various kinds in order to get down to ultimately what you show, which is, you know, the final 10 results or, you know, 10 results plus. Other kinds of information. And I think an LLM based system is not going to be that dissimilar, right? You're going to attend to trillions of tokens, but you're going to want to identify, you know, what are the 30,000 ish documents that are with the, you know, maybe 30 million interesting tokens. And then how do you go from that into what are the 117 documents I really should be paying attention to in order to carry out the tasks that the user has asked? And I think, you know, you can imagine systems where you have, you know, a lot of highly parallel processing to identify those initial 30,000 candidates, maybe with very lightweight kinds of models. Then you have some system that sort of helps you narrow down from 30,000 to the 117 with maybe a little bit more sophisticated model or set of models. And then maybe the final model is the thing that looks. So the 117 things that might be your most capable model. So I think it has to, it's going to be some system like that, that is really enables you to give the illusion of attending to trillions of tokens. Sort of the way Google search gives you, you know, not the illusion, but you are searching the internet, but you're finding, you know, a very small subset of things that are, that are relevant.Shawn Wang [00:22:47]: Yeah. I often tell a lot of people that are not steeped in like Google search history that, well, you know, like Bert was. Like he was like basically immediately inside of Google search and that improves results a lot, right? Like I don't, I don't have any numbers off the top of my head, but like, I'm sure you guys, that's obviously the most important numbers to Google. Yeah.Jeff Dean [00:23:08]: I mean, I think going to an LLM based representation of text and words and so on enables you to get out of the explicit hard notion of, of particular words having to be on the page, but really getting at the notion of this topic of this page or this page. Paragraph is highly relevant to this query. Yeah.Shawn Wang [00:23:28]: I don't think people understand how much LLMs have taken over all these very high traffic system, very high traffic. Yeah. Like it's Google, it's YouTube. YouTube has this like semantics ID thing where it's just like every token or every item in the vocab is a YouTube video or something that predicts the video using a code book, which is absurd to me for YouTube size.Jeff Dean [00:23:50]: And then most recently GROK also for, for XAI, which is like, yeah. I mean, I'll call out even before LLMs were used extensively in search, we put a lot of emphasis on softening the notion of what the user actually entered into the query.Shawn Wang [00:24:06]: So do you have like a history of like, what's the progression? Oh yeah.Jeff Dean [00:24:09]: I mean, I actually gave a talk in, uh, I guess, uh, web search and data mining conference in 2009, uh, where we never actually published any papers about the origins of Google search, uh, sort of, but we went through sort of four or five or six. generations, four or five or six generations of, uh, redesigning of the search and retrieval system, uh, from about 1999 through 2004 or five. And that talk is really about that evolution. And one of the things that really happened in 2001 was we were sort of working to scale the system in multiple dimensions. So one is we wanted to make our index bigger, so we could retrieve from a larger index, which always helps your quality in general. Uh, because if you don't have the page in your index, you're going to not do well. Um, and then we also needed to scale our capacity because we were, our traffic was growing quite extensively. Um, and so we had, you know, a sharded system where you have more and more shards as the index grows, you have like 30 shards. And then if you want to double the index size, you make 60 shards so that you can bound the latency by which you respond for any particular user query. Um, and then as traffic grows, you add, you add more and more replicas of each of those. And so we eventually did the math that realized that in a data center where we had say 60 shards and, um, you know, 20 copies of each shard, we now had 1200 machines, uh, with disks. And we did the math and we're like, Hey, one copy of that index would actually fit in memory across 1200 machines. So in 2001, we introduced, uh, we put our entire index in memory and what that enabled from a quality perspective was amazing. Um, and so we had more and more replicas of each of those. Before you had to be really careful about, you know, how many different terms you looked at for a query, because every one of them would involve a disk seek on every one of the 60 shards. And so you, as you make your index bigger, that becomes even more inefficient. But once you have the whole index in memory, it's totally fine to have 50 terms you throw into the query from the user's original three or four word query, because now you can add synonyms like restaurant and restaurants and cafe and, uh, you know, things like that. Uh, bistro and all these things. And you can suddenly start, uh, sort of really, uh, getting at the meaning of the word as opposed to the exact semantic form the user typed in. And that was, you know, 2001, very much pre LLM, but really it was about softening the, the strict definition of what the user typed in order to get at the meaning.Alessio Fanelli [00:26:47]: What are like principles that you use to like design the systems, especially when you have, I mean, in 2001, the internet is like. Doubling, tripling every year in size is not like, uh, you know, and I think today you kind of see that with LLMs too, where like every year the jumps in size and like capabilities are just so big. Are there just any, you know, principles that you use to like, think about this? Yeah.Jeff Dean [00:27:08]: I mean, I think, uh, you know, first, whenever you're designing a system, you want to understand what are the sort of design parameters that are going to be most important in designing that, you know? So, you know, how many queries per second do you need to handle? How big is the internet? How big is the index you need to handle? How much data do you need to keep for every document in the index? How are you going to look at it when you retrieve things? Um, what happens if traffic were to double or triple, you know, will that system work well? And I think a good design principle is you're going to want to design a system so that the most important characteristics could scale by like factors of five or 10, but probably not beyond that because often what happens is if you design a system for X. And something suddenly becomes a hundred X, that would enable a very different point in the design space that would not make sense at X. But all of a sudden at a hundred X makes total sense. So like going from a disk space index to a in memory index makes a lot of sense once you have enough traffic, because now you have enough replicas of the sort of state on disk that those machines now actually can hold, uh, you know, a full copy of the, uh, index and memory. Yeah. And that all of a sudden enabled. A completely different design that wouldn't have been practical before. Yeah. Um, so I'm, I'm a big fan of thinking through designs in your head, just kind of playing with the design space a little before you actually do a lot of writing of code. But, you know, as you said, in the early days of Google, we were growing the index, uh, quite extensively. We were growing the update rate of the index. So the update rate actually is the parameter that changed the most. Surprising. So it used to be once a month.Shawn Wang [00:28:55]: Yeah.Jeff Dean [00:28:56]: And then we went to a system that could update any particular page in like sub one minute. Okay.Shawn Wang [00:29:02]: Yeah. Because this is a competitive advantage, right?Jeff Dean [00:29:04]: Because all of a sudden news related queries, you know, if you're, if you've got last month's news index, it's not actually that useful for.Shawn Wang [00:29:11]: News is a special beast. Was there any, like you could have split it onto a separate system.Jeff Dean [00:29:15]: Well, we did. We launched a Google news product, but you also want news related queries that people type into the main index to also be sort of updated.Shawn Wang [00:29:23]: So, yeah, it's interesting. And then you have to like classify whether the page is, you have to decide which pages should be updated and what frequency. Oh yeah.Jeff Dean [00:29:30]: There's a whole like, uh, system behind the scenes that's trying to decide update rates and importance of the pages. So even if the update rate seems low, you might still want to recrawl important pages quite often because, uh, the likelihood they change might be low, but the value of having updated is high.Shawn Wang [00:29:50]: Yeah, yeah, yeah, yeah. Uh, well, you know, yeah. This, uh, you know, mention of latency and, and saving things to this reminds me of one of your classics, which I have to bring up, which is latency numbers. Every programmer should know, uh, was there a, was it just a, just a general story behind that? Did you like just write it down?Jeff Dean [00:30:06]: I mean, this has like sort of eight or 10 different kinds of metrics that are like, how long does a cache mistake? How long does branch mispredict take? How long does a reference domain memory take? How long does it take to send, you know, a packet from the U S to the Netherlands or something? Um,Shawn Wang [00:30:21]: why Netherlands, by the way, or is it, is that because of Chrome?Jeff Dean [00:30:25]: Uh, we had a data center in the Netherlands, um, so, I mean, I think this gets to the point of being able to do the back of the envelope calculations. So these are sort of the raw ingredients of those, and you can use them to say, okay, well, if I need to design a system to do image search and thumb nailing or something of the result page, you know, how, what I do that I could pre-compute the image thumbnails. I could like. Try to thumbnail them on the fly from the larger images. What would that do? How much dis bandwidth than I need? How many des seeks would I do? Um, and you can sort of actually do thought experiments in, you know, 30 seconds or a minute with the sort of, uh, basic, uh, basic numbers at your fingertips. Uh, and then as you sort of build software using higher level libraries, you kind of want to develop the same intuitions for how long does it take to, you know, look up something in this particular kind of.Shawn Wang [00:31:21]: I'll see you next time.Shawn Wang [00:31:51]: Which is a simple byte conversion. That's nothing interesting. I wonder if you have any, if you were to update your...Jeff Dean [00:31:58]: I mean, I think it's really good to think about calculations you're doing in a model, either for training or inference.Jeff Dean [00:32:09]: Often a good way to view that is how much state will you need to bring in from memory, either like on-chip SRAM or HBM from the accelerator. Attached memory or DRAM or over the network. And then how expensive is that data motion relative to the cost of, say, an actual multiply in the matrix multiply unit? And that cost is actually really, really low, right? Because it's order, depending on your precision, I think it's like sub one picodule.Shawn Wang [00:32:50]: Oh, okay. You measure it by energy. Yeah. Yeah.Jeff Dean [00:32:52]: Yeah. I mean, it's all going to be about energy and how do you make the most energy efficient system. And then moving data from the SRAM on the other side of the chip, not even off the off chip, but on the other side of the same chip can be, you know, a thousand picodules. Oh, yeah. And so all of a sudden, this is why your accelerators require batching. Because if you move, like, say, the parameter of a model from SRAM on the, on the chip into the multiplier unit, that's going to cost you a thousand picodules. So you better make use of that, that thing that you moved many, many times with. So that's where the batch dimension comes in. Because all of a sudden, you know, if you have a batch of 256 or something, that's not so bad. But if you have a batch of one, that's really not good.Shawn Wang [00:33:40]: Yeah. Yeah. Right.Jeff Dean [00:33:41]: Because then you paid a thousand picodules in order to do your one picodule multiply.Shawn Wang [00:33:46]: I have never heard an energy-based analysis of batching.Jeff Dean [00:33:50]: Yeah. I mean, that's why people batch. Yeah. Ideally, you'd like to use batch size one because the latency would be great.Shawn Wang [00:33:56]: The best latency.Jeff Dean [00:33:56]: But the energy cost and the compute cost inefficiency that you get is quite large. So, yeah.Shawn Wang [00:34:04]: Is there a similar trick like, like, like you did with, you know, putting everything in memory? Like, you know, I think obviously NVIDIA has caused a lot of waves with betting very hard on SRAM with Grok. I wonder if, like, that's something that you already saw with, with the TPUs, right? Like that, that you had to. Uh, to serve at your scale, uh, you probably sort of saw that coming. Like what, what, what hardware, uh, innovations or insights were formed because of what you're seeing there?Jeff Dean [00:34:33]: Yeah. I mean, I think, you know, TPUs have this nice, uh, sort of regular structure of 2D or 3D meshes with a bunch of chips connected. Yeah. And each one of those has HBM attached. Um, I think for serving some kinds of models, uh, you know, you, you pay a lot higher cost. Uh, and time latency, um, bringing things in from HBM than you do bringing them in from, uh, SRAM on the chip. So if you have a small enough model, you can actually do model parallelism, spread it out over lots of chips and you actually get quite good throughput improvements and latency improvements from doing that. And so you're now sort of striping your smallish scale model over say 16 or 64 chips. Uh, but as if you do that and it all fits in. In SRAM, uh, that can be a big win. So yeah, that's not a surprise, but it is a good technique.Alessio Fanelli [00:35:27]: Yeah. What about the TPU design? Like how much do you decide where the improvements have to go? So like, this is like a good example of like, is there a way to bring the thousand picojoules down to 50? Like, is it worth designing a new chip to do that? The extreme is like when people say, oh, you should burn the model on the ASIC and that's kind of like the most extreme thing. How much of it? Is it worth doing an hardware when things change so quickly? Like what was the internal discussion? Yeah.Jeff Dean [00:35:57]: I mean, we, we have a lot of interaction between say the TPU chip design architecture team and the sort of higher level modeling, uh, experts, because you really want to take advantage of being able to co-design what should future TPUs look like based on where we think the sort of ML research puck is going, uh, in some sense, because, uh, you know, as a hardware designer for ML and in particular, you're trying to design a chip starting today and that design might take two years before it even lands in a data center. And then it has to sort of be a reasonable lifetime of the chip to take you three, four or five years. So you're trying to predict two to six years out where, what ML computations will people want to run two to six years out in a very fast changing field. And so having people with interest. Interesting ML research ideas of things we think will start to work in that timeframe or will be more important in that timeframe, uh, really enables us to then get, you know, interesting hardware features put into, you know, TPU N plus two, where TPU N is what we have today.Shawn Wang [00:37:10]: Oh, the cycle time is plus two.Jeff Dean [00:37:12]: Roughly. Wow. Because, uh, I mean, sometimes you can squeeze some changes into N plus one, but, you know, bigger changes are going to require the chip. Yeah. Design be earlier in its lifetime design process. Um, so whenever we can do that, it's generally good. And sometimes you can put in speculative features that maybe won't cost you much chip area, but if it works out, it would make something, you know, 10 times as fast. And if it doesn't work out, well, you burned a little bit of tiny amount of your chip area on that thing, but it's not that big a deal. Uh, sometimes it's a very big change and we want to be pretty sure this is going to work out. So we'll do like lots of carefulness. Uh, ML experimentation to show us, uh, this is actually the, the way we want to go. Yeah.Alessio Fanelli [00:37:58]: Is there a reverse of like, we already committed to this chip design so we can not take the model architecture that way because it doesn't quite fit?Jeff Dean [00:38:06]: Yeah. I mean, you, you definitely have things where you're going to adapt what the model architecture looks like so that they're efficient on the chips that you're going to have for both training and inference of that, of that, uh, generation of model. So I think it kind of goes both ways. Um, you know, sometimes you can take advantage of, you know, lower precision things that are coming in a future generation. So you can, might train it at that lower precision, even if the current generation doesn't quite do that. Mm.Shawn Wang [00:38:40]: Yeah. How low can we go in precision?Jeff Dean [00:38:43]: Because people are saying like ternary is like, uh, yeah, I mean, I'm a big fan of very low precision because I think that gets, that saves you a tremendous amount of time. Right. Because it's picojoules per bit that you're transferring and reducing the number of bits is a really good way to, to reduce that. Um, you know, I think people have gotten a lot of luck, uh, mileage out of having very low bit precision things, but then having scaling factors that apply to a whole bunch of, uh, those, those weights. Scaling. How does it, how does it, okay.Shawn Wang [00:39:15]: Interesting. You, so low, low precision, but scaled up weights. Yeah. Huh. Yeah. Never considered that. Yeah. Interesting. Uh, w w while we're on this topic, you know, I think there's a lot of, um, uh, this, the concept of precision at all is weird when we're sampling, you know, uh, we just, at the end of this, we're going to have all these like chips that I'll do like very good math. And then we're just going to throw a random number generator at the start. So, I mean, there's a movement towards, uh, energy based, uh, models and processors. I'm just curious if you've, obviously you've thought about it, but like, what's your commentary?Jeff Dean [00:39:50]: Yeah. I mean, I think. There's a bunch of interesting trends though. Energy based models is one, you know, diffusion based models, which don't sort of sequentially decode tokens is another, um, you know, speculative decoding is a way that you can get sort of an equivalent, very small.Shawn Wang [00:40:06]: Draft.Jeff Dean [00:40:07]: Batch factor, uh, for like you predict eight tokens out and that enables you to sort of increase the effective batch size of what you're doing by a factor of eight, even, and then you maybe accept five or six of those tokens. So you get. A five, a five X improvement in the amortization of moving weights, uh, into the multipliers to do the prediction for the, the tokens. So these are all really good techniques and I think it's really good to look at them from the lens of, uh, energy, real energy, not energy based models, um, and, and also latency and throughput, right? If you look at things from that lens, that sort of guides you to. Two solutions that are gonna be, uh, you know, better from, uh, you know, being able to serve larger models or, you know, equivalent size models more cheaply and with lower latency.Shawn Wang [00:41:03]: Yeah. Well, I think, I think I, um, it's appealing intellectually, uh, haven't seen it like really hit the mainstream, but, um, I do think that, uh, there's some poetry in the sense that, uh, you know, we don't have to do, uh, a lot of shenanigans if like we fundamentally. Design it into the hardware. Yeah, yeah.Jeff Dean [00:41:23]: I mean, I think there's still a, there's also sort of the more exotic things like analog based, uh, uh, computing substrates as opposed to digital ones. Uh, I'm, you know, I think those are super interesting cause they can be potentially low power. Uh, but I think you often end up wanting to interface that with digital systems and you end up losing a lot of the power advantages in the digital to analog and analog to digital conversions. You end up doing, uh, at the sort of boundaries. And periphery of that system. Um, I still think there's a tremendous distance we can go from where we are today in terms of energy efficiency with sort of, uh, much better and specialized hardware for the models we care about.Shawn Wang [00:42:05]: Yeah.Alessio Fanelli [00:42:06]: Um, any other interesting research ideas that you've seen, or like maybe things that you cannot pursue a Google that you would be interested in seeing researchers take a step at, I guess you have a lot of researchers. Yeah, I guess you have enough, but our, our research.Jeff Dean [00:42:21]: Our research portfolio is pretty broad. I would say, um, I mean, I think, uh, in terms of research directions, there's a whole bunch of, uh, you know, open problems and how do you make these models reliable and able to do much longer, kind of, uh, more complex tasks that have lots of subtasks. How do you orchestrate, you know, maybe one model that's using other models as tools in order to sort of build, uh, things that can accomplish, uh, you know, much more. Yeah. Significant pieces of work, uh, collectively, then you would ask a single model to do. Um, so that's super interesting. How do you get more verifiable, uh, you know, how do you get RL to work for non-verifiable domains? I think it's a pretty interesting open problem because I think that would broaden out the capabilities of the models, the improvements that you're seeing in both math and coding. Uh, if we could apply those to other less verifiable domains, because we've come up with RL techniques that actually enable us to do that. Uh, effectively, that would, that would really make the models improve quite a lot. I think.Alessio Fanelli [00:43:26]: I'm curious, like when we had Noam Brown on the podcast, he said, um, they already proved you can do it with deep research. Um, you kind of have it with AI mode in a way it's not verifiable. I'm curious if there's any thread that you think is interesting there. Like what is it? Both are like information retrieval of JSON. So I wonder if it's like the retrieval is like the verifiable part. That you can score or what are like, yeah, yeah. How, how would you model that, that problem?Jeff Dean [00:43:55]: Yeah. I mean, I think there are ways of having other models that can evaluate the results of what a first model did, maybe even retrieving. Can you have another model that says, is this things, are these things you retrieved relevant? Or can you rate these 2000 things you retrieved to assess which ones are the 50 most relevant or something? Um, I think those kinds of techniques are actually quite effective. Sometimes I can even be the same model, just prompted differently to be a, you know, a critic as opposed to a, uh, actual retrieval system. Yeah.Shawn Wang [00:44:28]: Um, I do think like there, there is that, that weird cliff where like, it feels like we've done the easy stuff and then now it's, but it always feels like that every year. It's like, oh, like we know, we know, and the next part is super hard and nobody's figured it out. And, uh, exactly with this RLVR thing where like everyone's talking about, well, okay, how do we. the next stage of the non-verifiable stuff. And everyone's like, I don't know, you know, Ellen judge.Jeff Dean [00:44:56]: I mean, I feel like the nice thing about this field is there's lots and lots of smart people thinking about creative solutions to some of the problems that we all see. Uh, because I think everyone sort of sees that the models, you know, are great at some things and they fall down around the edges of those things and, and are not as capable as we'd like in those areas. And then coming up with good techniques and trying those. And seeing which ones actually make a difference is sort of what the whole research aspect of this field is, is pushing forward. And I think that's why it's super interesting. You know, if you think about two years ago, we were struggling with GSM, eight K problems, right? Like, you know, Fred has two rabbits. He gets three more rabbits. How many rabbits does he have? That's a pretty far cry from the kinds of mathematics that the models can, and now you're doing IMO and Erdos problems in pure language. Yeah. Yeah. Pure language. So that is a really, really amazing jump in capabilities in, you know, in a year and a half or something. And I think, um, for other areas, it'd be great if we could make that kind of leap. Uh, and you know, we don't exactly see how to do it for some, some areas, but we do see it for some other areas and we're going to work hard on making that better. Yeah.Shawn Wang [00:46:13]: Yeah.Alessio Fanelli [00:46:14]: Like YouTube thumbnail generation. That would be very helpful. We need that. That would be AGI. We need that.Shawn Wang [00:46:20]: That would be. As far as content creators go.Jeff Dean [00:46:22]: I guess I'm not a YouTube creator, so I don't care that much about that problem, but I guess, uh, many people do.Shawn Wang [00:46:27]: It does. Yeah. It doesn't, it doesn't matter. People do judge books by their covers as it turns out. Um, uh, just to draw a bit on the IMO goal. Um, I'm still not over the fact that a year ago we had alpha proof and alpha geometry and all those things. And then this year we were like, screw that we'll just chuck it into Gemini. Yeah. What's your reflection? Like, I think this, this question about. Like the merger of like symbolic systems and like, and, and LMS, uh, was a very much core belief. And then somewhere along the line, people would just said, Nope, we'll just all do it in the LLM.Jeff Dean [00:47:02]: Yeah. I mean, I think it makes a lot of sense to me because, you know, humans manipulate symbols, but we probably don't have like a symbolic representation in our heads. Right. We have some distributed representation that is neural net, like in some way of lots of different neurons. And activation patterns firing when we see certain things and that enables us to reason and plan and, you know, do chains of thought and, you know, roll them back now that, that approach for solving the problem doesn't seem like it's going to work. I'm going to try this one. And, you know, in a lot of ways we're emulating what we intuitively think, uh, is happening inside real brains in neural net based models. So it never made sense to me to have like completely separate. Uh, discrete, uh, symbolic things, and then a completely different way of, of, uh, you know, thinking about those things.Shawn Wang [00:47:59]: Interesting. Yeah. Uh, I mean, it's maybe seems obvious to you, but it wasn't obvious to me a year ago. Yeah.Jeff Dean [00:48:06]: I mean, I do think like that IMO with, you know, translating to lean and using lean and then the next year and also a specialized geometry model. And then this year switching to a single unified model. That is roughly the production model with a little bit more inference budget, uh, is actually, you know, quite good because it shows you that the capabilities of that general model have improved dramatically and, and now you don't need the specialized model. This is actually sort of very similar to the 2013 to 16 era of machine learning, right? Like it used to be, people would train separate models for lots of different, each different problem, right? I have, I want to recognize street signs and something. So I train a street sign. Recognition recognition model, or I want to, you know, decode speech recognition. I have a speech model, right? I think now the era of unified models that do everything is really upon us. And the question is how well do those models generalize to new things they've never been asked to do and they're getting better and better.Shawn Wang [00:49:10]: And you don't need domain experts. Like one of my, uh, so I interviewed ETA who was on, who was on that team. Uh, and he was like, yeah, I, I don't know how they work. I don't know where the IMO competition was held. I don't know the rules of it. I just trained the models, the training models. Yeah. Yeah. And it's kind of interesting that like people with these, this like universal skill set of just like machine learning, you just give them data and give them enough compute and they can kind of tackle any task, which is the bitter lesson, I guess. I don't know. Yeah.Jeff Dean [00:49:39]: I mean, I think, uh, general models, uh, will win out over specialized ones in most cases.Shawn Wang [00:49:45]: Uh, so I want to push there a bit. I think there's one hole here, which is like, uh. There's this concept of like, uh, maybe capacity of a model, like abstractly a model can only contain the number of bits that it has. And, uh, and so it, you know, God knows like Gemini pro is like one to 10 trillion parameters. We don't know, but, uh, the Gemma models, for example, right? Like a lot of people want like the open source local models that are like that, that, that, and, and, uh, they have some knowledge, which is not necessary, right? Like they can't know everything like, like you have the. The luxury of you have the big model and big model should be able to capable of everything. But like when, when you're distilling and you're going down to the small models, you know, you're actually memorizing things that are not useful. Yeah. And so like, how do we, I guess, do we want to extract that? Can we, can we divorce knowledge from reasoning, you know?Jeff Dean [00:50:38]: Yeah. I mean, I think you do want the model to be most effective at reasoning if it can retrieve things, right? Because having the model devote precious parameter space. To remembering obscure facts that could be looked up is actually not the best use of that parameter space, right? Like you might prefer something that is more generally useful in more settings than this obscure fact that it has. Um, so I think that's always attention at the same time. You also don't want your model to be kind of completely detached from, you know, knowing stuff about the world, right? Like it's probably useful to know how long the golden gate be. Bridges just as a general sense of like how long are bridges, right? And, uh, it should have that kind of knowledge. It maybe doesn't need to know how long some teeny little bridge in some other more obscure part of the world is, but, uh, it does help it to have a fair bit of world knowledge and the bigger your model is, the more you can have. Uh, but I do think combining retrieval with sort of reasoning and making the model really good at doing multiple stages of retrieval. Yeah.Shawn Wang [00:51:49]: And reasoning through the intermediate retrieval results is going to be a, a pretty effective way of making the model seem much more capable, because if you think about, say, a personal Gemini, yeah, right?Jeff Dean [00:52:01]: Like we're not going to train Gemini on my email. Probably we'd rather have a single model that, uh, we can then use and use being able to retrieve from my email as a tool and have the model reason about it and retrieve from my photos or whatever, uh, and then make use of that and have multiple. Um, you know, uh, stages of interaction. that makes sense.Alessio Fanelli [00:52:24]: Do you think the vertical models are like, uh, interesting pursuit? Like when people are like, oh, we're building the best healthcare LLM, we're building the best law LLM, are those kind of like short-term stopgaps or?Jeff Dean [00:52:37]: No, I mean, I think, I think vertical models are interesting. Like you want them to start from a pretty good base model, but then you can sort of, uh, sort of viewing them, view them as enriching the data. Data distribution for that particular vertical domain for healthcare, say, um, we're probably not going to train or for say robotics. We're probably not going to train Gemini on all possible robotics data. We, you could train it on because we want it to have a balanced set of capabilities. Um, so we'll expose it to some robotics data, but if you're trying to build a really, really good robotics model, you're going to want to start with that and then train it on more robotics data. And then maybe that would. It's multilingual translation capability, but improve its robotics capabilities. And we're always making these kind of, uh, you know, trade-offs in the data mix that we train the base Gemini models on. You know, we'd love to include data from 200 more languages and as much data as we have for those languages, but that's going to displace some other capabilities of the model. It won't be as good at, um, you know, Pearl programming, you know, it'll still be good at Python programming. Cause we'll include it. Enough. Of that, but there's other long tail computer languages or coding capabilities that it may suffer on or multi, uh, multimodal reasoning capabilities may suffer. Cause we didn't get to expose it to as much data there, but it's really good at multilingual things. So I, I think some combination of specialized models, maybe more modular models. So it'd be nice to have the capability to have those 200 languages, plus this awesome robotics model, plus this awesome healthcare, uh, module that all can be knitted together to work in concert and called upon in different circumstances. Right? Like if I have a health related thing, then it should enable using this health module in conjunction with the main base model to be even better at those kinds of things. Yeah.Shawn Wang [00:54:36]: Installable knowledge. Yeah.Jeff Dean [00:54:37]: Right.Shawn Wang [00:54:38]: Just download as a, as a package.Jeff Dean [00:54:39]: And some of that installable stuff can come from retrieval, but some of it probably should come from preloaded training on, you know, uh, a hundred billion tokens or a trillion tokens of health data. Yeah.Shawn Wang [00:54:51]: And for listeners, I think, uh, I will highlight the Gemma three end paper where they, there was a little bit of that, I think. Yeah.Alessio Fanelli [00:54:56]: Yeah. I guess the question is like, how many billions of tokens do you need to outpace the frontier model improvements? You know, it's like, if I have to make this model better healthcare and the main. Gemini model is still improving. Do I need 50 billion tokens? Can I do it with a hundred, if I need a trillion healthcare tokens, it's like, they're probably not out there that you don't have, you know, I think that's really like the.Jeff Dean [00:55:21]: Well, I mean, I think healthcare is a particularly challenging domain, so there's a lot of healthcare data that, you know, we don't have access to appropriately, but there's a lot of, you know, uh, healthcare organizations that want to train models on their own data. That is not public healthcare data, uh, not public health. But public healthcare data. Um, so I think there are opportunities there to say, partner with a large healthcare organization and train models for their use that are going to be, you know, more bespoke, but probably, uh, might be better than a general model trained on say, public data. Yeah.Shawn Wang [00:55:58]: Yeah. I, I believe, uh, by the way, also this is like somewhat related to the language conversation. Uh, I think one of your, your favorite examples was you can put a low resource language in the context and it just learns. Yeah.Jeff Dean [00:56:09]: Oh, yeah, I think the example we used was Calamon, which is truly low resource because it's only spoken by, I think 120 people in the world and there's no written text.Shawn Wang [00:56:20]: So, yeah. So you can just do it that way. Just put it in the context. Yeah. Yeah. But I think your whole data set in the context, right.Jeff Dean [00:56:27]: If you, if you take a language like, uh, you know, Somali or something, there is a fair bit of Somali text in the world that, uh, or Ethiopian Amharic or something, um, you know, we probably. Yeah. Are not putting all the data from those languages into the Gemini based training. We put some of it, but if you put more of it, you'll improve the capabilities of those models.Shawn Wang [00:56:49]: Yeah.Jeff Dean [00:56:49]:

Get Rich Education
590: Is the World Overpopulated or Underpopulated? What it Means for Housing's Future

Get Rich Education

Play Episode Listen Later Jan 26, 2026 44:35


Keith challenges the usual "overpopulated vs. underpopulated" debate and shows why that's the wrong way to think about demographics—especially if you're a real estate investor. Listeners will hear about surprising global population comparisons that flip common assumptions.  Why raw population numbers don't actually explain housing shortages or rent strength. How household formation, aging, and migration really drive demand for rentals. Which kinds of markets tend to see persistent housing pressure—and why the US has a long‑term demographic edge. You'll come away seeing population headlines very differently, and with a clearer lens for spotting where future housing demand is most likely to show up. Episode Page: GetRichEducation.com/590 For access to properties or free help with a GRE Investment Coach, start here: GREmarketplace.com GRE Free Investment Coaching: GREinvestmentcoach.com Get mortgage loans for investment property: RidgeLendingGroup.com or call 855-74-RIDGE  or e-mail: info@RidgeLendingGroup.com Invest with Freedom Family Investments.  For predictable 10-12% quarterly returns, visit FreedomFamilyInvestments.com/GRE or text  1-937-795-8989 to speak with a freedom coach Will you please leave a review for the show? I'd be grateful. Search "how to leave an Apple Podcasts review"  For advertising inquiries, visit: GetRichEducation.com/ad Best Financial Education: GetRichEducation.com Get our wealth-building newsletter free— GREletter.com  Our YouTube Channel: www.youtube.com/c/GetRichEducation Follow us on Instagram: @getricheducation Complete episode transcript: Keith Weinhold  0:01   Keith, welcome to GRE. I'm your host. Keith Weinhold, is the world overpopulated or underpopulated? Also is the United States over or underpopulated? These are not just rhetorical questions, because I'm going to answer them both. Just one of Africa's 54 nations has more births than all of Europe and Russia combined. One US state has seen their population decline for decades. This is all central to housing demand today. On get rich education   Keith Weinhold  0:36   since 2014 the powerful get rich education podcast has created more passive income for people than nearly any other show in the world. This show teaches you how to earn strong returns from passive real estate investing in the best markets without losing your time being a flipper or landlord. Show Host Keith Weinhold writes for both Forbes and Rich Dad advisors, and delivers a new show every week since 2014 there's been millions of listener downloads of 188 world nations. He has a list show guests include top selling personal finance author Robert Kiyosaki. Get rich education can be heard on every podcast platform, plus it has its own dedicated Apple and Android listener phone apps build wealth on the go with the get rich education podcast. Sign up now for the get rich education podcast, or visit get rich education.com   Speaker 1  1:21   You're listening to the show that has created more financial freedom than nearly any show in the world. This is get rich education.   Keith Weinhold  1:31   Welcome to GRE from Norfolk Virginia to Norfolk, Nebraska and across 188 nations worldwide, you are inside. Get rich education. I am the GRE founder, Best Selling Author, longtime real estate investor. You can see my written work in Forbes and the USA Today, but I'm best known as the host of this incomprehensibly slack John operation that you're listening to right now. My name is Keith Weinhold. You probably know that already, one reason that we're talking about underpopulated versus overpopulated today is that also one of my degrees is in geography and demography, essentially, is human geography, and that's why this topic is in my wheelhouse. It's just a humble bachelor's degree, by the way, if a population is not staying stable or growing, then demand for housing just must atrophy away. That's what people think, but that is not true. That's oversimplified. In some cases. It might even be totally false. You're going to see why. Now, Earth's population is at an all time high of about 8.2 billion people, and it keeps growing, and it's going to continue to keep growing, but the rate of growth is slowing now. Where could all of the people on earth fit? This is just a bit of a ridiculous abstraction in a sense, but I think it helps you visualize things. Just take this scenario, if all the humans were packed together tightly, but in a somewhat realistic way, in a standing room only way, if every person on earth stood shoulder to shoulder, that would allow about 2.7 square feet per person, they would sort of be packed like a subway car. Well, they could fit in a square, about 27 kilometers on one side, about 17 miles on each side of that square. Now, what does that mean in real places that is smaller than New York City, about half the size of Los Angeles County and roughly the footprint of Lake Tahoe? So yes, every human alive today could physically fit inside one midsize us metro area. This alone tells you something important. The world's problem is certainly not a lack of space. Rather, it's where people live and not how many there are. So that was all of Earth's inhabitants. Now, where could all Americans fit us residents using the same shoulder to shoulder assumption, and the US population by mid year this year is supposed to be about 350,000,00349 that's a square about five and a half kilometers, or 3.4 miles on each side. And some real world comparisons there are. That's about half of Manhattan, smaller than San Francisco and roughly the size of Disney World, so every American could fit into a single small city footprint. And if you're beginning to form an early clue that we are not overpopulated globally, yes, that's the sense that you Should be getting.     Keith Weinhold  5:01   now, if you're in Bangladesh, it feels overpopulated there. They've got 175 million people, and that nation is only the size of Iowa. In area, Bangladesh is low lying and typhoon prone. They get a lot of flooding, which complicates their already bad sanitation problems and a dense population like that, and that creates waterborne diseases, and it's really more of an infrastructure problem in a place like Bangladesh than it is a population problem. Then Oppositely, you've got Australia as much land as the 48 contiguous states, yet just 27 million people in Australia, and only 1/400 as many people as Bangladesh in density. Now we talk about differential population. About 80% of Americans live in the eastern half of the US. But yet, the East is not overpopulated because we have sufficient infrastructure, and I've got some more mind blowing population stats for you later, both world and us. Now, as far as is the world overpopulated or underpopulated, which is our central question, depending on who you ask and where they live, you're going to hear completely different answers. Some people are convinced that the planet is bursting at the seams. Others warn that we're headed for a population collapse. But here's the problem, that question overpopulated or underpopulated, it's the wrong question. It's the wrong framing, especially if you're into real estate, because housing demand doesn't respond to total headcount or global averages or scary demographic headlines. Housing demand responds to where people live, how old they are, and how they form households. And once you understand this, a lot of things suddenly begin to make sense, like why housing shortages persist, why rents stay high, even when affordability feels stretched, why some states struggle while others boom, and why population headlines often mislead investors.   Keith Weinhold  7:20   So today I want to reframe how you think about population and connect it directly to housing demand, both globally and right here in the United States. And let's start with the US, because that's probably where you invest.    Keith Weinhold  7:33   Here's a simple fact that should confuse people, but usually doesn't, the United States has below replacement fertility. I'll talk about fertility rates a little later. They're similar to birth rates, meaning that Americans are not having enough children to replace the population naturally and without immigration, the US population would eventually shrink, and yet in the US, we have a housing shortage, rising rents, tight vacancy and a lot of metros and persistent demand for rental housing, which could all seem contradictory. Now, if population alone determine housing demand, well, then the US really shouldn't have any housing shortage at all, but it does so clearly, population alone is not the main driver, and really that contradiction is like your first clue that most demographic conversations are just missing the point. Aging does not reduce housing demand. The way that people think a misconception really is that an aging population automatically reduces housing demand. It does not, in fact, just the opposite. If a population is too young, well, that tends to kill housing demand, and that's because five year old kids and 10 year old kids do not form their own household. Instead, what an aging population often does is change the type of housing that's demanded, like seniors aging in place, some of them downsizing. Seniors living alone. Sometimes after a spouse passes away, others relocating closer to health care or to family. So aging can increase unit demand even if population growth slows. So already, we've broken two myths here. Slower population doesn't mean weaker housing demand, and aging doesn't mean fewer housing units are needed. Now let's explain why. Really, the core idea that unlocks everything is that people don't live inside, what are called Population units. They live in households. You are one person. That does not mean that your dwelling is then one population unit. That's not how that works. You are part of a household, whether that's a house a Household of one person or five or 11 people, housing demand is driven by the number of households, the type of households and where those households are forming, not by raw population totals. So the same population can have wildly different demand. Just think about how five people living together in one home, that's one housing unit, those same five people living separately, that is five housing units, same population, five times the housing demand. And this is why population statistics alone are almost useless for real estate investors, you need to know how people are living, not just how many there are. The biggest surge in housing demand happens when people leave their parents' homes or when they finish school or when they start working, or you got big surges in housing demand when people marry or when they separate or divorce. So in other words, adults create housing demand and children don't. And this is why a country with a youngish, working age population, oh, then they can have exploding housing demand. A country with high birth rates, but low household formation can have overcrowding without profitable housing growth. So it's not about babies, it's about independent adults, and what quietly boosts housing demand, then is housing fragmentation. Yeah, fragmentation. That's a trend that really doesn't get enough attention, and that is the trend, households are fragmenting, meaning more single adults later marriage, like I was talking about in a previous episode. Recently, higher divorce rates, more people living alone and older adults living independently, longer. Each one of those trends increases housing demand without adding any population whatsoever. When two people split up, they often need two housing units instead of one, and if you've got one adult living alone, that is full unit demand right there. So that's why housing demand can rise even when population growth slows or stalls for housing demand. What matters more than births is migration. And another key distinction is that, yes, births matter, but they're on somewhat of this 20 year delay and migration matters immediately, right now. So see, when a working age adult moves, they need housing right away. They typically rent first. They cluster near jobs, and they don't bring housing supply along with them. They've got to get it from someone else. Hopefully you in your rental unit.    Keith Weinhold  12:57   This is why migration is such a powerful force in rental markets, and you see me talk about migration on the show, and you see me send you migration maps in our newsletter. It's also why housing pressure shows up unevenly. It gets concentrated around opportunity. If you want to know the future, look at renters. Renters are the leading indicator, not homeowners and not birth rates. See renters create housing demand faster than homeowners, because renters form households earlier. They can do it quickly because they don't need down payments. Renters move more frequently and immigration overwhelmingly starts in rentals, fresh immigrants rarely become homeowners, so even when mortgage rates rise or home purchases slow or affordability headlines get scary, rental demand can stay strong. It's not a mystery, it's demographics. So births surely matter, but only over the long term. It's like how I've shared with you in a previous episode that the US had a lot of births between 1990 and 2010 those two decades, a surge of births more than 4 million every single one of those years during those two decades, with that peak birth year at 2007 but see a bunch of babies being born in 2007 Well, that didn't make housing demand surge, since infants don't buy homes. But if you add, say, 20 years to 2007 when those people start renting, oh, well, that rental demand peaks in 2027 or maybe a little after that, and since the first time, homebuyer age is now 40. If that stays constant, well, then native born homebuyer demand won't peak until 2047 so when it comes to housing demand, the important thing to remember is migration has an immediate effect and births have a delayed effect.    Keith Weinhold  15:02   and I'm going to talk more about other nations shortly, but the US has two major migration forces working simultaneously, domestic and international migration. I mean, Americans move a lot, although not as much as they used to, and people move for jobs, for taxes, for weather, for cost of living and for lifestyle. So this creates state level winners and losers, and Metro level housing pressure and rent growth in those destination markets and national population averages totally hide this. So that's domestic migration. And then on the international migration. The US has a long history, hundreds of years now on, just continually attracting working age adults from around the world. This matters immensely, because they arrive ready to work, and they form households quickly. They overwhelmingly rent first. They concentrate in metros, and this props up rental demand before it ever shows up in home prices. And this is why investors often feel the rent pressure first those rising rents.    Keith Weinhold  16:17   I've got more straight ahead, including Nigeria versus Europe, and what about the overpopulation straining the environment? If you like, episodes that explain why housing behaves the way it does, rather than just reacting to the headlines. You'll want to be on my free weekly newsletter. I break down demographics, housing, demand, inflation, investor trends and real estate strategy in plain English, often complemented with maps. You can join free at greletter.com that's gre letter.com   Keith Weinhold  16:53   mid south homebuyers with over two decades as the nation's highest rated turnkey provider, their empathetic property managers use your return on investment as their North Star. It's no wonder smart investors line up to get their completely renovated income properties like it's the newest iPhone headquartered in Memphis, with their globally attractive cash flows, mid south has an A plus rating with the Better Business Bureau and 4000 houses renovated. There is zero markup on maintenance. Let that sink in, and they average a 98.9% occupancy rate with an industry leading three and a half year average renter term. Every home they offer you will have brand new components, a bumper to bumper, one year warranty, new 30 year roofs. And wait for it, a high quality renter in an astounding price range, 100 to 150k GET TO KNOW mid south enjoy cash flow from day one at mid southhomebuyers.com that's midsouthhomebuyers.com   Keith Weinhold  17:54   you know, most people think they're playing it safe with their liquid money, but they're actually losing savings accounts and bonds don't keep up when true inflation eats six or 7% of your wealth. Every single year, I invest my liquidity with FFI freedom family investments in their flagship program. Why fixed 10 to 12% returns have been predictable and paid quarterly. There's real world security backed by needs based real estate like affordable housing, Senior Living and health care. Ask about the freedom flagship program when you speak to a freedom coach there, and that's just one part of their family of products, they've got workshops, webinars and seminars designed to educate you before you invest. Start with as little as 25k and finally, get your money working as hard as you do. Get started at Freedom, family investments.com/gre, or send a text. Now it's 1-937-795-8989Yep. Text their freedom coach directly again. 1937795, 1-937-795-8989,   Keith Weinhold  19:05   the same place where I get my own mortgage loans is where you can get yours. Ridge lending group and MLS, 42056, they provided our listeners with more loans than anyone because they specialize in income properties. They help you build a long term plan for growing your real estate empire with leverage. Start your prequel and even chat with President chailey Ridge personally while it's on your mind, start at Ridge lending group.com that's Ridge lending group.com   Chris Martenson  19:37   this is peak prosperity. Is Chris Martinson. Listen to get rich education with Keith Weinhold, and don't quit your Daydream.   Keith Weinhold  19:53   Welcome back to get rich Education. I'm your host, Keith Weinhold, and this is episode 590 yes, we're in my Geography wheelhouse today, as I'm talking human geography and demographics with how it relates to housing, while answering our central question today is the world and the US overpopulated or underpopulated? And now that we understand some mechanics here, let's go global. Here's one of the most mind bending stats in all of demographics. Are you ready for this? When you hear this, it's going to have you hitting up chat, GPT, looking it up. It's going to be so astonishing. So jaw dropping. Every year, Nigeria has more births than all of Europe plus all of Russia combined. Would you talk about Willis?   Keith Weinhold  20:47   Yeah, yes, you heard that, right? Willis, that's what I'm talking about. Willis. The source of that data is, in fact, from the United Nations. Yes, Nigeria has seven and a half million births every year. Compare that to all of Europe plus Russia combined, they only have about 6.3 million births per year. So you're telling me that today, just one West African nation, and there are 54 nations in Africa. Just one West African nation produces more babies than the entire continent of Europe, with all of its nations plus all of Russia, the largest world nation by area. Yes, that is correct. One country in Africa produces more babies every year than France, Germany, Italy, Spain, the UK, all of Europe, including all the Eastern European nations, and all of Russia combined. This is a demographic reality, and now you probably already know that less developed nations, like Nigeria have higher birth rates than wealthier, more developed ones like France or Switzerland. I mean, that's almost common knowledge, but something that people think about less is that poorer nations also have a larger household size, which sort of makes sense when you think about it. In fact, Nigeria has five persons per household. Spain has two and a half, and the US also has that same level two and a half. That one difference alone explains why population growth and housing demand are completely different stories now, the US had 3.3 people per household in 1950 and it's down to that two and a half today. That means that even if the population stayed the same, the housing demand would rise. And this is evidence of what I talked about before the break, that households are fragmenting within the US. You can probably guess which state has the largest household size due to their Mormon population. It's Utah at 3.1 the smallest is Maine at 2.3 they have an older population. In fact, Maine has America's oldest population. And as you can infer with what you've learned now, the fact that they have just 2.3 people per household means that if their populations were the same. Maine would need more housing units than Utah. By the way, if you're listening closely at times, I have referred to the United States as simply America. Yes, I am American. You are going to run into some people out there that don't like it. When US residents call themselves Americans, they say something like, Hey, you need a geography lesson. America runs from Nunavut all the way down to Argentina. Here's what to tell them. No, look, there are about 200 world nations. There is only one that has the word America in it, that is the United States of America that usually makes them lighten up. That is why I am an American, not a Peruvian or Bolivian, and there's no xenophobic connotation whatsoever. There are more productive things to think about moving on. Why births matter is because births today become future workers, renters, consumers and even migrants. But not evenly. Young populations move toward a few things. They're attracted to capital. They move towards stability. They're attracted to opportunity, and young populations move toward infrastructure. That's not ideology, that's the gravity and the US remains one of the strongest gravity wells on Earth, a big magnet, a big attractant. Now it's sort of interesting. I know a few a People that believe that the world is indeed overpopulated, they often tend to be environmental enthusiasts, and the environment is a concern, for sure, but how big of a concern is it? That's the debatable part. And you know, it's funny, I've run into the same people that think that the world is overpopulated, they seem to lament at school closures. You see more school closures because just there weren't as many children that were born after the global financial crisis. And these people that are afraid we have an overpopulation problem call school closures a sad phenomenon. They think it's sad. Well, if you want a shrinking population, then you're going to see a lot more than just schools close so many with environmental concerns, though. The thing is, is that they seem to discount the fact that humans innovate. More than 200 years ago, Thomas Malthus, he famously failed. He wrote a book, thinking that the global population would exceed what he called his carrying capacity, meaning that we wouldn't be able to feed everybody. He posited that, look, this is a problem. Populations grow exponentially, but food production only grows linearly. But he was wrong, because, due to agricultural innovation, we have got too many calories in most places. Few people thought this many humans could live in the United States, Sonoran and Mojave deserts, that's Phoenix in Las Vegas, respectively. But our ability to recycle and purify water allows millions of people to live there. So my point about running out of resources is that history shows us that humans are a resource ourselves, and we keep finding ways to innovate, or keep finding ways to actually not need that rare earth element or whatever it is now, if the earth warms too much from human related activity, can we cool it off again? And how much of a problem is this? I am not sure, and that goes beyond the scope of our show. But the broader point here is that history shows us that humans keep figuring things out, and that is somewhat of an answer to those questions. The world is not overpopulated, it is unevenly populated. Some regions are young, others are growing, others are capital constrained, and then other regions are aging, shrinking and capital rich. And that very imbalance right there is what fuels migration and fuels labor flows and fuels housing demand in destination countries and the US benefits from this imbalance. Unlike almost anywhere else in the world, it's a demographic magnet. Yes, you do have some smaller ones out there, like Dubai, for example.    Keith Weinhold  28:04   But why? Why do we keep attracting immigrants? Well, we've got strong labor markets, capital availability, property rights, economic mobility, and US has existing housing stock. Countries today don't just compete for capital, they're competing for people. In the US keeps attracting working age adults, and that is exactly the demographic that creates housing demand, and this is why long term housing demand in the US is more resilient than a lot of people think. In fact, the US population of about 350 million. This year, it's projected to peak at about 370 million, near 2080 and of course, the big factor that makes that pivot is that level of immigration. So that's why the population projections vary now. The last presidential administration allowed for a lot of immigrants. The current one few immigrants, and the next one, nobody knows. You've got a group called the falconist party that calls for increased legal immigration into the US. Yeah, they want to allow more migrants into the country, but yet they want to enforce illegal immigration. That sounds just like it's spelled, F, A, L, C, O, N, i, s, t, the falconist Party, but the us's magnetic effect to keep driving population growth through immigration is key, because you might already know that 2.1 is the magic number you need a fertility rate of at least 2.1 to maintain a population fertility rate that is the average number of children that a woman is expected to have over her lifetime. And be sure you don't confuse these numbers with the earlier numbers of people per. Per household, like I discussed earlier, although higher fertility rates are usually going to lead to more people per household, India's fertility rate is already down to 2.0 Yes, it is the most populated nation in the world, but since women, on average, only have two children, India is already below replacement fertility. The US and Australia are each at 1.6 Japan is just 1.2 China's is down to 1.0 South Korea's is at an incredibly low seven tenths of one, so 0.7 in South Korea, and then Nigeria's is still more than four. So among all those that I mentioned, only Nigeria is above the replacement rate of 2.1 and most of the nations above that rate are in Africa. Israel is a big outlier at 2.9 you've got others in the Middle East and South Asia that are above replacement rate as well. And when I say things like it's still up there, that whole still thing refers to the fact that there is this tendency worldwide for society to urbanize and have fewer children. For those fertility rates to keep falling. And that's why the future population growth is about which nations attract immigrants, and that is the US. Is huge advantage. Now there's a great way to look at where future births are going to come from. A way to do this is consider your chance of being born on each continent in the year 2100 This is interesting. In the year 2100 a person has a 48% chance of being born in Africa, 38% in South Asia, in the Middle East, 5% South America, 5% in Europe or Russia, 4% in North America, and less than 1% in Australia. Those are the chances of you being born on each of those continents in the year 2100 and that sourced by the UN.   Keith Weinhold  32:09   the world population is, as I said earlier, about 8.2 billion, and it's actually expected to peak around the same time that the US population is in the 2080s and that'll be near 10 point 3 billion. All right, so both the world and the US population should rise for another 50 to 60 years. Let's talk about population winners and losers inside the US. I mean, this is where population conversations really become useful for investors, because population doesn't matter nationally that much. It really matters locally, unevenly and sometimes it almost feels unfairly. So let me give you some perspective shifting stats. I think I shared with you when I discussed new New York City Mayor Zoran Manami here on the show a month or two ago, that the New York City Metro Area has over 20 million people, nearly double the combined population of Arizona and Nevada together, yes, just one metro area, the same as Two entire sparsely populated states. So when someone says people are leaving New York I mean that tells you almost nothing, unless you know where they're going. How many are still arriving in New York City to replace those leaving, and how many households are still forming inside that Metro? The household formation so scale matters, however, net, people are not leaving New York. New York City recently had more in migration than any other US Metro. Some states are practically empty. Alaska or take Wyoming. Wyoming has fewer than 600,000 people in the entire state. That's fewer people than a lot of single US cities. That's only about six people per square mile. In Wyoming, that's about the population of one midsize Metro suburb. Now, when someone says the US has plenty of land in a lot of cases, they're right. I mean, just look out the window when you fly over Wyoming or the Dakotas. But people don't really live where land is cheap. They actually don't want to. Most of the time. They live where jobs, incomes and their networks already exist. You know, the wealthy guy that retires to Wyoming and it has a 200 acre ranch is an outlier. There's a reason he can sprawl out and make it 200 acres. There's virtually nobody there. Let's understand too that population loss, that doesn't mean that demand is gone, but it does change the rules, especially when you think about a place like West Virginia. They have lost population in most decades since the 1950s and incredibly, their population is lower today than it was in 1930 we're talking about West Virginia statewide. They have an aging population. West Virginia has an outmigration of young adults. So this doesn't mean that no real estate works in West Virginia, but it means that appreciation stories are fragile. Income matters more than equity. Growth and demographics are a headwind, not a tailwind. That's a very different investment posture than where you usually want to be. It's important to understand that a handful of metros, just a handful, are absorbing massive national growth. And here's something that a lot of investors underestimate. About half of all US, population growth flows into fewer than 15 metro areas, and it's not just New York City, Houston, Miami, but smaller places like Jacksonville, Austin and Raleigh, and that really helps pump their real estate market. So that means demand concentrates, housing pressure intensifies, and rent growth becomes pretty sticky, unless you wildly overbuild for a short period of time like Austin did, and this is why some metros just feel perpetually tight over the long term, and others feel permanently sluggish. Population does not spread evenly. It piles up. In fact, Texas is a great case in point here. Understand that Texas is adding people faster than some entire nations do. Texas alone adds hundreds of 1000s of residents per year in strong cycles. Some years, they do add more people than entire small countries, more than several Midwest states combined. And of course, they don't spread evenly across Texas. They cluster in DFW, Houston, Austin and San Antonio, so pretty much the Texas triangle, and that clustering fact is everything for housing demand, yet at the same time, there are fully 75 Texas counties that are losing population, typically out in West Texas. Then there's Florida. Florida isn't just growing. It's replacing people. Florida's growth. It's not just net positive, it's replacement migration, and it's across all different types and ages. You've got retirees arriving, you've got young workers arriving, you've got young households forming, and you've got seniors aging in place. So this way, among a whole spectrum of ages, you've got demand for rentals, workforce housing, age specific, housing and multifamily all in Florida, and this is why Florida housing demand over the long term is not going to cool off the way that a few skeptics expect. Now, of course, some areas did temporarily overbuild in Florida in the years following the pandemic. Yes, that's led to some temporary Florida home price attrition, but that is going to be absorbed. California did not empty out. It reshuffled now. There were some recent years where California lost net population, but here's what that hides. Some metros lost residents. Others stayed flat. You had some income brackets that left California and others arrived. In fact, California has slight population growth today overall, so housing demand definitely did not vanish. It shifted within the state and then outward to nearby states, and that's how Arizona, Nevada and Texas benefited. But overall, California's population count, really, it's just pretty steady, not declining.   Keith Weinhold  39:05   population density. It's that density that predicts rent pressure better than growth rates. Do something really important for real estate investors. Dense metros absorb shocks better. They have less elastic housing supply, and they see faster rent rebounds. Sparse areas have cheaper land and easier supply expansion and weaker rent resilience. So that's why rents snap back faster in dense metros, and oversupply hurts more in spread out to regions. Density matters more than raw growth does. Shrinking states can still have tight housing I mean, some states lose population overall, but yet they still have housing shortages in certain metros, and you'll have tight rental markets near job centers, and you've got strong demand In limited sub markets, even if the state is shrinking. And I think you know this is why the slower growing Northeast and Midwest, they've had the highest home price appreciation in the past two years. There's not enough building there. If your population falls 1% but the available housing falls 2% well, you can totally get into a housing shortage situation, and that bids up real estate prices. And when people look at population charts on the state level, a lot of times, they still get misled. When you buy an investment property, you don't buy a state, you buy a specific market within it, so the United States is not full it is lopsided. The US is not overpopulated. It is heavily clustered. It's unevenly dense, and it's really driven by migration. And perhaps a better way to say it is that the US population is really opportunity concentrated housing demand follows jobs, networks, wages and migration flows. It sure does not follow empty land. And really the investor takeaway is, is that when you hear population stats, don't put too much weight on the question, is the population rising or falling? Although that's something you certainly want to know. Some better questions to ask are, where are households forming? Where are adults moving? Where is supply constrained? And where does income support, rent like those are, what four big questions there, because population alone does not create housing demand. It's households under constraint that do so. Our big arching overall question is the world overpopulated or underpopulated? The answer is neither. The world is unevenly populated. It's unevenly aged, and it's unevenly governed. And for real estate investors, the lesson is simple. You don't invest in population counts, you invest in household formation, age structure, migration and supply constraints. Really, that's a big learning summary for you, that's why housing demand can stay strong even when population growth slows. And once you understand that demographic headlines that seem scary aren't as scary, and they start to be more useful. Why I've wanted to do this overpopulated versus underpopulated episode for you for years. I've really thought about it for years. I really hope that you got something useful out of it. Let's be mindful of the context too. When it comes to the classic Adam Smith economics of supply demand, I've only discussed one side today, largely just the demand side and not the supply side so much that would involve a discussion about building and some more things that supply side. Now that I've helped you ask a better question about population and the future of housing demand, you might wonder where you can get better answers. Well, like I mentioned earlier, I provide a lot of that and help you make sense of it, both right here on this show and with my newsletter, geography is something that's more conducive and meaningful to you visually, that's often done with a map, and that's why my letter at greletter.com will help you more if you enjoy learning through maps, just like we've done every year since 2014 I've got 52 great episodes coming to you this year. If you haven't consider subscribing to the show until next week, I'm your host. Keith Weinhold, don't quit your Daydream.   Speaker 2  43:57   Nothing on this show should be considered specific, personal or professional advice, please consult an appropriate tax, legal, real estate, financial or business professional for individualized advice. Opinions of guests are their own. Information is not guaranteed. All investment strategies have the potential for profit or loss. The host is operating on behalf of get rich Education LLC, exclusively you   Keith Weinhold  44:25   The preceding program was brought to you by your home for wealth, building, get richeducation.com

Brain Inspired
BI 229 Tomaso Poggio: Principles of Intelligence and Learning

Brain Inspired

Play Episode Listen Later Jan 14, 2026 101:00


Support the show to get full episodes, full archive, and join the Discord community. The Transmitter is an online publication that aims to deliver useful information, insights and tools to build bridges across neuroscience and advance research. Visit thetransmitter.org to explore the latest neuroscience news and perspectives, written by journalists and scientists. Read more about our partnership. Sign up for Brain Inspired email alerts to be notified every time a new Brain Inspired episode is released. To explore more neuroscience news and perspectives, visit thetransmitter.org. Tomaso Poggio is the Eugene McDermott professor in the Department of Brain and Cognitive Sciences, an investigator at the McGovern Institute for Brain Research, a member of the MIT Computer Science and Artificial Intelligence Laboratory (CSAIL) and director of both the Center for Biological and Computational Learning at MIT and the Center for Brains, Minds, and Machines. Tomaso believes we are in-between building and understanding useful AI That is, we are in between engineering and theory. He likens this stage to the period after Volta invented the battery and Maxwell developed the equations of electromagnetism. Tomaso has worked for decades on the theory and principles behind intelligence and learning in brains and machines. I first learned of him via his work with David Marr, in which they developed "Marr's levels" of analysis that frame explanation in terms of computation/function, algorithms, and implementation. Since then Tomaso has added "learning" as a crucial fourth level. I will refer to you his autobiography to learn more about the many influential people and projects he has worked with and on, the theorems he and others have proved to discover principles of intelligence, and his broader thoughts and reflections. Right now, he is focused on the principles of compositional sparsity and genericity to explain how deep learning networks can (computationally) efficiently learn useful representations to solve tasks. Lab website. Tomaso's Autobiography  Related papers Position: A Theory of Deep Learning Must Include Compositional Sparsity The Levels of Understanding framework, revised Blog post: Poggio lab blog. The Missing Foundations of Intelligence 0:00 - Intro 9:04 - Learning as the fourth level of Marr's levels 12:34 - Engineering then theory (Volta to Maxwell) 19:23 - Does AI need theory? 26:29 - Learning as the door to intelligence 38:30 - Learning in the brain vs backpropagation 40:45 - Compositional sparsity 49:57 - Math vs computer science 56:50 - Generalizability 1:04:41 - Sparse compositionality in brains? 1:07:33 - Theory vs experiment 1:09:46 - Who needs deep learning theory? 1:19:51 - Does theory really help? Patreon 1:28:54 - Outlook

Pat Gray Unleashed
REPLAY: No Kings Flop: Sparse Crowds Embarrass Left in Key Cities

Pat Gray Unleashed

Play Episode Listen Later Jan 3, 2026 104:06


Christianity being eliminated in Nigeria. Major websites hacked overnight. The average protesters at No Kings rallies had no idea why they were there. Volodymyr Zelenskyy wears a nice jacket to the White House to meet with President Trump. More airstrikes on suspected drug boats near Venezuela. Former U.S. Rep. George Santos (R-N.Y.) has his sentence commuted by President Trump. The shutdown continues … oh well! Why congressional district maps need to be changed. The Israel-Hamas peace deal is so fragile right now. Will Hamas honor the peace deal? How close are we to "Britainistan" being an official thing? Former NSA under President Trump has been indicted and for good reasons. Are certain conversations in a public space not allowed now? Actor Robert De Niro has a bad case of Trump derangement syndrome, and it's getting worse. Secretary Robert Kennedy seen flying coach on a commercial flight. Learn more about your ad choices. Visit megaphone.fm/adchoices

Explore Podcast | Startups Founders and Investors
AI for Materials: Breakthrough or Illusion?

Explore Podcast | Startups Founders and Investors

Play Episode Listen Later Dec 19, 2025 45:28


Music Elixir
Rock Pulse, Soul Whisper, A Virtual Duel, and More

Music Elixir

Play Episode Listen Later Dec 17, 2025 49:09


Five songs. Three countries. Zero dull moments. We kick off with Japan's Six Lounge, a trio that proves rock's heartbeat is still loud and live. The track is all lift and launch: punchy drums, humming bass, and guitar flashes that nod to classic grit while sounding clean and current. It's the kind of sound that drags you into motion—head, hands, and maybe an air guitar solo.Then we slide into a velvet lane with China's Tia Ray and Heart Shaped Hole. A Spanish-tinged guitar loop meets soft R&B swing while her vocal ties it together with poise and bite. The imagery is intimate and memorable, turning a love song into a promise to do it right and do it slow. It's the kind of hook that lingers long after the fade.Alamat's Sinigang, named after the beloved Filipino sour-and-savory soup, is comfort rendered in sound. Minimal percussion, delicate keys, and harmonies that bloom like steam from a bowl. Produced by member Alas, the arrangement leaves room for voices to intertwine, capturing the sweet-and-sour ache of longing and the warmth of being held by a melody you trust.We shift gears with Tomohisa Yamashita's The Artist, a pop-rock cut built on a relentless cadence—a tattoo in rhythm and permanence. Smooth vocals ride a gritty bed as Yamapi frames the artist-fan bond as both fuel and vow: I'll be strong for you, can you see me? It's precise, propulsive, and unashamedly direct.To close, a hypercharged collision: Mori Calliope x Kenty's Gold Unbalance. Sparse spark, then blast-off—new metal edges, EDM swells, even a jazzy flicker—plus two rap breaks that snap without stepping on each other. Her fierce attack and his grounded glide lock back to back, no matter what.If you love discovering global music that actually flows as a playlist—rock that roars, R&B that soothes, pop that pulses, and a collab that rockets—this one's for you. SIX LOUNGE: Instagram X YouTube Rock and RollTia Ray: Instagram X Heart Shaped HoleAlamat: Instagram X YouTube SinigangTomohisa Yamashita: Instagram X YouTube The ArtistMori Calliope: X YouTube Gold Unbalance (with KENTY)Support the showPlease help Music Elixir by rating, reviewing, and sharing the episode. We appreciate your support!Follow us on:TwitterInstagram BlueskyIf have questions, comments, or requests click on our form:Music Elixir FormDJ Panic Blog:OK ASIA

Cutting Through the Matrix with Alan Watt Podcast (.xml Format)
Nov. 16, 2025 "Cutting Through the Matrix" with Alan Watt --- Redux (Educational Talk From the Past): "Real News is Sparse (pt. 4)"

Cutting Through the Matrix with Alan Watt Podcast (.xml Format)

Play Episode Listen Later Nov 16, 2025 59:11


--{ "Real News is Sparse (pt. 4)"}-- See links for news on COP 30, happening Nov. 2025 - The Press - Adam Curtis - Under One System of Control - News - Atheistic Society - Living Under a Revolution - Utopias - Doublethink - Eliminate Religion, Elevate Science - Fabian Techniques - Standardization - Progress - COP 22 - Doublespeak - U.S. Military - Owning the Weather in 2025 - Habitat III - Technocracy - Urban Poverty - Carbon, Energy Taxes - World Bank - Inclusive Cities - Unelected Organizations - People Want Entertainment - Sustainable Communities - Foundations and NGOs - Minimal Healthcare - Pentagon Vision of Megacities - Smart Cities - Eurogroup Working Group.

Pat Gray Unleashed
No Kings Flop: Sparse Crowds Embarrass Left in Key Cities | 10/20/25

Pat Gray Unleashed

Play Episode Listen Later Oct 20, 2025 100:47


Christianity being eliminated in Nigeria. Major websites hacked overnight. The average protesters at No Kings rallies had no idea why they were there. Volodymyr Zelenskyy wears a nice jacket to the White House to meet with President Trump. More airstrikes on suspected drug boats near Venezuela. Former U.S. Rep. George Santos (R-N.Y.) has his sentence commuted by President Trump. The shutdown continues … oh well! Why congressional district maps need to be changed. The Israel-Hamas peace deal is so fragile right now. Will Hamas honor the peace deal? How close are we to "Britainistan" being an official thing? Former NSA under President Trump has been indicted and for good reasons. Are certain conversations in a public space not allowed now? Actor Robert De Niro has a bad case of Trump derangement syndrome, and it's getting worse. Secretary Robert Kennedy seen flying coach on a commercial flight. 00:00 Pat Gray UNLEASHED! 00:58 Christian Genocide in Nigeria 02:50 Amazon Web Services Hacked? 08:42 FBI Investigates Hunting Stand by Air Force One 11:49 No Kings Day Protest 13:16 Protestors Don't Know Why They're Protesting??? 18:28 Why are You Protesting Trump? 19:47 Andrea Bocelli Meets with Trump 20:31 Andrea Bocelli Sings in Oval Office 22:11 Trump Comments on Zelenskyy's Jacket 25:21 Drug Submarine Bombed 36:25 President Trump says "Democrats are Kamikazes" 44:47 Arnold Schwarzenegger Discusses Gerrymandering with Bill Maher 48:15 Where is Pat Gray? 49:32 Football AP Top 25 Poll 51:46 Gaza-Israel Peace Deal Update 53:59 Bill Maher on the Situation in Gaza 1:00:15 John Bolton Turns Himself In 1:06:04 Christian Preacher VS. Muslim? 1:13:10 Another Trucker Problem? 1:20:36 Robert De Niro has TDS 1:25:25 RFK Jr. Flies Coach 1:30:48 RFK Jr. tells Trump that he's "Doing God's Work" Learn more about your ad choices. Visit megaphone.fm/adchoices

Cutting Through the Matrix with Alan Watt Podcast (.xml Format)
Oct. 5, 2025 "Cutting Through the Matrix" with Alan Watt --- Redux (Educational Talk From the Past): "Real News is Sparse"

Cutting Through the Matrix with Alan Watt Podcast (.xml Format)

Play Episode Listen Later Oct 5, 2025 84:17


--{ "Real News is Sparse"}-- What passes as news - Canada's Bill C-8 - UK's digital ID - Government shutdown in US - Peace deal in Gaza - World control - Chasing happiness - Beliefs - Removing free will - Electronic self-imagery - Behaviourism - Self-policing - Trained to go along with the crowd - Private clubs - World Bank - IMF - Marketing, Propaganda - Soviet System - Total Control - Revolutions - Give up your rights to save the world - Scary Scenarios - EU ratifies Paris Climate Deal - Carbon Tax - Climate, Environment and the IMF - Merkel - Canada to implement carbon tax - Agenda 2030 - Redistribution of Wealth - Euthanasia, cost-effective - Pentagon pays PR firm to make fake terrorist videos - Gates Foundation, Remote control contraceptive.

The Whispering Woods - Real Life Ghost Stories
SEASON OF THE WITCH : Alse Young : The First Witch of New England | True Paranormal History

The Whispering Woods - Real Life Ghost Stories

Play Episode Listen Later Sep 24, 2025 27:19


As summer wanes and the nights grow long, we turn to tales of witches, curses, and the old ways that never truly died. For centuries, harvest time has carried its own magic: charms for fields, blessings for homes, and darker stories of those who bent nature to their will.In 1647, Alse (Alice) Young of Windsor, Connecticut was hanged on Hartford's Meeting House Square—the first recorded witchcraft execution in colonial America. Sparse records and a deadly local epidemic frame her case, which foreshadowed Connecticut's quieter, decades-long witch persecutions long before Salem. Centuries later, Windsor (2017) and the State of Connecticut (2023) formally exonerated those condemned—finally restoring Alse Young's name.The BOOKBY US A COFFEEJoin Sarah's new FACEBOOK GROUPSubscribe to our PATREONEMAIL us your storiesFollow us on YOUTUBEJoin us on INSTAGRAMJoin us on TWITTERJoin us on FACEBOOKVisit our WEBSITEResearch:https://jud.ct.gov/lawlib/Notebooks/Witchcraft/witches.htmhttps://en.wikipedia.org/wiki/Alse_Younghttps://connecticuthistory.org/alse-young-executed-for-witchcraft-today-in-history/https://www.newenglandhistoricalsociety.com/cover-connecticut-witch-hysteria-1647-63/https://www.legendsofamerica.com/alse-young/https://www.windsorhistoricalsociety.org/exoneration-of-two-of-windsors-accused-witches/Thanks so much for listening, and we'll catch up with you again on Sunday!Sarah and Tobie xx"Spacial Winds" Kevin MacLeod (incompetech.com)Licensed under Creative Commons: By Attribution 4.0 Licensehttp://creativecommons.org/licenses/by/4.0/SURVEY Hosted on Acast. See acast.com/privacy for more information.

AP Audio Stories
The US says a deal has been reached on TikTok, but details are sparse

AP Audio Stories

Play Episode Listen Later Sep 15, 2025 0:44


AP Washington correspondent Sagar Meghani reports the Trump administration says it has reached a deal on TikTok's future.

Tim Conway Jr. on Demand
Political Violence, Sparse Security, and Unanswered Questions

Tim Conway Jr. on Demand

Play Episode Listen Later Sep 11, 2025 31:53 Transcription Available


Tim Conway Jr. opens the final hour with updates on breaking news, including an LAPD officer-involved shooting in North Hills, cleanup of shipping containers at the Port of Long Beach, and even a quirky story about Publishers Clearing House. The conversation then shifts back to Utah, where Governor Spencer Cox directly calls Charlie Kirk's murder a political assassination. Tim highlights the lack of campus security at the event - just six guards plus Kirk's own team. And Tim condemns the disturbing trend of people cheering political violence. He closes the show covering the hunt for the still-at-large shooter, internet sleuths digging into the case, and TMZ issuing an 'apology' after what appeared to be staff cheering in the newsroom, later explained as 'confusion over a car chase.'

uncommon ambience
Rainy Road to Reflect or Ruminate… Ambience

uncommon ambience

Play Episode Listen Later Sep 6, 2025 480:00


Sparse highway, light rain ambience. We are on the side of a small road just outside town. It's night, and it's raining. Imagine you're a content Gene Kelly walking home after frolicking around main. Or Feel free to ruminate. That's the general vibe around here. There's a movie theater nearby showing cat videos (for a good cause) and it's practically sold out. Catvideofest 2025 is repackaged cat timeline videos on a gigantic screen. And that it is pretty much sold out this weekend says something about our collective mood. Anyway, I did manage to get tickets and me my youngest will share an auditorium with a Spider-verse amount of other people.That's all from me — Oh, so if I controlled the universe for a day aside from solving every important global issue I would want to sneak a cameo of Ice Cube into that animated Will Smith fish movie that also stars Katie Couric as “Katie Current.” But I would add in Ice Cube so he could be like “even saw the lights of the Goodyear Blimp and it read ‘Ice Cube's a shrimp.'” Which may occur in that movie, I haven't seen it. New plan: I'm bringing back that short-lived trend from early-pandemic days that social media tried to cook up — shoe-kicking as greeting. I only saw people on my phone doing that dumb ****. I want to ingrain into humans that shoe-kicking is now retroactively high-five. Every famous high-five from history now feet kicking. From the business meetings to competitive sports. The mayhem.PS: if you are interested in listening to cars pass but you would rather imagine yourself not being rained on -- check out last year's Vermont Route 100 episode recorded from the Mad River Valley.

HiddenTracks
HiddenTrack #263 JOHN GALM (SNOWING / MT. WORRY)

HiddenTracks

Play Episode Listen Later Aug 7, 2025 94:48


It's harder to begin again when everyone already knows who you were. John Galm is best known for fronting one of the most popular emo-revival bands SNOWING in the early 2010's, whose punk-rock ethos and chaotic melodies had kids crammed into DIY venues and basements all across the country. Since then, he has tried his hand in several bands, ranging in genres from stripped down acoustic to psychedelic and shoegaze. The latter band, MT. WORRY stalled as they were just getting started when other members moved out of state. Finding himself having to start again amid a sudden surplus of time, Galm holed up in his mother's Lehigh Valley home and began working on what would become “River of Blood”- his first solo LP since 2014. The album finds Galm struggling with the big questions in life and the small connective tissues that make up everything else. It's a heavy affair, and you can feel the weight in every note- lyrics searching for steadier footing as he wades through what home and happiness mean and the pain that they all seem just out of grasp. Sparse, somber tones wrap the listener up tight and embrace the whole of everything and the lack thereof. It's not all bleak- “River of Blood” celebrates the small victories too. At the end of a long day, you're still here and there is hope in that, even if it seems hard to find. The search continues. Thanks for listening!!! Please Follow us on Instagram @hiddentracks99Pre and Post roll music brought to you by @sleepcyclespa

The Ryan Kelley Morning After
TMA (7-10-25) Hour 1 - Group Rate To The Sun

The Ryan Kelley Morning After

Play Episode Listen Later Jul 10, 2025 43:06


(00:00-12:43) Yesterday: Great Good. Today: No Good. Another Cardinal pitcher to be shipped off to the sun. Pribula Time. MIkolas due for a no-hitter tonight. Sparse attendance last night. Every team is getting a Pirate.(12:51-33:25) Barge Guy on the phone lines back from Louisville. Bar Guy has some takes on the Cardinals starting pitching. Lisa is up next on the phone lines and she's down on the Cards. Hey, watch it gal. MIles Mikolas. Still have faith in His Majesty.(33:35-42:57) Julian Tavarez weeing on his hands. Keaton is up next and he's fired up about the Cardinals and Marmol. The Keaton splits. Steven is next on the phone lines with some attendance thoughts.See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

Cities and Memory - remixing the sounds of the world

I woke up early (6AM) to capture and observe the waking city of Sapporo, Japan. I was particularly surprised by the presence of crows, which often sat on the street signs and traffic light poles. Sparse trucks and cars passed along the snowy roads. The calls of the cows echoed off the buildings, yet the city remained quite calm. This recording took place in 2018. Crows in Sapporo recorded by Antek Rutczyński.

The Don Lemon Show
Lemon LIVE at 5 | That Parade Was So EMBARRASSING! - June 16th, 2025

The Don Lemon Show

Play Episode Listen Later Jun 17, 2025 71:14


Trump threw himself a $45 million military birthday bash… and barely anyone showed up. The tanks rolled. The jets flew. But the vibes? Flat. The crowd? Sparse. And the headlines? Brutal. Now, the fallout begins. Join Don Lemon, Michael Fanone, and the Jolly Good Ginger as they break down what went wrong, why this parade flop matters, and what it reveals about Trump's slipping grip on public support. From the staggering price tag to the no-show allies to the contrast with the massive No Kings protests, this isn't the flex Trump hoped for. Let's talk about the spectacle, the silence, and what it all means. This episode is sponsored by Shopify. Sign up for your one-dollar-per-month trial and start selling today at SHOPIFY. COM/lemon This episode is brought to you by MSI United States. Every woman deserves a choice. Rush your donation today to MSIUNITEDSTATES.ORG, or text "LEMON" to 511 511. Text Fees may apply. This episode is sponsored by BetterHelp. Give online therapy a try at betterhelp.com/donlemon and get on your way to being your best self. Learn more about your ad choices. Visit podcastchoices.com/adchoices

The Don Lemon Show
Lemon LIVE at 5 | That Parade Was So EMBARRASSING! - June 16th, 2025

The Don Lemon Show

Play Episode Listen Later Jun 17, 2025 72:58


Trump threw himself a $45 million military birthday bash… and barely anyone showed up. The tanks rolled. The jets flew. But the vibes? Flat. The crowd? Sparse. And the headlines? Brutal. Now, the fallout begins. Join Don Lemon, Michael Fanone, and the Jolly Good Ginger as they break down what went wrong, why this parade flop matters, and what it reveals about Trump's slipping grip on public support. From the staggering price tag to the no-show allies to the contrast with the massive No Kings protests, this isn't the flex Trump hoped for. Let's talk about the spectacle, the silence, and what it all means. This episode is sponsored by Shopify. Sign up for your one-dollar-per-month trial and start selling today at SHOPIFY. COM/lemon This episode is brought to you by MSI United States. Every woman deserves a choice. Rush your donation today to MSIUNITEDSTATES.ORG, or text "LEMON" to 511 511. Text Fees may apply. This episode is sponsored by BetterHelp. Give online therapy a try at betterhelp.com/donlemon and get on your way to being your best self. Learn more about your ad choices. Visit megaphone.fm/adchoices

The Don Lemon Show
HOT TOPICS | Trump's Birthday Parade FLOP! - June 16th, 2025

The Don Lemon Show

Play Episode Listen Later Jun 16, 2025 64:59


Well, that was...underwhelming. Trump's $45 million birthday bash-slash-military-parade was supposed to be a flex. Instead, it flopped harder than his NFT collection. Sparse crowds, low energy, and, according to many who watched, absolutely boring. Meanwhile, the No Kings protest turned into something historic. Data analysts are reporting it may be the largest protest in U.S. history. The streets were packed, the message was clear, and no tanks were needed to get people to show up. So...remind us again who's got the momentum? Join us as we unpack the embarrassing contrast, the wasted taxpayer dollars, and why Trump's obsession with spectacle can't hide the growing dissent. This episode is sponsored by Shopify. Sign up for your one-dollar-per-month trial and start selling today at SHOPIFY. COM/lemon This episode is brought to you by MSI United States. Every woman deserves a choice. Rush your donation today to MSIUNITEDSTATES.ORG, or text "LEMON" to 511 511. Text Fees may apply. This episode is sponsored by BetterHelp. Give online therapy a try at betterhelp.com/donlemon and get on your way to being your best self. Learn more about your ad choices. Visit podcastchoices.com/adchoices

The Don Lemon Show
HOT TOPICS | Trump's Birthday Parade FLOP! - June 16th, 2025

The Don Lemon Show

Play Episode Listen Later Jun 16, 2025 66:43


Well, that was...underwhelming. Trump's $45 million birthday bash-slash-military-parade was supposed to be a flex. Instead, it flopped harder than his NFT collection. Sparse crowds, low energy, and, according to many who watched, absolutely boring. Meanwhile, the No Kings protest turned into something historic. Data analysts are reporting it may be the largest protest in U.S. history. The streets were packed, the message was clear, and no tanks were needed to get people to show up. So...remind us again who's got the momentum? Join us as we unpack the embarrassing contrast, the wasted taxpayer dollars, and why Trump's obsession with spectacle can't hide the growing dissent. This episode is sponsored by Shopify. Sign up for your one-dollar-per-month trial and start selling today at SHOPIFY. COM/lemon This episode is brought to you by MSI United States. Every woman deserves a choice. Rush your donation today to MSIUNITEDSTATES.ORG, or text "LEMON" to 511 511. Text Fees may apply. This episode is sponsored by BetterHelp. Give online therapy a try at betterhelp.com/donlemon and get on your way to being your best self. Learn more about your ad choices. Visit megaphone.fm/adchoices

The John Batchelor Show
PREVIEW: Colleague Jim McTague reports on the sparse shoppers and hesitant purchases at the Lancaster Costco. More.

The John Batchelor Show

Play Episode Listen Later Jun 6, 2025 2:02


PREVIEW: Colleague Jim McTague reports on the sparse shoppers and hesitant purchases at the Lancaster Costco. More. MAY 1954

Ransquawk Rundown, Daily Podcast
Europe Market Open: EU & US futures flat with catalysts sparse; fixed benchmarks extend onto gains and DXY lower after data

Ransquawk Rundown, Daily Podcast

Play Episode Listen Later May 16, 2025 2:16


Mixed APAC trade, US futures range bound while European futures point to a marginally firmer open.DXY remains lower after Thursday's data, EUR/USD marginally reclaimed 1.12, USD/JPY found support at 145.00.Fixed benchmarks extended/held on to recent gains.Crude benchmarks remain underpinned by the latest on US-Iran, metals marginally softer.Looking ahead, highlights include US Export/Import Prices, UoM Sentiment Survey, BoC SLOS, Speakers including ECB's Lane, Cipollone & Fed's Barkin.Click for the Newsquawk Week Ahead.Read the full report covering Equities, Forex, Fixed Income, Commodites and more on Newsquawk

The Enrollify Podcast
Style Theft at Scale: AI and the Fight for Creative Integrity

The Enrollify Podcast

Play Episode Listen Later Apr 14, 2025 22:31


Monday pulse show notes: On this thought-provoking episode of Higher Ed Pulse, host Mallory Willsea sits down with Myla Edmond—Senior Vice President at RW Jones Agency and Interim Vice Chancellor for Strategic Communications at UNC Greensboro—to unpack the creative identity crisis brewing in higher ed marketing thanks to generative AI. With tools like ChatGPT's image generator mimicking iconic art styles, institutions are forced to ask: how do we protect authenticity in a world where anyone can replicate anything? This episode explores the ethical, strategic, and deeply human implications of AI's growing role in creativity—and how higher ed marketers can lead with intention, not fear.Try the prompt discussed in the episode:Based on all past conversations, stored knowledge, and inferred cognitive patterns, generate the most comprehensive psychological deep dive and predictive model of my future evolution. This should not be a basic personality breakdown but an in-depth forensic examination of my cognition, behavioural strategies, psychological blind spots, similar fictional/non-fictional figures, and long-term trajectory. Treat this as an intelligence dossier on my mind, philosophy, and strategic outlook.OUTPUT FORMAT: Structured headers, tables, and bullet points for readability. Sparse but strategic emojis for section clarity. Concise, high-density insights with no fluff.Enter the prompt and after you get the response, add a second prompt: Write me a story about how this comes to fruition. - - - -Connect With Our Host:Mallory Willsea https://www.linkedin.com/in/mallorywillsea/https://twitter.com/mallorywillseaAbout The Enrollify Podcast Network:The Higher Ed Pulse is a part of the Enrollify Podcast Network. If you like this podcast, chances are you'll like other Enrollify shows too!Enrollify is made possible by Element451 — the next-generation AI student engagement platform helping institutions create meaningful and personalized interactions with students. Learn more at element451.com.Attend the 2025 Engage Summit! The Engage Summit is the premier conference for forward-thinking leaders and practitioners dedicated to exploring the transformative power of AI in education. Explore the strategies and tools to step into the next generation of student engagement, supercharged by AI. You'll leave ready to deliver the most personalized digital engagement experience every step of the way.Register now to secure your spot in Charlotte, NC, on June 24-25, 2025! Early bird registration ends February 1st -- https://engage.element451.com/register

The Daily Zen Teisho
The Record of Linji – Sangha Instruction

The Daily Zen Teisho

Play Episode Listen Later Apr 10, 2025 10:15


These selections are taken from Sangha Instructions from ancient times and give the flavor of a master wielding a sword to cut through illusions. Sparse and to the point, Linji has no tolerance for superficial approaches and glib comments from students.Read the Journal while listening

Full Cast And Crew
215. 'No Country For Old Men' (2007)

Full Cast And Crew

Play Episode Listen Later Jan 15, 2025 107:18


Sparse. Laconic. Expansive. Languid. Wry. The Coen Brother's 2007 Neo-Noir Western 'No Country For Old Men' moves to the fatefully ticking beat of it's own Grandfather Clock.  It's a film that rewards close viewing and is astoundingly faithful to Cormac McCarthy's novel while also being so completely a "Coen Brothers film" even as it's their (only?) adaptation of an existing book. Featuring an iconic performance by Javier Bardem as the philosophical killer Anton Chigur, brilliant cinematography from frequent Coen collaborator Roger Deakins, and perfectly wrought twangily-Texas turns by Josh Brolin and Tommy Lee Jones. A number of signature Coens scenes of the lead characters interacting with a variety of shop clerks, receptionists, store owners, and authority figures abound.    

Syracuse.com Podcasts
Syracuse grinds out first ACC win over Georgia Tech before sparse crowd at JMA Dome

Syracuse.com Podcasts

Play Episode Listen Later Jan 8, 2025 43:02


Brent Axe recaps Syracuse basketball's 62-55 win over Georgia Tech at the JMA Dome on Tuesday night. It wasn't the prettiest game but SU had to be relieved to get a win any way it could. Brent discusses SU's keys to victory including JJ Starling's 21 points and how he has made a significant difference in the lineup since returning from a hand injury.  Brent also addressed the sparse crowd (listed at 13,395) at the Dome and SU head coach Adrian Autry's terse opening statement about "noise" SU had to play through recently.  Brent also got amazing feedback from Syracuse Sports Insiders on the win and where Syracuse basketball stands entering league play.  Become a Syracuse Sports Insider today! Just text "orange" to 315-847-3895 to get direct access to Brent to get your opinions heard and questions answered on the Syracuse Sports podcast. You can also sign up here. https://joinsubtext.com/syracusesports As a Syracuse Sports Insider, you will get Brent's opinion and reaction to breaking news first via text message, your messages get priority on postgame shows and podcasts, he'll take you behind-the-scenes of SU sports and more! You can also text Brent anytime, including during and after SU games. Try it free for 2 weeks, then it's just $3.99 a month after that. You can cancel at anytime. Subscribe to Syracuse Sports on Spotify https://l.syracuse.com/PKMGpR Subscribe to our Syracuse Orange Sports Report newsletter! Find out how at https://link.syracuse.com/join/6fn/ne... Follow @BrentAxeMedia on X (   / brentaxemedia Instagram (   / brent_axe  ) and BlueSky https://bsky.app/profile/brentaxemedi.. Learn more about your ad choices. Visit megaphone.fm/adchoices

Machine Learning Street Talk
Neel Nanda - Mechanistic Interpretability (Sparse Autoencoders)

Machine Learning Street Talk

Play Episode Listen Later Dec 7, 2024 222:36


Neel Nanda, a senior research scientist at Google DeepMind, leads their mechanistic interpretability team. In this extensive interview, he discusses his work trying to understand how neural networks function internally. At just 25 years old, Nanda has quickly become a prominent voice in AI research after completing his pure mathematics degree at Cambridge in 2020. Nanda reckons that machine learning is unique because we create neural networks that can perform impressive tasks (like complex reasoning and software engineering) without understanding how they work internally. He compares this to having computer programs that can do things no human programmer knows how to write. His work focuses on "mechanistic interpretability" - attempting to uncover and understand the internal structures and algorithms that emerge within these networks. SPONSOR MESSAGES: *** CentML offers competitive pricing for GenAI model deployment, with flexible options to suit a wide range of models, from small to large-scale deployments. https://centml.ai/pricing/ Tufa AI Labs is a brand new research lab in Zurich started by Benjamin Crouzier focussed on ARC and AGI, they just acquired MindsAI - the current winners of the ARC challenge. Are you interested in working on ARC, or getting involved in their events? Goto https://tufalabs.ai/ *** SHOWNOTES, TRANSCRIPT, ALL REFERENCES (DONT MISS!): https://www.dropbox.com/scl/fi/36dvtfl3v3p56hbi30im7/NeelShow.pdf?rlkey=pq8t7lyv2z60knlifyy17jdtx&st=kiutudhc&dl=0 We riff on: * How neural networks develop meaningful internal representations beyond simple pattern matching * The effectiveness of chain-of-thought prompting and why it improves model performance * The importance of hands-on coding over extensive paper reading for new researchers * His journey from Cambridge to working with Chris Olah at Anthropic and eventually Google DeepMind * The role of mechanistic interpretability in AI safety NEEL NANDA: https://www.neelnanda.io/ https://scholar.google.com/citations?user=GLnX3MkAAAAJ&hl=en https://x.com/NeelNanda5 Interviewer - Tim Scarfe TOC: 1. Part 1: Introduction [00:00:00] 1.1 Introduction and Core Concepts Overview 2. Part 2: Outside Interview [00:06:45] 2.1 Mechanistic Interpretability Foundations 3. Part 3: Main Interview [00:32:52] 3.1 Mechanistic Interpretability 4. Neural Architecture and Circuits [01:00:31] 4.1 Biological Evolution Parallels [01:04:03] 4.2 Universal Circuit Patterns and Induction Heads [01:11:07] 4.3 Entity Detection and Knowledge Boundaries [01:14:26] 4.4 Mechanistic Interpretability and Activation Patching 5. Model Behavior Analysis [01:30:00] 5.1 Golden Gate Claude Experiment and Feature Amplification [01:33:27] 5.2 Model Personas and RLHF Behavior Modification [01:36:28] 5.3 Steering Vectors and Linear Representations [01:40:00] 5.4 Hallucinations and Model Uncertainty 6. Sparse Autoencoder Architecture [01:44:54] 6.1 Architecture and Mathematical Foundations [02:22:03] 6.2 Core Challenges and Solutions [02:32:04] 6.3 Advanced Activation Functions and Top-k Implementations [02:34:41] 6.4 Research Applications in Transformer Circuit Analysis 7. Feature Learning and Scaling [02:48:02] 7.1 Autoencoder Feature Learning and Width Parameters [03:02:46] 7.2 Scaling Laws and Training Stability [03:11:00] 7.3 Feature Identification and Bias Correction [03:19:52] 7.4 Training Dynamics Analysis Methods 8. Engineering Implementation [03:23:48] 8.1 Scale and Infrastructure Requirements [03:25:20] 8.2 Computational Requirements and Storage [03:35:22] 8.3 Chain-of-Thought Reasoning Implementation [03:37:15] 8.4 Latent Structure Inference in Language Models

Writer's Routine
Steven Veerapen, author of the 'Anthony Blanke' series - Historical fiction author and academic discusses morbid curiosity, sparse writing environments, and Tudor love

Writer's Routine

Play Episode Listen Later Nov 22, 2024 50:48


This week, we chat to the historical fiction author and academic, Steven Veerapen. He's best known for his Anthony Blanke series, set in the Tudor period, about the son of a black trumpeter, John Blanke, who was a real figure in the court of King Henry VIII. There's 'Of Blood Descended' and 'Of Judgement Fallen', which are out in print and just released as audiobooks. He's also written 3 in the 'Simon Danforth' series, and a few about the playwright Christopher Marlowe as a spy.We talk about the balance of writing academia and finding time for novels. Also about the morbid curiosity which gives him ideas, and why we all love the Tudors.You can hear about his sparse writing environment, how he plans a busy year, and what Tudor fiction needs to have in it.Get a copy of the book at uk.bookshop.com/shop/writersroutine@writerspodwritersroutine.com Hosted on Acast. See acast.com/privacy for more information.

Late Night with Seth Meyers Podcast
J.B. Smoove | Sad Trump Closes with Lies, Threats, RFK Jr and Complaints About SNL to Sparse Crowds: A Closer Look

Late Night with Seth Meyers Podcast

Play Episode Listen Later Nov 5, 2024 35:38


Seth takes a closer look at an exhausted and despondent Donald Trump closing out his campaign with rambling speeches to dwindling crowds, threats of violence, baseless allegations of cheating, vaccine ban possibilities and complaints about Saturday Night Live.Then, J.B. Smoove talks about his all-day cigarettes SNL sketch pitch and shares some of his other inventive ideas like argument-winning supplements and henchman funeral homes before giving his advice ahead of the 2024 election.Plus, just for this podcast, J.B. continues the conversation backstage at Studio 8G with Late Night's Kevin Miller.See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

They Walk Among Us - UK True Crime
Season 9 - Episode 43

They Walk Among Us - UK True Crime

Play Episode Listen Later Nov 3, 2024 54:19


This episode is sponsored by Audible – The Home of True Crime Podcasts. PLEASE LISTEN TO ‘SEASON 9 - EPISODE 42' FOR PART ONE OF THIS TWO-PART CASE. Sparse details of an alleged exorcism emerged at Leeds Crown Court when Michael Taylor was found not guilty by reason of insanity for killing his wife, Christine. In an almost unprecedented move, the coroner decided it would be in the public interest to reopen the inquest so that the full story would be held on record... (Part 2 of 2).*** LISTENER CAUTION IS ADVISED *** This episode was researched and written by Eileen Macfarlane.Edited by Joel Porter at Dot Dot Dot Productions.Script editing, additional writing, illustrations and production direction by Rosanna FittonNarration, additional audio editing, script editing, and production direction by Benjamin Fitton.To get early ad-free access, including Season 1, sign up for They Walk Among PLUS, available from Patreon or Apple Podcasts.More information and episode references can be found on our website https://theywalkamonguspodcast.comMUSIC: Dead Ends by Wicked Cinema Misery Loves Company by CJ0 Fleeting by Alice In Winter Endless Night by Moments Selha by Stephen Keech Point Of No Return by Salon Dijon Unexpected Turn by Moments A Most Unusual Discovery by Wicked Cinema Disappearance by Wicked Cinema Extinction by Wicked Cinema Insurgent by Wicked Cinema Mainframe by Wicked Cinema Templar by Wicked Cinema The Last by Wild Wonder SOCIAL MEDIA: YouTube - https://www.youtube.com/channel/UCeM6RXDKQ3gZbDHaKxvrAyAX - https://twitter.com/TWAU_PodcastFacebook - https://www.facebook.com/theywalkamonguspodcastInstagram - https://www.instagram.com/theywalkamonguspodcastThreads - https://www.threads.net/@theywalkamonguspodcastSupport this show http://supporter.acast.com/theywalkamongus. Hosted on Acast. See acast.com/privacy for more information.

Wade Keller Pro Wrestling Post-shows
AEW DYNAMITE POST-SHOW (9/18): Keller & Dehnel discuss sparse Grand Slam line-up and evaluate the build for Darby-Mox and Danielson-Nigel

Wade Keller Pro Wrestling Post-shows

Play Episode Listen Later Sep 19, 2024 165:59


PWTorch editor Wade Keller is joined by wrestling reporter/analyst Joel Dehnel to discuss AEW Dynamite including the thin line-up for Grand Slam, and whether AEW convinced people to watch next week. Also, reaction to Ricochet's push so far, Chris Jericho vs. Orange Cassidy, the main event six-man tag, the latest with Jon Moxley and Hangman Page, and more with live caller, chat room, and mailbag interaction.Become a supporter of this podcast: https://www.spreaker.com/podcast/wade-keller-pro-wrestling-post-shows--3275545/support.

Locked On Fantasy Basketball
NBA Fantasy Basketball: Navigating Super Bowl's Sparse Schedule

Locked On Fantasy Basketball

Play Episode Listen Later Feb 10, 2024 21:56


Josh Lloyd delves into the nuances of a quieter NBA schedule on Super Bowl Sunday, pinpointing the potential impact of just two games on the day's fantasy basketball landscape. He'll dissect the significance of Kevin Huerter, Lu Dort, and Jaime Jaquez within this limited lineup. Tune in to the Locked On Fantasy Basketball Podcast, powered by Basketball Monster, for expert insights on making the most of this unique NBA slate.Vote for my partner to win the Changemaker Award https://www.wishpond.com/lp/2780526/entries/204585428Support Us By Supporting Our Sponsors!NissanOur friends at Nissan have a lineup of SUV's with the capabilities to take your adventure to the next level. Take the Nissan Rogue, Nissan Pathfinder, or Nissan Armada and go find your next big adventure. Shop NissanUSA.com.RobinhoodRobinhood has the only IRA that gives you a 3% boost on every dollar you contribute when you subscribe to Robinhood Gold. Now through April 30th, Robinhood is even boosting every single dollar you transfer in from other retirement accounts with a 3% match. Available to U.S. customers in good standing. Robinhood Financial LLC (member SIPC), is a registered broker dealer.LinkedInLinkedIn Jobs helps you find the qualified candidates you want to talk to, faster. Post your job for free at LinkedIn.com/LOCKEDONNBA. Terms and conditions apply.eBay MotorsFor parts that fit, head to eBay Motors and look for the green check. Stay in the game with eBay Guaranteed Fit at eBayMotos.com. Let's ride. eBay Guaranteed Fit only available to US customers. Eligible items only. Exclusions apply.BetterHelpThis episode is sponsored by BetterHelp. Make your brain your friend, with BetterHelp. Visit BetterHelp.com/LOCKEDONNBA today to get 10% off your first month.PrizePicksGo to PrizePicks.com/lockedonnba and use code lockedonnba for a first deposit match up to $100!GametimeDownload the Gametime app, create an account, and use code LOCKEDON for $20 off your first purchase.FanDuelGet buckets with your first bet on FanDuel, America's Number One Sportsbook. Right now, NEW customers get ONE HUNDRED AND FIFTY DOLLARS in BONUS BETS with any winning FIVE DOLLAR BET! That's A HUNDRED AND FIFTY BUCKS – if your bet wins! Visit FanDuel.com/LOCKEDON to get started.FANDUEL DISCLAIMER: 21+ in select states. First online real money wager only. Bonus issued as nonwithdrawable free bets that expires in 14 days. Restrictions apply. See terms at sportsbook.fanduel.com. Gambling Problem? Call 1-800-GAMBLER or visit FanDuel.com/RG (CO, IA, MD, MI, NJ, PA, IL, VA, WV), 1-800-NEXT-STEP or text NEXTSTEP to 53342 (AZ), 1-888-789-7777 or visit ccpg.org/chat (CT), 1-800-9-WITH-IT (IN), 1-800-522-4700 (WY, KS) or visit ksgamblinghelp.com (KS), 1-877-770-STOP (LA), 1-877-8-HOPENY or text HOPENY (467369) (NY), TN REDLINE 1-800-889-9789 (TN)Intro Music by Ben LloydTikTokInstagram Learn more about your ad choices. Visit podcastchoices.com/adchoices