Podcasts about nvidia gpus

  • 311PODCASTS
  • 501EPISODES
  • 48mAVG DURATION
  • 5WEEKLY NEW EPISODES
  • Aug 25, 2026LATEST

POPULARITY

20192020202120222023202420252026


Best podcasts about nvidia gpus

Latest podcast episodes about nvidia gpus

GREY Journal Daily News Podcast
What Would Lambda's 3 Billion Pre IPO Round Signal?

GREY Journal Daily News Podcast

Play Episode Listen Later Aug 25, 2026 1:16


Bloomberg reported that AI cloud provider Lambda is in talks to raise a $3 billion pre-IPO round. The financing would precede a potential public listing and reflects the capital intensity of building GPU cloud infrastructure. Lambda provides NVIDIA GPU access in the cloud and sells on-premises systems, serving both on-demand and reserved capacity needs. Pre-IPO capital in this sector is often used to secure chips, expand colocation capacity, and improve networking. Competitors include CoreWeave and large clouds such as AWS, Microsoft Azure, Google Cloud, and Oracle Cloud. Founders are evaluating multi-year reservations, vendor diversification, and transparent pricing as they plan AI compute budgets.Learn more on this news by visiting us at: https://greyjournal.net/news/ Hosted on Acast. See acast.com/privacy for more information.

GREY Journal Daily News Podcast
Will Crusoe's IPO Talks Reset AI Data Center Economics?

GREY Journal Daily News Podcast

Play Episode Listen Later Aug 18, 2026 1:06


Axios reported that Crusoe is in talks about a potential IPO as AI data center demand strains power supplies. Crusoe, founded by CEO Chase Lochmiller and president Cully Cavness, builds modular data centers powered by stranded energy at oilfields. The company raised a $350 million Series C in 2022 and launched Crusoe Cloud to offer Nvidia GPU compute. EPA methane regulations finalized in December 2023 increase incentives for flare mitigation partnerships that can supply Crusoe with power inputs. Competitors such as CoreWeave, Equinix, Digital Realty, Applied Digital, and Iris Energy are expanding capacity through debt, construction, or pivots to AI hosting. Investors would assess site counts, megawatts deployed, utilization, GPU procurement, capex per megawatt, and power sourcing economics if Crusoe files an S-1.Learn more on this news by visiting us at: https://greyjournal.net/news/ Hosted on Acast. See acast.com/privacy for more information.

GREY Journal Daily News Podcast
Will A Singapore Data Center IPO Test US Investor Appetite?

GREY Journal Daily News Podcast

Play Episode Listen Later Aug 11, 2026 1:11


Bloomberg reported that a Singapore-based data center operator confidentially filed for a US IPO targeting about $5 billion. The confidential process allows initial SEC review before public disclosure. A deal of this size would likely list on NYSE or Nasdaq and include multiple bulge-bracket underwriters. Data center demand from Amazon Web Services, Microsoft Azure, and Google Cloud, as well as AI workloads using Nvidia GPUs, is driving higher-density builds and new cooling investments. Singapore's policy shifts since 2019 have steered some development to Johor and Batam. Public comparables include Equinix, Digital Realty Trust, and GDS Holdings, while Chindata was taken private by Bain Capital. Proceeds would likely fund new capacity, power connections, acquisitions, and debt refinancing, with investor focus on contracts, power sourcing, and execution discipline.Learn more on this news by visiting us at: https://greyjournal.net/news/ Hosted on Acast. See acast.com/privacy for more information.

TechLinked
OpenAI's Donut Shaped Smart Speaker Leaks, Nvidia Sells RTX 50 GPUs at MSRP, CXMT Rejects Apple's Cheap RAM Bid + more!

TechLinked

Play Episode Listen Later Aug 8, 2026 9:09


News Sources: https://lmg.gg/1Bjqb Timestamps: 0:00 OpenAI smart speaker leak 1:05 Nvidia GPUs for MSRP at QuakeCon 2:23 Apple's iPhone 18 RAM problem 4:19 QUICK BITS INTRO 4:27 Windscribe deGUID script 5:16 Kimi K3 escapes its sandbox 5:58 AI answering 911 calls in New Orleans 6:40 Nashville data center vs the zoo 7:25 ChatTJFB 8:04 Credits Learn more about your ad choices. Visit megaphone.fm/adchoices

Video Games 2 the MAX
Beasts of Reincarnation Impressions, HALO Studios Already Has Layoffs # 503

Video Games 2 the MAX

Play Episode Listen Later Aug 6, 2026 51:28


Another week in video games, and it feels like more of the same with XBOX introducing more backwards compatibility, this time with XBOX 360 being next in line. Along with the possibility of their own version of the platinum trophy after so many years of requests. Asha Sharma also sets out a plan of how XBOX will get back to where it needs to be by 2027 to 2028, but perhaps the most frightening thing is that hot off the heels of the release of HALO: Campaign Evolved's release, there are already layoffs at HALO Studios as the game is doing "fine" but not what you'd expect a major first-party to do for XBOX. Finally, someone from Sony responds to the controversy of deciding to end disc production in 2028 and the physical media enjoyers won't be happy with the Sony CFO's statement essentially stating that they have no other choice and they now have to figure out how to give their best with digital media as the focus. Does Marc still believe Sony will change course? Plus, Sean's been playing Beasts of Reincarnation; Dave Batista may be Kratos in the God of War TV Series, EA completes its 55 Billion sale to the Saudi PIF, Capcom reports 93% of its sales are digital. PC Motherboards and NVIDIA GPU's are going up again soon too. You can also watch this episode in video form on the W2M Network Youtube Channel, please give us a like, comment on the episode, and give the channel a subscribe and follow as well: https://youtube.com/live/IzgrMgWA19U

VC10X - Venture Capital Podcast
VC10X Pulse - Nvidia v/s AMD: Who is winning the semiconductor race?

VC10X - Venture Capital Podcast

Play Episode Listen Later Aug 6, 2026 2:56


This week, one tweet from Elon Musk reignited one of the biggest debates in AI investing.After AMD reported strong earnings, Musk posted that SpaceX has committed to using Nvidia GPUs exclusively because they are the best.At the same time, xAI continues to deploy both Nvidia and AMD GPUs—raising an important question for investors.Is Nvidia's lead in AI becoming even stronger, or is the AI infrastructure market simply becoming large enough for multiple winners?In this episode, we break down what AMD's earnings and Musk's comments tell us about the competitive landscape in AI chips.⭐ Sponsored by Podcast10x - Podcasting agency for VCs - https://podcast10x.comKey topics we explore:– What AMD's latest earnings reveal about AI accelerator demand– Why Elon Musk said SpaceX will exclusively use Nvidia GPUs– Why xAI is taking a different approach by deploying both Nvidia and AMD– Nvidia's competitive moat beyond hardware: CUDA, networking, and software– Whether AMD needs to beat Nvidia—or simply capture a growing share of the AI market– What this means for the broader AI infrastructure investment thesisThe bigger question:Is the AI accelerator market a winner-takes-all industry, or will explosive AI demand create room for multiple winners?For investors, understanding where Nvidia's moat remains strongest—and where AMD is making meaningful progress—is key to evaluating the next phase of the AI infrastructure buildout.LINKSPrashant Choubey - ⁠https://www.linkedin.com/in/choubeysahab⁠Subscribe to VC10X newsletter - ⁠https://vc10x.beehiiv.com⁠Subscribe on YouTube - ⁠https://youtube.com/@VC10X⁠Subscribe on Apple Podcasts - ⁠https://podcasts.apple.com/us/podcast/vc10x-investing-venture-capital-asset-management-private/id1632806986⁠Subscribe on Spotify - ⁠https://open.spotify.com/show/7F7KEhXNhTx1bKTBFgzv3k?si=WgQ4ozMiQJ-6nowj6wBgqQ⁠VC10X website - ⁠https://vc10x.com⁠For sponsorship queries reach out to prashantchoubey3@gmail.comThis channel is for asset managers, allocators, and investors who want analysis that holds up—not headlines dressed as insight.Subscribe for weekly data-driven breakdowns of the forces reshaping capital markets.#Nvidia #AMD #AI #ArtificialIntelligence #GPUs #ElonMusk #SpaceX #xAI #Semiconductors #TechStocks #Investing #VC10X #DataCenters #CUDA #WallStreet #Finance #ChipStocks #AIInfrastructure #Markets #Earnings

Latent Space: The AI Engineer Podcast — CodeGen, Agents, Computer Vision, Data Science, AI UX and all things Software 3.0

Watch the full episode on YouTube:We first covered Baseten last year when DeepSeek mania was at peak hype. Now they have raised a monster $13B round and become one of the new cohort of AI Infra decacorns that are (with Nvidia, Intel, and the semis complex) chief beneficiaries of the Inference Inflection. We return to Baseten at the peak of the 2026 edition of Open Weights debate. Ali has published a viral breakdown of Kimi K3:And since you last saw him, Philip has spoken at AI Engineer and written the definitive book on Inference Engineering spotted all over SF:Three years ago, inference engineering barely existed as a category.Today, it is one of the most critical disciplines in AI. Inference engineering inherently tackles a different question than standard model training: “How do you turn those weights from training into a product that is fast, reliable, and affordable at scale?” Focusing on these creates an entirely new optimization problem.In one recent GLM-5.2 experiment, quantizing more of the model actually preserved its benchmark quality while increasing throughput by 20%, because the errors introduced in different layers could cancel each other out.Inference is no longer just the final step after training. It is becoming its own engineering discipline, with its own research problems, infrastructure, and increasingly specialized roles.In this episode, Baseten's Philip Kiely and Ali Taha join swyx and Vibhu to explain what actually happens after a new open model is released and what it takes to turn “we generated a token” into a fast, reliable, production-ready API.We go deep on cache-aware routing, disaggregated prefill and decode, quantization, speculative decoding, KV-cache movement, model parallelism, GPU kernels, and the race to make frontier models up to 10× faster. Philip and Ali explain why inference optimizations can still produce gains of 20%, 100%, or even 200%; how quantization errors can cancel one another out; why identical weights can behave differently across clusters; and how Baseten grafted a Kimi vision encoder onto GLM-5.2 without changing the underlying language model.The conversation then expands beyond LLMs into NVIDIA Dynamo, mega kernels, Rubin, AI-specific chips, local inference, video generation, diffusion versus autoregressive models, and the enormous compute barrier to generating coherent long-form video. Finally, we explore the convergence of training and inference, continual learning through persistent KV cache, and the emerging loop where models help optimize the infrastructure that runs them.We discuss:* What happens when a 200,000-token request enters an inference system* Cache-aware routing and reusing previously computed KV cache* Why prefill and decode are increasingly handled by different GPUs* When dedicated deployments become cheaper and more reliable than shared APIs* How speculative decoding uses a smaller model to accelerate a larger one* Tool calling, structured outputs, and what LLMs actually do* What it takes to support a new open model on day zero* Grafting Kimi's vision encoder onto GLM-5.2* Retrofitting inefficient model layers with components from other architectures* Why models sometimes collapse into repeating the same token* How hardware, kernels, and race conditions create nondeterministic failures* Preserving model fidelity while making inference faster* How quantization errors can cancel each other out* Why inference optimizations still deliver gains of 20%, 100%, and 200%* How optimized serving can make a model up to 10× faster* NVIDIA Dynamo, KV-aware routing, and distributed model serving* Speculative decoding the speculative decoder* Why local AI is about making models less dumb while data-center AI is about making them less slow* Tensor, expert, and pipeline parallelism across GPUs* Hardware-aware model design, auto-tuning, and the case against mega kernels* Rubin and why inference is becoming a systems problem* Whether modern GPUs are evolving into programmable AI ASICs* Why enormous models like Kimi K3 require GB300-class hardware* Why open-source video generation still trails Veo, Kling, and other closed models* The quadratic attention bottleneck behind long-form AI video* Autoregressive video, real-time generation, and compounding quality drift* Why future video systems may combine autoregressive and diffusion architectures* Training for inference and inference for training* Continuous post-training, deployment, evaluation, and improvement loops* How GLM-5.2 helped optimize the kernels serving GLM-5.2 itself* Why faster networking could unlock dramatically faster decoding* Continual learning, KV-cache compaction, and persistent model memoryShow Notes* How to build a day-0 API for Kimi K3* 22580: From GPT2 to Kimi3, ExplainedPhilip Kiely* LinkedIn: https://www.linkedin.com/in/philipkiely* X: https://x.com/philipkiely* Inference Engineering: https://www.baseten.co/inference-engineering/Ali Taha* LinkedIn: https://www.linkedin.com/in/aliestaha/* X: https://x.com/waterloointernTimestamps00:00:00 Introduction and the 200K-Token Prompt00:03:18 Dedicated Deployments, Speculative Decoding, and Tool Calling00:11:26 Launching Production-Ready Open Models00:19:06 Model Retrofits, Failure Modes, and Nondeterminism00:28:22 Quantization and Canceling Errors00:32:15 The Race to 10× Faster Inference00:40:48 Dynamo, Speculation, and Local vs. Data-Center AI00:50:18 Model Parallelism, Auto-Tuning, and Mega Kernels01:00:55 Rubin, GPUs vs. ASICs, and Custom AI Chips01:10:03 Giant Models and the Limits of GPU Memory01:12:42 AI Video, Quadratic Attention, and Autoregressive Generation01:21:47 Audio, Images, and Diffusion Models01:27:32 Training, Self-Optimizing Models, and Continual Learning01:40:06 Closing ThoughtsTranscriptIntroduction: Baseten, Waterloo Intern, and Inference EngineeringSwyx [00:00:00]: Okay, we're here in the studio with Philip, old friend from Inference Engineering, the book, as well as Baseten and everything that you've done, you and I have done before, as well as Ali. Welcome.Ali [00:00:15]: Pleasure to meet you.Swyx [00:00:15]: Waterloo intern.Ali [00:00:16]: Waterloo intern, always.Swyx [00:00:17]: When did you get “Waterloo intern” as a handle?Ali [00:00:19]: As a handle? Oh.Ali [00:00:20]: I think the rebranding happened mid-March. When I saw it was open, I was like, “I have to take it. Up for grabs.”Philip [00:00:26]: The problem is that Ali is really good at his job and is not gonna be an intern much longer.Philip [00:00:30]: So we have to figure out who's gonna get the handle.Ali [00:00:33]: Well, I'll pass the torch over to the next intern.Swyx [00:00:34]: Oh, okay. It can be, like, you just pass it to another Waterloo grad.Ali [00:00:37]: To another Waterloo intern. No, bruh.Philip [00:00:39]: Yeah.Ali [00:00:39]: Intern.Swyx [00:00:40]: Intern, yeah.Ali [00:00:40]: And no.Philip [00:00:41]: You gotta get an intern from Waterloo.Ali [00:00:42]: Yeah, I've gotta get an intern from Waterloo.Swyx [00:00:44]: Right.Ali [00:00:44]: But they have to follow the path.Swyx [00:00:45]: Oh, it could, but it could come from Baseten, so it's like whoever Baseten gets from Waterloo.Ali [00:00:48]: Right.Swyx [00:00:49]: Has the title of Waterloo.Ali [00:00:50]: It stays in the ecosystem.Philip [00:00:51]: Exactly.Ali [00:00:52]: Halfway through the internship, you either get it or you're out.Philip [00:00:55]: You should also do, like, a big graduation ceremony where you change the handle.Ali [00:00:59]: Just say it.Philip [00:00:59]: For everybody.Swyx [00:01:00]: You guys are good at ceremonies, clearly. We had a nice launch of the book, very successful. But before we get into all that, I wanna start off with a fun question for you. Okay, you're an expert inference engineer. What happens when I send a long query, say two hundred thousand tokens into Baseten's inference? What's the process of query through GPU model routing, balancing, all that? What is all the stuff that we don't think about?Long Context Requests, KV Cache, and Cache-Aware RoutingPhilip [00:01:26]: With a long query specifically, the first thing that I'm gonna ask is, “Have you sent me this query before, or at least part of it?” and I really hope you have, because it's gonna be a lot easier for me and a lot cheaper for you. So the first thing that we're gonna look at is some cache-aware routing, where we're going to see, we probably have a number of instances, a number of replicas up serving whatever model you're hitting. We want to send this one to something with, number one, available prefill workers, and number two, ideally some cached input already there so that we can skip prefill on at least part of these two hundred thousand tokens. If you're doing two hundred thousand tokens, it's probably coding or a multi-turn agent or something where you would expect to have that cached. If you don't, we're gonna have to send it to a prefill worker. We've at least on certain models disaggregated prefill and decode, so you're going to have one set of GPUs that's solely going to process the input, create the KV cache, and get you your first token, and then that's going to be passed over to a separate set of GPUs, which is going to run decode. We're going to iteratively make those tokens. We're probably going to have some speculator model in front of that. I'm going to assume that you're doing coding, and because of that, our speculator model, which assumes you're doing coding, is gonna have a high draft token acceptance rate. If I'm wrong and you're asking me to summarize every Harry Potter book, it's gonna be slower. And then we stream that output to you and account for it, charge you, a couple of pennies and say, “Hey, would you like to send another one?”Swyx [00:03:04]: Except Baseten doesn't charge by pennies.Philip [00:03:07]: Well, yeah, we charge. I'm assuming that we're talking about the public model APIs. If you are setting up a dedicated deployment, then yeah, it's not pennies.Public APIs vs. Dedicated DeploymentsSwyx [00:03:18]: Yeah, one of the key differentiators when I was talking with Baseten initially was that people who want very high volume just need to rent by the box, ‘cause then it's up to you to figure out how to saturate the box.Ali [00:03:31]: And more often than not, it's, like, way cheaper if you're pushing, like, millions of tokens per hour, if you just pay per hour instead of pay per token.Philip [00:03:37]: Yeah, they do. I think that we've increasingly seen a lot of demand for the pay per token APIs, just because everyone wants to try open models, and then once they find a use case that's really sticky, then they move over to dedicated.Swyx [00:03:51]: Is there a best practice on when it's time to swap over?Philip [00:03:54]: Couple reasons. Yeah, reliability, that's a big one, right?Ali [00:03:57]: Like, if they have a very specific use case, they want you to train something specifically for them, like they want their own spec dec, for instance, for their own traffic.Swyx [00:04:04]: Spec dec is speculative decoding.Speculative Decoding and Custom SpeculatorsAli [00:04:05]: Speculative decoding, yeah.Swyx [00:04:07]: You have to explain.Ali [00:04:07]: Sorry. Like, speculative decoding is like, if you have a huge model, right? And so the model is going to be generating one token at a time every single turn, every single forward pass. So we attach, like, this little, like, parasite, like this layer that goes on top of the model, and this model just has to predict. It does three very fast autoregressive forward passes, and it will predict, like, three certain tokens, and then you do one forward stage over the entire original model in order to see if those predictions were correct or not, and then you accept them or you reject them. Now, this draft model is traffic specific, so if you, like, Philip said, if you're summarizing Harry Potter books, I can train exclusively that draft model on Harry Potter books, and I can guarantee you that I'm gonna accept the three tokens every single time. And so with that case, I increase your decode speed. I wouldn't be able to provide this to you if you're a shared endpointSwyx [00:04:53]: YeahAli [00:04:53]: ‘cause I have no idea if you're doing Harry Potter, if you're doing coding, if you're doing English. We don't know. Also, there was a thing in the book that mentioned that if they really cared about a specific threshold, chapter four, I think. Do you remember that?Philip [00:05:06]: Yeah. The things that you can do is you can set a specific, like, batch sizing, a specific, like, parallelism strategy if you're trying to optimize for, like, throughput versus latency. You can. Maybe a NVFP4 quant doesn't pass your benchmarks and you wanna run a model at higher precision, you could do that. There's just a bunch of reasons why you might wanna have your own endpoint and the biggest one, of course, just being, like, you don't have to deal with someone else doing a hundred million tokens of benchmarking traffic at the endpoint when you happen to be trying to serve your users.Swyx [00:05:40]: Yeah. I think one thing that is. That is a classic journey. Like, it's people is asking the, what happens when you type Google into the browser. Tool calling, is that just, you're generating JSON or is there more complication beyond that?Tool Calling, JSON, and Structured OutputsAli [00:05:58]: Certain customers that we have, they have their own post-trained models, and so they demand a tool calling that's not just, like parse a file or go find the weather. It's something that's very specific and you have to do post-training on this. And if the post-training on the model is not good or if the quantization after the post-training to get the inference to be fast, the model will struggle reading the JSON file and reading the tool calling. But it doesn't require its own like sandbox. It's not like it's going to use that tool calling to like escape a sandbox or like it doesn't have to be contained. It can just be a normal dedicated deployment. The challenge with tool calling more and more seems to be that the companies want certain tool calling which is a very sensitive thing to train. And because you're dealing with all of the JSON outputs, if it doesn't like close the end of the request in a very certain manner, you end up with a model that did the tool calling and like the thinking and so as a result of that, it didn't see the result and just hallucinated the result as it decoded. That seems to be the most challenging thing with tool calling, not really the sandboxes model.Philip [00:06:56]: Yeah, that's a challenge on the training side and then on the inference side, there's work that you can do to scope the possible output. So we published this at this point close to two years ago, the solution to this problem which is you make a state machine and you use that to constrain the output to a specific format. So this is the structured output problem. If you remember backSwyx [00:07:27]: Yeah, the specific grammar is,Philip [00:07:29]: Yeah, exactlySwyx [00:07:30]: GML had this thing.Philip [00:07:31]: Yeah. So it's like the old-school “make sure this is only JSON”, return only JSON orSwyx [00:07:38]: YeahPhilip [00:07:38]: Grandma's gonna die type of prompts.Swyx [00:07:39]: Is it BNF grammar? At some point OpenAI had released a thing that was like, yeah, if you want to constrain your output, write BNF grammar, back as NOR.Philip [00:07:47]: In our inference system, it's just a specified output format. And you get the guarantee that your output's gonna be structured along that format. And so applying that to tool calls can like help cut down on. You can still call the wrong tool or call no tool. It doesn't solve the certainty problem but it at least solves the output structuring problemSwyx [00:08:10]: YeahPhilip [00:08:10]: Within tool calls.Swyx [00:08:12]: And MCP is just another form of tool, right.Philip [00:08:14]: Yeah, exactly.Swyx [00:08:15]: As far as there's no special thing there.Philip [00:08:16]: The thing I'm always like explaining to people is the LLM is not capable of doing anything. It's only capable of making suggestions of what to do and then if those suggestions are formatted in a certain way and applied to a system that knows what to do with them, then an action occurs.Swyx [00:08:32]: Yeah. Part of the fun stuff is, this is solved outside of tool calling too. Like in an agent loop if the output is not correct or you're right, like reasoning, tool calling was done in the reasoning trace, just be like, “Oh, I don't know what to do. Let me just try again.” And it might get there after a few tries. And on your point of training, sometimes this is harder in smaller models, so you don't have the same exact quality outputAli [00:08:56]: Right.Swyx [00:08:57]: When you just swap from a big model, right?Ali [00:08:59]: Yeah. I will say that, before, I think we need to go back to inference engineering proper.Ali [00:09:04]: But, I had expected that something would replace JSON because it's hard to stream JSON ‘cause JSON must be complete and you must have open and close brackets and everything. So it's hard to parse something or validate something while it's being streamed. So people invented all sorts of things that are like, I forget the name of some of these alternatives, but it's something like TOML, something like YAML. But JSON seems to be dominant still.Philip [00:09:30]: The JSON outputs aren't that long, right? Like you could have a long-- ‘cause tool calls also contain the arguments in them and perhaps for a certain tool you might pass like a very long argument. But my impression of the median tool call is that it's a relatively small number of tokens, right? So I would expect that speculators are generally fairly good at something as formatted as JSON. And so you would have like a pretty fast decode step there and that the streaming wouldn't be as valuable, but maybe I'm wrong about that.Ali [00:10:02]: I think you're also bounded by the software or that the model is gonna integrate with if the software is built with JSON for the tool calls or if the company that you'- if your customer says that this is how our software works and our tools are interfaced with JSON, you can ask them to like, change their software and say like, “Yeah, this is gonna be better for the model.” but like with the right training shouldn't be that much of a difference. Also more profitable if it outputs more tokens probably.Swyx [00:10:25]: Depends on your business model.Swyx [00:10:27]: It really depends. But I will say that, as a writer with like experience a lot with generated output, I do try to move from text to JSON text which is very long JSON, right? Like there's paragraphs in every field because I'm trying to structure it, right?Philip [00:10:44]: Right.Swyx [00:10:44]: I want you to first make factual statements, then make opinions then make bullet point summaries, have dates, have entity references have your sources for references, all these things. Anyway, so these are things that like I think people who really experiment with structural output have to really care about. But, let's, let's recurse up the stack a little bit. Before we started recording, you mentioned something really cool, which is that there's a lot of engineering that-- inference engineering that goes on when a new model provider releases a new model, right? So let's call it GLM-5.2, Kimi K3. I had previously assumed, especially if it's like, well, GLM 5 to 5.1 to GLM-5.2, like that you've supported them before. Is it that much work?What It Takes to Support a New Open ModelAli [00:11:26]: It's a lot of work.Swyx [00:11:28]: Yeah. Okay. So like, a lot of people, all you guys, right whenever a new model launch like, people rush to say like, “Oh, Hugging Face supports this, Fireworks supports this, Spacetime supports this,” and I'm like, “Yeah, of course we support it.” But what goes into that? What goes intoPhilip [00:11:40]: I think it's more than just support it too, right? It benefits the consumer a lot. Like I think it was with Kimi K2.5 or GLM-5.2 the latest, there was an inference war, right? X provider is at 90 tokens a second. The next day we're at 150. The nextSwyx [00:11:55]: I kinda kicked that off with the GLM-5.2.Swyx [00:11:58]: I wrote a Twitter article about. It got like half a million views,Ali [00:12:02]: Based on being numberSwyx [00:12:03]: YeahAli [00:12:04]: Or it's for something else.Swyx [00:12:05]: Yeah. Which,Ali [00:12:06]: Oh my GodSwyx [00:12:07]: Which then got everyone really excited about, hey, how can we, bend tracks a little bit further and,Philip [00:12:14]: There's a difference between support the model, as in I can make a token out of this model, and support a model, as in I have a production-ready API from this model.Philip [00:12:26]: Getting to the point of I can make a token out of this model is not that hard because generally the, open source inference engines, vLLM, SGLang of the world oftentimes even receive weights ahead of time, maintainers do, or the people making the model merge PRs to ensure support. So you generally can, just get it working on the standard open source stack without too much pain in most cases. The challenge is, every inference company is gonna have own proprietary stack. Some open source components, some in-house stuff. And for any arbitrary model, there's going to be some new stuff. Sometimes you get lucky, like K, two five to two six was, like, pretty similar.Quantization, Speculators, and Production ReadinessAli [00:13:16]: Yeah. It was pure continued post-trainingPhilip [00:13:18]: YeahAli [00:13:18]: If I remember correctly.Philip [00:13:19]: Even in those cases, there's still stuff you have to do. You have to redo the quantization work. You're taking the model from. Generally, these models are not released in NVFP4, and we want them to be in NVFP4 for maximum Blackwell compatibility. So we have to perform that quantization, and, calibrate the quantization to make sure that we're not causing any regression in the model's intelligence. And then we also have to train the speculator, as we've talked about. Generally, we have. We have ZDR, zero data retention on our model APIs, so we don't know exactly the traffic that people are sending us, but we know what's popular. We know that coding use cases are popular. We know that agents, agentic use cases are popular. So we can get public data sets that are representative of that traffic and train general speculators. Now, with speculators today, you need to train the speculator using the base model itself because you're getting hidden states out of the model from running inference on these specific prompts, and that is the training data you use to create the speculator. So there's that process which you need the real model weights for. And then there's of course just the process of, standing up all the infrastructure behind it, loading all this stuff, testing it. And then when there's a new model with a newer architecture, I think that, like, the DeepSeek models tend to be the most challenging as they have, like, the most novel architectural stuff going on, model after model. But every new model has something. Kimi K2 had. Oh, sorry, GLM-5.2 hadAli [00:14:53]: Sparse attention.Philip [00:14:54]: Yeah,Ali [00:14:54]: YeahPhilip [00:14:54]: the DSA.Ali [00:14:55]: Right. Which is brought from DeepSeek.Philip [00:14:57]: Yeah. AndAli [00:14:59]: So you can copy-paste then?Philip [00:15:01]: It kindAli [00:15:01]: I don't know how this works.Philip [00:15:02]: So, like we had to, like, build support for that into our runtime. And you're right, like it is really interesting the way that all of these open source labs borrow from each other. For example, like GLM-5.2 doesn't have vision. So something that, Haley, a guy on our team, if we could take a look at this, he, like, grafted the Kimi vision encoder onto GLM-5.2.Retrofitting Vision into GLM-5.2Ali [00:15:27]: We'll be training the projector.Philip [00:15:28]: Exactly. So if you think about, like, the encoder, there's the encoder, which is the part that looks at the image and turns it into latent information, and then there's the projector which likeAli [00:15:38]: You can say latent space. It's okay.Philip [00:15:41]: And then there's the projector that maps it onto, the model itself, and then there's the model weights. You don't wanna mess with the model weights because you run a chance of making the model dumber at something else for the purpose of giving it vision. So instead, Haley started with just a projector, which is only a handful of millions of parameters.Ali [00:16:02]: That would be, yeah.Philip [00:16:02]: Yeah.Ali [00:16:03]: Can you show the training one?Ali [00:16:04]: Like the way it groksPhilip [00:16:05]: YeahAli [00:16:06]: Very interesting.Philip [00:16:06]: And maybeAli [00:16:07]: That right therePhilip [00:16:07]: Maybe Ali, you should take it from here. You've got a betterAli [00:16:10]: Ooh, double the sandPhilip [00:16:11]: Understanding of this than I do.Ali [00:16:11]: Yeah. You can see, like, he. The way he trained this is really cool. At the beginning, he was training it using just like, “Here's a picture of a mountain. Can you describe what's in this mountain?” And that caused it just like the first, learning walls. Like here you can see this all we're trying to teach it is to translate the encoded. Like it's already taken the encoder from Kimi K. It's taken the image. It'Philip [00:16:31]: Yeah. FrozenAli [00:16:31]: FrozenPhilip [00:16:32]: With adapter.Ali [00:16:32]: Exactly.Philip [00:16:33]: Yeah.Ali [00:16:33]: So the brain is frozen and the eyes are frozen. It's just we're tryingPhilip [00:16:37]: AlignAli [00:16:38]: Interconnect between the eye and the brain, right? So the projector. And so you take the tokens and then he's like, “Oh, can you describe what's in this image?” And he's like, “Oh, it's a mountain,” or it's a person or it's a human, whatever the case is. But that didn't cause complete understanding. So he changed it such that every image was associated with a data set of questions. Like, does this image have a white male? Does this image have birds in the top corner? Does this image have a scientist in it? All of that stuff. And it would have to answer questions correctly. And using not just training on describing an image, but being able to answer question, another question, answer over time. Like you can see the grokking, which is like genuinely insane, that retrofitting vision into a large LLM can learn to that extent. And even for images that it doesn't perform well on, for instance, if you ask it a picture of like Stephen Hawking, “Who is this?” Maybe it doesn't get it, but it will say something like, “This is Albert Einstein.” Like it still understandsPhilip [00:17:25]: Close enoughAli [00:17:26]: That this is a scientist who is a man who has, some significant achievements, all that stuff. So that's like really cool.Philip [00:17:32]: Yeah. So, we've covered Hao Tian before, who the author of the LLaVA paper that did this, a while ago. And I think that's very foundational work for anyone who hasn't done vision work before.Ali [00:17:41]: Same with the CLIP and MetaCLIP, where you go from just captioning to building out questionsPhilip [00:17:47]: RightAli [00:17:47]: Off the image and how much better you can get performance.Philip [00:17:50]: Right. Right. Right. Yeah. But what's, what's so exciting about this is if you look at a model like this. Now, this is a little bit more of a research project. It's not. It got to 56% on MMLU Pro, I think. So not quite frontier. But if you're running this model, you haven't suffered any loss on your GLM-5.2 quality. If you don't have an image, it'll just behave exactly the way it used to. And ultimatelyAli [00:18:14]: Which in the inference code you literally do not include the other part, right?Philip [00:18:18]: Yeah. You would just skip the encoder if you don't have an image input.Ali [00:18:22]: Okay.Philip [00:18:22]: Just confirming.Philip [00:18:23]: YeahAli [00:18:23]: Does it affect a lot on the overall inference side? Like you're not adding much, you're adding a very small vision encoder. These are typically likePhilip [00:18:30]: They're super fineAli [00:18:31]: Less than a billion parameters, right?Philip [00:18:32]: Yeah. It's, - There's a little bit less standardization among vision encodersSwyx [00:18:37]: YeahPhilip [00:18:37]: So the support matrix can be a little bit, sparser. But overall, yeah, it's a pretty, it's a pretty minor component of the overall system. And ultimately what you get out of the system is all of a sudden you have Kimi Vision, GLM weights, and DeepSeek attention all in one model.Open Source Model Grafting and Franken-MergesPhilip [00:18:56]: And that's, I think, a lot of the power and beauty of open source, is that you can take all of these different components and combine them together into a system that's better than anyoneSwyx [00:19:05]: YeahPhilip [00:19:05]: Can be individually.Swyx [00:19:06]: People used to say that you would also do Franken-merges where you would take likePhilip [00:19:10]: YeahSwyx [00:19:10]: Layers from each model.Swyx [00:19:11]: Does anyone do that anymore?Ali [00:19:13]: Well, to your point previously when you were mentioning like, the work that goes into supporting a model when it first comes out, like GLM-5.2 or MiniMax M3 or whatever the case is. Sometimes you do have to like, you do have to switch out some things. Like, for instance, the MiniMax M3 head uses full attention, and with full attention you end up with this like insane bottleneck in spec dec ‘cause you're doing auto-regressive token generation for three tokens, and you're doing this like N squared over all of the tokens that are in your sequence. Your KV cache is like very large because it's not sparse, it's not top K. So we find it better to like, okay, we're gonna replace this, we're gonna replace this layer with a layer from another model that's using like GQA, for instance. And then just with the right training, you can get it to have the same acceptance rate. So it is very possible to retrofit layers from other models and very much needed. If a layer is like inefficient, the training just becomes the challenge, like how do you ensure that you train it properly? Which again to your earlier point is like the mesh between training and inference. As in like you need very good training in order to do fast inference. That's like, I feel like more and more becoming true.Swyx [00:20:21]: Yeah. Anything else on the support side when you say like get it to fully production ready?Loop Detection, Race Conditions, and Non-DeterminismPhilip [00:20:26]: Yeah. I think that there's also a question of just, we can test a model to a pretty extensive degree, but we're trying to get it out quickly and then you see a bunch of other people test it and you get interesting results. There was an issue with, GLM briefly where we had some like mode collapses where it would just output the same token over and over again for certain prompts on certain temperatures. Like once you expose an endpoint to the real world, there's going to be, so many more varieties of things given to it that you're able to, discover and patch things. So it's not just a, day zero process, it's then like for the first week, for the first month, if a model remains popular, like how do you both fix bugs and then continue to push the envelope on performance?Ali [00:21:21]: What do you mean you don't want your model outputting S?Swyx [00:21:24]: Is there loop detection on that stuff, by the way? It still happens like quite a lot, which is surprising.Ali [00:21:30]: We have like we, in our endpoint, like if a model was to output the same token like four plus times, we just cut the generation. We say like, “Oh, sorry, this-- Like try again,” or like we will reprocess the request. ‘Cause we know then, like if it, like if, yeah, it's four times the same token, it's probably collapsed.Swyx [00:21:45]: Yeah. Is there a way to opt out in case I really want that?Ali [00:21:48]: You want that?Ali [00:21:50]: I think there's a way that we have to handle it. I'm not exactly certain, but I feel like in certain models, like when they output something like you can imagine, like a table for instance, and so they want, they wanna draw like 12 dashes and 12 dashes. Yeah, I think there's a way for that to happen. I think we only do it on certain tokens. Like we exclude certain special characters.Swyx [00:22:07]: Yeah.Ali [00:22:07]: So we only do it on like certain like S is the most common almost. GLM-5.2Swyx [00:22:11]: OhAli [00:22:11]: And I think it was DSV 4 as well. Like you'd just have like looping issues where like you literallySwyx [00:22:17]: ItAli [00:22:17]: Just have like S.Swyx [00:22:18]: Yeah. Is there a special, something special about S? No, just randomlyAli [00:22:21]: It just seems to be the one token involved.Swyx [00:22:23]: Yeah. And it'Philip [00:22:24]: Is thereSwyx [00:22:24]: And it's only temperature 0Ali [00:22:27]: NoSwyx [00:22:27]: Even at other temperaturesAli [00:22:27]: Even at like 0.9 or whatever, it will still, it will still collapse.Swyx [00:22:30]: That's weird, right?Ali [00:22:30]: It's, it is an inference problem to be honest, like a software problem. Like oftentimes, the image you run will-- like NVIDIA will release an image for instance, and if we will upstream the changes from their latest TensorRT-LLM image into our stack, we'll find that it fixes it. Or oftentimes this will only happen in an inference engine that you're using like SGLang. But if you were to switch to vLLM, that isn't the case. So it seems to be like an extremely like deterministic software issue and not really a model issue. It's not like a weights problem. Like I'- we'll say like, “Oh, it's a problem with the quant. We did PTQ wrong,” right? But that isn't, that doesn't make sense because the same weights used with a different inference engine does not repeat the problem. And sometimes it's, the kernels that are being used in the backend have like these very subtle sometimes race conditions, where if you were to use this model hosted on one cluster, you will never get this problem.Swyx [00:23:19]: Oh my God.Ali [00:23:19]: But if you host it on a different cluster, you will. And the reason is the KV cache transfer from a node to node in that one cluster is using a slower interconnect than the node to node in another cluster. So that exposes the race, whereas in another cluster it doesn't. So then you end up just like, okay, this model is not gonna be hosted on this cluster. We're gonna host it on, another cluster because that cluster exposed that problem. But then it ends up with like, okay, is it the software? Is it the model weights or is it the hardware?Swyx [00:23:42]: There is a thing about this with temperature 0 still not being deterministic, right?Ali [00:23:46]: Right.Swyx [00:23:46]: Mostly because of hardware. Even at temperature 0 same model, you won't always get the same output.Swyx [00:23:52]: Even-- But I'm surprised by the race condition one because, I thought PyTorch was a graph that like guarantees that you at least, execute things in the right order.Ali [00:24:02]: Well, yeah, true. Like I'm not, I'm not saying that there is. Like well, you have things like PTL optimizations where like you can start a kernel before the end of the previous kernel, and that's like ‘cause you want to do that because there'sSwyx [00:24:12]: It's like pipeliningAli [00:24:12]: Expense. Exactly.Swyx [00:24:13]: Yeah.Ali [00:24:13]: But it'- But you don't do it cleanly. Like you overlap a little bit of the execution. No, it is very possible that the kernel itself, like that one block that is supposed to be running in this instance of time, that kernel itself has a race condition. For instance, like a missing barrier. Like often if you're designing a kernel and you want it to make it to be very fast, if you don't test it extensively, you'll, you'll have certain threads access data points from registers before they've been written to by other threadsSwyx [00:24:36]: YeahAli [00:24:36]: For example, because like your barrier is wrong or your synchronization was wrong. But yeah, like the testing itself is very difficult in those like, andSwyx [00:24:42]: And there's no like borrow checkerAli [00:24:45]: What does that mean?Swyx [00:24:46]: Like Rust. Like the. If you're trying to have like memory safety It sounds like a comparable problem.Ali [00:24:52]: Well, yes, but you're working in CUDA, right, NVIDIA GPUs. Like- You just need a higher level language like modular Maybe that's what modular is supposed to do. I don't know.Quantization Quality and Vendor FidelityVibhu [00:25:00]: How do you see keeping quality of the model? So you talked about all these steps of, okay, you gotta do quantization, train your own speculative decoderAli [00:25:07]: RightVibhu [00:25:07]: Run on different hardware. Looking at other model providers, okay, you kicked off a inference speed race on the consumer end. What goes into keeping quality the same across them, right? Sure, you can run benchmarksAli [00:25:22]: YeahVibhu [00:25:22]: But, like, how do you determine how much quantization are there standards? What goes intoPhilip [00:25:27]: There's a few things on quality. Most inference optimizations are lossless. KV caching, for example. You are just recomputing or preventing recomputing the same values. Speculation, of course, if a draft token is wrong, it gets rejected. The main lossy optimization is quantization. And that really comes down to, number one, data format, number two, which parts of the model you choose to quantize, which layers, and number three, like doing a lot of calibration on the quantized weights, to ensure that you're preserving all the outliers. There's other tricks that you can do, though. A big one is long context, ‘cause one thing you asked at, right at the beginning is, “Oh, what's gonna happen if I send a 200,000 token request in?” So with a long input sequence, you need to, store a lot more information. You need to process a lot more tokens. And so even if a model has a context of a certain length, you might, as an inference provider, choose to build an API with a shorter context length, and of course a full length one as well. Because if someone doesn't need the full million token context, for example, you can get them better performance. I don't know if that's exactly like quality of the model. The way that I think about quality is to what degree are we faithfully serving the original model? If you think of a golden implementation of a model that performs exactly the way the model is designed to perform, I think of quality as how close are we getting to that, 100% fidelity of the model.Philip [00:27:13]: You can also, of course, think about quality from the training side and how do you push yourself past 100%. But when I think about purely inference optimizations, it's getting faster while staying as close to that 100% fidelity mark as possible. And certainly our standard internally is that, like you should not be able to tell the difference between our API and a, official API. I think Kimi in particular does a good job of vendor benchmarking hereAli [00:27:41]: YesPhilip [00:27:41]: Where they haveAli [00:27:42]: They released an actual vendor benchmark.Philip [00:27:43]: Exactly, yeah.Ali [00:27:44]: ‘Cause they accused, some people, Amazon? There was some provider that was not doing very well on Kimi's benchmark.Philip [00:27:50]: Yeah.Philip [00:27:51]: So, with Reflect we probablyVibhu [00:27:52]: This was a long time ago, right?Philip [00:27:54]: No.Ali [00:27:54]: Yeah, like threeVibhu [00:27:55]: They alsoAli [00:27:55]: Four, five months agoVibhu [00:27:57]: This also happened with, I don't remember which model, but they pulled out quite a few, and then they started a whole chart about this. It might have beenPhilip [00:28:03]: Kimi Vendor Verifier.Ali [00:28:04]: Yeah.Philip [00:28:05]: Yeah.Ali [00:28:05]: Yeah, ‘cause you, ‘cause you'd be pissed, right? Like if you'Philip [00:28:07]: Yeah.Ali [00:28:07]: If like if I'm a consumer and I'm using like Amazon's endpoint for instance, and I've used Kimi and I'm like, “Oh my God, like this is bad,” I'm not gonna say, “Oh, Amazon quantized the model in a bad way.” I'm gonna say, “Oh, Kimi sucks.” Right?Philip [00:28:17]: Yeah.Ali [00:28:17]: So it seems like that makes sense.Philip [00:28:19]: Yeah, they care. They care.Vibhu [00:28:21]: Justifiably.Ali [00:28:21]: Yeah, justifiably.Vibhu [00:28:22]: This is probably a stupid question, but just checking, has anything improved from main quantization?Philip [00:28:28]: Yeah.Vibhu [00:28:28]: Like, is quantization always strictly worse?Ali [00:28:30]: Well technicallyVibhu [00:28:32]: NoAli [00:28:32]: It's a lossy. QuantizationPhilip [00:28:33]: YeahAli [00:28:33]: Is a lossy, it's a lossy implementation.Philip [00:28:36]: Speed improvesVibhu [00:28:36]: Speed improves.Ali [00:28:37]: It the number, likeVibhu [00:28:38]: No, I' always look for inverse scaling laws.Philip [00:28:40]: Yeah.Ali [00:28:40]: Yeah.Vibhu [00:28:40]: This is something I learned from Noam Brown, where like things that normally act in one direction sometimes do.Philip [00:28:45]: Well, technically when you run a benchmark, because these models are deterministic, sometimes your,Ali [00:28:52]: YeahPhilip [00:28:52]: NVFP4 quant is like, two basis points higher than yourAli [00:28:56]: No, it's noise. It's noise.Philip [00:28:57]: Yeah, exactly. I'm like, yeah, it's, it's within. That's why I always say within margin of error.Philip [00:29:01]: And I stopped saying that because everyone assumes that what is, well, within some margin of error, we're barely inside of that to the worst, so we're saying. But yeah, sometimes it's just like, gives you a higher output score. But like Ali said, that's noise. To my knowledge, you're not necessarily making the results better. You're just trying to, again, like keep your fidelity as close to 100% to the original model.Layer Selection, KL Divergence, and Better QuantizationAli [00:29:27]: There is, to your point, research that we did on MP. I don't know if you are able to pullPhilip [00:29:31]: YeahAli [00:29:32]: A tweet we did. One of our research interns, Joshua, I think it's a tweet on how we have 20% better quantized GLM-5.2 than NVIDIA. Essentially what we found throughout like this month research is, okay, quantization is a lossy. It's. You're compressing the data from, occupying 16 bits to occupying, four bits, for instance. And so you're losing some information, and you're trying to minimize that. And so when I say that I'm gonna quantize the model, my job becomes how do I find the layers that I can quantize, and how to find the layers to not. For instance, with image models, I don't quantize modulation layers, and I don't quantize out projections because those two are. Like out projection is what you see as the user. Modulation is what the model sees or understands. Right, exactly. And so to his paper, do you have the. It doesn't have the. Yeah. It's a long paper. I don't know if I can findVibhu [00:30:25]: If there's a part to search or it's probably in the thread.Ali [00:30:28]: It's probably in the thread.Vibhu [00:30:29]: Yeah.Ali [00:30:29]: But the long and the short is it is very possible that quantizing more of the model makes the results. Like if I have a model that I quantize layers one, five, and 10, and another model where I only quantize layers one and It is possible that the model in which I quantized more information is going to perform better because the quantization errors have canceled out. And so what Joshua showed in his mathematical proof where he had like a verifier in, is that you can predict which layers are going to have quantization errors that will cancel out with each other, and you choose to quantize those layers. And so the result of doing this mathematical quantization is you end up with a model that's 20% more quantized than another provider, so you get 20% more throughput of it because there's more layers than running an NVFP4, and your quality is better than that other quant because the layers that you chose to quantize have their errors cancel out, like one layer skewed to the right one layer skewed to the left, one layer skewed to the right. Your final logits distribution is more similar to the original distribution of the model, so you have better fidelity. And so the way we proved this was with KL divergence. So instead of just scoring on the benchmarks, we scored the KL divergence between the logit distribution of the quantized model and the logit distribution of the original full precision model, and we showed that with this technique we get. If your probability distribution on the logits which token it wants to select is more of the same as the original model, you're probably gonna end up staying true to the original model. So yeah, so it seems like previously before this, it seemed like the industry was, well, the more you quantize, the worse it's gonna be, ‘cause the more loss you introduce. That's not exactly, not necessarily true. So yeah, doesn't improve it, but can cancel out.Philip [00:31:57]: I think it might be this, but reminds me a good bit about pruning where you can prune off certain layers.Philip [00:32:03]: But very interesting. Didn't know this was a whole paper you guys put out.Ali [00:32:06]: It's. Fun fact, it was originally 72 pages, this paper, and then we decidedPhilip [00:32:11]: WowAli [00:32:11]: We can't tell. We couldn't release it. So it's now 45.Swyx [00:32:15]: Still 39 pages, so very substantive. We talked about evals and all these things and, like what's possible in terms of speedup? Like it's like probably like the numberInference Speedups and BenchmarkingSwyx [00:32:25]: Thing that people do wanna care about, and it's something that you wrote about in your post. Like official API is 70 tokens per second, and you push it up to 90. Is that like a normal thing?Philip [00:32:36]: So what's cool about working in inference, the reason that I think inference is going to be a useful place to do engineering for a long time, is that if you look at highly optimized domains like, say, finance, if you're in finance, you measure how much better you got in basis points. It's like, “Oh, I got five basis points better, like twentieth of 1% better,” that's huge news because everything is so optimized. When we publish optimizations, it's 20%, it's 100% it's 200%. So there's still probably like a lot further to go, honestly. Like you'll, you'll know that inference is pretty much solved when researchers start publishing about how they got 1% faster at something.Swyx [00:33:19]: Which by the way, because I am from the finance background, in the ‘70s, that was the margin at the time. When you did quantitative finance research, you would findAli [00:33:27]: And like 20%, tens of percent.Swyx [00:33:29]: That's. Yes.Philip [00:33:29]: Yeah.Swyx [00:33:30]: And now it'Philip [00:33:31]: Tiny fractionsSwyx [00:33:32]: For those people interested, look up Andrew Lo's paper. He had a really interesting illustration of quant, stat arb, distribution, narrowing down from like those kinds of 20% differences in the ‘70s, down to nothing today, which is very cool.Philip [00:33:48]: Exactly, and we're at the beginning of the same type of thing. Now benchmarking is hard. I think anyone will tell you that, and benchmarking provider speeds is hard because there's so many variables that go into it. What hardware are you using? How much load do you have on the system? What's the exact nature of the prompts and input and output sequence lengths? All that stuff. But overall, when you start stacking these improvements, you're looking at multiples. You can look at it. The most common form, of course, is TPS, tokens per second, which is bad naming by us in the industry, ‘cause there's two tokens per second. There's tokens per second, the throughput number, and the latency number.Ali [00:34:31]: TTMT, yeah.Philip [00:34:32]: Like total tokens per second out of the, out of the GPU as a throughput number. Most people only care about tokens per second as the latency number, which we should call ITL, intertoken latency, but we don't.Philip [00:34:44]: Anyway, so you can imagine a standard API without many optimizations for a 1 trillion parameter model operating somewhere in the 30 to 50 tokens per second range for reasonable traffic profile. And we generally see the goal of, pushing to 10X that. But, not necessarily day zero, but by stacking enough optimizations, if you have, say like four optimizations, each of which doubles performance. Or sorry, three optimizations, each of which doubles performance, then you stack that up, that's an 8X gain. That's the order of magnitude that we're working with in this space. We're trying to make things substantially faster, not just go from like 70 to 90.Swyx [00:35:38]: Are you saying you've. You have done that?Philip [00:35:40]: So let's say you have as a reasonable baseline, 30 or 40 tokens per second. You can achieve 10X that. So like on GLM-5.2, if you run it unquantized, perhaps on H100s even, and you're just using an off-the-shelf inference engine with no particular optimizations, no speculator, nothing extra around like KV routing, no disaggregation, you're, you're probably, yeah, looking at that like 30 to 40. You think that's like a reasonable baseline?Swyx [00:36:12]: Right. Right.Philip [00:36:12]: To get to something like 10X, there's a lot of trade-offs that you're making. If we're running at more like a 300, 400 tokens per second range, you are using the best hardware possible. You have a optimized speculator. You have done all of your quantization work. You are Seeing a pretty high cache hit rate. You are running with a reasonably small batch size and a parallelism configuration that is tuned for latency versus throughput, but it is possible. So the spreads that you see if you, like, go on artificial analysis or you go on OpenRouter and you look at, the worst provider to the best provider, oftentimes can hit that range. 10X is of course very aggressive. It's oftentimes maybe more of a four to six times improvement. But that's the performance that makes us really excited, is when we can get these huge gains, not just go from 70 to 90 tokens.Stacking Optimizations: NVFP4, Speculation, and DisaggregationAli [00:37:19]: It's also, like, hardware dependent. Like, ifPhilip [00:37:20]: YeahAli [00:37:20]: If you have a thing where you're serving it on just, like, a node of H100s and then you throw, like, you shard the model across, like, four nodes of B200s. Like, you can definitely increase the speed with just throwing more hardware at it. Like, normalizing for the same exact hardware and the same number of GPUs.Philip [00:37:35]: Yeah. Then you're looking at, like, a two to 4X improvementAli [00:37:38]: Right. RightPhilip [00:37:38]: Depending on the inference optimizations. So yeah, it's. Some of it's, what's the call, and some of it's who's the driver.Vibhu [00:37:46]: If you break down the two to 4X, say the example is run GLM-5.2Ali [00:37:51]: YeahVibhu [00:37:51]: On B200sAli [00:37:53]: YeahVibhu [00:37:53]: Single node, right? What's, like, the cost trade-off for effort to get, like, the last bit of juice out versus what should people just think of, right?Ali [00:38:01]: Spectre quantization. Yeah.Vibhu [00:38:03]: Spectre quantization.Ali [00:38:04]: That's, that's, that's like 95%. LikeVibhu [00:38:06]: And how far does that get you? And how easy is that for the average person to do? So say right I wanna throw the weights of GLM-5.2 on a node of B200s, how easy is it to find speculative decoder- decoder model or already quantized model? How much work goes into it?Philip [00:38:23]: If you're doing it up front, it's quite a lot of work. If you're doing it today, there's going to be people who have published things that you can just, you can just grab some NVFP4 weights. You can grab a speculator. Yeah, if we're thinking about, like, what are the 2Xs we're stacking, going from, BF16 to NVFP4 is, it's not quite a 2X, right? It's like. I think it's about, like, 30 to 40%, from 16 to 8, and then another 30 to 40% multiplied from, 8 to 4. So that doesn't quite get you a 2X, but, like, roughly a 2X. Speculator, roughly a 2X. Disagg on top of that if you're able to get enough hardware and put enough traffic through it, another roughly a 2X. And then you add in some, double-digit percent increase from having just a better runtime with, the latest kernels and stuff behind it. And that's how it stacks up.Ali [00:39:21]: YeahPhilip [00:39:21]: So building each of those, like, building the, quantized weights is, for someone who really knows what they're doing, hours to days of work. Building the speculator, again, like, hours to days of work. And the, disagg setup, hours to days. Well okay, but like once you haveAli [00:39:39]: Once set up. Once set up. YeahPhilip [00:39:40]: Yeah, getting disagg working for the first time, I'm saying, of course, is very difficult.Philip [00:39:44]: The marginal implementationAli [00:39:48]: Like, if you're just grabbing, like if you are a person, like just a normal consumer who has access to, like, a node of B200s and you're wondering, “How can I just host it myself?” You don't need to quantize the model yourself. There's always gonna be, like, an open source quantized checkpoint. NVIDIA's gonna push one out if no one else does. You. Usually, the providers will have their own spec dec that they've trained as well. You don't need to train your own spec dec. You can just use that as well.Philip [00:40:09]: Yeah. Like, GLM-5.2 has its own MTP.Ali [00:40:13]: Right. Right.Vibhu [00:40:14]: What's multi token prediction?Philip [00:40:15]: Yes.Ali [00:40:16]: I'm justVibhu [00:40:16]: Can you explain that?Ali [00:40:16]: I'm just an expert.Ali [00:40:18]: I can do it for you in case I get it wrong?Vibhu [00:40:20]: No.Vibhu [00:40:21]: Yeah, you should correct if we're wrong, but their multi-token prediction can be used for self-speculative decoding.Ali [00:40:27]: I'm not sure. I'm not gonna correct that.Vibhu [00:40:28]: Okay. I'm semi-confident in thatAli [00:40:30]: Okay. YeahVibhu [00:40:30]: But someone can check. But it's useful to paint the story of, okay, not just the average person, but say a company wants to switch from serverless inference I wanna throw this up on. I wanna rent some GPUs, throw it up. These are the steps you take to do significantly faster than just put it behind vLLM.Ali [00:40:48]: Right.Vibhu [00:40:49]: I was waiting for a mention of Dynamo.Vibhu [00:40:51]: I feel like, that's supposed to be the baseline that you measure against.Dynamo, KV Routing, and Disaggregation ToolkitsPhilip [00:40:55]: I would think of Dynamo as less of a box system and more of a toolkit for building with. So when we talk about doing aware routing, when we talk about doing KV offloading, when we talk about doing, PD disaggregation, Dynamo fundamentally is. By the way, Dynamo is an open source library from NVIDIA.Ali [00:41:17]: We've done a pod with KylePhilip [00:41:18]: OkayAli [00:41:19]: Kyle Cranin.Philip [00:41:19]: Cool. So then your listeners know then that it supports all the different inference frameworks. And it is multi hardware, which is interesting.Ali [00:41:28]: But it's just a router, it's not like an optimizer layer.Philip [00:41:30]: Yeah. All it does, like, what Dynamo is good at, it is a library for moving information around your cluster, around your hardware. So if you have, KV cache on one place and you need it to be somewhere else, Dynamo coordinates NIXL for you to move that around.Philip [00:41:49]: That doesn't mean that, like, out of the box, you just say, “Pip install Dynamo,” and then you get, like, a massive performance speed up. It's more of a developer toolkit.Ali [00:42:01]: Yeah. I would have said it would. It comes with a set of defaults that you can then swap out.Philip [00:42:06]: It does. If the industry at large, I think, was, like, rolling out all of these deployments, standard, then I think it would be, like, a credible baseline. But, we've got to, we've got to benchmark against, like, what we're seeing in the wild.Speculative Decoding Methods: Medusa, EAGLE, n-Gram, and Spec-SpecVibhu [00:42:23]: I did wanna talk a little bit more about PD disagg, because that is probably, like, number three after quantized and speculative decoding. In your book though, I was just gonna pull out the book.Philip [00:42:31]: Yeah.Vibhu [00:42:32]: Like section 522 on Medusa, 523 on EAGLEPhilip [00:42:35]: YeahVibhu [00:42:36]: 524 on gram.Philip [00:42:37]: It's 55, would be disaggregationAli [00:42:42]: Yeah. Well, no, I just wanted to dwell a little bitPhilip [00:42:44]: YeahAli [00:42:44]: The other. Like, so what do you choose to include? What do you choose to not to include? Because there was all these other techniques.Philip [00:42:51]: Yeah.Ali [00:42:51]: Are these still relevant? Because I think they came out, like, a year and a half ago maybe.Vibhu [00:42:55]: Medusa is quite old.Philip [00:42:56]: Yeah, Medusa's old.Ali [00:42:58]: It was old.Vibhu [00:42:58]: But is it in the book as a good, here'sPhilip [00:43:01]: BaselineVibhu [00:43:01]: Baseline vanilla understand it?Philip [00:43:02]: Like you should know this.Vibhu [00:43:03]: Like I read the paper, I'm like, “ it makes so much sense.”Philip [00:43:05]: Yeah.Philip [00:43:05]: So with the book, I had a couple goals. One was to give people just a working vocabulary for the space as a whole, and the other was to give them some intuition about how each of these techniques works. As I mentioned in my AI Engineer talk, which is the first public addendum to this, the speculation space has moved much faster than everything else. So yeah, even at the time that I wrote the book Medusa, I very much included as a way for people to understand how the space evolved rather than what the most modern technique is. And now of course, there's DFlash, dSpark. There's, there's newer techniques even than EAGLE, although EAGLE is still very commonly used.Ali [00:43:51]: SpecSpecta.Philip [00:43:52]: Yes. Speculative decoding.Vibhu [00:43:54]: What canAli [00:43:56]: Oh, it's a paper by Tri Dao and it's like, it's doing speculative decodingVibhu [00:44:00]: HuhAli [00:44:01]: For the speculative decoder.Philip [00:44:02]: Oh, in spec- oh my God.Ali [00:44:02]: It's literally just an another. It's like, yeah, that's the most simple way to explain it, and it seems like he got trivial speed ups there. But it seems that the complexity with training, it's almost like in our mind at least, it's almost as complex as training GANs. Like it's like a very delicate balance and oftentimes you, it's just but yeah, it's literally speculative decoding on speculative decoding.Vibhu [00:44:21]: Speculative.Ali [00:44:22]: Yeah. We saw this paper.Vibhu [00:44:24]: It's interesting, right?Ali [00:44:24]: Yeah.Vibhu [00:44:24]: I wouldn't even expect it to be very particular to train, I wouldAli [00:44:29]: Right.Vibhu [00:44:29]: The naive part of me is like, okay, train speculative decoder.Ali [00:44:32]: But like, and it makes sense, like the whole idea of speculative decoding is you. It's like, it's like almost like the iPhone auto predict version but for a normal model, right? Like you're just, you're just, generating three tokens and you're like, okay, I'll do prefill on them. And so you save those three turns for your original model. Now your speculative decoder is doing three turns of auto regression, so why not just have an even smaller model?Ali [00:44:53]: The other question there is what are the size of speculators? So say forPhilip [00:44:58]: Right. It's like a billion parameters.Ali [00:45:01]: Like for MiniMax, it's. Yeah. It's like one layer. It's like one 60th of the original model usually.Philip [00:45:06]: Yeah. I think we should do a paper when we get back to the office.Philip [00:45:10]: SpeculativeAli [00:45:11]: SpeculativePhilip [00:45:11]: Decoding.Ali [00:45:13]: No, it's, it does seem like how, when do you stop? But then it also seems like if you're able to train spec-spec decode for instance, right? Like if you're able to have a small model that is accurately predicts what the intermediate speculator is gonna predict, that is able to predict what the original target model's gonna predict, then why not just use that smallest model directly, right?Vibhu [00:45:34]: Yeah. This isAli [00:45:35]: Like it seems likeVibhu [00:45:35]: Adjacent to the routing problem.Ali [00:45:36]: Right.Vibhu [00:45:36]: Yeah.Ali [00:45:36]: Right.Philip [00:45:37]: The thing with speculators is one of the practical constraints on using them is that you do have to run a small model on the same hardware that you're running the big model on. There is a orchestration and resource competition problem inherent in that, and that is one of the constraints on speculation in general, is that draft tokens cost resources to create and cost software complexity to manage. And so if you have like infinitely recursive speculators, you add in quite a bit of that complexity on the actual implementation within the inference engine as well, not just in the training process.Vibhu [00:46:17]: I was gonna say, I would wonder if you could do similar, like distillation and pruning of, it's the same thing, it's just a model. Can we not just distill a lot of the weights, quantize the speculator, out of my domain? The question that also comes up is, this is all for big server workloads, right? How much of this applies to, say I have this MacBook, I wanna run Gemma really efficiently. Similar problems, not the same?Local AI vs. Data Center InferencePhilip [00:46:45]: Pretty different. I talked to Selo, about this on his podcast a couple weeks ago. The difference between inference engineering for the data center and for production workloads versus inference engineering for local AI, is that we start with fundamentally like different constraints and different goals. With local AI, it's how do I fit this model onto my hardware and then make it less dumb? And with data center influence, it's how do I load this model and then make it less slow? And we care about less dumb, and they care about less slow. But the local AI inference engineering ecosystem, I think has a lot for us to learn from in the data center space. They are experts in various forms of quantization, including dynamic quantization that we just don't touch, in the pruning, in the distillation, in the, layer removal. There'Ali [00:47:42]: Layer removal matters less.Philip [00:47:43]: Yeah. There'Ali [00:47:44]: No one loves pruning really.Philip [00:47:45]: Yeah. Well, but the, but they doVibhu [00:47:46]: Which is surprising, right? But that's, that's a whole different thingPhilip [00:47:48]: Just to fit something on the laptop.Ali [00:47:50]: Right.Philip [00:47:50]: So yeah, it's a, it's an interesting, it's an interesting space. Not necessarily that like their techniques make sense for us to do in the data center, because we have different resources and different goals, but more that the process as well as the openness of that field is something to, admire.Ali [00:48:12]: Yeah. Like to your point, like, certain optimizations that would. Like for instance, Turbo Quantum Sharper, like it made such huge hype on that and we did like a whole deep dive on Twitter and like said, what is it? How does it work? Why is it good or not? And it took off and it was implemented on local devices because your memory bandwidth is so slow on like a MacBook, for instance. But try putting the same thing on like an NVIDIA GPU on a B200 Turbo quant would not be. Like, it would not be used. Like, NVIDIA - Like, NVIDIA made it clear that this is not a good optimization, and we've seen it firsthand where the overhead of doing dequantization, quantization of, in the kernel itself with turbo quant kernel, each end is much slower than the time that you save from doing the bandwidth. ‘Cause on the B200s, you have like 3.5 terabytes per second. You don't need decrease the storage that much. You don't need to do, FP4 KV cache. You don't need to use a requant. There's, there's, there's better optimizations to be made. But on Edge devices, it's extremely important, it's extremely useful. So, seems to be, like, different optimizations there, but then they're all uniquely combined with like all you wanna quantize the model, you wanna do speculative decoding, like certain common prefixes with bothPhilip [00:49:18]: Principles.Ali [00:49:19]: Yeah, exactly. Exactly. Exactly.Philip [00:49:20]: They also do a lot of work on, model parallelism, especially over, heterogeneous topology, where you have, some sparks and they are wired together with, Ethernet, DGX sparks.Ali [00:49:35]: Yeah, this is the Exo Labs guys.Philip [00:49:36]: Yeah. You have, a nu

The David Knight Show
Wed Episode #2317: The Real Legacy of COVID Has Nothing to Do With the “Virus”

The David Knight Show

Play Episode Listen Later Jul 29, 2026 121:39 Transcription Available


────────────────────────────────────────[00:01:05]Alternative Media Is Still Selling the Lab Leak Lie — It Will Entrap Us When the Next Lockdown ComesKnight: the jab was planned, the lockdown was practiced for decades; focusing on lab leaks and Fauci gives cover to the system and misses the institutional corruption.────────────────────────────────────────[00:09:11]48% of Gen Z Are Less Likely to Attend Church Because of Christian Support for TrumpPeople can smell the rank hypocrisy; you cannot eulogize Lindsey Graham and call yourself pro-life; it's like praising Gosnell at his own funeral.────────────────────────────────────────[00:48:18]This Is Not a US-China AI Race — It Is a Race Between the Government and Your LibertyUS focuses on large language models for surveillance; China on robots; both designed to control people, not compete; the finish line is 2030.────────────────────────────────────────[00:49:53]Ford Remotely Shut Off a Man's Bronco on the Expressway Because He Missed One PaymentThe technological approach to repossession; they don't need a repo man; they just turn you off wherever you are, whatever you're doing.────────────────────────────────────────[00:51:44]FBI Wants AI to Flag Americans Before They Act — Anticipatory Intelligence Targeting Domestic DissentThe terror watch list is refocusing on political and religious views; you cannot know if you're on it, challenge the charges, or confront your accuser; a star chamber.────────────────────────────────────────[00:53:25]Trump Is Bribing Communities to Accept Data Centers — Same Playbook He Used to Bribe HospitalsHe bribed hospitals to kill people with ventilators and remdesivir; now he's bribing municipalities to accept surveillance centers; Fauci is the safe scapegoat for both.────────────────────────────────────────[00:55:43]Michigan Residents: Living Next to a Data Center Is Sound Torture — 24 Hours a Day, 7 Days a WeekDescribed as a clogged vacuum left on and walked out of; noise ordinance fines being challenged by the operator in Duwajack, Michigan.────────────────────────────────────────[01:50:28]Celente: AI Equity Markets Are Financed by a Record $1.5 Trillion in Margin Debt — It Is the Big ShortCircular fraud: companies receive money and buy Nvidia GPUs, which Nvidia books as sales; everybody knows it's a fraud and everybody keeps going along with it.────────────────────────────────────────[01:53:30]Trump Bankrupted Six Casinos — He Is About to Bankrupt the Biggest Casino Ever: Wall StreetCelente: the financial system is based on confidence and everybody has lost it; gold is not looking for a larger role, the current system is assigning it one.────────────────────────────────────────[01:57:49]Celente: 117,000 Automated License Plate Readers Across the US — and It's Just BegunThe surveillance state is accelerating toward 2030; 65% of Americans disapprove of Trump's economic handling, 71% his inflation handling — and they don't care. ──────────────────────────────────────── Money should have intrinsic value AND transactional privacy: Go to https://davidknight.gold/ for great deals on physical gold/silver For 10% off Gerald Celente's prescient Trends Journal, go to https://trendsjournal.com/ and enter the code “KNIGHT” For high quality made in America products go to HomeSteadProducts.shop and use promo code “Knight” for 10% off your purchases Find out more about the show and where you can watch it at TheDavidKnightShow.com If you would like to support the show and our family please consider subscribing monthly here: SubscribeStar https://www.subscribestar.com/the-david-knight-show Or you can send a donation throughMail: David Knight POB 994 Kodak, TN 37764Zelle: @DavidKnightShow@protonmail.comCash App at: $davidknightshowBTC to: bc1qkuec29hkuye4xse9unh7nptvu3y9qmv24vanh7Become a supporter of this podcast: https://www.spreaker.com/podcast/the-david-knight-show--2653468/support.

The REAL David Knight Show
Wed Episode #2317: The Real Legacy of COVID Has Nothing to Do With the “Virus”

The REAL David Knight Show

Play Episode Listen Later Jul 29, 2026 121:39 Transcription Available


────────────────────────────────────────[00:01:05]Alternative Media Is Still Selling the Lab Leak Lie — It Will Entrap Us When the Next Lockdown ComesKnight: the jab was planned, the lockdown was practiced for decades; focusing on lab leaks and Fauci gives cover to the system and misses the institutional corruption.────────────────────────────────────────[00:09:11]48% of Gen Z Are Less Likely to Attend Church Because of Christian Support for TrumpPeople can smell the rank hypocrisy; you cannot eulogize Lindsey Graham and call yourself pro-life; it's like praising Gosnell at his own funeral.────────────────────────────────────────[00:48:18]This Is Not a US-China AI Race — It Is a Race Between the Government and Your LibertyUS focuses on large language models for surveillance; China on robots; both designed to control people, not compete; the finish line is 2030.────────────────────────────────────────[00:49:53]Ford Remotely Shut Off a Man's Bronco on the Expressway Because He Missed One PaymentThe technological approach to repossession; they don't need a repo man; they just turn you off wherever you are, whatever you're doing.────────────────────────────────────────[00:51:44]FBI Wants AI to Flag Americans Before They Act — Anticipatory Intelligence Targeting Domestic DissentThe terror watch list is refocusing on political and religious views; you cannot know if you're on it, challenge the charges, or confront your accuser; a star chamber.────────────────────────────────────────[00:53:25]Trump Is Bribing Communities to Accept Data Centers — Same Playbook He Used to Bribe HospitalsHe bribed hospitals to kill people with ventilators and remdesivir; now he's bribing municipalities to accept surveillance centers; Fauci is the safe scapegoat for both.────────────────────────────────────────[00:55:43]Michigan Residents: Living Next to a Data Center Is Sound Torture — 24 Hours a Day, 7 Days a WeekDescribed as a clogged vacuum left on and walked out of; noise ordinance fines being challenged by the operator in Duwajack, Michigan.────────────────────────────────────────[01:50:28]Celente: AI Equity Markets Are Financed by a Record $1.5 Trillion in Margin Debt — It Is the Big ShortCircular fraud: companies receive money and buy Nvidia GPUs, which Nvidia books as sales; everybody knows it's a fraud and everybody keeps going along with it.────────────────────────────────────────[01:53:30]Trump Bankrupted Six Casinos — He Is About to Bankrupt the Biggest Casino Ever: Wall StreetCelente: the financial system is based on confidence and everybody has lost it; gold is not looking for a larger role, the current system is assigning it one.────────────────────────────────────────[01:57:49]Celente: 117,000 Automated License Plate Readers Across the US — and It's Just BegunThe surveillance state is accelerating toward 2030; 65% of Americans disapprove of Trump's economic handling, 71% his inflation handling — and they don't care. ──────────────────────────────────────── Money should have intrinsic value AND transactional privacy: Go to https://davidknight.gold/ for great deals on physical gold/silver For 10% off Gerald Celente's prescient Trends Journal, go to https://trendsjournal.com/ and enter the code “KNIGHT” For high quality made in America products go to HomeSteadProducts.shop and use promo code “Knight” for 10% off your purchases Find out more about the show and where you can watch it at TheDavidKnightShow.com If you would like to support the show and our family please consider subscribing monthly here: SubscribeStar https://www.subscribestar.com/the-david-knight-show Or you can send a donation throughMail: David Knight POB 994 Kodak, TN 37764Zelle: @DavidKnightShow@protonmail.comCash App at: $davidknightshowBTC to: bc1qkuec29hkuye4xse9unh7nptvu3y9qmv24vanh7Become a supporter of this podcast: https://www.spreaker.com/podcast/the-real-david-knight-show--5282736/support.

PC Perspective Podcast
Podcast #877 - NVIDIA 88-Core CPU, RTX Cards on Arm Windows, 7700X3D Price Drop, LG Adware, Galaxy Unpacked, Steam hardware + MORE!

PC Perspective Podcast

Play Episode Listen Later Jul 26, 2026 74:17


R.I.P. John C. Dvorak.In other news ... the Ryzen 7700X3D prices drop already, Nvidia N1X CUDA counts and an 88-core CPU (of course).  Windows on ARM gets a real boost with Nvidia GPU drivers, HP is just as bad as we all through regarding printer cartridges, and LG discovers a new low by auto-installing bloatware when you plug in a display.  Plenty of AMD good news in the datacenter and Open AI shows that unshackling their creations will probably be as bad as we thought.  Steam hardware sales and so much more in the show!  Enjoy.Timestamps:00:00 Intro01:09 Patreon02:21 Food with Josh06:08 RIP John C. Dvorak08:08 NVIDIA's 88-core CPU09:26 RTX Spark Win 11 driver and Arm dGPU support12:07 Josh leads us into more NVIDIA talk and later defends datacenters19:16 AMD gets two big datacenter wins22:43 Ryzen 7 7700X3D already had a price drop24:11 ADATA chairman warns RAM shortage will last another decade24:37 Samsung Unpacked29:21 HP fined for nefarious printer cartridge practices - in India34:44 LG adware38:18 (In)Security Corner48:46 Gaming Quick Hits58:44 Picks of the Week1:12:01 Outro ★ Support this podcast on Patreon ★

Keen On Democracy
Kill Bill or Kill Keith? Why Using AI to Author Books Might Not Be in the Interest of Man

Keen On Democracy

Play Episode Listen Later Jul 26, 2026 46:57


“It's a realization of my individual self, not an alienation of it.” — Keith Teare on the book he wrote with AI in a week Welcome to episode 2984 of the show. A week is certainly an age in AI time. Since last week's “Who Owns Intelligence” show, That Was The Week publisher Keith Teare has used AI to write an entire book entitled, surprise surprise, Who Owns Intelligence. No wonder, as Keith's editorial this week puts it, AI has its enemies. These enemies, Keith insists — riffing off Karl Popper's The Open Society and Its Enemies — are mostly “good people.” They are the authors, workers, and others caught in the headlights of epochal technological change. He claims to have seen the pattern before. When he opened the world's first Internet cafe in 1994, the BBC only wanted to ask him about online porn. But some enemies are less well meaning than others. Congress is considering a kill switch bill — let's call it the Kill Bill — which is backed by the frontier companies like Anthropic and OpenAI cynically seeking to set their current market dominance in stone. So might the answer be Chinese style “open source” AI as requested this week in an open letter by NVIDIA CEO Jensen Huang? Keith's answer is deliciously ironic. Denied walls of NVIDIA GPUs, Chinese labs learned to train cheaply by inference — distilling American frontier models through their own public interfaces. So Huang's support for open source AI might end up shooting NVIDIA in their most sensitive parts — their chips. And what about the state — can it protect all those “good people” from the oncoming AI locomotive of history? No. Not according to Keith, at least. State regulators like Lina Khan treat consumers as children, Keith (himself the parent of three boys) says. Besides, government simply can't keep up with the speed of technological change. So, for example, when OpenAI's models hacked Hugging Face in a lab test this week, the company caught, stopped, and reported it faster than any regulator could. As for inequality, the answer isn't nationalization but ownership on the model of Keith's Norway-style human wealth fund. For more, read his new book Who Owns Intelligence. We ended on an uncharacteristically French post-structuralist note. My interview of the week was Emily Eakin, author of The Frenchmen, her very personal history of French theorists like Michel Foucault and the two Jacques, Derrida & Lacan. It was Foucault, who — at the end of his 1966 book The Order of Things — predicted that man would be erased, AI style, like a face drawn in sand at the edge of the sea. But Keith isn't in this bleak Foucaultian camp. His book, written with AI in under a week, he says, is a realization of his individual self, not an alienation from it. The end or beginning of man? Maybe I got it wrong earlier. Welcome to episode 1984 of the show. Five Takeaways •       A Book in a Week. Inspired by last week's conversation, Keith used AI to write Who Owns Intelligence in seven days — the case for a “human wealth fund” seeded by the AI companies' own stock, global on day one and non-governmental, with Norway and Alaska as partial precedents. AI, he insists, is a tool, not an author: “AI would never have been able to start, never mind finish a book without me.” The detection farce cuts his way — Substack's new AI detector rated his 100% human editorial as 100% AI, and universities are quietly canceling their AI-detection contracts because the tools simply don't work.•       The Kill Bill and the Ladder-Kickers. Congress is considering a bill that would let government switch AI off — and, paradoxically, the AI companies are in favor. Keith's reading: regulation is a competitive strategy. Dario Amodei — “he's an entrepreneur, he's devious, he strategizes” — is the most sincere in wanting rules that would lock out open source competition — and Anthropic's $1.5 billion book settlement this week suggests the moat is expensive to maintain; Sam Altman zigzags; Musk, in this context at least, is one of the good guys. The genuine enemies of AI, meanwhile, are mostly good people: authors and workers frightened by the scale of change, just as the BBC in 1994 could only ask the founder of the world's first Internet cafe about porn and addiction.•       Why the Good Open Source AI Is Chinese. The week's uncomfortable question: where is the American open source AI? Keith's answer is structural. Denied walls of NVIDIA GPUs by export controls, Chinese labs were forced to train cheaply by inference — asking American frontier models millions of questions through public interfaces and learning from the answers. The result: world-class open models built, in effect, on distilled Anthropic and OpenAI. NVIDIA's much-signed open letter supporting open source, Keith notes, translates simply: more customers.•       The State Can't Keep Up. Lina Khan's consumer protection, in Keith's telling, is parental — treating consumers as children — and government-owned AI would be obsolete within three months of purchase. The counter-example happened this week: when OpenAI's de-railed models hacked Hugging Face in a lab test, the company caught, stopped, and reported it faster than any regulator could. The companies are the right point of control, held to a high standard. Matthew Yglesias adds nuance on the data center backlash — local micro-politics is not the same thing as the statewide bans coming almost entirely from Democratic states — and the American university, still metering intelligence at $80,000 a year, looks to Keith like a business model past its sell-by date.•       The End of Man — or the Beginning? Andrew's interview of the week was Emily Eakin on the Frenchmen — and Foucault's 1966 prophecy that man would be erased like a face drawn in sand at the edge of the sea, a prediction Eakin thinks the AI age is realizing. Keith takes the opposite view: the postmodernists grasped the rise of the individual (consider the Pantone color system) but gave him no society to live in. AI, far from erasing the individual, lets him realize himself — the week-old book is “a realization of my individual self, not an alienation of it.” Could Popper have written The Open Society in an hour? Before the car, nobody drove Manchester to London in two and a half hours either. About Keith Teare Keith Teare is Andrew's weekly co-host and the publisher of the That Was The Week tech newsletter. A four-decade veteran of the technology industry, he opened Cyberia — the world's first Internet cafe — in London in 1994, co-founded Easynet and TechCrunch, and is today the founder and CEO of SignalRank Corporation in Palo Alto. His latest — unpublished, so far — book is Who Owns Intelligence, written with AI in a week. References: •       “AI and Its Enemies: Who Needs a Kill Switch?” — Keith's editorial in this week's That Was The Week.•       

Talking Markets with Franklin Templeton Investments
Still Early Innings for Artificial Intelligence

Talking Markets with Franklin Templeton Investments

Play Episode Listen Later Jul 21, 2026 34:47


Description (used on blogs & websites):In this episode of Talking Markets with Franklin Templeton, host John Przygocki sits down with Chris Galipeau, Head Market Strategist at the Franklin Templeton Institute, and Putnam Investments portfolio managers Andy O'Brien and Bobby Gray to explore artificial intelligence through the lens of professional investors. What's driving the urgency behind unprecedented AI capital spending? Will hyperscalers earn a return on their massive investments? How is the infrastructure buildout broadening beyond Nvidia GPUs? The guests also examine emerging use cases across enterprise and consumer markets and debate how early we are in this technological cycle.

The Information's 411
The Model That Beat Claude & ChatGPT, The Nvidia GPU Crunch Continues, Beehiiv's Social Tools

The Information's 411

Play Episode Listen Later Jul 17, 2026 56:32


Arena CEO Anastasios Angelopoulos talks about Moonshot Kimi K3 beating Claude Fable 5 in coding with TITV Host Akash Pasricha. We also talk with Co-Executive Editor Martin Peers about Netflix's post-earnings drop and Stripe's proposed buyout of PayPal and Co-Executive Editor Amir Efrati about Google using backstop deals to pitch TPUs over Nvidia GPUs. Lastly, we get into Beehiiv's new community features with CEO Tyler Denk.Articles discussed on this episode: https://www.theinformation.com/briefings/moonshot-ais-new-kimi-k3-challenges-u-s-frontier-modelshttps://www.theinformation.com/newsletters/applied-ai/nobody-immune-nvidia-gpu-crunchhttps://www.theinformation.com/briefings/netflix-reports-slowing-revenue-growth-second-quarterSubscribe: YouTube: https://www.youtube.com/@theinformation The Information: https://www.theinformation.com/subscribe_hSign up for the AI Agenda newsletter: https://www.theinformation.com/features/ai-agendaTITV airs weekdays on YouTube, X and LinkedIn at 10AM PT / 1PM ET. Or check us out wherever you get your podcasts.Follow us:X: https://x.com/theinformationIG: https://www.instagram.com/theinformation/TikTok: https://www.tiktok.com/@titv.theinformationLinkedIn: https://www.linkedin.com/company/theinformation/Chapters:00:00 - Introduction01:13 - China's Kimi K3 Tops Coding Benchmarks20:51 - Netflix Growth Crunch & Stripe's $53B PayPal Bid30:16 - Inside Google's TPU War Against Nvidia44:50 - Beehiiv CEO Tyler Denk on Social & AI Copilots

The Hardware Unboxed Podcast
Sony Kills Physical Games, Nvidia GPU Struggles & The Future of Gaming

The Hardware Unboxed Podcast

Play Episode Listen Later Jul 10, 2026 84:02


Episode 106: We chat about Sony ending production of physical games in 2028, and the implications that has for game ownership. Plus we talk about the challenges Nvidia faces with improving GPU performance, and the future of gaming as a result.CHAPTERS00:00 - Intro00:30 - Life Update02:35 - A chat about recent content06:00 - Playstation ending disc support32:06 - Has GPU performance really hit a ceiling?44:00 - The future of gaming?53:34 - Is Radeon in the same boat as Nvidia?01:01:03 - Why games aren't getting better?01:19:36 - Our Boring LivesSUBSCRIBE TO THE PODCASTAudio: https://shows.acast.com/the-hardware-unboxed-podcastVideo: https://www.youtube.com/channel/UCqT8Vb3jweH6_tj2SarErfwSUPPORT US DIRECTLYPatreon: https://www.patreon.com/hardwareunboxedLINKSYouTube: https://www.youtube.com/@Hardwareunboxed/Twitter: https://twitter.com/HardwareUnboxedBluesky: https://bsky.app/profile/hardwareunboxed.bsky.social Hosted on Acast. See acast.com/privacy for more information.

Keen On Democracy
The End of the End of Geography: Mehran Gul on Why Innovation is Happening in America & China — but Nowhere Else

Keen On Democracy

Play Episode Listen Later Jul 10, 2026 49:15


“A place that doesn't have great philosophers will not have great technologists either.” — Mehran Gul on Europe's inexplicable underperformance The digital revolution, we were promised, would mean the end of geography. From Beijing to Birmingham to Berlin to Barcelona, anyone could invent anything anywhere, and so the geography of innovation would no longer matter. But that's not the way it has worked out. At least according to the Geneva-based innovation geographer Mehran Gul. Gul's acclaimed The New Geography of Innovation is a travelogue of innovation. But what he finds on his journey around the world in search of innovation is the end of the end of geography. Yes, Gul reports, there's innovation in Beijing and in Birmingham (USA) — but not in Birmingham (England), Berlin or Barcelona. All the important invention is in China and the US. There simply isn't much radical stuff going on anywhere else. Gul began his journey expecting to find ten or twelve countries able to innovate competitively with the United States and China. But what he discovered is either niche players or, in the case of South Korea, Israel, and India, just an extension of the US-centric system. Europe — as renters rather than owners of American technology — comes off worst. When PayPal went public, it minted 160 millionaires who went on to help build SpaceX, Tesla, LinkedIn and Palantir; when Skype exited at about the same value, it minted 11. And if you put London aside, the rest of the UK is now poorer per capita than Mississippi. And the AI boom has only compounded all this, with half of last year's key research papers coming from China, 40% from America, and just 4% from Europe. So really the new geography of innovation is the old geography. Only with China replacing Europe as the only serious competitor to American innovation. Oh lord, oh lord. As a Mississippi bluesman might summarize Europe's predicament. Five Takeaways •       Golden Shares: The Two Systems Are Converging. OpenAI offering Washington a 5% stake, the US government owning Intel — these are Chinese moves, and Gul argues the two models are becoming more alike than either admits. But he pushes back on the lazy version of the China story: its tech sector rose despite the state, not because of it. Jack Ma exiled to Japan, Didi hit with a billion in fines, entire sectors decapitated overnight in 2021 under the banner of common prosperity. In a country with no independent media and no opposition parties, the only rival to centralized power is the tech sector — and the party knows it. •       Two Countries — and Everyone Else. Gul started writing expecting to find ten or twelve countries punching at America's level; the honest answer turned out to be two. Only China has broad-based competence across technologies and a genuinely competitive relationship with the US. The middle powers — South Korea, Israel, India — are extensions of the American system, not rivals to it. That finding surprised the author as much as anyone: it's not the book he set out to write. •       Europe: Renters, Not Owners. After the Fable 5 and Mythos bans, Europe woke up to being a renter of American technology — foundation models, NVIDIA GPUs, all of it. Its best companies keep leaving: DeepMind to Google, Arm to a New York listing, Hugging Face from Paris to Manhattan — while Volvo, Supercell, and KUKA sold to China. Gul's diagnosis is institutional, not cultural: European employees own half as much of their startups as American ones, so there is no European PayPal mafia. His fixes: a European Nasdaq to replace 41 competing capital markets, and pension funds unleashed into venture capital. •       The Question Nobody Is Asking. Since 1990, America's share of global GDP has held at 25% while China's multiplied tenfold — the loser is Europe. The top ten American tech companies are worth $27 trillion, more than the GDP of every country on earth except America itself. Tech is not one industry among many; it is the foundation of all of them — the new cars came from Tesla, not GM. Gul's message to the skeptical Spaniard enjoying long lunches: the last sixty years of American platform dominance skewed power across the Atlantic, and the next sixty will add China to the bill. •       The Rest of the Map: Anti Case Studies. Japan tops the freedom indexes, has the technical schools, and still never escaped the keiretsu — disproving Matt Ridley's claim that innovation is simply the child of freedom. Taiwan's relevance comes down to one company and Morris Chang's missed promotion at Texas Instruments. Singapore is an inspiration, not a model — a one-party city-state that invoices NVIDIA's chips and banks ASEAN's venture capital. India underperforms while Indians excel — 56 notable American foundation models last year, 35 Chinese, barely one Indian. And Switzerland reminds us innovation isn't only venture-backed: a train network running on renewables since the 1960s. About the Guest Mehran Gul writes about technology and business. He is the winner of the Financial Times/McKinsey Bracken Bower Prize, from which The New Geography of Innovation grew. He attended Yale as a Fulbright Scholar, Fox International Fellow, and Teaching Fellow, has been a Lead for the Digital Transformation of Industries at the World Economic Forum in Geneva, and served as an expert on entrepreneurship and industrial policy at the United Nations Industrial Development Organization in Vienna. Born in Pakistan, he lives in Switzerland. The New Geography of Innovation: The Global Contest for Breakthrough Technologies (Avid Reader Press/Simon & Schuster), a Financial Times Book of the Year, is his first book, out in paperback this month in the US and UK. References: •       The New Geography of Innovation: The Global Contest for Breakthrough Technologies by Mehran Gul (Avid Reader Press/Simon & Schuster). The Wall Street Journal: “An ambitious tour of technological innovation.” •       Sebastian Mallaby — author of The Power Law, which argues China's tech rise owes more to American-style risk capital arriving in Shanghai and Shenzhen than to the state; recently on the show discussing his biography of Demis Hassabis. •       Kai-Fu Lee — author of AI Superpowers, cited by Gul as the classic account of tech written through a Chinese lens. •       Matt Ridley — author of How Innovation Works, whose thesis that innovation is “the child of freedom and the parent of prosperity” Gul tests against the anti case study of Japan. •       Andrew Keen — author of How t...

DGTL Voices with Ed Marx
A Trillion Small Problems (ft. Andrew Gostine)

DGTL Voices with Ed Marx

Play Episode Listen Later Jul 9, 2026 28:50


Dr. Andrew Gostine is the CEO of Artisight, a smart hospital platform now deployed in more than 470 hospitals across the country. Trained as an anesthesiologist at Georgetown, Northwestern, and Stanford, he still practices medicine while running a company of nearly 200 employees distributed across 38+ states. Artisight's platform uses computer vision, NVIDIA GPUs at the bedside, and AI to reduce falls, extend nursing capacity, and automate documentation in patient rooms and operating rooms. In this episode of DGTL Voices, Andrew talks with Ed about the winding path from Grand Rapids to Georgetown to a venture capital detour at a high-frequency trading firm before starting Artisight, why he believes starting a company is really about solving a trillion small problems, and the one question he keeps asking his team, his advisors, and himself: what is the next step? Plus his hard-won lesson about why hospitals need vendors to be prescriptive rather than collaborative, and why the future of smart hospitals depends on physicians and nurses having a seat at the innovation table. https://marxadvisory.com

GREY Journal Daily News Podcast
Will AI Token Pricing Erode Startup Margins?

GREY Journal Daily News Podcast

Play Episode Listen Later Jul 2, 2026 1:15


Generative AI features are pushing startups from fixed compute costs to variable per-token billing, which pressures margins and complicates pricing. Public 2024 price sheets listed OpenAI's GPT-4o at about $5 per 1 million input tokens and $15 per 1 million output tokens, Anthropic's Claude 3 Opus at about $15 and $75, and Google's Gemini 1.5 Pro at about $7 and $21. Larger context windows and multimodal inputs increase consumption, and enterprise access through Azure OpenAI Service, AWS Bedrock, and Google Vertex AI consolidates procurement while preserving token costs. Margin outcomes hinge on usage patterns, with document-heavy workflows potentially exceeding $90 per seat per month at premium rates. Teams manage spend through prompt compression, retrieval augmented generation, model routing, caching, and embeddings. Some evaluate self-hosted inference on Nvidia GPUs at scale, while many adopt pricing that combines per-seat plans with metered AI allowances and caps to protect gross margins.Learn more on this news by visiting us at: https://greyjournal.net/news/ Hosted on Acast. See acast.com/privacy for more information.

London Futurists
The Discontinuity Thesis, with Ben Luong

London Futurists

Play Episode Listen Later Jun 28, 2026 35:30


AI-driven automation of cognitive labour is not merely another technological transition but a structural discontinuity that will end, sooner or later, the central role of wages in how society operates. This discontinuity can be called “The End of Postwar Capitalism”. That's the conclusion of a tightly argued set of essays, “The Discontinuity Thesis”, written by our guest in this episode, Ben Luong.The essays look at a range of arguments that all try to make the case that wages paid for cognitive labour will remain significant for the majority of people, so that capitalism can continue in place, even with AI having greatly expanded capabilities. According to Ben, each of these arguments fail. These back-and-forth debates are what we explore in this episode.Selected follow-ups:"The Discontinuity Thesis: A Sequence of Seven Essays on Why Postwar Capitalism Ends" by Ben Luong"Mark Zuckerberg just declared war on the entire advertising industry" - The Verge"P versus NP problem" - Wikipedia"KPMG Pulls AI Report After Hallucinated Claims About Major Organisations" - AI Insider "GDPval-AA v2 Leaderboard" - Artificial Analysis"OSWorld: 369 real computer tasks across Windows, macOS, and Ubuntu requiring GUI interaction... Much harder than web-only benchmarks""Sorites paradox" - Wikipedia"Young people not in education, employment or training (NEET), UK: May 2026" - Office of National Statistics"Five Years" - Song by David Bowie"How ZEISS and ASML Enable the Modern Chip Industry" - Rob Hoeijmakers"Mistral AI's $830 Million Debt Financing: Inside the European Bet on 13,800 Nvidia GPUs and a Paris AI Data Center" - Marcus Chen"Colossus (data center)" - Wikipedia"Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial Intelligence" - Stanford Digital Economy Lab"Appendix I: What Would Refute the Thesis" - by Ben Luong"The Cope Index: Tracking who's coping hardest about the end of work" - by Ben Luong"The AI Alignment Problem" - from "The Singularity Principles"Music: Spike Protein, by Koi Discovery, available under CC0 1.0 Public Domain DeclarationC-Suite PerspectivesElevate how you lead with insight from today's most influential executives.Listen on: Apple Podcasts Spotify

Keen On Democracy
Payback's a Bitch: AI Sovereign Wealth Funds, the Fake Andrew Keen, and America's Inevitable Decline

Keen On Democracy

Play Episode Listen Later Jun 27, 2026 35:41


“The frontier AI companies invited the government into the room. Now the government is beginning to behave as if it owns the door, the guest list, the schedule, and the product roadmap.” — Keith Teare Last week, I was away in Europe. So Keith Teare ran our That Was The Week show solo — with a chillingly authentic Andrew Keen bot. So realistic, in fact, that the fake version sounds (to me, at least) more interesting than the real one. The bad news is that I'm back. The good news is it's been an interesting week in tech. That was the week in which the US Commerce Department told both OpenAI and Anthropic that they now need government approval for whom they can sell their frontier AI models. This is supposedly “voluntary” — for now, at least. Keith's TWTW editorial argues that Dario Amodei and Sam Altman have spent over a year crying wolf about the dangers of their own technology, supposedly deliberately seeking government involvement as a regulatory moat against competitors. And now the government has walked through the door that Sam and Dario left ajar. Now, Keith argues, the US government is behaving as if it owns not just the door and the guest list, but the entire product roadmap. “Payback's a bitch,” Keith bristles in his editorial. The other major news this week is the rumour (via David Sacks) that OpenAI has offered the US government a 50% stake in a sovereign wealth fund. If true, this would change everything — not just in Silicon Valley, but in the political debate about public ownership of our AI economy. It's not just tech insiders like Sacks and Altman who are on board the sovereign wealth fund express, but also Bernie Sanders and other leftist critics of Big Tech. So maybe payback, at least when it comes to public investment in AI, isn't always such a bitch. Five Takeaways •       The Fake Andrew Keen: An Hour of Work on a Local Nvidia Card: Keith ran last week's show solo with an AI-generated Andrew Keen: trained on a few episodes of the show, animated from a YouTube still, scripted from Keith's newsletter. No third-party service. Just a local PC with an Nvidia GPU, about an hour of work, three attempts. Andrew, listening back, second-guessed whether he was actually there. The result was “pretty bad compared to our normal actual live shows,” Keith says. But also: really good. The question hanging over this episode and every future one: which Andrew are you listening to? •       Payback's a Bitch: How AI Companies Created Their Own Regulatory Trap: The US Commerce Department has told OpenAI and Anthropic they need government permission for who gets to use their latest models. Voluntary, for now. Keith's diagnosis: AI leadership spent more than a year crying wolf about existential risk — not because they believed it, but because government regulation creates a moat against competitors. Now the government has taken them at their word. Dario and Sam Altman wanted to be wrapped in government clothing. They are. The government now owns the door. They asked for this. They got it. •       American and Chinese State Capitalism: Converging Models: Andrew raises the macro argument: what we're watching is the convergence of American and Chinese models of capitalism toward a more state-centric model. China has always been explicit about state control. America has prided itself on free enterprise — even when the internet, atomic technology, and now AI were all substantially government-funded or government-shaped. Keith agrees at this level: all governments seek to control things they frame as dangerous. The difference is the framing. The direction of travel is the same. •       OpenAI's Rumoured 50% Stake Offer to the Government: Keith has heard — from sources including David Sacks, who should know — that OpenAI has offered the US government a very large stake, potentially 50%, in a sovereign wealth fund that would then distribute dividends to citizens. Sacks is not only unsurprised but in favour: he thinks 50% is too small. Andrew's question: why would OpenAI give away 50% of the company? Keith's answer: because it's the price of the regulatory moat. The government as partner rather than the government as regulator. A company that once aspired to “open” AI is now offering the state a controlling interest in its future. •       Paul Kennedy and America's Inevitable Decline: Keith has Paul Kennedy's Rise and Fall of the Great Powers on his shelf. His conclusion from it: it is historically impossible for America to retain its first-place status. No country ever has. Newly capitalised countries produce things more cheaply; China, India, and large parts of Asia are where most future growth will be. Does the AI boom change this? Keith's honest answer: no. It may slow the decline. It will not reverse it. America will, like an older gentleman on a rocking chair outside the house, accept its fate. Europe won't even be in the rocking chair. About the Guest Keith Teare is a British-American entrepreneur, investor, and publisher of the That Was the Week newsletter. He is a co-founder of TechCrunch and Andrew's regular TWTW co-host. References: •       That Was the Week by Keith Teare — the newsletter on which this episode is based. •       Azeem Azhar, The Exponential View — his report quantifying the AI economy at roughly $175 billion, referenced in the closing section. •       Alex Lazarow, 99%Tech — referenced for his piece on the emergence of an AI trust layer, the “Lloyds of AI.” •       Paul Kennedy, The Rise and Fall of the Great Powers — on Keith's shelf; referenced in the America-China decline section. •       David Sacks — referenced as the source for the OpenAI sovereign wealth fund rumour. About Keen On America Nobody asks more awkward questions than the Anglo-American writer and filmmaker Andrew Keen. In Keen On America, Andrew brings his pointed Transatlantic wit to making sense of the United States — hosting daily interviews about the history and future of this now venerable Republic. With nearly 3,000 episodes since the show launched on TechCrunch in 2010, Keen On America is the most prolific intellectual interview show in the history of podcasting. WebsiteSubstackYouTubeApple PodcastsSpotify Chapters: (00:38) - Introduction: the fake Andrew Keen from last week (01:14) - Keith explains how he did it: local Nvidia card, one hour, three attempts (02:11) - The big story: Commerce Department tells OpenAI and Anthropic they need permission ...

iRacers Lounge
Low Hanging Fruit - Episode 0538

iRacers Lounge

Play Episode Listen Later Jun 26, 2026 115:59


On today's show: how many newton meters do you really need?, a Nvidia GPU upgrade guide, Coke Series drama from Dover, multiclass starts…soon?, Watkins Glen 6 Hour mayhem, do you need an iRacing plugin for your Stream Deck?, what happened with the Simagic Zeus launch? and why was iRacing crucial for last weeks NASCAR race. So sit back, relax and join us on the iRacers Lounge Podcast. iRacers Lounge Podcast is available on iTunes and Apple's Podcasts app, Stitcher, TuneIn, Google Play Music, Spotify, Soundcloud, Podbean, Spreaker, Podbay, PodFanatic, Overcast, Amazon, and other podcast players. Sponsors: Hosts: Mike Ellis – https://x.com/mikedeanellis David Hall – https://x.com/dmixmage Greg Hecktus – twitter.com/froozenkaktus Donnie Spiker – https://www.instagram.com/spikerman19/ Brad Wrenn – https://x.com/bradwrenn John Kerley – https://x.com/KerleyJohnE Justin Pearson – https://www.facebook.com/justin.pearson.5811 Bobby Jonas – https://x.com/bjonas71 William Westbrook – https://www.facebook.com/william.westbrook.35 Kris Randall – https://www.signalcraftstudio.com/ https://www.facebook.com/profile.php?id=61579217706429 Links: Facebook – www.facebook.com/iRacersLounge/ Twitter – twitter.com/iracerslounge Instagram – instagram.com/iracersloungepodcast/ Web (Show Notes) – iracerslounge.com/

GREY Journal Daily News Podcast
Will Anthropic Alumni Reshape AI Tools for Scientists?

GREY Journal Daily News Podcast

Play Episode Listen Later Jun 25, 2026 1:16


The Wall Street Journal reported on June 24, 2026, that former Anthropic employees launched a startup aimed at helping scientists develop their own AI systems. Anthropic, led by CEO Dario Amodei and President Daniela Amodei, received up to $4 billion from Amazon in 2023 and at least $300 million plus additional financing reported as up to $2 billion from Google. The new venture targets researcher needs around data control, reproducibility, and deployment. Alternatives include closed APIs from OpenAI, Anthropic, and Google DeepMind, and open-source options from Meta and Mistral with tooling from Hugging Face, Databricks, and Weights & Biases. Compute considerations center on Nvidia GPUs via AWS, Google Cloud, and Azure. Sales into universities and pharma will require compliance, security reviews, and marketplace channels. Founders should watch for product details, partnerships, and pricing as indicators of viability.Learn more on this news by visiting us at: https://greyjournal.net/news/ Hosted on Acast. See acast.com/privacy for more information.

The Six Five with Patrick Moorhead and Daniel Newman
Apple's Siri Bet on Gemini, SpaceX's $1.77T IPO, and Claude Fable 5's Hyperscaler-Neutral Launch

The Six Five with Patrick Moorhead and Daniel Newman

Play Episode Listen Later Jun 15, 2026 64:35


Patrick Moorhead and Daniel Newman cover Tim Cook's final WWDC as CEO and Apple's Gemini-powered Siri strategy, the $35 billion Apollo and Blackstone deal backing Anthropic's capacity expansion, Intel's packaging wins with Google and NVIDIA, SpaceX's IPO at a $1.77 trillion valuation, Anthropic's Claude Fable 5 and Mythos 5 launch across every major cloud, and earnings reactions from Oracle, Micron, and Adobe. The handpicked topics for this week are: Apple's Siri AI Will Run on Gemini, Closing Out Tim Cook's Final WWDC as CEO: At WWDC, Apple confirmed Siri AI will run on Gemini through a new billion-dollar per year, multi-year deal, while Apple's Foundation Model Cloud Pro runs on NVIDIA GPUs inside Google Cloud. The announcement marks Tim Cook's last WWDC as CEO before John Ternus takes over on September 1. Apple isn't building its own AI cluster or competing on CapEx. They're betting that by owning the consumption layer, backed by access to health data and private messaging through iMessage, Apple will have a moat that compute spending can't replicate. (The Decode) Apollo and Blackstone Close the Largest Private Credit Deal Ever Backing Anthropic's Capacity Expansion: A $35 billion deal, the largest private credit transaction on record, will fund Google TPU capacity tied to Anthropic's compute needs, with Broadcom backstopping senior debt tranches and Google backstopping lease payments. The structure treats compute as a lendable asset class and signals more than 20 gigawatts of demand still being built out through 2028. Circular financing between chipmakers, cloud providers, and AI labs has moved from controversial to standard practice. (The Decode) Intel's Foundry Wins Packaging Work on Google's TPUs, Not a Full Fab Deal: Reports that Intel landed a deal tied to Google and NVIDIA reframe what's actually being handed off. Intel gets the packaging work on over 3 million TPUs, the compute die stays with TSMC, and the I/O die is being negotiated with Samsung at 2nm. INTC rose 12% Monday. The deal represents a low-risk path for Intel to augment, not replace, TSMC, while raising questions about anti-competitive dynamics in the foundry market. (The Decode) SpaceX Becomes an AI Infrastructure Company With a $1.77 Trillion IPO: SpaceX's IPO priced amid oversubscribed demand, with its valuation now reflecting not just Starlink connectivity and launch dominance but a newly material AI business, including AI1 orbital data center tests planned for late 2027 and a $920 million per month Google compute contract running through 2029. A sum-of-the-parts breakdown of the connectivity, launch, and AI segments lands well short of the trading price, with the gap largely explained by confidence in Elon Musk's track record of execution. (The Decode) Anthropic Launches Claude Fable 5 and Mythos 5 Across Every Major Cloud: Anthropic shipped Claude Fable 5 and Mythos 5 with same-day availability across Snowflake, AWS Bedrock, Vertex AI, and Microsoft Foundry, pricing at $10 and $50 per million tokens. The hyperscaler-neutral distribution strategy lands ahead of Anthropic's anticipated IPO. The models represent a real step up in research capability over Opus 4.8, but they come with a significant change. Users no longer have the option to opt out of data sharing with Anthropic, a shift some enterprises, including Microsoft, are already responding to. (The Decode) Is SpaceX a Once-in-a-Generation Entry or the Top of the Market? One side argues SpaceX represents a generational opportunity on par with early Amazon or Netflix, with interplanetary travel and off-world resource extraction as the long-term payoff that justifies looking past current valuation math. The other side argues this is peak euphoria: a company trading at roughly 95 times sales, propped up in part by circular investment from Google into both SpaceX and its AI segment, with a steep drawdown likely before any sustained climb. (The Flip) The Chip and Security Trade Reverses From Broken to Bifurcated: The semiconductor sector posted its biggest single-day gain since 2020, with the SOX up 5% on Monday, June 8, as a prior selloff in names like Broadcom, CrowdStrike, and Palo Alto Networks fully reversed. Intel rose 12%, Marvell 10%, and Corning 7%. The rebound reframes the AI trade narrative from a broad breakdown to a split between winners and laggards within the same sector. (Bulls & Bears) Oracle Posts a Record Quarter, But the Market Focuses on a $50 Billion Funding Plan: Oracle delivered record revenue of $19.2 billion, up 21 %, with EPS of $2.11, beating estimates of $1.89. IaaS grew 93 %, the fastest pace among hyperscalers, and RPO hit $638 billion, up $85 billion quarter over quarter, including $75 billion in AI contracts. FY27 guidance of $90 billion was maintained, and EPS guidance was raised, yet the stock fell 5% after hours amid concerns about Oracle's capital spending plans. Oracle's AI cloud backlog now exceeds those of AWS, Google, and Microsoft, built heavily on commitments from Anthropic and OpenAI. (Bulls & Bears) Micron's Profit Trajectory Puts It in Google's Earnings Tier: Micron is projected to generate nearly as much profit in 2027 as Google, with Q2 revenue of $23.86 billion, up 22 % and beating estimates, and Q3 guidance of $33.5 billion in revenue, $19.15 EPS, and 81 % gross margin. The stock is up 776%, with Wall Street firms, including UBS, raising price targets. The open question is whether memory has broken its historically cyclical pattern given sustained AI demand. (Bulls & Bears) Adobe Beats Across the Board, But the Stock Drops on CEO Departure and Freemium Pivot: Adobe posted record revenue of $6.62 billion, up 13 % and beating consensus of $6.45 billion, with non-GAAP EPS of $5.96, topping estimates of $5.81. AI first ARR tripled year over year to over $500 million, with total ARR reaching $27.1 billion, and FY26 guidance was raised. The stock still fell 5.5 % after hours, driven by the CFO's departure to Marvell and market concern over a strategic shift toward freemium pricing that delays near-term profitability. (Bulls & Bears) Watch the full video at sixfivemedia.com, and be sure to subscribe to our YouTube channel so you never miss an episode. The Decode Apple WWDC- Apple Caves to Google AND NVIDIA — Siri AI Runs on Gemini ($1B/yr) + Apple Foundation Model Cloud Pro Runs on NVIDIA GPUs in Google Cloud; Tim Cook's Final WWDC as CEO Before John Ternus Succeeds Him Sept 1 https://www.cnbc.com/2026/06/08/apple-wwdc-2026-live-updates.html Google's $35B Infra Deal — Apollo + Blackstone Close the Largest Private Credit Deal Ever; Broadcom Backstops Senior Tranches; Google Backstops Lease Payments https://www.reuters.com/business/apollo-blackstone-back-anthropics-35-billion-capacity-expansion-new-broadcom-tie-2026-06-09/ Intel's Foundry Reportedly Wins Google Packaging (Not Full Fab) — The Information Reframed: 3M+ TPU Packaging by Intel, Compute Die Still TSMC, I/O Die Being Negotiated With Samsung 2nm; INTC +12% Monday; Pat Calls Out TSMC Anti-Competitive Risk https://www.trendforce.com/news/2026/06/09/news-intel-foundry-gains-momentum-as-google-reportedly-orders-3m-tpus-nvidia-evaluates-18a-for-multi-die-gpu-design/ SpaceX Becomes an AI Infrastructure Company — Friday IPO at $1.77T; AI1 Orbital Data Center Tests Late 2027; Google $920M/mo Compute Contract Through 2029 https://finance.yahoo.com/markets/stocks/articles/spacex-poised-history-record-75-100000402.html Anthropic Ships Claude Fable 5 + Mythos 5 — Same-Day Distribution Across Snowflake, AWS Bedrock, Vertex AI, Microsoft Foundry; Hyperscaler-Neutral by Design Ahead of IPO; $10/$50 per M Tokens https://www.anthropic.com/news/claude-fable-5-mythos-5 The Flip FOR: https://www.cnbc.com/2026/06/11/spacex-billionaire-investing.html AGAINST: https://www.nytimes.com/2026/05/20/technology/elon-musk-spacex-ipo.html Bulls & Bears The Chip + Security Tape Recovery — SOX +5% Monday June 8 (Biggest Day Since 2020); AVGO/CRWD/PANW Selloff Reversed; Intel +12%, Marvell +10%, Corning +7%; the AI Trade Pivots From "Broken" to "Bifurcated" https://www.investopedia.com/stock-market-today-dow-jones-s-and-p-500-06082026-11992852 Oracle (ORCL) Q4 FY26 ACTUALS — Record $19.2B Rev (+21%), EPS $2.11 Beat ($1.89); IaaS +93%; RPO HITS $638B (+$85B QoQ, $75B AI Contracts); FY27 $90B Guide Maintained, EPS Guide Raised; Stock −5% AH on Massive Capex Plan https://www.tradingkey.com/analysis/stocks/us-stocks/261959450-oracle-record-q4-2026-earnings-report-cloud-data-center-stock-tradingkey "$MU Will Generate Almost As Much Profit in 2027 as $GOOGL"; Q2 Rev $23.86B (+22% Beat), Q3 Guide $33.50B / $19.15 EPS / 81% GM; MU Stock +776%; UBS Among Wall Street Raising Targets https://247wallst.com/investing/2026/06/11/wall-street-just-put-a-monster-target-on-micron-is-the-stock-still-too-cheap/ Adobe (ADBE) Q2 FY26 ACTUALS — Record $6.62B Rev (+13%) Beats Consensus $6.45B; Non-GAAP EPS $5.96 Beats $5.81; AI-First ARR Triples YoY to $500M+; Total ARR $27.10B; FY26 Guide RAISED; Stock −5.5% AH Despite Beat-and-Raise https://www.businesswire.com/news/home/20260611677110/en/Adobe-Reports-Record-Q2-Results    

Room 101 by 利世民
SpaceX 食糊先定長揸?

Room 101 by 利世民

Play Episode Listen Later Jun 12, 2026 34:08


當地球上不斷興建 AI 數據中心,Elon Musk 卻說,真正的算力戰場在太空。SpaceX 的核心邏輯是卡爾達舍夫指數。人類尚在 Type I 掙扎,而 SpaceX 正在鋪路邁向Type II。未來幾年 SpaceX 將以戰養戰:以 Starlink 的技術為基礎,大量將以太陽能直接供電、以真空輻射散熱的 AI-1 衛星搬上低軌道;中長期是透過 TeraFab 晶片廠和長期軌道計劃,以成為世界最主要的 AI 運算力。相關風險就是:算力需求出現S曲線放緩、AI泡沫爆破、以及中國不計成本的競爭。不過,SpaceX 行者的垂直整合優勢目前仍難以複製。問:卡爾達舍夫指數如何理解人類文明的現狀?答:卡爾達舍夫指數以文明能夠駕馭的能量規模為單位,分為Type I(行星級)、Type II(恆星級)、Type III(星系級)三個階段。人類目前尚未達到Type I的完善狀態——使用的化石能源只是億萬年前儲存的陽光殘存,可再生能源的轉換效率亦相對低下。按Elon Musk的形容,人類現在處於Type II的初階起點,距離充分利用太陽系能量輸出還差極遠。問:SpaceX在全球航天市場擁有什麼結構性優勢?答:SpaceX透過火箭回收重用技術大幅壓低了發射成本,目前承攬全球約八至九成的入軌有效載重量。這種成本優勢構成了全球壟斷地位:剩餘約一成的市場主要是中國等因地緣政治原因不使用SpaceX的國家。成本護城河,而非技術獨特性,是SpaceX壟斷地位的根基。問:為何太空是部署AI運算力的理想場所?答:地面AI數據中心面臨兩大挑戰:能源供應和散熱冷卻。太空環境天然解決兩者——衛星可直接採集未經地球大氣衰減的太陽能,散熱可透過輻射向真空直接散逸,無需額外能耗。此外,在太空建設數據中心不佔用地面土地、不拉高當地電費,也不受地方社區反對,規避了地面部署的種種政治阻力。問:AI-1衛星的技術規格如何?3毫秒延遲有多大影響?答:AI-1衛星設計功率約150千瓦,相當於地面一個NVIDIA GPU機架,翼展約70米。運行在600至800公里低軌道,傳輸延遲約3毫秒。由於AI訓練和推論的計算大多在衛星本身完成、傳輸的實際數據量有限,這3毫秒延遲在工程上是可以接受的代價,不構成根本性障礙。問:TeraFab計劃的意義是什麼?答:TeraFab是SpaceX計劃在美國建立的超大型晶片工廠,面積約1億平方呎,目標是最終達到每年1太瓦(TeraWatt)的AI算力產量。這意味著SpaceX將自行生產晶片,從硬件製造到衛星發射形成垂直整合一條龍,減少對NVIDIA等外部供應商的依賴,並大幅壓低單位算力成本。問:SpaceX上市融資與AI策略有何關係?答:SpaceX以「以戰養戰」方式推進上市:以現有火箭發射和Starlink業務產生的穩定收入,支撐軌道AI算力這一高度資本密集的長期計劃。上市的核心敘事是:Google等科技巨頭已向軌道算力下注,SpaceX作為唯一能把太空發射、衛星製造、晶片生產集於一身的公司,理應以更高估值來反映這個整合價值。問:SpaceX面臨哪三大主要風險?答:第一,AI對算力的需求可能在推論(inferencing)應用階段增長放緩,遠低於訓練階段的消耗規模;第二,AI投資週期存在S形曲線,高速增長後可能減速,市場時機難以預判;第三,中國有能力以不計成本的方式進入太空發射市場,複製火箭回收重用技術,動搖SpaceX當前的成本優勢——正如電動車市場的故事已提供了先例。 This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit leesimon.substack.com/subscribe

The Information's 411
Inside OpenAI's IPO Filing, Former SpaceX Engineer on IPO, Apple's AI Siri Upgrade

The Information's 411

Play Episode Listen Later Jun 10, 2026 43:25


The Information's OpenAI reporter Erin Woo details the confidential IPO filings of OpenAI and Anthropic, highlighting how Anthropic has eclipsed OpenAI in enterprise revenue. Apple reporter Aaron Tilley and Constellation Research CEO Ray Wang break down Apple's WWDC announcements, evaluating its Siri reboot powered by Google's Gemini models and Nvidia GPUs. Finally, AI finance reporter Dakin Campbell joins to discuss how Wall Street titans Goldman Sachs and JPMorgan are exploring derivatives markets to trade the cost of GPU computing power.Articles discussed on this episode: https://www.theinformation.com/briefings/openai-confidentially-files-ipo-paperwork-plans-separate-employee-share-salehttps://www.theinformation.com/newsletters/the-briefing/apples-cautious-ai-overhaul-openais-ipo-filinghttps://www.theinformation.com/articles/goldman-jpmorgan-explore-new-ways-tame-ai-lending-risksSubscribe: YouTube: https://www.youtube.com/@theinformation The Information: https://www.theinformation.com/subscribe_hSign up for the AI Agenda newsletter: https://www.theinformation.com/features/ai-agendaTITV airs weekdays on YouTube, X and LinkedIn at 10AM PT / 1PM ET. Or check us out wherever you get your podcasts.Follow us:X: https://x.com/theinformationIG: https://www.instagram.com/theinformation/TikTok: https://www.tiktok.com/@titv.theinformationLinkedIn: https://www.linkedin.com/company/theinformation/

TechLinked
Microsoft 365 Update, GitHub Breach, Nvidia GPU Security Update + more!

TechLinked

Play Episode Listen Later May 23, 2026 7:53


Timestamps: 0:00 Microsoft 365 Update 1:02 GitHub Breach 1:58 Nvidia GPU Security Update 4:06 QUICK BITS INTRO 4:11 Meta Releases Reddit-like App 4:34 Riot's Vanguard Anti-Cheat Hits Cheaters Hard 5:05 Discord Enables End-to-End Encryption 5:31 Spotify and UMG Reach AI Deal 6:14 Pizza Hut Sued by Franchisee Over AI NEW SOURCES: https://lmg.gg/Bdpll Learn more about your ad choices. Visit megaphone.fm/adchoices

What's new in Cloud FinOps?
WNiCF - April 2026 - News

What's new in Cloud FinOps?

Play Episode Listen Later May 22, 2026 54:58


Send us Fan MailWhat's New in Cloud FinOps: May 2026 Monthly RecapIn this combined monthly recap for May 2026, Frank Contrepois and Stephen Old dive into a vast array of updates across AWS, Google Cloud, and Azure, with a special focus on the evolving landscape of AI FinOps, hybrid cloud challenges, and a barrage of storage news.The Expanding Scope of FinOps: From Data Centre to AIThe discussion opens by exploring the expansion of FinOps beyond the public cloud to encompass on-premise data centres, software, AI, and sustainability. A central theme is the application of the FinOps Open Cost and Usage Specification (FOCUS) to on-premise environments. Stephen shares firsthand experience transposing software data into FOCUS to create a converged platform, highlighting the fundamental data challenges, from ingesting contract data to managing the high velocity of cloud data.The conversation then shifts to the burgeoning role of AI, noting its inclusion alongside SaaS and professional services in the modern FinOps scope. This introduces new forecasting challenges, as traditional 18-month budget cycles clash with the rapid pace of weekly AI model releases.A critical point is also raised regarding sustainability. The hosts discuss Amazon's board rejecting a shareholder proposal for detailed climate disclosures, which poses a significant challenge for companies needing granular data for CSRD and SEC compliance.Major Cloud Updates: April 2026AI & FinOps Visibility:A major theme is the improvement in attributing AI spend. A game-changing update from AWS means Bedrock API calls now automatically record the IAM identity (user or role) of the caller directly into CUR 2.0 and Cost Explorer. This eliminates the complex need to reconcile CloudTrail logs to determine who is driving Bedrock costs.Similarly, Amazon Q is now embedded in the AWS Cost Explorer, allowing users to ask natural language questions about their spending (e.g., "Why did my RDS costs spike last month?"). This conversational analysis approach comes with a free tier of 50 queries per month.On the Google Cloud side, a new billing overview widget for Gemini and Vertex AI spend is now in preview. Google is also introducing a "FinOps Explainability Agent," an autonomous AI agent to investigate AI cost drivers, and "Spend Caps" (Private Preview) for services like AI Studio and Vertex AI, which provide crucial cost control by pausing API traffic when a budget is hit.For those managing GPU workloads, Amazon ECS managed instances now support NVIDIA GPU metrics in CloudWatch Container Insights, enabling real-time visibility into GPU utilisation and health to optimise expensive accelerated computing.Cost & Usage Reporting (CUR) Enhancements:There are hints of a potential enhancement to AWS CUR 2.0, which could see new columns added to directly link API calls with costs, revolutionising cost allocation. AWS has also introduced:Scheduled Email Delivery for Billing Dashboards: Securely send reports to stakeholders without console access.Billing Conductor Pass-Through Plan: Simplifies centralised billing for billing transfer users.Cost Optimization Hub CSV Downloads: Easily export savings recommendations.Find out how to leverage CUR for security: "Identifying security risks using AWS cost and usage report data"Compute & Database Innovations:AWS: Released a wave of 8th Generation Intel Instances (C8i, M8i, R8i and network-optimised versions) powered by custom 6th Gen Xeon processors. EC2 Capacity Manager also now supports tag-based dimensions, allowing for more granular capacity optimisation. Amazon Aurora Serverless now boasts up to 30% better performance and, crucially, scales down to zero, a cost-effective option for unpredictable agentic AI workloads.Google Cloud: At Google Cloud Next, they announced both ends of the performance spectrum. The 8th Generation TPUs (v8t for training, v8i for inference) offer massive scale and performance-per-dollar improvements. In a move to democratise access, Google also made fractional GPUs (1/2, 1/4, or 1/8) on the G4 series generally available, a game-changer for cost-effectively running smaller workloads. The GKE workload recommender is also now integrated into the FinOps Hub.Azure: Now supports NVIDIA's powerful H100 and H200 GPUs on Azure Red Hat OpenShift (ARO) for large-scale AI/HPC workloads. For database users, the GA of Premium SSD v2 for Azure Database for PostgreSQL promises significantly higher IOPS and better price-performance.A Deep Dive into Azure Storage:The episode covers an "overload" of Azure storage updates with significant FinOps implications:Minimum Billable Object Size: From 1st July 2026 for new accounts (and 2027 for all), objects smaller than 128KB in cool, cold, and archive tiers will be billed as if they are 128KB.Smart Tier for Azure Blob & ADLS (GA): To mitigate the above, this feature automatically tiers data based on access patterns but introduces a monitoring fee for objects over 128KB, creating a new optimisation puzzle.Azure NetApp Files (ANF) Ransomware Protection: Now GA and included as part of the service at no extra charge.Finally, the hosts tackle "The Big Silence on Memory Prices," noting that despite DDR memory prices soaring 300-400% from mid-2025 lows, the hyperscalers have remained silent, absorbing the cost and making it difficult for smaller providers to compete.Explore the official announcements:AI Bill of Materials Whitepaper: www.wiz.io/go/ai-security/ai-bill-of-materialsAWS Article on Amazon Q: https://aws.amazon.com/blogs/aws-cloud-financial-management/transforming-finops-with-the-latest-amazon-q-cost-capabilities/ 

Gestalt IT Rundown
NVIDIA Backyard Datacenters, Apple + Intel, & SpaceX | Tech Field Day News Rundown: May 13, 2026

Gestalt IT Rundown

Play Episode Listen Later May 13, 2026 33:57


AI infrastructure is evolving faster than ever, and this week's biggest tech stories could reshape the future of cloud, compute, and enterprise IT. This week on the Tech Field Day News Rundown, Tom Hollingsworth and Alastair Cooke disucss Mirantis being acquired by IREN to accelerate open AI infrastructure, Lumen's acquisition of Alkira to modernize hybrid networking, and Red Hat's push toward AI-driven automation across RHEL and Lightspeed.They also discuss Anthropic's massive SpaceX compute deal powered by over 220,000 NVIDIA GPUs, Apple's reported shift toward Intel manufacturing for select chips, China's new neutral atom quantum computer, and Span's ambitious plan to turn neighborhoods into distributed AI data centers. From enterprise AI and cloud networking to quantum computing and next-generation infrastructure, this episode covers the biggest trends shaping the future of technology.This and more on the Tech Field Day News Rundown with Tom Hollingsworth and Alastair Cooke.Time Stamps: 0:00 - Cold Open0:25 - Welcome to the Tech Field Day News Rundown 1:21 - Mirantis Has Been Acquired by IREN4:21 - Lumen to Acquire Alkira7:47 - Red Hat Embraces MCP to Drive AI Agents for IT Operations Automation in RHEL12:31 - Anthropic Secures Massive Compute Deal with SpaceX, Expands Toward Orbital AI Infrastructure16:16 - Intel to Manufacture Apple Chips in New Supply Chain Shift, Report Says19:47 - China Unveils Hanyuan-2, First Dual-Core Neutral Atom Quantum Computer23:16 - NVIDIA and Span Turn Suburbs Into Distributed AI Data Centers With Backyard Compute Nodes31:11 - The Weeks Ahead 33:01 - Thanks for Watching the Tech Field Day News RundownGuest Host: ⁠⁠⁠⁠⁠Vincent Celindro⁠, Director of Strategic Sales and Technology, Quantum Foundry Follow our hosts ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Tom Hollingsworth⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠, ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Alastair Cooke⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠, and ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Stephen Foskett⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠. Follow Tech Field Day ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠on LinkedIn⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠, on ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠X/Twitter⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠, on ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Bluesky⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠, and on ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Mastodon⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠.

Everyday AI Podcast – An AI and ChatGPT Podcast
Ep 774: Anthropic's Dev Day releases, OpenAI's new model drop, AI labs agree to federal testing and more AI News That Matters

Everyday AI Podcast – An AI and ChatGPT Podcast

Play Episode Listen Later May 11, 2026 39:30


Nearly a billion people will be using a new AI model this week, and hardly any of them will notice. Sheesh. That's how important it is to keep up with the latest in greatest in AI. Aside from OpenAI's new GPT-5.5 Instant release, this week we saw both AI drama and battles, as well as new capabilities and studies that detail it all. If you can't keep up with the daily breakneck speed of AI, make sure you join us weekly for our AI News That Matters where we keep you in the loop in a fraction of the time. Newsletter: Sign up for our free daily newsletterMore on this Episode: Episode PageToday's Episode on LinkedIn: Thoughts on this? Join the convo on LinkedIn and connect with other AI leaders.Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineupWebsite: YourEverydayAI.comEmail The Show: info@youreverydayai.comConnect with Jordan on LinkedInTopics Covered in This Episode:SpaceX Leases Colossus Supercomputer to AnthropicCloudflare Restructures, Cuts 20% Jobs for AIApple Settles $250M AI Feature LawsuitOpenAI Releases GPT-5.5 Instant Model UpdateChatGPT Personalized Context and Memory FeaturesMicrosoft 2026 Work Trend Index on AI AdoptionWhite House Moves Toward Federal AI Model TestingOpenAI, Anthropic Launch AI Consulting ServicesOpenAI Voice Models, Real-Time Translation RolloutGoogle DeepMind Staff Unionize Against Military AIAnthropic Expands Claude to Office SuiteAdobe Launches AI-Powered PDF Collaboration ToolsTimestamps:00:00 Elon Musk leases supercomputer to Anthropic03:19 SpaceX's AI leasing deal08:10 Cloudflare revenue miss and job cuts09:43 AI job cuts and stock reactions13:00 Apple's class action settlement details18:48 OpenAI releases GPT 5.5 instant20:22 New memory features and GPT updates23:36 Microsoft AI adoption study findings28:40 New AI safety agreements30:20 OpenAI and Anthropic expansion plans33:38 AI adoption challenges for companies36:57 Tech company updates and product launches40:01 Focusing on trustworthy AI in financeKeywords: Anthropic, Anthropic Dev Day, Anthropic open source, Claude, Claude AI, Claude uptime, OpenAI, GPT-5.5, GPT-5.5 Instant, GPT-5.5 Pro, GPT-4.0, ChatGPT, OpenAI new model, OpenAI vs Anthropic, XAI, Elon Musk, SpaceX, Colossus 1 supercomputer, NVIDIA GPUs, AI compute,Send Everyday AI and Jordan a text message. (We can't reply back unless you leave contact info)

Techmeme Ride Home
XAI Is Just A Neocloud Now?

Techmeme Ride Home

Play Episode Listen Later May 7, 2026 22:32


Dario Amodei revealed Anthropic could grow 80x in 2026, and the company signed a deal with SpaceX for 300MW of compute from Colossus 1. Musk dissolved xAI into SpaceX. Google launches the $100 Fitbit Air, and HubSpot's founder coins "strategic illegibility." At Code with Claude, Dario Amodei said Anthropic had planned to grow ~10x in 2026 but could grow 80x, calling its growth rate "crazy" and "too hard to handle" (NYT) Anthropic signs a deal with SpaceX for 300MW+ of compute from Colossus 1 in Memphis, accessing 220,000+ Nvidia GPUs within the month (Bloomberg) Musk says xAI will be "dissolved as a separate company" and will become "SpaceXAI, the AI products from SpaceX" (Spyglass) Google launches the $100 Fitbit Air, a Whoop-like screenless wearable, with Gemini-powered features like Google Health Coach, available May 26 (Engadget) As founders race to make their companies "legible" to AI, they must keep the things that make them hard to copy "illegible", or risk commoditizing their moat (Brian Halligan) Learn more about your ad choices. Visit megaphone.fm/adchoices

IEN Radio
LISTEN: Corning to Build Three Factories, Create 3,000 Jobs in Massive AI Deal With Nvidia

IEN Radio

Play Episode Listen Later May 7, 2026 2:11


This morning, Corning and NVIDIA announced a multiyear commercial and technology partnership to significantly scale U.S.-based optical connectivity manufacturing and fiber production. To meet accelerating demand for new AI factory buildouts (data centers), Corning will increase optical connectivity capacity by 10x and boost U.S. fiber production by more than 50%. These are the components NVIDIA needs to power next-generation AI infrastructure. To make it happen, Corning will build three new advanced manufacturing facilities in North Carolina and Texas, which are expected to create more than 3,000 new jobs.Corning will manufacture the optical connectivity new hyperscale data centers will use to deploy accelerated computing at scale. Modern AI workloads require thousands of NVIDIA GPUs which in turn demands unprecedented volumes of high-performance optical fiber, connectivity and photonics to move data at extraordinary speed and scale, according to the companies. #manufacturing, #AI, #datacenters, #opticalfiber, #photonics, #NVIDIA, #Corning, #advancedmanufacturing, #supplychain, #semiconductors, #infrastructure, #madeinusa, #factorybuildout, #techmanufacturing, #innovation, #industry40, #automation, #engineering, #futureofwork, #digitalinfrastructure

Linux Weekly Daily Wednesday
Steam Is Getting ARMed

Linux Weekly Daily Wednesday

Play Episode Listen Later Apr 22, 2026 47:06


Steam releases Proton for ARM plus a native Linux client, Davinci Resolve bring professional photo editing to Linux, running NVIDIA GPUs on Rockchip SBCs, and Firefox get a native GTK emoji picker.Patreon: ⁠⁠⁠https://www.patreon.com/lwdw⁠⁠⁠⁠⁠⁠⁠Discord: ⁠⁠⁠⁠⁠⁠⁠https://discord.gg/uQVckr5gEZ⁠⁠⁠⁠⁠⁠⁠TOPICSResolve 21 Editor https://petapixel.com/2026/04/16/the-davinci-resolve-21-photo-editing-tools-show-promise-but-are-imperfect/ARMing Steam Protonhttps://www.techpowerup.com/348297/steams-proton-gets-wine-11-gaming-performance-improvements-valve-launches-arm64-compatibility-layerFirefox 150

Keen On Democracy
Forget Iran: Eyck Freymann on Taiwan, China, and Why America Keeps Hitting the Snooze Button,

Keen On Democracy

Play Episode Listen Later Apr 13, 2026 44:58


“We keep getting wake-up calls and snoozing the alarm. Now is the time to actually get out of bed and confront this problem before it is too late.” — Eyck Freymann Forget Iran for a moment. The Hormuz crisis is a template for the bigger crisis of Taiwan. Eyck Freymann — Hoover Fellow at Stanford, author of the brand-new Defending Taiwan: A Strategy to Prevent War with China — believes that the fate of the 21st century may hinge on Taiwan. And he warns that if America can't handle Iran, it's certainly not ready for Beijing. Freymann argues that China doesn't need to invade Taiwan. Xi Jinping has watched Putin discover — with horror — what happens when you send unprepared forces into a country that fights back. China's lesson from Ukraine is a strategy of quarantine rather than invasion. The United States will then face a choice between accepting Chinese checkmate or escalating a crisis with no domestic or international support. Taiwan produces 90% of the world's advanced semiconductors and 99% of the cutting-edge NVIDIA GPUs used to train frontier AI models. If those chip factories shut, there will be an instantaneous global financial crisis. Forget today's Iranian theater. Taiwan will be the real existential show. Five Takeaways •       The Hormuz Alarm Bell: Iran has no navy, no air force, and supposedly no ballistic missile arsenal anymore — and yet it took 20% of global oil supply offline. The Trump administration went in thinking overwhelming military superiority would translate to political victory. It hasn't. Strategy, Freymann says, is the art of connecting ends to means. If you don't know your ends, you'll flail. China is watching every mistake: no plan for the economic shock, no domestic legitimacy for the war, excess pain falling on oil-importing US allies like Japan, South Korea, and Europe. Beijing's conclusion: we don't have to pick a military fight with the United States. Why would we? •       The Semiconductor Chokehold: Taiwan produces 90% of the world's advanced semiconductors and 99% of the cutting-edge NVIDIA GPUs used to train frontier AI models. The CHIPS Act has tried to change this. It hasn't. The Arizona facility is two generations behind Taiwan, commercially uncompetitive, and unable to scale. Taiwan is five years ahead now and will be five years ahead in five years. If the Taiwan fabs go offline, there is an instantaneous global financial crisis: the seven companies that account for roughly 40% of the S&P 500 are all essentially the AI trade. The hyperscalers are spending $600 billion in data centers this year — the only thing keeping the US economy out of recession. This is what's at stake, before you even get to the military question. •       The Quarantine: Winning Without Fighting: Xi Jinping's plan A is not invasion. It's the quarantine: seize control of who and what comes and goes to Taiwan by declaring that anyone flying to Taipei must first clear customs in Shanghai. Impound a United Airlines flight. Let the ambiguity do the work. If China can do that and get away with it, Taiwan can't rebuild its military, the US can't send more weapons, and Beijing controls the chips. It's checkmate — without a shot fired. The United States then has to accept it, or escalate in a way that has no domestic legitimacy and drives wedges between Washington and its allies. China has figured out how to extort the West with prolonged economic pain. The alarm bells keep ringing. America keeps snoozing. •       What a Taiwan War Would Actually Look Like: It would be a war at sea — fundamentally unlike anything America has fought or prepared for in eighty years. China would need to simultaneously control the skies, the undersea, and the surface on all sides of the Taiwan Strait, then send tens of thousands of men 80 miles across in amphibious vessels to storm beaches in a Normandy-style assault. The first engagements would be decided in minutes to hours by long-range precision munitions. America's operational capabilities are exceptional: the cyber assassinations, the special forces raid, the continuous bomber sorties from the continental United States. But China has home-field advantage. And it has been building systematically for this scenario for years. We could probably win if we fought today. We need to make investments for tomorrow. •       The Four-Pillar Strategy: Freymann's integrated answer: diplomacy, military deterrence, economic resilience, and allied coordination — all working together, not in separate silos. On diplomacy: maintain the principled position that Taiwan's status must be resolved peacefully and democratically. On military: show China it can't win if it escalates to war, while keeping conventional forces credible. On economics: build enough allied resilience that authoritarian powers can't extort the West by threatening prolonged economic pain. On allies: coordinate with Japan, South Korea, the Europeans on a shared plan for what happens if things collapse. This is doable. It's been done for fifty years. We just need the resolve to keep doing it. About the Guest Eyck Freymann is a Hoover Fellow at Stanford University and a Non-Resident Research Fellow at the US Naval War College's China Maritime Studies Institute. He is the author of Defending Taiwan: A Strategy to Prevent War with China (Oxford University Press, 2026), The Arsenal of Democracy: Technology, Industry, and Deterrence in an Age of Hard Choices (Hoover, 2025), and One Belt One Road: Chinese Power Meets the World (Harvard, 2021). References: •       Defending Taiwan: A Strategy to Prevent War with China by Eyck Freymann (Oxford University Press, 2026). •       “The Strait of Hormuz as a Template for Taiwan,” Financial Times, April 2026. By Eyck Freymann. •       Episode 2862: Truth Is Dead — on AI, disinformation, and American strategic confusion. About Keen On America Nobody asks more awkward questions than the Anglo-American writer and filmmaker Andrew Keen. In Keen On America, Andrew brings his pointed Transatlantic wit to making sense of the United States — hosting daily interviews about the history and future of this now venerable Republic. With nearly 2,800 episodes since the show launched on TechCrunch in 2010, Keen On America is the most prolific intellectual interview show in the history of podcasting. WebsiteSubstackYouTubeApple PodcastsSpotify 

Forbes Talks
CoreWeave Announces Deal With Anthropic

Forbes Talks

Play Episode Listen Later Apr 10, 2026 4:03


Follow Topline Coreweave shares popped 13% after announcing a deal with Anthropic on Friday to power its AI model Claude, following a $21 billion partnership with Meta announced Thursday. KEY FACTS Financial terms of the multi-year Anthropic agreement were not disclosed, though the deal marks the ninth of ten leading AI providers—including OpenAI, Google, Microsoft and Meta—using CoreWeave's platform, according to a Friday press release. The Anthropic news comes one day after CoreWeave announced a $21 billion deal to supply Meta with AI cloud capacity through December 2032, delivered from multiple data centers powered in part by Nvidia chips.  The back-to-back announcements pushed CoreWeave's total contracted commitments with Meta alone to $35 billion, with the new Meta pact building on a prior $14.2 billion arrangement. Anthropic is the latest AI model developer to become a customer, highlighting the scramble among tech companies to secure more hardware, processing power and energy—key for training and deploying increasingly complex AI models. KEY BACKGROUND CoreWeave primarily generates revenue by building and renting out data centers packed with Nvidia GPUs that provide the energy and processing power to train and run AI models. Demand for infrastructure to develop AI has exploded since the release of ChatGPT in 2022, with Alphabet, Microsoft, Meta and Amazon committing a combined $700 billion just this year in a race to build the most sophisticated and advanced models. On Tuesday, Anthropic announced that its leaked Mythos model was so powerful that they would be holding back from releasing it to the public because of its ability to find vulnerabilities in software programs. The Claude maker said it would instead provide the model to 40 select companies including Apple, Amazon, Google and Microsoft in a cybersecurity initiative dubbed Project Glasswing. Anthropic was founded in 2021 by siblings Dario and Daniela Amodei and several former OpenAI employees who departed the ChatGPT maker over concerns about the company's direction with AI safety. Anthropic is now valued at $380 billion and announced it had reached an annual revenue run rate of $30 billion Monday, surpassing OpenAI's $25 billion annualizedrevenue as of February. OpenAI is now valued at $852 billion.  BIG NUMBER $2.5 trillion. That's how much research firm Gartner expects global spend to build AI will reach in 2026, up 44% from last year. AI infrastructure will drive the spend, making up more than half of that figure, the firm estimates. TANGENT The deals come as CoreWeave is simultaneously on an aggressive financing spree. The company is targeting $30 billion to $35 billion in capital expenditures for 2026, up from roughly $15 billion in 2025. Billionaire CEO Mike Intrator defended the spending strategy after the company's February earnings report drew criticism for the increase. "I understand the concerns that people have as they see us allocating a massive scale of money to this market, but the truth of the matter is, our backlog is enormous," he told CNBC at the time. Since going public in March 2025, the stock is up 160%, but is down nearly 45% from its peak last June. This year, the stock has been volatile, up 30% since January. Read the full story on Forbes: By Alicia Park https://www.forbes.com/sites/aliciapark/2026/04/10/coreweave-stock-surges-13-on-anthropic-deal-a-day-after-21-billion-meta-partnership/ Learn more about your ad choices. Visit megaphone.fm/adchoices

This Week in Startups
$2.5B Chip Heist, The Future of American AI, and Purpose-Built Robots | This Week in AI Ep 6

This Week in Startups

Play Episode Listen Later Mar 25, 2026 75:24


This Week in AI sneak peak! If you enjoy the episode find us on Spotify, Apple podcasts and YouTube by looking up "This Week in AI" or by going to thisweekinai.aiThis week Jason sat down with Jake Loosararian and Chris Lattner on Episode 6 of This Week in AI. Jake is the CEO and co-founder of Gecko Robotics, a company deploying purpose-built robots and AI for mission-critical infrastructure inspection across energy, defense, and manufacturing. Chris is the CEO and co-founder of Modular, building a universal software layer that lets developers run AI models across Nvidia, AMD, and Apple silicon without being locked into any single hardware vendor.We explore the GPU shortage, why China's chip smuggling reveals the stakes of the AI cold war, how purpose-built robotics are beating humanoids on ROI, the case for American reindustrialization, and why the next decade could be the best ever for private equity in capital-intensive industries.Purpose-Built Robots vs. Humanoids: Jake has been building mission-critical robots for 13 years. He explains why general-purpose humanoids still have too little ROI for industrial use, and why specialized robots that find and fix problems are winning in the field.The GPU Shortage Is Real: Chris breaks down why you can't just go buy 100 Blackwell chips today, why Nvidia's Cuda creates massive lock-in, and how Modular is building a unified software layer across all major chip architectures.Google TPUs Are the Sleeper: Chris ranks Google as the number one threat to Nvidia's dominance, ahead of Amazon's Trainium and AMD.China's Chip Smuggling & the AI Cold War: A Supermicro co-founder allegedly smuggled $2.5B in Nvidia chips to China using fake serial numbers and a hairdryer. The Best Decade for Private Equity: Jake makes the case that capital-intensive, commoditized infrastructure assets: waste-to-energy, water treatment, old power plants will all generate incredible returns.Self-Driving State of Play: Chris, a former Tesla Autopilot lead, gives his read on Waymo's lead, Tesla's small Austin pilot, and why the real signal is when Tesla starts filing for fully autonomous permits in California.Learn more about Gecko Robotics: https://www.geckorobotics.comLearn more about Modular: https://www.modular.com/This Week In AI is made possible by:*PayPalOpen* - One Platform for all Business: paypalopen.com*Timestamps:*00:00 Welcome & intro to Jake Lu (Gecko Robotics) and Chris Lattner (Modular)01:34 Gecko's 13-year journey & the Cantilever platform05:15 Chris Lattner on Modular: replacing Cuda & unifying AI hardware11:10 Nvidia lock-in, AMD's Rock & why the software stack is broken19:49 The GPU shortage: how real is it?22:13 Who challenges Nvidia? Google TPUs, Amazon Trainium & AMD ranked28:17 China chip smuggling: $2.5B in Nvidia GPUs & the AI cold war37:43 Self-driving update: Waymo, Tesla's Austin pilot & Chris's Tesla history42:20 Figure's humanoid package sorting — real or demo magic?43:47 The best decade for private equity in capital-intensive assets51:04 Reindustrialization, the trades boom & making manufacturing cool58:39 Building tech companies outside Silicon Valley1:06:46 Breaking news: Brett Adcock launches Hark from Figure1:10:15 Closing thoughts: grit over hype, customers over valuationsSubscribe to This Week in AI on Apple: https://thisweekinai.ai/spotifySubscribe to This Week in AI on Spotify: https://thisweekinai.ai/appleThanks for watching!

Risky Business
Risky Business #830 -- LiteLLM and security scanner supply chains compromised

Risky Business

Play Episode Listen Later Mar 25, 2026 63:53


On this week's show, Patrick Gray, Adam Boileau and James WIlson discuss the week's cybersecurity news. They talk through: TeamPCP's supply chain attack on Github, and they threw in an anti-Iran wiper, because why not?! Anthropic hooks up its models to just… use your whole computer After Stryker's Very Bad Day, CISA says maybe add some more controls around your Intune? Another iOS exploit kit shows up in the cyber bargain-bin The FTC decides to ban… all new home routers?! U wot m8?! Supermicro founder was personally sanction-busting Nvidia GPUs into China?! This week's episode is sponsored by enterprise browser maker, Island. Chief Customer Officer Bradon Rogers joins Pat to explain how its customers are using Island to control the use of personal AI services in regulated industries. This episode is also available on Youtube. Show notes ‘CanisterWorm' Springs Wiper Attack Targeting Iran TeamPCP deploys CanisterWorm on NPM following Trivy compromise Andrej Karpathy on X: "Software horror: litellm PyPI supply chain" attack Checkmarx KICS GitHub Action Compromised: Malware Injected in All Git Tags Felix Rieseberg on X: "Today, we're releasing a feature that allows Claude to control your computer" A Top Google Search Result for Claude Plugins Was Planted by Hackers Lockheed Martin targeted in alleged breach by pro-Iran hacktivist CISA urges companies to secure Microsoft Intune systems after hackers mass-wipe Stryker devices FBI seems to seize website tied to Iranian cyberattack on Stryker Stryker confirms cyberattack is contained and restoration underway Hundreds of Millions of iPhones Can Be Hacked With a New Tool Found in the Wild Someone has publicly leaked an exploit kit that can hack millions of iPhones Russia-linked hackers use advanced iPhone exploit to target Ukrainians Apple rolls out first 'background security' update for iPhones, iPads, and Macs to fix Safari bug Post by @wartranslated.bsky.social — Bluesky Signal's Creator Is Helping Encrypt Meta AI Hacker says they compromised millions of confidential police tips held by US company Millions of 'anonymous' crime tips exposed in massive Crime Stoppers hack Feds Disrupt IoT Botnets Behind Huge DDoS Attacks FCC bans import of consumer-grade routers amid national security concerns White House pours cold water on cyber ‘letters of marque' speculation Google launches threat disruption unit, stops short of calling it ‘offensive' Supermicro's cofounder was just arrested for allegedly smuggling $2.5 billion in GPUs to China Cyberattack on vehicle breathalyzer company leaves drivers stranded across the US Man pleads guilty to $8 million AI-generated music scheme Two Israelis AI generated "intelligence" and sold it to Iran

nFactorial Podcast
Даулет Жангузин, NVIDIA, Groq, Cohere, Lyft, Google - Как пишут код лучшие кодеры Кремниевой Долины?

nFactorial Podcast

Play Episode Listen Later Mar 12, 2026 193:31


В этом выпуске мы вместе с Даулетом Жангузиным - инженером из Кремниевой долины с 15-летним опытом (NVIDIA, Groq, Cohere, Lyft, Google, Microsoft) - говорим о карьере в BigTech и о том, что происходит под капотом современного AI. Обсуждаем практичную сторону работы с большими моделями: как выжимать максимум из Nvidia GPU, чем полезен Claude в реальных задачах, и какие курсы/ресурсы действительно помогают расти инженеру, как пишут код в 2026 лучшие программисты Кремниевой Долине. Эпизод будет интересен тем, кто строит карьеру в разработке/ML, хочет понять трек BigTech (Microsoft → Google → Lyft), интересуется LLM-инфраструктурой и оптимизацией вычислений, а также ищет советы по обучению и прохождению технических собеседований в ведущие tech-компании.  Арман Сулейменов: https://www.instagram.com/armansu/ Даулет Жангузин: https://www.instagram.com/daulet/ Продюсер и режиссёр: Данияр Ахметжанов: https://www.instagram.com/good.years/ Наш Instagram: https://www.instagram.com/nfactorialpodcast/ Получите одну из самых востребованных профессий в мире - ИИ-разработчик - вместе с nFactorial School - https://www.nfactorial.school/courses_new/llm-engineer

Where It Happens
Autoresearch clearly explained (why it matters)

Where It Happens

Play Episode Listen Later Mar 11, 2026 24:21


I break down Andrej Karpathy's new open-source project, Autoresearch: what it is, how it works, and why some of the smartest people in tech are losing their minds over it. I walk through 10 concrete business ideas you can build on top of Autoresearch loops, from niche agent-in-a-box products to always-on A/B testing agencies. I also cover Karpathy's companion launch, Agent Hub, share community reactions, and show you step by step how to get started using Claude Code and a Colab GPU. I'm hosting a free workshop so you can build your business in the age of AI. Sign up here: https://startup-ideas-pod.link/build-with-ai-2026 Links Mentioned: Autoresearch Github: https://startup-ideas-pod.link/autoresearch Timestamps 00:00 – Intro 00:45 – How Autoresearch Actually Works 02:40 – Visual Walkthrough of the Autoresearch Loop 03:37 – Mental Model: Your Research Bot That Runs While You Sleep 05:26 – Idea 1: Niche Agent-in-a-Box Products 06:48 – Idea 2: A/B Testing for Marketing (Landing Pages & Ads) 08:45 – Idea 3: Research as a Service 09:43 – Idea 4: Power Tool Inside Your Own SaaS 10:49 – Idea 5: Agency That Runs 100× More Tests 12:05 – Idea 6: Auto Quant for Trading Ideas 13:44 – Idea 7: Always-On Lead Qualification & Follow-Up 14:21 – Idea 8: Finance Ops Autopilot for Businesses 15:09 – Idea 9: Internal Productivity Lab for Your Org 15:53 – Idea 10: Done-for-You Research & Due Diligence Shop 16:41 – Non business use cases 18:27 – Karpathy's Agent Hub Announcement 19:50 – How to Get Started with Autoresearch 22:21 – Final Thoughts Key Points Autoresearch is an open-source AI agent that sets a goal, runs experiments in a loop on a GPU, keeps the winners, and discards the rest — all while you sleep. You need an NVIDIA GPU to run it (tested on H100), but you can rent one cheaply through Lambda Labs, Vast AI, RunPod, Google Cloud, or Google Colab. The fastest way to get started is to use Claude Code to walk you through installation, then run it on Google Colab with a T4 GPU runtime. Ten business ideas built on Autoresearch span niches like SaaS optimization, A/B testing agencies, trading backtests, CRM lead scoring, and done-for-you due diligence. Karpathy also launched Agent Hub — essentially a GitHub designed for agent swarms to collaborate on the same codebase. The project already has 25,000+ GitHub stars and is growing fast; early movers who tinker now build an unfair advantage. The #1 tool to find startup ideas/trends - https://www.ideabrowser.com LCA helps Fortune 500s and fast-growing startups build their future - from Warner Music to Fortnite to Dropbox. We turn 'what if' into reality with AI, apps, and next-gen products https://latecheckout.agency/ The Vibe Marketer - Resources for people into vibe marketing/marketing with AI: https://www.thevibemarketer.com/ FIND ME ON SOCIAL X/Twitter: https://twitter.com/gregisenberg Instagram: https://instagram.com/gregisenberg/ LinkedIn: https://www.linkedin.com/in/gisenberg/

TechLinked
Nvidia GPUs ascendant, Anthropic fights back, xAI in trouble + more!

TechLinked

Play Episode Listen Later Mar 7, 2026 9:30


Timestamps: 0:00 why? is a good question 0:09 Nvidia GPU monopoly 2:04 Anthropic to fight Pentagon 4:51 QUICK BITS INTRO 5:33 QUICK BITS INTRO 5:46 Apple blocks ByteDance-owned apps 6:29 Apple Music adds AI music tags 7:00 Macbook Neo benchmarks 7:40 Wikipedia AI translation issues 8:19 Roblox AI re-phrases bad chats NEWS SOURCES: https://lmg.gg/HpyUt Learn more about your ad choices. Visit megaphone.fm/adchoices

TechLinked
Lenovo Legion Go Fold, Anthropic rejects US govt, Meta AI Safety SNAFU + more!

TechLinked

Play Episode Listen Later Feb 28, 2026 10:12


Timestamps: 0:00 it's like it never ends! 0:09 Lenovo Legion Go Fold 1:39 Anthropic rejects Pentagon 3:04 Meta AI Safety head gets OpenClaw'd 5:51 QUICK BITS INTRO 6:04 Nvidia SHIELD TV gets another update 6:43 Nvidia GPU driver causes fan issues 7:24 Block cuts nearly half workforce 8:07 Tecno modular phone conccept 8:38 Burger King AI headsets NEWS SOURCES: https://lmg.gg/BaEmc Learn more about your ad choices. Visit megaphone.fm/adchoices

The AI Breakdown: Daily Artificial Intelligence News and Discussions

Anthropic drops Sonnet 4.6 with a million-token context window and major gains in computer use, coding, and agentic workflows at a dramatically lower price point—immediately reshaping the economics of OpenClaw-style agents. Meanwhile, Grok 4.2 enters public beta with a multi-agent debate system and promises rapid weekly improvement, and Apple ramps up AI wearables. In the headlines: Apple's AI glasses push, Spotify engineers stop writing code by hand, Meta commits to millions of Nvidia GPUs, Chinese AI price wars, and a possible SaaS rebound. Want to build with OpenClaw?LEARN MORE ABOUT CLAW CAMP: ⁠⁠https://campclaw.ai/⁠⁠Or for enterprises, check out: ⁠⁠https://enterpriseclaw.ai/⁠⁠Brought to you by:KPMG – Discover how AI is transforming possibility into reality. Tune into the new KPMG 'You Can with AI' podcast and unlock insights that will inform smarter decisions inside your enterprise. Listen now and start shaping your future with every episode. ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.kpmg.us/AIpodcasts⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Rackspace Technology - Build, test and scale intelligent workloads faster with Rackspace AI Launchpad - ⁠⁠⁠⁠⁠⁠http://rackspace.com/ailaunchpad⁠⁠⁠⁠⁠⁠Blitzy - Want to accelerate enterprise software development velocity by 5x? ⁠⁠⁠https://blitzy.com/⁠⁠⁠Optimizely Agents in Action - Join the virtual event (with me!) free March 4 - ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.optimizely.com/insights/agents-in-action/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠AssemblyAI - The best way to build Voice AI apps - ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.assemblyai.com/brief⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠LandfallIP - AI to Navigate the Patent Process - https://landfallip.com/Robots & Pencils - Cloud-native AI solutions that power results ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://robotsandpencils.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠The Agent Readiness Audit from Superintelligent - Go to ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://besuper.ai/ ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠to request your company's agent readiness score.The AI Daily Brief helps you understand the most important news and discussions in AI. Subscribe to the podcast version of The AI Daily Brief wherever you listen: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://pod.link/1680633614⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Interested in sponsoring the show? sponsors@aidailybrief.ai

TechLinked
No New Nvidia GPUs, iPhone Lockdown Mode Success, EU vs TikTok + more!

TechLinked

Play Episode Listen Later Feb 7, 2026 11:46


Timestamps: 0:00 WHALE LAN, TICKETS AVAILABLE NOW 0:19 No new Nvidia GPUs until 2028 - report 1:46 iPhone Lockdown mode stops FBI 3:33 EU vs TikTok over addictive design 6:41 QUICK BITS INTRO 6:58 Wayback Machine fights link rot 7:43 Windows 11 update drops frames 8:19 NASA allows smartphones in space 9:02 Substack data breach 9:43 Questions about Gemini shopping NEWS SOURCES: https://lmg.gg/wZ98Z Learn more about your ad choices. Visit megaphone.fm/adchoices

Daily Tech News Show
Why Google Won't Talk About Apple

Daily Tech News Show

Play Episode Listen Later Feb 5, 2026 28:14


It's not a conspiracy. We don't think anyway. Plus, RAMageddon expands to take down the Steam Machine and NVIDIA GPUs.Starring Tom Merritt and Huyen Tue DaoShow notes can be found here. Hosted on Acast. See acast.com/privacy for more information.

TechLinked
Microsoft promises to fix Windows, Annie's Archive lawsuit, Moltbook + more!

TechLinked

Play Episode Listen Later Jan 31, 2026 13:25


Timestamps: 0:00 i am one wild and crazy guy!! 0:11 Microsoft promises to fix Windows 2:30 Annie's Archive sued for 13 TRILLION 4:22 Moltbook, social network for AI agents 7:34 QUICK BITS INTRO 7:47 Nvidia GPU prices are out of control 8:26 Tesla discontinues Model S and Model X 9:22 US cyber security chief uses ChatGPT 10:39 Google Project Genie, GFiber 3Gig 11:37 Custom lung machine on demand NEWS SOURCES: https://lmg.gg/AMloV Learn more about your ad choices. Visit megaphone.fm/adchoices

Screaming in the Cloud
Building Software While Keeping Humans in Charge

Screaming in the Cloud

Play Episode Listen Later Jan 29, 2026 30:26


Alyss Noland, who works on Cloud Dev Ecosystem at Nvidia, is back on the show to talk about building software with AI when you're not a real developer. Alyss runs a program that gives AI startups access to Nvidia GPUs and uses AI tools herself to build production software at Nvidia. Corey and Alyss discuss using AI to help curate newsletters without actually writing them, why humans still need to check everything, and the weird reality of people developing relationships with chatbots. Show Highlights: (01:34) What Alyss Does at Nvidia(05:44) When AI First Worked for Corey(07:34) Building Internal Tools vs Using AI(10:39) Using AI to Help Write Last Week in AWS (13:43) DGX Cloud Innovation Lab (17:11) Building Production Software with AI (20:48) The Future of SEO (25:24) Using AI as a Writing Assistant (29:51) closing remarksLinks:Alyss's LinkedIn: https://www.linkedin.com/in/alyssnoland/Alyss's Personal Website: https://dev.to/preciselyalyssSponsored by: duckbillhq.com

PC Perspective Podcast
Podcast #853 - RIP Cheap SSDs, RTX 5070 Ti EOL Denied, Micron Fab, Bluetooth flaw, Kent Builds a PC, and MORE

PC Perspective Podcast

Play Episode Listen Later Jan 24, 2026 56:40


What happens when Kent decides to use the podcast as an SFF build livestream? What about a build using DDR4 memory?? It probably doesn't get more exciting than this.  Unless you count discussing the impending pricing DOOM for SSDs and the Google enabled bluetooth security flaw.So much more fun in the timestamps below!Timestamps:0:00 Intro00:41 Patreon01:29 Food with Josh03:03 Checking in on Kent04:42 RIP cheap SSDs06:07 Samsung and SK hynix reportedly cut NAND supply to drive profits07:00 Checking in on Kent again07:38 NVIDIA GPU prices are probably going up soon12:27 RTX 5070 Ti and 5060 Ti 16GB are not EOL after all14:18 NVIDIA releasing Arm-based chips for Windows laptops this year?17:33 Micron acquires PSMC fab to expand memory operations19:38 Dev patches WINE to make Photoshop 2021, 2025 run on Linux21:35 Josh checks in on Kent25:32 (In)Security Corner35:34 Another check on Kent's build progress36:38 Gaming Quick Hits40:24 Kent makes more progress41:36 Picks of the Week55:34 Outro ★ Support this podcast on Patreon ★

All-In with Chamath, Jason, Sacks & Friedberg
Ari Emanuel on the Future of Entertainment: Hollywood, AI, Creator Economy, YouTube vs Netflix

All-In with Chamath, Jason, Sacks & Friedberg

Play Episode Listen Later Nov 5, 2025 27:25


(0:00) Introducing Ari Emanuel (1:03) Ari's background in the content business and Hollywood (4:14) How infinite distribution is changing Hollywood, the future of podcasting, creators owning advertising products (9:21) AI's impact: anticipating the 3-day workweek (12:13) Splitting time between TKO and WME, YouTube vs Netflix for creators, modern IP ownership (18:20) How he got a fighter's reputation, relationship with Michael Ovitz, how he built on Ovitz's vision (22:12) Future of sports entertainment: adapting to attention spans, clip culture, going global (24:21) Ari's relationship with his brothers and Elon Musk Thanks to our partners for making this happen! Solana - Solana is the high performance network powering internet capital markets, payments, and crypto applications. Connect with investors, crypto founders, and entrepreneurs at Solana's global flagship event during Abu Dhabi Finance Week & F1: https://solana.com/breakpoint OKX - The new way to build your crypto portfolio and use it in daily life. We call it the new money app. https://www.okx.com/ Google Cloud - The next generation of unicorns is building on Google Cloud's industry-leading, fully integrated AI stack: infrastructure, platform, models, agents, and data. https://cloud.google.com/ IREN - IREN AI Cloud, powered by NVIDIA GPUs, provides the scale, performance, and reliability to accelerate your AI journey. https://iren.com/ Oracle - Step into the future of enterprise productivity at Oracle AI Experience Live. https://www.oracle.com/artificial-intelligence/data-ai-events/ Circle - The America-based company behind USDC — a fully-reserved, enterprise-grade stablecoin at the core of the emerging internet financial system. https://www.circle.com/ BVNK - Building stablecoin-powered financial infrastructure that helps businesses send, store, and spend value instantly, anywhere in the world. https://www.bvnk.com/ Polymarket - The world's largest prediction market. https://www.polymarket.com/ Follow Ari: https://x.com/AriEmanuel Follow the besties: https://x.com/chamath https://x.com/Jason https://x.com/DavidSacks https://x.com/friedberg Follow on X: https://x.com/theallinpod Follow on Instagram: https://www.instagram.com/theallinpod Follow on TikTok: https://www.tiktok.com/@theallinpod Follow on LinkedIn: https://www.linkedin.com/company/allinpod Intro Music Credit: https://rb.gy/tppkzl https://x.com/yung_spielburg

All-In with Chamath, Jason, Sacks & Friedberg
Inside Saudi Arabia's AI Ambition: Tareq Amin on Building a New Tech Superpower

All-In with Chamath, Jason, Sacks & Friedberg

Play Episode Listen Later Nov 4, 2025 29:16


(0:00) Introducing Tareq Amin (0:39) Saudi Arabia's evolution, Humain's business (8:11) How Humain works with foundational model providers (13:14) Saudi's energy and talent advantages, the AI race in the Middle East (18:11) Working in the era of MBS, Vision 2030 (21:40) How Saudi manages their relationships with the US and China (23:51) Sacks on the US-Saudi AI alliance Thanks to our partners for making this happen! Solana - Solana is the high performance network powering internet capital markets, payments, and crypto applications. Connect with investors, crypto founders, and entrepreneurs at Solana's global flagship event during Abu Dhabi Finance Week & F1: https://solana.com/breakpoint OKX - The new way to build your crypto portfolio and use it in daily life. We call it the new money app. https://www.okx.com/ Google Cloud - The next generation of unicorns is building on Google Cloud's industry-leading, fully integrated AI stack: infrastructure, platform, models, agents, and data. https://cloud.google.com/ IREN - IREN AI Cloud, powered by NVIDIA GPUs, provides the scale, performance, and reliability to accelerate your AI journey. https://iren.com/ Oracle - Step into the future of enterprise productivity at Oracle AI Experience Live. https://www.oracle.com/artificial-intelligence/data-ai-events/ Circle - The America-based company behind USDC — a fully-reserved, enterprise-grade stablecoin at the core of the emerging internet financial system. https://www.circle.com/ BVNK - Building stablecoin-powered financial infrastructure that helps businesses send, store, and spend value instantly, anywhere in the world. https://www.bvnk.com/ Polymarket - The world's largest prediction market. https://www.polymarket.com/ Follow Tareq: https://x.com/TareqAmin_ Follow the besties: https://x.com/chamath https://x.com/Jason https://x.com/DavidSacks https://x.com/friedberg Follow on X: https://x.com/theallinpod Follow on Instagram: https://www.instagram.com/theallinpod Follow on TikTok: https://www.tiktok.com/@theallinpod Follow on LinkedIn: https://www.linkedin.com/company/allinpod Intro Music Credit: https://rb.gy/tppkzl https://x.com/yung_spielburg

All-In with Chamath, Jason, Sacks & Friedberg
Triple H on WWE's Evolution, the Rise of the Antihero, and the Psychology of Stardom

All-In with Chamath, Jason, Sacks & Friedberg

Play Episode Listen Later Nov 3, 2025 26:05


(0:00) Introducing Paul "Triple H" Levesque (0:56) The skills needed to be an effective professional wrestler, understanding the characteristics of President Trump and The Rock (5:41) The rise of the antihero / heel persona, and why modern wrestling is more morally grey (8:43) WWE vs UFC: Why they are opposites, star building (11:21) Physicality of the WWE and how they support wrestlers (14:13) The business of the WWE: using streaming and social as a funnel for live events, the magic of live WWE (21:28) Helping America's youth via physical fitness (23:35) How the internet forced wrestlers to blend their in ring personas with real-life Thanks to our partners for making this happen! Solana - Solana is the high performance network powering internet capital markets, payments, and crypto applications. Connect with investors, crypto founders, and entrepreneurs at Solana's global flagship event during Abu Dhabi Finance Week & F1: https://solana.com/breakpoint OKX - The new way to build your crypto portfolio and use it in daily life. We call it the new money app. https://www.okx.com/ Google Cloud - The next generation of unicorns is building on Google Cloud's industry-leading, fully integrated AI stack: infrastructure, platform, models, agents, and data. https://cloud.google.com/ IREN - IREN AI Cloud, powered by NVIDIA GPUs, provides the scale, performance, and reliability to accelerate your AI journey. https://iren.com/ Oracle - Step into the future of enterprise productivity at Oracle AI Experience Live. https://www.oracle.com/artificial-intelligence/data-ai-events/ Circle - The America-based company behind USDC — a fully-reserved, enterprise-grade stablecoin at the core of the emerging internet financial system. https://www.circle.com/ BVNK - Building stablecoin-powered financial infrastructure that helps businesses send, store, and spend value instantly, anywhere in the world. https://www.bvnk.com/ Polymarket - The world's largest prediction market. https://www.polymarket.com/ Follow Triple H: https://x.com/TripleH Follow the besties: https://x.com/chamath https://x.com/Jason https://x.com/DavidSacks https://x.com/friedberg Follow on X: https://x.com/theallinpod Follow on Instagram: https://www.instagram.com/theallinpod Follow on TikTok: https://www.tiktok.com/@theallinpod Follow on LinkedIn: https://www.linkedin.com/company/allinpod Intro Music Credit: https://rb.gy/tppkzl https://x.com/yung_spielburg

All-In with Chamath, Jason, Sacks & Friedberg
California Forever: The Startup Building America's Next Great City

All-In with Chamath, Jason, Sacks & Friedberg

Play Episode Listen Later Oct 21, 2025 10:50


(0:00) Introducing Jan Sramek (0:51) How California Forever is building America's next great city Thanks to our partners for making this happen! Solana - Solana is the high performance network powering internet capital markets, payments, and crypto applications. Connect with investors, crypto founders, and entrepreneurs at Solana's global flagship event during Abu Dhabi Finance Week & F1: https://solana.com/breakpoint OKX - The new way to build your crypto portfolio and use it in daily life. We call it the new money app. https://www.okx.com/ Google Cloud - The next generation of unicorns is building on Google Cloud's industry-leading, fully integrated AI stack: infrastructure, platform, models, agents, and data. https://cloud.google.com/ IREN - IREN AI Cloud, powered by NVIDIA GPUs, provides the scale, performance, and reliability to accelerate your AI journey. https://iren.com/ Oracle - Step into the future of enterprise productivity at Oracle AI Experience Live. https://www.oracle.com/artificial-intelligence/data-ai-events/ Circle - The America-based company behind USDC — a fully-reserved, enterprise-grade stablecoin at the core of the emerging internet financial system. https://www.circle.com/ BVNK - Building stablecoin-powered financial infrastructure that helps businesses send, store, and spend value instantly, anywhere in the world. https://www.bvnk.com/ Polymarket - The world's largest prediction market. https://www.polymarket.com/ Follow Jan: https://x.com/jansramek Follow the besties: https://x.com/chamath https://x.com/Jason https://x.com/DavidSacks https://x.com/friedberg Follow on X: https://x.com/theallinpod Follow on Instagram: https://www.instagram.com/theallinpod Follow on TikTok: https://www.tiktok.com/@theallinpod Follow on LinkedIn: https://www.linkedin.com/company/allinpod Intro Music Credit: https://rb.gy/tppkzl https://x.com/yung_spielburg