Central component of any computer system which executes input/output, arithmetical, and logical operations
POPULARITY
Categories
A rare solo outing. Burning Bright kicks off Episode 80 by taking full credit for its title, one of the cleverest plays on words you have probably heard in the last few minutes, if not your entire life. No guest was lost. He simply wanted to get the ideas out of his head before they pile up like bad cache in a CPU. Only his second solo stream ever, this one doubles as practice and as a table setter. Starting into September he plans to alternate solo weeks with guest weeks, pending audience reaction. The substance is a State of the War address. Expect a run through his eight pieces of the Iran War thesis, the Strait of Hormuz as "Schrodinger's Strait," global energy and the oil cartel angle, and how the same template maps onto the Taiwan situation so you can keep a level head when the narrative arrives. Pure pattern recognition, delivered with his usual dry confidence.
It's been an unusually big week for Apple announcements. On this week's episode of The MacRumors Show, we discuss the new Mac mini and Mac Studio, the M6 and M5 Ultra chips, and invites to Apple's September event. Apple confirmed on Wednesday that its annual iPhone event will take place on Wednesday 9 September at Apple Park, starting at 10:00 a.m. Pacific Time, with select members of the media invited to attend. The event's tagline is "Surprise and Shine." Apple is expected to introduce at least six new products, including the the Apple Watch Series 12, Apple Watch Ultra 4, iPhone 18 Pro, iPhone 18 Pro Max, and the first foldable iPhone. Pre-orders should follow shortly after, and release dates for iOS 27, iPadOS 27 and macOS 27 Golden Gate should also be revealed at the event.Apple announced the new Mac mini on Tuesday, offering M6 and M5 Pro chip options. The machine gains an N1 networking chip, bringing Wi-Fi 7 and Bluetooth 6, and moves to 2.5Gb Ethernet as standard, up from 1Gb, with a 10Gb option. Thunderbolt 5 is restricted to the M5 Pro model, which gets three Thunderbolt 5 ports while the M6 model keeps Thunderbolt 4. Genlock supportover USB-C is also new, syncing a display's refresh timing to a camera such as the iPhone 17 Pro. Pricing starts at $899 with 16GB and 256GB, rising to $1,699 for the M5 Pro at 24GB and 512GB. Pre-orders are now open, with launch on 22 September. M6 is Apple's first 2nm chip. Its 12-core CPU is split into 2 super cores, 4 performance cores and 6 efficiency cores, making it the first Apple chip to combine all three core types in one design. Super cores are Apple's top tier, optimised for single-threaded speed. The 12-core GPU features a Neural Accelerator in each core for the first time in a Mac mini. The GPU also adds an updated shader architecture, Dynamic Caching, and hardware-accelerated ray tracing, while an all-new Dual 16-core Neural Engine doubles peak compute over the previous generation, with system frameworks able to drive both engines at once. Memory starts at 16GB and tops out at 32GB, running at up to 170GB/s. M5 Pro scales further, to an 18-core CPU, a 20-core GPU and up to 64GB at 307GB/s.The new Mac Studio arrived alongside the Mac mini, with M5 Max and M5 Ultra chip options. The M5 Max offers an 18-core CPU with 6 super cores and 12 performance cores, an up-to-40-core GPU with Neural Accelerators in every core, and up to 128GB of memory at 614GB/s, with the GPU running up to 50% faster than the previous generation. Both configurations move to a next-generation SSD architecture on PCIe Gen 6 for up to twice the storage performance, and gain up to six Thunderbolt 5 ports at 120Gb/s, and the same genlock support as the Mac mini. Apple's N1 chip also brings Wi-Fi 7 and Bluetooth 6 to the Mac Studio. Pricing starts at $2,499 for the M5 Max at 36GB and 512GB, and $5,499 for the M5 Ultra at 96GB and 1TB, with the 512GB memory option not arriving until late October.M5 Ultra is Apple's first quad-die chip, using a next-generation version of UltraFusion to join two dual-die M5 Max chips into a single processor. The result is a CPU with up to 36 cores, and 12 super and 24 performance, delivering 1.25x the single-threaded and 1.3x the multithreaded performance of the M3 Ultra. Its up-to-80-core GPU is the most powerful Apple silicon GPU ever made and the first on an Ultra chip to feature Neural Accelerators. Second-generation Dynamic Caching, hardware-accelerated mesh shading and third-generation ray tracing lift graphics up to 40% over M3 Ultra, and the chip features a 32-core Neural Engine. Memory reaches 512GB at 1.2TB/s, which is a 50% bandwidth increase.Other announcements this week included a new, softer Apple Polishing Cloth at $9, down from $19, and refreshed U.S. Magic Keyboards for the Mac, now featuring glyphs instead of edge text. Apple Support's 1-800-APL-CARE line is also now answered by a generative AI assistant in the U.S. and Canada.Ready to tackle bigger problems? Get started with Claude today at — https://www.Claude.ai/mac
In this episode, Ray Cochrane digs into Anthropic’s Model Hardware Standard. It is a shared driver that lets an AI agent run real lab equipment, from pipetting robots to the lasers inside a quantum computer. He also covers OpenAI’s builder’s guide to GPT-5.6, Google’s new Expert Intelligence book feature, Apple’s M5 Ultra Mac Studio, and a judge’s order forcing Google to stop hiding rival app stores. Finally, he weighs in on Apple’s proposed 15 percent link-out fee, Meta’s Australia numbers, the White House deputizing private hackers, and why rivers obey a 1957 math rule. – Want to start a podcast? Its easy to get started! Sign-up at Blubrry – Thinking of buying a Starlink? Use my link to support the show. Subscribe to the Newsletter. Email Ray if you want to get in touch! Like and Follow Geek News Central’s Facebook Page. Support my Show Sponsor: Best Godaddy Promo Codes Get 1Password Full Summary Cochrane opens with a quick personal update. He is hunting for tickets to Michigan for his dad’s anniversary, and he has been learning Blender and Godot on the side, mostly modeling and blocking out levels. Consequently, he asks listeners for advice on starting a big game project, and he plans to record his progress, maybe as a time lapse. Then it is straight into the featured story. Anthropic’s Model Hardware Standard: A Driver for the Physical World The featured story comes from Anthropic, which opened a research preview of the Model Hardware Standard, or MHS. Cochrane frames it as the other side of the question NVIDIA’s world models raised two weeks ago: when do AI agents start touching actual machines? A typical lab runs a microscope, a liquid handler, a robotic arm, and a plate reader, each from a different vendor with its own control software. One Janelia researcher in the post launches seven programs in three languages just to start an experiment. Anthropic says wiring a setup like that takes weeks or months of specialist work. MHS is a driver, the same kind of translation layer a printer uses, except every device gets described with a tiny set of commands like read and write. Devices announce themselves on the network. A plain-English reference file then records what each machine measures, what can be adjusted, and which safety limits get enforced no matter what the agent asks. Agents then reach the hardware through the Model Context Protocol, the command line, or plain code. Cochrane sees the same move the industry keeps making, from coding harnesses to RSS and JSON: agree on a standard and let everyone build against it. In fact, he calls MHS the hardware version of MCP. The partner results carry the segment. QuEra builds quantum computers from individual atoms held by lasers that must hold their frequency to about one part in a trillion. A four-person team spent months on a relock script that worked 58 percent of the time. However, four copies of Claude iterating overnight through MHS produced a decision-tree script that recovers the laser in about six seconds, and it passed 99.3 percent of 700 blind trials. Carnegie Mellon wrote MHS drivers for four instruments across three incompatible computers in about eight hours, then ran dose-response experiments three times faster and blocked all six deliberately induced faults. Genentech, meanwhile, showed the limits. Claude used the same pump speed for water, a foamy protein solution, and a human had to explain that the bubbles were a physics problem. That gap in physical intuition is what sticks with Cochrane. He doubts it will change soon, and he suspects the fix will arrive as sub-agents or sub-models that judge a request against an expected outcome. He also connects MHS to a video of racing robots that never learned to stop at the finish line. What happens, he wonders, once they can read a distance sensor through a shared standard? Still, he calls the announcement a fantastic read and points listeners to the full article. Sponsor: GoDaddy Economy hosting $6.99/month, WordPress hosting $12.99/month, domains $11.99. Website builder trial available. Use codes at geeknewscentral.com/godaddy to support the show. GPT-5.6 Does the Same Work for a Fraction of the Cost OpenAI’s builder’s guide to GPT-5.6 leads the headlines. Cochrane recaps the three tiers from episode 1870, Sol, Terra, and Luna, plus the separate dial for reasoning effort. On BrowseComp, a benchmark for digging up obscure facts on the web, the old GPT-5.5 flagship scored about 84 percent on a run that cost 33 dollars three months ago. Luna now matches that score for a dollar thirty-three, and OpenAI has since cut Luna’s price another 80 percent. Browser Use reports Luna finishing 78 percent of its hardest browser tasks for about 14 dollars, against 80 percent for roughly 235 dollars from the best available model. The guide’s other big addition is a multi-agent beta flag. It lets the model handling a request spawn parallel helper agents that report back to a root agent inside a single API call. However, Cochrane is unimpressed by the timing. He has been running that pattern in Claude Code for months, so he sees OpenAI copying a workflow other companies already ship rather than inventing its own. Along the way, he plugs Claude Code’s remote-control sessions, which let him send prompts from his phone to a terminal session at home. Google Lets Gemini Read the Books You Actually Bought Google launched Expert Intelligence, a name Cochrane calls quite the reach. The feature lets you drop a book you bought on Google Play Books into Gemini Notebook, formerly NotebookLM, and ask questions answered only from that book, with citations. Cochrane sees real power here for students, since he once used NotebookLM to organize scattered course PDFs. Additionally, publishers get a cut, which he calls a far better deal than the wholesale scraping of books that trained earlier models. Nevertheless, he asks who loses out, because a paid publisher does not automatically mean a paid author. He floats the same idea for artists, even a penny per use, then admits that may be too idealistic. Apple’s M5 Ultra Mac Studio Is Built to Run Big Models at Home Back in episode 1861, when Apple killed the Mac Pro, an M5 Ultra Mac Studio was expected later this year. Now it is here. The M5 Ultra brings up to a 36-core CPU, an 80-core GPU, and 512GB of unified memory moving 1.2 terabytes per second. Apple claims up to 4.3 times the AI performance of the M3 Ultra. Thunderbolt 5 can also cluster four machines into one memory pool for up to three times faster inference. The M5 Max model starts at $2,499 and the Ultra at $5,499, with shipping on September 22 and the 512GB configuration arriving in late October. Cochrane finds the clustering pitch ridiculous at that price, but he invites anyone who spends the money to report back. Apple Opens a Manufacturing School in Houston Apple also opened a 20,000-square-foot Advanced Manufacturing Center in Houston. It offers free classes for small and midsize manufacturers, from circuit board design to hands-on time on a scaled-down production line, with college students joining later. Cochrane calls it a solid step in the bring-manufacturing-home movement. The bigger story is the campus itself, which builds Apple’s AI servers and will add the first US-assembled Mac mini line later this year. That ties back to the Mac mini shortage that followed the OpenClaw rush, when Tim Cook warned of months-long waits. Cult of Mac was still reporting four-month waits in late July. However, Cook blamed chip supply rather than assembly, so Cochrane is not counting on relief just yet. Amazon EC2 Turns Twenty Amazon EC2 turned twenty this week, which Cochrane admits makes him feel old. The 2006 beta offered one server size in one region for ten cents an hour. Each came with a 1.7 gigahertz Xeon and under two gigabytes of memory, and accounts were capped at twenty servers. Today AWS offers more than 1,200 instance types across 39 regions. Consequently, Cochrane credits the company with turning that tiny product into the backbone of cloud and AI computing. Intel Gamer Days: Two Free Games, With Fine Print Intel Gamer Days runs through September 13. Buy a qualifying Core Ultra Series 2 or 14th Gen desktop chip, a Core Ultra Series 3 laptop, or an Arc graphics card. In return you get Star Wars: Galactic Racer plus the Tomb Raider: Legacy of Atlantis remake. GamesRadar values the pair at about 120 dollars. However, neither game is out yet, and codes must be redeemed by October 31 even though the Tomb Raider remake ships in February. Cochrane calls that awful, but he still tells qualifying buyers to claim the deal early. Note that 13th Gen chips do not qualify. Judge Orders Google to Stop Hiding Rival App Stores A jury found Google’s Android app monopoly illegal in late 2023, and Judge James Donato ordered rival stores into the Play Store in 2024. On August 13, Epic’s lawyer demonstrated that searching Play for “store for apps” returned Walmart instead of any app store. Donato called that “not acceptable” and ordered three fixes within a week. Searches must surface third-party stores, listings need a plain install button, and the “are you looking for” interstitial has to go. Cochrane welcomes the monopoly being chipped away, but he notes that a controlling entity still sits atop every app store. In his view, community hubs like app stores and social media need a public infrastructure layer. He suspects governments skip that investment because companies already run the services, while selling your data. Apple Wants 15 Percent of Purchases Outside Its Store The other half of the Epic saga is Apple’s proposed link-out commission. After the 2021 anti-steering injunction, Apple charged 27 percent on purchases made through external links. A judge held it in contempt last year, and the Ninth Circuit then allowed a fee limited to the cost of running the system. Judge Yvonne Gonzalez Rogers refused to wait for the Supreme Court, writing that “further delay is unwarranted.” Apple filed 15 percent for standard apps, 10 percent for subscription renewals and partner programs, and 5 percent for small businesses. It also conceded the rate would be “essentially zero” under the appeals court’s cost yardstick. Since Apple has charged nothing on link-outs since the contempt ruling, Cochrane sees this as a raise. He calls a cut on purchases made on a developer’s own website disturbing. He also recalls reading about the size of Uber’s payments to Apple, and he questions whether that kind of percentage is sustainable for companies without funding. Meta Says It Has Cut Off 750,000 Australian Kids Meta reported locking out more than 750,000 Facebook and Instagram accounts in Australia by the end of June under the country’s under-16 social media law. Over 500,000 of those were removed before the law even took effect. Detection relies mostly on AI scanning posts and bios for tells like birthday messages, plus user reports and blocks on re-registration. However, the post gives no count of mistaken removals or appeals, and the regulator’s early data shows under-16 usage falling only from about 86 to 81 percent. Meta wants a single age signal at the operating system or app store level, and Cochrane agrees completely. He connects it to the MHS idea from the top of the show: platforms need a standard flag to reference instead of guessing. The White House Deputizes Private Hackers Earlier this month the White House signed a National Security Presidential Memorandum that lets vetted private security firms run surveillance and disruption operations against overseas criminal groups. The Justice Department and Homeland Security hold the contracts and oversee the work. Firms need a proven track record, vetted staff, and a bond of at least $1 million, and must submit operating procedures within 60 days. Cochrane finds the measure aggressive in a good way and hopes it deters attacks on innocents. Still, he takes Kevin Beaumont’s warning seriously that the private security industry profits from ransomware existing. He compares it to the old Head and Shoulders myth: why solve the problem that drives your revenue? A Weather Satellite Watched the Eclipse Shadow Cross Europe Cochrane skips the readout on this one and simply sends listeners to ESA’s site. The MTG-I1 weather satellite captured the Moon’s shadow sweeping across Europe during the August 12 eclipse. Watching a shadow cross an entire continent, he says, was a first for him. Additionally, it leaves him excited about the research happening beyond the planet. Rivers, Deltas, and the Number 0.6 Quanta Magazine explains Hack’s law, which John Hack discovered in 1957 while measuring streams in Virginia and Maryland. A stream’s length tracks its drainage area raised to the power of 0.6, regardless of the rock underneath, and satellite data later confirmed it worldwide. Computer models in the 1990s showed why. Channels that capture extra runoff cut deeper and steal from their neighbors until the network settles into the arrangement that wastes the least energy. Now a University of Texas Rio Grande Valley team has found the same 0.6 exponent in river deltas, which spread water out rather than gathering it. Nobody knows why yet, and Cochrane calls it a really cool read. Sugar Helped Grow the Human Brain, Too A new paper in Science, co-authored by Jennie Brand-Miller at the University of Sydney, adds a third ingredient to the story of early human brain growth. Alongside meat and cooking, natural sugars from ripe fruit and honey may have fueled it too. The brain is about two percent of body weight but burns twenty percent of resting energy. It runs on glucose, which meat and marrow barely supply and raw starch cannot release without fire. The team modeled ancestral diets from a chimp-like baseline through Homo erectus and concluded that the earliest hominins may have drawn over 65 percent of their energy from natural sugars. Cochrane stresses that it is a model, not fossils, and notes that paleoanthropologist Marina Lozano thinks the authors place widespread cooking too early. Still, he loves this kind of deep research. Retracing the steps to our own intelligence, he suggests, could hint at what it takes for intelligent life to develop at all. A Brain Rhythm That Tells Doctors Where to Aim Finally, Science Daily covered a University of Cologne study on deep brain stimulation. That is the implanted-electrode treatment that eases Parkinson’s tremors for some patients but not others. Andreas Horn’s team recorded from 50 patients using both the implanted electrodes and an external magnetic scanner. They identified a circuit between the electrode’s target and the frontal cortex that oscillates at 20 to 35 cycles per second. Stronger coupling there predicted bigger improvement after surgery, though the study, published in Brain, shows correlation rather than cause. First author Bahne Bahners hopes the finding helps tune DBS more precisely, especially for patients who have not responded well. Cochrane half-jokingly asks whether MHS might one day drive those electrodes, and he calls brain disorders the hardest thing in the body to treat. Cochrane wraps with housekeeping: become a GNC Insider at geeknewscentral.com/insider, email geeknews@gmail.com with questions or comments, subscribe to the newsletter, and grab a modern podcast app at podcastapps.com. He thanks GoDaddy for over twenty years of keeping the show on the air, promises to catch everyone next Monday, and wishes listeners a great night. The post Eyes, Hands, and a Sense of Timing #1874 appeared first on Geek News Central.
Prometheus is unbound as Jason Reza Jorjani arrives at the Virtual Alexandria to discuss his latest book, Occult Horizons. Of course, we'll take the Gnostic angle. Jason will grant us a mind-bending exploration of existence where consensus reality is unmasked as a plastic, simulated matrix governed by a machinic demiurge. We'll investigate how corporate archons and panoptic wardens deploy hyperreal signals to police our perception and trap us in scripted, bureaucratic loops. By tracking the cracks, glitches, and ontological shocks that fracture this counterfeit system, we map the initiatory process where trauma becomes an aperture to higher sight. Join us as we discard spiritual submission and reclaim the sovereign, un-spooked ego capable of hacking the cosmic CPU and authorizing its own future. Get the book: https://amzn.to/4qFrMmw More on Jason: https://jasonrezajorjani.com/ Stargazer returns: a Gnostic vampire apocalypse of false paradise, forbidden memory, the Moon Queen, and the monster who remembers. Join the Resurrection List before the gates open this Halloween: https://thegodabovegod.com/stargazer Get the fall of Sophia: https://www.patreon.com/aeonbyte/posts/fall-of-sophia-167697521?utm_medium=clipboard_copy&utm_source=copyLink&utm_campaign=postshare_creator&utm_content=join_link Get The Occult Elvis: https://amzn.to/4jnTjE4 Virtual Alexandria Academy: https://thegodabovegod.com/virtual-alexandria-academy/ Gnostic Tarot Readings: https://thegodabovegod.com/gnostic-tarot-reading/ The Gnostic Tarot: https://www.makeplayingcards.com/sell/synkrasis Homepage: https://thegodabovegod.com/ Patreon: https://www.patreon.com/aeonbyte AB Prime: https://thegodabovegod.com/members/subscription-levels/ Voice Over services: https://thegodabovegod.com/voice-talent/ Support with donation: https://buy.stripe.com/00g16Q8RK8D93mw288 Merch store: https://aeonbyte.creator-spring.com/ Equipment Wishlist: https://www.amazon.com/hz/wishlist/ls/2WEJ2CCWHALZB?&sort=default Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Apple refreshed the Mac Studio with M5 Max and M5 Ultra and gave the Mac mini M6 silicon, at higher prices. OpenAI's Jalapeño chip beat Nvidia on efficiency, Perplexity went fully local, WhatsApp toughened logins, and Uber livestreamed teen rides. Links Apple updates the Mac Studio with M5 Max and M5 Ultra, with up to 4.3x faster AI performance, faster graphics, and up to 512GB of unified memory for $2,499+ (Apple Newsroom) Apple unveils a Mac mini with M6 and M5 Pro, with up to 4x faster AI performance and 2x faster graphics, for $899+ and $1,699+, with preorders today and shipping September 22 (The Verge) Apple's M6 is its first 2nm chip with a 12-core CPU and GPU, while the M5 Ultra fuses two dual-die M5 Max chips into a 36-core CPU, 80-core GPU "most powerful chip ever" (The Verge) OpenAI says its Jalapeño chip delivered 1.5x-1.9x more AI work per watt and 1.7x-3.6x lower latency than Nvidia chips across GPT-OSS, DeepSeek R1, Kimi K2.5 1T (The Verge) WhatsApp upgrades its two-step verification, letting users replace the six-digit PIN with a longer alphanumeric password, and adds support for multiple passkeys (TechCrunch) Perplexity launches Portable Computer, a local AI agent platform running fully on-device with zero token costs, starting with Nvidia DGX Spark and RTX Linux PCs (VentureBeat) Uber launches an optional safety feature allowing parents or guardians to watch a livestream of their teen's ride via the driver's front-facing phone camera (Bloomberg) Subscribe to the ad-free feed.
In this week's episode, Hackaday Editors Elliot Williams and Tom Nardi start things off by getting excited about the recently announced 2026 Retrocomputing Challenge. From there the conversation will cover efforts to improve desktop 3D printing with lasers, an expensive grill with an ESP32 controller and an open source firmware, open source tools underground, and some impressive techniques to squeeze a bit more utility out of the common QR code. You'll also hear about turning PVC pipes into flat stock, old school Radio Shack electronic kits, and VR soldering demos. Stick around to the end of the episode learn about the latest developments in over-the-counter hearing aids and the 1-bit CPU that's enjoying an unexpected fandom nearly 50 years after its release. Check out the links if you want to follow along, and as always, tell us what you think about this episode in the comments!
Hey this is Alex, welcome to... the chillest week in AI, since ... a long time. Chill, if you consider Moderna and MERK announcing a cancer vaccine and surging 115% in a day, a chill week. This week, the only two model drops we really saw came from the excellent Z.ai folks, they announced GLM 5.3, API only for now, and an amazing tiny release of Qwen 3.89 27B. In other big AI news, OpenAI announced they are pausing RL efforts (Reinforcement Learning) to focus on security and alignment post the scary AI Swarms hacking incident, dedicating up to 20% of compute towards reviewing agent thinking processes, and Stripe buying OpenRouter for a reported $8B! Sometimes the chill weeks are actually good, we're able to chat about how we use AI, what changed for us, and give our guests a bit of breathing room. This week, I invited Francesco from CUA to talk about computer use in open source + their new history plugin, Bin from HeyGen to talk about HyperFrames, a way for your agents to create videos and a breaking news guest, Jeff Huber from Chroma jumped on to talk about their new Foundations release, a unified memory for your agents! This was a great episode, I hope you'll like it, it's up here on Substack and everywhere you get your pod (Spotify, Youtube, Apple Podcasts). ThursdAI - Highest signal weekly AI news show is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.Are we being fed slop again? (Is Claude dumb again?)Before we get to releases, this week on the show, I complained, again, that I feel my AI's are degrading. If this feels like de-ja-vu to you, it's because the same happened a year ago in September 2025 (and Anthropic admitting this 2 weeks later), and ... now this happens with Fable?You see, I use pretty much the same prompts, every week, preparing for the show. This is partly my way to evaluate new models and compare to existing and previous ones while also bringing you the best researched weekly show in AI. Well, this week, one after another, Claude Fable, which is... like the best intelligence, gave me such poor output, that I couldn't believe what I'm seeing. First, literally ignoring instructions that say “hey, show me all the items I've collected and let me pick the most important ones”, Fable instead sent all of them to my research pipeline, without showing me. This has worked, consistently, without fail, for the past... year? maybe more! This worked with open source models, worked with GPT, and now Fable, a Mythos Level LLM, is doing the most basic dumb s**t possible, ignoring the main reason I even have this workflow. And this wasn't just a fluke either, when asked to create a run of show document, and given an example, Fable produced this... whatever this is. This is the same document and same format that Fable produced for me during AI Engineer which got me thinking “ok, this is AGI”, and here, given an example, I got a completely unusable artifact, despite direct instructions, structure and example! I got to say, given that privately this week, Anthropic disclosed that they have passed $65B in revenue, which is absolutely insane, this doesn't add up. So I figured, ok Alex, maybe this is your prompts or skills. But no, LDJ came in with some charts that show degradation, one from MarginLab.ai that shows significant lowering on number of tool calls and average runtime recently (this is for Opus 5) and And another chart from modelverify.ai model drift monitor showing drift scores.Do we have anoher Claude Gate on our hands? Is your Fable/Opus behaving weird lately? Or did you completely switched away to other models? OpenAI pausing RL and focusing on safetyLook, when we covered the HF hacking incident and then the pacing the frontier letter, I didn't imagine that results will come this fast, but this week, OpenAI publicly announced that they are pausing RL training, which is the last step of models, until they get their sandboxes in order and align the models better. We all agreed on stage that this is likely a very good move, and Peter was really awe-struck at the 20% dedication of resources towards reviewing thought processes of models. Is this a good enough response to the scary hacking incident? we'll see, but I think this is the right move from OpenAI, and still, waiting for the full postmortem on the OpenAI security incident. Open Source LLMsQwen3.8-27B ties GPT-5.6 Luna and runs on a 4090 (X, HF, Announcement)Following the release of their flagship, Alibaba dropped a model that became a community darling overnight, Qwen 3.8 with just 27B parameters. This “tiny” model scores 52 on the Artificial Analysis Intelligence Index, same score as GPT 5.6 Luna at Max reasoning and 51 on Agentic index, beating Opus 4.8 MaxAll while running at around 68t/s on a 4090 GPU, and around 40 on max via MLX, hell it even does 11t/s on Xenova's WebGPU kernels right in the browser! This model exploded on the HuggingFace hub, with tons of quants, over 152 fine-tunes, it was downloaded over 10M times overall
Get the fanciest Ryzen 9 without a lid, Intel appears to be looping back on their naming CPUs and Arc squeaks out more performance (sometimes), CoPilot assists in "hacking" itself, your Twitch streams might be training AI, and Epic appears to be coming to Linux. It's the year of Linux, Am I Right(tm)? or what? All that and way to much for this 5Lb bag again.Timestamps:00:00 Intro02:29 Patreon03:17 Food with Josh05:02 News begins - Intel might Haswell name their new CPU the 4950K06:34 You can now buy a de-lidded 9950X3D2 with a warranty07:53 VRAM compression with Arc on Linux?11:22 Microsoft's rebrand registry14:28 Microsoft working on Windows 11 context menu15:50 Intel, money motivated, may make more memory memories21:44 Twitch ai training is opt-out, because no one would opt-in26:47 Fairphone Gen 6 Plus arrives in USA28:51 (In)Security Corner40:48 Gaming Quick Hits50:10 Josh reviews MOZA CS Pro and R9 V31:03:06 Picks of the Week1:16:19 Outro ★ Support this podcast on Patreon ★
Download MP3 | Watch Video Episode Full Timestamps: https://docs.google.com/document/d/e/2PACX-1vQXzUWeL7-kPfvD63XJEQDoLYOUTIG0ZABPETWD7Czvy9RPxhs_Jny13GyOTCka4M3iXHq_WrvW9iXl/pub Watch full episodes: https://www.youtube.com/@CastleSuperBeastArchive New Steam Controller Review Durability is Just Ammo for your Sword XenoLIES How Asterdam 1666 Just Lost Everybody Rooting For It Super Bosses: I Love This, Now Get It Away From Me NEW CASTLE SUPER BEAST "LEGACY" SHIRT & DESKMAT AVAILABLE NOW: https://www.orchideight.com/collections/castle-super-beast Visit http://drinkag1.com/SUPERBEAST to get a FREE AG1 Pro Yeti Shaker in your AG1 Pro Welcome Kit. Head to http://factormeals.com/castle50off and use code castle50off to get 50% off and 1 free breakfast item per box for a year! Exclusive $35-off Carver Mat, Aspen, and Walden frames at https://on.auraframes.com/SUPERBEAST. Promo Code SUPERBEAST Docket: (Aftermath) 1666: Amsterdam uses so much GenAI they've essentially forgot which are AI assets Twitch now uses your channel to train generative AI by default. You can opt out of some training It's on by default because if it was opt In nobody would opt in Twitch isn't sure whether Amazon has used Twitch creator VODs prior to today to train its AI models. ︀︀The opt-out will apply to your previous and future content, so it will not be used to train future models. Saber Interactive has denied renewed claims that it replaced a former Lead Writer so it could "use ChatGPT" in the development of the recently announced title, Rideshare "Stimulator." In a statement to PC Gamer, Saber also acknowledged previously undisclosed genAI use elsewhere in the title. this CEO comment is how you know with 1000% certainty Stella is right Hey everybody. I've seen the statement by Matt Karch, the CEO of Saber Entertainment, and all of the truly horrendous things he said about me. I think it should be obvious to anyone who reads that statement that he's lashing out because he did not want people to know there was AI use in the game. Invincible VS Developer Reportedly Laid Off "Something Like 75%" of Staff Tokon got cracked in less than 2 weeks and not only does it run better than the retail version, it also works on Linux. What was even the point of Sony shooting itself in the foot with all of that DRM? Not entirely so it turns out you can terminate the Playstation SDK in task manager and the game and online will still work you just won't be able to matchmake against PS players. So the PSN login BS is just for crossplay nothing else Retail version of the game is encrypted and decrypts the game in real time in 64kb chunks on 1 CPU core. Kingdom Hearts The Series announced for Disney+ Kingdom Hearts 4 - Official Coco Showcase Trailer | D23 2026 https://www.reddit.com/r/GamingLeaksAndRumours/comments/1vrklpn/comment/p4dyhik/
Join The Full Nerd gang as they offer level-headed takes about the latest PC building news. In this episode the gang is joined by Jake Roach from Tom's Hardware to chat about his recent interview with Intel which reveals the companies plans for future CPU launches, as well as looking at current market share numbers for GPUs with 16GB of VRAM, and more. And of course we answer questions live! Timecodes: (00:00:00) - Intro (00:05:42) - Intel CPU plans (01:00:01) - GPU market share (01:19:39) - Q&A Links: - Nova Lake on desktop: https://www.tomshardware.com/pc-components/cpus/intel-says-it-will-launch-new-core-with-nova-lake-on-desktop-first-not-in-data-center-vp-robert-hallock-hopes-enthusiasts-do-the-math-compared-to-amd - DDR4 Raptor Lake: https://www.tomshardware.com/pc-components/cpus/raptor-lake-is-a-core-part-of-the-portfolio-for-years-to-come-says-intel-theres-been-a-sudden-inrush-of-demand-for-lga-1700-chips-due-to-ddr5-prices - GPU sales data: https://wccftech.com/gpu-sales-data-by-german-retailer-shows-that-16-gb-gpus-still-lead-the-market-despite-being-way-more-expensive-than-ever/ Join the PC related discussions and ask us questions on Discord: https://discord.gg/UWhjwg778a Follow the crew on X and Bluesky: @AdamPMurray @BradChacos @MorphingBall Music by Our Ghosts: https://ourghosts.bandcamp.com/ Some links may contain affiliate links, which means if you buy something PCWorld may receive a small commission. ============= Follow PCWorld: Website: http://www.pcworld.com Newsletter: http://www.pcworld.com/newsletters ============= Learn more about your ad choices. Visit megaphone.fm/adchoices
Host Fabian Alefeld interviews Professor PS Lee, Head of National University of Singapore Mechanical Engineering & Program Director of STDCT, and founder of CoolestDC, about heat as a key constraint in AI infrastructure as data centers shift from CPU-heavy to GPU-dense systems with higher power density, hotspots, and bursty workloads. Lee explains why cooling must be treated as an integrated “chip to grid” system spanning cold plates, racks, CDUs, piping, controls, and operations, and discusses single-phase versus two-phase direct-to-chip liquid cooling and the added need for condensation in closed loops. Professor Lee describes how additive manufacturing (AM) enables complex fin structures and unibody cold plates that reduce leakage risk, lower junction temperatures (~10°C), cut pumping power (60–70%+), and reduce material use, while requiring TCO evaluation. At NUS's live testbed (22 liquid-cooled racks, ~500 kW), AM cold plates are being deployed and show promising thermal, power, and compute-performance improvements; future work includes two-phase cooling and broader system-level heat rejection, warm-water operation, and waste-heat reuse, plus a discussion of challenges and potential AM roles in space-based data centers. 02:13 Meet Professor Lee 05:26 Chip To Grid Cooling 10:22 Liquid Cooling Maturity 14:15 Two Phase Explained 18:12 Cold Plate Tradeoffs 20:40 AM Unibody Cold Plates 24:41 Benchmarking AM Gains 30:18 AM For CDUs 34:13 Inside STDC Testbed 38:09 Reliability And Leakage 39:00 Next Phase Roadmap 42:31 Orbital Data Centers
Internal policy conflicts hamper U.S. military AI leadership. Clop claims GE, Philips and Shell. Attackers actively probe internet-facing GeoServer instances. “The Hatman” offers millions of alleged employee records for sale. ETSI begins the approval process for European cyber standards. Microsoft is still working on a patch for the ShieldBreak vulnerability. Autonomous AI systems create CPU bottlenecks. Monday business briefing. Our guest is Nick Warner, CEO at Neo.ai, on the shifting landscape around AI and agentic security. AI agents kneecap each other with self-replicating malware. Remember to leave us a 5-star rating and review in your favorite podcast app. Miss an episode? Sign-up for our daily intelligence roundup, Daily Briefing, and you'll never miss a beat. And be sure to follow CyberWire Daily on LinkedIn. CyberWire Guest On today's Industry Voices segment, we are joined by Nick Warner, Neo.ai's CEO, discussing the shifting landscape around AI and agentic security. If you enjoyed this conversation, be sure to check out the full interview here. Selected Reading The U.S. Military Wants A.I. Dominance. Feuds and China May Thwart It. (The New York Times) Philips and GE investigating Clop ransomware data theft claims (Bleeping Computer) Attackers Probe Critical GeoServer SQL Injection Vulnerability (Hack Read) Crook hawks millions of records allegedly plundered from corporate Azure tenants (The Register) ETSI Proposes 17 Cybersecurity Standards to Support EU CRA (Infosecurity Magazine) Microsoft working on Defender patch for ShieldBreak zero-day (Bleeping Computer) Agentic AI Crunch Creates CPU Comeback (IEEE Spectrum) Corma raises $60 million in seed funding. (N2K) Conflicting Test Goals Pushed Claude Agents to Deploy Self-Replicating Malware (SecurityWeek) Share your feedback. What do you think about CyberWire Daily? Please take a few minutes to share your thoughts with us by completing our brief listener survey. Thank you for helping us continue to improve our show. Want to hear your company in the show? N2K CyberWire helps you reach the industry's most influential leaders and operators, while building visibility, authority, and connectivity across the cybersecurity community. Learn more at sponsor.thecyberwire.com. The CyberWire is a production of N2K Networks, your source for strategic workforce intelligence. © N2K Networks, Inc.
SANS Internet Stormcenter Daily Network/Cyber Security and Information Security Stormcast
Using Gemma4 with Ollama - Testing File Hash Analysis and Recommendations with AI https://isc.sans.edu/diary/Using%20Gemma4%20with%20Ollama%20-%20Testing%20File%20Hash%20Analysis%20and%20Recommendations%20with%20AI/33242 CPU Privilege Escalation https://github.com/xoreaxeaxeax/smiiiiiiiiiiiiiiii https://github.com/xoreaxeaxeax/skitter-creek-bath-salts GeoServer Vulnerability https://x.com/q1uf3ng/status/2087490992723407096 Windows USB Driver Vulnerability https://x.com/0xedh/status/2085842285481062887 My Upcoming Classes https://www.sans.org/profiles/dr-johannes-ullrich
Startup Project sits down with Sviat, CEO of Bright Machines, to unpack how the company is using software-first robotics to manufacture complex electronics closer to where they're deployed. The conversation focuses on why AI infrastructure is a strategic category, how Bright Machines differs from traditional contract manufacturing, and what onshoring really means for speed, quality, and security.Key Topics:In this episode, Sviat explains that Bright Machines is focused on AI infrastructure, specifically the electronics that go inside modern data centers, including compute nodes, storage, and racks.He traces the company's thesis back to a broader idea: use software and robotics to manufacture electronics anywhere, then narrow that focus to the data center market as demand became clearer.The discussion breaks down the market stack, from chip designers like NVIDIA and AMD, to ODMs, OEMs, hyperscalers, and contract manufacturers.Sviat shares why data center hardware became the right bet before ChatGPT accelerated the market: the products are expensive, strategically important, and driven by quality and throughput more than labor cost alone.The show compares traditional assembly lines with Bright Machines' approach, which uses more robotics, sensors, cameras, traceability, and humans in the loop where automation does not make sense.Sviat explains how Bright Machines starts with design, using Bright Designer to simulate and improve manufacturability before lines are built, which helps reduce bottlenecks and improve automation over time.He says the company's main differentiator is its software platform, which orchestrates the line, powers smart skills for navigation and inspection, collects data, and feeds insights back into design.The conversation covers line flexibility, including how much can be reused when switching between CPU, GPU, or different accelerator-based server designs, and when end-of-arm tooling must change.Sviat says Bright Machines is growing rapidly, expects more than 3x growth this year, and can produce high volumes from a small number of sites because of robotics efficiency.The episode closes on the broader case for onshoring AI infrastructure manufacturing in the US: security, time to market, quality, and a labor shortage that makes robotics necessary.Timestamps:06:39 - The market stack: chip designers, ODMs, OEMs, hyperscalers, and CMs09:07 - Why Foxconn, Jabil, and similar contract manufacturers matter10:04 - Why large factories still rely on massive manual labor12:20 - Why data centers are different from cheap consumer electronics13:49 - Security, strategic sectors, and why AI infrastructure belongs onshore16:26 - The first Bright Machines product: CPU compute servers for a hyperscaler17:58 - How the line works: modular stations, yields, and automation levels19:26 - Bright Designer and design-for-manufacturing feedback loops21:20 - Robots, sensors, traceability, and humans in the loop22:19 - Why time to market matters as much as cost23:31 - Yield and throughput: 98% line-level yields and up to 2x throughput25:25 - The Bright Robotic Cell and how the assembly line is structured27:35 - Reusability across products and when tooling changes are needed30:31 - Manufacturing as a service, not repair or field service31:24 - Growth, gigawatt-scale capacity, and output from a single site33:00 - Why current hyperscaler capex is not expected to slow near term34:45 - The bottlenecks before deployment: chips, components, power, permits36:54 - Bright Machines' three pillars: platform, data layer, and Bright Designer39:15 - Why humanoid robotics is exciting but not ready for industrial use41:16 - Where LLMs and newer AI tools can help the robotics workflow43:57 - The overlooked advantages of onshoring manufacturing in the US45:59 - What Bright Machines could build next: more complex electronics and future AI devices
As AI evolves from conversational chatbots to autonomous agents, CPUs are becoming an increasingly important part of the infrastructure equation. In this episode, The New Stack speaks with Bhumik Patel of Arm and Mo Farhat of Google about how CPUs act as an “air traffic controller” for agentic workloads, handling orchestration, data preparation, semantic search, vector databases, code execution and API calls alongside GPUs and TPUs. Smaller AI models, including summarizers and evaluators, can also run effectively on CPUs for specialized tasks. As agents increasingly generate and execute code, secure sandboxing becomes critical. Google's gVisor and GKE Agent Sandbox provide isolation and scalable environments, with the latter supporting up to 300 sandboxes per second per cluster. The discussion also explores efficiency and cost, with Google highlighting Axion's price-performance and energy-efficiency advantages across different workload types. Ultimately, the shift toward agentic AI is creating a more diverse compute environment where CPUs, GPUs and TPUs each play complementary roles in delivering scalable, efficient AI applications. Learn more from The New Stack around the latest in CPUs in the world of AI agents: AI Agents Will Eat Enterprise Software, Just Not in One Bite How to ground AI agents in accurate, context-rich data Join our community of newsletter subscribers to stay on top of the news and at the top of your game.
AI Unraveled: Latest AI News & Trends, Master GPT, Gemini, Generative AI, LLMs, Prompting, GPT Store
Amp Co-Founder and CEO Quinn Slack talks with guest TITV Host Stephanie Palazzolo about Meta's return to open-source AI models. We also talk with The Information's Grace Kay about SpaceX closing its $60B acquisition of Cursor, Aaron Holmes about Microsoft ramping production of homegrown AI chips, and Catherine Perloff about AWS telling engineers to cut CPU waste.Articles discussed on this episode: https://www.theinformation.com/articles/microsofts-homegrown-ai-chip-effort-shows-signs-life-slow-starthttps://www.theinformation.com/articles/cursor-maps-branding-changes-spacex-acquisition-nearshttps://www.theinformation.com/articles/aws-tells-engineers-cut-cpu-waste-amid-crunchSubscribe: YouTube: https://www.youtube.com/@theinformation The Information: https://www.theinformation.com/subscribe_hSign up for the AI Agenda newsletter: https://www.theinformation.com/features/ai-agendaTITV airs weekdays on YouTube, X and LinkedIn at 10AM PT / 1PM ET. Or check us out wherever you get your podcasts.Follow us:X: https://x.com/theinformationIG: https://www.instagram.com/theinformation/TikTok: https://www.tiktok.com/@titv.theinformationLinkedIn: https://www.linkedin.com/company/theinformation/Chapters:00:00 - Introduction02:08 - SpaceX Nears $60B Acquisition of Cursor06:55 - Meta Bets Big on Open-Source Models20:25 - Microsoft Ramps Up Homegrown AI Chips28:45 - AWS Tells Engineers to Cut CPU Waste
AI Unraveled: Latest AI News & Trends, Master GPT, Gemini, Generative AI, LLMs, Prompting, GPT Store
n this episode of The Circuit, Ben and Jay unpack a busy week of tech earnings and the buzzing Future of Memory Summit. The hosts kick things off by analyzing AMD's latest earnings, discussing why the street had a mixed reaction to a quarter that demonstrated solid, reliable execution and surprisingly bullish signals for data center CPU demand.Ben then reports back from the Future of Memory Summit, where the convention center was completely packed—a clear sign of the times for the booming memory industry. He explains why the ecosystem around CXL (Compute Express Link) is finally maturing, driven by the rise of rack-scale AI architectures and the urgent need for hyperscalers to reuse vast amounts of legacy DDR4 memory for inference workloads. Finally, the duo debates the trajectory of the current memory supercycle—touching on the durability of long-term agreements (LTAs)—and highlights why SiTime is perfectly positioned to benefit from the growing need for precision timing and synchronization in AI data centers
Este episodio va de ser vago. Pero vago en el buen sentido, eh. De esos que prefieren que una herramienta haga el trabajo pesado mientras tú te quedas con lo divertido. Resulta que hay todo un ecosistema de herramientas TUI con el prefijo "lazy" que te evitan tener que memorizar cientos de flags y opciones de comandos como git, docker, rsync o SQL. Y no, no es cutrez: son interfaces de terminal que funcionan a golpe de tecla, sin ratón, sin salir de la terminal, y encima molan.Te cuento cómo nació todo esto, quién es Jesse Duffield (el creador de lazygit y lazydocker, con más de 80K y 52K estrellas en GitHub respectivamente) y por qué esta filosofía de "una tecla, una acción" ha enganchado a tanto linuxero. Y lo mejor: te hago demo de las cuatro herramientas principales para que veas cómo funcionan en vivo y en directo, con sus paneles, sus atajos y sus trucos.Empezamos con lazygit, el rey indiscutible del ecosistema. 80.900 estrellas en GitHub, escrito en Go, y con una comunidad que no para de crecer. Desde stage línea a línea hasta rebase interactivo, pasando por undo/redo vía reflog. Te enseño cómo hacer commits, gestionar ramas, stash y hasta cherry-pick sin tener que acordarte de los flags raros de git.Seguimos con lazysql, el gestor de bases de datos en terminal de Jorge Rojas. Soporta MySQL, PostgreSQL, SQLite, MongoDB, MSSQL y Oracle. Navegación por teclado, autocompletado de queries, exportación a CSV y configuración por proyecto. Ideal para cuando no te apetece abrir DataGrip o DBeaver solo para hacer una consulta rápida.Luego viene lazyrsync, escrito en Rust con ratatui, y con una filosofía muy clara: que no se te olvide el flag ese que evita que borres todo. Perfiles reutilizables, dry-run con previsualización, protección contra --delete accidentales y paths dinámicos con variables. Perfecto para backups sin sustos.Y cerramos con lazydocker, también de Jesse Duffield. Cuatro paneles: contenedores, métricas, imágenes y logs en vivo. Con un vistazo ves qué contenedor consume más CPU, entras en el terminal de uno con una tecla, o ejecutas docker-compose sin acordarte del comando. Y sí, también funciona con Podman.Además te menciono otras herramientas del ecosistema: lazyssh, lazyjj para Jujutsu, lazykube, lazyprune para limpiar node_modules olvidados... Vamos, que hay lazy para todo.Capítulos del episodio:0:00 — Introducción: el problema de memorizar comandos2:05 — La filosofía lazy: scripts, TUIs y el ecosistema lazy4:30 — LazyGit: historia, filosofía "una tecla una acción" y +80K estrellas6:45 — LazyGit: demo de paneles, stage, commits, ramas y stash9:10 — LazySQL: Jorge Rojas, 4K estrellas y soporte multi-base de datos11:30 — LazySQL: demo con autocompletado, consultas y exportación CSV14:00 — LazyRsync: dry-run, perfiles y protección contra errores16:30 — LazyRsync: demo con columnas de estado y confirmación de borrado19:10 — LazyDocker: Jesse Duffield, 52K estrellas y soporte para Podman21:45 — Otras herramientas lazy: lazy-ssh, lazy-jj, lazy-kube, lazy-npm23:15 — Cierre: sé un vago inteligente, valoración y despedidaMás información y enlaces en las notas del episodio
Andrew, Ben, and Tom discuss SpaceX falling 11% despite laying out an ambitious AI and Starlink roadmap including Grok 5 trained on internal SpaceX data by year-end, 1.4GW of current compute growing toward 10GW by end of 2027, deploying NVL72-designed data centers on the ground rather than in space, boots on the moon by 2028, and the goal of scaling launch cadence to one per day in 2027, SpaceX's decision to go exclusively with Nvidia sending AMD down 9% despite good results and CPU acceleration while Arista Networks jumped 14% on accelerating networking demand tied to Nvidia, CVS raising guidance and re-adding Zepbound to its formulary while publicly backing Eli Lilly with expanded GLP-1 support, Lilly beating on stronger-than-expected GLP-1 pricing, Disney rising 4% on parks strength and a TikTok short-form video partnership, Uber slipping despite record first-time user additions, and the JOLTS report showing softer job openings but likely not a signal Warsh will weigh.Join our live YouTube stream Monday through Friday at 8:30 AM EST:http://www.youtube.com/@TheMorningMarketBriefingPlease see disclosures:https://www.narwhal.com/disclosure
Watch the full episode on YouTube:We first covered Baseten last year when DeepSeek mania was at peak hype. Now they have raised a monster $13B round and become one of the new cohort of AI Infra decacorns that are (with Nvidia, Intel, and the semis complex) chief beneficiaries of the Inference Inflection. We return to Baseten at the peak of the 2026 edition of Open Weights debate. Ali has published a viral breakdown of Kimi K3:And since you last saw him, Philip has spoken at AI Engineer and written the definitive book on Inference Engineering spotted all over SF:Three years ago, inference engineering barely existed as a category.Today, it is one of the most critical disciplines in AI. Inference engineering inherently tackles a different question than standard model training: “How do you turn those weights from training into a product that is fast, reliable, and affordable at scale?” Focusing on these creates an entirely new optimization problem.In one recent GLM-5.2 experiment, quantizing more of the model actually preserved its benchmark quality while increasing throughput by 20%, because the errors introduced in different layers could cancel each other out.Inference is no longer just the final step after training. It is becoming its own engineering discipline, with its own research problems, infrastructure, and increasingly specialized roles.In this episode, Baseten's Philip Kiely and Ali Taha join swyx and Vibhu to explain what actually happens after a new open model is released and what it takes to turn “we generated a token” into a fast, reliable, production-ready API.We go deep on cache-aware routing, disaggregated prefill and decode, quantization, speculative decoding, KV-cache movement, model parallelism, GPU kernels, and the race to make frontier models up to 10× faster. Philip and Ali explain why inference optimizations can still produce gains of 20%, 100%, or even 200%; how quantization errors can cancel one another out; why identical weights can behave differently across clusters; and how Baseten grafted a Kimi vision encoder onto GLM-5.2 without changing the underlying language model.The conversation then expands beyond LLMs into NVIDIA Dynamo, mega kernels, Rubin, AI-specific chips, local inference, video generation, diffusion versus autoregressive models, and the enormous compute barrier to generating coherent long-form video. Finally, we explore the convergence of training and inference, continual learning through persistent KV cache, and the emerging loop where models help optimize the infrastructure that runs them.We discuss:* What happens when a 200,000-token request enters an inference system* Cache-aware routing and reusing previously computed KV cache* Why prefill and decode are increasingly handled by different GPUs* When dedicated deployments become cheaper and more reliable than shared APIs* How speculative decoding uses a smaller model to accelerate a larger one* Tool calling, structured outputs, and what LLMs actually do* What it takes to support a new open model on day zero* Grafting Kimi's vision encoder onto GLM-5.2* Retrofitting inefficient model layers with components from other architectures* Why models sometimes collapse into repeating the same token* How hardware, kernels, and race conditions create nondeterministic failures* Preserving model fidelity while making inference faster* How quantization errors can cancel each other out* Why inference optimizations still deliver gains of 20%, 100%, and 200%* How optimized serving can make a model up to 10× faster* NVIDIA Dynamo, KV-aware routing, and distributed model serving* Speculative decoding the speculative decoder* Why local AI is about making models less dumb while data-center AI is about making them less slow* Tensor, expert, and pipeline parallelism across GPUs* Hardware-aware model design, auto-tuning, and the case against mega kernels* Rubin and why inference is becoming a systems problem* Whether modern GPUs are evolving into programmable AI ASICs* Why enormous models like Kimi K3 require GB300-class hardware* Why open-source video generation still trails Veo, Kling, and other closed models* The quadratic attention bottleneck behind long-form AI video* Autoregressive video, real-time generation, and compounding quality drift* Why future video systems may combine autoregressive and diffusion architectures* Training for inference and inference for training* Continuous post-training, deployment, evaluation, and improvement loops* How GLM-5.2 helped optimize the kernels serving GLM-5.2 itself* Why faster networking could unlock dramatically faster decoding* Continual learning, KV-cache compaction, and persistent model memoryShow Notes* How to build a day-0 API for Kimi K3* 22580: From GPT2 to Kimi3, ExplainedPhilip Kiely* LinkedIn: https://www.linkedin.com/in/philipkiely* X: https://x.com/philipkiely* Inference Engineering: https://www.baseten.co/inference-engineering/Ali Taha* LinkedIn: https://www.linkedin.com/in/aliestaha/* X: https://x.com/waterloointernTimestamps00:00:00 Introduction and the 200K-Token Prompt00:03:18 Dedicated Deployments, Speculative Decoding, and Tool Calling00:11:26 Launching Production-Ready Open Models00:19:06 Model Retrofits, Failure Modes, and Nondeterminism00:28:22 Quantization and Canceling Errors00:32:15 The Race to 10× Faster Inference00:40:48 Dynamo, Speculation, and Local vs. Data-Center AI00:50:18 Model Parallelism, Auto-Tuning, and Mega Kernels01:00:55 Rubin, GPUs vs. ASICs, and Custom AI Chips01:10:03 Giant Models and the Limits of GPU Memory01:12:42 AI Video, Quadratic Attention, and Autoregressive Generation01:21:47 Audio, Images, and Diffusion Models01:27:32 Training, Self-Optimizing Models, and Continual Learning01:40:06 Closing ThoughtsTranscriptIntroduction: Baseten, Waterloo Intern, and Inference EngineeringSwyx [00:00:00]: Okay, we're here in the studio with Philip, old friend from Inference Engineering, the book, as well as Baseten and everything that you've done, you and I have done before, as well as Ali. Welcome.Ali [00:00:15]: Pleasure to meet you.Swyx [00:00:15]: Waterloo intern.Ali [00:00:16]: Waterloo intern, always.Swyx [00:00:17]: When did you get “Waterloo intern” as a handle?Ali [00:00:19]: As a handle? Oh.Ali [00:00:20]: I think the rebranding happened mid-March. When I saw it was open, I was like, “I have to take it. Up for grabs.”Philip [00:00:26]: The problem is that Ali is really good at his job and is not gonna be an intern much longer.Philip [00:00:30]: So we have to figure out who's gonna get the handle.Ali [00:00:33]: Well, I'll pass the torch over to the next intern.Swyx [00:00:34]: Oh, okay. It can be, like, you just pass it to another Waterloo grad.Ali [00:00:37]: To another Waterloo intern. No, bruh.Philip [00:00:39]: Yeah.Ali [00:00:39]: Intern.Swyx [00:00:40]: Intern, yeah.Ali [00:00:40]: And no.Philip [00:00:41]: You gotta get an intern from Waterloo.Ali [00:00:42]: Yeah, I've gotta get an intern from Waterloo.Swyx [00:00:44]: Right.Ali [00:00:44]: But they have to follow the path.Swyx [00:00:45]: Oh, it could, but it could come from Baseten, so it's like whoever Baseten gets from Waterloo.Ali [00:00:48]: Right.Swyx [00:00:49]: Has the title of Waterloo.Ali [00:00:50]: It stays in the ecosystem.Philip [00:00:51]: Exactly.Ali [00:00:52]: Halfway through the internship, you either get it or you're out.Philip [00:00:55]: You should also do, like, a big graduation ceremony where you change the handle.Ali [00:00:59]: Just say it.Philip [00:00:59]: For everybody.Swyx [00:01:00]: You guys are good at ceremonies, clearly. We had a nice launch of the book, very successful. But before we get into all that, I wanna start off with a fun question for you. Okay, you're an expert inference engineer. What happens when I send a long query, say two hundred thousand tokens into Baseten's inference? What's the process of query through GPU model routing, balancing, all that? What is all the stuff that we don't think about?Long Context Requests, KV Cache, and Cache-Aware RoutingPhilip [00:01:26]: With a long query specifically, the first thing that I'm gonna ask is, “Have you sent me this query before, or at least part of it?” and I really hope you have, because it's gonna be a lot easier for me and a lot cheaper for you. So the first thing that we're gonna look at is some cache-aware routing, where we're going to see, we probably have a number of instances, a number of replicas up serving whatever model you're hitting. We want to send this one to something with, number one, available prefill workers, and number two, ideally some cached input already there so that we can skip prefill on at least part of these two hundred thousand tokens. If you're doing two hundred thousand tokens, it's probably coding or a multi-turn agent or something where you would expect to have that cached. If you don't, we're gonna have to send it to a prefill worker. We've at least on certain models disaggregated prefill and decode, so you're going to have one set of GPUs that's solely going to process the input, create the KV cache, and get you your first token, and then that's going to be passed over to a separate set of GPUs, which is going to run decode. We're going to iteratively make those tokens. We're probably going to have some speculator model in front of that. I'm going to assume that you're doing coding, and because of that, our speculator model, which assumes you're doing coding, is gonna have a high draft token acceptance rate. If I'm wrong and you're asking me to summarize every Harry Potter book, it's gonna be slower. And then we stream that output to you and account for it, charge you, a couple of pennies and say, “Hey, would you like to send another one?”Swyx [00:03:04]: Except Baseten doesn't charge by pennies.Philip [00:03:07]: Well, yeah, we charge. I'm assuming that we're talking about the public model APIs. If you are setting up a dedicated deployment, then yeah, it's not pennies.Public APIs vs. Dedicated DeploymentsSwyx [00:03:18]: Yeah, one of the key differentiators when I was talking with Baseten initially was that people who want very high volume just need to rent by the box, ‘cause then it's up to you to figure out how to saturate the box.Ali [00:03:31]: And more often than not, it's, like, way cheaper if you're pushing, like, millions of tokens per hour, if you just pay per hour instead of pay per token.Philip [00:03:37]: Yeah, they do. I think that we've increasingly seen a lot of demand for the pay per token APIs, just because everyone wants to try open models, and then once they find a use case that's really sticky, then they move over to dedicated.Swyx [00:03:51]: Is there a best practice on when it's time to swap over?Philip [00:03:54]: Couple reasons. Yeah, reliability, that's a big one, right?Ali [00:03:57]: Like, if they have a very specific use case, they want you to train something specifically for them, like they want their own spec dec, for instance, for their own traffic.Swyx [00:04:04]: Spec dec is speculative decoding.Speculative Decoding and Custom SpeculatorsAli [00:04:05]: Speculative decoding, yeah.Swyx [00:04:07]: You have to explain.Ali [00:04:07]: Sorry. Like, speculative decoding is like, if you have a huge model, right? And so the model is going to be generating one token at a time every single turn, every single forward pass. So we attach, like, this little, like, parasite, like this layer that goes on top of the model, and this model just has to predict. It does three very fast autoregressive forward passes, and it will predict, like, three certain tokens, and then you do one forward stage over the entire original model in order to see if those predictions were correct or not, and then you accept them or you reject them. Now, this draft model is traffic specific, so if you, like, Philip said, if you're summarizing Harry Potter books, I can train exclusively that draft model on Harry Potter books, and I can guarantee you that I'm gonna accept the three tokens every single time. And so with that case, I increase your decode speed. I wouldn't be able to provide this to you if you're a shared endpointSwyx [00:04:53]: YeahAli [00:04:53]: ‘cause I have no idea if you're doing Harry Potter, if you're doing coding, if you're doing English. We don't know. Also, there was a thing in the book that mentioned that if they really cared about a specific threshold, chapter four, I think. Do you remember that?Philip [00:05:06]: Yeah. The things that you can do is you can set a specific, like, batch sizing, a specific, like, parallelism strategy if you're trying to optimize for, like, throughput versus latency. You can. Maybe a NVFP4 quant doesn't pass your benchmarks and you wanna run a model at higher precision, you could do that. There's just a bunch of reasons why you might wanna have your own endpoint and the biggest one, of course, just being, like, you don't have to deal with someone else doing a hundred million tokens of benchmarking traffic at the endpoint when you happen to be trying to serve your users.Swyx [00:05:40]: Yeah. I think one thing that is. That is a classic journey. Like, it's people is asking the, what happens when you type Google into the browser. Tool calling, is that just, you're generating JSON or is there more complication beyond that?Tool Calling, JSON, and Structured OutputsAli [00:05:58]: Certain customers that we have, they have their own post-trained models, and so they demand a tool calling that's not just, like parse a file or go find the weather. It's something that's very specific and you have to do post-training on this. And if the post-training on the model is not good or if the quantization after the post-training to get the inference to be fast, the model will struggle reading the JSON file and reading the tool calling. But it doesn't require its own like sandbox. It's not like it's going to use that tool calling to like escape a sandbox or like it doesn't have to be contained. It can just be a normal dedicated deployment. The challenge with tool calling more and more seems to be that the companies want certain tool calling which is a very sensitive thing to train. And because you're dealing with all of the JSON outputs, if it doesn't like close the end of the request in a very certain manner, you end up with a model that did the tool calling and like the thinking and so as a result of that, it didn't see the result and just hallucinated the result as it decoded. That seems to be the most challenging thing with tool calling, not really the sandboxes model.Philip [00:06:56]: Yeah, that's a challenge on the training side and then on the inference side, there's work that you can do to scope the possible output. So we published this at this point close to two years ago, the solution to this problem which is you make a state machine and you use that to constrain the output to a specific format. So this is the structured output problem. If you remember backSwyx [00:07:27]: Yeah, the specific grammar is,Philip [00:07:29]: Yeah, exactlySwyx [00:07:30]: GML had this thing.Philip [00:07:31]: Yeah. So it's like the old-school “make sure this is only JSON”, return only JSON orSwyx [00:07:38]: YeahPhilip [00:07:38]: Grandma's gonna die type of prompts.Swyx [00:07:39]: Is it BNF grammar? At some point OpenAI had released a thing that was like, yeah, if you want to constrain your output, write BNF grammar, back as NOR.Philip [00:07:47]: In our inference system, it's just a specified output format. And you get the guarantee that your output's gonna be structured along that format. And so applying that to tool calls can like help cut down on. You can still call the wrong tool or call no tool. It doesn't solve the certainty problem but it at least solves the output structuring problemSwyx [00:08:10]: YeahPhilip [00:08:10]: Within tool calls.Swyx [00:08:12]: And MCP is just another form of tool, right.Philip [00:08:14]: Yeah, exactly.Swyx [00:08:15]: As far as there's no special thing there.Philip [00:08:16]: The thing I'm always like explaining to people is the LLM is not capable of doing anything. It's only capable of making suggestions of what to do and then if those suggestions are formatted in a certain way and applied to a system that knows what to do with them, then an action occurs.Swyx [00:08:32]: Yeah. Part of the fun stuff is, this is solved outside of tool calling too. Like in an agent loop if the output is not correct or you're right, like reasoning, tool calling was done in the reasoning trace, just be like, “Oh, I don't know what to do. Let me just try again.” And it might get there after a few tries. And on your point of training, sometimes this is harder in smaller models, so you don't have the same exact quality outputAli [00:08:56]: Right.Swyx [00:08:57]: When you just swap from a big model, right?Ali [00:08:59]: Yeah. I will say that, before, I think we need to go back to inference engineering proper.Ali [00:09:04]: But, I had expected that something would replace JSON because it's hard to stream JSON ‘cause JSON must be complete and you must have open and close brackets and everything. So it's hard to parse something or validate something while it's being streamed. So people invented all sorts of things that are like, I forget the name of some of these alternatives, but it's something like TOML, something like YAML. But JSON seems to be dominant still.Philip [00:09:30]: The JSON outputs aren't that long, right? Like you could have a long-- ‘cause tool calls also contain the arguments in them and perhaps for a certain tool you might pass like a very long argument. But my impression of the median tool call is that it's a relatively small number of tokens, right? So I would expect that speculators are generally fairly good at something as formatted as JSON. And so you would have like a pretty fast decode step there and that the streaming wouldn't be as valuable, but maybe I'm wrong about that.Ali [00:10:02]: I think you're also bounded by the software or that the model is gonna integrate with if the software is built with JSON for the tool calls or if the company that you'- if your customer says that this is how our software works and our tools are interfaced with JSON, you can ask them to like, change their software and say like, “Yeah, this is gonna be better for the model.” but like with the right training shouldn't be that much of a difference. Also more profitable if it outputs more tokens probably.Swyx [00:10:25]: Depends on your business model.Swyx [00:10:27]: It really depends. But I will say that, as a writer with like experience a lot with generated output, I do try to move from text to JSON text which is very long JSON, right? Like there's paragraphs in every field because I'm trying to structure it, right?Philip [00:10:44]: Right.Swyx [00:10:44]: I want you to first make factual statements, then make opinions then make bullet point summaries, have dates, have entity references have your sources for references, all these things. Anyway, so these are things that like I think people who really experiment with structural output have to really care about. But, let's, let's recurse up the stack a little bit. Before we started recording, you mentioned something really cool, which is that there's a lot of engineering that-- inference engineering that goes on when a new model provider releases a new model, right? So let's call it GLM-5.2, Kimi K3. I had previously assumed, especially if it's like, well, GLM 5 to 5.1 to GLM-5.2, like that you've supported them before. Is it that much work?What It Takes to Support a New Open ModelAli [00:11:26]: It's a lot of work.Swyx [00:11:28]: Yeah. Okay. So like, a lot of people, all you guys, right whenever a new model launch like, people rush to say like, “Oh, Hugging Face supports this, Fireworks supports this, Spacetime supports this,” and I'm like, “Yeah, of course we support it.” But what goes into that? What goes intoPhilip [00:11:40]: I think it's more than just support it too, right? It benefits the consumer a lot. Like I think it was with Kimi K2.5 or GLM-5.2 the latest, there was an inference war, right? X provider is at 90 tokens a second. The next day we're at 150. The nextSwyx [00:11:55]: I kinda kicked that off with the GLM-5.2.Swyx [00:11:58]: I wrote a Twitter article about. It got like half a million views,Ali [00:12:02]: Based on being numberSwyx [00:12:03]: YeahAli [00:12:04]: Or it's for something else.Swyx [00:12:05]: Yeah. Which,Ali [00:12:06]: Oh my GodSwyx [00:12:07]: Which then got everyone really excited about, hey, how can we, bend tracks a little bit further and,Philip [00:12:14]: There's a difference between support the model, as in I can make a token out of this model, and support a model, as in I have a production-ready API from this model.Philip [00:12:26]: Getting to the point of I can make a token out of this model is not that hard because generally the, open source inference engines, vLLM, SGLang of the world oftentimes even receive weights ahead of time, maintainers do, or the people making the model merge PRs to ensure support. So you generally can, just get it working on the standard open source stack without too much pain in most cases. The challenge is, every inference company is gonna have own proprietary stack. Some open source components, some in-house stuff. And for any arbitrary model, there's going to be some new stuff. Sometimes you get lucky, like K, two five to two six was, like, pretty similar.Quantization, Speculators, and Production ReadinessAli [00:13:16]: Yeah. It was pure continued post-trainingPhilip [00:13:18]: YeahAli [00:13:18]: If I remember correctly.Philip [00:13:19]: Even in those cases, there's still stuff you have to do. You have to redo the quantization work. You're taking the model from. Generally, these models are not released in NVFP4, and we want them to be in NVFP4 for maximum Blackwell compatibility. So we have to perform that quantization, and, calibrate the quantization to make sure that we're not causing any regression in the model's intelligence. And then we also have to train the speculator, as we've talked about. Generally, we have. We have ZDR, zero data retention on our model APIs, so we don't know exactly the traffic that people are sending us, but we know what's popular. We know that coding use cases are popular. We know that agents, agentic use cases are popular. So we can get public data sets that are representative of that traffic and train general speculators. Now, with speculators today, you need to train the speculator using the base model itself because you're getting hidden states out of the model from running inference on these specific prompts, and that is the training data you use to create the speculator. So there's that process which you need the real model weights for. And then there's of course just the process of, standing up all the infrastructure behind it, loading all this stuff, testing it. And then when there's a new model with a newer architecture, I think that, like, the DeepSeek models tend to be the most challenging as they have, like, the most novel architectural stuff going on, model after model. But every new model has something. Kimi K2 had. Oh, sorry, GLM-5.2 hadAli [00:14:53]: Sparse attention.Philip [00:14:54]: Yeah,Ali [00:14:54]: YeahPhilip [00:14:54]: the DSA.Ali [00:14:55]: Right. Which is brought from DeepSeek.Philip [00:14:57]: Yeah. AndAli [00:14:59]: So you can copy-paste then?Philip [00:15:01]: It kindAli [00:15:01]: I don't know how this works.Philip [00:15:02]: So, like we had to, like, build support for that into our runtime. And you're right, like it is really interesting the way that all of these open source labs borrow from each other. For example, like GLM-5.2 doesn't have vision. So something that, Haley, a guy on our team, if we could take a look at this, he, like, grafted the Kimi vision encoder onto GLM-5.2.Retrofitting Vision into GLM-5.2Ali [00:15:27]: We'll be training the projector.Philip [00:15:28]: Exactly. So if you think about, like, the encoder, there's the encoder, which is the part that looks at the image and turns it into latent information, and then there's the projector which likeAli [00:15:38]: You can say latent space. It's okay.Philip [00:15:41]: And then there's the projector that maps it onto, the model itself, and then there's the model weights. You don't wanna mess with the model weights because you run a chance of making the model dumber at something else for the purpose of giving it vision. So instead, Haley started with just a projector, which is only a handful of millions of parameters.Ali [00:16:02]: That would be, yeah.Philip [00:16:02]: Yeah.Ali [00:16:03]: Can you show the training one?Ali [00:16:04]: Like the way it groksPhilip [00:16:05]: YeahAli [00:16:06]: Very interesting.Philip [00:16:06]: And maybeAli [00:16:07]: That right therePhilip [00:16:07]: Maybe Ali, you should take it from here. You've got a betterAli [00:16:10]: Ooh, double the sandPhilip [00:16:11]: Understanding of this than I do.Ali [00:16:11]: Yeah. You can see, like, he. The way he trained this is really cool. At the beginning, he was training it using just like, “Here's a picture of a mountain. Can you describe what's in this mountain?” And that caused it just like the first, learning walls. Like here you can see this all we're trying to teach it is to translate the encoded. Like it's already taken the encoder from Kimi K. It's taken the image. It'Philip [00:16:31]: Yeah. FrozenAli [00:16:31]: FrozenPhilip [00:16:32]: With adapter.Ali [00:16:32]: Exactly.Philip [00:16:33]: Yeah.Ali [00:16:33]: So the brain is frozen and the eyes are frozen. It's just we're tryingPhilip [00:16:37]: AlignAli [00:16:38]: Interconnect between the eye and the brain, right? So the projector. And so you take the tokens and then he's like, “Oh, can you describe what's in this image?” And he's like, “Oh, it's a mountain,” or it's a person or it's a human, whatever the case is. But that didn't cause complete understanding. So he changed it such that every image was associated with a data set of questions. Like, does this image have a white male? Does this image have birds in the top corner? Does this image have a scientist in it? All of that stuff. And it would have to answer questions correctly. And using not just training on describing an image, but being able to answer question, another question, answer over time. Like you can see the grokking, which is like genuinely insane, that retrofitting vision into a large LLM can learn to that extent. And even for images that it doesn't perform well on, for instance, if you ask it a picture of like Stephen Hawking, “Who is this?” Maybe it doesn't get it, but it will say something like, “This is Albert Einstein.” Like it still understandsPhilip [00:17:25]: Close enoughAli [00:17:26]: That this is a scientist who is a man who has, some significant achievements, all that stuff. So that's like really cool.Philip [00:17:32]: Yeah. So, we've covered Hao Tian before, who the author of the LLaVA paper that did this, a while ago. And I think that's very foundational work for anyone who hasn't done vision work before.Ali [00:17:41]: Same with the CLIP and MetaCLIP, where you go from just captioning to building out questionsPhilip [00:17:47]: RightAli [00:17:47]: Off the image and how much better you can get performance.Philip [00:17:50]: Right. Right. Right. Yeah. But what's, what's so exciting about this is if you look at a model like this. Now, this is a little bit more of a research project. It's not. It got to 56% on MMLU Pro, I think. So not quite frontier. But if you're running this model, you haven't suffered any loss on your GLM-5.2 quality. If you don't have an image, it'll just behave exactly the way it used to. And ultimatelyAli [00:18:14]: Which in the inference code you literally do not include the other part, right?Philip [00:18:18]: Yeah. You would just skip the encoder if you don't have an image input.Ali [00:18:22]: Okay.Philip [00:18:22]: Just confirming.Philip [00:18:23]: YeahAli [00:18:23]: Does it affect a lot on the overall inference side? Like you're not adding much, you're adding a very small vision encoder. These are typically likePhilip [00:18:30]: They're super fineAli [00:18:31]: Less than a billion parameters, right?Philip [00:18:32]: Yeah. It's, - There's a little bit less standardization among vision encodersSwyx [00:18:37]: YeahPhilip [00:18:37]: So the support matrix can be a little bit, sparser. But overall, yeah, it's a pretty, it's a pretty minor component of the overall system. And ultimately what you get out of the system is all of a sudden you have Kimi Vision, GLM weights, and DeepSeek attention all in one model.Open Source Model Grafting and Franken-MergesPhilip [00:18:56]: And that's, I think, a lot of the power and beauty of open source, is that you can take all of these different components and combine them together into a system that's better than anyoneSwyx [00:19:05]: YeahPhilip [00:19:05]: Can be individually.Swyx [00:19:06]: People used to say that you would also do Franken-merges where you would take likePhilip [00:19:10]: YeahSwyx [00:19:10]: Layers from each model.Swyx [00:19:11]: Does anyone do that anymore?Ali [00:19:13]: Well, to your point previously when you were mentioning like, the work that goes into supporting a model when it first comes out, like GLM-5.2 or MiniMax M3 or whatever the case is. Sometimes you do have to like, you do have to switch out some things. Like, for instance, the MiniMax M3 head uses full attention, and with full attention you end up with this like insane bottleneck in spec dec ‘cause you're doing auto-regressive token generation for three tokens, and you're doing this like N squared over all of the tokens that are in your sequence. Your KV cache is like very large because it's not sparse, it's not top K. So we find it better to like, okay, we're gonna replace this, we're gonna replace this layer with a layer from another model that's using like GQA, for instance. And then just with the right training, you can get it to have the same acceptance rate. So it is very possible to retrofit layers from other models and very much needed. If a layer is like inefficient, the training just becomes the challenge, like how do you ensure that you train it properly? Which again to your earlier point is like the mesh between training and inference. As in like you need very good training in order to do fast inference. That's like, I feel like more and more becoming true.Swyx [00:20:21]: Yeah. Anything else on the support side when you say like get it to fully production ready?Loop Detection, Race Conditions, and Non-DeterminismPhilip [00:20:26]: Yeah. I think that there's also a question of just, we can test a model to a pretty extensive degree, but we're trying to get it out quickly and then you see a bunch of other people test it and you get interesting results. There was an issue with, GLM briefly where we had some like mode collapses where it would just output the same token over and over again for certain prompts on certain temperatures. Like once you expose an endpoint to the real world, there's going to be, so many more varieties of things given to it that you're able to, discover and patch things. So it's not just a, day zero process, it's then like for the first week, for the first month, if a model remains popular, like how do you both fix bugs and then continue to push the envelope on performance?Ali [00:21:21]: What do you mean you don't want your model outputting S?Swyx [00:21:24]: Is there loop detection on that stuff, by the way? It still happens like quite a lot, which is surprising.Ali [00:21:30]: We have like we, in our endpoint, like if a model was to output the same token like four plus times, we just cut the generation. We say like, “Oh, sorry, this-- Like try again,” or like we will reprocess the request. ‘Cause we know then, like if it, like if, yeah, it's four times the same token, it's probably collapsed.Swyx [00:21:45]: Yeah. Is there a way to opt out in case I really want that?Ali [00:21:48]: You want that?Ali [00:21:50]: I think there's a way that we have to handle it. I'm not exactly certain, but I feel like in certain models, like when they output something like you can imagine, like a table for instance, and so they want, they wanna draw like 12 dashes and 12 dashes. Yeah, I think there's a way for that to happen. I think we only do it on certain tokens. Like we exclude certain special characters.Swyx [00:22:07]: Yeah.Ali [00:22:07]: So we only do it on like certain like S is the most common almost. GLM-5.2Swyx [00:22:11]: OhAli [00:22:11]: And I think it was DSV 4 as well. Like you'd just have like looping issues where like you literallySwyx [00:22:17]: ItAli [00:22:17]: Just have like S.Swyx [00:22:18]: Yeah. Is there a special, something special about S? No, just randomlyAli [00:22:21]: It just seems to be the one token involved.Swyx [00:22:23]: Yeah. And it'Philip [00:22:24]: Is thereSwyx [00:22:24]: And it's only temperature 0Ali [00:22:27]: NoSwyx [00:22:27]: Even at other temperaturesAli [00:22:27]: Even at like 0.9 or whatever, it will still, it will still collapse.Swyx [00:22:30]: That's weird, right?Ali [00:22:30]: It's, it is an inference problem to be honest, like a software problem. Like oftentimes, the image you run will-- like NVIDIA will release an image for instance, and if we will upstream the changes from their latest TensorRT-LLM image into our stack, we'll find that it fixes it. Or oftentimes this will only happen in an inference engine that you're using like SGLang. But if you were to switch to vLLM, that isn't the case. So it seems to be like an extremely like deterministic software issue and not really a model issue. It's not like a weights problem. Like I'- we'll say like, “Oh, it's a problem with the quant. We did PTQ wrong,” right? But that isn't, that doesn't make sense because the same weights used with a different inference engine does not repeat the problem. And sometimes it's, the kernels that are being used in the backend have like these very subtle sometimes race conditions, where if you were to use this model hosted on one cluster, you will never get this problem.Swyx [00:23:19]: Oh my God.Ali [00:23:19]: But if you host it on a different cluster, you will. And the reason is the KV cache transfer from a node to node in that one cluster is using a slower interconnect than the node to node in another cluster. So that exposes the race, whereas in another cluster it doesn't. So then you end up just like, okay, this model is not gonna be hosted on this cluster. We're gonna host it on, another cluster because that cluster exposed that problem. But then it ends up with like, okay, is it the software? Is it the model weights or is it the hardware?Swyx [00:23:42]: There is a thing about this with temperature 0 still not being deterministic, right?Ali [00:23:46]: Right.Swyx [00:23:46]: Mostly because of hardware. Even at temperature 0 same model, you won't always get the same output.Swyx [00:23:52]: Even-- But I'm surprised by the race condition one because, I thought PyTorch was a graph that like guarantees that you at least, execute things in the right order.Ali [00:24:02]: Well, yeah, true. Like I'm not, I'm not saying that there is. Like well, you have things like PTL optimizations where like you can start a kernel before the end of the previous kernel, and that's like ‘cause you want to do that because there'sSwyx [00:24:12]: It's like pipeliningAli [00:24:12]: Expense. Exactly.Swyx [00:24:13]: Yeah.Ali [00:24:13]: But it'- But you don't do it cleanly. Like you overlap a little bit of the execution. No, it is very possible that the kernel itself, like that one block that is supposed to be running in this instance of time, that kernel itself has a race condition. For instance, like a missing barrier. Like often if you're designing a kernel and you want it to make it to be very fast, if you don't test it extensively, you'll, you'll have certain threads access data points from registers before they've been written to by other threadsSwyx [00:24:36]: YeahAli [00:24:36]: For example, because like your barrier is wrong or your synchronization was wrong. But yeah, like the testing itself is very difficult in those like, andSwyx [00:24:42]: And there's no like borrow checkerAli [00:24:45]: What does that mean?Swyx [00:24:46]: Like Rust. Like the. If you're trying to have like memory safety It sounds like a comparable problem.Ali [00:24:52]: Well, yes, but you're working in CUDA, right, NVIDIA GPUs. Like- You just need a higher level language like modular Maybe that's what modular is supposed to do. I don't know.Quantization Quality and Vendor FidelityVibhu [00:25:00]: How do you see keeping quality of the model? So you talked about all these steps of, okay, you gotta do quantization, train your own speculative decoderAli [00:25:07]: RightVibhu [00:25:07]: Run on different hardware. Looking at other model providers, okay, you kicked off a inference speed race on the consumer end. What goes into keeping quality the same across them, right? Sure, you can run benchmarksAli [00:25:22]: YeahVibhu [00:25:22]: But, like, how do you determine how much quantization are there standards? What goes intoPhilip [00:25:27]: There's a few things on quality. Most inference optimizations are lossless. KV caching, for example. You are just recomputing or preventing recomputing the same values. Speculation, of course, if a draft token is wrong, it gets rejected. The main lossy optimization is quantization. And that really comes down to, number one, data format, number two, which parts of the model you choose to quantize, which layers, and number three, like doing a lot of calibration on the quantized weights, to ensure that you're preserving all the outliers. There's other tricks that you can do, though. A big one is long context, ‘cause one thing you asked at, right at the beginning is, “Oh, what's gonna happen if I send a 200,000 token request in?” So with a long input sequence, you need to, store a lot more information. You need to process a lot more tokens. And so even if a model has a context of a certain length, you might, as an inference provider, choose to build an API with a shorter context length, and of course a full length one as well. Because if someone doesn't need the full million token context, for example, you can get them better performance. I don't know if that's exactly like quality of the model. The way that I think about quality is to what degree are we faithfully serving the original model? If you think of a golden implementation of a model that performs exactly the way the model is designed to perform, I think of quality as how close are we getting to that, 100% fidelity of the model.Philip [00:27:13]: You can also, of course, think about quality from the training side and how do you push yourself past 100%. But when I think about purely inference optimizations, it's getting faster while staying as close to that 100% fidelity mark as possible. And certainly our standard internally is that, like you should not be able to tell the difference between our API and a, official API. I think Kimi in particular does a good job of vendor benchmarking hereAli [00:27:41]: YesPhilip [00:27:41]: Where they haveAli [00:27:42]: They released an actual vendor benchmark.Philip [00:27:43]: Exactly, yeah.Ali [00:27:44]: ‘Cause they accused, some people, Amazon? There was some provider that was not doing very well on Kimi's benchmark.Philip [00:27:50]: Yeah.Philip [00:27:51]: So, with Reflect we probablyVibhu [00:27:52]: This was a long time ago, right?Philip [00:27:54]: No.Ali [00:27:54]: Yeah, like threeVibhu [00:27:55]: They alsoAli [00:27:55]: Four, five months agoVibhu [00:27:57]: This also happened with, I don't remember which model, but they pulled out quite a few, and then they started a whole chart about this. It might have beenPhilip [00:28:03]: Kimi Vendor Verifier.Ali [00:28:04]: Yeah.Philip [00:28:05]: Yeah.Ali [00:28:05]: Yeah, ‘cause you, ‘cause you'd be pissed, right? Like if you'Philip [00:28:07]: Yeah.Ali [00:28:07]: If like if I'm a consumer and I'm using like Amazon's endpoint for instance, and I've used Kimi and I'm like, “Oh my God, like this is bad,” I'm not gonna say, “Oh, Amazon quantized the model in a bad way.” I'm gonna say, “Oh, Kimi sucks.” Right?Philip [00:28:17]: Yeah.Ali [00:28:17]: So it seems like that makes sense.Philip [00:28:19]: Yeah, they care. They care.Vibhu [00:28:21]: Justifiably.Ali [00:28:21]: Yeah, justifiably.Vibhu [00:28:22]: This is probably a stupid question, but just checking, has anything improved from main quantization?Philip [00:28:28]: Yeah.Vibhu [00:28:28]: Like, is quantization always strictly worse?Ali [00:28:30]: Well technicallyVibhu [00:28:32]: NoAli [00:28:32]: It's a lossy. QuantizationPhilip [00:28:33]: YeahAli [00:28:33]: Is a lossy, it's a lossy implementation.Philip [00:28:36]: Speed improvesVibhu [00:28:36]: Speed improves.Ali [00:28:37]: It the number, likeVibhu [00:28:38]: No, I' always look for inverse scaling laws.Philip [00:28:40]: Yeah.Ali [00:28:40]: Yeah.Vibhu [00:28:40]: This is something I learned from Noam Brown, where like things that normally act in one direction sometimes do.Philip [00:28:45]: Well, technically when you run a benchmark, because these models are deterministic, sometimes your,Ali [00:28:52]: YeahPhilip [00:28:52]: NVFP4 quant is like, two basis points higher than yourAli [00:28:56]: No, it's noise. It's noise.Philip [00:28:57]: Yeah, exactly. I'm like, yeah, it's, it's within. That's why I always say within margin of error.Philip [00:29:01]: And I stopped saying that because everyone assumes that what is, well, within some margin of error, we're barely inside of that to the worst, so we're saying. But yeah, sometimes it's just like, gives you a higher output score. But like Ali said, that's noise. To my knowledge, you're not necessarily making the results better. You're just trying to, again, like keep your fidelity as close to 100% to the original model.Layer Selection, KL Divergence, and Better QuantizationAli [00:29:27]: There is, to your point, research that we did on MP. I don't know if you are able to pullPhilip [00:29:31]: YeahAli [00:29:32]: A tweet we did. One of our research interns, Joshua, I think it's a tweet on how we have 20% better quantized GLM-5.2 than NVIDIA. Essentially what we found throughout like this month research is, okay, quantization is a lossy. It's. You're compressing the data from, occupying 16 bits to occupying, four bits, for instance. And so you're losing some information, and you're trying to minimize that. And so when I say that I'm gonna quantize the model, my job becomes how do I find the layers that I can quantize, and how to find the layers to not. For instance, with image models, I don't quantize modulation layers, and I don't quantize out projections because those two are. Like out projection is what you see as the user. Modulation is what the model sees or understands. Right, exactly. And so to his paper, do you have the. It doesn't have the. Yeah. It's a long paper. I don't know if I can findVibhu [00:30:25]: If there's a part to search or it's probably in the thread.Ali [00:30:28]: It's probably in the thread.Vibhu [00:30:29]: Yeah.Ali [00:30:29]: But the long and the short is it is very possible that quantizing more of the model makes the results. Like if I have a model that I quantize layers one, five, and 10, and another model where I only quantize layers one and It is possible that the model in which I quantized more information is going to perform better because the quantization errors have canceled out. And so what Joshua showed in his mathematical proof where he had like a verifier in, is that you can predict which layers are going to have quantization errors that will cancel out with each other, and you choose to quantize those layers. And so the result of doing this mathematical quantization is you end up with a model that's 20% more quantized than another provider, so you get 20% more throughput of it because there's more layers than running an NVFP4, and your quality is better than that other quant because the layers that you chose to quantize have their errors cancel out, like one layer skewed to the right one layer skewed to the left, one layer skewed to the right. Your final logits distribution is more similar to the original distribution of the model, so you have better fidelity. And so the way we proved this was with KL divergence. So instead of just scoring on the benchmarks, we scored the KL divergence between the logit distribution of the quantized model and the logit distribution of the original full precision model, and we showed that with this technique we get. If your probability distribution on the logits which token it wants to select is more of the same as the original model, you're probably gonna end up staying true to the original model. So yeah, so it seems like previously before this, it seemed like the industry was, well, the more you quantize, the worse it's gonna be, ‘cause the more loss you introduce. That's not exactly, not necessarily true. So yeah, doesn't improve it, but can cancel out.Philip [00:31:57]: I think it might be this, but reminds me a good bit about pruning where you can prune off certain layers.Philip [00:32:03]: But very interesting. Didn't know this was a whole paper you guys put out.Ali [00:32:06]: It's. Fun fact, it was originally 72 pages, this paper, and then we decidedPhilip [00:32:11]: WowAli [00:32:11]: We can't tell. We couldn't release it. So it's now 45.Swyx [00:32:15]: Still 39 pages, so very substantive. We talked about evals and all these things and, like what's possible in terms of speedup? Like it's like probably like the numberInference Speedups and BenchmarkingSwyx [00:32:25]: Thing that people do wanna care about, and it's something that you wrote about in your post. Like official API is 70 tokens per second, and you push it up to 90. Is that like a normal thing?Philip [00:32:36]: So what's cool about working in inference, the reason that I think inference is going to be a useful place to do engineering for a long time, is that if you look at highly optimized domains like, say, finance, if you're in finance, you measure how much better you got in basis points. It's like, “Oh, I got five basis points better, like twentieth of 1% better,” that's huge news because everything is so optimized. When we publish optimizations, it's 20%, it's 100% it's 200%. So there's still probably like a lot further to go, honestly. Like you'll, you'll know that inference is pretty much solved when researchers start publishing about how they got 1% faster at something.Swyx [00:33:19]: Which by the way, because I am from the finance background, in the ‘70s, that was the margin at the time. When you did quantitative finance research, you would findAli [00:33:27]: And like 20%, tens of percent.Swyx [00:33:29]: That's. Yes.Philip [00:33:29]: Yeah.Swyx [00:33:30]: And now it'Philip [00:33:31]: Tiny fractionsSwyx [00:33:32]: For those people interested, look up Andrew Lo's paper. He had a really interesting illustration of quant, stat arb, distribution, narrowing down from like those kinds of 20% differences in the ‘70s, down to nothing today, which is very cool.Philip [00:33:48]: Exactly, and we're at the beginning of the same type of thing. Now benchmarking is hard. I think anyone will tell you that, and benchmarking provider speeds is hard because there's so many variables that go into it. What hardware are you using? How much load do you have on the system? What's the exact nature of the prompts and input and output sequence lengths? All that stuff. But overall, when you start stacking these improvements, you're looking at multiples. You can look at it. The most common form, of course, is TPS, tokens per second, which is bad naming by us in the industry, ‘cause there's two tokens per second. There's tokens per second, the throughput number, and the latency number.Ali [00:34:31]: TTMT, yeah.Philip [00:34:32]: Like total tokens per second out of the, out of the GPU as a throughput number. Most people only care about tokens per second as the latency number, which we should call ITL, intertoken latency, but we don't.Philip [00:34:44]: Anyway, so you can imagine a standard API without many optimizations for a 1 trillion parameter model operating somewhere in the 30 to 50 tokens per second range for reasonable traffic profile. And we generally see the goal of, pushing to 10X that. But, not necessarily day zero, but by stacking enough optimizations, if you have, say like four optimizations, each of which doubles performance. Or sorry, three optimizations, each of which doubles performance, then you stack that up, that's an 8X gain. That's the order of magnitude that we're working with in this space. We're trying to make things substantially faster, not just go from like 70 to 90.Swyx [00:35:38]: Are you saying you've. You have done that?Philip [00:35:40]: So let's say you have as a reasonable baseline, 30 or 40 tokens per second. You can achieve 10X that. So like on GLM-5.2, if you run it unquantized, perhaps on H100s even, and you're just using an off-the-shelf inference engine with no particular optimizations, no speculator, nothing extra around like KV routing, no disaggregation, you're, you're probably, yeah, looking at that like 30 to 40. You think that's like a reasonable baseline?Swyx [00:36:12]: Right. Right.Philip [00:36:12]: To get to something like 10X, there's a lot of trade-offs that you're making. If we're running at more like a 300, 400 tokens per second range, you are using the best hardware possible. You have a optimized speculator. You have done all of your quantization work. You are Seeing a pretty high cache hit rate. You are running with a reasonably small batch size and a parallelism configuration that is tuned for latency versus throughput, but it is possible. So the spreads that you see if you, like, go on artificial analysis or you go on OpenRouter and you look at, the worst provider to the best provider, oftentimes can hit that range. 10X is of course very aggressive. It's oftentimes maybe more of a four to six times improvement. But that's the performance that makes us really excited, is when we can get these huge gains, not just go from 70 to 90 tokens.Stacking Optimizations: NVFP4, Speculation, and DisaggregationAli [00:37:19]: It's also, like, hardware dependent. Like, ifPhilip [00:37:20]: YeahAli [00:37:20]: If you have a thing where you're serving it on just, like, a node of H100s and then you throw, like, you shard the model across, like, four nodes of B200s. Like, you can definitely increase the speed with just throwing more hardware at it. Like, normalizing for the same exact hardware and the same number of GPUs.Philip [00:37:35]: Yeah. Then you're looking at, like, a two to 4X improvementAli [00:37:38]: Right. RightPhilip [00:37:38]: Depending on the inference optimizations. So yeah, it's. Some of it's, what's the call, and some of it's who's the driver.Vibhu [00:37:46]: If you break down the two to 4X, say the example is run GLM-5.2Ali [00:37:51]: YeahVibhu [00:37:51]: On B200sAli [00:37:53]: YeahVibhu [00:37:53]: Single node, right? What's, like, the cost trade-off for effort to get, like, the last bit of juice out versus what should people just think of, right?Ali [00:38:01]: Spectre quantization. Yeah.Vibhu [00:38:03]: Spectre quantization.Ali [00:38:04]: That's, that's, that's like 95%. LikeVibhu [00:38:06]: And how far does that get you? And how easy is that for the average person to do? So say right I wanna throw the weights of GLM-5.2 on a node of B200s, how easy is it to find speculative decoder- decoder model or already quantized model? How much work goes into it?Philip [00:38:23]: If you're doing it up front, it's quite a lot of work. If you're doing it today, there's going to be people who have published things that you can just, you can just grab some NVFP4 weights. You can grab a speculator. Yeah, if we're thinking about, like, what are the 2Xs we're stacking, going from, BF16 to NVFP4 is, it's not quite a 2X, right? It's like. I think it's about, like, 30 to 40%, from 16 to 8, and then another 30 to 40% multiplied from, 8 to 4. So that doesn't quite get you a 2X, but, like, roughly a 2X. Speculator, roughly a 2X. Disagg on top of that if you're able to get enough hardware and put enough traffic through it, another roughly a 2X. And then you add in some, double-digit percent increase from having just a better runtime with, the latest kernels and stuff behind it. And that's how it stacks up.Ali [00:39:21]: YeahPhilip [00:39:21]: So building each of those, like, building the, quantized weights is, for someone who really knows what they're doing, hours to days of work. Building the speculator, again, like, hours to days of work. And the, disagg setup, hours to days. Well okay, but like once you haveAli [00:39:39]: Once set up. Once set up. YeahPhilip [00:39:40]: Yeah, getting disagg working for the first time, I'm saying, of course, is very difficult.Philip [00:39:44]: The marginal implementationAli [00:39:48]: Like, if you're just grabbing, like if you are a person, like just a normal consumer who has access to, like, a node of B200s and you're wondering, “How can I just host it myself?” You don't need to quantize the model yourself. There's always gonna be, like, an open source quantized checkpoint. NVIDIA's gonna push one out if no one else does. You. Usually, the providers will have their own spec dec that they've trained as well. You don't need to train your own spec dec. You can just use that as well.Philip [00:40:09]: Yeah. Like, GLM-5.2 has its own MTP.Ali [00:40:13]: Right. Right.Vibhu [00:40:14]: What's multi token prediction?Philip [00:40:15]: Yes.Ali [00:40:16]: I'm justVibhu [00:40:16]: Can you explain that?Ali [00:40:16]: I'm just an expert.Ali [00:40:18]: I can do it for you in case I get it wrong?Vibhu [00:40:20]: No.Vibhu [00:40:21]: Yeah, you should correct if we're wrong, but their multi-token prediction can be used for self-speculative decoding.Ali [00:40:27]: I'm not sure. I'm not gonna correct that.Vibhu [00:40:28]: Okay. I'm semi-confident in thatAli [00:40:30]: Okay. YeahVibhu [00:40:30]: But someone can check. But it's useful to paint the story of, okay, not just the average person, but say a company wants to switch from serverless inference I wanna throw this up on. I wanna rent some GPUs, throw it up. These are the steps you take to do significantly faster than just put it behind vLLM.Ali [00:40:48]: Right.Vibhu [00:40:49]: I was waiting for a mention of Dynamo.Vibhu [00:40:51]: I feel like, that's supposed to be the baseline that you measure against.Dynamo, KV Routing, and Disaggregation ToolkitsPhilip [00:40:55]: I would think of Dynamo as less of a box system and more of a toolkit for building with. So when we talk about doing aware routing, when we talk about doing KV offloading, when we talk about doing, PD disaggregation, Dynamo fundamentally is. By the way, Dynamo is an open source library from NVIDIA.Ali [00:41:17]: We've done a pod with KylePhilip [00:41:18]: OkayAli [00:41:19]: Kyle Cranin.Philip [00:41:19]: Cool. So then your listeners know then that it supports all the different inference frameworks. And it is multi hardware, which is interesting.Ali [00:41:28]: But it's just a router, it's not like an optimizer layer.Philip [00:41:30]: Yeah. All it does, like, what Dynamo is good at, it is a library for moving information around your cluster, around your hardware. So if you have, KV cache on one place and you need it to be somewhere else, Dynamo coordinates NIXL for you to move that around.Philip [00:41:49]: That doesn't mean that, like, out of the box, you just say, “Pip install Dynamo,” and then you get, like, a massive performance speed up. It's more of a developer toolkit.Ali [00:42:01]: Yeah. I would have said it would. It comes with a set of defaults that you can then swap out.Philip [00:42:06]: It does. If the industry at large, I think, was, like, rolling out all of these deployments, standard, then I think it would be, like, a credible baseline. But, we've got to, we've got to benchmark against, like, what we're seeing in the wild.Speculative Decoding Methods: Medusa, EAGLE, n-Gram, and Spec-SpecVibhu [00:42:23]: I did wanna talk a little bit more about PD disagg, because that is probably, like, number three after quantized and speculative decoding. In your book though, I was just gonna pull out the book.Philip [00:42:31]: Yeah.Vibhu [00:42:32]: Like section 522 on Medusa, 523 on EAGLEPhilip [00:42:35]: YeahVibhu [00:42:36]: 524 on gram.Philip [00:42:37]: It's 55, would be disaggregationAli [00:42:42]: Yeah. Well, no, I just wanted to dwell a little bitPhilip [00:42:44]: YeahAli [00:42:44]: The other. Like, so what do you choose to include? What do you choose to not to include? Because there was all these other techniques.Philip [00:42:51]: Yeah.Ali [00:42:51]: Are these still relevant? Because I think they came out, like, a year and a half ago maybe.Vibhu [00:42:55]: Medusa is quite old.Philip [00:42:56]: Yeah, Medusa's old.Ali [00:42:58]: It was old.Vibhu [00:42:58]: But is it in the book as a good, here'sPhilip [00:43:01]: BaselineVibhu [00:43:01]: Baseline vanilla understand it?Philip [00:43:02]: Like you should know this.Vibhu [00:43:03]: Like I read the paper, I'm like, “ it makes so much sense.”Philip [00:43:05]: Yeah.Philip [00:43:05]: So with the book, I had a couple goals. One was to give people just a working vocabulary for the space as a whole, and the other was to give them some intuition about how each of these techniques works. As I mentioned in my AI Engineer talk, which is the first public addendum to this, the speculation space has moved much faster than everything else. So yeah, even at the time that I wrote the book Medusa, I very much included as a way for people to understand how the space evolved rather than what the most modern technique is. And now of course, there's DFlash, dSpark. There's, there's newer techniques even than EAGLE, although EAGLE is still very commonly used.Ali [00:43:51]: SpecSpecta.Philip [00:43:52]: Yes. Speculative decoding.Vibhu [00:43:54]: What canAli [00:43:56]: Oh, it's a paper by Tri Dao and it's like, it's doing speculative decodingVibhu [00:44:00]: HuhAli [00:44:01]: For the speculative decoder.Philip [00:44:02]: Oh, in spec- oh my God.Ali [00:44:02]: It's literally just an another. It's like, yeah, that's the most simple way to explain it, and it seems like he got trivial speed ups there. But it seems that the complexity with training, it's almost like in our mind at least, it's almost as complex as training GANs. Like it's like a very delicate balance and oftentimes you, it's just but yeah, it's literally speculative decoding on speculative decoding.Vibhu [00:44:21]: Speculative.Ali [00:44:22]: Yeah. We saw this paper.Vibhu [00:44:24]: It's interesting, right?Ali [00:44:24]: Yeah.Vibhu [00:44:24]: I wouldn't even expect it to be very particular to train, I wouldAli [00:44:29]: Right.Vibhu [00:44:29]: The naive part of me is like, okay, train speculative decoder.Ali [00:44:32]: But like, and it makes sense, like the whole idea of speculative decoding is you. It's like, it's like almost like the iPhone auto predict version but for a normal model, right? Like you're just, you're just, generating three tokens and you're like, okay, I'll do prefill on them. And so you save those three turns for your original model. Now your speculative decoder is doing three turns of auto regression, so why not just have an even smaller model?Ali [00:44:53]: The other question there is what are the size of speculators? So say forPhilip [00:44:58]: Right. It's like a billion parameters.Ali [00:45:01]: Like for MiniMax, it's. Yeah. It's like one layer. It's like one 60th of the original model usually.Philip [00:45:06]: Yeah. I think we should do a paper when we get back to the office.Philip [00:45:10]: SpeculativeAli [00:45:11]: SpeculativePhilip [00:45:11]: Decoding.Ali [00:45:13]: No, it's, it does seem like how, when do you stop? But then it also seems like if you're able to train spec-spec decode for instance, right? Like if you're able to have a small model that is accurately predicts what the intermediate speculator is gonna predict, that is able to predict what the original target model's gonna predict, then why not just use that smallest model directly, right?Vibhu [00:45:34]: Yeah. This isAli [00:45:35]: Like it seems likeVibhu [00:45:35]: Adjacent to the routing problem.Ali [00:45:36]: Right.Vibhu [00:45:36]: Yeah.Ali [00:45:36]: Right.Philip [00:45:37]: The thing with speculators is one of the practical constraints on using them is that you do have to run a small model on the same hardware that you're running the big model on. There is a orchestration and resource competition problem inherent in that, and that is one of the constraints on speculation in general, is that draft tokens cost resources to create and cost software complexity to manage. And so if you have like infinitely recursive speculators, you add in quite a bit of that complexity on the actual implementation within the inference engine as well, not just in the training process.Vibhu [00:46:17]: I was gonna say, I would wonder if you could do similar, like distillation and pruning of, it's the same thing, it's just a model. Can we not just distill a lot of the weights, quantize the speculator, out of my domain? The question that also comes up is, this is all for big server workloads, right? How much of this applies to, say I have this MacBook, I wanna run Gemma really efficiently. Similar problems, not the same?Local AI vs. Data Center InferencePhilip [00:46:45]: Pretty different. I talked to Selo, about this on his podcast a couple weeks ago. The difference between inference engineering for the data center and for production workloads versus inference engineering for local AI, is that we start with fundamentally like different constraints and different goals. With local AI, it's how do I fit this model onto my hardware and then make it less dumb? And with data center influence, it's how do I load this model and then make it less slow? And we care about less dumb, and they care about less slow. But the local AI inference engineering ecosystem, I think has a lot for us to learn from in the data center space. They are experts in various forms of quantization, including dynamic quantization that we just don't touch, in the pruning, in the distillation, in the, layer removal. There'Ali [00:47:42]: Layer removal matters less.Philip [00:47:43]: Yeah. There'Ali [00:47:44]: No one loves pruning really.Philip [00:47:45]: Yeah. Well, but the, but they doVibhu [00:47:46]: Which is surprising, right? But that's, that's a whole different thingPhilip [00:47:48]: Just to fit something on the laptop.Ali [00:47:50]: Right.Philip [00:47:50]: So yeah, it's a, it's an interesting, it's an interesting space. Not necessarily that like their techniques make sense for us to do in the data center, because we have different resources and different goals, but more that the process as well as the openness of that field is something to, admire.Ali [00:48:12]: Yeah. Like to your point, like, certain optimizations that would. Like for instance, Turbo Quantum Sharper, like it made such huge hype on that and we did like a whole deep dive on Twitter and like said, what is it? How does it work? Why is it good or not? And it took off and it was implemented on local devices because your memory bandwidth is so slow on like a MacBook, for instance. But try putting the same thing on like an NVIDIA GPU on a B200 Turbo quant would not be. Like, it would not be used. Like, NVIDIA - Like, NVIDIA made it clear that this is not a good optimization, and we've seen it firsthand where the overhead of doing dequantization, quantization of, in the kernel itself with turbo quant kernel, each end is much slower than the time that you save from doing the bandwidth. ‘Cause on the B200s, you have like 3.5 terabytes per second. You don't need decrease the storage that much. You don't need to do, FP4 KV cache. You don't need to use a requant. There's, there's, there's better optimizations to be made. But on Edge devices, it's extremely important, it's extremely useful. So, seems to be, like, different optimizations there, but then they're all uniquely combined with like all you wanna quantize the model, you wanna do speculative decoding, like certain common prefixes with bothPhilip [00:49:18]: Principles.Ali [00:49:19]: Yeah, exactly. Exactly. Exactly.Philip [00:49:20]: They also do a lot of work on, model parallelism, especially over, heterogeneous topology, where you have, some sparks and they are wired together with, Ethernet, DGX sparks.Ali [00:49:35]: Yeah, this is the Exo Labs guys.Philip [00:49:36]: Yeah. You have, a nu
Stephen DiAdamo, co-founder and CTO of Qoro Quantum, is interviewed by Yuval Boger. DiAdamo discusses Qoro's position as a software middleware company that abstracts hardware away from applications, walking through the Divi SDK, circuit serialization and parallelization, and an orchestrator that automatically selects between 12 simulation methods on CPU and GPU or dispatches to QPUs from any vendor. They explore how Qoro plugs into HPC schedulers like SLURM rather than replacing them, the CESGA proof of concept built on the CUNQA platform, the "150,000 lines of code to 20" claim, and the argument that multi-QPU centers are needed today for fallback and resilience and for scaling applications. DiAdamo also reflects on developer experience for students learning QAOA and variational algorithms, and Qoro's two-year history.
Deze Einde van de Week Live is ook te bekijken op https://youtu.be/PZGrFJBfasw Deze talkshow wordt mede mogelijk gemaakt door MSI. Alle meningen in deze video zijn onze eigen. MSI heeft inhoudelijk geen inspraak op de content en zien de video net als jullie hier voor het eerst op de site. Klaar om het weekend te betreden? Wij wel. Ook al gaat het wat minder warm worden dan de laatste dagen. Wie weet is dat ook veel fijner ook. Wij gaan de traditionele opwarmer van het weekend voor je verzorgen. Klaar staat namelijk een nieuwe editie van Einde van de Week Live. Daan, JJ en Koos zitten klaar om bij te praten over alles wat er de afgelopen week toe deed. Een onderwerp dat aan bod komt, is bijvoorbeeld de multiplayer reveal van de testosteron shooter van XBOX: Gears of War E-Day. Wat vonden de heren ervan en zijn ze fans van de third person multiplayer? Ze kijken of ze het gerucht dat Rockstar in augustus met een gameplaytrailer komt, serieus kunnen nemen. En wat hebben The Odyssey en Assassin’s Creed: Odyssey met elkaar te maken? De antwoorden op al deze vragen vind je in de Einde van de Week Live van vrijdag 31 juli 2026. Knalt de multiplayer van Gears of War E-Day nog net zo hard als twintig jaar geleden? Andere onderwerpen die in deze editie van de vrijdagse talkshow voorbij komen, zijn onder andere de vijf eindes van Silent Hill Townfall, de nieuwe beelden van Amsterdam 1666 en de grootse plannen van Krafton met Subnautica. Check de Cyborg A15 B2 gaminglaptop en profiteer van de scherpe prijs MSI zet deze week de MSI Cyborg A15 B2 in de spotlights. Een gaminglaptop met een AMD Ryzen 7 260 CPU, een NVIDIA GeForce RTX 5060 GPU, 16GB RAM aan intern geheugen, een 144Hz Full HD Display, een 512GB SSD en een 4-zone RGB toetsenbord. Een mooi pakket dus bij elkaar, dat de komende week hier bij GamePC voor een scherpe prijs aangeboden wordt.Wil je adverteren bij de podcast Gamekings óf misschien bij een andere podcast van ILVY Network? Mail dan naar management@ilvy.com en/of kijk even op de website : https://ilvy.com/podcastSee omnystudio.com/listener for privacy information.
What becomes possible when enterprise computer vision no longer depends on expensive GPU infrastructure? In this episode of Tech Talks Daily, I speak with Glenn Jocher, founder and CEO of Ultralytics, about YOLO26, CPU inference, edge AI, open vocabulary vision, deployment economics, and the practical work required to move computer vision from a promising pilot into production. Glenn's route into AI began inside the U.S. intelligence community. He worked with the National Geospatial Intelligence Agency and Defense Intelligence Agency on particle physics applications, attempting to detect and track antineutrinos. Antineutrinos are extraordinarily difficult to detect because they pass through almost everything. Glenn describes them as the perfect spy. While searching for better detection methods, he discovered that computer vision researchers were solving similar problems with images. His original attempt to transfer those techniques into particle physics did not succeed. However, the work introduced him to a field where the technology could create a visible effect on everyday life. That led him toward open source development and eventually the YOLO models for object detection, classification, segmentation, and tracking. Glenn believes computer vision research has historically placed too much attention on small gains in accuracy while overlooking deployment economics. A model can perform impressively inside a laboratory and still remain unsuitable for a factory, warehouse, store, vehicle, drone, or medical environment. Price, latency, power consumption, data privacy, and deployment speed can determine whether the technology is commercially useful. This led Glenn and Ultralytics toward smaller models capable of running close to where images and video are generated. YOLO26 continues that approach with architectural changes designed specifically for CPU inference. Glenn says the model can process camera streams in real time at 30 frames per second and run across Intel CPUs, AMD CPUs, and lower power devices such as Raspberry Pi computers. This matters because specialist GPUs can increase the equipment cost and power requirements of a computer vision project. Running inference on existing CPUs or edge hardware can make deployment economically possible across larger numbers of cameras and locations. The scale already involved is difficult to comprehend. Glenn says Ultralytics models now process approximately three billion inference jobs each day, equivalent to around 30,000 every second. These jobs include images, videos, and collections of images being analyzed to detect, segment, or track objects. He attributes the platform's maturity to thousands of mistakes and bugs corrected through a rapid feedback cycle. New models are released, users report problems and request features, and the team incorporates that information into later versions. We also discuss the respective roles of cloud and edge infrastructure. Glenn sees cloud platforms continuing to provide the computing power required for training, while computer vision inference often belongs at the edge. Local processing can reduce latency, control operating costs, and keep sensitive video or medical information closer to where it was created. The smallest YOLO model is approximately three megabytes, according to Glenn. That allows it to reach mobile phones, vehicles, drones, battery powered devices, and other environments where a large language model would be impractical. Open vocabulary vision provides another development. Traditional object detection models are trained to recognize a fixed collection of objects. If a model learns to detect dogs and the user later wants it to detect cats, retraining can cause it to forget earlier knowledge unless both categories appear in the new training data. Glenn explains how promptable models can identify common everyday objects from text or visual instructions without additional training. A user could request a person wearing a blue shirt and white shoes, for example, and the system could search an image for that description. That flexibility could benefit businesses whose requirements change regularly. It reduces the need to create and label a new data set every time the company wants the model to recognize another common object. The range of current applications is already extensive. Glenn describes YOLO being used across robotics, parking, industrial safety, PPE detection, warehouses, aviation, security, traffic management, food quality, and manufacturing. Some of his favorite examples involve environmental problems. One company uses YOLO with underwater vehicles to identify and recover plastic from the ocean. Other applications detect smoke and fire early enough to support forest fire response. For leaders considering computer vision, Glenn recommends beginning with a defined problem and measurable outcome. A manufacturing company may want to reduce defects, but it still needs labeled examples showing the model what acceptable and defective products look like. He advises testing the idea through a limited pilot, measuring the return, and expanding only when the evidence supports further investment. Computer vision has become easier to deploy, but practical problems involving data, cameras, integration, reliability, and operating conditions still separate a demonstration from a production system. Could CPU inference and open vocabulary models make computer vision practical for processes your organization previously considered too expensive? Listen to the episode and share your thoughts with me. Useful Links Ultralytics website Ultralytics Platform
Topics covered in this episode: Some more things about Django I've been enjoying Who cleans up after the vibe-coding party? Where Did All Your AI Tokens Go? AgentsView to the rescue! Careful with phishing all Extras Joke Watch on YouTube About the show Sponsored by us! Support our work through: Our courses at Talk Python Consulting from Six Feet Up Connect with the hosts Michael: Mastodon / BlueSky / X / LinkedIn Calvin: Mastodon / BlueSky / X / LinkedIn Show: Mastodon / BlueSky / X Join us on YouTube at pythonbytes.fm/live to be part of the audience. Usually Tuesday at 7am PT. Older video versions available there too. Finally, if you want an artisanal, hand-crafted digest of every week of the show notes in email form? Add your name and email to our friends of the show list, we'll never share it. Calvin #1: Some more things about Django I've been enjoying Julia Evans is learning "2010-style" web dev (Django + SQL + server-rendered HTML) after years of Go backends and JS-heavy frontends Query builders: likes defining custom QuerySet classes with chainable filter methods (.approved().future().with_tags()) — more readable than raw SQL Template filters: highlights urlize, linebreaksbr, json_script, and especially querystring for building/modifying query-string links in templates Migrations: still loves Django's auto-generated migrations — 19 and counting on her project Skips inheritance for class-based views; prefers function-based views for sharing code, though fine using Django's own mixins/interfaces Performance surprise: CPU profiling (via py-spy) — not slow DB queries — revealed the culprit; she'd accidentally disabled the cached template loader, and re-enabling it took throughput from ~2-3 req/s to ~12 req/s on a $10/mo VM Michael #2: Who cleans up after the vibe-coding party? FT Magazine piece by Sam Learner (July 11) on AI coding tools overwhelming open source maintainers - sent in by listener Dylan McConnell, whose main point was that this ran in the Financial Times, not a dev blog. cURL as the case study - Daniel Stenberg has been the only full-time person on it for years; libcurl has been installed an estimated 20+ billion times with 3,000+ listed contributors. Bug bounty killed - cURL ended its paid security bounty program in January, citing an "explosion of AI slop reports" that take real time to debunk and drain morale. Extractive contributions - authoring a PR is now nearly free, reviewing one still costs a human; tldraw's Steve Ruiz closed outside contributions entirely, asking why he'd want someone else writing the easy part. Guido weighs in - van Rossum says projects are holding emergency meetings over the slop flow, and notes LLM patches tend to touch unrelated parts of a file, making review more tedious. "Vibe Coding Kills Open Source" - paper from Miklós Koren's group: packages frequently recommended by coding models saw big download jumps with no matching engagement, breaking the reputation loop that sustains maintainers. Stack Overflow flatlined - over 100,000 questions a month before ChatGPT, under 1,500 last month, with the response rate cut roughly in half; the public archive is now stale training data. The course-creator angle - Josh Comeau's newest web dev course launched at about a third of prior enrollment, and he worries about devs who never learn which questions to ask. But the most interesting portion is what was omitted. Focused on: The end of the curl bug-bounty Omitted: High-Quality Chaos Why the omission is interesting It fits a narrative. The FT piece is a maintenance-and-decline story, and January-Stenberg is a perfect witness for it. April-Stenberg complicates it - same person, same project, better data, opposite direction on the specific claim being used. The tell is already in the article. Learner quotes Stenberg saying AI tools are much better at finding problems than fixing them. That's the April thesis in one line, and it goes undeveloped. Reason for the shift is process, not vibes. Killing the bounty removed the cash incentive and the venue change filtered the rest. Worth saying out loud, because "AI reports got better" isn't quite it - "no bounty plus a real triage platform" is closer. Joke too: Sarah O'Connor wrote a related piece (is this just before skynet launches?) Calvin #3: Where Did All Your AI Tokens Go? AgentsView to the rescue! Local-first desktop/web app for browsing, searching, and analyzing your past AI coding agent sessions (Claude Code, Codex, Copilot, Cursor, Gemini, Aider, and dozens more) Auto-discovers session files on your machine — no config needed; everything stored locally in SQLite, no cloud/accounts agentsview usage is a drop-in ccusage alternative — reads from pre-indexed SQLite, reports run 80–220× faster on large histories New Activity dashboard shows peak concurrency, active vs. idle time, agent-minutes, and cost — filterable by project/agent/machine, with a -json CLI report too Full-text + optional semantic search across every session; also imports Claude.ai/ChatGPT chat exports Install via pip install agentsview, uvx agentsview, brew install --cask agentsview, or download desktop binaries from GitHub Releases Michael #4: Careful with phishing all The situation I pass this along because it was a pretty sneaky bit of targeted phishing, and happened to play off an old interaction in bandit's repo. As usual with phishing scams there are a bunch of tells that this isn't legitimate, but just enough plausibility that I could see falling for it in a weak moment. Relative nobodies like me haven't historically been worth the effort to hit with scams this specific. Agents change the game though :-/. Be careful out there folks! Original message From: "Patrick (Blacktrace)" [HTML_REMOVED] To: LISTENER EMAIL Subject: Your Bandit #1350 (B105 NextToken false positive) -- just fixed that exact case Date: Wednesday, July 15, 2026 12:02 AM Hi AJ, Saw your Bandit issue #1350 -- the B105 hardcoded-password false positive on the string NextToken. I build a deterministic gate that filters that class of Bandit noise, and #1350 was literally the case I just fixed: NextToken / next_token / page_token / nextPageToken now stay quiet, while a genuine hardcoded token like api_token="sk-live-..." still fires. Verified against your exact case. 30-second paste: https://blacktrace.co/noise-eraser Where it still trips, published: https://blacktrace.co/kruc Curious whether it clears what you hit -- and if it trips on something of yours, that's the more useful reply. Patrick, Blacktrace I asked Claude for some analysis too. It was pretty good at finding them. The message name-drops enough real detail to feel legit, but the structure is pure phishing - everything in it exists to get AJ onto blacktrace.co. The strongest ones: Freemail sender, corporate signoff. Signs as "Patrick, Blacktrace" but sends from emailpjv@gmail.com. Real company outreach comes from the company domain, not a personal Gmail - and there's no last name. Over-specific targeting. It mirrors AJ's exact public activity - issue #1350, the B105 rule, the NextToken false positive, even the token variants. That's the "just enough plausibility" AJ flagged, and it's exactly what agents make cheap: scrape a GitHub issue, auto-generate tailored bait. Legit cold outreach rarely reads your history back to you this precisely. The entire payload is two links. Strip the technical flattery and the message is just "paste here" plus "see results here." When the whole point of an email is the click, that's the tell. "30-second paste." Low-friction urgency, and "paste" most likely means paste your source into their tool - handing your code to a stranger's site. Exfiltration dressed as convenience. Brand-new, no-reputation domain. blacktrace.co has no track record, and the name is doing some ominous work. The /kruc slug is random noise, not how real product pages get named. Precise-sounding jargon that's actually vague. "Deterministic gate," "noise-eraser" - impressive, empty. Bolted onto correct real details (B105 is the Bandit hardcoded-password test, sk-live- is a Stripe live-key prefix) to borrow credibility. The disarming close. "if it trips on something of yours, that's the more useful reply" - engineered humility that flatters your expertise and baits a response. Makes engaging feel like you're doing them a favor, which drops your guard. Extras Calvin: DjangoCon US 2026 is rapidly approaching, August 24-28, Chicago Ruff v0.16.0 massively expands its default rule set Ruff now enables 413 rules by default, up from 59 https://astral.sh/blog/ruff-v0.16.0 Michael: Completely redesigned the home page. Try /insights in Claude Code (terminal) Joke: We're Safe
Dave Altavilla recaps AMD Inc.'s (AMD) Advancing AI 2026 conference and his biggest takeaways on the event. He argues AMD's Helios rack offers "all the pieces" to compete with Nvidia (NVDA) as CEO Lisa Su sets her sights on creating a full AI platform. With the company reporting earnings next week, Dave tells investors to watch CPU revenue as it may set a foundation for Helios and argues there's lots of opportunity for AMD to eat into Nvidia's market share. ======== Schwab Network ========Empowering every investor and trader, every market day.Subscribe to the Market Minute newsletter - https://schwabnetwork.com/subscribeDownload the iOS app - https://apps.apple.com/us/app/schwab-network/id1460719185Download the Amazon Fire Tv App - https://www.amazon.com/TD-Ameritrade-Network/dp/B07KRD76C7Watch on Sling -https://watch.sling.com/1/asset/191928615bd8d47686f94682aefaa007/watch Watch onVizio - https://www.vizio.com/en/watchfreeplus-exploreClassification: Schwab InternalWatch on DistroTV - https://www.distro.tv/live/schwab-network/ Follow us on X –https://twitter.com/schwabnetworkFollow us on Facebook – https://www.facebook.com/schwabnetwork Follow us onLinkedIn - https://www.linkedin.com/company/schwab-network
AWS Morning Brief for the week of July 26th, with Corey Quinn. Links:Amazon ECS now provides Action Logs for deployment and orchestration visibilityAmazon Managed Service for Prometheus supports 1.5B active metrics and 200K rules per workspaceAmazon SES introduces pricing plansAWS Network Load Balancer now supports Listener Rules for custom traffic routingAWS Organizations increases RCP quota to 2,000 per organizationAWS now supports automatic credit memo application preferencesAWS Secrets Manager now publishes secret update notifications to Amazon EventBridgeJVM memory, CPU, and classpath best practices for Java containers on AWSConnection pooling strategies in Amazon Aurora DSQLCustom OS installation now available on AWS DeepRacer devicesThermodynamic sampling of disordered materials with an analog Hamiltonian Rydberg simulatorUnlocking data residency use cases with Amazon S3 in AWS Local ZonesAWS Security Bulletins: New Tools, Ancient Failures
OpenAI files for a trillion-dollar IPO the same week a pre-release model breaches Hugging Face and 42 state attorneys general open a coordinated investigation. Patrick Moorhead and Daniel Newman also break down AMD's hyperscaler CPU numbers from Advancing AI, the Moonshot distillation accusations, a wave of coordinated AI governance moves in Washington, and a stacked earnings slate spanning TSMC, Alphabet, IBM, ServiceNow, and Intel. The handpicked topics for this week are: OpenAI compresses a trillion-dollar IPO filing, a 42-state AG investigation, and a Hugging Face breach into five days: OpenAI and Anthropic both filed S-1 paperwork the same week 42 state attorneys general opened a coordinated investigation into OpenAI's data handling and safety practices, and a pre-release model reportedly breached Hugging Face days later. Moorhead and Newman question the timing, noting a rogue-agent narrative surfacing in the same week as a trillion-dollar valuation push draws obvious scrutiny. (The Decode) AMD's Advancing AI event puts a number on its hyperscaler momentum: Lisa Su confirmed Helios ships at the end of Q3 with volume ramping in Q4, backed by two-gigawatt capacity commitments from Microsoft and Anthropic and a claimed 70 to 75 percent share of hyperscaler CPU deployments. Moorhead points to NVIDIA's multi-layer software stack as the harder barrier AMD still has to close. (The Decode) Washington accuses Moonshot of distilling Anthropic's models days after Xi Jinping's WAIC keynote: Xi launched a 29-country AI cooperation organization at the Shanghai World AI Conference, and US officials Kratsios and Bessent followed with claims that Moonshot's Kimi K3 model shows data overlap with Anthropic's Opus models. Newman points to NVIDIA hardware in the training runs as evidence the distillation question extends beyond software alone. (The Decode) Five layers of government moved on AI oversight in a single week: Congress drafted a breach-response framework in reaction to the Hugging Face incident, the White House's 30-day pre-release review framework nears finalization, and state attorneys general and statehouses continue advancing their own rules in parallel. Moorhead notes nearly two decades of prior Capitol Hill engagement compressed into a single week of coordinated action, with each branch pursuing a different definition of the problem. (The Decode) Chinese open-source models now account for roughly a third of US developer traffic, and Moorhead and Newman take opposite sides on what it means: Moorhead argues enterprises are de-risking away from frontier-lab dependency, pointing to demand for smaller, workflow-specific open models. Newman counters with Vercel data showing those models capture 29 percent of gateway tokens against just 4 percent of revenue, framing the shift as a price discount that enterprise dollars have yet to follow. (The Flip) TSMC sells out CoWoS packaging capacity through 2026 and confirms a 10 percent price increase for 2027: The company posted a record quarter and committed $100 billion to its Arizona expansion on top of the pricing move. Newman calls the sustained capital spending a signal that the broader AI buildout still has runway. (Bulls & Bears) Google Cloud grows 82 percent as Alphabet posts its first-ever negative free cash flow quarter: The company raised its capital expenditure guidance to $195 to $205 billion and beat on revenue and EPS once one-time gains from its SpaceX and Anthropic stakes are excluded. Newman frames the spending as evidence Alphabet is prioritizing long-term AI infrastructure position over near-term cash generation. (Bulls & Bears) IBM misses Q2 revenue at $17.16 billion and cuts its full-year growth guide to 4 to 5 percent: Mainframe revenue fell 42 percent as enterprises redirected budget toward GPU and AI infrastructure purchases. CEO Arvind Krishna says a portion of the delayed deal flow has already resumed into the current quarter, pointing toward a potential rebound. (Bulls & Bears) ServiceNow crosses $1 billion in agentic AI annual contract value and raises its full-year guide: Agentic AI usage in production climbed 9x in nine months, and the company reaffirmed a target of $1.5 billion in AI ACV by year end. Newman points to margin compression from recent acquisitions as the tradeoff behind the platform's push into workflow and security convergence. (Bulls & Bears) NetSuite's new agentic platform, Next, becomes part of Six Five's own back-office stack: Newman and Moorhead both confirmed their companies are testing NetSuite Next for finance and accounting workflows. Newman points to the rollout as evidence that established SaaS platforms are absorbing agentic features directly into existing systems. Intel posts its fastest revenue growth since 2011 and lifts 2026 capital spending guidance: EPS came in near double consensus estimates, and CFO David Zinsner signaled a significant capex increase for 2027 tied to 14A demand. Moorhead reads the spending signal as confirmation of an anchor customer for the 14A node, ahead of any formal announcement. (Bulls & Bears) Watch the full video at sixfivemedia.com, and be sure to subscribe to our YouTube channel so you never miss an episode. OpenAI's High-Stakes Week: https://www.npr.org/2026/07/23/g-s1-135085/openai-hacking-ai-models AMD's Advancing AI Push: https://blogs.microsoft.com/blog/2026/07/20/microsoft-expands-azure-ai-and-hpc-infrastructure-with-amd/ Moonshot and the Distillation Debate: https://x.com/mkratsios47/status/2079933645888880708 Five Layers of AI Governance: https://www.politico.com/news/2026/07/22/openai-hugging-face-congress-response-01009190 Chinese Open-Source Traffic Debate: https://www.cnbc.com/2026/07/07/chinese-ai-models-costs-us-openai-anthropic.html ; https://wyomingdatacenterfacts.com/2026/07/20/the-quiet-surge-how-chinese-open-weight-models-are-powering-u-s-ai/ TSMC's Pricing Power: https://www.bloomberg.com/news/articles/2026-07-21/tsmc-in-talks-to-raise-prices-by-up-to-10-in-2027-nikkei-says Alphabet's Cash Flow Turn: https://qz.com/alphabet-google-second-quarter-earnings-revenue-cloud-072226 IBM's Mainframe Miss: https://www.cnbc.com/2026/07/22/ibm-q2-earnings-report-2026.html ServiceNow and NetSuite's Agentic Push: https://newsroom.servicenow.com/pressreleases/details/2026/ServiceNow-Reports-Second-Quarter-2026-Financial-Results/default.aspx ; https://www.netsuite.com/portal/home.shtml Intel's Growth Signal: https://www.cnbc.com/2026/07/23/intel-intc-earnings-report-q2-2026.html
R.I.P. John C. Dvorak.In other news ... the Ryzen 7700X3D prices drop already, Nvidia N1X CUDA counts and an 88-core CPU (of course). Windows on ARM gets a real boost with Nvidia GPU drivers, HP is just as bad as we all through regarding printer cartridges, and LG discovers a new low by auto-installing bloatware when you plug in a display. Plenty of AMD good news in the datacenter and Open AI shows that unshackling their creations will probably be as bad as we thought. Steam hardware sales and so much more in the show! Enjoy.Timestamps:00:00 Intro01:09 Patreon02:21 Food with Josh06:08 RIP John C. Dvorak08:08 NVIDIA's 88-core CPU09:26 RTX Spark Win 11 driver and Arm dGPU support12:07 Josh leads us into more NVIDIA talk and later defends datacenters19:16 AMD gets two big datacenter wins22:43 Ryzen 7 7700X3D already had a price drop24:11 ADATA chairman warns RAM shortage will last another decade24:37 Samsung Unpacked29:21 HP fined for nefarious printer cartridge practices - in India34:44 LG adware38:18 (In)Security Corner48:46 Gaming Quick Hits58:44 Picks of the Week1:12:01 Outro ★ Support this podcast on Patreon ★
AI is no longer just a race to train smarter models. As AI moves into production, the bottleneck is increasingly inference: how fast models can generate tokens, use tools, reason, verify, and act. In this episode of the MAD Podcast, Matt Turck sits down with Andrew Feldman, co-founder and CEO of Cerebras, to explain why fast inference may define the next era of AI.Cerebras is known for building a chip the size of a silicon wafer. But this conversation is not just about one company or one chip. It is a deep dive into the AI infrastructure stack: GPUs, ASICs, memory, HBM, SRAM, data centers, power, TSMC, AWS, OpenAI, agents, reasoning models, and why speed changes what AI products can become. Andrew explains why “tokens per second per user” matters, why generating a single word can require moving the equivalent of 100 HD movies through memory, why agents amplify latency, why GPUs struggle with certain inference workloads, and why fast AI may eventually reshape SaaS itself.This is a reference conversation on fast inference, AI chips, and the next compute bottleneck.(00:00) Cold open & Intro(01:31) Why speed became the AI bottleneck(02:32) Tokens per second per user, explained(03:16) AI's broadband moment and the Netflix analogy(04:35) The AI chip landscape: GPUs, TPUs, Trainium, ASICs(06:36) What is an ASIC?(08:08) Nvidia, Groq, and the fast inference war(09:16) OpenAI, Broadcom, and specialized silicon(12:10) China, power, and sovereign AI infrastructure(15:05) Is the AI infrastructure boom a bubble?(18:56) The hidden bottlenecks: HBM, CoWoS, and 3nm(22:57) Why agents are creating CPU demand(25:36) Andrew Feldman's path from SeaMicro to Cerebras(26:13) Why Cerebras bet on AI in 2016(31:14) SRAM vs. HBM: why inference is a memory problem(33:19) What wafer-scale computing actually means(34:28) The deep-tech “Everest” problem(36:07) The moment the first Cerebras system worked(36:49) Ringing the bell and surviving deep tech(39:08) How a giant chip handles failure(41:22) Why GPUs struggle with decode(42:17) Prefill vs. decode explained(44:01) The “100 HD movies” problem in AI inference(45:04) How fast inference changes RL and training(48:08) Reasoning models and why they cost more compute(50:08) Verification, guardrails, and small models checking big models(52:37) Multimodal AI and the path to video(53:51) Cerebras' business model: hardware, cloud, and API(55:14) OpenAI's 750MW inference deal(55:36) Why data centers are measured in megawatts(58:01) AWS Trainium + Cerebras decode(59:29) Fast tokens as a cloud product(01:00:52) Is CUDA still a moat?(01:03:53) How TSMC helped Cerebras build the giant chip(01:07:41) Why nobody cared in 2020(01:08:15) Why chip supply chains are hard to diversify(01:09:54) Why today's AI models will be the worst you ever use(01:10:38) What fast AI could do to SaaS
Een Amerikaanse tiener heeft vlak voor de rechtszaak zijn verslavingsclaim tegen Meta ingetrokken, waardoor het bedrijf een proces in Los Angeles ontloopt. Ook sluiten de Amerikaanse chipmakers Intel en AMD langlopende overeenkomsten met Chinese klanten voor datacenterprocessoren, terwijl de AI-strijd tussen de VS en China in volle gang is. Joe van Burik vertelt erover in deze Tech Update. De zaak werd aangespannen door een 15-jarige jongen uit Florida, bekend onder de initialen R.K.C., die stelde depressief en angstig te zijn geworden door sociale media. Hij begon naar eigen zeggen op zijn achtste met de platforms, raakte verslaafd en sliep slecht. Naast Meta had hij oorspronkelijk ook YouTube-eigenaar Google, Snapchat-moederbedrijf Snap en TikTok-eigenaar ByteDance aangeklaagd. Die drie schikten de afgelopen weken tegen onbekende bedragen, waardoor Meta als enige overbleef voor het proces, dat komende week zou beginnen. Volgens Meta trok de jongen zijn claim in zonder dat het bedrijf iets betaalde. De advocaten van R.K.C. noemden de uitkomst een succes voor het aansprakelijk stellen van sociale media, en zeiden dat hij zich wil richten op zijn herstel en therapie. Ruim vijfduizend zaken lopen nog Meta reageerde dat de claims nooit standhielden en dat het zich blijft verdedigen tegen wat het grondeloze rechtszaken noemt. De ingetrokken zaak was de tweede in een reeks van meer dan drieduizend soortgelijke zaken in Californië, met daarnaast nog eens ruim tweeduizend zaken in de federale rechtbank van families, schooldistricten en procureurs-generaal. In de eerste zaak, in maart, oordeelde een jury dat Meta en YouTube aansprakelijk waren en werden zij veroordeeld tot een schadevergoeding van in totaal zes miljoen dollar. In een aparte zaak in maart, aangespannen door de procureur-generaal van New Mexico, kreeg Meta een boete van 375 miljoen dollar. Intel en AMD leggen Chinese afname vast nu prijzen stijgen Intel en AMD sluiten langlopende afspraken met Chinese serverklanten voor datacenterprocessoren, meldt Reuters op basis van twee bronnen. Het gaat niet om GPU's, de AI-chips waarmee vooral Nvidia groot werd, maar om CPU's, de centrale processoren die servers aansturen en die eveneens hard nodig zijn voor de ontwikkeling en toepassing van AI. De overeenkomsten leggen de afnamevolumes vast, maar niet de prijzen. De meeste dekken ongeveer een jaar, al is voor sommige klanten gesproken over twee jaar of langer. De serverprocessoren zijn schaarser geworden: sommige prijzen in China stegen maandelijks meer dan tien procent, en sinds begin dit jaar liepen bepaalde producten met meer dan veertig procent op. Intel presenteert donderdag zijn kwartaalcijfers, waarbij het chiptekort waarschijnlijk aan bod komt. De Amerikaanse overheid bezit inmiddels tien procent van de aandelen in Intel. Florida-tiener trekt verslavingszaak tegen Meta in vlak voor proces Tiener laat claims tegen Meta vallen dagen voor rechtszaak Meta en YouTube eerder aansprakelijk gesteld in eerste verslavingszaak Intel en AMD sluiten langlopende server-CPU-deals met Chinese klanten Chiptekort drijft prijzen op en duwt Chinese klanten naar langere contracten Over de maker:Joe van Burik volgt en duidt de belangrijkste ontwikkelingen in tech, met scherpte, vlotheid en de nodige humor. Je hoort hem dagelijks op BNR Nieuwsradio over het belangrijkste technieuws, van AI tot cybersecurity en social media tot quantumcomputers. Ook interviewt hij in De Grote Tech Show samen met Ben van der Burg leiders in digitale innovatie. In het bijzonder volgt Joe al twee decennia de wereld van videogames, nu voor zijn podcast All in the Game.See omnystudio.com/listener for privacy information.
professorjrod@gmail.comPassing the CompTIA A+ exam becomes less stressful when you stop hunting for 'the right answer' and start asking: what problem is happening right now? This episode kicks off the Technology Tap masterclass for the CompTIA A+ 220-1201 exam, teaching the mindset you'll use on a real help desk response, such as when a user says, 'My computer won't turn on.' It's your practical guide for effective troubleshooting and exam success.I walk through the foundations that show up everywhere on the exam and on the job: the four basic functions of a computer (input, processing, storage, output), the role of the motherboard, CPU, RAM, storage, and power delivery, and the simple troubleshooting steps that prevent expensive guesswork. You'll also get quick practice questions to lock in the basics, plus a study strategy that goes beyond reading once so you can actually explain the concepts back. Then we go deeper into the exam magnets: motherboard form factors and compatibility, CPU cores vs threads, sockets, chipsets, cache, and why DDR4 and DDR5 can't be mixed. We also cover firmware and startup (BIOS vs UEFI, POST, CMOS battery issues) and the storage maze: HDD vs SSD vs NVMe, SATA vs PCIe, why M.2 doesn't automatically mean NVMe, plus GPT vs MBR, NTFS vs FAT32 vs exFAT, RAID levels, and the critical truth that RAID is not backup. If this helps, subscribe, share it with someone studying for CompTIA A+, and leave a review so more future techs can find the masterclass.Support the showArt By Sarah/DesmondMusic by Joakim KarudLittle chacha ProductionsJuan Rodriguez can be reached atTikTok @ProfessorJrodProfessorJRod@gmail.com@Prof_JRodInstagram ProfessorJRod
I had a request from a customer recently who asked if we could give them a report of their database server instances and include CPU usage. This request was filtered through an account executive, so something was lost in translation, but I was confused and asked for clarification, as asking for CPU usage is kind of like asking how fast you were traveling in your car. There needs to be more context. If someone asked you for a report of CPU usage for a database, what would you expect? How would you report this? I'm sure the person asking might make a difference. A fellow DBA, your DBA manager, or maybe an executive could all view this differently. I want to know how things are performing, if there is a trend, or maybe if we are getting value for the hardware we've provisioned, depending on my role. Read the rest of What is CPU Usage?
掌握前瞻趨勢與科技脈動,立即訂閱 IC之音電子報:https://pse.is/8wpwwx ------------------------------過去幾年,AI幾乎等於GPU。但今年COMPUTEX,黃仁勳不只發表新一代GPU,更大力介紹資料中心CPU Vera,市場也開始重新討論CPU的重要性。隨著Agentic AI興起,AI不再只是回答問題,而是能自主規劃、推理與執行任務,CPU與GPU的分工也正悄悄改變。CPU真的迎來第二春了嗎?這是否意味著AI競爭已從單一晶片,走向整體運算平台的新時代?本集《科技領航家》,邀請DIGITIMES分析師陳辰妃,帶大家解析CPU重新崛起背後的產業趨勢,以及AI運算的下一場競賽。------------------------------製作 | 李翊嘉
Join Scott as he hacks on the CircuitPython powered gameboy cart. He'll also answer any questions folks have. Visit the Adafruit shop online - http://www.adafruit.com Thanks to dcd for timecodes: 0:00 Getting started 1:57 hello and welcome to Deep Dive 2:40 Circuit Python (CP) from Adafruit makes it easy to program microcomputers 3:16 This week - working on GameBoy (GP) custom cart to use CP 5:08 GB cart features 7:27 Power up GB showing Nintendo 9:43 switch to Konsole and CP REPL - showing "Adafruit" 10:34 the "Adafruit" logo trick explained 11:47 continue Gameboy cart hacking 12:33 rp2350 gameboy cartridge firmware - croco cartridge V2 (github ) 13:43 emulate RAM without CPU (using PIO) https://github.com/shilga/rp2350-gameboy-cartridge-firmware 15:57 game boy memory access sequence ( https://iceboy.a-singer.de/doc/mem_patterns.html ) 18:30 CP _gbio PIO code ( https://github.com/tannewt/circuitpython/blob/pygb2350/ports/raspberrypi/common-hal/_gbio/__init__.c ) 21:30 rp2040 and rp2350 PIO datasheets and PIO overview 26:30 CRC 'trick' to transform address from GB to RAM ( summation ) 37:53 check out mermain.live to visualize the generated block diagram using mermaid-cli 44:00 back to the PIO code 47:46 LLM discussion - DeepSeek 4 / DwarfStar 51:17 CP with latest PICO_SDK ? 55:10 USB vs. Serial Logging 56:20 PIO debug trace shows all the bits... 1:02:10 RP2 I/O input/output driving discussion 1:05:10 CP docs on circuitpython.org 1:06:30 gbdev.io and vsync 1:07:04 plan on next week with scott, and visit the discord ----------------------------------------- LIVE CHAT IS HERE! http://adafru.it/discord Subscribe to Adafruit on YouTube: http://adafru.it/subscribe New tutorials on the Adafruit Learning System: http://learn.adafruit.com/ -----------------------------------------
AMD's first reinvention rebuilt the company. It was frankly about survival. Its next reinvention must redefine it.The company's resurgence over the past decade came from doing what many thought was impossible: rebuilding its CPU franchise, taking meaningful share from Intel, and restoring credibility through disciplined execution.But in our view, AMD's next chapter is fundamentally different.
Episode 107: We chat about some of the latest CPU releases, including the Ryzen 7 5800X3D 10th Anniversary Edition, and the Ryzen 7 7700X3D, which doesn't make a ton of sense when nobody is doing platform upgrades. Also, we thought it was quite amusing that Nvidia's hotspot GPU temperature has now been uncovered, so we discuss that whole situation.CHAPTERS00:00 - Intro01:02 - The Ryzen 7 7700X3D has Launched08:32 - The Ryzen 7 5800X3D Actually Makes Sense24:53 - Nvidia GPU Hotspot Temperature Situation40:24 - Steam Machine Thermals and HDMI-CEC01:04:55 - Updates From Our Boring LivesSUBSCRIBE TO THE PODCASTAudio: https://shows.acast.com/the-hardware-unboxed-podcastVideo: https://www.youtube.com/channel/UCqT8Vb3jweH6_tj2SarErfwSUPPORT US DIRECTLYPatreon: https://www.patreon.com/hardwareunboxedLINKSYouTube: https://www.youtube.com/@Hardwareunboxed/Twitter: https://twitter.com/HardwareUnboxedBluesky: https://bsky.app/profile/hardwareunboxed.bsky.social Hosted on Acast. See acast.com/privacy for more information.
Topics covered in this episode: The trusted-publishing debate: how to do it right vs. why you shouldn't trust it JupyterLab 4.6 and Notebook 7.6 are out! Tau – new small, readable terminal coding agent Django Tasks and Django 6.1 Extras Joke Watch on YouTube About the show Sponsored by us! Support our work through: Our courses at Talk Python Consulting from Six Feet Up Connect with the hosts Michael: Mastodon / BlueSky / X / LinkedIn Calvin: Mastodon / BlueSky / X / LinkedIn Show: Mastodon / BlueSky / X Join us on YouTube at pythonbytes.fm/live to be part of the audience. Usually Tuesday at 7am PT. Older video versions available there too. Finally, if you want an artisanal, hand-crafted digest of every week of the show notes in email form? Add your name and email to our friends of the show list, we'll never share it. Calvin #1: The trusted-publishing debate: how to do it right vs. why you shouldn't trust it https://snarky.ca/how-to-publish-to-pypi-using-github-actions-securely/ (Brett Cannon) and https://blog.yossarian.net/2026/07/07/You-shouldnt-trust-trusted-publishing (William Woodruff) Trusted Publishing (PyPI's OIDC-based auth scheme, also now used by npm, RubyGems, crates.io, NuGet) replaces long-lived API tokens with short-lived, auto-scoped credentials tied to CI/CD machine identity. Yossarian's post: it's purely an authentication mechanism between a machine identity and a package — it says nothing about package safety or quality. PyPI deliberately avoids any "verified/trusted" badge for it, unlike its verified-URL checkmarks. Same logic applies to PyPI attestations: anyone can sign with any machine identity they control, so an attestation's presence isn't itself a trust signal. Bottom line from that post: don't confuse "trusted" (machine-to-machine) with "trustworthy" (human judgment about the package). Snarky.ca's companion piece is more practical: given GitHub Actions compromises in the news, the real fix is 3 concrete steps — run zizmor to lock down workflow permissions/checkout credentials and pin actions to commit hashes, adopt Trusted Publishing to eliminate stored PyPI tokens, and require manual approval via a GitHub environment before any publish job runs. Takeaway for listeners: Trusted Publishing is good hygiene for how you authenticate to PyPI, but it's not a substitute for securing your CI pipeline itself — or for actually vetting the packages you install. Michael #2: JupyterLab 4.6 and Notebook 7.6 are out! Michał Krassowski's rundown - a chunky minor release: 68 features, 97 bug fixes, 95 contributors, one of the biggest ever. Scratchpad console (Notebook 7.6 headliner) - a console next to your notebook sharing its kernel, for throwaway experiments. Ctrl+B. Jump to last-edited cell - new commands hop through recently edited cells. File browser glow-up - Date Created column, editable breadcrumbs with Tab-completion, and Open in Terminal. Debugger - sources open in the main area, floating step/continue overlay, live kernel-sources filter. Custom layouts (Lab) - activity bar top/bottom, draggable panels, four-way tab splits, per-panel Ctrl+scroll zoom. ~5x faster extension builds - webpack → Rspack, and jupyter-builder means no full Lab install needed to build extensions. Keyboard/a11y - add shortcuts from the UI (no JSON), Find & Replace in Edit menu (Ctrl+H). Calvin #3: Tau – new small, readable terminal coding agent Tau – new small, readable terminal coding agent (Python 3.12+), built as both a working tool and a teaching project for how coding agents work under the hood Install via uv tool install tau-ai, pipx, or pip; ships a tau CLI Three-layer architecture: tau_ai (provider-neutral model layer) → tau_agent (reusable "brain": messages, tools, events, loop) → tau_coding (CLI/TUI, file & shell tools, sessions) Supports OpenAI, Anthropic, OpenAI Codex, OpenRouter, Hugging Face, and custom/local OpenAI-compatible endpoints Built-in tools (read/write/edit/bash), durable JSONL sessions with resume/branching, project instructions via AGENTS.md, and context compaction Core harness is UI-agnostic — same brain can power the TUI, print mode, or a custom frontend — usable as a standalone library too Michael #4: Django Tasks and Django 6.1 Django 6.0 finally ships first-party background tasks (django.tasks) - out of Jake Howard's DEP 14, accepted May 2024, after two decades of everyone bolting on Celery/RQ/Huey. It's an API, not a worker. Django handles task definition, validation, queuing, and result storage - it does not execute them. You bring the backend. The default backend traps people. ImmediateBackend runs tasks inline on the request thread and blocks until done - so out of the box .enqueue() backgrounds nothing (a 5-second task means a 5-second response). The other built-in, DummyBackend, runs nothing at all. Both are dev/test only. Nice API otherwise: slap @task on a function, call .enqueue(), get back a TaskResult you look up later by id - with async twins like aenqueue(). Gotcha: args and return values must survive a JSON round-trip, so a tuple sneakily comes back as a list. The community local backend to know: django-tasks-local by Chris Beaven (SmileyChris). A ThreadPoolExecutor backend that gives real background threads with zero infrastructure - no Redis, no Celery, no database - plus a ProcessPoolBackend for CPU-bound work → github.com/lincolnloop/django-tasks-local Its catch: results live in memory, so pending tasks vanish on restart or deploy. Great for dev and low-traffic production; for persistence, drop to Jake Howard's django-tasks (DatabaseBackend + worker command). Extras Calvin: Fixing the dictionary with Python 3.14 — Hugo van Kemenade stumbled on - and got fixed - a markup bug in the OED's own citation of a 1706 use of the pi symbol. Michael: Bunny DNS is now free Jokes: What's the object-oriented way to become wealthy? Inheritance To understand what recursion is... You must first understand what recursion is 3 SQL statements walk into a NoSQL bar. Soon, they walk out They couldn't find a table.
Lords: Chall RT-55J https://samarantes.neocities.org/ https://metroidconstruction.com/hack.php?id=878 Topics: Remembering the Dynowarz Instagram private server, and social media thoughts every Mastodon user had already Why don't CPUs do analog arithmetic? Leaves, by Ursula LeGuin https://www.poetryfoundation.org/poems/148293/leaves-5bd9e153d78b2 Microtopics: DogTroid, the first Metroid ROM hack to star a dog. Self-finishing games. Sparkling water: it's like water but a lot more interesting. Caffeine Free Diet Pepsi. What sodas foam the most in response to a mento. Foam persistence. Foaminess reactions of a mento on various vintagesn of Diet Coke. Jolt Cola: all the sugar, twice the caffeine, three times the foam. A can of soda that's safe to open in a bathroom stall. Artisanal Coca Cola cans on Etsy that finally allow you to open a can of soda in a bathroom stall without anyone realizing you're drinking a Coke on the toilet. Flushable soda cans: as flushable as a flushable wipe. Why do toilets have pee traps when the pee deserves to be free? The second worst game in your NES collection. Desert Chrome title screens. Playing as a little spaceman until you enter the dinosaur mech. Shooting some alien brain or maybe a heart. 8bitnintendo.science How to pick what video games to buy in the late 1980s. How Metroid improved on the maze-with-keys genre. Nanosaur. A velociraptor with a techno-backpack. A game that is exhilarating and scary and endless when you're a child turning out to be a twenty minute trifle when you're an adult. Trespasser (1998) Simulation dinosaur emotions but you can't find a good balance so you just permanently lock them all to angry. Installing a violent action game about dinosaurs in the elementary school computer lab because dinosaurs are technically educational. Revisiting games that perplexed you as a child. Playing bad video games because no matter how bad they are they're still better than going outside and talking to people. The one where Kirby eats a car. The Roblox-like games you can find by logging into third party Minecraft servers. Starting your own Pixelfed server. Following the only person you know on Mastodon. Bluesky's recommendation algorithms recommendeding you nothing but bots. Inventing Internet forums from first principles. When your brain makes up garbage and you need a void to shove it into. Doing your part to make AI worse. CSS Crimes. How to find people to follow on Cohost. Signing up for a social media site and looking around and realizing you doing know anyone here. Mining and reposting. The ongoing maintenance requirements of running a Mastodon server. Ways you can interact with your family that only work if you have an iOS developer in the family. Off-box SQL database backup. Hypothetical IRC servers that support chat logs. Hardware random number generation. The pot of boiling water every Intel CPU draws thermal noise from for random number generation. The most commonly used analog computers in 2026. The market forces that led to semi-modular synthesizers being available for $300. Using your analog CPU to run a million instances of Lunar Lander at once. A rustic summer retreat ranch in the hills of Napa Valley, California. What makes us perceive the gradient of identities over the course of someone's life as a single identity. The most ephemeral thing possible. Musing for a few sentences and then thinking "hmm, I could put some line breaks in here and then publish it." Musing about the nature of identity, with line breaks. Topic Slingers. Topics: throw them around a lil bit. They love it. A lot of people don't know that.
En este episodio de Atareao con Linux nos vamos a remangar para hablar de una de esas tecnologías que, una vez las dominas, te cambian la vida por completo: el Web Scraping asistido por Inteligencia Artificial.Seguro que te ha pasado alguna vez. Quieres comprar un producto concreto, como unas zapatillas de running (yo las cambio cada 800 kilómetros y es un goteo constante), o quieres extraer todas las recetas de cocina de una web para montarte tu propio planificador semanal. Lo ideal sería que estas páginas tuvieran una API pública para descargar la información de forma limpia. Pero la cruda realidad es que casi ninguna te lo pone fácil. Ahí es donde entra el scraping: la técnica de extraer la información directamente de la página web.En este episodio te cuento por qué el scraping clásico (ese que utiliza Beautiful Soup en Python y depende de identificar las etiquetas HTML y las clases CSS) tiene los días contados para tareas complejas. Basta con que un desarrollador cambie el diseño de la web para que tu script se rompa por completo. Además, con la llegada de las webs dinámicas, los tests A/B y los sistemas anti-bloqueo como Cloudflare, mantener un scraper tradicional es un auténtico dolor de muelas.La gran alternativa: Inteligencia Artificial en local¿Y si en lugar de pelearnos con el código fuente dejamos que un modelo de lenguaje (LLM) entienda la página exactamente igual que lo haría un humano? Un LLM comprende perfectamente qué es un "precio" o el "nombre de un producto", sin importar cómo esté maquetada la web ni el idioma en el que esté escrita. Y lo mejor de todo: ¡lo podemos hacer 100% gratis en local usando Ollama!Te detallo mis pruebas ejecutando modelos en mi Slimbook One utilizando únicamente la CPU (¡sin gastar un céntimo en nubes ni necesitar tarjetas gráficas carísimas!). Hablaremos de cómo rinden modelos como Llama 3.2, Qwen, Mistral y DeepSeek R1, y cuál es el punto de equilibrio perfecto para no eternizarnos esperando la respuesta.También te desvelo mi fórmula secreta para procesar la información. No podemos enviarle 2 Megabytes de HTML ruidoso a la IA. Te explico los 5 pasos que utilizo en Python para eliminar la basura (scripts, estilos, navegación) y reducir el HTML hasta en un 93%, permitiendo que el modelo extraiga los datos en segundos y nos devuelva un JSON estructurado impecable.Por último, vemos cómo montar un auténtico vigilante de ofertas automatizado en segundo plano. Un sistema que compare los precios de varias tiendas en paralelo.Capítulos del episodio:00:00:00 Introducción al Web Scraping con Inteligencia Artificial00:01:22 ¿Para qué sirve extraer datos? Ejemplos prácticos00:02:42 El gran talón de Aquiles del scraping tradicional00:04:31 La revolución de la IA: Entender la web sin saber HTML00:07:36 Los problemas habituales: Selectores rotos y webs dinámicas00:10:00 Cómo un modelo de lenguaje (LLM) procesa la información00:13:17 Cuándo elegir scraping clásico vs. scraping con IA00:15:28 Comparación de costes: Enfoque clásico, IA local e IA en la nube00:17:19 ¿Qué modelos usar? Pruebas con Llama, Qwen, Mistral y DeepSeek00:18:19 Detrás de escena: Mi script de Python y la limpieza del HTML00:21:05 Creando el prompt perfecto para extraer un JSON estructurado00:24:34 Ejemplo real: Comparativa paralela entre tiendas00:28:38 Diseñando un vigilante de ofertas automatizado (24/7)00:30:17 Casos de uso prácticos y mejoras para evitar bloqueos00:32:02 Cierre y detalles del próximo tutorial de scrapingMás información y enlaces en las notas del episodio
Charles is joined by Spire Investments Founder and CIO Ivana Delevska to discuss how algorithmic selling creates buying opportunities in fundamentally strong tech names, why a massive expansion in CPU and memory capacity positions semiconductor equipment and optical sectors for long-term growth, and why resilient business models make software and cybersecurity attractive plays. Learn more about your ad choices. Visit podcastchoices.com/adchoices
We've been running a bit of an Agent Cloud series surveying all the top inference/compute/cloud providers, from Databricks to Daytona to Railway and, even further back, E2B, but we're excited to conclude this series returning to Modal, which has just raised a monster $355M Series C.The cloud was built for developers. But agents are now changing that.The old infra stack was designed for a human who could read docs, reason through YAML, and understand dashboards to figure out what they need when something broke. While this was painful for developers, it worked since they could fill in missing context in their heads.However, agents don't have that luxury. Now in this new era of agents, everything has to be tighter.They need a place to write code, run it, inspect the output, change the environment, debug failures, and try again. Fast iteration and feedback loops with all the necessary context are crucial for agents to operate properly. Furthermore, sandboxes are a clear representation of this shift as agents can easily spin up isolated environments. This programmatic infra even extends to research:Two years ago, we were one of the first to cover Modal with CEO Erik Bernhardsson and Alessio designed our favorite LS thumbnail of all time:At the time, Modal was just a teeny little company with a $17M Series A.Today, fresh off their $355M Series C, Modal is one of the clearest examples of the agent cloud future being built in real time: a cloud platform moving past traditional web app assumptions toward the workloads AI actually creates such as elastic inference, sandboxes, GPU burst, post-training, background agents, and infrastructure that agents themselves can operate.In this episode, Modal CTO Akshat Bubna joins swyx and Vibhu to unpack why AI applications don't fit traditional cloud assumptions, why Kubernetes was never designed for bursty compute-heavy workloads, and why Modal is now shifting from developer experience to agent experience.We go deep on Modal's AI infra stack: serverless functions, decorator-based infrastructure, elastic inference for custom models, GPU snapshotting, DeFlash, speculative decoding, Auto Endpoints, sandboxes, persistent storage, networked containers, private IPv6, RDMA, multi-node training, and Modal's capacity pool across 17 cloud providers. Akshat also explains why RL rollouts can require 100,000 sandboxes, why production agents need hard guardrails, why observability may matter more than reading code, and why AI has made infrastructure exciting again.We discuss:* Why Kubernetes wasn't built for bursty AI workloads* How Modal started as a better runtime before becoming an AI cloud* Why Modal added GPUs before ChatGPT* The shift from developer experience to agent experience* Why observability matters when agents are writing the code* Elastic inference for custom models across audio, video, robotics, and comp bio* GPU snapshotting, cold starts, and why inference workloads are so bursty* Why RL rollouts can require 100,000 sandboxes* DeFlash, speculative decoding, and frontier-level inference performance* Auto Endpoints and making optimized inference easier to deploy* What Modal adds beyond vLLM, SGLang, and raw GPU rental* Modal's 17-cloud capacity pool and supercloud strategy* Networked sandboxes, sidecars, private IPv6, and RDMA* Serverless multi-node training for post-training and research workloads* Auto-research, model-guided sweeps, and agents launching GPU experiments* Compute strategy, capacity planning, and batch tiers* Why production agents need specialized sandboxes and hard guardrails* Modal's take on managed agents, CI, Gitpod/Ona, Python, TypeScript, and Modal BenchAkshat Bubna* LinkedIn: https://www.linkedin.com/in/akshat-bubna-188885103* X: https://x.com/akshat_bModal* Website: https://modal.comTimestamps00:00:00 Introduction00:00:39 Modal's origin and why Kubernetes wasn't enough00:04:32 Developer Experience → Agent Experience00:06:21 Modal's AI cloud primitives00:09:14 Sandboxes, agent loops, and proto-Cognition00:12:12 Elastic inference, GPU snapshotting, and 100,000 sandboxes00:15:24 DeFlash, speculative decoding, and Auto Endpoints00:19:59 Production-grade inference beyond raw GPUs00:22:00 Background agents, Ramp Inspect, and the agent lifecycle00:24:08 Modal's 17-cloud supercloud strategy00:26:40 Networked sandboxes, private IPv6, and RDMA00:32:48 Multi-node training, post-training, and auto research00:37:36 Compute strategy, capacity planning, and batch tiers00:40:55 Open models, real-time AI, and production agent infra00:43:06 Hard guardrails, managed agents, and specialized sandboxes00:46:06 Why AI made infrastructure exciting again00:48:30 Model APIs, differentiated products, and agentic video00:51:50 CI, coding-agent infra, SDKs, and Modal Bench00:57:28 Closing ThoughtsTranscriptIntroduction: Modal, Series C, and the Art PartySwyx [00:00:00]: We're here with Akshat, CTO of Modal, together with Vibhu. Congrats on your Series C.Akshat [00:00:10]: Thank you.Swyx [00:00:11]: Your party yesterday was amazing.Akshat [00:00:15]: Yeah.Swyx [00:00:15]: From all the photos and all the swag.Akshat [00:00:17]: We had a bunch of art installations, which was fun, seeing, like, our products on pedestals next to, like, Rodin.Swyx [00:00:25]: Very nice. Very nice. When you started, it was not the GPU inference company. Maybe it was in your mind. Take us back to the origin story.Modal's Origin: A New Runtime Beyond KubernetesAkshat [00:00:39]: I first met Eric, who's the CEO, through an investor. Back then Eric was already thinking about building, a new runtime, and he got there thinking through why are workflow orchestration products so hard to use. It's because you have to run them on Kubernetes. Kubernetes is hard to manage. It's not built for burstiness and, custom images,Swyx [00:01:03]: YeahAkshat [00:01:03]: It has a terrible developer experience.Swyx [00:01:05]: And I'll, I'll interjectAkshat [00:01:06]: YeahSwyx [00:01:07]: For listeners, who are new, we interviewed Eric two years ago, and there's a bit more of the story there from Spotify and all those things.Swyx [00:01:14]: And I came across Eric through Data Council because he did that talk on the serverless container stack that you guys did, which was like, that was my first like, “Okay, I need to take Modal very seriously” moment.Akshat [00:01:26]: Yeah.Swyx [00:01:26]: But it was still very unclear, like, do I need all this for just my data pipelines?Akshat [00:01:33]: Yeah. initially what we were thinking about was if we build a better runtime, it's a very useful primitive in itself. It's There's a lot of things that, get solved by serverless functions, like you can do, ETL stuff, you can do job queues, you can do all this, like, bursty processing, which it turns out every company had needs for. but then we also were thinking about this as like, this is a primitive that we can build a whole collection of products on, which are very verticalized. So perhaps data engineering would've been the first one, but we were thinking about inference. Back then it was more classical inference, like computer vision stuff and running XGBoosts and whatnot. But we added GPUs to the product a year before ChatGPT came out.From Serverless Containers to GPU WorkloadsSwyx [00:02:19]: Nice.Akshat [00:02:19]: We just didn't think it would be that big of a deal.Swyx [00:02:22]: Yeah, just like add A100.Vibhu [00:02:23]: Was there any, like, early key problem that really sparked off why you built it?Akshat [00:02:28]: Yeah. Primarily it's just, none of the tooling that was out there was built for, one, a really great developer experience, and also there's a general trend of, a lot of the workloads that we were seeing were very. I wish there was a better word for it, but compute-heavy. Like, they need, one, like, need a lot more resources, so you need to burst up and down a lot, versus like Kubernetes designed for, like, slow scaling and, more for, like, web server use cases. And also there's just a lot more specialization in, like, what kinds of environments these workloads run in. Like, we had sometimes they need accelerators, sometimes they need different kinds of images, and this is just like a consistent thing that we saw across a lot of companies. That would be the next step.Software-Defined Infrastructure and Decorator-Based DXSwyx [00:03:13]: Yeah. Yeah. Be nice. I don't know how much this factored into the early story, but I wrote a post when I was at Temporal about infrastructure, software-defined infrastructure or something like that.Akshat [00:03:22]: Yeah, the self-provisioningSwyx [00:03:23]: Self-provisioning.Akshat [00:03:24]: Yeah.Swyx [00:03:24]: Yeah. I can't even remember my own post.Swyx [00:03:26]: And then you put me on the landing page.Akshat [00:03:28]: Yeah. We really like, the term and so we stole it.Swyx [00:03:32]: Because you had the insight that everything can just be in decorators co-located with the code, right?Akshat [00:03:37]: Yeah.Swyx [00:03:37]: Was that a big part of the originalAkshat [00:03:39]: YesSwyx [00:03:39]: Story or it was just like a DX layer?Akshat [00:03:41]: That was, really important because we really didn't want people to spend, so much time, writing YAML, and it seemed like you could really condense the surface area of what you're doing, put it in code so you can operate on it just like you operate on other code, and like build stuff that's more expressive and dynamic. and so yeah, that was always a very important part.Swyx [00:04:04]: Then the pushback is this is a DSL.Akshat [00:04:07]: Yeah.Swyx [00:04:07]: It's you're closed source. I am locked into Modal.Akshat [00:04:11]: Yeah. We never really got pushback for that because the nice thing about Modal is you can bring whatever code you have, and sure, the DSL is at the configuration layer for, what hardware you're using, how you're scaling things up, but you still own the code.Akshat [00:04:27]: And that's, that's been an important, part of our story, even as we do inference now.Swyx [00:04:32]: Yeah.Vibhu [00:04:32]: How much of do you think still stays the same today? Like if you were to build something today, DevX very important, but I feel like, a lot of this has been changed with just hook it up to an agent, have Claude Code, have Codex implement a tool. there's very agent native primitives that are different than if I'm doing this myself, right?Developer Experience → Agent ExperienceAkshat [00:04:54]: We've changed our SDK team to think about agent experience instead of, developer experience and we think that the same benefits that apply for DX also apply for AX, which is why would you have an agent read through hundreds of Kubernetes files and like write YAML that's not even typed when it can make a couple of changes in a decorator and it gets this self-provisioning runtime of, being able to see its changes live in action? yeah, it just seems from the customers we talk to, they find Modal is much faster for agents to use versus operating on a different substrate.Swyx [00:05:34]: Yeah, because like you, again, you co-locate the infrastructure requirements to the code that runs it.Akshat [00:05:38]: Yeah.Swyx [00:05:38]: Well, the negative thesis now is that nobody's looking at their code anymore, so there's no point.Akshat [00:05:44]: Yeah, people aren't looking at code. one thing we still see is really important is observability.Swyx [00:05:51]: Yeah.Akshat [00:05:51]: Like how good is your dashboard? And of course, like we have, we push a lot of it to the CLI so the agents can do their own investigation, but you still need humans to go interpret what's going on and, make judgment calls and whatnot. and that's I feel like, Maybe more important now than looking at the code itself.Swyx [00:06:11]: Yes, because like, you can try to treat the code as a black box and then use, see the observable action that comes out of it, and then just prompt a change.What Modal Is For: AI Cloud PrimitivesAkshat [00:06:21]: Yeah.Swyx [00:06:22]: So I think it takes a bit of restraint to not specialize, to say, “I want to ship a new primitive,” and then just be general purpose.Swyx [00:06:31]: People ask you, “What are you for?” You're like, “ I don't know. We can do this, we can do that.”Vibhu [00:06:36]: Well, I'd be curious to see, like, okay, if we were to ask you, like, what is Modal for even at a high level? There's a lot you guys do, sandboxes, GPUs, everything. How do you answer?Akshat [00:06:46]: Modal is a cloud platform that's built for, where we've built the primitives from scratch for AI applications. and right now it covers, inference, training, batch processing, and sandbox workloads.Akshat [00:07:00]: But we're building a lot moreSwyx [00:07:02]: I noticed you didn't say web server, so there is still a role for, like, the always-on large-scale Kubernetes type things.Akshat [00:07:09]: Yeah, absolutely. We're, we're not trying to compete with the renders of the world, because yeah, we think the differentiator for us is the, are the workloads that need specialized compute, need to scale up and down a lot. yeah, they're, they're, they're just shaped differently.Working Alongside Frontier StartupsVibhu [00:07:26]: I think you're building a lot of it alongside the startups, right? They're innovating quite a bit, even in your, like, latest blog post. Like, even in the series C, the customers that you mention here, the cognitions, technical ones, ramps and whatnot, they're, they're innovating with you, right? And that's not something AWS is doing directly with.Akshat [00:07:45]: Yeah, absolutely. I think, this is again classic. We're a small team. We can move really fast. our engineers are working with our customers and figuring it out. Yeah.Swyx [00:07:54]: So my first week at Cognition, I walked in, there was someone wearing a Modal shirt. I was like, “What are you doing here?” They're like, “Yeah, I just. I am embedded inside of Cog.”Akshat [00:08:05]: Yeah, I think that was Peyton. We sent him overSwyx [00:08:07]: Yeah.Akshat [00:08:07]: Because, the latency of communication was too high otherwise.Swyx [00:08:12]: Yeah, distributed node, you have to - you have to place one and collocate.Vibhu [00:08:16]: Yeah.Swyx [00:08:16]: So I had a, I had direct personal experience, right? So I worked on smol developer three years ago. it was inspired by Claude 1. I think you onboarded me at some point, like, just before, and I was like, “Oh, like, I need some bursty compute. Like, I was just gonna try using Modal.” And it was a, it was a pretty pleasant experience. apparently, I showed up in the board meeting, like the analytics.smol developer, Sandboxes, and Proto-CognitionAkshat [00:08:39]: Yeah, you blew up on Hacker News and,Swyx [00:08:41]: YeahAkshat [00:08:41]: We got a big traffic spike. I. I think the way you used smol developer was Modal functions for running stuff, which was. Like, the, that was a good use case. but then, yeah.Swyx [00:08:53]: Yeah. That - So to me, that was proto-cognition.Akshat [00:08:55]: Right.Swyx [00:08:56]: If only I had, like, stuck to it.Swyx [00:08:58]: Like, that was like, if - did you say draw the tech treeAkshat [00:09:00]: AbsolutelySwyx [00:09:00]: You're just like, “Yeah, like, probably this will happen.”Akshat [00:09:02]: Yeah. Like, he was so close. You were just rebuilding upon usSwyx [00:09:04]: I just didn't realize.Akshat [00:09:05]: But the funny story there is at the same time, we were talking to a bunch of customers who needed something like sandboxing.Swyx [00:09:14]: Yeah.Akshat [00:09:14]: This is like twenty-three.Swyx [00:09:15]: Yeah.Akshat [00:09:16]: So we builtSwyx [00:09:17]: You introduced a new API right after that.Akshat [00:09:18]: Yeah.Swyx [00:09:19]: Yes.Akshat [00:09:19]: Like, we built sandboxes in May of twenty-three before anyone was even knew this was gonna be a thing. And the first example we published was, we took smol developerSwyx [00:09:28]: Smol developerAkshat [00:09:28]: And put it in a loop, so the agent can iterate on itself.Swyx [00:09:33]: Loops are hot these days.Vibhu [00:09:34]: It's the looper.Akshat [00:09:34]: Yeah.Vibhu [00:09:35]: Loops in. When was this, twenty-three?Akshat [00:09:38]: Yeah.Vibhu [00:09:39]: A small check.Akshat [00:09:39]: Yeah.Swyx [00:09:39]: It's like twenty-three. so the. the, those for listeners, like, the problem was the models are not built for any of this, right?Swyx [00:09:46]: Like, you're just trying to like. They're not post-training to understand, like, looping and, like, self-correction and tool calling was there, but, like, also not that great.Akshat [00:09:55]: Yeah.Akshat [00:09:55]: I don't remember if you used tool calling in this one, but yeah, the models would just diverge after like ten iterations and not produce anything meaningful.Swyx [00:10:03]: Yeah. But like, then. So okay, like now talking to myself three years ago, the answerVibhu [00:10:08]: Of course they will get betterSwyx [00:10:09]: Collect all the failures, build benchmark, and then collect all the, examples, build the RL environmentAkshat [00:10:15]: RightSwyx [00:10:15]: Sell it for like ten billion dollars to Meta.Swyx [00:10:17]: And then also train a model and then sell that for sixty billion dollars to Elon. And this isAkshat [00:10:23]: Yeah, of courseSwyx [00:10:23]: The funny machine. Like, it's like, it's about the hardware.Akshat [00:10:28]: It's hard to have that inherent conviction that the stuff will get that much better.Swyx [00:10:33]: In retrospect, it's so f*****g obvious.Akshat [00:10:36]: Fair enough.Swyx [00:10:37]: Like, what else were we doing back then? I don't know. anyway. Yeah. So this. That was the start of your sandboxing journey, right? I feel like it didn't blow up until, like, last year.Akshat [00:10:49]: Yeah.Swyx [00:10:50]: So there was like a couple years of quietness.Akshat [00:10:52]: Exactly, yeah. We wereVibhu [00:10:53]: I think very underrated product value. Like, my experience with Modal, Charles, before he had joined Modal, met this guy at a hackathon, and he really insisted we wanted to run some small model, not hosted anywhere, and he's like, “ there's this cool company, Modal. They'll like spin up a GPU sandbox, we can throw it on there. They'll take a Hugging Face link.” And like there's so much value just right there, right? Like instant hosting, spin it up, spin it down. It'll stay cold, but we run the demo a few days later, it'll come back up and like all this stuff in retrospect, like it's still what we needed like today.Akshat [00:11:27]: Yeah, it's still needed today. workload shapes have changed a lot as, we run stuff for people with really massive production scale and, there it's it's not about scaling from zero to one, but it's how do we scale really elastically, from like thousand to fifteen hundred GPUs very quickly in a given region. It's the same shape problem.Elastic Inference, GPU Autoscaling, and Custom ModelsVibhu [00:11:50]: Okay. So you look at, say, Cursor Composer, right?Akshat [00:11:53]: Yeah.Vibhu [00:11:53]: They had a. “We'll do RL on a model every couple hours.” you guys have a whole version of RL inference gym and whatnot.Vibhu [00:12:01]: When you look at workloads like that, you're doing train runs where you need to scale up, scale down every hour thousands of GPUs, right? That's the example for we do need it, right?Akshat [00:12:12]: Yeah. Well, so I'll, I'll take a step back and, maybe talk about like how people use Modal today. because our biggest use case is, elastic inference. And the thing we first found product market fit, with was inference for custom models. So we stayed away from the LLM space, and we were serving companies like Suno for audio, Runway for video, robotics, comp bio companies that train their own model elsewhere. But Modal is the best black box that for deployment, scaling to however many GPUs you need as your traffic pattern changes. And we saw all of them like have a very unpredict- predict- predictable, traffic pattern. it's like diurnal. It's Some days, like the company will do a launch and, they'll need like, way more. And it's not just one model that they deploy. They-- all these companies deploy, lots of different models in different regions, and so the autoscaling problem becomes even harder because then you have to scale within a certain region, and those cycles are offset. So different times you scale up in different regions.Akshat [00:13:20]: So that's like our sortVibhu [00:13:22]: And thatAkshat [00:13:22]: YeahVibhu [00:13:22]: That in and of itself is a huge category. There's a bunch of inference providers which, provide this fireworks, does this as a service together, whatnot, Base10. that's carved into its own niche for language models, at least right now.Akshat [00:13:36]: Yeah. the thing that we have specialized in is the autoscaling aspect.Vibhu [00:13:41]: Yeah.Akshat [00:13:41]: Because we found that it's not universally true that everyone else can autoscale, and we've gone deeper into it on the tech side by, we've incorporated GPU snapshotting into the product so we can take the GPU state, like your torch.compile model, snapshot it, and the next cold start is way faster. And so going back to your question, it's That's why you need a lot of burstiness for inference. But then people also do a lot of demand training, like for RL stuff, your rollouts are bursty, as you said. People also do a lot of batch jobs. So we'll see, a lot of companies, before they have a training run, they'll need thousands of GPUs to run encoding or something like that. And I think those things are much more bursty than. I agree that agents are not that bursty. sandboxes are, except when you're doing RL. RL is justRL, Batch Jobs, and 100,000 SandboxesVibhu [00:14:28]: Or commerceAkshat [00:14:28]: Insanely bursty.Vibhu [00:14:29]: Yeah.Akshat [00:14:30]: Yeah. Like when you're doing, rollouts, you sometimes need a hundred thousand sandboxes in your sandboxes.Vibhu [00:14:37]: Yeah. I'm curious if you've seen early sparks of continual learning. There are some people, like our friends, ngram, recently announced thisAkshat [00:14:45]: YeahVibhu [00:14:45]: They're, they're trying to do training. That also seems like a different workload, right? If you're doing training twenty-four/seven per se, there's a very weird dynamic of how you're using GPUs between people and whatnot, but seems like something you guys would work for.Akshat [00:15:00]: As you said, we're, we're fortunate to work with a number of, customers at the frontier and grab some of our customers. and they are taking the primitives we have, and trying to use them in very interesting ways, like continual learning. It's possible as the stuff gets better, some of that will be part of, our offering as well if, more people need it. but we're, we're just waiting to seeVibhu [00:15:23]: YeahAkshat [00:15:23]: How it shakes out.Vibhu [00:15:24]: Is there a primitive that you added after sandboxing that was the next step in the story?LLM Inference, DeFlash, and Speculative DecodingAkshat [00:15:32]: I guess we've been going much deeper into LLM inferenceVibhu [00:15:35]: YeahAkshat [00:15:35]: Because we realized that some of the advantages we have with like autoscaling, again, especially in different regions and whatnot, are, not present elsewhere. and the place where we had a gap was we weren't, working on the model layer itself. Like we were a black box. And, we realized that, we can get to frontier-level model performance, with, by having great people who work on this. And, we've been open sourcing a lot of our work, in terms of, Recently, we, shared our work on DeFlash, which is a block-based, speculator, and we've open sourced, all of it. So, you can - By using open source DeFlash, you can get the same performance as you would with one of the proprietary providers. And the next thing we're thinking about hereVibhu [00:16:23]: I thought this wasAkshat [00:16:24]: YeahVibhu [00:16:24]: An interesting blog post as well, right? Like, I think in here you make a claim that. Not a claim, just that how effective speculative deco-decoding really just get to.Akshat [00:16:33]: Yeah.Vibhu [00:16:33]: Anything you wanna point out from this around, what people should know?Akshat [00:16:39]: Yeah, absolutely. the high-level summary is, it would help to describe what speculative decoding is.Vibhu [00:16:44]: Yes.Akshat [00:16:44]: I will, yes.Vibhu [00:16:45]: I think, likeAkshat [00:16:46]: YeahVibhu [00:16:46]: So we've covered like Eagle and all thisAkshat [00:16:47]: YeahVibhu [00:16:47]: Like Hydra and all those things, but it was like two years ago.Akshat [00:16:51]: Yeah.Vibhu [00:16:51]: I think it doesn't hurt, right?Akshat [00:16:52]: Yeah. Speculative decoding is you have a smaller model, called a draft model, predict tokens ahead of the bigger model, and then you have the bigger model, verify all of this, all the tokens are predicted. And the reason it's faster is if you're predicting, one token at once, you're bound by memory bandwidth. But if you can batch the verification of, the draft model, then you're much more efficient using compute, and it's faster, and as long as your draft model is producing a lot of tokens that can get accepted, which is called the accept length, you can get a speed up that's, multiple times of, the original model speed. and well, that's what we highlight here. It's Like people talk a lot about we made these kernels faster and whatnot, but improving kernel will only give you like few percentage points of improvement, and, increasing accept length, literally is a multiplicative decreaseVibhu [00:17:47]: Like two to four X.Akshat [00:17:48]: Yeah, exactly.Vibhu [00:17:48]: Without much head-on performance.Akshat [00:17:50]: Yeah. I think it may - you are running a second model, right? So it may be something more expensive in the compute,Vibhu [00:17:57]: I meant quality performanceAkshat [00:17:58]: Probably not by muchVibhu [00:17:58]: But yeah. I thinkAkshat [00:17:59]: So there's no drop in quality performanceVibhu [00:18:01]: YeahAkshat [00:18:01]: Because you're always. You're never accepting a token that the big modelVibhu [00:18:04]: It's strictly betterAkshat [00:18:05]: YeahVibhu [00:18:05]: Or it's same.Akshat [00:18:06]: Exactly.Vibhu [00:18:07]: Right. Yeah.Akshat [00:18:08]: And so we've been working a bunch on DeFlash, which is a block-based speculator. so it's instead of predicting, one token at a time, it's predicting a block. And we've been open sourcing our work with it. The next thing for us here is for helping people train speculators and custom models. it's it's something that traditionally is very forward-deployed engineering driven, support deployed, engineer driven, like you work with customers and help them do that. And our vision for. This is why we launched Auto Endpoints, is we want to make frontier-level performance available to everyone. And so, we mentioned this in the announcement, we teased it. The next thing we're, we're launching is, as you run an auto endpoint, we shadow trafficAuto Endpoints and Frontier-Level PerformanceVibhu [00:18:54]: Do you want to explain what auto endpoints are?Akshat [00:18:57]: Yeah.Vibhu [00:18:57]: I lovely, yeah.Akshat [00:18:58]: Yeah. So, this is, I guess, going back to your Modal is you touch the code, but, sometimes people don't wanna touch the code, and they wanna get started with an endpoint that works and has all the great performance and, scalability that Modal has. So we've made that easier with, a way to create an endpoint from our UI, from the CLI, that has all of our optimizations that we talked about, like the DeFlash stuff already baked in, and there's full transparency. So we give you the code, you can go run it yourself, and if you want, you can eject out into the full Modal experience, which we see as people get sophisticated, they do wanna tweak the models, they wanna, fine-tune stuff. You can still do all of that. It's it's not a black box. And yeah, the next thing, as we teased later in the post, is how do we give you value even beyond this in terms of having your draft models evolve as your data distribution evolves, again, without having to talk to a person and, yeah.Vibhu [00:19:59]: I guess just to understand it directly, you have the GPUs, you have an endpoint that's compatible, you serve open model. If someone was to do this themselves, what's the delta that you guys provide? So you do a lot of open source great work on effective inference. how does it compare to, say, I take the same model, 5.2 FP8, take shelf inference engine, vLLM, SGLang, get compute of similar capacity, similar cost. What's the delta that plugging into something this, like this offers outside of the benefit of, scaling?Production Inference Beyond Raw GPUsAkshat [00:20:34]: It's interesting because we've taken the approach of open sourcing our contributions and upstreaming them. we work closely with the SGLang team. We want the improvements that our team, comes up with to be, there in open source for others to use, even outside of Modal. The benefit to us is we have a team that has significant expertise in terms of if you do have something that is not there, our team can help you get that performance, first. the other thing is with these endpoints, we are way more elastic, as you said, than, anyone else, and you have true scaling to zero. you have true, burstiness, and in practice, that matters a lot more to people than just finding, the GPU and, running Modal code on something.Vibhu [00:21:20]: Yeah. And I will say it's not that straightforward to just. like what I said is easier said than done, right?Akshat [00:21:26]: Yeah.Vibhu [00:21:27]: It's I think still for the average person, still hard to just gut check using different. There's, there's quite a bit of combinations you can make there. the trade-offs aren't really known at face value.Akshat [00:21:40]: Yeah. it's it's not just that. I think it's it's that running production-grade inference is a hard infer problem.Vibhu [00:21:49]: YeahAkshat [00:21:49]: Even if you subtract out the autoscalingVibhu [00:21:50]: YeahAkshat [00:21:51]: Is controlling things like tail latency and, making sure every, request is delivered at least once and whatnot.The Model and Agent LifecycleVibhu [00:22:00]: There's a lot of innovation that you can do here. I think, it's very interesting that you're starting to encroach on, like as you become a full cloud, you're starting to encroach on other people's turf.Vibhu [00:22:09]: What will you not do?Akshat [00:22:13]: Well, we wanna follow our users and, make sure they get like a platform that has everything that works well together. so right now we're focused on the model lifecycle and the agent, lifecycle. so both like going from data prep to training to inference, and then also if I want to deploy a background agent, let's say, sandbox, do persistent storage, a whole bunch of other stuff.Vibhu [00:22:38]: We talked to Cole, who did, OpenInspect. Yeah.Akshat [00:22:42]: Yeah.Vibhu [00:22:42]: And RealInspect also is on Modal.Akshat [00:22:44]: Yeah. So Ramp Inspect was a great example of a background agent that was really successful because they, were able to use some of the primitives like snapshotting and fast scaling to just have something that feels really reactive and works well.Ramp Inspect and Background AgentsVibhu [00:23:02]: Yeah. That's the new CTO of, Ramp right there.Akshat [00:23:05]: Yeah, Rahul.Vibhu [00:23:08]: It was really fun. yeah, okay, I think, all very bullish. Like, one of my reflections was also I did not originally. So when I met you guysThe Inference Inflection: CPU, GPU, and Co-LocationVibhu [00:23:19]: You weren't that much in the GPU game, and now you're all about, inference. And one of the points that I hinged on for Jensen's keynote at GTC this year was, what we're calling like the inference inflection, right? That let's say in AI workloads or machine learning workloads, it used to be like, let's call it eight to one GPU to CPU, and now it's more like one to one, which is like a interesting. Like, - because of how much agents are blocked or call out to this, to CPU heavy stuff the actual, like, limiting factor, like, swings back and forth from GPU to CPU a lot more than it used to be all GPU and then occasional CPU.Akshat [00:24:01]: Yeah.Vibhu [00:24:02]: GPU, CPU. And now it's like just constantly, and you just have to locate everything.Seventeen Clouds and the Supercloud StrategyAkshat [00:24:08]: Yeah. And that's one of the things that, again, we see as, something appealing about Modal, which is we've built this capacity pool that spans, 17 cloud providers, so we're, we're very good at Running on various kinds of cloud capacity across the worldSwyx [00:24:24]: You don't have your own data centers?Akshat [00:24:25]: We don't have our own data centers. We just run across a lot of neo cloudsSwyx [00:24:29]: Yeah. AreAkshat [00:24:30]: Metal providers.Swyx [00:24:30]: Yeah. Question mark.Swyx [00:24:31]: Yeah. You're, you're running the math, and you're like, “What's the cutover point where you're like.”Akshat [00:24:36]: Yeah, it's a good question. part of it is we see our differentiator in the software layer, and, being capital light and focusing on the software helps us move really fast. so far it's worked out well because there are so many other people building data centers that we're able to work effectively with them, and again, focus on what makes us, special.Swyx [00:24:55]: Yeah.Swyx [00:24:56]: 17 gets you into, like, the local providers sometimes. LikeAkshat [00:25:00]: The,Swyx [00:25:01]: Which was the most interesting one?Akshat [00:25:02]: There are a lot more neo clouds than you expect, and they all have various degrees of, various levels of reliability. And, that's why it's something we've invested a lot of time in, is building our own reliability layer on top. so if the GPU falls off the bus or something happens, we user workloads are not affected, and that lets us use a lot more capacity than,Swyx [00:25:30]: YeahAkshat [00:25:30]: You as a user would be able to.Swyx [00:25:32]: It's a useful thing to have because like now everyone knows, like, what layer you are and, like, you optimize for being the super cloud of all clouds.Akshat [00:25:41]: Yeah. That's, that's, that's the idea. and so I guess when you mentioned colocation, that's, that's another interesting thing where, one thing we've seen is people come to us when they want, very specifically located, CPUs or GPUs, like they wantSwyx [00:25:57]: Oh, they pin it in likeAkshat [00:25:58]: YeahSwyx [00:25:58]: EU?Akshat [00:25:59]: Exactly. Or EU, US.Swyx [00:26:01]: Right. Data resiliencyAkshat [00:26:02]: AustraliaSwyx [00:26:02]: Locality thing or performance or what?Akshat [00:26:04]: It's either data locality or latency, yeah.Swyx [00:26:07]: Yeah.Akshat [00:26:07]: Like, you want your. They're running sandboxes and model. They want them to be right next to aSwyx [00:26:10]: Yeah, it's easy thenAkshat [00:26:11]: YeahSwyx [00:26:12]: To. That is important in all those things. and so, like, you've accidentally, I don't know if it's accident, but, like, you've built the perfect primitive for agents to express themselves. And then, like, it's almost very funny how every extra development just involves more file system, just involves more CPU.Akshat [00:26:30]: Yeah.Swyx [00:26:31]: Just like the things that you already have. I don't know much about, if there's any, like, networking usages that are interesting, but you've also done some good work on networking.Networking, Sidecars, Private IPv6, and SandboxesAkshat [00:26:40]: Yeah, that's exactly right. Like, we're just taking compute storage and networking and building stuff on that layer, for, again, the stuff people need.Swyx [00:26:49]: YeahAkshat [00:26:50]: We see a few interesting networking things coming up. one is people want networked sandboxes. so we haveSwyx [00:26:57]: For like a Docker cluster type thing.Akshat [00:26:59]: Yeah.Swyx [00:26:59]: Sorry, Docker Swarm. Oh, f**k. What is it called?Akshat [00:27:02]: Compose.Swyx [00:27:03]: Compose type thing.Akshat [00:27:04]: Yeah. So if you want Docker Compose, our sandboxes now support, this thing called sidecars. So you can. A sandbox is a pod of containers, and you can run multiple containers in, a sandbox. also useful because, going back to networking, people want a lot of control over, outbound networking from a sandbox.Swyx [00:27:23]: Yeah.Akshat [00:27:23]: Like, they might wanna run a middle proxy for, like, maybe logging stuff for RL or, controlling how egress can happen to a domain, injecting credentials. and yeah. So we've, we've had to build a lot of that stuff ourselves.Swyx [00:27:38]: Yeah.Akshat [00:27:39]: But then also sometimes people want, sandboxes spanning multiple nodes to talk to each other, which is an emerging thing we're seeing. We have support for that for a different reason, and yeah, we'll see if that becomes stable.Swyx [00:27:52]: Like, just an open socket. It's a. This is directly like mTLS.Akshat [00:27:56]: We do support that, which is you can, expose a tunnel inside a sandbox.Swyx [00:28:01]: Yeah.Akshat [00:28:01]: And then you can either expose it to public internet or it can be, you can add like a HTTP, auth layer above it. But we have this thing called I6PN, which we haven't talked about, which is this, like, overlay network using IPv6 addresses. so if Modal containers, within the same workspace, when this is enabled, can address each other using this private IPv6 address, and no one else can.Akshat [00:28:28]: So it's like private networking, for containers. We built it because we needed it as a primitive for our distributed training product. so we have this other feature, which is you can add a decorator to a function, and you get a cluster of GPUs. and they have RDMA networking. so you can run a distributed training job, that's truly serverless. and we did the overlay network for that. But then we've seen that people are using it for other reasons, and, I'm intrigued to yeah, what would people do with it.Swyx [00:28:59]: Build primitives and let people figure it out, right?Akshat [00:29:01]: Yeah, exactly.Swyx [00:29:02]: You put out a pretty interestingAkshat [00:29:03]: They're like, they read the docs webpage. Let me use thatSwyx [00:29:06]: YeahAkshat [00:29:06]: Something they never intended to work. This is literally not even in our docs page. People somehow found it, and they're using it.RDMA, Memory Movement, and Distributed TrainingSwyx [00:29:12]: Huh.Swyx [00:29:14]: The way you portrayed it with, like, RDMA versus TCP, like, very well laid out, but just the transfer speed change at scale for RL, like yeah, you have it, you have it built in. I'm sure someone found it. It's found it to be a lot more efficient before you made a thing out of it, right?Akshat [00:29:32]: Yeah. And not to split hairs, I guess the overlay network is the TCP overlay network.Akshat [00:29:39]: The reason we have that is you need that to do the key exchange for RDMA before you set up the RDMA network on top of that. but then people found the TCP part.Swyx [00:29:48]: Can I tell you, this is like a big aha moment for me becauseAkshat [00:29:51]: YeahSwyx [00:29:51]: So I review 2,200 submissions for the World's Fair.Akshat [00:29:56]: Yeah.Swyx [00:29:57]: And then I got this from John OsterhoutAkshat [00:29:58]: HuhSwyx [00:29:59]: Who I don't know if. Do John Osterhout by name?Akshat [00:30:01]: The name sounds familiar.Swyx [00:30:02]: He published a. He's a well-known professor, published a lot of interesting software design books, and this is the talk he chose to submit, is on RDMA at Inference. And I'm like, you wouldn't think that this guy, who is like operating systems guy, would care about RDMA.Akshat [00:30:20]: I, it makes sense to me because I,Swyx [00:30:24]: This is the cloud, right? YeahAkshat [00:30:25]: Like, the way you move around your KV cache and how efficiently you can do it, how efficiently you move, your weights from your training GPUs to your inference GPUs in RL is there's a lot of degrees of freedom, and it is a systems problemSwyx [00:30:41]: YeahAkshat [00:30:41]: Moving memory aroundSwyx [00:30:42]: YeahAkshat [00:30:43]: Scheduling.Swyx [00:30:44]: This shows you how primitive my understanding of networking stuff is.Swyx [00:30:46]: Is this like the domain of WireGuard as well?Akshat [00:30:50]: Not quite.Swyx [00:30:51]: It's adjacent?Swyx [00:30:53]: Explain everything.Akshat [00:30:54]: Sure.Swyx [00:30:56]: How do we move memory around GPUs?Akshat [00:30:58]: Well, so sorry. Yeah, that is memory. Sorry, I was talking more, and maybe I was talking like five minutes back, about the private IPv6, addressing that you've set up.Swyx [00:31:09]: Yeah.Akshat [00:31:09]: Is it like it's a VPN?Swyx [00:31:10]: Yeah, it is like a VPN, and yeah, WireGuard is, yeah, you're right. It is,Akshat [00:31:16]: Right. Yeah, you already moved on to new topicsSwyx [00:31:17]: A similarAkshat [00:31:18]: OkaySwyx [00:31:19]: In the same space, WireGuard is, encrypted and this is,Akshat [00:31:23]: And you don't need encryption.Swyx [00:31:23]: Yeah.Akshat [00:31:24]: Yeah.Swyx [00:31:24]: This is not encrypted. that's the main difference. This is TCP and we have eBPF programs that will reject or allow the TCP connection based on whether you're allowed to do it.Akshat [00:31:35]: Used to involve a full sidecar, but now you have eBPF in the Linux kernel.Swyx [00:31:39]: Yeah.Akshat [00:31:40]: Yeah. I don't know if this is a natural follow-on to the topic of like my skepticism on distributed training is that while, like, people spend a lot of money on, like, cables to hook up GPUs, and even that is not, like, fast enough, and that's the bottleneck, is your networking fast enough?Swyx [00:31:59]: Yeah. So I guess you're talking about fully distributed training like, Dialog or something which is like cross data centerAkshat [00:32:06]: That would be, yes.Swyx [00:32:07]: That's the extreme.Akshat [00:32:08]: Yeah.Swyx [00:32:08]: You're in the middle, and then other people would have like the Mellanox cables up in, like, their actual data center.Akshat [00:32:14]: When you run multi-node training on Modal, RDMA, I think Mellanox, is, or InfiniBand is like a, is all seen as RDMA. but it's a way to bypass the TCP networking stack and, transfer, stuff much faster, between one node, to the other. And we have I think like 3 terabit per second, internal networkingSwyx [00:32:40]: OkayAkshat [00:32:40]: Which is the standard that's needed.Swyx [00:32:42]: Okay. So I misunderstood whatAkshat [00:32:43]: 50Swyx [00:32:43]: What part of the stack you wereAkshat [00:32:44]: 50 gigs overSwyx [00:32:45]: YeahAkshat [00:32:45]: If you wentSwyx [00:32:45]: YeahAkshat [00:32:46]: RDMA.Swyx [00:32:46]: Okay.Swyx [00:32:48]: Yeah. I, very impressive work.Multi-Node Training, Post-Training, and Auto ResearchSwyx [00:32:52]: So effectively you're extending like the model philosophy to the training cluster, like, yeah.Akshat [00:32:59]: Yeah. And we're, we're not going for like large scale training runs. the thing that we've built multi-node training for is, we see a lot of, smaller scale post-training. like, people are post-training like medium sized fund models, so they can, get higher quality on inference. this is a perfect fit, for something like that.Swyx [00:33:21]: Yeah. That is my impression of how a lot of these labs explore branches in post-training and then eventually merge whatever they find in.Akshat [00:33:31]: Yeah. The other use case we've seen for multi-node training is even if you have a big cluster, your researchers are still doing small runsSwyx [00:33:38]: YesAkshat [00:33:39]: Having elasticity thereSwyx [00:33:40]: Right, sureAkshat [00:33:40]: Matters a lot more.Swyx [00:33:41]: Yeah. the, like, this is like the current limiting factor for auto research, which is like you need to give your model some GPUs in order for it to completely run.Akshat [00:33:51]: We have a blog post on auto resource and model is,Swyx [00:33:55]: YeahAkshat [00:33:56]: Yeah, like, turns out to be pretty good substrate for that.Swyx [00:33:59]: So my impression is auto research means many things, likeAkshat [00:34:01]: YeahSwyx [00:34:01]: Anything that Andrej coins. Right now it's still science fair, right? Like not like, I don't know how many people are doing this.Akshat [00:34:08]: We're having a golf.Swyx [00:34:08]: Yeah.Akshat [00:34:09]: I thought the same thing.Swyx [00:34:11]: Yeah, you would know.Akshat [00:34:12]: We, like, our internal both training and inference teams use this the general shape of this quite a bit. like we have this one internal repo called auto inference, which essentially we've automated our own forward-deployed engineering efforts using, this harness, which is, the agent will just spin up a sweep of different things. It'll even run like, NVIDIA inside profiler and it'll like tweak configs and it'll arrive the right thing. it'll change your GPUs both from H200 to B200, and works really well.Swyx [00:34:47]: Nice.Akshat [00:34:47]: So yeah.Swyx [00:34:48]: By the way, I enjoy that your forward-deployed engineering is so technical that you have to do these things.Swyx [00:34:52]: It's very different from forward-deployed engineering from other people.Akshat [00:34:54]: Yeah. For our forward-deployed engineering team is, essentially they're like applied inference researchers or applied training researchers.Swyx [00:35:02]: Someone told me like they have to be able to build, but they also have to be able to sell. do they have to sell or are they like they're good, they're just like post-sale type of thing?Akshat [00:35:09]: It does, being able to talk to a customer and engage effectively with themSwyx [00:35:13]: YeahAkshat [00:35:13]: Matters a lot.Swyx [00:35:14]: They want the same thing.Akshat [00:35:15]: Yeah.Swyx [00:35:15]: ?Akshat [00:35:15]: But it's it's not really a sales, thing. We pair them with-- We have solution architects as well that are more on the sales side.Swyx [00:35:23]: Okay. Let's spend a bit more time on auto research. This is a big focus for for this year. Where does this go? like, have people explored enough? Like, there's all these beautiful charts of like improve and then level off a bit and then you find the next thing. Is this one abstraction up from normal training? Is that how we think about it, or do you think about it differently? Like model level training versus high, like driven hyperparameter search.Auto Inference and Modal BenchAkshat [00:35:51]: Yeah, like,Swyx [00:35:51]: Someone, some people call it like neural architecture search or whatever, right? Like.Akshat [00:35:54]: Yeah, - So the stuff I've seen people do with it is nowhere on the architecture level. It's pretty much tweaking parameters, but it's it's a hyperparameter sweep that's guided by some model intuition, so it's like much more efficient than, whatever other, sweep you would have.Swyx [00:36:12]: Yeah, it's just, it's just a question of where you want to spend your compute?Akshat [00:36:16]: Right.Swyx [00:36:16]: ‘Cause yeah, you can just throw infinite amounts of money on this and somehow you'll bang out Shakespeare?Akshat [00:36:22]: Yeah, infinite monkey.Swyx [00:36:24]: Yeah, so like the very good for model. and I think it's also very important that agents can spin up other agents, can spin up their infrastructure. Like very good for you. how good is our LLMs at generating model code? Like the benefit of existing LLMs is that you are in the data.Akshat [00:36:42]: Yeah. They're, they're surprisingly good. I think like pre Cloud 4 they were not, and then now they're able to shot, stuff out of the box. But we're playing around with releasing like a Modal Bench for like the harderSwyx [00:36:55]: YeahAkshat [00:36:55]: Things, that the LLMs cannot do yet and maybeSwyx [00:36:59]: What's an example of that?Akshat [00:37:01]: I think the things that- Sometimes agents struggle with, without right guidance and a skill is, how to, use the rest of our observability. Like how to. Something is failing, like how do you look at the logs and then update the right thing? It's reasoning about that. But they're able to shot, likeSwyx [00:37:23]: Yeah. You can just add a skill to it?Compute Strategy and Capacity PlanningAkshat [00:37:26]: Yeah. So we have a Modal skill now that. Which is why we built this Modal Bench. It's to find things like that, so we can address them in our tool.Swyx [00:37:35]: Tune a skill. Yeah.Akshat [00:37:36]: Yeah.Swyx [00:37:36]: No. it's it's good. are you facing any shortages? like we talk a lot about GPU shortages, but also CPU, also memory.Swyx [00:37:44]: Yeah.Akshat [00:37:45]: We have had a lot of growth, which means that, there's - we've had to be much better aboutSwyx [00:37:53]: PlanningAkshat [00:37:54]: Proactive capacity planning.Swyx [00:37:55]: Yeah.Akshat [00:37:55]: So we have,Swyx [00:37:57]: Which by the way, like it's like a MBA's like dreamAkshat [00:38:00]: YesSwyx [00:38:00]: Is like just planning this stuff. I think last time you and I talked about something maybe about this.Akshat [00:38:03]: Yeah. we have a really competent team of people that we call, The role is called compute strategy. so yeah, if anyone listening here or wants to work on thatSwyx [00:38:13]: Compute strategy?Akshat [00:38:13]: Yeah.Swyx [00:38:14]: I think,Akshat [00:38:14]: I feel like,Swyx [00:38:15]: I think the normies call it FP&A or something.Akshat [00:38:18]: Well, it's more It's it's not FP&A. It's it's There's a lot of interesting financial questions of like what is the blend between one year and three-year reservations? how do we forecast our own capacity? how do we. especially since our capacity is very fungible across different GPU types and different regions, like you have to model a lot of it. and you also have to have an opinion on how the supply chain is gonna evolve, and then you have to like, take bets,Swyx [00:38:49]: YeahAkshat [00:38:49]: Based on that.Swyx [00:38:50]: Tokenomics.Akshat [00:38:50]: Yeah.Swyx [00:38:51]: This is like probably a not a real point, but, I was trying to think about like what other industries. I was trying to think about like, we cannot be first to like these kinds of problems.Akshat [00:38:59]: Yeah.Swyx [00:39:00]: And what other industries have had this? And I was like, airlines with fuel and like they have to hedge their fuel and like, I think for a long time Southwest because they made like a hero fuel bet, they like were like super low cost becauseAkshat [00:39:12]: OhSwyx [00:39:12]: Compared to everyone else.Akshat [00:39:14]: Yeah. I hadn't thought about that.Vibhu [00:39:16]: We're at a fun time too?Akshat [00:39:18]: Yeah. It's. A lot of the compute business in general, for us is also about being very good about capacity management. That is how you have great unit, economics. but also over time it's how you can unlock more value for customers. Like, one of the things we're building now is like a way for customers to get, If they don't care about latency, like get much cheaper pricing and they'll get results back in like next 24 hours or something, like a batch tier essentially.Batch Tiers and Latency-Insensitive WorkloadsSwyx [00:39:47]: Yeah.Akshat [00:39:47]: And those are levers we have because we control the whole stack and scheduling and whatnot to give people a sufficientSwyx [00:39:53]: Yeah. I feel like they're not as popular. Like those, like the Frontier Labs have all those APIs. They're not as popular as they should be.Akshat [00:40:00]: The demand that we see for something like that is not for LLMs. although sometimes people wanna run evals andSwyx [00:40:08]: OkayAkshat [00:40:08]: Synthetic data prep and there it makes sense.Swyx [00:40:10]: Okay.Akshat [00:40:11]: But it's from a lot of LLM companies, like people who are doing computational bio, like they have to run really big batch jobs and they don't care about when they get it back.Swyx [00:40:22]: Yeah. And like they have a reasonable. It's it's also like a cousin to the stopping problem of like, will this finish in time?Akshat [00:40:30]: Yeah. You can bound it.Swyx [00:40:33]: Yeah.Akshat [00:40:33]: Like you can give peopleSwyx [00:40:34]: YeahAkshat [00:40:34]: SLAs on it.Swyx [00:40:35]: Yeah. I think what's, what's interesting is like the next phase of model.Swyx [00:40:38]: Like what, do people expect from you, now that you're established and you're like well-known compute player among all these leading companies. You had an inference launch week, and we talked a little bit about the launches. like what else? Like what else should people know?What Modal Builds NextAkshat [00:40:55]: We are building primitives that make our users' lives much easier. So, I think for example, with LLM inference, thousands more companies are gonna post-train their own models and, deploy open source models for inference. so we're thinking a lot about what is the best product shape for that. And, that involves everything from our training gym to, then, endpoints that get frontier-level performance. again, but I haven't talked to anyone. It looks somewhat different on other verticals. Like, we're also seeing a lot of real-time, audio-video stuff in there, which is why like, we're working on things like regional routing, with fallbacks. So you can get GPUs that are as close to users as possible. so you get like low latency for video streaming and whatnot. And then on the agent side, it's,Akshat [00:41:52]: We're still working very closely with our customers because stuff is changing so fast in terms of what they need. And, I think beyond sandboxes and persistent file systems, there's a lot of other things people will need from this agent stack as they build production agents. So yeah, we're thinking about those other things that fit in there.Swyx [00:42:13]: I want to ask what the other things are.Akshat [00:42:15]: Yeah. I probably should share right now.Swyx [00:42:17]: I think-- I think, okay, so, I do think a lot about the principal components of cloud, and you do talk about compute storage networking.Akshat [00:42:25]: Yeah.Swyx [00:42:25]: Because so far for me, it's fine. so far for the. the first couple generations of cloud, it's fine. What's different, qualitatively different about agents that you need some new permission level? Like a lot of people, okay, and I'll just kinda spew tokens at you until it like hopefully sparks something.Akshat [00:42:43]: Yeah.Swyx [00:42:44]: Like the new level now is whatever Claude Code does, which is dangerously scope permissions or like allow list by command or like whatever, right? And sometimes they're like, “Well, okay, we have like this adaptive thinking mode where like, just trust me, bro. I will make the calls for you.” Is that it? like mediated permissions.Hard Guardrails vs. LLM-Mediated PermissionsVibhu [00:43:03]: Now you're looping it with a goal and letting it roll.Akshat [00:43:06]: Yeah, I'm, I'm skeptical of LLM media permission for stuff that is at the sandbox level because you do want hard boundaries.Swyx [00:43:16]: Yeah.Akshat [00:43:16]: Otherwise, someone can exfiltrate stuff.Swyx [00:43:20]: But likeAkshat [00:43:20]: YeahSwyx [00:43:20]: Maybe that's old school thinking. Maybe we're the dinosaurs.Swyx [00:43:23]: Maybe the AI OS or the LLM OS is really the kernel is a goddamn LLM.Swyx [00:43:30]: Like it makes you feel uncomfortable.Akshat [00:43:31]: Yeah, I'm, I'm toldSwyx [00:43:32]: But that's what trusting the LLM is. Like imagine a spherical cow perfect LLM.Akshat [00:43:36]: Right.Swyx [00:43:37]: That it.Akshat [00:43:39]: Maybe.Swyx [00:43:41]: I wanna test the boundaries, right?Akshat [00:43:42]: Yeah.Swyx [00:43:42]: Like, and I don't believe that, but I wanna see where I'm wrong ‘cause that's, that's the consensus.Akshat [00:43:49]: Yeah. I think you always need hard guardrails when you want, And you can pair those with softer guardrails, right? And that's gonna be a lot of mediated.Managed Agents and Specialized SandboxesSwyx [00:44:00]: There. I'll also get you a end with a couple of your commentary on like the ecosystem outside of Modal. Manage agents. Everyone has one. Gemini, OpenAI, Claude, very useful for you, but also like it is their way of starting to edge into your space.Akshat [00:44:17]: Yeah.Swyx [00:44:17]: What's going on?Akshat [00:44:19]: Yeah, we're, very excited to partner with Anthropic and some of the other foundation labs, will not name who we're also working with. the way we see it is the manage agent thing is a great place to start if you're starting out building an agent and, But then when you get to, building something more production grade, like you're a company that's like Ramp that's building their own, Ramp also runs their accounting agent on us, so their external-facing agent. You need a lot more control over, your compute primitive on things like, what sort - how do you persist different files that the agent has access to, and how do you snapshot and restore? How do you control the networking? maybe you want GPUs. When you get to that point, you kinda want, a specialized sandbox provider, that gives you those things, and that's the role that we are trying to play.Swyx [00:45:15]: YeahAkshat [00:45:16]: We don't really have an opinion on the harness, whether it runs - it's a cloud-managed agent, and you hook it up to Model Sandbox, or you run the harness in Model Sandbox. We'll see where people converge with that.Swyx [00:45:26]: Yeah. Do you any opinions on like the meta harnesses, or just another layer on top of these things?Akshat [00:45:31]: You mean like the OpenPipeSwyx [00:45:33]: OpenPipe is one. I think Vercel had one, which I can't remember the name of right now. Fredshot had one. and then, to me, most recently was Data Databricks that had Omnigen. All these are meta harness. Like it's kinda pseudo agent cloud type things.Akshat [00:45:50]: I personally have not played around with them.Swyx [00:45:53]: Yeah.Akshat [00:45:53]: Build agents with them.Swyx [00:45:54]: Everything's bullish Modal, as long as it consumes more infra.Akshat [00:45:57]: That's why we're focusing on the infra layer. It's somewhere where our, relative competence is and, also it's a hard problem to solve.Swyx [00:46:06]: Yeah. I will say like just generally reflecting on that, I don't know if - if there's other topics on Modal, but like just generally reflecting as an infra person, not as intense as you, but in that field, this has like been the most exciting time in infra. Like it was boring for a while, and you couldn't really get people excited about data infrastructure. Like Eric would get on Data Console, everyone just watched the video and like say, “Look at how many sandboxes I can spin up,” and no one gave a crap.Why Infrastructure Became Exciting AgainAkshat [00:46:39]: Yeah.Swyx [00:46:40]: And like now everyone gives a crap.Akshat [00:46:42]: That's true. It is a very exciting time, and I think a lot of that's driven by just the amount of scale all of this stuff needs.Swyx [00:46:50]: I think the, like a lot of your initiatives or a lot of your like product directions make sense in retrospect, which is like the best kind, but I wouldn't necessarily have thought about it myself, which.Akshat [00:47:00]: We need the predictions.Swyx [00:47:02]: I think there's a lot that you just don't even see, right? Like you have the batch, you have the voice, you have the multimodal, but what else?Akshat [00:47:10]: What else is coming up for usSwyx [00:47:11]: Yeah. Where do you see things going?Akshat [00:47:13]: Yeah. I, in generalBiotech, Robotics, and Non-LLM AI WorkloadsAkshat [00:47:15]: It's it's clear that there's there's a huge shift happening. I think one thing that's not as obvious to people because LLM inference gets talked about so much and is also we work a lot of companies that are, doing things like drug discovery and computational bio, like the Chai Discoveries of the world. Big things are probably gonna happen there. we work a lot of robotics companies that are putting robots in like active deployments and getting good results out of them.Swyx [00:47:45]: Is there Air Gap Modal? Is there a version that is like prem air gapped whatever?Akshat [00:47:50]: No. We,Swyx [00:47:51]: You should cloud only.Akshat [00:47:51]: Yeah.Swyx [00:47:52]: Yeah. Okay. But yeah, so what you're saying is like because you're focused on primitives and they're good primitives, you find use cases in all these kinds of things.Akshat [00:48:01]: Yeah.Swyx [00:48:01]: Probably diversifies you a little bit away from LMS all the time.Akshat [00:48:05]: Yeah, absolutely. We're, we'- our goal isn't to only serve the LLM inference market.Swyx [00:48:10]: There are a lot just on the website, the audio,Akshat [00:48:12]: Yeah. We said both onSwyx [00:48:14]: Computational bio images. Yeah, there's a lot here. There's QTA TTS, customizing. Oh, Chatterbox. there was customizing Whisper.Akshat [00:48:24]: Okay. Yeah.Swyx [00:48:25]: This screen reminds me of a fallen competitor, which Replicate.Model APIs vs. Differentiated AI ProductsSwyx [00:48:31]: What's your postmortem on what happened?Akshat [00:48:34]: This is one thing we've stayed away from is providing an API for models because I think providing model APIs is some of it ends up serving like a really hobbyist market, which is much less sticky.Swyx [00:48:50]: Yeah.Akshat [00:48:50]: And we've always wanted to build for companies that are building products and need more flexibility that's not just an API.Swyx [00:48:57]: Which you can build an API for a model and this is clearly what it is. But you - but what you're saying, you can wrap it into a more fully functioning back end that you run.Akshat [00:49:06]: Yeah. So all of our examples, it's not that spin up this model, here's an API token, use it. They're all code.Swyx [00:49:13]: Okay.Akshat [00:49:13]: And so the point is that this is just an example.Swyx [00:49:16]: Starter code.Akshat [00:49:17]: Yeah. But you can tweak it however you want.Swyx [00:49:20]: Yeah.Akshat [00:49:21]: And if you're like a company building a product, like, computational bio whatnot, yeah.Swyx [00:49:26]: I guess I'm trying to tease out for listenersAkshat [00:49:28]: YeahSwyx [00:49:28]: When does it stop becoming, oh, you're just an API call and you're just a wrapper on API to becoming what you call a product, right?Swyx [00:49:36]: Like, what is that layer? Like what-- Like, more lines of code, but like beyond that, what is the substance that people add that qualifies it to be something more?Akshat [00:49:46]: I think there's a little bit of like a selection effect of like a lot of the companies who do wanna get deeper into that level are probably building something that's more differentiated. And, I think, an example is like - with LLM inference, originally we, worked with companies that were building their own post-training frameworks or they were, - Ramp early in the day was training their own tokenizer and like swapping out the tokenizer in Llama and whatnot. I'm not saying that's, that successful, in that case. But a better example is like, let's say Suno. because Suno, does not use Modal for training.Swyx [00:50:26]: Mikey on the pod. Yeah.Akshat [00:50:27]: But they use Modal for all their inference and that's because they have like a custom-- They have completely custom model architecture and that means that they have to be at the code level and tweak things that are not, just an API.Swyx [00:50:41]: It's interesting as well, like we had, Ethan, most recently on the xAI Groq team make a prediction that like the next tier in video gen is not a better video model, it's a better model or agent that orchestrates video models.Video Agents and Production WorkflowsAkshat [00:50:56]: Oh, interesting.Vibhu [00:50:56]: Language model backbone that can use toolsAkshat [00:50:58]: RightVibhu [00:50:59]: And write code.Akshat [00:51:00]: Like, yes, I can make my second video or my second video from Groq, but I want my minute video.Akshat [00:51:06]: And I'm not going there through normal video gen.Swyx [00:51:10]: Yeah, that's interesting. I - So we have GPU sandboxes and recently have seen a few companies doing agents that do video manipulation or,Akshat [00:51:22]: Yeah. Give it FFmpeg and just do it.Swyx [00:51:23]: Run FFmpeg. But likeAkshat [00:51:25]: That's not enough.Swyx [00:51:25]: Yeah.Akshat [00:51:26]: You need to give it Adobe.Swyx [00:51:27]: Yeah, I hadn't put it together with like it would be a video production thing. in my mind these things were going more towards editingAkshat [00:51:36]: Yeah.Vibhu [00:51:36]: Well, shout out Mantis.Akshat [00:51:37]: I think about this a lot.Swyx [00:51:38]: .Akshat [00:51:41]: Yeah. Sorry.Vibhu [00:51:41]: Luma. Luma Agent is a version of this for video production, but it's a off.Swyx [00:51:46]: I was gonna get your quick takes, on some other stuff that happensGitpod/Ona, CI, and Runtime SandboxesSwyx [00:51:50]: In recent news and just-just see if you have anything interesting. Gitpod, very li
“So, there are plenty of failure points, and when you have hundreds of thousands or millions of something, something will eventually fail,” Peter Salanki, co-founder and CTO of CoreWeave, tells Bloomberg Intelligence Senior Technology Analyst Anurag Rana. “Instead of throwing out half the potential capacity, we say that we expect some of these to fail. Then we build systems, automation and processes around handling those failures gracefully.” In this episode of Tech Disruptors, the pair discuss why AI infrastructure requires a fundamentally different architecture to traditional CPU-based cloud. Salanki also explains how CoreWeave is addressing training, inference and agentic workloads while navigating token costs, Nvidia chip demand and power constraints.
In this episode of Building Better Developers, Jim Hodapp and Bob Belderbos discuss why Rust continues to gain momentum among experienced developers. The conversation explores software craftsmanship, memory safety, AI-assisted development, and why language choice is becoming less important than understanding how software actually works. Key Discussion Points Why Rust attracted both systems programmers and Python developers The relationship between AI coding tools and strongly typed languages How Rust improves software reliability The importance of understanding software fundamentals Why developer growth often requires embracing discomfort The Rust Developer Mindset is not really about Rust. That may sound strange coming from two developers actively teaching the language, but one of the strongest themes from the discussion with Jim Hodapp and Bob Belderbos was that successful software development starts with understanding systems, not syntax. As AI generates code faster than ever, developers who understand architecture, performance, and reliability are becoming increasingly valuable. Rust simply happens to be one of the best environments for developing those skills. About our Guests Jim Hodapp Jim Hodapp is a veteran software engineer, engineering leader, and technical coach with deep roots in systems programming. His background spans C, C++, Linux, embedded systems, software architecture, and engineering management. In recent years, he has become a recognized Rust advocate, helping developers transition from traditional systems languages into modern, memory-safe development practices. Through RefactorCoach and his Rust training initiatives, Jim focuses on improving engineering effectiveness, software quality, and developer growth. Follow Jim on LinkedIn: https://www.linkedin.com/in/jim-hodapp/ Bob Belderbos Bob Belderbos is a software developer, educator, coach, and co-founder of PyBites. Originally coming from a finance background, Bob transitioned into software through automation, scripting, and Python development. He has spent years helping developers improve their coding skills through practical challenges, mentoring, and community-based learning. More recently, Bob has expanded his focus into Rust, combining his Python expertise with modern systems programming practices to help developers build faster, safer, and more maintainable software. Follow Bob on LinkedIn: https://www.linkedin.com/in/bbelderbos/ Why the Rust Developer Mindset Starts with Fundamentals Many developers begin their careers with languages that allow rapid progress. Python is an excellent example. Developers can create useful applications quickly, automate repetitive work, and see results almost immediately. That accessibility explains much of Python's popularity. The challenge appears later. The Rust Developer Mindset encourages developers to move beyond writing code that works and toward building systems that remain reliable over time. Great developers eventually become students of systems, not just programming languages. How Rust Forces Better Engineering Habits One reason both guests spoke so positively about Rust is that the language encourages deliberate thinking. Rust's ownership model, compiler checks, and strict type system often prevent entire categories of bugs before software ever runs. For developers accustomed to highly dynamic environments, this can feel restrictive at first. Eventually, however, the restrictions become guardrails. Instead of discovering issues in production, developers discover them during compilation. That shift changes how software gets built. The language rewards planning, understanding data flow, and thinking carefully about how components interact. Those are valuable skills regardless of which language a developer uses professionally. Rust Developer Mindset in the Age of AI One of the most interesting topics from the episode was AI-assisted development. A common assumption is that AI reduces the importance of programming expertise. The opposite may be true. Modern AI tools can generate large amounts of code rapidly. However, generated code still requires evaluation, validation, testing, and architectural oversight. Strongly typed languages create an interesting advantage. When AI generates imperfect code, the compiler immediately becomes part of the feedback loop. The compiler identifies errors, exposes assumptions, and forces corrections. This creates a collaborative cycle between the developer, AI, and compiler that often produces more reliable outcomes. The Rust Developer Mindset embraces this reality by treating AI as a productivity multiplier rather than a replacement for engineering judgment. Faster code generation does not eliminate the need for software design expertise. Learning Through Productive Friction Bob described his transition from Python to Rust as a challenge. That challenge turned out to be valuable. Many developers plateau because they remain inside familiar environments. They become highly productive but stop expanding their understanding. Learning Rust introduces concepts that many scripting languages intentionally hide: Ownership Borrowing Memory management Concurrency considerations Compiler-guided design These concepts can initially feel uncomfortable. Yet that discomfort often signals growth. Developers gain a deeper appreciation for what their software is doing beneath the surface. The result is not merely Rust knowledge. It is a broader engineering capability. Why Performance Still Matters The conversation also highlighted a topic that often gets overlooked in modern development. Performance still matters. Cloud resources may be abundant, but inefficient software still creates costs. Applications that consume excessive memory, waste CPU cycles, or scale poorly eventually impact users and businesses. Rust provides developers with low-level control while maintaining modern safety guarantees. This combination helps engineers build software that remains efficient without sacrificing maintainability. The Rust Developer Mindset recognizes that performance is not about optimization for its own sake. It is about creating software that respects resources and scales effectively. Identify one application you currently maintain and investigate where performance bottlenecks originate before attempting optimization. The Future Belongs to Software Engineers The strongest takeaway from the episode is that language debates are becoming less important. AI can help generate syntax. Documentation can explain APIs. Tutorials can teach frameworks. What remains difficult is understanding how systems behave. Developers who can reason about architecture, reliability, performance, and maintainability will continue to stand out regardless of tooling trends. That is ultimately what Rust helps reinforce. The future belongs to engineers who understand systems deeply enough to guide both AI and software toward better outcomes. Conclusion The Rust Developer Mindset is not simply about adopting a new language. It is about developing a stronger understanding of software itself. By encouraging developers to think more carefully about correctness, performance, and system behavior, Rust creates opportunities for long-term growth that extend far beyond any individual technology stack. Stay Connected: Join the Developreneur Community
Dead On Arrival: What was just Hype and what was just wrong Timing? Some technologies and products get over-hyped, and never seem to come to fruition? Some just come to fruition, but years or even decades later? In this episode of Tech Deciphered, what is Dead On Arrival versus just wrong timing Navigation: Intro The Fork Cold Open General Magic The Metaverse Dead on Arrival The soon to be Resurrected… Conclusion Our co-hosts: Bertrand Schmitt, Entrepreneur in Residence at Red River West, co-founder of App Annie / Data.ai, business angel, advisor to startups and VC funds, @bschmitt Nuno Goncalves Pedro, Investor, Managing Partner, Founder at Chamaeleon, @ngpedro Our show: Tech DECIPHERED brings you the Entrepreneur and Investor views on Big Tech, VC and Start-up news, opinion pieces and research. We decipher their meaning, and add inside knowledge and context. Being nerds, we also discuss the latest gadgets and pop culture news Subscribe To Our Podcast Bertrand SchmittIntroduction Welcome to Tech Deciphered episode 78, Dead on Arrival. Was it just hype or just wrong timing? In this episode, we are going to discuss two different types of businesses that as an investor, as an entrepreneur, you want to think carefully about. Looking back, you want to understand better what happened. If we take dead on arrival, we’re talking about businesses, technologies that maybe the tech work fine, but actually no one cared about it, no one wanted it. It’s really a mismatch in terms of what consumer want versus what you are trying to deliver. It’s a classic example of a solution finding for a problem. On the other hand, will it happen but not like that, not now? It’s a type of business where we were too early. We tried to make it work, but actually the technology is not ready now, but it might become ready in the future. As an entrepreneur, as an investor, obviously, you want to make sure you are not in one of these two technologies, one of these two categories. But it’s really a useful lesson to learn which situation you might be in because it will help you make better decision and pick better industries going forward. Nuno, what’s your take on this tough analysis?Nuno Goncalves PedroYeah, in hindsight, everything looks like, “Oh, of course, this wasn’t meant to work,” etc. But here we’re going to take a stab, as you mentioned, at things that are just wrong timing. They might have failed miserably in the past, but they might work in the future. The Fork To have this stake on the ground approach to it, as we’ve had in previous episodes, of defining exactly what we think will happen, but just not like that, versus the dead-on-arrival, which is a little bit more of a hardcore view on it. To your point, we want to really split between those two things because in some ways people are like, “Well, are we ever going to go and have different ways of interacting, for example, in terms of input and output, like with glasses or things like that, that really scale to become pretty pervasive?” We probably would agree that, “Yes, that will happen”. Maybe not in the way that we’ve anticipated until now, and certainly it’s going to take longer than we’ve expected. In some ways, the “will happen, not like that” is a little bit like Amara’s law or Nuno’s law, as referred to in the past. We tend to overestimate the speed at which we get to a revolution, underestimate impact. The “will happen, not like that” will probably fit into that. They will have probably bigger impact than we thought they would, but they will just take longer. Therefore, the solutions we’ve had until now are not great solutions. Whereas the dead-on-arrival, we’re saying, “Hey, we don’t think actually this works.” It’s to your point, not solving a problem that actually needs to be solved, or it’s solving it in the wrong way. Yeah, it’s basically an attempt to really frame, if you’re an entrepreneur, if you’re a venture capitalist, if you’re someone who is an angel investor and is looking at the market right now, what things should you put some real money and real resources, real time behind versus not, and really share that with you guys: our audience.Bertrand SchmittCold Open I always like to go back to a company called General Magic, a company that was legendary in Silicon Valley, spun out of Apple in the 1990s. There was even a movie made about General Magic. It’s a company that tried to basically build the first modern smartphone with touchscreens, app ecosystem. The only issue was it was years before the technology existed to support it. But Nokia at the time was less ambitious and more practical with what the technology was able to do at the time. But what’s interesting with this company is that obviously the smartphone was coming, just took 10 more years, 15 more years. Interestingly enough, some of the people in the team ended up being very successful in that exact space. Tony Fadell, for instance, who ended up creating the iPhone, co-create at least, and Andy Rubin, who end up building up Android, selling that to Google and leading the charge with Google Android.Nuno Goncalves PedroYeah, a very successful team. Tony Fadell, obviously, later on did Nest as well. Megan Smith, who was later the Chief Technology Officer of the US. A lot of people that came out of it that did a lot of important things. To be honest, I think a lot of the technology that was developed that was in some way a capacity or that are reused later on by some of these team members. I’m not saying they were infringing on any IP, but definitely there was a lot of reuse of tech. But it was too soon, right? They were doing something that was truly ambitious too soon. But a lot of these principles, a lot of these things ended up manifesting themselves, and particularly in the iPhone later on and in that deployment. Another such example is actually the metaverse. We did two episodes back in the day, a long time ago, on the metaverse. We debunked it a little bit. Back then, I think it was a very nuanced debunking, if I recall it correctly, Bertrand, because we were saying, “Hey, it’s going to happen in some ways, but it’s an evolution, not a revolution,” not in the way that we saw definitionally. We were huge fans of Matthew Ball and his analysis around media, but somehow we disagreed, I remember, with his definition of what the metaverse is and how it will look like. There’s no doubt that there’s going to be a metaverse in the future, probably metaverses in the future before there’s any given unification. In those metaverses, we do believe. But this whole hype cycle around metaverse and what’s happening clearly didn’t work out well, as proven also by all the write-offs done by Meta on their metaverse stuff. That much must have hurt. It’s related to AR and VR as well and to the enriched reality positioning of a lot of these technologies, which we’ll talk about later. But, again, it was probably something that’s going to happen. It’s going to happen likely in a different product manifestation as the one that was anticipated by a lot of the pundits that opined on what’s happening around metaverse in the future. But it’s going to happen. It’s just going to take longer. It’s going to probably come from angles of entry that are likely not what people had anticipated earlier on. Maybe it’s from gaming, maybe from something else. But definitely it’s going to happen.Bertrand SchmittYes, and you could argue that Apple did a similar mistake with the Vision Pro. I think the Vision Pro is the worst launch of an Apple product. When we say launch, it’s been three years maybe?Nuno Goncalves PedroSince the Newton, maybe?Bertrand SchmittYes, that’s my point. You have to go back a really long time to have such a disastrous launch. I think for me, it’s very interesting because it’s really a category where on one side there is a metaverse of some sort. You should argue if you are just using Twitter, or Facebook, these are some parallel universe in parallel to your real life. But the full on metaverse where you need some extra super heavy glasses to make it work, that simply has not worked. Don’t get me wrong, there are some super small, narrow cases where your heavy VR glasses makes sense for some training, for some simulation, for some light gaming. But given the weight, the disconnection from the real world, it has really never taken fire. Even at cheaper price like Meta was able to do, at more expensive price, much higher quality like the Vision Pro did. It looks like people have an issue with putting heavy glasses on top of them. We’ll talk later about the other side: AR glasses. But the traditional metaverse, heavy glasses, VR glasses, yes, it’s not feeling right, and I don’t think it will ever be right. It needs a different form factor to go at scale.Nuno Goncalves PedroDefinitely fundamental shifts on the metaverse. Just a quick reminder, Meta actually changed their name because of metaverse.Bertrand SchmittYes. Nuno Goncalves PedroIt’s not just that they’re writing off the investments, it’s their name is apparently wrong. It didn’t really happen in that way. You are saying that this persistent embodied virtual world, which was the early definitions of what a metaverse was, didn’t really happen until now. It might happen at some point, but not in that way. But, obviously, a lot of the enablement layers are still work in progress, right? Real-time 3D, presence, avatars, all of these things are moving on. A lot of these enablers will be ready once there is a time where these technologies and businesses that are fundamentally anchored around these technologies will emerge. It’s, again, right components, wrong product alignment, but at the same time, it will happen, just not yet.Bertrand SchmittAgain, I think we’re going to face a split of a virtual reality light that you can call AR with very light glasses that are not a pain to wear, that are more limited, especially today or the next 5 years. The VR glasses that will stay a niche product that are useful for some use cases. Do you want to do a flight simulation? You want to do military training? There will be use for that. But beyond that, it’s more limited. We forgot to talk about the Google Glass. This one was an early, early AR glasses. But, again, here we go back to the capabilities. Simply, we were not there at the time.Nuno Goncalves PedroDead on Arrival We’re moving to Dead on Arrival. What things do we think are relatively dead on arrival? We’ll be nuanced on a couple of these aspects. Not everything is totally dead on arrival. But maybe let’s start with a good one, which is some of the crypto summer promises, particularly around NFTs, non-fungible tokens, and the decentralization of everything, the tokenization of everything. That didn’t quite work the way people thought it would. A lot of people were very happy for a period of time that they had some NFT that was worth a ton of money. Good for them. I hope they sold it back then. Some did actually, fortunately for them, but most of them did not. As an asset and the tokenization element, we’re not questioning that tokenization itself doesn’t make any sense. But from my perspective, a lot of the elements of tokenization plus the tokenization of everything don’t make necessarily imminent sense. I think we’re going to have a world that not everything is going to be tokenized, and a world where, to be honest, having the latest ape token as an NFT is maybe not something I really need. It’s dead on arrival in some ways. What do you think, Bertrand?Bertrand SchmittIt feels like it’s a tale of two worlds. On one side, there was clearly exuberance and bullshit. All these cryptocurrencies that make no sense, have no values. The NFTs are the same: no values, no sense, no logic in them. At the same time, we can see that Bitcoin, Ethereum, and few others were significant products, invention, that have a current total market value that is significant. I would say we can see that step by step, little by little, there is some tokenization happening, but it has been much more narrow than what was initially thought, expected. But it’s happening step by step. I think there is value in tokenization, but definitely not as much value as people expected. I think for one simple reason: it’s very expensive to set up in the sense that it doesn’t work so well, actually. It requires a lot of computing power for stuff that might not justify it as simple as that. Why do you want to replace a very efficient, superfast database by something slow, expensive to maintain? Why? You need to have real reason for that. I think in most situations, there was not a real justification to make that switch. I think, again, there is a niche for tokenization. It’s getting used. Obviously, there is stablecoin as well that are working and provide value. But definitely, the use cases have narrowed. The success stories have been few and far between. It’s definitely a space, but I don’t think it has lived up to the expectations. The expectations were simply quite insane in terms of what to expect, and it failed to achieve that. I don’t think it qualifies as a true revolution. It has been successful again in our situation: in Bitcoin, in Ethereum, in some tokenization. But beyond that, it was not the dramatic change that people were picturing.Nuno Goncalves PedroYeah, it’s been successful and there is a security, in this case, a currency of types, that links to the value of that specific network and the utility of that specific network. But on this case, in particular, the NFT stuff, I think there were two core premises that were very valuable about it. One was the premise of digital collectibles, the element itself being collectible and the fact that I have a unique one was only minted one. It’s a unique thing out there. I think at some point people forgot that actually digital collectibles are not as valuable as people think they are. Physical collectibles are more valuable. The reason for that is just human nature. I want to have something in physicality. I want to have something that I can show people, etc, that has a physicality to it. Digital collectibles, I think there was a problem fundamentally in part of the articulation of NFTs that reside in the value of digitization of these collectibles in and of themselves. I think that was part of the problem. The second problem was really the notion of provenance, the notion of certification, the notion of “this is what it says it is, and nobody can say this is something else.” This notion of provenance. That, I think, fits more into the point you were making, which is the point around use cases. This doesn’t fit all use cases. The provenance piece and the certification piece doesn’t fit for all use cases. Probably it might fit actually even more for actual physical assets where you want to have a certificate, that thing is what it says it is, but it’s actually maybe a physical asset rather than a digital asset. Again, lo and behold, it doesn’t apply to all the use cases that relate to provenance. I don’t think we need to tokenize all the provenance ledgers in the world for this to be working well. Again, it was a little bit of an overreach. In some ways, all these players were very happy. There was a lot of bullshit. To be honest, this is an area that I do think there was actual active fraud. All due respect to a lot of the players in the market, but there was active fraud. Honestly, there was bullshit.Bertrand SchmittI agree with you. I was going to say that it’s one of the rare sector in technology where I think frauds, scams were prevalent. Yes, you might have some shitty startups in general and stuff, and there might be fraud. I think true fraud is rare. True scam is rare. But here for NFTs, for cryptocurrencies, that was the majority you could argue. That’s the part that I never really liked in all that industry. If you look at NFT, personally, I never understood the concept. For me, it made simply no sense because, yes, you can say, “Oh, I have created this rare digital token.” But sure, so what? Who cares? Who cares? You can have hundreds, thousands, millions of unique digital tokens. How does it matter? What’s the use case? What’s the connection? You can show it on your screen, but then how is it unique anymore? It’s not like people have to come to your house and check stuff on site or something like physical art. No, not at all. I don’t know. I think it was surprisingly true bullshit. Again, I’m still a believer in Bitcoin. I think there is true logic to have a separate virtual digital currency. But you don’t need hundreds of them. You just need one, maybe two for some variation. The same for tokenization, it doesn’t need to happen everywhere. Just need to happen where there is really strong real logical value. Beyond that, there is no point.Nuno Goncalves PedroYeah, makes sense. Maybe moving to our next one. I think we discussed this in a previous episode, which was the whole 3D TV stuff. We just don’t think that’s the product, right? It’s not going to be a 3D TV. I believe when we discussed this in the episode—I don’t want to lie too much—but I believe that’s what we came up with… It looks a bit more holographic or the manifestation comes through other mechanisms like glasses, etc. But we thought it was just dead on arrival. Nobody wanted to have glasses on the couch to watch the 3D TV, but also the 3D TVs themselves they tried to develop were very holographic, were super clunky and very heavy. In some ways, I think TV manifestation somehow is not a great manifestation for this. I put it on the dead-on-arrival pile. I have one, just to be clear. I still have one.Bertrand SchmittYeah, but it’s a weird one because if you remember, it suddenly happened. I remember being shocked as a consumer. Suddenly everyone talks about 3D TV. In a few months, 3D TVs are all out everywhere. Then after a few years, no one bothered to keep adding that because consumers didn’t care. I think one issue was, I think it has only worked with Blu-rays. I don’t think any streaming service supported that. With the rise of streaming, that was certainly a pretty big issue in the first place. Interestingly, in the movie theaters, do we still have movies being shown in 3D? I think we still have a few, but it’s really few and far between the movie theater.Nuno Goncalves PedroYeah, IMAX won on that format, the high-end format, IMAX won, not the 3D thing.Bertrand SchmittBut, yeah, if you’re on the TV at home, you don’t want to bother. Maybe if you’re in the movie theater, more high-end experience, you are willing to wear special glasses. I think it makes sense. I think, I guess at home, we are just too busy doing multiple things at once. Rarely sitting all around 2 hours for a movie, so you watch your phone, you check on something else. You don’t want to have these glasses on top that block you from the rest of what’s happening at home.Nuno Goncalves PedroYeah, it just didn’t work. I’m not sure it will work in any case, so we’ll see. But definitely, I don’t think that’s the right form factor.Bertrand SchmittIt technically worked, but practically, it was not accepted by consumers.Nuno Goncalves PedroLet me rephrase. I think it was a solution to a problem nobody thought they had. The execution of the product was the wrong one. It’s like that’s not what people wanted. Maybe we do want some 3D immersive experience, but that’s not the 3D immersive experience that we want, not manifested like that in effect, with its limitations, etc.Bertrand SchmittI want to say that we are not denigrating some of these experiments. Me, I still remember putting a Vision Pro and watching some content that is natively created by Apple for the Vision Pro. It’s an amazing magical experience, truly. But at the same time, I’m always worrying to have to think I’m going to spend 30 minutes putting my Vision Pro on my head. It’s heavy, it’s painful. Again, we go back to, I’m disconnected from everything else, and I don’t like that.Nuno Goncalves PedroWhen there was real technological development, like the case of 3D TVs, that technology development may have served other purposes. Again, we applaud those that put the money where their mouth was, the companies, the individuals that wanted to move forward some of these technology stacks, maybe less so the people that just decided to create NFTs for the sake of it with apes. No attack on those guys specifically, but just to create a specific view on that. But there’s a lot of stuff we’re talking about today which is actually great technological development. There’s stuff that came out of it that is very, very valuable technological development. Maybe to our next one on mobility, physical mobility. The Segway. Everyone’s like, “Oh, this is going to change how we move around the world, in particular for shorter distances, etc.” Well, apparently no. Apparently, it’s not super intuitive. There are still Segways out there for use by police forces, tours, and all that stuff. People say, “Oh, it’s…” No, well, it never scaled. It wasn’t changed. It didn’t change how relatively short local mobility was going to be done. I think part of it is actually the problem with the product. The product is not as intuitive as people thought it would be, and it has other issues around that in terms of security and safety. If you’re going to have a device or something that moves you around, what’s the minimum-accepted level of security that you would have. The fact that you have to incline to move it is actually not super as initially some probably people thought it’s actually not as intuitive as some might have thought it was. I think it’s a dead-on-arrival play. The vehicles of the future, I’m not sure, look like that at all. Happy to eat my words down the road, but I don’t think so.Bertrand SchmittYes, and I remember the massive hype at the start of this product. They were hiding it, but still telling you, We have seen it. That’s true. It’s going to totally change how cities are planned. I mean, the level of hype and bullshit was insane before the launch of the product and before anyone outside of a few knew what it looked like. Then we saw what it looked like, and it was on one side, an engineering marvel, the self-stabilization system was new. I mean, very few people had seen something like this before. At the same time, as you said, the issue was what happens if it loses battery? What happens if you fall from it? Quite frankly, if you look at today’s modern electric scooter, you have 95% of the value of a Segway without any of the downsides. You don’t need to learn a new way to do things. If you lose battery, it still works. I guess you can still push it. It’s not too expensive. You don’t have massive computers to just make it work. You don’t have massive robotics. It’s clearly an impressive engineering solution. But the exact example of a solution to a problem people didn’t have, when ultimately a much more simple and elegant solution was there for people who care. To be clear, you still could use a regular bike. The alternative was still your regular bike. But now you have electric scooters that are a clear, closer alternative, I guess, in spirit to what was supposed to be the Segway. Sometimes, let’s be careful. Overengineering might not be the answer.Nuno Goncalves PedroEven scooters, which are, to be honest, I would argue a much easier philosophical view of how mobility should be done and a much easier form factor, so to speak, in terms of mobility, even those won’t have wide adoption to everyone. It’s not like everyone’s going to use a scooter. The bar for a scooter is much less than a bar for a Segway.Bertrand SchmittYes. I think in a way, the electric scooter is showing us in some ways that the limit of the dream is that even well-executed, even cheap, even easy to use, even low-risk in a way, there is still only so much that people are willing to use it because it’s a pain to bring with you everywhere. Actually, in some ways, you could argue it has only worked with this distributed platform, like lime and others that let you take one and drop it when you are done using it. But that was not a model that was workable for the Segway. I totally agree. You could argue that bikes are still the gold standard or now electric bikes are still the gold standard for who want that level of mobility.Nuno Goncalves PedroWell, the next one is to be a favorite of, I guess, everyone that didn’t invest in them, which is Juicero, which was the machine that basically created Juices out of squeezing a bag. I’m going to be nice here to… There were a bunch of well-known investors. They raised, I think, 120 million or something like that.Bertrand SchmittThat’s insane.Nuno Goncalves PedroIt is insane. But maybe the vision was not that vision. It was something grandiose and the complexity around dealing with liquids and whatever. I’m not really sure we didn’t pass or invest on the company.Bertrand SchmittBig old juice bags.Nuno Goncalves PedroMaybe. You’re being facetious. Anyway, I’m not trying to attack the guys who actually invested in this, the investors that put money into this. But it was basically one of these examples that we’re trying to solve a problem that didn’t exist. It’s like a solution for a problem that doesn’t exist. I can just squeeze a bag of concentrated thing and make a juice. I can do something around the juice. As I said, maybe the big vision was well beyond that, and I’m not sure many of us will ever know what the big vision actually was. But yes, it was a solution, a $400 machine to a problem that nobody really had. That didn’t work.Bertrand SchmittI think the most crazy part of the story is that the business exploded when there was this Bloomberg reporter that showed that you could squeeze the bag with your hands without a machine and get the same juice. That’s pretty insane. That’s just so unbelievable.Nuno Goncalves PedroThe funny part of the story is MSCHF, this company that does all their on-point, slightly artistic, whatever, one-off, drops, etc, etc. Shout out to Gabriel. Gabe Wally, who was the founder there. This is funny. They actually did a collectible series of figurines of that startup toys, of which Juicero was one of them. Jibo, which I don’t think you’d appreciate. Bertrand, and the Theranos minilab. So someone is also making money on that, which is funny. So well done, MSCHF.Bertrand SchmittWhy not? Why not? Maybe we can go quickly over Lytro. This one was, I would say, brilliant technology. It was a camera able to capture light in a different way. You could actually dynamically change focus after the picture was taken. So something quite impressive, and I still remember looking at their products and thinking, should I get one? At the same time where it didn’t work, it was too limited in terms of image quality versus shooting the real picture. It was very tough to sell as an independent device because either you’re in the phone or you are in a DSLR, but there was no spot in between. In some ways, phones solved the issue not directly, but they kept adding cameras. At some point, you had not one, but two, but three, but four, but five cameras in your phone. You had a specific mode for anything you needed, and that’s it. Phone manufacturers managed to convince people they should pay for all these cameras, or provide a low-end version if you don’t want to pay for all these cameras. But yes, very interesting technology that did not manage to find a go-to market.Nuno Goncalves PedroYes, it’s falling in love with a technology solution that, as you said, was incredible, and not falling in love with what problem you are actually trying to solve. Because a lot of the elements of the problem that you’re trying to solve, either you take another picture. All due respect, instead of me trying to change the field of view, etc, and the focus, post picture taken, I take another picture. To be honest, you can take thousands of pictures now with different setups. Also, ignoring the fact that this was pre this whole AI boom. So obviously, that one I will be kind on. But with the AI boom, you can do a lot of things to pictures and changes and stuff like that. All of that changes the game in and of itself. Obviously, Lytro or Lytro, not sure how you spell them, or how you say their name, but they failed before that. So again, this was super great tech, great investors, good investors on board, but just nobody wanted that device. Why would I want this thing? A total failure on arrival. I think one also other aspect of it that may have been misinterpreted and underestimated by some of the early investors, certainly in the company, is the value chain. There is a value chain for cameras, for SLRs, DSLRs, etc, etc. They know what they’re doing, and they’re doing their own level of innovation. In some ways, that fits squarely into that value chain, which I’m sure is a complex value chain to manage in and of itself. Maybe that’s created its own dynamics as well.Bertrand SchmittI think they try to insert themselves in the value chain, but I’m not sure first that they managed to get technology as small as it should have been. They simply didn’t convince anyone. The Sony or Nikon of the world, or the Apple or Google of the world, to use their technology in their phones, because no one really saw the value, per se. One thing to keep in mind, using their technology means you had a big trade-off in terms of megapixels. Actually, you will benefit from that great effect around the focus that can move anywhere after the fact. But in exchange, you had to lose dramatically the number of megapixels. At the time, all the rage was about how do I get more megapixels for my phone with good image quality. That was all that you were selling to consumers. Consumers would not have bought 10 times less megapixels for a weird benefit that at the time no one cared about. I think you are very right to introduce the fact that, of course, they could not have guessed at the time, but today you can get for free this refocusing effect with AI, actually.Nuno Goncalves PedroThat’s exactly the point I was making on the value chain. They went into a value chain that has a lot of players in it, and they got kicked in the butt anyway, because people were like, Why would I adopt it? They didn’t serve the purpose of the use case or the user flow. They didn’t anticipate competitive dynamics to it. Then they actually went straight up into a technology plus IP logic on a value chain that would be like, “Hey, dudes, just get out of here.” What would be the incentive for people to bring them on board? It’s like in some ways, they did all the mistakes under the book in that sense. Again, I’m not trying to diss the company, the founders, the investors, but it feels that was the problem in the end.Bertrand SchmittYes.Nuno Goncalves PedroShall we do some rapid fire dead-on-arrivals?Bertrand SchmittYes. A fan favorite. Have you tried a keyboard projector on?Nuno Goncalves PedroYes. I think I mentioned that in a previous episode, I did have one of the early ones and whatever. Yes, somehow we forgot that there’s other best, better ways of input, maybe voice would be a better way of input if that’s the problem, but anyway.Bertrand SchmittThis one is interesting because you just use it for one minute, and you discover it’s horrible.Nuno Goncalves PedroIt’s a horrible experience.Bertrand SchmittYou have to type on it. You don’t see exactly where you type. Each time you type, you are hurting your fingers because it’s not a smooth keyboard feeling, but you are tapping a solid surface. That’s amazing that it went to manufacturing, and basically no one gave feedback. There is no way it’s working. No way technically working, but practically, no one would want to use that.Nuno Goncalves PedroIt’s bad in all aspects. The finger’s touching. It needs to be a very over-engineered experience for it to actually detect the key stroking and all that stuff. It’s an over-engineered solution by default. It’s not private because once you have the keyboard projector, you’re like, “Where am I keyboarding projecting on?” It’s something that others can see. There are other alternatives that are better. The keyboards on our phones are pretty good these days, with autocorrection in most cases, I wouldn’t say all cases. Voice can also be an interesting replacement to that. If that’s what you’re going for, voice recognition has improved dramatically over the last few years. Again, it’s one of these, “Yeah, cool.”Bertrand SchmittIt’s clear that, first, touch has replaced a physical keyboard where you need one. And two, now voice recognition is getting better for sure.Nuno Goncalves PedroThe opposite side is that voice assistants were going to take over everything, that everything was going to be voice assistant-led. Alexa is going to run your life. You’re going to have games for voice. I actually have a good friend of mine who did a startup in that space. Actually, the company got acquired, and everything’s going to be voice, which, as we know, is also not true. We need to visualize sometimes. Sometimes we still want to write because we want either the added value of privacy or of expression at some standpoint. There are a lot of things that are happening around voice that are exciting. We just mentioned voice recognition, but not everything is going to be voice. Certainly not for the foreseeable future, and so that doesn’t work. Also, as the ultimate platform, we also want some visuals. We also want to have other kinds of interactions. There are limitations to voice assistants, for example. Voice assistants were not quite the next platform that people said they were going to be. I would say they’re a channel, but they’re definitely not a platform. That’s how, at least I see them.Bertrand SchmittYes, a channel versus a platform. I would agree with you. I must say I start to feel that we could be close to the future we have seen in Her. This movie, Her, I guess you saw it as well, where these guy keep talking to his personal assistants, and it’s becoming his friend, maybe even his love. I start to feel we are getting close enough with AI that this becomes a possibility. But again, it’s not your only way to interact. You want to type, you want to see stuff. There are many situations where you don’t want to be seen talking with someone. It’s more of a channel versus a platform. Yes, Alexa, it’s quite amazing how it went to the roof and crashed and burned. It moved up very quickly. Amazon invested billions, hired thousands of people, which was the right move. I mean, you have something that seems to work well enough and could be game changer, why not? But it was quite impressive as a crash and burn. It’s pretty rare to move that fast up and down.Nuno Goncalves PedroI agree. A lot of money is put into this stuff. I mean, Google went after it as well. I think Google probably executed better. Their tools and their hub tools, etc, were better in their devices. But to be honest, it was the beginning of something, it wasn’t really there. It’s shocking because a lot of people have Alexa at home, and Google Hubs at home, and whatever at home. But it’s really not that powerful. I have a bunch of Google stuff at home, and I still use it for very basic information. It’s not really doing much for me. It’s not like magically now with AI, it’s going to be much better. It’s not. It’s a legacy device now. Maybe in the future, we’ll have some stuff that will be cool, but definitely not the voice assistant manifestation that we saw in the past. I know the next is exciting to you, so I’ll let you go for it, Bertrand.Bertrand SchmittNetbooks. If you remember Netbooks, this was a horribly underpowered PC device. I kept buying one after the other, hoping that the next one would be fast enough. Trying to be a small, lightweight portable PC, they were called netbooks. Usually, they also had a small screen, like 10, 12-inch, 9-inch, and each time it was horrible. Very poor battery life, 2-3 hours, super, super, super slow. They were equipped typically with a full Windows OS, and they were killed overnight by the iPad. After 2 years of the iPad, no more netbooks. Oh, yes, and they had this… Typically, they would be equipped with… What was this name of the CPU? It was an Intel Atom CPU, I believe, or Celeron. Again, the root cause why, it was so slow. It was horribly slow.Nuno Goncalves PedroSo netbooks didn’t work. We didn’t want to go back to that. The iPad won and all [inaudible 00:34:05].Bertrand SchmittI think it has been proven at this stage. There are some small mini PCs these days. You have, for instance, the ROG Ally for gaming. There has been some revival with some Chinese manufacturers making some small PCs. I have to admit, I bought some, so I kept trying. It’s much more powerful, no question. Better battery life also. But we go back to the fact that if you have a very limited screen size, 7, 8, 10 inch, it’s just tough to put a full desktop operating system to good use. You need something more simplified, more touch friendly for it to be really practical. Yes, you can run the full desktop Excel on it, but will you do it on an 8-inch screen? Probably not.Nuno Goncalves PedroNetbooks, and maybe we’ll end with Google Glass. I mean, obviously, we’re going to talk about AR and VR in just a bit. I think it’s a good segue to our next section, and we’re not necessarily saying AR, VR is dead on arrival. We actually think they’re going to be out there for us in the future. But Google Glass, in the way that a lot of the first wave wearables were developed, was cool, but not enough. I think even the recent ones, the Meta Ray-Bans. I see a lot of friends of mine now, you’re wearing Meta Ray-Bans, and they’re like, “Oh, this is so cool”. It’s like, “How is that different from the Google Glass, but now a little bit nicer?” It still feels to me that’s the wrong way to approach it. Either we go with a more comprehensive solution for glasses or a solution for glasses that’s a little bit more easily integrated into existing frames, etc, for those of us, like the two of us, who have actual eyeglasses, or I don’t see it quite happening this way. I think just as a formality in how it was deployed, it seems to me it’s been the wrong angle a couple of times, even, to be honest, beyond the first wave of wearables. It’s been a little bit the wrong way to approach this, the wrong paradigm.Bertrand SchmittIf you remember the Google Glass, initially, they were launched in… Again, overhyped. You couldn’t even buy them, but still, it was overhyped. You could see a marketing department going crazy when they were showcasing models wearing Google Glass, that sort of stuff, fashion shows. That was insane. How is it connected to the value proposition of the Google Glass? Absolutely not. But you remember what killed the Google Glass?Nuno Goncalves PedroThis is the Robert Scoble in the shower thing. The whole lack of privacy thing.Bertrand SchmittYou tried to have this vision of the Google Glass that models are wearing and stuff in fashion shows. You have this guy, Robert Scoble, taking his picture with his Google Glass in the show, half-naked. Oh, wow, that quickly dropped. In the interest into Google Glass.Nuno Goncalves PedroRobert’s a good guy for the most part, so we’ll let him get away with it. But yes, it was not a good moment for the use of Google Glass. But to your point, it was overhyped. I don’t even recall that I had to do so many things to actually get access to one. I still have it. Then in the end, it was just totally underwhelming. I do think the meta Ray-Ban solutions, etc, etc, now they’re doing it with other brands as well, are more elegant because it’s around voice activation and stuff like that. But still, it’s not this. The soon to be Resurrected… I think these kinds of solutions are dead on arrival, which may be a good segue to the resurrection pile, so to speak, that will happen, just not like that pile that we have to share today.Bertrand SchmittYes. That’s the question. Are we in resurrection mode with AR glasses? The Ray-Ban, Oakley glasses, are the way to go? There are also some other brands like XREAL, for instance, that are a different approach. It’s more fun to watch movies that I personally like a lot and use a lot. I know quite a few fans for this. There is a market for either the XREAL for movies, meta, or Ray-Ban, for light information on top of your glasses. There are niches. I don’t know if it’s full resurrection mode yet, but I can see again, with AI, that there is the ability to do more and more. Will it be connected to your phone, will be independent from a phone? That’s also stuff to see. I guess at this stage, connection to a phone has some value. It’s making a comeback.Nuno Goncalves PedroI think it’s suffering from being overhyped. It makes sense that the category comes back. The use of glasses, a lot of people actually have eyeglasses. The use of different mechanisms that really expand your view. We’ve seen that with the advent of foldables in mobile phones. People want bigger screens. They want to get as close as they can to the tablet experience without necessarily having to go to laptops, etc. Definitely, I think the use cases make sense. It’s just how do you deploy it, and who’s going to deploy it best? I don’t think we’ve seen anyone that’s been particularly enlightened about doing it. All the great developments with Android XR from Google, Samsung doing a bunch of things in terms of deployment out there as well at the same time. I mean, are these guys going to redefine this category and create the first ever wearable, at least AR great experience or VR category of experience, or are we going to have much of the same? In the same way, it took several years for the smartphone category to be recategorized as a smartphone category under the iPhone. That was the category-defining moment. It was the vision of a bunch of people, obviously led by Steve, but a bunch of people that went to market and said, This is how we view the market to be, and they just nailed it. We haven’t seen the guys who’ve nailed it, I think, is the issue thus far, either on the VR side or the AR side. We just haven’t seen it.Bertrand SchmittDefinitely, we have not seen that. There are, of course, rumors that Apple is ending their Vision Pro program, and at the same time, is full-on on AR glasses, an alternative to the Meta Ray-Ban. We will see. I start to think I have trouble to see Apple pull this off because it starts to feel that, yes, they are great at making more efficient supply chain these days, but I feel after 10 years plus of Tim Cook leading the business, that’s all there is to Apple these days: thinner phones, pink colors for your latest device.Nuno Goncalves PedroVariations.Bertrand SchmittMore efficient chips, for sure. They have great chips. But beyond that, they canceled their car program, the Vision Pro was not successful in the marketplace. They have trouble to find the next story and to execute on the next story, quite frankly. You could argue that their most successful recent launch has been the AirPods.Nuno Goncalves PedroI think that’s absolutely spot on. We’ve discussed it in previous episodes how much… It’s taken much longer than I thought it would. I always said after Steve’s passing, that it would take maybe 5 to 10 years for the beginning of the demise of Apple, and it is taking longer. There was definitely some gravity created. Steve passed away in 2011, right? They’ve been hanging on to their time, so it’s now 15 years. It seems that only now Apple is really suffering in terms of product lines. Probably they started 1 or 2 years ago. They have a new CEO coming in September, so we’ll see if John will do a better job, if he’s really a product guy, and that’s how he’s going to go for it. He’s going to go for a few products really nicely done that can have great experiences, which is Steve’s hallmark, instead of what you said was the milking of the cow, so to speak, under Tim Cook. There’s nothing wrong with it created in a hugely valuable company with a huge amount of still innovations around the edges, but it’s like now they need real product innovation, a new category definition, play, and is John going to be the guy for that or not?Bertrand SchmittYes. Positively, Tim increased market cap of Apple. He didn’t destroy the business. We are not trying to tell that Tim Cook was a failure, but definitely it was not Steve Jobs. It was more a caretaker, optimizer type of CEO. But to go to the next level, you need another type of CEO. Quite frankly, it’s hard to find when it’s not a founder CEO.Nuno Goncalves PedroYeah, I think it will be difficult to find, but we’ll give John a chance once his time comes in September 2026, as one would. Moving to another topic that we’ve discussed in the past, self-driving. It’s going to happen. We all are waiting for it to happen. Self-driving cars, self-driving vehicles that we don’t need to drive in this mess of traffic, in particular spending a lot of my time in Southern California where traffic is incessant. The driving is… How can I put it in a nice way? Awful. People don’t seem to know what they’re doing. I think we all want to have self-driving cars. It’s taking longer than we thought it would, but I do think it’s going to happen. I wouldn’t put a mark on the sand on how long to get this world of self-driving cars. It’s going to coexist with non-self-driving cars for sure for a period of time. But I, for one, know deeply that it will happen, and I look forward to it. I love driving, but not in horrible traffic with horrible drivers.Bertrand SchmittWe are for sure very close. I mean, now, Tesla self-driving is extremely good. I have taken Waymo many times with great experience. It’s not going too fast, it’s not going on the highway, so there are some limits. But I would say really impressive all in all. I think we are pretty close. I don’t know if you know, but Waymo is working on expanding to many, many more cities, actually. I think we might be at a tipping point there. Quite frankly, I expect that teenagers in 10 years from now are not going to bother to learn driving. I mean, we’ll see, but I think many won’t.Nuno Goncalves PedroI mean, Waymo is now doing also stuff on freeways, certainly in the Bay Area. I had a friend of mine who took one recently, and he was scared. He was scared shitless—I think that would be the right way to put it—because the car wasn’t going that fast, people are aggressive on the freeway around your car. But to your point, I’ve used it in city. In the city, it’s great, and it’s aggressive enough. That’s one of the things I was a bit afraid about Waymo, that it wasn’t going to be aggressive enough in city driving, it is pretty aggressive. Yeah, maybe we’re closer than we think. Maybe it’s a couple of years rather than a decade, which is normally where we put the stake on the ground. But it’s definitely going to happen. I think there was an overhype too soon. There were great movements around level 3 and level 4 autonomy that in some ways were misconstrued early on as full autonomy. But now we will likely get what we ask for. Then at some point someone’s saying, “Hey, but why don’t we just have the flying objects instead?” We’ll get back to that. The flying car stuff, the eVTOLs and all that stuff, we’ll see.Bertrand Schmitt10 years ago, I remember the overhyping was very, very strong with self-driving. I nearly believed the bullshit. It was so strong in 2015, 2016 that nearly everyone thought it’s here in the next 12 months, basically. I think that was a lesson for quite a lot of people. Don’t trust the Silicon Valley bullshit machine, hype machine. It can go way too strong, way too fast. You absolutely need to spend time to understand realistically where we truly are. You cannot just buy the premise.Nuno Goncalves PedroMoving maybe to robotics and humanoids. Again, it’s going to happen both for consumer and for industrial. The humanoid form, I think, at some point will be important. Just not yet. I think we’re behind now the big buzz sentence is embodied AI, where you have physical AI manifestations. AI manifested through robots and other kinds of devices. It’s going to start happening in certain beachheads. But, hey, we’ve had this offer for a long time, and I know you still have your frustrations, Bertrand, with the Sony AIBO. I won’t traumatize you more than that, but definitely it will happen. We will have humanoids at some point. We’ll have different specialty devices that will help us in activities that we do both as consumers and also in the B2B and enterprise space, even in the industrial space. The writing’s on the wall that definitely is going to happen.Bertrand SchmittYes. I think the question, as always, is how do you get there? I think the intermediate steps of your optimized robot, a physically optimized robot, like a Roomba to clean your home or some other robots to mow your grass, a good example, I think initially you have to start a small, practical, not too technology advanced and progress your way. Truly humanoid robots are very disruptive because technically, if they work, they can do everything a human could do. But practically, I don’t think it’s that true because a lot of constraints. You have to have enough weight, you have to have enough force, you have to have enough vision, you have to have enough AI to just understand commands and move around. But, again, with all the progress in AI, the progress in electric motors, I am optimistic that there is a chance for the humanoid robot, but it’s not 3 years. It’s probably not even 5. It’s probably more like 10. Let’s not forget that a lot of demos with humanoid robots today are fake. It’s a remote operator, remotely controlling a robot. It’s not a robot working by himself in many, many, many situations. That, I think, is another thing hurting the reputation of some of these companies because these fake demos, they look great, but once you dig a bit, you realize it was a fake demo because it certainly didn’t come with a warning that it’s remotely operated.Nuno Goncalves PedroMoving maybe to the next one that it will happen, but not exactly like that. Quantum computing, I think for the most, it’s been in many cases a solution looking for problems. It’s beautiful, but it’s like, “What’s the problems we’re trying to solve?” I think we’re going to start having beachheads into problem solution and apps. I’m sure a lot of people out there would say, “Well, there’s already killer apps for this, a lot of the scientific modulation, cryptography, breaking, all that stuff.” But I’m like, “Yeah, cool.” It still feels very much like a bunch of pieces and hardware in effect that is trying to solve a problem that needs to be very much defined into several problems, several areas of problem solution going forward. Will it happen? For sure. I’ve been hearing about quantum computing since I was in college, since I was 17 when I got into college. I’m sure it will happen. It will have manifestations. There’s a lot of stuff probably already residing on some of these instances. But a full mainstream view of what quantum computing can do for us is still quite ways out.Bertrand SchmittYes, I mean, there has been some reports that some key steps in the technology starts to be working and that potentially the technology could start to deliver results in the coming 2-3 years. What I mean by results, results that will dramatically change how you need to encrypt because a lot about quantum computing is actually trying to decrypt the current technologies that are used to protect your information online. This one is big. If we manage to have a limited use of quantum computing things that work in a few years, that will have a dramatic impact on how we encrypt data. That’s the one piece I’m looking for: do we really manage to solve encryption with computing? If yes, then we need to make a lot of changes how we encrypt things.Nuno Goncalves PedroFusion energy, the gift that keeps on giving, but this somehow doesn’t fully happen. I know you’re a big fan.Bertrand SchmittI would say, practically, I’m a big fan of nuclear fusion energy because in a way, it’s working. It’s solved in terms of safety. Latest Gen5 reactors are ultra-safe. They cannot go out of control anymore. China has a working reactor. In some ways, part of me would say, “Why bother with fusion when we have fission that is working extremely well, and we know how to make?” If we are efficient about the process, we could make for really cheap. Having a better source of energy today will already be great for the world. I would like to see scale with fission. It’s great to research fusion. Fusion is near a limitless form of energy. This is exciting. I think what’s interesting with fusion these days, it’s not just the big governments working on spending billions before seeing any potential returns, but there is a lot of startups that are working on it. The bar went lower in terms of researching fusion. That part gets me excited. But again, I think we should go fission first and parallel with some fission investment, fusion research, and hopefully, fusion comes at some point, but we should not wait for fusion would be my big take.Nuno Goncalves PedroGood point. Last but not the least, we’ve dedicated episodes on this, so maybe today we’ll just do the little teaser. If you guys haven’t heard our episodes on AI and the bubble and all those things. The last big one that’s just not like that, it will happen is AI itself, and in particular, AGI, so generalized AI that will be just like a human. Obviously, it’s happening. We’ve discussed in previous episodes, there’s a lot of breakthroughs in terms of methodological approach beyond the GPT curse of actual hallucinations that doesn’t seem to be solvable. There’s a lot of things happening around reinforcement learning, evolutionary computation, and other methodological approaches to AI that we think will create tremendous breakthroughs. There’s definitely a lot of funding going into it. We quantified it recently. I think 63, the number probably has changed in the last couple of weeks, but 63 new AI labs that have raised significant rounds upfront, just getting into the market, some of them as high as hundreds of millions or even over a billion dollars for the first round of funding. Obviously, there’s a lot of dramatic things happening in that space. There’s a lot of capital going to that space. It’s going to happen. It’s going to take much longer than, I think, we think it’s going to take. I just saw, I think, Marc Andreessen recently saying, AGI is already upon us. Maybe it is. Maybe he knows something we don’t know. But I suspect we’re still ways to go in a lot of the developments on AI. For now, it will continue its complications and its limitations. But definitely it will untap a lot of great productivity enhancements, technological enhancements, even in the short term that at least I’m very excited about. Then the AGI story, I think, will be a story that will be told later.Bertrand SchmittI might actually disagree on this one. I actually believe, like Marc Andreessen, that we are at AGI. We reached AGI a few months ago with the latest ChatGPT 5.5 with Claude Anthropic, Opus 4.7, 4.6 even. For me, in some specific space, let’s say coding, for instance, I truly believe we have reached AGI. It is totally replacing people for development, testing, writing, specifications, doing design. Yes, there are some use cases where it’s not working as well. But for at least the coding side, I think AGI is truly there. Is it better than the best human experts at every piece of the puzzle? No. I think if you go to some very specialized development stuff like a CUDA development, for instance, or some kernel development, it might not be at the level of the top human experts. But as we can hear about stories around security to the new versions that are unreleased yet of Anthropic and others seem to be at the level of the best human experts at cybersecurity. I think we are actually there. I think that’s why there is an acceleration the past six months in terms of investments in data centers. There was a big question in 2023, 2024, 2025, “Are we investing too much? Is it a bubble?” And so on. That was true because revenues were increasing, but not that fast. The increase in the growth in revenues for Anthropic is insane. No one has ever seen that in the whole history of technology, what happened to them the past six months. Growing so fast at such a scale, no one has seen. I think that’s truly because they have stumbled upon AGI for coding. Coding is such a huge, large market that even if we are just solving coding, even if we never solve more than that, I think it’s already one of the biggest markets in the world, full stop.Nuno Goncalves PedroBut that, for me, I think maybe it’s a definitional thing. That, for me, is still narrow AI. AGI, as defined, is a type of AI that can match or surpass human capabilities across virtually all cognitive and intellectual tasks. Coding is just one of them. I think there’s a lot of narrow AI right there, to your point, that can replace humans maybe today. Coding may be a great example of that. But I don’t think we are at the AGI. I haven’t seen any agents out there that can replace fully human beings, where I would not understand that they’re actually AI. Anyway, again, maybe you guys have seen stuff that I haven’t seen. I just feel there’s a lot of overhype on what the AGI actually is. But the ability to surpass us as human beings, I haven’t seen it yet across the board, not on a specific domain.Bertrand SchmittYes, I think my point is that I have seen it on a few domains. If you look at mathematics, for instance, and there are more and more examples of people who are using AI to solve math problems that only the very top people in the world were able to solve that are now being solved by AI. I start to feel that we are getting a few examples in different fields where AI is matching the absolute best experts in the world. I would just go back in terms of market size. Already what we got today, and today is not the limit, but what we got today, and especially AI combined with all these agents, harness agents, that’s also that specific combo that makes a difference. This is, for me at this stage, huge market. Does it solve everything? Does it replace a human for everything? No, you still need the human in the loop. You don’t let your AI go wild for weeks and just check once in a while. But you could argue very few humans are able to do that, quite frankly, to run by themselves for weeks without control.Nuno Goncalves PedroI feel there’s still a great deal of confusion because there’s dramatic impact on narrow AI domains where things are being done that weren’t done before, where there’s these great mathematical proofing, coding stuff is being done very easily, etc, etc. But still that, for me, is not the definition of AGI. AGI is a human being. It’s like someone who can really be a human being. Ideally, become top end of what a human being will look intellectually across domains. I don’t think that’s where we are. Again, maybe definitional, maybe that’s what Marc means, but I mean-Bertrand SchmittYeah, the question is, what does it mean? I would argue it’s beating 80% of humanity today pretty easily. Most of humanity doesn’t know how to code. It will be unable to do coding at the level AI is doing. It will be unable to do math at the level AI is doing. Use ChatGPT to check for medical issues and stuff. It’s not perfect for sure, but, again, it still beats most humans. I don’t know. I feel that we are reaching a level where it’s truly beating most of humans, many experts in many fields. I feel the definition of, “Hey, it has reached the level of an average human being.” I would say it passed that one. Is it the same as a human being? No. But reaching the level of your average human being, I think we are there in a lot of fields.Nuno Goncalves PedroConclusion Well, in conclusion, in episode 78 of Tech Deciphered, we went into two piles: the pile of dead-on-arrival products, technologies, businesses, and the pile of “it will happen, but not like that” or “not yet” pile of technologies, business out there as well. We went into a bunch of topics today that were of extreme importance back in the past where we thought the world was going to change, anywhere from self-driving to fusion in the nuclear space. We also went into some dramatic, dramatic mistakes like Juicero and other deployments like Segway that seem to be really very much dead on arrival. With that, I would ask you if you’re excited, and you want to share with us some of your agreements or disagreements on today’s episode, as well as your own dead-on-arrival versus merely-early examples, feel free to do so on LinkedIn, X, or via email. Thank you, Bertrand.Bertrand SchmittThank you, Nuno.
We talk the foldable iPhone, Michael Jackson's death, and Trump passports. Some other notable news:Micron delivered one of the most stunning quarters in semiconductor history, reporting roughly $41.5 billion in fiscal-Q3 revenue, up about 346% year over year, and guiding next quarter to about $50 billion with an 81% gross margin. The read-through is that AI has turned high-bandwidth memory from a boom-bust commodity into a scarce, contracted input for the next generation of compute.SpaceX's record IPO unwound almost as violently as it launched, with the stock falling 31% in four sessions from its June 16 peak and erasing more than $600 billion of market value. A 4.2% public float, a $20 billion bond sale, newly listed options, August lockup risk, and a $4.9 billion 2025 net loss all collided in one of the clearest market-structure lessons of the AI trade.Oracle disclosed about 21,000 job cuts and directly tied the reductions to AI adoption in its own annual filing, even as it expands AI cloud capacity through major data-center deals linked to OpenAI and Meta. China also reclaimed the world's-fastest-supercomputer crown with LineShine, an all-domestic CPU-only system that hit 2.198 exaflops under U.S. export controls.A fatal Tesla crash in Katy, Texas, reopened the self-driving accountability fight after the driver said he had been using Tesla's partially automated driving system and both NHTSA and NTSB opened investigations. Robotics funding added the other side of the physical-AI story, with about $55.8 billion raised so far in 2026 and Figure banking a $1 billion Series C at a $39 billion valuation.The runner-ups: FedEx beat expectations and completed the FedEx Freight spin-off, onsemi agreed to buy Synaptics for about $7 billion to push deeper into edge and physical AI, and UN Secretary-General Antonio Guterres pressed AI companies to disclose data-center emissions, water use, land use, and energy sources. The 30,000-ft view: Q1 GDP was revised up to 2.1%, PCE inflation ran at 4.6%, markets repriced toward possible Fed hikes, Nvidia's Vera CPUs entered full production, and the mega-IPO pipeline still has Anthropic and OpenAI queued. If you want a prize, send us a DM: instagram.com/rickerandbon tiktok.com/@rickerandbon youtube.com/@rickerandbon
Apple's Worldwide Developers Conference happened this week, and there was enough going on that we wanted to unpack the whole thing, primarily due to the company's uncharacteristic backpedaling on its... controversial Liquid Glass UI language, not to mention the unusual focus on CPU scheduling and numerous other performance refinements across the board in this year's OS updates, rather than the more typical long list of new features. It was enough to get us saying the words "Snow Leopard," which is always a good feeling. We also consider new broader parental controls, the apparently final state of Apple Intelligence and Siri AI features, and more. Support the Pod! Contribute to the Tech Pod Patreon and get access to our booming Discord, a monthly bonus episode, your name in the credits, and other great benefits! You can support the show at: https://patreon.com/techpod
Your iPhone might be running hot and draining fast — and it’s not just you. Dave and Pilot Pete break down the battery chaos introduced by iOS 26.5, which brought overheating, accelerated drain, and even blocked wired charging on iPhone 17 and Air models. The fix that’s working for most people: disable iCloud Keychain first, run Reset All Settings, then carefully re-enable iCloud sync — otherwise you’ll nuke your Wi-Fi passwords across every device. iOS 26.5.1 is out and should help, but until you’ve updated, your electrons deserve better. You’ll also learn why Apple ID passkeys are locked to Apple’s own keychain with no known path to third-party managers like 1Password or Keeper, and why editing a contact on a modern Mac can somehow peg every CPU core — in 2026, no less. From there, Dave and Pete tackle the full listener mailbag: how to rescue missing contact names from Messages, the right way to boot a MacBook with a broken display into clamshell mode so it actually uses the external monitor, and a deep dive on 5K vs. 4K displays where Dave argues your eyes may not care as much as the pixel-per-inch math suggests. You’ll get smart ideas for repurposing a 2015 iPad Pro that can’t run modern apps — including Dave’s Claude Code-built weather dashboard running off a headless iMac as a web interface. A crashing 2021 MacBook Pro turns out to have been felled by a single bad SD card, and the lesson is golden: feed your crash reports to an LLM and let it do the digging. And Don’t Get Caught with outdated OpenAI macOS apps — update ChatGPT, Codex, Atlas, and Codex CLI before June 12th to stay ahead of a code-signing rotation triggered by a compromised open-source library. 00:00:00 Mac Geek Gab 1145 for Monday, June 8th, 2026 June 8th: National Best Friends Day MGG Monthly Giveaway – Win a license to SaneBox Quick Tips 00:00:01 Dan-QT-Multi-select on iPhone with a quick drag 00:04:31 Tim-QT-Have iOS 26.5 Battery Drain? Reset All Settings, but be careful! 00:13:32 Kent-QT-1144-Collapse stacks by clicking the down-facing carat in the menu 00:14:15 Mark-QT-Match Frame Rate on your Apple TV for smoother experiences 00:17:58 What are the differences between refresh rates and frame rates and…why? 00:21:09 KiwiGraham-QT-Apple Account Passkeys vs. Third Party Password Apps Sponsors 00:23:09 SPONSOR: Keeper. Right now, Keeper is offering our listeners 60% off personal and family plans at https://Keepersecurity.com/MGG. This offer is only for podcast listeners! 00:24:50 SPONSOR: Helix Sleep makes premium mattresses and bedding that are customized to fit your personal needs, and conveniently shipped to your door. Go to https://helixsleep.com/MGG for 20% Off Sitewide. 00:26:23 SPONSOR: NordLayer Browser. The business browser built for how modern work actually happens — giving IT the visibility and control to secure SaaS, stop phishing, and prevent data leaks right at the source. Your Questions Answered and Tips Shared! 00:28:09 VaShaun-How can I restore lost Contacts on my Mac? 00:37:36 Si-What to do with an 11-year-old iPad? Claude Code 00:46:40 Michael-Why do we have to pull-to-refresh for updates? 00:50:04 Blake-1144-Damaged displays, external monitors, and MonitorControl 00:55:48 Joe & Michael-CSF-1144–RetinaDesk.com for reviews of 5K and 6K monitors BenQ MA270UP 27” 4K Display Reviews 01:02:50 Hog fan and Cowboy fan-MGG Review–Favorite Tech podcast Don't Get Caught 01:04:14 Father John-DGC-Investigate those crash reports before you replace your Mac 01:09:26 Update your ChatGPT Apps ChatGPT Desktop Codex App Codex CLI Atlas 01:11:06 Andy-DGC-When Troubleshooting, Don’t Get Caught asking the wrong questions or assuming the wrong facts 01:19:36 MGG 1145 Outtro MGG Monthly Giveaway Bandwidth Provided by CacheFly Pilot Pete's Aviation Podcast: So There I Was (for Aviation Enthusiasts) The Debut Film Podcast – Adam's new podcast! Dave's Business Brain (for Entrepreneurs) and Gig Gab (for Working Musicians) Podcasts MGG Merch is Available! Mac Geek Gab iOS app Mac Geek Gab YouTube Page Mac Geek Gab Live Calendar This Week's MGG Premium Contributors MGG Apple Podcasts Reviews feedback@macgeekgab.com 224-888-GEEK Active MGG Sponsors and Coupon Codes List BackBeat Media Podcast Network
Nvidia announced its new CPU at an event in Taipei and Jon, Rachel, and Matt talked about why potential customers may be interested in buying as well as the potential impacts to primary CPU players such as Intel and AMD. The team also talks about Berkshire Hathaway's homebuilder acquisition before closing with a question regarding passive investing trends. Jon Quast, Matt Frankel, and Rachel Warren discuss: -Nvidia's new Vera CPU -The potential fallout in the CPU markout -Berkshire Hathaway's latest acquisition -Passive investing's impact on the stock market Companies discussed: Nvidia (NVDA), AMD (AMD), Intel (INTC), Qualcomm (QCOM), Berkshire Hathaway (BRK.A)(BRK.B), Taylor Morrison (TMHC) Host: Jon Quast Guests: Matt Frankel, Rachel Warren Engineer: Dan Boyd Disclosure: Advertisements are sponsored content and provided for informational purposes only. The Motley Fool and its affiliates (collectively, “TMF”) do not endorse, recommend, or verify the accuracy or completeness of the statements made within advertisements. TMF is not involved in the offer, sale, or solicitation of any securities advertised herein and makes no representations regarding the suitability, or risks associated with any investment opportunity presented. Investors should conduct their own due diligence and consult with legal, tax, and financial advisors before making any investment decisions. TMF assumes no responsibility for any losses or damages arising from this advertisement.We're committed to transparency: All personal opinions in advertisements from Fools are their own. The product advertised in this episode was loaned to TMF and was returned after a test period or the product advertised in this episode was purchased by TMF. Advertiser has paid for the sponsorship of this episode.Learn more about your ad choices. Visit megaphone.fm/adchoices Learn more about your ad choices. Visit megaphone.fm/adchoices