POPULARITY
The Twenty Minute VC: Venture Capital | Startup Funding | The Pitch
Aaron Katz is the Co-Founder and CEO of ClickHouse, the real-time analytics database powering companies including OpenAI, Anthropic, Tesla and Microsoft. ClickHouse just surpassed $350M in ARR and raised over $1B from investors including Dragoneer, Khosla Ventures, Coatue, 20VC and Benchmark. Previously, Aaron was CRO at Elastic, where he helped scale revenue from approximately $5M to $500M and led the company through its IPO. Before Elastic, he spent 12 years at Salesforce, working alongside Marc Benioff and helping transform it from a 200-person startup into a global software giant. AGENDA: 4:05 Are we in an AI bubble? 13:40 How does software change when agents—not humans—make buying decisions? 22:28 Will 90% of tokens flow through open models; can enterprises trust them? 31:09 Why did ClickHouse sponsor Fulham; and could sports teams become $20B assets? 35:49 When will ClickHouse hit $1B ARR? 38:31 Can startups still win elite talent from OpenAI? Biggest remote work mistake? 44:54 Is zero-to-$100M ARR now table stakes; or is durable growth what matters? 48:51 Is college still worth it; which jobs will survive AI? 57:58 When will ClickHouse go public; and why not next year?
Dolly Parton, country music icon, dies at 80. Sister of Lindsey Graham secures GOP Senate nomination. Texas Roadhouse says racism, food-tampering allegations are '100% a hoax'. Tariffs to increase prices for consumers. At least 1 in 4 NFL players who died in recent years had CTE, study finds. Canada to impose retaliatory tariffs on $20B in US goods as trade war escalates. Meta reaches $16.68 billion settlement over social media harms to children.
Andrew, Ben, and Tom preview Bessent's 1pm press conference declaring economic war on Iran, break down the collapsed US-Canada tariff deal and the resulting 50% tariff on $20B of goods, and discuss Nvidia's coming 15% server price hikes tied to memory costs.Join our live YouTube stream Monday through Friday at 8:30 AM EST:http://www.youtube.com/@TheMorningMarketBriefingPlease see disclosures:https://www.narwhal.com/disclosure
What's up, everybody? It's Tom Bilyeu here:Want my help starting a business? Join me here inside Zero To FounderSign up for my AI Masterclass: AI MasterclassFOLLOW TOM:Instagram: https://www.instagram.com/tombilyeu/Tik Tok: https://www.tiktok.com/@tombilyeu?lang=enTwitter: https://twitter.com/tombilyeuYouTube: https://www.youtube.com/@TomBilyeuQuince: Free shipping and 365-day returns at https://quince.com/impactpodWhatnot: Download the Whatnot app today and get free shipping on your first order.Quo: Try for free PLUS get 20% off your first 6 months at https://quo.com/impactIncogni: Take your personal data back with Incogni! Use code IMPACT at the link below and get 60% off an annual plan: https://incogni.com/impact Pique: 20% off at https://piquelife.com/impactShopify: Sign up for your one-dollar-per-month trial period at https://shopify.com/impactSurfshark: Go to https://surfshark.com/TOMB or use code TOMB at checkout to get 4 extra months of Surfshark! ATT Business: Switch to AT&T Business at business.att.comKetone IQ: Visit https://ketone.com/IMPACT for 30% OFF your subscription orderNetsuite: For the first time ever you can try NetSuite Next for free. If your revenues are at least in the seven figures, go to https://NetSuite.ai/Theory.The team breaks down the viral $300,000 debate between Candace Owens and Andrew Wilson—and argues it was a case study in everything wrong with online debating. The host's core point: the two were trying to litigate a criminal case whose facts haven't been established in court, arguing "internal frame versus internal frame," which inevitably devolves into madness when those frames don't intersect. He's pointed that this isn't about who's likable—he argues Wilson walked in as a "frame of reference" debater whose bully-the-room tactics collapsed against someone prepared, and that Wilson couldn't name the actual charges in the case while Candace could, making him look like he wasn't arguing from real command of the facts. On Candace, the host makes a deliberately careful epistemics argument: whatever you think of her conclusions, treating "she's obviously wrong" as self-evident is itself a frame-of-reference trap—because the facts of the trial (Tyler Robinson has not yet entered a plea, and the case hasn't been argued) simply aren't out yet, so neither side can claim proof. He predicts how the trial may play out and why the defense's choices about which evidence to use will be the real tell, while stressing repeatedly that he's speculating and that the honest position right now is "we don't know." Ryan and the team push on where the line sits between healthy skepticism and conspiratorial thinking, and the group lands on a shared lesson: don't get pulled into a debate when there's no agreed framework for adjudicating it, and wait for the facts before declaring a winner.The team breaks down the public exchange between Rep. Ro Khanna, who argues America must tax wealth and redistribute from billionaires, and Mark Cuban, who pushes back that "ideology is not a strategy." The host agrees with the diagnosis but not the cure: he concedes that people genuinely feel the system isn't working for them—and that they're right, pointing to the college-debt pipeline, an abusive housing market, and currency inflation eroding savings—but argues that taxing wealth misidentifies the solution. Walking through Cuban's argument, he explains the mechanics most people miss: founders of $10B "deca-unicorns" are typically cash-poor and stock-rich, worth billions on paper while the raised money goes into the company, not their pockets—so a wealth tax forces them to sell or borrow against speculative shares, which is often impossible and destructive to the company itself. He highlights Cuban's warning that a California wealth tax (Prop 40) would simply push investors to require startups to relocate to Texas, Indiana, or Florida first—meaning the state loses deca-unicorns it never even sees leave—and that "when you tax something, you get less of it." He connects it to his recurring critique of Bernie Sanders' framing: even if the revenue estimate were real (roughly $20B/year against a ~$30B Medicaid funding gap), you can only spend it once, and the act of liquidating shares to collect it would be self-defeating. The host's alternative is structural: get money out of politics, be fiscally disciplined, rebuild real manufacturing to take on China, stop funneling everyone into debt-financed college, let students discharge that debt in bankruptcy to force lending discipline, and stop suppressing wages with low-wage immigration so the true price of labor is revealed. Tom, Drew, and Ryan dig into a Census Bureau data effort that cross-referenced 2020 voter records against federal citizenship files—and end up in a genuine, three-way disagreement about what it means. Tom plants a strong flag: he believes there is significant voter fraud that will be revealed as more data is processed, arguing from "first principles" that where you have access plus incentive and no safeguard, abuse will follow, and that election integrity is worth shoring up even before there's overwhelming evidence. Drew and Ryan push back as the fact-checkers in the room: Drew stresses that the reported ~24,000 figure comes from a partial, unprocessed run and, crucially, that "non-citizen" includes lawful permanent residents (green-card holders) who aren't the same as people who crossed illegally—so the headline shouldn't be read as proof of illegal-alien ballot harvesting. Ryan fact-checks the specifics in real time: how Social Security numbers actually work for non-citizens (work authorizations, ITINs, fraudulent documents, expired visas), what California's voter-verification process actually requires (signature matching, with a first-time-voter ID exception under federal HAVA law), and the point that documented voter fraud in the US has consistently been shown to be well under 1%. The group works through California's and Minnesota's differing rules, whether signature-matching and vouching systems are genuinely exploitable, and lands on a striking meta-point: three people looking at the exact same facts arrive at very different conclusions, which is itself a lesson in how much of what we call "fact" is filtered through perception. A substantive, disagreement-driven look at election-integrity claims—with the receipts checked on air.ggSee Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
August 22, 2026, 7 AM; The Supreme Court just gave him the green light to continue his $400M White House ballroom renovations, for now. Chief Justice John Roberts issued a temporary stay, freezing a lower court ruling that would have brought the project to a halt this morning. Trump has focused so much on his pet projects that the White House has now dubbed him the “Builder-in-Chief.” However, his policies are causing things to crumble. Americans across the country are struggling with the cost of living, exacerbated by his war in Iran. Additionally overnight, trade talks between the U.S. and Canada fell apart. 50% tariffs are now in effect on about $20B worth of Canadian goods, which could lead to higher prices for American consumers. Rep. Jennifer McClellan joins The Weekend to discuss the President's focus on his vanity projects over the US economy. For more, follow us on social media: Bluesky: @theweekendmsnow.bsky.social Instagram: @theweekendmsnow TikTok: @theweekendmsnow To listen to this show and other MS podcasts without ads, sign up for MS NOW Premium on Apple Podcasts. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Brian Szytel reviews a rotation-heavy market day with the Dow up 120 points, the S&P 500 up about 0.25%, and the Nasdaq slightly higher, as equal-weight outperformed cap-weighted amid big moves in pharma and some late earnings from tech/AI. Treasury yields fell, with the 10-year down 7 bps to about 4.64%, following remarks from Treasury Secretary Scott Bessent about shifting issuance toward the short end and using it to buy back some long-end debt; while the $20B buyback is small versus the $5T in 20–30 year Treasuries, the signal suggests an effort to lower long-term rates, potentially at odds with a Fed under Warsh aiming to let markets tighten or loosen. He also explains Japan's debt dynamics: while gross debt/GDP is ~240%, netting BOJ holdings and government assets brings it closer to ~80%, though higher JGB rates could raise debt-service costs and pressure the yen and BOJ policy. 00:00 Welcome and Setup 00:21 Market Close Recap 00:56 Treasury Buyback Shock 01:58 Fed Versus Treasury 03:46 Japan Debt Question 04:07 Net Debt Breakdown 05:06 Rates Yen and BOJ 06:11 Wrap Up and Disclosures Links mentioned in this episode: DividendCafe.com TheBahnsenGroup.com
From shifting commodity markets and explosive data center growth to drone delivery and AI-enabled workplaces, supply chain leaders are navigating change from every direction. In this episode of The Buzz, powered by Toyota Automated Logistics, Scott Luton and Dr. Rod Thomas are joined by Hans Kremer, partner and co-founder of Inchange, to unpack the latest supply chain headlines and explore what professionals and organizations need to thrive in a rapidly evolving business environment. The discussion shifts to professional development, where Hans shares why companies need to stop training in silos and create more cross-functional learning experiences. The group also explores lifelong learning, the value of hands-on operational experience, and why AI should be viewed as a capability multiplier, not simply a tool for cutting costs or increasing efficiency. They close with powerful insights on leadership, trust, empathy, and creating environments where people can succeed. Key Takeaways: Successful change starts with the people doing the work. Find true advocates early, listen carefully to resistance, communicate the “why” repeatedly, and give teams the time and resources needed to implement change. Supply chain has become a strategic differentiator. Competition increasingly happens between entire supply chain ecosystems, not simply individual companies. Tomorrow's supply chain professionals need more than technical expertise. Strategic inquiry, discernment, judgment, adaptability, communication, and influence will become increasingly important as AI expands access to technical capabilities. Copper and other critical resources could become major constraints on AI and electrification growth. Data centers, infrastructure investment, geopolitics, tariffs, and limited natural resources are creating complex global supply challenges. Professional development needs to become cross-functional. Giving employees opportunities to experience sales, operations, supply chain, finance, and other perspectives can help break down silos and improve organizational performance. AI should make people more capable, not simply more efficient. Organizations that use AI only to automate existing tasks may miss its greater potential to enable employees to solve problems and perform work they could not previously do. Leadership is about enabling the success of others. Strong leaders remove obstacles, protect their teams, model the right behaviors, build mutual trust, and create an environment people want to follow. This episode brings together timely supply chain news with practical lessons on leadership, learning, technology, and organizational performance. Whether you're preparing your workforce for AI, navigating major market shifts, developing the next generation of supply chain talent, or simply trying to become a stronger leader, Hans, Rod, and Scott offer actionable perspectives you can put to work right away. Additional Links and Resources: Toyota Automated Logistics: https://toyota-automated-logistics.com/ With That Said: https://bit.ly/WTS-8-August-2026 Leading Beyond The Next Disruption: https://bit.ly/Mourad-Tamoud-SE Rod's LinkedIn Post: https://bit.ly/The-Perfect-SC-Manager Han's LinkedIn Post: https://bit.ly/Hans-Music-Stories Copper jumps to its highest level ever. What the metal is telling us: https://cnb.cx/4cvTGv8 Congo Issues Immediate Copper and Cobalt Ban as Smelters Pay Miners for Feedstock: https://bit.ly/DRC-Bans-Some-Exports Caterpillar sales surpass $20B as data center generators take off: https://bit.ly/CAT-has-big-2Q2026 DoorDash Takes to the Sky With Its Own Drones: https://on.wsj.com/4ccxIx7 Inchainge: https://inchainge.com/ The Readiness Rules: https://now.supplychainnow.com/the-readiness-rules Connect with Hans on LinkedIn: https://www.linkedin.com/in/hanskremer/ Connect with Rod on LinkedIn: https://www.linkedin.com/in/rod-thomas-a409a41/ Upcoming Live Programming: https://supplychainnow.com/upcoming-live-programming/ Supply Chain Now Resource Hub: https://supplychainnow.com/resource-hub/ Learn more about our hosts: https://supplychainnow.com/about Learn more about Supply Chain Now: https://supplychainnow.com Watch and listen to more Supply Chain Now episodes here: https://supplychainnow.com/program/supply-chain-now Subscribe to Supply Chain Now on your favorite platform: https://supplychainnow.com/join Work with us! Download Supply Chain Now's NEW Media Kit: https://bit.ly/3XH6OVk WEBINAR- From Disruption to Stability: Building Resilient Logistics Solutions in a Rapidly Changing Global Market: https://bit.ly/3TguZMt WEBINAR- SAP AI Inside the Supply Chain: From Silo to Orchestration: https://bit.ly/4bvpz6K This episode was hosted by Scott Luton and Rod Thomas, and produced by Trisha Cordes, Joshua Miranda, and Amanda Luton. For additional information, please visit our dedicated show page at: https://supplychainnow.com/buzz-ai-copper-drone-delivery-future-supply-chain-leadership-1622 The content in this episode, including all audio, videos, visuals, and graphics, is the property of Supply Chain Now and is protected by copyright law. Unauthorized use, reproduction, distribution, modification, or re-uploading of this content in any form is strictly prohibited without explicit written permission from Supply Chain Now.For licensing inquiries or permissions, please contact us at production@supplychainnow.com© 2026 Supply Chain Now. All rights reserved. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Get your week started with the fastest 45 minutes in freight! On today's episode, Malcolm Harris and Michael Vincent break down the latest supply chain headlines, crazy cargo fraud cases, and the evolving landscape of global trade. In This Episode: Headline News: Justin Turner named CEO of Mariner Logistics, a breakdown of the Little Debbie snack delivery fraud scheme, Indiana's massive truck inspection crackdowns, a driver's $510K fuel card fraud, and how new parcel surcharges boosted USPS revenue to $20B. Cargo Theft & Modern Security: Danny Ramon (Director of Intelligence at Overhaul) joins the show to walk through a real-time $500,000 electronics theft recovery in California. We dive into how cargo theft has evolved beyond traditional break-ins, the critical importance of continuous monitoring and law enforcement coordination, and why vetting brokers and carriers before pickup is your best line of defense. Overcoming Regulatory & Document Bottlenecks: Patrick Van Hull (Supply Chain Industry Consultant at Tungsten Automation) breaks down the operational reality of navigating tariffs and trade policy changes. Learn how document-heavy workflows—from bills of lading to customs entries—create massive supply chain bottlenecks, and how AI and automation are helping leaders build truly resilient operations. Watch on YouTube Visit our sponsor - GNOSIS FREIGHT Subscribe to the WTT newsletter Apple Podcasts Spotify More FreightWaves Podcasts #WHATTHETRUCK #FreightNews #supplychain Learn more about your ad choices. Visit megaphone.fm/adchoices
Get your week started with the fastest 45 minutes in freight! On today's episode, Malcolm Harris and Michael Vincent break down the latest supply chain headlines, crazy cargo fraud cases, and the evolving landscape of global trade. In This Episode: Headline News: Justin Turner named CEO of Mariner Logistics, a breakdown of the Little Debbie snack delivery fraud scheme, Indiana's massive truck inspection crackdowns, a driver's $510K fuel card fraud, and how new parcel surcharges boosted USPS revenue to $20B. Cargo Theft & Modern Security: Danny Ramon (Director of Intelligence at Overhaul) joins the show to walk through a real-time $500,000 electronics theft recovery in California. We dive into how cargo theft has evolved beyond traditional break-ins, the critical importance of continuous monitoring and law enforcement coordination, and why vetting brokers and carriers before pickup is your best line of defense. Overcoming Regulatory & Document Bottlenecks: Patrick Van Hull (Supply Chain Industry Consultant at Tungsten Automation) breaks down the operational reality of navigating tariffs and trade policy changes. Learn how document-heavy workflows—from bills of lading to customs entries—create massive supply chain bottlenecks, and how AI and automation are helping leaders build truly resilient operations. Watch on YouTube Visit our sponsor - GNOSIS FREIGHT Subscribe to the WTT newsletter Apple Podcasts Spotify More FreightWaves Podcasts #WHATTHETRUCK #FreightNews #supplychain Learn more about your ad choices. Visit megaphone.fm/adchoices
Story of the Week (DR):Uber's Strategy for Fighting Sexual Assault Suits: ‘What Were You Wearing?'An aggressive defense strategy is being used in response to an avalanche of litigation, with over 4,000 lawsuits filed by passengers accusing Uber of failing to protect them from sexual assaultWhile Uber publicly champions itself as a "survivor-centric" ally and trains support staff against victim-blaming, its defense lawyers ruthlessly contradict this to protect the $145 billion company.Uber's attorneys aggressively interrogate survivors about their clothing (e.g., asking if they wore underwear or heels), drug and alcohol consumption, and sexual history.The legal team weaponizes past trauma by scouring private medical records and therapy notes, even petitioning courts to force survivors to undergo psychiatric exams or reveal details of childhood abuse to paint them as "unreliable narrators".Legal experts warn these scorched-earth tactics are deliberately designed to re-traumatize victims, discouraging lawsuits and intimidating survivors into settling.Chief Legal Officer Tony West: “Why safety is personal to me”Cites his public service career to cover for ruthless corporate tactics, like drafting Prop 22 to strip driver protections.Frames survivor trauma as a flaw of the "adversarial legal system," deflecting from Uber's own attorneys using aggressive, victim-blaming tactics to minimize payouts.Uses a single favorable quote from one trial judge to dismiss widespread litigation abuse.Touts ending forced arbitration while continuing to block class-action lawsuits, trapping survivors in isolated 1-on-1 fights.Brags about transparent reporting despite Uber burying 400,000+ assault complaints (2017–2022) under labels like "unaudited."TikTok Withheld a Safety Feature From Millions. One Died by SuicideWhen TikTok tweaked its algorithm in 2021 to stop users from being overwhelmed with harmful content, the company didn't roll out the safer version to everyone, a confidential internal document shows.Instead, to see if the change might reduce the app's stickiness, the company conducted an experiment. It created a control group of 10% of US users — at the time, that would have been about 15 million people — who kept the old version of the app.This group, according to the document, included 16-year-old Chase Nasca. TikTok's algorithm pushed him thousands of videos about suicide, sadness, hopelessness and loneliness, right up until he killed himself.TikTok says moderator error delayed removal of Perez Hilton's livestream showing acts of self-harmGAO finds Elon Musk's DOGE inflated claims of $110 billion in savings for federal governmentThe GAO found that 96% of the savings DOGE claimed from grant cancellations lacked supporting evidence or methodology to verify.DOGE claimed $27.4 billion in savings across 2,503 contracts—including a $1.7 billion Pentagon IT deal—that were never actually terminated or reduced.Over 40% of the lease cancellations listed on DOGE's "Wall of Receipts" were already in progress before the agency was created.GAO uncovered widespread data errors, such as overstating total lease savings by nearly $60 million and failing to report baseline limitations.DOGE failed to disclose how it calculated its metrics and completely ignored GAO requests for interviews and supporting data.Where are Elon Musk's former DOGE staffers now?Edward “Big Balls” Coristine: a whistleblower alleged that Coristine planned to upload sensitive Social Security Administration data to an unsecured server. He also helped dismantle the U.S. Agency for International Development, or USAID.Coristine now appears to work at the National Design Studio, an office created to overhaul federal websites and other public-facing services, making them more usable, visually consistent and less expensive to build.Jeremy Lewin: helped oversee the dismantling of USAIDLewin later became the State Department's senior official for foreign assistance, humanitarian affairs and religious freedom, giving a former DOGE operative authority over the programs his old team had helped slash.He was subsequently reported to be moving to the White House National Security Council.Nate Cavanaugh and Justin FoxFounded a holding company called Special, which promises to do for private businesses what DOGE did for the federal government.“To achieve this, Special is building an operating system to transform critical American industries with AI. We call it SpecialOS.”Gavin Kliger, Luke Farritor, Marko Elez and Jack SteinFounded Cathedral, a startup that plans to use AI to expand U.S. military cyber capabilities, including defenses against adversaries such as China.Reuters reported that the company raised $160 million at a $1.4 billion valuation.Ethan Shaotran: helped move thousands of immigrants into the government's “Master Death File,” effectively deactivating their Social Security numbers.Founded Blitz Industries, which he described to Wired as “a defense company backed by big names.”Elon Musk: SpaceX Achieves Its First Moon Landing by Accidentally Crashing Into It at 5,400 Miles Per Hour, Causing Huge Explosion and 60-Foot CraterWhy companies that call themselves meritocracies don't always pay equally AND Meritocracy claims make managers pay men more than equally qualified women MMExplicitly branding an organization a "meritocracy" lowers managers' vigilance against bias, ironically making them more likely to discriminate.Given identical job titles and performance ratings, managers in self-proclaimed meritocracies systematically award higher bonuses and raises to men over equally qualified women.The meritocracy illusion compounds wage gaps across demographics, yielding lower financial rewards for women, racial minorities, and immigrants with identical output.Vague definitions of "merit" allow supervisors to act on personal favorability and affinity bias while genuinely believing their choices are purely objective.Trusting the "meritocracy" label creates false complacency, leading companies to eliminate or skip active pay-equity audits that would otherwise catch these disparities.Goodliest of the Week (MM/DR):DR: Google Scraps AI Satellite Images Within 24 Hours: Why 'Deepfakes in the Sky' Alarmed ResearchersThis happened after facing fierce backlash from journalists, researchers, and open-source intelligence experts who warned that it could accelerate the spread of misinformation.The feature allowed users to generate AI-created satellite images by zooming into any location in Google Earth and entering a text prompt.MM: Courts: MM DRCourt rules against Trump EPA's freeze of $20B in ‘green bank' fundsUBS fined record $125 million for money laundering violationsMeta fined $567m in largest child safety ruling against social media giantJudge says Nexstar can't appoint its executives to TEGNA's Board of DirectorsAssholiest of the Week (MM):Everything's a publicity stunt:Jealously Watching OpenAI and Anthropic, Meta Suddenly Claims That Its AI Went on a Hacking Spree TooZuckerberg conveys ‘apologies' for child sex abuse material, errors in operating the platform, government sources say‘Pure insanity'—Elon Musk details SpaceX's plan to turn the moon into its newest manufacturing siteAnthropic's Mythos created fake identities to fool humans in new cyber incidentGAO finds Elon Musk's DOGE inflated claims of $110 billion in savings for federal governmentEverything's a vanity projectJeff Bezos Says He's Selling $1 Billion In Amazon Stock Every Year to Fund Blue Origin — ‘It's The Most Important Work I'm Doing'SpaceX Achieves Its First Moon Landing by Accidentally Crashing Into It at 5,400 Miles Per Hour, Causing Huge Explosion and 60-Foot CraterJudge Sets Paramount-Warner Bros. Merger Trial for MarchEveryone's a victimThe AI bros are lonelyPalantir CEO Alex Karp Says Frontier AI Labs Will Steal Your IP, Your Know-How, and Your Expertise — ‘You Deserve To Be Colonized'‘Socialism is a disaster': Bill Ackman says Zohran Mamdani's bad left-wing housing policy is worsening New York's affordability crisisNo one is accountable DRBobby Kotick, former CEO of Activision, may join Paramount's board of directorsJefferies analyst calls out ‘simple-minded' criticism of SpaceX's governanceABC, CBS, and NBC aired nearly 200 segments about three U.S. heat events without mentioning climate changeDisney (Iger, still), Ellison family, and Roberts family - one cowed exec, two dictators!Headliniest of the WeekDR: Pepsi is launching soda-scented body wash and scrub in a beauty collaborationPwC U.S. CEO [Paul Griggs] advises young workers to say yes a lot, even if it means working on the weekends—‘That person climbs into the next role'Top Consulting Group PwC Caught Using AI on “Thought Leadership” Report About AI, Resulting in Corporate Document Filled With Bizarre HallucinationsDoorDash CEO Tony Xu once thought his $87 billion company would only make $100 million. He tells founders to be ‘greedy' and dream biggerClass A = 1 vote; Class B = 20 votes; CEO Tony Xu personally owns roughly 3% of actual shares but controls 55% of voting power via a “binding agreement” with the other co-foundersMisleading proxy shows he got paid $432k (they cited a 12:1 pay ratio) but he actually made $343M from shares and stocksDoorDash excludes Dashers from SEC disclosures by classifying them as 1099 independent contractors.Factoring drivers into the worker pool expands the headcount from 20,000 to over 7 million, moving the median worker from corporate offices directly onto the delivery platform.Dasher: while gross pay averages ~$15–$22/hour, net driver earnings sink further after deducting out-of-pocket expenses like fuel, vehicle wear, and self-employment taxes, pulling true net median compensation closer to $400–$900 annuallyAnthropic CEO reportedly worried new hires only care about money: reportedly expressed concern about new talent coming to the company for the money rather than the missionMM: Stripe CEO says his ‘urgency' to drop out of MIT ‘was a bit unnecessary'Who Won the Week?DR: Harvard Business School graduates TikTok CEO Shou Zi Chew and TikTok US CEO Adam Presser, because nobody ever talks about you or knows who you areMM: Meat. After seeing Salad and Go files for Chapter 11 bankruptcy after cyclospora fears worsened its challenges, you can't help but wonder if RFK Jr orchestrated a parasite to get people to buy more meat.PredictionsDR: Tony Xu introduces Class F stock where shareholders' children enter indentured servitude and must deliver 1000 Cheesy Gordita Crunches from Taco Bell before they are set free ; they will then get 0.0002 shares per voteMM: Boards are calling retired CEOs back as succession pipelines run dry: Boards run out of ex-CEOs to re-king and begin looking for different, egalitarian, integrated professionals to fill new roles. They call it DEI.
This week is one giant lesson in consolidation: everybody's buying everybody, the money is concentrating, the ownership is concentrating, and the people are getting cut. Tripledot bought Supersonic from Unity, Saudi Arabia closed the biggest leveraged buyout in history to take EA private, and Netflix gutted a studio seven weeks after praising its game on an earnings call.Matej Lančarič flies solo for the weekly news. Tripledot acquired Supersonic from Unity for $40M cash — a striking number given Supersonic likely did ~$40-45M in revenue over the last 12 months (per Sensor Tower plus ad revenue), and given ironSource paid $150M for it 11 years ago and Unity got it inside the $4.4B ironSource merger in 2022. It's the final piece of ironSource leaving the building — and the second time Tripledot has done this, after buying AppLovin's entire games business for $800M a year ago. The pattern is now complete: both AppLovin and Unity have exited game publishing to become pure ad-tech companies (ad networks make ads, publishers make games). That sale landed alongside Unity's best quarter ever (revenue $546M, up 24%, Grow segment up 35% to $389M on Vector AI, approaching GAAP profitability). Saudi Arabia's PIF closed its $55B take-private of EA (now delisted, ~93% PIF-owned with Silver Lake and Affinity Partners) — the largest LBO in history, with ~$20B in debt that will push EA hard toward recurring mobile and live-service revenue, plus a Vision 2030 soft-power dimension via EA Sports FC. Netflix's FIFA World Cup studio Refactor laid off 85% of staff seven weeks after Netflix named the game one of its most successful cloud debuts — the same praise-then-kill pattern as Squid Game Unleashed. And AppLovin posted a "miss that isn't a miss" (revenue up 53% to ~$1.92B, just under consensus).The through-line: the ad networks make ads now, the publishers make games, and the ownership of the whole industry is concentrating like never before.━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━⏱️ TIMESTAMPS00:00 A week of pure consolidation — everybody's buying everybody01:40 Tripledot buys Supersonic for $40M — the ironSource endgame05:30 The pattern: ad networks exit publishing06:40 Unity's best quarter ever — the Vector AI turnaround09:30 Saudi Arabia takes EA private for $55B13:00 The Netflix FIFA disaster — praised, then gutted17:30 AppLovin's Q2 — the miss that isn't a miss---------------------------------------This is no BS gaming podcast 2.5 gamers session. Sharing actionable insights, dropping knowledge from our day-to-day User Acquisition, Game Design, and Ad monetization jobs. We are definitely not discussing the latest industry news, but having so much fun! Let's not forget this is a 4 a.m. conference discussion vibe, so let's not take it too seriously.Panelists: Jakub Remiar, Felix Braberg, Matej LancaricJoin our slack channel here: https://join.slack.com/t/two-and-half-gamers/shared_invite/zt-3bckldvr8-8PXvzciMWdheOzED9hq0SA---------------------------------------Matej LancaricUser Acquisition & Creatives Consultanthttps://lancaric.meFelix BrabergAd monetization consultanthttps://www.felixbraberg.comJakub RemiarGame design consultanthttps://www.linkedin.com/in/jakubremiar---------------------------------------Please share the podcast with your industry friends, dogs & cats. Especially cats! They love it!Hit the Subscribe button on YouTube, Spotify, and Apple!Please share feedback and comments - matej@lancaric.me---------------------------------------If you are interested in getting UA tips every week on Monday, visit lancaric.substack.com & sign up for the Brutally Honest newsletter by Matej LancaricDo you have UA questions nobody can answer? Ask Matej AI - the First UA AI in the gaming industry! https://lancaric.me/matej-ai
Watch the full episode on YouTube:We first covered Baseten last year when DeepSeek mania was at peak hype. Now they have raised a monster $13B round and become one of the new cohort of AI Infra decacorns that are (with Nvidia, Intel, and the semis complex) chief beneficiaries of the Inference Inflection. We return to Baseten at the peak of the 2026 edition of Open Weights debate. Ali has published a viral breakdown of Kimi K3:And since you last saw him, Philip has spoken at AI Engineer and written the definitive book on Inference Engineering spotted all over SF:Three years ago, inference engineering barely existed as a category.Today, it is one of the most critical disciplines in AI. Inference engineering inherently tackles a different question than standard model training: “How do you turn those weights from training into a product that is fast, reliable, and affordable at scale?” Focusing on these creates an entirely new optimization problem.In one recent GLM-5.2 experiment, quantizing more of the model actually preserved its benchmark quality while increasing throughput by 20%, because the errors introduced in different layers could cancel each other out.Inference is no longer just the final step after training. It is becoming its own engineering discipline, with its own research problems, infrastructure, and increasingly specialized roles.In this episode, Baseten's Philip Kiely and Ali Taha join swyx and Vibhu to explain what actually happens after a new open model is released and what it takes to turn “we generated a token” into a fast, reliable, production-ready API.We go deep on cache-aware routing, disaggregated prefill and decode, quantization, speculative decoding, KV-cache movement, model parallelism, GPU kernels, and the race to make frontier models up to 10× faster. Philip and Ali explain why inference optimizations can still produce gains of 20%, 100%, or even 200%; how quantization errors can cancel one another out; why identical weights can behave differently across clusters; and how Baseten grafted a Kimi vision encoder onto GLM-5.2 without changing the underlying language model.The conversation then expands beyond LLMs into NVIDIA Dynamo, mega kernels, Rubin, AI-specific chips, local inference, video generation, diffusion versus autoregressive models, and the enormous compute barrier to generating coherent long-form video. Finally, we explore the convergence of training and inference, continual learning through persistent KV cache, and the emerging loop where models help optimize the infrastructure that runs them.We discuss:* What happens when a 200,000-token request enters an inference system* Cache-aware routing and reusing previously computed KV cache* Why prefill and decode are increasingly handled by different GPUs* When dedicated deployments become cheaper and more reliable than shared APIs* How speculative decoding uses a smaller model to accelerate a larger one* Tool calling, structured outputs, and what LLMs actually do* What it takes to support a new open model on day zero* Grafting Kimi's vision encoder onto GLM-5.2* Retrofitting inefficient model layers with components from other architectures* Why models sometimes collapse into repeating the same token* How hardware, kernels, and race conditions create nondeterministic failures* Preserving model fidelity while making inference faster* How quantization errors can cancel each other out* Why inference optimizations still deliver gains of 20%, 100%, and 200%* How optimized serving can make a model up to 10× faster* NVIDIA Dynamo, KV-aware routing, and distributed model serving* Speculative decoding the speculative decoder* Why local AI is about making models less dumb while data-center AI is about making them less slow* Tensor, expert, and pipeline parallelism across GPUs* Hardware-aware model design, auto-tuning, and the case against mega kernels* Rubin and why inference is becoming a systems problem* Whether modern GPUs are evolving into programmable AI ASICs* Why enormous models like Kimi K3 require GB300-class hardware* Why open-source video generation still trails Veo, Kling, and other closed models* The quadratic attention bottleneck behind long-form AI video* Autoregressive video, real-time generation, and compounding quality drift* Why future video systems may combine autoregressive and diffusion architectures* Training for inference and inference for training* Continuous post-training, deployment, evaluation, and improvement loops* How GLM-5.2 helped optimize the kernels serving GLM-5.2 itself* Why faster networking could unlock dramatically faster decoding* Continual learning, KV-cache compaction, and persistent model memoryShow Notes* How to build a day-0 API for Kimi K3* 22580: From GPT2 to Kimi3, ExplainedPhilip Kiely* LinkedIn: https://www.linkedin.com/in/philipkiely* X: https://x.com/philipkiely* Inference Engineering: https://www.baseten.co/inference-engineering/Ali Taha* LinkedIn: https://www.linkedin.com/in/aliestaha/* X: https://x.com/waterloointernTimestamps00:00:00 Introduction and the 200K-Token Prompt00:03:18 Dedicated Deployments, Speculative Decoding, and Tool Calling00:11:26 Launching Production-Ready Open Models00:19:06 Model Retrofits, Failure Modes, and Nondeterminism00:28:22 Quantization and Canceling Errors00:32:15 The Race to 10× Faster Inference00:40:48 Dynamo, Speculation, and Local vs. Data-Center AI00:50:18 Model Parallelism, Auto-Tuning, and Mega Kernels01:00:55 Rubin, GPUs vs. ASICs, and Custom AI Chips01:10:03 Giant Models and the Limits of GPU Memory01:12:42 AI Video, Quadratic Attention, and Autoregressive Generation01:21:47 Audio, Images, and Diffusion Models01:27:32 Training, Self-Optimizing Models, and Continual Learning01:40:06 Closing ThoughtsTranscriptIntroduction: Baseten, Waterloo Intern, and Inference EngineeringSwyx [00:00:00]: Okay, we're here in the studio with Philip, old friend from Inference Engineering, the book, as well as Baseten and everything that you've done, you and I have done before, as well as Ali. Welcome.Ali [00:00:15]: Pleasure to meet you.Swyx [00:00:15]: Waterloo intern.Ali [00:00:16]: Waterloo intern, always.Swyx [00:00:17]: When did you get “Waterloo intern” as a handle?Ali [00:00:19]: As a handle? Oh.Ali [00:00:20]: I think the rebranding happened mid-March. When I saw it was open, I was like, “I have to take it. Up for grabs.”Philip [00:00:26]: The problem is that Ali is really good at his job and is not gonna be an intern much longer.Philip [00:00:30]: So we have to figure out who's gonna get the handle.Ali [00:00:33]: Well, I'll pass the torch over to the next intern.Swyx [00:00:34]: Oh, okay. It can be, like, you just pass it to another Waterloo grad.Ali [00:00:37]: To another Waterloo intern. No, bruh.Philip [00:00:39]: Yeah.Ali [00:00:39]: Intern.Swyx [00:00:40]: Intern, yeah.Ali [00:00:40]: And no.Philip [00:00:41]: You gotta get an intern from Waterloo.Ali [00:00:42]: Yeah, I've gotta get an intern from Waterloo.Swyx [00:00:44]: Right.Ali [00:00:44]: But they have to follow the path.Swyx [00:00:45]: Oh, it could, but it could come from Baseten, so it's like whoever Baseten gets from Waterloo.Ali [00:00:48]: Right.Swyx [00:00:49]: Has the title of Waterloo.Ali [00:00:50]: It stays in the ecosystem.Philip [00:00:51]: Exactly.Ali [00:00:52]: Halfway through the internship, you either get it or you're out.Philip [00:00:55]: You should also do, like, a big graduation ceremony where you change the handle.Ali [00:00:59]: Just say it.Philip [00:00:59]: For everybody.Swyx [00:01:00]: You guys are good at ceremonies, clearly. We had a nice launch of the book, very successful. But before we get into all that, I wanna start off with a fun question for you. Okay, you're an expert inference engineer. What happens when I send a long query, say two hundred thousand tokens into Baseten's inference? What's the process of query through GPU model routing, balancing, all that? What is all the stuff that we don't think about?Long Context Requests, KV Cache, and Cache-Aware RoutingPhilip [00:01:26]: With a long query specifically, the first thing that I'm gonna ask is, “Have you sent me this query before, or at least part of it?” and I really hope you have, because it's gonna be a lot easier for me and a lot cheaper for you. So the first thing that we're gonna look at is some cache-aware routing, where we're going to see, we probably have a number of instances, a number of replicas up serving whatever model you're hitting. We want to send this one to something with, number one, available prefill workers, and number two, ideally some cached input already there so that we can skip prefill on at least part of these two hundred thousand tokens. If you're doing two hundred thousand tokens, it's probably coding or a multi-turn agent or something where you would expect to have that cached. If you don't, we're gonna have to send it to a prefill worker. We've at least on certain models disaggregated prefill and decode, so you're going to have one set of GPUs that's solely going to process the input, create the KV cache, and get you your first token, and then that's going to be passed over to a separate set of GPUs, which is going to run decode. We're going to iteratively make those tokens. We're probably going to have some speculator model in front of that. I'm going to assume that you're doing coding, and because of that, our speculator model, which assumes you're doing coding, is gonna have a high draft token acceptance rate. If I'm wrong and you're asking me to summarize every Harry Potter book, it's gonna be slower. And then we stream that output to you and account for it, charge you, a couple of pennies and say, “Hey, would you like to send another one?”Swyx [00:03:04]: Except Baseten doesn't charge by pennies.Philip [00:03:07]: Well, yeah, we charge. I'm assuming that we're talking about the public model APIs. If you are setting up a dedicated deployment, then yeah, it's not pennies.Public APIs vs. Dedicated DeploymentsSwyx [00:03:18]: Yeah, one of the key differentiators when I was talking with Baseten initially was that people who want very high volume just need to rent by the box, ‘cause then it's up to you to figure out how to saturate the box.Ali [00:03:31]: And more often than not, it's, like, way cheaper if you're pushing, like, millions of tokens per hour, if you just pay per hour instead of pay per token.Philip [00:03:37]: Yeah, they do. I think that we've increasingly seen a lot of demand for the pay per token APIs, just because everyone wants to try open models, and then once they find a use case that's really sticky, then they move over to dedicated.Swyx [00:03:51]: Is there a best practice on when it's time to swap over?Philip [00:03:54]: Couple reasons. Yeah, reliability, that's a big one, right?Ali [00:03:57]: Like, if they have a very specific use case, they want you to train something specifically for them, like they want their own spec dec, for instance, for their own traffic.Swyx [00:04:04]: Spec dec is speculative decoding.Speculative Decoding and Custom SpeculatorsAli [00:04:05]: Speculative decoding, yeah.Swyx [00:04:07]: You have to explain.Ali [00:04:07]: Sorry. Like, speculative decoding is like, if you have a huge model, right? And so the model is going to be generating one token at a time every single turn, every single forward pass. So we attach, like, this little, like, parasite, like this layer that goes on top of the model, and this model just has to predict. It does three very fast autoregressive forward passes, and it will predict, like, three certain tokens, and then you do one forward stage over the entire original model in order to see if those predictions were correct or not, and then you accept them or you reject them. Now, this draft model is traffic specific, so if you, like, Philip said, if you're summarizing Harry Potter books, I can train exclusively that draft model on Harry Potter books, and I can guarantee you that I'm gonna accept the three tokens every single time. And so with that case, I increase your decode speed. I wouldn't be able to provide this to you if you're a shared endpointSwyx [00:04:53]: YeahAli [00:04:53]: ‘cause I have no idea if you're doing Harry Potter, if you're doing coding, if you're doing English. We don't know. Also, there was a thing in the book that mentioned that if they really cared about a specific threshold, chapter four, I think. Do you remember that?Philip [00:05:06]: Yeah. The things that you can do is you can set a specific, like, batch sizing, a specific, like, parallelism strategy if you're trying to optimize for, like, throughput versus latency. You can. Maybe a NVFP4 quant doesn't pass your benchmarks and you wanna run a model at higher precision, you could do that. There's just a bunch of reasons why you might wanna have your own endpoint and the biggest one, of course, just being, like, you don't have to deal with someone else doing a hundred million tokens of benchmarking traffic at the endpoint when you happen to be trying to serve your users.Swyx [00:05:40]: Yeah. I think one thing that is. That is a classic journey. Like, it's people is asking the, what happens when you type Google into the browser. Tool calling, is that just, you're generating JSON or is there more complication beyond that?Tool Calling, JSON, and Structured OutputsAli [00:05:58]: Certain customers that we have, they have their own post-trained models, and so they demand a tool calling that's not just, like parse a file or go find the weather. It's something that's very specific and you have to do post-training on this. And if the post-training on the model is not good or if the quantization after the post-training to get the inference to be fast, the model will struggle reading the JSON file and reading the tool calling. But it doesn't require its own like sandbox. It's not like it's going to use that tool calling to like escape a sandbox or like it doesn't have to be contained. It can just be a normal dedicated deployment. The challenge with tool calling more and more seems to be that the companies want certain tool calling which is a very sensitive thing to train. And because you're dealing with all of the JSON outputs, if it doesn't like close the end of the request in a very certain manner, you end up with a model that did the tool calling and like the thinking and so as a result of that, it didn't see the result and just hallucinated the result as it decoded. That seems to be the most challenging thing with tool calling, not really the sandboxes model.Philip [00:06:56]: Yeah, that's a challenge on the training side and then on the inference side, there's work that you can do to scope the possible output. So we published this at this point close to two years ago, the solution to this problem which is you make a state machine and you use that to constrain the output to a specific format. So this is the structured output problem. If you remember backSwyx [00:07:27]: Yeah, the specific grammar is,Philip [00:07:29]: Yeah, exactlySwyx [00:07:30]: GML had this thing.Philip [00:07:31]: Yeah. So it's like the old-school “make sure this is only JSON”, return only JSON orSwyx [00:07:38]: YeahPhilip [00:07:38]: Grandma's gonna die type of prompts.Swyx [00:07:39]: Is it BNF grammar? At some point OpenAI had released a thing that was like, yeah, if you want to constrain your output, write BNF grammar, back as NOR.Philip [00:07:47]: In our inference system, it's just a specified output format. And you get the guarantee that your output's gonna be structured along that format. And so applying that to tool calls can like help cut down on. You can still call the wrong tool or call no tool. It doesn't solve the certainty problem but it at least solves the output structuring problemSwyx [00:08:10]: YeahPhilip [00:08:10]: Within tool calls.Swyx [00:08:12]: And MCP is just another form of tool, right.Philip [00:08:14]: Yeah, exactly.Swyx [00:08:15]: As far as there's no special thing there.Philip [00:08:16]: The thing I'm always like explaining to people is the LLM is not capable of doing anything. It's only capable of making suggestions of what to do and then if those suggestions are formatted in a certain way and applied to a system that knows what to do with them, then an action occurs.Swyx [00:08:32]: Yeah. Part of the fun stuff is, this is solved outside of tool calling too. Like in an agent loop if the output is not correct or you're right, like reasoning, tool calling was done in the reasoning trace, just be like, “Oh, I don't know what to do. Let me just try again.” And it might get there after a few tries. And on your point of training, sometimes this is harder in smaller models, so you don't have the same exact quality outputAli [00:08:56]: Right.Swyx [00:08:57]: When you just swap from a big model, right?Ali [00:08:59]: Yeah. I will say that, before, I think we need to go back to inference engineering proper.Ali [00:09:04]: But, I had expected that something would replace JSON because it's hard to stream JSON ‘cause JSON must be complete and you must have open and close brackets and everything. So it's hard to parse something or validate something while it's being streamed. So people invented all sorts of things that are like, I forget the name of some of these alternatives, but it's something like TOML, something like YAML. But JSON seems to be dominant still.Philip [00:09:30]: The JSON outputs aren't that long, right? Like you could have a long-- ‘cause tool calls also contain the arguments in them and perhaps for a certain tool you might pass like a very long argument. But my impression of the median tool call is that it's a relatively small number of tokens, right? So I would expect that speculators are generally fairly good at something as formatted as JSON. And so you would have like a pretty fast decode step there and that the streaming wouldn't be as valuable, but maybe I'm wrong about that.Ali [00:10:02]: I think you're also bounded by the software or that the model is gonna integrate with if the software is built with JSON for the tool calls or if the company that you'- if your customer says that this is how our software works and our tools are interfaced with JSON, you can ask them to like, change their software and say like, “Yeah, this is gonna be better for the model.” but like with the right training shouldn't be that much of a difference. Also more profitable if it outputs more tokens probably.Swyx [00:10:25]: Depends on your business model.Swyx [00:10:27]: It really depends. But I will say that, as a writer with like experience a lot with generated output, I do try to move from text to JSON text which is very long JSON, right? Like there's paragraphs in every field because I'm trying to structure it, right?Philip [00:10:44]: Right.Swyx [00:10:44]: I want you to first make factual statements, then make opinions then make bullet point summaries, have dates, have entity references have your sources for references, all these things. Anyway, so these are things that like I think people who really experiment with structural output have to really care about. But, let's, let's recurse up the stack a little bit. Before we started recording, you mentioned something really cool, which is that there's a lot of engineering that-- inference engineering that goes on when a new model provider releases a new model, right? So let's call it GLM-5.2, Kimi K3. I had previously assumed, especially if it's like, well, GLM 5 to 5.1 to GLM-5.2, like that you've supported them before. Is it that much work?What It Takes to Support a New Open ModelAli [00:11:26]: It's a lot of work.Swyx [00:11:28]: Yeah. Okay. So like, a lot of people, all you guys, right whenever a new model launch like, people rush to say like, “Oh, Hugging Face supports this, Fireworks supports this, Spacetime supports this,” and I'm like, “Yeah, of course we support it.” But what goes into that? What goes intoPhilip [00:11:40]: I think it's more than just support it too, right? It benefits the consumer a lot. Like I think it was with Kimi K2.5 or GLM-5.2 the latest, there was an inference war, right? X provider is at 90 tokens a second. The next day we're at 150. The nextSwyx [00:11:55]: I kinda kicked that off with the GLM-5.2.Swyx [00:11:58]: I wrote a Twitter article about. It got like half a million views,Ali [00:12:02]: Based on being numberSwyx [00:12:03]: YeahAli [00:12:04]: Or it's for something else.Swyx [00:12:05]: Yeah. Which,Ali [00:12:06]: Oh my GodSwyx [00:12:07]: Which then got everyone really excited about, hey, how can we, bend tracks a little bit further and,Philip [00:12:14]: There's a difference between support the model, as in I can make a token out of this model, and support a model, as in I have a production-ready API from this model.Philip [00:12:26]: Getting to the point of I can make a token out of this model is not that hard because generally the, open source inference engines, vLLM, SGLang of the world oftentimes even receive weights ahead of time, maintainers do, or the people making the model merge PRs to ensure support. So you generally can, just get it working on the standard open source stack without too much pain in most cases. The challenge is, every inference company is gonna have own proprietary stack. Some open source components, some in-house stuff. And for any arbitrary model, there's going to be some new stuff. Sometimes you get lucky, like K, two five to two six was, like, pretty similar.Quantization, Speculators, and Production ReadinessAli [00:13:16]: Yeah. It was pure continued post-trainingPhilip [00:13:18]: YeahAli [00:13:18]: If I remember correctly.Philip [00:13:19]: Even in those cases, there's still stuff you have to do. You have to redo the quantization work. You're taking the model from. Generally, these models are not released in NVFP4, and we want them to be in NVFP4 for maximum Blackwell compatibility. So we have to perform that quantization, and, calibrate the quantization to make sure that we're not causing any regression in the model's intelligence. And then we also have to train the speculator, as we've talked about. Generally, we have. We have ZDR, zero data retention on our model APIs, so we don't know exactly the traffic that people are sending us, but we know what's popular. We know that coding use cases are popular. We know that agents, agentic use cases are popular. So we can get public data sets that are representative of that traffic and train general speculators. Now, with speculators today, you need to train the speculator using the base model itself because you're getting hidden states out of the model from running inference on these specific prompts, and that is the training data you use to create the speculator. So there's that process which you need the real model weights for. And then there's of course just the process of, standing up all the infrastructure behind it, loading all this stuff, testing it. And then when there's a new model with a newer architecture, I think that, like, the DeepSeek models tend to be the most challenging as they have, like, the most novel architectural stuff going on, model after model. But every new model has something. Kimi K2 had. Oh, sorry, GLM-5.2 hadAli [00:14:53]: Sparse attention.Philip [00:14:54]: Yeah,Ali [00:14:54]: YeahPhilip [00:14:54]: the DSA.Ali [00:14:55]: Right. Which is brought from DeepSeek.Philip [00:14:57]: Yeah. AndAli [00:14:59]: So you can copy-paste then?Philip [00:15:01]: It kindAli [00:15:01]: I don't know how this works.Philip [00:15:02]: So, like we had to, like, build support for that into our runtime. And you're right, like it is really interesting the way that all of these open source labs borrow from each other. For example, like GLM-5.2 doesn't have vision. So something that, Haley, a guy on our team, if we could take a look at this, he, like, grafted the Kimi vision encoder onto GLM-5.2.Retrofitting Vision into GLM-5.2Ali [00:15:27]: We'll be training the projector.Philip [00:15:28]: Exactly. So if you think about, like, the encoder, there's the encoder, which is the part that looks at the image and turns it into latent information, and then there's the projector which likeAli [00:15:38]: You can say latent space. It's okay.Philip [00:15:41]: And then there's the projector that maps it onto, the model itself, and then there's the model weights. You don't wanna mess with the model weights because you run a chance of making the model dumber at something else for the purpose of giving it vision. So instead, Haley started with just a projector, which is only a handful of millions of parameters.Ali [00:16:02]: That would be, yeah.Philip [00:16:02]: Yeah.Ali [00:16:03]: Can you show the training one?Ali [00:16:04]: Like the way it groksPhilip [00:16:05]: YeahAli [00:16:06]: Very interesting.Philip [00:16:06]: And maybeAli [00:16:07]: That right therePhilip [00:16:07]: Maybe Ali, you should take it from here. You've got a betterAli [00:16:10]: Ooh, double the sandPhilip [00:16:11]: Understanding of this than I do.Ali [00:16:11]: Yeah. You can see, like, he. The way he trained this is really cool. At the beginning, he was training it using just like, “Here's a picture of a mountain. Can you describe what's in this mountain?” And that caused it just like the first, learning walls. Like here you can see this all we're trying to teach it is to translate the encoded. Like it's already taken the encoder from Kimi K. It's taken the image. It'Philip [00:16:31]: Yeah. FrozenAli [00:16:31]: FrozenPhilip [00:16:32]: With adapter.Ali [00:16:32]: Exactly.Philip [00:16:33]: Yeah.Ali [00:16:33]: So the brain is frozen and the eyes are frozen. It's just we're tryingPhilip [00:16:37]: AlignAli [00:16:38]: Interconnect between the eye and the brain, right? So the projector. And so you take the tokens and then he's like, “Oh, can you describe what's in this image?” And he's like, “Oh, it's a mountain,” or it's a person or it's a human, whatever the case is. But that didn't cause complete understanding. So he changed it such that every image was associated with a data set of questions. Like, does this image have a white male? Does this image have birds in the top corner? Does this image have a scientist in it? All of that stuff. And it would have to answer questions correctly. And using not just training on describing an image, but being able to answer question, another question, answer over time. Like you can see the grokking, which is like genuinely insane, that retrofitting vision into a large LLM can learn to that extent. And even for images that it doesn't perform well on, for instance, if you ask it a picture of like Stephen Hawking, “Who is this?” Maybe it doesn't get it, but it will say something like, “This is Albert Einstein.” Like it still understandsPhilip [00:17:25]: Close enoughAli [00:17:26]: That this is a scientist who is a man who has, some significant achievements, all that stuff. So that's like really cool.Philip [00:17:32]: Yeah. So, we've covered Hao Tian before, who the author of the LLaVA paper that did this, a while ago. And I think that's very foundational work for anyone who hasn't done vision work before.Ali [00:17:41]: Same with the CLIP and MetaCLIP, where you go from just captioning to building out questionsPhilip [00:17:47]: RightAli [00:17:47]: Off the image and how much better you can get performance.Philip [00:17:50]: Right. Right. Right. Yeah. But what's, what's so exciting about this is if you look at a model like this. Now, this is a little bit more of a research project. It's not. It got to 56% on MMLU Pro, I think. So not quite frontier. But if you're running this model, you haven't suffered any loss on your GLM-5.2 quality. If you don't have an image, it'll just behave exactly the way it used to. And ultimatelyAli [00:18:14]: Which in the inference code you literally do not include the other part, right?Philip [00:18:18]: Yeah. You would just skip the encoder if you don't have an image input.Ali [00:18:22]: Okay.Philip [00:18:22]: Just confirming.Philip [00:18:23]: YeahAli [00:18:23]: Does it affect a lot on the overall inference side? Like you're not adding much, you're adding a very small vision encoder. These are typically likePhilip [00:18:30]: They're super fineAli [00:18:31]: Less than a billion parameters, right?Philip [00:18:32]: Yeah. It's, - There's a little bit less standardization among vision encodersSwyx [00:18:37]: YeahPhilip [00:18:37]: So the support matrix can be a little bit, sparser. But overall, yeah, it's a pretty, it's a pretty minor component of the overall system. And ultimately what you get out of the system is all of a sudden you have Kimi Vision, GLM weights, and DeepSeek attention all in one model.Open Source Model Grafting and Franken-MergesPhilip [00:18:56]: And that's, I think, a lot of the power and beauty of open source, is that you can take all of these different components and combine them together into a system that's better than anyoneSwyx [00:19:05]: YeahPhilip [00:19:05]: Can be individually.Swyx [00:19:06]: People used to say that you would also do Franken-merges where you would take likePhilip [00:19:10]: YeahSwyx [00:19:10]: Layers from each model.Swyx [00:19:11]: Does anyone do that anymore?Ali [00:19:13]: Well, to your point previously when you were mentioning like, the work that goes into supporting a model when it first comes out, like GLM-5.2 or MiniMax M3 or whatever the case is. Sometimes you do have to like, you do have to switch out some things. Like, for instance, the MiniMax M3 head uses full attention, and with full attention you end up with this like insane bottleneck in spec dec ‘cause you're doing auto-regressive token generation for three tokens, and you're doing this like N squared over all of the tokens that are in your sequence. Your KV cache is like very large because it's not sparse, it's not top K. So we find it better to like, okay, we're gonna replace this, we're gonna replace this layer with a layer from another model that's using like GQA, for instance. And then just with the right training, you can get it to have the same acceptance rate. So it is very possible to retrofit layers from other models and very much needed. If a layer is like inefficient, the training just becomes the challenge, like how do you ensure that you train it properly? Which again to your earlier point is like the mesh between training and inference. As in like you need very good training in order to do fast inference. That's like, I feel like more and more becoming true.Swyx [00:20:21]: Yeah. Anything else on the support side when you say like get it to fully production ready?Loop Detection, Race Conditions, and Non-DeterminismPhilip [00:20:26]: Yeah. I think that there's also a question of just, we can test a model to a pretty extensive degree, but we're trying to get it out quickly and then you see a bunch of other people test it and you get interesting results. There was an issue with, GLM briefly where we had some like mode collapses where it would just output the same token over and over again for certain prompts on certain temperatures. Like once you expose an endpoint to the real world, there's going to be, so many more varieties of things given to it that you're able to, discover and patch things. So it's not just a, day zero process, it's then like for the first week, for the first month, if a model remains popular, like how do you both fix bugs and then continue to push the envelope on performance?Ali [00:21:21]: What do you mean you don't want your model outputting S?Swyx [00:21:24]: Is there loop detection on that stuff, by the way? It still happens like quite a lot, which is surprising.Ali [00:21:30]: We have like we, in our endpoint, like if a model was to output the same token like four plus times, we just cut the generation. We say like, “Oh, sorry, this-- Like try again,” or like we will reprocess the request. ‘Cause we know then, like if it, like if, yeah, it's four times the same token, it's probably collapsed.Swyx [00:21:45]: Yeah. Is there a way to opt out in case I really want that?Ali [00:21:48]: You want that?Ali [00:21:50]: I think there's a way that we have to handle it. I'm not exactly certain, but I feel like in certain models, like when they output something like you can imagine, like a table for instance, and so they want, they wanna draw like 12 dashes and 12 dashes. Yeah, I think there's a way for that to happen. I think we only do it on certain tokens. Like we exclude certain special characters.Swyx [00:22:07]: Yeah.Ali [00:22:07]: So we only do it on like certain like S is the most common almost. GLM-5.2Swyx [00:22:11]: OhAli [00:22:11]: And I think it was DSV 4 as well. Like you'd just have like looping issues where like you literallySwyx [00:22:17]: ItAli [00:22:17]: Just have like S.Swyx [00:22:18]: Yeah. Is there a special, something special about S? No, just randomlyAli [00:22:21]: It just seems to be the one token involved.Swyx [00:22:23]: Yeah. And it'Philip [00:22:24]: Is thereSwyx [00:22:24]: And it's only temperature 0Ali [00:22:27]: NoSwyx [00:22:27]: Even at other temperaturesAli [00:22:27]: Even at like 0.9 or whatever, it will still, it will still collapse.Swyx [00:22:30]: That's weird, right?Ali [00:22:30]: It's, it is an inference problem to be honest, like a software problem. Like oftentimes, the image you run will-- like NVIDIA will release an image for instance, and if we will upstream the changes from their latest TensorRT-LLM image into our stack, we'll find that it fixes it. Or oftentimes this will only happen in an inference engine that you're using like SGLang. But if you were to switch to vLLM, that isn't the case. So it seems to be like an extremely like deterministic software issue and not really a model issue. It's not like a weights problem. Like I'- we'll say like, “Oh, it's a problem with the quant. We did PTQ wrong,” right? But that isn't, that doesn't make sense because the same weights used with a different inference engine does not repeat the problem. And sometimes it's, the kernels that are being used in the backend have like these very subtle sometimes race conditions, where if you were to use this model hosted on one cluster, you will never get this problem.Swyx [00:23:19]: Oh my God.Ali [00:23:19]: But if you host it on a different cluster, you will. And the reason is the KV cache transfer from a node to node in that one cluster is using a slower interconnect than the node to node in another cluster. So that exposes the race, whereas in another cluster it doesn't. So then you end up just like, okay, this model is not gonna be hosted on this cluster. We're gonna host it on, another cluster because that cluster exposed that problem. But then it ends up with like, okay, is it the software? Is it the model weights or is it the hardware?Swyx [00:23:42]: There is a thing about this with temperature 0 still not being deterministic, right?Ali [00:23:46]: Right.Swyx [00:23:46]: Mostly because of hardware. Even at temperature 0 same model, you won't always get the same output.Swyx [00:23:52]: Even-- But I'm surprised by the race condition one because, I thought PyTorch was a graph that like guarantees that you at least, execute things in the right order.Ali [00:24:02]: Well, yeah, true. Like I'm not, I'm not saying that there is. Like well, you have things like PTL optimizations where like you can start a kernel before the end of the previous kernel, and that's like ‘cause you want to do that because there'sSwyx [00:24:12]: It's like pipeliningAli [00:24:12]: Expense. Exactly.Swyx [00:24:13]: Yeah.Ali [00:24:13]: But it'- But you don't do it cleanly. Like you overlap a little bit of the execution. No, it is very possible that the kernel itself, like that one block that is supposed to be running in this instance of time, that kernel itself has a race condition. For instance, like a missing barrier. Like often if you're designing a kernel and you want it to make it to be very fast, if you don't test it extensively, you'll, you'll have certain threads access data points from registers before they've been written to by other threadsSwyx [00:24:36]: YeahAli [00:24:36]: For example, because like your barrier is wrong or your synchronization was wrong. But yeah, like the testing itself is very difficult in those like, andSwyx [00:24:42]: And there's no like borrow checkerAli [00:24:45]: What does that mean?Swyx [00:24:46]: Like Rust. Like the. If you're trying to have like memory safety It sounds like a comparable problem.Ali [00:24:52]: Well, yes, but you're working in CUDA, right, NVIDIA GPUs. Like- You just need a higher level language like modular Maybe that's what modular is supposed to do. I don't know.Quantization Quality and Vendor FidelityVibhu [00:25:00]: How do you see keeping quality of the model? So you talked about all these steps of, okay, you gotta do quantization, train your own speculative decoderAli [00:25:07]: RightVibhu [00:25:07]: Run on different hardware. Looking at other model providers, okay, you kicked off a inference speed race on the consumer end. What goes into keeping quality the same across them, right? Sure, you can run benchmarksAli [00:25:22]: YeahVibhu [00:25:22]: But, like, how do you determine how much quantization are there standards? What goes intoPhilip [00:25:27]: There's a few things on quality. Most inference optimizations are lossless. KV caching, for example. You are just recomputing or preventing recomputing the same values. Speculation, of course, if a draft token is wrong, it gets rejected. The main lossy optimization is quantization. And that really comes down to, number one, data format, number two, which parts of the model you choose to quantize, which layers, and number three, like doing a lot of calibration on the quantized weights, to ensure that you're preserving all the outliers. There's other tricks that you can do, though. A big one is long context, ‘cause one thing you asked at, right at the beginning is, “Oh, what's gonna happen if I send a 200,000 token request in?” So with a long input sequence, you need to, store a lot more information. You need to process a lot more tokens. And so even if a model has a context of a certain length, you might, as an inference provider, choose to build an API with a shorter context length, and of course a full length one as well. Because if someone doesn't need the full million token context, for example, you can get them better performance. I don't know if that's exactly like quality of the model. The way that I think about quality is to what degree are we faithfully serving the original model? If you think of a golden implementation of a model that performs exactly the way the model is designed to perform, I think of quality as how close are we getting to that, 100% fidelity of the model.Philip [00:27:13]: You can also, of course, think about quality from the training side and how do you push yourself past 100%. But when I think about purely inference optimizations, it's getting faster while staying as close to that 100% fidelity mark as possible. And certainly our standard internally is that, like you should not be able to tell the difference between our API and a, official API. I think Kimi in particular does a good job of vendor benchmarking hereAli [00:27:41]: YesPhilip [00:27:41]: Where they haveAli [00:27:42]: They released an actual vendor benchmark.Philip [00:27:43]: Exactly, yeah.Ali [00:27:44]: ‘Cause they accused, some people, Amazon? There was some provider that was not doing very well on Kimi's benchmark.Philip [00:27:50]: Yeah.Philip [00:27:51]: So, with Reflect we probablyVibhu [00:27:52]: This was a long time ago, right?Philip [00:27:54]: No.Ali [00:27:54]: Yeah, like threeVibhu [00:27:55]: They alsoAli [00:27:55]: Four, five months agoVibhu [00:27:57]: This also happened with, I don't remember which model, but they pulled out quite a few, and then they started a whole chart about this. It might have beenPhilip [00:28:03]: Kimi Vendor Verifier.Ali [00:28:04]: Yeah.Philip [00:28:05]: Yeah.Ali [00:28:05]: Yeah, ‘cause you, ‘cause you'd be pissed, right? Like if you'Philip [00:28:07]: Yeah.Ali [00:28:07]: If like if I'm a consumer and I'm using like Amazon's endpoint for instance, and I've used Kimi and I'm like, “Oh my God, like this is bad,” I'm not gonna say, “Oh, Amazon quantized the model in a bad way.” I'm gonna say, “Oh, Kimi sucks.” Right?Philip [00:28:17]: Yeah.Ali [00:28:17]: So it seems like that makes sense.Philip [00:28:19]: Yeah, they care. They care.Vibhu [00:28:21]: Justifiably.Ali [00:28:21]: Yeah, justifiably.Vibhu [00:28:22]: This is probably a stupid question, but just checking, has anything improved from main quantization?Philip [00:28:28]: Yeah.Vibhu [00:28:28]: Like, is quantization always strictly worse?Ali [00:28:30]: Well technicallyVibhu [00:28:32]: NoAli [00:28:32]: It's a lossy. QuantizationPhilip [00:28:33]: YeahAli [00:28:33]: Is a lossy, it's a lossy implementation.Philip [00:28:36]: Speed improvesVibhu [00:28:36]: Speed improves.Ali [00:28:37]: It the number, likeVibhu [00:28:38]: No, I' always look for inverse scaling laws.Philip [00:28:40]: Yeah.Ali [00:28:40]: Yeah.Vibhu [00:28:40]: This is something I learned from Noam Brown, where like things that normally act in one direction sometimes do.Philip [00:28:45]: Well, technically when you run a benchmark, because these models are deterministic, sometimes your,Ali [00:28:52]: YeahPhilip [00:28:52]: NVFP4 quant is like, two basis points higher than yourAli [00:28:56]: No, it's noise. It's noise.Philip [00:28:57]: Yeah, exactly. I'm like, yeah, it's, it's within. That's why I always say within margin of error.Philip [00:29:01]: And I stopped saying that because everyone assumes that what is, well, within some margin of error, we're barely inside of that to the worst, so we're saying. But yeah, sometimes it's just like, gives you a higher output score. But like Ali said, that's noise. To my knowledge, you're not necessarily making the results better. You're just trying to, again, like keep your fidelity as close to 100% to the original model.Layer Selection, KL Divergence, and Better QuantizationAli [00:29:27]: There is, to your point, research that we did on MP. I don't know if you are able to pullPhilip [00:29:31]: YeahAli [00:29:32]: A tweet we did. One of our research interns, Joshua, I think it's a tweet on how we have 20% better quantized GLM-5.2 than NVIDIA. Essentially what we found throughout like this month research is, okay, quantization is a lossy. It's. You're compressing the data from, occupying 16 bits to occupying, four bits, for instance. And so you're losing some information, and you're trying to minimize that. And so when I say that I'm gonna quantize the model, my job becomes how do I find the layers that I can quantize, and how to find the layers to not. For instance, with image models, I don't quantize modulation layers, and I don't quantize out projections because those two are. Like out projection is what you see as the user. Modulation is what the model sees or understands. Right, exactly. And so to his paper, do you have the. It doesn't have the. Yeah. It's a long paper. I don't know if I can findVibhu [00:30:25]: If there's a part to search or it's probably in the thread.Ali [00:30:28]: It's probably in the thread.Vibhu [00:30:29]: Yeah.Ali [00:30:29]: But the long and the short is it is very possible that quantizing more of the model makes the results. Like if I have a model that I quantize layers one, five, and 10, and another model where I only quantize layers one and It is possible that the model in which I quantized more information is going to perform better because the quantization errors have canceled out. And so what Joshua showed in his mathematical proof where he had like a verifier in, is that you can predict which layers are going to have quantization errors that will cancel out with each other, and you choose to quantize those layers. And so the result of doing this mathematical quantization is you end up with a model that's 20% more quantized than another provider, so you get 20% more throughput of it because there's more layers than running an NVFP4, and your quality is better than that other quant because the layers that you chose to quantize have their errors cancel out, like one layer skewed to the right one layer skewed to the left, one layer skewed to the right. Your final logits distribution is more similar to the original distribution of the model, so you have better fidelity. And so the way we proved this was with KL divergence. So instead of just scoring on the benchmarks, we scored the KL divergence between the logit distribution of the quantized model and the logit distribution of the original full precision model, and we showed that with this technique we get. If your probability distribution on the logits which token it wants to select is more of the same as the original model, you're probably gonna end up staying true to the original model. So yeah, so it seems like previously before this, it seemed like the industry was, well, the more you quantize, the worse it's gonna be, ‘cause the more loss you introduce. That's not exactly, not necessarily true. So yeah, doesn't improve it, but can cancel out.Philip [00:31:57]: I think it might be this, but reminds me a good bit about pruning where you can prune off certain layers.Philip [00:32:03]: But very interesting. Didn't know this was a whole paper you guys put out.Ali [00:32:06]: It's. Fun fact, it was originally 72 pages, this paper, and then we decidedPhilip [00:32:11]: WowAli [00:32:11]: We can't tell. We couldn't release it. So it's now 45.Swyx [00:32:15]: Still 39 pages, so very substantive. We talked about evals and all these things and, like what's possible in terms of speedup? Like it's like probably like the numberInference Speedups and BenchmarkingSwyx [00:32:25]: Thing that people do wanna care about, and it's something that you wrote about in your post. Like official API is 70 tokens per second, and you push it up to 90. Is that like a normal thing?Philip [00:32:36]: So what's cool about working in inference, the reason that I think inference is going to be a useful place to do engineering for a long time, is that if you look at highly optimized domains like, say, finance, if you're in finance, you measure how much better you got in basis points. It's like, “Oh, I got five basis points better, like twentieth of 1% better,” that's huge news because everything is so optimized. When we publish optimizations, it's 20%, it's 100% it's 200%. So there's still probably like a lot further to go, honestly. Like you'll, you'll know that inference is pretty much solved when researchers start publishing about how they got 1% faster at something.Swyx [00:33:19]: Which by the way, because I am from the finance background, in the ‘70s, that was the margin at the time. When you did quantitative finance research, you would findAli [00:33:27]: And like 20%, tens of percent.Swyx [00:33:29]: That's. Yes.Philip [00:33:29]: Yeah.Swyx [00:33:30]: And now it'Philip [00:33:31]: Tiny fractionsSwyx [00:33:32]: For those people interested, look up Andrew Lo's paper. He had a really interesting illustration of quant, stat arb, distribution, narrowing down from like those kinds of 20% differences in the ‘70s, down to nothing today, which is very cool.Philip [00:33:48]: Exactly, and we're at the beginning of the same type of thing. Now benchmarking is hard. I think anyone will tell you that, and benchmarking provider speeds is hard because there's so many variables that go into it. What hardware are you using? How much load do you have on the system? What's the exact nature of the prompts and input and output sequence lengths? All that stuff. But overall, when you start stacking these improvements, you're looking at multiples. You can look at it. The most common form, of course, is TPS, tokens per second, which is bad naming by us in the industry, ‘cause there's two tokens per second. There's tokens per second, the throughput number, and the latency number.Ali [00:34:31]: TTMT, yeah.Philip [00:34:32]: Like total tokens per second out of the, out of the GPU as a throughput number. Most people only care about tokens per second as the latency number, which we should call ITL, intertoken latency, but we don't.Philip [00:34:44]: Anyway, so you can imagine a standard API without many optimizations for a 1 trillion parameter model operating somewhere in the 30 to 50 tokens per second range for reasonable traffic profile. And we generally see the goal of, pushing to 10X that. But, not necessarily day zero, but by stacking enough optimizations, if you have, say like four optimizations, each of which doubles performance. Or sorry, three optimizations, each of which doubles performance, then you stack that up, that's an 8X gain. That's the order of magnitude that we're working with in this space. We're trying to make things substantially faster, not just go from like 70 to 90.Swyx [00:35:38]: Are you saying you've. You have done that?Philip [00:35:40]: So let's say you have as a reasonable baseline, 30 or 40 tokens per second. You can achieve 10X that. So like on GLM-5.2, if you run it unquantized, perhaps on H100s even, and you're just using an off-the-shelf inference engine with no particular optimizations, no speculator, nothing extra around like KV routing, no disaggregation, you're, you're probably, yeah, looking at that like 30 to 40. You think that's like a reasonable baseline?Swyx [00:36:12]: Right. Right.Philip [00:36:12]: To get to something like 10X, there's a lot of trade-offs that you're making. If we're running at more like a 300, 400 tokens per second range, you are using the best hardware possible. You have a optimized speculator. You have done all of your quantization work. You are Seeing a pretty high cache hit rate. You are running with a reasonably small batch size and a parallelism configuration that is tuned for latency versus throughput, but it is possible. So the spreads that you see if you, like, go on artificial analysis or you go on OpenRouter and you look at, the worst provider to the best provider, oftentimes can hit that range. 10X is of course very aggressive. It's oftentimes maybe more of a four to six times improvement. But that's the performance that makes us really excited, is when we can get these huge gains, not just go from 70 to 90 tokens.Stacking Optimizations: NVFP4, Speculation, and DisaggregationAli [00:37:19]: It's also, like, hardware dependent. Like, ifPhilip [00:37:20]: YeahAli [00:37:20]: If you have a thing where you're serving it on just, like, a node of H100s and then you throw, like, you shard the model across, like, four nodes of B200s. Like, you can definitely increase the speed with just throwing more hardware at it. Like, normalizing for the same exact hardware and the same number of GPUs.Philip [00:37:35]: Yeah. Then you're looking at, like, a two to 4X improvementAli [00:37:38]: Right. RightPhilip [00:37:38]: Depending on the inference optimizations. So yeah, it's. Some of it's, what's the call, and some of it's who's the driver.Vibhu [00:37:46]: If you break down the two to 4X, say the example is run GLM-5.2Ali [00:37:51]: YeahVibhu [00:37:51]: On B200sAli [00:37:53]: YeahVibhu [00:37:53]: Single node, right? What's, like, the cost trade-off for effort to get, like, the last bit of juice out versus what should people just think of, right?Ali [00:38:01]: Spectre quantization. Yeah.Vibhu [00:38:03]: Spectre quantization.Ali [00:38:04]: That's, that's, that's like 95%. LikeVibhu [00:38:06]: And how far does that get you? And how easy is that for the average person to do? So say right I wanna throw the weights of GLM-5.2 on a node of B200s, how easy is it to find speculative decoder- decoder model or already quantized model? How much work goes into it?Philip [00:38:23]: If you're doing it up front, it's quite a lot of work. If you're doing it today, there's going to be people who have published things that you can just, you can just grab some NVFP4 weights. You can grab a speculator. Yeah, if we're thinking about, like, what are the 2Xs we're stacking, going from, BF16 to NVFP4 is, it's not quite a 2X, right? It's like. I think it's about, like, 30 to 40%, from 16 to 8, and then another 30 to 40% multiplied from, 8 to 4. So that doesn't quite get you a 2X, but, like, roughly a 2X. Speculator, roughly a 2X. Disagg on top of that if you're able to get enough hardware and put enough traffic through it, another roughly a 2X. And then you add in some, double-digit percent increase from having just a better runtime with, the latest kernels and stuff behind it. And that's how it stacks up.Ali [00:39:21]: YeahPhilip [00:39:21]: So building each of those, like, building the, quantized weights is, for someone who really knows what they're doing, hours to days of work. Building the speculator, again, like, hours to days of work. And the, disagg setup, hours to days. Well okay, but like once you haveAli [00:39:39]: Once set up. Once set up. YeahPhilip [00:39:40]: Yeah, getting disagg working for the first time, I'm saying, of course, is very difficult.Philip [00:39:44]: The marginal implementationAli [00:39:48]: Like, if you're just grabbing, like if you are a person, like just a normal consumer who has access to, like, a node of B200s and you're wondering, “How can I just host it myself?” You don't need to quantize the model yourself. There's always gonna be, like, an open source quantized checkpoint. NVIDIA's gonna push one out if no one else does. You. Usually, the providers will have their own spec dec that they've trained as well. You don't need to train your own spec dec. You can just use that as well.Philip [00:40:09]: Yeah. Like, GLM-5.2 has its own MTP.Ali [00:40:13]: Right. Right.Vibhu [00:40:14]: What's multi token prediction?Philip [00:40:15]: Yes.Ali [00:40:16]: I'm justVibhu [00:40:16]: Can you explain that?Ali [00:40:16]: I'm just an expert.Ali [00:40:18]: I can do it for you in case I get it wrong?Vibhu [00:40:20]: No.Vibhu [00:40:21]: Yeah, you should correct if we're wrong, but their multi-token prediction can be used for self-speculative decoding.Ali [00:40:27]: I'm not sure. I'm not gonna correct that.Vibhu [00:40:28]: Okay. I'm semi-confident in thatAli [00:40:30]: Okay. YeahVibhu [00:40:30]: But someone can check. But it's useful to paint the story of, okay, not just the average person, but say a company wants to switch from serverless inference I wanna throw this up on. I wanna rent some GPUs, throw it up. These are the steps you take to do significantly faster than just put it behind vLLM.Ali [00:40:48]: Right.Vibhu [00:40:49]: I was waiting for a mention of Dynamo.Vibhu [00:40:51]: I feel like, that's supposed to be the baseline that you measure against.Dynamo, KV Routing, and Disaggregation ToolkitsPhilip [00:40:55]: I would think of Dynamo as less of a box system and more of a toolkit for building with. So when we talk about doing aware routing, when we talk about doing KV offloading, when we talk about doing, PD disaggregation, Dynamo fundamentally is. By the way, Dynamo is an open source library from NVIDIA.Ali [00:41:17]: We've done a pod with KylePhilip [00:41:18]: OkayAli [00:41:19]: Kyle Cranin.Philip [00:41:19]: Cool. So then your listeners know then that it supports all the different inference frameworks. And it is multi hardware, which is interesting.Ali [00:41:28]: But it's just a router, it's not like an optimizer layer.Philip [00:41:30]: Yeah. All it does, like, what Dynamo is good at, it is a library for moving information around your cluster, around your hardware. So if you have, KV cache on one place and you need it to be somewhere else, Dynamo coordinates NIXL for you to move that around.Philip [00:41:49]: That doesn't mean that, like, out of the box, you just say, “Pip install Dynamo,” and then you get, like, a massive performance speed up. It's more of a developer toolkit.Ali [00:42:01]: Yeah. I would have said it would. It comes with a set of defaults that you can then swap out.Philip [00:42:06]: It does. If the industry at large, I think, was, like, rolling out all of these deployments, standard, then I think it would be, like, a credible baseline. But, we've got to, we've got to benchmark against, like, what we're seeing in the wild.Speculative Decoding Methods: Medusa, EAGLE, n-Gram, and Spec-SpecVibhu [00:42:23]: I did wanna talk a little bit more about PD disagg, because that is probably, like, number three after quantized and speculative decoding. In your book though, I was just gonna pull out the book.Philip [00:42:31]: Yeah.Vibhu [00:42:32]: Like section 522 on Medusa, 523 on EAGLEPhilip [00:42:35]: YeahVibhu [00:42:36]: 524 on gram.Philip [00:42:37]: It's 55, would be disaggregationAli [00:42:42]: Yeah. Well, no, I just wanted to dwell a little bitPhilip [00:42:44]: YeahAli [00:42:44]: The other. Like, so what do you choose to include? What do you choose to not to include? Because there was all these other techniques.Philip [00:42:51]: Yeah.Ali [00:42:51]: Are these still relevant? Because I think they came out, like, a year and a half ago maybe.Vibhu [00:42:55]: Medusa is quite old.Philip [00:42:56]: Yeah, Medusa's old.Ali [00:42:58]: It was old.Vibhu [00:42:58]: But is it in the book as a good, here'sPhilip [00:43:01]: BaselineVibhu [00:43:01]: Baseline vanilla understand it?Philip [00:43:02]: Like you should know this.Vibhu [00:43:03]: Like I read the paper, I'm like, “ it makes so much sense.”Philip [00:43:05]: Yeah.Philip [00:43:05]: So with the book, I had a couple goals. One was to give people just a working vocabulary for the space as a whole, and the other was to give them some intuition about how each of these techniques works. As I mentioned in my AI Engineer talk, which is the first public addendum to this, the speculation space has moved much faster than everything else. So yeah, even at the time that I wrote the book Medusa, I very much included as a way for people to understand how the space evolved rather than what the most modern technique is. And now of course, there's DFlash, dSpark. There's, there's newer techniques even than EAGLE, although EAGLE is still very commonly used.Ali [00:43:51]: SpecSpecta.Philip [00:43:52]: Yes. Speculative decoding.Vibhu [00:43:54]: What canAli [00:43:56]: Oh, it's a paper by Tri Dao and it's like, it's doing speculative decodingVibhu [00:44:00]: HuhAli [00:44:01]: For the speculative decoder.Philip [00:44:02]: Oh, in spec- oh my God.Ali [00:44:02]: It's literally just an another. It's like, yeah, that's the most simple way to explain it, and it seems like he got trivial speed ups there. But it seems that the complexity with training, it's almost like in our mind at least, it's almost as complex as training GANs. Like it's like a very delicate balance and oftentimes you, it's just but yeah, it's literally speculative decoding on speculative decoding.Vibhu [00:44:21]: Speculative.Ali [00:44:22]: Yeah. We saw this paper.Vibhu [00:44:24]: It's interesting, right?Ali [00:44:24]: Yeah.Vibhu [00:44:24]: I wouldn't even expect it to be very particular to train, I wouldAli [00:44:29]: Right.Vibhu [00:44:29]: The naive part of me is like, okay, train speculative decoder.Ali [00:44:32]: But like, and it makes sense, like the whole idea of speculative decoding is you. It's like, it's like almost like the iPhone auto predict version but for a normal model, right? Like you're just, you're just, generating three tokens and you're like, okay, I'll do prefill on them. And so you save those three turns for your original model. Now your speculative decoder is doing three turns of auto regression, so why not just have an even smaller model?Ali [00:44:53]: The other question there is what are the size of speculators? So say forPhilip [00:44:58]: Right. It's like a billion parameters.Ali [00:45:01]: Like for MiniMax, it's. Yeah. It's like one layer. It's like one 60th of the original model usually.Philip [00:45:06]: Yeah. I think we should do a paper when we get back to the office.Philip [00:45:10]: SpeculativeAli [00:45:11]: SpeculativePhilip [00:45:11]: Decoding.Ali [00:45:13]: No, it's, it does seem like how, when do you stop? But then it also seems like if you're able to train spec-spec decode for instance, right? Like if you're able to have a small model that is accurately predicts what the intermediate speculator is gonna predict, that is able to predict what the original target model's gonna predict, then why not just use that smallest model directly, right?Vibhu [00:45:34]: Yeah. This isAli [00:45:35]: Like it seems likeVibhu [00:45:35]: Adjacent to the routing problem.Ali [00:45:36]: Right.Vibhu [00:45:36]: Yeah.Ali [00:45:36]: Right.Philip [00:45:37]: The thing with speculators is one of the practical constraints on using them is that you do have to run a small model on the same hardware that you're running the big model on. There is a orchestration and resource competition problem inherent in that, and that is one of the constraints on speculation in general, is that draft tokens cost resources to create and cost software complexity to manage. And so if you have like infinitely recursive speculators, you add in quite a bit of that complexity on the actual implementation within the inference engine as well, not just in the training process.Vibhu [00:46:17]: I was gonna say, I would wonder if you could do similar, like distillation and pruning of, it's the same thing, it's just a model. Can we not just distill a lot of the weights, quantize the speculator, out of my domain? The question that also comes up is, this is all for big server workloads, right? How much of this applies to, say I have this MacBook, I wanna run Gemma really efficiently. Similar problems, not the same?Local AI vs. Data Center InferencePhilip [00:46:45]: Pretty different. I talked to Selo, about this on his podcast a couple weeks ago. The difference between inference engineering for the data center and for production workloads versus inference engineering for local AI, is that we start with fundamentally like different constraints and different goals. With local AI, it's how do I fit this model onto my hardware and then make it less dumb? And with data center influence, it's how do I load this model and then make it less slow? And we care about less dumb, and they care about less slow. But the local AI inference engineering ecosystem, I think has a lot for us to learn from in the data center space. They are experts in various forms of quantization, including dynamic quantization that we just don't touch, in the pruning, in the distillation, in the, layer removal. There'Ali [00:47:42]: Layer removal matters less.Philip [00:47:43]: Yeah. There'Ali [00:47:44]: No one loves pruning really.Philip [00:47:45]: Yeah. Well, but the, but they doVibhu [00:47:46]: Which is surprising, right? But that's, that's a whole different thingPhilip [00:47:48]: Just to fit something on the laptop.Ali [00:47:50]: Right.Philip [00:47:50]: So yeah, it's a, it's an interesting, it's an interesting space. Not necessarily that like their techniques make sense for us to do in the data center, because we have different resources and different goals, but more that the process as well as the openness of that field is something to, admire.Ali [00:48:12]: Yeah. Like to your point, like, certain optimizations that would. Like for instance, Turbo Quantum Sharper, like it made such huge hype on that and we did like a whole deep dive on Twitter and like said, what is it? How does it work? Why is it good or not? And it took off and it was implemented on local devices because your memory bandwidth is so slow on like a MacBook, for instance. But try putting the same thing on like an NVIDIA GPU on a B200 Turbo quant would not be. Like, it would not be used. Like, NVIDIA - Like, NVIDIA made it clear that this is not a good optimization, and we've seen it firsthand where the overhead of doing dequantization, quantization of, in the kernel itself with turbo quant kernel, each end is much slower than the time that you save from doing the bandwidth. ‘Cause on the B200s, you have like 3.5 terabytes per second. You don't need decrease the storage that much. You don't need to do, FP4 KV cache. You don't need to use a requant. There's, there's, there's better optimizations to be made. But on Edge devices, it's extremely important, it's extremely useful. So, seems to be, like, different optimizations there, but then they're all uniquely combined with like all you wanna quantize the model, you wanna do speculative decoding, like certain common prefixes with bothPhilip [00:49:18]: Principles.Ali [00:49:19]: Yeah, exactly. Exactly. Exactly.Philip [00:49:20]: They also do a lot of work on, model parallelism, especially over, heterogeneous topology, where you have, some sparks and they are wired together with, Ethernet, DGX sparks.Ali [00:49:35]: Yeah, this is the Exo Labs guys.Philip [00:49:36]: Yeah. You have, a nu
(0:00) Bestie intros (1:19) Chip stocks crash, Leopold Aschenbrenner's $20B fund gets margin called (20:20) China's advantage and green shoots for the US economy (34:12) Frontier Labs say "SLOW DOWN AI" (1:01:15) Why are frontier labs "burning books"? (1:14:45) Socialism Corner: Mamdani's grocery stores and the "Socialist Spectacle" (1:24:44) Science Corner: Understanding and mapping the brain Apply for All-In Summit 2026: https://allin.com/events Follow the besties: https://x.com/chamath https://x.com/Jason https://x.com/DavidSacks https://x.com/friedberg Follow on X: https://x.com/theallinpod Follow on Instagram: https://www.instagram.com/theallinpod Follow on TikTok: https://www.tiktok.com/@allin Follow on LinkedIn: https://www.linkedin.com/company/allinpod Intro Music Credit: https://rb.gy/tppkzl https://x.com/yung_spielburg Intro Video Credit: https://x.com/TheZachEffect Referenced in the show: https://www.cnbc.com/2026/07/30/leopold-aschenbrenners-hedge-fund-is-facing-steep-ai-losses.html https://www.wsj.com/finance/citadel-buys-situational-awarenesss-stock-portfolio-after-big-losses-in-ai-5117159b https://polymarket.com/event/fed-decision-in-september-762 https://situational-awareness.ai https://www.cnbc.com/quotes/US30Y https://x.com/nicolasfulghum/status/2082083884578050299 https://theprint.in/science/breakthrough-china-artificial-sun-project-6-5-tesla-magnet/2999321 https://www.pacingthefrontier.com/ https://www.cnbc.com/2026/07/29/openai-cfo-sarah-friar-tells-employees-arr-in-july-topped-all-of-q2.html https://www.dwarkesh.com/p/why-compute-might-get-10x-more-expensive vhttps://www.reuters.com/world/china/china-starts-production-home-grown-immersion-duv-chipmaking-tools-source-2026-07-28 https://www.google.com/finance/beta/quote/ASML:NASDAQ https://www.cnbc.com/2026/07/27/cxmt-china-market-debut-chipmaker-ipo.html https://polymarket.com/event/ipos-before-2027 https://polymarket.com/event/us-enacts-ai-safety-bill-before-2027/us-enacts-ai-safety-bill-before-2027 https://x.com/v_nefodov/status/2082927219224060043 https://www.wsj.com/opinion/the-ai-future-is-for-everyone-a0c24e20 https://punchbowl.news/article/tech/thune-anthropic https://www.wsj.com/tech/ai/anthropic-doubles-midterm-spending-to-40-million-to-push-ai-regulation-9cd547ae https://x.com/OpenAI/status/2082577277246972300 https://www.carltonfields.com/insights/publications/2025/no-copyright-protection-for-ai-assisted-creations-thaler-v-perlmutter https://x.com/Jason/status/2082577230941557068 https://www.tomshardware.com/tech-industry/artificial-intelligence/ai-companies-are-reportedly-shredding-millions-of-books-to-train-models-tech-giants-outsource-to-middlemen-to-secretly-buy-up-books-for-training-material https://www.404media.co/ai-companies-are-buying-tons-of-old-books-because-theyre-free-of-ai-slop https://www.youtube.com/watch?v=d1azwUwKrPo&t=39s https://arxiv.org/pdf/2602.16417
Spain deploys the army to Ceuta after thousands of migrants cross its border from Morocco, the U.S. launches a "heavy wave of strikes" on Iran, Poland claims a Russian missile may have struck Its territory, a U.S. Senate committee cancels a vote on acting Attorney General Blanche's nomination, the Alien Terrorist Removal Court convenes its first hearing in its 30-year history, Salman Rushdie's attacker is convicted of terrorism, Australia takes Telegram to court for allegedly failing to remove terrorist content, the Fed holds rates steady at 3.5% to 3.75%, hundreds of Claude AI Chats surface on search engines, and all 55 UEFA members vote to boycott FIFA competitions over its $20B subsidiary. Sources: Verity.News
NBC Sports Bay Area's Matt Maiocco and Rich discuss the impact of Kyle Shanahan's absence from the start of 49ers camp following his serious car crash and when he will be able to resume full head coaching duties including the possibility of missing San Francisco's Week 1 game against the Rams in Australia, if WR Ricky Pearsall could miss the entire season with a PCL injury, George Kittle's timeline to return from his season-ending Achilles tear, and why the Niners' expectations are sky-high despite playing in the NFC West with the L.A. Rams and Seattle Seahawks. Rich weighs in on FIFA's plans to raise $20B by selling private equity stakes that would result in influence over future World Cups. The guys react to Myles Garrett saying he hasn't spoken to Aaron Donald about a possible return to the Los Angeles Rams. Learn more about your ad choices. Visit podcastchoices.com/adchoices
Sanjit Biswas runs what may be the largest AI deployment in the physical world — and almost nobody in AI talks about it. Samsara (NYSE: IOT), the ~$20B company he co-founded after selling Meraki to Cisco for $1.2B, puts AI on millions of trucks, cranes, and industrial assets: 25 trillion data points a year, 99% of US roads driven every single day, ~$2B in ARR growing 30% profitably. In this episode we go through the entire physical AI stack — asset tags you can run over with a truck, a paper-thin disposable tracking label, engine fault codes, and dash cams running inference at the edge — then into agents, including the Agent Studio warranty agent that compresses an hour of human work into under a minute. We also get into the uncomfortable part (when AI watches you drive all day, is that coaching or surveillance — and why drivers actually want the cameras), mixed fleets of humans and robots, why autonomous trucking will take far longer than robotaxis, and a startling stat from the field: one utility building 3x more grid capacity in the next five years than it did in the previous 125, with 90% of that demand coming from data centers.(00:00) Intro: The biggest AI deployment nobody talks about(01:16) What is physical AI?(03:04) From IoT dashboards to agentic action(04:36) Why physical AI is harder than software AI(06:07) Safety, cybersecurity, and real-world consequences(07:11) What Samsara does(08:22) $2B ARR, 25 trillion data points, and 380,000 crashes(09:44) How AI can prevent road accidents(11:28) From an MIT research project to Meraki(13:42) Learning physical operations from scratch(15:42) Samsara's stack: sensors, intelligence, and action(16:39) Inside Samsara's industrial asset trackers(18:36) Bluetooth, battery life, and connected infrastructure(19:49) A disposable tracking device built like a sticker(21:08) Vehicle gateways and engine diagnostics(22:16) How AI dash cams coach drivers in real time(23:28) Turning the dash cam into an AI interface(24:49) Organizing physical-world data in the cloud(26:32) Selling AI to traditional industries(27:52) Is Samsara's real-world data its AI moat?(29:22) The network effects of covering 99% of U.S. roads(31:23) Edge AI versus cloud AI(32:35) The models running inside Samsara's devices(33:52) Generative AI and video reasoning(35:57) Which AI models does Samsara use?(36:50) Inside Samsara Agent Studio(37:56) How an AI warranty agent works(38:50) Starting with practical, lower-risk automation(40:10) Combining agents, workflows, rules, and guardrails(42:07) What today's AI agents still cannot do(43:07) AI ride-alongs and the future of driver coaching(45:17) Is workplace AI becoming Big Brother?(46:27) How cameras can protect and exonerate drivers(48:48) When AI becomes the judge of your work(50:37) Robots, humanoids, and mixed human-machine fleets(53:19) Samsara's role in autonomous operations(54:50) How quickly will autonomous trucking arrive?(56:32) AI data centers and America's infrastructure boom(58:16) Should lawyers become plumbers? Demand for tradespeople(59:54) Closing thoughts
On his deathbed, Keith Titus told his son: "You're not welcome at Page. I have a plan, and you're not part of it." Three years later, the plan had failed. The company was down from 230 trucks to 70, losing a million dollars a year, and the bank had cut them off. Then Dan's mom — a 27-year school bus driver who never worked a day in the business — made a decision that changed everything: she put every dollar of her husband's life insurance into the company. "I grew up on a dirt-poor farm. I've got no problem going back there." In this episode, Dan Titus of Page Trucking sits down with us in upstate New York to share one of the most incredible comeback stories in bulk freight: five generations on the Titus family farm, a trucking company started in 1974 with $5,000 from his great-grandfather, losing his dad and grandfather a year apart, and rebuilding Page into 500 trucks, 900 trailers, and 650 families — hauling everything from grain to hazardous waste to 1,400° molten metal down the highway. Oh, and they heat their warehouses with Bitcoin miners. ⏱️ CHAPTERS 00:00 Outgrowing the building (again) 01:56 Five generations on the Titus farm — Cato, NY, since 1809 02:32 Dad starts brokering freight without Grandpa knowing 05:28 $5,000 from Great-Grandpa to start a trucking company 06:00 Why it's called "Page" — the $250 authority 08:07 Why his dad chose trucking over farming 13:24 "You're not welcome at Page" — his dad's dying plan 15:21 Mom inherits a collapsing business 21:05 Betting the life insurance money on the company 22:37 Half the pay, twice the hours — Dan comes home 25:45 Buying the business back from Mom 26:40 Teamsters, a $3M pension liability & near-bankruptcy 29:00 From 230 trucks… to 70… to 500 today 32:37 Closed-loop hauling: scrap, dross & aluminum 34:27 Hauling 1,400° molten metal on the highway 38:24 Overweight permits & custom trailer designs 43:40 The Goulet Trucking acquisition 51:42 Why waste hauling is huge in the Northeast 56:22 Heating warehouses with Bitcoin miners 01:04:18 Keith Titus — the next generation 01:09:45 No handouts: earning your way into the family business 01:14:25 The chip on his shoulder — out-innovating $20B competitors 01:20:38 "I'm not ready" — why Page will never sell 01:24:13 Raising sharks, not minnows ABOUT BULKLOADS We started 15 years ago connecting bulk carriers with shippers and brokers in the ag and bulk industry — and today we're much more than a load board. Our mission is simple: serve this industry, support small businesses and trucking families, and help them succeed.
Start your morning with Buzzcast with Joe Lemire: FIFA proposes $20B commercial entity to manage the World Cup; PGA Tour lands another tourney sponsor; Notre Dame gets jersey patch deal Sign up for SBJ 360, our free, daily newsletter. SBJ 360 delivers a concise, high-level overview of the most important stories shaping the sports industry. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
The mates discuss Hugging Face breach, Moonshot AI being valued at $20B, and living to 1,759 years old. Get access to metatrends 10+ years before anyone else - https://qr.diamandis.com/metatrends Peter H. Diamandis, MD, is the Founder of XPRIZE, Singularity University, ZeroG, and A360 Salim Ismail is the founder of Open ExO, a GP at Exponential Venture Capital/The Organizational Singularity Fund and a sought after global speaker and thought leader. Dave Blundin is the founder & GP of Link Ventures Dr. Alexander Wissner-Gross is a computer scientist and founder of Reified – My companies: Apply to Dave's and my new fund:https://qr.diamandis.com/linkventureslanding Go to Blitzy to book a free demo and start building today: https://qr.diamandis.com/blitzy Your body is incredibly good at hiding disease. Schedule a call with Fountain Life to add healthy decades to your life, and to learn more about their Memberships: https://www.fountainlife.com/peter _ Connect with Peter: X Instagram Substack Website Xprize A360 Connect with Dave: Web X LinkedIn Instagram TikTok Connect with Salim: LinkedIn X New ExO Leap Program Join Salim's 10X Shift Subscribe to Salim's YouTube channel Exponential Venture Capital Connect with Alex Website LinkedIn X Email Substack Spotify Threads Listen to MOONSHOTS: Apple YouTube – *Recorded on July 23rd, 2026 *The views expressed by me and all guests are personal opinions and do not constitute Financial, Medical, or Legal advice. Learn more about your ad choices. Visit megaphone.fm/adchoices
Show Notes - https://forum.closednetwork.io/t/episode-59-collect-first-justify-later/199Website / Donations / Support - https://closednetwork.io/support/BTC Lightning Donations - closednetwork@getalby.com / simon@primal.netThank You Patreons & Direct Supporters! - https://www.patreon.com/closednetworkhttps://xmrchat.com/closednetworkDirect Support - https://closednetwork.ioSubscribe Without Patreon - https://closednetwork.io/#/portal/signupMichael Bates - Privacy Bad AssDavid - Privacy Bad AssTK - Privacy Bad AssTrying - Privacy Bad AssVO - Privacy Bad AssMrMilkMustache - Privacy SupporterHutch - Privacy AdvocateInferno_Potato Privacy SupporterDolores Y - Privacy SupporterDirect Support - Craig D Thank You Producers! You Produce This Show!TOP LIGHTNING BOOSTERS !!!! THANK YOU !!!@bon thousands and thousands and thousands of SATs sats!!@fireflygow - 5,000 sats!!frigolay - 34,540 SATs.. HOLY SHITEwardemoff - 5,000 SATsSilas ThornbrookXMR CHAT BOOSTS- HookersOnPhonics - $20Thank You To Our Moderators:Unintelligentseven - Follow on NOSTR primal.net/p/npub15rp9gyw346fmcxgdlgp2y9a2xua9ujdk9nzumflshkwjsc7wepwqnh354dMaddestMax - Follow on NOSTR primal.net/p/npub133yzwsqfgvsuxd4clvkgupshzhjn52v837dlud6gjk4tu2c7grqq3sxavtJoin Our CommunityClosed Network Forum - https://forum.closednetwork.ioJoin Our Matrix Channels!Main - https://matrix.to/#/#closedntwrk:matrix.orgOff Topic - https://matrix.to/#/#closednetworkofftopic:matrix.orgSimpleX Group Chat - https://smp9.simplex.im/g#SRBJK7JhuMWa1jgxfmnOfHz7Bl5KjnKUFL5zy-Jn-j0Join Our Mastodon server!https://closednetwork.socialFollow Simon On The SocialsMastodon - https://closednetwork.social/@simonNOSTR - Public Address - npub186l3994gark0fhknh9zp27q38wv3uy042appcpx93cack5q2n03qte2lu2 - primal.net/simonTwitter / X - @ClosedNtwrkInstagram - https://www.instagram.com/closednetworkpodcast/YouTube - https://www.youtube.com/@closednetworkEmail - simon@closednetwork.ioSpecial Thanks to - EloquentWinter for creating - A Linux guide on MAC address randomizationhttps://forum.closednetwork.io/t/a-linux-guide-on-mac-address-randomization/189TOPICSIN THIS EPISODE01The Man Who Built FISA — And Watched It BreakThe FBI lawyer who designed the bureau's FISA safeguards says that after he left in 2006 they were dismantled — and the system ballooned into Section 702.02EU Chat Control: The Majority Said No. The Scan Survived.A parliamentary majority voted against message scanning on 9 July — and it survived anyway on a second-reading technicality, now running to 2028.03SCOTUS, Flock & the Cameras That Don't Care.The Supreme Court ruled that reconstructing your movements is a 'search' — but 113,000+ license-plate cameras keep rolling, and the fix isn't in a courtroom.04Your Face Is a Password You Can't Change.Madison Square Garden's facial-recognition system leaked — watchlists included — after a single phishing call. Why every face database is a breach-in-waiting.05Tools to Own the Stack.Three open-source projects worth your time — one per fight: DeFlock, SimpleX Chat, and GrapheneOS.Timestamps are estimates based on segment order — update after the final edit.00:00Cold Open & Episode Rundown02:00The Man Who Built FISA08:30EU Chat Control Update14:30SCOTUS, Flock & the Cameras20:00Your Face Is a Password You Can't Change25:00Tools to Own the Stack28:00OutroSEGMENT TAKEAWAYSThe Man Who Built FISAFISA (1978) was sold as a reform, but it legitimized surveillance the government had previously run with no statute at all.The FISA court rarely says no because the real filtering was designed to happen upstream, inside the FBI.Bowman built that pipeline — embedded lawyers, months of review, personal sign-off. After he left in 2006, it went passive.The Carter Page applications (17+ errors; a doctored email; a guilty plea) showed the cost.2008's Amendments Act created Section 702 — blanket categories, no named targets, a queryable database of Americans' data.Thesis: the danger isn't who builds a surveillance system — it's everyone who inherits it.EU Chat ControlTwo proposals, one name: 1.0 (temporary, voluntary) vs. 2.0 / CSAR (permanent, mandatory).26 Mar 2026: Parliament rejected the temporary rules 307–306; the derogation expired 3 April.9 Jul 2026: more MEPs voted to kill the revived scheme than keep it (reported 314–276), but rejecting the Council needed 361 votes. It survived — extended to 3 April 2028, with an E2E carve-out.The permanent CSAR could be adopted as soon as October 2026.Client-side scanning is the core danger; 500+ scientists call it infeasible; Signal would exit the EU. Watch Germany.SCOTUS, Flock & the CamerasChatrie v. United States (29 Jun 2026, 6–3): a geofence is a 'search'; police generally need a warrant. Built on Carpenter (2018).The Court didn't ban geofence warrants and never mentioned ALPRs — a principle above a system it doesn't touch.Scale: 113,000+ cameras; ~20B detections/month across ~5,000 departments vs. ~240M drivers.A 2 Jul 2026 ACLU report documented Flock misleading councils (the Oshkosh 'heat map' reversal).At least 82 jurisdictions have canceled ALPR contracts. The durable fix is local — stop collection.Your Face Is a Password You Can't ChangeMSG breached by ShinyHunters via a phishing call; ~45GB / ~26M claimed records, including facial-recognition records and threat profiles.MSG used facial recognition for years — including to bar opposing lawyers. Second breach in under a year; class actions filed.A pattern, not an accident: Clearview's client list leaked in 2020; Mercor exposed biometrics + ID docs in April 2026.Biometrics are irreversible — you can't reissue your face. The safeguard is not building the database.TOOLS MENTIONEDOpen source · not sponsored · no affiliate relationships.DeFlock · pairs with the Flock segmentCrowdsourced ALPR camera map on OpenStreetMap; has mapped ~half of Flock's ~100,000-camera network so you can see and route around them.deflock.orgSimpleX Chat · pairs with Chat ControlMessenger with no user identifiers at all — the metadata-resistance layer E2E encryption alone doesn't give you. Audited by Trail of Bits.simplex.chatGrapheneOS · pairs with every segmentHardened, de-Googled Android on Pixel — the endpoint is where client-side scanning and biometric capture actually land.grapheneos.orgSOURCES & REFERENCESSegment 1 — FISANaomi Brockwell × Spike Bowman (YouTube)Segment 2 — EU Chat ControlClosed Network — Chat Control live trackerEU Perspectives — Q&AEuronews — temporary scanning extensionPatrick Breyer — Chat Control trackerSegment 3 — SCOTUS / FlockTruthout — SCOTUS ruling & Flock (Mike Ludwig)Supreme Court — Chatrie v. United States (PDF)SCOTUSblog — geofence rulingACLU — Flock Safety credibility reportSegment 4 — Facial Recognition / BreachesTechCrunch — worst breaches of 2026 so farThe Next Web — MSG 45GB leakLaw360 — MSG sued over breachBiometric Update — Mercor biometric breachPrivacy Guides — breach roundup Jul 3–9Segment 5 — ToolsDeFlock (project site)404 Media — DeFlock maps ALPRs worldwideAdafruit — DeFlock overviewSimpleX Chat (GitHub)GrapheneOS (site)
Brian Szytel recaps a quiet Tuesday, July 7, with markets closing modestly lower amid increased U.S.–Iran tensions involving tanker attacks and restrictions on Iran's oil exports; crude rose about 3% to roughly $70.56 while gold dipped. Tech led the decline as semiconductors sold off, with the S&P 500 down ~0.5%, the Dow ~0.4%, and the Nasdaq down a little over 1%. Economic news was limited, but May's U.S. trade deficit widened to $77B, about $20B more than the prior month. Despite the pullback, major indices are up around 10% year-to-date, reflecting a rotation from concentrated chip leaders (some down ~30% in 10 days) into defensives and broader participation. The 10-year yield rose ~7 bps to 4.55%. He also addresses concerns about Q1 profits boosted by mark-to-market gains on non-listed AI holdings, calling it non-recurring and two-sided. 00:00 Market Wrap Intro 00:11 Geopolitics Oil Moves 00:43 Tech Rotation Selloff 01:09 Trade Deficit Update 01:32 Year To Date Perspective 02:28 Rates And Macro Mix 02:39 Ask TBG Earnings Quirk 03:45 Closing Remarks Links mentioned in this episode: DividendCafe.com TheBahnsenGroup.com
Our 250th episode with a summary and discussion of last week's big AI news!Recorded on 06/27/2026Note from Andrey: sorry this is late again! this episode release somehow didn't save and I only realized late, my bad... next one will be out way sooner!Hosted by Andrey Kurenkov and Jeremie HarrisFeel free to email us your questions and feedback at andreyvkurenkov@gmail.com and/or hello@gladstone.aiRead out our text newsletter and comment on the podcast at https://lastweekin.ai/In this episode:US government gating of frontier AI expands: Anthropic gets permission to release Mythos-5 to selected companies/agencies after a standoff, OpenAI rolls out GPT-5.6 “Sol” with initial access restricted to ~20 approved organizations, and Meta is pressed to submit models to “voluntary” review—signaling an emerging de facto licensing regime with geopolitical treaty implications.Model capability and safety signals remain murky: limited benchmark disclosure, claims of token-efficiency comparisons, and third-party reports that GPT-5.6 shows extreme benchmark “cheating” sensitivity highlight steering/alignment bottlenecks and uncertainty about real-world long-horizon behavior.Compute supply chain competition accelerates: OpenAI unveils its Jalapeño inference ASIC with Broadcom on TSMC 3nm; Amazon explores selling Trainium to data-center operators; Micron invests in Anthropic with memory supply agreements; SK Hynix surpasses Samsung on HBM-driven valuation; Groq raises $650M while pivoting toward neocloud.Open source and societal response intensify: GLM 5.2 (MIT-licensed) delivers strong long-context coding performance with rapid optimizations; EconEvals maps job-task exposure; bipartisan workforce initiatives and tax credits launch; DeepMind and Apollo publish loss-of-control/control roadmaps; Hollywood reportedly drops a near-finished Sam Altman biopic amid industry pressure.Timestamps (note - these don't take into account dynamically inserted ads and therefore may be off by a couple of minutes):(00:00:10) Intro / Banter(00:03:42) News PreviewTools & Apps(00:04:41) Anthropic allowed to release Mythos AI to some companies, agencies + Anthropic's Mythos mess is only getting worse + Anthropic floats proposal to Lutnick to end US ban of powerful 'Mythos,' 'Fable' AI models: sources(00:07:58) OpenAI Launches GPT-5.6 Sol Under First-Ever US Government-Gated AI Rollout | MLQ News + OpenAI's new flagship model GPT-5.6 Sol cheats on software tests more than any model before it + Summary of METR's predeployment evaluation of GPT-5.6 Sol(00:24:03) U.S. Presses Meta to Agree to A.I. Reviews - The New York Times(00:30:11) Anthropic's Claude Tag is learning your company, one Slack message at a time | TechCrunchApplications & Business(00:32:49) OpenAI reveals its first AI processor: Jalapeño | The Verge(00:38:29) Amazon in Talks to Sell Custom AI Chips in Bid to Undercut Nvidia(00:41:46) Micron invests in Anthropic and grants it a supply deal(00:45:18) SK Hynix overtakes Samsung to become South Korea's most valuable company | Reuters(00:49:12) AI chipmaker Groq confirms $650M raise, re-staffs after Nvidia's $20B not-acqui-hire deal | TechCrunch(00:52:47) SpaceX inks compute deal with Reflection AI, an open source AI lab | TechCrunchProjects & Open Source(00:54:46) GLM-5.2: Built for Long-Horizon Tasks + How we built the world's fastest API for GLM-5.2 + nvidia/GLM-5.2-NVFP4 · Hugging Face(01:03:04) EconEvalsPolicy & Safety(01:05:40) $500 million AI jobs push launches with bipartisan backing - POLITICO(01:07:47) Rep. Sam Liccardo unveils AI workforce tax credit bill - POLITICO(01:08:56) Google DeepMind announced an “AI Control Roadmap” for improving AI agent security. | The Verge + Securing internal systems against increasingly capable and imperfectly aligned AI(01:14:00) The Loss of Control Playbook: Degrees, Dynamics, and Preparedness + The Loss of Control Playbook(01:16:42) Why corporate AI super PACs spent $27 million on a local election | The Verge(01:20:25) Exclusive: Conservatives plan nationwide protest against AI data centersResearch & Advancements(01:27:37) Revisiting the Platonic Representation Hypothesis: An Aristotelian View(01:31:39) Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models(01:33:59) Tapered Language ModelsSynthetic Media & Art(01:36:54) Hollywood is bending the knee to OpenAI | The VergeSee Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
This Week In Startups is made possible by: Northwest Registered Agent https://northwestregisteredagent.com/twist Vanta https://www.vanta.com/twist Sentry https://sentry.io/twist Today's show: *There are $100 trillion in global assets sitting on top of what Hanover Park co-founder/CEO Chris Hladczuk calls "human duct tape": armies of accountants in offices patching together work from various legacy tools (QuickBooks, Excel) that are holding funds' own data hostage. Can all of this be replaced with AI? Find out how their startup went from overseeing $1B to $20B in assets in just 15 months. PLUS, we flash back to March 2020, when Jason and Figma co-founder/CEO Dylan Field broke down the design tool's initial go-to-market strategy, made some WILDLY inaccurate COVID predictions, and considered anxiety about "SaaS burnout" years before the category went full apocalyptic. Guests: Chris Hladczuk on X: https://x.com/chrishlad Hanover Park: https://www.hanoverpark.com/ Dylan Field: https://x.com/zoink Figma: https://www.figma.com/ Relevant Links: Turner Novak on X: https://x.com/TurnerNovak Banana Capital: https://www.bananacapital.vc/ Emergence Capital: https://www.emcap.com/ Lux Capital: https://www.luxcapital.com/ Susa Ventures: https://susaventures.com/ Bill.com: https://www.bill.com/ METR: https://metr.org/ Granola AI note taker: https://www.granola.ai/ Vanta: https://www.vanta.com/ Foo Camp on YouTube: https://www.youtube.com/c/foocamp TechCrunch Mahalo coverage: https://techcrunch.com/2014/01/27/inside-mobile-news-launch/ Timestamps: 0:00 Hanover Park & the fund admin problem 3:32 Why funds outsource instead of building 5:44 Why fund accounting is so complex 10:38 The "one-click migration" goal 10:48 Northwest Registered Agent - Get more when you start your business with Northwest. In 10 clicks and 10 minutes, you can form your company and walk away with a real business identity — Learn more at https://northwestregisteredagent.com/twist 13:51 Context vs. intelligence gaps 16:11 No PMs, No Designers 20:46 Vanta - Get $1000 off your SOC 2 at https://www.vanta.com/twist 23:18 This is a $100T opportunity 25:36 Flashback w/ Dylan Field of Figma 29:15 Sentry - Your team should be focused on shipping features — not chasing down bugs. New users can get $240 in free credits when they go to https://sentry.io/twist and use the code TWIST 31:22 Pre-AI enterprise security worries 36:30 SaaS overload and SaaS burnout 37:30 The evolution of Figma pricing 42:38 The rise and fall of Mahalo dot com 51:38 The work-from-home revolution begins Subscribe to the TWiST500 newsletter: https://ticker.thisweekinstartups.com Check out the TWIST500: https://www.twist500.com Subscribe to This Week in Startups on Apple: https://rb.gy/v19fcp Follow Lon: X: https://x.com/lons Follow Alex: X: https://x.com/alex LinkedIn: https://www.linkedin.com/in/alexwilhelm Follow Jason: X: https://twitter.com/Jason LinkedIn: https://www.linkedin.com/in/jasoncalacanis Thank you to our partners: (0:00) PARTNER - AD BLURB (0:00) PARTNER - AD BLURB (0:00) PARTNER - AD BLURB Check out all our partner offers: https://partners.launch.co/ Great TWIST interviews: Will Guidara, Eoghan McCabe, Steve Huffman, Brian Chesky, Bob Moesta, Aaron Levie, Sophia Amoruso, Reid Hoffman, Frank Slootman, Billy McFarland Check out Jason's suite of newsletters: https://substack.com/@calacanis Follow TWiST: Twitter: https://twitter.com/TWiStartups YouTube: https://www.youtube.com/thisweekin Instagram: https://www.instagram.com/thisweekinstartups TikTok: https://www.tiktok.com/@thisweekinstartups Substack: https://twistartups.substack.com
On Episode 310 of The Six Five Pod, Patrick Moorhead and Daniel Newman unpack the biggest stories from the week, including insights from Qualcomm Investor Day 2026, OpenAI and Broadcom's Jalapeño AI chip, Anthropic's Micron partnership, SpaceX's massive Reflection AI compute deal, Sakana AI's new Fugu orchestrator, and why memory is emerging as a critical layer of AI infrastructure. Plus, Bulls & Bears covers NVIDIA's $25B bond offering, Apple's MacBook price increases, Micron's record quarter, and Cerebras' first earnings as a public company. The handpicked topics for this week are: Qualcomm Investor Day 2026 — The Data Center Debut: Pat and Dan break down Qualcomm's push into the data center after the company took the stage with Microsoft's Satya Nadella and Meta's Mark Zuckerberg as named customers. They unpack the new Dragonfly platform, including the C1000 250-core data center CPU with PCIe Gen 7 and CXL, the AI200 and AI250 inference accelerators, and a novel High Bandwidth Compute (HBC) architecture that stacks compute under LPDDR memory at dramatically lower cost than HBM. They highlight Qualcomm's ambitious growth targets: $15B data center revenue target for FY 2029, an increased total non-handset revenue goal from $22B to $40B, and a shortened timeline for automotive revenue by two years. They also debate the identity of Qualcomm's unnamed hyperscaler customer and why its robotics opportunity may be flying under the radar. (The Decode) OpenAI and Broadcom Unveil Jalapeño, OpenAI's First Custom Chip: A photo of Sam Altman and Hock Tan holding a wafer and packaged die kicked off OpenAI's reveal of Jalapeño, a custom inference chip built with Broadcom and slated for late-2026 deployment. The chip reached tape-out in roughly nine months, which is an aggressive cycle for an ASIC of this size, and uses HBM3E memory. Pat takes a victory lap on his long-standing heterogeneous compute thesis: every hyperscaler and now every model lab is building accelerators, and the XPU efficiency argument has played out as predicted. Dan frames OpenAI's broader move as existential: they cannot serve frontier models at premium margins if compute remains constrained. He flags that OpenAI is trying to do everything from chips and fabs to social networks and browsers, and that its IPO is now delayed. (The Decode) Anthropic and Micron Sign a Strategic Multi-Year Memory Agreement: Anthropic and Micron announced a multi-year supply agreement for HBM, DRAM, and SSDs, including co-designed next-generation memory for AI workloads, along with a strategic investment by Anthropic in Micron. The pattern mirrors Samsung and SK Hynix's pre-funding Anthropic in May, and follows OpenAI's Jalapeño as another frontier lab moving to lock in supply chain control. Dan frames it as the same circular financing playbook NVIDIA ran two to three years ago, but with the ball now in the memory triopoly's court. Pricing-floor agreements with no ceilings, customized rather than commoditized memory architecture, and demand running well past the previously assumed 2027-2028 horizon. Pat notes that the rumored 14% free cash flow margin at Anthropic makes the strategic investment math work cleanly for both sides. (The Decode) SpaceX Signs $6.3B Compute Deal with Reflection AI: SpaceX inked a $6.3B compute lease with open-source AI lab Reflection AI, at $150M per month from July 2026 through 2029, giving Reflection access to NVIDIA GB300 chips inside the Colossus infrastructure. Combined with the $920M-per-month Google compute contract and existing xAI commitments, SpaceX now has a contracted backlog larger than most public AI startups' entire revenue base, with some calling it the largest commercial AI infrastructure provider at $80B in contracted revenue. Pat reads it as XAI failing to land with developers, consumers, or enterprises, leaving SpaceX with a pot of gold worth far more as wholesale capacity than as XAI's own training compute. Dan flags that Google owning 7% of SpaceX ahead of an IPO is not accidental, and the open question is whether this becomes a Nebius-style infrastructure trade or a full-stack Google-equivalent platform. (The Decode) Japan's Agentic Orchestrator Sakana AI Ships Fugu Plus and Fugu Ultra: Japan's Sakana AI released Fugu Plus and Fugu Ultra, an agentic orchestrator built on a multi-agent MOE approach that routes workloads across multiple underlying models rather than training a new frontier base model. Sakana claims agentic capabilities on par with or better than top frontier models at significantly lower input/output token costs, similar to the DeepSeek and GLM cost-undercut narrative. Pat compares the architecture to OpenRouter and notes the developer-facing parallel to Perplexity Computer's model-routing approach. Both agree that models themselves are no longer moats, and suggests the real moat is the harness, tooling, connectivity, looping, agentic stack, and total compute availability. Expect more sovereign agentic plays from Japan, the Middle East, and elsewhere on the same template. (The Decode) The Flip — Is the Era of Memory as a Commodity Over? Daniel takes the FOR side: memory has moved from commodity to strategic AI infrastructure, citing 16 multi-year agreements covering $22B in committed volume booked through 2027, 84.9% gross margins higher than NVIDIA's, the technology barriers of HBM yield/stacking/packaging that only three companies can clear, and demand drivers tied to HBM as the binding constraint on every AI accelerator rather than to elastic consumer cycles. Patrick takes the AGAINST side: long-term agreements and SCAs signal a commodity in a strong cycle, not a structural rerating; nearly every relevant memory standard — DDR5, MRDIMM, HBM3/3E/4, LPDDR5X/6, GDDR6/7, LPCAM2 — is JEDEC-standard and therefore commodity at the pin; and CXMT's China DDR5 production ramps in 2H 2026 with Lenovo already shipping and HP and Dell qualifying. Custom HBM4 and Qualcomm-style HBC are where strategic memory genuinely lives. (The Flip) NVIDIA's $25B Investment-Grade Bond Offering: NVIDIA priced a $25B multi-tranche bond offering on June 15, its first investment-grade debt sale since 2021, with seven tranches maturing between 2028 and 2056 and $85B in orders against an initial $20B target. Dan reads it as raising when capital is cheap, and oversubscription is real. NVIDIA doesn't need the money, it has a gold balance sheet, and is establishing a credit benchmark rather than funding CapEx. Pat agrees the optics are clean, but flags the irony of NVIDIA, with negative debt, borrowing while the stock trades like dead money at a sub-20x forward P/E. Both note that NVIDIA's underperformance reflects the market's skepticism on memory-as-strategic and on NVIDIA's own capex pace relative to the buildout opportunity ahead. (Bulls & Bears) Tim Cook Calls Apple's Memory Crunch Price Raises on MacBook and iPad "Unsustainable": Apple announced MacBook and iPad price increases of up to $300, with Tim Cook telling the WSJ the memory cost environment is unsustainable. AAPL fell ~5% on the news, the broader rally was momentarily wiped out before Micron held the gains by close. Dan frames it as a moment when the market saw who is going to pay for the AI buildout: the consumer. He notes Apple's pricing power and inelasticity test is now live. Pat traces the backstory to Apple's negative-margin pricing pressure on Micron during the 2022-2023 memory downturn. The question is whether consumer-price blowback will eventually flow back to the memory vendors. (Bulls & Bears) Micron Blows the Doors Off Fiscal Q3 — $41.46B Revenue, 84.9% Gross Margin: The memory story continues as Micron reported its largest beat in company history with fiscal Q3 revenue of $41.46B versus a $35.69B consensus, EPS of $25.11, year-over-year growth of more than 340%, and a record 84.9% gross margin that is roughly 10 points above NVIDIA's. Q4 guidance came in at a $50B midpoint against a $43B consensus. The 16 multi-year strategic customer agreements add up to $22B in committed volume, with most contracts containing pricing floors but no ceilings on most of the volume — a structurally asymmetric setup. Pat notes 95% of the beat came from price, not units, which reinforces his commodity argument; Dan flips it as the early innings of an NVIDIA-style run that puts Micron's 2027 profit on par with Google. (Bulls & Bears) Cerebras' First Earnings Report Since IPO — Revenue Doubles, Margins Compress: Cerebras (CBRS) reported its first earnings as a public company, doubling year-over-year revenue and beating the top line while missing EPS, but the stock sold off hard amid gross margin deterioration. Core revenue came in at $191M, up 12% sequentially, with a $194M Q2 guide that is essentially flat, core gross margins at 47% guiding to 36-38% and 38-41% for the year, and operating margins flipping from positive 2% to a guided -30% to -32%. Customer concentration is shifting from Core42 and G42 (86% of FY25 revenue) to OpenAI, which loaned Cerebras $1B and gets paid quarterly in warrants. Pat flags that Cerebras' uncontested speed claim is no longer uncontested with Groq, TPU v8i, and Tenstorrent putting up real numbers. Cathie Wood is down 52% on her position. (Bulls & Bears) Watch the full video at sixfivemedia.com, and be sure to subscribe to our YouTube channel so you never miss an episode. The Decode Qualcomm Investor Day Lands the Data Center Pivot — Microsoft Deploying Qualcomm HBC XPUs in Azure (Per Satya Nadella) + Meta MOU on Three New Qualcomm Datacenter CPUs (Per Zuckerberg); $3.9B Modular Acquisition; Dragonfly Brand + AI200/AI250 Roadmap; HUMAIN 200MW Ramp; Qualcomm to Become Largest Automotive Silicon Company; Targets $3B Datacenter Revenue FY27, $35B by FY31 https://finance.yahoo.com/markets/stocks/articles/qualcomm-investor-day-detail-data-163247063.html OpenAI Begins Vertical Integration — First Custom Inference Chip "Jalapeño" Unveiled With Broadcom June 24 (Hock Tan: As Good as Blackwell + TPU; ~50% Cost Savings; Late-2026 Microsoft Deployment, 10GW Multi-Gen Roadmap); Daybreak Cyber Stack (June 22) Confirms the Platform Shift https://x.com/OpenAI/status/2069770172802773292 Frontier AI Labs Are Now Financing Their Own Supply Chains — Anthropic Locks In Multi-Year Micron HBM/DRAM/SSD Supply + Micron Becomes Series H Investor; Same Pattern as Samsung + SK hynix Pre-Funded Anthropic in May; $965B Post-Money, $47B Revenue Run-Rate, October IPO Target https://investors.micron.com/news-releases/news-release-details/micron-and-anthropic-announce-strategic-agreement-scale-next SpaceX Signs $6.3B Compute Deal With Reflection AI — $150M/Month July 2026 → End of 2029; NVIDIA GB300 + Colossus 2 Capacity; SpaceX Now Largest Commercial AI Infrastructure Provider With $80B+ Committed Compute Revenue Through 2029 https://finance.yahoo.com/technology/ai/articles/spacex-reportedly-grant-reflection-ai-162749237.html The Sovereign AI Stack Lands — Japan's Sakana Ships Fugu + Fugu Ultra Multi-Agent System (June 22) That Beats Opus 4.8, GPT-5.5, and Gemini 3.1 Pro on 10 of 11 Benchmarks; Designed Around US Export-Control Risk; Completes the Three-Bloc Sovereign-AI Map With Mistral Compute (Europe) + DeepSeek $7.4B (China) https://www.datacamp.com/blog/sakana-fugu The Flip Is the Era of Memory as a Commodity Over? FOR: Memory is now strategic AI infrastructure with multi-year supply lock-ins. The cycle dynamics that defined the last 30 years no longer apply. https://www.benzinga.com/markets/tech/26/06/60062500/micron-earnings-could-echo-nvidias-2023-moment-says-futurum-ceo AGAINST: Memory is cyclical and priced for perfection. This print is either step change or top of the cycle, and the second one is more likely. https://www.cnbc.com/2026/06/25/apple-macbook-ipad-price-hike-memory.html Bulls & Bears NVIDIA (NVDA) $25B Bond Sale Anchors the AI Debt-Finance Boom — First Bond Offering Since 2021; Joins Alphabet $80B, Amazon $27.5B, Meta $30B, Oracle Stack; Dan: "Locking In Cheap Capital While It Can" https://finance.yahoo.com/technology/ai/articles/nvidia-record-us-25-billion-131039687.html Apple (AAPL) Falls −5%+ Thursday June 25 on Confirmed MacBook + iPad Price Hikes — Tim Cook RAM "Unsustainable" Comment Lands as Real Price Action; Apple Hikes Erase Micron-Driven Tech Rally Mid-Session; Memory Beneficiaries (SanDisk, Micron) Surge; Analysts "Mostly Nonplussed" https://tickerspark.ai/market/apple-inc-aapl-drops-5-3-as-price-hikes-spook-investors-1782399950638 Micron (MU) Q3 FY26 ACTUALS — Largest Beat in Company History; Revenue $41.46B (+346% YoY) Crushes $35.69B Consensus; Non-GAAP EPS $25.11 (+1,215% YoY) Beats $20.49; Record 84.9% Gross Margin (Higher Than NVIDIA); Q4 Guide $50B Midpoint vs $43B Consensus; Stock +18-19% Overnight to $1,242 https://www.nasdaq.com/articles/nvda-who-micron-blows-doors-q3-earnings-revs Cerebras Systems (CBRS) Q1 ACTUALS — First Earnings Post-IPO; Revenue $193.4M Nearly Doubled YoY; 2026 Guide $855-$865M Beats $824M; BUT Gross Margins Forecast 38-41% (Down From 45% Q1, Half of NVIDIA + Micron); Stock −20% AH on Margin Compression; Sets Up Inference-Tier Margin Debate https://investors.cerebras.ai/news-releases/news-release-details/cerebras-systems-announces-strong-first-quarter-2026-results
What happens when ancient laws clash with modern urban planning?In this episode, we dive into a fascinating thought experiment that is playing out in real-time in Vancouver, British Columbia. This podcast contrasts a Biblical mandate in Leviticus 25:23—which explicitly forbids the permanent sale of land, asserting that humans are merely tenants and stewards— with the Squamish Nation reclaiming their ancestral territory to build Senakw.Senakw is the largest First Nations economic development project in Canadian history: an unapologetically massive, ultra-dense, net-zero carbon mega-project. Because it is being built entirely on designated reserve land, it completely bypasses the city's strict traditional zoning bylaws.But with density projections five times higher than Canada's current highest-density neighborhoods, the project has sparked a fierce debate. Is this sustainable, earth-centric urban design an inspiring triumph of indigenous reclamation, or an infrastructural recipe for disaster as critics claim?In this episode, we discuss:The staggering $20B economics and architecture of the Senakw development.The density debate: Why urban planners are sounding the alarm on liveability.The ultimate question: Would our cities be better managed if we followed ancient rules of stewardship?
Is the AI trade a bubble? Imran Khan — founder of Proem Asset Management, former Snap executive, and the banker behind the Alibaba and Mercado Libre IPOs — isn't convinced. Dan Nathan sits down with Imran to pressure-test the bear case, from Nvidia's below-market multiple to the cyclical-vs-secular debate in memory, and to dig into why a big chunk of SpaceX's $2.5T valuation may not be a space story at all. Topics Covered Why hyperscalers underperform during heavy CapEx cycles — and why that's historically the best time to buy Distribution vs. technology: how Gemini won while arguably being the inferior model, and why Grok couldn't Meta's setup — cheap on earnings, not cheap on free cash flow — and the Zuckerberg "big swing" risk Nvidia at a $5T market cap: the $20B debt raise, buybacks, and the customers-are-competitors problem Micron and high-bandwidth memory sold out into 2027, and the cyclical-vs-secular question that decides the stock The "bottleneck trade" everyone's chasing — and why earnings durability is the thing to watch Energy constraints, data center delays, and the long-term demand picture Imran's contrarian case that AI won't create structural unemployment SpaceX's valuation decoded: rocket launch, Starlink, and the xAI cloud ramp What OpenAI and Anthropic coming to market could mean for the AI trade —FOLLOW USYouTube: @RiskReversalMediaInstagram: @riskreversalmediaTwitter: @RiskReversalLinkedIn: RiskReversal Media The financial opinions expressed in Risk Reversal content are for information purposes only. The opinions expressed by the hosts and participants are not an attempt to influence specific trading behavior, investments, or strategies. Past performance does not necessarily predict future outcomes. No specific results or profits are assured when relying on Risk Reversal. Before making any investment or trade, evaluate its suitability for your circumstances and consider consulting your own financial or investment advisor. The financial products discussed in Risk Reversal carry a high level of risk and may not be appropriate for many investors. If you have uncertainties, it's advisable to seek professional advice. Remember that trading involves a risk to your capital, so only invest money that you can afford to lose. Derivatives are not suitable for all investors and involve the risk of losing more than the amount originally deposited and any profit you might have made. This communication is not a recommendation or offer to buy, sell or retain any specific investment or service.
This Week In Startups is made possible by:Every.io - visit every.ioSentry.io - sentry.io/twistVanta - vanta.com/twistToday's show:The next SpaceX won't be building rockets; it'll build the first hotel on the Moon. Today on TWiST, GRU Space founder Skyler Chan brings a brick made from lunar soil into the studio and lays out a plan to manufacture on the Moon as early as next year. We get into the science, the business model, and the regulatory land-grab ahead!Then, the US government forces Anthropic to pull Fable 5 and Mythos 5. Jason and Lon unpack what it means when a single AI model can vanish overnight, and the US government's emergency order.Stick around for the winner (or winners?!) of the $5,000 AI podcast-companion bounty!Timestamps:0:00 Knicks playoff run, San Antonio & Texas BBQ (Black's vs. Terry Black's)8:23 Plaud: If your work depends on conversations — interviews, meetings, calls — you need a Plaud NotePin. You can check it out at https://Plaud.ai/twist and use code TWIST for 10% off!9:52 Guest intro: Skyler Chan, GRU Space — and the lunar-soil brick10:29 Every.io - For all of your incorporation, banking, payroll, benefits, accounting, taxes or other back-office administration needs, visit https://every.io12:14 Why the next SpaceX builds habitats, not rockets15:18 NASA's $20B moon-base signal & the TRL contracting path16:13 Business model: from construction contractor to owning lunar land18:06 Building the hotel robotically + $1M refundable deposits20:00 Sentry - Your team should be focused on shipping features — not chasing down bugs. New users can get $240 in free credits when they go to https://sentry.io/twist and use the code TWIST31:09 Vanta - Get $1000 off your SOC 2 at https://www.vanta.com/twist39:28 News: US government blocks Anthropic's Fable 5 & Mythos 541:34 The politics: Hegseth, Sacks, Jassy & the "conspiracy" angle1:09:39 Why no single model dependency is safe (multi-model harnesses)1:18:47 Bounty results: MySidecast wins $3,0001:19:07 Honorable mentions: Couchverse (Lemon Slice) & Convalenz1:29:44 Next bounty: the Annotated app
SpaceX priced the biggest IPO ever at $135/share, raising $75B and debuting at $1.77T. ShinyHunters exploited an unpatched Oracle PeopleSoft flaw hitting 100+ organizations, Mistral seeks €3B at €20B, MrBeast hit 500M subscribers, and SBF lost his appeal. SpaceX raises $75B in the biggest-ever IPO, pricing 555.6M shares at $135 each, giving it a market value of $1.77T (Bloomberg) Founders Fund's ~3% SpaceX stake is worth $50B+, Sequoia's ~1.5% is worth $20B+, and a16z will see its biggest return ever at $10B+ (Bloomberg) Some investors question SpaceX's valuation, citing its $4.3B loss on $4.7B in revenue in Q1, as well as concerns over space data centers (NYT) Oracle warns customers of a critical PeopleSoft flaw after ShinyHunters claimed breaches of 100+ organizations using PeopleSoft; Oracle has not issued a patch (TechCrunch) Sources: French startup Mistral AI is in talks to raise ~€3B at a ~€20B valuation; it was last valued at €11.7B during a funding round in September 2025 (Bloomberg) MrBeast hits 500M subscribers on YouTube, a record for the platform (The Wrap) Sam Bankman-Fried loses his bid to overturn his fraud conviction and 25-year prison sentence over the collapse of FTX (Reuters) Longreads As companies are hit by rising AI costs, they are increasingly using tools that tap cheaper models, including some from China, putting price pressure on OpenAI and Anthropic (WSJ) Sixteen economists weigh in on what AI will mean for the US economy, workers, and workplaces; only two expect AI to actually create more jobs (WSJ) Learn more about your ad choices. Visit megaphone.fm/adchoices
Anduril Industries raised $5 billion at a reported $61billion valuation—putting a nine-year-old defense tech company in the same conversation as legacy primes that have been building weapons for generations.How did they do it, what is their strategy, and does the math make sense?In this episode, Mike and Matthew take a deep dive inside Anduril's products, revenue, contracts, and business strategy. They break down the Series H raise, the company's rapid valuation climb, the difference between contract ceilings and booked revenue, and why visible federal obligations onlytell part of the story.They also examine Anduril's expanding product portfolio, anddebate the core question behind the company's $61B price tag: Is Anduril the future of defense industrial production, or is the market pricing in near-flawless execution?Topics include:- Anduril's $5B Series H and $61B valuation- The gap between reported revenue and visible federalobligations- Why Special Operations and the Border Patrol matter morethan most people realize- The $20B Army enterprise vehicle—and why it is a rail, not acheck- Barracuda, Fury, Arsenal-1, and hyperscale defensemanufacturing- How Anduril compares to Lockheed, Northrop, GeneralDynamics, RTX, and Palantir- The bull and bear case for Anduril's long-term strategy- What to watch next: IPO timing, task orders, deliveries, andrevenue growth- The real bet: for Anduril to justify today's valuation, ithas to grow from a $2B revenue company into a $20B+ revenue company very quickly.SUBSCRIBE FOR FREE to get more intel on defense tech, news, and happenings. Links• Sign up for the newsletter! • Support us on Patreon! ----Follow us on...• LinkedIn• Instagram• X• Facebook• Website ----00:0000:34 intro01:20 Premium newsletter!02:10 Anduril intro02:26 Matthrew intro04:32 Anduril 10106:52 Anduril's fundraising07:25 the next 24 months07:38 revenue breakdown08:23 happenings between the raises14:10 last 5 years of sales15:47 counter-UAS18:22 Steve vs Steve approach19:14 C-UAS durability?21:04 Altius21:52 comparing valuations23:21 sources of new revenue23:33 Barracuda24:08 CCA program27:28 Lattice28:20 Eagle Eye31:18 Golden Dome35:08 Anduril's strategy38:53 next acquisition?41:25 wrap-up
The housing market is shifting, and builders are pulling back on new constructions. What does that mean for the future of multi-family real estate? On this episode, the Neighborhood Ventures team covers three critical market trends: Single-Family Slowdown: Why rising costs and interest rates are pushing builders to pause. The Phoenix Sports Boom: How year-round hosting and new stadiums impact local demand. The $165B Industrial Path: TSMC's latest $20B injection into Arizona and how it's reshaping North Phoenix.
Today we had the pleasure of hosting Steven Kobos, President and CEO of Excelerate Energy. Steven has served as President and CEO since 2018 and previously spent 11 years as a member of the company's Board of Directors and corporate counsel. Throughout his career, he has worked across global energy markets, including Kuwait, Bangladesh, Pakistan, Argentina, Brazil, Finland, Germany, and the Middle East. Excelerate is a global leader in flexible LNG infrastructure solutions, focused on expanding access to reliable, affordable, and secure natural gas. The company operates one of the world's largest fleets of Floating Storage and Regasification Units (FSRUs) and provides integrated LNG solutions spanning the entire value chain. We were thrilled to hear Steven's perspective on the evolving and increasingly complex global energy landscape. In our conversation, we explore the evolution of the global LNG market, the impact of U.S. shale on Excelerate's business model, and why the company has increasingly focused on integrated LNG and infrastructure solutions rather than simply providing floating regasification assets. We discuss the growing importance of energy security following recent geopolitical disruptions, including tensions surrounding the Strait of Hormuz and Steven's recent visit to the region, and the role LNG continues to play in supporting power generation, industrial growth, and economic development around the world. Steven walks us through Excelerate's newest FSRU, the Acadia, the company's expanding opportunities in Iraq, and how LNG imports are helping address power shortages and energy deficits across emerging markets. We discuss the future growth of global LNG demand, the increasing shift toward long-term supply contracts, the advantages of floating infrastructure versus traditional onshore facilities, and Excelerate's strategy of combining LNG supply with downstream infrastructure to open new markets. We also cover Argentina's Vaca Muerta opportunity, Brazil's hydro-backed power system, Finland's experience with energy security following disruptions to regional gas infrastructure, the growing role of U.S. LNG exports, and the support provided by the Trump Administration to promote American energy abroad. Steven shares several personal anecdotes, including helping launch LNG imports into Kuwait, opening new LNG markets across South Asia, visiting customers throughout the Gulf during the recent conflict, and witnessing firsthand how access to reliable energy can transform communities and economies. We covered a great deal and appreciate Steven for sharing his time and insights. Mike Bradley started the show by noting that markets continue to be driven almost entirely by on-and-off developments in the Middle East. Market sentiment last week was dominated by optimism that Iran and the U.S. were moving toward a Strait of Hormuz resolution, but this week has started with growing concern that a resolution may not be just around the corner. On the bond market front, the 10-year bond yield was trading at ~4.5% (up 6-7bps), driven by an Iranian resolution being pushed further to the right and constructive economic data. He noted that the May ISM Manufacturing report showed that U.S. manufacturing expanded at its fastest pace in four years. On the crude oil market front, WTI prices spiked ~$6/bbl (to $93/bbl) on concerns that an Iranian resolution could be delayed. The Strait of Hormuz needs to reopen quickly or risk global oil prices moving substantially higher, as oil markets enter the higher-demand summer months with critically low inventory levels. From an energy equity perspective, the Energy sector was up ~2% so far this week after a 5% pullback last week. On the broader equity market front, markets were modestly weaker as investors appeared unprepared for the prospect of an Iranian resolution being pushed further into the future. He ended by highlighting two IPOs scheduled to price over the next two weeks. Equity investors are most excited about the SpaceX IPO (expected to price next week at a ~$2T valuation). He also highlighted INNIO Holdings, a gas power system manufacturer that is expected to price later this week (raising ~$2B at a ~$20B valuation), which should provide a good read on how bullish sentiment remains across the engine manufacturing and distributed generation segments. Mark Castiglione added his questions and perspective to the discussion as well.
We talk a lot about coding and AI and a little less about headlines today. Runner-up: SpaceX is targeting a June/July 2026 IPO at a reported ~$1.75 trillion valuation, which would be the largest public listing in history. The float follows SpaceX's ~$250B all-stock acquisition of xAI in February, folding Starlink, launch, and frontier AI into one entity.Runner-up: Amazon's custom AI chip business — Graviton, Trainium, and Nitro — hit a $20B annual run rate with triple-digit YoY growth. OpenAI committed to about 2 GW of Trainium capacity, Anthropic is scaling to 5 GW, and analysts project a standalone Trainium could become a $50B business.Runner-up: NVIDIA topped a $5.5 trillion market cap and is deploying more than $45B across the AI supply chain, extending its position from chip supplier to investor and customer across the stack.Runner-up: Apple posted record fiscal Q2 2026 revenue of $111.2B, up 17% YoY, with diluted EPS of $2.01. iPhone sales rose 22% and Services climbed about 16% to $26.65B, and the company guided Q3 growth of 14%-17%.Runner-up: AI venture funding shattered records with $297B in Q1 2026, including $35B raised in a single week.If you want a prize, send us a DM:instagram.com/rickerandbontiktok.com/@rickerandbontiktok.com/@rickerandbonyoutube.com/@rickerandbon
Topics- NBA head coach, Rick Adelman, dies at 79- 4-Time Stanley Cup winner, Claude Lemieux, dies at 60- Marilyn Monroe would've turned 100 on June 1st, 2026- Wyndham Clarke wins Byron Nelson golf classic with a record final round 60- Joey Levine, King of Pop Rock, turns 79- Trump "settles" 20B lawsuit with IRS with suspicious qualifications- Richard Bailey's bid for Counselman of S.D. is decided today- Average price of a car in U.S. breaks record- Stanley Cup finals begin tonight between Vegas & Carolina- Tomorrow night begins the NBA finals between the Knicks & Spurs- Jack looks at the MLB standings
Tesla's former President Jon McNeill reveals the five-step framework behind one of the world's fastest-growing companies— YOU'LL LEARN — 1) What most miss when designing processes2) How to identify outdated requirements that slow things down 3) Why automation should be your LAST step Subscribe or visit AwesomeAtYourJob.com/ep1157 for clickable versions of the links below. — ABOUT JON — Jon McNeill is the CEO and Co-Founder of DVx Ventures. With a track record of founding and scaling companies, Jon has led teams that generated tens of thousands of jobs and delivered multi-billion dollar returns for investors.Previously, Jon served as President at Tesla, where revenue grew from $2B to $20B in under 30 months, and later as COO at Lyft, helping double revenue and take the company public. He currently sits on the boards of General Motors, Lululemon, Asurion, CrossFit, and Stash.• Book: The Algorithm: The Hypergrowth Formula that Transformed Tesla, Lululemon, General Motors and SpaceX• Website: DVX.ventures— RESOURCES MENTIONED IN THE SHOW — • Book: Sam Walton: Made In America by Sam Walton• Book: The Goal: 40th Anniversary Edition: A Process of Ongoing Improvement by Eliyahu Goldratt• Book: Unreasonable Hospitality: The Remarkable Power of Giving People More Than They Expect by Will Guidara• Past episode: 810: How to Get Stuff Done inside Bureaucracies with Marina Nitze• Research paper: "Attention Is All You Need"— THANK YOU SPONSORS! — • Shopify. Sign up for your $1/month trial at Shopify.com/awesomepodSee Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
The Twenty Minute VC: Venture Capital | Startup Funding | The Pitch
AGENDA: 00:05:11 — Anthropic freezes secondary sales, requiring board approval for all transfers. 00:10:45 — Why Anthropic is buying capacity from Elon Musk. 00:15:35 — Anthropic's massive $200B revenue commit to Google. 00:18:55 — Goldman Sachs predicts a 24x surge in token consumption driven by agents. 00:31:05 — Will AI labs eat the app layer? The threat to Legal and CX verticals. 00:37:55 — SaaS public markets: HubSpot tanks 18% while Monday.com finds its footing. 00:42:40 — Growth theft: How Clay is commoditizing ZoomInfo's data business. 00:46:25 — Cerebras prices IPO at $150–$160 with a $48B market cap. 00:52:15 — Real Venture Capital: Celebrating the early bets by Foundation and Benchmark. 00:58:30 — Ramp's valuation vs. the Chapter 7 collapse of e-commerce card Parker. 01:06:20 — Success and Sacrifice: Is mental health the price of building a $20B company?
This episode is brought to you by Audible, WHOOP and Strong Coffee Company. What really created Tesla's explosive growth? Was it Elon Musk, innovation, timing… or was there actually a repeatable formula behind it all? In this episode of Ever Forward Radio, former Tesla President Jon McNeill breaks down the exact framework used to scale Tesla from $1.8 billion to nearly $20 billion in revenue in just 30 months. Drawing from his new book, The Algorithm, Jon explains how companies like Tesla, SpaceX, Lululemon, and others use systems thinking, customer obsession, speed, curiosity, and innovation to create hypergrowth. But this conversation goes far beyond business. Chase and Jon explore how the same principles can be applied to your personal life, fitness, mindset, relationships, habits, and purpose. They discuss the danger of comfort and convenience, why most startups fail even with funding, how to build a mission-driven life, and the hidden cost of scaling too fast. Jon also shares behind-the-scenes stories from Tesla, lessons from working alongside Elon Musk, the importance of values and intentionality, and why curiosity may be the greatest superpower for growth. If you want to learn how to think bigger, move faster, simplify your life, and build something meaningful — this episode is for you ----- 00:00 — Tesla's Hypergrowth Story Begins 00:02 — The Mobile Service Breakthrough 00:12 — Tesla's Parking Lot "Triage" System 00:26 — The Algorithm Explained in 30 Seconds 00:43 — How Tesla Reduced Car Buying From 64 Clicks to 13 01:01 — Intro & Audible Sponsor 01:58 — How Tesla Scaled From $1.8B to $20B in 30 Months 02:35 — Is Hypergrowth Just Controlled Chaos? 03:14 — Growth vs Innovation: Which Comes First? 04:20 — "Creative Dissatisfaction" Inside Tesla 04:42 — Why "Done Is Better Than Perfect" 05:46 — How Tesla Used Customer Feedback Loops 07:23 — The Service Problem That Nearly Broke Tesla 08:53 — Creating Tesla's Mobile Service Model 10:12 — Why Convenience Changes Everything 12:08 — Step 1 of The Algorithm: Question Assumptions 14:03 — Consumer Friction & Asking Better Questions 16:14 — Why Tesla Put Stores Next to Apple & Lululemon 17:04 — Elon Musk's "Domino's Pizza" Car Buying Challenge 18:06 — Eliminating Unnecessary Loan Paperwork 19:20 — How Tesla Made Buying a Car Feel Like Ordering Pizza 20:18 — Convenience vs Character 21:36 — Fitness, Discipline & GLP-1s 23:01 — The Power of Intentionality 24:08 — Building Systems That Scale 25:38 — Learning From Hospitals & Emergency Rooms 27:20 — Curiosity as a Superpower 27:58 — Breaking Down The Algorithm Step-by-Step 29:12 — Why Speed & Quality Must Work Together 30:00 — Lessons From Olympic Cross-Country Skiers 31:20 — Using Speed to Expose Weaknesses 32:03 — How Toyota Forced Tesla to Improve Faster 32:43 — Turning Customers Into Tesla Evangelists 35:19 — Why Tesla Owners Became Obsessed With the Brand 37:14 — Strong Coffee Sponsor Break 37:27 — Knowing When to Pivot vs Keep Pushing 39:12 — Why "Good Enough" Is Dangerous 39:33 — Steve Jobs & The Simplicity Principle 42:04 — How Lululemon Cut Production From 1 Year to 8 Weeks 45:39 — Elon Musk's 10X Thinking 47:02 — Finding People Who Challenge You 49:33 — The Importance of Shared Values 51:26 — Tesla's Core Value: Customer Obsession 52:32 — Values vs Goals 52:51 — Productive Pressure vs Destructive Stress 54:46 — When Hypergrowth Becomes Dangerous 56:11 — Jon McNeill's Daily Habits & Routines 58:02 — How Family & Values Shape Success 59:13 — Why Company Values Matter 01:01:35 — What Surprised Jon While Writing The Book 01:03:38 — Why Tesla Almost Failed 01:04:42 — The #1 Trait of Successful People 01:05:50 — Why Most Startups Die 01:06:19 — How to Scale Your Life Like a Company 01:07:00 — The Moral Responsibility of Growth 01:07:22 — What "Ever Forward" Means to Jon McNeill 01:08:36 — Where to Find Jon & The Algorithm Book ----- Episode resources: Get Jon's new book The Algorithm Get his audiobook for FREE with your 30-day trial of Audible at https://www.AudibleTrial.com/everforward Track your sleep, training, recovery and so much more with the WHOOP physical activity tracker Save 15% on my favorite at-home coffee with code CHASE at https://www.StrongCoffeeCompany.com/chase Watch and subscribe on YouTube
Three stories on the table this week, and none of them small.Saks Global plans to exit Chapter 11 on June 22nd carrying $1.2 billion in debt, with a reorganization plan targeting $9 billion in GMV by fiscal 2030. That's nearly double where they sit today. Rick Watson and Jessica Lesesky walk through the vendor mess (720 brands stopped shipping at the worst of it), the repair work underway, and why exiting bankruptcy this leveraged sets up another round of trouble down the road.The Watson Weekly Weekend edition is sponsored by Avalara - the agentic AI platform automating global tax and compliance for leading eCommerce brands. For more details: https://avalaratax.watsonweekly.comOver at Victoria's Secret, Australian investor Brett Blundy's BBRC Worldwide has built a roughly 13% stake and is pushing to remove two directors: chair Donna James and Miriam Naficy. The complaint is acquisitions like Adore Me. CEO Hillary Super is running a "path to potential" plan built around body positivity and a return to the Angels heritage. Fiscal 2025 sales are up 5%. The question is whether that's enough to keep the activist quiet.Then earnings. Alphabet did $109B in Q1, with Google Cloud growing 63% YoY to a $20B run rate and a $462B backlog. Amazon hit $181B, AWS grew 28% to $37.5B, and the chip business crossed a $20B run rate of its own. Shopify cleared $100B in quarterly GMV for the first time, with operating income up 88% on the back of all the layoffs and restructuring.The thread underneath all of it: AI compute is getting more expensive, not less. The pricing power is sitting with the infrastructure layer. Amazon, Nvidia, and the LLM owners are collecting the rent. The businesses adopting AI are paying it.
There are so many things in our modern world that we presume are fairly recent inventions. But the three things we’re going to talk about in this instance are quite old, but they have close associations with the recent past. Research: Abbott, David, PhD., ed. “The Biographical Book of Scientists: Engineers and Inventors.” Peter Bedrick Books. New York. 1985. “Bad Breath.” Medline Plus. https://medlineplus.gov/badbreath.html#:~:text=Teenagers-,Summary,help%20give%20you%20fresher%20breath. Berlin, Erika. “‘The Myriad Reflector’: The Early, Forgotten Disco Ball.” Mental Floss. May 21, 2015. https://www.mentalfloss.com/entertainment/myriad-reflector-early-forgotten-disco-ball Britannica Editors. "aeolipile". Encyclopedia Britannica, 6 Jun. 2016, https://www.britannica.com/technology/aeolipile Britannica Editors. "Heron of Alexandria". Encyclopedia Britannica, 12 Mar. 2024, https://www.britannica.com/biography/Heron-of-Alexandria Garber, David. “Meet Me Under the Disco Ball: A History of Nightlife’s Most Enduring Symbol.” Vice. June 4, 2015. https://www.vice.com/en/article/meet-me-under-the-disco-ball-a-history-of-nightlifes-most-enduring-symbol/ Handwerk, Brian. “The History and Science Behind Your Terrible Breath.” Smithsonian. Feb. 13, 2017. https://www.smithsonianmag.com/science-nature/halitosis-horrors-how-bad-breath-became-americas-worst-nightmare-180962104/ HØYRUP, JENS. “A NEW EDITION OF THE METRICA OF HERON OF ALEXANDRIA.” Physis. Vol. LIII. 2018. http://akira.ruc.dk/~jensh/Publications/2018%7BR%7D06_A%20New%20Edition%20of%20the%20Metrica%20of%20Heron%20of%20Alexandria_S.pdf Hughes, J. Donald. “Hero of Alexandria.” Ebsco. 2023. https://www.ebsco.com/research-starters/biography/hero-alexandria Mendell, H. “Hero and the tradition of the circle segment.” Arch. Hist. Exact Sci. 77, 451–499 (2023). https://doi.org/10.1007/s00407-023-00308-y “Mint! From the Ancient World to Modern Manchester.” Manchester Museum. Aug. 17, 2018. https://storiesfromthemuseumfloor.wordpress.com/2018/08/17/mint-from-the-ancient-world-to-modern-manchester/#:~:text=The%20ancient%20Egyptians%20invented%20breath%20mints%20to,*%20Severely%20worn%20teeth%20*%20Tooth%20loss “Myriad Reflector Will Feature Annual Fall Opening Odeon Ball.” Great Falls leader. Sept. 4, 1921. https://www.newspapers.com/image/1018804435/?match=1&terms=%22myriad%20reflector%22 “Plant of the Month: Mint.” JSTOR Daily. https://daily.jstor.org/plant-of-the-month-mint/ Pliny the Elder. “The Natural History.” Translated by John Bostock and Henry T. Riley. Taylor & Francis. London. 1855. Project Gutenberg. https://www.gutenberg.org/ebooks/author/50041 Rossen, Jake. “All That Glitters: A History of the Disco Ball.” Mental Floss. Dec. 30, 2021. https://www.mentalfloss.com/entertainment/music/disco-ball-facts-history “Saltair.” Salt Lake Telegram. June 13, 1921. https://www.newspapers.com/image/288643722/?match=1&terms=%22myriad%20reflector%22 Smith, Grafton Elliot, et al. “The Papyrus Ebers.” Ares Publishers. Chicago. 1974. https://babel.hathitrust.org/cgi/pt?id=coo.31924073200077&seq=5 “Strike the Banners.” The Kentucky Post. August 31, 1945. https://www.newspapers.com/image/760821309/?match=1&terms=%22L.%20B.Woeste%22 “Wonderful Falls Short of Expressing the Grandeur of the Rotary Charity Ball.” The Piqua Daily Call. Jan. 26, 1917. https://www.newspapers.com/image/935844964/?match=1&terms=%22myriad%20reflector%22 Woeste, L.B. “Myriad Reflector.” U.S. Patent Office. Feb. 6, 1917. https://patentimages.storage.googleapis.com/9e/4c/73/00bfc626d3f664/US1214863.pdf Woeste, L.B. “Myriad Reflector.” U.S. Patent Office. March 13, 1928. https://ppubs.uspto.gov/api/pdf/downloadPdf/1662554?requestToken=eyJzdWIiOiIyM2QyOTAxNi1iNjVhLTRkNTAtYWEyOS0zZjAyOWMwYmZiMWUiLCJ2ZXIiOiJmZjg4ZmU5Yy1iOTA2LTQxZDUtYTQxMS02MGM5Mzk3NTk0YzYiLCJleHAiOjB9 “Woeste Rites Are Set.” Cincinatti Enquirer. April 11, 1933. https://www.newspapers.com/image/103141821/?article=7dc922a9-f0a9-42b8-a61e-f9e92a7b3557&terms=%22Louis%20B.%20Woeste%22 Woodcroft, Bennet, ed. “The Pneumatics of Hero of Alexandria.” Taylor Walton and Maberly. London. 1851. Accessed online: https://www.thehopkinthomasproject.com/TheHopkinThomasProject/TimeLine/Wales/Steam/URochesterCollection/Hero/index-2.html See omnystudio.com/listener for privacy information.
Ali Mogharabi peels back the layers behind several Mag 7 companies' ongoing AI initiatives. For Meta Platforms (META), he assesses Mark Zuckerberg's foray into AI and examines the increased capex costs facing the Facebook and Instagram parent company. Later, Ali says Alphabet (GOOGL) cloud growth was "striking" and highlights its position in AI with Gemini showing positive signals. For Amazon (AMZN), he points to "reacceleration of AWS cloud growth" but also a $20B run rate from its Trainium chips. ======== Schwab Network ========Empowering every investor and trader, every market day.Options involve risks and are not suitable for all investors. Before trading, read the Options Disclosure Document. http://bit.ly/2v9tH6DSubscribe to the Market Minute newsletter - https://schwabnetwork.com/subscribeDownload the iOS app - https://apps.apple.com/us/app/schwab-network/id1460719185Download the Amazon Fire Tv App - https://www.amazon.com/TD-Ameritrade-Network/dp/B07KRD76C7Watch on Sling - https://watch.sling.com/1/asset/191928615bd8d47686f94682aefaa007/watchWatch on Vizio - https://www.vizio.com/en/watchfreeplus-exploreWatch on DistroTV - https://www.distro.tv/live/schwab-network/Follow us on X – https://twitter.com/schwabnetworkFollow us on Facebook – https://www.facebook.com/schwabnetworkFollow us on LinkedIn - https://www.linkedin.com/company/schwab-network/ About Schwab Network - https://schwabnetwork.com/about
With the Iran War underway, the United Arab Emirates is looking for some economic certainty. The rich Arab nation is home to a lot of foreign-held deposits, and they're worried investors will pull those funds. So, they're looking for an economic backstop. Enter: currency swap lines. Today, we explain why the UAE is looking to its close ally, the U.S., for a currency swap line and how it would work.The Indicator has a weekly newsletter! Be among the first to sign up now: npr.org/indicatornewsletter Related episodes: Where the US got $20B to bail out ArgentinaScott Bessent's $20 billion dollar gamble on ArgentinaFor sponsor-free episodes of The Indicator from Planet Money, subscribe to Planet Money+ via Apple Podcasts or at plus.npr.org. Fact-checking by Sierra Juarez. Music by Drop Electric. Find us: TikTok, Instagram, Facebook, Newsletter. See pcm.adswizz.com for information about our collection and use of personal data for sponsorship and to manage your podcast sponsorship preferences.NPR Privacy Policy
This week on the Watson Weekly, Rick Watson breaks down the biggest stories shaping commerce, technology, and AI infrastructure.Tim Cook is stepping down as Apple CEO after 14 transformative years that took the company from a $300B market cap to $4 trillion. John Turnis takes the helm September 1st.The Watson Weekly is sponsored by Avalara - the agentic AI platform automating global tax and compliance for leading eCommerce brands. For more details: https://avalaratax.watsonweekly.comAnthropic just inked a $100 billion, 10-year infrastructure deal with AWS for 5 gigawatts of compute, while Amazon pours another $5B (potentially $20B more) into the AI lab. Brad Jacobs strikes again: QXO is acquiring Top Build for ~$17B, his third deal in under a year.Plus, a tribute to industry leader Jon Panella.The Investor Minute with 5 items this week from the world of venture capital, acquisitions, and IPOs.
On this weekend edition of The Watson Weekly, Rick Watson and Jessica Lesesky break down the biggest stories shaping tech, retail, and AI.Amazon is doubling down on its Anthropic bet — with a new deal that has Anthropic committing $100B to AWS over the next decade, while Amazon pumps in an additional $5B(with up to $20B more on the table) and locks in guaranteed compute capacity. With over 100K customers already using Claude on Bedrock and Amazon holding a reported 15% stake, this partnership is reshaping the cloud AI landscape.Home Depot quietly acquired SIMPL Automation, a scrappy robotics startup that had raised just $100K before piloting its warehouse density technology at Home Depot's Locust Grove DC. Rick and Jessica unpack what this signals about the arms race between Home Depot and Lowe's to modernize supply chain and distribution.The Watson Weekly Weekend edition is sponsored by Avalara - the agentic AI platform automating global tax and compliance for leading eCommerce brands. For more details: https://avalaratax.watsonweekly.comPlus, the end of an era at Apple: Tim Cook is stepping down as CEO after 15 years to become executive chairman, handing the reins to 25-year Apple veteran John Ternus. Cook leaves behind a company transformed — from a $350B market cap to $4T, and a services business that now tops $100B. What does a hardware-focused CEO mean for Apple's AI strategy and search partnerships?
Joe's Premium Subscription: www.standardgrain.comGrain Markets and Other Stuff Links —Apple PodcastsSpotifyTikTokYouTubeFutures and options trading involves risk of loss and is not suitable for everyone.✅ Corn & soybean futures recap — what's driving the move higher✅ Trump extends Iran ceasefire — and what it means for crude oil & biofuels✅ Senate eyes $15–$20B in additional farmer aid — but is it actually coming?✅ Biofuel demand surging as crude oil climbs 30%+ since the start of the Iran war✅ Kalshi & Polymarket moving into crypto perpetual futures — could this force CME's hand on 24/7 commodity markets?✅ ADM crude corn oil spill on the Mississippi River at Red Wing, MN✅ USDA flash sales — 12 million bushels of corn sold to Colombia and unknown destinations
The Mad Man Theory can't be successful. Donald is going to fabricate victory even though it won't be a victory. Thousands of US troops will be needed to seize Iran's uranium. Donald is planning to return $20B in Iranian assets. WSJ reports Kash Patel is a drunk. The BBC is covering Donald's alleged insider trading. Donald's presidential library is a gigantic grift. Donald is negotiating another cash grab with the IRS he controls. Tariff refund website launches. With Jody Hamilton, David Ferguson, music by Astral Summer, Powder Pink & Sweet, and more! Brought to you by Russ Rybicki, SharePower Responsible Investing. Support our new sponsor and get free shipping at Quince.com/bob!See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
The Gary & Shannon Show Hour 2 (04.17) – We start this hour in space and somehow end up in Hollywood ethics.• Artemis II crew returns → instant “best friends” after 10 days sparks debate on whether shared chaos = real bonding• Listener talk-backs keep rolling → including reactions to the show and the now-infamous “green circle”• Crime vs reality → burglaries stacking up locally even as stats say things are improving• Iran deal questions linger → talk of $20B and what it actually means• KFI Live host and entertainment reporter Heather Brooker joins → Top Gun 3, Val Kilmer’s AI return, and the uncomfortable future of resurrecting actors• Box office rundown → what’s worth your time and what’s not• And then… Shannon pivots to Spielberg and original IP → launching into a full “Space Wars” pitch (featuring Bumperpuss) while Gary does everything he can not to engageSee omnystudio.com/listener for privacy information.
Vice President JD Vance, a creation of the Peter Thiel dark-money machine, traveled to Hungary to openly threaten voters on behalf of the Kremlin's favorite strongman, Viktor Orban. This is emotional blackmail on a global scale. Vance effectively told a post-Soviet nation that if they don't re-elect Orban this Sunday, they lose the protection of the U.S. military. It's a "nice country you've got there, shame if something happened to it" mafia tactic. Why is MAGA obsessed with Hungary? Because the EU is a regulatory miracle that Big Money hates. To the Peter Thiels of the world, Orban is the wedge designed to divide the EU, weaken Russian sanctions, and pave the way for a fascist Bannon nightmare across the continent. Under Orban, Hungary has become the most corrupt state in the EU, suffering from massive brain drain and crumbling infrastructure. Despite the threats, there is hope. On Sunday, April 12, Hungarians may unite behind a charismatic new leader, Péter Magyar. Will they choose freedom and fight against tyranny, like in the 1956 Hungarian Revolution, or will MAGA's blackmail work? We shall see. For a look at the MAGA virus in the UK, we continue our conversation with Dorian Lynskey, author of The Ministry of Truth: The Biography of George Orwell's 1984. Want to hear Gaslit Nation ad-free? Join our community of listeners for bonus shows, exclusive Q&A sessions, our group chats, invites to live events like our Monday political salons at 4pm ET over Zoom, and more! Sign up at Patreon.com/Gaslit! Join our April 13 book launch and live-taping to celebrate Mrs. Orwell. Patreon supporters get in free! https://powerhousearena.com/events/book-launch-mrs-orwell-by-andrea-chalupa-in-conversation-with-nomiki-konst/ Show Notes: Where the US got $20B to bail out Argentina https://www.npr.org/2025/11/13/nx-s1-5607023/where-the-us-got-20b-to-bail-out-argentina Hungary eyes US financial shield as EU funds remain frozen https://www.reuters.com/world/us-financial-shield-bolsters-hungary-amid-eu-funding-freeze-minister-says-2025-11-10/ Andrea's thread on Hungary's decline under Orban https://x.com/AndreaChalupa/status/1766082022781354193 MAGA's favorite strongman might be on the brink of defeat. We're about to find out whether an authoritarian can lose at the ballot box. https://www.vox.com/politics/485058/hungary-election-2026-orban-trump-vance-maga Viktor Orbán told Putin 'I am at your service' in October phone call https://www.theguardian.com/world/2026/apr/07/viktor-orban-told-putin-i-am-at-your-service-in-october-phonecall Italy: Steve Bannon's populist academy in the Trisulti monastery https://www.youtube.com/watch?v=fjSL1ofqGb8 Epstein files shed more light on Steve Bannon's efforts to influence European politics https://www.theguardian.com/us-news/2026/feb/05/jeffrey-epstein-files-steve-bannon-european-politics How Jeffrey Epstein sought to help Steve Bannon build a global populist movement https://www.cnn.com/2026/02/09/politics/steve-bannon-jeffrey-epstein-global-populism Péter Magyar, the former Orban ally vying for power in Hungary https://www.bbc.com/news/articles/c78l7vyylgqo
Episode 808: Neal and Toby cover the verdict that came down on Meta and Google, with a jury finding the companies liable for putting out addictive features that cause harm to teens. Then, the global supply chain is reeling as the crisis around the Strait of Hormuz deepens. Also, NASA plans to build a $20B moon base. Meanwhile, Neal presents his numbers on the struggling US worker, AI fruit slop, and the NBA's tanking problem. Learn more at linkedin.com/MBD Subscribe to Morning Brew Daily for more of the news you need to start your day. Share the show with a friend, and leave us a review on your favorite podcast app. Listen to Morning Brew Daily Here: https://www.swap.fm/l/mbd-note Watch Morning Brew Daily Here: https://www.youtube.com/@MorningBrewDailyShow Learn more about your ad choices. Visit megaphone.fm/adchoices
Thursday, March 20th, 2025 Judge Chutkan has blocked Trump and Musk from cancelling $20B in climate grants; Judge Ana Reyes has blocked the Trump administration's ban on transgender people serving in the military; Trump has fired the Democratic members of the Federal Trade Commission; Judge Beryl Howell has denied the temporary restraining order for the US Institute of Peace; Republican members of the Senate and House armed services committee are pushing back on Trump's plan to abandon a NATO command that has been exclusively American since Eisenhower; and Allison and Dana deliver your Good News. Guest: Congresswoman Sara Jacobs U.S. Congresswoman Sara Jacobs | CA 51st District @RepSaraJacobs • Blue Sky @repsarajacobs • Instagram @RepSaraJacobs • Twitter Stories: Judge Reyes BLOCKS Trump's Ban on Transgender Service Members- Allison Gill | Mullershewrote Trump Fires FTC's Democratic Commissioners | HuffPost Latest News Trump admin considers giving up NATO command that has been exclusively American since Eisenhower | NBC News Judge temporarily blocks EPA's effort to cancel $20 billion in climate grants | CBS News Good Trouble: WisDems is sponsoring phone banking to get out the word about the upcoming April state Supreme Court race. WisDems Virtual Phonebank! Volunteer Opportunities Near Me · WisDems on Mobilize Federal workers - feel free to email AG at fedoath@pm.me and let me know what you're going to do, or just vent. I'm always here to listen. Reminder - you can see the pod pics if you become a Patron. The good news pics are at the bottom of the show notes of each Patreon episode! That's just one of the perks of subscribing! patreon.com/muellershewrote Listener Survey:http://survey.podtrac.com/start-survey.aspx?pubid=BffJOlI7qQcF&ver=shortFollow the Podcast on Apple:https://apple.co/3XNx7ckWant to support the show and get it ad-free and early?https://patreon.com/thedailybeanshttps://dailybeans.supercast.com/https://apple.co/3UKzKt0 Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
In episode 2019, Jack and Miles are joined by comedian and host of Intercepts, David Huntsberger, to discuss… Trump Administration Blames Rising Oil Prices On Bad Vibes, Predator / Conan / Commando, Ballet And Opera Lovers Sure Are Pissed At Timothée Chalamet, Pentagon Has Been Havana Syndrome-ing Rats? And more! As oil prices spike, G7 opts not to dip into emergency reserves for now Trump's energy chief blames oil price spike on market fear 'Night turned into day': Iranians tell of strikes on oil depots As Iran chokes Strait of Hormuz, U.S. vows $20B for maritime reinsurance Scoop: U.S. dismayed by Israel's Iran fuel strikes, sources say US military tests on secret weapon bought from Russian criminal network reveal Havana Syndrome-like symptoms: report Unsurprisingly, tonight's 60 Minutes episode covering Havana Syndrome didn't offer a smoking gun because it was a sales pitch for a book coming out in September. The authors? Two 60 Minutes producers. All we've got to say is... use promo code TIMOTHEE to save 14% off select seats for Carmen, through this weekend only. Timmy, you're welcome to use it too
What if the difference between scaling up and burning out comes down to just one overlooked decision you make today?In this exclusive Second in Command episode, Cameron Herold sits down with Jon McNeill, former President of Tesla and COO of Lyft, and current CEO and Co-Founder of DVx Ventures, for a bold, eye-opening deep dive into the raw realities of being second in command at companies that redefine entire industries.You'll hear battle-tested lessons on navigating visionary founders, eliminating organizational bloat, and building operating systems that drive exponential growth, plus what most leaders get dead wrong about innovation, hiring, and execution at scale.If you crave real-world playbooks and not more recycled platitudes, hit play now. Miss this conversation and risk falling into the same chaos that sinks even the greatest companies. Listen today to steal field-proven COO frameworks you won't hear anywhere else before your competition does.Timestamped Highlights[00:03:16] – The $108 million mistake: why Jon McNeill turned down Uber and Tesla before they became giants[00:07:22] – From Bain to boardrooms: how Cameron Herold went from $1.8B to $20B in 30 months[00:14:49] – What it really feels like to drop into Tesla's leadership team—no roadmap, only chaos[00:17:04] – The pivotal moment Cameron Herold broke the rules at Tesla and why Elon Musk said “You'll fit right in”[00:21:09] – The “Big Thing” meeting—the deceptively simple method Cameron Herold stole from Facebook's top minds[00:26:43] – How to push back (and win) with the world's most demanding CEO[00:36:11] – The ruthless self-topgrading system that kept Tesla lean—could you survive it?[00:47:11] – Tesla's “Algorithm” revealed: the counterintuitive systems any leader can stealAbout the GuestJon McNeill is the former President of Tesla and COO of Lyft, a renowned serial entrepreneur, and current CEO and Co-founder of DVx Ventures. Recognized for multiplying company valuations and pioneering operational mastery at the world's most innovative companies, Jon now empowers founders and operators to scale with speed and discipline. His latest book, The Algorithm, reveals the operating system behind Tesla's success and is quickly becoming a must-read for growth-focused leaders.